mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new example + bug fix <img width="2085" height="1239" alt="image" src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3" /> Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation** tutorial for [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`. It complements the existing Nemotron-3-Nano tutorial (pruning + distillation + FP8). Here the model is unpruned and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is what QAD exists to close. **Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than BF16* in 10 of 12 shapes, because a BF16 activation forces vLLM onto the Marlin dequant fallback and never reaches the Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to 1.30x) and shrinks the checkpoint 67 GiB -> 22 GiB (3.1x). **What the study found:** only 2 of 6 benchmarks show a statistically significant PTQ deficit, so those are the only two QAD can recover. IFBench is recovered to parity with BF16 (-2.62 pp -> -0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains a significant gap. The other four are lossless under W4A4 to begin with. Also added: - `modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml` — the PTQ recipe used as the QAD student, usable via `--recipe`. - `data_blend.yaml` — the token-budgeted blend config for the distillation data. - `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark. tau2-bench is separate because it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `deployment.command` is global to a config. **Two export fixes found while producing these checkpoints** (both change library/example behaviour, both have changelog entries under 0.48.0 Bug Fixes): - `unified_export_megatron.py` — MCore builds `embedding` on the MTP stage as well as the first, so gating export on `hasattr(model, "embedding")` wrote a **second, unreferenced copy of the vocab embedding** whenever an MTP model was exported with PP > 1. The index mapped the key to the later shard, so the extra copy never loaded but still shipped — ~1 GB for this model. Now gated on `model.pre_process`, MCore's own "this rank owns the input embedding" flag. - `export_quantized_megatron_to_hf.py` — stopped passing Megatron's `moe_router_dtype` as the router's *storage* dtype. It is a routing *compute* dtype; the parameter is bf16 in a bf16 model, so the export was widening bf16 to fp32. All 21,495,808 router values in the exported checkpoint have their low 16 bits zero, and vLLM builds the gate at the model dtype and rounds on load, so the dropped bytes carried no information. `export_mcore_gpt_to_hf` still accepts the override. ### Usage ```bash # 1. PTQ (2 GB200 nodes for EP=8) srun ... python examples/megatron_bridge/quantize.py \ --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \ --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \ --tp_size 1 --ep_size 8 --pp_size 1 \ --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \ --seq_length 8192 --skip_generate \ --export_megatron_path /path/to/qwen36_w4a4_megatron # 2. QAD (32 nodes x 4 GB200) python -u examples/megatron_bridge/distill.py \ --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \ --student_megatron_path /path/to/qwen36_w4a4_megatron \ --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \ --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \ --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \ --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \ --no_async_save --eval_iters 0 --save_interval 50 \ --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output ``` ### Testing **Library changes.** `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains `test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports a PP=2 model built with `mtp_num_layers=1` and asserts no tensor lands in more than one shard. Verified to **fail without the fix**: ``` AssertionError: tensors written to more than one shard: {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors', 'model-00002-of-00002.safetensors')} ``` The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch this — it fakes `_get_mtp_state_dict` on a model with no real MTP, so the last stage never builds an embedding. Ran the whole `tests/gpu_megatron/torch/export/` suite with and without the fix: identical failure sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local `ImportError: FLA is not installed`), 58 passed with vs 56 without — the +2 being the new test's two workers. `tests/unit/recipe` passes 368/368 after the recipe path move. `model.pre_process` is always present: `GPTModelExporter.__init__` raises unless the model is `GPTModel` or `HybridModel`, and both set it unconditionally. Both export fixes were also applied to the real 23 GB checkpoints and re-validated end to end: every retained tensor md5-identical, index/shard integrity re-checked, and a **full GPQA re-evaluation of the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question t-test over the same 198 questions x 16 repeats: -0.76 pp, p=0.21, not significant). **Numbers in the tutorial** come from real runs, not estimates: - **253 evaluation runs** across BF16, the published W4A16 checkpoint, W4A4 PTQ, and QAD at 50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench; GPQA is one `num_repeats: 16` run). - The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was re-evaluated under this same harness (36 runs) rather than quoted from its card, so the W4A16 row is same-harness. - Every figure and results-table value is generated from the collected `results.yml` files by a script, and I verified the README table cell-by-cell against that data after each edit. - Throughput rows were cross-checked against the recorded AIPerf sweeps; the QAD wall-clock figures against the two jobs' Slurm records (`03:34:49` + `02:09:49`). - All CLI flags in the tutorial were verified to exist in `quantize.py` / `distill.py` / `export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against `dataset_utils.py`. The tutorial also records the non-obvious constraints found the hard way: QAD on this model requires `TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP starves Qwen3-VL's M-RoPE of `position_ids`, CP hits a rope shard mismatch), EP must match the PTQ checkpoint, and `--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]` fp32 logits are 30.31 GiB per tensor on a 248,320-token vocabulary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP export dedup test --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes --> - Did you get Claude approval on this PR?: ✅ <!-- not yet run --> ### Additional Information Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0` label has been removed. Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface` to `model_type` (it is now a compatibility symlink), so the recipe moved to `model_type/qwen3_6_moe/ptq/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4 quantization and quantization-aware distillation. * Added checkpoint export, accuracy evaluation, and vLLM throughput benchmarking workflows. * Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro, SciCode, and tau2 Telecom. * Added a token-budgeted supervised fine-tuning data configuration. * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE. * **Documentation** * Added benchmark results, deployment guidance, hardware requirements, reproduction steps, HTTPS endpoint guidance, announcement filters, and tutorial links. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>