Files
Model-Optimizer/tests
Keval MorabiaandClaude Opus 5 a448ba9757 Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 02:10:12 +05:30
..