### What does this PR do? Type of change: new example + bug fix <img width="2085" height="1239" alt="image" src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3" /> Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation** tutorial for [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`. It complements the existing Nemotron-3-Nano tutorial (pruning + distillation + FP8). Here the model is unpruned and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is what QAD exists to close. **Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than BF16* in 10 of 12 shapes, because a BF16 activation forces vLLM onto the Marlin dequant fallback and never reaches the Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to 1.30x) and shrinks the checkpoint 67 GiB -> 22 GiB (3.1x). **What the study found:** only 2 of 6 benchmarks show a statistically significant PTQ deficit, so those are the only two QAD can recover. IFBench is recovered to parity with BF16 (-2.62 pp -> -0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains a significant gap. The other four are lossless under W4A4 to begin with. Also added: - `modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml` — the PTQ recipe used as the QAD student, usable via `--recipe`. - `data_blend.yaml` — the token-budgeted blend config for the distillation data. - `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark. tau2-bench is separate because it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `deployment.command` is global to a config. **Two export fixes found while producing these checkpoints** (both change library/example behaviour, both have changelog entries under 0.48.0 Bug Fixes): - `unified_export_megatron.py` — MCore builds `embedding` on the MTP stage as well as the first, so gating export on `hasattr(model, "embedding")` wrote a **second, unreferenced copy of the vocab embedding** whenever an MTP model was exported with PP > 1. The index mapped the key to the later shard, so the extra copy never loaded but still shipped — ~1 GB for this model. Now gated on `model.pre_process`, MCore's own "this rank owns the input embedding" flag. - `export_quantized_megatron_to_hf.py` — stopped passing Megatron's `moe_router_dtype` as the router's *storage* dtype. It is a routing *compute* dtype; the parameter is bf16 in a bf16 model, so the export was widening bf16 to fp32. All 21,495,808 router values in the exported checkpoint have their low 16 bits zero, and vLLM builds the gate at the model dtype and rounds on load, so the dropped bytes carried no information. `export_mcore_gpt_to_hf` still accepts the override. ### Usage ```bash # 1. PTQ (2 GB200 nodes for EP=8) srun ... python examples/megatron_bridge/quantize.py \ --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \ --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \ --tp_size 1 --ep_size 8 --pp_size 1 \ --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \ --seq_length 8192 --skip_generate \ --export_megatron_path /path/to/qwen36_w4a4_megatron # 2. QAD (32 nodes x 4 GB200) python -u examples/megatron_bridge/distill.py \ --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \ --student_megatron_path /path/to/qwen36_w4a4_megatron \ --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \ --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \ --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \ --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \ --no_async_save --eval_iters 0 --save_interval 50 \ --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output ``` ### Testing **Library changes.** `tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains `test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports a PP=2 model built with `mtp_num_layers=1` and asserts no tensor lands in more than one shard. Verified to **fail without the fix**: ``` AssertionError: tensors written to more than one shard: {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors', 'model-00002-of-00002.safetensors')} ``` The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch this — it fakes `_get_mtp_state_dict` on a model with no real MTP, so the last stage never builds an embedding. Ran the whole `tests/gpu_megatron/torch/export/` suite with and without the fix: identical failure sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local `ImportError: FLA is not installed`), 58 passed with vs 56 without — the +2 being the new test's two workers. `tests/unit/recipe` passes 368/368 after the recipe path move. `model.pre_process` is always present: `GPTModelExporter.__init__` raises unless the model is `GPTModel` or `HybridModel`, and both set it unconditionally. Both export fixes were also applied to the real 23 GB checkpoints and re-validated end to end: every retained tensor md5-identical, index/shard integrity re-checked, and a **full GPQA re-evaluation of the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question t-test over the same 198 questions x 16 repeats: -0.76 pp, p=0.21, not significant). **Numbers in the tutorial** come from real runs, not estimates: - **253 evaluation runs** across BF16, the published W4A16 checkpoint, W4A4 PTQ, and QAD at 50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench; GPQA is one `num_repeats: 16` run). - The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was re-evaluated under this same harness (36 runs) rather than quoted from its card, so the W4A16 row is same-harness. - Every figure and results-table value is generated from the collected `results.yml` files by a script, and I verified the README table cell-by-cell against that data after each edit. - Throughput rows were cross-checked against the recorded AIPerf sweeps; the QAD wall-clock figures against the two jobs' Slurm records (`03:34:49` + `02:09:49`). - All CLI flags in the tutorial were verified to exist in `quantize.py` / `distill.py` / `export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against `dataset_utils.py`. The tutorial also records the non-obvious constraints found the hard way: QAD on this model requires `TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP starves Qwen3-VL's M-RoPE of `position_ids`, CP hits a rope shard mismatch), EP must match the PTQ checkpoint, and `--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]` fp32 logits are 30.31 GiB per tensor on a 248,320-token vocabulary. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP export dedup test --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes --> - Did you get Claude approval on this PR?: ✅ <!-- not yet run --> ### Additional Information Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0` label has been removed. Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface` to `model_type` (it is now a compatibility symlink), so the recipe moved to `model_type/qwen3_6_moe/ptq/`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4 quantization and quantization-aware distillation. * Added checkpoint export, accuracy evaluation, and vLLM throughput benchmarking workflows. * Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro, SciCode, and tau2 Telecom. * Added a token-budgeted supervised fine-tuning data configuration. * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE. * **Documentation** * Added benchmark results, deployment guidance, hardware requirements, reproduction steps, HTTPS endpoint guidance, announcement filters, and tutorial links. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Qwen3.6-35B-A3B: W4A4 NVFP4 PTQ + QAD + vLLM Deployment
End-to-end W4A4 optimization of Qwen/Qwen3.6-35B-A3B demonstrating how to push past weight-only quantization: NVFP4 W4A4 post-training quantization → Quantization-Aware Distillation (QAD) to recover the accuracy it costs → evaluation benchmarking → vLLM deployment. This document covers:
- Data Preparation — tokenizing the SFT blend for distillation
- Quantization — W4A4 NVFP4 PTQ with
examples/megatron_bridge/quantize.py - QAD — recovering accuracy with
examples/megatron_bridge/distill.py - Export — converting the quantized Megatron checkpoint to a deployable HF checkpoint
- Evaluation — benchmarking with NeMo Evaluator across MMMU-Pro, GPQA Diamond, SciCode, and more
- vLLM Inference Benchmarking — throughput comparison against BF16 on GB200
It closes with potential further improvements — the levers this run did not explore, and what the evidence says each might buy.
Note
This tutorial complements the Nemotron-3-Nano-30B-A3B tutorial, which covers pruning + distillation + FP8 PTQ. Here the starting point is an unpruned model and the technique under test is W4A4 — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is exactly what QAD exists to close.
Results
Main results — all models evaluated with the same setup. Values are mean ± sem across repeats; intermediate QAD checkpoints are in the figure above.
| Model | MMMU-Pro | GPQA Diamond | SciCode (Subtask) | AA-LCR | IFBench | tau2-bench Telecom | Average |
|---|---|---|---|---|---|---|---|
| BF16 (teacher) | 74.6 ± 0.2 | 84.7 | 39.9 ± 0.6 | 69.1 ± 0.8 | 60.0 ± 0.5 | 94.2 ± 1.0 | 70.4 |
| W4A16 NVFP4 PTQ (published)1 | 73.8 ± 0.2 | 84.4 | 38.5 ± 0.5 | 66.3 ± 1.0 | 59.4 ± 0.5 | 96.4 ± 0.4 | 69.8 |
| W4A4 NVFP4 PTQ (QAD student) | 73.4 ± 0.2 | 84.7 | 39.1 ± 0.7 | 70.0 ± 1.1 | 57.9 ± 0.5 | 94.2 ± 1.2 | 69.9 |
| ↳ + QAD 500 iters | 73.9 ± 0.2 | 84.2 | 40.2 ± 0.6 | 69.4 ± 1.5 | 59.6 ± 0.5 | 93.4 ± 0.4 | 70.1 |
1 The W4A16 NVFP4 row is the published checkpoint re-evaluated here with the same evaluation harness.
Where QAD actually helps
W4A4 quantization only measurably hurt two of the six benchmarks, so those are the only two where QAD has anything to recover:
| Benchmark | W4A4 PTQ vs BF16 | After 500 QAD iters | Outcome |
|---|---|---|---|
| IFBench | −2.6 pp (real loss) | −0.3 pp | ✅ recovered — back to BF16 |
| MMMU-Pro | −1.2 pp (real loss) | −0.7 pp | ⚠️ about 40% recovered, gap remains |
| GPQA Diamond, SciCode, AA-LCR, tau2-bench | no loss | no change | — nothing to recover |
"Real loss" means the gap is larger than the run-to-run noise, so it is not a measurement artifact. The four benchmarks in the last row land within ±1.2 pp of BF16 both before and after QAD, which is inside their noise — W4A4 is effectively lossless there.
Statistical detail (click to expand)
Comparisons are paired per-question t-tests: the same question answered by both models is compared directly, which cancels question-difficulty variance and is far more sensitive than comparing run averages. p is the probability of seeing a gap this large if the two models were actually equal; below 0.05 is conventionally "real".
| Benchmark | PTQ vs BF16 | QAD 500 vs PTQ | QAD 500 vs BF16 |
|---|---|---|---|
| MMMU-Pro | −1.23 (p=0.0002) | +0.51 (p=0.078) | −0.72 (p=0.023) |
| GPQA Diamond | +0.03 (p=0.96) | −0.41 (p=0.60) | −0.38 (p=0.58) |
| SciCode (Subtask) | −0.81 (p=0.40) | +1.15 (p=0.22) | +0.33 (p=0.72) |
| AA-LCR | +0.75 (p=0.73) | −0.62 (p=0.68) | +0.12 (p=0.95) |
| IFBench | −2.62 (p=0.017) | +2.33 (p=0.036) | −0.29 (p=0.79) |
| tau2-bench Telecom | −0.00 (p=1.00) | −0.88 (p=0.36) | −0.88 (p=0.44) |
See Potential further improvements for what would likely close the remaining MMMU-Pro gap.
vLLM Throughput (4× GB200, vLLM 0.28.0, TP=4 + EP)
| shape (ISL/OSL) | concurrency | BF16 | W4A16 NVFP4 | W4A4 NVFP4 | W4A4 / BF16 |
|---|---|---|---|---|---|
| decode 128/2048 | 1 | 351 | 225 | 307 | 0.88× |
| decode 128/2048 | 32 | 6,544 | 5,602 | 7,320 | 1.12× |
| decode 128/2048 | 128 | 17,636 | 15,222 | 20,076 | 1.14× |
| chat 8000/1000 | 32 | 3,649 | 2,373 | 4,079 | 1.12× |
| chat 8000/1000 | 128 | 5,571 | 6,138 | 6,370 | 1.14× |
| prefill 32000/400 | 128 | 854 | 1,092 | 1,107 | 1.30× |
Output tokens/s; 6 of the 12 measured shapes shown. W4A4 beats BF16 in 9 of 12 and beats W4A16 in all 12. Checkpoint size drops from 67 GiB → 22 GiB (3.1×).
The W4A16 column is why this tutorial targets W4A4 at all: weight-only NVFP4 is slower than BF16 in 10 of 12 shapes. A BF16 activation forces vLLM onto the Marlin dequant-to-BF16 fallback, which never reaches the Blackwell FP4 tensor cores. Only quantizing activations too unlocks them.
Steps to Reproduce
Environment: Results were produced with container nvcr.io/nvidia/nemo:26.08, ModelOpt 0.47.0 and nemo-evaluator-launcher 0.2.6 (with nemo-evaluator 0.2.8) on GB200 (aarch64). See the Megatron-Bridge README for environment setup (including ModelOpt mount path) and container usage. Deployment and evaluation of NVFP4 checkpoints require a Blackwell GPU.
1. Data Preparation
QAD currently supports language-model distillation with text-only data, so the blend is SFT text: nvidia/Nemotron-Cascade-2-SFT-Data, all 8 splits, 17.3B tokens.
Tokenize with the token-budgeted blend workflow using data_blend.yaml — set output_dir in that file first:
| Split | Tokens | Weight |
|---|---|---|
| chat | 9.73B | 56.2 |
| math | 3.65B | 21.1 |
| science | 1.89B | 10.9 |
| instruction_following | 0.57B | 3.3 |
| conversational_agent | 0.57B | 3.3 |
| terminal_agent | 0.57B | 3.3 |
| swe | 0.31B | 1.8 |
| safety | 0.003B | 0.02 |
Weights are each split's share of the dataset's own token count — a single pass over the natural mixture, not a tuned blend. The 500-iteration schedule consumes 8.4B tokens, about half an epoch, so nothing is resampled. Note the agent/SWE splits total under 9%, which is relevant to the agentic coverage discussed under potential further improvements.
2. Quantization
W4A4 NVFP4 PTQ using the recipe at modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml. See examples/megatron_bridge/README.md for full PTQ documentation.
PTQ takes ~9 min on 2 GB200 nodes (1.2 GPU-hours) at EP=8.
What gets quantized (read from the exported checkpoint's quantization_config):
| Component | Precision | Tensors |
|---|---|---|
MoE routed experts (down/gate/up_proj) |
NVFP4 W4A4 (block 16) | 30,720 |
| Shared expert | NVFP4 W4A4 | 120 |
lm_head |
NVFP4 W4A4 | 1 |
Linear attention (in_proj_qkv/in_proj_z/out_proj, 30 layers) |
FP8 W8A8 | 90 |
Full attention (q/k/v/o_proj, 10 layers) |
FP8 W8A8 | 40 |
| KV cache | FP8 (static) | — |
MTP head, MoE router, conv1d, in_proj_a/in_proj_b, embeddings, vision tower |
BF16 | — |
99.2% of quantized tensors are MoE experts — where the FP4 throughput win comes from. Attention stays at FP8: only 130 tensors, and numerically more sensitive.
W4A4 NVFP4 PTQ command (click to expand)
--ep_size 8 needs 8 ranks, which is 2 GB200 nodes (4 GPUs each) — launched with srun, one task per GPU. On 8-GPU nodes a single-node torchrun --nproc_per_node 8 is equivalent.
# SBATCH --nodes=2 --ntasks-per-node=4 --gpus-per-node=4
srun ... python /opt/Model-Optimizer/examples/megatron_bridge/quantize.py \
--hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
--recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
--tp_size 1 --ep_size 8 --pp_size 1 \
--calib_dataset_name cnn_nemotron_v2_mix \
--calib_num_samples 1024 \
--calib_batch_size 1 \
--seq_length 8192 \
--export_megatron_path /path/to/qwen36_w4a4_megatron \
--skip_generate
Set RANK/WORLD_SIZE/LOCAL_RANK from SLURM_PROCID/SLURM_NTASKS/SLURM_LOCALID inside the srun body, as in the QAD command below.
Important
Set
--ep_sizeto the expert-parallel size you will run QAD at. QAD loads this checkpoint directly and EP must match — the expert layout is baked into the distributed checkpoint. Re-exporting at a different EP later means re-running PTQ.
3. Quantization-Aware Distillation (QAD)
QAD fine-tunes the quantized student against the BF16 teacher, so the student learns weights that survive 4-bit rounding. See the QAD section of the Megatron-Bridge README.
Minimum hardware: the student and teacher are both resident, plus an fp32 gradient buffer and optimizer state — roughly 124 GB/GPU of the 185 GiB on a GB200 at the settings below. We used 32 nodes × 4 GB200 (128 GPUs); 500 iterations took 5.7 hours wall-clock (~735 GB200 GPU-hours). Steady-state training is ~37 s/iter (~5.1 h of the total); the remaining ~35 min is per-job startup, loading the 67 GiB teacher and the student — the run was split across two 4-hour jobs, so that cost is paid twice.
QAD command (click to expand)
Launched with srun, one task per GPU, across 32 nodes (128 ranks). Set RANK / WORLD_SIZE /
LOCAL_RANK from SLURM_PROCID / SLURM_NTASKS / SLURM_LOCALID in the srun body; python -u
keeps the multi-node logs unbuffered.
# SBATCH --nodes=32 --ntasks-per-node=4 --gpus-per-node=4
srun ... python -u /opt/Model-Optimizer/examples/megatron_bridge/distill.py \
--teacher_hf_path Qwen/Qwen3.6-35B-A3B \
--student_hf_path Qwen/Qwen3.6-35B-A3B \
--student_megatron_path /path/to/qwen36_w4a4_megatron \
--tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
--data_paths "${DATA_BLEND}" \
--data_path_to_cache /path/to/cache \
--seq_length 32768 \
--mbs 1 \
--gbs 512 \
--train_iters 500 \
--lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 \
--logit_kl_topk 4096 \
--recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
--no_async_save \
--eval_iters 0 \
--save_interval 50 \
--output_dir /path/to/qad_output
Non-default arguments:
--student_megatron_path— the quantized checkpoint from Section 2;--student_hf_pathstill points at the BF16 model, which supplies the architecture.--tp_size 1 --pp_size 1 --cp_size 1— required, not chosen (see below).--ep_size 8must match the PTQ checkpoint.--seq_length 32768 --gbs 512— 16.8M tokens/iteration, 1.7B per 100 iterations.--lr 1e-5 --min_lr 1e-6— an order of magnitude below typical distillation LRs: the job is to adapt weights to quantization, not to learn the task.--logit_kl_topk 4096— restricts the KD loss to the teacher's top-4096 vocab entries. With a 248,320-token vocabulary the dense[seq, vocab]fp32 logits are 30.31 GiB per tensor at 32K, which OOMs on its own.--recompute_*/--no_async_save/--eval_iters 0— all needed to fit. Async save spawns a worker needing its own CUDA context; the validation path computes full-vocab LM and MTP cross-entropy (top-k applies to training only), so eval OOMs at 32K even though training fits.
4. Export
Convert the quantized Megatron checkpoint to a deployable unified HuggingFace checkpoint (~7 min on 1 node, 0.5 GB200 GPU-hours). The unified exporter loads at TP=1, so use pipeline parallelism to shard across GPUs.
Export command (click to expand)
torchrun --nproc_per_node 4 /opt/Model-Optimizer/examples/megatron_bridge/export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
--megatron_path /path/to/qad_output/checkpoints/iter_0000500 \
--pp_size 4 \
--export_unified_hf_path /path/to/qwen36_w4a4_qad_hf
The exported checkpoint is directly deployable with vLLM, TensorRT-LLM and SGLang. Sanity-check it with generate_vllm.py --model <path>.
5. Evaluation
One config per benchmark in eval_configs/, so each can be launched independently with its own walltime and sampling:
| Benchmark | Config | Temp | Runs | Walltime | Metric name |
|---|---|---|---|---|---|
| MMMU-Pro | mmmu_pro.yaml | 1.0 | 8 | 2:30 | mmmu-pro_pass_at_1_symbolic_correct |
| GPQA Diamond | gpqa.yaml | 1.0 | 1 (avg-of-16) | 4:00 | gpqa_pass_at_1_avg-of-16_symbolic_correct |
| SciCode (Subtask) | scicode.yaml | 0.6 | 8 | 2:00 | scicode_pass_at_1_subtask_accuracy |
| AA-LCR | aa_lcr.yaml | 1.0 | 8 | 2:00 | aalcr_pass_at_1_judge_correct |
| IFBench | ifbench.yaml | 1.0 | 8 | 1:30 | ifbench_pass_at_1_average_score |
| tau2-bench Telecom | tau2_telecom.yaml | 1.0 | 3 | 1:30 | tau2_bench_telecom_pass_at_1_pass_at_1 |
Sampling follows the nvidia/Qwen3.6-35B-A3B-NVFP4 card (T=1.0, top_p=0.95, max_new_tokens=131072), except SciCode, which the card specifies at T=0.6. parallelism varies (32 default, 16 for AA-LCR's long contexts and for tau2-bench's multi-turn tool round-trips, 8 for SciCode's long generations) as a throughput knob only. The same configs serve BF16 and NVFP4 unchanged — vLLM reads the FP8 KV-cache setting from the checkpoint's own quantization_config.
Important
Two benchmarks need a second model endpoint besides the one under test. AA-LCR scores answers with a judge (
NS_JUDGE_URL/LCR_JUDGE_MODEL_ID); tau2-bench drives a simulated user (TAU2_ENDPOINT_URL/TAU2_USER_MODEL_ID). Both readINFERENCE_API_KEY, and neither can produce its metric without that endpoint. Point both athttps://URLs for the final endpoint: the key rides every request, and a redirect can downgrade it tohttpor send it to another origin. tau2-bench additionally needs its own deployment —--enable-auto-tool-choice --tool-call-parser qwen3_coderand--max-num-seqs 16— which is why it cannot share a config with the others (deployment.commandis global).qwen3_coderis what the model card specifies; the template's<tool_call>markers resemblehermes, which would mis-parse tool calls and silently invalidate the benchmark. Itstop_k/presence_penaltygo through tau2'sagent_argspassthrough, which NEL's top-levelparamsblock does not accept — that is why the other five configs omit them.
Important
Every task sets
num_repeats: 1; the repeat counts above come from launching a config that many times. N launches give N independentpass@1values, which is whatmean ± semand the paired tests need —num_repeats: Ninstead yields a singlepass@1[avg-of-N]with no spread. GPQA is the deliberate exception.
Important
Run enough repeats. These benchmarks are noisy, and 3 was not enough to tell signal from noise in either direction. At 3 repeats, AA-LCR (only 100 questions) showed a 3 pp swing that looked real and vanished completely at 8; IFBench's genuine +2.3 pp gain was invisible at 3 and only appeared at 8. Three runs also understates how noisy a benchmark is, so the measured spread looks tighter than it is.
Evaluation launch steps (click to expand)
Set execution.hostname, execution.account and deployment.checkpoint_path in the config, or override with -o <option>=<value>.
pip install "nemo-evaluator-launcher[all]==0.2.6"
# Required by every config:
export HF_TOKEN=<your_huggingface_token>
export SLURM_JOB_DIR=<path_to_slurm_job_output_dir>
# Required only by AA-LCR (judge) and tau2-bench (user simulator):
export INFERENCE_API_KEY=<key_for_those_endpoints>
export NS_JUDGE_URL=<judge_endpoint_url>
export LCR_JUDGE_MODEL_ID=<judge_model_id>
export TAU2_ENDPOINT_URL=<user_simulator_url>
export TAU2_USER_MODEL_ID=<user_simulator_model>
# One benchmark. To verify the pipeline first, add
# `-o ++evaluation.nemo_evaluator_config.config.params.limit_samples=8`
nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
# The 8 independent runs behind the results table: launch the same config 8 times and pool the
# per-run pass@1 values. Each launch returns its own invocation id; record them so you can select
# exactly those runs when pooling (each eval directory also holds a small canary run).
for i in $(seq 1 8); do
nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
done
# All six benchmarks at their reported repeat counts
for cfg_runs in mmmu_pro:8 gpqa:1 scicode:8 aa_lcr:8 ifbench:8 tau2_telecom:3; do
cfg=${cfg_runs%%:*}; runs=${cfg_runs##*:}
for i in $(seq 1 "$runs"); do
nemo-evaluator-launcher run --config "eval_configs/${cfg}.yaml"
done
done
For more details on NeMo Evaluator, see the GitHub repo and documentation.
Output length
Accuracy is only half the serving cost — a model that scores the same while emitting more tokens is slower end to end. Mean completion tokens per benchmark (response_stats.avg_completion_tokens, same runs as the results table):
| Benchmark | BF16 | W4A16 NVFP4 | W4A4 NVFP4 PTQ | + QAD 500 iters | runs/side |
|---|---|---|---|---|---|
| SciCode (Subtask) | 5,348 | +22.9% | +25.5% | +90.9% | 8 |
| IFBench | 12,177 | +7.0% | +9.3% | +17.4% | 8 |
| AA-LCR | 2,560 | +7.1% | +6.9% | +9.8% | 8 |
| MMMU-Pro | 9,382 | +3.0% | +6.2% | −3.8% | 8 |
| GPQA Diamond | 13,561 | +17.1% | +6.6% | −8.5% | 1 |
4-bit weights lengthen SciCode outputs by ~23-25% on their own — the published W4A16 checkpoint does it too, so it is not something QAD or W4A4 introduced. QAD then pushes SciCode to +90.9%, for an unchanged score (40.2 vs 39.9), while pulling GPQA and MMMU-Pro back toward BF16. The throughput table is measured at fixed output length, so it does not capture this.
What the SciCode number actually is — worth reading before running QAD on another model
It is not verbosity. It is a failure to terminate on a small fraction of sub-steps:
- Sub-steps that hit the 131,072-token cap inside
<think>go from 0.7% (20/2704, BF16) to 3.6% (96/2704, QAD 500). Almost all return zero answer tokens (93 of those 96) — the model writes a complete solution, says "I think I've been going in circles", and writes it again. In the case we inspected, a 20-word window repeats 352 times and 97.7% of the trace's 20-word windows are duplicates. - Those 3.6% of sub-steps burn 45.7% of all completion tokens, so they dominate the mean: excluding them it is +30.4% rather than +90.9%.
- The median also roughly doubles (+91.6%), so the whole distribution shifted right — this is not only a tail effect.
- Capped rate peaks at iteration 50 (4.3%) and settles at 3.1% / 3.6% by 300 / 500; it is not gradual drift.
The obvious suspect — that --logit_kl_topk 4096 leaves the stop tokens outside the loss — did not hold up. Probing the BF16 teacher over one runaway trace: </think> does fall outside top-4096 at 35% of positions overall, but in the looping region the teacher gives <|im_end|> a median rank of 5 and </think> ~570, both well inside top-k. The teacher is signalling "stop here" at positions the loss did cover, and the student still does not stop. More likely: the blend has few "the answer is written, now stop" positions in this style, and a teacher-forced loss never exercises free-running generation 10K+ tokens deep.
For the next QAD run, three things follow: track the length-capped rate as a first-class metric alongside accuracy (a benchmark score can stay flat while 3.6% of responses return nothing); consider top-p instead of top-k for the KD loss so coverage adapts to the teacher's entropy rather than a fixed rank; and if memory allows, full-vocab KL — at 32K on this 248,320-token vocabulary the dense fp32 logits are 30.31 GiB per tensor, which is why top-k was used here, but more GPU memory or a smaller model or shorter sequence may afford it.
6. vLLM Inference Benchmarking
Throughput was measured with AIPerf against a served vLLM endpoint on 4× GB200 — three ISL/OSL shapes × four concurrencies = the 12 shapes in the results table.
Serve + benchmark commands (click to expand)
The same vllm serve command works for BF16 and NVFP4; the quantized checkpoint carries its own format and FP8 KV-cache settings in quantization_config.
vllm serve <checkpoint_path> --served-model-name bench \
--host 127.0.0.1 --port 8000 \
--tensor-parallel-size 4 --data-parallel-size 1 --enable-expert-parallel \
--max-model-len 262144 --reasoning-parser qwen3 \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}' \
--max-num-batched-tokens 8192 --enable-chunked-prefill
# decode-bound, chat, and prefill-bound shapes at four concurrencies each
for shape in "128 2048 decode" "8000 1000 chat" "32000 400 prefill"; do
set -- $shape; ISL=$1; OSL=$2; NAME=$3
for C in 1 8 32 128; do
aiperf profile -m bench --endpoint-type chat --streaming -u localhost:8000 \
--synthetic-input-tokens-mean $ISL --output-tokens-mean $OSL \
--concurrency $C --request-count $(( C * 5 )) \
--tokenizer <checkpoint_path> \
--extra-inputs ignore_eos:true --random-seed 42 \
--artifact-dir bench/${NAME}_isl${ISL}_osl${OSL}/c${C}
done
done
Tip
Check the endpoint answers a known question correctly before benchmarking it. A checkpoint served with the wrong KV dtype or an unexpected kernel still produces tokens at a plausible rate — the throughput number looks fine and is meaningless. We gate every benchmark on a short prompt with a verifiable answer.
Tip
To deploy the model with vLLM, refer to the vLLM Quickstart documentation.
Potential further improvements
This is a single 500-iteration run at 32K on a text-only blend. Each of those three choices is an unexplored lever.
1. Continue QAD longer. The recovery had not saturated: IFBench's gain arrived late (−0.11 pp vs the student at iteration 300, +2.33 pp at 500) and MMMU-Pro was still climbing (+0.51 pp, p=0.078). Training also used only 8.4B of the blend's 17.3B tokens, so it can roughly double before repeating a sample. At ~735 GPU-hours per 500 iterations this is the most direct experiment, and IFBench + MMMU-Pro alone (~55 GPU-hours) are enough to read the result.
2. Longer sequence length (64K). The model serves at 262K context, so 32K trains on a fraction of it; AA-LCR, the one long-context benchmark, showed no movement either way. At 32K the run already needs top-k KD plus full recompute, and memory is dominated by the resident student + teacher + fp32 gradient buffer rather than by sequence length — so 64K needs a different memory lever.
3. A better data blend. The clearest signal in the study: MMMU-Pro is multimodal, the blend is text-only, and MMMU-Pro is the one deficit QAD did not close. Since QAD supports language-model distillation with text-only data, recovering a multimodal benchmark that way asks the method to do something it is not set up for. Agentic coverage is also thin (agent/SWE splits under 9%, and tau2-bench showed the transient 8× agent_error spike), and weights here are simply proportional to token count — contrast the Nemotron-3-Nano blend, which is deliberately designed against target benchmarks and ablated.
If you run only one, run (1): no engineering risk, the trajectory says the gain is still arriving, and it produces the evidence to judge whether (2) and (3) are worth their cost.
