Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)

### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Keval Morabia
2026-09-17 02:10:12 +05:30
committed by GitHub
co-authored by Claude Opus 5
parent f68bf83cfa
commit a448ba9757
21 changed files with 1432 additions and 6 deletions
+6
View File
@@ -12,6 +12,10 @@ Changelog
- Add support for quantizing and calibrating enabled operators outside the transformer layers, such as ``lm_head``, when using layerwise calibration.
- Add an end-to-end BEVFormer ONNX PTQ example with temporal calibration data generation, INT8 and FP8 quantization, TensorRT engine building, and nuScenes accuracy evaluation. See `examples/onnx_ptq/bevformer/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/onnx_ptq/bevformer>`_ for details.
*Megatron Framework (M-LM / M-Bridge)*
- Add an end-to-end W4A4 NVFP4 PTQ and QAD tutorial for Qwen3.6-35B-A3B also covering evaluation and vLLM throughput benchmarking. See `examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/>`_ for details.
*Misc*
- A tracked ``examples/hf_ptq/hf_ptq.py`` run now writes ``.experiment.json`` into ``--export_path`` and uploads the same file with the run, so a checkpoint on disk names the experiment and MLflow run id that produced it. The pointer is written only once the export completes, and an export that is not tracked removes one it would otherwise inherit from a reused ``--export_path`` or from a quantized source checkpoint.
@@ -29,6 +33,8 @@ Changelog
**Bug Fixes**
- Fix ``examples/megatron_bridge/export_quantized_megatron_to_hf.py`` storing the MoE router at Megatron's ``moe_router_dtype``, which is a routing *compute* dtype, not a storage one. The router now exports at the export ``dtype`` like every other unquantized weight, matching what ``hf_ptq.py`` and the released NVFP4 checkpoints contain; pass ``moe_router_dtype`` to ``export_mcore_gpt_to_hf`` explicitly if you want the old fp32 storage.
- Fix unified Megatron export writing a second, unreferenced copy of the vocab embedding when a model with MTP layers is exported with pipeline parallelism. The duplicate was never loaded but inflated the checkpoint by the size of the embedding (about 1 GB for Qwen3.6-35B-A3B); re-export to reclaim the space.
- Fix ONNX INT8 entropy calibration failing or producing invalid quantization parameters for FP16 activations.
- Fix ``--use_fsdp2`` HuggingFace checkpoint export gathering the whole model onto rank 0, which made export the dominant phase of a PTQ run and could exhaust host memory on large models. The model is now split into per-decoder-layer units dealt round-robin across ranks; each rank gathers every unit but keeps, packs, and writes only the ones it owns, so a rank buffers roughly ``model / world_size`` instead of the whole checkpoint, and rank 0 writes the combined index. Export configurations that cannot be split this way now raise instead of producing a mismatched checkpoint: FSDP2 combined with another DTensor parallelism (for example FSDP2 + tensor parallel on a 2-D mesh; HSDP is supported), models whose decoder layers cannot be discovered, a decoder layer object reused across layers, and a module that holds the decoder layers while owning parameters of its own.
- Speed up ``mtq.quantize`` on FSDP2-sharded fused-MoE models. Promoting static-block weight quantizers gathered each expert's slice of the fused weight across ranks even though only quantizer state is read, adding a collective per expert to calibration.
+1
View File
@@ -27,6 +27,7 @@ Model Optimizer is also integrated with [NVIDIA Megatron-Bridge](https://github.
## Latest News
- [2026/09/16] [**End-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B**](./examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B): NVFP4 W4A4 PTQ plus quantization-aware distillation, reaching up to 1.30x vLLM throughput over BF16 and 3.1x smaller checkpoints while recovering the accuracy W4A4 costs.
- [2026/08/24] [BLOG: AutoQuantize: A Fast Automatic Mixed-Precision Assignment](https://nvidia.github.io/Model-Optimizer/announcements/autoquantize.html)
- [2026/08/17] [BLOG: Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer](https://developer.nvidia.com/blog/developing-nemotron-3-5-lightning-nvfp4-with-qad-using-nvidia-model-optimizer/): Learn how quantization-aware distillation recovers accuracy from aggressive NVFP4 quantization while reducing model size and increasing throughput.
- [2026/06/26] [BLOG: Creating the NVIDIA Nemotron 3 Ultra NVFP4 Checkpoint with NVIDIA Model Optimizer](https://developer.nvidia.com/blog/creating-the-nvidia-nemotron-3-ultra-nvfp4-checkpoint-with-nvidia-model-optimizer/): How we quantized Nemotron 3 Ultra (550B) to NVFP4 with Model Optimizer — up to 5.9× higher decode-heavy inference throughput than GLM-5.1 754B FP4 while matching BF16 accuracy. [NVFP4 Checkpoint](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4) on Hugging Face.
Binary file not shown.

After

Width:  |  Height:  |  Size: 240 KiB

@@ -0,0 +1,157 @@
:orphan:
Recovering W4A4 NVFP4 Accuracy with Quantization-Aware Distillation
###################################################################
:Author: Model Optimizer Team
:Date: September 16, 2026
:Tags: quantization, nvfp4, w4a4, qad, distillation, megatron-bridge
Weight-only NVFP4 does not make this model faster. On Blackwell, a BF16 activation forces vLLM onto
the Marlin dequant-to-BF16 fallback, which never reaches the FP4 tensor cores: W4A16 measured
*slower* than BF16 in 10 of 12 shapes. Quantizing activations as well (W4A4) unlocks those kernels
and beats BF16 in 9 of 12 shapes, but costs accuracy that post-training quantization alone does not
recover. Quantization-Aware Distillation (QAD) is what closes that gap.
We ran the full flow on `Qwen/Qwen3.6-35B-A3B <https://huggingface.co/Qwen/Qwen3.6-35B-A3B>`_ with
Megatron-Bridge: W4A4 NVFP4 PTQ, then 500 QAD iterations against the BF16 teacher, evaluated across
six benchmarks. The complete reproduction steps, configs and recipe are in the
`tutorial <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B>`__.
Highlights
**********
* **W4A4 is the configuration worth targeting.** It beats BF16 in 9 of 12 measured shapes (up to
1.30x) and shrinks the checkpoint from 67 GiB to 22 GiB (3.1x). Weight-only W4A16 is slower than
BF16 in 10 of 12.
* **Only 2 of 6 benchmarks lost measurable accuracy to W4A4**, so those are the only two where QAD
has anything to recover.
* **IFBench is fully recovered.** PTQ costs 2.6 pp; 500 QAD iterations return the model to
statistical parity with BF16.
* **MMMU-Pro recovers about 40%** of its deficit and retains a measurable gap — the blend is
text-only, and MMMU-Pro is multimodal.
* **Repeat counts matter more than expected.** At 3 repeats one benchmark showed a 3 pp swing that
vanished at 8, and IFBench's real gain was invisible until 8.
Results
*******
.. image:: assets/qwen36-w4a4-qad-learning-curves.png
:alt: W4A4 NVFP4 accuracy across PTQ and QAD iterations 50, 300 and 500 for six benchmarks
:width: 100%
**Figure 1. W4A4 NVFP4 accuracy vs. QAD training iteration, Qwen3.6-35B-A3B.** Each panel starts at
the PTQ student (x=0) and follows it through 500 QAD iterations. The dashed line is the BF16
teacher; the dotted line is the PTQ starting point. Bars are ±1 sem.
Accuracy, ``mean ± sem`` across repeats:
.. list-table::
:header-rows: 1
:widths: 26 12 12 12 12 12 14
* - Model
- MMMU-Pro
- GPQA-D
- SciCode
- AA-LCR
- IFBench
- tau2-bench
* - BF16 (teacher)
- 74.6 ± 0.2
- 84.7
- 39.9 ± 0.6
- 69.1 ± 0.8
- 60.0 ± 0.5
- 94.2 ± 1.0
* - W4A4 NVFP4 PTQ
- 73.4 ± 0.2
- 84.7
- 39.1 ± 0.7
- 70.0 ± 1.1
- 57.9 ± 0.5
- 94.2 ± 1.2
* - **+ QAD 500 iters**
- **73.9 ± 0.2**
- **84.2**
- **40.2 ± 0.6**
- **69.4 ± 1.5**
- **59.6 ± 0.5**
- **93.4 ± 0.4**
Where QAD Helps
***************
Only IFBench and MMMU-Pro show a W4A4 deficit larger than run-to-run noise. On the other four,
W4A4 is effectively lossless and QAD neither helps nor hurts.
.. list-table::
:header-rows: 1
:widths: 22 26 26 26
* - Benchmark
- W4A4 PTQ vs BF16
- After 500 QAD iters
- Outcome
* - IFBench
- −2.6 pp
- −0.3 pp
- Recovered to BF16 parity
* - MMMU-Pro
- −1.2 pp
- −0.7 pp
- ~40% recovered, gap remains
Throughput
**********
Measured with AIPerf against a served vLLM endpoint on 4x GB200, three ISL/OSL shapes at four
concurrencies each (output tokens/s):
.. list-table::
:header-rows: 1
:widths: 26 14 14 16 16 14
* - Shape (ISL/OSL)
- Conc.
- BF16
- W4A16 NVFP4
- W4A4 NVFP4
- W4A4 / BF16
* - decode 128/2048
- 128
- 17,636
- 15,222
- **20,076**
- **1.14x**
* - chat 8000/1000
- 32
- 3,649
- 2,373
- **4,079**
- **1.12x**
* - prefill 32000/400
- 128
- 854
- 1,092
- **1,107**
- **1.30x**
W4A4 loses to BF16 only at concurrency 1, by a near-constant 0.87–0.88x across all three shapes.
That is fixed per-call overhead in the FP4 kernel path, which batch-1 decode has no arithmetic
intensity to amortize — not a recipe knob.
What It Costs
*************
QAD dominates: 500 iterations at 32K sequence length on 32 nodes (128 GB200) took 5.7 hours,
about 735 GPU-hours. PTQ is 9 minutes and the HF export a few more. Evaluating one checkpoint
across all six benchmarks at the repeat counts above is roughly 100 GPU-hours.
Learn More
**********
The `tutorial <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B>`__
has the full reproduction path — data blend, PTQ recipe, QAD command, export, and one NeMo Evaluator
config per benchmark — plus the parallelism constraints QAD imposes on this model and what we would
try next to close the remaining MMMU-Pro gap.
+9
View File
@@ -13,6 +13,9 @@ Release notes, technical updates, examples, and deployment stories from the Mode
<button class="announcement-tag is-active" type="button" data-tag="all" aria-pressed="true">All</button>
<button class="announcement-tag" type="button" data-tag="release" aria-pressed="false">Release</button>
<button class="announcement-tag" type="button" data-tag="autoquantize" aria-pressed="false">AutoQuantize</button>
<button class="announcement-tag" type="button" data-tag="quantization" aria-pressed="false">Quantization</button>
<button class="announcement-tag" type="button" data-tag="nvfp4" aria-pressed="false">NVFP4</button>
<button class="announcement-tag" type="button" data-tag="qad" aria-pressed="false">QAD</button>
<button class="announcement-tag" type="button" data-tag="speculative-decoding" aria-pressed="false">Speculative decoding</button>
<button class="announcement-tag" type="button" data-tag="dflash" aria-pressed="false">DFlash</button>
<button class="announcement-tag" type="button" data-tag="dspark" aria-pressed="false">DSpark</button>
@@ -24,6 +27,12 @@ Release notes, technical updates, examples, and deployment stories from the Mode
</div>
<div class="announcement-grid" id="announcement-grid">
<article class="announcement-card" data-date="2026-09-16" data-title="Recovering W4A4 NVFP4 Accuracy with Quantization-Aware Distillation" data-summary="Weight-only NVFP4 is slower than BF16 on Blackwell; W4A4 unlocks the FP4 kernels, and QAD recovers the accuracy it costs." data-tags="quantization nvfp4 w4a4 qad distillation megatron-bridge">
<div class="announcement-card-meta">September 16, 2026 &middot; Model Optimizer Team</div>
<h2><a href="announcements/qwen36-w4a4-qad.html">Recovering W4A4 NVFP4 Accuracy with Quantization-Aware Distillation</a></h2>
<p>Weight-only NVFP4 is <em>slower</em> than BF16 on Blackwell. W4A4 beats it in 9 of 12 shapes, and QAD recovers the accuracy W4A4 costs on Qwen3.6-35B-A3B.</p>
<div class="announcement-card-tags"><span>quantization</span><span>nvfp4</span><span>w4a4</span><span>qad</span><span>distillation</span><span>megatron-bridge</span></div>
</article>
<article class="announcement-card" data-date="2026-08-24" data-title="AutoQuantize: A Fast Automatic Mixed-Precision Assignment" data-summary="AutoQuantize finds low-sensitivity mixed-precision assignments with gradient-based scoring under a modeled effective-bits budget." data-tags="autoquantize quantization mixed-precision modelopt">
<div class="announcement-card-meta">August 24, 2026 &middot; Model Optimizer Team</div>
<h2><a href="announcements/autoquantize.html">AutoQuantize: A Fast Automatic Mixed-Precision Assignment</a></h2>
+2
View File
@@ -17,6 +17,8 @@ This directory contains examples of using Model Optimizer with the [NeMo Megatro
> [!TIP]
> Checkout the [Nemotron-3-Nano-30B-A3B pruning + distillation (with data blend prep) + quantization tutorial](tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) for a complete end-to-end workflow using Megatron-Bridge!
>
> Or the [Qwen3.6-35B-A3B W4A4 NVFP4 + QAD tutorial](tutorials/Qwen3.6-35B-A3B/README.md) for an end-to-end quantization-aware distillation workflow.
## Pre-Requisites
@@ -166,7 +166,6 @@ def main(args: argparse.Namespace):
export_extra_modules=export_extra_modules,
dtype=torch.bfloat16,
export_dir=args.export_unified_hf_path,
moe_router_dtype=getattr(unwrapped_model.config, "moe_router_dtype", None),
trust_remote_code=trust_remote_code,
)
print_rank_0(f"Exported HuggingFace checkpoint to {args.export_unified_hf_path}")
@@ -0,0 +1,382 @@
# Qwen3.6-35B-A3B: W4A4 NVFP4 PTQ + QAD + vLLM Deployment
End-to-end W4A4 optimization of [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) demonstrating how to push past weight-only quantization: NVFP4 W4A4 post-training quantization → Quantization-Aware Distillation (QAD) to recover the accuracy it costs → evaluation benchmarking → vLLM deployment. This document covers:
1. **[Data Preparation](#1-data-preparation)** — tokenizing the SFT blend for distillation
2. **[Quantization](#2-quantization)** — W4A4 NVFP4 PTQ with `examples/megatron_bridge/quantize.py`
3. **[QAD](#3-quantization-aware-distillation-qad)** — recovering accuracy with `examples/megatron_bridge/distill.py`
4. **[Export](#4-export)** — converting the quantized Megatron checkpoint to a deployable HF checkpoint
5. **[Evaluation](#5-evaluation)** — benchmarking with NeMo Evaluator across MMMU-Pro, GPQA Diamond, SciCode, and more
6. **[vLLM Inference Benchmarking](#6-vllm-inference-benchmarking)** — throughput comparison against BF16 on GB200
It closes with [potential further improvements](#potential-further-improvements) — the levers this run did not explore, and what the evidence says each might buy.
> [!NOTE]
> This tutorial complements the [Nemotron-3-Nano-30B-A3B tutorial](../NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md), which covers pruning + distillation + FP8 PTQ. Here the starting point is an unpruned model and the technique under test is **W4A4** — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is exactly what QAD exists to close.
## Results
![W4A4 NVFP4 accuracy recovery during QAD](figures/qad_learning_curves.png)
<b>Main results</b> — all models evaluated with the same [setup](#5-evaluation). Values are `mean ± sem` across repeats; intermediate QAD checkpoints are in the figure above.
| Model | MMMU-Pro | GPQA Diamond | SciCode (Subtask) | AA-LCR | IFBench | tau2-bench Telecom | Average |
| --- | --- | --- | --- | --- | --- | --- | --- |
| **BF16** (teacher) | 74.6 ± 0.2 | 84.7 | 39.9 ± 0.6 | 69.1 ± 0.8 | 60.0 ± 0.5 | 94.2 ± 1.0 | 70.4 |
| [W4A16 NVFP4 PTQ](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4) (published)<sup>1</sup> | 73.8 ± 0.2 | 84.4 | 38.5 ± 0.5 | 66.3 ± 1.0 | 59.4 ± 0.5 | 96.4 ± 0.4 | 69.8 |
| **W4A4 NVFP4 PTQ** (QAD student) | 73.4 ± 0.2 | 84.7 | 39.1 ± 0.7 | 70.0 ± 1.1 | 57.9 ± 0.5 | 94.2 ± 1.2 | 69.9 |
| **↳ + QAD 500 iters** | **73.9 ± 0.2** | **84.2** | **40.2 ± 0.6** | **69.4 ± 1.5** | **59.6 ± 0.5** | **93.4 ± 0.4** | **70.1** |
><sup>1</sup> The W4A16 NVFP4 row is the published checkpoint re-evaluated here with the same evaluation harness.
### Where QAD actually helps
W4A4 quantization only measurably hurt **two** of the six benchmarks, so those are the only two where QAD has anything to recover:
| Benchmark | W4A4 PTQ vs BF16 | After 500 QAD iters | Outcome |
| --- | --- | --- | --- |
| **IFBench** | **−2.6 pp** (real loss) | **−0.3 pp** | ✅ recovered — back to BF16 |
| **MMMU-Pro** | **−1.2 pp** (real loss) | **−0.7 pp** | ⚠️ about 40% recovered, gap remains |
| GPQA Diamond, SciCode, AA-LCR, tau2-bench | no loss | no change | — nothing to recover |
"Real loss" means the gap is larger than the run-to-run noise, so it is not a measurement artifact. The four benchmarks in the last row land within ±1.2 pp of BF16 both before and after QAD, which is inside their noise — W4A4 is effectively lossless there.
<details>
<summary>Statistical detail (click to expand)</summary>
Comparisons are **paired per-question t-tests**: the same question answered by both models is compared directly, which cancels question-difficulty variance and is far more sensitive than comparing run averages. `p` is the probability of seeing a gap this large if the two models were actually equal; below 0.05 is conventionally "real".
| Benchmark | PTQ vs BF16 | QAD 500 vs PTQ | QAD 500 vs BF16 |
| --- | --- | --- | --- |
| MMMU-Pro | −1.23 (p=0.0002) | +0.51 (p=0.078) | −0.72 (p=0.023) |
| GPQA Diamond | +0.03 (p=0.96) | −0.41 (p=0.60) | −0.38 (p=0.58) |
| SciCode (Subtask) | −0.81 (p=0.40) | +1.15 (p=0.22) | +0.33 (p=0.72) |
| AA-LCR | +0.75 (p=0.73) | −0.62 (p=0.68) | +0.12 (p=0.95) |
| IFBench | −2.62 (p=0.017) | +2.33 (p=0.036) | −0.29 (p=0.79) |
| tau2-bench Telecom | −0.00 (p=1.00) | −0.88 (p=0.36) | −0.88 (p=0.44) |
</details>
See [Potential further improvements](#potential-further-improvements) for what would likely close the remaining MMMU-Pro gap.
### vLLM Throughput (4× GB200, vLLM 0.28.0, TP=4 + EP)
| shape (ISL/OSL) | concurrency | BF16 | W4A16 NVFP4 | **W4A4 NVFP4** | W4A4 / BF16 |
| --- | --- | --- | --- | --- | --- |
| decode 128/2048 | 1 | 351 | 225 | 307 | 0.88× |
| decode 128/2048 | 32 | 6,544 | 5,602 | **7,320** | **1.12×** |
| decode 128/2048 | 128 | 17,636 | 15,222 | **20,076** | **1.14×** |
| chat 8000/1000 | 32 | 3,649 | 2,373 | **4,079** | **1.12×** |
| chat 8000/1000 | 128 | 5,571 | 6,138 | **6,370** | **1.14×** |
| prefill 32000/400 | 128 | 854 | 1,092 | **1,107** | **1.30×** |
Output tokens/s; 6 of the 12 measured shapes shown. **W4A4 beats BF16 in 9 of 12 and beats W4A16 in all 12.** Checkpoint size drops from **67 GiB → 22 GiB (3.1×)**.
The W4A16 column is why this tutorial targets W4A4 at all: **weight-only NVFP4 is *slower* than BF16** in 10 of 12 shapes. A BF16 activation forces vLLM onto the Marlin dequant-to-BF16 fallback, which never reaches the Blackwell FP4 tensor cores. Only quantizing activations too unlocks them.
---
## Steps to Reproduce
**Environment:** Results were produced with container `nvcr.io/nvidia/nemo:26.08`, ModelOpt 0.47.0 and `nemo-evaluator-launcher` 0.2.6 (with `nemo-evaluator` 0.2.8) on GB200 (aarch64). See the [Megatron-Bridge README](../../README.md) for environment setup (including ModelOpt mount path) and container usage. Deployment and evaluation of NVFP4 checkpoints require a Blackwell GPU.
### 1. Data Preparation
QAD currently supports **language-model distillation with text-only data**, so the blend is SFT text: [nvidia/Nemotron-Cascade-2-SFT-Data](https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-SFT-Data), all 8 splits, 17.3B tokens.
Tokenize with the [token-budgeted blend workflow](../../../dataset/MEGATRON_DATA_PREP.md#prepare-token-budgeted-data-blends) using [data_blend.yaml](data_blend.yaml) — set `output_dir` in that file first:
| Split | Tokens | Weight |
| --- | --- | --- |
| chat | 9.73B | 56.2 |
| math | 3.65B | 21.1 |
| science | 1.89B | 10.9 |
| instruction_following | 0.57B | 3.3 |
| conversational_agent | 0.57B | 3.3 |
| terminal_agent | 0.57B | 3.3 |
| swe | 0.31B | 1.8 |
| safety | 0.003B | 0.02 |
Weights are each split's share of the dataset's own token count — a single pass over the natural mixture, not a tuned blend. The 500-iteration schedule consumes **8.4B tokens, about half an epoch**, so nothing is resampled. Note the agent/SWE splits total under 9%, which is relevant to the agentic coverage discussed under [potential further improvements](#potential-further-improvements).
---
### 2. Quantization
W4A4 NVFP4 PTQ using the recipe at [`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`](../../../../modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml). See [examples/megatron_bridge/README.md](../../README.md) for full PTQ documentation.
PTQ takes **~9 min on 2 GB200 nodes (1.2 GPU-hours)** at EP=8.
**What gets quantized** (read from the exported checkpoint's `quantization_config`):
| Component | Precision | Tensors |
| --- | --- | --- |
| MoE routed experts (`down`/`gate`/`up_proj`) | **NVFP4 W4A4** (block 16) | 30,720 |
| Shared expert | **NVFP4 W4A4** | 120 |
| `lm_head` | **NVFP4 W4A4** | 1 |
| Linear attention (`in_proj_qkv`/`in_proj_z`/`out_proj`, 30 layers) | **FP8 W8A8** | 90 |
| Full attention (`q`/`k`/`v`/`o_proj`, 10 layers) | **FP8 W8A8** | 40 |
| KV cache | **FP8** (static) | — |
| MTP head, MoE router, `conv1d`, `in_proj_a`/`in_proj_b`, embeddings, vision tower | BF16 | — |
99.2% of quantized tensors are MoE experts — where the FP4 throughput win comes from. Attention stays at FP8: only 130 tensors, and numerically more sensitive.
<details>
<summary>W4A4 NVFP4 PTQ command (click to expand)</summary>
`--ep_size 8` needs 8 ranks, which is **2 GB200 nodes** (4 GPUs each) — launched with `srun`, one task per GPU. On 8-GPU nodes a single-node `torchrun --nproc_per_node 8` is equivalent.
```bash
# SBATCH --nodes=2 --ntasks-per-node=4 --gpus-per-node=4
srun ... python /opt/Model-Optimizer/examples/megatron_bridge/quantize.py \
--hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
--recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
--tp_size 1 --ep_size 8 --pp_size 1 \
--calib_dataset_name cnn_nemotron_v2_mix \
--calib_num_samples 1024 \
--calib_batch_size 1 \
--seq_length 8192 \
--export_megatron_path /path/to/qwen36_w4a4_megatron \
--skip_generate
```
Set `RANK`/`WORLD_SIZE`/`LOCAL_RANK` from `SLURM_PROCID`/`SLURM_NTASKS`/`SLURM_LOCALID` inside the `srun` body, as in the [QAD command](#3-quantization-aware-distillation-qad) below.
> [!IMPORTANT]
> Set `--ep_size` to the expert-parallel size you will run QAD at. QAD loads this checkpoint directly and **EP must match** — the expert layout is baked into the distributed checkpoint. Re-exporting at a different EP later means re-running PTQ.
</details>
---
### 3. Quantization-Aware Distillation (QAD)
QAD fine-tunes the quantized student against the BF16 teacher, so the student learns weights that survive 4-bit rounding. See the [QAD section of the Megatron-Bridge README](../../README.md#quantization-aware-distillation-qad).
Minimum hardware: the student and teacher are both resident, plus an fp32 gradient buffer and optimizer state — roughly 124 GB/GPU of the 185 GiB on a GB200 at the settings below. We used **32 nodes × 4 GB200 (128 GPUs)**; 500 iterations took **5.7 hours wall-clock (~735 GB200 GPU-hours)**. Steady-state training is ~37 s/iter (~5.1 h of the total); the remaining ~35 min is per-job startup, loading the 67 GiB teacher and the student — the run was split across two 4-hour jobs, so that cost is paid twice.
<details>
<summary>QAD command (click to expand)</summary>
Launched with `srun`, one task per GPU, across 32 nodes (128 ranks). Set `RANK` / `WORLD_SIZE` /
`LOCAL_RANK` from `SLURM_PROCID` / `SLURM_NTASKS` / `SLURM_LOCALID` in the `srun` body; `python -u`
keeps the multi-node logs unbuffered.
```bash
# SBATCH --nodes=32 --ntasks-per-node=4 --gpus-per-node=4
srun ... python -u /opt/Model-Optimizer/examples/megatron_bridge/distill.py \
--teacher_hf_path Qwen/Qwen3.6-35B-A3B \
--student_hf_path Qwen/Qwen3.6-35B-A3B \
--student_megatron_path /path/to/qwen36_w4a4_megatron \
--tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
--data_paths "${DATA_BLEND}" \
--data_path_to_cache /path/to/cache \
--seq_length 32768 \
--mbs 1 \
--gbs 512 \
--train_iters 500 \
--lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 \
--logit_kl_topk 4096 \
--recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
--no_async_save \
--eval_iters 0 \
--save_interval 50 \
--output_dir /path/to/qad_output
```
Non-default arguments:
- `--student_megatron_path` — the quantized checkpoint from Section 2; `--student_hf_path` still points at the BF16 model, which supplies the architecture.
- `--tp_size 1 --pp_size 1 --cp_size 1` — **required, not chosen** (see below). `--ep_size 8` must match the PTQ checkpoint.
- `--seq_length 32768 --gbs 512` — 16.8M tokens/iteration, 1.7B per 100 iterations.
- `--lr 1e-5 --min_lr 1e-6` — an order of magnitude below typical distillation LRs: the job is to adapt weights to quantization, not to learn the task.
- `--logit_kl_topk 4096` — restricts the KD loss to the teacher's top-4096 vocab entries. With a 248,320-token vocabulary the dense `[seq, vocab]` fp32 logits are **30.31 GiB per tensor** at 32K, which OOMs on its own.
- `--recompute_*` / `--no_async_save` / `--eval_iters 0` — all needed to fit. Async save spawns a worker needing its own CUDA context; the validation path computes full-vocab LM and MTP cross-entropy (top-k applies to training only), so eval OOMs at 32K even though training fits.
</details>
---
### 4. Export
Convert the quantized Megatron checkpoint to a deployable unified HuggingFace checkpoint (**~7 min on 1 node, 0.5 GB200 GPU-hours**). The unified exporter loads at TP=1, so use pipeline parallelism to shard across GPUs.
<details>
<summary>Export command (click to expand)</summary>
```bash
torchrun --nproc_per_node 4 /opt/Model-Optimizer/examples/megatron_bridge/export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
--megatron_path /path/to/qad_output/checkpoints/iter_0000500 \
--pp_size 4 \
--export_unified_hf_path /path/to/qwen36_w4a4_qad_hf
```
</details>
The exported checkpoint is directly deployable with [vLLM](https://github.com/vllm-project/vllm), [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) and [SGLang](https://github.com/sgl-project/sglang). Sanity-check it with `generate_vllm.py --model <path>`.
---
### 5. Evaluation
One config per benchmark in [eval_configs/](eval_configs/), so each can be launched independently with its own walltime and sampling:
| Benchmark | Config | Temp | Runs | Walltime | Metric name |
| --- | --- | --- | --- | --- | --- |
| MMMU-Pro | [mmmu_pro.yaml](eval_configs/mmmu_pro.yaml) | 1.0 | 8 | 2:30 | `mmmu-pro_pass_at_1_symbolic_correct` |
| GPQA Diamond | [gpqa.yaml](eval_configs/gpqa.yaml) | 1.0 | 1 (avg-of-16) | 4:00 | `gpqa_pass_at_1_avg-of-16_symbolic_correct` |
| SciCode (Subtask) | [scicode.yaml](eval_configs/scicode.yaml) | **0.6** | 8 | 2:00 | `scicode_pass_at_1_subtask_accuracy` |
| AA-LCR | [aa_lcr.yaml](eval_configs/aa_lcr.yaml) | 1.0 | 8 | 2:00 | `aalcr_pass_at_1_judge_correct` |
| IFBench | [ifbench.yaml](eval_configs/ifbench.yaml) | 1.0 | 8 | 1:30 | `ifbench_pass_at_1_average_score` |
| tau2-bench Telecom | [tau2_telecom.yaml](eval_configs/tau2_telecom.yaml) | 1.0 | 3 | 1:30 | `tau2_bench_telecom_pass_at_1_pass_at_1` |
Sampling follows the [nvidia/Qwen3.6-35B-A3B-NVFP4](https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4) card (`T=1.0, top_p=0.95, max_new_tokens=131072`), except SciCode, which the card specifies at `T=0.6`. `parallelism` varies (32 default, 16 for AA-LCR's long contexts and for tau2-bench's multi-turn tool round-trips, 8 for SciCode's long generations) as a throughput knob only. The same configs serve BF16 and NVFP4 unchanged — vLLM reads the FP8 KV-cache setting from the checkpoint's own `quantization_config`.
> [!IMPORTANT]
> **Two benchmarks need a second model endpoint besides the one under test.** AA-LCR scores answers
> with a judge (`NS_JUDGE_URL` / `LCR_JUDGE_MODEL_ID`); tau2-bench drives a simulated user
> (`TAU2_ENDPOINT_URL` / `TAU2_USER_MODEL_ID`). Both read `INFERENCE_API_KEY`, and neither can
> produce its metric without that endpoint. Point both at `https://` URLs for the final endpoint:
> the key rides every request, and a redirect can downgrade it to `http` or send it to another
> origin. tau2-bench additionally needs its own **deployment** —
> `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and `--max-num-seqs 16` — which is why
> it cannot share a config with the others (`deployment.command` is global). `qwen3_coder` is what
> the model card specifies; the template's `<tool_call>` markers resemble `hermes`, which would
> mis-parse tool calls and silently invalidate the benchmark. Its `top_k` / `presence_penalty` go
> through tau2's `agent_args` passthrough, which NEL's top-level `params` block does not accept —
> that is why the other five configs omit them.
> [!IMPORTANT]
> Every task sets `num_repeats: 1`; the repeat counts above come from **launching a config that many times**. N launches give N independent `pass@1` values, which is what `mean ± sem` and the paired tests need — `num_repeats: N` instead yields a single `pass@1[avg-of-N]` with no spread. GPQA is the deliberate exception.
> [!IMPORTANT]
> **Run enough repeats.** These benchmarks are noisy, and 3 was not enough to tell signal from noise in either direction. At 3 repeats, AA-LCR (only 100 questions) showed a 3 pp swing that looked real and **vanished completely** at 8; IFBench's genuine +2.3 pp gain was **invisible** at 3 and only appeared at 8. Three runs also *understates* how noisy a benchmark is, so the measured spread looks tighter than it is.
<details>
<summary>Evaluation launch steps (click to expand)</summary>
Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` in the config, or override with `-o <option>=<value>`.
```bash
pip install "nemo-evaluator-launcher[all]==0.2.6"
# Required by every config:
export HF_TOKEN=<your_huggingface_token>
export SLURM_JOB_DIR=<path_to_slurm_job_output_dir>
# Required only by AA-LCR (judge) and tau2-bench (user simulator):
export INFERENCE_API_KEY=<key_for_those_endpoints>
export NS_JUDGE_URL=<judge_endpoint_url>
export LCR_JUDGE_MODEL_ID=<judge_model_id>
export TAU2_ENDPOINT_URL=<user_simulator_url>
export TAU2_USER_MODEL_ID=<user_simulator_model>
# One benchmark. To verify the pipeline first, add
# `-o ++evaluation.nemo_evaluator_config.config.params.limit_samples=8`
nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
# The 8 independent runs behind the results table: launch the same config 8 times and pool the
# per-run pass@1 values. Each launch returns its own invocation id; record them so you can select
# exactly those runs when pooling (each eval directory also holds a small canary run).
for i in $(seq 1 8); do
nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
done
# All six benchmarks at their reported repeat counts
for cfg_runs in mmmu_pro:8 gpqa:1 scicode:8 aa_lcr:8 ifbench:8 tau2_telecom:3; do
cfg=${cfg_runs%%:*}; runs=${cfg_runs##*:}
for i in $(seq 1 "$runs"); do
nemo-evaluator-launcher run --config "eval_configs/${cfg}.yaml"
done
done
```
</details>
For more details on NeMo Evaluator, see the [GitHub repo](https://github.com/NVIDIA-NeMo/evaluator) and [documentation](https://docs.nvidia.com/nemo/evaluator/latest/).
#### Output length
Accuracy is only half the serving cost — a model that scores the same while emitting more tokens is slower end to end. Mean completion tokens per benchmark (`response_stats.avg_completion_tokens`, same runs as the results table):
| Benchmark | BF16 | W4A16 NVFP4 | W4A4 NVFP4 PTQ | + QAD 500 iters | runs/side |
| --- | --- | --- | --- | --- | --- |
| SciCode (Subtask) | 5,348 | +22.9% | +25.5% | **+90.9%** | 8 |
| IFBench | 12,177 | +7.0% | +9.3% | +17.4% | 8 |
| AA-LCR | 2,560 | +7.1% | +6.9% | +9.8% | 8 |
| MMMU-Pro | 9,382 | +3.0% | +6.2% | −3.8% | 8 |
| GPQA Diamond | 13,561 | +17.1% | +6.6% | −8.5% | 1 |
**4-bit weights lengthen SciCode outputs by ~23-25% on their own** — the published W4A16 checkpoint does it too, so it is not something QAD or W4A4 introduced. **QAD then pushes SciCode to +90.9%**, for an unchanged score (40.2 vs 39.9), while pulling GPQA and MMMU-Pro back toward BF16. The throughput table is measured at fixed output length, so it does not capture this.
<details>
<summary><b>What the SciCode number actually is — worth reading before running QAD on another model</b></summary>
It is not verbosity. It is a **failure to terminate on a small fraction of sub-steps**:
- Sub-steps that hit the 131,072-token cap inside `<think>` go from **0.7% (20/2704, BF16) to 3.6% (96/2704, QAD 500)**. Almost all return **zero answer tokens** (93 of those 96) — the model writes a complete solution, says *"I think I've been going in circles"*, and writes it again. In the case we inspected, a 20-word window repeats **352 times** and 97.7% of the trace's 20-word windows are duplicates.
- Those 3.6% of sub-steps burn **45.7% of all completion tokens**, so they dominate the mean: excluding them it is **+30.4%** rather than +90.9%.
- The **median** also roughly doubles (+91.6%), so the whole distribution shifted right — this is not *only* a tail effect.
- Capped rate peaks at **iteration 50** (4.3%) and settles at 3.1% / 3.6% by 300 / 500; it is not gradual drift.
The obvious suspect — that `--logit_kl_topk 4096` leaves the stop tokens outside the loss — **did not hold up**. Probing the BF16 teacher over one runaway trace: `</think>` does fall outside top-4096 at 35% of positions overall, but *in the looping region* the teacher gives `<|im_end|>` a median rank of **5** and `</think>` ~570, both well inside top-k. The teacher is signalling "stop here" at positions the loss did cover, and the student still does not stop. More likely: the blend has few "the answer is written, now stop" positions in this style, and a teacher-forced loss never exercises free-running generation 10K+ tokens deep.
**For the next QAD run**, three things follow: track the **length-capped rate** as a first-class metric alongside accuracy (a benchmark score can stay flat while 3.6% of responses return nothing); consider **top-p instead of top-k** for the KD loss so coverage adapts to the teacher's entropy rather than a fixed rank; and if memory allows, **full-vocab KL** — at 32K on this 248,320-token vocabulary the dense fp32 logits are 30.31 GiB per tensor, which is why top-k was used here, but more GPU memory or a smaller model or shorter sequence may afford it.
</details>
---
### 6. vLLM Inference Benchmarking
Throughput was measured with [AIPerf](https://github.com/ai-dynamo/aiperf) against a served vLLM endpoint on 4× GB200 — three ISL/OSL shapes × four concurrencies = the 12 shapes in the [results table](#vllm-throughput-4-gb200-vllm-0280-tp4--ep).
<details>
<summary>Serve + benchmark commands (click to expand)</summary>
The same `vllm serve` command works for BF16 and NVFP4; the quantized checkpoint carries its own format and FP8 KV-cache settings in `quantization_config`.
```bash
vllm serve <checkpoint_path> --served-model-name bench \
--host 127.0.0.1 --port 8000 \
--tensor-parallel-size 4 --data-parallel-size 1 --enable-expert-parallel \
--max-model-len 262144 --reasoning-parser qwen3 \
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}' \
--max-num-batched-tokens 8192 --enable-chunked-prefill
# decode-bound, chat, and prefill-bound shapes at four concurrencies each
for shape in "128 2048 decode" "8000 1000 chat" "32000 400 prefill"; do
set -- $shape; ISL=$1; OSL=$2; NAME=$3
for C in 1 8 32 128; do
aiperf profile -m bench --endpoint-type chat --streaming -u localhost:8000 \
--synthetic-input-tokens-mean $ISL --output-tokens-mean $OSL \
--concurrency $C --request-count $(( C * 5 )) \
--tokenizer <checkpoint_path> \
--extra-inputs ignore_eos:true --random-seed 42 \
--artifact-dir bench/${NAME}_isl${ISL}_osl${OSL}/c${C}
done
done
```
</details>
> [!TIP]
> Check the endpoint answers a known question correctly before benchmarking it. A checkpoint served with the wrong KV dtype or an unexpected kernel still produces tokens at a plausible rate — the throughput number looks fine and is meaningless. We gate every benchmark on a short prompt with a verifiable answer.
> [!TIP]
> To deploy the model with vLLM, refer to the [vLLM Quickstart documentation](https://docs.vllm.ai/en/stable/getting_started/quickstart/).
---
## Potential further improvements
This is a single 500-iteration run at 32K on a text-only blend. Each of those three choices is an unexplored lever.
**1. Continue QAD longer.** The recovery had not saturated: IFBench's gain arrived *late* (−0.11 pp vs the student at iteration 300, +2.33 pp at 500) and MMMU-Pro was still climbing (+0.51 pp, p=0.078). Training also used only **8.4B of the blend's 17.3B tokens**, so it can roughly double before repeating a sample. At ~735 GPU-hours per 500 iterations this is the most direct experiment, and IFBench + MMMU-Pro alone (~55 GPU-hours) are enough to read the result.
**2. Longer sequence length (64K).** The model serves at 262K context, so 32K trains on a fraction of it; AA-LCR, the one long-context benchmark, showed no movement either way. At 32K the run already needs top-k KD plus full recompute, and memory is dominated by the resident student + teacher + fp32 gradient buffer rather than by sequence length — so 64K needs a different memory lever.
**3. A better data blend.** The clearest signal in the study: **MMMU-Pro is multimodal, the blend is text-only, and MMMU-Pro is the one deficit QAD did not close.** Since QAD supports language-model distillation with text-only data, recovering a multimodal benchmark that way asks the method to do something it is not set up for. Agentic coverage is also thin (agent/SWE splits under 9%, and tau2-bench showed the transient 8× `agent_error` spike), and weights here are simply proportional to token count — contrast the [Nemotron-3-Nano blend](../NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md#data-blend), which is deliberately designed against target benchmarks and [ablated](../NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/ABLATIONS.md#effect-of-data-blend-tool_calling).
If you run only one, run **(1)**: no engineering risk, the trajectory says the gain is still arriving, and it produces the evidence to judge whether (2) and (3) are worth their cost.
@@ -0,0 +1,51 @@
# Token-budgeted data blend for QAD on Qwen3.6-35B-A3B.
# Consumed by the token-budgeted blend workflow in examples/dataset/MEGATRON_DATA_PREP.md
# ("Prepare token-budgeted data blends"). Weights are each split's share of the dataset's own
# token count, i.e. a single pass over the natural mixture rather than a tuned blend.
#
# QAD supports language-model distillation with text-only data, so this blend is SFT text only.
tokenizer: Qwen/Qwen3.6-35B-A3B
output_dir: /path/to/tokenized_cascade2_qwen36
target_tokens: 17_300_000_000
sources:
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: chat
split: train
content_field: messages
weight: 56.2
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: math
split: train
content_field: messages
weight: 21.1
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: science
split: train
content_field: messages
weight: 10.9
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: instruction_following
split: train
content_field: messages
weight: 3.3
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: conversational_agent
split: train
content_field: messages
weight: 3.3
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: terminal_agent
split: train
content_field: messages
weight: 3.3
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: swe
split: train
content_field: messages
weight: 1.8
- hf_dataset: nvidia/Nemotron-Cascade-2-SFT-Data
config: safety
split: train
content_field: messages
weight: 0.02
@@ -0,0 +1,110 @@
# AA-LCR eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=1.0, top_p=0.95, max_new_tokens=131072.
# Long-context, only 100 questions -- the noisiest benchmark here.
# JUDGE-SCORED: needs a separate judge endpoint (NS_JUDGE_URL / LCR_JUDGE_MODEL_ID /
# INFERENCE_API_KEY). Without it the `judge_correct` metric cannot be produced. NS_JUDGE_URL must
# be an https:// URL for the final endpoint -- INFERENCE_API_KEY rides every judge request, and a
# redirect can downgrade it to http or send it to another origin.
#
# Results use 8 independent run(s): launch this config 8 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/aa_lcr.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 02:00:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--max-num-seqs 32
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
# The judge is a separate endpoint and needs its own key.
INFERENCE_API_KEY: host:INFERENCE_API_KEY
nemo_evaluator_config:
config:
params:
parallelism: 16
request_timeout: 3600
max_retries: 10
max_new_tokens: 131072
temperature: 1.0
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: ns_aa_lcr
env_vars:
HF_TOKEN: host:HF_TOKEN
INFERENCE_API_KEY: host:INFERENCE_API_KEY
nemo_evaluator_config:
config:
params:
extra:
num_repeats: 1
# AA-LCR scores answers with a judge model behind its own endpoint.
judge_support: true
judge:
url: ${oc.env:NS_JUDGE_URL}
model_id: ${oc.env:LCR_JUDGE_MODEL_ID}
api_key: INFERENCE_API_KEY
@@ -0,0 +1,96 @@
# GPQA Diamond eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=1.0, top_p=0.95, max_new_tokens=131072.
# num_repeats: 16 in a SINGLE run -- that is what `pass@1[avg-of-16]` means, so GPQA is the one
# benchmark reported without a +/- figure. Longest task at ~2.5 h.
#
# Results use 1 independent run(s): launch this config 1 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/gpqa.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 04:00:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--max-num-seqs 32
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
nemo_evaluator_config:
config:
params:
parallelism: 32
request_timeout: 3600
max_retries: 10
max_new_tokens: 131072
temperature: 1.0
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: ns_gpqa
nemo_evaluator_config:
config:
params:
extra:
num_repeats: 16
@@ -0,0 +1,94 @@
# IFBench eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=1.0, top_p=0.95, max_new_tokens=131072.
#
# Results use 8 independent run(s): launch this config 8 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/ifbench.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 01:30:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--max-num-seqs 32
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
nemo_evaluator_config:
config:
params:
parallelism: 32
request_timeout: 3600
max_retries: 10
max_new_tokens: 131072
temperature: 1.0
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: ns_ifbench
nemo_evaluator_config:
config:
params:
extra:
num_repeats: 1
@@ -0,0 +1,94 @@
# MMMU-Pro eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=1.0, top_p=0.95, max_new_tokens=131072.
#
# Results use 8 independent run(s): launch this config 8 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 02:30:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--max-num-seqs 32
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
nemo_evaluator_config:
config:
params:
parallelism: 32
request_timeout: 3600
max_retries: 10
max_new_tokens: 131072
temperature: 1.0
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: ns_mmmu_pro
nemo_evaluator_config:
config:
params:
extra:
num_repeats: 1
@@ -0,0 +1,95 @@
# SciCode eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=0.6, top_p=0.95, max_new_tokens=131072.
# T=0.6 per the model card for this task (all others use 1.0); lower parallelism for long generations.
#
# Results use 8 independent run(s): launch this config 8 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/scicode.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 02:00:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--max-num-seqs 32
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
nemo_evaluator_config:
config:
params:
parallelism: 8
request_timeout: 3600
max_retries: 10
max_new_tokens: 131072
temperature: 0.6
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: ns_scicode
nemo_evaluator_config:
config:
params:
extra:
num_repeats: 1
@@ -0,0 +1,131 @@
# tau2-bench Telecom eval for Qwen3.6-35B-A3B. See README section 5 for the task/metric table.
#
# Set `execution.hostname`, `execution.account` and `deployment.checkpoint_path` (BF16, W4A4 PTQ or QAD)
# before running. Serves BF16 and NVFP4 unchanged: vLLM reads the FP8 KV-cache setting from the
# checkpoint's own `quantization_config`. NVFP4 needs a Blackwell GPU.
#
# Sampling per the nvidia/Qwen3.6-35B-A3B-NVFP4 card: T=1.0, top_p=0.95, max_new_tokens=131072.
# AGENTIC. Needs a separate USER-SIMULATOR endpoint (tau2 drives a simulated user with a second
# model): set TAU2_ENDPOINT_URL / TAU2_USER_MODEL_ID / INFERENCE_API_KEY. We used Qwen-235B.
# TAU2_ENDPOINT_URL must be an https:// URL for the final endpoint -- INFERENCE_API_KEY rides every
# request, and a redirect can downgrade it to http or send it to another origin.
#
# Needs --enable-auto-tool-choice --tool-call-parser qwen3_coder in the deployment, which is
# why it cannot share a config with the others (`deployment.command` is global). `qwen3_coder` is what
# the Qwen3.6-35B-A3B card specifies; the template's <tool_call> markers resemble `hermes`, which would
# mis-parse tool calls and silently invalidate the benchmark. Task-level token/parallelism budgets are
# lower than this file's own defaults because each task holds a multi-turn conversation plus tool
# round-trips open. `top_k`/`presence_penalty` go through tau2's `agent_args` passthrough -- NEL's
# top-level params block does not accept them, which is why the other five configs omit them.
#
# Results use 3 independent run(s): launch this config 3 times and pool the per-run pass@1.
#
# pip install "nemo-evaluator-launcher[all]==0.2.6"
# export HF_TOKEN=<token>; export SLURM_JOB_DIR=<dir>; export HF_HOME=<cache>
# nemo-evaluator-launcher run --config eval_configs/tau2_telecom.yaml
defaults:
- execution: slurm/default
- deployment: vllm
- _self_
execution:
type: slurm
hostname: ???
username: ${oc.env:USER}
account: ???
partition: batch
num_nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
gres: "gpu:4"
walltime: 01:30:00
output_dir: ${oc.env:SLURM_JOB_DIR}
mode: sequential
mounts:
mount_home: false
deployment:
checkpoint_path: ???
served_model_name: qwen3.6-35b-a3b
port: 8000
image: vllm/vllm-openai:v0.28.0
env_vars:
HF_TOKEN: host:HF_TOKEN
command: >-
vllm serve /checkpoint
--served-model-name ${deployment.served_model_name}
--host 0.0.0.0
--port ${deployment.port}
--tensor-parallel-size 4
--data-parallel-size 1
--enable-expert-parallel
--max-model-len 262144
--reasoning-parser qwen3
--model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}'
--max-num-batched-tokens 8192
--enable-chunked-prefill
--enable-auto-tool-choice
--tool-call-parser qwen3_coder
--max-num-seqs 16
endpoints:
chat: /v1/chat/completions
completions: /v1/completions
health: /health
evaluation:
env_vars:
HF_TOKEN: host:HF_TOKEN
DUMMY_API_KEY: lit:dummy
# The user simulator is a separate endpoint, so it needs its own key.
INFERENCE_API_KEY: host:INFERENCE_API_KEY
nemo_evaluator_config:
config:
params:
parallelism: 64
request_timeout: 3600
max_retries: 50
max_new_tokens: 131072
temperature: 1.0
top_p: 0.95
target:
api_endpoint:
api_key_name: DUMMY_API_KEY
adapter_config:
use_caching: true
use_reasoning: true
log_failed_requests: true
use_request_logging: true
max_logged_requests: 2
use_response_logging: true
max_logged_responses: 2
params_to_add:
chat_template_kwargs:
enable_thinking: true
tasks:
- name: tau2_bench_telecom
env_vars:
HF_TOKEN: host:HF_TOKEN
INFERENCE_API_KEY: host:INFERENCE_API_KEY
nemo_evaluator_config:
config:
params:
max_new_tokens: 65536
parallelism: 16
extra:
# 114 tasks x n_samples trials per run.
n_samples: 3
skip_failed_samples: true
# tau2's --agent-llm-args template has an agent_args passthrough, so the Qwen3.6
# card's thinking-mode top_k / presence_penalty can be set here (NEL's top-level
# params block cannot carry them).
agent_args:
top_k: 20
presence_penalty: 1.5
# The simulated user is a second model behind its own endpoint.
user:
url: ${oc.env:TAU2_ENDPOINT_URL}
model_id: ${oc.env:TAU2_USER_MODEL_ID}
api_key: INFERENCE_API_KEY
judge:
enabled: false
num_repeats: 1
Binary file not shown.

After

Width:  |  Height:  |  Size: 240 KiB

@@ -8,3 +8,4 @@ Each one walks through a complete workflow using the scripts in [examples/megatr
| Tutorial | What it covers |
| --- | --- |
| [NVIDIA-Nemotron-3-Nano-30B-A3B-BF16](NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) | End-to-end optimization of the Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) model: Minitron structured pruning (31.6B/A3.6B → 22B/A3.0B) → two-phase knowledge distillation (100B tokens, 8K then 32K seq length) → quantization → vLLM deployment. Includes data-blend preparation, evaluation setup, and detailed pruning / data-blend / long-context ablations. |
| [Qwen3.6-35B-A3B](Qwen3.6-35B-A3B/README.md) | End-to-end **W4A4 NVFP4** optimization of the Qwen3.6-35B-A3B (MoE + hybrid linear/full attention VLM) model: NVFP4 W4A4 post-training quantization → quantization-aware distillation (QAD) to recover the accuracy W4A4 costs → evaluation → vLLM deployment. |
@@ -125,7 +125,10 @@ class GPTModelExporter:
eagle_module. Otherwise, only export the base model.
dtype: The weights data type to export the unquantized layers.
trust_remote_code: Whether to trust remote code in the HuggingFace pretrained model.
moe_router_dtype: The data type of the MoE router. Can be "fp32", "fp64", or None (default to the model dtype).
moe_router_dtype: Storage dtype override for the exported MoE router weight, "fp32",
"fp64" or None to store it at ``dtype`` like every other unquantized weight.
This is not Megatron's ``moe_router_dtype``, which is a routing *compute* dtype;
HF checkpoints conventionally store the router at the model dtype.
clamp_kv_cache_scales: Whether to clamp FP8 KV cache scaling factors to at least 1.0.
"""
@@ -551,8 +554,9 @@ class GPTModelExporter:
def _get_state_dict(self):
model = self.model
# Embedding
if hasattr(model, "embedding"):
# Embedding. MCore also builds `embedding` on the MTP stage, so gating on hasattr
# alone would emit a second, orphaned copy of the vocab embedding when PP > 1.
if model.pre_process and hasattr(model, "embedding"):
self.rules["word_embeddings"](model.embedding.word_embeddings)
# Decoder layers
@@ -0,0 +1,142 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# W4A4 NVFP4 for Qwen3.6-MoE under **Megatron-Core** (examples/megatron_bridge/quantize.py).
# Used as the QAD student in examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B.
#
# This is architecture-level, not a mirror of a released checkpoint, so it lives under
# model_type/<model_type>/ rather than models/<org>/<checkpoint>/. It is the W4A4 counterpart of
# model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast, which covers the same family for
# the HuggingFace flow.
#
# `_mcore` in the filename is load-bearing: the selectors below are **Megatron-Core leaf names**,
# so nothing matches under hf_ptq.py, where `mtq.quantize` then raises because no weight-quantizer
# pattern matched. Use the qwen3_5_moe w4a16 recipe for that flow.
#
# Qwen3.6-35B-A3B is a VLM whose language model is a 40-layer hybrid: 30 GatedDeltaNet
# (linear-attention) layers interleaved with 10 full-attention layers (indices 3, 7, ... 39),
# 256 routed experts plus a shared expert, and an MTP head.
#
# What this quantizes:
# routed + shared experts NVFP4 W4A4 (block 16, dynamic input scales)
# output layer (lm_head) NVFP4 W4A4
# self-attention FP8 W8A8
# linear-attention FP8 W8A8 (in_proj / out_proj)
# KV cache FP8 (cast mode)
# MTP, routers, conv1d, vision tower, embeddings -- BF16
#
# W4A4 rather than W4A16 because a BF16 activation forces vLLM onto the Marlin dequant fallback,
# which measured 0.64-0.86x BF16 throughput; only FP4 activations reach the Blackwell FP4 tensor
# cores. lm_head is included because it is 34.6% of per-token weight traffic.
#
# Megatron-Core fuses several projections that are separate in HuggingFace, so the selectors below
# are Megatron leaf names: `mlp.experts.linear_fc{1,2}` (TEGrouped), `shared_experts.linear_fc1`
# (gate+up), `self_attention.linear_qkv` (q+k+v). Both attention kinds are called `self_attention`,
# so they are told apart by leaf name: `in_proj`/`out_proj`/`conv1d` vs `linear_qkv`/`linear_proj`.
#
# `in_proj` fuses HF's [query, key, value, z, beta, alpha], so the reference recipe's in_proj_a /
# in_proj_b BF16 exclusion cannot be expressed here; `_gated_delta_net_slicing` in
# modelopt/torch/export/unified_export_megatron.py re-splits it on export and writes those two
# out in BF16.
imports:
base_disable_all: configs/ptq/units/base_disable_all
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
fp8: configs/numerics/fp8
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
nvfp4: configs/numerics/nvfp4
metadata:
recipe_type: ptq
description: >-
Megatron-Core W4A4 (NVFP4 weights and activations, block 16) on MoE routed experts, shared
experts and dense MLP; W4A4 NVFP4 on the output layer; FP8 W8A8 on softmax-attention and
GatedDeltaNet projections; FP8 KV cache with constant amax; MTP block, routers, conv1d and
the vision tower left in BF16. Max calibration.
quantize:
algorithm: max
quant_cfg:
- $import: base_disable_all
# ---- W4A4 NVFP4 (weights AND activations) ------------------------------
# Two expert layouts must both match: TEGroupedMLP (`mlp.experts.linear_fc1.weight_quantizer_<N>`)
# and SequentialMLP (`mlp.experts.local_experts.<N>.linear_fc1.weight_quantizer`). The wildcard
# between `experts.` and `linear_fc` covers the `local_experts.<N>.` segment; without it the
# SequentialMLP layout silently matches nothing.
- quantizer_name: '*mlp.experts.*linear_fc1*weight_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.experts.*linear_fc1*input_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.experts.*linear_fc2*weight_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.experts.*linear_fc2*input_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.shared_experts.linear_fc1*weight_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.shared_experts.linear_fc1*input_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.shared_experts.linear_fc2*weight_quantizer*'
cfg: {$import: nvfp4}
- quantizer_name: '*mlp.shared_experts.linear_fc2*input_quantizer*'
cfg: {$import: nvfp4}
# Dense MLP: no-op on this MoE model, kept so the file also covers dense smoke models.
- quantizer_name: '*.mlp.linear_fc1.weight_quantizer'
cfg: {$import: nvfp4}
- quantizer_name: '*.mlp.linear_fc1.input_quantizer'
cfg: {$import: nvfp4}
- quantizer_name: '*.mlp.linear_fc2.weight_quantizer'
cfg: {$import: nvfp4}
- quantizer_name: '*.mlp.linear_fc2.input_quantizer'
cfg: {$import: nvfp4}
# ---- FP8 W8A8 softmax attention (layers 3, 7, 11, ...) -----------------
- quantizer_name: '*self_attention.linear_qkv.weight_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.linear_qkv.input_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.linear_proj.weight_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.linear_proj.input_quantizer'
cfg: {$import: fp8}
# ---- FP8 W8A8 GatedDeltaNet linear attention (all other layers) --------
- quantizer_name: '*self_attention.in_proj.weight_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.in_proj.input_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.out_proj.weight_quantizer'
cfg: {$import: fp8}
- quantizer_name: '*self_attention.out_proj.input_quantizer'
cfg: {$import: fp8}
# ---- FP8 KV cache, constant amax --------------------------------------
- $import: kv_fp8_cast
# ---- Standard exclusions ----------------------------------------------
# Disables routers, conv1d, output_layer, embeddings and vision-tower patterns.
# output_layer is re-enabled below.
- $import: default_disabled_quantizers
# Keep the whole MTP subtree in BF16. Must follow the expert / attention selectors above,
# which would otherwise match its inner layers.
- {quantizer_name: '*mtp*', enable: false}
# ---- W4A4 NVFP4 on the output layer (HF `lm_head`) --------------------
# Must follow default_disabled_quantizers, which disables `*output_layer*`.
- quantizer_name: '*output_layer.weight_quantizer'
cfg: {$import: nvfp4}
- quantizer_name: '*output_layer.input_quantizer'
cfg: {$import: nvfp4}
+12 -2
View File
@@ -273,7 +273,7 @@ that baseline. The deviations come in four kinds:
| Kind | What changes vs. the general recipe | Examples |
|------|-------------------------------------|----------|
| **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama` |
| **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `qwen3_6_moe`, `vit`, `nemotron_llama` |
| **Algorithm override** | Same numerics & scope, but the *calibration algorithm* is tweaked because the default breaks or regresses | `gemma`, `gemma4`, `mpt` |
| **Extra exclusions** | Adds disabled-quantizer patterns so non-language branches stay full precision | `nemotron_vl`, `diffusion_gemma` |
| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/NVIDIA-Nemotron-3-*`, `models/mistralai/Mistral-Medium-3.5-128B` |
@@ -282,7 +282,7 @@ The numerics and standard exclusions are still inherited from `configs/`
wherever possible — the model folder captures *only* the delta. Each `<task>/`
folder may carry a `README.md` spelling out that delta.
### Architecture-aware `quant_cfg` — `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama`
### Architecture-aware `quant_cfg` — `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `qwen3_6_moe`, `vit`, `nemotron_llama`
**`minimax_m3_vl/ptq/mxfp8_nvfp4_experts`** applies MXFP8 to the language-model
linear layers and MSE-calibrated NVFP4 to routed experts, with expert
@@ -353,6 +353,16 @@ checkpoint with `quant_algo: null`. These select `*moe*` instead and disable the
router (`moe.gate`) and `share_expert` on top. Use them, not the general
recipes, for Step-3.7 checkpoints; Step-3.5 has its own recipe above.
- **`qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore`** is the **Megatron-Core** W4A4
counterpart of `qwen3_5_moe`'s W4A16 recipe, used as the QAD student in the
[Qwen3.6-35B-A3B tutorial](../examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/README.md).
MoE routed + shared experts and `lm_head` → NVFP4 W4A4 (block 16); softmax attention and the
GatedDeltaNet `in_proj`/`out_proj` → FP8 W8A8; KV cache → FP8 cast; MTP block, routers, `conv1d`,
the vision tower and embeddings stay BF16. W4A4 rather than W4A16 because a BF16 activation keeps
vLLM on the Marlin dequant fallback, which measured *slower* than BF16. The `_mcore` suffix is
load-bearing: selectors are Megatron-Core leaf names (`mlp.experts.linear_fc1`,
`self_attention.linear_qkv`), so under `hf_ptq.py` nothing matches and `mtq.quantize` raises.
### Algorithm overrides — `gemma`, `gemma4`, `mpt`
These quantize the **same layers** as the general recipes; only the
@@ -532,6 +532,48 @@ def test_unified_export_megatron_pp2_mtp_metadata_matches_shards(dist_workers_si
)
def _test_export_pp2_mtp_no_duplicate_tensors(tmp_path, model_dir, rank, size):
"""MCore builds an embedding on the MTP stage too; export must still write it exactly once."""
config = transformers.AutoConfig.from_pretrained(model_dir)
model = get_mcore_gpt_model(
tensor_model_parallel_size=1,
pipeline_model_parallel_size=size,
initialize_megatron=True,
num_layers=config.num_hidden_layers,
hidden_size=config.hidden_size,
num_attention_heads=config.num_attention_heads,
num_query_groups=config.num_key_value_heads,
ffn_hidden_size=config.intermediate_size,
max_sequence_length=config.max_position_embeddings,
vocab_size=config.vocab_size,
activation_func="swiglu",
normalization="RMSNorm",
transformer_impl="modelopt",
mtp_num_layers=1,
).cuda()
export_dir = tmp_path / "export_pp2_mtp_dedup"
export_mcore_gpt_to_hf(model, model_dir, dtype=torch.bfloat16, export_dir=str(export_dir))
if rank == 0:
shard_of = {}
duplicated = {}
for shard in sorted(export_dir.glob("model-*.safetensors")):
with safe_open(str(shard), framework="pt", device="cpu") as sf:
for key in sf.keys(): # noqa: SIM118
if key in shard_of:
duplicated[key] = (shard_of[key], shard.name)
shard_of[key] = shard.name
assert not duplicated, f"tensors written to more than one shard: {duplicated}"
assert "model.embed_tokens.weight" in shard_of
def test_unified_export_megatron_pp2_mtp_no_duplicate_tensors(dist_workers_size_2, tmp_path):
model_dir = create_tiny_llama_dir(tmp_path)
dist_workers_size_2.run(partial(_test_export_pp2_mtp_no_duplicate_tensors, tmp_path, model_dir))
def test_qkv_slicing_records_hf_excludes_for_unquantized_fused_qkv():
"""Unquantized fused MCore linear_qkv should become HF q/k/v excludes."""
exporter = object.__new__(GPTModelExporter)