Files
Keval MorabiaandClaude Opus 5 9c1cf80f1b test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do?

Type of change: new tests

Runs the VLM case of `test_qad` under context parallelism, so QAD on a
Qwen3-VL model is covered on the path that until now could not run at
all.

`Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the
batch has to hand it full-length `position_ids`. Megatron-Bridge's
`get_batch` was sharding them too, leaving the rotary embedding at `seq
/ cp**2` against hidden states at `seq / cp`. The fix is upstream in
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243);
this PR is the coverage that would have caught it.

The VLM case moves from tensor to context parallelism and the PTQ step
is sized to the same TP, so QAD still loads a matching checkpoint. The
LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and
`test_distill_vlm` still covers TP for a VLM, so nothing loses coverage.

### Usage

```bash
# Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243
# it now works for VLMs, where it previously died in the rotary embedding.
python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ...
```

### Testing

On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on
`PYTHONPATH`:

- `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD
across 2 CP ranks, export; quantizers survive and the vision tower is
byte-identical. **1 passed (183 s).** Without the upstream fix the same
run dies with `AttributeError: 'NoneType' object has no attribute
'ndim'` in `rope.py:175`.
- `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).**
- Gate check: on today's `nemo:26.08` (no #6243) the probe resolves
`False` and the VLM case stays at `cp_size=1`, byte-identical to current
CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a
1-GPU runner it degenerates to today's config either way.
- `pre-commit run --files ...` clean (ruff check, ruff format, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests extended
rather than new ones added.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test coverage and one doc line; no feature, break, deprecation, or
fix for a released bug.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

- Depends on
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243).
Safe to merge before it lands: the gate keeps the VLM case at
`cp_size=1` until a container ships the fix, at which point the coverage
switches on by itself.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the Qwen3.6 QAD instructions to keep tensor and pipeline
parallelism set to 1, while allowing context parallelism to increase for
longer sequences with the `nemo:26.10` container.
* **Tests**
* QAD validation now selects parallelism settings based on whether the
Megatron-Bridge context-parallel fix is available, and reports when
multi-GPU VLM coverage is reduced.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-01 23:05:55 +05:30

25 KiB
Raw Permalink Blame History

Qwen3.6-35B-A3B: W4A4 NVFP4 PTQ + QAD + vLLM Deployment

End-to-end W4A4 optimization of Qwen/Qwen3.6-35B-A3B demonstrating how to push past weight-only quantization: NVFP4 W4A4 post-training quantization → Quantization-Aware Distillation (QAD) to recover the accuracy it costs → evaluation benchmarking → vLLM deployment. This document covers:

  1. Data Preparation — tokenizing the SFT blend for distillation
  2. Quantization — W4A4 NVFP4 PTQ with examples/megatron_bridge/quantize.py
  3. QAD — recovering accuracy with examples/megatron_bridge/distill.py
  4. Export — converting the quantized Megatron checkpoint to a deployable HF checkpoint
  5. Evaluation — benchmarking with NeMo Evaluator across MMMU-Pro, GPQA Diamond, SciCode, and more
  6. vLLM Inference Benchmarking — throughput comparison against BF16 on GB200

It closes with potential further improvements — the levers this run did not explore, and what the evidence says each might buy.

Note

This tutorial complements the Nemotron-3-Nano-30B-A3B tutorial, which covers pruning + distillation + FP8 PTQ. Here the starting point is an unpruned model and the technique under test is W4A4 — aggressive enough that PTQ alone leaves a measurable accuracy gap, which is exactly what QAD exists to close.

Results

W4A4 NVFP4 accuracy recovery during QAD

Main results — all models evaluated with the same setup. Values are mean ± sem across repeats; intermediate QAD checkpoints are in the figure above.

Model MMMU-Pro GPQA Diamond SciCode (Subtask) AA-LCR IFBench tau2-bench Telecom Average
BF16 (teacher) 74.6 ± 0.2 84.7 39.9 ± 0.6 69.1 ± 0.8 60.0 ± 0.5 94.2 ± 1.0 70.4
W4A16 NVFP4 PTQ (published)1 73.8 ± 0.2 84.4 38.5 ± 0.5 66.3 ± 1.0 59.4 ± 0.5 96.4 ± 0.4 69.8
W4A4 NVFP4 PTQ (QAD student) 73.4 ± 0.2 84.7 39.1 ± 0.7 70.0 ± 1.1 57.9 ± 0.5 94.2 ± 1.2 69.9
↳ + QAD 500 iters 73.9 ± 0.2 84.2 40.2 ± 0.6 69.4 ± 1.5 59.6 ± 0.5 93.4 ± 0.4 70.1

1 The W4A16 NVFP4 row is the published checkpoint re-evaluated here with the same evaluation harness.

Where QAD actually helps

W4A4 quantization only measurably hurt two of the six benchmarks, so those are the only two where QAD has anything to recover:

Benchmark W4A4 PTQ vs BF16 After 500 QAD iters Outcome
IFBench −2.6 pp (real loss) −0.3 pp ✅ recovered — back to BF16
MMMU-Pro −1.2 pp (real loss) −0.7 pp ⚠️ about 40% recovered, gap remains
GPQA Diamond, SciCode, AA-LCR, tau2-bench no loss no change — nothing to recover

"Real loss" means the gap is larger than the run-to-run noise, so it is not a measurement artifact. The four benchmarks in the last row land within ±1.2 pp of BF16 both before and after QAD, which is inside their noise — W4A4 is effectively lossless there.

Statistical detail (click to expand)

Comparisons are paired per-question t-tests: the same question answered by both models is compared directly, which cancels question-difficulty variance and is far more sensitive than comparing run averages. p is the probability of seeing a gap this large if the two models were actually equal; below 0.05 is conventionally "real".

Benchmark PTQ vs BF16 QAD 500 vs PTQ QAD 500 vs BF16
MMMU-Pro −1.23 (p=0.0002) +0.51 (p=0.078) −0.72 (p=0.023)
GPQA Diamond +0.03 (p=0.96) −0.41 (p=0.60) −0.38 (p=0.58)
SciCode (Subtask) −0.81 (p=0.40) +1.15 (p=0.22) +0.33 (p=0.72)
AA-LCR +0.75 (p=0.73) −0.62 (p=0.68) +0.12 (p=0.95)
IFBench −2.62 (p=0.017) +2.33 (p=0.036) −0.29 (p=0.79)
tau2-bench Telecom −0.00 (p=1.00) −0.88 (p=0.36) −0.88 (p=0.44)

See Potential further improvements for what would likely close the remaining MMMU-Pro gap.

vLLM Throughput (4× GB200, vLLM 0.28.0, TP=4 + EP)

shape (ISL/OSL) concurrency BF16 W4A16 NVFP4 W4A4 NVFP4 W4A4 / BF16
decode 128/2048 1 351 225 307 0.88×
decode 128/2048 32 6,544 5,602 7,320 1.12×
decode 128/2048 128 17,636 15,222 20,076 1.14×
chat 8000/1000 32 3,649 2,373 4,079 1.12×
chat 8000/1000 128 5,571 6,138 6,370 1.14×
prefill 32000/400 128 854 1,092 1,107 1.30×

Output tokens/s; 6 of the 12 measured shapes shown. W4A4 beats BF16 in 9 of 12 and beats W4A16 in all 12. Checkpoint size drops from 67 GiB → 22 GiB (3.1×).

The W4A16 column is why this tutorial targets W4A4 at all: weight-only NVFP4 is slower than BF16 in 10 of 12 shapes. A BF16 activation forces vLLM onto the Marlin dequant-to-BF16 fallback, which never reaches the Blackwell FP4 tensor cores. Only quantizing activations too unlocks them.


Steps to Reproduce

Environment: Results were produced with container nvcr.io/nvidia/nemo:26.08, ModelOpt 0.47.0 and nemo-evaluator-launcher 0.2.6 (with nemo-evaluator 0.2.8) on GB200 (aarch64). See the Megatron-Bridge README for environment setup (including ModelOpt mount path) and container usage. Deployment and evaluation of NVFP4 checkpoints require a Blackwell GPU.

1. Data Preparation

QAD currently supports language-model distillation with text-only data, so the blend is SFT text: nvidia/Nemotron-Cascade-2-SFT-Data, all 8 splits, 17.3B tokens.

Tokenize with the token-budgeted blend workflow using data_blend.yaml — set output_dir in that file first:

Split Tokens Weight
chat 9.73B 56.2
math 3.65B 21.1
science 1.89B 10.9
instruction_following 0.57B 3.3
conversational_agent 0.57B 3.3
terminal_agent 0.57B 3.3
swe 0.31B 1.8
safety 0.003B 0.02

Weights are each split's share of the dataset's own token count — a single pass over the natural mixture, not a tuned blend. The 500-iteration schedule consumes 8.4B tokens, about half an epoch, so nothing is resampled. Note the agent/SWE splits total under 9%, which is relevant to the agentic coverage discussed under potential further improvements.


2. Quantization

W4A4 NVFP4 PTQ using the recipe at modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml. See examples/megatron_bridge/README.md for full PTQ documentation.

PTQ takes ~9 min on 2 GB200 nodes (1.2 GPU-hours) at EP=8.

What gets quantized (read from the exported checkpoint's quantization_config):

Component Precision Tensors
MoE routed experts (down/gate/up_proj) NVFP4 W4A4 (block 16) 30,720
Shared expert NVFP4 W4A4 120
lm_head NVFP4 W4A4 1
Linear attention (in_proj_qkv/in_proj_z/out_proj, 30 layers) FP8 W8A8 90
Full attention (q/k/v/o_proj, 10 layers) FP8 W8A8 40
KV cache FP8 (static) —
MTP head, MoE router, conv1d, in_proj_a/in_proj_b, embeddings, vision tower BF16 —

99.2% of quantized tensors are MoE experts — where the FP4 throughput win comes from. Attention stays at FP8: only 130 tensors, and numerically more sensitive.

W4A4 NVFP4 PTQ command (click to expand)

--ep_size 8 needs 8 ranks, which is 2 GB200 nodes (4 GPUs each) — launched with srun, one task per GPU. On 8-GPU nodes a single-node torchrun --nproc_per_node 8 is equivalent.

# SBATCH --nodes=2 --ntasks-per-node=4 --gpus-per-node=4
srun ... python /opt/Model-Optimizer/examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix \
    --calib_num_samples 1024 \
    --calib_batch_size 1 \
    --seq_length 8192 \
    --export_megatron_path /path/to/qwen36_w4a4_megatron \
    --skip_generate

Set RANK/WORLD_SIZE/LOCAL_RANK from SLURM_PROCID/SLURM_NTASKS/SLURM_LOCALID inside the srun body, as in the QAD command below.

Important

Set --ep_size to the expert-parallel size you will run QAD at. QAD loads this checkpoint directly and EP must match — the expert layout is baked into the distributed checkpoint. Re-exporting at a different EP later means re-running PTQ.


3. Quantization-Aware Distillation (QAD)

QAD fine-tunes the quantized student against the BF16 teacher, so the student learns weights that survive 4-bit rounding. See the QAD section of the Megatron-Bridge README.

Minimum hardware: the student and teacher are both resident, plus an fp32 gradient buffer and optimizer state — roughly 124 GB/GPU of the 185 GiB on a GB200 at the settings below. We used 32 nodes × 4 GB200 (128 GPUs); 500 iterations took 5.7 hours wall-clock (~735 GB200 GPU-hours). Steady-state training is ~37 s/iter (~5.1 h of the total); the remaining ~35 min is per-job startup, loading the 67 GiB teacher and the student — the run was split across two 4-hour jobs, so that cost is paid twice.

QAD command (click to expand)

Launched with srun, one task per GPU, across 32 nodes (128 ranks). Set RANK / WORLD_SIZE / LOCAL_RANK from SLURM_PROCID / SLURM_NTASKS / SLURM_LOCALID in the srun body; python -u keeps the multi-node logs unbuffered.

# SBATCH --nodes=32 --ntasks-per-node=4 --gpus-per-node=4
srun ... python -u /opt/Model-Optimizer/examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --data_paths "${DATA_BLEND}" \
    --data_path_to_cache /path/to/cache \
    --seq_length 32768 \
    --mbs 1 \
    --gbs 512 \
    --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 \
    --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save \
    --eval_iters 0 \
    --save_interval 50 \
    --output_dir /path/to/qad_output

Non-default arguments:

  • --student_megatron_path — the quantized checkpoint from Section 2; --student_hf_path still points at the BF16 model, which supplies the architecture.
  • --tp_size 1 --pp_size 1 — required, not chosen (see below). --ep_size 8 must match the PTQ checkpoint. You may increase --cp_size to enable context parallelism for longer sequence lengths (nemo:26.10 container onwards).
  • --seq_length 32768 --gbs 512 — 16.8M tokens/iteration, 1.7B per 100 iterations.
  • --lr 1e-5 --min_lr 1e-6 — an order of magnitude below typical distillation LRs: the job is to adapt weights to quantization, not to learn the task.
  • --logit_kl_topk 4096 — restricts the KD loss to the teacher's top-4096 vocab entries. With a 248,320-token vocabulary the dense [seq, vocab] fp32 logits are 30.31 GiB per tensor at 32K, which OOMs on its own.
  • --recompute_* / --no_async_save / --eval_iters 0 — all needed to fit. Async save spawns a worker needing its own CUDA context; the validation path computes full-vocab LM and MTP cross-entropy (top-k applies to training only), so eval OOMs at 32K even though training fits.

4. Export

Convert the quantized Megatron checkpoint to a deployable unified HuggingFace checkpoint (~7 min on 1 node, 0.5 GB200 GPU-hours). The unified exporter loads at TP=1, so use pipeline parallelism to shard across GPUs.

Export command (click to expand)
torchrun --nproc_per_node 4 /opt/Model-Optimizer/examples/megatron_bridge/export_quantized_megatron_to_hf.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --megatron_path /path/to/qad_output/checkpoints/iter_0000500 \
    --pp_size 4 \
    --export_unified_hf_path /path/to/qwen36_w4a4_qad_hf

The exported checkpoint is directly deployable with vLLM, TensorRT-LLM and SGLang. Sanity-check it with generate_vllm.py --model <path>.


5. Evaluation

One config per benchmark in eval_configs/, so each can be launched independently with its own walltime and sampling:

Benchmark Config Temp Runs Walltime Metric name
MMMU-Pro mmmu_pro.yaml 1.0 8 2:30 mmmu-pro_pass_at_1_symbolic_correct
GPQA Diamond gpqa.yaml 1.0 1 (avg-of-16) 4:00 gpqa_pass_at_1_avg-of-16_symbolic_correct
SciCode (Subtask) scicode.yaml 0.6 8 2:00 scicode_pass_at_1_subtask_accuracy
AA-LCR aa_lcr.yaml 1.0 8 2:00 aalcr_pass_at_1_judge_correct
IFBench ifbench.yaml 1.0 8 1:30 ifbench_pass_at_1_average_score
tau2-bench Telecom tau2_telecom.yaml 1.0 3 1:30 tau2_bench_telecom_pass_at_1_pass_at_1

Sampling follows the nvidia/Qwen3.6-35B-A3B-NVFP4 card (T=1.0, top_p=0.95, max_new_tokens=131072), except SciCode, which the card specifies at T=0.6. parallelism varies (32 default, 16 for AA-LCR's long contexts and for tau2-bench's multi-turn tool round-trips, 8 for SciCode's long generations) as a throughput knob only. The same configs serve BF16 and NVFP4 unchanged — vLLM reads the FP8 KV-cache setting from the checkpoint's own quantization_config.

Important

Two benchmarks need a second model endpoint besides the one under test. AA-LCR scores answers with a judge (NS_JUDGE_URL / LCR_JUDGE_MODEL_ID); tau2-bench drives a simulated user (TAU2_ENDPOINT_URL / TAU2_USER_MODEL_ID). Both read INFERENCE_API_KEY, and neither can produce its metric without that endpoint. Point both at https:// URLs for the final endpoint: the key rides every request, and a redirect can downgrade it to http or send it to another origin. tau2-bench additionally needs its own deployment — --enable-auto-tool-choice --tool-call-parser qwen3_coder and --max-num-seqs 16 — which is why it cannot share a config with the others (deployment.command is global). qwen3_coder is what the model card specifies; the template's <tool_call> markers resemble hermes, which would mis-parse tool calls and silently invalidate the benchmark. Its top_k / presence_penalty go through tau2's agent_args passthrough, which NEL's top-level params block does not accept — that is why the other five configs omit them.

Important

Every task sets num_repeats: 1; the repeat counts above come from launching a config that many times. N launches give N independent pass@1 values, which is what mean ± sem and the paired tests need — num_repeats: N instead yields a single pass@1[avg-of-N] with no spread. GPQA is the deliberate exception.

Important

Run enough repeats. These benchmarks are noisy, and 3 was not enough to tell signal from noise in either direction. At 3 repeats, AA-LCR (only 100 questions) showed a 3 pp swing that looked real and vanished completely at 8; IFBench's genuine +2.3 pp gain was invisible at 3 and only appeared at 8. Three runs also understates how noisy a benchmark is, so the measured spread looks tighter than it is.

Evaluation launch steps (click to expand)

Set execution.hostname, execution.account and deployment.checkpoint_path in the config, or override with -o <option>=<value>.

pip install "nemo-evaluator-launcher[all]==0.2.6"

# Required by every config:
export HF_TOKEN=<your_huggingface_token>
export SLURM_JOB_DIR=<path_to_slurm_job_output_dir>

# Required only by AA-LCR (judge) and tau2-bench (user simulator):
export INFERENCE_API_KEY=<key_for_those_endpoints>
export NS_JUDGE_URL=<judge_endpoint_url>
export LCR_JUDGE_MODEL_ID=<judge_model_id>
export TAU2_ENDPOINT_URL=<user_simulator_url>
export TAU2_USER_MODEL_ID=<user_simulator_model>

# One benchmark. To verify the pipeline first, add
# `-o ++evaluation.nemo_evaluator_config.config.params.limit_samples=8`
nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml

# The 8 independent runs behind the results table: launch the same config 8 times and pool the
# per-run pass@1 values. Each launch returns its own invocation id; record them so you can select
# exactly those runs when pooling (each eval directory also holds a small canary run).
for i in $(seq 1 8); do
    nemo-evaluator-launcher run --config eval_configs/mmmu_pro.yaml
done

# All six benchmarks at their reported repeat counts
for cfg_runs in mmmu_pro:8 gpqa:1 scicode:8 aa_lcr:8 ifbench:8 tau2_telecom:3; do
    cfg=${cfg_runs%%:*}; runs=${cfg_runs##*:}
    for i in $(seq 1 "$runs"); do
        nemo-evaluator-launcher run --config "eval_configs/${cfg}.yaml"
    done
done

For more details on NeMo Evaluator, see the GitHub repo and documentation.

Output length

Accuracy is only half the serving cost — a model that scores the same while emitting more tokens is slower end to end. Mean completion tokens per benchmark (response_stats.avg_completion_tokens, same runs as the results table):

Benchmark BF16 W4A16 NVFP4 W4A4 NVFP4 PTQ + QAD 500 iters runs/side
SciCode (Subtask) 5,348 +22.9% +25.5% +90.9% 8
IFBench 12,177 +7.0% +9.3% +17.4% 8
AA-LCR 2,560 +7.1% +6.9% +9.8% 8
MMMU-Pro 9,382 +3.0% +6.2% −3.8% 8
GPQA Diamond 13,561 +17.1% +6.6% −8.5% 1

4-bit weights lengthen SciCode outputs by ~23-25% on their own — the published W4A16 checkpoint does it too, so it is not something QAD or W4A4 introduced. QAD then pushes SciCode to +90.9%, for an unchanged score (40.2 vs 39.9), while pulling GPQA and MMMU-Pro back toward BF16. The throughput table is measured at fixed output length, so it does not capture this.

What the SciCode number actually is — worth reading before running QAD on another model

It is not verbosity. It is a failure to terminate on a small fraction of sub-steps:

  • Sub-steps that hit the 131,072-token cap inside <think> go from 0.7% (20/2704, BF16) to 3.6% (96/2704, QAD 500). Almost all return zero answer tokens (93 of those 96) — the model writes a complete solution, says "I think I've been going in circles", and writes it again. In the case we inspected, a 20-word window repeats 352 times and 97.7% of the trace's 20-word windows are duplicates.
  • Those 3.6% of sub-steps burn 45.7% of all completion tokens, so they dominate the mean: excluding them it is +30.4% rather than +90.9%.
  • The median also roughly doubles (+91.6%), so the whole distribution shifted right — this is not only a tail effect.
  • Capped rate peaks at iteration 50 (4.3%) and settles at 3.1% / 3.6% by 300 / 500; it is not gradual drift.

The obvious suspect — that --logit_kl_topk 4096 leaves the stop tokens outside the loss — did not hold up. Probing the BF16 teacher over one runaway trace: </think> does fall outside top-4096 at 35% of positions overall, but in the looping region the teacher gives <|im_end|> a median rank of 5 and </think> ~570, both well inside top-k. The teacher is signalling "stop here" at positions the loss did cover, and the student still does not stop. More likely: the blend has few "the answer is written, now stop" positions in this style, and a teacher-forced loss never exercises free-running generation 10K+ tokens deep.

For the next QAD run, three things follow: track the length-capped rate as a first-class metric alongside accuracy (a benchmark score can stay flat while 3.6% of responses return nothing); consider top-p instead of top-k for the KD loss so coverage adapts to the teacher's entropy rather than a fixed rank; and if memory allows, full-vocab KL — at 32K on this 248,320-token vocabulary the dense fp32 logits are 30.31 GiB per tensor, which is why top-k was used here, but more GPU memory or a smaller model or shorter sequence may afford it.


6. vLLM Inference Benchmarking

Throughput was measured with AIPerf against a served vLLM endpoint on 4× GB200 — three ISL/OSL shapes × four concurrencies = the 12 shapes in the results table.

Serve + benchmark commands (click to expand)

The same vllm serve command works for BF16 and NVFP4; the quantized checkpoint carries its own format and FP8 KV-cache settings in quantization_config.

vllm serve <checkpoint_path> --served-model-name bench \
    --host 127.0.0.1 --port 8000 \
    --tensor-parallel-size 4 --data-parallel-size 1 --enable-expert-parallel \
    --max-model-len 262144 --reasoning-parser qwen3 \
    --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 128}' \
    --max-num-batched-tokens 8192 --enable-chunked-prefill

# decode-bound, chat, and prefill-bound shapes at four concurrencies each
for shape in "128 2048 decode" "8000 1000 chat" "32000 400 prefill"; do
    set -- $shape; ISL=$1; OSL=$2; NAME=$3
    for C in 1 8 32 128; do
        aiperf profile -m bench --endpoint-type chat --streaming -u localhost:8000 \
            --synthetic-input-tokens-mean $ISL --output-tokens-mean $OSL \
            --concurrency $C --request-count $(( C * 5 )) \
            --tokenizer <checkpoint_path> \
            --extra-inputs ignore_eos:true --random-seed 42 \
            --artifact-dir bench/${NAME}_isl${ISL}_osl${OSL}/c${C}
    done
done

Tip

Check the endpoint answers a known question correctly before benchmarking it. A checkpoint served with the wrong KV dtype or an unexpected kernel still produces tokens at a plausible rate — the throughput number looks fine and is meaningless. We gate every benchmark on a short prompt with a verifiable answer.

Tip

To deploy the model with vLLM, refer to the vLLM Quickstart documentation.


Potential further improvements

This is a single 500-iteration run at 32K on a text-only blend. Each of those three choices is an unexplored lever.

1. Continue QAD longer. The recovery had not saturated: IFBench's gain arrived late (−0.11 pp vs the student at iteration 300, +2.33 pp at 500) and MMMU-Pro was still climbing (+0.51 pp, p=0.078). Training also used only 8.4B of the blend's 17.3B tokens, so it can roughly double before repeating a sample. At ~735 GPU-hours per 500 iterations this is the most direct experiment, and IFBench + MMMU-Pro alone (~55 GPU-hours) are enough to read the result.

2. Longer sequence length (64K). The model serves at 262K context, so 32K trains on a fraction of it; AA-LCR, the one long-context benchmark, showed no movement either way. At 32K the run already needs top-k KD plus full recompute, and memory is dominated by the resident student + teacher + fp32 gradient buffer rather than by sequence length — so 64K needs a different memory lever.

3. A better data blend. The clearest signal in the study: MMMU-Pro is multimodal, the blend is text-only, and MMMU-Pro is the one deficit QAD did not close. Since QAD supports language-model distillation with text-only data, recovering a multimodal benchmark that way asks the method to do something it is not set up for. Agentic coverage is also thin (agent/SWE splits under 9%, and tau2-bench showed the transient 8× agent_error spike), and weights here are simply proportional to token count — contrast the Nemotron-3-Nano blend, which is deliberately designed against target benchmarks and ablated.

If you run only one, run (1): no engineering risk, the trajectory says the gain is still arriving, and it produces the evidence to judge whether (2) and (3) are worth their cost.