### What does this PR do? Type of change: documentation + minor example-script tweaks Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16 tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md) results using the **new shared calibration loop (sequence packing)** and to **fix tool calling in evaluation**. - Refreshed the prune → distill → eval → **FP8** results (accuracy + vLLM throughput tables) with the new calibration loop. - **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now run the Python sandbox tool. The tutorial reports both **with-tools** and **no-tools** GPQA/AIME and shows `mean ± std_dev`. - Script tweaks: `quantize.py` calibration now uses sequence packing (`pack=True`) which leads to slight improvement in PTQ; `prune_minitron.py` defaults `inference_batch_size` to `calib_batch_size`. ### Testing Documentation + small example-script changes; tutorial relative links resolve and the results tables / figure were verified consistent. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: Yes - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ (will run `/claude review`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the main guide and evaluator instructions for prune + distill + FP8/NVFP4 quantization, including refreshed vLLM deployment tips, benchmark/noise presentation, and long-context tool-calling attribution notes. * Refreshed README technique examples/links, reordered the model support matrix rows, and improved pruning overview/support-matrix text. * **Changes to Examples** * NAS pruning now documents higher GPU memory usage vs manual pruning; pruning batching defaults were improved. * Quantization PTQ calibration uses packed document packing; quantized checkpoint export messaging was streamlined. * Updated pruning/distillation/quantization tutorial guidance, metrics/tables, command parameters, and evaluator YAML settings (KV-cache dtype, generation defaults, task behavior). <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
17 KiB
Ablations: Nemotron-3-Nano-30B-A3B-BF16
Note
Every benchmark table in this document is from evaluations run without eval-time tool calling (no Python sandbox). The main README reports both with-tools and no-tools numbers for GPQA and AIME. Consequently, the gains attributed to the
tool_callingtraining data in the data-blend ablation reflect training-time exposure to agentic/function-call reasoning, not tools invoked during evaluation.
Pruning
Note
The search space analysis below is specific to Nemotron hybrid (Mamba + Attention + MoE) models. Standard transformers expose only layers/hidden/attention/FFN dimensions, but these models add Mamba-specific dimensions (
mamba_num_heads,mamba_head_dim) and MoE dimensions (num_moe_experts,moe_ffn_hidden_size,moe_shared_expert_intermediate_size). The resulting search space is significantly larger, and the default 40% width / 20% depth constraints include many dead-zone architectures that waste scoring compute. The tighter constraints recommended here were derived from this analysis.
- Score function:
mmlu_10pct_bs32(zero-shot MMLU on 10% subset, no distillation applied) - Candidates analyzed: ~50 across multiple NAS runs, all with
--prune_target_active_params 3e9 - Random baseline (5-way MMLU): ~0.25
- Search space note:
num_attention_headswas skipped in all runs.
Key Findings Summary
| Dimension | Good range | Avoid | Key finding |
|---|---|---|---|
num_layers |
≥ 48 | ≤ 46 | 42L fails universally; 46L avg MMLU 0.261 |
mamba_state_dim |
≥ 3072 (56×56 or 64×64) | < 3072; asymmetric pairs (56×64) | Symmetric reduction — both heads and head_dim must shrink together |
hidden_size |
2304–2560 | 2688 (original); 2048 | Original hidden_size is actively harmful when other dims are pruned |
| MoE dims | experts 96–128; shared 3072–3712 | experts = 88 | Weak independent signal after controlling for dominant dims |
Original Model Dimensions
Params: 31.6B, Active: 3.6B
| Dimension | Value |
|---|---|
num_hidden_layers |
52 |
hidden_size |
2688 |
mamba_num_heads |
64 |
mamba_head_dim |
64 |
num_moe_experts |
128 |
moe_ffn_hidden_size |
1856 |
moe_shared_expert_intermediate_size |
3712 |
Top Candidates (best seen across all runs)
All candidates below have active_params = 3.00B.
| Score | Layers | Hidden | Heads | HeadDim | Experts | FFN | Shared | Total Parameters |
|---|---|---|---|---|---|---|---|---|
| 0.4783 | 52 | 2304 | 64 | 64 | 104 | 1856 | 3072 | 22.3B |
| 0.4727 | 52 | 2560 | 64 | 64 | 80 | 1280 | 3712 | 14.2B |
| 0.4650 | 48 | 2560 | 56 | 56 | 112 | 1792 | 3072 | 25.4B |
| 0.4608 | 52 | 2304 | 64 | 64 | 96 | 1856 | 3072 | 20.7B |
| 0.4329 | 48 | 2560 | 56 | 56 | 104 | 1792 | 3072 | 23.7B |
| 0.4119 | 50 | 2304 | 64 | 64 | 80 | 1856 | 3712 | 17.6B |
| 0.3762 | 52 | 2560 | 48 | 64 | 96 | 1536 | 3712 | 19.3B |
Two architecture families dominate the top results — see Design Recipe below.
Dimension Sensitivity
Sensitivity is ranked by strength of signal across all candidates.
1. mamba_state_dim = mamba_num_heads × mamba_head_dim — strongest width signal
mamba_state_dim sensitivity: threshold ≥3072, best families 56×56 and 64×64 (click to expand)
Analyzing num_heads and head_dim jointly is more predictive than either alone:
| state_dim | Formula | Avg MMLU | Max MMLU | n |
|---|---|---|---|---|
| 1920 | 48×40 | 0.254 | 0.257 | 2 |
| 2304 | 48×48 | 0.249 | 0.260 | 8 |
| 2688 | 48×56 or 56×48 | 0.264 | 0.310 | 7 |
| 3072 | 48×64 | 0.342 | 0.376 | 3 |
| 3136 | 56×56 | 0.400 | 0.465 | 4 |
| 3584 | 56×64 or 64×56 | 0.261 | 0.340 | 9 |
| 4096 | 64×64 | 0.399 | 0.478 | 7 |
Threshold: state_dim ≥ 3072 to escape the near-random zone. Below 3072, no candidate has ever exceeded 0.31 regardless of other settings. The two reliable good families are symmetric configurations: 56×56 = 3136 and 64×64 = 4096 (original).
Why 3584 has a poor average despite being large: All 3584 candidates in the data are at 46L (depth-limited). The low average is a depth confound, not an inherent failure of 3584.
Asymmetric reductions hurt: Reducing only one of {num_heads, head_dim} while keeping the other at 64 (giving 3584) performs worse than symmetric reduction of both to 56 (giving 3136). The 56×56 pattern is consistently more reliable.
head_dim=48 is uniformly bad: Across all candidates with head_dim=48, every single one scored 0.240–0.260. This holds across varying layers (48L–52L) and hidden sizes. head_dim=48 is the effective lower bound under 30% width pruning and it is never viable.
2. num_layers — hard lower bound
num_layers sensitivity: hard floor at 48L, 42L universally fails (click to expand)
| Layers | Avg MMLU | Max MMLU | n |
|---|---|---|---|
| 42 | 0.232 | 0.234 | 7 |
| 46 | 0.261 | 0.340 | 12 |
| 48 | 0.409 | 0.465 | 4 |
| 50 | 0.257 | 0.412 | 9 |
| 52 | 0.336 | 0.478 | 19 |
42L is a universal failure — 7/7 candidates at 42L scored near-random, with no other dimension able to compensate. Eliminated by the 15% depth constraint.
46L is still suboptimal — avg 0.261, no candidate above 0.340. The effective floor for good performance is 48L. The 15% depth constraint (min 45L) is correct but 46L candidates still appear in results and are reliably mediocre.
50L avg is pulled down by head_dim=48 candidates — when controlling for head_dim≥56, 50L performs comparably to 52L. The 50L failure is a head_dim confound, not a genuine depth issue.
3. hidden_size — bad at both extremes, joint constraint with depth
hidden_size sensitivity: 2304–2560 good, original 2688 consistently bad (click to expand)
| hidden_size | Avg MMLU | Max MMLU | n |
|---|---|---|---|
| 2304 | 0.445 | 0.478 | 4 |
| 2560 | 0.307 | 0.473 | 26 |
| 2688 | 0.258 | 0.268 | 10 |
hidden_size=2688 (the original) is definitively bad in pruned configurations. This is confirmed by candidates at 52L with hidden=2688 — sufficient depth — all scoring 0.255–0.260. It is not a depth confound. Keeping hidden size un-pruned while reducing other dimensions means the active param budget is consumed inefficiently, leaving too little capacity in the MoE and Mamba layers.
hidden_size=2304 requires num_layers ≥ 48. When paired with 42L, it scores 0.232. Paired with 48–52L, it produces the best candidates. This is a joint constraint.
hidden_size=2048 (seen in runs without an active param constraint): consistently undershoots the 3B active target, capping MMLU at ~0.30. Not a viable option.
4. num_moe_experts, moe_ffn_hidden_size, moe_shared_expert_intermediate_size — weak independent signal
MoE dimension sensitivity: weak signal, experts=88 bad, shared 3072–3712 preferred (click to expand)
No strong monotonic pattern after controlling for the dominant dimensions above.
experts=88: consistently bad (avg 0.260)moe_shared_expert_intermediate_sizein 3072–3712: preferred; 2560 and 3328 tend to be worsemoe_ffn_hidden_size=1280: produced the most parameter-efficient good architecture (0.4727, 14.15B total) but is a 31% reduction from the original 1856 — just outside the 30% width constraint. Use 32% width to include this family if total-param efficiency matters.
Pruning Constraint Recommendations
--max_depth_pruning 0.15 # eliminates 42L dead zone; no good arch needs >7 layers removed
--max_width_pruning 0.30 # eliminates head_dim≤40, shared=2560; 0.32 to also include ffn=1280
--prune_target_params 28e9 # not very critical constraint as active params are the primary target
--prune_target_active_params 3e9 # required; omitting this causes active params to undershoot, capping MMLU at ~0.30
--hparams_to_skip num_attention_heads
Remaining dead zones within 15%/30% search space — the following are still reachable by the NAS but consistently fail:
hidden_size=2688— all 4 candidates at 52L scored 0.255–0.260; consider hardcoding min hidden to 2304num_layers=46— avg 0.261 across 12 candidates, none above 0.340mamba_head_dim=48— all 7 candidates scored 0.240–0.260; consider hardcoding min head_dim to 56
Design Recipe
Two confirmed high-quality architecture families based on all candidates:
Family 1 — Best MMLU (hidden=2304, state_dim=4096):
52L | hidden=2304 | mamba_num_heads=64 | mamba_head_dim=64 | num_moe_experts=96–104 | moe_ffn_hidden_size=1856 | shared=3072
active=3.00B, total=20.7–22.3B, MMLU=0.461–0.478
Family 2 — Good MMLU, larger total params (hidden=2560, state_dim=3136):
48L | hidden=2560 | mamba_num_heads=56 | mamba_head_dim=56 | num_moe_experts=104–112 | moe_ffn_hidden_size=1792 | shared=3072
active=3.00B, total=23.7–25.4B, MMLU=0.433–0.465
Required conditions for any good candidate:
| Condition | Threshold | Failure rate below threshold |
|---|---|---|
num_layers |
≥ 48 | 42L: 7/7 fail; 46L: 12/12 below 0.35 |
mamba_state_dim |
≥ 3072 | 0/24 candidates below 3072 exceed 0.31 |
hidden_size |
2304–2560 | 2688: 10/10 below 0.27 |
mamba_head_dim |
≥ 56 | 15/15 candidates with head_dim≤48 score 0.24–0.26 |
Distillation results: 1st best vs 2nd best pruning candidate
To show that the pre-distillation MMLU score is a reasonable proxy for downstream quality, we also distilled the 2nd-best candidate by MMLU score (within the 30% width / 15% depth constraints) and evaluated at 2.5B and 20B tokens.
| Rank | Score | Layers | Hidden | Heads × HeadDim | Experts | FFN | Shared | Total | Family |
|---|---|---|---|---|---|---|---|---|---|
| 1st | 0.4783 | 52 | 2304 | 64 × 64 | 104 | 1856 | 3072 | 22.3B | Family 1 |
| 2nd | 0.4650 | 48 | 2560 | 56 × 56 | 112 | 1792 | 3072 | 25.4B | Family 2 |
Evaluation after distillation (same blend, same training schedule):
| Checkpoint | Candidate | MMLU | MMLU Pro | GPQA Diamond | LiveCodeBench v6 | AIME 2025 | IFBench | SciCode (Subtask) |
|---|---|---|---|---|---|---|---|---|
| 2.5B tokens | 1st best | 68.6 | 73.6 | 62.5 | 57.5 | 79.1 | 58.3 | 22.6 |
| 2.5B tokens | 2nd best | 68.1 | 72.6 | 60.0 | 48.1 | 63.0 | 50.7 | 15.2 |
| 20B tokens | 1st best | 70.8 | 74.6 | 65.3 | 61.0 | 79.8 | 63.6 | 22.5 |
| 20B tokens | 2nd best | 71.0 | 74.6 | 61.6 | 49.6 | 54.0 | 57.3 | 20.5 |
Observations:
- In this run the 1st-best candidate wins on most benchmarks at both checkpoints. The pre-distillation MMLU proxy score gap (0.4783 vs 0.4650) translated to meaningful downstream gaps on reasoning benchmarks here (AIME: 79.1 vs 63.0 at 2.5B; 79.8 vs 54.0 at 20B).
- The NAS proxy score is a useful signal but not a guarantee. We have seen cases in other experiments where ranks flip after distillation — the proxy is computed without any recovery training, and small differences in architecture can interact non-trivially with the distillation blend. When the top few candidates have similar scores (e.g. within ~2-3% MMLU), running short (~2B-token) distillation on each before committing to the full run is the safer practice. Picking only the 1st-best blindly works often; but may not always be the best choice.
Distillation
Effect of data blend (tool_calling)
Two distillation data blends were tried on the same pruned 22B/A3.0B candidate with identical hyperparameters during the 8K-seq-length phase.
- Old blend (no tool_calling): 30% Pretraining (Code 5, General 20, MATH 5) + 70% Post-training v1/v3 (Math 30, Coding 20, Science 15, IF 5). Schedule: 100B @ 8K then +25B @ 32K (125B total).
- New blend (used in main README): adds 5%
Nemotron-Agentic-v1 / tool_calling, reducing Math 30→27 and Science 15→13 to make room. Same pretraining split. Schedule: 80B @ 8K then +20B @ 32K (100B total).
The blend comparison below is capped at 80B tokens — beyond that the two runs diverge in seq_length (new blend switches to 32K at 80B; old blend continues at 8K until 100B), so per-checkpoint deltas at higher token counts would conflate blend differences with long-context-phase differences.
Old blend results (8K seq length only). Average is over the 6 reasoning benchmarks (excluding MMLU as we already have MMLU Pro), matching the main README convention:
| Tokens (iters at 8K) | MMLU | MMLU Pro | GPQA Diamond | LiveCodeBench v6 | AIME 2025 | IFBench | SciCode (Subtask) | Average |
|---|---|---|---|---|---|---|---|---|
| 2.5B (100 iters) | 68.6 | 73.6 | 62.5 | 57.5 | 79.1 | 58.3 | 22.6 | 58.9 |
| 20B (800 iters) | 70.8 | 74.6 | 65.3 | 61.0 | 79.8 | 63.6 | 22.5 | 61.1 |
| 40B (1600 iters) | 71.6 | 75.6 | 64.5 | 61.6 | 76.8 | 67.4 | 28.3 | 62.4 |
| 60B (2400 iters) | 71.5 | 76.0 | 67.5 | 63.0 | 77.5 | 68.4 | 29.4 | 63.6 |
| 80B (3200 iters) | 71.7 | 76.5 | 68.4 | 64.2 | 80.2 | 66.5 | 28.7 | 64.1 |
Per-benchmark difference (Δ = new − old), at matching iter counts during the 8K phase. Average Δ uses the 6-benchmark average (no MMLU):
| Benchmark | 2.5B Δ | 20B Δ | 40B Δ | 60B Δ | 80B Δ |
|---|---|---|---|---|---|
| MMLU | +0.1 | 0.0 | -0.3 | +0.2 | 0.0 |
| MMLU Pro | -0.3 | +0.2 | +0.8 | +0.1 | 0.0 |
| GPQA Diamond | +1.2 | +0.7 | +2.7 | +0.6 | +0.7 |
| LiveCodeBench v6 | -2.2 | +1.3 | +0.7 | +0.6 | -0.3 |
| AIME 2025 | -1.5 | -0.2 | +3.0 | +1.3 | +0.5 |
| IFBench | +0.8 | +1.8 | -1.4 | -1.1 | 0.0 |
| SciCode (Subtask) | +2.5 | +3.6 | -1.7 | -2.4 | +0.3 |
| Average | +0.1 | +1.3 | +0.7 | -0.1 | +0.2 |
Summary — new blend is preferred:
- GPQA Diamond is the most consistent winner — positive Δ at every single checkpoint (+1.2/+0.7/+2.7/+0.6/+0.7). The tool_calling allocation helps throughout training, not just at the end. Since these evals run without eval-time tools, the gain comes from training-time exposure to agentic/function-call reasoning traces transferring to GPQA — not from the model invoking a tool during evaluation.
- SciCode trends positive early (+2.5/+3.6 at 2.5B/20B) but neutral-to-slightly-negative mid-training (-1.7/-2.4 at 40B/60B) and positive again at 80B (+0.3). Even averaged-of-8, SciCode has the largest residual checkpoint-to-checkpoint noise of any benchmark.
- AIME 2025 trends positive from 40B onward (+3.0/+1.3/+0.5). The early dip is consistent with the math share dropping from 30%→27%; the model catches up once enough tokens have been seen.
- IFBench is small-positive early (+0.8/+1.8) then dips mid-training (-1.4/-1.1) before recovering to neutral at 80B (0.0). Net roughly flat over training.
- MMLU / MMLU Pro / LiveCodeBench v6 are essentially flat — the upweighted General pretraining split keeps knowledge metrics steady despite the math/science share decrease.
- Average is non-negative at 4 of 5 checkpoints (+0.1/+1.3/+0.7/-0.1/+0.2). Net slightly positive overall, with most of the win coming from GPQA and AIME at later checkpoints.
Effect of long context training
After the 8K-seq-length phase of the old-blend run, training was continued with the same blend but with seq_length increased from 8192 to 32768. The longer-context phase is short (200–1000 additional iters) but disproportionately impactful. Numbers below use the old-blend run because it has the longer LC sweep (1000 iters / +25B tokens). The new-blend run's shorter LC phase (+800 iters / +20B tokens, ending at 100B) is in the main README and shows the same qualitative finding (large AIME jump immediately at the start of the LC phase) — though at higher absolute AIME scores there, since the README evals enable tools.
Average is over the 6 reasoning benchmarks (excluding MMLU), matching the main README:
| Tokens (additional iters at 32K) | MMLU | MMLU Pro | GPQA Diamond | LiveCodeBench v6 | AIME 2025 | IFBench | SciCode (Subtask) | Average |
|---|---|---|---|---|---|---|---|---|
| 100B (end of 8K phase) | 71.8 | 76.6 | 68.4 | 64.5 | 81.0 | 68.4 | 28.0 | 64.5 |
| 105B (+200 iters at 32K) | 71.9 | 76.8 | 69.7 | 65.6 | 87.3 | 69.3 | 29.5 | 66.4 |
| 125B (+1000 iters at 32K) | 71.9 | 76.7 | 70.5 | 65.7 | 88.0 | 68.5 | 27.6 | 66.2 |
Summary — high-leverage, low-cost addition:
- AIME 2025 sees the largest jump: +6.3 points from 100B → 105B after only 200 additional iters at 32K. AIME chains-of-thought routinely exceed 8K tokens, so the 8K-trained student was being truncated mid-reasoning; allowing 32K immediately unlocks the latent capability.
- GPQA Diamond (+2.1) and LiveCodeBench v6 (+1.2) also benefit meaningfully — both produce long reasoning traces.
- MMLU / MMLU Pro / IFBench are flat — these benchmarks comfortably fit within 8K and gain nothing from longer context.
- Overall Average lifts from 64.5 → 66.4 (+1.9) after the first 200 iters at 32K. Going further to 1000 iters slightly regresses (66.2), so most of the LC benefit is captured early.