Files
Keval MorabiaandClaude Sonnet 4.6 bb08094ff1 Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary

End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how
ModelOpt techniques stack:

- **Pruning** — Minitron structured pruning 9B → 7B
- **Distillation** — Megatron-Bridge knowledge distillation up to 80B
tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math
500, IFEval, SciCode
- **Evaluation** - using nemo-evaluator
- **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on
vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from
\`config.json\`)
- **vLLM Throughput** — BF16 vs FP8 benchmark on single H100

<img width="2085" height="1740" alt="image"
src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87"
/>

<img width="2085" height="810" alt="image"
src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843"
/>


### Files changed

- `examples/pruning/minitron/README.md` — index page for Minitron
end-to-end tutorials
- `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` —
full repro doc with 6 sections: data prep, pruning, distillation,
evaluation, FP8 quantization, vLLM benchmarking
-
`examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml`
— NeMo Evaluator config used for all benchmark numbers
- `examples/pruning/puzzletron/README.md` — index page for Puzzletron
distillation results
- `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron
distillation results (renamed from puzzletron.md)
- `examples/pruning/README.md` — updated Results section with direct
links to new locations
- `examples/megatron_bridge/README.md` — updated results link to point
to `examples/pruning/`
- `examples/puzzletron/README.md` — updated distillation results link
- `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for
all datasets used in the data blend

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Documentation

* **New end-to-end tutorial** for model optimization covering Minitron
pruning, knowledge distillation, FP8 quantization, and vLLM deployment
with reproducibility steps and benchmark results
* **Dataset preparation guide** with ready-to-run tokenization templates
for Nemotron HuggingFace datasets
* **Evaluation configuration** and results documentation including
ablation studies across multiple benchmarks
* **Updated navigation** across pruning, distillation, and dataset
examples to streamline user workflows

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 14:47:16 +05:30

2.6 KiB

Puzzletron Distillation Results

The following MMLU results demonstrate knowledge distillation on student models that were first compressed using Puzzletron. The original (uncompressed) model serves as the teacher, and distillation recovers accuracy lost during compression.

Qwen3-8B compressed to 80% of original

The student was created by compressing Qwen3-8B to 80% of its original size using Puzzletron.

Model MMLU Humanities Other Social Sci STEM
Student (before distillation) 0.5910 0.5046 0.6363 0.6831 0.5855
Student (after distillation) 0.6921 0.5906 0.7316 0.7975 0.7016
Teacher (original Qwen3-8B) 0.7493 0.6648 0.7856 0.8385 0.7526

MMLU accuracy improved from 59.10% to 69.21% (+10.11 pp) after distillation with just 100 iterations on WikiText-103, recovering 64% of the gap to the teacher model.

Llama-3.1-8B-Instruct compressed to 50% of original

The student was created by compressing Llama-3.1-8B-Instruct to 50% of its original size using Puzzletron.

Model MMLU Humanities Other Social Sciences STEM
Student (before distillation) 0.2316 0.2462 0.2292 0.2250 0.2274
Student (after distillation) 0.2960 0.3146 0.3085 0.2925 0.2768
Teacher (original Llama-3.1-8B-Instruct) 0.6839 0.7231 0.7038 0.7667 0.5911

Llama-3.1-8B-Instruct compressed to 69% of original (regression)

The student was created by compressing Llama-3.1-8B-Instruct to ~69% of its original size using Puzzletron. This example shows regression due to overfitting on the small WikiText-103 dataset (100 iterations). MMLU was evaluated on a subset of 100 samples per task:

Model MMLU Humanities Other Social Sciences STEM
Student (before distillation) 0.6626 0.7069 0.6892 0.7525 0.5574
Student (after distillation) 0.6496 0.6862 0.6677 0.7433 0.5532
Teacher (original Llama-3.1-8B-Instruct) 0.6839 0.7231 0.7038 0.7667 0.5911

MMLU decreased from 66.26% to 64.96% (-1.30 pp) -- the model overfitted to WikiText-103. This highlights the importance of using larger, more diverse datasets for distillation.

Recommendations

  • Use larger datasets for production distillation (e.g., Nemotron-Pretraining-SFT-v1) to avoid overfitting as shown in the regression case above.
  • Train for more iterations to ensure proper convergence.