Files
Model-Optimizer/examples/megatron_bridge/tutorials
Keval MorabiaandClaude Opus 4.8 21d0069e4d MBridge VLM distillation / QAD support (#1938)
## What

Adds VLM (e.g. Qwen3.5-VL, Gemma3-VL) knowledge-distillation / QAD
support to the Megatron-Bridge examples:

- `distill.py` distills only the **language model** submodule (vision
tower + projector untouched), reusing the LLM training path.
- New `export_distilled_megatron_to_hf.py` converts a distilled Megatron
checkpoint (**any** iteration) to HF. Required especially for VLM
distilled ckpt as it only has LM weights so we need to initialize full
VLM, swap LLM weights then save to HF
- Renames `export.py` → `export_quantized_megatron_to_hf.py`.

## Related upstream Megatron-Bridge PRs to be available in nemo:26.08
container:

- NVIDIA-NeMo/Megatron-Bridge#4707 — `DistillationProvider` submodule
distillation (non-blocking; added temporary WAR)
- NVIDIA-NeMo/Megatron-Bridge#4706 — MoE expert weight-mapping fix
(Qwen3.5-VL-MoE with moe_grouped_gemm=False). Also removed ModelOpt side
WAR previously added as it was not accurate; better to wait till next
container release or mount latest MBridge into the 26.06 container.

## Testing

- Qwen3.6-35B-A3B Pruning + Distillation with MMLU evaluation sanity
check (results below in comments)
- Cosmos 2 Reason 2B valiadted by SAs (results below in comments)
- Validated end-to-end on `nemo:26.06` (distill → separate HF export;
LLM + VLM, incl. TP→TP/PP reshard). `test_distill_vlm` runs the export
script as a CI e2e step.
- Many CICD tests for wide coverage of all mbridge scripts

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **New Features**
* Added a dedicated HuggingFace exporter for distilled Megatron
checkpoints, with distinct LLM vs VLM conversion flows.

* **Bug Fixes**
* Improved Megatron-Bridge distillation/export consistency, including
safer handling of VLMs and targeted submodule distillation.

* **Documentation**
* Updated Megatron-Bridge READMEs and tutorials to reference the new
quantized and distilled export scripts and revised CLI guidance.

* **Tests**
* Expanded distillation, QAD, and quantization/export tests to cover
LLM/VLM variants, with conditional skipping for unsupported MoE setups.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 15:16:30 +00:00
..

Megatron-Bridge Tutorials

End-to-end tutorials that combine ModelOpt optimization techniques on NVIDIA Megatron-Bridge models. Each one walks through a complete workflow using the scripts in examples/megatron_bridge (prune_minitron.py, distill.py, quantize.py, export_quantized_megatron_to_hf.py, export_distilled_megatron_to_hf.py).

Available tutorials

Tutorial What it covers
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 End-to-end optimization of the Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) model: Minitron structured pruning (31.6B/A3.6B → 22B/A3.0B) → two-phase knowledge distillation (100B tokens, 8K then 32K seq length) → quantization → vLLM deployment. Includes data-blend preparation, evaluation setup, and detailed pruning / data-blend / long-context ablations.