mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What Adds VLM (e.g. Qwen3.5-VL, Gemma3-VL) knowledge-distillation / QAD support to the Megatron-Bridge examples: - `distill.py` distills only the **language model** submodule (vision tower + projector untouched), reusing the LLM training path. - New `export_distilled_megatron_to_hf.py` converts a distilled Megatron checkpoint (**any** iteration) to HF. Required especially for VLM distilled ckpt as it only has LM weights so we need to initialize full VLM, swap LLM weights then save to HF - Renames `export.py` → `export_quantized_megatron_to_hf.py`. ## Related upstream Megatron-Bridge PRs to be available in nemo:26.08 container: - NVIDIA-NeMo/Megatron-Bridge#4707 — `DistillationProvider` submodule distillation (non-blocking; added temporary WAR) - NVIDIA-NeMo/Megatron-Bridge#4706 — MoE expert weight-mapping fix (Qwen3.5-VL-MoE with moe_grouped_gemm=False). Also removed ModelOpt side WAR previously added as it was not accurate; better to wait till next container release or mount latest MBridge into the 26.06 container. ## Testing - Qwen3.6-35B-A3B Pruning + Distillation with MMLU evaluation sanity check (results below in comments) - Cosmos 2 Reason 2B valiadted by SAs (results below in comments) - Validated end-to-end on `nemo:26.06` (distill → separate HF export; LLM + VLM, incl. TP→TP/PP reshard). `test_distill_vlm` runs the export script as a CI e2e step. - Many CICD tests for wide coverage of all mbridge scripts <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **New Features** * Added a dedicated HuggingFace exporter for distilled Megatron checkpoints, with distinct LLM vs VLM conversion flows. * **Bug Fixes** * Improved Megatron-Bridge distillation/export consistency, including safer handling of VLMs and targeted submodule distillation. * **Documentation** * Updated Megatron-Bridge READMEs and tutorials to reference the new quantized and distilled export scripts and revised CLI guidance. * **Tests** * Expanded distillation, QAD, and quantization/export tests to cover LLM/VLM variants, with conditional skipping for unsupported MoE setups. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Megatron-Bridge Tutorials
End-to-end tutorials that combine ModelOpt optimization techniques on NVIDIA Megatron-Bridge models.
Each one walks through a complete workflow using the scripts in examples/megatron_bridge (prune_minitron.py, distill.py, quantize.py, export_quantized_megatron_to_hf.py, export_distilled_megatron_to_hf.py).
Available tutorials
| Tutorial | What it covers |
|---|---|
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | End-to-end optimization of the Nemotron-3-Nano-30B-A3B-BF16 (MoE + Mamba-Transformer hybrid) model: Minitron structured pruning (31.6B/A3.6B → 22B/A3.0B) → two-phase knowledge distillation (100B tokens, 8K then 32K seq length) → quantization → vLLM deployment. Includes data-blend preparation, evaluation setup, and detailed pruning / data-blend / long-context ablations. |