mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: New example Add a Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data. It performs 4 steps 1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a Megatron-Core BF16 checkpoint 2. PTQ: quantize the Megatron-Core checkpoint to `MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config 3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat` data. To train on a different subset or load the entire dataset, you may modify `--finetune-data-split` and `--finetune-data-files` flags. 4. Export: export the QAD checkpoint to HuggingFace format so it is ready for local inference All steps use the TE (Transformer Engine) spec, which with the new TEGroupedMLP per-expert quantizer is approximately 10-15% faster than the previous local ModelOpt spec (which used SequentialMLP) on Hybrid-MoE models. ### Usage ``` # Usage from tools/launcher: source .env-slurm uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added a launcher configuration for NVFP4 quantization-aware distillation of the Nemotron 3 Nano 30B-A3B model. - Added support for selecting training or fine-tuning workflows through `MLM_TRAIN_SCRIPT`. - Improved forwarding of additional training arguments. - **Updates** - Updated the Megatron-LM launcher component to a newer revision. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>