Files
Model-Optimizer/tools
Jenny Chen f2bfe63183 Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data (#2134)
### What does this PR do?

Type of change: New example

Add a Nemotron Nano 3 QAD Launcher Example on OSS
Nemotron-Post-Training-V2 data. It performs 4 steps
1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a
Megatron-Core BF16 checkpoint
2. PTQ: quantize the Megatron-Core checkpoint to
`MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config
3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to
the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat`
data. To train on a different subset or load the entire dataset, you may
modify `--finetune-data-split` and `--finetune-data-files` flags.
4. Export: export the QAD checkpoint to HuggingFace format so it is
ready for local inference


All steps use the TE (Transformer Engine) spec, which with the new
TEGroupedMLP per-expert quantizer is approximately 10-15% faster than
the previous local ModelOpt spec (which used SequentialMLP) on
Hybrid-MoE models.

### Usage

```
# Usage from tools/launcher:
source .env-slurm
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, backward breaking changes,
deprecations, or fixes for critical bugs present in previous releases.
-->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Added a launcher configuration for NVFP4 quantization-aware
distillation of the Nemotron 3 Nano 30B-A3B model.
- Added support for selecting training or fine-tuning workflows through
`MLM_TRAIN_SCRIPT`.
  - Improved forwarding of additional training arguments.

- **Updates**
  - Updated the Megatron-LM launcher component to a newer revision.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-08-10 13:52:52 -07:00
..