Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data (#2134)

### What does this PR do?

Type of change: New example

Add a Nemotron Nano 3 QAD Launcher Example on OSS
Nemotron-Post-Training-V2 data. It performs 4 steps
1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a
Megatron-Core BF16 checkpoint
2. PTQ: quantize the Megatron-Core checkpoint to
`MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config
3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to
the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat`
data. To train on a different subset or load the entire dataset, you may
modify `--finetune-data-split` and `--finetune-data-files` flags.
4. Export: export the QAD checkpoint to HuggingFace format so it is
ready for local inference


All steps use the TE (Transformer Engine) spec, which with the new
TEGroupedMLP per-expert quantizer is approximately 10-15% faster than
the previous local ModelOpt spec (which used SequentialMLP) on
Hybrid-MoE models.

### Usage

```
# Usage from tools/launcher:
source .env-slurm
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, backward breaking changes,
deprecations, or fixes for critical bugs present in previous releases.
-->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Added a launcher configuration for NVFP4 quantization-aware
distillation of the Nemotron 3 Nano 30B-A3B model.
- Added support for selecting training or fine-tuning workflows through
`MLM_TRAIN_SCRIPT`.
  - Improved forwarding of additional training arguments.

- **Updates**
  - Updated the Megatron-LM launcher component to a newer revision.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
This commit is contained in:
Jenny Chen
2026-08-10 13:52:52 -07:00
committed by GitHub
parent 41b18a8568
commit f2bfe63183
3 changed files with 170 additions and 7 deletions
+19 -6
View File
@@ -24,8 +24,9 @@ trap 'error_handler $0 $LINENO' ERR # ERROR HANDLER
###################################################################################################
# Quantization-Aware Distillation / SFT training on a quantized MCore checkpoint
# (the output of common/megatron_lm/quantize/quantize.sh). Wraps the Megatron-LM
# ModelOpt post-training train.sh example (QAT: --modelopt-enabled).
# (the output of common/megatron_lm/quantize/quantize.sh). By default, wraps the
# Megatron-LM ModelOpt post-training train.sh example (QAT: --modelopt-enabled).
# Set MLM_TRAIN_SCRIPT=finetune to use finetune.sh for HF chat-template datasets.
#
# Extra flags (data/train/optim/eval, e.g. --sft, --seq-length, --lr,
# --dist-ckpt-strictness) are forwarded as CLI args and assembled into
@@ -37,6 +38,7 @@ trap 'error_handler $0 $LINENO' ERR # ERROR HANDLER
# MLM_MODEL_CKPT Quantized MCore ckpt to load (default: /cicd/megatron-lm/${MLM_MODEL_CFG})
# MLM_MODEL_SAVE Where to save the trained ckpt (default: MLM_MODEL_CKPT)
# HF_MODEL_CKPT HF source ckpt for tokenizer/config (default: /hf-local/${MLM_MODEL_CFG})
# MLM_TRAIN_SCRIPT Megatron-LM post_training/modelopt script: train (default) or finetune
# DP CP TP PP EP ETP Parallelism (defaults from train.sh)
if [[ -z ${MLM_MODEL_CFG} ]]; then
@@ -52,11 +54,22 @@ fi
export MLM_MODEL_SAVE=${MLM_MODEL_SAVE:-${MLM_MODEL_CKPT}}
export MLM_SKIP_INSTALL=1
TRAIN_EXE="bash modules/Megatron-LM/examples/post_training/modelopt/train.sh"
case "${MLM_TRAIN_SCRIPT:-train}" in
train)
TRAIN_EXE=(bash modules/Megatron-LM/examples/post_training/modelopt/train.sh)
;;
finetune)
TRAIN_EXE=(bash modules/Megatron-LM/examples/post_training/modelopt/finetune.sh)
;;
*)
echo "[ERROR] MLM_TRAIN_SCRIPT must be 'train' or 'finetune'." >&2
exit 1
;;
esac
export MLM_EXTRA_ARGS=${@}
echo "=== QAD/SFT training ${MLM_MODEL_CFG} (load ${MLM_MODEL_CKPT}) ==="
${TRAIN_EXE} ${MLM_MODEL_CFG}
export MLM_EXTRA_ARGS="$*"
echo "=== QAD/SFT training ${MLM_MODEL_CFG} via ${MLM_TRAIN_SCRIPT:-train} (load ${MLM_MODEL_CKPT}) ==="
"${TRAIN_EXE[@]}" "${MLM_MODEL_CFG}"
###################################################################################################
@@ -0,0 +1,150 @@
# NVIDIA Nemotron 3 Nano 30B-A3B NVFP4 quantization-aware distillation (QAD).
#
# The pipeline converts the Hugging Face BF16 model to an MCore teacher,
# quantizes a separate MCore student, distills the student for 400 iterations,
# and exports the resulting checkpoint. The training task uses an explicit
# Nemotron-Post-Training-Dataset-v2 chat shard so Hugging Face Datasets does not
# prepare the repository's other large splits.
#
# Training topology: 2 nodes x 8 GPUs, TP=2, PP=1, CP=1, EP=4, ETP=1.
# This gives DP=8 and EDP=4. With micro-batch-size=1 and global-batch-size=16,
# train-samples=6400 produces 400 iterations. To run at 16K instead, change both
# --seq-length and --max-position-embeddings to 16384.
#
# Requirements:
# - The BF16 model is mounted at /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.
# - HF_TOKEN can access the gated nvidia/Nemotron-Post-Training-Dataset-v2 dataset.
#
# Usage from tools/launcher:
# source .env-slurm
# uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes
job_name: Nemotron-3-Nano-30B-A3B_QAD_32k_400iter
pipeline:
allow_to_fail: false
skip: false
note: "NVFP4 TEGroupedMLP QAD at 32K for 400 iterations on one explicit Nemotron post-training chat shard"
# Import the BF16 Hugging Face checkpoint as the MCore teacher checkpoint.
task_0:
script: common/megatron_bridge/import/import.sh
environment:
- HF_MODEL_ID: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- OUTPUT_DIR: /cicd/megatron-lm-bf16/nvidia
- TORCH_DTYPE: bfloat16
slurm_config:
_factory_: "slurm_factory"
container: nvcr.io/nvidia/nemo:26.06.00
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
nodes: 1
ntasks_per_node: 1
gpus_per_node: 4
# Quantize the teacher checkpoint into the NVFP4 student checkpoint.
task_1:
script: common/megatron_lm/quantize/quantize.sh
args:
- --seq-length 4096 --max-position-embeddings 32768
- --calib-size 32
- --skip-generate
- --export-default-te-spec
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- QUANT_CFG: MAMBA_MOE_NVFP4_CONSERVATIVE_CFG
- MLM_MODEL_CKPT: /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-MCore
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- RUN_MMLU: "false"
- RUN_EXPORT: "false"
- DP: "1"
- CP: "1"
- TP: "2"
- PP: "1"
- EP: "4"
- ETP: "1"
slurm_config:
_factory_: "slurm_factory"
container: nvcr.io/nvidia/nemo:26.06.00
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
nodes: 1
ntasks_per_node: 8
gpus_per_node: 8
# Distill the quantized student from the BF16 teacher on chat-template data.
task_2:
script: common/megatron_lm/train/sft.sh
args:
# Data
- --seq-length 32768 --max-position-embeddings 32768
- --micro-batch-size 1 --global-batch-size 16
- --train-samples 6400
- --lr-decay-samples 6400
- --lr-warmup-samples 0
- --split 99,1,0
- --finetune-data-split chat
- --finetune-data-files data/chat-00000-of-00012.parquet
# QAD
- --modelopt-enabled
- --export-default-te-spec
- --export-kd-teacher-load /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-MCore
- --attention-dropout 0.0 --hidden-dropout 0.0
- --no-check-for-nan-in-loss-and-grad
- --recompute-granularity selective
- --recompute-modules layernorm moe
- --sequence-parallel
- --ckpt-fully-parallel-load --ckpt-fully-parallel-save
# Optimizer
- --lr 5.0e-6
- --lr-decay-style constant
- --clip-grad 1.0 --weight-decay 0.0
- --adam-beta1 0.9 --adam-beta2 0.95
- --init-method-std 0.010
- --use-distributed-optimizer
# Evaluation and checkpoints
- --eval-iters 2 --eval-interval 25
- --save-interval 50 --log-interval 10
- --dist-ckpt-strictness log_all
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- MLM_MODEL_CKPT: /cicd/megatron-lm/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- MLM_MODEL_SAVE: /cicd/megatron-lm-qad/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- MLM_TRAIN_SCRIPT: finetune
- DATASET: nvidia/Nemotron-Post-Training-Dataset-v2
- DP: "1"
- CP: "1"
- TP: "2"
- PP: "1"
- EP: "4"
- ETP: "1"
slurm_config:
_factory_: "slurm_factory"
container: nvcr.io/nvidia/nemo:26.06.00
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
nodes: 2
ntasks_per_node: 8
gpus_per_node: 8
task_3:
script: common/megatron_lm/export/export.sh
args:
- --export-default-te-spec
- --dist-ckpt-strictness log_all
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- QUANT_CFG: MAMBA_MOE_NVFP4_CONSERVATIVE_CFG
- MLM_MODEL_CKPT: /cicd/megatron-lm-qad/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- EXPORT_DIR: /cicd/export/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16_NVFP4_QAD
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- DP: "1"
- CP: "1"
- TP: "1"
- PP: "4"
- EP: "1"
- ETP: "1"
slurm_config:
_factory_: "slurm_factory"
container: nvcr.io/nvidia/nemo:26.06.00
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
nodes: 1
ntasks_per_node: 4
gpus_per_node: 4