mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data (#2134)
### What does this PR do? Type of change: New example Add a Nemotron Nano 3 QAD Launcher Example on OSS Nemotron-Post-Training-V2 data. It performs 4 steps 1. Teacher conversion: Convert the HuggingFace BF16 checkpoint to a Megatron-Core BF16 checkpoint 2. PTQ: quantize the Megatron-Core checkpoint to `MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` quant config 3. QAD (Quantization Aware Distillation): distill the BF16 checkpoint to the PTQ checkpoint on a subset of the Nemotron-Post-Training-V2 `chat` data. To train on a different subset or load the entire dataset, you may modify `--finetune-data-split` and `--finetune-data-files` flags. 4. Export: export the QAD checkpoint to HuggingFace format so it is ready for local inference All steps use the TE (Transformer Engine) spec, which with the new TEGroupedMLP per-expert quantizer is approximately 10-15% faster than the previous local ModelOpt spec (which used SequentialMLP) on Hybrid-MoE models. ### Usage ``` # Usage from tools/launcher: source .env-slurm uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added a launcher configuration for NVFP4 quantization-aware distillation of the Nemotron 3 Nano 30B-A3B model. - Added support for selecting training or fine-tuning workflows through `MLM_TRAIN_SCRIPT`. - Improved forwarding of additional training arguments. - **Updates** - Updated the Megatron-LM launcher component to a newer revision. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
This commit is contained in:
@@ -24,8 +24,9 @@ trap 'error_handler $0 $LINENO' ERR # ERROR HANDLER
|
||||
###################################################################################################
|
||||
|
||||
# Quantization-Aware Distillation / SFT training on a quantized MCore checkpoint
|
||||
# (the output of common/megatron_lm/quantize/quantize.sh). Wraps the Megatron-LM
|
||||
# ModelOpt post-training train.sh example (QAT: --modelopt-enabled).
|
||||
# (the output of common/megatron_lm/quantize/quantize.sh). By default, wraps the
|
||||
# Megatron-LM ModelOpt post-training train.sh example (QAT: --modelopt-enabled).
|
||||
# Set MLM_TRAIN_SCRIPT=finetune to use finetune.sh for HF chat-template datasets.
|
||||
#
|
||||
# Extra flags (data/train/optim/eval, e.g. --sft, --seq-length, --lr,
|
||||
# --dist-ckpt-strictness) are forwarded as CLI args and assembled into
|
||||
@@ -37,6 +38,7 @@ trap 'error_handler $0 $LINENO' ERR # ERROR HANDLER
|
||||
# MLM_MODEL_CKPT Quantized MCore ckpt to load (default: /cicd/megatron-lm/${MLM_MODEL_CFG})
|
||||
# MLM_MODEL_SAVE Where to save the trained ckpt (default: MLM_MODEL_CKPT)
|
||||
# HF_MODEL_CKPT HF source ckpt for tokenizer/config (default: /hf-local/${MLM_MODEL_CFG})
|
||||
# MLM_TRAIN_SCRIPT Megatron-LM post_training/modelopt script: train (default) or finetune
|
||||
# DP CP TP PP EP ETP Parallelism (defaults from train.sh)
|
||||
|
||||
if [[ -z ${MLM_MODEL_CFG} ]]; then
|
||||
@@ -52,11 +54,22 @@ fi
|
||||
export MLM_MODEL_SAVE=${MLM_MODEL_SAVE:-${MLM_MODEL_CKPT}}
|
||||
export MLM_SKIP_INSTALL=1
|
||||
|
||||
TRAIN_EXE="bash modules/Megatron-LM/examples/post_training/modelopt/train.sh"
|
||||
case "${MLM_TRAIN_SCRIPT:-train}" in
|
||||
train)
|
||||
TRAIN_EXE=(bash modules/Megatron-LM/examples/post_training/modelopt/train.sh)
|
||||
;;
|
||||
finetune)
|
||||
TRAIN_EXE=(bash modules/Megatron-LM/examples/post_training/modelopt/finetune.sh)
|
||||
;;
|
||||
*)
|
||||
echo "[ERROR] MLM_TRAIN_SCRIPT must be 'train' or 'finetune'." >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
export MLM_EXTRA_ARGS=${@}
|
||||
echo "=== QAD/SFT training ${MLM_MODEL_CFG} (load ${MLM_MODEL_CKPT}) ==="
|
||||
${TRAIN_EXE} ${MLM_MODEL_CFG}
|
||||
export MLM_EXTRA_ARGS="$*"
|
||||
echo "=== QAD/SFT training ${MLM_MODEL_CFG} via ${MLM_TRAIN_SCRIPT:-train} (load ${MLM_MODEL_CKPT}) ==="
|
||||
"${TRAIN_EXE[@]}" "${MLM_MODEL_CFG}"
|
||||
|
||||
###################################################################################################
|
||||
|
||||
|
||||
+150
@@ -0,0 +1,150 @@
|
||||
# NVIDIA Nemotron 3 Nano 30B-A3B NVFP4 quantization-aware distillation (QAD).
|
||||
#
|
||||
# The pipeline converts the Hugging Face BF16 model to an MCore teacher,
|
||||
# quantizes a separate MCore student, distills the student for 400 iterations,
|
||||
# and exports the resulting checkpoint. The training task uses an explicit
|
||||
# Nemotron-Post-Training-Dataset-v2 chat shard so Hugging Face Datasets does not
|
||||
# prepare the repository's other large splits.
|
||||
#
|
||||
# Training topology: 2 nodes x 8 GPUs, TP=2, PP=1, CP=1, EP=4, ETP=1.
|
||||
# This gives DP=8 and EDP=4. With micro-batch-size=1 and global-batch-size=16,
|
||||
# train-samples=6400 produces 400 iterations. To run at 16K instead, change both
|
||||
# --seq-length and --max-position-embeddings to 16384.
|
||||
#
|
||||
# Requirements:
|
||||
# - The BF16 model is mounted at /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16.
|
||||
# - HF_TOKEN can access the gated nvidia/Nemotron-Post-Training-Dataset-v2 dataset.
|
||||
#
|
||||
# Usage from tools/launcher:
|
||||
# source .env-slurm
|
||||
# uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/megatron_lm_qad.yaml --yes
|
||||
|
||||
job_name: Nemotron-3-Nano-30B-A3B_QAD_32k_400iter
|
||||
pipeline:
|
||||
allow_to_fail: false
|
||||
skip: false
|
||||
note: "NVFP4 TEGroupedMLP QAD at 32K for 400 iterations on one explicit Nemotron post-training chat shard"
|
||||
|
||||
# Import the BF16 Hugging Face checkpoint as the MCore teacher checkpoint.
|
||||
task_0:
|
||||
script: common/megatron_bridge/import/import.sh
|
||||
environment:
|
||||
- HF_MODEL_ID: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- OUTPUT_DIR: /cicd/megatron-lm-bf16/nvidia
|
||||
- TORCH_DTYPE: bfloat16
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
container: nvcr.io/nvidia/nemo:26.06.00
|
||||
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
|
||||
nodes: 1
|
||||
ntasks_per_node: 1
|
||||
gpus_per_node: 4
|
||||
|
||||
# Quantize the teacher checkpoint into the NVFP4 student checkpoint.
|
||||
task_1:
|
||||
script: common/megatron_lm/quantize/quantize.sh
|
||||
args:
|
||||
- --seq-length 4096 --max-position-embeddings 32768
|
||||
- --calib-size 32
|
||||
- --skip-generate
|
||||
- --export-default-te-spec
|
||||
environment:
|
||||
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- QUANT_CFG: MAMBA_MOE_NVFP4_CONSERVATIVE_CFG
|
||||
- MLM_MODEL_CKPT: /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-MCore
|
||||
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- RUN_MMLU: "false"
|
||||
- RUN_EXPORT: "false"
|
||||
- DP: "1"
|
||||
- CP: "1"
|
||||
- TP: "2"
|
||||
- PP: "1"
|
||||
- EP: "4"
|
||||
- ETP: "1"
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
container: nvcr.io/nvidia/nemo:26.06.00
|
||||
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
|
||||
nodes: 1
|
||||
ntasks_per_node: 8
|
||||
gpus_per_node: 8
|
||||
|
||||
# Distill the quantized student from the BF16 teacher on chat-template data.
|
||||
task_2:
|
||||
script: common/megatron_lm/train/sft.sh
|
||||
args:
|
||||
# Data
|
||||
- --seq-length 32768 --max-position-embeddings 32768
|
||||
- --micro-batch-size 1 --global-batch-size 16
|
||||
- --train-samples 6400
|
||||
- --lr-decay-samples 6400
|
||||
- --lr-warmup-samples 0
|
||||
- --split 99,1,0
|
||||
- --finetune-data-split chat
|
||||
- --finetune-data-files data/chat-00000-of-00012.parquet
|
||||
# QAD
|
||||
- --modelopt-enabled
|
||||
- --export-default-te-spec
|
||||
- --export-kd-teacher-load /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16-MCore
|
||||
- --attention-dropout 0.0 --hidden-dropout 0.0
|
||||
- --no-check-for-nan-in-loss-and-grad
|
||||
- --recompute-granularity selective
|
||||
- --recompute-modules layernorm moe
|
||||
- --sequence-parallel
|
||||
- --ckpt-fully-parallel-load --ckpt-fully-parallel-save
|
||||
# Optimizer
|
||||
- --lr 5.0e-6
|
||||
- --lr-decay-style constant
|
||||
- --clip-grad 1.0 --weight-decay 0.0
|
||||
- --adam-beta1 0.9 --adam-beta2 0.95
|
||||
- --init-method-std 0.010
|
||||
- --use-distributed-optimizer
|
||||
# Evaluation and checkpoints
|
||||
- --eval-iters 2 --eval-interval 25
|
||||
- --save-interval 50 --log-interval 10
|
||||
- --dist-ckpt-strictness log_all
|
||||
environment:
|
||||
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- MLM_MODEL_CKPT: /cicd/megatron-lm/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- MLM_MODEL_SAVE: /cicd/megatron-lm-qad/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- MLM_TRAIN_SCRIPT: finetune
|
||||
- DATASET: nvidia/Nemotron-Post-Training-Dataset-v2
|
||||
- DP: "1"
|
||||
- CP: "1"
|
||||
- TP: "2"
|
||||
- PP: "1"
|
||||
- EP: "4"
|
||||
- ETP: "1"
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
container: nvcr.io/nvidia/nemo:26.06.00
|
||||
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
|
||||
nodes: 2
|
||||
ntasks_per_node: 8
|
||||
gpus_per_node: 8
|
||||
|
||||
task_3:
|
||||
script: common/megatron_lm/export/export.sh
|
||||
args:
|
||||
- --export-default-te-spec
|
||||
- --dist-ckpt-strictness log_all
|
||||
environment:
|
||||
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- QUANT_CFG: MAMBA_MOE_NVFP4_CONSERVATIVE_CFG
|
||||
- MLM_MODEL_CKPT: /cicd/megatron-lm-qad/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- EXPORT_DIR: /cicd/export/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16_NVFP4_QAD
|
||||
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
|
||||
- DP: "1"
|
||||
- CP: "1"
|
||||
- TP: "1"
|
||||
- PP: "4"
|
||||
- EP: "1"
|
||||
- ETP: "1"
|
||||
slurm_config:
|
||||
_factory_: "slurm_factory"
|
||||
container: nvcr.io/nvidia/nemo:26.06.00
|
||||
modelopt_install_path: /opt/venv/lib/python3.12/site-packages/modelopt
|
||||
nodes: 1
|
||||
ntasks_per_node: 4
|
||||
gpus_per_node: 4
|
||||
Submodule tools/launcher/modules/Megatron-LM updated: 0429f9fe62...4a279f3552
Reference in New Issue
Block a user