### What does this PR do? Type of change: new tests Runs the VLM case of `test_qad` under context parallelism, so QAD on a Qwen3-VL model is covered on the path that until now could not run at all. `Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the batch has to hand it full-length `position_ids`. Megatron-Bridge's `get_batch` was sharding them too, leaving the rotary embedding at `seq / cp**2` against hidden states at `seq / cp`. The fix is upstream in [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243); this PR is the coverage that would have caught it. The VLM case moves from tensor to context parallelism and the PTQ step is sized to the same TP, so QAD still loads a matching checkpoint. The LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and `test_distill_vlm` still covers TP for a VLM, so nothing loses coverage. ### Usage ```bash # Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243 # it now works for VLMs, where it previously died in the rotary embedding. python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ... ``` ### Testing On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on `PYTHONPATH`: - `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD across 2 CP ranks, export; quantizers survive and the vision tower is byte-identical. **1 passed (183 s).** Without the upstream fix the same run dies with `AttributeError: 'NoneType' object has no attribute 'ndim'` in `rope.py:175`. - `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).** - Gate check: on today's `nemo:26.08` (no #6243) the probe resolves `False` and the VLM case stays at `cp_size=1`, byte-identical to current CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a 1-GPU runner it degenerates to today's config either way. - `pre-commit run --files ...` clean (ruff check, ruff format, mypy, bandit, markdownlint). ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — existing tests extended rather than new ones added. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — test coverage and one doc line; no feature, break, deprecation, or fix for a released bug. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information - Depends on [NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243). Safe to merge before it lands: the gate keeps the VLM case at `cp_size=1` until a container ships the fix, at which point the coverage switches on by itself. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the Qwen3.6 QAD instructions to keep tensor and pipeline parallelism set to 1, while allowing context parallelism to increase for longer sequences with the `nemo:26.10` container. * **Tests** * QAD validation now selects parallelism settings based on whether the Megatron-Bridge context-parallel fix is available, and reports when multi-GPU VLM coverage is reduced. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Megatron Bridge
This directory contains examples of using Model Optimizer with the NeMo Megatron-Bridge framework for quantization, pruning, and distillation. These workflows can be used on their own or combined.
| Section | Description | Link |
|---|---|---|
| Pre-Requisites | Development environment setup | [Link] |
| Post-Training Quantization | Quantizing a model | [Link] |
| Distillation | Distilling a pruned or quantized model | [Link] |
| Pruning | Pruning a model using Minitron algorithm | [Link] |
| Sanity-Check Generation | Quick generation check with vLLM | [Link] |
| Resources | Extra links to relevant resources | [Link] |
Tip
Checkout the Nemotron-3-Nano-30B-A3B pruning + distillation (with data blend prep) + quantization tutorial for a complete end-to-end workflow using Megatron-Bridge!
Or the Qwen3.6-35B-A3B W4A4 NVFP4 + QAD tutorial for an end-to-end quantization-aware distillation workflow.
Pre-Requisites
Running these examples requires many additional dependencies to be installed (e.g., Megatron-Bridge, Megatron-core, etc.), hence we strongly recommend directly using the NeMo container (e.g., nvcr.io/nvidia/nemo:26.08) which has all the dependencies installed.
To get the ModelOpt examples scripts, mount your Model-Optimizer repo to the container as follows:
export MODELOPT_DIR=${PWD}/Model-Optimizer # or set to your local Model-Optimizer repository path if you have cloned it
if [ ! -d "${MODELOPT_DIR}" ]; then
git clone https://github.com/NVIDIA/Model-Optimizer.git ${MODELOPT_DIR}
fi
export DOCKER_IMAGE=nvcr.io/nvidia/nemo:26.08
docker run \
--gpus all \
--shm-size=16GB \
--net=host \
--ulimit memlock=-1 \
--rm -it \
-v ${MODELOPT_DIR}:/opt/Model-Optimizer \
-v ${MODELOPT_DIR}/modelopt:/opt/venv/lib/python3.12/site-packages/modelopt \
-v ${MODELOPT_DIR}/modelopt_recipes:/opt/venv/lib/python3.12/site-packages/modelopt_recipes \
-w /opt/Model-Optimizer/examples/megatron_bridge \
${DOCKER_IMAGE} bash
Warning
Use
python -m pipinstead ofpipto avoid conflicts with the system-wide installed packages in the NeMo containers. You may also refer to this doc on how to correctly install packages in the NeMo containers without breaking existing torch installation.
You also need to login with your HuggingFace token to download gated datasets / models.
Note that the default dataset for pruning and quantization is nemotron-post-training-dataset-v2, which is gated.
hf auth login --token <your token>
Importing a HuggingFace Checkpoint (optional)
The scripts below take a HuggingFace checkpoint directly, but if you want a Megatron distributed
checkpoint (e.g. to reuse across runs), convert it with Megatron-Bridge's conversion script — use
--tp / --pp / --ep to shard a model that does not fit on one GPU (--ep for MoE models), or
--device cpu to convert in a single process without GPUs:
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
--executor local \
--device gpu \
--gpus-per-node 8 \
--hf-model Qwen/Qwen3-8B \
--megatron-path /tmp/Qwen3-8B-megatron
Post-Training Quantization
This section shows how to quantize a HuggingFace model using ModelOpt in the Megatron-Bridge framework. Quantization is a two-step flow:
- quantize.py applies post-training quantization (PTQ) with calibration and saves a Megatron checkpoint (with ModelOpt state). Tensor / pipeline / expert parallelism are all supported, and the checkpoint can be reloaded for further training (Quantization Aware Training / Quantization Aware Distillation).
- export_quantized_megatron_to_hf.py converts that Megatron checkpoint to a HuggingFace (unified) checkpoint that deploys directly with TensorRT-LLM, vLLM, or SGLang.
quantize.py supports the following formats via --quant_cfg (e.g. fp8, nvfp4, int8_smoothquant, int4_awq, w4a8_awq_beta, ...). You can also pass any full config name exposed by ModelOpt (e.g. NVFP4_DEFAULT_CFG) or a YAML --recipe (e.g. general/ptq/nvfp4_default-kv_fp8, authoritative for quant_cfg + algorithm + KV-cache). KV-cache quantization can be enabled on top via --kv_cache_quant (e.g. fp8, nvfp4).
Step 1 — quantize Qwen3-8B to NVFP4 on 2 GPUs (Tensor Parallelism = 2) using 1024 samples from default dataset (Mix of cnn_dailymail and nemotron-post-training-dataset-v2) for calibration (sequence length = 4096):
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--quant_cfg nvfp4 \
--tp_size 2 \
--calib_batch_size 1 \
--seq_length 4096 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron
Note
Data parallelism is implicit:
DP = world_size / (tp_size * pp_size * cp_size). Launching with more GPUs thantp_size * pp_size * cp_sizeshards calibration across the extra data-parallel ranks (e.g.torchrun --nproc_per_node 8 quantize.py --tp_size 2runs with DP=4).
Step 2 — export the Megatron checkpoint to a deployable HuggingFace checkpoint:
torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--pp_size 2 \
--export_unified_hf_path /tmp/Qwen3-8B-NVFP4-hf
Note
The HuggingFace unified exporter can't split weights across GPUs with tensor parallelism. For large models, use
--pp_sizeonexport_quantized_megatron_to_hf.pyto shard the export across GPUs with pipeline parallelism instead.
Tip
To recover the accuracy lost during quantization, fine-tune the quantized Megatron checkpoint (from step 1) with Quantization Aware Distillation (QAD) before running the step 2 export.
To see the full usage for advanced configurations, run torchrun --nproc_per_node 1 quantize.py --help (or export_quantized_megatron_to_hf.py --help).
Vision-Language Models (VLMs)
For a vision-language model (e.g. Qwen3.5-VL, Gemma3-VL), quantize.py automatically quantizes only the language model and leaves the vision tower and vision-language projector in full precision, then saves the full VLM back as a Megatron checkpoint. The calibration modality is inferred from --calib_dataset_name:
- An image-text dataset (the default for VLMs,
nemotron_vlm_dataset_v2) drives the full VLM forward, so the language model is calibrated on vision-conditioned activations. - A text dataset runs text-only calibration of the language model (vision tower idle).
Note
HuggingFace unified export (
export_quantized_megatron_to_hf.py) of a quantized VLM covers Qwen3-VL and Qwen3.5-VL. Other VLMs such as Gemma3-VL are saved in Megatron checkpoint format only.
Tracking runs with MLflow
Set MLflow's own MLFLOW_TRACKING_URI, or pass --mlflow <tracking-uri>, to record a run on an MLflow server. Every script here that writes a checkpoint takes the flag — prune_minitron.py, quantize.py, distill.py, export_quantized_megatron_to_hf.py and export_distilled_megatron_to_hf.py — and they share one experiment-name convention, so a pruning, a quantization, the distillation that refines its checkpoint and the export that deploys it can be found together. (One gap remains: distill.py --hf_export_path writes its HuggingFace checkpoint from rank 0, which is not the rank that owns that run, so it carries no pointer.)
torchrun --nproc_per_node 2 quantize.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--recipe general/ptq/nvfp4_default-kv_fp8 \
--tp_size 2 \
--export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--mlflow https://<your-mlflow-server>/
The run opens before the model loads, so a bad URI fails in seconds rather than after a full calibration. Only the master rank uploads: the invocation, every argument as a searchable param, the resolved recipe, that rank's log and the quantizer summary — plus .experiment.json written into --export_megatron_path once the checkpoint is saved, so a checkpoint on disk names the run that produced it. A failed run is still recorded, with its traceback.
--mlflow_experiment defaults to $USER/<script>/<model basename>-<variant> — the variant being the pruning target for prune_minitron.py, the recipe name for quantize.py, the student checkpoint for distill.py, and the Megatron checkpoint for the export — and --mlflow_run_name to the UTC start time. Authentication uses MLflow's own environment variables. See the hf_ptq README for the full artifact list and the $MLFLOW_TRACKING_URI semantics.
distill.py — whether it is distilling a pruned BF16 student or running QAD from a quantized one — works differently under the hood. It is a training loop, and Megatron-Bridge logs to MLflow from inside it, so the run is opened here and Megatron-Bridge joins it: you get the invocation, every argument and a run log, plus the per-iteration training metrics and the full resolved config that only Megatron-Bridge can see, alongside the existing --wandb_project and TensorBoard logging. Two differences worth knowing:
- Uploading checkpoints as MLflow artifacts is off by default here, where Megatron-Bridge turns it on: a distilled checkpoint is tens to hundreds of GB and would be pushed over HTTP on every save. Pass
--mlflow_log_checkpointsto opt in. - The run is opened on the last rank, because that is the rank Megatron-Bridge looks at for one to join. The uploaded
logs/distill.logis therefore that rank's output;print_rank_0lines, which is most of what the script itself prints, stay on rank 0 and do not reach it.
Distillation
This section shows how to distill a student model from a teacher model in the Megatron-Bridge framework.
This can be used stand-alone or after Pruning / Post-Training Quantization to recover accuracy of the model by distilling from the original model (teacher).
The distill.py script supports both standard HuggingFace checkpoints and Puzzletron AnyModel checkpoints as student/teacher inputs. Just pass the checkpoint path via --student_hf_path / --teacher_hf_path. The distilled model is saved to <output_dir>/checkpoints in Megatron distributed checkpoint format.
To distill a student whose weights live in a Megatron checkpoint (e.g. a quantized checkpoint from quantize.py for Quantization Aware Distillation, or a pruned checkpoint), additionally pass --student_megatron_path — --student_hf_path is still required to build the student architecture.
Data Preparation
The distillation script expects pre-tokenized data in Megatron's binary format (.bin / .idx files).
See the Dataset Preparation README
for full instructions on tokenizing JSONL files and Hugging Face datasets and get the list of output prefixes that you can use for --data_paths argument.
Alternatively, pass --sft --sft_dataset_root <dir> to distill on raw prompt-completion JSONL
with the loss masked to the completion. The directory must hold training.jsonl (and
validation.jsonl when --eval_iters > 0) of {"input": <prompt>, "output": <response>} records, which are tokenized with
the model's own HuggingFace tokenizer. Both fields are tokenized as written, except that
leading and trailing spaces on each field are stripped — no chat template is applied. So if your
model expects role/turn markers, include them in the "input" field yourself, and express any
significant separator as a newline rather than a trailing space. A BOS token is prepended
automatically when the tokenizer prepends one at inference, so do not add it yourself; an EOS
token is appended after the response. A record longer than --seq_length is truncated from the
start of "input", which drops any system prompt or opening role marker baked in there, so
pre-filter or pre-truncate the corpus if that matters.
Teacher and student must share a tokenizer — distillation scores the teacher on the student's token ids, and the KD losses compare the two models' logits elementwise over the vocab dimension.
Distillation with Real Data
Example usage to distill a 4B student (HF) from an 8B teacher (HF) on 8 GPUs (TP=8, PP=1):
torchrun --nnodes 1 --nproc_per_node 8 distill.py \
--tp_size 8 \
--teacher_hf_path Qwen/Qwen3-8B \
--student_hf_path Qwen/Qwen3-4B \
--data_paths 1.0 tokenized_qwen3/data1_text_document 1.0 tokenized_qwen3/data2_text_document \
--data_path_to_cache /path/to/cache/dataset_indices_qwen3 \
--seq_length 8192 \
--mbs 1 \
--gbs 768 \
--train_iters 15000 \
--lr 1e-4 \
--min_lr 1e-5 \
--lr_warmup_iters 50 \
--eval_interval 100 \
--eval_iters 32 \
--log_interval 10 \
--output_dir /output/qwen3_8b_to_4b_distill
Tensorboard logging is enabled by default and logs are saved to <output_dir>/tensorboard directory.
To use Weights & Biases for logging, set the WANDB_API_KEY environment variable and pass the --wandb_project argument.
Optionally, you can also pass --wandb_entity and --wandb_exp_name arguments to group runs under a project and experiment name.
To measure the initial student's CE and distillation losses, add --validate_only to the command.
This skips training and evaluates the student at iteration 0.
To see all available arguments:
torchrun --nproc_per_node 1 distill.py --help
Quick Test with Mock Data
Example usage with mock data for quick testing (no pre-tokenized data needed):
torchrun --nproc_per_node 8 distill.py \
--tp_size 8 \
--teacher_hf_path Qwen/Qwen3-0.6B \
--student_hf_path Qwen/Qwen3-0.6B \
--use_mock_data \
--seq_length 512 \
--mbs 1 \
--gbs 8 \
--train_iters 100 \
--eval_interval 10 \
--eval_iters 4 \
--output_dir /tmp/test_distill
Vision-Language Models (VLMs)
For a vision-language model (e.g. Qwen3.5-VL, Gemma3-VL), distill.py distills only the language model (on text data) and leaves the vision tower and projector untouched — matching the pruning and quantization behavior. It composes with pruning and QAD (--student_megatron_path) exactly as for LLMs, and the HF export reuses --student_hf_path (no --student_hf_model needed).
torchrun --nproc_per_node 8 distill.py \
--tp_size 8 \
--teacher_hf_path Qwen/Qwen3-VL-2B-Thinking \
--student_hf_path Qwen/Qwen3-VL-2B-Thinking \
...
Converting to Hugging Face format (optional)
A non-quantized distilled checkpoint (LLM or VLM) is saved in Megatron distributed format. If you need a HuggingFace checkpoint, there are two ways to convert it (for a QAD checkpoint, which retains quantization state, use export_quantized_megatron_to_hf.py instead — see QAD):
Inline -- add --hf_export_path to the distill.py command to automatically convert the final checkpoint after distillation:
torchrun --nnodes 1 --nproc_per_node 8 distill.py \
... \
--hf_export_path /path/to/save/distilled_hf_ckpt
--student_hf_path builds the student and provides the exported config / tokenizer. --student_hf_model is a reference HF model with a homogeneous architecture, used as the export template only for heterogeneous (Puzzletron/NAS) students; for homogeneous models and VLMs, omit it -- it defaults to --student_hf_path.
Separate conversion -- convert any saved iteration (intermediate or final) with export_distilled_megatron_to_hf.py:
torchrun --nproc_per_node 1 export_distilled_megatron_to_hf.py \
--student_hf_path <student_hf_model_or_path> \
--megatron_path <distill_output_dir>/checkpoints/iter_<iter_number> \
--hf_export_path /path/to/save/distilled_hf_ckpt
Use --export_iterations to export multiple saved checkpoints, for example to evaluate how model
quality changes during distillation. To export multiple iterations, keep those Megatron checkpoints
during distillation. The default is to keep the last 5 checkpoints; set --checkpoint_keep_last -1
to keep all saved checkpoints.
Then export all retained checkpoints, with one Hugging Face checkpoint written per
iter_<iteration> subdirectory:
torchrun --nproc_per_node 1 export_distilled_megatron_to_hf.py \
--student_hf_path <student_hf_model_or_path> \
--megatron_path <distill_output_dir>/checkpoints \
--hf_export_path /path/to/save/hf_validation_checkpoints \
--export_iterations all
The export path contains one loadable Hugging Face checkpoint per exported iteration:
hf_validation/
├── iter_0000100/
├── iter_0000200/
└── iter_0000300/
To export selected iterations instead, use --export_iterations 200 400 600.
Quantization Aware Distillation (QAD)
To recover the accuracy lost during Post-Training Quantization, distill the quantized model (student) from the original, unquantized model (teacher). Pass the quantized Megatron checkpoint produced by quantize.py via --student_megatron_path (the ModelOpt quantizers are restored automatically, so distillation trains the fake-quantized student), while --student_hf_path provides the student architecture and --teacher_hf_path points to the original unquantized model.
If you do not already have a suitable QAD dataset, start with data/nemotron-cascade-2-blend.yaml. It defines a general-purpose mixture of SFT data for QAD. Copy it, set the tokenizer for the target model, and adjust the output directory, sources, and weights as needed before preparing data. Its default 17.3-billion-token budget covers 1000 iterations at global batch size 512 and sequence length 32768, including a 1% validation holdout and margin. Recalculate the budget when changing those settings, and keep the prepared data unchanged when resuming.
We also use a smaller learning rate for QAD:
torchrun --nproc_per_node 8 distill.py \
--tp_size 8 \
--teacher_hf_path Qwen/Qwen3-8B \
--student_hf_path Qwen/Qwen3-8B \
--student_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
--data_paths 1.0 tokenized_qwen3/data1_text_document 1.0 tokenized_qwen3/data2_text_document \
--data_path_to_cache /path/to/cache/dataset_indices_qwen3 \
--seq_length 8192 \
--gbs 768 \
--train_iters 1000 \
--lr 1e-5 \
--min_lr 5e-6 \
--output_dir /output/qwen3_8b_nvfp4_qad
The distilled checkpoint retains the ModelOpt quantization state, so it can be converted to a deployable HuggingFace checkpoint with export_quantized_megatron_to_hf.py (point --megatron_path at /output/qwen3_8b_nvfp4_qad/checkpoints), exactly like the PTQ checkpoint in step 2 above.
Slurm Usage
To run the distillation script on a Slurm cluster for multi-node training, you just need use python instead of torchrun and set the number of nodes using #SBATCH --nodes=<num_nodes> clause in your Slurm script.
Distillation Results
See examples/pruning/ for distillation experiment results covering Minitron and Puzzletron pruning algorithms.
Pruning
This section shows how to prune a HuggingFace model using Minitron algorithm in Megatron-Bridge framework. Checkout other available pruning algorithms, supported frameworks and models, and general pruning getting-started in the pruning README.
The script supports three NAS-based pruning targets and one manual export mode:
| Mode | Flag | Description |
|---|---|---|
| NAS | --prune_target_params |
Prune to a target total parameter count |
| NAS | --prune_target_active_params |
Prune to a target active parameter count (useful for MoE models). For non-MoE models, this is equivalent to --prune_target_params. |
| NAS | --prune_target_memory_mb |
Prune to a target memory footprint in MB (weights + KV-cache) for a given batch size and sequence length assuming BF16 precision |
| Manual | --prune_export_config |
Prune directly to a specified architecture config (no NAS). Useful if you want to take top K candidates and do a short knowledge distillation before selecting the best model. |
Multiple NAS targets can be combined — e.g. --prune_target_params 6e9 --prune_target_memory_mb 12288 finds the best model with under 6B params and under 12GB memory footprint at (default) batch size 1 and sequence length 4096 assuming BF16 precision.
Prune by total parameter count — prune Qwen3-8B to 6B on 2-GPUs (Pipeline Parallelism = 2) while skipping pruning of num_attention_heads using following defaults:
1024 samples from nemotron-post-training-dataset-v2 for calibration,
at-most 20% depth (num_layers) and 40% width is pruned per prunable hparam (hidden_size, ffn_hidden_size, ...),
top-10 candidates are evaluated for MMLU score (5% sampled data) to select the best model.
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_target_params 6e9 \
--hparams_to_skip num_attention_heads \
--output_hf_path /tmp/Qwen3-8B-Pruned-6B
Prune by active parameter count — useful for MoE models where most experts are inactive per token (e.g. prune Nemotron-3-Nano-30B-A3B-BF16 (3.6B active params) to 3B active params):
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
--prune_target_active_params 3e9 \
--output_hf_path /tmp/Nemotron-3-Nano-30B-A3B-BF16-Pruned-3B-Active
Prune by memory footprint — prune to fit a target GPU memory budget (weights + KV-cache at the given sequence length and batch size, assuming BF16):
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_target_memory_mb 12288 \
--seq_length 4096 \
--calib_batch_size 1 \
--output_hf_path /tmp/Qwen3-8B-Pruned-12GB
Manual pruning — prune directly to a specified architecture (no NAS, no score evaluation):
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_export_config '{"hidden_size": 3584, "ffn_hidden_size": 9216}' \
--output_hf_path /tmp/Qwen3-8B-Pruned-6B-manual
To see the full usage for advanced configurations, run:
torchrun --nproc_per_node 1 prune_minitron.py --help
Tip
If number of layers in the model is not divisible by number of GPUs i.e. pipeline parallel (PP) size, you can configure uneven PP by setting
--num_layers_in_first_pipeline_stageand--num_layers_in_last_pipeline_stage. E.g. for Qwen3-8B with 36 layers and 8 GPUs, you can set both to 3 to get 3-5-5-5-5-5-5-3 layers per GPU.
Note
NAS-based pruning requires ~2x the GPU memory of Manual pruning because it needs to simultaneously hold original model while evaluating each pruned candidate.
Note
Multi-token-prediction (MTP) heads (e.g. Qwen3.5) are not pruned yet — they are dropped for the prune run and the saved checkpoint has no MTP. Autoregressive inference is unaffected; for speculative decoding, run a short MTP SFT on the pruned model.
Vision-Language Models (VLMs)
For a vision-language model (e.g. Qwen3.5-VL, Gemma3-VL), prune_minitron.py automatically prunes only the language model and leaves the vision tower intact, then saves the full VLM back. All the pruning modes above (parameter count, active parameter count, memory footprint, and manual export_config) work unchanged, with two VLM-specific caveats:
- The
--prune_target_params/--prune_target_active_params/--prune_target_memory_mbtargets (andexport_configdimensions) apply to the language model only — the (unpruned) vision tower's parameters are not counted, so the full saved VLM will be larger than the target. hidden_sizeis never pruned for VLMs (it is shared with the vision projector).
torchrun --nproc_per_node 2 prune_minitron.py \
--pp_size 2 \
--hf_model_name_or_path Qwen/Qwen3.5-4B \
--prune_target_params 3e9 \
--output_hf_path /tmp/Qwen3.5-4B-Pruned-3B
Sanity-Check Generation
generate_vllm.py runs a quick generation check on an exported HuggingFace checkpoint using vLLM — a useful smoke test for a quantized, pruned, or distilled model to confirm it still produces coherent text. For quantized checkpoints, vLLM auto-detects the ModelOpt quantization from the exported hf_quant_config.json, so no extra flags are needed:
# Quantized model
python generate_vllm.py --model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 --trust_remote_code
# Pruned model
python generate_vllm.py --model /tmp/Qwen3-8B-Pruned-6B
Note
--trust_remote_codeis only needed for models that ship custom modeling code (e.g. Nemotron); Qwen models don't require it.