33 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
Keval MorabiaandClaude Opus 5 f2f0d6958e Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?

Type of change: new feature

`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.

Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:

- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.

Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.

One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.

The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).

### Usage

```bash
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --recipe general/ptq/nvfp4_default-kv_fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
    --mlflow https://<your-mlflow-server>/

# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```

The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.

### Testing

- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.

- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:25:30 -07:00
Keval MorabiaandClaude Opus 5 a448ba9757 Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 02:10:12 +05:30
Keval MorabiaandClaude Opus 5 61757c9781 Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?

Type of change: Bug fix + new feature

Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.

Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.

#### Two blockers

1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.

#### Four silent-corruption bugs

3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.

#### Four more bugs, found only by running real checkpoints

The tiny fixtures could not reach these; each came from a real model or
a real quant format.

7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.

`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.

#### New: Qwen3.5-VL

`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.

#### New: the export path verifies itself

- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.

The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.

### Usage

```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --quant_cfg nvfp4 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron

torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
    --pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf

# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```

### Testing

All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.

| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |

The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.

#### Model coverage

`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.

| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |

| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |

QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.

#### Real-model validation

Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.

| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |

The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.

**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.

Two limitations worth stating plainly:

- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
  helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.

#### Guard verification

Each new guard was made to fire, not just to compile:

| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |

Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.

Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier

### Additional Information

**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.

Known gaps, unchanged by this PR:

- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
  * Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.

* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
  * Fixed Qwen3.5-VL GatedDeltaNet export handling.
  * Preserved visual-model weights exactly during export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 13:29:59 +00:00
Keval MorabiaandClaude Opus 5 d73278808b Bump nemo container requirement to 26.08 for MBridge examples (#2257)
### What does this PR do?

Type of change: Bug fix

Bumps the Megatron-Bridge examples, tests and launcher configs to
`nemo:26.08` and removes the version-gated fallbacks they carried, plus
the fixes needed to make the suites green on that container.

**26.08 bump and shim removal**

- Examples, CI workflows, `noxfile.py` and the `mbridge_*` launcher
configs move to `nemo:26.08`.
- `examples/megatron_bridge/_distillation_provider.py` is deleted —
26.08's Megatron-Bridge ships `convert_to_distillation_provider(...,
distill_submodule=...)` natively, so `distill.py` imports it directly.
- `prune_minitron.py` drops the `AutoBridge.from_hf_config` /
config-only-export probing; `--no_moe_grouped_gemm` is no longer needed
in the MoE pruning tests, and the Qwen3.5-MoE `skipif` is gone (native
MoE expert mappings are in 26.08).
- `_DynamicMambaMixer` targets only the raw `conv1d_weight` /
`conv1d_bias` parameters that replaced the `conv1d` module in
Megatron-Core.

**MambaModel / MambaModelProvider removal**

Megatron-Core has shipped `HybridModel` since 26.06 and `MambaModel` is
a deprecated subclass that shares its `forward`, so `DMRegistry`
resolves those instances to the `HybridModel` registration and the
separate entry is redundant. Same for `MambaModelProvider` vs
`HybridModelProvider` on the bridge side. `MambaMixer` / `MambaLayer` /
`ExtendedRMSNorm` are untouched — the layers still exist. The deprecated
`get_te_mamba_stack_spec` is removed; use `get_te_hybrid_stack_spec`.

**Bug fix: compressed output_layer extra state**

`mtq.compress` converts even a *disabled* `output_layer` into a
`RealQuantLinear` (its weight is left uncompressed, since
`pack_real_quantize_weight` skips disabled quantizers). The guard added
in #2112 excluded `RealQuantLinear` by class, so the layer emitted extra
state and every worker died in `GPTModel.sharded_state_dict`:

```
RuntimeError: Boolean value of Tensor with more than one value is ambiguous
  megatron/core/models/gpt/gpt_model.py:896, in sharded_state_dict
    output_extra_state and output_extra_state.data
```

The guard now keys off whether the weight was actually compressed
(`QTensorWrapper`) instead of the class. This took out all 12
`test_homogeneous_compressed_sharded_state_dict` params, and the crashed
workers poisoned the pool, which surfaced as unrelated timeouts and NCCL
errors in `test_layer_sync_moe_local_experts_amax`,
`test_kv_cache_quant`, `test_kv_cache_amax_sync`,
`test_convert_mcore_te_gpt_model` and
`test_homogeneous_sharded_state_dict_te_spec` — 21 tests in total. The
e2e coverage is `skip_flaky_on_blackwell`, so CI never ran it;
`test_output_layer_extra_state_empty_when_nothing_quantized` now asserts
the contract directly and is not skipped.

**Checkpoint import entry point**

26.08 replaced `examples/conversion/convert_checkpoints.py` with
`scripts/conversion/convert.sh`, so
`tools/launcher/common/megatron_bridge/import/import.sh` and the three
README snippets are retargeted. `import.sh` uses the distributed GPU
backend with `GPUS_PER_NODE` / `TP` / `PP` / `EP` knobs.

**Megatron-LM on nemo:26.06** keeps working: `_get_mamba_conv1d` still
dispatches between the `conv1d` module (26.06 and earlier) and the raw
parameters (26.08+), so `import_mcore_gpt_from_hf` /
`export_mcore_gpt_to_hf` handle NemotronH on both. Only the
Megatron-Bridge examples and Minitron pruning of Mamba/hybrid models
require 26.08.

**Test consolidation**

`test_export_distilled_megatron_to_hf.py` is merged into
`test_distill.py`: `test_distill_llm` becomes
`test_distill_llm_hf_export` and covers the standalone
`--export_iterations all` run on the checkpoints it already produces,
saving one full distillation (~185 s of CI time). The two mamba-named
gpu test files are renamed to `hybrid`.

### Usage

```bash
# HF -> Megatron import, via Megatron-Bridge's 26.08 conversion entry point
bash /opt/Megatron-Bridge/scripts/conversion/convert.sh import \
    --executor local \
    --device gpu \
    --gpus-per-node 8 \
    --hf-model Qwen/Qwen3-8B \
    --megatron-path /tmp/Qwen3-8B-megatron
```

### Testing

All on `nvcr.io/nvidia/nemo:26.08`, 2x RTX 6000 Ada, no timeout
overrides:

- `tests/examples/megatron_bridge`: 16 passed, 1 skipped (28m14s). The
skip is the `gemma3vl` QAD param, now `@pytest.mark.manual` since
`qwen3_5_moe_vl` covers the VLM QAD path.
- `tests/gpu_megatron` (`_extensions`, `distill`, `export`, `opt`,
`peft`, `sparsity`, `speculative`, `utils`): 61 passed, 5 xpassed.
- `tests/gpu_megatron/torch/export` re-run after the conv1d dispatch
change: 27 passed.
- The 21 previously failing/hanging quantization tests: 21 passed (12 +
9).
- `tests/gpu_megatron/torch/{nas,prune}`: verified separately.

`import.sh` equivalence on a toy `qwen3_moe`, comparing all 12 weight
tensors after flattening each dist checkpoint with `dcp_to_torch_save` —
the GPU backend at 1 GPU, `--tp 2`, `--pp 2`, `--ep 2`, and `import.sh`
end-to-end (`GPUS_PER_NODE=2 EP=2`) are all byte-identical to `--device
cpu`.

`nemo:26.06` compatibility was checked directly in that image:
`megatron.core.models.hybrid.HybridModel`, the modelopt hybrid spec and
`hybrid_layer_pattern` are all present, while
`megatron.bridge.models.hybrid` and the bridge's `distill_submodule` are
not. The NemotronH round-trip test failed there before the conv1d
dispatch was restored and the dispatch is back in place; per project
convention the suites themselves only run on 26.08.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ⚠️ Megatron-Bridge examples plus
Minitron pruning of Mamba/hybrid models now require `nemo:26.08`.
Megatron-LM quantization and checkpoint export still run on
`nemo:26.06`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`test_output_layer_extra_state_empty_when_nothing_quantized` for the
compress fix; existing tests extended for the merged export coverage.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added guidance for importing Hugging Face checkpoints into Megatron
distributed format.
* Expanded distillation workflows to export selected or all checkpoint
iterations.

* **Improvements**
  * Expanded Hybrid model support across Megatron workflows.
  * Updated distributed import tooling with GPU and parallelism options.
  * Updated supported environments and examples to NVIDIA NeMo 26.08.

* **Bug Fixes**
* Corrected output-layer quantization state handling when quantization
is disabled.

* **Documentation**
  * Added compatibility guidance for current and legacy NeMo containers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:31:41 +05:30
Keval MorabiaandClaude Opus 5 fbcdc16c2d Remove deprecations marked in 0.45 and 0.46 (#2182)
### What does this PR do?

Type of change: Backward breaking change (deprecation removal)

Ahead of the 0.47 code freeze, this removes every deprecation still
outstanding from the previous two releases (0.45 and 0.46). Two are
intentionally left in place: the **Python 3.10** drop and the
**transformers 4.x** drop

| Deprecation | Marked in | Replacement |
| --- | --- | --- |
| `--auto_quantize_bits` / `_method` / `_score_size` / `_cost_model` /
`_active_moe_expert_ratio` | 0.46 | AutoQuantize `--recipe` |
| `examples/llm_ptq` symlink + `examples/vlm_ptq/` forwarder | 0.46 |
`examples/hf_ptq` (`--vlm` for VLMs) |
| `QuantizationArgumentsWithConfig` alias | 0.45 |
`QuantizationArguments` |
| `QFORMAT_ALIASES` short names | 0.45 | canonical preset basenames |
| `layerwise` bool + flat `layerwise_checkpoint_dir` | 0.45 | nested
`layerwise: {enable, checkpoint_dir}` |
| in-trainer `quant_cfg` / `--quant_cfg` | 0.45 | `--recipe` |

#### Two things worth a closer look

**1. The `use_sequential` alias goes too.** It is the pre-#1251 alias on
`QuantizeAlgorithmConfig.layerwise` and only ever carried a bool. Once
the bool form is rejected it cannot accept a valid value, so keeping it
would only produce a differently-worded validation error. Note the
direction is breaking either way (`extra="forbid"`): a pre-0.45
`modelopt_state` carrying `use_sequential: True` or a top-level
`layerwise_checkpoint_dir` now fails validation instead of being
migrated.

**2. Removing in-trainer `--quant_cfg` required two new recipes.** The
`examples/gpt-oss` QAT flow ran on `--quant_cfg
MXFP4_MLP_WEIGHT_ONLY_CFG` and no `general/ptq/` recipe covered it. This
PR adds `general/ptq/mxfp4_mlp_weight_only` and
`general/ptq/nvfp4_mlp_weight_only`, verified to `model_dump` identical
to `mtq.MXFP4_MLP_WEIGHT_ONLY_CFG` / `mtq.NVFP4_MLP_WEIGHT_ONLY_CFG`,
and migrates the gpt-oss README, both SFT configs, `sft.py` and
`tests/examples/gpt-oss/test_gpt_oss_qat.py`. `examples/llm_qat` was
already recipe-only.

### Usage

```bash
# AutoQuantize: --auto_quantize_* flags -> an AutoQuantize recipe
scripts/huggingface_example.sh --model $HF_PATH \
  --recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --calib_batch_size 4

# --qformat / --quant_cfg: short name -> canonical preset basename
#   int8_sq -> int8_smoothquant                nvfp4_mse           -> nvfp4_w4a4_weight_mse_fp8_sweep
#   int8_wo -> int8_weight_only                nvfp4_local_hessian -> nvfp4_w4a4_weight_local_hessian
#   w4a8_awq -> w4a8_awq_beta                  fp8_pb_wo           -> fp8_2d_blockwise_weight_only
#   nvfp4_awq -> nvfp4_awq_lite                fp8_pc_pt           -> fp8_per_channel_per_token
scripts/huggingface_example.sh --model $HF_PATH --quant int8_smoothquant

# VLM PTQ: examples/vlm_ptq -> examples/hf_ptq with --vlm
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --vlm

# gpt-oss QAT: --quant_cfg <CFG name> -> --recipe <recipe path>
accelerate launch --config_file configs/zero3.yaml sft.py \
  --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b \
  --recipe general/ptq/mxfp4_mlp_weight_only --output_dir gpt-oss-20b-qat
```

```python
# Layerwise calibration: bool / flat key -> nested LayerwiseConfig
quant_cfg["algorithm"] = {"method": "gptq", "layerwise": {"enable": True, "checkpoint_dir": "/path"}}
```

### Testing

- `tests/unit/recipe` (229 passed),
`tests/unit/torch/quantization/test_config_validation.py` (79 passed),
`tests/examples/hf_ptq/test_hf_ptq_args.py` (23 passed).
- Verified the two new recipes `model_dump` identical to the `mtq.*_CFG`
constants they replace.
- `ruff check modelopt/ examples/ tests/` clean; `ruff format --check`
clean on all changed Python files.
- GPU suites
(`tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py`,
`test_accelerate_gpu.py`, `test_gptq.py`) had their preset / layerwise
literals updated but were not run locally — relying on CI.
- `examples/llm_qat/ARGUMENTS.md` is hand-edited to match what the
`generate-arguments-md` hook emits; the generator could not run locally
(missing `transformers` package metadata in this environment).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — that is the point of the PR:
it removes shims deprecated in 0.45/0.46. Callers must move to the
replacements in the table above. Additionally, a pre-0.45
`modelopt_state` carrying `use_sequential` or a top-level
`layerwise_checkpoint_dir` will now fail config validation.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests migrated to
the surviving APIs;
`TestLayerwiseNestedConfig::test_legacy_forms_rejected` pins that the
bool form, the `use_sequential` alias and the flat checkpoint-dir key
are all rejected. Tests covering the removed shims were deleted.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up: the transformers 4.x drop deprecated in 0.46 is still
outstanding and will need its own PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added MXFP4 and NVFP4 weight-only quantization recipes for MLP and MoE
layers.
- Added shared layer exclusions for more accurate effective-bits
calculations.

## Improvements
- Updated PTQ, QAT, GPT-OSS, deployment, and quantization-format
examples with current recipe names and configuration formats.
- Standardized layerwise settings under nested configuration fields.

## Breaking Changes
- Removed deprecated AutoQuantize options, `quant_cfg` usage, format
aliases, legacy layerwise settings, and compatibility example paths.
- Recipe-based and nested configuration forms are now required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 11:43:47 +05:30
yueshen2016 686da8d893 feat(megatron-bridge): SFT-masked data support in distillation (#2113)
## What does this PR do ?

**Type of change:** New feature

**Overview:** Adds SFT-masked data support to the Megatron-Bridge
distillation example, so a
model can be distilled on prompt/response pairs with the loss masked to
the response.

Today `examples/megatron_bridge/distill.py` only consumes
pretraining-style data — `GPTDataset`
over pre-tokenized blends with `NullTokenizer` — so the loss is computed
over every token. When
distilling an instruction-tuned model it is usually preferable to train
on prompt/response pairs
and mask the loss to the response, matching how the model was
fine-tuned.

## Usage

```bash
python examples/megatron_bridge/distill.py \
  --teacher_hf_path <teacher> --student_hf_path <student> \
  --sft --sft_dataset_root /path/to/data \
  ...
```

where `/path/to/data` holds `training.jsonl` / `validation.jsonl` of
records:

```json
{"input": "<prompt>", "output": "<response>"}
```

## How it works

Switches the data path to Bridge's `FinetuningDatasetConfig` (NeMo-style
`GPTSFTDataset`):

* `prompt_template="{input}{output}"` tokenizes input+output verbatim —
adjacent placeholders,
  no separator — so the text is fed exactly as provided
* `label_key="output"` with `answer_only_loss=True` masks the loss to
the response
  (`answer_start_idx == len(context_ids)`)
* `truncation_field="input"` truncates the context when a pair exceeds
`seq_length`

Two supporting changes, both scoped to `--sft`:

* **Tokenizer.** SFT reads raw text, so it uses the model's real
HuggingFace tokenizer. The
pretraining path consumes pre-tokenized data and keeps `NullTokenizer`.
* **Loss reduction.** A response-only mask requires per-token loss to
combine correctly across
context-parallel ranks, so `calculate_per_token_loss` is enabled and
`average_in_collective`
  is disabled. Both are untouched on the pretraining path.

## Testing

Used for quantization-aware distillation of Nemotron-Nano-3 (W4A16
NVFP4) at `seq_length=32768`
with CP>1: 200 iterations, logits-distillation loss `3.37e-2 -> 1.91e-2`
monotonically, router
`seq_load_balancing_loss` steady, and the resulting checkpoint exports
and serves correctly.

Opt-in: without `--sft` the existing mock/blend data path is unchanged.

## Before your PR is "Ready for review"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes — purely additive and
opt-in behind `--sft`.
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added supervised fine-tuning (SFT) support for distillation workflows.
* Added configuration and validation for SFT dataset locations and
supported inputs.
* Added raw prompt/completion JSONL datasets with response-only loss
masking.
* Added truncation, end-of-sequence handling, and student-tokenizer
support without automatic chat templates or BOS tokens.
  * Added matching student and teacher vocabulary validation.
  * Preserved existing mock and GPT dataset modes for non-SFT runs.

* **Documentation**
* Documented required filenames, record format, tokenizer behavior, and
completion-only loss masking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-08-13 01:10:29 +00:00
Keval MorabiaandClaude Opus 5 f99523279a Minitron pruning fixes for Nemotron-3.5-Lightning-30B-A3B and Deepseek (#2159)
### What does this PR do?

Type of change: Bug fix + new feature

Two model families that could not be pruned end-to-end now can:

- **Nemotron-3.5-Lightning**
(`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`) — a native
`NemotronHForCausalLM` that ships without remote code and carries MTP
heads. Fixes a calibration crash and HF-export failures on the modern
Megatron-Bridge / transformers stack.
- **DeepSeek-V3** — fixes an MLA Q-LoRA crash during calibration, and
adds a `candidate_filter` search option to `mcore_minitron` so its
MoE-FFN dimensions stay prunable while remaining representable in HF.

Also makes a rank-local failure under pipeline parallelism fail fast
instead of stalling.

#### 1. Nemotron Lightning: prune + HF export
(`examples/megatron_bridge/prune_minitron.py`)

1. **MTP calibration crash.** On newer Megatron-LM, `mtp_process` is
derived from the hybrid *pattern*, not from `mtp_num_layers`. Setting
only `mtp_num_layers=0` in the calibration provider overrides was
insufficient: the provider's `finalize()` re-appended the MTP suffix to
`hybrid_layer_pattern` (because `mtp_hybrid_override_pattern` was still
set and `mtp_use_repeated_layer=True`), so `mtp_process=True` while
`mtp_num_layers=0` and the calibration forward hit `assert
self.config.mtp_num_layers > 0`. Fix: also clear
`mtp_hybrid_override_pattern` in the calibration overrides so MTP is
fully disabled (MTP heads are dropped from the pruned model, as before).

2. **HF export via a config-only bridge (hybrid models only).** The old
export built a *dummy* HF model to obtain the bridge, then streamed
weights. This breaks on native NemotronH because (a) native
`NemotronHConfig` makes `hybrid_override_pattern` a read-only property,
and (b) transformers 5.12 saves the input embedding under a different
key than the bridge mapping expects (`backbone.embedding` vs
`backbone.embeddings`); the mismatch made `build_conversion_tasks` drop
the embedding task on its owning rank, leaving an owner-less PP
placeholder that crashed `save_hf_weights` with `Object must exist on at
least one PP rank`. Fix: stream weights through a **config-only** bridge
(`AutoBridge.from_hf_config(hf_cfg).save_hf_pretrained(...)`), available
since Megatron-Bridge 0.5.0 (nemo:26.06). A config-only bridge has
`hf_keys=None`, so the embedding task is never dropped, no dummy model
is built, and the output uses the canonical HF key names.

This is **restricted to hybrid providers**, which are the only models
that need it; non-hybrids keep the dummy-model path that CI has always
exercised.

Writing the source artifacts is now rank-0-only. Every rank used to
write the source `config.json`, which races with the pruned
`config.json` that `save_hf_pretrained` writes from rank 0 alone: a late
write from another rank leaves a checkpoint whose config does not match
its weights. This reproduced intermittently on both Qwen3 and NemotronH
before the fix, and 3/3 clean runs after.

`save_hf_pretrained` takes no `trust_remote_code` argument — it reads
the flag **off the bridge** to fetch the source checkpoint's artifacts,
and `from_hf_config` cannot infer it because
`AutoConfig.from_pretrained` consumes the kwarg rather than storing it
on the config. So the flag is set explicitly on the bridge instance;
otherwise remote-code models would silently lose it.

3. **Config write-back correctness:**
- `hybrid_override_pattern` is only written for older remote-code
configs that lack `layer_types`; native configs carry the cadence in
`layer_types` (read-only `hybrid_override_pattern` is skipped).
- `n_shared_experts` is preserved (a fixed count) instead of being
re-derived by `moe_shared_expert_intermediate_size //
moe_ffn_hidden_size`, which is DeepSeek-style logic that would corrupt
NemotronH's count.

Non-hybrids, VLMs, and Megatron-Bridge builds without config-only export
keep the dummy-model path, with a `warn_rank_0` when a hybrid has to
fall back. The README's `transformers<5` workaround is **removed**: it
existed because the dummy-model path broke on transformers 5, and the
config-only path handles NemotronH on every supported container.

#### 2. `candidate_filter` for `mcore_minitron`
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

DeepSeek-style MoE configs have no explicit shared-expert-size field:
they size the shared expert as `n_shared_experts *
moe_intermediate_size`, where `moe_intermediate_size` is the (also
prunable) **routed** expert size. So only candidates with
`moe_shared_expert_intermediate_size % moe_ffn_hidden_size == 0` can be
written back to HF at all.

Candidates come from a Cartesian `product()` of independent per-hparam
choice lists, so no per-hparam restriction can express a constraint
*between* two hparams. New optional `candidate_filter` search-config key
(default `None`, so existing behaviour is unchanged): a callable that
rejects candidate configs before the metric computation, making the
search cheaper rather than more expensive. It receives **every**
supported hparam, with non-searched ones filled in from the model
config, so a filter still works when one of its hparams was skipped or
had a single choice.

Rejected candidates are not cached, so — like `score_func`, whose cached
scores are reused without re-validation — the filter is assumed
unchanged when resuming from a `checkpoint`.

`prune_minitron.py` wires this up for DeepSeek-style configs, so
**both** `moe_ffn_hidden_size` and `moe_shared_expert_intermediate_size`
stay prunable (the search then only picks shared sizes that are a
multiple of the routed one). A `--prune_export_config` that violates the
constraint never reaches the filter, so the export path now raises
`ValueError` instead of writing a checkpoint whose config disagrees with
its weights.

#### 3. MLA Q-LoRA pruning
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

Pruning any MLA model with `q_lora_rank` set died during calibration
with `AttributeError: 'tuple' object has no attribute 'view'`.

`hidden_size` importance estimation blanket-patches every
`TELayerNormColumnParallelLinear` with `return_layernorm_output=True` to
capture post-layernorm activations. When `q_lora_rank` is set, MCore
builds `linear_q_up_proj` as a `TELayerNormColumnParallelLinear` — the
Q-LoRA layernorm is fused into it, which is why `q_layernorm` is
`IdentityOp` — so it was patched too, even though its layernorm is over
the **latent rank**, not `hidden_size`. TE then returns `((out, ln_out),
bias)` and MCore's `q, _ = self.linear_q_up_proj(...)` leaves `q` a
tuple.

Isolated by probing the module before and after dynamic conversion:

| Setup | `linear_q_up_proj` returns | Forward |
| --- | --- | --- |
| Before conversion | `tuple(Tensor, NoneType)` | — |
| After conversion, no hooks | `tuple(Tensor, NoneType)` | OK |
| After conversion **+ importance hooks** | `tuple(tuple(Tensor,
Tensor), NoneType)` | AttributeError |

So conversion is innocent; registering the importance hooks is the
trigger. Fix: exclude MLA's Q/KV up-projections from both the patch and
unpatch loops. `test_mcore_mla_pruning` did not catch this because it
builds MLA without `q_lora_rank`, where MCore uses a plain
`linear_q_proj` and nothing is patched.

#### 4. Fail fast instead of stalling on a rank-local error under PP
(`modelopt/torch/utils/distributed.py`)

A rank raising inside a distributed entrypoint left the whole job
stalled until the process group timed out, with **no diagnostic output
at all**: the failing rank blocked in `cleanup()`'s barrier while its
peers blocked in `recv_from_prev_pipeline_rank_`, and Python only prints
a traceback once the enclosing `finally` returns. A crash on one rank
was indistinguishable from a slow job.

- `dist.cleanup()` skips the barrier when unwinding from an exception.
- New `dist.abort()` prints the traceback, flushes and exits
immediately. Skipping the barrier alone is **not** enough — a stack dump
showed the failing rank then blocking in `destroy_process_group` for the
same reason — so the error path must not tear the process group down at
all. `SystemExit` is re-raised rather than aborted, so an intentional
exit (e.g. the `--score_lower_bound` gate) keeps its exit code and
prints no traceback. Kept out of `cleanup()` so no library caller gets a
surprise process exit.
- Called from the entrypoints that wrap `main()` in `try/finally`: the
five `examples/megatron_bridge` scripts.

Measured on a 2-GPU PP run whose rank 0 raises during calibration: **10
min timeout kill with no visible error → 31s, exit 1, real traceback.**
This is a latent, pre-existing issue (the `try/finally` predates this
PR); it only surfaces on a failing PP run, which is why CI never hit it.

### Usage

```bash
torchrun --nproc_per_node 4 examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
    --pp_size 4 \
    --prune_target_active_params 3e9 \
    --output_hf_path /path/to/Nemotron-3.5-Lightning-30B-A3B-Pruned-A3.0B
```

### Testing

- **End-to-end on nemo:26.08.rc6** (4× GB300, transformers 5.12.1,
Megatron-Bridge with config-only export): pruning + export complete
(`EXIT=0`, "Saved pruned model … Done!"). The exported checkpoint has
canonical **plural** `backbone.embeddings.weight` keys, **0 MTP
tensors**, and a config that reloads correctly (`num_hidden_layers=52`
from `layers_block_type`, `n_shared_experts=1`,
`num_nextn_predict_layers=0`, pruned `hidden_size`/`mamba_*`/MoE dims,
reconstructed `hybrid_override_pattern`).

  <details>
<summary>Pruning search log (<code>--prune_target_active_params
3e9</code>)</summary>

  ```text
Top 10 Candidates with Scores

┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃

┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.49B │ 0.5406
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 2 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2427 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 3 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2643
│
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.4552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3712} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.5860
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2294 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3328} │ │ │ │
│ 7 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.5231
│
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 8 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 21.81B │ 0.5042 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 9 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2462 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 10 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.5685 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │

└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘


╭────────────────────────────────────────────────────────────────────────
Best Subnet
─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304,
'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,
'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.5860 │

╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭────────────────────────────────────────────────────── Pruned Model
Stats ───────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=8) weights: 42489.7 MB,
kv_cache: 384.0 MB, mamba_state: 190.5 MB, Total: 43064.2 MB │

╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
  ```

  </details>

- **`tests/examples/megatron_bridge/test_prune_minitron.py`** —
`nemotron_h` now exports to HF and reloads (previously it stopped at a
Megatron checkpoint, since the dummy-model path needed
`transformers<5`), plus an `n_shared_experts` config assertion; the dead
`megatron_format` branch is gone. It runs on the CI container:
**verified on nemo:26.06.01 (transformers 5.8.1) and nemo:26.08.rc6.**
-
**`tests/gpu_megatron/torch/prune/plugins/test_mcore_mamba_minitron_pruning.py`**
— the `nas_memory_mb` search test now passes a `candidate_filter` and
asserts the exact number of rejected candidates (256 of the 512-combo
grid) plus the surviving candidates' validity; its `expected_top_k`
goldens are regenerated accordingly. Because
`moe_shared_expert_intermediate_size` is in that test's skip list, this
also covers the model-config fallback for hparams that are not in the
search space.

Verified on 2 GPUs, on both the CI container (nemo:26.06.01) and
nemo:26.08.rc6:

| Test | Result |
| --- | --- |
| `test_prune_minitron[qwen3]` | PASSED on 26.06.01 and 26.08.rc6 |
| `test_prune_minitron[deepseek_v3]` | PASSED (52s) — MLA Q-LoRA +
`candidate_filter` end-to-end |
| `test_prune_minitron[nemotron_h]` | PASSED on 26.06.01 (58s) and
26.08.rc6 (61s) |
| `test_mcore_mamba_hybrid_pruning_nas_memory_mb` | PASSED |
| `test_mcore_mamba_hybrid_pruning_nas_params` | PASSED (unchanged
sibling, run to check the regenerated goldens did not disturb it) |
| 2-GPU PP run failing on rank 0 | fails in 31s with a real traceback
(was a 10 min stall) |


### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `candidate_filter` defaults
to `None` (existing searches unchanged), and the config-only export is
limited to hybrid providers on nemo:26.08+, so dense / MoE / VLM exports
keep the path they use today.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ✅

### Additional Information

Enables the Prune + Distill workflow for
`NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` (native, no-remote-code
`NemotronHForCausalLM` with MTP heads). Pruning-time MTP support was
scoped and intentionally deferred — MTP heads are dropped and can be
re-derived via a short SFT with `mtp_num_layers=1` on the
pruned+distilled model.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:05:56 +05:30
mxinO 9d360af34f Add ModelOpt QAD skill for Slurm workflows (#2010)
### What does this PR do?

Type of change: new feature

Adds a general Slurm-only QAD skill based on the supported Megatron
Bridge
workflow. The skill:

- starts from a measured BF16-to-PTQ benchmark gap and preserves the
preceding
  PTQ configuration or recipe;
- gates QAD on exact Megatron Bridge model support and successful
Megatron PTQ,
using its master-rank quantizer summary as a scoped `amax` sanity check;
- requires model- and hardware-derived TP/PP/CP/EP/ETP topology
selection;
- streams and randomly samples only the required
`nvidia/Nemotron-Cascade-2-SFT-Data` token budget and uses Megatron
sequence
  packing;
- defaults to 32K sequences, LR `1e-5` with cosine decay, a 1000-step
cap, and
  GBS 512;
- requires explicit user authorization because QAD is costly, validates
two
batches every 25 steps, saves every 50 steps, and monitors a decreasing
  smoothed loss trend;
- evaluates an early checkpoint around step 150 and continues only when
  benchmark recovery and the loss trend justify more training;
- follows the established common Slurm and remote-execution guidance
instead of
  duplicating mutable commands from the Megatron Bridge README.

Also exposes Megatron Bridge `save_interval`, `exit_interval`, and
`exit_duration_in_mins` through `examples/megatron_bridge/distill.py`,
with example-test coverage for checkpoint
and ModelOpt-state preservation at an early exit.

### Usage

```text
Use the QAD skill to recover the measured BF16-to-PTQ benchmark gap for
<model> on <Slurm cluster>, preserving the validated PTQ recipe.
```

### Testing

- `PYTHONPATH=$PWD pre-commit run --all-files`
- Passed every hook on the rebased branch, including Ruff, Ruff format,
mypy,
YAML/recipe validation, launcher reference validation, Bandit, generated
    arguments, symlink synchronization, and Markdown lint.
- `python
~/.codex/skills/.system/skill-creator/scripts/quick_validate.py
.agents/skills/qad`
  - `Skill is valid!`

Qwen3-0.6B result-bearing validation:

- Resources: one exclusive node, 8 H100 GPUs
- Container: `nvcr.io/nvidia/nemo:26.06`
- Quantization: NVFP4, group size 16, embedding excluded
- QAD topology: TP=1, PP=1, CP=4, EP=1, DP=2
- Training validation configuration: sequence length 32768, MBS=1,
GBS=8,
`train_iters=1000`, LR `1e-5` / minimum LR `1e-6`, 50 warmup iterations,
  cosine decay, `eval_interval=150`, `exit_interval=150`,
  `exit_duration_in_mins=220`
- This result-bearing run used the then-current coupled eval/save
cadence. The
final skill now validates two batches every 25 steps and saves every 50;
the
  example test covers the independent checkpoint cadence.
- The reduced GBS 8 is intentionally validation-only; the skill retains
GBS 512
  as the production default.
- Data: exactly 10,000,000 sampled tokens from four
  `nvidia/Nemotron-Cascade-2-SFT-Data` configs:
  - math: 2,306,011 tokens / 364 documents
  - science: 1,191,257 tokens / 285 documents
  - chat: 6,142,077 tokens / 1,800 documents
  - instruction following: 360,655 tokens / 411 documents
- Megatron built packed 32K GPT samples from the materialized prefixes;
the full
  dataset was not downloaded.
- QAD loss was finite and decreased from `0.2640341` at iteration 10 to
`0.1060580` at iteration 150. Final gradient norm was `0.747`, with zero
  skipped and zero NaN iterations. Validation distillation loss was
  `0.09715855`.
- The iteration-150 checkpoint saved successfully with `modelopt_state`,
and
  both PTQ and QAD-150 exported to unified Hugging Face format.
- Identical full MMLU 0-shot comparison through the Megatron evaluator:

  | Model | Accuracy |
  | --- | ---: |
  | BF16 | 0.39517164 |
  | PTQ | 0.32851446 |
  | QAD-150 | 0.38740921 |

QAD-150 recovered `0.05889475 / 0.06665718 = 88.35%` of the measured PTQ
gap,
so validation stopped at the early evidence gate rather than continuing
  blindly toward 1000 iterations.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did
  you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — this adds an agent skill and
example-only
  lifecycle flags.
- Did you get Claude approval on this PR?: N/A

### Additional Information

All seven branch commits are cryptographically signed and include a
`Signed-off-by` trailer.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Updated the QAD skill documentation with a clear “Execute in this
order” workflow, including a revised default recovery training policy.
* Added a new `nemotron-cascade-2` dataset blend configuration with an
increased token budget.
* Enhanced the MeGatron Bridge distillation CLI with stricter interval
argument validation and support for configurable save-and-exit controls.
* **Documentation**
* Expanded Megatron Bridge README guidance for dataset preparation,
token-budget recalculation, and resume expectations.
* **Tests**
* Improved distillation and QAD tests to validate early-exit behavior
and checkpoint expectations.
  * Added unit tests covering distillation CLI interval validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Meng Xin <mxin@nvidia.com>
2026-08-01 19:56:50 +08:00
Daniel Korzekwa d39f385fb8 Save hf checkpoint at every valitation iteration during distillation. (#1897)
### What does this PR do?

Save hf checkpoint at every valitation iteration during distillation.

Addition functionality:
- added `--validate_only` in `examples/megatron_bridge/distill.py` to
enable computing validation losses for iter 0
- added `--reset_optimizer` in `examples/megatron_bridge/distill.py` to
enable not using presaved optimizer, e.g., when changing the number of
train iters.

### Usage

- examples/megatron_bridge/distill.py
- examples/megatron_bridge/README.md (line 228)


### Testing

- tests/examples/megatron_bridge/test_distill.py

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?: ✅ 

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added `--validate_only` for student-only validation at iteration 0.
* Added `--hf_validation_export_path` with
`--hf_validation_export_interval` to export validated student
HuggingFace artifacts during distillation.
* Added `prepare_data_blend.py` for YAML-driven token-budgeted data
blends.
* Added `--max_tokens` to stop Megatron preprocessing after a token
budget.
* **Bug Fixes**
* Validation exports avoid duplicate checkpoints, preserve the student
architecture/config, and export only student artifacts.
* **Documentation**
* Expanded researcher and tutorial guides for iterative workflows and
token-budgeted blends.
* **Tests**
* Updated distillation/blend/max_tokens test coverage, including
validate-only and interval-based exports.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
2026-07-23 19:48:42 +02:00
Keval MorabiaandClaude Opus 4.8 21d0069e4d MBridge VLM distillation / QAD support (#1938)
## What

Adds VLM (e.g. Qwen3.5-VL, Gemma3-VL) knowledge-distillation / QAD
support to the Megatron-Bridge examples:

- `distill.py` distills only the **language model** submodule (vision
tower + projector untouched), reusing the LLM training path.
- New `export_distilled_megatron_to_hf.py` converts a distilled Megatron
checkpoint (**any** iteration) to HF. Required especially for VLM
distilled ckpt as it only has LM weights so we need to initialize full
VLM, swap LLM weights then save to HF
- Renames `export.py` → `export_quantized_megatron_to_hf.py`.

## Related upstream Megatron-Bridge PRs to be available in nemo:26.08
container:

- NVIDIA-NeMo/Megatron-Bridge#4707 — `DistillationProvider` submodule
distillation (non-blocking; added temporary WAR)
- NVIDIA-NeMo/Megatron-Bridge#4706 — MoE expert weight-mapping fix
(Qwen3.5-VL-MoE with moe_grouped_gemm=False). Also removed ModelOpt side
WAR previously added as it was not accurate; better to wait till next
container release or mount latest MBridge into the 26.06 container.

## Testing

- Qwen3.6-35B-A3B Pruning + Distillation with MMLU evaluation sanity
check (results below in comments)
- Cosmos 2 Reason 2B valiadted by SAs (results below in comments)
- Validated end-to-end on `nemo:26.06` (distill → separate HF export;
LLM + VLM, incl. TP→TP/PP reshard). `test_distill_vlm` runs the export
script as a CI e2e step.
- Many CICD tests for wide coverage of all mbridge scripts

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **New Features**
* Added a dedicated HuggingFace exporter for distilled Megatron
checkpoints, with distinct LLM vs VLM conversion flows.

* **Bug Fixes**
* Improved Megatron-Bridge distillation/export consistency, including
safer handling of VLMs and targeted submodule distillation.

* **Documentation**
* Updated Megatron-Bridge READMEs and tutorials to reference the new
quantized and distilled export scripts and revised CLI guidance.

* **Tests**
* Expanded distillation, QAD, and quantization/export tests to cover
LLM/VLM variants, with conditional skipping for unsupported MoE setups.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 15:16:30 +00:00
Keval MorabiaandClaude Opus 4.8 2fc352be2d Add VLM pruning and PTQ with image-text calibration (Megatron-Bridge) (#1792)
### What does this PR do?

Type of change: New feature

Adds **vision-language model (VLM) support** to the Megatron-Bridge
examples for both **Minitron pruning** (`prune_minitron.py`) and **PTQ**
(`quantize.py`). Only the **language model** is pruned/quantized — the
vision tower and vision→language projector are left in full precision —
and the full VLM is saved back. `hidden_size` is skipped for pruning
when it is shared with the vision→LM projector.

Supported VLMs (tested e2e): **Qwen{3,3.5}-VL** (dense; hybrid
GatedDeltaNet + gated attention) and **Gemma3-VL** (sliding/full
attention).

### Calibration (image-text)

Calibration is conditioned on real **image-text** data so the language
model's pruning importance / quantizer statistics see vision-conditioned
activations. The modality is inferred from `--calib_dataset_name`:

- an **image-text** dataset (default for VLMs,
`nemotron_vlm_dataset_v2`) drives the **full VLM forward**;
- a **text** dataset runs text-only calibration of the language model
(for text-vs-image ablations).

A shared `get_megatron_vlm_calibration_forward_loop` (built on
`megatron_prefill`) drives the full VLM forward over image-text pairs
from `vlm_dataset_utils` (`scienceqa`, `nemotron_vlm_dataset_v2`, with
config-driven subset/shard caps to bound downloads). It shards across
**data-parallel (DP)** ranks like the text loop (#1804); **context
parallelism (CP)** applies to text-only VLM calibration (the shared text
loop), not the multimodal forward — splitting the sequence would
misalign the merged vision embeddings.

### Results - Cosmos-Reason2-2B

Validated end-to-end on **Cosmos-Reason2-2B** (Qwen3-VL). Minitron NAS
prunes the language-model tower **1.72B → ~1.59B** (vision encoder +
projector frozen), top_k=1. Calibration data drives pruning importance;
image-text calibration runs the full VLM forward.

| Model | Calibration | MMLU | BLINK Rel-Depth | RealWorldQA |
|---|---|---|---|---|
| Baseline (1.72B) | — | 0.58 | 0.76 | 0.61 |
| Pruned (1.59B) | text (`nemotron-post-training-dataset-v2`) | 0.51\* |
~0.69 | ~0.57 |
| Pruned (1.59B) | image+text (`nemotron_vlm_dataset_v2`) | 0.49\* |
**0.77** | **0.61** |

\* Pruned MMLU on the 10% split (the pruning score function); baseline
MMLU is the full set. The VLM-benchmark numbers for the text row were
measured with a different text calibration set and are expected to be
similar for `nemotron-post-training-dataset-v2` (marked `~`).

> [!NOTE]
> These numbers come from short single runs on small eval splits — read
them for **high-level trends only**, not as exact values.

Takeaways: pruning the LM tower of a VLM works end-to-end. **Image-text
calibration** (this PR's feature) preserves the VLM benchmarks better
than text-only — BLINK Rel-Depth ~0.77 vs ~0.69 and RealWorldQA ~0.61 vs
~0.57, both close to the unpruned baseline (0.76 / 0.61) — which is the
motivation for calibrating on vision-conditioned activations.

### Results - Qwen3.5-9B

| Model                      | MMLU   | MMStar |
|----------------------------|:------:|:------:|
| Qwen3.5-9B      | 0.7003 | 0.6117 |
| Pruned-7B (text calib)        | 0.5527 | 0.4411 |
| Pruned-7B (image+text calib)  | 0.5107 | 0.3941 |

### Key changes

- `quantize.py`: quantizes the **root** model with non-LM (vision)
quantizers disabled, so the ModelOpt state lives on the root (required
by the Megatron save) while only the language model is quantized.
- `prune_minitron.py`: image-text (or text) calibration for VLM pruning
importance.
- Shared VLM calibration forward loop (`megatron_prefill`-based, unwraps
tuple outputs, DP-sharded) + `vlm_dataset_utils`.
- Tiny VLM test fixtures (Qwen3.5-VL, Gemma3-VL) with vision tokens
derived dynamically from the reference processor; VLM prune + quantize
example tests.
- README + CHANGELOG.

### Usage

```bash
# Prune the language model of a VLM (image-text calibration by default)
torchrun --nproc_per_node 2 prune_minitron.py \
    --pp_size 2 \
    --hf_model_name_or_path <vlm> \
    --prune_target_params 3e9 \
    --output_hf_path /tmp/vlm-pruned

# PTQ the language model of a VLM
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path <vlm> \
    --quant_cfg fp8 \
    --export_megatron_path /tmp/vlm-fp8-megatron
```

### Testing

- `test_prune_minitron.py::test_prune_minitron_vlm` — Gemma3-VL,
image-text (ScienceQA) calibration; full load → prune (depth + ffn) →
save → reload.
- `test_quantize_export.py::test_quantize_vlm` — Qwen3.5-VL, text
calibration; quantize LM → save Megatron checkpoint.
- LM regression tests (`test_prune_minitron`,
`test_quantize_and_export`) unchanged and passing.

### Not in scope

- **HF unified export of a quantized VLM** is not yet supported;
`export.py` saves the Megatron checkpoint only for VLMs (tracked by a
TODO in `export.py`). The recommended path is to route the megatron→HF
quant export through Megatron-Bridge's
`AutoBridge.export_hf_weights_quant(quantization_checker, quant_fn,
quant_block_size)`, which reuses the bridge's per-model mcore↔HF mapping
— covering Qwen3.5-VL / Gemma3-VL and the vision tower/projector (left
full precision) for free — so modelopt supplies only the checker +
pack/scale fn + `hf_quant_config` (KV-cache scales need a separate
path). This avoids re-authoring per-model mappings in modelopt (cf.
#1482's Qwen3-VL-only `mcore_qwen3vl.py`).

> [!NOTE]
> Qwen3.5-VL **MoE** is not tested e2e: the Megatron-Bridge weight
conversion expects packed (`gate_up_proj`) experts that transformers'
tiny checkpoint doesn't emit. MoE pruning itself is covered by
`test_mcore_qwen35_gdn_moe_pruning`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Follow-up to the GatedDeltaNet/MLA/latent-MoE pruning PR (#1747).
Rebased on `main` to pick up CP/DP calibration (#1804); the VLM
calibration loop now shards across DP ranks the same way. `hidden_size`
pruning for VLMs (requires resizing the vision projector) is left for a
future PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added VLM-aware Minitron pruning and post-training quantization that
target only the language-model portion, keeping the vision
tower/projector in full precision.
* Calibration now auto-selects text vs image-text datasets based on
model type, with modality validation.
* Expanded Megatron-Core CP/DP guidance and introduced a `--cp_size`
flag in quantization examples.
* **Bug Fixes**
* Improved VLM generation/prefill output handling and made vocabulary
sizing more robust for VLM wrappers.
* **Tests / Documentation**
* Updated pruning/quantization docs and refreshed/added VLM-focused
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 02:47:02 +05:30
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Keval MorabiaandClaude Opus 4.8 aa2a6a1b5d Add context-parallel (CP) and data-parallel (DP) support to Megatron calibration, and MMLU (#1804)
### What does this PR do?

Type of change: new feature

Adds **context-parallel (CP)** and **data-parallel (DP)** support to the
shared Megatron-Core inference/calibration utilities so PTQ calibration
and the MMLU sanity check work across these parallelisms (in addition to
the existing TP/PP/SP/EP).

**Context parallelism (CP):**
- **`megatron_calibration` / `megatron_mmlu`** — partition each sequence
across CP ranks (zigzag load-balanced, via `get_batch_on_this_cp_rank`).
MMLU gathers the per-rank logits back to the full sequence for
last-token scoring.
- **`megatron_prefill`** — accepts a CP-partitioned `position_ids`, and
under CP passes `attention_mask=None` so the CP-aware causal attention
builds the mask itself (a local triu mask would be wrong for the
per-rank zigzag chunks). Also wrapped in `torch.no_grad()` (pure
inference; lets MMLU run at larger batch sizes without retaining the
autograd graph).
- **`examples/megatron_bridge/quantize.py`** — new `--cp_size` flag.

**Data parallelism (DP):**
- **`get_dataset_dataloader`** — new `distributed` / `sampler_kwargs` to
shard the dataset across ranks with a `DistributedSampler`.
- **`megatron_calibration`** — shards calibration data across the DP
group (amax is max-reduced across DP inside `mtq` calibration, so the
per-rank shards combine correctly).
- **`megatron_mmlu`** — shards whole batches across DP ranks and
all-reduces the per-subject counts back to full-dataset accuracy.
- DP is **implicit**: `DP size = world_size / (tp * pp * cp)` —
launching with more GPUs than `tp * pp * cp` engages it. (No `--dp_size`
flag.)

RoPE is applied by the model per CP rank, so `position_ids` are only
needed for models with absolute/learned position embeddings.

### Usage

```bash
# Context parallelism = 2
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
    --cp_size 2 --export_megatron_path /tmp/Qwen3-8B-NVFP4-cp2

# Data parallelism = 2 (implicit: tp*pp*cp = 1, 2 GPUs)
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-dp2
```

### Testing

Added `cp` and `dp` cases to
`tests/gpu_megatron/torch/utils/plugins/test_utils_megatron.py::test_megatron_generate_and_mmlu`
(Qwen3-0.6B). All four parallelisms pass on 2 GPUs over the full
1430-example MMLU shard set (confirming the DP all-reduce reconstructs
the full count): tp=0.373, pp=0.375, cp=0.371, dp=0.375.

End-to-end PTQ on **Qwen3-8B → NVFP4**, comparing MMLU before and after
PTQ across parallelisms:

| Parallelism | MMLU before PTQ (bf16) | MMLU after PTQ (NVFP4) |
| :--- | :---: | :---: |
| TP=2 | 0.7294 | 0.7058 |
| CP=2 | 0.7292 | 0.7101 |
| DP=2 | 0.7292 | 0.7099 |

> **Common PTQ args used for the runs above:** `--hf_model_name_or_path
Qwen/Qwen3-8B --quant_cfg nvfp4 --seq_length 1024 --calib_num_samples
512 --calib_batch_size 16`, default calibration dataset
(`cnn_nemotron_v2_mix` = cnn_dailymail +
nemotron-post-training-dataset-v2 mix). MMLU evaluated at `fraction=1.0,
batch_size=16` on the full test set. Runs on 2× RTX 6000 Ada — TP=2 uses
`--tp_size 2`, CP=2 uses `--cp_size 2`, DP=2 uses all model-parallel
sizes = 1 (implicit DP over the 2 GPUs).

CP-, DP-, and TP-calibrated models all land within ~0.4% MMLU of each
other both before and after PTQ, confirming CP/DP calibration yields an
equivalently-quantized model.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all CP/DP logic is gated on
`cp_size > 1` / `dp_size > 1`; non-CP/DP (TP/PP/SP/EP) behavior is
unchanged. `megatron_prefill` gains an optional `position_ids` arg
(defaults to the previous behavior); `get_dataset_dataloader` gains
optional `distributed`/`sampler_kwargs` (default off).
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: ✅ — `cp` and `dp`
parametrizations added to the existing Megatron generate/MMLU test.
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅ (run `/claude review`)

### Additional Information

`torch.no_grad()` on `megatron_prefill` is a shared change (also
benefits the calibration / generate / PEFT-test callers) — pure
inference, so no behavioral change beyond lower memory.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes
* **New Features**
* Added context-parallel (CP) and data-parallel (DP) support across
shared inference, calibration, and MMLU evaluation.
* Introduced CP-aware prefill with per-rank logits gathering for correct
last-token scoring.
* Added `--cp_size` to the quantization example (DP is derived
automatically).
* **Improvements**
* Extended dataset dataloader utilities with optional distributed
sampling controls.
* **Bug Fixes**
* Calibration and evaluation now work when CP is enabled (no longer
restricted to CP=1).
* **Tests**
  * Expanded Megatron generate/MMLU coverage to include CP and DP modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 22:39:22 +05:30
Keval Morabia b6bf6b7997 Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-23 21:42:37 +05:30
Keval MorabiaandClaude Opus 4.8 55a2101e2f Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do?

Type of change: documentation + minor example-script tweaks

Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR
was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16
tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md)
results using the **new shared calibration loop (sequence packing)** and
to **fix tool calling in evaluation**.

- Refreshed the prune → distill → eval → **FP8** results (accuracy +
vLLM throughput tables) with the new calibration loop.
- **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now
run the Python sandbox tool. The tutorial reports both **with-tools**
and **no-tools** GPQA/AIME and shows `mean ± std_dev`.
- Script tweaks: `quantize.py` calibration now uses sequence packing
(`pack=True`) which leads to slight improvement in PTQ;
`prune_minitron.py` defaults `inference_batch_size` to
`calib_batch_size`.

### Testing

Documentation + small example-script changes; tutorial relative links
resolve and the results tables / figure were verified consistent.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the main guide and evaluator instructions for prune + distill
+ FP8/NVFP4 quantization, including refreshed vLLM deployment tips,
benchmark/noise presentation, and long-context tool-calling attribution
notes.
* Refreshed README technique examples/links, reordered the model support
matrix rows, and improved pruning overview/support-matrix text.

* **Changes to Examples**
* NAS pruning now documents higher GPU memory usage vs manual pruning;
pruning batching defaults were improved.
* Quantization PTQ calibration uses packed document packing; quantized
checkpoint export messaging was streamlined.
* Updated pruning/distillation/quantization tutorial guidance,
metrics/tables, command parameters, and evaluator YAML settings
(KV-cache dtype, generation defaults, task behavior).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 18:32:13 +00:00
Keval MorabiaandClaude Opus 4.8 5584ce4558 Migrate Nemotron-3-Nano tutorial PTQ to MBridge scripts and move under examples/megatron_bridge (#1601)
### What does this PR do?

Type of change: documentation (+ minor test fixes)

Migrates the Nemotron-3-Nano-30B-A3B-BF16 tutorial quantization step
from `examples/llm_ptq/hf_ptq.py` to the Megatron-Bridge quantize +
export, and relocates the tutorial next to the scripts it now uses. Now
that the whole tutorial is Megatron-Bridge based, it lives under
`examples/megatron_bridge/`.

- **Quantization migration:** replace the single `hf_ptq.py` call with
`examples/megatron_bridge/quantize.py` (calibrate + save a Megatron
checkpoint) → `examples/megatron_bridge/export.py` (deployable unified
HF checkpoint). The FP8 results table is refreshed with the
`quantize.py` numbers (same defaults, slightly better on average).
- **Relocation:** moved
`examples/pruning/minitron/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/` →
`examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/`.
A **redirect-stub `README.md`** remains at the old path (a directory
symlink isn't traversable in the GitHub web UI), and all in-repo
references (root README, CHANGELOG, pruning READMEs, megatron_bridge
README) plus the tutorial's own relative links are updated.
- **Evaluation:** per-format vLLM benchmark commands (BF16 / FP8), FP8
deployment notes documented in `nemo_evaluator.yaml`, reduced
LiveCodeBench/AIME `num_repeats` (were too slow), and bumped the
`nemo-evaluator-launcher` pin.
- **Misc:** drop the `examples/megatron_bridge/requirements.txt`
`transformers<5` pin in favor of an inline "downgrade `transformers<5`
to save pruned Nemotron checkpoints" note; guard the hybrid Mamba-MoE
sharded-state-dict test behind `HAS_MAMBA` (requires `mamba_ssm`);
shrink the tiny Gemma3 test fixture's attention heads.

> **Note:** the **NVFP4 + QAD** experiments (formerly the focus of this
PR) are split out — their accuracy/throughput results are still in
progress — and will follow in a separate PR on top of this one.

### Testing

Docs-only + test-guard changes. Pre-commit hooks (markdownlint, RST
checks, ruff, mypy) pass. The tutorial's relative links and the old-path
redirect stub were verified to resolve to real files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (old tutorial path still
resolves via a redirect-stub README; `quantize.py`/`export.py` already
exist in `examples/megatron_bridge`)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (adjusts/guards existing
tests only)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (existing tutorial entry updated to the new path)
- Did you get Claude approval on this PR?: ✅

### Additional Information

Supersedes the previous "Part 3 of 4 (NVFP4 + QAD docs)" scope of this
PR; the NVFP4 + QAD tutorial additions will land in a follow-up.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Moved the Nemotron-3-Nano-30B-A3B tutorial into the Megatron-Bridge
tutorials and replaced the old file with a pointer to the new location.
* Updated vLLM throughput numbers to 2.6× and expanded
results/throughput tables.
* Reworked the FP8 quantization/export workflow and added a note to use
transformers<5 when saving pruned models.
* Added a tutorials index and adjusted evaluator launcher pin and repeat
counts.

* **Tests**
* Tests now detect optional Mamba support and skip related tests when
unavailable.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 18:55:42 +00:00
Keval MorabiaandClaude Opus 4.8 54ce4e09d8 Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do?

Type of change: new example

**Note:** This is **part 2 of 4** (builds on #1589):

- **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
- **Part 2 (this PR):** extend `distill.py` for quantization-aware
distillation (QAD) — load a quantized Megatron checkpoint as the
student.
- **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601
- **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

Extends `examples/megatron_bridge/distill.py` to initialize the student
from a **Megatron checkpoint** (a quantized checkpoint from
`quantize.py`, or a pruned one) via `--student_megatron_path`, enabling
**Quantization Aware Distillation (QAD)**:

- `--student_hf_path` still builds the student architecture;
`--student_megatron_path` supplies the (optionally quantized) weights.
- For a quantized checkpoint, the ModelOpt quantize mode + base weights
are restored onto the **plain student before the knowledge-distillation
conversion** (`restore_sharded_modelopt_state` is a no-op once a model
is already converted), so the distilled checkpoint stays exportable as a
quantized model with `export.py`.

**Upstream dependency / workaround:** `DistillationProvider.provide()`
has no seam to transform the student before the KD conversion, so this
patches `provide()` at the class level (via an `id()`-keyed registry,
because the provider proxies instance-attribute assignment to its
teacher once the teacher is set). A companion Megatron-Bridge PR adds a
first-class `DistillationProvider.student_pre_conversion_hook`; from
nemo:26.06 onwards the workaround should be removed and replaced with
that hook (a removal note in `distill.py` documents exactly how).

### Usage

```bash
# 1) PTQ -> quantized Megatron checkpoint (part 1)
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# 2) QAD: distill the quantized student from the unquantized teacher
torchrun --nproc_per_node 8 distill.py \
    --teacher_hf_path Qwen/Qwen3-8B \
    --student_hf_path Qwen/Qwen3-8B \
    --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --data_paths 1.0 tokenized/data_text_document \
    --train_iters 1000 --output_dir /output/qwen3_8b_qad

# 3) export the distilled quantized checkpoint (part 1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /output/qwen3_8b_qad/checkpoints \
    --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf
```

### Testing

`tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo
`26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the
quantized student → `export.py` to a unified HF checkpoint, asserting
`hf_quant_config.json` is written (proves the quantize mode survived
QAD). Includes a commented-out vLLM deployment check, validated locally
(full flow passes; vLLM loads the export as `quantization=modelopt`).
Existing normal/Puzzletron distillation tests still pass.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example feature; default
behavior unchanged when `--student_megatron_path` is not set)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Depends on a companion Megatron-Bridge PR adding
`DistillationProvider.student_pre_conversion_hook` (the upstream
replacement for the class-level `provide()` workaround). The Nemotron-3
tutorial NVFP4 + QAD experiments ship in part 3.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization Aware Distillation (QAD) workflow to recover accuracy of
quantized Megatron students and distill from quantized checkpoints.
* CLI option to initialize a distillation student from a Megatron
checkpoint and a structure-only load path for bridging.

* **Documentation**
* Expanded runnable quantize → QAD → export guidance and best-practice
tips.

* **Tests**
  * End-to-end test validating quantize → QAD → export artifacts.

* **Chores / UX**
* Clearer rank-aware messages, improved tokenizer padding handling, and
more consistent export behavior (fixed export dtype).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 21:28:25 +00:00
Keval MorabiaandClaude Opus 4.8 f21977a5fc Add Megatron-Bridge PTQ quantize + export example scripts (#1589)
### What does this PR do?

Type of change: new example

Adds a two-step post-training quantization (PTQ) flow for
**Megatron-Bridge** models under `examples/megatron_bridge/`, mirroring
the Megatron-LM `quantize.sh` / `export.sh` split:

- **`quantize.py`** — loads an HF model via Megatron-Bridge, applies
ModelOpt PTQ (via a `--quant_cfg` alias / full config name, or a
`--recipe` YAML), with optional KV-cache quant, weight-only,
compression, and MoE expert-ratio calibration, then saves a **Megatron
checkpoint** (with ModelOpt state). Tensor / pipeline / expert
parallelism are all supported, and the checkpoint can later be reloaded
for further training (QAT / distillation).
- **`export.py`** — loads the quantized Megatron checkpoint, **re-shards
to TP=1**, and exports a **HuggingFace (unified)** checkpoint deployable
with TensorRT-LLM / vLLM / SGLang.

**Why the split?** The unified HF exporter (`export_mcore_gpt_to_hf`)
does not gather tensor-parallel-sharded weights — Megatron-LM likewise
forces `TP=1` during its export step. Saving a TP-sharded Megatron
checkpoint first lets us calibrate at TP>1 (to fit large models) and
then reload re-sharded to TP=1 for the HF export. A combined
single-script flow silently produced corrupt HF checkpoints under TP>1
(collided per-rank shards), which this split avoids.

> **Note:** This is **part 1 of 4**:
> - **Part 1 (this PR):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
> - **Part 2:** extend `distill.py` for quantization-aware distillation
(QAD) — load a quantized Megatron checkpoint as the student.
> - **Part 3:** add NVFP4 + QAD-on-pruned-checkpoint experiments to the
Nemotron-3-Nano-30B-A3B tutorial.
> - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

### Usage

```bash
# Step 1: quantize (TP/PP/EP supported) -> Megatron checkpoint
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --quant_cfg fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# Step 2: export -> deployable HuggingFace (unified) checkpoint (re-shards to TP=1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --export_unified_hf_path /tmp/Qwen3-8B-FP8-hf
```

### Testing

`tests/examples/megatron_bridge/test_quantize.py` (validated on a 2-GPU
NeMo `26.04` container):

- `test_quantize_export_and_vllm_deployment` — quantize a tiny Qwen3 via
a recipe at TP=2 → `export.py` re-shards to TP=1 → load + generate with
**vLLM** (skipped if vLLM absent).
- `test_quantize_megatron_checkpoint_reload` — quantize at TP=2 → reload
the Megatron checkpoint via the bridge and assert ModelOpt quantizers
were restored.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

The Nemotron-3 tutorial update to use these scripts is intentionally
**not** included here — it ships with the part 3 PR alongside the NVFP4
+ QAD experiments.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded post-training quantization (PTQ) workflow documentation with
detailed step-by-step examples and configuration guidance for the
Megatron-Bridge framework.

* **New Features**
* Added quantization tool for applying PTQ to Megatron models with
calibration support.
* Added export tool for converting quantized models to a deployable
format.

* **Tests**
* Added integration tests validating the complete
quantization-export-deployment workflow, including inference validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 19:36:15 +00:00
Keval MorabiaandClaude Sonnet 4.6 999c99913e Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill + Quantize + Nemo Evaluator + vLLM deployment (#1376)
### What does this PR do?

Type of change: example/tutorial <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

Add Nemotron-3-Nano-30B-A3B-BF16 e2e tutorial: Prune + Distill +
Quantize + Nemo Evaluator + vLLM deployment

<img width="2079" height="1613" alt="image"
src="https://github.com/user-attachments/assets/19b6ab82-7f01-45df-a0a5-d1c3282b384a"
/>



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* End-to-end Nemotron-3-Nano-30B tutorial: pruning, two‑phase
distillation, FP8 PTQ, evaluation, and vLLM deployment; new ablations
and long‑context analyses.
* Distillation CLI: configurable seed and activation‑recomputation
options.

* **Bug Fixes**
* Preprocessing hardened to skip malformed JSONL and normalize tool‑call
argument formats.

* **Documentation**
* Many README/examples/evaluator docs updated (news list, tokenization
guides, tutorials, configs, and deployment notes).

* **Tests**
* Added test verifying preprocessing handles stringified tool‑call
arguments.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1376?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-28 18:35:59 +00:00
Keval MorabiaandClaude Opus 4.7 5d0441ae3d Create shared Megatron calibration forward loop for prune / quantize with megatron pretraining data style sequence packing (#1501)
## Summary

Replaces the bespoke calibration loops in Megatron-LM and
Megatron-Bridge prune / quantize example scripts with a single shared
utility,
`modelopt.torch.utils.plugins.megatron_calibration.get_megatron_calibration_forward_loop`.

The shared loop iterates a packed calibration dataloader built via
`get_dataset_dataloader(pack=True)` and drives a logits-free prefill
pass through the model so activation hooks fire on every layer.
`pack=True` produces Megatron-LM pretraining-style **global-stream
packing**: all raw samples are concatenated into one EOS-separated token
stream and sliced into uniform-length rows. The trained model has seen
this distribution extensively during pretraining, so the activations
produced during calibration are representative of the model's natural
behavior.

Migrates four call sites:
- `examples/megatron_bridge/prune_minitron.py`
- `Megatron-LM/examples/post_training/modelopt/{prune,quantize}.py`
(separate PR:
[NVIDIA/Megatron-LM#4881](https://github.com/NVIDIA/Megatron-LM/pull/4881))
- `Megatron-Bridge/examples/quantization/quantize.py` (separate PR)

Each call site passes `pack=True` explicitly with an inline comment so
users see the option and know when to flip it. The function-level
defaults (`get_dataset_dataloader(pack=False)`,
`get_megatron_calibration_forward_loop(pack=False)`) remain
back-compat-safe.

Unified defaults across all four sites: `--calib-dataset
nemotron-post-training-dataset-v2`, `--calib-size 1024`,
`--calib-max-sequence-length 4096`, `--calib-batch-size 1`.

## Experimental results

Qwen3-8B on full 100% MMLU (n=14042; binomial 2σ noise floor ≈ ±0.78 pt
at acc ≈ 0.7), 0-shot, eval batch_size=4. Calibration on the default
workload: nemotron-post-training-dataset-v2, seq_length=4096,
calib_batch_size=8.

**Three calibration data shapes compared:**

- **Padded**: one doc per row, padded to `seq_length`, pad tokens flow
through the forward (legacy `get_calib_dataloader` pad+truncate
behavior).
- **Trimmed**: one doc per row, each row trimmed to its real content
length via `attention_mask`, with EOS forced at the last real position;
pad never enters the forward. This is no longer part of this PR.
- **Packed**: global-stream slicing — all docs concatenated
EOS-separated into one token stream, sliced into uniform `seq_length`
rows. Matches Megatron's `.bin`/`.idx` pretraining distribution. Enabled
via `pack=True` in `get_megatron_calibration_forward_loop`.

| Workload | Padded | Trimmed | **Packed** |
|---|---|---|---|
| M-LM NVFP4 quantize (`NVFP4_DEFAULT_CFG`) | 0.707 | 0.708 | 0.709 |
| M-Bridge Minitron prune (Qwen3-8B → 30L / 3584 / 11776 ≈ 6B params) |
0.576 | 0.573 | **0.589** |

### Key findings

- **M-LM quantize quality is calibration-mode-insensitive** for dense
Qwen3-8B NVFP4 — all three modes are nearly identical.
- **M-Bridge prune**: More sensitive to calibration data shape. Packed
wins over Padded on full MMLU.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-28 18:34:37 +00:00
Keval Morabia f0eaa198df Enable active-param and memory based Minitron pruning constraint (#1377)
### What does this PR do?

Type of change: New feature, new tests, documentation.

OMNIML-4108: Extends the Minitron NAS pruner to support pruning by
**active parameter count** (`active_params`) and **memory footprint**
(`memory_mb`) in addition to the existing total parameter count
(`params`) constraint. Also adds standalone utilities for analytical
model stats.

#### Changes

**New pruning constraint keys**
- `active_params`: prune to a target number of active (routed) params —
useful for MoE models where total ≫ active; when present,
`active_params` is the **primary sort/display metric** for candidates
(priority: `active_params` > `params` > `memory_mb`)
- `memory_mb`: prune to fit a memory budget (BF16 weights + KV-cache +
Mamba state at a given sequence length and batch size)
- Constraints can be combined (AND logic): e.g. `{"params": 6e9,
"memory_mb": 12288}`

**New standalone utilities**
(`modelopt.torch.nas.plugins.megatron_model_stats`)
- `mcore_param_count`: analytically computes total and active parameter
counts for GPT and Mamba/hybrid MCore models
- `mcore_memory_footprint_mb`: estimates memory in MB (weights +
KV-cache + Mamba state)
- `print_mcore_model_stats`: rich-formatted model stats panel

**Rich-formatted pruning logs** — search space, top-k candidate tables,
and best subnet panel printed on rank 0

**`prune_score_func` format update** — now `mmlu_<N>pct_bs<bs>` (e.g.
`mmlu_10pct_bs32`) to explicitly control batch size for MMLU evaluation;
old `mmlu_<N>pct` format removed

**Infrastructure**
- NeMo container bumped to `nvcr.io/nvidia/nemo:26.04` in CI and docs
- Added `examples/megatron_bridge/requirements.txt` with
`transformers<5.0` (required for saving some Nemotron-3-Nano models)

### Usage

```python
# Prune to 3B active params (MoE-aware) — active_params is the primary sort metric
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"active_params": 3e9}, config=pruning_config)

# Prune to fit a 12 GB memory budget
mtp.prune(model, mode=[("mcore_minitron", ss_config)], constraints={"memory_mb": 12288}, config=pruning_config)
```

### Testing

Pruned Nemotron-3-Nano-30B-A3B (31.6B, A3.6B) --> A3.0B. Takes <1hr on
8x H100 (more details in #1376)

```bash
torchrun --nproc_per_node 8 examples/megatron_bridge/prune_minitron.py \
    --pp_size 8 \
    --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 \
    --trust_remote_code \
    --prune_target_params 28e9 \
    --prune_target_active_params 3e9 \
    --hparams_to_skip num_attention_heads \
    --seq_length 8192 \
    --output_hf_path pruned/Nemotron-3-Nano-30B-A3B-Pruned-28B-A3B-top20-max15depth-max30width-mmlu_10pct_bs32 \
    --top_k 20 \
    --max_depth_pruning 0.15 \
    --max_width_pruning 0.30 \
    --prune_score_func mmlu_10pct_bs32 \
    --num_layers_in_first_pipeline_stage 5 \
    --num_layers_in_last_pipeline_stage 5
```

```
╭──────────────────────────────────────────────────── Original Model Stats ─────────────────────────────────────────────────────╮
│ Total Parameters                              31.58B                                                                          │
│ Active Parameters                             3.58B                                                                           │
│ Memory (BF16, seq_length=8192, batch_size=1)  weights: 60230.1 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 60301.9 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

                                                                 Top 20 Candidates with Scores
┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃  # ┃ export_config                                                                                                         ┃ active_params ┃ params ┃  score ┃
┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│  1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 120,          │         3.00B │ 27.06B │ 0.3399 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  2 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.4650 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  3 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2343 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  4 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'mamba_head_dim': 48, 'num_moe_experts': 96,           │         3.00B │ 20.09B │ 0.2552 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  5 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 21.61B │ 0.2601 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 64, 'num_moe_experts': 96,           │         3.00B │ 19.28B │ 0.3762 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│  7 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,          │         3.00B │ 22.28B │ 0.4783 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│  8 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 96,           │         3.00B │ 21.99B │ 0.2420 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│  9 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2399 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 10 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 112,          │         3.00B │ 26.17B │ 0.2601 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│ 11 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 112,          │         3.00B │ 25.37B │ 0.2503 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 12 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.4329 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 13 │ {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 128,          │         3.00B │ 26.17B │ 0.2587 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 2816}                                             │               │        │        │
│ 14 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 64, 'mamba_head_dim': 56, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2336 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 15 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'mamba_head_dim': 56, 'num_moe_experts': 96,           │         3.00B │ 20.09B │ 0.2559 │
│    │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 16 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 96,           │         3.00B │ 20.70B │ 0.4608 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
│ 17 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2455 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 18 │ {'num_layers': 50, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 104,          │         3.00B │ 24.42B │ 0.2503 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328}                                             │               │        │        │
│ 19 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 48, 'mamba_head_dim': 48, 'num_moe_experts': 120,          │         3.00B │ 27.92B │ 0.2587 │
│    │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712}                                             │               │        │        │
│ 20 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56, 'mamba_head_dim': 64, 'num_moe_experts': 104,          │         3.00B │ 23.68B │ 0.2469 │
│    │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3072}                                             │               │        │        │
└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘

╭──────────────────────────────────────────────────────────────────────── Best Subnet ─────────────────────────────────────────────────────────────────────────╮
│ export_config  {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1856,     │
│                'moe_shared_expert_intermediate_size': 3072}                                                                                                  │
│ active_params  3.00B                                                                                                                                         │
│ params         22.28B                                                                                                                                        │
│ score          0.4783                                                                                                                                        │
╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭───────────────────────────────────────────────────── Pruned Model Stats ──────────────────────────────────────────────────────╮
│ Total Parameters                              22.28B                                                                          │
│ Active Parameters                             3.00B                                                                           │
│ Memory (BF16, seq_length=8192, batch_size=1)  weights: 42489.7 MB, kv_cache: 48.0 MB, mamba_state: 23.8 MB, Total: 42561.6 MB │
╰───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
```

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-05 09:58:47 +05:30
Keval MorabiaandClaude Sonnet 4.6 bb08094ff1 Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary

End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how
ModelOpt techniques stack:

- **Pruning** — Minitron structured pruning 9B → 7B
- **Distillation** — Megatron-Bridge knowledge distillation up to 80B
tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math
500, IFEval, SciCode
- **Evaluation** - using nemo-evaluator
- **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on
vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from
\`config.json\`)
- **vLLM Throughput** — BF16 vs FP8 benchmark on single H100

<img width="2085" height="1740" alt="image"
src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87"
/>

<img width="2085" height="810" alt="image"
src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843"
/>


### Files changed

- `examples/pruning/minitron/README.md` — index page for Minitron
end-to-end tutorials
- `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` —
full repro doc with 6 sections: data prep, pruning, distillation,
evaluation, FP8 quantization, vLLM benchmarking
-
`examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml`
— NeMo Evaluator config used for all benchmark numbers
- `examples/pruning/puzzletron/README.md` — index page for Puzzletron
distillation results
- `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron
distillation results (renamed from puzzletron.md)
- `examples/pruning/README.md` — updated Results section with direct
links to new locations
- `examples/megatron_bridge/README.md` — updated results link to point
to `examples/pruning/`
- `examples/puzzletron/README.md` — updated distillation results link
- `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for
all datasets used in the data blend

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Documentation

* **New end-to-end tutorial** for model optimization covering Minitron
pruning, knowledge distillation, FP8 quantization, and vLLM deployment
with reproducibility steps and benchmark results
* **Dataset preparation guide** with ready-to-run tokenization templates
for Nemotron HuggingFace datasets
* **Evaluation configuration** and results documentation including
ablation studies across multiple benchmarks
* **Updated navigation** across pruning, distillation, and dataset
examples to streamline user workflows

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 14:47:16 +05:30
361f7e391b Merge puzzletron compression algorithm (#1121)
### What does this PR do?

Implement puzzletron compression algorithm based on Puzzle paper
(https://arxiv.org/abs/2411.19146)

<details>
<summary> Th list of reviewed and merged MRs that resulted in the
feature/puzzletron branch</summary>

Merging dkorzekwa/any_model to feature/puzzletron

[Add anymodel directories to feature/puzzletron by danielkorzekwa · Pull
Request #974 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/974)
- merged

[Draft: anymodel activation scoring by danielkorzekwa · Pull Request
#989 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/989)
- merged

[Draft: Merge anymodel pruning by danielkorzekwa · Pull Request #990 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/990/)
- merged

[Draft: Merging anymodel:build_library_and_stats by danielkorzekwa ·
Pull Request #993 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/993)
- merged

[Dkorzekwa/any model calc one block scores by danielkorzekwa · Pull
Request #994 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/994)
- merged

[Draft: merge any_model: mip_and_realize_models by danielkorzekwa · Pull
Request #995 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/995)
- merged

[Dkorzekwa/any model other modeqls by danielkorztiekwa · Pull Request
#1007 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1007/)
- merged

PR to 1007: https://github.com/NVIDIA/Model-Optimizer/pull/1039 - merged

[Dkorzekwa/anymodel gptoss by danielkorzekwa · Pull Request #1020 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1020)
- merged

[Merge any_model tutorial by danielkorzekwa · Pull Request #1035 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1035)
- merged

[Merge mbridge distillation for any_model by danielkorzekwa · Pull
Request #1036 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1036)
- merged

[MR branch for the remaining difference between dkorzekwa/any_model an…
by danielkorzekwa · Pull Request #1047 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1047)
- merged

[Dkorzekwa/decilm hf code cleanup by danielkorzekwa · Pull Request #1071
·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1071)
- merged

[Dkorzekwa/decilm hf code cleanup 2 by danielkorzekwa · Pull Request
#1073 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1073)
- merged

[Dkorzekwa/anymodel subblock stats by danielkorzekwa · Pull Request
#1085 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1085)
- merged

[Dkorzekwa/anymodel subblock stats nodecilm by danielkorzekwa · Pull
Request #1102 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1102)
- merged

[Dkorzekwa/decilm cleanup post subblockstats by danielkorzekwa · Pull
Request #1103 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1103)
- merged

[code clean up by danielkorzekwa · Pull Request #1110 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1110)
- merged

Merging into main:

[Activation hooks redesign (reuse hooks component across both minitron
and puzzletron) by danielkorzekwa · Pull Request #1022 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1022)
- merged

[Dkorzekwa/puzzletron use importance hooks from prune by danielkorzekwa
· Pull Request #1115 ·
NVIDIA/Model-Optimizer](https://github.com/NVIDIA/Model-Optimizer/pull/1115)
- merged

</details>

<!-- Details about the change. -->

### Usage

Puzzletron tutorial:

https://github.com/NVIDIA/Model-Optimizer/tree/feature/puzzletron/examples/puzzletron

### Testing
The main e2e test for compressing 9 models with Puzzletron:

https://github.com/NVIDIA/Model-Optimizer/blob/feature/puzzletron/tests/gpu/torch/puzzletron/test_puzzletron.py

2-gpu nightly tests: 

-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24468209205/job/71501061203
-
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24470214159/job/71508152952

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Puzzletron: end-to-end heterogeneous pruning & NAS workflow with
AnyModel support, example pipelines, deployment and evaluation
utilities, and tools for converting/pruning and exporting compressed
checkpoints.

* **Documentation**
* Comprehensive Puzzletron tutorials, model-specific guides, evaluator
instructions, example configs, and changelog entry.

* **Chores**
* CI/workflow updates (extras installation, longer GPU test timeout),
pre-commit hook exclusion updated, and CODEOWNERS entries added.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Signed-off-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Signed-off-by: Daniel Korzekwa <daniel.korzekwa@gmail.com>
Signed-off-by: jrausch <jrausch@nvidia.com>
Signed-off-by: root <root@pool0-00848.cm.cluster>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
Co-authored-by: Liana Mikaelyan <45925959+LianaMikael@users.noreply.github.com>
Co-authored-by: J Rausch <38429553+j-rausch@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-16 00:48:11 +05:30
Keval MorabiaandClaude Sonnet 4.6 9b4f43a9cc Improve Megatron Tokenization: streaming, reasoning_content support, HF in-memory tokenization, etc (#1221)
### What does this PR do?

Type of change: New feature

Improvements to `megatron_preprocess_data` for Nemotron v3 post-training
datasets and Megatron-Bridge distillation workflows:

- **`--reasoning_content`** flag (`strip` / `inline` / `native`) to
handle the `reasoning_content` field in Nemotron Post-Training v3
assistant messages
- **No intermediate JSONL** for HuggingFace datasets — load directly
from Arrow cache via `_iter_hf_as_json` + `process_hf_split`
- **Return output prefixes** (`list[str]`) from the Python API so
callers can build `--data_paths` without hardcoding paths; also printed
at end of run
- **Gzip input support** — `.jsonl.gz` files accepted directly;
`--input_dir` globs both `*.jsonl` and `*.jsonl.gz`
- **`--strip_newlines`** flag (opt-in) to replace newlines with spaces
in plain-text values; default preserves newlines (no breaking change for
code/structured-text datasets)
- **`--hf_streaming`** flag for very large datasets — only consumed rows
are downloaded; automatically falls back to non-streaming (with a
warning) if `--hf_max_samples_per_split` is not set, since streaming
without a cap is slower than cached non-streaming
- **Auto-shuffle** when `--hf_max_samples_per_split` is set — reservoir
sampling (buffer=10,000, seed=42) applied before capping to avoid biased
prefix sampling
- Remove `_document` suffix from output filenames (`_text.bin` instead
of `_text_document.bin`)
- Fix duplicate BOS token for chat-template data
(`add_special_tokens=False`)
- Fix `TypeError` in `process_hf_split` (`sum(list)` not `sum(int)`)
- Suppress duplicate prints across pool workers via
`_is_main_or_first_worker()`
- Raise `KeyError` instead of warning for missing JSON keys
- Default `hf_split=None` (all splits) instead of `"train"`

### Usage

```python
from modelopt.torch.utils.plugins.megatron_preprocess_data import megatron_preprocess_data

# Nemotron v3 with reasoning content preserved inline as <think>...</think>
prefixes = megatron_preprocess_data(
    hf_dataset="nvidia/Nemotron-Post-Training-Dataset-v3",
    json_keys=["messages"],
    tokenizer_name_or_path="Qwen/Qwen3-0.6B",
    output_dir="tokenized/",
    workers=32,
    reasoning_content="inline",
)
# prefixes == ["tokenized/nvidia--Nemotron-Post-Training-Dataset-v3_..._messages"]
data_paths = [x for p in prefixes for x in ("1.0", p)]

# Large pretraining dataset — stream + cap (auto-shuffled before capping)
prefixes = megatron_preprocess_data(
    hf_dataset="nvidia/Nemotron-CC-v2.1",
    hf_name="High-Quality",
    hf_max_samples_per_split=5_000_000,
    hf_streaming=True,
    json_keys=["text"],
    tokenizer_name_or_path="Qwen/Qwen3-0.6B",
    output_dir="tokenized/",
    workers=32,
    append_eod=True,
    strip_newlines=True,
)
```

### Testing

- New unit tests
- Tested tokenization on Nemotron Pretraining and Post-training v3
datasets

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (output filename changed:
`_text_document` → `_text`; existing callers need to re-tokenize or
rename files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Preprocessing tool: configurable reasoning-content modes
(strip|inline|native), optional newline stripping, gzip (.jsonl.gz)
input support, HF streaming mode, auto-shuffle when per-split max
samples set, returns output-file prefixes, and processes all HF splits
by default without writing intermediate JSONL.

* **Documentation**
* Consolidated and reformatted dataset preparation and tokenization
guidance; updated install/auth instructions and examples.

* **Tests**
* Tests updated to validate returned prefixes, reasoning-content
behaviors, gzip input handling, HF streaming warnings, and HF output
assertions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 01:08:11 +05:30
Keval Morabia d0bf0bef96 Remove deprecated Nemo 2.0 references / examples (#1098)
### What does this PR do?

- Remove `examples/nemo_run` and other deprecated Nemo 2.0 references
- Add Megatron-Bridge example links where missing

<!-- Details about the change. -->

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Removed deprecated NeMo 2.0 support and related example flows and
utilities.

* **Documentation**
* Updated docs and examples to emphasize Megatron-Bridge / Megatron-LM
and refreshed technique/deployment guidance and links.

* **New Features**
* Added CLI options for additional parallelism (context/expert
tensor/expert model) in Megatron-Bridge distillation.

* **Chores**
* Removed legacy CI configs and refreshed container image tags across
examples and docs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-24 14:11:13 +05:30
Keval Morabia f22f4f5522 Add Full TE Spec support for Megatron Pruning DynamicModules + MoE bug fixes (#1024)
### What does this PR do?

Type of change: Improvement + Bug Fix <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

Quantization recently added support for Full TE spec. Adding same for
Pruning as well so we can retire ModelOpt spec and just use standard TE
spec.
**NOTE: We still dont support TEGroupedGemm and instead use TE
SequentialMLP for now (but this can be configured in standard TE Spec so
we dont need modelopt spec)**

Note that this does not affect the usage of the pruning workflow but
makes pruning slightly faster and may result in slightly different
pruned model because of different kernel and numerics.

[Bug fix]: Previously NAS-based pruning for MoE models would hang when
evaluating MMLU for pruned candidate models because of a bug. Fixed in
this PR as well

[Bug fix]: Previously hidden size importance hooks were not applied to
pre_mlp_layernorm for MoE layers. Fixed in this PR as well resulting in
a significant improvement in MMLU for Qwen3-30B-A3B


### Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] Unit tests updated and passing
- [x] Compare pruning results for Qwen3-8B -> 6B. ⚠️ Difference in MMLU
scores resulting in a different best picked model. But scores more or
less in similar range - difference may be because of different kernel
for TE layers
```
Least important 6 layers:
    ModelOpt Spec: 27, 28, 29, 31, 32, 33
    TE Spec: 27, 28, 30, 31, 32, 33

Top 10 pruned candidates:
| num_layers | hidden_size | ffn_hidden_size | Params (B) | MMLU (ModelOpt Spec) | MMLU (TE Spec) |
|------------|-------------|-----------------|------------|----------------------|----------------|
| 34         | 3328        | 11264           | 5.99       | 0.390                | 0.397          |
| 30         | 3584        | 11776           | 5.99       | 0.572 [BEST]         | 0.575          |
| 36         | 3840        | 8192            | 5.98       | 0.511                | 0.511          |
| 36         | 3584        | 9216            | 5.98       | 0.477                | 0.497          |
| 36         | 3072        | 11776           | 5.97       | 0.278                | 0.252          |
| 32         | 3584        | 10752           | 5.96       | 0.542                | 0.541          |
| 36         | 3328        | 10240           | 5.92       | 0.365                | 0.412          |
| 34         | 3840        | 8704            | 5.91       | 0.537                | 0.539          |
| 30         | 4096        | 9216            | 5.90       | 0.566                | 0.591 [BEST]   |
| 34         | 3584        | 9728            | 5.89       | 0.499                | 0.510          |
```
- [x] Compare pruning results for Nemotron-Nano-9B-v2 -> 7B. MMLU scores
slight difference but best pruned model selection same
```
Least important 8 layers (Before and After): [43, 44, 45, 46, 47, 48, 50, 52]

Top 10 pruned candidates:
| num_layers | hidden_size | mamba_num_heads | mamba_head_dim | ffn_hidden_size | Params (B) | MMLU (ModelOpt Spec) | MMLU (TE Spec) |
|------------|-------------|------------------|---------------|-----------------|------------|----------------------|----------------|
| 50         | 4480        | 128              | 56            | 15680           | 7.00       | 0.211                | 0.202          |
| 56         | 4096        | 96               | 80            | 14336           | 7.00       | 0.438                | 0.436          |
| 48         | 4352        | 120              | 80            | 13824           | 7.00       | 0.679 [BEST]         | 0.679 [BEST]   |
| 56         | 4352        | 112              | 80            | 10240           | 7.00       | 0.516                | 0.520          |
| 54         | 4480        | 104              | 80            | 11264           | 7.00       | 0.263                | 0.262          |
| 46         | 4480        | 128              | 72            | 14848           | 7.00       | 0.610                | 0.617          |
| 50         | 4480        | 112              | 64            | 15680           | 7.00       | 0.426                | 0.421          |
| 54         | 4096        | 112              | 80            | 13312           | 7.00       | 0.579                | 0.589          |
| 56         | 4352        | 120              | 72            | 10752           | 7.00       | 0.466                | 0.469          |
| 52         | 4352        | 120              | 72            | 12800           | 7.00       | 0.561                | 0.560          |
```
- [x] Compare pruning results for Qwen3-30B-A3B -> 24B. Previously there
was a bug in hooks added so now we see a big improvement
```
Top 10 pruned candidates (~1 hour per candidate MMLU computation so skipped after 3):
| num_layers | hidden_size | num_attention_heads | num_moe_experts | Params (B)| MMLU (ModelOpt Spec) | MMLU (TE Spec) |
|------------|-------------|---------------------|-----------------|-----------|----------------------|----------------|
| 46         | 2048        | 28                  | 104             | 23.98B    | 0.663                | 0.698          |
| 40         | 2048        | 28                  | 120             | 23.95B    | 0.577                | 0.668          |
| 46         | 1792        | 24                  | 120             | 23.94B    | 0.435                | 0.500          |
| 46         | 2048        | 24                  | 104             | 23.88B    |                      |                |
| 40         | 2048        | 24                  | 120             | 23.87B    |                      |                |
| 46         | 1792        | 20                  | 120             | 23.85B    |                      |                |
| 40         | 2048        | 20                  | 120             | 23.78B    |                      |                |
| 46         | 2048        | 20                  | 104             | 23.78B    |                      |                |
| 42         | 2048        | 32                  | 112             | 23.62B    |                      |                |
| 48         | 1792        | 32                  | 112             | 23.54B    |                      |                |
``` 
- [x] Run pruning experiments for gptoss-20b (21B actually) -> 18B with
TESpec. Seems like GPTOSS MMLU is dropping steeply even in 15% pruning
```
"Only considering atmost 40% for width and 20% for depth pruning hparams
Skipping hparams_to_skip=['num_attention_heads'] during search space generation...
        Search space for num_layers: [20, 22, 24]
        Search space for hidden_size: [2048, 2304, 2560, 2816, 2880]
        Search space for num_moe_experts: [24, 32]
        Search space for moe_ffn_hidden_size: [2048, 2304, 2560, 2816, 2880]
        Total search space in consideration: 150

Top 10 candidates with scores:
        {'num_layers': 20, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.62B params, 0.3780 score
        {'num_layers': 22, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2560} -> 17.32B params, 0.4160 score [BEST SUBNET]
        {'num_layers': 20, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 17.27B params, 0.3523 score
        {'num_layers': 20, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.23B params, 0.3848 score
        {'num_layers': 22, 'hidden_size': 2560, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 17.13B params, 0.3062 score
        {'num_layers': 24, 'hidden_size': 2880, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2304} -> 17.09B params, 0.3984 score
        {'num_layers': 22, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2560} -> 16.94B params, 0.3957 score
        {'num_layers': 20, 'hidden_size': 2816, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 16.88B params, 0.3835 score
        {'num_layers': 22, 'hidden_size': 2560, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2816} -> 16.78B params, 0.2154 score
        {'num_layers': 24, 'hidden_size': 2304, 'num_moe_experts': 32, 'moe_ffn_hidden_size': 2880} -> 16.73B params, 0.0014 score"
```
- [x] Run pruning experiments for Nemotron-3-Nano-30B-A3B (31.5B
actually) -> 24B with TESpec
```
Top 10 candidates with scores:
        {'num_layers': 46, 'hidden_size': 2688, 'mamba_num_heads': 64, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} -> 24.00B params, 0.0000 score
        {'num_layers': 52, 'hidden_size': 2048, 'mamba_num_heads': 64, 'num_moe_experts': 128, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 24.00B params, 0.2764 score
        {'num_layers': 48, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3712} -> 24.00B params, 0.6098 score
        {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 64, 'num_moe_experts': 104, 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size': 3328} -> 24.00B params, 0.6233 score
        {'num_layers': 48, 'hidden_size': 2688, 'mamba_num_heads': 64, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 24.00B params, 0.6301 score [BEST SUBNET]
        {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 2560} -> 23.99B params, 0.6125 score
        {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3328} -> 23.99B params, 0.4255 score
        {'num_layers': 50, 'hidden_size': 2048, 'mamba_num_heads': 64, 'num_moe_experts': 128, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3584} -> 23.99B params, 0.2859 score
        {'num_layers': 50, 'hidden_size': 2688, 'mamba_num_heads': 56, 'num_moe_experts': 96, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3072} -> 23.99B params, 0.6125 score
        {'num_layers': 42, 'hidden_size': 2304, 'mamba_num_heads': 40, 'num_moe_experts': 120, 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size': 3584} -> 23.99B params, 0.0366 score
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ⚠️ TE has different kernels so
pruned model may be slightly different because of different numerics
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
OMNIML-3504


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Full Transformer Engine support for Minitron pruning; no custom model
spec required.

* **Bug Fixes**
* Resolved pruning hang on MoE models by correcting importance-hook
behavior.

* **Documentation**
* Updated changelog and example README; bumped recommended container tag
and expanded Docker run/mount guidance; adjusted release dates.

* **Improvements**
* Per-rank local activation handling, broader candidate caching, runtime
router scoring mitigation, clearer pruning status messages.

* **Tests**
* Updated tests to validate Transformer Engine backend and adjusted
dynamic-module expectations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-20 16:26:29 +05:30
Keval Morabia 4eacb0da72 Add trust_remote_code cli option for mbridge distillation (#934)
## What does this PR do?

Nemotron2/3 need `AutoBridge(..., trust_remote_code=True)` which was
missing previously

## Testing
<!-- Mention how have you tested your change if applicable. -->

Nemotron-nano-v2 can be distilled using tokenized
Nemotron-Pretraining-SFT-v1 data

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-25 22:30:53 +01:00
Keval Morabia c689ea1176 flush print megatron tokenization stats and update readme (#927)
## What does this PR do?

When running the script, I often see the print stats for tokenization
(every `log_interval`) not showing up or showing up very very delayed.
Hence using `print(..., flush=True)` to fix this.

Also update README that the example shown for tokenization takes too
long to run, split into multiple .jsonl files for efficiently running
the tokenization; and try out a smaller dataset first to test the script

## Testing
<!-- Mention how have you tested your change if applicable. -->

Split Nemotron-pretraining-SFT-v1 dataset into multiple .jsonl splits
and then tokenize them parallelly in different slurm jobs.

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-24 17:41:38 +01:00
Keval Morabia f78385e70e Improve megatron dataset preprocessing script and update docs (#918)
## What does this PR do?

Improve megatron dataset preprocessing script and update docs

## Usage
<!-- You can potentially add a usage example below. -->

```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
    --hf_dataset nvidia/Nemotron-Pretraining-SFT-v1 \
    --hf_name Nemotron-SFT-General \
    --hf_split train \
    --hf_max_samples_per_split 10_000_000 \
    --json_keys text \
    --tokenizer Qwen/Qwen3-0.6B \
    --output_dir /path/to/tokenized/data/qwen3 \
    --workers 32 \
    --max_sequence_length 256_000
```

```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
    --jsonl_paths /path/to/data1.jsonl /path/to/data2.jsonl ... \
    --json_keys text \
    --tokenizer Qwen/Qwen3-0.6B \
    --output_dir /path/to/tokenized/data/qwen3 \
    --workers 32 \
    --max_sequence_length 256_000
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- Downloaded and tokenized Nemotron-Pretraining-SFT-v1 with
Nemotron-Nano-v2 tokenizer

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated data preparation guides with new CLI patterns and Hugging Face
Hub integration instructions.

* **New Features**
* Added batch tokenization via directory input and direct Hugging Face
dataset downloads with flexible subset/split filtering.

* **Configuration Updates**
* Optimized distillation settings: adjusted optimizer parameters and
increased checkpoint retention.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-24 01:12:55 +05:30
Keval Morabia 8d67cb08d5 Add Megatron-Bridge recipe-free distillation example script (#861)
## What does this PR do?

**Type of change:** New example script <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

- [x] M-Bridge recipe-free distillation script so its more easier to run
and can support pruned models
- [x] Fix resuming distillation run

## Usage
<!-- You can potentially add a usage example below. -->

```python
torchrun --nproc_per_node 8 distill.py \
    --teacher_hf_path Qwen/Qwen3-8B \
    --student_hf_path Qwen3-8B-NAS-Pruned-6B \
    --tp_size 8 \
    --data_paths <climbmix 25% tokenized (~90B tokens)> \
    --data_path_to_cache /path/to/cache/climbmix_dataset_indices_qwen3 \
    --seq_length 4096 \
    --mbs 8 \
    --gbs 768 \
    --train_iters 28500 \
    --lr 1e-4 \
    --min_lr 1e-5 \
    --lr_warmup_iters 100 \
    --eval_interval 500 \
    --eval_iters 32 \
    --log_interval 10 \
    --output_dir qwen3_8b_6b_mbridge_distill
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] Re-ran Qwen3 8B -> 6B experiments and compare with Nemo2 results
from blog

Best subnet from NAS: `{'num_layers': 30, 'hidden_size': 3584,
'ffn_hidden_size': 11776} -> 5.99B params, 0.5718 score`

| Model | MMLU | GSM8K - flexible, strict | MBPP (coding) |
| ------- | ------ | ------- | ------- |
| Qwen3-8B | 74.9 | 87.5, 84.6 | 65.4 |
| Qwen3-8B-Pruned-6B | 57.6 | 11.6, 10.0 | 4.8 |
| Qwen3-8B-Pruned-6B (Distilled for 16k steps i.e. 50B tokens ~3k GPU
hours) | 71.6 | 78.0, 64.7 | 43.4 |
| Qwen3-8B-Pruned-6B (Distilled for 28.5k steps i.e. 90B tokens ~5.2k
GPU hours) | 71.9 | 78.1, 64.8 | 44.2 |
| Qwen3-4B | 70.0 | 81.1, 84.7 | 62.8 |

Previous Nemo2 experiments on depth pruned Qwen3 8B -> 6B (24 layers)
had MMLU ~72.0 so more or less similar. No hparam tuning done for
current M-Bridge distillation run

- [ ] (Separate PR) GitHub CI/CD test for example script with NeMo 26.02
container

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: N/A
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added complete distillation workflow and example for Megatron-Bridge
optimization.

* **Documentation**
* Enhanced setup guide with Docker workflows, data preparation steps,
and detailed distillation instructions.
  * Improved usage documentation and help references.

* **Improvements**
* Better data preprocessing output with human-readable formatting for
metrics.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-11 22:29:51 +00:00
Keval Morabia 944dd1a284 Move parallel_state init and warnings to Quant DynamicModule + MBridge pruning doc update (#854)
## What does this PR do?

**Type of change:** Minor improvement <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

Only quantization DynamicModules use the parallel_state attribute so for
all other model opt methods, we see a parallel state not initialized
warning which could be confusing hence moving it to QuantModule class
instead

Minor update to MBridge pruning docs

## Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Distributed parallel state support is now available in quantization
workflows for multi-GPU training.

* **Bug Fixes**
* Improved resource cleanup in distributed training to ensure proper
environment finalization.

* **Documentation**
* Updated example paths and added new manual pruning configuration
examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-04 22:23:21 +00:00
Keval Morabia fb5923c89e Add Megatron-Bridge pruning example scripts (#800)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

Megatron-Bridge pruning example scripts (HF input, HF / Megatron
output). Also defined some utility functions we can reuse for adding
examples for quantization or other optimizations:
- `modelopt.torch.utils.plugins.mbridge.load_mbridge_model_from_hf`:
Load HF to MBridge with ModelOpt spec in desired TP/PP/etc configuration
-
`modelopt.torch.utils.plugins.mbridge.get_hf_mbridge_calibration_loop`:
Create `forward_loop` for calibration on a HF dataset
- Supports all datasets available in
`modelopt.torch.utils.dataset_utils` (`cnn_dailymail`,
`nemotron-post-training-dataset-v2`, etc)
  - Support applying chat template for chat-based data

## Usage
<!-- You can potentially add a usage example below. -->

From `nvcr.io/nvidian/nemo:26.02.rc1` container (mount latest code to
`/opt/Megatron-Bridge` and `/opt/Model-Optimizer`)

```python
torchrun --nproc_per_node 2 /opt/Model-Optimizer/examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --prune_target_params 6e9 \
    --hparams_to_skip num_attention_heads \
    --output_hf_path /tmp/Qwen3-8B-Pruned-6B
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] Manually ran pruning script in nemo:25.11 container (plus modelopt
and mbridge mounted to latest) for Qwen3-8B and Nemotron-Nano-9B-v2 with
PP=8 and PP=4
- [ ] Added per-PR CI/CD test for example script

Results when pruning Qwen3 8B -> 6B (10 different configurations) with
and without chat template on the dataset samples. Perhaps MMLU is not
the right metric to look at.

| Layers | Hidden Size | FFN Hidden Size | Params | MMLU (Concatenated
messages) | MMLU (Applied Chat Template) |

|--------|-------------|-----------------|--------|------------------|----------------------|
| 34 | 3328 | 11264 | 5.99B | 0.401 | 0.393 |
| 30 | 3584 | 11776 | 5.99B | 0.588 | 0.576 |
| 36 | 3840 | 8192 | 5.98B | 0.507 | 0.518 |
| 36 | 3584 | 9216 | 5.98B | 0.477 | 0.469 |
| 36 | 3072 | 11776 | 5.97B | 0.255 | 0.249 |
| 32 | 3584 | 10752 | 5.96B | 0.554 | 0.549 |
| 28 | 4096 | 10240 | 5.94B | 0.400 | 0.438 |
| 36 | 4096 | 7168 | 5.93B | 0.461 | 0.438 |
| 36 | 3328 | 10240 | 5.92B | 0.362 | 0.359 |
| 34 | 3840 | 8704 | 5.91B | 0.515 | 0.546 |

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: ‼️ TODO
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

**New Features**
* Added new Megatron-Bridge pruning example demonstrating Minitron-based
model optimization with advanced pruning configurations.

**Documentation**
* Updated core project documentation to highlight Megatron-Bridge as a
supported optimization framework.
* Added comprehensive example documentation for Megatron-Bridge
workflows including pruning, distillation, and quantization.
* Updated pruning guides with Megatron-Bridge integration examples and
best practices.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-03 04:30:05 +00:00