17 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
Keval MorabiaandClaude Opus 5 4eb86524f0 [1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do?

Type of change: refactor (no functional change)

**[1/2] of a split. Merge this first; #2514 is [2/2] and is based on
this branch.**

Three example scripts had each reimplemented the same MLflow wiring: the
flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the
params/tags/artifacts a run uploads, and the open/close dance with its
status. The copies had already drifted — only `hf_ptq` wrote a
provenance pointer, only `vllm_serve` republished the resolved URI — and
every new tracked script meant another copy.

What a script records is now one declarative `Tool` record, **declared
in the script itself, beside the flags it reads**:

```python
# examples/megatron_bridge/quantize.py
QUANTIZE = Tool(
    name="megatron_bridge_quantize",
    tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...",
    variant_help="recipe name, or --quant_cfg if no --recipe",
    variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"),
    model=lambda args: args.hf_model_name_or_path,
    checkpoint=lambda args: args.export_megatron_path,
    texts=lambda args: resolved_recipe_texts(args.recipe),
    outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"},
)
```

`tracked_run` takes that record and runs the whole thing, so a script
adds tracking in three lines: `add_mlflow_args(parser, TOOL)`,
`resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args,
TOOL):`. The shared module knows no script's flags.

`examples/hf_ptq`, `examples/vllm_serve` and
`examples/megatron_bridge/quantize.py` move onto it. Three helpers fall
away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s
two flag pass-throughs).

### Usage

No user-facing change. The flags, their spellings and the experiment
naming are exactly as before; a script author now writes a `Tool`
instead of four functions.

### Testing

- `tests/unit/torch/utils/test_mlflow.py`,
`tests/examples/hf_ptq/test_hf_ptq_args.py`,
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the
only lane that runs it), which drives `quantize.py` for real: **34
passed**, locally and in this PR's `megatron` lane.
- `pre-commit run --files <changed>`: all hooks pass.
- The four suites shared four copies of a stand-in for the `mlflow`
module, which had drifted — one recorded artifacts as a list, another as
a dict, a third made `log_artifact` a no-op, so a test asserting on an
upload asserted nothing. They now share one
`tests/_test_utils/mlflow.py`, which also emulates the fluent API's
habit of opening a run when none is active.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `track_run` and
`checkpoint_run_tags` are removed, but neither shipped in a release
(0.47.0's `__all__` is `MlflowRunLogger`, `command_text`,
`current_user`, `default_experiment_name`, `validate_tracking_uri`, all
unchanged here). Three deliberate behaviour changes, each in shared code
and each tested:
- The `source_checkpoint_path` tag resolves to an absolute path where it
recorded the raw argument, which a chain of runs needs to join on the
pair. `run_tags` is shared, so this applies to every script that writes
the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local
`--hf_model_name_or_path`. A source that names no directory, such as a
Hub `org/name` id, is still recorded as given.
- `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a
block ending in `SystemExit(0)` as `FINISHED` where it recorded
`FAILED`, since a script that ends by calling `sys.exit()` rather than
returning has still finished.
- `.experiment.json`'s `tracking_uri` and the `run_url` built from it
drop a trailing `/` from the tracking URI, so the link is
`https://host/#/...` rather than `https://host//#/...`. Only reachable
by constructing `MlflowRunLogger` directly; every CLI path already
stripped the slash in `resolve_tracking_uri`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-visible change; the entry is in [2/2].
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split out of #2514. This half is the enabling refactor with no behaviour
change; #2514 is the feature it unlocks and is based on this branch. At
~605 changed lines of core logic it is over the ~500 guideline; the
owner accepted a two-PR split rather than three, and everything #2514
alone consumes — `split_tracking_credentials`,
`log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands
there rather than here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 19:57:03 +00:00
Keval MorabiaandClaude Opus 5 f2f0d6958e Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?

Type of change: new feature

`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.

Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:

- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.

Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.

One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.

The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).

### Usage

```bash
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --recipe general/ptq/nvfp4_default-kv_fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
    --mlflow https://<your-mlflow-server>/

# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```

The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.

### Testing

- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.

- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:25:30 -07:00
Shengliang Xu 3c87751903 Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?

Type of change: deprecation

Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.

| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |

`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.

A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.

#### The warning fires only when a flag is actually passed

`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.

`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.

### Usage

```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast

# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```

### Testing

Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:

- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.

Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the removal release for these flags is not decided here, only the
deprecation.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-14 13:32:28 -07:00
Keval MorabiaandClaude Opus 5 4956213d67 Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do?

Type of change: Bug fix

Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE
checkpoint: two
Megatron-Core → HuggingFace export bugs that make it unservable (§1–2),
the dead code the first
leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock
(§5), and four
distillation / data-prep bugs that stopped QAD itself from running (§6).

#### 1. Routed experts were exported packed, and vLLM cannot load that

```
AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter
  'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2
```

`mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16
upstream** checkpoint, which
really is packed. But that mapping is only used for **quantized**
export, and vLLM's quantized MoE
loader needs per-expert scales — both released NVFP4 checkpoints
(`Qwen3.6-35B-A3B-NVFP4` via
hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are
per-expert.

`_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split
each expert's fused gate+up
and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires
it up. The Megatron
checkpoint layout is unchanged, so affected checkpoints need only a
**re-export**.

`_verify_exported_keys` is relaxed to match: exported modules now
contribute their ancestor
prefixes, so expanding one source module into many is not reported as
~82 dropped tensors. A module
genuinely absent still has nothing beneath its prefix and is still
caught.

#### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed

`GPTModel.sharded_state_dict` drops `output_layer._extra_state` and
asserts it is empty. ModelOpt
keeps quantizer state there, so saving raised and — since that method
also backs the load plan —
loading silently restored the layer **unquantized**.

`keep_gpt_output_layer_extra_state()` retains it, applied from
`megatron_replace_quant_module_hook` so **every** Megatron model gets it
(Megatron-LM and NeMo
users included, neither of whom can import `mbridge`, which needs
`megatron.bridge`). It matches
the upstream body by AST before replacing it and self-disables
otherwise.

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is **closed, not
merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose
`sharded_state_dict` has no
pop-and-assert, so this side keeps the workaround.

Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of
per-token weight traffic** on a
model with ~2.9B active params.

#### 3. Cleanup

Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is
removed with
`_grouped_mlp_packing` and the `quantize=` / `record_quant_config=`
parameters that existed only to
serve it. Llama-4's `PackNameRemapping` is unaffected.

Two smaller review-driven fixes: the gated-split shape checks raise
`ValueError` rather than
`assert` (stripped under `-O`), and per-expert quant metadata is
recorded for
`local_expert_indices` rather than every global id, fixing
non-contiguous EP.

#### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron
quantize example

`examples/megatron_bridge/quantize.py` accepted the flag and threaded it
into the `mtq` config, but
`_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9
refs) and never in
`plugins/megatron.py` (0); `mode.py:247` only assigns it to modules
already exposing the attribute.
On a Megatron MoE model it was accepted and silently ignored — a trap,
since on a 256-expert model
it reads like a major quality lever. `hf_ptq.py` keeps it, where it
works.

#### 5. Fix multi-GPU image-text (VLM) calibration deadlocking

VLM calibration hung for 30 minutes and died on a gloo timeout whenever
`world_size > 1`, with no
error until the timeout fired.

`NemotronTarPlusJsonlIterable` split its budget with truncating
division, so the stream supplied
fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**).
`_ShardedIterable` gives
rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of
`world_size` leaves the
trailing rank one short — it exits the forward loop early and the others
block on the next
collective. The arithmetic predicts both observed hangs exactly: 1024 →
stall at **255/256**,
512 (yielding 510) → **127/128**.

Fixed both ends: subset budgets are distributed with `divmod` so they
sum exactly, and
`_ShardedIterable` truncates every rank to `floor(len / world)` — which
also covers `num_samples`
not being divisible by `world_size`, as the first fix alone does not.

Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024
samples): the configuration that
hung twice now completes 256/256 and exports. Unit tests cover both
fixes and fail without them.

#### 6. Fix the distillation path so QAD can actually run

Four independent bugs, all hit while running QAD end to end on
Qwen3.6-35B-A3B. Each blocks a
different configuration, and together they made every sequence length
OOM or abort.

- **Context parallel aborts.** The DDP config derived
`average_in_collective` from `--sft` alone,
but context parallel also needs per-token loss reduction, so any
`--cp_size > 1` run died on
  `Cannot average in collective when calculating per-token loss`.
- **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting
"without gathering full
logits", it cast the *whole* vocabulary to FP32 before selecting the
top-k, allocating two
`[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k
vocab. Reducing before
the cast is equivalent: widening is exact and temperature scaling is
monotonic, so the selected
  entries and the loss are unchanged.
- **MTP cross-entropy ran when it had nothing to recover.**
`skip_lm_loss` exempts the MTP heads
unconditionally, so their CE materialised another FP32 `[seq, vocab]`
tensor even when the MTP
head is excluded from quantization — as it is in every recipe here (775
of 906
`exclude_modules`, zero MTP `weight_scale` tensors exported). It is now
skipped **only** when the
model is quantized and MTP is left out of it; plain distillation such as
pruning recovery still
trains the MTP head. `test_mtp_excluded_from_quantization` pins all four
cases.
- **One bad record deadlocked data prep.** `megatron_preprocess_data`
re-raised chat-template
failures out of a pool worker, stalling the whole job until it timed out
— three malformed
records cost a multi-hour tokenization run. They are now skipped with a
warning, matching the
  existing handling of malformed JSONL a few lines above.

Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported
for a while but the
example never passed through; `test_qad` now exercises that path.

§4, §5 and §6 are independent of §1–3; happy to split them out if
reviewers prefer.

### Usage

No API change. Exported names now match the released checkpoints:

```
model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2}
lm_head.{weight,weight_scale,weight_scale_2}
```

### Testing

- `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert
rules. Verified these fail
without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` /
`NemotronHForCausalLM` as
  controls.
- `test_unified_export_megatron.py` — the gate/up split, per-block scale
slicing, the 0-dim scalar
fallback, and both directions of the `_verify_exported_keys` relaxation.
- `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases:
payload detection, no-op second
call, warn-and-skip on an unrecognised `sharded_state_dict`, and
`test_patches_stock_megatron_core`
which installs a replica of the real pre-fix upstream body (verified
against `be08ce5b1~1`) so the
  patched path is exercised whichever megatron-core is installed.
- `test_qad.py` — CI caught that its reference comparison still assumed
packed experts; fixed.

**End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200,
nemo:26.08:**

| | before | after |
| --- | --- | --- |
| export self-check | `Export dropped 82 tensor(s)` | passes |
| expert tensors | `mlp.experts.gate_up_proj` (packed) |
`mlp.experts.<E>.{gate,up,down}_proj` |
| **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** |
**`Loading weights took 25.61 s`** |
| **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** |

### Results these fixes unblocked

The export fix is what made a Megatron-produced NVFP4 MoE checkpoint
servable at all, so it enabled
a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16
baseline measured on the
same harness, from **paired** per-question tests:

| recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench
|
| --- | --- | --- | --- | --- | --- |
| **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 |
+0.48 | −0.44 |
| **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) |
−0.53 |
| **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** |
**+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) |

Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per
recipe; MMMU-Pro is 3 runs per
side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR
(68.33 → 71.33, p=0.25, 3 runs
per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at
100 questions and 114
tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim
either way.

#### QAD status (what §6 unblocked)

With the §6 fixes in place, QAD runs end to end on this model: 32 nodes,
`TP=1 PP=1 CP=1 EP=8`,
seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read,
MMMU-Pro at iteration 50
(0.84 B tokens), 3 runs per side, paired per-question:

| | MMMU-Pro | vs BF16 |
| --- | --- | --- |
| BF16 | 74.55 | — |
| W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** |
| + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) |

The PTQ deficit that motivated this work is no longer statistically
significant after 50 QAD
iterations. The improvement itself (+0.39 vs PTQ) is **not** significant
at p=0.41, and 50
iterations is 10% of the planned budget, so this is a direction rather
than a result. A full
six-benchmark sweep at iterations 50 and 300 is running; these numbers
will be superseded.

Two findings worth flagging beyond this PR:

- **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves
activations in BF16, so vLLM
cannot use the FP4 tensor cores and falls back to
`MarlinNvFp4LinearKernel` / `'MARLIN'` MoE.
W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel`
and beats W4A16 in
**12/12** shapes. The Marlin line count tracks the recipe exactly (one
W4A16 layer ⇒ one Marlin
  line ⇒ zero once `lm_head` is W4A4).
- **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for
the fastest recipe,
confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench,
AA-LCR and τ²-Telecom show no
  significant regression.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Megatron checkpoints
unaffected; re-export to gain the loadable layout. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it
now errors instead of being ignored -->
- Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings
addressed, threads resolved. -->

### Additional Information

Both export bugs were found while reproducing
`nvidia/Qwen3.6-35B-A3B-NVFP4` through
`examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is closed — see §2.
Labeled `cherry-pick-0.47.0`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 23:19:56 +05:30
Keval MorabiaandClaude Opus 5 61757c9781 Support quantized Qwen3-VL / Qwen3.5-VL (dense + MoE) export from Megatron-Bridge and verify exported checkpoints (#2276)
### What does this PR do?

Type of change: Bug fix + new feature

Enables quantized **Qwen3-VL** and **Qwen3.5-VL** (dense and MoE) →
unified HuggingFace export from Megatron-Bridge, and fixes the bugs
found along the way (ten from testing, plus a further round from
review). Most of them produced a valid-looking checkpoint and a green
test run, so the PR also makes the export path verify its own output.

Review is easiest commit-by-commit — each of the eleven commits is
self-contained and independently green.

#### Two blockers

1. **The exporter rejected the Megatron-Bridge VLM wrapper.**
`GPTModelExporter` only unwrapped MCore's `LLaVAModel`, so
`Qwen3VLModel` raised `ValueError: Input to GPTModelExport must be a
megatron.core.models.GPTModel!`. It now unwraps any wrapper exposing
`.language_model`.
2. **A VLM QAD checkpoint couldn't be loaded back.** `distill.py` passes
`distill_submodule="language_model"`, so the checkpoint holds only the
language model and the load died on `KeyError:
vision_model.patch_embed.proj.weight`. The loader now reads the
checkpoint metadata and targets `.language_model` when there are no
vision weights.

#### Four silent-corruption bugs

3. **VLM QAD discarded all ModelOpt state** (shipped in 0.46).
`ModeloptStateManager` requires state on the **root** of whatever gets
checkpointed. `quantize.py` quantizes the VLM root, so PTQ anchors it
there — but QAD checkpoints only `language_model`, orphaning it. The
saved `modelopt_state_dict` was literally `[]`; the `*_quantizer._amax`
tensors were still present but got dropped on load
(`dist_ckpt_strictness="assume_ok_unexpected"`), and the export came out
plain BF16 with no `hf_quant_config.json`.
4. **Fused grouped-GEMM MoE experts were omitted entirely.** The MoE
dispatch had no `else`, so an architecture without an
`experts.linear_fc1` rule exported *zero routed experts*. This hit
**`Qwen3MoeForCausalLM`** — a registered, supported architecture with no
export test — not just VLMs. A tiny Qwen3-MoE exported 37 of 45 tensors,
exit 0, no warning.
5. **Qwen3.5's GatedDeltaNet output norm was off by exactly 1.0.**
Megatron stores that gamma zero-centered, HF centers it on 1. Correct
names, correct shapes, wrong values — invisible to any structural check.
Megatron-Bridge's importer confirms the convention
(`RMSNorm2ZeroCenteredRMSNormMapping`).
6. **The disabled-quantizer patterns silently no-op on Megatron paths.**
They are written against HuggingFace module names. `*mixer.conv1d*`
matches only because MCore and HF happen to agree on "mixer" for Mamba;
`*linear_attn.conv1d*` never matched (Megatron calls it
`self_attention.conv1d`), so the conv1d was calibrated.
`*linear_attn.in_proj_a/b*` **cannot** match at all — Megatron fuses all
six GDN sections behind one quantizer — so the alpha/beta gates the
recipe wants in BF16 were exported in FP8.

#### Four more bugs, found only by running real checkpoints

The tiny fixtures could not reach these; each came from a real model or
a real quant format.

7. **Routed experts were written in a layout no real Qwen3.5 checkpoint
uses.** Real Qwen3.5 stores experts packed as `[num_experts, out, in]`;
the mapping emitted per-expert names, so every routed expert was
dropped. The fixture actively hid this: transformers *unpacks* experts
on `save_pretrained`, so the saved reference agreed with the wrong
output. Fixed with a `transpose` kwarg on `_pack_name_remapping` plus a
`GroupedMLPPacking` rule, so fused `TEGroupedMLP` reaches the same
packed tensors — which is also what lets Qwen3.5 keep grouped GEMM
(**22.1 GB/GPU vs 38.9 GB/GPU** on a 20-layer, 256-expert model).
8. **`_grouped_mlp_packing` was broken for NVFP4.** It max-merged
`weight_scale`, but NVFP4 needs each expert's per-block scales *stacked*
with only the global `weight_scale_2` merged; it also dequantized packed
`uint8` against per-block scales, and passed `block_size=None`.
`weight_scale_2` is never populated in an FP8 run, so the whole branch
was dead code under FP8-only testing. `_grouped_mlp_slicing` gained
`quantize=False` so packing can quantize once over the stack, matching
`_pack_name_remapping`.
9. **`_mtp_prefix` corrupted every VLM's MTP tensor names.** It did
`prefix.replace("model", "mtp")` uncounted, so
`model.language_model.layers.{}` became `mtp.language_mtp.layers.0.*` —
tensors present and correctly valued, under names nothing loads.
LLM-only prefixes contain one occurrence, so this was invisible until a
VLM with MTP was exported.
10. **`load_multimodal_components` rejected HF repo ids.** `quantize.py
--hf_model_name_or_path Qwen/Qwen3.5-0.8B` worked, but the documented
export step failed with *"It should be a directory"*. Its sibling in the
same file already resolved repo ids via `snapshot_download`; now it does
too. This affected **every** VLM export.

`Qwen3_5ForConditionalGeneration` (dense Qwen3.5-VL) is now registered
for export and vision passthrough, which bugs 9 and 10 were blocking.

#### New: Qwen3.5-VL

`GatedDeltaNetSlicing` splits Megatron's fused `in_proj` (`[query, key,
value, z, beta, alpha]`) into HF's `in_proj_qkv` / `_z` / `_b` / `_a`,
taking sizes from the module's own `in_proj_split_sections` so TP
sharding falls out. Widening coverage to Qwen3.5's *gated
full-attention* layers then exposed a further split bug: gated attention
packs a per-head output gate beside each query head, so `_qkv_slicing`
split 192 rows as 96/48/48 instead of 128/32/32. It now derives the
group stride from `config.attention_output_gate`, matching
Megatron-Bridge's `split_qkv_weights`. The non-gated path is unchanged.

#### New: the export path verifies itself

- `assert_exported_checkpoint_matches` compares an exported checkpoint
against the model it came from — key set, shapes (accounting for NVFP4
`uint8` packing), safetensors index consistency, and values — replacing
existence-only assertions in all three export tests.
- `GPTModelExporter.save_pretrained` now raises if the export dropped
tensors the source checkpoint has, so *user* runs on architectures CI
never sees are protected too, not just tiny models.
- Loading a checkpoint whose quantizer tensors have no restorable state
now raises instead of silently loading unquantized.
- `assert_has_modelopt_state` replaces `rglob("modelopt_state")`, which
passes on an empty state; `assert_no_quantizers_matching` fails on
future HF↔Megatron name drift.

The mapping is also table-driven now: vision-tower prefixes live in
`all_mcore_hf_vision_passthrough_mapping` and
`with_language_model_prefix` is shared, so adding a VLM no longer means
editing `unified_export_megatron.py`. Five call sites that answered "is
this a VLM" three different ways now share `get_language_model` /
`is_vlm_config`.

### Usage

```bash
# Dense VLM (Qwen3-VL) -- no extra flags
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --quant_cfg nvfp4 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron

torchrun --nproc_per_node 2 export_quantized_megatron_to_hf.py \
    --hf_model_name_or_path Qwen/Qwen3-VL-8B-Instruct \
    --megatron_path /tmp/Qwen3-VL-8B-NVFP4-megatron \
    --pp_size 2 --export_unified_hf_path /tmp/Qwen3-VL-8B-NVFP4-hf

# Gated MoE (Qwen3.5-VL, Qwen3-MoE) -- no extra flags either. The scripts derive the
# expert layout from the model config, so quantize / distill / export all agree.
# --no_moe_grouped_gemm forces SequentialMLP if you want it explicitly.
```

### Testing

All in `nvcr.io/nvidia/nemo:26.08` on 2x RTX 6000 Ada.

| Suite | Result | Time |
|---|---|---|
| `tests/examples/megatron_bridge/` (full) | 18 passed | 27m58 |
| `tests/gpu_megatron/torch/export/` | 38 passed | 2m13 |
| `tests/unit/torch/export/` | 186 passed | 1.5s |
| pre-commit (ruff, ruff format, mypy, bandit) | clean | — |
| `tests/examples/megatron_bridge/test_quantize_export.py` on **2 GPUs**
(`pp_size=2`) | 3 passed | 5m |

The export leg of `test_quantize_and_export` now scales with `num_gpus`
like its quantize leg
already did. Previously it was hardcoded to one process, so the
collective checkpoint load ran at
PP=1 on both the 1-GPU PR runner and the 2-GPU nightly — which is how a
guard that raised on only
some pipeline stages (and therefore hung the job) reached review. The
dense `qwen3` case was dropped
in exchange: `qwen3_moe` already covers the non-VLM script path,
`qwen3vl` covers a dense decoder,
and that case was the one exceeding the 300s cap in CI.

#### Model coverage

`tests/gpu_megatron` runs in-process and is cheap, so it owns
per-architecture **mapping**
correctness. The example tests spawn `torchrun` per step and are ~50x
slower per case, so they
cover **script wiring** only — CLI flags, recipe resolution, and
checkpoint hand-off between steps.

| Suite | Models |
|---|---|
| `test_unified_export_megatron` | llama, nemotron, nemotron_h, qwen3vl,
qwen3_moe, qwen3_5_moe_vl x {none, FP8, NVFP4, +/-KV} x {grouped GEMM,
SequentialMLP} + eagle / medusa / MTP (29 params) |
| `test_megatron_importer` | nemotron_h, llama export->import round-trip
|
| `test_moe_layout_choice` | per-architecture grouped-GEMM exportability
(6 architectures) |
| `test_distill_megatron` | KD loss mechanics |

| Model | prune | quantize+export | QAD | distill+export |
|---|:--:|:--:|:--:|:--:|
| qwen3 | Y | Y | Y | Y |
| qwen3_moe | - | **Y (new)** | - | - |
| qwen3vl | - | **Y (moved from QAD)** | - | - |
| nemotron_h | Y | **Y (new)** | - | - |
| qwen3_5_vl | - | - | - | Y |
| qwen3_5_moe_vl | Y | **Y (new, both expert layouts)** | Y | - |
| deepseek_v3 | Y | - | - | - |
| gemma3vl | Y | - | ~~manual~~ removed | - |

QAD's unique property is that ModelOpt state survives distillation,
which needs one LLM and one
VLM rather than one case per architecture. Moving the rest to
quantize+export drops a `torchrun`
launch each: QAD went from 3 CI cases to 2 while quantize+export went
from 1 to 4, adding two
architectures for about a minute.

#### Real-model validation

Tiny fixtures cannot catch layout or scale bugs that only appear at real
dimensions, so the export
path was run end-to-end on released checkpoints. This is where bugs 7-10
came from.

| Model | Run | Result |
|---|---|---|
| Nemotron-3.5-Lightning-30B-A3B | NVFP4 4o6 PTQ → export → MMLU |
**0.7825 ± 0.0105** (gate 0.75) |
| Nemotron-3.5-Lightning-30B-A3B | Minitron pruning | 22.28B/3.00B
active, **0.5944** (gate 0.58) |
| Qwen3.5-0.8B (dense VLM) | FP8 PTQ → export → MMLU | BF16 0.4895 →
**0.4832** (±0.0127) |
| Qwen3.5-35B-A3B, half-depth (20 layers, 256 experts) | FP8 + NVFP4 PTQ
→ export | keys + shapes + **values** match reference |
| Qwen3.5-35B-A3B, full | FP8 PTQ | OOM on 2x48GB (see below) |

The half-depth model keeps real weights, real dims and all 256 experts.
Both expert layouts produce
identical key sets, and all exports pass
`assert_exported_checkpoint_matches(..., check_values=True)`
— every tensor, including all 20 x 256 experts, dequantizes to within
tolerance of the BF16
reference, so a transposed or mis-ordered expert stack would fail. NVFP4
lands in the correct packed
layout (`gate_up_proj [256, 1024, 1024]` U8, `weight_scale [256, 1024,
128]` E4M3,
`weight_scale_2 []` F32). Its *accuracy* is not meaningful — truncating
to 20 of 40 layers leaves a
chance-level model (BF16 0.2322, FP8 0.2538) — so it validates
correctness, not quality.

**Re-validated on the final code.** The numbers above were first taken
mid-review; since then the
NVFP4 block-scale merge changed on both packed paths, the vision-tower
download became two-stage,
and an expert-layout load guard was added. Both gating runs were
therefore repeated end to end:
Nemotron went 0.7748 → **0.7825 ± 0.0105** and Qwen3.5-0.8B went 0.4678
→ **0.4832 ± 0.0127**, with
the rest of the Nemotron pipeline reproducing exactly (3519 quantizers,
69GB checkpoint, 21GB
export). Both deltas are inside their own stderr, so the claim is that
the rework costs no accuracy
— not that it improved it. The Nemotron export also runs at `--pp_size
2`, exercising the new
collective layout guard on a real 30B MoE across pipeline stages.

Two limitations worth stating plainly:

- **No quantized accuracy number for a full-size MoE.** The full 35B
OOMs at 47.37 GiB while
*constructing* the model on 2x48GB, with grouped GEMM already enabled,
so no calibration knob
  helps. Needs more GPUs than this setup has.
- **vLLM cannot yet serve packed FP8 Qwen3.5 experts.** `vllm
0.24.1.dev0` builds its fused expert
mapping weight-only, rewriting `experts.down_proj_input_scale` to
`w2_weight_input_scale` while the
parameter it registers is `w2_input_scale`. This is upstream and
independent of how the checkpoint
is produced — both of our export paths fail it identically. The 0.8B
numbers above are unaffected
(dense), and the packed exports are verified against the reference
checkpoint instead.

#### Guard verification

Each new guard was made to fire, not just to compile:

| Guard | Verification |
|---|---|
| Export self-check | Disabled the MoE guard, re-exported Qwen3-MoE -
independently reported all 24 dropped tensors. No false positives across
llama, nemotron, qwen3, qwen3-moe, qwen3vl, qwen3.5-vl, deepseek_v3
incl. eagle / medusa / MTP |
| Dropped-state raise | Deleted `modelopt_state` from a checkpoint with
50 quantizer tensors - raised instead of loading unquantized |
| NVFP4 value check | Flipped a `q_proj` - failed at `max_rel_err=1.74`
against a 0.3 threshold |
| Zero-centered gamma | Reproduced the off-by-1.0 on a good export -
caught as "not bit-exact" |
| Exclusion guard | Asserts no calibrated quantizer matches `conv1d` /
`mlp.router` / `output_layer` |

Exported artifacts are validated, not just their existence: 0 missing
keys vs reference, vision
tower bitwise-identical, dequantized weights within FP8 E4M3 error
(<=4.6%). The
`in_proj_a`/`in_proj_b` check is load-bearing - swapped alpha/beta would
still match on shape but
show ~100% error.

Also ran a tiny-Qwen3 **LLM** control through both steps to confirm the
exporter changes are a
no-op off the VLM path.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the scripts now derive the
MoE expert layout from the model config, building SequentialMLP only for
architectures with no `experts.linear_fc1` rule, and the exporter raises
rather than dropping experts it has no rule for. Those runs previously
"succeeded" while writing a checkpoint containing no expert weights, so
no working behaviour is removed. `--no_moe_grouped_gemm` forces
SequentialMLP explicitly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — approved (round 10: 0
CRITICAL, 0 IMPORTANT, 0 new suggestions); CodeRabbit approved earlier

### Additional Information

**MoE expert layout is now chosen automatically.** Only Nemotron-H can
export fused grouped-GEMM experts, so every other MoE architecture would
otherwise need `--no_moe_grouped_gemm` on all four scripts or hit a wall
at export. The scripts derive the layout from the model config — grouped
GEMM unless it would not be exportable — so they agree without threading
a flag. This changes MoE activation scales from one shared scale to
per-expert for the affected architectures.

Known gaps, unchanged by this PR:

- **Gated MoE still cannot use fused grouped GEMM.**
`_grouped_mlp_slicing` emits one weight per expert with no gate/up split
— its only prior caller, Nemotron-H, is non-gated, so every other MoE
architecture is built as `SequentialMLP` (see below). Adding that split
would restore the faster layout, but it needs a deliberate call on
activation-scale semantics: grouped GEMM keeps **one shared** activation
scale across experts while `SequentialMLP` has **per-expert** scales, so
the two are not numerically equivalent. It also needs EP>1 coverage.
- **Qwen3.5's alpha/beta gates share Megatron's fused `in_proj`
quantizer,** so they can only be kept in BF16 at export, not excluded by
name. Full fidelity needs per-section quantizers on the fused
projection.
- **Anchoring ModelOpt state on `.language_model`** (which would let
`quantize.py` quantize the language model directly and drop its
name-based non-LM disabling) needs a coordinated Megatron-Bridge change:
`save_sharded_modelopt_state` is ModelOpt code, but the restore the
Bridge path uses is Bridge's own and unconditionally restores onto the
root.
- **Gemma3-VL** remains Megatron-checkpoint only (`OMNIML-5366`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added Muse Glimmer AutoQuantize and Alpamayo QAD workflows.
* Added streaming Kimi-K3 conversion and NVFP4 activation headroom
calibration.
  * Added SFT-masked distillation for Megatron-Bridge.
* Added unified Hugging Face export for quantized Qwen3-VL and
Qwen3.5-VL checkpoints.
* MoE expert layouts are selected automatically, with an option to force
sequential experts.

* **Bug Fixes**
* Improved export validation for tensor coverage, MoE mappings,
quantizer state, and NVFP4 scales.
  * Fixed Qwen3.5-VL GatedDeltaNet export handling.
  * Preserved visual-model weights exactly during export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-02 13:29:59 +00:00
Keval MorabiaandClaude Opus 5 fbcdc16c2d Remove deprecations marked in 0.45 and 0.46 (#2182)
### What does this PR do?

Type of change: Backward breaking change (deprecation removal)

Ahead of the 0.47 code freeze, this removes every deprecation still
outstanding from the previous two releases (0.45 and 0.46). Two are
intentionally left in place: the **Python 3.10** drop and the
**transformers 4.x** drop

| Deprecation | Marked in | Replacement |
| --- | --- | --- |
| `--auto_quantize_bits` / `_method` / `_score_size` / `_cost_model` /
`_active_moe_expert_ratio` | 0.46 | AutoQuantize `--recipe` |
| `examples/llm_ptq` symlink + `examples/vlm_ptq/` forwarder | 0.46 |
`examples/hf_ptq` (`--vlm` for VLMs) |
| `QuantizationArgumentsWithConfig` alias | 0.45 |
`QuantizationArguments` |
| `QFORMAT_ALIASES` short names | 0.45 | canonical preset basenames |
| `layerwise` bool + flat `layerwise_checkpoint_dir` | 0.45 | nested
`layerwise: {enable, checkpoint_dir}` |
| in-trainer `quant_cfg` / `--quant_cfg` | 0.45 | `--recipe` |

#### Two things worth a closer look

**1. The `use_sequential` alias goes too.** It is the pre-#1251 alias on
`QuantizeAlgorithmConfig.layerwise` and only ever carried a bool. Once
the bool form is rejected it cannot accept a valid value, so keeping it
would only produce a differently-worded validation error. Note the
direction is breaking either way (`extra="forbid"`): a pre-0.45
`modelopt_state` carrying `use_sequential: True` or a top-level
`layerwise_checkpoint_dir` now fails validation instead of being
migrated.

**2. Removing in-trainer `--quant_cfg` required two new recipes.** The
`examples/gpt-oss` QAT flow ran on `--quant_cfg
MXFP4_MLP_WEIGHT_ONLY_CFG` and no `general/ptq/` recipe covered it. This
PR adds `general/ptq/mxfp4_mlp_weight_only` and
`general/ptq/nvfp4_mlp_weight_only`, verified to `model_dump` identical
to `mtq.MXFP4_MLP_WEIGHT_ONLY_CFG` / `mtq.NVFP4_MLP_WEIGHT_ONLY_CFG`,
and migrates the gpt-oss README, both SFT configs, `sft.py` and
`tests/examples/gpt-oss/test_gpt_oss_qat.py`. `examples/llm_qat` was
already recipe-only.

### Usage

```bash
# AutoQuantize: --auto_quantize_* flags -> an AutoQuantize recipe
scripts/huggingface_example.sh --model $HF_PATH \
  --recipe general/auto_quantize/nvfp4_fp8_at_5p4bits --calib_batch_size 4

# --qformat / --quant_cfg: short name -> canonical preset basename
#   int8_sq -> int8_smoothquant                nvfp4_mse           -> nvfp4_w4a4_weight_mse_fp8_sweep
#   int8_wo -> int8_weight_only                nvfp4_local_hessian -> nvfp4_w4a4_weight_local_hessian
#   w4a8_awq -> w4a8_awq_beta                  fp8_pb_wo           -> fp8_2d_blockwise_weight_only
#   nvfp4_awq -> nvfp4_awq_lite                fp8_pc_pt           -> fp8_per_channel_per_token
scripts/huggingface_example.sh --model $HF_PATH --quant int8_smoothquant

# VLM PTQ: examples/vlm_ptq -> examples/hf_ptq with --vlm
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --vlm

# gpt-oss QAT: --quant_cfg <CFG name> -> --recipe <recipe path>
accelerate launch --config_file configs/zero3.yaml sft.py \
  --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b \
  --recipe general/ptq/mxfp4_mlp_weight_only --output_dir gpt-oss-20b-qat
```

```python
# Layerwise calibration: bool / flat key -> nested LayerwiseConfig
quant_cfg["algorithm"] = {"method": "gptq", "layerwise": {"enable": True, "checkpoint_dir": "/path"}}
```

### Testing

- `tests/unit/recipe` (229 passed),
`tests/unit/torch/quantization/test_config_validation.py` (79 passed),
`tests/examples/hf_ptq/test_hf_ptq_args.py` (23 passed).
- Verified the two new recipes `model_dump` identical to the `mtq.*_CFG`
constants they replace.
- `ruff check modelopt/ examples/ tests/` clean; `ruff format --check`
clean on all changed Python files.
- GPU suites
(`tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py`,
`test_accelerate_gpu.py`, `test_gptq.py`) had their preset / layerwise
literals updated but were not run locally — relying on CI.
- `examples/llm_qat/ARGUMENTS.md` is hand-edited to match what the
`generate-arguments-md` hook emits; the generator could not run locally
(missing `transformers` package metadata in this environment).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — that is the point of the PR:
it removes shims deprecated in 0.45/0.46. Callers must move to the
replacements in the table above. Additionally, a pre-0.45
`modelopt_state` carrying `use_sequential` or a top-level
`layerwise_checkpoint_dir` will now fail config validation.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests migrated to
the surviving APIs;
`TestLayerwiseNestedConfig::test_legacy_forms_rejected` pins that the
bool form, the `use_sequential` alias and the flat checkpoint-dir key
are all rejected. Tests covering the removed shims were deleted.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up: the transformers 4.x drop deprecated in 0.46 is still
outstanding and will need its own PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added MXFP4 and NVFP4 weight-only quantization recipes for MLP and MoE
layers.
- Added shared layer exclusions for more accurate effective-bits
calculations.

## Improvements
- Updated PTQ, QAT, GPT-OSS, deployment, and quantization-format
examples with current recipe names and configuration formats.
- Standardized layerwise settings under nested configuration fields.

## Breaking Changes
- Removed deprecated AutoQuantize options, `quant_cfg` usage, format
aliases, legacy layerwise settings, and compatibility example paths.
- Recipe-based and nested configuration forms are now required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 11:43:47 +05:30
Keval MorabiaandClaude Opus 5 f99523279a Minitron pruning fixes for Nemotron-3.5-Lightning-30B-A3B and Deepseek (#2159)
### What does this PR do?

Type of change: Bug fix + new feature

Two model families that could not be pruned end-to-end now can:

- **Nemotron-3.5-Lightning**
(`nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16`) — a native
`NemotronHForCausalLM` that ships without remote code and carries MTP
heads. Fixes a calibration crash and HF-export failures on the modern
Megatron-Bridge / transformers stack.
- **DeepSeek-V3** — fixes an MLA Q-LoRA crash during calibration, and
adds a `candidate_filter` search option to `mcore_minitron` so its
MoE-FFN dimensions stay prunable while remaining representable in HF.

Also makes a rank-local failure under pipeline parallelism fail fast
instead of stalling.

#### 1. Nemotron Lightning: prune + HF export
(`examples/megatron_bridge/prune_minitron.py`)

1. **MTP calibration crash.** On newer Megatron-LM, `mtp_process` is
derived from the hybrid *pattern*, not from `mtp_num_layers`. Setting
only `mtp_num_layers=0` in the calibration provider overrides was
insufficient: the provider's `finalize()` re-appended the MTP suffix to
`hybrid_layer_pattern` (because `mtp_hybrid_override_pattern` was still
set and `mtp_use_repeated_layer=True`), so `mtp_process=True` while
`mtp_num_layers=0` and the calibration forward hit `assert
self.config.mtp_num_layers > 0`. Fix: also clear
`mtp_hybrid_override_pattern` in the calibration overrides so MTP is
fully disabled (MTP heads are dropped from the pruned model, as before).

2. **HF export via a config-only bridge (hybrid models only).** The old
export built a *dummy* HF model to obtain the bridge, then streamed
weights. This breaks on native NemotronH because (a) native
`NemotronHConfig` makes `hybrid_override_pattern` a read-only property,
and (b) transformers 5.12 saves the input embedding under a different
key than the bridge mapping expects (`backbone.embedding` vs
`backbone.embeddings`); the mismatch made `build_conversion_tasks` drop
the embedding task on its owning rank, leaving an owner-less PP
placeholder that crashed `save_hf_weights` with `Object must exist on at
least one PP rank`. Fix: stream weights through a **config-only** bridge
(`AutoBridge.from_hf_config(hf_cfg).save_hf_pretrained(...)`), available
since Megatron-Bridge 0.5.0 (nemo:26.06). A config-only bridge has
`hf_keys=None`, so the embedding task is never dropped, no dummy model
is built, and the output uses the canonical HF key names.

This is **restricted to hybrid providers**, which are the only models
that need it; non-hybrids keep the dummy-model path that CI has always
exercised.

Writing the source artifacts is now rank-0-only. Every rank used to
write the source `config.json`, which races with the pruned
`config.json` that `save_hf_pretrained` writes from rank 0 alone: a late
write from another rank leaves a checkpoint whose config does not match
its weights. This reproduced intermittently on both Qwen3 and NemotronH
before the fix, and 3/3 clean runs after.

`save_hf_pretrained` takes no `trust_remote_code` argument — it reads
the flag **off the bridge** to fetch the source checkpoint's artifacts,
and `from_hf_config` cannot infer it because
`AutoConfig.from_pretrained` consumes the kwarg rather than storing it
on the config. So the flag is set explicitly on the bridge instance;
otherwise remote-code models would silently lose it.

3. **Config write-back correctness:**
- `hybrid_override_pattern` is only written for older remote-code
configs that lack `layer_types`; native configs carry the cadence in
`layer_types` (read-only `hybrid_override_pattern` is skipped).
- `n_shared_experts` is preserved (a fixed count) instead of being
re-derived by `moe_shared_expert_intermediate_size //
moe_ffn_hidden_size`, which is DeepSeek-style logic that would corrupt
NemotronH's count.

Non-hybrids, VLMs, and Megatron-Bridge builds without config-only export
keep the dummy-model path, with a `warn_rank_0` when a hybrid has to
fall back. The README's `transformers<5` workaround is **removed**: it
existed because the dummy-model path broke on transformers 5, and the
config-only path handles NemotronH on every supported container.

#### 2. `candidate_filter` for `mcore_minitron`
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

DeepSeek-style MoE configs have no explicit shared-expert-size field:
they size the shared expert as `n_shared_experts *
moe_intermediate_size`, where `moe_intermediate_size` is the (also
prunable) **routed** expert size. So only candidates with
`moe_shared_expert_intermediate_size % moe_ffn_hidden_size == 0` can be
written back to HF at all.

Candidates come from a Cartesian `product()` of independent per-hparam
choice lists, so no per-hparam restriction can express a constraint
*between* two hparams. New optional `candidate_filter` search-config key
(default `None`, so existing behaviour is unchanged): a callable that
rejects candidate configs before the metric computation, making the
search cheaper rather than more expensive. It receives **every**
supported hparam, with non-searched ones filled in from the model
config, so a filter still works when one of its hparams was skipped or
had a single choice.

Rejected candidates are not cached, so — like `score_func`, whose cached
scores are reused without re-validation — the filter is assumed
unchanged when resuming from a `checkpoint`.

`prune_minitron.py` wires this up for DeepSeek-style configs, so
**both** `moe_ffn_hidden_size` and `moe_shared_expert_intermediate_size`
stay prunable (the search then only picks shared sizes that are a
multiple of the routed one). A `--prune_export_config` that violates the
constraint never reaches the filter, so the export path now raises
`ValueError` instead of writing a checkpoint whose config disagrees with
its weights.

#### 3. MLA Q-LoRA pruning
(`modelopt/torch/prune/plugins/mcore_minitron.py`)

Pruning any MLA model with `q_lora_rank` set died during calibration
with `AttributeError: 'tuple' object has no attribute 'view'`.

`hidden_size` importance estimation blanket-patches every
`TELayerNormColumnParallelLinear` with `return_layernorm_output=True` to
capture post-layernorm activations. When `q_lora_rank` is set, MCore
builds `linear_q_up_proj` as a `TELayerNormColumnParallelLinear` — the
Q-LoRA layernorm is fused into it, which is why `q_layernorm` is
`IdentityOp` — so it was patched too, even though its layernorm is over
the **latent rank**, not `hidden_size`. TE then returns `((out, ln_out),
bias)` and MCore's `q, _ = self.linear_q_up_proj(...)` leaves `q` a
tuple.

Isolated by probing the module before and after dynamic conversion:

| Setup | `linear_q_up_proj` returns | Forward |
| --- | --- | --- |
| Before conversion | `tuple(Tensor, NoneType)` | — |
| After conversion, no hooks | `tuple(Tensor, NoneType)` | OK |
| After conversion **+ importance hooks** | `tuple(tuple(Tensor,
Tensor), NoneType)` | AttributeError |

So conversion is innocent; registering the importance hooks is the
trigger. Fix: exclude MLA's Q/KV up-projections from both the patch and
unpatch loops. `test_mcore_mla_pruning` did not catch this because it
builds MLA without `q_lora_rank`, where MCore uses a plain
`linear_q_proj` and nothing is patched.

#### 4. Fail fast instead of stalling on a rank-local error under PP
(`modelopt/torch/utils/distributed.py`)

A rank raising inside a distributed entrypoint left the whole job
stalled until the process group timed out, with **no diagnostic output
at all**: the failing rank blocked in `cleanup()`'s barrier while its
peers blocked in `recv_from_prev_pipeline_rank_`, and Python only prints
a traceback once the enclosing `finally` returns. A crash on one rank
was indistinguishable from a slow job.

- `dist.cleanup()` skips the barrier when unwinding from an exception.
- New `dist.abort()` prints the traceback, flushes and exits
immediately. Skipping the barrier alone is **not** enough — a stack dump
showed the failing rank then blocking in `destroy_process_group` for the
same reason — so the error path must not tear the process group down at
all. `SystemExit` is re-raised rather than aborted, so an intentional
exit (e.g. the `--score_lower_bound` gate) keeps its exit code and
prints no traceback. Kept out of `cleanup()` so no library caller gets a
surprise process exit.
- Called from the entrypoints that wrap `main()` in `try/finally`: the
five `examples/megatron_bridge` scripts.

Measured on a 2-GPU PP run whose rank 0 raises during calibration: **10
min timeout kill with no visible error → 31s, exit 1, real traceback.**
This is a latent, pre-existing issue (the `try/finally` predates this
PR); it only surfaces on a failing PP run, which is why CI never hit it.

### Usage

```bash
torchrun --nproc_per_node 4 examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 \
    --pp_size 4 \
    --prune_target_active_params 3e9 \
    --output_hf_path /path/to/Nemotron-3.5-Lightning-30B-A3B-Pruned-A3.0B
```

### Testing

- **End-to-end on nemo:26.08.rc6** (4× GB300, transformers 5.12.1,
Megatron-Bridge with config-only export): pruning + export complete
(`EXIT=0`, "Saved pruned model … Done!"). The exported checkpoint has
canonical **plural** `backbone.embeddings.weight` keys, **0 MTP
tensors**, and a config that reloads correctly (`num_hidden_layers=52`
from `layers_block_type`, `n_shared_experts=1`,
`num_nextn_predict_layers=0`, pruned `hidden_size`/`mamba_*`/MoE dims,
reconstructed `hybrid_override_pattern`).

  <details>
<summary>Pruning search log (<code>--prune_target_active_params
3e9</code>)</summary>

  ```text
Top 10 Candidates with Scores

┏━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┳━━━━━━━━┳━━━━━━━━┓
┃ # ┃ export_config ┃ active_params ┃ params ┃ score ┃

┡━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━╇━━━━━━━━╇━━━━━━━━┩
│ 1 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.49B │ 0.5406
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 2 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 56,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2427 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 3 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 21.61B │ 0.2643
│
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 4 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 19.28B │ 0.4552 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3712} │ │ │ │
│ 5 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 104, │ 3.00B │ 22.28B │ 0.5860
│
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 6 │ {'num_layers': 52, 'hidden_size': 2560, 'mamba_num_heads': 48,
'mamba_head_dim': 48, 'num_moe_experts': 96, │ 3.00B │ 21.99B │ 0.2294 │
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3328} │ │ │ │
│ 7 │ {'num_layers': 48, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 104, │ 3.00B │ 23.68B │ 0.5231
│
│ │ 'moe_ffn_hidden_size': 1792, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 8 │ {'num_layers': 46, 'hidden_size': 2560, 'mamba_num_heads': 56,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 21.81B │ 0.5042 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3584} │ │ │ │
│ 9 │ {'num_layers': 52, 'hidden_size': 2688, 'mamba_num_heads': 48,
'mamba_head_dim': 56, 'num_moe_experts': 96, │ 3.00B │ 20.09B │ 0.2462 │
│ │ 'moe_ffn_hidden_size': 1536, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │
│ 10 │ {'num_layers': 52, 'hidden_size': 2304, 'mamba_num_heads': 64,
'mamba_head_dim': 64, 'num_moe_experts': 96, │ 3.00B │ 20.70B │ 0.5685 │
│ │ 'moe_ffn_hidden_size': 1856, 'moe_shared_expert_intermediate_size':
3072} │ │ │ │

└────┴───────────────────────────────────────────────────────────────────────────────────────────────────────────────────────┴───────────────┴────────┴────────┘


╭────────────────────────────────────────────────────────────────────────
Best Subnet
─────────────────────────────────────────────────────────────────────────╮
│ export_config {'num_layers': 52, 'hidden_size': 2304,
'mamba_num_heads': 64, 'mamba_head_dim': 64, 'num_moe_experts': 104,
'moe_ffn_hidden_size': 1856, │
│ 'moe_shared_expert_intermediate_size': 3072} │
│ active_params 3.00B │
│ params 22.28B │
│ score 0.5860 │

╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯

╭────────────────────────────────────────────────────── Pruned Model
Stats ───────────────────────────────────────────────────────╮
│ Total Parameters 22.28B │
│ Active Parameters 3.00B │
│ Memory (BF16, seq_length=8192, batch_size=8) weights: 42489.7 MB,
kv_cache: 384.0 MB, mamba_state: 190.5 MB, Total: 43064.2 MB │

╰─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯
  ```

  </details>

- **`tests/examples/megatron_bridge/test_prune_minitron.py`** —
`nemotron_h` now exports to HF and reloads (previously it stopped at a
Megatron checkpoint, since the dummy-model path needed
`transformers<5`), plus an `n_shared_experts` config assertion; the dead
`megatron_format` branch is gone. It runs on the CI container:
**verified on nemo:26.06.01 (transformers 5.8.1) and nemo:26.08.rc6.**
-
**`tests/gpu_megatron/torch/prune/plugins/test_mcore_mamba_minitron_pruning.py`**
— the `nas_memory_mb` search test now passes a `candidate_filter` and
asserts the exact number of rejected candidates (256 of the 512-combo
grid) plus the surviving candidates' validity; its `expected_top_k`
goldens are regenerated accordingly. Because
`moe_shared_expert_intermediate_size` is in that test's skip list, this
also covers the model-config fallback for hparams that are not in the
search space.

Verified on 2 GPUs, on both the CI container (nemo:26.06.01) and
nemo:26.08.rc6:

| Test | Result |
| --- | --- |
| `test_prune_minitron[qwen3]` | PASSED on 26.06.01 and 26.08.rc6 |
| `test_prune_minitron[deepseek_v3]` | PASSED (52s) — MLA Q-LoRA +
`candidate_filter` end-to-end |
| `test_prune_minitron[nemotron_h]` | PASSED on 26.06.01 (58s) and
26.08.rc6 (61s) |
| `test_mcore_mamba_hybrid_pruning_nas_memory_mb` | PASSED |
| `test_mcore_mamba_hybrid_pruning_nas_params` | PASSED (unchanged
sibling, run to check the regenerated goldens did not disturb it) |
| 2-GPU PP run failing on rank 0 | fails in 31s with a real traceback
(was a 10 min stall) |


### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `candidate_filter` defaults
to `None` (existing searches unchanged), and the config-only export is
limited to hybrid providers on nemo:26.08+, so dense / MoE / VLM exports
keep the path they use today.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ✅

### Additional Information

Enables the Prune + Distill workflow for
`NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16` (native, no-remote-code
`NemotronHForCausalLM` with MTP heads). Pruning-time MTP support was
scoped and intentionally deferred — MTP heads are dropped and can be
re-derived via a short SFT with `mtp_num_layers=1` on the
pruned+distilled model.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 21:05:56 +05:30
Keval Morabia 01c708e792 Add HybridModel MBridge support for nemo:26.08 (#2005)
### Description

MBridge main (nemo:26.08) will initialize Nemotron-H as HybridModel
instead of MambaModel (subclassed of HybridModel). Also make minimum
nemo container 26.04

### Testing 

Tested Nemotron-3-Nano PTQ with MBridge main (fails otherwise)

Tested locally `tests/gpu_megatron` and `tests/examples/megatron_bridge`
with `nemo:26.06.01` + Mount latest MBridge/Mcore

GH CICD tests will be added with nemo:26.08 release

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added TE HybridModel stack-spec support and enabled Megatron-Bridge
hybrid providers/models for export/import and runtime handling.
* **Updates**
* Dataset packing now oversamples raw text at **16x** and improves the
packed-mode underflow warning.
* Quantization: `--quant_cfg` now defaults to `None` unless explicitly
set (or via `--recipe`).
* Distillation example: validation settings are provided via a dedicated
top-level validation configuration.
* Improved plugin import warnings to report the originating call
location; model stats now support HybridModel.
* **Deprecations**
* Megatron-Bridge / Megatron-LM optimization features now require NeMo
container `nemo:26.04` or newer (`nemo:26.06` recommended).
  * The Mamba stack specification helper is deprecated.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-07-23 18:26:15 +05:30
Keval MorabiaandClaude Opus 4.8 21d0069e4d MBridge VLM distillation / QAD support (#1938)
## What

Adds VLM (e.g. Qwen3.5-VL, Gemma3-VL) knowledge-distillation / QAD
support to the Megatron-Bridge examples:

- `distill.py` distills only the **language model** submodule (vision
tower + projector untouched), reusing the LLM training path.
- New `export_distilled_megatron_to_hf.py` converts a distilled Megatron
checkpoint (**any** iteration) to HF. Required especially for VLM
distilled ckpt as it only has LM weights so we need to initialize full
VLM, swap LLM weights then save to HF
- Renames `export.py` → `export_quantized_megatron_to_hf.py`.

## Related upstream Megatron-Bridge PRs to be available in nemo:26.08
container:

- NVIDIA-NeMo/Megatron-Bridge#4707 — `DistillationProvider` submodule
distillation (non-blocking; added temporary WAR)
- NVIDIA-NeMo/Megatron-Bridge#4706 — MoE expert weight-mapping fix
(Qwen3.5-VL-MoE with moe_grouped_gemm=False). Also removed ModelOpt side
WAR previously added as it was not accurate; better to wait till next
container release or mount latest MBridge into the 26.06 container.

## Testing

- Qwen3.6-35B-A3B Pruning + Distillation with MMLU evaluation sanity
check (results below in comments)
- Cosmos 2 Reason 2B valiadted by SAs (results below in comments)
- Validated end-to-end on `nemo:26.06` (distill → separate HF export;
LLM + VLM, incl. TP→TP/PP reshard). `test_distill_vlm` runs the export
script as a CI e2e step.
- Many CICD tests for wide coverage of all mbridge scripts

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **New Features**
* Added a dedicated HuggingFace exporter for distilled Megatron
checkpoints, with distinct LLM vs VLM conversion flows.

* **Bug Fixes**
* Improved Megatron-Bridge distillation/export consistency, including
safer handling of VLMs and targeted submodule distillation.

* **Documentation**
* Updated Megatron-Bridge READMEs and tutorials to reference the new
quantized and distilled export scripts and revised CLI guidance.

* **Tests**
* Expanded distillation, QAD, and quantization/export tests to cover
LLM/VLM variants, with conditional skipping for unsupported MoE setups.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-16 15:16:30 +00:00
Keval MorabiaandClaude Opus 4.8 2fc352be2d Add VLM pruning and PTQ with image-text calibration (Megatron-Bridge) (#1792)
### What does this PR do?

Type of change: New feature

Adds **vision-language model (VLM) support** to the Megatron-Bridge
examples for both **Minitron pruning** (`prune_minitron.py`) and **PTQ**
(`quantize.py`). Only the **language model** is pruned/quantized — the
vision tower and vision→language projector are left in full precision —
and the full VLM is saved back. `hidden_size` is skipped for pruning
when it is shared with the vision→LM projector.

Supported VLMs (tested e2e): **Qwen{3,3.5}-VL** (dense; hybrid
GatedDeltaNet + gated attention) and **Gemma3-VL** (sliding/full
attention).

### Calibration (image-text)

Calibration is conditioned on real **image-text** data so the language
model's pruning importance / quantizer statistics see vision-conditioned
activations. The modality is inferred from `--calib_dataset_name`:

- an **image-text** dataset (default for VLMs,
`nemotron_vlm_dataset_v2`) drives the **full VLM forward**;
- a **text** dataset runs text-only calibration of the language model
(for text-vs-image ablations).

A shared `get_megatron_vlm_calibration_forward_loop` (built on
`megatron_prefill`) drives the full VLM forward over image-text pairs
from `vlm_dataset_utils` (`scienceqa`, `nemotron_vlm_dataset_v2`, with
config-driven subset/shard caps to bound downloads). It shards across
**data-parallel (DP)** ranks like the text loop (#1804); **context
parallelism (CP)** applies to text-only VLM calibration (the shared text
loop), not the multimodal forward — splitting the sequence would
misalign the merged vision embeddings.

### Results - Cosmos-Reason2-2B

Validated end-to-end on **Cosmos-Reason2-2B** (Qwen3-VL). Minitron NAS
prunes the language-model tower **1.72B → ~1.59B** (vision encoder +
projector frozen), top_k=1. Calibration data drives pruning importance;
image-text calibration runs the full VLM forward.

| Model | Calibration | MMLU | BLINK Rel-Depth | RealWorldQA |
|---|---|---|---|---|
| Baseline (1.72B) | — | 0.58 | 0.76 | 0.61 |
| Pruned (1.59B) | text (`nemotron-post-training-dataset-v2`) | 0.51\* |
~0.69 | ~0.57 |
| Pruned (1.59B) | image+text (`nemotron_vlm_dataset_v2`) | 0.49\* |
**0.77** | **0.61** |

\* Pruned MMLU on the 10% split (the pruning score function); baseline
MMLU is the full set. The VLM-benchmark numbers for the text row were
measured with a different text calibration set and are expected to be
similar for `nemotron-post-training-dataset-v2` (marked `~`).

> [!NOTE]
> These numbers come from short single runs on small eval splits — read
them for **high-level trends only**, not as exact values.

Takeaways: pruning the LM tower of a VLM works end-to-end. **Image-text
calibration** (this PR's feature) preserves the VLM benchmarks better
than text-only — BLINK Rel-Depth ~0.77 vs ~0.69 and RealWorldQA ~0.61 vs
~0.57, both close to the unpruned baseline (0.76 / 0.61) — which is the
motivation for calibrating on vision-conditioned activations.

### Results - Qwen3.5-9B

| Model                      | MMLU   | MMStar |
|----------------------------|:------:|:------:|
| Qwen3.5-9B      | 0.7003 | 0.6117 |
| Pruned-7B (text calib)        | 0.5527 | 0.4411 |
| Pruned-7B (image+text calib)  | 0.5107 | 0.3941 |

### Key changes

- `quantize.py`: quantizes the **root** model with non-LM (vision)
quantizers disabled, so the ModelOpt state lives on the root (required
by the Megatron save) while only the language model is quantized.
- `prune_minitron.py`: image-text (or text) calibration for VLM pruning
importance.
- Shared VLM calibration forward loop (`megatron_prefill`-based, unwraps
tuple outputs, DP-sharded) + `vlm_dataset_utils`.
- Tiny VLM test fixtures (Qwen3.5-VL, Gemma3-VL) with vision tokens
derived dynamically from the reference processor; VLM prune + quantize
example tests.
- README + CHANGELOG.

### Usage

```bash
# Prune the language model of a VLM (image-text calibration by default)
torchrun --nproc_per_node 2 prune_minitron.py \
    --pp_size 2 \
    --hf_model_name_or_path <vlm> \
    --prune_target_params 3e9 \
    --output_hf_path /tmp/vlm-pruned

# PTQ the language model of a VLM
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path <vlm> \
    --quant_cfg fp8 \
    --export_megatron_path /tmp/vlm-fp8-megatron
```

### Testing

- `test_prune_minitron.py::test_prune_minitron_vlm` — Gemma3-VL,
image-text (ScienceQA) calibration; full load → prune (depth + ffn) →
save → reload.
- `test_quantize_export.py::test_quantize_vlm` — Qwen3.5-VL, text
calibration; quantize LM → save Megatron checkpoint.
- LM regression tests (`test_prune_minitron`,
`test_quantize_and_export`) unchanged and passing.

### Not in scope

- **HF unified export of a quantized VLM** is not yet supported;
`export.py` saves the Megatron checkpoint only for VLMs (tracked by a
TODO in `export.py`). The recommended path is to route the megatron→HF
quant export through Megatron-Bridge's
`AutoBridge.export_hf_weights_quant(quantization_checker, quant_fn,
quant_block_size)`, which reuses the bridge's per-model mcore↔HF mapping
— covering Qwen3.5-VL / Gemma3-VL and the vision tower/projector (left
full precision) for free — so modelopt supplies only the checker +
pack/scale fn + `hf_quant_config` (KV-cache scales need a separate
path). This avoids re-authoring per-model mappings in modelopt (cf.
#1482's Qwen3-VL-only `mcore_qwen3vl.py`).

> [!NOTE]
> Qwen3.5-VL **MoE** is not tested e2e: the Megatron-Bridge weight
conversion expects packed (`gate_up_proj`) experts that transformers'
tiny checkpoint doesn't emit. MoE pruning itself is covered by
`test_mcore_qwen35_gdn_moe_pruning`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Follow-up to the GatedDeltaNet/MLA/latent-MoE pruning PR (#1747).
Rebased on `main` to pick up CP/DP calibration (#1804); the VLM
calibration loop now shards across DP ranks the same way. `hidden_size`
pruning for VLMs (requires resizing the vision projector) is left for a
future PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added VLM-aware Minitron pruning and post-training quantization that
target only the language-model portion, keeping the vision
tower/projector in full precision.
* Calibration now auto-selects text vs image-text datasets based on
model type, with modality validation.
* Expanded Megatron-Core CP/DP guidance and introduced a `--cp_size`
flag in quantization examples.
* **Bug Fixes**
* Improved VLM generation/prefill output handling and made vocabulary
sizing more robust for VLM wrappers.
* **Tests / Documentation**
* Updated pruning/quantization docs and refreshed/added VLM-focused
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 02:47:02 +05:30
ZhiyuandClaude Opus 4.8 f335459dc0 refactor(examples): rename llm_ptq → hf_ptq (symlink for back-compat) (#1759)
## What does this PR do?

**Type of change:** refactor / deprecation (examples)

Follow-up to #1705 (which consolidated `examples/vlm_ptq` into
`examples/llm_ptq`). Since that example now covers Hugging Face **LLM
and VLM** PTQ, the `llm_ptq` name is a misnomer. This renames the
directory to `examples/hf_ptq` and leaves a relative symlink
`examples/llm_ptq → hf_ptq` so existing paths/commands keep working
during a deprecation window.

Requested by @kevalmorabia97 on #1705 (with the symlink-for-back-compat
approach), targeted for the **same 0.46 release** as the consolidation.

### Changes
- `git mv examples/llm_ptq → examples/hf_ptq` and
`tests/examples/llm_ptq → tests/examples/hf_ptq` (the CI runner maps the
matrix name to both `examples/<name>` and `tests/examples/<name>`).
- Add a tracked back-compat symlink `examples/llm_ptq → hf_ptq`.
- Update CI matrices and all repo **path references** (docs, READMEs,
agent skills, launcher/debugger tools, tests) from `llm_ptq` to
`hf_ptq`.
- Keep Python identifiers / test-util module names
(`run_llm_ptq_command`, `llm_ptq_utils`) — they name the LLM-PTQ task,
not the directory.
- Preserve the CODEOWNERS team slug
(`modelopt-examples-llm_ptq-codeowners`) and historical CHANGELOG
entries; add a CHANGELOG deprecation note.

### Back-compat caveats (inherent to git directory symlinks)
- ✅ Linux/macOS CLI usage and Python `cwd`/pytest resolution work
through the symlink.
- ⚠️ Windows git checkouts don't materialize symlinks by default (low
impact — this example is Linux-only in practice).
- ⚠️ GitHub web doesn't follow directory symlinks, so legacy external
deep-links to `examples/llm_ptq/...` won't navigate in. All **internal**
references are repointed to `hf_ptq`, so the symlink is only for legacy
external/CLI use.

### Usage (unchanged via symlink)
```bash
# New canonical path
cd examples/hf_ptq
scripts/huggingface_example.sh --model <hf_model> --quant fp8

# Old path still works (forwards via symlink)
cd examples/llm_ptq && scripts/huggingface_example.sh --model <hf_model> --quant fp8
```

### Testing
- `bash -n` on moved/edited shell scripts (new path + via symlink).
- `py_compile` on moved/edited Python; test re-export shim repointed to
`examples/hf_ptq/example_utils`.
- Verified git tracks `examples/llm_ptq` as a single symlink (mode
120000), not a duplicated tree (no pre-commit / pytest
double-processing).
- `pre-commit run` on all changed files passes.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (relative symlink keeps
`examples/llm_ptq` paths valid; see caveats above)
- Did you write any new necessary tests?: N/A (pure rename; existing
tests moved with the dir)
- Did you update Changelog?: ✅

### Additional Information
Follow-up (later release): remove the `examples/llm_ptq` symlink once
external references have migrated.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* PTQ guidance now directs to the unified Hugging Face PTQ flow,
including VLM quantization via the shared `--vlm` entry point.
* **Documentation**
* Updated README and guide links, references, and command snippets to
use `hf_ptq` (replacing `llm_ptq`).
* Deprecated and consolidated `vlm_ptq` into `hf_ptq`; removed
VILA/NVILA coverage from the Hugging Face PTQ examples.
* **Bug Fixes**
* Improved detection and routing so local/manual setup uses the correct
PTQ source.
* **Tests / Chores**
  * CI and example tests updated to run the `hf_ptq` variants.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 08:48:48 +00:00
Keval MorabiaandClaude Opus 4.8 aa2a6a1b5d Add context-parallel (CP) and data-parallel (DP) support to Megatron calibration, and MMLU (#1804)
### What does this PR do?

Type of change: new feature

Adds **context-parallel (CP)** and **data-parallel (DP)** support to the
shared Megatron-Core inference/calibration utilities so PTQ calibration
and the MMLU sanity check work across these parallelisms (in addition to
the existing TP/PP/SP/EP).

**Context parallelism (CP):**
- **`megatron_calibration` / `megatron_mmlu`** — partition each sequence
across CP ranks (zigzag load-balanced, via `get_batch_on_this_cp_rank`).
MMLU gathers the per-rank logits back to the full sequence for
last-token scoring.
- **`megatron_prefill`** — accepts a CP-partitioned `position_ids`, and
under CP passes `attention_mask=None` so the CP-aware causal attention
builds the mask itself (a local triu mask would be wrong for the
per-rank zigzag chunks). Also wrapped in `torch.no_grad()` (pure
inference; lets MMLU run at larger batch sizes without retaining the
autograd graph).
- **`examples/megatron_bridge/quantize.py`** — new `--cp_size` flag.

**Data parallelism (DP):**
- **`get_dataset_dataloader`** — new `distributed` / `sampler_kwargs` to
shard the dataset across ranks with a `DistributedSampler`.
- **`megatron_calibration`** — shards calibration data across the DP
group (amax is max-reduced across DP inside `mtq` calibration, so the
per-rank shards combine correctly).
- **`megatron_mmlu`** — shards whole batches across DP ranks and
all-reduces the per-subject counts back to full-dataset accuracy.
- DP is **implicit**: `DP size = world_size / (tp * pp * cp)` —
launching with more GPUs than `tp * pp * cp` engages it. (No `--dp_size`
flag.)

RoPE is applied by the model per CP rank, so `position_ids` are only
needed for models with absolute/learned position embeddings.

### Usage

```bash
# Context parallelism = 2
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
    --cp_size 2 --export_megatron_path /tmp/Qwen3-8B-NVFP4-cp2

# Data parallelism = 2 (implicit: tp*pp*cp = 1, 2 GPUs)
torchrun --nproc_per_node 2 examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg nvfp4 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-dp2
```

### Testing

Added `cp` and `dp` cases to
`tests/gpu_megatron/torch/utils/plugins/test_utils_megatron.py::test_megatron_generate_and_mmlu`
(Qwen3-0.6B). All four parallelisms pass on 2 GPUs over the full
1430-example MMLU shard set (confirming the DP all-reduce reconstructs
the full count): tp=0.373, pp=0.375, cp=0.371, dp=0.375.

End-to-end PTQ on **Qwen3-8B → NVFP4**, comparing MMLU before and after
PTQ across parallelisms:

| Parallelism | MMLU before PTQ (bf16) | MMLU after PTQ (NVFP4) |
| :--- | :---: | :---: |
| TP=2 | 0.7294 | 0.7058 |
| CP=2 | 0.7292 | 0.7101 |
| DP=2 | 0.7292 | 0.7099 |

> **Common PTQ args used for the runs above:** `--hf_model_name_or_path
Qwen/Qwen3-8B --quant_cfg nvfp4 --seq_length 1024 --calib_num_samples
512 --calib_batch_size 16`, default calibration dataset
(`cnn_nemotron_v2_mix` = cnn_dailymail +
nemotron-post-training-dataset-v2 mix). MMLU evaluated at `fraction=1.0,
batch_size=16` on the full test set. Runs on 2× RTX 6000 Ada — TP=2 uses
`--tp_size 2`, CP=2 uses `--cp_size 2`, DP=2 uses all model-parallel
sizes = 1 (implicit DP over the 2 GPUs).

CP-, DP-, and TP-calibrated models all land within ~0.4% MMLU of each
other both before and after PTQ, confirming CP/DP calibration yields an
equivalently-quantized model.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all CP/DP logic is gated on
`cp_size > 1` / `dp_size > 1`; non-CP/DP (TP/PP/SP/EP) behavior is
unchanged. `megatron_prefill` gains an optional `position_ids` arg
(defaults to the previous behavior); `get_dataset_dataloader` gains
optional `distributed`/`sampler_kwargs` (default off).
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: ✅ — `cp` and `dp`
parametrizations added to the existing Megatron generate/MMLU test.
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅ (run `/claude review`)

### Additional Information

`torch.no_grad()` on `megatron_prefill` is a shared change (also
benefits the calibration / generate / PEFT-test callers) — pure
inference, so no behavioral change beyond lower memory.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes
* **New Features**
* Added context-parallel (CP) and data-parallel (DP) support across
shared inference, calibration, and MMLU evaluation.
* Introduced CP-aware prefill with per-rank logits gathering for correct
last-token scoring.
* Added `--cp_size` to the quantization example (DP is derived
automatically).
* **Improvements**
* Extended dataset dataloader utilities with optional distributed
sampling controls.
* **Bug Fixes**
* Calibration and evaluation now work when CP is enabled (no longer
restricted to CP=1).
* **Tests**
  * Expanded Megatron generate/MMLU coverage to include CP and DP modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 22:39:22 +05:30
Keval MorabiaandClaude Opus 4.8 55a2101e2f Update Nemotron-3 Pruning, Distillation and PTQ results based on new shared calibration loop with seq packing and add tool-calling eval fix (#1660)
### What does this PR do?

Type of change: documentation + minor example-script tweaks

Follow-up to #1601. Originally scoped to add **NVFP4 + QAD**, this PR
was **repurposed** to refresh the [Nemotron-3-Nano-30B-A3B-BF16
tutorial](examples/megatron_bridge/tutorials/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/README.md)
results using the **new shared calibration loop (sequence packing)** and
to **fix tool calling in evaluation**.

- Refreshed the prune → distill → eval → **FP8** results (accuracy +
vLLM throughput tables) with the new calibration loop.
- **Tool-calling eval fix** (`nemo_evaluator.yaml`): GPQA and AIME now
run the Python sandbox tool. The tutorial reports both **with-tools**
and **no-tools** GPQA/AIME and shows `mean ± std_dev`.
- Script tweaks: `quantize.py` calibration now uses sequence packing
(`pack=True`) which leads to slight improvement in PTQ;
`prune_minitron.py` defaults `inference_batch_size` to
`calib_batch_size`.

### Testing

Documentation + small example-script changes; tutorial relative links
resolve and the results tables / figure were verified consistent.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: Yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the main guide and evaluator instructions for prune + distill
+ FP8/NVFP4 quantization, including refreshed vLLM deployment tips,
benchmark/noise presentation, and long-context tool-calling attribution
notes.
* Refreshed README technique examples/links, reordered the model support
matrix rows, and improved pruning overview/support-matrix text.

* **Changes to Examples**
* NAS pruning now documents higher GPU memory usage vs manual pruning;
pruning batching defaults were improved.
* Quantization PTQ calibration uses packed document packing; quantized
checkpoint export messaging was streamlined.
* Updated pruning/distillation/quantization tutorial guidance,
metrics/tables, command parameters, and evaluator YAML settings
(KV-cache dtype, generation defaults, task behavior).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 18:32:13 +00:00
Keval MorabiaandClaude Opus 4.8 54ce4e09d8 Add Quantization Aware Distillation (QAD) to Megatron-Bridge example (#1600)
### What does this PR do?

Type of change: new example

**Note:** This is **part 2 of 4** (builds on #1589):

- **Part 1 (#1589):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
- **Part 2 (this PR):** extend `distill.py` for quantization-aware
distillation (QAD) — load a quantized Megatron checkpoint as the
student.
- **Part 3:** https://github.com/NVIDIA/Model-Optimizer/pull/1601
- **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

Extends `examples/megatron_bridge/distill.py` to initialize the student
from a **Megatron checkpoint** (a quantized checkpoint from
`quantize.py`, or a pruned one) via `--student_megatron_path`, enabling
**Quantization Aware Distillation (QAD)**:

- `--student_hf_path` still builds the student architecture;
`--student_megatron_path` supplies the (optionally quantized) weights.
- For a quantized checkpoint, the ModelOpt quantize mode + base weights
are restored onto the **plain student before the knowledge-distillation
conversion** (`restore_sharded_modelopt_state` is a no-op once a model
is already converted), so the distilled checkpoint stays exportable as a
quantized model with `export.py`.

**Upstream dependency / workaround:** `DistillationProvider.provide()`
has no seam to transform the student before the KD conversion, so this
patches `provide()` at the class level (via an `id()`-keyed registry,
because the provider proxies instance-attribute assignment to its
teacher once the teacher is set). A companion Megatron-Bridge PR adds a
first-class `DistillationProvider.student_pre_conversion_hook`; from
nemo:26.06 onwards the workaround should be removed and replaced with
that hook (a removal note in `distill.py` documents exactly how).

### Usage

```bash
# 1) PTQ -> quantized Megatron checkpoint (part 1)
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B --quant_cfg fp8 --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# 2) QAD: distill the quantized student from the unquantized teacher
torchrun --nproc_per_node 8 distill.py \
    --teacher_hf_path Qwen/Qwen3-8B \
    --student_hf_path Qwen/Qwen3-8B \
    --student_megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --data_paths 1.0 tokenized/data_text_document \
    --train_iters 1000 --output_dir /output/qwen3_8b_qad

# 3) export the distilled quantized checkpoint (part 1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /output/qwen3_8b_qad/checkpoints \
    --export_unified_hf_path /tmp/qwen3_8b_qad_fp8_hf
```

### Testing

`tests/examples/megatron_bridge/test_qad.py` (validated on a 2-GPU NeMo
`26.04` container): quantize a tiny Qwen3 at TP=2 → QAD distill from the
quantized student → `export.py` to a unified HF checkpoint, asserting
`hf_quant_config.json` is written (proves the quantize mode survived
QAD). Includes a commented-out vLLM deployment check, validated locally
(full flow passes; vLLM loads the export as `quantization=modelopt`).
Existing normal/Puzzletron distillation tests still pass.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example feature; default
behavior unchanged when `--student_megatron_path` is not set)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Depends on a companion Megatron-Bridge PR adding
`DistillationProvider.student_pre_conversion_hook` (the upstream
replacement for the class-level `provide()` workaround). The Nemotron-3
tutorial NVFP4 + QAD experiments ship in part 3.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization Aware Distillation (QAD) workflow to recover accuracy of
quantized Megatron students and distill from quantized checkpoints.
* CLI option to initialize a distillation student from a Megatron
checkpoint and a structure-only load path for bridging.

* **Documentation**
* Expanded runnable quantize → QAD → export guidance and best-practice
tips.

* **Tests**
  * End-to-end test validating quantize → QAD → export artifacts.

* **Chores / UX**
* Clearer rank-aware messages, improved tokenizer padding handling, and
more consistent export behavior (fixed export dtype).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 21:28:25 +00:00
Shengliang Xu dbdff11a7f Drive PTQ example qformat choices from preset YAMLs (hf_ptq, multinode_ptq, megatron_bridge) (#1525)
### What does this PR do?

Type of change: Refactor

Replace the hardcoded `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` dicts
in the PTQ
example scripts with a small `_load_preset_cfg_choices()` helper that
discovers the
available qformat names by listing
`modelopt_recipes/configs/ptq/presets/{model,kv}/`
and **eagerly loads every preset YAML into a plain dict at import** via
the existing
`load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` path.
The directory listing becomes the source of truth for the `--qformat` /
`--kv_cache_qformat` CLI vocabulary.

> Note: an earlier revision used a lazy, copy-on-access `Mapping`. That
was overkill
> for these example scripts — the previous `mtq.*_CFG` module constants
were
> themselves eagerly-loaded shared dicts, and every call site that
mutates a config
> already deepcopies first — so it is now a plain eager dict. A lazy
variant can be
> reintroduced later if import time ever matters.

**Scope.** Three scripts carried the same hardcoded tables and all three
are migrated:

- `examples/llm_ptq/hf_ptq.py` — `--qformat` / `--kv_cache_qformat`.
- `examples/llm_ptq/multinode_ptq.py` — `--qformat` /
`--kv_cache_qformat`.
- `examples/megatron_bridge/quantize.py` — `--quant_cfg` /
`--kv_cache_quant`
  (still also accepts any full `mtq.config.choices` name).

All three scripts share the discovery helper (`load_quant_cfg_choices`),
the canonical
alias table (`QFORMAT_ALIASES`), the KV disable sentinel
(`KV_CACHE_NONE`), and the
ready-built `QUANT_CFG_CHOICES` / `KV_QUANT_CFG_CHOICES` mappings via
the new
**`modelopt.recipe.presets`** module — no copy lives in the example
scripts anymore. The
module is a standalone import, so `import modelopt.recipe` stays cheap;
only an explicit
`import modelopt.recipe.presets` triggers the eager preset load.

A small alias table preserves previously-supported short CLI names
(`int8_sq`,
`nvfp4_awq`, `fp8_pb_wo`, …, plus the Megatron-Bridge `fp8_blockwise`)
as deprecation
shims. It is documented as not-for-extension — new formats land as
preset YAMLs, and
longer term, configurations should be authored as full recipes
(`--recipe`). The alias
logic is fail-fast: an alias pointing at a missing preset raises
`ValueError` at import.

Also adds `presets/kv/fp8_cast.yaml` and `presets/kv/nvfp4_cast.yaml`,
composed from the
existing `kv_fp8_cast` / `kv_nvfp4_cast` unit fragments. This promotes
`fp8_cast` /
`nvfp4_cast` to first-class KV presets and lets us delete the runtime
`_set_kv_cache_constant_amax` helper and all its call sites —
`use_constant_amax` is now
authoritative in the YAML. The KV-calibration-skip decision is derived
from the config
(`_kv_cfg_uses_constant_amax`), not from hardcoded format names.

**⚠️ CLI surface expansion (owner sign-off requested).** Because the
directory listing
is now the CLI vocabulary, each script accepts **every** preset under
`presets/{model,kv}/`, not just its previously curated subset. For
`hf_ptq.py` this is
the same surface the prior table covered; for `multinode_ptq.py` and the
Megatron-Bridge
script it is broader (e.g. KV `fp8_affine` / `fp8_cast` / `nvfp4_cast` /
`nvfp4_rotate`
are now selectable). This is intended ("the directory is the policy"),
but please confirm
those two scripts are meant to expose all presets — if a given path has
not validated a
format, it should be gated explicitly.

### Usage

```bash
# Old short names still work via the alias shim
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_sq --kv_cache_qformat fp8_cast --export_path out/

# Canonical preset basenames work directly
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat int8_smoothquant --kv_cache_qformat fp8_cast --export_path out/

# A newly-added preset YAML is valid on the CLI of all three scripts with no code change
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path <model> --qformat nvfp4_awq_full --export_path out/
```

### Testing

- New: `tests/unit/recipe/test_presets.py` smoke tests for
`modelopt.recipe.presets` —
every discovered model/KV preset loads into a `quant_cfg` dict, the
directory listing is
fully covered, deprecation aliases resolve to their canonical preset,
the KV `none`
sentinel does not collide with a preset, and a stale alias raises. These
guard the eager
import-time load (one bad preset would otherwise break `import
modelopt.recipe.presets`
  and every PTQ example).
- Previously verified locally (uv `.venv` py3.13 + `dev-py310-modelopt`
conda):
all previously-supported `--qformat` / `--kv_cache_qformat` names
resolve to dicts
bit-equal to the corresponding `mtq.*_CFG` constants; `fp8_cast` /
`nvfp4_cast` carry
`use_constant_amax: true` while non-cast variants do not; argparse
accepts
`--kv_cache_qformat none` plus all variants; unknown qformats raise at
lookup / argparse.
- All pre-commit hooks pass (ruff, mypy, bandit, license, rst, yaml).
- Pre-merge manual checks recommended by review (run in an env with the
deps installed):
`python examples/llm_ptq/hf_ptq.py --help`, `… multinode_ptq.py --help`,
  `… megatron_bridge/quantize.py --help`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — all previously-valid CLI
values continue to work via the alias table; output configs are
bit-equivalent to the prior hardcoded path.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
deps.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_ptq/test_example_utils.py` preset-discovery smoke
tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ☐

### Additional Information

Out of scope / follow-up: the `_AUTO_QUANTIZE_QFORMATS` table and
`_canonical_qformat`
helper in `hf_ptq.py` are intentionally left hardcoded — auto_quantize
is being
refactored/reimplemented and they are expected to be removed soon.

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-06-05 12:44:16 -07:00
Keval MorabiaandClaude Opus 4.8 f21977a5fc Add Megatron-Bridge PTQ quantize + export example scripts (#1589)
### What does this PR do?

Type of change: new example

Adds a two-step post-training quantization (PTQ) flow for
**Megatron-Bridge** models under `examples/megatron_bridge/`, mirroring
the Megatron-LM `quantize.sh` / `export.sh` split:

- **`quantize.py`** — loads an HF model via Megatron-Bridge, applies
ModelOpt PTQ (via a `--quant_cfg` alias / full config name, or a
`--recipe` YAML), with optional KV-cache quant, weight-only,
compression, and MoE expert-ratio calibration, then saves a **Megatron
checkpoint** (with ModelOpt state). Tensor / pipeline / expert
parallelism are all supported, and the checkpoint can later be reloaded
for further training (QAT / distillation).
- **`export.py`** — loads the quantized Megatron checkpoint, **re-shards
to TP=1**, and exports a **HuggingFace (unified)** checkpoint deployable
with TensorRT-LLM / vLLM / SGLang.

**Why the split?** The unified HF exporter (`export_mcore_gpt_to_hf`)
does not gather tensor-parallel-sharded weights — Megatron-LM likewise
forces `TP=1` during its export step. Saving a TP-sharded Megatron
checkpoint first lets us calibrate at TP>1 (to fit large models) and
then reload re-sharded to TP=1 for the HF export. A combined
single-script flow silently produced corrupt HF checkpoints under TP>1
(collided per-rank shards), which this split avoids.

> **Note:** This is **part 1 of 4**:
> - **Part 1 (this PR):** Megatron-Bridge `quantize.py` + `export.py`
support and tests.
> - **Part 2:** extend `distill.py` for quantization-aware distillation
(QAD) — load a quantized Megatron checkpoint as the student.
> - **Part 3:** add NVFP4 + QAD-on-pruned-checkpoint experiments to the
Nemotron-3-Nano-30B-A3B tutorial.
> - **Part 4:** repeat the NVFP4 + QAD experiments on a non-Nemotron
model.

### Usage

```bash
# Step 1: quantize (TP/PP/EP supported) -> Megatron checkpoint
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --quant_cfg fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-FP8-megatron

# Step 2: export -> deployable HuggingFace (unified) checkpoint (re-shards to TP=1)
torchrun --nproc_per_node 1 export.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --megatron_path /tmp/Qwen3-8B-FP8-megatron \
    --export_unified_hf_path /tmp/Qwen3-8B-FP8-hf
```

### Testing

`tests/examples/megatron_bridge/test_quantize.py` (validated on a 2-GPU
NeMo `26.04` container):

- `test_quantize_export_and_vllm_deployment` — quantize a tiny Qwen3 via
a recipe at TP=2 → `export.py` re-shards to TP=1 → load + generate with
**vLLM** (skipped if vLLM absent).
- `test_quantize_megatron_checkpoint_reload` — quantize at TP=2 → reload
the Megatron checkpoint via the bridge and assert ModelOpt quantizers
were restored.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A (new example)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependencies)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

The Nemotron-3 tutorial update to use these scripts is intentionally
**not** included here — it ships with the part 3 PR alongside the NVFP4
+ QAD experiments.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded post-training quantization (PTQ) workflow documentation with
detailed step-by-step examples and configuration guidance for the
Megatron-Bridge framework.

* **New Features**
* Added quantization tool for applying PTQ to Megatron models with
calibration support.
* Added export tool for converting quantized models to a deployable
format.

* **Tests**
* Added integration tests validating the complete
quantization-export-deployment workflow, including inference validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-02 19:36:15 +00:00