Commit Graph
1251 Commits
Author SHA1 Message Date
Chenjie LuoandClaude Opus 5 b16356776b Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do?

Type of change: Code refactoring

`#2448` added the GGML IQ packing kernels as **two** torch extensions,
`modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges
them into one,
`modelopt_cuda_ext_ggml`.

The existing per-extension split in `extensions.py` exists for reasons
that don't apply
to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while
`_fp8`/`_mx` gate on
`>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the
base `tensor_quant`
kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none
of that —
same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split
only compiled the
shared header twice, ran nvcc twice, and grew the loader, `__getattr__`,
and
`precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly
following, that
scales badly.

Changes:

- New `ggml/ggml.cpp` holds both host-side validation wrappers and the
single
`PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously
each module
exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`;
the
  validation logic and docstrings carry over unchanged.
- `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`,
which builds
`ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The
retry-on-`raise_if_failed`
  semantics of the old getters are preserved.
- Each format keeps its kernels in its own translation unit, so adding a
format is a new
`.cu` plus one `module.def` — no new extension, loader, or
`precompile()` line.

No caller outside `extensions.py` and its tests referenced the old
getters on `main`, so
nothing else changes. **Note for the follow-up PRs in the `#2448` series
(`#2446`/`#2447`/`#2449`): the codec layer should call
`get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of
`get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.**

### Usage

```python
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml

ext = get_cuda_ext_ggml(raise_if_failed=True)
iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid)            # uint8 [numel / 256, 50]
iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74]
```

### Testing

Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000`
container), building the
merged extension from scratch:

- `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24
passed** (6:44). This
is the full existing IQ suite (zero-block layout, encode, dtype
rejection,
row-straddling rejection, invalid/negative-zero scales, byte-exact dtype
equivalence,
and the brute-force optimality round-trip) reparametrized onto the
merged module,
  plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests.
- Verified `precompile()` loads all four extensions and that the merged
module exports
  exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities.
- Off-GPU: compiled the three sources directly and linked them into one
`.so` to confirm
no duplicate-symbol collisions between the two `.cu` translation units.
- `pre-commit run --files ...` passes on all changed files (ruff, mypy,
clang-format,
  bandit, license headers).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the removed getters were
added in `#2448`
(merged today, unreleased) and have no callers outside this file's own
tests.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; the moved wrappers keep their original
attribution.
- Did you write any new necessary tests?: ✅ — existing coverage
reparametrized onto the merged module; no behavior change to test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal refactor of an unreleased, not-yet-wired-up API.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Follow-up to #2448. Merge before the remaining PRs in that series
(#2446, #2447, #2449)
land, so the codec layer is written against `get_cuda_ext_ggml` and no
rename is needed
afterwards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
  - Added IQ1_S packing support through the GGML CUDA extension.
  - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing.
- Improved extension loading reliability when a cached extension is
unavailable.

- **Changes**
  - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`.
- Consolidated IQ1_S and IQ2_XS extension access under the shared GGML
loader.
- Updated GPU validation and coverage to use the unified extension
interface.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 20:46:41 +00:00
Chenjie Luo 542012d4d5 Update CODEOWNERS (#2463)
Add modelopt-torch-kernels-codeowners



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added code ownership coverage for the `modelopt/torch/kernels`
directory.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-09-17 20:08:12 +00:00
yeyu-nvidiaandClaude Opus 5 ad1bad7817 fix(specdec): gather sharded hidden states in DFlash/DSpark AR generation (#2458)
### What does this PR do?

Type of change: Bug fix

Fixes the nightly
`tests/regression/torch/speculative/test_dflash.py::test_dflash_ar_validate`,
red every night since 2026-09-10 with:

```
WARNING: sample 0 (writing) failed: Expected all tensors to be on the same device,
but got tensors is on cuda:1, different from other tensors on cuda:0
(when checking argument in method wrapper_CUDA_cat)
RuntimeError: AR validation produced no results: all 3 samples failed.
```

**No PR caused this.** `ar_validate.py` loads with `device_map="auto"`,
and the regression runner (`…rtxpro6000-l-2-…`) has 2 GPUs, so even
Qwen3-0.6B gets sharded. `pseudo_speculative_generate()` then
concatenates hidden states from `target_layer_ids`, which spans early
*and* late layers — different devices once the base model is split:

```python
selected = [base_outputs.hidden_states[lid + hid_offset] for lid in self.target_layer_ids]
target_hidden = torch.cat(selected, dim=-1)   # cuda:0 ++ cuda:1 -> RuntimeError
```

The block tensors have the same problem from the other side: they follow
`input_ids.device`, while the draft module sits on the *last* base
layer's device (`_place_draft`).

What changed on 2026-09-10 was detection, not behaviour. #2288 made
`ar_validate.py` raise when every sample fails; before it, the report
block was guarded by `if results and …`, so zero results printed nothing
and exited **0**. The test had been green while validating nothing. Its
runtime is the tell: 23–26s in every green nightly, 21.6s in the first
red one, where it does no validation work at all.

The same defect is in #2288's own description, on an 8-GPU sharded
**EAGLE3** checkpoint: 80/80 samples dead on the identical error, job
exiting 0. So this is not DFlash-specific in principle — but EAGLE3's
generate path gathers separately (`pop_and_gather_aux_hiddens`) and
needs its own look, which is why this PR stops at the DFlash family.

### Fix

Gather everything the draft consumes onto the draft's device, and hand
results back on the caller's device since `validate_online` cats them
onto the running sequence. Every `.to()` is a no-op on a single device,
so single-GPU behaviour is bit-identical.

`HFDominoModel` inherits `HFDFlashModel.pseudo_speculative_generate`, so
it is fixed by the same change; `HFDSparkModel` overrides it and is
patched in parallel (including `prev_token` feeding `markov_step`).

### Usage

No API change.

```bash
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <dflash ckpt> --osl 10 --num_samples 3 --steps 7
```

### Testing

- `tests/unit/torch/speculative/plugins/test_hf_dflash.py` and
`test_hf_dspark.py`: **110 passed, 1 skipped**.
- The skip is the new `TestShardedBaseGeneration`, which is gated on
`torch.cuda.device_count() >= 2`. It stands in for accelerate's dispatch
— the base forward stays whole, but its hidden states are spread across
two devices the way a real `device_map="auto"` split spreads them — and
asserts both returned tensors come back on the caller's device. **I have
no 2-GPU box to hand, so this test is unverified locally; it will first
execute in CI.** It reproduces the reported error against the unpatched
code by construction, but please treat that claim as unconfirmed until
the run goes green.
- The real verification is the next nightly: `test_dflash_ar_validate`
should pass *and* print an AR number, which it has not done in any log I
can find.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — device moves only; no-ops
when the base model is on one device.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one multi-GPU regression
test (skipped without 2 GPUs).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — bug in an unreleased-cycle path that only ever produced a silent
no-op; no user-visible behaviour to describe.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Worth a maintainer view: `device_map="auto"` is a poor default for AR
validation, which is a single-process step-by-step loop. Every sharded
run of it that I can find has failed. Pinning it to one device would be
a smaller blast radius than making every drafting path device-safe — but
it would also cap validation at models that fit on one GPU, so I have
not changed it here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved speculative generation with sharded Hugging Face models
across multiple GPUs.
* Ensured generated draft tokens and outputs are returned on the
caller’s input device.
* Preserved Markov sampling behavior while improving device
coordination.

* **Tests**
* Added multi-GPU coverage for sharded model generation, including
output placement and draft shape validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:28:57 -07:00
yeyu-nvidiaandClaude Opus 5 216f28a6e0 Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)
### What does this PR do?

Type of change: documentation

Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.

**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:

```
skills/eagle3-new-model/    skills/eagle3-review-logs/
skills/eagle3-triage/       skills/eagle3-validate/
```

Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.

**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:

```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md                      # router: stage table x algorithm table
└── references/
    ├── stages/                   # the procedure — algorithm-independent
    │   ├── configure.md          # <- eagle3-new-model
    │   ├── review-logs.md        # <- eagle3-review-logs
    │   ├── triage.md             # <- eagle3-triage
    │   └── validate.md           # <- eagle3-validate
    └── algorithms/               # the data sheet — per-algorithm
        ├── README.md             # contract: 6 required sections
        ├── eagle3.md
        ├── dflash.md
        ├── dspark.md             # DFlash variant — delta only
        └── domino.md             # DFlash variant — delta only
```

Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.

DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.

Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:

- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.

**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.

### Usage

```
/speculative-decoding
```

Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:

```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun   # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes      # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```

### Testing

- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`

### Additional Information

Follows #2025, which moved the skill tree into the installable plugin.

Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):

1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.

Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.

* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:09:14 -07:00
2ff2e1bc80 [OMNIML-5899] Add CUDA kernels for IQ packing (#2448)
## Summary

- add native CUDA packing kernels for IQ1_S and IQ2_XS
- load both extensions through the quantization extension module
- validate caller metadata and launch bounds before contiguous
materialization
- normalize non-finite input elements consistently with the Python
reference path
- accept caller-computed IQ2_XS FP16 superblock scales to avoid a
duplicate reduction
- share common packing helpers and add direct extension compilation and
boundary tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only native kernel sources, shared packing helpers,
extension loading and build registration, and direct extension tests. It
does not contain Python codecs, export code, or recipes.

## GPU test coverage

Direct kernel-boundary coverage is included in this PR:

- [extension compilation, zero payloads, input validation, and
row-alignment
checks](https://github.com/NVIDIA/Model-Optimizer/blob/74e94db9601870e3569c7c8e73506a2f08c29da8/tests/gpu/_extensions/test_torch_extensions.py)

Pack/dequantize numerical, native/reference byte-parity, and
non-finite-policy tests are owned by the quantization PR:
[IQ1_S](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py)
and
[IQ2_XS](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py).

## Dependency behavior

On `main`, this PR provides optional CUDA extension loaders and direct
extension tests. The Python encoders and fallback dispatch land in
#2446. Until #2446 lands, no quantization path calls these getters, so a
load failure reports only that the extension is unavailable.

The IQ2_XS packer accepts one caller-computed FP16 scale per 256-value
block. #2446 owns that predictor and passes the same values to the
native and reference encoders.

## Provenance

- The CUDA kernels were independently written.
- They implement the packed-format contract and sign-parity convention
from the pinned [llama.cpp
definition](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h).
- `16.875 = 15 × (1 + 1/8)` is a derived IQ1_S constant.
- `0.125` is part of the encoded format.
- `0.61` is our empirical IQ1_S scale predictor, not copied from
upstream code.
- The IQ2_XS predictor constants are owned by #2446 and are not
duplicated in this kernel.

Human review is still required to confirm that the attribution and
license treatment are sufficient.

## Validation

- repository hooks, including native formatting, pass for all changed
files
- extension loader and test modules compile as Python
- all 20 direct-extension and CUDA integration test cases collect
locally
- CUDA runtime execution is delegated to GPU CI because the local host
is macOS
- restricted-term scan passes

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 17:56:19 +00:00
835c041c58 fix(specdec): resolve the eagle aux-layer preset in the vLLM hidden-state dump (#2410)
### What does this PR do?

Type of change: Bug fix

Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel.

The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the
flag's own default** — so the documented invocation aborted before
writing any state:

```
File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone
    ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()})
ValueError: invalid literal for int() with base 10: 'eagle'
```

**Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM
container, where importing `modelopt.torch` fails (the full init chain
pulls in omegaconf and friends). It therefore carries
`_resolve_aux_layers_standalone`, a local copy of the preset logic in
`common.resolve_aux_layers`. That copy implemented the `dflash` preset
and explicit id lists, but never `eagle` — while `add_aux_layers_args`
defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and
were unaffected; only the vLLM path forked, and nothing compared the
fork against its source.

This PR resolves `eagle` inline, mirroring
`hf_eagle.default_eagle_aux_layer_ids`.

It also fixes a second defect the bug exposes: the function already had
a message naming the accepted values, but it was unreachable, because
`int()` raised first. An unrecognised preset now reports what it accepts
instead of surfacing the raw `int()` error — which is what made the
original failure opaque.

### Usage

The previously-broken documented invocation now works:

```bash
cd examples/speculative_decoding
python collect_hidden_states/compute_hidden_states_vllm.py \
    --model Qwen/Qwen2.5-0.5B-Instruct \
    --input-data ../dataset/synthetic_conversations_1k.jsonl \
    --output-dir /tmp/hs_vllm \
    --max-seq-len 512 --tp 1
```

`--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8`
are unchanged.

### Testing

Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which
pins the standalone copy to the shared implementation it mirrors:

- `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer
counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately
including counts small enough that the `max(0, ...)` clamps collapse ids
together.
- A named regression case for `nvbugs/6753684`.
- `dflash` and explicit-list behaviour unchanged.
- Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the
actionable message.
- Out-of-range ids still rejected.

Divergence here is silent — the dump would write plausible-looking
hidden states from the *wrong* layers, surfacing much later as a poor
acceptance rate. Hence pinning to the reference rather than asserting
hardcoded lists alone.

All 20 assertions verified and every pre-commit hook passes (`ruff`,
`mypy`, `bandit`, RST lint, license headers).

One caveat worth stating plainly: **pytest could not be run locally.**
`tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`,
which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in
this machine's torch. Each assertion was executed directly against the
real module instead, but CI is the first genuine pytest run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — strictly widens accepted
input; `dflash` and explicit lists behave identically.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — bug fix for a defect present in a previous release.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

The underlying fragility is the duplicated implementation, not this one
missing branch. The function's own `TODO: drop this once
common.resolve_aux_layers is decoupled from the heavy modelopt.torch
import chain` is the real fix; the new test narrows the gap but does not
close it. Worth tracking separately if the vLLM dump is expected to keep
pace with new presets.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection.
* Added support for the documented `eagle` preset alongside `dflash` and
explicit layer IDs.
  * Improved invalid-option errors to clearly list accepted formats.
* Rejects `dflash` configurations when the target model has too few
layers.
  * Continues rejecting layer IDs outside the model’s available range.
* **Documentation**
  * Added a v0.48.0 changelog entry for the fix.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-17 17:54:30 +00:00
Wei-Ming Chen f377b77116 Fail fast on non-finite AutoQuantize output gradients (#2432)
### What does this PR do?

Type of change: Bug fix

Fail fast when AutoQuantize receives non-finite output gradients, before
accumulating sensitivity scores. The error names the affected module and
suggests checking the model, data, and loss; it also identifies cuDNN
SDPA backward on fully masked rows as one possible cause and gives an
explicit retry workaround.

Unlike the earlier revision, this does not disable cuDNN or change any
attention backend settings. Invalid gradients are not zeroed or ignored,
and there is no automatic retry.

### Usage

No API or recipe changes. For the reproduced cuDNN failure, the caller
can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before
a fresh AutoQuantize run.

### Testing

- AutoQuantize unit suite: **110 passed**. Coverage includes NaN and
positive/negative infinity, module diagnostics, preventing invalid score
accumulation, model-state cleanup, and unchanged SDPA backend settings.
- Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main
`8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size
8, 512 calibration samples, and
`w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`:
- Default backend: the new diagnostic fired at
`model.language_model.layers.39.self_attn.q_proj`; cuDNN remained
enabled and no quantized model was exported. Expected-error check
passed.
- Explicit cuDNN-disabled fresh run: both 64-batch calibration passes,
all 16 scoring batches, optimization at **5.99 effective bits**, and
checkpoint export completed with exit code 0. Verified all three indexed
safetensors shards and quantization configuration; the index contains
93,563 tensors.
- All applicable pre-commit checks and `git diff --check` passed.

### Before your PR is "*Ready for review*"

Contributor guidelines and security guidance followed; commits are
signed and signed off.

- Is this change backward compatible?: Yes; finite-gradient behavior and
backend settings are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither
added.
- Did you write any new necessary tests?: Yes.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes, 0.48.0 bug fixes.
- Did you get Claude approval on this PR?: No; awaiting review.

### Additional Information

This improves error reporting rather than fixing the upstream cuDNN
kernel. Blackwell-specificity is not established. Checkpoint
deployment/reload was not tested.

A calibration-only checkpoint-resume attempt completed scoring but
encountered a separate `candidate_stats` KeyError. That issue is outside
this patch; the successful export validation above used a fresh run
without search-state resume.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* AutoQuantize now fails fast when output gradients contain non-finite
values, with an actionable error identifying the affected module.
* Attention backend settings are preserved and restored after successful
runs and failures.
* Model state is restored when sensitivity scoring encounters an error
during setup or execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-17 17:52:46 +00:00
realAsma b9cfdce8dc docs: add Local Hessian NVFP4 weight-scale announcement blog (#2417)
### What does this PR do?

Type of change: documentation

Adds a Local Hessian announcement blog at
`docs/source/announcements/local-hessian.rst`, covering the NVFP4
per-block
weight-scale rule that minimizes output error instead of weight error.

Contents:

- Derivation of the per-block output-error objective and its `16x16`
local
  Hessian, with numbered equations.
- Results on Qwen3.5-9B: scale-setting comparison against max, MSE, and
  Four-over-six, plus composition with GPTQ.
- Figure 1, a grouped bar chart of the Qwen3.8-27B W4A4 candidate scores
  (BF16 in gray, the two scale rules in NVIDIA greens).
- A "Using Local Hessian" section with the config example and the
  end-to-end `hf_ptq.py` command.

Two supporting changes outside the blog:

- `docs/source/_static/announcements.css`: the `shibuya` theme has no
`span.eqno` rule, so Sphinx's default `float: right` on equation numbers
cannot share a line with MathJax's full-width display block and the
number
renders *above* the equation. This anchors it to the right of the
equation
  instead, and shrinks the table-note class.
-
`docs/source/announcements/assets/qwen3-27b-w4a4-scale-rule-accuracy.png`:
  the Figure 1 asset.

### Usage

```python
import modelopt.torch.quantization as mtq

config = {
    "quant_cfg": [...],  # quantizer configuration
    "algorithm": {"method": "local_hessian", "fp8_scale_sweep": True},
}

model = mtq.quantize(model, config, forward_loop)
```

### Testing

Documentation only; no code paths change. The `.rst` parses cleanly
under
docutils. The rendered page has not been checked with a full
`sphinx-build`,
so the equation-number CSS fix and the figure placement are worth an
eyeball
on the built docs before merge.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Two items to settle before this is ready to publish:

1. The `--recipe` example points at

`modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_local_hessian-fp8_attn-kv_fp8_cast.yaml`,
a placeholder path derived from the existing recipe naming convention.
It
   needs to match whatever lands in #2363.
2. The tables report single-run team measurements; the blog says so and
makes
   no significance claims.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added guidance on NVFP4 Local-Hessian weight-scale selection,
including mathematical details, accuracy comparisons, runtime
considerations, limitations, configuration examples, and reproduction
steps.
- Updated announcement labels, headings, metadata, descriptions, and
filtering text to use “Local-Hessian.”

- **Style**
- Improved announcement formatting for equation labels, display-equation
spacing, Hessian results, table headers, and explanatory notes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-16 21:48:11 +00:00
Keval MorabiaandClaude Opus 5 a448ba9757 Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 02:10:12 +05:30
Chad Voegele f68bf83cfa Add published MiniMax M2.7 NVFP4 PTQ recipe (#2437)
Chad's Agent

### What does this PR do?

Type of change: new example.

Backfill the recipe for
[nvidia/MiniMax-M2.7-NVFP4](https://huggingface.co/nvidia/MiniMax-M2.7-NVFP4):
max-calibrated NVFP4 W4A4 experts (block size 16), BF16
attention/router/LM head, and FP8 KV-cache cast mode.

The published metadata identifies ModelOpt
`0.43.0rc2.dev96+g3baa2da62.d20260410`, matching the original run. At
that revision, `nvfp4_experts_only` used max calibration and the CLI
defaulted to `fp8_cast`. The upload record links that run's checkpoint
to the published model.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path /path/to/MiniMax-M2.7 \
  --recipe models/MiniMaxAI/MiniMax-M2.7/ptq/nvfp4_experts_only-kv_fp8_cast \
  --moe_calib_experts_ratio 1.0 \
  --export_path /path/to/MiniMax-M2.7-NVFP4
```

### Testing

- Loaded the recipe and checked its resolved patterns/numerics against
the published `config.json`: 47,927 Linear module names and 124 KV
quantizers across all 62 layers.
- Confirmed `hf_quant_config.json` converts exactly to the published
`config.json` quantization config using ModelOpt's exporter conversion.
- Pre-commit checks and `git diff --check` passed.
- No unit tests or README added; full-model PTQ was not rerun.

The published config confirms formats and exclusions, not the
calibration algorithm or fixed KV amax; those come from the historical
run. This is scheme verification, not a weight-hash comparison or
bit-identical regeneration claim.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — recipe-only backfill;
validated directly.
- Did you update Changelog?: N/A — recipe backfill.
- Did you get Claude approval on this PR?: N/A

### Additional Information

[Published
config.json](https://huggingface.co/nvidia/MiniMax-M2.7-NVFP4/blob/main/config.json)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a post-training quantization recipe for the MiniMax-M2.7 model.
* Supports NVFP4 quantization for expert weights and activations with
max calibration.
* Adds FP8 casting for the KV cache while retaining BF16 precision for
attention, routing, and language-model output components.

* **Documentation**
  * Documented the recipe and its corresponding published checkpoint.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-16 15:37:28 -05:00
yeyu-nvidiaandClaude Opus 4.6 8025a3dc54 specdec(recipe): add MiniMax-M2.7-DFlash streaming multi-node pipeline (#1835)
## Summary

- Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE)
streaming DFlash training
- 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each)
over NIXL RDMA hidden-state transport
- Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)`
+ final layer output
- MiniMax-specific: trust_remote_code, FSDP2 via accelerate config,
mask_token=200054, YaRN rope_scaling factor=48
- Topology matches Kimi-K2.5 large-MoE streaming recipe

Resolves OMNIML-5221

## Test plan

- [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`)
- [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`)
- [ ] Full streaming training run

Signed-off-by: Ye Yu <yeyu@nvidia.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a multi-node launcher configuration for MiniMax-M2.7 streaming
training with speculative decoding.
* Added an end-to-end workflow for conversational data preparation,
distributed training, checkpoint export, and vLLM smoke testing.
* Supports configurable serving and training settings across multiple
nodes.

* **Enhancements**
* Applies Transformers version overrides only to trainer or single-node
environments.
* Preserves the serving environment’s Transformers version on dedicated
serving nodes.
* Uses the launcher’s built-in bfloat16 mixed-precision configuration
for training.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-16 12:54:57 -07:00
Chad Voegele c118d359c1 docs: replace legacy TensorRT-LLM engine deployment guidance (#2436)
Chad's Agent

### What does this PR do?

Type of change: documentation.

Replace legacy TensorRT checkpoint export, support matrix, and
engine-build instructions with `export_hf_checkpoint` and TensorRT-LLM's
PyTorch backend. Preserve the existing 0.48.0 deprecation / 0.49.0
removal notice and page URL. Update the customized-model guide and
deployment skill to match.

### Usage

Follow the linked unified HF export guide. No API changes.

### Testing

- `git diff --check` passed.
- `uvx pre-commit run --files docs/source/deployment/1_tensorrt_llm.rst
docs/source/guides/_customized_model_quantization.rst
plugins/modelopt/skills/deployment/references/trtllm.md` passed.
- Full Sphinx build delegated to the Docs workflow; preview expected
after deployment.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅ Documentation only; page URL
retained.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation only.
- Did you update Changelog?: N/A — guidance correction; no new API
deprecation.
- Did you get Claude approval on this PR?: ❌ Not requested yet.

### Additional Information

Removes instructions for the TensorRT backend that current TensorRT-LLM
releases no longer support.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated TensorRT-LLM deployment guidance to use `export_hf_checkpoint`
with the PyTorch backend.
- Clarified that this workflow does not require TensorRT engine
construction.
- Updated DBRX customization instructions for exporting and deploying
quantized models.
- Added TensorRT-LLM version requirements and links to unified Hugging
Face deployment guidance.
- Removed guidance for the legacy TensorRT-LLM checkpoint exporter and
outdated troubleshooting steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-16 13:43:37 -05:00
Chenjie LuoandClaude Opus 5 655f94c207 Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)
### What does this PR do?

Type of change: documentation

`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.

This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.

It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.

**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.

### Usage

No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:

```python
from modelopt.recipe import load_recipe

cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```

### Testing

Docs-only change; no code paths touched, so no new or updated tests.

- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
  including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
  and the existing `CHANGELOG.rst` entry, so the documented defaults
  (`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
  NVFP4-input-quantizer-only scope match the code.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->

### Additional Information

The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 10:56:23 -07:00
realAsma 6a4b3f147e [Fix] Calibrate non-decoder modules during layerwise quantization (#2339)
### What does this PR do?

Type of change: Bug fix

Layerwise calibration now calibrates enabled quantizers outside
transformer layers, such as `lm_head`, while hiding decoder layers from
the additional calibration traversal. It also fails early when these
quantizers are combined with progressive layerwise export, whose
in-place conversion makes the required model calibration unsafe.

### Usage

No API changes.

### Testing

- `pytest_pwd tests/unit/torch/quantization/test_calib.py
tests/unit/torch/quantization/test_layerwise_calibrate.py -q` — 68
passed
- Pre-commit on all six changed files — passed
- `git diff --check` — passed

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Layerwise calibration now supports quantizers outside transformer
decoder layers, including LM-head activation quantizers.
- Calibration supports models with CPU-offloaded components while
preserving model behavior.
  - The MSE calibration utility is now publicly available.
- Added warnings for calibration passes involving offloaded components.

- **Bug Fixes**
- Prevented unsupported export-mode calibration when outside-layer
quantizers are present.
- Ensured outside-layer calibration runs after transformer-layer
restoration.
- Preserved original forwards while hiding decoder subtrees from
traversal and state collection.
  - Ensured cleanup and alias restoration after errors or completion.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-15 22:10:50 +00:00
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00
Chad Voegele 30f89908f0 Audit missing labels before release cherry-picks (#2373)
Chad's Agent

## Summary

- query NVBugs MCP for the exact `Committed_ModelOpt_<VERSION>` keyword
and inspect linked PRs
- audit PRs merged after the latest RC tag for missing cherry-pick
labels
- verify audit result pagination, report findings, and require
confirmation before adding labels

## Testing

- `npx --yes markdownlint-cli2@0.18.1
plugins/modelopt/skills/release-cherry-pick/SKILL.md`
- exercised the 0.47.0 audit using `0.47.0rc1`; GitHub returned 13
complete results
- verified the exact NVBug keyword query returned 33 bugs and inspected
their comments
- `git diff --check`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the release cherry-pick workflow to exclude pull requests
whose merge commits are already included in the target release branch.
- Applied this audit to both NVBug-linked and recently merged pull
request candidates.
- Updated recent pull request queries to reference the selected version
and release branch.
  - Renumbered the remaining workflow steps for consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-15 11:36:12 -05:00
Shengliang Xu 3c87751903 Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?

Type of change: deprecation

Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.

| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |

`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.

A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.

#### The warning fires only when a flag is actually passed

`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.

`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.

### Usage

```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast

# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```

### Testing

Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:

- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.

Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the removal release for these flags is not decided here, only the
deprecation.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-14 13:32:28 -07:00
Jenny Chen f70991f36e Docs: Add QAT and QAD guide [OMNIML-4859] (#2255)
### What does this PR do?

Type of change: Documentation

Add detailed quantization aware training (QAT and QAD) guide in our
docs, featuring examples on how to run in HF, Megatron-Bridge, and
Megatron-LM

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive guide for quantization-aware training (QAT) and
distillation (QAD), including workflows, framework guidance, setup,
training, and export steps.
  * Updated the Guides navigation to include the new QAT/QAD guide.
* Expanded quantization guidance with QAT/QAD workflows, scale handling,
and NVFP4 references.
* Improved save and restore guidance with clearer references and an
explanatory note.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-09-14 07:19:00 -07:00
github-actions[bot]andgithub-actions[bot] 6878c0ebd2 [chore]: weekly bump of uv.lock on main (2026-09-14) (#2428)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

<details>
<summary>uv lock --upgrade output</summary>

```
Using CPython 3.12.14 interpreter at: /opt/hostedtoolcache/Python/3.12.14/x64/bin/python3
Resolved 206 packages in 9.26s
Updated accelerate v1.14.0 -> v1.15.0
Updated ast-serialize v0.9.0 -> v0.11.2
Updated databricks-sdk v0.136.0 -> v0.139.0
Updated filelock v3.32.5 -> v3.32.6
Updated google-auth v2.57.1 -> v2.58.0
Updated huggingface-hub v1.30.0 -> v1.31.0
Updated hydra-core v1.3.2 -> v1.3.6
Updated multidict v6.7.1 -> v6.8.0
Added onnxruntime-ep-nv-tensorrt-rtx-cu13 v0.4.0
Updated onnxscript v0.7.1 -> v0.7.2
Updated platformdirs v4.11.7 -> v4.11.8
Updated regex v2026.9.3 -> v2026.9.10
Updated tqdm v4.70.0 -> v4.70.1
Updated tzdata v2026.3 -> v2026.4
Updated uv v0.12.10 -> v0.12.13
Updated uvicorn v0.52.4 -> v0.53.0
Updated virtualenv v21.7.8 -> v21.7.9
```

</details>

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-14 12:45:03 +00:00
realAsma 700e188ce5 Document copy-PR testing authorization (#2423)
### What does this PR do?

Type of change: documentation.

Explain how authorized vetters start NVIDIA-runner checks for pull
requests from forks, including the SHA-qualified copy-PR command and
links to the centralized contributor and vetter guidance.

### Usage

N/A; documentation-only change.

### Testing

- `pre-commit run --files CONTRIBUTING.md`
- `git diff --check`
- Verified both copy-PR documentation links resolve

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update `CHANGELOG.rst`?: N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Clarifies the missing required-check state encountered on #2231.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated fork pull-request guidance to require an authorized reviewer’s
approval before NVIDIA-hosted CI runs.
- Added instructions for contributors without write access to obtain
GitHub workflow approval from a reviewer with write permission.
- Added the `/ok to test <full-head-sha>` command for authorizing
testing against a specific commit.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-13 22:08:19 +05:30
JoshuaandCursor 1030791f53 Fix distributed AutoQuantize scoring and share backward setup (#2231)
### What does this PR do?

Type of change: Bug fix

AutoQuantize can measure a group of quantized expert layers at their
enclosing MLP output. That enclosing module is often a plain PyTorch
container and does not carry distributed-group information, so its
sensitivity score was not combined across data- or expert-parallel
workers.

This PR obtains the distributed groups from the quantized layers when
the scoring module does not provide them. It also preserves construction
order for quantized modules, scoring modules, and their registered
hyperparameters so every worker accumulates scores in the same order.

The temporary state needed by backward-based scoring is now managed by
one shared session. The session installs and removes forward patches and
invocation-specific output-gradient hooks, controls parameter gradients,
and restores the active quantization recipes even when scoring raises an
exception. Scoring methods remain responsible for their own score
calculation.

### Usage

N/A — this fixes existing AutoQuantize behavior and does not add an API
or flag.

### Testing

- `pre-commit run --files modelopt/torch/quantization/algorithms.py
tests/unit/torch/quantization/test_autoquant.py`
- `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102
passed
- Added a real two-rank gradient AutoQuantize test covering MoE experts
scored at an enclosing MLP.
- Added regressions for deterministic hyperparameter registration,
per-invocation replay for reused score modules, exact
`forward`-attribute restoration, partial setup rollback, and cleanup
after a scoring failure.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅ — no API or checkpoint format
changes; distributed sensitivity values now include the missing
reduction.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no new feature, deprecation, breaking change, or critical
release-note item.
- Did you get Claude approval on this PR?: ❌ — pending review.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization scoring consistency through deterministic
ordering and invocation handling.
* Added more reliable distributed score aggregation, including support
for mixture-of-experts models.
* Improved gradient-based scoring for repeated evaluations, tuple
outputs, and checkpoint-compatible workflows.
* Ensured model behavior and scoring state are restored after successful
or failed evaluations.
  * Avoided unnecessary output replay when gradients are not required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Joshua Hill <joshua.hill@baseten.co>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 14:55:32 +00:00
Ajinkya RasaneandCodex 51de53e48c [6463897] Fix narrow FP16 histogram calibration (#2412)
### What does this PR do?

Type of change: Bug fix

FP16 entropy calibration can fail for sufficiently narrow activation
ranges because NumPy may construct the histogram bin edges at FP16
precision. NumPy 2.2 and later reject the resulting collapsed bin
spacing, while earlier versions can silently return invalid,
non-monotonic edges. The same precision issue can recur when ONNX
Runtime merges later calibration batches.

Losslessly widen FP16 activation values to FP32 while calculating and
merging histograms in both ONNX entropy calibration paths, then restore
the source dtype at the calibration-to-quantization boundary. This keeps
the histogram bins stable without changing FP16 Q/DQ scale or graph
dtype semantics. FP32 inputs and public APIs are unchanged.

For full-range FP16 activations, ONNX Runtime can overflow while
subtracting FP16 calibration endpoints before it widens the result.
Retry only a non-finite FP16 scale calculation with FP32 endpoints, then
cast the finite scale back to FP16. Existing finite FP16 calculations
and all non-FP16 calculations continue to use ONNX Runtime's original
result.

The fallback intentionally patches only the `qdq_quantizer` binding used
for calibrated activation ranges. Initializer and weight quantization
continue to use ONNX Runtime's existing `quant_utils` path unchanged;
full-range FP16 weight scaling is outside this calibration fix.

The regression tests exercise both collectors across initial collection,
an equal-range merge, and an expanding-range merge. They also verify the
internal FP32 histogram and external FP16 calibration-range contract,
including finite saturation when restoring sanitized values. A real
entropy calibration test covers full-range FP16 values and verifies
finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast
integration verifies the same FP16 scale-type contract through the
public quantization path.

### Usage

N/A — no API or usage change.

### Testing

All tests ran with CUDA hidden.

- NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- Public INT8 entropy quantization with full-range FP16 calibration
data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX
Runtime session loaded.
- Changed-file pre-commit hooks: passed.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up to #1558.

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 23:49:01 -04:00
Ajinkya RasaneandCodex 5b1f7e86cc [6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
### What does this PR do?

Type of change: documentation

Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:

- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.

This also removes an inaccurate source comment without changing runtime
behavior.

### Usage

```bash
python -m modelopt.onnx.quantization \
    --onnx_path=model.onnx \
    --quantize_mode=int8 \
    --calibration_data_path=calib.npy \
    --output_path=model.quant.onnx
```

### Testing

- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [6701308]

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.

- **Tests**
- Updated quantization command examples to use the current calibration
data path option.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 22:04:19 +00:00
ZhiyuandClaude Opus 5 c37a6948db feat(quantization): fail fast when a quant config matches no weight quantizer (#2203)
### What does this PR do?

Type of change: New feature (fail-fast guard; behavior change on a
previously silent path)

A `quant_cfg` whose module patterns don't match the model is not an
error to `set_quantizer_by_cfg` — every pattern simply matches nothing.
The run then calibrates, exports, and hands back a checkpoint that is
silently unquantized:

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Nothing in the run says so. It has only ever been caught by someone
reading the exported `hf_quant_config.json` afterwards — most recently
on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665),
after a full 8×B200 calibration), and before that on MiniMax-M3, where
fused-expert detection skipped the experts and an experts-only recipe
matched nothing.

`mtq.quantize` now compares the config's intent against the outcome and
raises **before calibration**:

```
RuntimeError: The quantization config asks for weight quantization but no weight quantizer was
enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no
weight quantizer:
  *.experts.*weight_quantizer
Either the patterns do not match this architecture's module names (check the model-specific
recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights
were never converted to quantized modules (an unsupported custom module, e.g. a
trust_remote_code MoE layout).
```

Scoped to avoid false positives:

- **Only configs that ask for weight quantization** are checked (an
entry with `enable` and `weight_quantizer` in its pattern), so
activation-only and KV-cache-only configs are unaffected.
- **Intent is read from each pattern's final entry**, since `quant_cfg`
entries apply in order: a pattern that is enabled and then disabled
later asks for nothing by the end.
- **Configs refining an already-quantized model** (weight quantizers
enabled by an earlier `mtq.quantize`) are left alone.

Matching goes through `conversion._match_quantizer` — the same matcher
`set_quantizer_by_cfg` used to apply the config — so "did this pattern
match anything?" is answered exactly as the applying code would. A local
`fnmatch` diverges on the two cases that matcher handles:
`SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and
fused-experts names (`..._weight_quantizers.0` normalizing to
`..._weight_quantizer`).

### Usage

No API change. A config that would previously have produced an
unquantized checkpoint now raises:

```python
mtq.quantize(model, {"quant_cfg": [
    {"quantizer_name": "*", "enable": False},
    {"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
]}, forward_loop)   # RuntimeError if the model has no `experts` modules
```

### Testing

Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one
per branch of the guard: patterns matching nothing raise; an
activation-only config still runs; weight patterns disabled by a later
entry still run; enabled-then-retracted patterns still run;
`SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer
names count as matched; and the already-quantized refinement path is
exercised. Each was checked to be non-vacuous by removing the
corresponding branch and confirming exactly that test fails.

**One existing test changed.**
`tests/gpu/torch/export/test_fsdp2_export.py` parametrized over
`NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case
ran the FSDP2 paths against an *unquantized* model, and the new guard
reported it (4 GPU failures on the first CI run, all `quant_config6`;
`NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The
parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps
the scoped-recipe coverage. **If reviewers would rather not change that
test's meaning, the alternative is to downgrade the guard to a warning —
flagging it explicitly as a decision.**

I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the
four MLP/experts-scoped configs raise, and the other three
(`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`,
`NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE
models (Qwen3-MoE, gpt-oss), so no other test is affected.

Ran locally after rebasing onto current `main` (torch 2.11, transformers
5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271
passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which
needs `hydra`): 2676 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`. GPU tests were not run locally (no suitable GPU); the
FSDP2 change above is reasoned from the CI failure, not re-run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — deliberately. A config that
previously produced a `quant_algo: null` checkpoint now raises. Any such
run was already not doing what it claimed; the three scoping rules above
keep intentional non-weight quantization working.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (Backward Breaking Changes)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes
the specific model that motivated this. Independent branches; either can
merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantization now detects enabled weight-quantization patterns that do
not apply to any model weights and reports a clear validation error
before calibration.
- Broad wildcard patterns and nested quantizers are now handled
correctly.
- Overlapping patterns respect the final matching setting, including
later disabling rules.
- Existing quantized models can be refined using the parsed
configuration.
- Activation-only and explicitly disabled weight-quantization
configurations remain supported.
- Pipeline-parallel stages without targeted weights can bypass this
validation when configured to do so.

- **Documentation**
- Documented the process-wide override for bypassing unmatched
weight-quantizer validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 13:26:35 -07:00
Ajinkya RasaneandCodex bd90a5ed51 [OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do?

Type of change: new example

Adds an end-to-end BEVFormer-tiny ONNX PTQ example under
`examples/onnx_ptq/bevformer` with:

- The [BEVFormer
Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile)
from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second
container definition in this repository.
- A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10
source patch, and TensorRT plugin compilation at container runtime.
- Ordered temporal calibration-data generation that propagates
`prev_bev`, resets state at scene boundaries, computes CAN bus deltas,
writes the exact requested sample count, and publishes output
atomically.
- INT8 and FP8 quantization with BEVFormer-specific calibration
defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback
policy.
- Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus
nuScenes evaluation instructions.
- Shared temporary-ONNX-copy handling for BEVFormer and VoVNet
quantization so shape inference cannot mutate the source model.
- Focused CPU-only tests for temporal state, exact-count cleanup,
quantization defaults, plugin configuration, and source-model
preservation.

The ONNX PTQ index continues to document the shared PETR/FAR3D
containers separately and links to the BEVFormer guide.

### Usage

Clone the pinned DL4AGX revision and build its BEVFormer image:

```bash
git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX
git -C /path/to/DL4AGX checkout --detach \
  9f7b29104c253d5bc68334e7b83b3eecb72d4572
docker build \
  --build-arg TORCH_CUDA_ARCH_LIST=8.9 \
  --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \
  --tag modelopt-onnx-bevformer \
  /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq
```

The example command targets compute capability 8.9; use the deployment
GPU's compute capability for another architecture.

After exporting the model and generating temporal calibration data,
quantize it with:

```bash
python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \
  --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \
  --calibration-dir=/artifacts/calibration \
  --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \
  --quantization-mode=fp8 \
  --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx
```

See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin
compilation, export, FP16 feedback-engine creation, temporal
calibration, INT8/FP8 engine builds, and evaluation.

### Validation

#### Current revision: three-sample smoke validation

The Dockerfile at the pinned DL4AGX revision built successfully, and the
current Model Optimizer checkout was mounted into it for the workflow
below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch
2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage.

A fresh three-frame A/B/B smoke passed without calculating partial NDS
or mAP:

- Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and
FP8 engine builds passed.
- Temporal calibration published exactly three batches with
`use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent
feedback on the third frame, and the expected CAN bus deltas.
- The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs
with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero
points.
- Every engine produced finite outputs and 300 non-empty decoded
detections for each frame; maximum scores ranged from 0.9489 to 0.9608.

CPU-only validation passed the 16-test focused example set and the full
ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the
warning-as-error documentation build also passed.

#### Historical full accuracy reference

The following results were collected previously with TensorRT 10.14.1.48
on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration
samples and all 6,019 nuScenes validation samples. They are reference
results, not a completed full evaluation of the current revision.

| Precision | NDS | mAP |
| :-- | --: | --: |
| FP16 | 0.3546 | 0.2515 |
| INT8/FP16 | 0.3512 | 0.2505 |
| FP8/FP16 | 0.3526 | 0.2489 |

#### Performance

Following the PETR and FAR3D convention, performance is reported only as
speedup normalized to FP16. The archived full-validation engines were
benchmarked with five interleaved TensorRT 10.14 trials on the same
NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU
Compute Time median with data transfers disabled, CUDA Graphs enabled,
spin-wait, a 1-second warmup, a 10-second measurement window, and one
inference stream.

| Precision | Speedup vs. FP16 |
| :-- | --: |
| FP16 | 1.00x |
| INT8/FP16 | 1.87x |
| FP8/FP16 | 1.18x |

TODO: Investigate why FP8 delivers less speedup than INT8 for
BEVFormer-tiny.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from another source or added a new PIP dependency,
did you follow the guidance in `CONTRIBUTING.md`?: ✅
- Did you write the necessary tests?: ✅
- Did you update the
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional information

Reference workflow: [NVIDIA DL4AGX BEVFormer INT8
example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq).

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end BEVFormer 3D detection workflow for ONNX
post-training quantization, including temporal calibration, INT8/FP8
quantization, TensorRT engine generation, and nuScenes evaluation.
* Added command-line tools for preparing calibration data and quantizing
BEVFormer models.
* **Bug Fixes**
* Prevented source ONNX models from being overwritten during
quantization and improved temporary model cleanup.
  * Fixed FP8 export for BF16 models during real-weight compression.
* **Documentation**
* Added setup, usage, compatibility, and performance guidance for the
BEVFormer workflow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 16:15:59 -04:00
Chenjie LuoandClaude Opus 5 b80e164472 Carry a checkpoint's ModelOpt PTQ run into its eval's MLflow run (#2407)
### What does this PR do?

Type of change: documentation (agent skill)

Follow-up to #2374, which made a tracked `hf_ptq.py --mlflow` run leave
`.experiment.json` in the checkpoint it writes, naming the MLflow run
that quantized it. Nothing on the eval side read that file, so an eval
of a quantized checkpoint recorded no link back to the quantization that
produced it.

The `evaluation` skill now reads it. **Step 3** `cat`s the file as soon
as `checkpoint_path` is known; **Step 4** carries it into
`export.mlflow`:

| `.experiment.json` field | goes to |
| --- | --- |
| `experiment_name` | `export.mlflow.experiment_name`, verbatim |
| `run_name` / `run_id` / `run_url` | tags `modelopt_run_name` /
`modelopt_run_id` / `modelopt_run_url` |
| `tracking_uri`, `experiment_id` | deliberately unmapped |

The `modelopt_` prefix keeps them from reading as the eval's own run.
Values are quoted, or an all-digit `run_id` (or a `run_name` like
`20260910`) is YAML-coerced to an int or a date.

**`tracking_uri` is deliberately not inherited**, and the consequence is
documented rather than implied. Evals go to whatever
`$MLFLOW_TRACKING_URI` names. When the PTQ tracked to a different server
— the usual case, since `modelopttools:eval-config` points evals at
`mlflow.frontier-evals` while `hf_ptq --mlflow` typically writes to
`mlflow-modelopt` — the inherited name creates a *same-named, empty*
experiment on the eval server, and `modelopt_run_url` is the only route
back to the PTQ run. Evals of one checkpoint still group under a stable
name. Both cases are spelled out in Step 4 so nobody goes looking for
the PTQ run beside the eval.

Also updated: `references/nel-next.md` (nel-next configures MLflow
through its own `export_config.mlflow`) and `accessing-mlflow` (the
`tags.modelopt_run_id` query that closes the loop — without it the tag
is write-only).

### Usage

```bash
cat "$CHECKPOINT_PATH"/.experiment.json   # absent → name the experiment as usual
```

```yaml
export:
  mlflow:
    tracking_uri: ${oc.env:MLFLOW_TRACKING_URI}        # NOT the file's
    experiment_name: alice/hf_ptq/Qwen3.8-27B-NVFP4    # verbatim from .experiment.json
    tags:
      modelopt_run_name: '20260910-175422'
      modelopt_run_id: '7bec239a3a154970b062f3024a5ff20e'
      modelopt_run_url: 'https://<modelopt-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e'
```

Finding every eval of a checkpoint a given PTQ run produced:

```python
MLflow:query_runs(experiment_id, "tags.modelopt_run_id = '<ptq_run_id>'")
```

### Testing

Docs-only, so it was tested by having an agent follow the new wording
end to end and checking what NEL actually submitted.

A GPQA Diamond config (single repeat) was generated against an NVFP4
checkpoint carrying a `.experiment.json`, dry-run clean, then submitted
on a SLURM cluster. Verified in the `export_config.yml` heredoc **inside
the submitted `run.sub`** — the file the export job actually consumes,
not just the source YAML:

```yaml
experiment_name: chenjiel/hf_ptq/Qwen3.8-27B-NVFP4        # inherited verbatim
tags:
  modelopt_run_name: 20260910-000000-synthetic-test-fixture
  modelopt_run_id: 00000000fake0000fake0000fake0000
  modelopt_run_url: https://<modelopt-mlflow-server>/#/experiments/999/runs/...
tracking_uri: https://<eval-mlflow-server>/    # env var, not the file's
```

This confirmed the open question behind the change: NEL accepts
arbitrary `modelopt_*` keys in `export.mlflow.tags`. The
`.experiment.json` used was a synthetic fixture, since the checkpoints
predate #2374.

Repeated on a **second cluster with a second checkpoint** (different
account, QoS-based scheduler, FP8 rather than NVFP4) and carried through
to completion, so the result is confirmed as *stored by MLflow* rather
than merely submitted. Canary scored `gpqa pass@1 symbolic_correct:
90.0` (card: 88.01) with `no_answer: 0.0`, then the auto-export wrote:

```
experiment        chenjiel/hf_ptq/Qwen3.8-27B-FP8   (id 2106)
run               eval-<invocation>-ns_gpqa          FINISHED
modelopt_run_name 20260910-000000-synthetic-test-fixture
modelopt_run_id   00000000fake0000fake0000fake0000
modelopt_run_url  https://<modelopt-mlflow-server>/#/experiments/999/runs/...
```

Experiment 2106 did not exist on the eval server beforehand — the export
created it. That is the shadow-experiment case this PR documents,
observed rather than predicted.

The same exercise corrected the first draft of the wording: it had
claimed the rule makes a quantization and its evals "sit together on one
server", which is false whenever the two servers differ. Confirmed by
API — `chenjiel/hf_ptq/Qwen3.8-27B-NVFP4` returns
`RESOURCE_DOES_NOT_EXIST` on `mlflow.frontier-evals`. Rewritten as
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — additive guidance; the
absent-file path is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — skill documentation;
validated by a real submission, above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — agent-skill guidance, not a library/example feature.
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up to #2374.

Three adjacent problems were found while validating and deliberately
**left out of scope** — happy to split them into their own PR:

1. `recipes/examples/example_eval.yaml` puts `sbatch_comment` under a
top-level `cluster:` key, but the executor reads
`cfg.execution.sbatch_comment` (`executors/slurm/executor.py:688`) and
SKILL.md Step 1 says the same. As shipped the template silently drops
the idle-GPU reaper exemption, so long evals get reaped.
2. Step 4's `execution.gres` bullet names the wrong key for
`internal/slurm/<cluster>` configs, which supply `gpus_per_node` and no
`gres`; following it literally emits a redundant flag.
3. Step 3's `max_new_tokens` rules can contradict each other — "take the
card's highest" can equal `max_model_len`, leaving no room for the
prompt.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Clarified MLflow evaluation configuration requirements, including
literal experiment and sampling values, CPU partition settings, and
ModelOpt provenance.
- Documented validation and fallback behavior for unresolved `${...}`
placeholders in provenance values.
- Added guidance for querying cross-server evaluations by run ID and
handling same-named local experiments.
- Clarified that provenance files are optional, independent of
quantization detection, and do not control deployment flags.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 13:08:56 -07:00
Jenny Chen 5f8c76e2aa Fix TEGroupedMLP quantizer checkpoint resharding (#2319)
### What does this PR do?

Type of change: Bug fix for
https://github.com/NVIDIA/Model-Optimizer/issues/2209

Fix TEGroupedMLP per-expert weight quantizer checkpoint resharding.

`TEGroupedMLP` now saves its per-expert quantizer state as singleton
local shards, allowing the distributed checkpoint format to retain each
expert's global identity. Restore also initializes scalar `_amax`
placeholders after ModelOpt extra-state restoration so distributed
checkpoint loading can populate quantizer state for experts that move
between ranks.

This fixes restoring quantized TEGroupedMLP checkpoints across
expert-parallel and tensor-parallel topology changes.


Previously there was a bug that had two parts
1. TEGroupedMLP did not mark its per-expert quantizer state as
singleton_local_shards. That meant the scalar
weight_quantizer.<expert>._amax state was not saved with the same
globally unique expert identity as the grouped-expert weights, so DCP
could not reliably redistribute it across EP layouts.

2. During restore, ModelOpt’s extra-state restoration can leave _amax
absent for experts that were not local on the checkpoint’s saving rank.
The subsequent distributed checkpoint load then had no destination
tensor to populate.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing

- `ruff format`, `ruff check`, `mypy`, `bandit`, and repository
pre-commit hooks
- Focused GPU regression:

  ```bash
python3 -m pytest
tests/gpu_megatron/torch/quantization/plugins/test_megatron.py \
    -k te_grouped_sharded_state_dict_reshard -v

Replaced the prior metadata-only TEGroupedMLP sharded-state test with an
end-to-end distributed-checkpoint save/restore regression test.

The new test:
- Quantizes a TEGroupedMLP with per-expert NVFP4 weight quantizers.
- Assigns each local expert a distinct, deterministic `_amax` based on
its global expert index.
- Saves both the model distributed checkpoint and sharded ModelOpt
state.
- Rebuilds the model under a different TP/EP topology.
- Restores ModelOpt state, loads the distributed checkpoint, and
verifies each target-local expert received the expected global-expert
`_amax`.

The parameterized test covers:
- EP=2 -> EP=1
- EP=1 -> EP=2
- TP=1 -> TP=2
- TP=2 -> TP=1
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved checkpoint restoration for grouped quantizers by initializing
missing quantization statistics with compatible shapes.
- Improved restoration across supported grouped quantizer
configurations, including sequential groups and parallel checkpoint
layouts.
- Extra module state is now finalized through supported post-load
callbacks when available.
- Preserved populated quantized output-layer state during checkpoint
operations while removing empty placeholders.

- **Tests**
- Expanded checkpoint resharding coverage across tensor- and
expert-parallel configurations.
- Added coverage for disabled, dynamic, and other grouped quantizer
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
2026-09-11 19:28:43 +00:00
h-guo18andShengliang Xu c5d1065331 Modeling Lib: per-model architecture kick off with spec registration (#1828)
### What does this PR do?

**Type of change:** New feature — per-model infrastructure (with
behavior changes, see below)

Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it
knows about a
model, keyed by HF model type (`config.model_type`), one package per
type
(`<model_type>/specs.py`). Importing the package registers every spec;
consumers call
`get_spec(model_type)` / `match_moe_block(module, model_type)` and read
the fields.

The point is the infrastructure, not the migration. Today ModelOpt's
per-model knowledge
is scattered across if/elif chains in whichever subsystem happened to
need it, so the same
fact gets restated per subsystem and drifts. A model's MoE block
classes, its expert
projection naming, which of its norms store `w - 1` — these are facts
about the *model*,
and several subsystems want them. This PR gives them one home and one
lookup.

It is a kick-off, so it is deliberately narrow: specs plus the first
consumer. **Export is
that first consumer**, which is why most of the changed lines are on the
export side —
not because this is an export refactor. Quantization, speculative
decoding and the rest
keep their own tables for now; the per-model directory is where their
sections land later,
and `specs.py` is just the first thing in it.

| Data the first consumer moved in | Spec field | Consumers |
|---|---|---|
| MoE block classes and expert linear naming | `MoESpec` |
`get_expert_linear_names`, `get_experts_list`, `is_moe`,
`sync_moe_gate_up_amax` |
| grouped expert-export support set | `ExportSpec.grouped_expert_export`
| `get_experts_list` |
| AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` |
`fuse_prequant_to_linear` |
| weight-plus-one norm class names |
`ExportSpec.weight_plus_one_norm_names` |
`_layernorm_uses_weight_plus_one` |

`ModelSpec` holds each concern as a separate, optional section, and the
split is what
makes it extensible: **topic sections** hold architecture facts any
subsystem can read
(`MoESpec` — block classes, expert projection naming, how the experts
are stored),
**subsystem sections** hold one subsystem's own data and policy
(`ExportSpec` — AWQ fusion
rules, weight-plus-one norms, whether grouped expert export is validated
for this model).
A new subsystem adds a section rather than a table.

A section is `None` when the model has nothing to say about it, so a
dense model carries
`moe_spec=None` rather than an empty one. Two further fields record
where a model's
classes come from at all: `modeling_source` (`transformers` or
`remote_code`) and
`min_transformers_version`.

### Behavior changes

**1. Lookup key: `type(root_model).__name__.lower()` →
`config.model_type`.** The old key
changes after `quantize` wraps the model, forcing substring matching;
`model_type` is
stable, so `ExportContext` carries it and lookups are exact. Matching is
also tightened to
case-insensitive **exact** names against the module's **MRO**, so
quantized subclasses
still match via their base class without substring false positives.

**2. ⚠️ `get_expert_linear_names` raises instead of guessing.**
Unmatched MoE blocks
previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they
now raise
`NotImplementedError` telling you to register a `ModelSpec`.

The blast radius depends on the transformers release, because only
*iterable* expert
layouts consult the spec — a fused expert container is resolved by the
structural
first-projection check and never asks for naming. Measured against both
ends of the
support matrix:

| transformers | Unregistered MoE families `is_moe` admits | Reach the
spec lookup | Actually affected |
|---|---|---|---|
| 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none |
| 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** |

On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`,
`olmoe`,
`qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`,
so the `w1`
fallback already failed with `AttributeError` on `main` — for them this
trades an
unhelpful error for one that names the fix.

The other two, **`minimax` and `phimoe`**, are Mixtral-derived with
`w1`/`w2`/`w3`
experts and did export on `main`, so they are now **registered** rather
than left to
regress. Both only take the per-expert path on transformers 4; on 5
their experts are
fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check
resolves them, which
their specs decline to answer for. **So no model that exported correctly
before this
change regresses.**

This is backward-incompatible and has a `CHANGELOG` entry.

**3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE
path.** Neither
was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and
`DeepseekV32MoE`) is
invisible to `is_moe` — its class name does not end in `SparseMoeBlock`
and it calls its
router `gate`, so neither the name test nor the structural
`router`+`experts` test
matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec
to resolve expert
naming from. Registering `block_names` is what puts V3 on the MoE path
at all, so this is
added coverage rather than a fix to existing behavior.

`deepseek` stays registered alongside them: it describes the remote-code
`DeepseekMoE`
block of DeepSeek-MoE/V1, and its spec is the only thing making that
block detectable.

**4. Fused expert containers are skipped in `get_experts_list` instead
of crashing.**
transformers 5 replaced several iterable expert `ModuleList`s with a
single module holding
3-D parameters, while the specs still describe the transformers 4
iterable layout — a spec
cannot tell the two apart, only the module can. The AWQ/SVDQuant
resmooth pass in
`requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on
a module with no
`__len__` and died with `TypeError` mid-export. **Mixtral already hits
this on `main`**;
registering `deepseek_v3` only made an existing bug easier to reach. The
skip is scoped to
specs that claim iterable experts, so a layout the spec calls
unsupported (DBRX, whose
per-expert linears live under `experts.mlp`) still fails loudly rather
than silently
dropping resmoothing.

Everything else is behavior-preserving, pinned by tests:

- **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization
admits MoE blocks
*structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach
it. With no
spec it falls back to every declared gate/up naming — what it did
pre-registry.
Skipping them would leave the fused `gate_up_proj` halves on
inconsistent
  `weight_scale_2`.
- **`ExportSpec.grouped_expert_export` matches the legacy
`get_experts_list` support set**,
except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy
behavior to
  match. Note
`qwen3_5_moe`: legacy keyed off the root class name and
`"qwen3_5moeforcausallm"`
matched none of its qwen substrings (the `_5` breaks
`qwen3moeforcausallm`), so it
raised — the spec keeps `False` to match. Enabling it belongs in its own
PR.
- **`is_moe` consults the model's own spec first**, before the generic
name and structural
fallbacks, so per-model data always wins. The three checks are or-ed, so
this is a
no-op; the same reordering is deliberately *not* applied to
`get_expert_linear_names`,
  where the structural fused-experts check must stay first.
- **Two values are corrected, not moved.** `gpt_oss` declared
`block_names=("GptOssMoE",)`, which matches no real module (transformers
names it
`GptOssMLP`); naming still resolved via the single-naming shortcut, so
this was
  invisible. Now `("GptOssMLP", "GptOssMoE")`.

**5. ⚠️ DBRX expert input amax is now populated, changing its exported
scales.** The
second corrected value, called out separately because it has a numerical
consequence. The
legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers
does not define — it
names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3`
default, every
`hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated
`False`, and the
handler wrote no expert input amax at all. `dbrx/specs.py` declares the
real names
(`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it
was written to
do.

Anyone diffing a re-exported DBRX checkpoint against an older one will
see different
expert activation scales. That is the fix working, not a regression from
the refactor. No
`CHANGELOG` entry: DBRX is not listed as a supported export target in
`docs/source/deployment/` or `examples/hf_ptq/README.md`.

### Scope

One consumer wired up: the unified HF export path.

The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after
#2365, and are
deprecated. This PR changes one import line there — `is_moe` moved to
`modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no
longer resolves —
and nothing else. Their `is_moe()` calls still pass no `model_type` and
fall back to the
all-specs search, reproducing today's behavior; their hardcoded tables,
including a second
copy of the weight-plus-one norm names, stay as they are and migrate
whenever that package
does.

Worth noting that #2365 landed the same boundary from the other
direction: the slimmed
`export/layer_utils.py` is now documented as "module-shape predicates
and MoE quantizer
helpers shared by every export backend", which is exactly the set of
functions this PR
rewires. The Megatron path has its own per-family registry and is
untouched; folding it in
is a later question, not this PR's.

`is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into
`modelopt/torch/models/moe.py` — whether a module is an MoE block is a
modeling question,
not an export one, and it is the first piece of shared modeling logic to
follow the specs
into the new package.

### Relationship to #1939 (why a second registry?)

Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module
structure**
(*which handler runs?*), this registry resolves **family data** (*what
are its values?*).
Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared
iterable-experts handler
(also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names`
resolves the
`mixtral` spec to `("w1", "w2", "w3")`. One handler serves many
families, one spec serves
many handlers. Sharing the matcher machinery is a planned follow-up.

### Usage

```python
# modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits
register(
    ModelSpec(
        model_type="qwen3_moe",
        min_transformers_version="4.57",
        moe_spec=MoESpec(
            block_names=("Qwen3MoeSparseMoeBlock",),
            expert_linear_names=("gate_proj", "down_proj", "up_proj"),
            gate_up_pair=("gate_proj", "up_proj"),
        ),
        export_spec=ExportSpec(
            grouped_expert_export=True,
            pqs_fuse_rules=(
                (("Qwen3MoeAttention",), "v_proj", "o_proj"),
                (("Qwen3MoeMLP",), "up_proj", "down_proj"),
            ),
        ),
    )
)
```

One layout per model. `block_names` is a tuple, so a model whose MoE
appears under several
class names is covered as long as they share a layout (`gpt_oss`'s
`GptOssMLP`/`GptOssMoE`).
`gemma4_text` imports its section from `gemma4` rather than restating
it.

### Testing

`tests/unit/torch/export/` — **202 passed**;
`tests/unit/torch/quantization/` — **903
passed**. The `test_export_diffusers.py` collection error and 6
`test_quant_aware_conversion.py` failures are pre-existing and reproduce
on untouched
`main`, as does the one failing pre-commit hook
(`generate-arguments-md`, no torch in its
venv).

`tests/unit/torch/models/test_model_specs.py`: registry matching (MRO /
quantized classes
/ model-type scoping); every registered MoE section's `(block_names,
expert_linear_names,
fused_expert_names, gate_up_pair)` as an exhaustive table that fails if
a spec is added
without a row; the exhaustive `grouped_expert_export` support set; the
structural
fused-expert shortcut and the spec's precedence over it; the
fused-container skip in
`get_experts_list`; the `sync_moe_gate_up_amax` fallback; and
legacy-equivalence of the
aggregated `pqs_fuse_rules` / `gate_up_pairs` /
`weight_plus_one_norm_names`.
Mutation-checked — flipping a flag, typo-ing a block name, or swapping
the precedence each
fail.

`tests/unit/torch/models/test_specs_vs_transformers.py` validates the
specs against the
*installed* transformers rather than a mirrored table: every registered
block class must
name a real class from its `min_transformers_version` on, and every
`remote_code` model
must still be absent. It names no model — which specs to check, and
which cannot be, both
come from the registry.

> The run counts above were measured before the follow-up commits in
this branch (spec
> precedence, the policy split, the `MoESpec`/`MoELayout` merge) and
have not been
> re-measured; CI on the current head is the authority.

**Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from
config, run
against this branch and base `main`; not checked in):

| model_type | real block class | this branch | base `main` |
|---|---|---|---|
| `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` |
`gate/down/up` | same |
| `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same |
| `nemotron_h` | `NemotronHMoE` | `up/down` | same |
| `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` |
| `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` |

Five types identical to base; both differences are base bugs the
registry surfaces.
`dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig`
missing
`rope_theta`). `sync_moe_gate_up_amax` was separately checked against
base across
registered, unregistered, and no-config models — identical in all cases.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `get_expert_linear_names` no
longer falls back to `w1/w2/w3`; see "Behavior changes" #2.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code, no new dependency (stdlib only).
- Did you write any new necessary tests?: ✅ —
`tests/unit/torch/export/test_model_specs.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — happy to add an entry given the backward-incompatible item.
- Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding
on `sync_moe_gate_up_amax` is fixed above.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added model-aware export support across a broad range of MoE and dense
architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic,
DBRX, GPT-OSS, Llama, and Mixtral variants.
- Improved quantization and export handling for model-specific expert
layouts, projection fusion, normalization, and gate/up synchronization.
- Added automatic architecture recognition using Hugging Face model
metadata.

- **Bug Fixes**
- Unsupported or ambiguous model layouts now produce explicit errors
instead of applying potentially incorrect defaults.

- **Documentation**
- Added documentation describing model specifications and supported
architecture metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-11 12:12:54 -07:00
realAsmaandKeval Morabia dbe28e1e05 Document the MSE calibration API (#2405)
### What does this PR do?

Type of change: documentation

Expose `mse_calibrate` through `model_calib.__all__` so Sphinx
autosummary includes the existing MSE calibration API on the generated
`model_calib` reference page.

The documentation configuration honors each module's curated `__all__`
surface. Although `mse_calibrate` was implemented and used by the
calibration dispatcher, it was missing from that surface and was
therefore filtered out during API generation.

### Usage

```python
from modelopt.torch.quantization.model_calib import mse_calibrate
```

### Testing

- `pre-commit run --files modelopt/torch/quantization/model_calib.py`
- `git diff --check -- modelopt/torch/quantization/model_calib.py`
- Generated the recursive autosummary API tree using the repository
Sphinx configuration and module template.
- Verified the generated RST contains `mse_calibrate`.
- Rendered the focused module page and verified its function-table link,
anchor, signature, and docstring.

The canonical full build was unavailable in the active environment
because the configured `shibuya` theme is not installed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing
documentation build directly exercises this declarative autosummary
contract.
- Did you update Changelog?: N/A — this is a documentation-visibility
repair for an existing API.
- Did you get Claude approval on this PR?: N/A

### Additional Information

No source implementation or calibration behavior changed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Made the MSE calibration capability publicly available for
quantization workflows.

- **Chores**
- Increased documentation build time limits to improve reliability for
longer-running builds.
- Increased multi-version test time limits to better accommodate
extended test runs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-11 18:18:23 +00:00
noeyy-mino a1bcda4727 Fix protobuf size-check failures in ONNX deployment (#2403)
### What does this PR do?

Type of change:  Bug fix:6701737

The ONNX deployment path assumed that ModelProto.ByteSize() would always
return a valid size. With newer protobuf versions, querying the size of
a model exceeding the protobuf serialization limit can itself raise
EncodeError: Failed to serialize proto.

Replaced both direct size checks with the existing
is_model_too_large_for_protobuf() helper. This helper handles size-query
failures conservatively and checks the protobuf size limit:

Shape inference now selects the external-data/file-based path when
ByteSize() fails or the model is too large.
Metadata creation uses the same safe check instead of raising another
serialization error.
The unused TWO_GB constant was also removed.

### Usage

```
python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image
```

### Testing
the above test case pass on B100

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved ONNX model size detection during shape inference and export
processing.
* Ensured large models consistently use the appropriate external-data
handling path.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
2026-09-11 10:46:10 -04:00
noeyy-mino 5d2d5a5d15 deprecate trtllm-build in weight_sparsity (#2371)
### What does this PR do?

Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190

<!-- Details about the change. -->

### Usage

```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt  --calib_size 128  --output_dir Llama-3.1-8B-Instruct_pts

python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts

trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
    --tp_size 1 \
    --pp_size 1 \
    --host 0.0.0.0 \
    --port 8000

```

### Testing
PTS and SAT tested

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
  * Corrected the PTS model restoration path.

* **Bug Fixes**
  * Model export now saves the tokenizer alongside the checkpoint.
  * Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
2026-09-11 16:51:23 +05:30
Wei-Ming Chen 59d93af064 [OMNIML-5570, OMNIML-5569] 1/2 Add layer-wise KV-cache AutoQuant with forward KL (#2272)
### What does this PR do?

Type of change: new feature.

Adds standalone layer-wise KV-cache AutoQuantize through the existing
public
`mtq.auto_quantize` API:

- dispatches KV search with
`constraints={"effective_bits": ..., "cost_model": "kv_cache"}` and
forward-KL
  sensitivity;
- selects one supported K/V format for every eligible causal-attention
layer;
- supports persistent/exportable FP8 K/V, NVFP4 K/V, and FP8-K/NVFP4-V
candidates;
- solves a K/V-width- and scale-storage-aware additive recipe with the
existing
  PuLP-backed constrained solver;
- uses `BaseSearcher` lifecycle and safe checkpoint restore/save
machinery;
- preserves existing non-KV execution while isolating K/V candidate
calibration;
- returns standard AutoQuantize state that can be re-solved at another
KV budget;
- produces a complete KV-only replay config that disables every non-KV
quantizer;
- saves JSON-safe sensitivity metadata and the exact selected layer
mapping; and
- invokes the public API from `examples/hf_ptq/hf_ptq.py` through a
standalone
  calibration-free recipe.

The implementation is architecture-driven. Plain and
conditional-generation Qwen
causal attention is supported, VLM vision attention is excluded through
the existing
language-model extraction boundary, hybrid full-attention mixers are
discovered through
their paired K/V quantizers, and nonattention/Mamba modules remain
outside the search.
Ambiguous language-model roots, unsupported distributed execution,
structural
algorithms, invalid storage declarations, nonpersistent scales, and
unsupported K/V
pairs fail closed.

KV-only unified HF exports leave weight-quantization fields unset.
Uniform all-FP8 or
all-NVFP4 selections retain their legacy KV scheme while also carrying
the complete
`kv_cache_quantized_layers` map and schema version; genuinely
layer-mixed selections use
the KV-side `MIXED_PRECISION` marker plus the same map. This keeps
weight-loader metadata
accurate and prevents disabled vision attention from making uniform
language-model KV
quantization appear partially quantized.

GEMM PTQ/AutoQuantize followed by KV AutoQuantize is intentionally
excluded and proposed
separately in stacked PR #2273.

### Why KV search has a dedicated backend

The user-facing entry point remains `mtq.auto_quantize`; no separate
public KV search API
is introduced. `AutoQuantizeKVSearcher` extends `BaseSearcher` and
reuses its reset,
checkpoint load/save, and search lifecycle, along with existing Pydantic
configuration,
calibration, safe checkpoint I/O, and PuLP-backed selection utilities.

The backend remains KV-specific because a decision owns paired K/V
quantizers on one
attention layer, its cost depends on separate K/V widths and data/scale
storage, BF16 is a
scoring reference but not a deployable solver choice, and the
optimization objective is
additive isolated forward KL under a KV-storage constraint. These
contracts do not match
the weight-domain hparam grouping, parameter-count cost, or
threshold-selection behavior
of the existing weight AutoQuant searchers. Keeping the specialization
behind the shared
API avoids changing established weight-search solver and scoring
behavior.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3.8-27B \
  --recipe general/auto_quantize/kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3.8-27b-mixed-kv
```

The search checkpoint is compatible only with the same model,
eligible-layer geometry,
candidate configurations, and scoring setup. Use a distinct checkpoint
path after any of
those inputs change.

KV-cache AutoQuantize rejects `--use_fsdp2` before model loading because
its sensitivity
scoring, selection, and checkpoint writes are single-process. Existing
weight
AutoQuantize retains its previous experimental FSDP2 warning and
behavior.

### Testing

- Focused coverage exercises candidate validation/calibration, paired
K/V scoring and
storage accounting, solving, checkpoint resume, failure atomicity,
disabled layers,
fresh-model replay, Qwen/VLM/hybrid boundaries, JSON-safe reports, and
unified export.
- Uniform FP8/NVFP4 KV-only exports retain the legacy KV scheme and
complete layer map
without claiming a weight algorithm; disabled VLM vision attention is
excluded from
  causal-KV eligibility.
- The shipped recipe runs end to end on a tiny offline Qwen fixture and
preserves
  exportable scale state.
- After merging current `main`: 432 focused recipe/KV/export/hf_ptq
tests passed, with one
unrelated optional-dependency skip; changed-file pre-commit hooks
passed.

### Deployment gate

The producer schema is covered here. Runtime consumption of
`kv_cache_quantized_layers` is tracked in vLLM PR
https://github.com/vllm-project/vllm/pull/52813. Do not treat a produced
checkpoint as
runtime-supported until that consumer lands and the target K/V kernels
are available.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow
  guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅

### Additional information

- This is split from the combined ground-truth implementation in draft
PR #2211 to reduce
  review scope; composition is isolated in #2273.
- The standalone core tree contains no composed GEMM→KV recipe schema or
orchestration.
- No model-name checks, checkpoint-specific layer lists, campaign data
contracts, cluster
  launch logic, or runtime-kernel implementations are included.

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-10 22:48:34 -07:00
Chenjie LuoandClaude Opus 5 d69e93a72b Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?

Type of change: new feature

A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.

A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.

Two deliberate behaviours:

- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.

A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.

### Usage

```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
    --export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```

```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
  "tracking_uri": "https://<your-mlflow-server>",
  "experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
  "experiment_id": "36",
  "run_id": "7bec239a3a154970b062f3024a5ff20e",
  "run_name": "20260910-175422",
  "run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```

```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```

### Testing

**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.

**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.

**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:

- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 22:24:37 +00:00
Chenjie LuoandClaude Opus 5 a74054ab2b Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?

Type of change: new feature

The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:

- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags

`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.

Two details worth a reviewer's attention:

**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.

**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.

`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —

```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```

arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.

### Usage

```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```

### Testing

Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.

End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):

```
modelopt_internal_version  '49fa29d5'
modelopt_version           '0.47.0rc0.post32+gd38ed5ead'
git_sha                    'd38ed5ead'
quant_cfg                  'NVFP4_DEFAULT_CFG'
```

The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌

### Additional Information

Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 20:29:16 +00:00
7f7c46d820 [6508436] Fix BF16 FP8 ONNX export (#2314)
### What does this PR do?

Type of change: Bug fix

Fix FP8 ONNX export for BF16 models during real-weight compression
without changing the public API or the `weights_dtype="fp32"` default.

The FP8 exporter preserves BF16 initializer bits when bridging
GraphSurgeon NumPy arrays to Torch, widens BF16 values exactly to FP32
for normalization, and leaves existing FP16/FP32 handling unchanged.
Conv scales and dequantized outputs retain the source dtype, and scales
round upward when needed so serialized values cannot cause FP8 overflow.

`weights_dtype="bf16"` is accepted as a no-op only for FP8-only models
whose floating parameters are all BF16. Registered buffers do not affect
this weight-focused decision and may preserve higher-precision regions
in the exported graph. Unsupported BF16 FP8-to-FP16 and FP32 or
mixed-parameter-to-BF16 conversions are rejected with `ValueError`
before temporary export paths are created. A narrow GraphSurgeon fix
preserves integer BF16 value-info dtypes.

### Usage

```python
onnx_bytes, metadata = get_onnx_bytes_and_metadata(
    quantized_fp8_model,
    (sample_input,),
    weights_dtype="bf16",
    onnx_opset=23,
)
```

### Testing

- Seven focused CPU regressions passed: BF16 QDQ compression and integer
dtype handling, BF16-to-BF16 and FP32-to-FP16 Conv/Linear export, and
four unsupported-conversion cases.
- QDQ utilities: 31 passed; pytest 2.25s, wall 29.88s.
- FP8 MHA exporter: 6 passed; pytest 2.05s, wall 32.71s.
- Torch deploy utilities: 51 passed; pytest 8.94s, wall 25.74s.
- Torch ONNX CPU export: 36 passed; pytest 4.85s, wall 32.38s.
- Changed-file pre-commit hooks: all passed; wall 8.07s.
- Exact-head FP8 BF16 GPU workflow at `f21d62a`: exit code 0; ONNX
checker passed; 6 FP8 initializers, 3 native `DequantizeLinear` nodes,
and 12 BF16 initializers.
- Refreshed GitHub CI at `f21d62a`: 50 passed and 1 skipped. Unit, GPU,
and regression required aggregates and Codecov passed. Two ONNX example
leaves failed because the runner could not load a cuDNN sublibrary;
their dependent example aggregate consequently failed.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information

- TODO: Deliver authoritative `native`/FP32/FP16/BF16 ONNX export across
all quantized formats in follow-up pull requests.

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-10 18:30:26 +00:00
Shengliang Xu d19925e446 simple refactor(export): split TensorRT-LLM-only code into modelopt/torch/export/trtllm (#2365)
### What does this PR do?

Type of change: refactor.

**The TensorRT-LLM checkpoint export format is deprecated.** Per
`docs/source/deployment/1_tensorrt_llm.rst`: *"The
`export_tensorrt_llm_checkpoint` API will be deprecated in future
releases. Users are encouraged to transition to the unified HF export
API, which provides enhanced functionality and flexibility for exporting
models to multiple inference frameworks including TensorRT-LLM, vLLM,
and SGLang."*

That deprecated code was not sitting off to one side — it was
**interleaved with the export path we actually want to grow.**
`modelopt/torch/export` mixed the deprecated TensorRT-LLM checkpoint
logic with the framework-agnostic HF/Megatron export code, in the same
modules:

- `layer_utils.py` was 1,986 lines, of which ~1,600 were TensorRT-LLM
`build_*_config` builders. The HF path imports this module for five
small predicates (`is_moe`, `is_quantlinear`, …) and dragged the whole
deprecated builder set in with them.
- `model_config.py` held the TensorRT-LLM `ModelConfig` dataclasses
*and* the `QUANTIZATION_*` / `KV_CACHE_*` constants that every backend
needs, so all of HF export imported the deprecated checkpoint schema to
get a format name string.
- `quant_utils.py` carried two helpers whose only caller is the
deprecated `postprocess.py`.

**This PR isolates the deprecated format so it stops polluting the
HuggingFace export path.** Everything reachable only from
`export_tensorrt_llm_checkpoint` now lives under
`modelopt/torch/export/trtllm/`, and the dependency is **one-way**:
`trtllm/` reaches into the parent through `quant_format`, `quant_utils`
and `layer_utils`, and **no implementation module in the parent imports
`trtllm/`.** The single exception is the deprecation re-export in
`modelopt/torch/export/__init__.py` described below, which is scheduled
for deletion in 0.49.0.

That one-way edge is the property worth protecting in review. It means
the deprecated format can be evolved, frozen, or eventually removed
without touching HF export, and HF export can no longer accidentally
grow a dependency on it.

### Deprecation handling

The format has carried a deprecation notice in the deployment docs since
`bc546943b4` (2025-10-08, first shipped in 0.39.0) — about 11 months.
But the deprecation policy in `README.md` also specifies *how* a
deprecation is communicated: a changelog entry, a source statement of
timing, and a runtime warning on use. **None of those existed**; only
one docs page ever said anything. So 0.48.0 is the first release that
gives users a signal they can act on, and this PR treats it as the
*start* of the migration period rather than the end:

- Both entry points now emit a `DeprecationWarning` naming 0.48.0 and
the 0.49.0 removal.
- `export_tensorrt_llm_checkpoint` and
`torch_to_tensorrt_llm_checkpoint` **remain importable from
`modelopt.torch.export`** for this release only, so existing callers
keep working *and* actually receive the warning. Removing the path in
the same release that first warns would mean callers hit `ImportError`
and never see it.
- The 0.49.0 removal date is stated in all four channels the policy
names: the runtime warning, the source (`.. deprecated:: 0.48.0` plus a
comment), the changelog, and the deployment doc.

The deeper module paths (`modelopt.torch.export.model_config_export`,
`modelopt.torch.export.model_config`) are **not** forwarded. Neither
appeared in a docs example, and `model_config.py` never declared
`__all__`, so by the `__all__` convention in `CONTRIBUTING.md` they were
never part of the public surface.

Eight modules had no non-TRT-LLM importer and moved whole:
`model_config_export`, `model_config_utils`, `postprocess`,
`distribute`, `tensorrt_llm_utils`, `tensorrt_llm_type`,
`hf_config_map`, `mcore_config_map`.

Three were genuinely mixed and were split by call-graph analysis rather
than by file:

| module | stayed shared (HF path) | moved to `trtllm/` (deprecated) |
|---|---|---|
| `model_config.py` | `QUANTIZATION_*`, `KV_CACHE_*`,
`FUSION_FREE_FORMATS` → new leaf module `quant_format.py` | the
`ModelConfig` dataclasses + `LINEAR_*`/`LAYERNORM_*` checkpoint-layout
constants |
| `layer_utils.py` | 9 module-shape predicates and MoE quantizer helpers
(`is_moe`, `is_quantlinear`, `get_experts_list`,
`sync_moe_gate_up_amax`, …) | the 39 `build_*_config` builders and
enc/dec helpers |
| `quant_utils.py` | everything else | `get_scaling_factor_from_weight`,
`resmooth_and_get_scale` (only caller is `trtllm/postprocess.py`) |

`adjust_attn_amax_values` was deliberately left in the shared
`quant_utils.py`: it has no production caller at all (only a test), so
"used only by TRT-LLM export" is not demonstrable for it.

Nothing was added or removed. `export_tensorrt_llm_checkpoint` behaves
exactly as before, just from a new import path and with a warning
attached.

### Usage

```python
# Deprecated TensorRT-LLM checkpoint export — new home, and warns on call
from modelopt.torch.export.trtllm import (
    export_tensorrt_llm_checkpoint,
    torch_to_tensorrt_llm_checkpoint,
)
from modelopt.torch.export.trtllm.model_config import ModelConfig

# The pre-0.48 path still works for one release, and warns — removed in 0.49.0
from modelopt.torch.export import export_tensorrt_llm_checkpoint

# Shared format constants — new home, still re-exported from the top level
from modelopt.torch.export.quant_format import QUANTIZATION_NVFP4, KV_CACHE_FP8
from modelopt.torch.export import QUANTIZATION_NVFP4  # still works

# The recommended path — unchanged
from modelopt.torch.export import export_hf_checkpoint, get_model_type
```

### Testing

- `pre-commit` on all changed files: passes (ruff, ruff-format,
**mypy**, bandit, markdownlint). mypy caught one implicit re-export of
`is_layernorm`, now imported from the shared module directly.
- `tests/unit/torch/export`: **189 passed**. With the new `trtllm/` test
dir: **193 passed**.
- Full `tests/unit/torch`: **2367 passed, 0 export failures**. The 45
failures are pre-existing environment issues — a deepspeed circular
import and a read-only HF cache — confirmed by reading their error text,
not assumed.
- `pytest tests/gpu/torch/export --collect-only`: 172 items, no
collection error.
- In-repo consumers updated and re-verified by an AST scan that imports
every `modelopt.torch.export*` module referenced anywhere in the tree
and checks each imported name still resolves: `hf_ptq.py`,
`export_trtllm_ckpt.py`, `deepseek_v3/ptq.py`, the AutoQuantize
notebook, `hf_ptq/README.md`, 2 docs pages, 4 tests.
- **Deprecation contract is covered by committed tests** (3 new, in the
`trtllm/` test dir): the pre-0.48 top-level import still resolves to the
same objects, `torch_to_tensorrt_llm_checkpoint` warns *at call time*
rather than on first `next()` (it returns a generator, so a naive
`warnings.warn` in the body would fire late or never), and one
`export_tensorrt_llm_checkpoint` call emits exactly one warning rather
than two. The first of these makes closing the migration window early a
test failure rather than a silent regression. `pyproject.toml` sets no
`filterwarnings = error`, so no suite fails on the new warning.
- **After merging `main`** (4 commits, incl. a 180-line rewrite of
`unified_export_megatron.py` that touches a file this PR also edits):
merged with no conflicts, then re-verified rather than trusted — import
scan clean across 24 export modules, `ruff check` clean repo-wide, 193
export tests passing, GPU collection still clean.

**Not run: the GPU suites** (`tests/gpu/torch/export`,
`tests/gpu_trtllm`) — no GPU in my environment.
`tests/gpu/torch/export/test_export.py` had its imports retargeted, so
it is the one most worth a GPU run before merge.

### Reviewer note: the deprecated path has no test coverage

Worth knowing before reviewing. **No test in the repo — including
`tests/examples/` — calls `export_tensorrt_llm_checkpoint`,
`torch_to_tensorrt_llm_checkpoint`, any `build_*_config`,
`convert_to_tensorrt_llm_config`, or `postprocess_model_config`.** So
~4,600 moved lines have no direct tests, and this refactor is validated
by import-graph reasoning, lint and mypy rather than by tests exercising
the moved code.

Given the format is deprecated and scheduled for removal in 0.49.0, **no
new coverage is planned for the conversion path itself** — writing fresh
tests for an API being removed next release isn't a good use of effort.
The gap is documented so reviewers can weigh the risk, not as a TODO.
(The deprecation *mechanism* is tested; see Testing.)

One caveat on how the gap was established: a runtime check showing all
12 `trtllm` modules in `sys.modules` after the export suites is *not*
evidence of coverage — importing any submodule runs
`trtllm/__init__.py`, which star-imports `model_config_export` and pulls
in the rest. Real line coverage could not be measured (`coverage`'s
tracer is incompatible with this venv's torch build: `ValueError: module
functions cannot set METH_CLASS or METH_STATIC`, on both the C tracer
and `sysmon`). The claim rests on a call-site audit generated from the
actual public symbols of those modules.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ for the public API —
`export_tensorrt_llm_checkpoint` and `torch_to_tensorrt_llm_checkpoint`
remain importable from `modelopt.torch.export` through the 0.49.0
migration period, now with a `DeprecationWarning`. The undocumented
submodule paths `modelopt.torch.export.model_config_export` and
`.model_config` did move; see **Usage**.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; existing code relocated.
- Did you write any new necessary tests?: ✅ — 3 tests covering the
deprecation contract (old import path, call-time warning, exactly-one
warning). One existing test also moved to mirror the source split. No
new coverage for the deprecated conversion path itself; see the note
above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 **Deprecations**, covering both the runtime warning and
the new import location.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Git detected the moves, so the diff stays reviewable: 8 files show as
pure renames (100%), the three split files as rename/copy at 94–99%
similarity, and only `layer_utils.py` as a 79% rewrite — expected, since
it shed 1,616 lines to `trtllm/`.

`examples/hf_ptq/hf_ptq.py` and
`examples/llm_sparsity/weight_sparsity/export_trtllm_ckpt.py` still call
the deprecated API, so those examples now print the warning. That is the
intended nudge, but happy to silence or migrate them if preferred. They
import from the new `.trtllm` path already, so they need no change at
0.49.0.

Two incidental changes, easy to revert if unwanted:
- `modelopt/torch/export/layer_utils.py` mode `100755 → 100644` (it was
needlessly executable).
- The new test is named `test_trtllm_quant_utils.py`, not
`test_quant_utils.py`: these directories have no `__init__.py`, so
pytest derives the module name from the bare filename and the shorter
name fails collection with `import file mismatch` against the existing
`test_quant_utils.py` one level up.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added shared quantization and KV-cache format definitions for export
workflows.
- Added expanded TensorRT-LLM export support, including broader model
architecture and quantization handling.
- Added distributed export utilities for coordinating checkpoint data
across processes.

- **Deprecation**
- TensorRT-LLM checkpoint export now emits a warning and is scheduled
for removal in version 0.49.0.
- Use the documented export module and save optimized model state
explicitly when needed.

- **Documentation**
- Updated guides and examples with new import paths and deprecation
guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-10 10:44:04 -07:00
Chad Voegele 28dc117594 Add Day 0 Sub-agent Roles (#2006)
### What does this PR do?

Type of change: new feature

Adding sub-agent definition for Day 0 workflows.

I'm adding subagents as an alternative, while we evaluate which is
better. They are duplicated because Claude & Codex have different
formats & different locations. Deriving from a shared vendor-neutral
format at build-time is too complicated.

### Usage

Tell your agent "Use the modelopt_model_quantizer agent to quantize the
model"

### Testing

Ran a trial using Qwen-2.5, Muse Glimmer, Qwen-2.8, and GLM-5.3-Flash.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
See design doc.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
- Added specialized Model Optimizer agents for downloading,
quantization, recipe search, evaluation, deployment, and performance
benchmarking.
- Added coordinated Codex and Claude Code access with validation
requirements, artifact preservation, and concise handoffs.
- Improved skill discovery across plugin and repository installations.

## Documentation
- Clarified agent discovery, layout, and canonical editing locations.
- Updated configuration guidance for supported agent definitions.

## Tests
- Added synchronization checks for agent definitions and links.
- Improved test compatibility for Python versions below 3.11.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-10 17:30:37 +00:00
yeyu-nvidiaandClaude Opus 4.8 9d0df45849 specdec: config_overrides for nested text_config checkpoints + load VLM-capable bases in merge_lora (#2289)
### What does this PR do?

Type of change: New feature + bug fix

Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.

**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.

**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.

Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.

### Usage

```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <ckpt> --trust_remote_code \
    --config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'

# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
    --model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
    --base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
    --config_overrides '{"num_hidden_layers": 36}'
```

```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36})  # default None
```

### Testing

Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:

- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.

No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
  - Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.

- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-09-10 10:06:35 -07:00
Ajinkya Rasane 635688d26e [6410139] Fix ONNX AutoCast for large external initializers (#2317)
### What does this PR do?

Type of change: Bug fix

Fix ONNX AutoCast for models whose external initializers exceed the
in-memory protobuf limit.

- Keep external tensor payloads reference-only during graph sanitization
and type inference, then materialize them once before value-dependent
classification and conversion.
- Duplicate shared initializers directly in `GraphProto`, preserving
external-data metadata without reading tensor bytes.
- Use file-backed ONNX paths for validation, shape inference, reference
execution, and custom-operator inspection when required.
- Avoid a redundant sanitizer pass in the fully sanitized AutoCast path
while preserving existing behavior for direct `PrecisionConverter` and
`convert_to_f16()` callers.

No CLI flags, dependencies, or public return types change.

### Usage

```bash
python -m modelopt.onnx.autocast \
    --onnx_path model.onnx \
    --output_path model_bf16.onnx \
    --low_precision_type bf16
```

### Testing

- Ran the complete CPU-only AutoCast unit suite with no GPU visible: 248
passed.
- Ran focused ONNX utility regressions covering shared initializer
duplication and file-backed protobuf routing: 8 passed.
- Ran pre-commit on all changed files.
- Ran CPU-only integration coverage with an exact 2,147,485,696-byte
external initializer using protobuf 7.35.1 and 6.33.6.
- Verified BF16, FP16, shared-initializer, and aggregate-external-data
runtime-cast cases.
- Verified every integration output with `onnx.checker.check_model(...,
full_check=True)`; outputs expected to remain external-data-backed did
so.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed ONNX AutoCast failures for models with external initializers
larger than 2 GiB.
- Improved handling of large or external-data models during shape
inference, conversion, and runtime validation.
- Preserved external initializer metadata while avoiding unnecessary
data materialization.
- Improved temporary-file cleanup when model loading or inference fails.
- Improved processing of shared initializers, nested graphs, and custom
nodes.
- **Tests**
- Added coverage for large models, external initializers, nested graphs,
custom nodes, and runtime cleanup.
- **Documentation**
- Updated the changelog with recent fixes and benchmarking information.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-09-10 16:16:43 +00:00
sugunav14andClaude Opus 5 079078de9d FSDP2 export optimizations (#2207)
### What does this PR do?

  **Type of change:** Performance enhancement

FSDP2 checkpoint export gathered the whole model onto rank 0 and wrote
it from a single process. On large models that made export the slowest
phase of a PTQ run and could exhaust host memory.

The model is now split into export units — one per decoder layer, plus
one for everything else that owns state — and the units are dealt
round-robin across ranks. Every rank enters the unshard window for every
unit, because the gather is a collective; only the owner keeps the
result, packs it, and streams it to its own shard files. Rank 0 names
the shards and writes the index once every rank has finished.

Two properties follow. Peak host memory per rank is ~model/world instead
of the whole checkpoint on rank 0. And packing runs on a gathered, full
weight, so it is byte-for-byte the single-process code path — no amax
narrowing, no block-quantized fallback, no scale-buffer heuristics.

Measured on B300 (bia), NVFP4, `hf_ptq.py --use_fsdp2`:

| model | world | export | artifact |
|---|---|---|---|
| Qwen3-235B-A22B | 8 (1 node) | **168.6s** | 146,361 keys / 16 shards /
0.122 TiB |
| oakhaven-max 4.5 TB | 64 (8 nodes) | **1359.7s** | 568,398 keys / 158
shards / 1.292 TiB |

Checkpoints were validated, not just timed: every index entry present,
no leftover `__shard_part*`, all ranks exit 0. A repeat of the oakhaven
configuration landed at 1411.7s, so treat ~1.36–1.41 ks as its range. No
speedup ratio is quoted: the old rank-0 path was last measured on
different hardware on a different day, and this cluster drifts up to 14%
run-to-run on an identical tree, so a cross-day ratio would not
reproduce. A same-day baseline of `main` is the outstanding item.

Two alternatives were considered and rejected.
`torch.distributed.checkpoint`'s `get_model_state_dict` full-gathers to
rank 0 only with no per-module owner knob, so it reproduces the exact
host-memory bound this removes, and per layer it was not faster than
`full_tensor()`. Packing each rank's shard in place was ~9% faster at
Qwen3-235B w8 on the trees measured at the time, but rebuilt `FSDPParam`
objects from torch's private FSDP internals and carried ~400 lines of
format-specific shard handling; gathering to the owner deletes all of
it.

Configurations that would silently corrupt a checkpoint now raise
instead: FSDP2 composed with another DTensor parallelism (FSDP2 + TP on
a 2-D mesh; HSDP is fine), models with no discoverable decoder layers, a
decoder layer object reused across layers, and a container that holds
the decoder layers while owning direct state of its own.

### Usage

No API change — `export_hf_checkpoint` selects the path automatically.

```python
from modelopt.torch.export import export_hf_checkpoint

# model is FSDP2-wrapped; every rank must call this (it unshards collectively)
export_hf_checkpoint(model, export_dir="./exported")
```

Testing

- `test_fsdp2_streaming_export_matches_reference` — world = GPU count
under real FSDP2; merged checkpoint compared tensor-by-tensor against a
single-process export of the same seeded model, index asserted to span
more than one shard file. Covers the ownership split and the cross-rank
index merge.
- `test_fsdp2_gathered_pack_matches_reference` — gathered packing vs
whole-weight packing over nvfp4 / fp8 / int8 / fp8-per-channel-per-token
/ int4-blockwise
- World=1 CPU tests for unit enumeration, shard sub-splitting by
`max_shard_size`, tied-alias dropping, and `extra_state_dict`
- `tests/unit/torch/utils/test_perf.py` — sampling and threshold logic
of `maybe_clear_cuda_cache`, which replaces an unconditional
`torch.cuda.empty_cache()` after every packed weight

Before your PR is "*Ready for review*"

Make sure you read and follow Contributor guidelines and your commits
are signed (git commit -s -S).

Make sure you read and follow the Security Best Practices (e.g. avoiding
hardcoded trust_remote_code=True, torch.load(..., weights_only=False),
pickle, etc.).

- Is this change backward compatible?: ✅ <!-- No API change; non-FSDP2
and offloaded export paths are unchanged. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ✅

### Additional Information

Pre-export time (load, quantizer insertion, calibration) is now the
dominant phase of an FSDP2 PTQ run. It is unaffected by this PR and
addressed separately in #2359.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved FSDP2 Hugging Face checkpoint export by streaming rank-owned
layers to shard files, reducing export time and host-memory usage for
large models.
- Improved export handling for multimodal models with missing
architecture metadata.
- Standardized weight packing across supported parallelism
configurations.

- **Performance**
  - Reduced repeated module lookups and unnecessary CUDA cache cleanup.
  - Improved shard generation, merging, and index completeness.

- **Compatibility**
- Expanded support for FSDP2 quantized export workflows and shard-local
packing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Signed-off-by: sugunav14 <178320438+sugunav14@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 07:20:45 +00:00
ZhiyuandClaude Opus 5 613e5e8b2e feat(quantization): PTQ support for Step-3.7 MoE checkpoints (#2202)
### What does this PR do?

Type of change: New feature

Adds PTQ support for **Step-3.7** (`stepfun-ai/Step-3.7-Flash`).
Follow-up to [NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665) /
[OMNIML-5583](https://jirasw.nvidia.com/browse/OMNIML-5583): with the
export crash fixed in #2071 the run completes, but the checkpoint it
writes is silently unquantized —

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Two independent causes, both from Step's `trust_remote_code` modeling
code.

**1. The expert weights were invisible to quantization.** Step-3.5 and
Step-3.7 ship the same custom `MoELinear`: a plain `nn.Module` holding
one 3-D `weight` of `[num_experts, out_features, in_features]`, whose
`forward(x, expert_id)` runs `F.linear` against the selected slice. It
is not an `nn.Linear`, and the weights sit on the projection submodule
rather than on the expert container, so neither the plain-linear path
nor `_fused_experts_wrapper_class` (which wants a 3-D `down_proj`
*Parameter*) claims it.

The `_QuantMoELinear` wrapper that handles exactly this layout has
existed since #1063, but its registration was gated on the Step-3.5
class names:

```python
if type(model).__name__ not in ("Step3p5ForCausalLM", "Step3p5Model"):
    return
for module in model.modules():
    if type(module).__name__ == "Step3p5MoEMLP":
```

Step-3.7's root is `Step3p7ForConditionalGeneration` and its container
is `Step3p7MoEMLP`, so it returned immediately and no expert ever got a
quantizer. Detection is now **structural** — a 3-D `weight` plus
`num_experts` / `in_features` / `out_features` and a
two-positional-argument forward — so any Step revision (or another model
shipping this layout) is picked up without a third hardcoded name.
`_reconstruct_fused_moe_linear` likewise matches the wrapper type
instead of the generated `QuantMoELinear` class name; a model whose
class is spelled differently would otherwise quantize fine but export
unusable per-expert keys.

**2. Step's module names don't match the general recipes.** The MoE
block is `moe` and the dense sibling is `share_expert`, so
`*.experts.*`, `*block_sparse_moe*` and `*mlp*` reach none of the routed
experts. This PR ships
`huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast` and
`huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`, which select `*moe*`
and disable the router (`moe.gate`) and the shared expert — mirroring
the existing Step-3.5 recipe — and documents the naming trap in
`modelopt_recipes/ptq.md`.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --model /local/Step-3.7-Flash --trust_remote_code \
    --recipe huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast \
    --dataset /local/cnn_dailymail --calib_size 32 --export_path /local/Step-3.7-Flash-nvfp4
```

### Testing

- `tests/unit/torch/quantization/plugins/test_moe_linear.py` —
structural detection (positive plus 2-D-weight / wrong-forward
negatives), registration on a Step-3.7-shaped model, per-expert
quantizers with calibrated amax, and reconstruction back to the 3-D
parameter. Plus two export-dispatch tests added from review: the
registry resolves a `Quant_SyntheticMoELinear` to `_export_moe_linear`
(this one fails against the old name-keyed registration), and the
handler fills an unrouted expert's input amax.
- `tests/unit/recipe/test_step3p7_recipes.py` — drives both shipped
recipes over a model mirroring Step's real paths
(`model.language_model.layers[i].{moe,share_expert,mlp}`): routed
experts NVFP4-quantized per expert, router / shared expert / dense MLP /
`lm_head` per recipe scope.

Ran locally (torch 2.11, transformers 5.5.4 — the version in the bug
report): the two new files (12 tests) plus `tests/unit/recipe`,
`tests/unit/torch/quantization/plugins/`,
`tests/unit/torch/export/test_export_weight.py` and
`test_export_registry.py` — 324 passed. Full `tests/unit` (minus
onnx/puzzletron): 2490 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`.

Not run end-to-end on the real Step-3.7-Flash checkpoint (1.4 TB /
8×B200) — QA can re-run against this branch with the recipe above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — Step-3.5 keeps working; the
name gate is replaced by a superset.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Review updates

- **Export dispatch keyed on the class name** (caught in review):
registration was made class-name-independent, but `_export_moe_linear`
in `hf_export_handlers.py` was still registered for the literal string
`"QuantMoELinear"`, so a compatible class under another name bypassed
the input-amax fallback for unrouted experts. The predicate now matches
`_QuantMoELinear` through the MRO (lazy import, since the wrapper lives
in the optional transformers plugin), keeping the name check as a
fallback so the synthetic stand-ins in `test_export_registry.py` /
`test_export_weight.py` still match. This was the same class-name
coupling already fixed in `_reconstruct_fused_moe_linear`, in the file I
hadn't looked at.
- Canonical 2026 license header on the new test file.
- Recipe comments corrected: `*moe*` matches the router, but **not**
`share_expert` (`layers.N.share_expert.*` contains no `moe` segment) —
that entry is an explicit guard, not an override.
- `ptq.md` recommendation scoped to Step-3.7, since Step-3.5 has its own
recipe.

### Additional Information

Pairs with #2203 (fail fast when a quant config matches no weight
quantizer), which turns this class of silent no-op into an error for any
model. Independent branches; either can merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization (PTQ) support for Step-3.7 Flash
models, including per-expert quantization for routed MoE layers.
- Added NVFP4 recipes for expert-only and MLP-only quantization, with
optional FP8 KV-cache support.

- **Bug Fixes**
- Improved detection, quantization, reconstruction, and export of
expert-indexed MoE layers across Step model revisions.
- Added safeguards to prevent quantization of unsupported offloaded
expert weights.

- **Documentation**
- Clarified Step-3.7 recipe selection, calibration guidance, and
quantization exclusions.

- **Tests**
- Added coverage for expert, MLP, router, attention, shared-expert,
export, and calibration behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 18:04:43 -07:00
sugunav14andClaude Opus 5 74a1dce390 Speed up FSDP2 MoE calibration by dropping redundant expert-weight gathers (#2359)
### What does this PR do?

Type of change: Perf enhancement

`promote_static_block_weight_quantizers` runs at the end of
`max_calibrate`. All it does is read
quantizer state -- it takes the weight the iterator hands it and throws
it away. For this it goes through `iter_weights_for_calibration`, and on
a fused-MoE module that iterator yields `weight[idx]`, once per expert.

Under FSDP2 the fused weight is a DTensor spread across ranks, so
`weight[idx]` isn't a cheap view.
Slicing it makes PyTorch pull the whole fused expert weight back from
every rank, just to hand over
one slice that the loop then drops. That happens once per expert, per
projection, per layer -- tens
of thousands of round-trips on a large MoE.

The fix adds `iter_weight_quantizers_for_calibration`:

- The base `QuantModule` implementation delegates to
`iter_weights_for_calibration` and drops
the weight, so every subclass gets a correct implementation for free and
the two iterators
  cannot drift apart.
- Only the fused-MoE class — the one where materializing the weight view
is itself expensive —
overrides it, walking the per-expert quantizer `ModuleList` directly. It
keeps the same skip
condition as the weight iterator, since fetching an attribute is free
and only indexing
  collectives.

`promote_static_block_weight_quantizers` then iterates quantizers
instead of `(weight, quantizer)`
pairs. No other caller changes: the other four call sites genuinely use
the weight.

Why not just wrap the promote loop in
`enable_weight_access_and_writeback`, the way
`weight_only_quantize` does? It would still gather per module for
weights the
loop never reads. `weight_only_quantize` needs the window because it
actually computes amax from the
weight; promote only needs the quantizer.

Also included: a warn-once check if `_amax` is ever a `DTensor`. It
should not be — `_amax` is a
registered buffer, and `fully_shard` shards parameters, not buffers —
which is why the
`reduce_amax` in this loop stays local. If that assumption ever breaks,
the reduction becomes a
per-quantizer collective too, and the warning says so rather than
letting it degrade silently.

### Results

Measured on Qwen-3.8 2.4T at world 64 (8 nodes, B300), NVFP4:

| | pre-export |
|---|---|
| before | 79.3 min |
| after | **19.0 min** |

60.3 min saved, 4.2x. The block is gone rather than shortened -- the run
logs 13 silent minutes
in total, all of it model load. Calibration itself is unchanged at 2:23,
so the saving comes from
the promote loop rather than from work moving elsewhere. End-to-end
projects to 41.7 min.

### Usage

No API change. `mtq.quantize` picks this up automatically.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Additive;
`iter_weights_for_calibration` and all its
  callers that use the weight are untouched.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ Under 0.47 Bug Fixes — the call site dates to 0.46.
- Did you get Claude approval on this PR?: ❌ Not yet.




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Performance**
- Reduced calibration overhead for FSDP2-sharded fused-MoE quantization
by avoiding unnecessary expert-weight gathering when only quantizer
state is needed.
- Preserved quantizer ordering and projection filtering during
calibration.

- **Bug Fixes**
- Added a warning when global amax reduction requires collective
processing for individual quantizers.

- **Tests**
- Added coverage for gated and non-gated fused-expert calibration,
including verification that fused weights are not unnecessarily indexed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 23:28:01 +00:00
yeyu-nvidiaandClaude Opus 4.8 279d510616 fix(specdec): correct resume and bound staging in the vLLM hidden-state dump (#2080)
### What does this PR do?

**Type of change:** Bug fix

Fixes two issues in the vLLM offline hidden-state dump

(`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py`).
Both are invisible on small dumps and only bite at scale, which is why
they survived until
now — they were found while dumping ~194k conversations for a MiniMax-M3
draft.

**1. Resume silently re-processed already-finished work.**

`keep_conversation` skips conversations whose `.pt` already exists, but
that predicate reads
**on-disk state**, which is not part of the fingerprint `datasets`
computes for `filter()`
(it hashes the function and the dataset). With a persistent HF cache
reused across a resumed
or requeued run, the cached *"keep everything"* result from an earlier
run — computed when
few or no `.pt` files existed — is replayed. The run then re-generates
and **overwrites**
conversations it had already completed, and reports `Removed 0
conversations due to existing
output files` while doing so.

Observed on a 194k-conversation dump: ~62k `.pt` rewritten over a
two-hour window with the
total output count completely flat.

Fix: pass `load_from_cache_file=False` so the filter re-checks the disk
on every run.

**2. Staging exhausted `/dev/shm` partway through large dumps.**

The script generated the **entire** dataset before saving anything. The
KV connector stages
each conversation's hidden states under its `shared_storage_path`
(`/dev/shm`, i.e. RAM, by
default) and they are only freed by `cleanup_hidden_states()` in the
save loop — so every
conversation stayed staged simultaneously. On a large dump this exhausts
the space and the
connector starts failing writes:

```
Hidden-states write failed for req_id=...:
  SafetensorError('Error while serializing: I/O error: No space left on device (os error 28)')
```

Fix: generate and save in chunks of `--save-chunk-size` (default 256),
so at most one chunk
is staged at a time. As a side benefit the dump becomes **incrementally
durable** — an
interrupted run (walltime limit, node failure) keeps its finished
conversations and the
resume path above continues from them, instead of losing the whole run's
work.

### Testing

- Reproduced both failures on a 194k-conversation MiniMax-M3 dump (8-way
DP, TP8), and
confirmed both fixes on the same workload: after the change the output
count advanced
monotonically across requeues (123k → 194k) with no rewrites, and
`/dev/shm` stayed bounded
  through completion.
- `pre-commit run --files ...` passes (ruff check/format, mypy, bandit,
license, rst checks).
- Behavior is unchanged for a fresh single-shot dump other than the
chunked generate calls;
  the default `--save-chunk-size 256` is the only new knob.

### Additional Information

Extracted from #1749, which is otherwise superseded by the streaming
DFlash/DSpark path — these
two fixes are model-agnostic and apply to any offline dump, so they are
worth landing on their
own.

### Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No — the failure modes are
multi-process/at-scale (datasets cache reuse across runs, connector RAM
staging) and are not reproducible in the unit-test harness.
- **Did you add or update any necessary documentation?**: Yes —
CHANGELOG entry.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added chunked hidden-state generation for large vLLM offline runs.
* Added a configurable save-chunk size, defaulting to 256 conversations.
  * Enabled incremental saving and resumption of hidden-state outputs.

* **Bug Fixes**
* Improved resume filtering to accurately detect existing output files.
  * Reduced memory usage by saving and releasing each generated chunk.
* Ensured temporary files are cleaned up after interrupted or skipped
saves.
  * Added atomic output-file replacement to prevent incomplete results.
* Added validation to prevent invalid conversation IDs from creating
unsafe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-09 20:08:42 +00:00
yeyu-nvidia 56af187565 ar_validate: fail loudly when every sample fails (#2288)
### What does this PR do?

Type of change: Bug fix

`validate_ar()` catches per-sample exceptions, prints a `WARNING`, and
returns whatever succeeded. When *every* sample failed it returned an
empty list, and the reporting block was guarded by `if results and
accelerator.is_main_process:` — so the script printed no results and
exited **0**. A run where 100% of samples failed was indistinguishable
from a successful one.

This bit us on a real run: an EAGLE3 checkpoint loaded with
`device_map="auto"` was sharded across 8 GPUs, every one of the 80
samples died with `Expected all tensors to be on the same device`, and
the job still exited 0 with no AR number anywhere in the log — the
wrapper stamped it PASS.

Now it raises, so the caller sees a non-zero exit. Any previously
"passing" run that printed no AR number was never meaningful.

### Usage

No API change. Existing invocations are unaffected when at least one
sample succeeds:

```bash
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <ckpt> --steps 3 --osl 1024 --num_samples 80
```

### Testing

Reproduced the silent-pass on a Cosmos3-Nano EAGLE3 checkpoint (80/80
samples failing): before this change the job exited 0 and stamped PASS;
after it, the job exits non-zero with the sample failures visible.
Confirmed the normal path is unchanged by a subsequent run that
completed 80/80 and printed AR 3.42.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — only affects the
all-samples-failed case, which previously produced no output and a
misleading exit 0.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — the failure path requires
a model that errors during AR validation; the existing suite has no
harness for that.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — behavior fix in an example script, not a released-feature change.
- Did you get Claude approval on this PR?: ❌ — not yet run.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Improved validation error handling when all samples fail.
  * Validation now rejects non-positive sample counts before processing.
* Empty validation results are clearly distinguished from cases where
all samples fail.
* Error messages report the actual number of validation samples
attempted, capped at the available dataset size.
  * Empty validation results are no longer reported as successful.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-09-09 12:40:55 -07:00
Keval MorabiaandClaude Opus 5 4956213d67 Unblock Qwen3.5/3.6 QAD: Megatron export, calibration, and distillation fixes (#2334)
### What does this PR do?

Type of change: Bug fix

Everything that blocked running QAD on a quantized Qwen3.5 / Qwen3.6 MoE
checkpoint: two
Megatron-Core → HuggingFace export bugs that make it unservable (§1–2),
the dead code the first
leaves behind (§3), a no-op flag (§4), a multi-GPU calibration deadlock
(§5), and four
distillation / data-prep bugs that stopped QAD itself from running (§6).

#### 1. Routed experts were exported packed, and vLLM cannot load that

```
AttributeError: Layer language_model.model.layers.23.mlp.experts has no parameter
  'w2_weight_weight_scale_2' for checkpoint weight ...experts.down_proj_weight_scale_2
```

`mcore_qwen35vl.py` used `GroupedMLPPacking`, mirroring the **BF16
upstream** checkpoint, which
really is packed. But that mapping is only used for **quantized**
export, and vLLM's quantized MoE
loader needs per-expert scales — both released NVFP4 checkpoints
(`Qwen3.6-35B-A3B-NVFP4` via
hf_ptq, `Nemotron-3.5-Lightning-30B-A3B-NVFP4` via Megatron-LM) are
per-expert.

`_grouped_mlp_slicing` gains `gate_proj_name` / `up_proj_name` to split
each expert's fused gate+up
and slice its per-block `weight_scale`; `GroupedGatedMLPSlicing` wires
it up. The Megatron
checkpoint layout is unchanged, so affected checkpoints need only a
**re-export**.

`_verify_exported_keys` is relaxed to match: exported modules now
contribute their ancestor
prefixes, so expanding one source module into many is not reported as
~82 dropped tensors. A module
genuinely absent still has nothing beneath its prefix and is still
caught.

#### 2. A quantized `output_layer` (`lm_head`) could not be checkpointed

`GPTModel.sharded_state_dict` drops `output_layer._extra_state` and
asserts it is empty. ModelOpt
keeps quantizer state there, so saving raised and — since that method
also backs the load plan —
loading silently restored the layer **unquantized**.

`keep_gpt_output_layer_extra_state()` retains it, applied from
`megatron_replace_quant_module_hook` so **every** Megatron model gets it
(Megatron-LM and NeMo
users included, neither of whom can import `mbridge`, which needs
`megatron.bridge`). It matches
the upstream body by AST before replacing it and self-disables
otherwise.

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is **closed, not
merged**: nemo:26.10 migrates `GPTModel` to `HybridModel`, whose
`sharded_state_dict` has no
pop-and-assert, so this side keeps the workaround.

Not cosmetic: `lm_head` is 248320×2048 = 509M params, **34.6% of
per-token weight traffic** on a
model with ~2.9B active params.

#### 3. Cleanup

Nothing maps `GroupedMLPPacking` once qwen3_5 is switched over; it is
removed with
`_grouped_mlp_packing` and the `quantize=` / `record_quant_config=`
parameters that existed only to
serve it. Llama-4's `PackNameRemapping` is unaffected.

Two smaller review-driven fixes: the gated-split shape checks raise
`ValueError` rather than
`assert` (stripped under `-O`), and per-expert quant metadata is
recorded for
`local_expert_indices` rather than every global id, fixing
non-contiguous EP.

#### 4. Remove the no-op `--moe_calib_experts_ratio` from the Megatron
quantize example

`examples/megatron_bridge/quantize.py` accepted the flag and threaded it
into the `mtq` config, but
`_moe_calib_experts_ratio` exists only in `plugins/huggingface.py` (9
refs) and never in
`plugins/megatron.py` (0); `mode.py:247` only assigns it to modules
already exposing the attribute.
On a Megatron MoE model it was accepted and silently ignored — a trap,
since on a 256-expert model
it reads like a major quality lever. `hf_ptq.py` keeps it, where it
works.

#### 5. Fix multi-GPU image-text (VLM) calibration deadlocking

VLM calibration hung for 30 minutes and died on a gloo timeout whenever
`world_size > 1`, with no
error until the timeout fired.

`NemotronTarPlusJsonlIterable` split its budget with truncating
division, so the stream supplied
fewer samples than requested (1024 over 3 subsets → 341×3 = **1023**).
`_ShardedIterable` gives
rank *r* items *r, r+W, r+2W…*, so a stream that is not a multiple of
`world_size` leaves the
trailing rank one short — it exits the forward loop early and the others
block on the next
collective. The arithmetic predicts both observed hangs exactly: 1024 →
stall at **255/256**,
512 (yielding 510) → **127/128**.

Fixed both ends: subset budgets are distributed with `divmod` so they
sum exactly, and
`_ShardedIterable` truncates every rank to `floor(len / world)` — which
also covers `num_samples`
not being divisible by `world_size`, as the first fix alone does not.

Verified on Qwen3.6-35B-A3B (EP=4, `nemotron_vlm_dataset_v2`, 1024
samples): the configuration that
hung twice now completes 256/256 and exports. Unit tests cover both
fixes and fail without them.

#### 6. Fix the distillation path so QAD can actually run

Four independent bugs, all hit while running QAD end to end on
Qwen3.6-35B-A3B. Each blocks a
different configuration, and together they made every sequence length
OOM or abort.

- **Context parallel aborts.** The DDP config derived
`average_in_collective` from `--sft` alone,
but context parallel also needs per-token loss reduction, so any
`--cp_size > 1` run died on
  `Cannot average in collective when calculating per-token loss`.
- **`TopKLogitsKLLoss` was not memory-efficient.** Despite documenting
"without gathering full
logits", it cast the *whole* vocabulary to FP32 before selecting the
top-k, allocating two
`[seq, vocab]` tensors — 30.3 GiB each at seq 32768 on this model's 248k
vocab. Reducing before
the cast is equivalent: widening is exact and temperature scaling is
monotonic, so the selected
  entries and the loss are unchanged.
- **MTP cross-entropy ran when it had nothing to recover.**
`skip_lm_loss` exempts the MTP heads
unconditionally, so their CE materialised another FP32 `[seq, vocab]`
tensor even when the MTP
head is excluded from quantization — as it is in every recipe here (775
of 906
`exclude_modules`, zero MTP `weight_scale` tensors exported). It is now
skipped **only** when the
model is quantized and MTP is left out of it; plain distillation such as
pruning recovery still
trains the MTP head. `test_mtp_excluded_from_quantization` pins all four
cases.
- **One bad record deadlocked data prep.** `megatron_preprocess_data`
re-raised chat-template
failures out of a pool worker, stalling the whole job until it timed out
— three malformed
records cost a multi-hour tokenization run. They are now skipped with a
warning, matching the
  existing handling of malformed JSONL a few lines above.

Also exposes `--logit_kl_topk`, which `DistillationConfig` has supported
for a while but the
example never passed through; `test_qad` now exercises that path.

§4, §5 and §6 are independent of §1–3; happy to split them out if
reviewers prefer.

### Usage

No API change. Exported names now match the released checkpoints:

```
model.language_model.layers.0.mlp.experts.<E>.{gate,up,down}_proj.{weight,weight_scale,weight_scale_2}
lm_head.{weight,weight_scale,weight_scale_2}
```

### Testing

- `test_mcore_export_mappings.py` — qwen3_5 mappings emit per-expert
rules. Verified these fail
without the fix (2 failed / 11 passed), with `Qwen3MoeForCausalLM` /
`NemotronHForCausalLM` as
  controls.
- `test_unified_export_megatron.py` — the gate/up split, per-block scale
slicing, the 0-dim scalar
fallback, and both directions of the `_verify_exported_keys` relaxation.
- `test_megatron.py::TestKeepGptOutputLayerExtraState` — 15 cases:
payload detection, no-op second
call, warn-and-skip on an unrecognised `sharded_state_dict`, and
`test_patches_stock_megatron_core`
which installs a replica of the real pre-fix upstream body (verified
against `be08ce5b1~1`) so the
  patched path is exercised whichever megatron-core is installed.
- `test_qad.py` — CI caught that its reference comparison still assumed
packed experts; fixed.

**End to end on `Qwen/Qwen3.6-35B-A3B` (35B MoE, 256 experts), 4×GB200,
nemo:26.08:**

| | before | after |
| --- | --- | --- |
| export self-check | `Export dropped 82 tensor(s)` | passes |
| expert tensors | `mlp.experts.gate_up_proj` (packed) |
`mlp.experts.<E>.{gate,up,down}_proj` |
| **vLLM v0.28.0 load** | **`AttributeError`, engine never starts** |
**`Loading weights took 25.61 s`** |
| **NEL eval (GPQA-D, MMMU-Pro)** | **FAILED** | **SUCCESS** |

### Results these fixes unblocked

The export fix is what made a Megatron-produced NVFP4 MoE checkpoint
servable at all, so it enabled
a full PTQ study on Qwen3.6-35B-A3B. Accuracy deltas are against a BF16
baseline measured on the
same harness, from **paired** per-question tests:

| recipe | throughput vs BF16 | GPQA-D | SciCode ×8 | MMMU-Pro | IFBench
|
| --- | --- | --- | --- | --- | --- |
| **W4A16** (weight-only) | **0.64–0.86×** — *slower* | −0.06 | −0.15 |
+0.48 | −0.44 |
| **W4A4** | 8/12 shapes faster | −0.60 | −0.70 | −1.48 (p=0.019) |
−0.53 |
| **W4A4 + 4-bit `lm_head`** | **9/12 shapes**, up to **1.30×** |
**+0.03** (p=0.96) | −0.81 | **−1.16** (p=0.016) | −1.65 (ns) |

Repeats: GPQA-D is `pass@1[avg-of-16]`; SciCode is 8 pooled runs per
recipe; MMMU-Pro is 3 runs per
side and IFBench 2–3 for BF16 and the last row, 1 elsewhere. AA-LCR
(68.33 → 71.33, p=0.25, 3 runs
per side) and τ²-Telecom (94.25 → 94.25, 3 runs per side) are on par; at
100 questions and 114
tasks they cannot resolve below ~5 pp and ~3 pp, so they carry no claim
either way.

#### QAD status (what §6 unblocked)

With the §6 fixes in place, QAD runs end to end on this model: 32 nodes,
`TP=1 PP=1 CP=1 EP=8`,
seq 32768, gbs 512, ~38 s/iter, 124 GB/GPU peak. First accuracy read,
MMMU-Pro at iteration 50
(0.84 B tokens), 3 runs per side, paired per-question:

| | MMMU-Pro | vs BF16 |
| --- | --- | --- |
| BF16 | 74.55 | — |
| W4A4 + 4-bit `lm_head` (PTQ) | 73.39 | **−1.16, p=0.016** |
| + QAD, iteration 50 | 73.78 | −0.77, p=0.089 (ns) |

The PTQ deficit that motivated this work is no longer statistically
significant after 50 QAD
iterations. The improvement itself (+0.39 vs PTQ) is **not** significant
at p=0.41, and 50
iterations is 10% of the planned budget, so this is a direction rather
than a result. A full
six-benchmark sweep at iterations 50 and 300 is running; these numbers
will be superseded.

Two findings worth flagging beyond this PR:

- **Weight-only NVFP4 is slower than BF16 on Blackwell.** W4A16 leaves
activations in BF16, so vLLM
cannot use the FP4 tensor cores and falls back to
`MarlinNvFp4LinearKernel` / `'MARLIN'` MoE.
W4A4 selects `FLASHINFER_TRTLLM` + `FlashInferCuteDslNvFp4LinearKernel`
and beats W4A16 in
**12/12** shapes. The Marlin line count tracks the recipe exactly (one
W4A16 layer ⇒ one Marlin
  line ⇒ zero once `lm_head` is W4A4).
- **The only accuracy cost is multimodal**: **−1.2 pp on MMMU-Pro** for
the fastest recipe,
confirmed over 3 runs per side (p=0.016). GPQA-D, SciCode, IFBench,
AA-LCR and τ²-Telecom show no
  significant regression.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Megatron checkpoints
unaffected; re-export to gain the loadable layout. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.47.0 → Bug Fixes; includes the removed flag, since passing it
now errors instead of being ignored -->
- Did you get Claude approval on this PR?: ✅ <!-- Reviewed; all findings
addressed, threads resolved. -->

### Additional Information

Both export bugs were found while reproducing
`nvidia/Qwen3.6-35B-A3B-NVFP4` through
`examples/megatron_bridge/`. Follow-up to #2332. Upstream counterpart

[NVIDIA/Megatron-LM#7086](https://github.com/NVIDIA/Megatron-LM/pull/7086)
is closed — see §2.
Labeled `cherry-pick-0.47.0`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-09 23:19:56 +05:30
2d35643452 LiLiCorr training (#2342)
### What does this PR do?

Type of change: new feature

Adds **LiLiCorr**, a candidate-lattice reranker for DFlash drafts, as a
new `projector_type` on the
existing `dflash` mode — plus three DFlash-wide improvements that apply
to every variant, and an
optional composition with DFlash2's grouped convolutions.

A DFlash drafter is trained on per-position marginals rather than on the
joint block distribution,
so its drafted tokens are individually plausible yet jointly incoherent.
LiLiCorr keeps the top-`k`
candidates the backbone already produces at each block position, scores
transitions between adjacent
candidates with a small two-layer transformer, and commits a path
through the lattice greedily.
Serving is unchanged in kind: verify still checks every drafted token
against the target, so the
emitted distribution is untouched and only acceptance length moves.

- Paper: [LiLiCorr: Lightweight Likelihood Correlation of Parallel
Drafts for Speculative Decoding](https://arxiv.org/abs/2608.20530)
(arXiv:2608.20530)
- Blog: https://research.nvidia.com/labs/nemotron/lilicorr/
- **Companion PR — serving support:**
[sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462)

This PR is the **training** half. It trains the drafters and exports
them; the companion PR above is
what serves the resulting checkpoints, and is what the comparison table
below was measured through.

**What is in the commits**

| | |
| --- | --- |
| LiLiCorr draft variant | `hf_lilicorr.py`, `modeling_lilicorr.py`,
conversion routing, config fields, export |
| Three DFlash-wide features | fp32 master weights for the draft, draft
activation checkpointing, and a DDP hang fix — all default-off or
behaviour-preserving, all applying to `dflash`, `domino`, `dspark` and
`dflash2` alike |
| Optional grouped convolutions | composes LiLiCorr with DFlash2's
`DFlashGroupedConv`; see the dependency note below |
| Two recipes | `lilicorr.yaml` and `lilicorr_conv.yaml` |
| CPU unit tests, CHANGELOG, one launcher example | |

**⚠️ The convolutions depend on the DFlash2 branch, and cannot run until
it merges.**

`modeling_lilicorr.py` imports `DFlashGroupedConv` from
`modeling_dflash2`, which today exists only
on `haoguo/dflash2-support`. The class is **imported rather than copied
on purpose** — it is the only
way the two variants cannot drift apart arithmetically — but the
consequence is that the
convolutional recipe cannot run against `main` as it stands.

So the import is **deferred into `_install_sublayer_convs`** rather than
taken at module scope.
Everything else in this PR, including the plain LiLiCorr reranker, has
no DFlash2 dependency at all
and works on `main` today; an eager import would have made the whole
plugin unimportable for the sake
of one optional feature. Requesting the convolutions without DFlash2
present raises an `ImportError`
naming the two config keys to remove, rather than failing at import
time.

**This PR carries two of @h-guo18's commits, with authorship and
sign-off preserved.** Both are
independent of DFlash2 itself and both are needed here:

- `1419d47e`, the no-op sublayer seam. Without it
`DFlashDecoderLayer.forward` never calls the
wrappers the convolutions install onto, so the modules would be built,
counted and exported while
  computing nothing. It is arithmetically an identity on its own.
- `ba377e7a`, the RoPE-θ fix. On Transformers 5 a config carries both a
top-level `rope_theta` and a
`rope_parameters` dict; the real base lives in the dict while the class
default (10,000 for Qwen3)
stays visible as the flat attribute. Reading the flat field first builds
a draft whose RoPE base is
100× off a Qwen3-8B target's, which trains and exports without
complaint. Both the training-side
enforcement and the exporter's `_get_rope_theta` are affected on `main`
today.

Both are @h-guo18's work and belong to their branches; they are carried
here only so that this PR
stands on its own. **If those branches land first, this PR can be
rebased onto them and the two
commits dropped**, and they can equally be split out now if that is
easier to review.

The same applies to `dflash_fp32_master_weights`, which is also in
flight on
`haoguo/dflash-fp32-master-weights`. The field name is shared
deliberately so that there is only
ever one knob rather than two spellings of it, and both versions default
to off. Whichever lands
first, this PR can be rebased onto it.

### Usage

Train with the shipped recipe:

```python
from modelopt.recipe import load_recipe

config = load_recipe("general/speculative_decoding/lilicorr.yaml")
# Qwen3-8B target, 6 epochs, block size 16 (15 drafted slots, 16 verified),
# DFlash decay objective at gamma 7.0, fp32 master weights for the draft.
```

Or convert directly:

```python
import modelopt.torch.speculative as mtsp

config = {
    "dflash_block_size": 16,
    "dflash_loss_objective": "decay",
    "dflash_loss_decay_factor": 7.0,
    "dflash_fp32_master_weights": True,
    "dflash_lilicorr_w_ce": 0.25,
    "dflash_lilicorr_w_margin": 0.0,
    "dflash_lilicorr_w_pen": 0.25,
    "dflash_architecture_config": {
        "num_hidden_layers": 5,
        "projector_type": "lilicorr",
        "lilicorr_candidate_topk": 8,
        # Optional, and all-or-nothing: adding these two keys wraps every draft
        # sublayer in DFlash2's grouped convolution. Requires the DFlash2 variant.
        # "conv_kernel_size": 2,
        # "conv_group_size": 16,
    },
}
mtsp.convert(model, [("dflash", config)])
```

### Results

Six drafters for a **Qwen3-8B** target, all trained **in ModelOpt on one
matched contract** — the
same corpus, schedule and block geometry for every arm, so no row
carries a training advantage.
Training data is NVIDIA's
[Nemotron Post-Training Dataset
v2](https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2)
with the multilingual split excluded, generated from the target with
**thinking disabled**;
**6 epochs**; block size 16 (15 drafted slots, 16 verified); DFlash
decay objective at gamma 7;
**8 nodes × 8 H100, global batch size 64** (one sequence per device, no
gradient accumulation).

All six were then exported and served through SGLang on a **single H100
80GB**, `tp_size 1`, at
concurrency 1, greedy, `fa3`, mean of two replicates, with the whole
node held exclusive per
benchmark. Speedup is output tokens/s against an autoregressive baseline
measured in the same
allocation.

Cells are `acceptance length / speedup-vs-AR`; **★ fastest, ☆ second
fastest**:

| benchmark | LiLiCorr+conv | LiLiCorr | DSpark | DFlash2 | Domino |
DFlash |
|---|---|---|---|---|---|---|
| gsm8k | ★ 7.715 / 5.26x | ☆ 7.557 / 5.22x | 7.375 / 4.86x | 7.252 /
5.06x | 7.225 / 4.87x | 6.341 / 4.59x |
| math500 | ★ 9.241 / 6.54x | ☆ 9.064 / 6.52x | 9.012 / 6.15x | 8.999 /
6.49x | 8.976 / 6.25x | 7.909 / 5.88x |
| aime25 | ★ 8.285 / 6.03x | ☆ 8.156 / 6.03x | 8.043 / 5.61x | 7.967 /
5.91x | 8.066 / 5.77x | 7.126 / 5.44x |
| humaneval | ★ 7.393 / 4.01x | 7.077 / 3.93x | 7.163 / 3.72x | ☆ 7.081
/ 3.95x | 6.864 / 3.73x | 6.156 / 3.68x |
| mbpp_sanitized | ★ 5.999 / 4.18x | ☆ 5.849 / 4.13x | 5.888 / 3.95x |
5.685 / 4.05x | 5.679 / 3.91x | 5.027 / 3.70x |
| livecodebench | ★ 7.975 / 5.40x | ☆ 7.754 / 5.33x | 7.775 / 5.10x |
7.601 / 5.26x | 7.553 / 5.04x | 6.808 / 4.88x |
| alpaca_eval | ☆ 3.697 / 2.69x | ★ 3.656 / 2.70x | 3.588 / 2.52x |
3.467 / 2.58x | 3.627 / 2.59x | 3.222 / 2.46x |
| mtbench | ★ 4.014 / 2.94x | ☆ 3.939 / 2.93x | 3.957 / 2.78x | 3.748 /
2.80x | 3.948 / 2.84x | 3.478 / 2.67x |

**Against every other approach in the table, LiLiCorr with convolutions
is the fastest on all eight
benchmarks.** Plain LiLiCorr is the fastest on seven of the eight; the
exception is humaneval, a
164-prompt slice, where DFlash2 is ahead by 0.5%.

`DFlash` is the deliberately head-free control; every head clears it by
+7.60% to +21.67% on
acceptance, which is the check that a head actually loaded. Reproducing
the `LiLiCorr+conv` column
additionally needs the DFlash2 variant.

Acceptance length is bit-reproducible under greedy decoding and its
replicate spread here was 0.00%
on every benchmark; throughput has a ~0.2% floor.

### What `dflash_fp32_master_weights` does, and what it is worth

Today the draft is cast to the frozen base model's dtype — bf16 — before
the optimizer is built.
AdamW then allocates its moments with `zeros_like(p)`, so the
**optimizer state becomes bf16 too**.
That is the problem: bf16 has too few mantissa bits to represent the
small updates Adam's second
moment accumulates, so those updates round away and the effective step
size decays on its own,
independently of the learning-rate schedule.

The flag is standard mixed precision instead: the draft's master weights
stay in fp32 while the
matmuls run in bf16. It requires a bf16 autocast around the forward,
which HF `Trainer` supplies
under `TrainingArguments.bf16`. Paths that do not go through the Trainer
— evaluation,
`pseudo_speculative_generate`, a plain `convert()` and forward —
currently need the caller to
supply it, and no shipped recipe exercises those (`estimate_ar: false`,
`do_eval: false`). Making
the draft supply its own autocast is a follow-up, held back from here on
review because it touches
every DFlash variant and wants e2e coverage of the existing recipes.

Compute speed is unchanged. The cost is memory, about 12 bytes per
parameter for the weight plus
Adam's two moments instead of 6, plus a doubled gradient all-reduce
under DDP, since fp32
parameters mean fp32 gradients. Under FSDP2 that second cost is what
`MixedPrecisionPolicy(reduce_dtype=...)` exists to control.

It is worth **7 to 14 percent of acceptance length**, measured at the
end of training on gsm8k, and
it helps every projector type:

| arm | bf16 | fp32 | Δ acceptance length |
| --- | ---: | ---: | ---: |
| LiLiCorr | 6.8670 | 7.5573 | **+10.05%** |
| DFlash2 | 6.7396 | 7.2518 | **+7.60%** |
| Domino | 6.5854 | 7.2252 | **+9.71%** |
| DSpark | 6.4621 | 7.3752 | **+14.13%** |
| DFlash | 5.9030 | 6.3412 | **+7.42%** |

Every arm in the comparison table above was trained with it on, and
**both shipped recipes set it
`true`**, so the documented path gets it.

It defaults to **off**, so no existing DFlash, Domino or DSpark run
changes behaviour. Both shipped
LiLiCorr recipes set it `true`, which is the arithmetic their numbers
were trained with. Flipping
the default is a reasonable follow-up once the autocast above is in.

The draft is drawn in fp32 and, under this flag, kept there; an
unpromoted run rounds the same draw
to the base model's dtype. So the bf16 and fp32 rows of the table above
start from the same
initialization at the precision each trains in, rather than from two
different draws. A unit test
pins that.

The flag also survives a resume. `modify()` runs under `from_pretrained`
with the base model still
on meta and cannot place the draft at all, so `restore_draft_precision`
re-applies the dtype, the
device and the rotary buffer once the weights are loaded and before the
Trainer builds the
optimizer — the last point that can still decide the Adam moment dtype.
It also reloads the draft's
tensors at the dtype they were saved in, since checkpoints store the
draft in fp32 while the base is
bf16 and `dtype="auto"` gives every tensor one dtype.

@h-guo18 has the same field in flight on
`haoguo/dflash-fp32-master-weights`, plus an
HF-format-resume fix this PR does not have. The name is shared
deliberately so there is only ever
one knob; whichever lands first, the other should be dropped rather than
merged.

### Testing

- **257 CPU unit tests pass** across `tests/unit/torch/speculative/`,
including the existing DFlash,
Domino, DSpark and Eagle suites. 48 of them are new and cover LiLiCorr
specifically: conversion
routing, head geometry, the required-field validation, the three-term
objective and its absolute
weights, gradient reach into both the head and the drafter body, and the
export contract.
- Both recipes load and validate through `modelopt.recipe.load_recipe`.
- The three DFlash-wide changes are covered behaviourally: the fp32 flag
is checked on the
optimizer's moment dtypes rather than only on parameters, since the
moments are the point of the
change, and on the initialization described above; activation
checkpointing is asserted to leave
draft gradients bit-identical with the flag on and off; and the rotary
buffer is asserted present
after `modify()` on a real device while still deferred on meta, which is
the case the laziness
  existed for.
- The resume path has its own test: after a `save_pretrained` /
`from_pretrained` round trip,
`restore_draft_precision` is asserted to return the draft to fp32 with
its stored weights intact
and its Adam moments in fp32. Without it the draft comes back in the
base dtype with the flag
  still set, which is the failure it exists to prevent.
- `TestDFlashLazyRotaryEmb` was updated rather than left passing: it
asserted the rotary buffer does
*not* exist after convert, and the DDP fix deliberately changes that on
non-meta devices. The
  replacement pins the refined invariant in both directions.
- The published checkpoints were trained with this arithmetic, verified
rather than assumed: a
fingerprint over draft initialisation, loss and gradients is compared
against the pre-review tree
for both `dflash` and `lilicorr`. Loss and gradients are **bitwise
identical**. Initialisation
moves, by less than bf16 resolution, and that is the single-dtype change
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — every addition is opt-in. The
new `projector_type` is
selected only by config, `dflash_fp32_master_weights` defaults to off,
and the
activation-checkpointing and DDP fixes preserve behaviour. No existing
default changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in
`CONTRIBUTING.md`: ✅ — no new dependencies. Four files carry `# Adapted
from
https://github.com/sgl-project/SpecForge/...` headers for the DFlash
backbone and loss they derive
from (Apache-2.0), matching the attribution already on `hf_dflash.py` in
this repo. The two
commits described above are @h-guo18's, cherry-picked with authorship
and sign-off preserved.
- Did you write any new necessary tests?: ✅ — 48 new CPU tests, plus the
updated rotary test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
once opened.

### Additional Information

The convolutional recipe is the memory worst case: at an 8B target,
combined with fp32 master
weights, it may need `training.gradient_checkpointing: true` to fit on
80 GiB, and it fits without at
4B. Checkpointing is mathematically neutral — same objective, same data
order, same resulting model —
but it trades step time for memory, so a run using it is not
step-time-comparable with one that does
not. The recipe header says so.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added LiLiCorr speculative decoding with candidate-lattice reranking,
configurable objectives, metrics, export support, and optional grouped
convolutions.
* Added FP32 master-weight support with improved mixed-precision
behavior and gradient checkpointing.
* Added LiLiCorr training recipes and a Qwen3-8B launcher configuration.
* **Bug Fixes**
* Improved rotary-embedding configuration handling and corrected DFlash
distributed-training hangs.
* Added validation for invalid LiLiCorr configurations and improved
exported reranking metadata.
* **Documentation**
* Expanded guidance for FP32 master weights, training workflows, and
LiLiCorr configuration.
* **Tests**
* Expanded coverage across training, evaluation, generation, export, and
checkpoint workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: mrusanovsky <mrusanovsky@nvidia.com>
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 00:35:31 +08:00
haoxiz-nvidia acdf330414 Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do?

Type of change: new feature

Adds opt-in support for using the standalone TensorRT-RTX ABI Execution
Provider during ModelOpt ONNX quantization.

Users select the ABI backend with:

`--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi`

When selected, ModelOpt imports and registers the installed TensorRT-RTX
ABI provider before creating the ONNX Runtime inference session. The
backend selection is propagated through INT8, FP8, and INT4 AWQ
calibration paths, including the Windows GenAI LLM quantization example.

The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward
compatible. The `legacy` backend is still the default and continues to
use TensorRT-RTX libraries supplied through `PATH`.

For Windows x64 with Python 3.11 or newer, the ONNX dependencies now
include:

- `onnxruntime-gpu~=1.26.0`
- `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0`

Keeping `onnxruntime-gpu` allows users to select either CUDA EP or
TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build
instructions are intentionally out of scope and will be documented
separately.

### Usage

```powershell
python -m modelopt.onnx.quantization `
  --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" `
  --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" `
  --quantize_mode=int8 `
  --output_path="C:\path\to\int8_abi\model.onnx" `
  --calibration_eps=NvTensorRtRtx `
  --trt_rtx_backend=abi `
  --use_external_data_format `
  --high_precision_dtype=fp32 `
  --log_level=INFO

### Testing
unit test have been added

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ 
- Did you get Claude approval on this PR?: pending



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

- **New Features**
  - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64.
  - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default.
  - Added validation for unsupported backends and incompatible TensorRT plugin configurations.
  - Updated Windows ARM64 installation support and platform-specific package configuration.

- **Documentation**
  - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps.
  - Documented the new TensorRT-RTX backend command-line option.

- **Tests**
  - Added coverage for ABI provider registration, backend validation, and compatibility checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-09 04:04:46 +00:00
Shengliang Xu fbd5e942e2 Add NVFP4 PTQ recipe for zai-org/GLM-5.3-Flash (experts + dense MLP) (#2312)
### What does this PR do?

Type of change: new feature (model recipe)

Adds an NVFP4 PTQ recipe for
**[zai-org/GLM-5.3-Flash](https://huggingface.co/zai-org/GLM-5.3-Flash)**.

GLM-5.3-Flash is a `glm5_next` VLM MoE — 45 decoder layers, 288 routed
experts, and hybrid attention: KDA (linear-attention) layers interleaved
with NoPE sparse-MLA layers. It requires `transformers >= 5.16.1`;
earlier releases cannot parse the config.

**`nvfp4_experts_dense_mlp-kv_fp8_cast`** applies:

| component | precision |
|---|---|
| routed experts (layers 3–44, 288 each) | NVFP4 W4A4 |
| dense MLP (layers 0–2, 9 modules) | NVFP4 W4A4 |
| KV cache | FP8 (cast mode, constant amax) |
| shared experts, router gate, KDA + MLA attention, vision tower,
embeddings, `lm_head` | BF16 |

`mlp_layer_types` marks only layers 0–2 `dense` and 3–44 `sparse`, so
the dense-MLP scope adds just 9 modules (`mlp.gate_proj` / `mlp.up_proj`
/ `mlp.down_proj`) on top of the routed experts. The recipe starts from
`base_disable_all`, so only the listed globs re-enable anything.

> **On scope / why only one recipe.** An earlier revision of this PR
also shipped a model-specific `nvfp4_experts_only-kv_fp8_cast`. It was
removed: on this model it enables the **identical** quantizer set as the
general `general/ptq/nvfp4_experts_only-kv_fp8_cast` (its
`*block_sparse_moe*` entries are no-ops here and
`default_disabled_quantizers` is redundant in experts-only scope), so it
wasn't a model-specific deviation. For plain experts-only NVFP4, use the
general recipe. The genuine model-specific delta — shipped here — is the
dense-MLP scope plus the vision-tower exclusion below.

#### The load-bearing `*visual*` disable

The vision tower reuses the language model's leaf names —
`model.visual.blocks.<N>.mlp.gate_proj` and friends, across 24 blocks —
so the dense-MLP patterns match **144 modules inside `model.visual.*`**.
Entries apply in order, so a trailing `{quantizer_name: '*visual*',
enable: false}` is what keeps them BF16, and it has to stay last.
(`*.experts.*` needs a literal `.experts.`, so it never reaches the
vision tower.)

The shared `default_disabled_quantizers` unit is deliberately not
imported: for this model only its `*visual*` pattern changes anything —
every other pattern either matches no module here, or matches one that
`base_disable_all` already left off (`lm_head`, the `mlp.gate.` routers)
and that nothing re-enables.

#### Two model-specific points, documented in the file header

- **`layerwise.enable=false` is required, not incidental.** This is a
VLM, so the decoder layers nest under `model.language_model.layers` and
`layerwise_calibrate` cannot locate them.
- **The MTP head is not built, so it is neither quantized nor
exported.** The config declares `num_hidden_layers: 45` (with
`num_nextn_predict_layers: 1`), so the HF model class instantiates
decoder layers 0–44 only and never constructs the MTP layer.

Filed under `modelopt_recipes/models/` per the split introduced in
#2219, keyed by the source hub model — alongside `moonshotai/Kimi-K3`
and `mistralai/Mistral-Medium-3.5-128B`. There is no published
`nvidia/GLM-5.3-Flash-NVFP4` yet; the `models/` section explicitly
covers "published **(or planned)**" checkpoints.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <zai-org/GLM-5.3-Flash checkpoint> \
    --recipe models/zai-org/GLM-5.3-Flash/ptq/nvfp4_experts_dense_mlp-kv_fp8_cast \
    --export_path <output>
```

### Testing

- **`tests/unit/recipe/test_glm_5_3_recipe.py`** (new) — applies the
recipe to a tiny `glm5_next`-like VLM MoE and asserts the
enabled/disabled state per module: routed experts + dense MLP → NVFP4;
vision tower, shared experts, router gate, KDA `conv1d`, MLA attention
and `lm_head` → BF16. This pins the wildcard precedence — in particular
that the trailing `*visual*` disable keeps the vision tower BF16 even
though it reuses the dense-MLP leaf names, and that `*mlp.gate_proj*`
doesn't catch the router `mlp.gate`.
- **`tests/unit/recipe/test_recipe_docs.py`** — all checks pass,
including `test_every_model_specific_ptq_dir_is_mentioned` (the
`models/zai-org/GLM-5.3-Flash/ptq/` folder appears in `ptq.md`).

The recipe's scope was also checked against the model's actual module
names: the dense-MLP patterns match 144 modules under `model.visual.*`,
which the trailing disable returns to BF16; an exported checkpoint
carries `input_scale` / `weight_scale` / `weight_scale_2` on
`layers.0–2.mlp.*_proj` while `visual.blocks.0.mlp.gate_proj` retains
only `.weight` / `.bias`; and `kv_cache_quant_algo: FP8` survives the
trailing disable.

### Before your PR is "*Ready for review*"

- Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed.
- Is this change backward compatible?: ✅ (one new recipe + one new unit
test; `ptq.md` updated)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_glm_5_3_recipe.py` pins the recipe's wildcard
precedence
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — consistent with the recent recipe additions (#2219, #2269, #2287),
which did not add entries
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Source model: https://huggingface.co/zai-org/GLM-5.3-Flash

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added post-training quantization recipes for GLM-5.3-Flash.
* Supports NVFP4 W4A4 quantization for routed experts, with an
additional configuration covering dense MLP layers.
  * Enables FP8 key-value cache casting.
* Uses maximum-based calibration with layerwise calibration disabled for
the VLM layout.
* Retains BF16 precision for shared experts, vision components,
attention, embeddings, routing, language head, and MTP components.

* **Documentation**
* Documented the available GLM-5.3-Flash quantization configurations and
precision assignments.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-08 18:00:59 -07:00
Shengliang Xu 3f4b95e8a4 Add the NVFP4 PTQ recipe for Qwen/Qwen3.8-2.4T-A95B (#2302)
### What does this PR do?

Type of change: new feature (model recipe)

Adds the NVFP4 PTQ recipe for **Qwen/Qwen3.8-2.4T-A95B** — the recipe
used to produce

[**nvidia/Qwen3.8-2.4T-A95B-NVFP4**](https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4).

Qwen/Qwen3.8-2.4T-A95B is a `qwen3_5_moe_text` MoE: 92 layers, 512
routed experts (top-10)
plus a shared expert, with **hybrid attention** — gated-delta
(linear-attention) layers
interleaved with full-attention layers. It is transformers-native from
>= 5.9 and its config
ships `base_model_ep_plan`, so no ModelOpt plugin is required.

The recipe applies:

| component | precision |
|---|---|
| routed experts | NVFP4 (MSE-searched static weight scales, dynamic
input scales) |
| self-attention | FP8 (W8A8, all projections) |
| linear-attention | FP8 (W8A8, the full gated-delta path — `conv1d` +
all in/out projections) |
| KV cache | FP8 (cast mode) |
| everything else | BF16 — including MTP, left unquantized |

Two things are documented in the file header because they affect how the
recipe should be
read:

- **The full gated-delta path is FP8, and that was validated
end-to-end.** The `conv1d` and
the in/out projections (`in_proj_qkv` / `in_proj_z` / `in_proj_a` /
`in_proj_b`, `out_proj`)
are all FP8; only the norms stay BF16. `nn.Conv1d` is a registered
ModelOpt quant module, so
the recipe's broad `*linear_attn*` rules reach `linear_attn.conv1d` too
— this is intentional
and matches the published `nvidia/Qwen3.8-2.4T-A95B-NVFP4`, whose
`hf_quant_config.json` lists
`linear_attn.conv1d` as FP8 on every gated-delta layer (the interleaved
full-attention layers
have no `conv1d`). (An earlier revision of the file header / ptq.md
wrongly stated the
  recurrent path is never quantized; corrected in this PR.)
- **The source ships as native block-FP8** (`quant_method=fp8`,
`weight_block_size [128,128]`,
dynamic activations). The loader dequantizes it to BF16 before
quantizers are inserted, so
the calibrated scales are against BF16 weights, not against the shipped
FP8.

Filed under `modelopt_recipes/models/` per the split introduced in #2219
(per-`model_type`
recipes vs model-hub checkpoint recipes); this one targets a published
checkpoint, alongside
`deepseek-ai/DeepSeek-V4-Pro-0813` and the Nemotron-3 entries.

### Usage

```bash
# The recipe is consumed by the PTQ entrypoint the same way as the other
# modelopt_recipes/models/ entries:
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <Qwen/Qwen3.8-2.4T-A95B checkpoint> \
    --recipe models/Qwen/Qwen3.8-2.4T-A95B/ptq/nvfp4_experts_mse-fp8_self_attn-fp8_linear_attn-kv_fp8_cast \
    --export_path <output>
```

### Testing

The exported checkpoint was evaluated against the BF16 baseline on
**GPQA, AA-LCR, SciCode,
IFBench and Terminal-Bench 2.1**, with no meaningful accuracy regression
on any of them. The
published `nvidia/Qwen3.8-2.4T-A95B-NVFP4` checkpoint is the artifact
this recipe produces —
its `hf_quant_config.json` is the ground truth for which modules are
quantized (routed experts
NVFP4; self-attention, all linear-attention projections **and**
`conv1d`, and KV cache FP8).

No new unit tests: this is a declarative recipe composed entirely of
existing units
(`base_disable_all`, `nvfp4`, `nvfp4_static`, `fp8`, `kv_fp8_cast`), all
already covered.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new file only; no existing
behaviour touched)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — declarative recipe over
existing, tested units
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — consistent with the recent recipe additions (#2219, #2269, #2287),
which did not add entries
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Model card: https://huggingface.co/nvidia/Qwen3.8-2.4T-A95B-NVFP4


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a post-training quantization recipe for the Qwen3.8-2.4T-A95B
model.
* Supports MSE-searched NVFP4 quantization for routed expert layers and
FP8 quantization across self-attention and gated-delta linear-attention
paths.
* Supports FP8 cast-mode key-value caching while retaining BF16
precision for multi-token prediction and gated-delta normalization
layers.

* **Documentation**
* Clarified the model’s hybrid precision configuration, including FP8
treatment of the gated-delta convolution path and the scope of broad
linear-attention patterns.
  * Documented source-checkpoint dequantization and validation behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-08 16:41:01 -07:00
Frida HouandClaude Opus 5 0688761ce9 feat(export): support multimodal and MTP models in layerwise export (#2303)
### What does this PR do?

Type of change: New feature

**Layerwise export now supports multimodal and MTP models.** Both were
refused outright, and
both were refused for the same reason: `finalize()` was called from
inside
`layerwise_calibrate`, which is the wrong scope for it.

**1. Calibration does not know which model the checkpoint describes.**
It only sees the
module it was handed. A VLM calibrates its *language model*, so the
shards, the exclusions
and `config.json` all came out describing that submodel rather than the
whole VLM. Moving the
call out lets the caller root the exporter at the parent — and without
the key prefixing,
tower collection or ambient parent handle an earlier attempt needed,
because the decoder
layers are the same objects from either root.

**2. Calibration runs before things the export needs exist.** Orphaned
MTP weights are loaded
*after* calibration, by which point every shard had already been
written, so they could not
be passed at all. After the move they are an ordinary argument to
`finalize()`, with no
staging attribute stashed on the model.

### How it works

The exporter is created by whoever owns the export and **announced on
the model** that
`mtq.quantize` is given. Calibration picks it up, binds it, and drives
it per layer; the
export that follows reads it back and finishes the checkpoint:

```python
LayerwiseExporter(full_model, export_path).announce(language_model)
mtq.quantize(language_model, quant_cfg, forward_loop=loop)
...
getattr(full_model, LAYERWISE_EXPORTER_ATTR).finalize(extra_state_dict=mtp_state_dict)
```

Calibration and export are handed *different* models, so `announce()`
publishes the exporter
on each end separately: the caller announces on the model being
calibrated, and `bind()`
announces on the export root. Neither side has to know where the other
looked, and the lookup
stays an O(1) `getattr` rather than a `named_modules()` scan — worth
avoiding at roughly
1.65 µs/module, or ~500 ms on a Kimi-K3-sized model. For a non-VLM both
roots are the same
object and the second announcement is a no-op. `finalize()` clears every
attachment it
recorded, so the module graph does not retain a live exporter
afterwards.

`mtq.quantize` and `mtq.calibrate` are **unchanged** — a layerwise-only
feature does not
belong in the public quantization API. The attribute follows
`_mtp_layer_prefixes`, which
crosses the same calibration→export boundary the same way
(`hf_ptq.py:538` sets it,
`unified_export_hf.py:870` reads it back).

Construction is inert: `__init__` records only the export root and the
directory, because the
caller builds it before `mtq.quantize`, when there are no quantizers yet
to validate or read
a config from. `bind()` does that, called from calibration after
quantizer insertion and
before any layer is converted — the only window where both hold, and the
same instant the
exporter used to be constructed, so unsupported models still fail in
seconds rather than
hours. Only the calibration pass that sets `export_dir` drives the
exporter: a list-form
algorithm runs one pass per entry, and an earlier one must not convert
layers a later one
still has to calibrate.

### Usage

Nothing changes for a plain layerwise-export recipe:
`layerwise.export_dir` still drives it.
Pre-attaching an exporter is the opt-in for the two cases that need it —
a checkpoint whose
root is wider than the calibrated model, and orphaned tensors to merge
at the end.

The one behaviour change for a config-only caller is that `mtq.quantize`
now writes the layer
shards but no longer finishes the checkpoint. Both exit paths warn with
what is still owed,
and `LayerwiseConfig.export_dir`'s description has been corrected — it
previously promised "a
complete, loadable checkpoint when the last layer lands" and still
listed multimodal and MTP
as raising `NotImplementedError`.

### Testing

`tests/gpu/torch/export/test_layerwise_export.py` — **29 passed**.
Beyond the 24 inherited
from #2136, five new ones, each with a negative control confirming it
fails without its fix:

- orphaned MTP tensors reach the tail shard *and* the index
- an exporter rooted at the parent widens the checkpoint's namespace
- the config-only path announces an exporter that can be finished, and
finalize clears it
- only the pass that sets `export_dir` drives the exporter
- an exporter whose root holds a different number of layers is refused
at `bind()`

Full suites: `tests/gpu/torch/export` + `tests/gpu/torch/quantization`
**1012 passed / 55
skipped**, `tests/unit` **3318 passed / 15 skipped**, pre-commit clean.
Both suites also
report failures in `test_implicit_gemm.py` (FP4 conv kernels),
`test_triton_fa_p_qdq.py`,
`test_autocast_quantize_int8` and `test_engine_builder.py` collection;
all reproduce unchanged
on `main` and none touch the paths in this diff.

Measured against the whole-model exporter on a tiny Gemma3-VL, towers
prepared exactly as
`hf_ptq` does:

```
keys: baseline=80  layerwise=80   only-baseline=[]  only-layerwise=[]
differing values: 0
vision tower present: True     VLM namespace: True
config.json is the VLM: True   hf_quant_config match: True
exclude_modules: ['language_model.lm_head', 'vision_tower.vision_model*']   (both sides)
```

#### End-to-end through `hf_ptq.py`

Same FP8 recipe both sides; the baseline drops `layerwise.export_dir`
and is exported by
`main`, so the diff isolates this PR. Every tensor matches in key,
dtype, shape and value,
and `config.json` / `hf_quant_config.json` match too.

| Model | Covers | Keys | Differing |
|---|---|---|---|
| Qwen3-VL-8B-Instruct | multimodal | 1254 = 1254 | 0 |
| GLM-4.7-Flash | MoE + MTP | 28119 = 28119 | 0 |

The VLM checkpoint keeps the vision tower unquantized (351
`model.visual.*` keys, no
`weight_scale` among them) while the language model is FP8. The MTP run
reports 212 orphaned
tensors; all 212 land in `model-tail.safetensors` and in the index, with
`model.layers.47*` in
`exclude_modules`.

**Not yet validated:** an accelerate-offloaded run, and a serving canary
on the exported
checkpoints.

### Refusals

`export_dir` without `enable`, and an exporting algorithm entry with no
calibration method,
are both refused before calibration starts — neither reaches the
per-layer pass, so both
would otherwise export nothing. The early gate is a heuristic on the
recipe, so `hf_ptq` also
raises a plain `RuntimeError` at export time if calibration turned out
not to have run; that
backstop, not the gate, is what makes the failure legible on paths the
recipe check cannot
predict.

`bind()` requires the layers calibration will drive and refuses a root
that discovers a
different number of them. Only the count is checked here: `export_layer`
already rejects a
reordering or a substituted module on its first call, and a length
difference is the one
mismatch it structurally cannot catch — every call would pass and
`_write_index` would then
open a shard that was never written, at the very end of the run.

Orphan tensors are merged into the tail with no collision check,
matching the whole-model path
(`unified_export_hf.py:1623`). `load_mtp_weights` returns exactly the
keys absent from
`model.state_dict()`, so a collision with an exported tensor is not
reachable through the only
producer, and a guard would only make the two export paths diverge.

### Why not reuse `export_hf_checkpoint`

It was the first idea and it is the most expensive one. Its transformers
path is whole-model
at every step — `_prepare_moe_inputs`,
`requantize_resmooth_fused_llm_layers` (which runs a
dummy forward that would fail on already-converted layers),
`_process_quantized_modules`, a
full `model.state_dict()` in host RAM, then `save_pretrained` rewriting
shards already on disk
— and it raises outright under `has_accelerate_offload`.
`save_pretrained(state_dict={})` is
not an escape either: safetensors' shared-storage check fires on MoE
even with an empty dict.

The natural consolidation target is the **streaming** exporter, which is
already most of
`finalize()`: 122 lines vs 74, sharing `decoder_owned_ids`,
`enable_weight_access_and_writeback`, `_dispatch_export_handler`,
`_reconstruct_fused_moe_linear`, `_add_mtp_exclusions`,
`_postprocess_single_tensor`,
`requires_weight_materialization` and `save_non_weight_artifacts`.
Folding them together needs
roughly four knobs: skip the whole-model prep, skip layers already
written, seed the index
with the existing shards, and inject the quant config. That is a
separate change and
deliberately not in this one.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ —
`mtq.quantize`/`mtq.calibrate` signatures are
unchanged, and a recipe that only sets `layerwise.export_dir` behaves as
before. The one
behaviour change is that `mtq.quantize` no longer finishes the
checkpoint on its own:
callers must now call `finalize()` on the exporter, which calibration
leaves on the model.
- If you copied code from any other sources or added a new PIP
dependency, did you follow
  guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ❌ — pending.
- Did you get Claude approval on this PR?: ❌ — the last review's
findings are all addressed;
  needs a re-run.

### Additional Information

Follow-ups this enables: #2259 (MTP) reduces to close to nothing, and
the multimodal work in
#2218 no longer needs `export_parent`, the key prefixing, or the tower
collection.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Layerwise export now supports clearer control over export locations
and calibrated layer handling.
- Export workflows provide improved support for resuming, sharded
checkpoints, mixture-of-experts models, and nested model namespaces.

- **Bug Fixes**
  - Improved handling of exported checkpoint shards and extra tensors.
- Added clearer warnings when exports require completion before loading.

- **Documentation**
- Clarified that layerwise exports write shards during calibration and
require an explicit finalization step.
- Documented that the in-memory model is not suitable for inference
after layerwise export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-08 20:48:47 +00:00