719 Commits
Author SHA1 Message Date
Chenjie LuoandClaude Opus 5.5 ad8cd63847 Share one CUDA encoder per IQ family (#2615)
### What does this PR do?

Type of change: refactor (no behaviour change)

The five GGML IQ CUDA encoders were five copies of the same search.
IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and
launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same
structure with a different choice space. Every scaled packer also
validated its scales twice, in the `ggml.cpp` pybind wrapper and again
in the CUDA entry point.

This PR keeps **one encoder per family**, as two templates:

- **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in
shared memory, the 16 local scales are scored per group, and each vector
then takes its best entry under the chosen scale. A format supplies its
group shape, whether it stores seven sign bits and recovers the eighth
from parity, and a `store()` that writes the chosen entries, sign masks
and local scales into its layout.
- **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary
grid. Each group picks one of `kChoices` options. With `kSharedShift`
the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`);
otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now
writes FP16 scales, so both IQ1 formats take the same input.

Each format file is now one `Format` struct, holding its layout
constants and `store()`, plus its entry point: 58–100 lines each.
Validation lives once in `common.cuh`, as `check_pack_inputs` and
`check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points
directly instead of through five wrappers. **The kernel sources shrink
from 1,536 to 1,241 lines** (+665 / −960).

This is the first of two PRs. #2604 builds on it: it adds CUDA decoders
as a `decode()` next to each format's `store()`, and makes export reuse
fake quant's packed payloads.

### Testing

**Nothing changes in the output.** Before the refactor I hashed 40
outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode
and decode, on a weight with zero, tiny, oversized and non-finite
blocks. All 40 hash the same afterwards.

**Encode speed is unchanged.** Old and new were timed alternately for
four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048
weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms,
IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 /
15.4.

- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed**
- **Validation reports the same errors in the same order.** Over 5
formats × 8 combinations of bad arguments (devices, dtype, width, grid
shape, scales dtype, length and sign), every first error matches main's.
- `tests/gpu/_extensions/test_torch_extensions.py`: the
validation-message tests pass. #2515's two Q8_0 tests fail identically
on a clean `main` on this GPU.
- IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`,
`test_convert_hf_config.py`, `test_presets.py`,
`test_export_weight.py`): **173 passed**

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Same bindings, messages and
bytes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: N/A. A refactor with no
behaviour change, verified by the hashes above and the existing GPU
tests.
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: **this** → #2604 (pack each IQ weight once and decode on
CUDA).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Quantization now checks that inputs and grids are CUDA tensors on the
same device, with compatible shapes. Scaled formats also validate scale
type, shape, and finite, non-negative values.
* **Improvements**
* IQ1 and IQ2 formats share common encoding paths while retaining their
format-specific output layouts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:14:42 -07:00
Chenjie LuoandClaude Opus 5.5 aa89722d38 [6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed
the PyTorch codec; this PR adds its **CUDA encoder** and makes the
format reachable. With it, ModelOpt supports all five GGML IQ formats at
one and two bits.

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq1_m`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

### The kernel

In the kernel the delta shift is free per group, so it sits above the
entry index in the sort key: a tie still prefers the lower shift and
then the lower entry, as the reference encoder does. The 2048-entry grid
IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory
limit, so both kernels read it from global memory and rely on the cache.

| 5632×2048 weight | torch | CUDA | |
|---|---|---|---|
| IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** |

### Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and
apply the ±1/8 shift the same way. So those three steps move into
`common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and
IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA
output on a 5632×2048 weight hashes the same before and after, and so
does IQ1_M's, compared against the pre-split version of this change.
IQ1_S encodes at the same speed (306 M elem/s).

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ1_M-specific test code: backend dispatch, weight caching, the
`num_bits` guard, `convert_hf_config` metadata, Megatron export and the
`TensorQuantizer` tests in the shared battery. The shared CUDA battery
gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **166 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **221 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO
6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch
encoder parity
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** (9 tests × 5 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed**
- reconstruction error falls monotonically across all five formats,
pinned by a test
- `general/ptq` now holds 31 recipes; `ptq.md` is updated.

Rebased onto `main` after #2513 merged. The resulting tree is identical
to the one the runs above tested, and the unit set was rerun on it: 166
passed.

On this GPU, two of #2515's Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M
codec), all merged → **this**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M weight-only quantization at 1.75 bits per weight, with
CUDA acceleration and a 256-value block size.
* Added an IQ1_M post-training quantization recipe for eligible linear
layers; calibration data is not required.
* Added IQ1_M to the supported GGML-compatible formats and recipe
listings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:23:48 -07:00
Chenjie LuoandClaude Opus 5.5 e5b63320ab [5/6] Add the IQ1_M codec (#2513)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 09:45:48 -07:00
kinjalpatel27 ad9ea97a4b Fix vLLM compilation guard for models without marker (#2518)
### What does this PR do?

Type of change: Bug fix

Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.

Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.

### Usage

```python
with disable_compilation(model):
    mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```

No caller changes are required.

### Testing

- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-09-29 14:28:32 -07:00
Chenjie LuoandClaude Opus 5.5 3091b8ff69 [4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512
landed the PyTorch codec; this PR adds its **CUDA encoder** and makes
the format reachable:

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq2_s`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B
parameters**.

### The kernel

IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its
search the most expensive in the family. The codebook and its norms take
36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB
static limit, so they are declared statically like the IQ2_XS and
IQ2_XXS kernels.

That cost is why the kernel matters more here than anywhere else:

| | torch | CUDA | |
|---|---|---|---|
| IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×**
|
| extrapolated to a 27B model | ~9.8 hours | **~37 s** | |

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ2_S-specific test code: backend dispatch and weight caching, the
`num_bits` guard, `convert_hf_config` metadata (uniform and mixed
precision), all 9 Megatron export tests, and the two `TensorQuantizer`
tests in the shared battery. The shared CUDA battery gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **134 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **192 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO
6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder
parity, determinism, reconstruction at scale, zero and non-finite
policy, float64 input and the fallback path.
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **36 passed** (9 tests × 4 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**.
TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`,
`block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)`
uint8.
- `general/ptq` now holds 30 recipes.
- The shared-memory change in `b7739d5d0` leaves the packed bytes
identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M
elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun:
42 passed.

All of the above was rerun after rebasing onto `main` at `c2aaa44f6`.
That base adds a Q8_0 packer to the same GGML extension (#2515), and
changes the hf_ptq example and the export code this format goes through.
The packed IQ2_S bytes still hash the same. On this RTX PRO 6000
(sm_120), two of #2515's own Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
#2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ2_S weight-only quantization for eligible linear layers, at
2.5625 bits per weight.
* Added a PTQ recipe that requires no calibration data. Weights must
meet the existing 256-value block-size constraint.
* Added CUDA-accelerated packing for CUDA weights, with a Python
fallback when the CUDA extension is unavailable.
* **Documentation**
  * Updated the PTQ recipe catalog and IQ-format size tradeoffs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 20:35:38 +00:00
sychen52andClaude Opus 5.5 834c90d7a1 Fix NVFP4 fake quant zeroing blocks with small scales (#2549)
### What does this PR do?

  Type of change: Bug fix

The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute
>= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block
scale below 1e-5 with
1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static
Triton kernel, the CUDA extension fallback and NVFP4 export have no such
floor; the floor was
only guarding division by zero. The Triton kernel had its own copy of
the scale/round code instead of the shared `nvfp4_scalar_quant`.

- Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block
scale).
- Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule.
- Both: a zero, inf or NaN global amax uses a unit block scale, like the
CUDA extension (the conv kernels returned NaN for inf/NaN before).

  Blocks with scale >= 1e-5 are unchanged.

- New tests: power-of-two scaling of input and global amax (2^-10,
2^-20) scales the output by the same factor (Triton, standalone conv
FP4, fused conv3d); invalid
global amax gives unit-scale rounding. The conv test's Python reference
drops the floor.
- B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization`
NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ).
- Negative control on `main`: the small-input tests fail on both
kernels; conv also fails for inf/NaN global amax.

  ### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (only blocks with scale < 1e-5
or an inf/NaN global amax change, from zeros/NaN to correct values)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
  - Did you write any new necessary tests?: ✅
  - Did you update Changelog?: N/A
  - Did you get Claude approval on this PR?: ❌




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* FP4 quantization now preserves proportional output when inputs and
valid global scales are reduced together, including at small scales.
* Zero, infinite, or NaN global scales use a safe fallback, preserving
inputs already representable in FP4.
* Small positive block scales are no longer discarded by an absolute
scale threshold; subnormal scale handling is covered across quantization
paths.
* **Tests**
* Added coverage for scale consistency across input types and block
sizes, invalid global scales, and subnormal FP8 block scales.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:31:31 -07:00
h-guo18andClaude Opus 5 c2aaa44f60 [Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do?

Type of change: new feature

Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft
variant of the existing DFlash mode, selected with
`dflash_architecture_config.projector_type="dflash2"` alongside
`domino`, `dspark` and `lilicorr`.

DFlash2 keeps DFlash's one-pass parallel backbone and adds two
components that recover the acceptance a purely parallel draft loses:

- **Grouped dynamic depthwise convolution** around every attention and
MLP sublayer, giving each block position a view of its predecessors
*inside* the block. Taps do not cross the block boundary, so the draft
stays one forward pass.
- **Low-rank candidate selector** scoring transitions between adjacent
positions' top-k candidates, so serving walks one coherent path instead
of taking an independent argmax per position.

Both start as exact no-ops — the convolution's `base_kernel` is an
identity and `kernel_projection` is zeroed; the selector's
`successor_codebook` is zeroed — so a freshly built DFlash2 draft *is*
its DFlash backbone, and enabling the variant is an extension rather
than a perturbation. This matches the reference implementation
([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772),
merged) and the way `modeling_lilicorr` already installs this same
convolution class.

**This unblocks a recipe already shipped on `main`.**
`modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv`
from `modeling_dflash2`, so
`modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml`
raises at model build today and its CHANGELOG entry documents a feature
that cannot run. Landing this makes it runnable.

Module and parameter names match the SGLang/vLLM `DFlash2DraftModel`
loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2`
checkpoint: **81 tensors, 21 name patterns, zero difference in either
direction**. The serving side,
[vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816),
has since merged (`b389ac294`) with no change to the checkpoint
contract.

### Usage

```bash
python examples/speculative_decoding/main.py \
  --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \
  model.model_name_or_path=Qwen/Qwen3-8B \
  data.data_path=<corpus>.jsonl \
  training.output_dir=<out>
```

```yaml
# modelopt_recipes/general/speculative_decoding/dflash2.yaml
dflash:
  dflash_selector_loss_alpha: 1.0      # weight of the candidate-selector CE term
  dflash_architecture_config:
    projector_type: dflash2
    conv_kernel_size: 2                # taps; must not exceed the block size
    conv_group_size: 16                # must divide hidden_size
    selector_rank: 256
    selector_top_k: 16
```

### Testing
<img width="2000" height="1320" alt="image"
src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c"
/>

**Unit** — 25 CPU tests in
`tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full
`tests/unit/torch/speculative/` suite passes with no regressions. The
ones worth keeping pin invariants that a decreasing loss does not catch:

- the convolution is an exact identity on the **default** construction,
and its taps stay inside the block while a position still sees its
predecessors;
- the block-offset contract shared by the training objective and
`CandidateSelector.greedy_path` — a misaligned objective still
converges;
- the export fields the vLLM loader requires, including the top-level
`block_size` that `DFlash2Exporter` derives the nested copy from;
- which selector factors receive gradient on the first step.
`successor_codebook` starts at zero, so `predecessor_codebook` and
`hidden_projection` take one step to begin moving. That is a warm start,
not a dead branch, and both sides are asserted.

**End-to-end** — trained on Qwen3-8B against a plain DFlash control with
every other argument identical (plot above). Monotonic convergence, no
NaN/divergence, no DDP unused-parameter issues. Note the losses are
**not comparable** across arms: DFlash2's includes the selector CE term.

**Serving (vLLM)** — the exported drafter loads and drafts under the
merged DFlash2 path (`RESOLVED draft architectures:
['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes
the convolution from `1 + num_speculative_tokens` at runtime rather than
from the checkpoint, so a `block_size=16` drafter is only correct at
`num_speculative_tokens=15`; and at that value the upstream path
currently hits an illegal memory access in `_cache_draft_logits`
([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)),
independent of which checkpoint is used.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — additive. New
`projector_type`, its own registry and exporter, one new config field;
DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents
are untouched.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ —
`modeling_dflash2.py` is adapted from
[SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and
carries its MIT notice. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under `0.48.0`.
- Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all
review threads addressed and resolved.

### Additional Information

Rebased onto current `main`. Two commits from the original branch were
dropped because
[#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them
first, with authorship preserved: the no-op sublayer seam in
`modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix —
`main`'s version of the latter is stricter, so this PR no longer touches
`hf_dflash.py` at all.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash2 speculative decoding with grouped dynamic convolutions
and low-rank candidate selection.
* Added configurable selector-loss weighting, including an option to
disable it.
  * Added DFlash2 model conversion and export support.
  * Added checkpoints compatible with SGLang and vLLM DFlash2 serving.

* **Documentation**
* Added training recipes and a Qwen3-8B online DFlash2 training
configuration.

* **Tests**
* Added coverage for conversion, training, metrics, gradients, and
export compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 11:43:02 +08:00
Wei-Ming Chen 1d392999b4 [OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do?

Type of change: new feature.

Follow-up to merged #2272. Adds composition of existing GEMM
quantization with KV-cache AutoQuantize:

- fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize;
- gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent
mixed-KV AutoQuantize;
- an optional `kv_auto_quantize` recipe stage with independent method,
constraints, candidates, and checkpoint path;
- ordered `hf_ptq.py` orchestration that keeps selected
weight/activation QDQ active while its calibration state remains frozen
during KV candidate calibration;
- fail-closed validation when a preceding stage leaves actual K/V
quantizers enabled; and
- unified export of a uniform-weight or mixed-weight checkpoint together
with the selected per-layer KV map.

The KV search still uses the public `mtq.auto_quantize(...,
constraints={"cost_model": "kv_cache", ...})` API from #2272. On a
converted model, the API preserves existing non-KV quantizers and
requires K/V to be disabled before search. Fresh-model behavior is
unchanged and starts from a deny-all quantizer baseline.

#### Why a follow-up field instead of a generic stage list?

This PR deliberately supports the two composition forms required by
`hf_ptq.py` without replacing the stable recipe schema. Existing recipes
already express a fixed `quantize` baseline plus one primary
`auto_quantize` search. A generic ordered `stages` list would require a
broader recipe/API migration, indexed checkpoint semantics, and
compatibility rules for arbitrary stage sequences. There is not yet a
demonstrated third search stage that justifies that surface-area change.

The two searches are not combined inside `mtq.auto_quantize`: each
invocation owns one search domain, constraint model, scoring method, and
resumable checkpoint. Their ordering and independent checkpoint paths
are orchestration concerns, while candidate calibration, scoring,
selection, and state application remain in the shared public API. A
general stage pipeline can be considered separately if more than this
one optional KV follow-up is needed.

Both solvers and scoring protocols are unchanged. The KV checkpoint
compatibility signature additionally fingerprints the preceding
quantizer configuration and calibrated state. Unsupported uniform-weight
plus mixed-KV exports record `kv_cache_deployment_supported: false` in
both ModelOpt and converted HF metadata.

### Usage

Fixed FP8 GEMM PTQ followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-fp8-and-mixed-kv
```

Weight AutoQuantize followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --auto_quantize_checkpoint /path/to/weight_autoquant.pth \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv
```

KV checkpoint resume requires identical preceding non-K/V quantizer
configuration and calibrated state. If rerunning the preceding stage
changes that state, use a new KV checkpoint path to recompute
sensitivities; configuration identity alone is insufficient to reuse the
scores safely.

### Testing

- Latest changed-area validation: 126 tests passed across `hf_ptq.py`
orchestration, KV checkpoint compatibility, export metadata, and HF
configuration conversion.
- A broader local run had 604 passes, one skip, and six failures: two
socket-binding failures under the sandbox and four local Transformers
API incompatibilities. This is not a full-suite pass.
- The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen
fixture.
- Public API coverage verifies that composed KV search preserves
preceding weight quantization and rejects enabled K/V state.
- Changed-file pre-commit hooks passed; the isolated recipe validator
also passed after dependency bootstrap.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅ (0.48.0 composition feature and KV
checkpoint flag deprecation)
- Did you get Claude approval on this PR?: ❌

### Additional Information

- This follow-up targets `main`, which contains merged #2272.
- `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are
intentionally separate because KV sensitivities depend on the preceding
GEMM state.
- Uniform-weight plus mixed-KV exports are for artifact inspection until
the runtime's uniform-weight ModelOpt configuration consumes
`kv_cache_quantized_layers`. Export emits an actionable warning and
records `kv_cache_deployment_supported: false`; this marker does not
itself add runtime support.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added staged post-training quantization workflows for weights and KV
caches, including dedicated KV-cache checkpoints.
- Added FP8/NVFP4 recipes with configurable bit constraints and scoring.
  - KV-cache quantization now supports pre-quantized models.

- **Bug Fixes**
- Mixed weight and KV-cache quantization now exports with a warning
instead of failing.
- Improved validation and checkpoint compatibility for staged
configurations.
- Added safeguards for configurations without enabled weight quantizers.

- **Documentation**
- Clarified staged KV-cache workflows, checkpoint options, configuration
behavior, and unsupported deployment combinations.
- Documented deprecated legacy quantization options and their
replacement behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-29 00:15:40 +00:00
hychiang de2e810219 [OMNIML-5899] Add Q8_0 CUDA packing kernel (#2515)
### What does this PR do?

Type of change: new feature

Adds the Q8_0 CUDA packing layer for the three-PR Q8_0 series:

- packs 32-value blocks into the 34-byte GGML-compatible payload;
- exposes `q8_0_pack` through the shared GGML extension;
- validates device, dtype, row alignment, and CUDA launch bounds;
- tests byte layout, accepted dtypes, non-finite handling, float64
narrowing, FP16 scale boundaries, reconstruction error, and invalid
inputs.

This PR contains only the kernel and extension boundary. The
codec/backend and export/recipe layers remain in the later PRs.

### Usage

```python
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml

extension = get_cuda_ext_ggml(raise_if_failed=True)
packed = extension.q8_0_pack(weight)
```

### Testing

- Combined #2515 -> #2516 -> #2517 stack: 108 focused CPU codec,
backend, export, and recipe tests passed.
- Ruff, formatting, and whitespace checks passed for the changed Python
test.
- The shared-extension Q8_0 and existing IQ tests require CUDA CI; the
latest run is pending on the current PR head.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: yes; no
dependency was added and no implementation code was copied
- Did you write any new necessary tests?: yes
- Did you update `CHANGELOG.rst`?: N/A; the user-facing entry is in
#2517
- Did you get Claude approval on this PR?: pending

### Provenance

The CUDA encoder was independently written for ModelOpt. It implements
the packed-format contract and scalar quantization formula documented by
the pinned llama.cpp definitions:

- [Q8_0 packed
structure](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h)
- [Q8_0 scalar reference
formula](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c)

No llama.cpp implementation code is incorporated into this CUDA source.
Human code-owner confirmation of this provenance and attribution is
requested before merge.

### Related PRs

Merge order:

1. **Kernel - this PR**
2. [#2516 - Q8_0 quantization codec and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2516)
3. [#2517 - Q8_0 checkpoint export and
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2517)

All three PRs target `main`.

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
2026-09-28 15:22:25 -07:00
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
Chenjie LuoandClaude Opus 5.5 767ef5533e [3/5] Add the IQ2_S codec (#2512)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at
one and two bits (2.5625 bits per weight). This one lands the **PyTorch
codec**: the encoder, the decoder and the 1024-entry codebook. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2565 adds the CUDA encoder,
registers the format and adds its recipe.

### What's distinctive about it

**IQ2_S is the one format llama.cpp's own tooling gives no head start
on**, so the search is written against the GGML layout directly.

The interesting difference from IQ2_XS and IQ2_XXS is sign handling.
IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit
parity-coded index. The encoder therefore takes the input signs as they
are instead of flipping the weakest element to fix parity, and the
search compares magnitudes directly, which is simpler than its siblings.

### Why the codec lands before the kernel

The CUDA encoder's tests use this codec as their reference. They compare
against the PyTorch encoder byte for byte and draw the grid and scale
predictor from it. So the kernel cannot be tested before the codec
exists, and it follows in #2565 together with the registration. Every
registered format therefore keeps a CUDA encoder.

### Test changes that make the split possible

A codec can now land before it is registered, so two test contracts in
`test_iq_formats.py` are stated precisely:

- The two tests that go through `TensorQuantizer` (pass-through
gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`.
Every other battery test calls the codec directly and covers IQ2_S here.
- The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <=
set(FORMATS)` instead of equality. That is what its docstring already
said: a registered format must be listed, or it escapes the contract.
- `test_registry_lists_every_exported_encoder` is unchanged, and it is
why this PR leaves the package exports alone: an exported encoder must
be registered.

The error-by-bit-width failure message also labels errors by the order
they were measured in; it previously zipped them with alphabetical
names.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry.
Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so
CI keeps checking bytes we did not produce.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S
codec cases, including the llama.cpp conformance check
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by
this PR

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new
codebook is a GGML table, carried in `codebooks.py` with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
**this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added GGML-compatible IQ2_S quantization and dequantization support,
including access to its magnitude grid.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 14:03:01 -07:00
Edwardssss 57f929e358 Bound the stale-capture warning to once per capture and name its cause (#2492)
### What does this PR do?

Type of change: Bug fix

The activation-capture hooks latch `_intermediate_output` on every
forward, and only
`DistillationModel.compute_kd_loss()` ever clears it. Any training loop
that does not call
`compute_kd_loss()` therefore warns on every forward after the first,
once per hooked module:

```
plain transformers Trainer, 16 micro-batches
  before: 30 warnings (Teacher 15 + Student 15)
  after:   2 warnings (Teacher  1 + Student  1)
```

The count scales with forwards, not optimizer steps —
`gradient_accumulation_steps` of 1, 4 and
16 all produced 30 warnings for the same 16 forwards — so the
accumulator the report blamed is
not involved. A plain `transformers.Trainer` reproduces it because its
default `compute_loss`
reads the student's own CE from `outputs.loss` and never calls
`compute_kd_loss()`; the
warning is a symptom of that, but neither message said so. The student
message attributed the
re-forward to Activation Checkpointing only, and the teacher message
called the situation
"expected" while still raising `UserWarning`.

This PR:
- warns once per captured activation via
`_warn_once_about_stale_output`, re-armed wherever a
capture is cleared, so a consumed capture that is followed by another
unconsumed one is still
reported (Activation Checkpointing re-runs forwards and relies on that);
- states the cause and the consequence in both messages:
`compute_kd_loss()` did not run since
  the previous forward, so no KD loss is applied;
- moves the three `_intermediate_output = None` reset sites onto one
`_clear_captured_output`
helper so the capture and its warning state cannot drift apart, and has
the layerwise teacher
hook share both helpers rather than keeping a second copy of the
warning.

### Usage

No API change. A loop that applies KD should call `compute_kd_loss()`
once per forward:

```python
class KDLossTrainer(Trainer):
    def compute_loss(self, model, inputs, return_outputs=False, **kwargs):
        outputs = model(**inputs)
        loss = model.compute_kd_loss(student_loss=outputs.loss)  # consumes the captures
        return (loss, outputs) if return_outputs else loss
```

### Testing

`tests/unit/torch/distill/test_distill.py`:
- `test_duplicate_fwd_hook_call` now pins the bound under
`warnings.simplefilter("always")`
(3 forwards -> exactly 2 warnings) instead of relying on the
interpreter's per-message
  deduplication, and asserts the message names `compute_kd_loss`;
- `test_stale_output_warning_rearms_after_consuming` covers the re-arm:
warn, consume, warn
again -> 4 warnings. This one passes before and after the change; it
guards against a future
"warn once ever" simplification silently muting the Activation
Checkpointing case.

`tests/unit/torch/distill/test_layerwise.py`:
- `test_layerwise_stale_output_warning_is_bounded` covers the layerwise
hooks (3 forwards ->
  exactly 2 warnings).

The first and third tests fail on `main` (`assert 4 == 2`), verified in
a worktree of `main`
carrying these test files.

Measured with a `transformers.Trainer` over a 16-micro-batch run,
counting
`"already has an intermediate output stored"`:

| setup | before | after |
|---|---|---|
| plain `Trainer`, `gradient_accumulation_steps=1` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=4` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=16` | 30 | 2 |
| `Trainer` calling `compute_kd_loss()`, accum=4 | 0 | 0 |

```
$ python -m pytest tests/unit/torch/distill -q
34 passed

$ python -m pytest tests/unit/torch -q --ignore=tests/unit/torch/deploy
1 failed, 2673 passed, 18 skipped in 217.17s
```

The single failure is
`tests/unit/torch/quantization/plugins/test_huggingface.py::test_dbrx`,
which fails identically on `main` in this environment because
`transformers 5.17` is outside the
`transformers>=4.57,<5.15` range pinned in `pyproject.toml`. It is
unrelated to this change.

`pre-commit run --files <changed files>` passes every hook (ruff check,
ruff format, mypy,
bandit, insert-license, large files, line endings).

Note on severity, since it affects how the issue reads: Python already
deduplicates a warning
by (message, location), so with the default filters this shows up twice
rather than 30 times.
The repetition is user-visible under `-W always`,
`PYTHONWARNINGS=always`, pytest, or one
warning registry per DDP rank, and the count is what scales with epoch
length. The behavioural
fix that matters for all filters is the latch.

### Before your PR is "*Ready for review*"

Is this change backward compatible?: ✅
If you copied code from any other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md: N/A
Did you write any new necessary tests?: ✅
Did you update Changelog?: N/A
Did you get Claude approval on this PR?: N/A

### Additional Information

Fixes #2487.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved handling of captured activations during distillation,
including clearing stale outputs and resetting warning behavior after
outputs are consumed.
- Stale-output warnings now appear once per affected activation and
clarify when knowledge-distillation loss was not applied.
- Warnings distinguish expected cases involving activation checkpointing
or teacher evaluation.

- **Tests**
- Expanded coverage for warning counts, warning reset behavior, and
stale-output handling in standard and layerwise distillation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Edwardssss <ed_129@qq.com>
2026-09-28 22:01:07 +02:00
Keval MorabiaandClaude Opus 5 4eb86524f0 [1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do?

Type of change: refactor (no functional change)

**[1/2] of a split. Merge this first; #2514 is [2/2] and is based on
this branch.**

Three example scripts had each reimplemented the same MLflow wiring: the
flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the
params/tags/artifacts a run uploads, and the open/close dance with its
status. The copies had already drifted — only `hf_ptq` wrote a
provenance pointer, only `vllm_serve` republished the resolved URI — and
every new tracked script meant another copy.

What a script records is now one declarative `Tool` record, **declared
in the script itself, beside the flags it reads**:

```python
# examples/megatron_bridge/quantize.py
QUANTIZE = Tool(
    name="megatron_bridge_quantize",
    tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...",
    variant_help="recipe name, or --quant_cfg if no --recipe",
    variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"),
    model=lambda args: args.hf_model_name_or_path,
    checkpoint=lambda args: args.export_megatron_path,
    texts=lambda args: resolved_recipe_texts(args.recipe),
    outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"},
)
```

`tracked_run` takes that record and runs the whole thing, so a script
adds tracking in three lines: `add_mlflow_args(parser, TOOL)`,
`resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args,
TOOL):`. The shared module knows no script's flags.

`examples/hf_ptq`, `examples/vllm_serve` and
`examples/megatron_bridge/quantize.py` move onto it. Three helpers fall
away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s
two flag pass-throughs).

### Usage

No user-facing change. The flags, their spellings and the experiment
naming are exactly as before; a script author now writes a `Tool`
instead of four functions.

### Testing

- `tests/unit/torch/utils/test_mlflow.py`,
`tests/examples/hf_ptq/test_hf_ptq_args.py`,
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the
only lane that runs it), which drives `quantize.py` for real: **34
passed**, locally and in this PR's `megatron` lane.
- `pre-commit run --files <changed>`: all hooks pass.
- The four suites shared four copies of a stand-in for the `mlflow`
module, which had drifted — one recorded artifacts as a list, another as
a dict, a third made `log_artifact` a no-op, so a test asserting on an
upload asserted nothing. They now share one
`tests/_test_utils/mlflow.py`, which also emulates the fluent API's
habit of opening a run when none is active.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `track_run` and
`checkpoint_run_tags` are removed, but neither shipped in a release
(0.47.0's `__all__` is `MlflowRunLogger`, `command_text`,
`current_user`, `default_experiment_name`, `validate_tracking_uri`, all
unchanged here). Three deliberate behaviour changes, each in shared code
and each tested:
- The `source_checkpoint_path` tag resolves to an absolute path where it
recorded the raw argument, which a chain of runs needs to join on the
pair. `run_tags` is shared, so this applies to every script that writes
the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local
`--hf_model_name_or_path`. A source that names no directory, such as a
Hub `org/name` id, is still recorded as given.
- `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a
block ending in `SystemExit(0)` as `FINISHED` where it recorded
`FAILED`, since a script that ends by calling `sys.exit()` rather than
returning has still finished.
- `.experiment.json`'s `tracking_uri` and the `run_url` built from it
drop a trailing `/` from the tracking URI, so the link is
`https://host/#/...` rather than `https://host//#/...`. Only reachable
by constructing `MlflowRunLogger` directly; every CLI path already
stripped the slash in `resolve_tracking_uri`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-visible change; the entry is in [2/2].
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split out of #2514. This half is the enabling refactor with no behaviour
change; #2514 is the feature it unlocks and is based on this branch. At
~605 changed lines of core logic it is over the ~500 guideline; the
owner accepted a two-PR split rather than three, and everything #2514
alone consumes — `split_tracking_credentials`,
`log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands
there rather than here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 19:57:03 +00:00
hychiang ed7e87953c Fix grouped expert quantizer checkpoint replicas (#2500)
### What does this PR do?

Type of change: Bug fix

Fix distributed checkpoint saving for quantized Transformer Engine
grouped MoE experts when tensor parallelism and expert parallelism are
both greater than one.

#### Observed error

With NeMo 26.08 and TP=2, EP=4, ETP=1, quantization and calibration
complete, but the run crashes while saving the distributed checkpoint:

```text
megatron.core.dist_checkpointing.core.CheckpointingException:
Invalid sharding pattern validation.
Invalid access pattern for ShardedTensor(
    key='decoder.layers.1.mlp.experts.experts.32.linear_fc1.weight_quantizer._amax',
    ...
)
```

#### Root cause

`_QuantMegatronTEGroupedLinear.sharded_state_dict` created each expert
quantizer's sharded tensors without passing the expert tensor-parallel
and expert data-parallel process groups to
`make_sharded_tensors_for_checkpoint`. It then manually replaced only
the expert-data-parallel component of `replica_id`.

That manual rewrite is insufficient when tensor and expert parallelism
are both enabled: replica ownership is derived using the wrong
process-group topology, so ranks can publish an inconsistent access
pattern for the same globally indexed expert quantizer key.
Megatron-Core correctly rejects that checkpoint during sharding
validation.

#### Fix

Resolve the expert model-, tensor-, and data-parallel groups from one
`_pg_collection`-aware helper, then pass the expert TP/DP groups to
`make_sharded_tensors_for_checkpoint` for per-expert quantizer buffers.
This avoids mixing process-group sources when models use a non-global
`ProcessGroupCollection` or grouped child modules convert without their
parent MLP, and removes the manual `replica_id` rewrite. Shared,
whole-linear quantizer buffers deliberately retain the default dense
TP/DP replica groups because their keys and offsets carry no expert
identity; using only expert TP/DP groups would make replica IDs collide
across EP ranks.

#### Relationship to PR #2319

[PR #2319](https://github.com/NVIDIA/Model-Optimizer/pull/2319) fixed
two earlier TEGroupedMLP checkpoint problems:

1. Per-expert quantizer state did not retain the same globally unique
expert identity as the grouped-expert weights, so it could not be
redistributed reliably when the EP layout changed.
2. After ModelOpt extra-state restoration, an expert that moved to a
different rank could be missing the `_amax` or `_global_amax`
destination buffer required by the subsequent distributed checkpoint
load.

That PR therefore fixed resharding and restore correctness across
topology changes. Its regression matrix changed one parallel dimension
at a time: EP changed while TP=1, or TP changed while EP=1.

The remaining issue was on the save path when TP and EP were
simultaneously greater than one. Although the expert keys and restore
buffers were correct after #2319, the per-expert quantizer shards were
still constructed without the expert TP/DP process groups and then had
only part of their `replica_id` rewritten manually. Under TP=2, EP=4,
ETP=1, this produced the invalid access pattern rejected by
Megatron-Core before the checkpoint could be saved.

This PR complements #2319 by fixing that process-group/replica mapping
and adding combined TP+EP coverage.

A four-rank save-and-restore regression test covers combined TP=2, EP=2,
and ETP=1.

### Usage

N/A. This fixes checkpoint behavior without changing the public API.

### Testing

- [x] Focused Ruff check and format validation
- [x] Focused mypy validation
- [x] `git diff --check origin/main..HEAD`
- [x] Four-rank grouped-expert checkpoint save/restore test with TP=2,
EP=2, ETP=1: 1 passed on the original fix
- [x] End-to-end NeMo 26.08 quantization and checkpoint save on 8 B200
GPUs with TP=2, EP=4, ETP=1 on the original fix
- [x] Review-response validation on `dbfa1cdb5`: Ruff 0.15.20
format/check, `git diff --check`, and Python compilation
- [x] Final four-rank grouped-expert checkpoint regression on
`c2be0f2c6`: TP=2, EP=2, ETP=1 with both ordinary and overridden
ModelOpt TP state; 2 passed in 459.38s on 4 B200 GPUs with NeMo 26.08

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

The end-to-end validation produced a complete eight-shard distributed
checkpoint and exited successfully.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed checkpoint saving for quantized grouped MoE experts when tensor
and expert parallelism are enabled.
- Improved handling of shared and per-expert quantizer state across
parallel configurations to support more reliable checkpoint save and
restore.

- **Tests**
- Added NVFP4 grouped expert checkpoint round-trip coverage with tensor
and expert parallelism, including configurations with the
tensor-parallel group override enabled and disabled.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
2026-09-24 20:01:56 +00:00
JoshuaandCursor 63c4b660bd Add Aumann-Shapley sensitivity scoring method to auto_quantize (#2183)
[Paper](https://arxiv.org/abs/2607.12266) ·
[Overview](https://x.com/waterloo_intern/status/2076460984475263401) ·
[Implementation
thread](https://x.com/the_joshua_hill/status/2076427869388255635)

Depends on #2231, which provides the shared AutoQuantize
backward-scoring infrastructure. Until that PR merges, the focused PR B
diff is available
[here](https://github.com/joshua-hill/Model-Optimizer/compare/fix/autoquant-scoring-infrastructure...feat/aumann-shapley-autoquant).

### What does this PR do?

Type of change: new feature

This PR adds `method="aumann_shapley"` to `mtq.auto_quantize`. It is a
label-free scoring method that measures how each candidate quantization
format affects the model across the path from full precision to
quantized.

The new method is opt-in; existing `gradient` and `kl_div` behavior is
unchanged.

For each calibration batch, the method:

1. Runs the baseline model and saves its next-token distribution.
2. Measures each candidate format at a configurable number of points
along the quantization path.
3. Uses the KL-divergence gradients at those points to assign a damage
contribution to every runtime group and candidate format.
4. Measures the most aggressive candidate configuration once and uses
that value to calibrate the per-group damage model.

The resulting scores use the existing AutoQuantize linear-program
solver. The search can either:

- choose the least damaging configuration that meets an `effective_bits`
target; or
- choose the smallest configuration whose predicted damage stays below
`max_predicted_damage`.

The selected recipe records `predicted_damage` in mean per-token KL
units together with its validity and fit diagnostics.

### Public API

`auto_quantize` gains an optional `method_options` dictionary. For
`method="aumann_shapley"`, it accepts:

| Option | Default | Meaning |
|---|---:|---|
| `num_path_nodes` | `2` | Number of points used to average gradients
along the quantization path. |
| `damage_link` | `"coverage"` | How per-group scores combine.
`"coverage"` uses `damage = c * (1 - exp(-sum(b)))`; `"additive"` sums
the path contributions. |
| `max_predicted_damage` | `None` | Replaces the bit target with a
maximum predicted mean per-token KL. |

Method options are validated before the model is modified. Unknown
options and incompatible targets fail early.

### Usage

Select a configuration for a target effective bit width:

```python
import modelopt.torch.quantization as mtq

model, search_state = mtq.auto_quantize(
    model,
    constraints={"effective_bits": 4.8},
    quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"],
    data_loader=calib_loader,
    forward_step=lambda model, batch: model(**batch),
    method="aumann_shapley",
)

print(search_state["best"]["predicted_damage"])
print(search_state["best"]["predicted_damage_valid"])
```

Or let the search choose the bit width for a predicted-damage target:

```python
model, search_state = mtq.auto_quantize(
    model,
    constraints={},
    quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"],
    data_loader=calib_loader,
    forward_step=forward_step,
    method="aumann_shapley",
    method_options={"max_predicted_damage": 0.05},
)
```

### Implementation

- Reuses the candidate-replay and backward-scoring lifecycle introduced
in #2231.
- Reuses the existing AutoQuantize linear-program solver for both search
directions.
- Numerically integrates the coverage path when converting measured
contributions into per-group damage costs.
- Preserves deterministic runtime-group and candidate ordering.
- Keeps raw measurements, solver scores, and damage-model diagnostics
distinct in the search state.
- Rejects incompatible checkpoint resumes while allowing the same scores
to be re-solved for a new bit budget.
- Retains the shared MoE score-module rules so routed experts are scored
at their enclosing block.

Recipe integration will follow separately.

### Testing

Focused tests:

```text
pytest -q \
  tests/unit/torch/quantization/test_autoquant.py::test_backward_scoring_session_restores_partial_setup \
  tests/unit/torch/quantization/test_autoquant_shapley.py
```

Result: **51 passed**.

The tests cover:

- end-to-end scoring and configuration generation;
- agreement between path contributions and measured quantization damage;
- exact allocation checks against exhaustive search;
- effective-bits and predicted-damage search modes;
- checkpoint resume and offline re-solving;
- custom formats and heterogeneous candidate ladders;
- distributed reductions and nested MoE score modules;
- reused score modules and model-specific backward support;
- non-finite measurements and invalid-fit reporting; and
- input validation before model conversion.

End-to-end checks with NVFP4 and FP8 candidates at a 6.0-bit target:

- `Qwen/Qwen2.5-0.5B-Instruct` reaches 5.998 effective bits, with the
summed path contributions reproducing 98% of the directly measured
lowest-precision KL.
- `Qwen/Qwen3-30B-A3B` reaches 6.000 effective bits and reproduces 99%,
with all 48 MoE layers scored once at the sparse-MoE block rather than
per expert.

### Production use

We use this method in production for NVFP4 checkpoints of Kimi-K3,
MiniMax-M3, and GLM-5.2.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow the contributing guidance?: ✅ No copied code
and no new dependencies.
- Did you write the necessary tests?: ✅
- Did you update `CHANGELOG.rst`?: ✅
- Are the commits signed and signed off?: ✅


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
  - Added label-free Aumann–Shapley scoring for automatic quantization.
- Added configurable path sampling, damage modeling, effective-bit
targets, and predicted-damage bounds.
- Added temporary weight-folding support with automatic state
restoration.
- Added method-specific search options, checkpoint resumption, and
distributed scoring.

- **Bug Fixes**
- Improved cleanup and restoration of quantizer state, gradients, hooks,
and forward behavior after scoring or failures.
- Added validation and clearer handling for unsupported configurations
and invalid measurements.

- **Tests**
- Expanded coverage for scoring, solver behavior, distributed execution,
checkpointing, and custom quantization formats.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Joshua Hill <joshua.hill@baseten.co>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-23 20:51:55 -07:00
Chenjie LuoandClaude Opus 5.5 400498d82d [2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do?

Type of change: refactor (no behaviour change)

Addresses review feedback on #2511. Backend dispatch and export each
kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the
backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in
export. All four listed the same formats. Adding a format meant a row in
each, and the lists could drift apart. That had already happened twice
in #2511: `convert_hf_config.py` kept its own upper-case spelling of the
family and dropped IQ2_XXS metadata, and the Megatron export tests were
hard-wired to two formats.

Each format module now declares **one `IQFormat` record** beside its
encoder and decoder: name, block geometry, `quantize`, `dequantize`, and
its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists
them.

- Backend dispatch looks formats up in the registry.
- Both exporters take the packer and block geometry from it.
- Export's `IQ_FORMATS` is derived from it instead of being written out
again.
- `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed.
- The per-format fake-quant wrappers collapse into one
`IQFormat.fake_quant`, which does the `num_bits` check and calls the
existing cache helper.

Codebooks, searches, payload layouts and CUDA encoders stay in each
format's module.

#### Series and merge order

This is one slice of the IQ format series. It targets `main` so unit CI
runs, and **its diff includes #2511's commits until #2511 merges**.

1. #2511 — IQ2_XXS format
2. **this PR** — one registration per format
3. #2512 — IQ2_S format
4. #2513 — IQ1_M format

After this lands, #2512 and #2513 are restacked onto it, so each adds a
format module and a single registry entry instead of rows in four
tables.

#### Design choices

- **An explicit list, not self-registration at import.** If formats
registered themselves when their module was imported, the registry's
contents would depend on import order.
- **Backward compatible, with one behaviour change.** `iq1_s_fake_quant`
and `iq2_xs_fake_quant` are public on main, so each format keeps its
`<fmt>_fake_quant` name as an alias of its record's method. The three
removed tables were introduced by #2511 and never released. The
behaviour change: on main, the alias looked the encoder up at call time,
so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record
captures the encoder and decoder when it's built, so patching those
module functions reaches neither dispatch nor the alias. Substitute
through `IQ_FORMAT_REGISTRY` instead.
- **Registering a format declares it exportable, and that's intended.**
Export's `IQ_FORMATS` is derived from the registry, so a format
registered for dispatch is also claimed by both exporters and
`convert_hf_config`. That can't be wrong for an IQ format: fake quant is
`dequantize(quantize(w))`, so a format can't be dispatched without the
packer and block geometry, and those are all export reads. A QAT-only IQ
format can't exist. If one ever needs to land ahead of its export path,
an `exportable` flag on the record is a one-line addition.
- **The registry is the substitution seam.** Dispatch now reads the
registry, so tests that swap an encoder or decoder swap the registry
entry. Patching the format module's function would no longer reach
dispatch.
- **Test expectations stay independent of the registry.** Tests take the
*list* of formats from the registry, but their expected values come from
each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A
mis-wired registry entry therefore can't make both sides of an assertion
agree.

#### What it does not unify

The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list,
codebook sizes in `common.cuh`) and the recipes and docs remain per
format. "One registration" holds for the Python side, which is where all
four tables lived.

### Usage

Adding a format after this PR (for example IQ2_S in #2512) needs its
module and one line in the registry:

```python
# modelopt/torch/quantization/ggml/iq2_s.py
IQ2_S_FORMAT = IQFormat(
    name="iq2_s",
    block_size=IQ2_S_BLOCK_SIZE,
    block_bytes=IQ2_S_BLOCK_BYTES,
    quantize=quantize_iq2_s,
    dequantize=dequantize_iq2_s,
    block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE,
    decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE,
)

# modelopt/torch/quantization/ggml/registry.py
IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)}
```

Looking up a format:

```python
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY

fmt = IQ_FORMAT_REGISTRY["iq2_xxs"]
packed, shape = fmt.quantize(weight)          # GGML blocks
fmt.block_bytes, fmt.effective_bits           # 66, 2.0625
```

### Testing

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py` — **94 passed**
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed**
(RTX PRO 6000)
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses
for that suite
- broader sweep of IQ, export and recipe unit tests — **153 passed**,
none failed

**New guards on the registry itself:**
- every encoder the package exports is registered
- each record points at its own format's codec, geometry and chunk
defaults
- the public `<fmt>_fake_quant` alias is the registered record's method
- export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the
registry
- a format's `fake_quant` refuses a quantizer configured for another
format. Dispatch picks the record by `num_bits`, so it never reaches
this guard; the test covers direct callers of a record or alias. The
three per-format guards it replaced were untested on main.
- every registered format is listed in the shared test batteries

Checked by mutation: leaving IQ2_XXS out of the registry, or registering
it with the IQ2_XS encoder, each fails the guard written for that case.

**Coverage gap closed along the way:** `test_ggml_backend.py` was
hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or
packed-once coverage. Those tests now run over the registry.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — public per-format fake-quant
names are kept as aliases; the removed tables were never released.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — internal refactor with no
user-visible change
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`,
`IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently
enumerate the same formats. A common pack/dequantize/fake_quant
interface would let backend dispatch and export consume one
registration."*

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* IQ quantization formats are available through a shared format
registry, keeping format details and quantization behavior consistent
across supported workflows.
* IQ-format model exports use registered format information for
quantization metadata and weight packing.
* **Tests**
* Expanded checks to cover registered IQ formats and verify consistent
format support across quantization and export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 23:32:11 +00:00
Frida HouandClaude Opus 5 1c4cde7788 Fix: Release per-layer expert weights in layerwise export under offload (#2466)
### What does this PR do?

Type of change: Bug fix

Layerwise export leaks one layer's worth of quantized tensors per layer
on offloaded models, and dies of OOM partway through a large MoE. Two
things accumulate, both because the offload window cannot reclaim what
the export pass adds inside it.

**1. Per-expert holder modules.** `_export_fused_experts` splits a fused
MoE experts module into per-expert holders and attaches them to the live
model:

```python
proj = nn.Module();  proj.weight = wrapper.weight   # packed U8
expert.add_module(proj_name, proj)
module.add_module(str(idx), expert)
```

They are plain `nn.Module`s built *inside* the weight-access window, so
they carry no accelerate `_hf_hook`.
`weight_access_and_writeback_context` closes by iterating the modules it
collected at entry and calling `hook.post_forward()` through each one's
own offload hook — the holders satisfy neither condition, so nothing
returns them to meta.

**2. Scale buffers.** `_export_quantized_weight` registers
`weight_scale` / `weight_scale_2` / `input_scale` on the layer's
pre-existing, hooked sub-modules, and `AlignDevicesHook.post_forward`
runs with `offload_buffers=False`:

```
after post_forward: {'weight': 'meta', 'weight_scale': 'cpu'}
```

The packed weight goes back to meta; the scales do not. For
`_QuantMoELinear` models this half is the whole leak on its own —
`_reconstruct_fused_moe_linear` restacks every expert's scales into one
`register_buffer` on the hooked wrapper.

A whole-model export never notices: one pass, write the state dict,
exit. Layerwise runs the same pass once per decoder layer, so every
finished layer stays resident.

### Measurements

Through unmodified `examples/hf_ptq/hf_ptq.py` on a Qwen3.5-MoE-shaped
model (10 layers, 64 experts), printing `torch.cuda.memory_allocated()`
after each exported layer:

| placement | per layer | over 9 layers |
| --- | --- | --- |
| offload | +0.052 GiB | 0.579 → 1.047 GiB |
| resident | -0.135 GiB | falls, as designed |

+0.052 GiB is exactly one layer's quantized experts: packed U8 100.7M/2
= 0.047 GiB plus FP8 block scales 100.7M/16 = 0.006 GiB.

At scale it is fatal rather than wasteful. Qwen/Qwen3.8-2.4T-A95B (92
layers, 512 experts) leaks 12.9 GB packed + 1.6 GB scales per layer, so
92 layers want 1.33 TB that no budget on a 283 GB card or 952 GB host
absorbs. The run died of CUDA OOM at layer 16/92 with
`--max_gpu_memory_gb 240`, and at `--max_gpu_memory_gb 30` leaked the
same 15 GB/layer onto the host instead.

### The fix

`_release_exported_tensors` (`model_utils.py`) is a context manager that
snapshots each sub-module's child-module and buffer names on entry, and
on exit drops whatever appeared. Persisting happens *inside* the block,
so "release only once it is on disk" is structural rather than a
comment.

Both packing sites use it: `LayerwiseExporter.export_layer` and the
offload decoder loop in `_export_transformers_checkpoint_streaming`.

The streaming writer had solved the same leak inline with a heuristic —
null every CUDA buffer, and every CUDA parameter on a hook-less module —
and that block is deleted in favour of the shared helper. Keying on
*what the pass added* rather than on device and hook presence drops two
assumptions that only held for a terminal, offloaded export: it no
longer nulls buffers the layer already had, nor parameters of
sub-modules accelerate simply did not hook. That is also what makes it
safe for the layerwise path, where resident models are supported and the
model outlives the export.

Deliberately out of scope: the FSDP2 per-unit loop in
`collect_export_tensors` keeps its per-unit holders. That predates this
PR, this PR does not touch that loop, and closing it needs its own
change and its own test.

### Usage

No API change. Existing layerwise export under offload simply stops
growing:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3.8-2.4T-A95B \
    --qformat nvfp4 \
    --export_path /path/to/export \
    --max_gpu_memory_gb 240
```

### Testing

- `tests/unit/torch/export` and `tests/unit/torch/quantization` — 1246
passed
- `tests/gpu/torch/export/test_layerwise_export.py` +
`test_offload_export.py` — 35 passed
- `cuda_alloc` over 10 offloaded layers: 0.526 → 0.518 GiB (-0.008), was
+0.468
- Exported checkpoint byte-identical to the unfixed run: 7841 tensors, 0
mismatches, max abs diff 0.0; `hf_quant_config.json` / `config.json` /
index identical
- Full Qwen3.8-2.4T-A95B PTQ then completed all 92 layers: peak GPU 117
GB of 283, peak host RSS 71 GB of 952, flat across 50 consecutive layers
at 23-26 s/layer

Coverage added: the existing `test_export_creates_per_expert_submodules`
now runs the export inside the context manager and asserts the holders
are gone on exit, and a new test in `test_offload_export.py` pins the
`offload_buffers=False` behaviour the buffer half exists for —
export-registered scales dropped, pre-existing buffers untouched.

`tests/gpu/torch/export/test_fsdp2_export.py` reports 34 failures in my
environment. They are **pre-existing and unrelated**: the same 34 fail
identically on this branch and on the merge-base (`2b1f33d0ef`), with
byte-identical failure sets and runtimes within 3 s. All 34 are `Failed:
Timeout (>120.0s)` from hung NCCL collectives, with zero assertion
failures.

The GPU export suites and the whole-model measurements above were run at
`2999d7cfb7`. The two commits since — swapping the holder marker for a
child-name diff, and moving the helper to `model_utils` — are covered by
the unit suites; a GPU re-run before merge is worthwhile.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one new offload test for
the buffer half, and the existing fused-experts export test now covers
holder release
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — `layerwise.export_dir` is new in the unreleased 0.48.0, so this
bug was introduced and fixed within the same cycle
- Did you get Claude approval on this PR?: ✅

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 14:43:12 -07:00
Chenjie LuoandClaude Opus 5 a21411adde Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do?

Type of change: new feature

llama.cpp defines five GGML IQ formats at one and two bits; we ship two.
This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and
IQ2_XS, and is the **first of three**.

On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`,
`Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84
B parameters — 10.6% of the file**, which a reader limited to
IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats
account for 17.3%.

| format | bpw | bytes/256 | codebook | |
|---|---|---|---|---|
| `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing |
| **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this
PR** |
| `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing |

The encoder follows the existing single-pass grid search at a fixed
anchored super-block scale, and the CUDA kernel the existing per-block
structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a
4-bit sub-block scale into the same 32-bit word as four 7-bit sign
indices, and its 256-entry grid needs no high index bits.

### Groundwork the next two reuse

Two things land here because IQ2_XXS is the first format to need them:

- **Export registry.** The IQ family was spelled as a two-element tuple
at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and
`unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset
plus per-format packer and block-geometry tables, so a format is a row
rather than a sweep through the exporters.
- **Shared test contract.** The per-format test files had drifted apart
— each of `iq1_s` and `iq2_xs` tested things the other did not. They
become one parametrized module per layer (unit and CUDA), so every
format is held to the same contract and a new one inherits it.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs
```

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared
against `dequantize_row_iq2_xxs` from `ggml-quants.c`:

```
IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry, as
does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint
ship as conformance vectors so CI keeps checking bytes we did not
produce; mutation testing confirms they catch a wrong sign-field width.

The CUDA encoder is byte-identical to the PyTorch reference on a fixed
input and runs at **1047.9 M elem/s against the torch search's 10.7** on
a 5632×2048 weight.

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99
passed
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7
checks × 3 formats)
- `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds
29 recipes, `ptq.md` updated
- reconstruction error decreases monotonically with bit width, pinned by
a test

Pre-existing failures in `tests/unit/torch/export/` and
`test_autoquant.py` are `transformers`/`torchvision` import problems in
my environment — identical counts with and without this change.

### A finding about already-merged code

Checking the new kernel against its PyTorch reference at 4096 blocks
showed that **CUDA and torch encoders disagree on roughly 1 block in
6000 — including the already-merged `iq2_xs`**, at 0.0163% against
IQ2_XXS's 0.0000%.

Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA
fuses it with `fmaf` while torch uses separate ops; where two local
scales fall within a float32 ULP the roundings pick different sides.
Adjudicated against float64, neither path is better (5 to 6). Worst-case
cost is **1.48e-08** relative reconstruction error, and run-to-run
determinism on a given device holds.

This is pre-existing, not introduced here — `test_iq2_xs_cuda.py`
asserts exact byte parity but on a 16-block weight where ties
essentially never arise. I have **not** changed that test; rewording a
guarantee on merged code belongs in its own change. The new shared GPU
tests assert exact parity on a small fixed input and compare
reconstruction error at scale.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new
codebook is a GGML table, carried in `codebooks.py` beside the existing
ones so the MIT-licensed surface stays in that one file, with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

First of three; **IQ2_S** and **IQ1_M** follow and build on this branch.
Replaces #2505, which carried all three at once. Follows #2446 / #2447 /
#2448 / #2449, which landed IQ1_S and IQ2_XS.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added IQ2_XXS weight-only quantization, including CUDA acceleration
and support for Hugging Face and Megatron exports.
- Added the `general/ptq/iq2_xxs` recipe. It requires no calibration
data and supports eligible layers with a weight dimension divisible by
256.
- Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at
approximately 1.56, 2.06, and 2.31 bits per weight, respectively.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 18:04:37 +00:00
Keval MorabiaandClaude Opus 5 f2f0d6958e Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?

Type of change: new feature

`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.

Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:

- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.

Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.

One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.

The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).

### Usage

```bash
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --recipe general/ptq/nvfp4_default-kv_fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
    --mlflow https://<your-mlflow-server>/

# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```

The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.

### Testing

- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.

- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:25:30 -07:00
h-guo18andClaude Opus 5 87f7d1432f fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do?

Type of change: Bug fix

**Follow-up to #2342**, which split this out on review (commit
`c67784d9`), and a rethink of how
the flag is implemented.

`dflash_fp32_master_weights` exists because the DFlash draft is cast to
the frozen bf16 target's
dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse
to hold them. At
`beta2=0.999` a single step changes `v` by at most **0.100%**, while the
smallest change bf16 can
represent near `v` is **0.164% mean / 0.388% max** (measured): every
decrease rounds away, `v` only
grows, and the effective step size decays on its own from step 1.

#2342 fixed that by **promoting the draft model to fp32**. Everything
else followed from giving the
model a dtype the rest of it does not have — a bf16 autocast at every
entry point, two transformers
loader hints so `from_pretrained(dtype="auto")` would not round the
draft away, a post-condition
check because those hints fail silently, and a doubled DDP gradient
all-reduce.

**This PR puts the fp32 in the optimizer instead**, where Megatron-LM,
DeepSpeed and apex put it.
`MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter
plus fp32 moments in
`self.state[p]`, steps on the master, and copies back at the parameter's
dtype. The model is never
anything but the base dtype, so every one of those follow-on pieces is
deleted, gradients stay
bf16, and the exported drafter is unchanged. What the placement costs is
that wiring the optimizer
becomes the training loop's job:
`EagleTrainerWithAccLog.create_optimizer` builds it, and
`VerifyMasterWeightsCallback` raises at the end of step 1 if the moments
are not fp32.

**The default flips to `True`** — the flag now changes optimizer memory
and optimizer arithmetic
and nothing else. Flipping it on the model-promoted implementation turns
**25 of 259** unit tests
red; flipping it here is **259 passed**. Set it to `False` to reclaim
the memory, about 12 bytes per
draft parameter instead of 4.

<details>
<summary>Three drive-by fixes, independent of the above</summary>

- `_place_draft` is folded back into `modify()` — it fused the draft's
dtype, its device and an
  eager rotary buffer behind one meta guard.
- The module docstring's claim that `DFlashModule` has an `_apply`
meta-buffer fix is removed
  (`grep "def _apply"` matches nothing, and never did).
- #2342's field description no longer lists `evaluation` as a broken
path — `forward`
short-circuits to the base model when `not self.training`, so the draft
never runs there.

</details>

### Usage

No API change. `dflash_fp32_master_weights` now means the *optimizer*
holds fp32 master weights
rather than the draft model being fp32.

### Testing

**1 · The refactor is arithmetically a no-op.** Both implementations run
AdamW on an fp32 tensor,
so given the same starting values and the same gradients the
trajectories are identical — 1000
steps, `weight_decay=0.01`:

```
old fp32 parameter  vs  new fp32 master : bitwise equal = True  (max |diff| 0.0e+00)
exp_avg / exp_avg_sq                    : bitwise equal = True
optimizer state dtypes                  : ['torch.float32']
model parameter dtype                   : torch.bfloat16
```

Initial values have to be matched at bf16 first, or the bf16 arm's
one-time rounding of the draw
shows up as a 2e-4 "difference" that is not arithmetic. With that
controlled, the two
implementations differ only in their *inputs*: gradient precision (fp32
vs bf16 — torch 2.10
requires `grad.dtype == param.dtype`) and that one-time rounding.

**2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B
base, real corpus, one GPU
per arm, three arms — pure bf16 (flag off), the #2342 implementation,
and this one — on two
algorithms trained independently, sharing seed, data order and
initialisation within an algorithm.

<img width="2925" height="960" alt="image"
src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156"
/>


The two fp32 arms sit on top of each other for the whole run while bf16
stays above both, and the
old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to
be compared against.

**Acceptance length says the same thing, and settles what the loss could
not.** All six drafters at
the end of those curves were exported and served under vLLM against the
same base, and measured on
MT-Bench (80 prompts, 8 categories, greedy, one request at a time,
`num_speculative_tokens` =
trained `block_size` − 1, every knob but the drafter held fixed):

| | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) |
new − old | fp32 − bf16 |
|---|---|---|---|---|---|
| `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 &nbsp;`t=−0.90` |
+0.0619 &nbsp;`t=+13.8` |
| `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 &nbsp;`t=+1.42` |
+0.0312 &nbsp;`t=+9.2` |

Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI
straddles zero (`dflash`
[−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16
does not come close to it,
and the sign of new-vs-old **flips between the two algorithms** — what a
rounding difference looks
like, not a bias. This is also the comparison the training loss could
not give: all three arms are
**exported and served in bf16**, so the old implementation's fp32 draft
weights are rounded at
export exactly as they would be for deployment, and the "its loss was
computed on a more precise
forward" caveat below does not apply. `lilicorr` needs

[vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934),
applied as an overlay so
that both algorithms are measured on one engine build.

The right panel is the mechanism, and the one signal that depends on
neither the seed nor the
choice of loss statistic: Adam's updates to the draft's RMSNorm gains
are smaller than the bf16 ULP
at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and
the gains never move — not
one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by
3e-06. Both fp32 arms move
all of them, by the same amount.

Two results behind the figure rather than in it. **fp32-vs-bf16 grows
with the horizon** while
old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000
→ −0.341 at 30000, and on
`lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference
that stays near 0.02 at every
horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`).
That is what a compounding
bias and a rounding difference respectively should look like, and it is
the reason the longer runs
were worth doing. And
**across seeds**, the paired old-vs-new difference at 1500 steps is
+0.0003 (n=6) on `dflash` and
+0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero.

<details>
<summary>Limits of the above, stated rather than smoothed over</summary>

At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557
(t=+2.89, 4/5 seeds in the same
direction) — nominally significant, suggesting the new implementation
was genuinely worse there.
Four further `lilicorr` seeds were run against that pre-declared
question; two came back strongly
negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]).
The earlier reading was
small-sample noise.

At 1500 steps on `lilicorr` that CI is *not* narrower than the
fp32-vs-bf16 effect it is being
compared against, so the 1500-step sweep alone cannot certify
equivalence there — `lilicorr` is
still at loss 9.3 and deep in its early transient, and it is the long
runs that resolve it. On
`dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect
(−0.129).

One asymmetry the loss comparison cannot separate: the old
implementation held the draft weights in
fp32 *at forward time*, so its training loss was computed on a more
precise forward, while both
implementations export bf16. Any residual advantage it appears to have
is therefore an upper bound.

</details>

**3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on
**transformers 5.0.0** and
**5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10).
`TestDFlashFp32MasterWeights`
is rewritten for the new mechanism; the two that would have caught the
traps in this design are
`test_resume_does_not_round_the_master_back_down`
(`Optimizer.load_state_dict` casts float state to
its parameter's dtype, so a naive subclass rounds the master and both
moments to bf16 on *every*
resume, silently, with the loss still falling) and
`test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest
cover the draft's dtype with
the flag either way, that no forward path needs an autocast any more,
that plain AdamW really does
leave the moments in bf16, and that an fp32 model allocates no redundant
master. A sharded FSDP2
`DTensor` keeps an fp32 master and fp32 moments through a step.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ for artifacts, with one
intentional default change.
The draft's stored dtype goes back to matching the base, as it was
before #2342; existing
checkpoints load unchanged and the exported drafter is unaffected. The
flag now defaults to
**`True`** — the measurements above are the reason, and the cost is fp32
master + fp32 moments for
the draft only. A training loop that builds its own optimizer instead of
using the shipped
`create_optimizer` gets plain AdamW and none of this;
`VerifyMasterWeightsCallback` makes that
  fail loudly at step 1 rather than skip the feature quietly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — draft; will run `/claude
review` before marking ready.

### Additional Information

**On the 7–14% acceptance-length gain quoted in #2342:** that was
measured on the fp32-model
arithmetic and is not re-derived here. What is measured above is the
like-for-like comparison this
PR has to answer — same corpus, same horizon, same serving path, one
implementation swapped.

**History:** commits 1–3 restore the autocast design as it was split
out; commits 4–6 replace it.
Happy to squash before review.

<details>
<summary>Alternatives measured and rejected, so they do not get
re-proposed</summary>

- **Swapping `p.data` to the master and calling `super().step()`**
(reuses all of AdamW, ~20 lines
instead of ~50): bit-identical on ordinary parameters over 25 steps, but
silently wrong under
FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's
reported dtype while the
local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype`
reads bf16, and
`zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass
either way.
- **Narrowing the autocast from `__call__` to `forward`** (while it
still existed): turns 10
Domino/DSpark tests red — the variants apply their heads in their own
`forward` overrides,
  outside `DFlashModule.forward`.
- **Building the rotary buffer on meta and letting the loader
materialise it**: makes RoPE
correctness depend on transformers selecting a branch by class-name
substring
(`"RotaryEmbedding" in module.__class__.__name__`), and the `if not
hasattr` guard is then
permanently satisfied, so a later `to_empty()` leaves garbage forever —
measured `4.56e-41`, i.e.
  cos=1 / sin=0, no positional encoding at all.
- **Building it eagerly in `DFlashModule.__init__`**: lands before the
dtype cast, so `Module.to`
rounds the RoPE frequencies to bf16 on the default path — measured
`0.8659643530845642` →
  `0.8671875`, loss `3.47230935097` → `3.47114777565`.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- DFlash now uses FP32 optimizer master weights and Adam moments by
default while keeping draft parameters in the base model’s dtype.
- Master-weight training preserves optimizer precision when restoring
checkpoints.
  - The feature can be disabled to reduce optimizer memory usage.
  - Draft models consistently follow the base model’s dtype and device.

- **Bug Fixes**
  - DFlash workflows now support operation without autocast.
- Added validation for compatible AdamW-family optimizers and
master-weight precision, including resumed training runs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 15:39:33 +08:00
1b4e7dfb14 [OMNIML-5899] Add IQ post-training quantization recipes (#2449)
## Summary

- add numerics configs for IQ1_S and IQ2_XS
- add model presets and general PTQ recipes
- document the supported weight-shape requirement
- reuse the packed IQ weight across forwards instead of re-encoding it
every time
- add end-to-end `hf_ptq` coverage for both formats
- add the release-note entry

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

Sixteen files. The PR started as eight recipe config, documentation and
changelog files; the end-to-end test added for them surfaced two
performance
bugs in the already-merged codec, and fixing those pulled in the codec
files
and their tests.

Beyond the original recipe set it now touches four files owned by #2446
—
`ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and
`ggml/backend.py` —
plus three test files. It still contains no kernel or export files.

Keeping those fixes here rather than moving them to #2446 is deliberate
and
confirmed with the stack owner: #2446 is already merged, and both bugs
are only
observable through the end-to-end test this PR adds, so splitting them
would
separate each fix from the test that demonstrates it.

## Packed-weight cache fix

Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S
`hf_ptq`
run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing
the same
154 weights 15400 times — 100 times each.

The repacking is not calibration. These recipes set `algorithm: null`
and
`hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100
passes are
the sample `generate()` calls `hf_ptq` makes before and after
quantization: one
decode step re-runs weight fake-quant on every linear, and the packed
payload
was thrown away each time.

`_PackedWeightCache` was already there to prevent exactly this, and it
never
hit. It keyed on the identity of the tensor the backend was handed, but
`TensorQuantizer` passes a fresh *view* of the weight on every forward,
so the
identity check never matched twice.

The fix anchors the entry to `inputs._base` — the parameter the view is
taken
from — held as a weakref. The parameter is stable across forwards, so
the cache
hits; the reference stays weak, so the payload is released with the
weight and
offloaded/meta-device flows are unaffected. (A strong reference does
make the
cache hit, but pins full-precision storage for the life of the
quantizer, which
is the opposite of what those flows need.)

Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and
after:

| | packer calls | packing time | wall clock |
|---|---|---|---|
| before | 15400 | 499.0 s | 8m57s |
| after | 154 (one per weight) | 6.2 s | 3m21s |

`test_ggml_weight_is_packed_once_across_forwards` pins this: it counts
encoder
calls across five forwards under `torch.inference_mode()` (what
`generate()`
runs under) and asserts exactly one.

## Decode chunk fix

With packing cached, the end-to-end cost moved entirely into the decode,
and
IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case).
Instrumenting
both showed packing was no longer the cost at all — IQ2_XS packs
*faster*:

| | pack calls | packing time | wall clock |
|---|---|---|---|
| IQ1_S | 154 | 6.2 s | 3m21s |
| IQ2_XS | 154 | 1.6 s | 17m59s |

The cause was one constant serving two loops with opposite
characteristics.
`_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which
holds the
large codebook-search temporaries and runs once per weight; IQ2_XS sets
it to
256 rather than IQ1_S's 1024 because its search sweeps sixteen local
scales per
grid tile. But the *decode* shared it — and the decode has tiny
temporaries,
runs on every forward, and is never cached, so a small chunk only
multiplies
kernel launches.

Decoding a 2048x5632 weight:

| chunk | IQ2_XS decode | transient peak |
|---|---|---|
| 256 (was) | 91.4 ms | +24 MiB |
| 1024 | 22.9 ms | +31 MiB |
| 4096 (now) | 5.8 ms | +56 MiB |

The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded
through
`fake_quantize_with_cache`. The encode bounds are untouched, so the
memory
ceiling stays where it was aimed. The seven-iteration sign-parity loop
in
`dequantize_iq2_xs` is also folded into three XOR steps, off the same
per-forward path.

Per case in `tests/examples/hf_ptq` on 2xH100:

| | before | after |
|---|---|---|
| IQ1_S | 259.65 s | 94.64 s |
| IQ2_XS | 1081.05 s | 102.77 s |

Both now fit the 300s `tests/examples` default, so the cases carry no
explicit
timeout.

## Integration contracts

- `block_sizes: {-1: 256}` records the native packed-block contract; it
does not drive the GGML fake-quant scale search. The export path reads
`TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`,
validates it against the format block size, and records `group_size:
256` in checkpoint metadata.
- The model presets intentionally expose `--qformat iq1_s` and
`--qformat iq2_xs` in `hf_ptq`.

## Testing

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed,
1 skipped
- `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load
with the expected backend and no unused search option
- `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2
passed on 2×H100 (TinyLlama, both formats end to end through export),
3m17s for the pair
- pre-commit hooks pass on all changed files
- larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers)
quantizes
and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself
192s

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:46:50 +00:00
yingguo-trt 5bb7343592 [https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect disabled quantization during calibration (#2434)
### What does this PR do?

Type of change: Bug fix

During max calibration, `enable_stats_collection()` calls
`disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ
attention paths bypass `TensorQuantizer.forward()` and previously
selected the Triton/Kitchen path from the configured enabled state
alone, so P quant-dequant could still execute while quantization was
inactive.

This change:

- enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is
true;
- bypasses all fused P-QDQ paths when the P quantizer is disabled or
quantization is inactive;
- adds focused parameterized regression tests covering Triton and
Kitchen dispatch across enabled, `disable()`, and `disable_quant()`
states.

### Usage

N/A. This restores the existing `disable_quant()` contract and does not
introduce a new API.

### Testing

- `pytest
tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14
passed
- Targeted pre-commit checks passed
- Kitchen coverage verifies the full `{disable(), disable_quant()} ×
{lazy, already initialized}` dispatch matrix remains bypassed while
inactive and resumes after re-enabling
- Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides
on B300: [module regression build
#121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/)
- Controlled B300 isolation passed when only the P-BMM quantizers were
hard-disabled during calibration, and also passed when the existing
single Triton attention configuration was forced
- Fixed-code B300 validation passed in three independent runs: [Jenkins
build
123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/),
[Jenkins build
124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/),
and [Jenkins build
125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/).
Jenkins build 123 compared all 178 module outputs byte-identically and
found 0/377 `amax` and 0/377 `scale` changes.

The confirmed impact is incorrect calibration behavior plus unstable
quantizer state and module outputs; downstream benchmark accuracy impact
has not been established.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

The failure was isolated to fused P-QDQ runtime-state dispatch.
Quantizer topology and configuration were identical in the failing
same-wheel comparison.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Corrected attention dispatch so disabled or inactive quantization uses
the original attention implementation.
- Preserved optimized quantized attention when quantization is enabled.
- Ensured attention masks remain unchanged when using the original
attention implementation.
- Improved fallback behavior for fused attention paths, including
correct initialization and reuse when quantization is re-enabled.

- **Tests**
- Added coverage for enabled and disabled quantization states,
attention-mask handling, fallback selection, and result consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com>
2026-09-22 06:51:23 +00:00
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00
Shengliang Xu f1abc75626 Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?

Type of change: bug fix + new tests

Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.

#### The problem

`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.

#### The fix

`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.

#### Two mechanisms, disjoint by construction

Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.

| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |

A tensor is never both copied and carried; that would export it twice,
and a test asserts it.

#### Removed as redundant

`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.

#### Two deliberate behavioural changes

**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.

**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.

### Usage

No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:

```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```

### Testing

`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.

The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.

All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.

Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.

### The algorithm: which weights get carried, and how

Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.

**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.

**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.

#### The index is not an inventory of the checkpoint

This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.

`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:

| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |

So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.

`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.

#### Flow

1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.

#### What is deliberately excluded

- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.

### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`

Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.

The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.

The two cases now differ only in *why* the module is invisible to the
quantizer walk:

| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |

`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.

An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.

Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.

`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.

* **Behavior Changes**
  * Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
  * MTP-specific export exclusions are no longer applied.

* **Bug Fixes**
  * Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-21 14:00:34 -07:00
yueshen2016andClaude Opus 5 fc4c40fcbe Fix HF export crash when a dynamic-block quantizer has zero amax (#2438)
### What does this PR do?

Type of change: Bug fix

`TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized
for dynamic-block
quantizers, while the static path immediately below it has always
substituted `maxbound` for
zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a
recipe that applies it to
an *activation* quantizer — e.g.
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets
`*mlp*input_quantizer` — feeds a raw `0.0` into
`NVFP4QTensor.get_activation_scaling_factor`,
whose assert aborts the entire export:

```
AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj'
(type=QuantLinear):  activation scaling factor 0.0 not positive.
```

Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE
expert — saw only zeros, so
one dead layer costs the whole run at the final export step.

This factors the substitution into `_sanitize_export_amax()` and calls
it from both branches. Two
details beyond de-duplication:

- **Branch-free, so it survives a meta `amax`.** `torch.where` +
`nan_to_num` both have meta
kernels; `bool()` on a meta tensor raises. The layerwise and streaming
export flows carry meta
`amax` — `validate_attr` short-circuits on `is_meta` for exactly that
reason — so only the
  warning is gated on a materialized tensor.
- **No longer mutates calibrated state.** The old in-place `amax[amax ==
0] = ...` wrote through a
view of `self._amax`; `torch.where` returns a fresh tensor, so that
hazard disappears.
- **Warns, with a count.** The fix turns a loud failure into a silent
one, and a zero amax means
calibration never activated that layer — worth surfacing rather than
papering over. The message
reports how many entries were substituted, since per-location dedup
otherwise collapses many
dead experts into one uninformative message. A healthy model emits none.

Scope: only the activation path is data-dependent and reachable this
way. Weight-side `_amax` uses
are left alone, since a weight amax of 0 would require an all-zero
weight matrix.

**Knowingly left as follow-up:**
`export/quant_utils.py::get_scaling_factor` discards the sanitized
`amax` when `num_bits == (2, 1)` and recomputes via
`get_weights_scaling_factor_2_from_quantizer`,
which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input*
quantizer on a module whose
*weight* quantizer is a different format (or disabled) can still trip
`assert torch.all(scaling_factor > 0)`. Format dispatch is
weight-driven, so the reported recipe
does not reach that branch; fixing it properly changes a signature
shared with the weight-side
callers and is out of scope here.

Not a regression. The dynamic early return, the `type: dynamic` numerics
unit, and the recipe that
combines them all ship in released 0.46.0 / 0.46.1.

### Usage

No new or changed API. Exports that previously aborted now complete and
warn:

```python
# Recipe applies dynamic NVFP4 to *mlp*input_quantizer; layer 37 never activated during calibration.
mtq.quantize(model, quant_cfg, forward_loop)
export_hf_checkpoint(model, export_dir=out)   # before: AssertionError; now: exports + UserWarning
```

### Testing

- New `test_amax_export_unusable_amax`, parametrized over zero and NaN,
covering the
dynamic-NVFP4 and static per-tensor configs; asserts the exported scale
is positive and that
export leaves the calibrated `amax` untouched. Plus
`test_amax_export_meta_amax`, pinning that
a meta `amax` survives export rather than raising. Both run on CPU and
CUDA via the shared tester.
- `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40
passed.
`tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed
(GB300).
- End-to-end repro on GB300, small Llama with one MLP fed all-zero
activations under
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()`
`0.0` → `6.0`, live layer
unchanged at `3.921875`, and `export_hf_checkpoint` goes from the
`AssertionError` above to
  writing `model.safetensors`.
- Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on
a healthy model
(Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal
path is unaffected.
- `pre-commit run` clean on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; its
one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs
addressed or answered in 251f2e3d

### Additional Information

Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter
also notes it passed on
0.47.0rc0; that is not explained by code — `git diff
0.47.0rc0..0.47.0rc1` touches
`export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new
`clamp_fp8_scales` argument
whose default preserves the old behaviour) and the INT4-AWQ packing
path, neither of which is on
the dense-HF NVFP4 activation-scale path. Whether `amax` lands on
exactly 0 is
calibration/model-state dependent, which is what makes it look
version-flaky.

Worth flagging separately: in the reported log the **pre-PTQ** sample
output is already
gibberish, so that BF16 checkpoint looks broken independently of
quantization. This change stops
the crash, but such a run will now export a valid-but-garbage checkpoint
— the new warning is the
signal to investigate.

Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing
release.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed Hugging Face checkpoint export when dynamic-block quantizers
have zero or invalid calibration scales.
* Exports now use a positive fallback scale and issue a warning instead
of failing when applicable.
* Export operations no longer modify the original calibrated quantizer
state.
* Meta-device exports remain non-erroring and preserve device placement.
* **Tests**
* Added coverage for zero- and invalid-scale exports across dynamic and
static quantization modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Yue <yueshen@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:47:34 +00:00
Keval MorabiaandClaude Opus 5 b311c054de Fix hybrid stack spec serialization in Megatron-Bridge checkpoints (#2452)
### What does this PR do?

Type of change: Bug fix

Hybrid (e.g. Nemotron-H) checkpoints saved by the
`examples/megatron_bridge` scripts could not be reloaded or exported to
HuggingFace:

TypeError: MLPSubmodules.__init__() missing 2 required positional
arguments: 'linear_fc1' and 'linear_fc2'

**Root cause.** Megatron-LM's YAML writer represents a
`functools.partial` via `_partial_representer`, which passes each
keyword value through `represent_data`. A dataclass instance has no
representer, so it falls through to `_safe_object_representer`, which
emits only `{_target_, _call_}` and drops every field. The default
hybrid stack spec builds its dense-MLP and MoE layers as exactly such
partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec`
on the provider — which is serialized into every checkpoint's
`run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written
with no fields at all. Both export paths hit it: `convert.sh` via
`from_auto_config`, and `export_distilled_megatron_to_hf.py` via
`export_ckpt → load_megatron_model`. Every hybrid provider is affected,
dense or MoE.

**Fix.** `set_moe_expert_layout()` stores a named, zero-argument
*factory* instead. The provider already calls a callable spec at build
time (`_resolve_hybrid_stack_spec`), so model construction is unchanged
— only the serialized form differs:

```yaml
  hybrid_stack_spec:
    _call_: false
    _target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec
```

The grouped-GEMM factory is Megatron-Bridge's own
`transformer_engine_hybrid_stack_spec`, so stock tooling
(`scripts/conversion/convert.sh`) resolves it without importing
ModelOpt.

**Known limitation — the two layouts are not symmetric.** The
SequentialMLP layout has no bridge-side equivalent (the upstream TE
hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt
target, which `instantiate` only accepts in a process that has imported
`mbridge.py` and thereby run `register_allowed_target_prefix`. A
SequentialMLP hybrid checkpoint therefore converts through the ModelOpt
entrypoints but not through stock `convert.sh`, where it fails on the
disallowed prefix instead of on `MLPSubmodules` — no regression, but
that path stays broken for this one layout. The reach is narrow:
`use_moe_grouped_gemm()` returns True for any architecture with a
grouped-expert export rule, NemotronH included, so SequentialMLP
requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs
an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge.

Both spec builders also move out of `nas/plugins/megatron.py`, which
never used them, into a new `utils/plugins/megatron_layer_specs.py`
beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that
module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is
reached by 16 test files through
`tests/_test_utils/torch/megatron/models.py`, which is bridge-free.

The underlying defect is upstream in
`megatron/training/config/yaml_utils.py`; this only stops ModelOpt from
stepping on it, so it is worth a separate Megatron-LM issue.

### Usage

No API change — hybrid checkpoints saved after this fix convert with the
existing commands:

```bash
torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \
    --student_hf_path <student_hf_model_or_path> \
    --megatron_path   <distill_out>/checkpoints \
    --hf_export_path  <hf_out> \
    --export_iterations all
```

### Testing

Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B
Nemotron-3.5-Lightning pruned+distilled run:

- **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a
real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml`
(the writer used for `run_config.yaml`), reloaded via `instantiate`,
resolved. `moe_grouped_gemm=True` →
`TELayerNormColumnParallelLinear`/`TERowParallelLinear` +
`TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays
callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a
saved config cannot regress.
- Applying the equivalent `run_config.yaml` fix to 32 iteration
checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104
experts) with populated `MLPSubmodules` / `MoESubmodules`.
- **End-to-end exports**, 6 iterations, all rc 0, each producing exactly
the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards /
41.5 GiB, all weights finite, drift from the base rising monotonically
with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU,
`convert.sh` GPU (4×GB300, TP=4), and
`export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU
123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still
builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side
conversion.
- `ruff check` / `ruff format --check` passed on the source files before
the module move.

**Not yet run:**
`tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) —
the GPU allocation expired. It asserts the round-trip property verified
manually above, but its `HybridModelProvider(num_layers=2,
hidden_size=64, num_attention_heads=4)` construction is unverified. The
module move is verified only by reference grep and syntax check, so
please also run one test that uses
`tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run
either (unavailable in the environment used).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ behavior; note
`get_te_hybrid_stack_spec` moved module
(`modelopt.torch.nas.plugins.megatron` →
`modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint
from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the
changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (added, not yet executed —
see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
before marking ready.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Hybrid checkpoints now record complete layer specifications in
`run_config.yaml`.
  * Recorded specifications support conversion to Hugging Face format.
* Hybrid MoE configurations support grouped-GEMM and sequential-MLP
modes.
* Configuration-based reconstruction preserves the selected MoE layout.

* **Compatibility**
* Checkpoints from earlier releases may require manually setting the
hybrid layer specification before conversion.

* **Tests**
* Added coverage confirming hybrid specifications survive configuration
serialization and can be recreated successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 13:45:21 +05:30
ed5c5ed369 [OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary

- add IQ format metadata and packed-weight export
- support Hugging Face and TP=1 Megatron export paths
- reject fused-MoE IQ export until a deployment loader owns its packed
layout
- document the shaped `uint8` weight contract and the fused-expert
boundary
- add Hugging Face, Megatron, metadata, and fused-expert export tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only export code, deployment documentation, and export
tests. It targets `main` and should merge after #2448 and #2446. It does
not contain kernel, codec/backend, or recipe files.

## Deployment consumer boundary

Dense weights and individually named expert weights use the documented
shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally
rejected with `NotImplementedError`: its payload would have shape
`[num_experts, out_features, in_features // 256, payload_bytes]`, and no
deployment loader in this stack currently owns that layout. Support
should be enabled only with a loader integration test.

## Test coverage

- [Hugging Face packed-weight
export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py)
- [quantization
metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py)
- [Megatron unified export and fused-MoE
rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py)

## Validation

- all pre-commit hooks pass for the changed files
- 89 focused Hugging Face export, metadata, and fused-expert tests pass
locally
- direct checks cover both fused-MoE export entry points for IQ1_S and
IQ2_XS
- Megatron GPU execution remains delegated to GPU CI
- restricted-term scan passes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for IQ1_S and IQ2_XS GGML quantization formats in
unified Hugging Face and Megatron exports.
* Added quantization metadata, tensor-shape recovery, packing details,
and IQ2_XS size documentation.
* Added validation for required block sizes and tensor parallelism
settings.

* **Limitations**
  * Fused-MoE and GPT-OSS IQ expert packing are not supported.
* IQ exports require standard `weight` attributes in Hugging Face
models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 00:00:33 +00:00
kaix-nv d23030f91d [1/n] Adds skip-softmax calibration through the vLLM serving path (#1992)
### What does this PR do?

Type of change: new feature

Calibrates skip-softmax thresholds through the vLLM V1 execution path
for FlashAttention and FlashInfer. Calibration measures the paged
KV-cache path used at serving time, aggregates raw skipped/total tile
counts across tensor-parallel head shards, fits separate prefill and
decode curves, and exports the existing `sparse_attention_config`
checkpoint schema.

This uses raw counts rather than averaging per-rank sparsity ratios
because TP ranks can contribute different tile populations; summing
numerators and denominators before division preserves the global
tile-weighted result. The vLLM adapter lives in
`plugins/sparse_attn_calibration.py` rather than
`SparseAttentionStatsManager`: the latter records module-local ratios
for the HF calibration flow and has no aligned cross-process merge
contract, while this path must merge per-sample raw counts from vLLM
workers. Fitting and export still reuse `DynamicThresholdCalibrator` and
the canonical conversion helpers so the model and checkpoint schema do
not fork.

Skip decisions depend on tile geometry. The common Triton launch
boundary fixes the KV tile at 128 tokens and the prefill query tile at
128 tokens, including for direct kernel callers. Single-query decode can
use a 16x128 compute tile without changing its skip decision.
Measurement bypasses autotuning; serving still tunes warp and
pipeline-stage counts while keeping the decision geometry fixed.

### Usage

```bash
python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \
  --prompts_file prompts.txt \
  --target_sparse_ratio 0.7 \
  --fit_logspace \
  --tensor_parallel_size 4 \
  --decode_tokens 32 \
  --update_checkpoint_config
```

Calibration supports tensor parallelism and requires pipeline-parallel
and data-parallel sizes of 1. It always writes
`sparse_attention_config.json`; `--update_checkpoint_config` also merges
the result into `<CKPT>/config.json`.

### Testing

Latest revision `1e969cb380` (rebased onto main `02b58eb146`,
2026-09-17):

- Calibration/count-fitting unit tests: **33 passed**
(`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`).
- Paged and contiguous calibration GPU suite: **33 passed**
(`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including
NHD/HND equivalence, partial query tiles, decode counts, and
malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX
A6000; local GPU 0 was unavailable.
- Calibration CLI tests: **21 passed**
(`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`).
- `pre-commit run --files <four changed files>`: passed, including Ruff,
mypy, and Bandit.
- The new regression tests reproduced the skipped-counter truncation and
missing cache-boundary checks before the fix. Calibration arithmetic and
the 20-point threshold grid are unchanged.

Historical validation from earlier revisions (not rerun end-to-end for
this update):

- `PYTHONPATH="$PWD" pytest -q
tests/examples/vllm_serve/test_calibrate_sparse_attn.py
tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py`
— 37 passed.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
— 65 passed, including kv-first, blocks-first, and packed FlashAttention
cache layouts.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py
tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
— 33 passed.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py
tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py
tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
— 31 passed, 1 skipped because the GPU lacks enough shared memory for
the fp32 tile.
- `pre-commit run --files <changed files>` — passed.
- Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4,
48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill
`(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus
the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy
fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by
only -0.201%, while `a` retains the known grid-weighting shift.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ Active skip-softmax fixes the
calibrated decision geometry (serving still tunes warp/stage counts),
and sparse-only vLLM installs fail fast for unsupported DCP,
DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead
of installing silently.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code or new dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pipeline parallelism is rejected during calibration because the current
count-merging contract aligns records across tensor-parallel head
shards, not across pipeline stages with disjoint attention layers. The
unrelated HF padded-query behavior change was removed from this PR so it
can be reviewed independently with its own compatibility test.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added vLLM skip-softmax calibration for paged attention, including
prefill/decode support and checkpoint configuration generation.
- Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3
conversion, and NVFP4 activation headroom calibration recipes.
- Added calibration statistics aggregation, phase-specific fitting,
threshold validation, and preservation of existing sparse-attention
settings.

- **Bug Fixes**
  - Improved NVFP4 CPU/ONNX scale validation and clamping.
- Added clearer handling for unsupported quantization, cache, CUDA
graph, and engine configurations.
  - Standardized serving and calibration tile behavior.

- **Documentation**
- Expanded vLLM serving guidance, calibration instructions,
compatibility requirements, and sparse-attention limitations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-09-18 12:31:10 -07:00
Ajinkya RasaneandCodex cf1f48fa0f [5565357] Fix SDXL NVFP4 export and performance (#2336)
### What does this PR do?

Type of change: Bug fix

Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:

- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.

For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.

This PR also changes shared exporter behavior:

- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.

Other model recipe configurations remain unchanged.

### Usage

```bash
python quantize.py \
    --model sdxl-1.0 \
    --model-dtype Half \
    --trt-high-precision-dtype Half \
    --format fp4 \
    --block-size 16 \
    --batch-size 2 \
    --calib-size 128 \
    --n-steps 20 \
    --quantized-torch-ckpt-save-path ./sdxl-fp4 \
    --onnx-dir ./onnx-sdxl-fp4
```

### Testing

- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
  - 302 native block-scaled NVFP4 GEMM tactics;
  - 38 native FP8 Conv tactics;
  - no FP4 Q/K/V projections;
  - all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [5565357]

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.

- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-18 17:22:28 +00:00
Ajinkya Rasane 17ef5b6c39 Split calibrated ONNX graph capabilities (#2468)
### What does this PR do?

Type of change: new feature

Splits the calibrated ONNX graph utilities into four cohesive capability
owners:

- `graph_indexing.py` owns read-only graph indexing and pattern
matching.
- `graph_selection.py` owns calibrated placement decisions and the
runtime probes needed to make those decisions;
`get_extended_model_outputs` stays with selection for that reason.
- `graph_rewrites.py` owns graph transformations that are independent of
Q/DQ policy.
- `qdq_graph.py` owns graph-level Q/DQ analysis and policy, including
policy-specific rewrites such as `remove_partial_input_qdq`;
`qdq_utils.py` remains the lower-level Q/DQ node, tensor, and format
utility layer.

This four-way split is the intended long-term layout. All production and
test callers now import the owning modules directly, and the former
`graph_utils.py` catch-all module is removed. The extraction preserves
the existing graph-selection and rewrite algorithms; all 43 moved
functions were verified to have identical ASTs before and after the
split.

### Usage

```python
# N/A: behavior-preserving internal refactor.
```

### Testing

- `pytest -q tests/unit/onnx/quantization`
  - 376 passed, 13 intentionally xfailed
- Focused tests after rebasing onto latest `main`
  - 47 passed, 6 intentionally xfailed
- GPU ONNX quantization suite, excluding the existing AutoTune
integration test
  - 46 passed, 3 skipped
- Broader ONNX unit suite
  - 658 passed, 1 skipped, 13 intentionally xfailed
- One ONNX Runtime `CopyTensorAsync is not implemented` failure was
reproduced unchanged on the base revision.
- Review-feedback validation
  - Public `insert_matmul_casts` star-import assertion passed.
- `pytest -q tests/unit/onnx/quantization/test_graph_selection.py
tests/unit/onnx/test_partitioning.py`: 42 passed.
- Changed-file pre-commit hooks passed, including Ruff, formatting,
mypy, Bandit, RST checks, and license checks.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: No -- direct imports from
`modelopt.onnx.quantization.graph_utils` must migrate to the new
capability modules. No shim is retained because removing the catch-all
path is the purpose of this ownership split; the package-level
quantization entry points are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: Yes
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes -- the backward-breaking section contains a symbol-level migration
map.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Follows #2457, which added focused characterization coverage for the
calibrated ONNX quantization paths.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added expanded ONNX graph analysis for tensor relationships, pattern
detection, FP8 attention matching, and quantization candidate selection.
- Added safer Q/DQ transformations, custom-operator casting,
mixed-precision configuration, and redundant-cast cleanup.
- Improved quantization support for INT4, INT8, FP8, convolution,
attention, and weight-selection scenarios.

- **Refactor**
  - Reorganized ONNX quantization utilities into specialized components.
- Removed the legacy graph utility module; related functionality is now
available through dedicated modules.

- **Tests**
- Updated coverage for graph selection, partitioning, insertion points,
and legacy-module removal.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-09-18 10:01:14 -04:00
9e3d555aa1 [OMNIML-5899] Add IQ quantization codecs and backend (#2446)
## Summary

- add IQ1_S and IQ2_XS reference codecs and a weight-only fake-quant
backend
- register and export both formats from the quantization package
- cache compact packed weights across unchanged forwards and invalidate
on tensor or config changes
- use one Python-side IQ2_XS FP16 scale predictor for both reference and
CUDA packing
- validate packed payload metadata, normalize CUDA cache keys, and
define a shared non-finite policy

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns the Python codecs, backend dispatch, package registration,
license attribution, CPU codec/backend tests, and CUDA
numerical/reference-path tests. The native CUDA layer and direct
extension tests remain in #2448; export and recipes remain in their own
PRs.

## Why the codecs are separate from `qtensor`

The new `ggml/` package contains stateless reference codecs and
fake-quant backend functions. They transform ordinary tensors into
packed format payloads and reconstruct tensors for fake quantization;
they do not define persistent runtime quantized-tensor objects.

`BaseQuantizedTensor` subclasses under `qtensor/` own runtime tensor
objects and execution dispatch. Keeping the codecs separate avoids
claiming a runtime tensor contract that these formats do not yet
provide. A `qtensor` type can be added later if a runtime execution path
requires one.

## Compatibility boundary

The Python encoders intentionally use fixed-scale, unweighted searches.
They are not intended to reproduce another encoder's bytes for every
input when that encoder performs iterative scale refinement or
importance weighting. Compatibility is defined by the canonical
codebooks, 50/74-byte payload layouts, and pinned dequantization
formulas.

IQ2_XS computes the FP16 superblock scale once in the Python predictor
and passes it to the CUDA packer. This removes a duplicate
floating-point reduction and makes native/reference byte parity use the
same scale. Non-finite input elements are treated as zero during packing
in both implementations.

The unit tests construct nonzero payload fields independently and
validate metadata, signs, local scales, and global scales. The CUDA
tests compare native packed bytes with this Python reference encoder.

## Test coverage

- [IQ1_S CPU codec
tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_iq1_s.py)
- [IQ2_XS CPU codec
tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_iq2_xs.py)
- [registered backend and cache
tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/unit/torch/quantization/test_ggml_backend.py)
- [IQ1_S CUDA byte-parity, numerical, non-finite, zero-payload, and
fallback
tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py)
- [IQ2_XS CUDA byte-parity, numerical, non-finite, zero-payload,
underflow, and fallback
tests](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py)

## Licensing

The embedded codebook data cites the pinned upstream MIT source, carries
its license notice, and uses the repository's third-party license
mechanism. Human OSRB/code-owner confirmation is still required; this PR
does not claim that approval.

## Validation

- focused lint, format, and type checks pass for all changed Python
files
- 36 focused CPU codec and backend tests pass locally
- all 20 direct-extension and CUDA integration test cases collect
locally; runtime CUDA execution remains delegated to GPU CI
- restricted-term scan passes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
  - Added GGML quantization support for IQ1_S and IQ2_XS formats.
- Added quantization, dequantization, and fake-quantization workflows
with pass-through gradients.
  - Added CPU fallback when CUDA acceleration is unavailable.
- Added validation for packed weights, tensor shapes, formats, and
backend options.
- Added configurable chunk processing and caching for repeated
quantization.

- **Tests**
- Added comprehensive CPU and CUDA coverage for accuracy, validation,
caching, fallback behavior, and edge cases.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 06:27:58 +00:00
Ajinkya RasaneandAjinkya Rasane 9895d6f129 [6771663] Preserve ONNX API output types when wiring casts (#2451)
### What does this PR do?

Type of change: Bug fix

Preserves the public ONNX graph I/O types captured at the API boundary
when `PrecisionConverter` wires output casts. Type inference can change
the working graph's output declaration before conversion; consulting
that mutated declaration caused the required cast back to the original
public type to be discarded and metadata restoration to fail.

The converter now derives its I/O type map from the preserved boundary
metadata and uses that map when deciding whether a cast should become a
public graph output. A regression test covers an FP32 output whose
working declaration is inferred as FP16, and the changelog records the
corrected behavior.

### Usage

```python
# No API changes are required. Existing conversions now preserve the original
# public I/O declarations when keep_io_types=True.
converted = convert_to_f16(model, keep_io_types=True)
```

### Testing

- Ran `pytest tests/unit/onnx/autocast/test_precisionconverter.py` (186
passed).
- Ran `pytest tests/unit/onnx/autocast` (249 passed).
- Ran Ruff check and format validation on the changed Python files.
- Verified the original minimal end-to-end reproduction with Python 3.12
and TensorRT 10.16.1.11; conversion now completes without the
output-metadata restoration error.
- Verified the full CLI quantization path proceeds through the formerly
failing one-Q/DQ scheme and successfully benchmarks the generated
TensorRT engine.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [6771663]


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Bug Fixes**
- Fixed ONNX FP16 conversion to preserve public graph output types when
type inference changes internal declarations.
- Ensured output casts are inserted correctly when preserving
input/output types is enabled.

- **Tests**
- Added regression coverage confirming preserved output types, correct
cast insertion, and valid ONNX model generation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ajinkya Rasane <ajinkyaashwin@gmail.com>
Co-authored-by: Ajinkya Rasane <ajinkyaashwin@gmail.com>
2026-09-18 03:47:25 +00:00
Chenjie LuoandClaude Opus 5 b16356776b Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do?

Type of change: Code refactoring

`#2448` added the GGML IQ packing kernels as **two** torch extensions,
`modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges
them into one,
`modelopt_cuda_ext_ggml`.

The existing per-extension split in `extensions.py` exists for reasons
that don't apply
to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while
`_fp8`/`_mx` gate on
`>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the
base `tensor_quant`
kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none
of that —
same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split
only compiled the
shared header twice, ran nvcc twice, and grew the loader, `__getattr__`,
and
`precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly
following, that
scales badly.

Changes:

- New `ggml/ggml.cpp` holds both host-side validation wrappers and the
single
`PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously
each module
exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`;
the
  validation logic and docstrings carry over unchanged.
- `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`,
which builds
`ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The
retry-on-`raise_if_failed`
  semantics of the old getters are preserved.
- Each format keeps its kernels in its own translation unit, so adding a
format is a new
`.cu` plus one `module.def` — no new extension, loader, or
`precompile()` line.

No caller outside `extensions.py` and its tests referenced the old
getters on `main`, so
nothing else changes. **Note for the follow-up PRs in the `#2448` series
(`#2446`/`#2447`/`#2449`): the codec layer should call
`get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of
`get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.**

### Usage

```python
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml

ext = get_cuda_ext_ggml(raise_if_failed=True)
iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid)            # uint8 [numel / 256, 50]
iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74]
```

### Testing

Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000`
container), building the
merged extension from scratch:

- `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24
passed** (6:44). This
is the full existing IQ suite (zero-block layout, encode, dtype
rejection,
row-straddling rejection, invalid/negative-zero scales, byte-exact dtype
equivalence,
and the brute-force optimality round-trip) reparametrized onto the
merged module,
  plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests.
- Verified `precompile()` loads all four extensions and that the merged
module exports
  exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities.
- Off-GPU: compiled the three sources directly and linked them into one
`.so` to confirm
no duplicate-symbol collisions between the two `.cu` translation units.
- `pre-commit run --files ...` passes on all changed files (ruff, mypy,
clang-format,
  bandit, license headers).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the removed getters were
added in `#2448`
(merged today, unreleased) and have no callers outside this file's own
tests.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; the moved wrappers keep their original
attribution.
- Did you write any new necessary tests?: ✅ — existing coverage
reparametrized onto the merged module; no behavior change to test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal refactor of an unreleased, not-yet-wired-up API.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Follow-up to #2448. Merge before the remaining PRs in that series
(#2446, #2447, #2449)
land, so the codec layer is written against `get_cuda_ext_ggml` and no
rename is needed
afterwards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
  - Added IQ1_S packing support through the GGML CUDA extension.
  - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing.
- Improved extension loading reliability when a cached extension is
unavailable.

- **Changes**
  - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`.
- Consolidated IQ1_S and IQ2_XS extension access under the shared GGML
loader.
- Updated GPU validation and coverage to use the unified extension
interface.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 20:46:41 +00:00
yeyu-nvidiaandClaude Opus 5 ad1bad7817 fix(specdec): gather sharded hidden states in DFlash/DSpark AR generation (#2458)
### What does this PR do?

Type of change: Bug fix

Fixes the nightly
`tests/regression/torch/speculative/test_dflash.py::test_dflash_ar_validate`,
red every night since 2026-09-10 with:

```
WARNING: sample 0 (writing) failed: Expected all tensors to be on the same device,
but got tensors is on cuda:1, different from other tensors on cuda:0
(when checking argument in method wrapper_CUDA_cat)
RuntimeError: AR validation produced no results: all 3 samples failed.
```

**No PR caused this.** `ar_validate.py` loads with `device_map="auto"`,
and the regression runner (`…rtxpro6000-l-2-…`) has 2 GPUs, so even
Qwen3-0.6B gets sharded. `pseudo_speculative_generate()` then
concatenates hidden states from `target_layer_ids`, which spans early
*and* late layers — different devices once the base model is split:

```python
selected = [base_outputs.hidden_states[lid + hid_offset] for lid in self.target_layer_ids]
target_hidden = torch.cat(selected, dim=-1)   # cuda:0 ++ cuda:1 -> RuntimeError
```

The block tensors have the same problem from the other side: they follow
`input_ids.device`, while the draft module sits on the *last* base
layer's device (`_place_draft`).

What changed on 2026-09-10 was detection, not behaviour. #2288 made
`ar_validate.py` raise when every sample fails; before it, the report
block was guarded by `if results and …`, so zero results printed nothing
and exited **0**. The test had been green while validating nothing. Its
runtime is the tell: 23–26s in every green nightly, 21.6s in the first
red one, where it does no validation work at all.

The same defect is in #2288's own description, on an 8-GPU sharded
**EAGLE3** checkpoint: 80/80 samples dead on the identical error, job
exiting 0. So this is not DFlash-specific in principle — but EAGLE3's
generate path gathers separately (`pop_and_gather_aux_hiddens`) and
needs its own look, which is why this PR stops at the DFlash family.

### Fix

Gather everything the draft consumes onto the draft's device, and hand
results back on the caller's device since `validate_online` cats them
onto the running sequence. Every `.to()` is a no-op on a single device,
so single-GPU behaviour is bit-identical.

`HFDominoModel` inherits `HFDFlashModel.pseudo_speculative_generate`, so
it is fixed by the same change; `HFDSparkModel` overrides it and is
patched in parallel (including `prev_token` feeding `markov_step`).

### Usage

No API change.

```bash
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <dflash ckpt> --osl 10 --num_samples 3 --steps 7
```

### Testing

- `tests/unit/torch/speculative/plugins/test_hf_dflash.py` and
`test_hf_dspark.py`: **110 passed, 1 skipped**.
- The skip is the new `TestShardedBaseGeneration`, which is gated on
`torch.cuda.device_count() >= 2`. It stands in for accelerate's dispatch
— the base forward stays whole, but its hidden states are spread across
two devices the way a real `device_map="auto"` split spreads them — and
asserts both returned tensors come back on the caller's device. **I have
no 2-GPU box to hand, so this test is unverified locally; it will first
execute in CI.** It reproduces the reported error against the unpatched
code by construction, but please treat that claim as unconfirmed until
the run goes green.
- The real verification is the next nightly: `test_dflash_ar_validate`
should pass *and* print an AR number, which it has not done in any log I
can find.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — device moves only; no-ops
when the base model is on one device.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one multi-GPU regression
test (skipped without 2 GPUs).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — bug in an unreleased-cycle path that only ever produced a silent
no-op; no user-visible behaviour to describe.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Worth a maintainer view: `device_map="auto"` is a poor default for AR
validation, which is a single-process step-by-step loop. Every sharded
run of it that I can find has failed. Pinning it to one device would be
a smaller blast radius than making every drafting path device-safe — but
it would also cap validation at models that fit on one GPU, so I have
not changed it here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved speculative generation with sharded Hugging Face models
across multiple GPUs.
* Ensured generated draft tokens and outputs are returned on the
caller’s input device.
* Preserved Markov sampling behavior while improving device
coordination.

* **Tests**
* Added multi-GPU coverage for sharded model generation, including
output placement and draft shape validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:28:57 -07:00
2ff2e1bc80 [OMNIML-5899] Add CUDA kernels for IQ packing (#2448)
## Summary

- add native CUDA packing kernels for IQ1_S and IQ2_XS
- load both extensions through the quantization extension module
- validate caller metadata and launch bounds before contiguous
materialization
- normalize non-finite input elements consistently with the Python
reference path
- accept caller-computed IQ2_XS FP16 superblock scales to avoid a
duplicate reduction
- share common packing helpers and add direct extension compilation and
boundary tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only native kernel sources, shared packing helpers,
extension loading and build registration, and direct extension tests. It
does not contain Python codecs, export code, or recipes.

## GPU test coverage

Direct kernel-boundary coverage is included in this PR:

- [extension compilation, zero payloads, input validation, and
row-alignment
checks](https://github.com/NVIDIA/Model-Optimizer/blob/74e94db9601870e3569c7c8e73506a2f08c29da8/tests/gpu/_extensions/test_torch_extensions.py)

Pack/dequantize numerical, native/reference byte-parity, and
non-finite-policy tests are owned by the quantization PR:
[IQ1_S](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq1_s_cuda.py)
and
[IQ2_XS](https://github.com/NVIDIA/Model-Optimizer/blob/e8d937081d8cd01cf8e44d43915df443b79deb17/tests/gpu/torch/quantization/test_iq2_xs_cuda.py).

## Dependency behavior

On `main`, this PR provides optional CUDA extension loaders and direct
extension tests. The Python encoders and fallback dispatch land in
#2446. Until #2446 lands, no quantization path calls these getters, so a
load failure reports only that the extension is unavailable.

The IQ2_XS packer accepts one caller-computed FP16 scale per 256-value
block. #2446 owns that predictor and passes the same values to the
native and reference encoders.

## Provenance

- The CUDA kernels were independently written.
- They implement the packed-format contract and sign-parity convention
from the pinned [llama.cpp
definition](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h).
- `16.875 = 15 × (1 + 1/8)` is a derived IQ1_S constant.
- `0.125` is part of the encoded format.
- `0.61` is our empirical IQ1_S scale predictor, not copied from
upstream code.
- The IQ2_XS predictor constants are owned by #2446 and are not
duplicated in this kernel.

Human review is still required to confirm that the attribution and
license treatment are sufficient.

## Validation

- repository hooks, including native formatting, pass for all changed
files
- extension loader and test modules compile as Python
- all 20 direct-extension and CUDA integration test cases collect
locally
- CUDA runtime execution is delegated to GPU CI because the local host
is macOS
- restricted-term scan passes

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 17:56:19 +00:00
Wei-Ming Chen f377b77116 Fail fast on non-finite AutoQuantize output gradients (#2432)
### What does this PR do?

Type of change: Bug fix

Fail fast when AutoQuantize receives non-finite output gradients, before
accumulating sensitivity scores. The error names the affected module and
suggests checking the model, data, and loss; it also identifies cuDNN
SDPA backward on fully masked rows as one possible cause and gives an
explicit retry workaround.

Unlike the earlier revision, this does not disable cuDNN or change any
attention backend settings. Invalid gradients are not zeroed or ignored,
and there is no automatic retry.

### Usage

No API or recipe changes. For the reproduced cuDNN failure, the caller
can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before
a fresh AutoQuantize run.

### Testing

- AutoQuantize unit suite: **110 passed**. Coverage includes NaN and
positive/negative infinity, module diagnostics, preventing invalid score
accumulation, model-state cleanup, and unchanged SDPA backend settings.
- Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main
`8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size
8, 512 calibration samples, and
`w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`:
- Default backend: the new diagnostic fired at
`model.language_model.layers.39.self_attn.q_proj`; cuDNN remained
enabled and no quantized model was exported. Expected-error check
passed.
- Explicit cuDNN-disabled fresh run: both 64-batch calibration passes,
all 16 scoring batches, optimization at **5.99 effective bits**, and
checkpoint export completed with exit code 0. Verified all three indexed
safetensors shards and quantization configuration; the index contains
93,563 tensors.
- All applicable pre-commit checks and `git diff --check` passed.

### Before your PR is "*Ready for review*"

Contributor guidelines and security guidance followed; commits are
signed and signed off.

- Is this change backward compatible?: Yes; finite-gradient behavior and
backend settings are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither
added.
- Did you write any new necessary tests?: Yes.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes, 0.48.0 bug fixes.
- Did you get Claude approval on this PR?: No; awaiting review.

### Additional Information

This improves error reporting rather than fixing the upstream cuDNN
kernel. Blackwell-specificity is not established. Checkpoint
deployment/reload was not tested.

A calibration-only checkpoint-resume attempt completed scoring but
encountered a separate `candidate_stats` KeyError. That issue is outside
this patch; the successful export validation above used a fresh run
without search-state resume.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* AutoQuantize now fails fast when output gradients contain non-finite
values, with an actionable error identifying the affected module.
* Attention backend settings are preserved and restored after successful
runs and failures.
* Model state is restored when sensitivity scoring encounters an error
during setup or execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-17 17:52:46 +00:00
realAsma b9cfdce8dc docs: add Local Hessian NVFP4 weight-scale announcement blog (#2417)
### What does this PR do?

Type of change: documentation

Adds a Local Hessian announcement blog at
`docs/source/announcements/local-hessian.rst`, covering the NVFP4
per-block
weight-scale rule that minimizes output error instead of weight error.

Contents:

- Derivation of the per-block output-error objective and its `16x16`
local
  Hessian, with numbered equations.
- Results on Qwen3.5-9B: scale-setting comparison against max, MSE, and
  Four-over-six, plus composition with GPTQ.
- Figure 1, a grouped bar chart of the Qwen3.8-27B W4A4 candidate scores
  (BF16 in gray, the two scale rules in NVIDIA greens).
- A "Using Local Hessian" section with the config example and the
  end-to-end `hf_ptq.py` command.

Two supporting changes outside the blog:

- `docs/source/_static/announcements.css`: the `shibuya` theme has no
`span.eqno` rule, so Sphinx's default `float: right` on equation numbers
cannot share a line with MathJax's full-width display block and the
number
renders *above* the equation. This anchors it to the right of the
equation
  instead, and shrinks the table-note class.
-
`docs/source/announcements/assets/qwen3-27b-w4a4-scale-rule-accuracy.png`:
  the Figure 1 asset.

### Usage

```python
import modelopt.torch.quantization as mtq

config = {
    "quant_cfg": [...],  # quantizer configuration
    "algorithm": {"method": "local_hessian", "fp8_scale_sweep": True},
}

model = mtq.quantize(model, config, forward_loop)
```

### Testing

Documentation only; no code paths change. The `.rst` parses cleanly
under
docutils. The rendered page has not been checked with a full
`sphinx-build`,
so the equation-number CSS fix and the figure placement are worth an
eyeball
on the built docs before merge.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

Two items to settle before this is ready to publish:

1. The `--recipe` example points at

`modelopt_recipes/models/Qwen/Qwen3.8-27B/ptq/nvfp4_local_hessian-fp8_attn-kv_fp8_cast.yaml`,
a placeholder path derived from the existing recipe naming convention.
It
   needs to match whatever lands in #2363.
2. The tables report single-run team measurements; the blog says so and
makes
   no significance claims.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added guidance on NVFP4 Local-Hessian weight-scale selection,
including mathematical details, accuracy comparisons, runtime
considerations, limitations, configuration examples, and reproduction
steps.
- Updated announcement labels, headings, metadata, descriptions, and
filtering text to use “Local-Hessian.”

- **Style**
- Improved announcement formatting for equation labels, display-equation
spacing, Hessian results, table headers, and explanatory notes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-16 21:48:11 +00:00
Keval MorabiaandClaude Opus 5 a448ba9757 Add end-to-end W4A4 NVFP4 + QAD tutorial for Qwen3.6-35B-A3B (#2411)
### What does this PR do?

Type of change: new example + bug fix

<img width="2085" height="1239" alt="image"
src="https://github.com/user-attachments/assets/b9ced215-ce8c-4dbe-be74-a75c1c4714b3"
/>


Adds an end-to-end **W4A4 NVFP4 + Quantization-Aware Distillation**
tutorial for
[Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) at
`examples/megatron_bridge/tutorials/Qwen3.6-35B-A3B/`.

It complements the existing Nemotron-3-Nano tutorial (pruning +
distillation + FP8). Here the model
is unpruned and the technique under test is **W4A4** — aggressive enough
that PTQ alone leaves a
measurable accuracy gap, which is what QAD exists to close.

**Why W4A4 rather than weight-only NVFP4:** W4A16 measured *slower than
BF16* in 10 of 12 shapes,
because a BF16 activation forces vLLM onto the Marlin dequant fallback
and never reaches the
Blackwell FP4 tensor cores. W4A4 beats BF16 in 9 of 12 shapes (up to
1.30x) and shrinks the
checkpoint 67 GiB -> 22 GiB (3.1x).

**What the study found:** only 2 of 6 benchmarks show a statistically
significant PTQ deficit, so
those are the only two QAD can recover. IFBench is recovered to parity
with BF16 (-2.62 pp ->
-0.29 pp, gain of +2.33 pp, p=0.036); MMMU-Pro recovers ~40% and retains
a significant gap. The
other four are lossless under W4A4 to begin with.

Also added:

-
`modelopt_recipes/model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore.yaml`
— the PTQ
  recipe used as the QAD student, usable via `--recipe`.
- `data_blend.yaml` — the token-budgeted blend config for the
distillation data.
- `eval_configs/*.yaml` — one NeMo Evaluator config per benchmark.
tau2-bench is separate because
it needs `--enable-auto-tool-choice --tool-call-parser qwen3_coder` and
`deployment.command` is
  global to a config.

**Two export fixes found while producing these checkpoints** (both
change library/example behaviour,
both have changelog entries under 0.48.0 Bug Fixes):

- `unified_export_megatron.py` — MCore builds `embedding` on the MTP
stage as well as the first, so
gating export on `hasattr(model, "embedding")` wrote a **second,
unreferenced copy of the vocab
embedding** whenever an MTP model was exported with PP > 1. The index
mapped the key to the later
shard, so the extra copy never loaded but still shipped — ~1 GB for this
model. Now gated on
`model.pre_process`, MCore's own "this rank owns the input embedding"
flag.
- `export_quantized_megatron_to_hf.py` — stopped passing Megatron's
`moe_router_dtype` as the
router's *storage* dtype. It is a routing *compute* dtype; the parameter
is bf16 in a bf16 model,
so the export was widening bf16 to fp32. All 21,495,808 router values in
the exported checkpoint
have their low 16 bits zero, and vLLM builds the gate at the model dtype
and rounds on load, so
the dropped bytes carried no information. `export_mcore_gpt_to_hf` still
accepts the override.

### Usage

```bash
# 1. PTQ (2 GB200 nodes for EP=8)
srun ... python examples/megatron_bridge/quantize.py \
    --hf_model_name_or_path Qwen/Qwen3.6-35B-A3B \
    --recipe model_type/qwen3_6_moe/ptq/w4a4_nvfp4-fp8_attn-kv_fp8_cast_mcore \
    --tp_size 1 --ep_size 8 --pp_size 1 \
    --calib_dataset_name cnn_nemotron_v2_mix --calib_num_samples 1024 --calib_batch_size 1 \
    --seq_length 8192 --skip_generate \
    --export_megatron_path /path/to/qwen36_w4a4_megatron

# 2. QAD (32 nodes x 4 GB200)
python -u examples/megatron_bridge/distill.py \
    --teacher_hf_path Qwen/Qwen3.6-35B-A3B --student_hf_path Qwen/Qwen3.6-35B-A3B \
    --student_megatron_path /path/to/qwen36_w4a4_megatron \
    --tp_size 1 --pp_size 1 --cp_size 1 --ep_size 8 \
    --seq_length 32768 --mbs 1 --gbs 512 --train_iters 500 \
    --lr 1e-5 --min_lr 1e-6 --lr_warmup_iters 50 --logit_kl_topk 4096 \
    --recompute_granularity full --recompute_method uniform --recompute_num_layers 1 \
    --no_async_save --eval_iters 0 --save_interval 50 \
    --data_paths "${DATA_BLEND}" --output_dir /path/to/qad_output
```

### Testing

**Library changes.**
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py` gains
`test_unified_export_megatron_pp2_mtp_no_duplicate_tensors`: it exports
a PP=2 model built with
`mtp_num_layers=1` and asserts no tensor lands in more than one shard.
Verified to **fail without
the fix**:

```
AssertionError: tensors written to more than one shard:
  {'model.embed_tokens.weight': ('model-00001-of-00002.safetensors',
                                 'model-00002-of-00002.safetensors')}
```

The pre-existing `..._pp2_mtp_metadata_matches_shards` test cannot catch
this — it fakes
`_get_mtp_state_dict` on a model with no real MTP, so the last stage
never builds an embedding.

Ran the whole `tests/gpu_megatron/torch/export/` suite with and without
the fix: identical failure
sets (3 failures both ways, all `qwen3_5_moe_vl_*` from a local
`ImportError: FLA is not installed`),
58 passed with vs 56 without — the +2 being the new test's two workers.
`tests/unit/recipe` passes
368/368 after the recipe path move. `model.pre_process` is always
present: `GPTModelExporter.__init__`
raises unless the model is `GPTModel` or `HybridModel`, and both set it
unconditionally.

Both export fixes were also applied to the real 23 GB checkpoints and
re-validated end to end: every
retained tensor md5-identical, index/shard integrity re-checked, and a
**full GPQA re-evaluation of
the fixed checkpoint** scored 83.49 vs 84.25 before (paired per-question
t-test over the same 198
questions x 16 repeats: -0.76 pp, p=0.21, not significant).

**Numbers in the tutorial** come from real runs, not estimates:

- **253 evaluation runs** across BF16, the published W4A16 checkpoint,
W4A4 PTQ, and QAD at
50 / 300 / 500 iterations — 8 repeats per benchmark (3 for tau2-bench;
GPQA is one
  `num_repeats: 16` run).
- The published `nvidia/Qwen3.6-35B-A3B-NVFP4` checkpoint was
re-evaluated under this same harness
(36 runs) rather than quoted from its card, so the W4A16 row is
same-harness.
- Every figure and results-table value is generated from the collected
`results.yml` files by a
script, and I verified the README table cell-by-cell against that data
after each edit.
- Throughput rows were cross-checked against the recorded AIPerf sweeps;
the QAD wall-clock figures
  against the two jobs' Slurm records (`03:34:49` + `02:09:49`).
- All CLI flags in the tutorial were verified to exist in `quantize.py`
/ `distill.py` /
`export_quantized_megatron_to_hf.py`, and `cnn_nemotron_v2_mix` against
`dataset_utils.py`.

The tutorial also records the non-obvious constraints found the hard
way: QAD on this model requires
`TP=PP=CP=1` (TP breaks quantizer `_amax` dist-checkpoint sharding, PP
starves Qwen3-VL's M-RoPE of
`position_ids`, CP hits a rope shard mismatch), EP must match the PTQ
checkpoint, and
`--logit_kl_topk` is mandatory at 32K because the dense `[seq, vocab]`
fp32 logits are 30.31 GiB per
tensor on a 248,320-token vocabulary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ <!-- gpu_megatron PP=2+MTP
export dedup test -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!-- 0.48.0: Megatron Framework + two Bug Fixes -->
- Did you get Claude approval on this PR?: ✅ <!-- not yet run -->

### Additional Information

Changelog entries are filed under **0.48.0**; the `cherry-pick-0.47.0`
label has been removed.
Rebased onto `main` after #2328 renamed `modelopt_recipes/huggingface`
to `model_type` (it is now a
compatibility symlink), so the recipe moved to
`model_type/qwen3_6_moe/ptq/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end Qwen3.6-35B-A3B tutorial for W4A4 NVFP4
quantization and quantization-aware distillation.
* Added checkpoint export, accuracy evaluation, and vLLM throughput
benchmarking workflows.
* Added evaluation configurations for AA-LCR, GPQA, IFBench, MMMU-Pro,
SciCode, and tau2 Telecom.
  * Added a token-budgeted supervised fine-tuning data configuration.
  * Added a Megatron-Core NVFP4/FP8 quantization recipe for Qwen3.6-MoE.

* **Documentation**
* Added benchmark results, deployment guidance, hardware requirements,
reproduction steps, HTTPS endpoint guidance, announcement filters, and
tutorial links.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 02:10:12 +05:30
realAsma 6a4b3f147e [Fix] Calibrate non-decoder modules during layerwise quantization (#2339)
### What does this PR do?

Type of change: Bug fix

Layerwise calibration now calibrates enabled quantizers outside
transformer layers, such as `lm_head`, while hiding decoder layers from
the additional calibration traversal. It also fails early when these
quantizers are combined with progressive layerwise export, whose
in-place conversion makes the required model calibration unsafe.

### Usage

No API changes.

### Testing

- `pytest_pwd tests/unit/torch/quantization/test_calib.py
tests/unit/torch/quantization/test_layerwise_calibrate.py -q` — 68
passed
- Pre-commit on all six changed files — passed
- `git diff --check` — passed

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Layerwise calibration now supports quantizers outside transformer
decoder layers, including LM-head activation quantizers.
- Calibration supports models with CPU-offloaded components while
preserving model behavior.
  - The MSE calibration utility is now publicly available.
- Added warnings for calibration passes involving offloaded components.

- **Bug Fixes**
- Prevented unsupported export-mode calibration when outside-layer
quantizers are present.
- Ensured outside-layer calibration runs after transformer-layer
restoration.
- Preserved original forwards while hiding decoder subtrees from
traversal and state collection.
  - Ensured cleanup and alias restoration after errors or completion.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-09-15 22:10:50 +00:00
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00
Shengliang Xu 3c87751903 Deprecate the single-format quantization CLI flags in favour of --recipe (#2426)
### What does this PR do?

Type of change: deprecation

Deprecates the single-format quantization CLI flags in favour of
`--recipe`. Passing one now emits a `DeprecationWarning`; nothing else
changes.

| script | flags |
|---|---|
| `examples/hf_ptq` | `--qformat`, `--kv_cache_qformat` |
| `examples/megatron_bridge/quantize.py` | `--quant_cfg`,
`--kv_cache_quant`, `--weight_only` |
| `examples/torch_onnx` | `--qformat` |

`--recipe` was already authoritative over all six — silently on
`hf_ptq`, and with a runtime warning on `megatron_bridge` — and
`modelopt/recipe/presets.py` already records the intent in a comment:
*"the long-term direction is to retire `--qformat` /
`--kv_cache_qformat` in favour of `--recipe`"*. This makes that a real
deprecation.

A recipe carries the quantization config, the calibration algorithm and
the KV-cache setting in one file, so they cannot drift apart the way
separate flags can. That drift is not hypothetical: the preset path
applies no MTP exclusion while the recipe unit
`default_disabled_quantizers` disables `mtp.*`, so the same model
quantizes differently depending on which entry point was used.

#### The warning fires only when a flag is actually passed

`RecipeSupersededAction` is an `argparse.Action`, and argparse invokes
an action only for options present on the command line — never for a
default. That matters because several of these default to a *quantizing*
value (`--qformat fp8`, `--kv_cache_qformat fp8_cast`); warning on the
defaults would fire on every run, including runs that correctly use
`--recipe` and never mention the flag.

`examples/speculative_decoding/scripts/quantize_drafter.py` keeps
`--qformat` undeprecated: it has no `--recipe`, so there would be
nothing to migrate to.

### Usage

```bash
# deprecated
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> --qformat nvfp4 --kv_cache_qformat fp8_cast

# replacement
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <ckpt> \
    --recipe general/ptq/nvfp4_experts_only-kv_fp8_cast
```

### Testing

Three tests in `tests/examples/hf_ptq/test_hf_ptq_args.py`, all passing:

- passing `--qformat` / `--kv_cache_qformat` raises `DeprecationWarning`
and still parses the value;
- omitting them raises nothing and leaves the defaults (`fp8`,
`fp8_cast`) untouched;
- the action stays wired to both flags, so a future edit cannot drop it
while leaving the help text.

Defaults and parsed values were diffed against `main` and are unchanged
— the action stores exactly what `store` / `store_true` would have.
`ruff` findings are at parity with `main` on every changed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the flags still work, they
only warn.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Deprecations.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the removal release for these flags is not decided here, only the
deprecation.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Deprecations**
* Legacy quantization CLI options now issue visible `FutureWarning`
messages only when explicitly provided.
* Use `--recipe` instead of deprecated options in Hugging Face PTQ,
Megatron-Bridge, and torch-to-ONNX workflows.
* Existing option values, defaults, and parsing behavior remain
unchanged.
* Weight AutoQuantize recipes without an explicit `kv_cache` setting
continue to use `--kv_cache_qformat` as a fallback.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-14 13:32:28 -07:00
JoshuaandCursor 1030791f53 Fix distributed AutoQuantize scoring and share backward setup (#2231)
### What does this PR do?

Type of change: Bug fix

AutoQuantize can measure a group of quantized expert layers at their
enclosing MLP output. That enclosing module is often a plain PyTorch
container and does not carry distributed-group information, so its
sensitivity score was not combined across data- or expert-parallel
workers.

This PR obtains the distributed groups from the quantized layers when
the scoring module does not provide them. It also preserves construction
order for quantized modules, scoring modules, and their registered
hyperparameters so every worker accumulates scores in the same order.

The temporary state needed by backward-based scoring is now managed by
one shared session. The session installs and removes forward patches and
invocation-specific output-gradient hooks, controls parameter gradients,
and restores the active quantization recipes even when scoring raises an
exception. Scoring methods remain responsible for their own score
calculation.

### Usage

N/A — this fixes existing AutoQuantize behavior and does not add an API
or flag.

### Testing

- `pre-commit run --files modelopt/torch/quantization/algorithms.py
tests/unit/torch/quantization/test_autoquant.py`
- `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102
passed
- Added a real two-rank gradient AutoQuantize test covering MoE experts
scored at an enclosing MLP.
- Added regressions for deterministic hyperparameter registration,
per-invocation replay for reused score modules, exact
`forward`-attribute restoration, partial setup rollback, and cleanup
after a scoring failure.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅ — no API or checkpoint format
changes; distributed sensitivity values now include the missing
reduction.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no new feature, deprecation, breaking change, or critical
release-note item.
- Did you get Claude approval on this PR?: ❌ — pending review.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization scoring consistency through deterministic
ordering and invocation handling.
* Added more reliable distributed score aggregation, including support
for mixture-of-experts models.
* Improved gradient-based scoring for repeated evaluations, tuple
outputs, and checkpoint-compatible workflows.
* Ensured model behavior and scoring state are restored after successful
or failed evaluations.
  * Avoided unnecessary output replay when gradients are not required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Joshua Hill <joshua.hill@baseten.co>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-13 14:55:32 +00:00
Ajinkya RasaneandCodex 51de53e48c [6463897] Fix narrow FP16 histogram calibration (#2412)
### What does this PR do?

Type of change: Bug fix

FP16 entropy calibration can fail for sufficiently narrow activation
ranges because NumPy may construct the histogram bin edges at FP16
precision. NumPy 2.2 and later reject the resulting collapsed bin
spacing, while earlier versions can silently return invalid,
non-monotonic edges. The same precision issue can recur when ONNX
Runtime merges later calibration batches.

Losslessly widen FP16 activation values to FP32 while calculating and
merging histograms in both ONNX entropy calibration paths, then restore
the source dtype at the calibration-to-quantization boundary. This keeps
the histogram bins stable without changing FP16 Q/DQ scale or graph
dtype semantics. FP32 inputs and public APIs are unchanged.

For full-range FP16 activations, ONNX Runtime can overflow while
subtracting FP16 calibration endpoints before it widens the result.
Retry only a non-finite FP16 scale calculation with FP32 endpoints, then
cast the finite scale back to FP16. Existing finite FP16 calculations
and all non-FP16 calculations continue to use ONNX Runtime's original
result.

The fallback intentionally patches only the `qdq_quantizer` binding used
for calibrated activation ranges. Initializer and weight quantization
continue to use ONNX Runtime's existing `quant_utils` path unchanged;
full-range FP16 weight scaling is outside this calibration fix.

The regression tests exercise both collectors across initial collection,
an equal-range merge, and an expanding-range merge. They also verify the
internal FP32 histogram and external FP16 calibration-range contract,
including finite saturation when restoring sanitized values. A real
entropy calibration test covers full-range FP16 values and verifies
finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast
integration verifies the same FP16 scale-type contract through the
public quantization path.

### Usage

N/A — no API or usage change.

### Testing

All tests ran with CUDA hidden.

- NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- Public INT8 entropy quantization with full-range FP16 calibration
data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX
Runtime session loaded.
- Changed-file pre-commit hooks: passed.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up to #1558.

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 23:49:01 -04:00
Ajinkya RasaneandCodex 5b1f7e86cc [6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
### What does this PR do?

Type of change: documentation

Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:

- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.

This also removes an inaccurate source comment without changing runtime
behavior.

### Usage

```bash
python -m modelopt.onnx.quantization \
    --onnx_path=model.onnx \
    --quantize_mode=int8 \
    --calibration_data_path=calib.npy \
    --output_path=model.quant.onnx
```

### Testing

- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [6701308]

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.

- **Tests**
- Updated quantization command examples to use the current calibration
data path option.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 22:04:19 +00:00
ZhiyuandClaude Opus 5 c37a6948db feat(quantization): fail fast when a quant config matches no weight quantizer (#2203)
### What does this PR do?

Type of change: New feature (fail-fast guard; behavior change on a
previously silent path)

A `quant_cfg` whose module patterns don't match the model is not an
error to `set_quantizer_by_cfg` — every pattern simply matches nothing.
The run then calibrates, exports, and hands back a checkpoint that is
silently unquantized:

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Nothing in the run says so. It has only ever been caught by someone
reading the exported `hf_quant_config.json` afterwards — most recently
on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665),
after a full 8×B200 calibration), and before that on MiniMax-M3, where
fused-expert detection skipped the experts and an experts-only recipe
matched nothing.

`mtq.quantize` now compares the config's intent against the outcome and
raises **before calibration**:

```
RuntimeError: The quantization config asks for weight quantization but no weight quantizer was
enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no
weight quantizer:
  *.experts.*weight_quantizer
Either the patterns do not match this architecture's module names (check the model-specific
recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights
were never converted to quantized modules (an unsupported custom module, e.g. a
trust_remote_code MoE layout).
```

Scoped to avoid false positives:

- **Only configs that ask for weight quantization** are checked (an
entry with `enable` and `weight_quantizer` in its pattern), so
activation-only and KV-cache-only configs are unaffected.
- **Intent is read from each pattern's final entry**, since `quant_cfg`
entries apply in order: a pattern that is enabled and then disabled
later asks for nothing by the end.
- **Configs refining an already-quantized model** (weight quantizers
enabled by an earlier `mtq.quantize`) are left alone.

Matching goes through `conversion._match_quantizer` — the same matcher
`set_quantizer_by_cfg` used to apply the config — so "did this pattern
match anything?" is answered exactly as the applying code would. A local
`fnmatch` diverges on the two cases that matcher handles:
`SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and
fused-experts names (`..._weight_quantizers.0` normalizing to
`..._weight_quantizer`).

### Usage

No API change. A config that would previously have produced an
unquantized checkpoint now raises:

```python
mtq.quantize(model, {"quant_cfg": [
    {"quantizer_name": "*", "enable": False},
    {"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
]}, forward_loop)   # RuntimeError if the model has no `experts` modules
```

### Testing

Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one
per branch of the guard: patterns matching nothing raise; an
activation-only config still runs; weight patterns disabled by a later
entry still run; enabled-then-retracted patterns still run;
`SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer
names count as matched; and the already-quantized refinement path is
exercised. Each was checked to be non-vacuous by removing the
corresponding branch and confirming exactly that test fails.

**One existing test changed.**
`tests/gpu/torch/export/test_fsdp2_export.py` parametrized over
`NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case
ran the FSDP2 paths against an *unquantized* model, and the new guard
reported it (4 GPU failures on the first CI run, all `quant_config6`;
`NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The
parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps
the scoped-recipe coverage. **If reviewers would rather not change that
test's meaning, the alternative is to downgrade the guard to a warning —
flagging it explicitly as a decision.**

I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the
four MLP/experts-scoped configs raise, and the other three
(`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`,
`NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE
models (Qwen3-MoE, gpt-oss), so no other test is affected.

Ran locally after rebasing onto current `main` (torch 2.11, transformers
5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271
passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which
needs `hydra`): 2676 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`. GPU tests were not run locally (no suitable GPU); the
FSDP2 change above is reasoned from the CI failure, not re-run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — deliberately. A config that
previously produced a `quant_algo: null` checkpoint now raises. Any such
run was already not doing what it claimed; the three scoping rules above
keep intentional non-weight quantization working.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (Backward Breaking Changes)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes
the specific model that motivated this. Independent branches; either can
merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantization now detects enabled weight-quantization patterns that do
not apply to any model weights and reports a clear validation error
before calibration.
- Broad wildcard patterns and nested quantizers are now handled
correctly.
- Overlapping patterns respect the final matching setting, including
later disabling rules.
- Existing quantized models can be refined using the parsed
configuration.
- Activation-only and explicitly disabled weight-quantization
configurations remain supported.
- Pipeline-parallel stages without targeted weights can bypass this
validation when configured to do so.

- **Documentation**
- Documented the process-wide override for bypassing unmatched
weight-quantizer validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 13:26:35 -07:00
Jenny Chen 5f8c76e2aa Fix TEGroupedMLP quantizer checkpoint resharding (#2319)
### What does this PR do?

Type of change: Bug fix for
https://github.com/NVIDIA/Model-Optimizer/issues/2209

Fix TEGroupedMLP per-expert weight quantizer checkpoint resharding.

`TEGroupedMLP` now saves its per-expert quantizer state as singleton
local shards, allowing the distributed checkpoint format to retain each
expert's global identity. Restore also initializes scalar `_amax`
placeholders after ModelOpt extra-state restoration so distributed
checkpoint loading can populate quantizer state for experts that move
between ranks.

This fixes restoring quantized TEGroupedMLP checkpoints across
expert-parallel and tensor-parallel topology changes.


Previously there was a bug that had two parts
1. TEGroupedMLP did not mark its per-expert quantizer state as
singleton_local_shards. That meant the scalar
weight_quantizer.<expert>._amax state was not saved with the same
globally unique expert identity as the grouped-expert weights, so DCP
could not reliably redistribute it across EP layouts.

2. During restore, ModelOpt’s extra-state restoration can leave _amax
absent for experts that were not local on the checkpoint’s saving rank.
The subsequent distributed checkpoint load then had no destination
tensor to populate.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing

- `ruff format`, `ruff check`, `mypy`, `bandit`, and repository
pre-commit hooks
- Focused GPU regression:

  ```bash
python3 -m pytest
tests/gpu_megatron/torch/quantization/plugins/test_megatron.py \
    -k te_grouped_sharded_state_dict_reshard -v

Replaced the prior metadata-only TEGroupedMLP sharded-state test with an
end-to-end distributed-checkpoint save/restore regression test.

The new test:
- Quantizes a TEGroupedMLP with per-expert NVFP4 weight quantizers.
- Assigns each local expert a distinct, deterministic `_amax` based on
its global expert index.
- Saves both the model distributed checkpoint and sharded ModelOpt
state.
- Rebuilds the model under a different TP/EP topology.
- Restores ModelOpt state, loads the distributed checkpoint, and
verifies each target-local expert received the expected global-expert
`_amax`.

The parameterized test covers:
- EP=2 -> EP=1
- EP=1 -> EP=2
- TP=1 -> TP=2
- TP=2 -> TP=1
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved checkpoint restoration for grouped quantizers by initializing
missing quantization statistics with compatible shapes.
- Improved restoration across supported grouped quantizer
configurations, including sequential groups and parallel checkpoint
layouts.
- Extra module state is now finalized through supported post-load
callbacks when available.
- Preserved populated quantized output-layer state during checkpoint
operations while removing empty placeholders.

- **Tests**
- Expanded checkpoint resharding coverage across tensor- and
expert-parallel configurations.
- Added coverage for disabled, dynamic, and other grouped quantizer
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
2026-09-11 19:28:43 +00:00
h-guo18andShengliang Xu c5d1065331 Modeling Lib: per-model architecture kick off with spec registration (#1828)
### What does this PR do?

**Type of change:** New feature — per-model infrastructure (with
behavior changes, see below)

Starts `modelopt/torch/models/`: somewhere for ModelOpt to keep what it
knows about a
model, keyed by HF model type (`config.model_type`), one package per
type
(`<model_type>/specs.py`). Importing the package registers every spec;
consumers call
`get_spec(model_type)` / `match_moe_block(module, model_type)` and read
the fields.

The point is the infrastructure, not the migration. Today ModelOpt's
per-model knowledge
is scattered across if/elif chains in whichever subsystem happened to
need it, so the same
fact gets restated per subsystem and drifts. A model's MoE block
classes, its expert
projection naming, which of its norms store `w - 1` — these are facts
about the *model*,
and several subsystems want them. This PR gives them one home and one
lookup.

It is a kick-off, so it is deliberately narrow: specs plus the first
consumer. **Export is
that first consumer**, which is why most of the changed lines are on the
export side —
not because this is an export refactor. Quantization, speculative
decoding and the rest
keep their own tables for now; the per-model directory is where their
sections land later,
and `specs.py` is just the first thing in it.

| Data the first consumer moved in | Spec field | Consumers |
|---|---|---|
| MoE block classes and expert linear naming | `MoESpec` |
`get_expert_linear_names`, `get_experts_list`, `is_moe`,
`sync_moe_gate_up_amax` |
| grouped expert-export support set | `ExportSpec.grouped_expert_export`
| `get_experts_list` |
| AWQ `pre_quant_scale` fusion rules | `ExportSpec.pqs_fuse_rules` |
`fuse_prequant_to_linear` |
| weight-plus-one norm class names |
`ExportSpec.weight_plus_one_norm_names` |
`_layernorm_uses_weight_plus_one` |

`ModelSpec` holds each concern as a separate, optional section, and the
split is what
makes it extensible: **topic sections** hold architecture facts any
subsystem can read
(`MoESpec` — block classes, expert projection naming, how the experts
are stored),
**subsystem sections** hold one subsystem's own data and policy
(`ExportSpec` — AWQ fusion
rules, weight-plus-one norms, whether grouped expert export is validated
for this model).
A new subsystem adds a section rather than a table.

A section is `None` when the model has nothing to say about it, so a
dense model carries
`moe_spec=None` rather than an empty one. Two further fields record
where a model's
classes come from at all: `modeling_source` (`transformers` or
`remote_code`) and
`min_transformers_version`.

### Behavior changes

**1. Lookup key: `type(root_model).__name__.lower()` →
`config.model_type`.** The old key
changes after `quantize` wraps the model, forcing substring matching;
`model_type` is
stable, so `ExportContext` carries it and lookups are exact. Matching is
also tightened to
case-insensitive **exact** names against the module's **MRO**, so
quantized subclasses
still match via their base class without substring false positives.

**2. ⚠️ `get_expert_linear_names` raises instead of guessing.**
Unmatched MoE blocks
previously fell through to `["w1", "w2", "w3"]` (Mixtral naming); they
now raise
`NotImplementedError` telling you to register a `ModelSpec`.

The blast radius depends on the transformers release, because only
*iterable* expert
layouts consult the spec — a fused expert container is resolved by the
structural
first-projection check and never asks for naming. Measured against both
ends of the
support matrix:

| transformers | Unregistered MoE families `is_moe` admits | Reach the
spec lookup | Actually affected |
|---|---|---|---|
| 5.14.1 (`tf_latest`) | 19 | **0** — all fused | none |
| 4.57.6 (`tf_min`) | 9 | 7 iterable | **2** |

On 4.57, five of the seven (`ernie4_5_moe`, `flex_olmo`, `jamba`,
`olmoe`,
`qwen3_omni_moe`) name their experts `gate_proj`/`up_proj`/`down_proj`,
so the `w1`
fallback already failed with `AttributeError` on `main` — for them this
trades an
unhelpful error for one that names the fix.

The other two, **`minimax` and `phimoe`**, are Mixtral-derived with
`w1`/`w2`/`w3`
experts and did export on `main`, so they are now **registered** rather
than left to
regress. Both only take the per-expert path on transformers 4; on 5
their experts are
fused (`MiniMaxExperts`, `PhimoeExperts`) and the structural check
resolves them, which
their specs decline to answer for. **So no model that exported correctly
before this
change regresses.**

This is backward-incompatible and has a `CHANGELOG` entry.

**3. DeepSeek-V3 and V4 are now registered, and newly visible to the MoE
path.** Neither
was reachable before: `DeepseekV3MoE` (like `DeepseekV2Moe` and
`DeepseekV32MoE`) is
invisible to `is_moe` — its class name does not end in `SparseMoeBlock`
and it calls its
router `gate`, so neither the name test nor the structural
`router`+`experts` test
matches. `DeepseekV4SparseMoeBlock` was detected by name but had no spec
to resolve expert
naming from. Registering `block_names` is what puts V3 on the MoE path
at all, so this is
added coverage rather than a fix to existing behavior.

`deepseek` stays registered alongside them: it describes the remote-code
`DeepseekMoE`
block of DeepSeek-MoE/V1, and its spec is the only thing making that
block detectable.

**4. Fused expert containers are skipped in `get_experts_list` instead
of crashing.**
transformers 5 replaced several iterable expert `ModuleList`s with a
single module holding
3-D parameters, while the specs still describe the transformers 4
iterable layout — a spec
cannot tell the two apart, only the module can. The AWQ/SVDQuant
resmooth pass in
`requantize_resmooth_fused_llm_layers` reached `len(module.experts)` on
a module with no
`__len__` and died with `TypeError` mid-export. **Mixtral already hits
this on `main`**;
registering `deepseek_v3` only made an existing bug easier to reach. The
skip is scoped to
specs that claim iterable experts, so a layout the spec calls
unsupported (DBRX, whose
per-expert linears live under `experts.mlp`) still fails loudly rather
than silently
dropping resmoothing.

Everything else is behavior-preserving, pinned by tests:

- **`sync_moe_gate_up_amax` keeps generic coverage.** Quantization
admits MoE blocks
*structurally*, so unregistered families (Olmoe, Jamba, MiniMax…) reach
it. With no
spec it falls back to every declared gate/up naming — what it did
pre-registry.
Skipping them would leave the fused `gate_up_proj` halves on
inconsistent
  `weight_scale_2`.
- **`ExportSpec.grouped_expert_export` matches the legacy
`get_experts_list` support set**,
except for `deepseek_v3`/`deepseek_v4`, which are new and had no legacy
behavior to
  match. Note
`qwen3_5_moe`: legacy keyed off the root class name and
`"qwen3_5moeforcausallm"`
matched none of its qwen substrings (the `_5` breaks
`qwen3moeforcausallm`), so it
raised — the spec keeps `False` to match. Enabling it belongs in its own
PR.
- **`is_moe` consults the model's own spec first**, before the generic
name and structural
fallbacks, so per-model data always wins. The three checks are or-ed, so
this is a
no-op; the same reordering is deliberately *not* applied to
`get_expert_linear_names`,
  where the structural fused-experts check must stay first.
- **Two values are corrected, not moved.** `gpt_oss` declared
`block_names=("GptOssMoE",)`, which matches no real module (transformers
names it
`GptOssMLP`); naming still resolved via the single-naming shortcut, so
this was
  invisible. Now `("GptOssMLP", "GptOssMoE")`.

**5. ⚠️ DBRX expert input amax is now populated, changing its exported
scales.** The
second corrected value, called out separately because it has a numerical
consequence. The
legacy branch keyed on `DBRXMoeSparseMoeBlock`, a class transformers
does not define — it
names the block `DbrxFFN`. So DBRX fell through to the `w1`/`w2`/`w3`
default, every
`hasattr(experts_mlp, linear_name)` in `_prepare_dbrx_experts` evaluated
`False`, and the
handler wrote no expert input amax at all. `dbrx/specs.py` declares the
real names
(`w1_linear`, `w2_linear`, `v1_linear`), so the handler now does what it
was written to
do.

Anyone diffing a re-exported DBRX checkpoint against an older one will
see different
expert activation scales. That is the fix working, not a regression from
the refactor. No
`CHANGELOG` entry: DBRX is not listed as a supported export target in
`docs/source/deployment/` or `examples/hf_ptq/README.md`.

### Scope

One consumer wired up: the unified HF export path.

The TRT-LLM builders now live in `modelopt/torch/export/trtllm/` after
#2365, and are
deprecated. This PR changes one import line there — `is_moe` moved to
`modelopt/torch/models/moe.py`, so importing it from `..layer_utils` no
longer resolves —
and nothing else. Their `is_moe()` calls still pass no `model_type` and
fall back to the
all-specs search, reproducing today's behavior; their hardcoded tables,
including a second
copy of the weight-plus-one norm names, stay as they are and migrate
whenever that package
does.

Worth noting that #2365 landed the same boundary from the other
direction: the slimmed
`export/layer_utils.py` is now documented as "module-shape predicates
and MoE quantizer
helpers shared by every export backend", which is exactly the set of
functions this PR
rewires. The Megatron path has its own per-family registry and is
untouched; folding it in
is a later question, not this PR's.

`is_moe` itself moved out of `modelopt/torch/export/layer_utils.py` into
`modelopt/torch/models/moe.py` — whether a module is an MoE block is a
modeling question,
not an export one, and it is the first piece of shared modeling logic to
follow the specs
into the new package.

### Relationship to #1939 (why a second registry?)

Orthogonal layers: #1939's `ExportModuleRegistry` dispatches on **module
structure**
(*which handler runs?*), this registry resolves **family data** (*what
are its values?*).
Exporting a `QuantMixtralSparseMoeBlock`, #1939 picks the shared
iterable-experts handler
(also serving Qwen/DeepSeek/Gemma4); inside, `get_expert_linear_names`
resolves the
`mixtral` spec to `("w1", "w2", "w3")`. One handler serves many
families, one spec serves
many handlers. Sharing the matcher machinery is a planned follow-up.

### Usage

```python
# modelopt/torch/models/qwen3_moe/specs.py — adding a model needs no engine edits
register(
    ModelSpec(
        model_type="qwen3_moe",
        min_transformers_version="4.57",
        moe_spec=MoESpec(
            block_names=("Qwen3MoeSparseMoeBlock",),
            expert_linear_names=("gate_proj", "down_proj", "up_proj"),
            gate_up_pair=("gate_proj", "up_proj"),
        ),
        export_spec=ExportSpec(
            grouped_expert_export=True,
            pqs_fuse_rules=(
                (("Qwen3MoeAttention",), "v_proj", "o_proj"),
                (("Qwen3MoeMLP",), "up_proj", "down_proj"),
            ),
        ),
    )
)
```

One layout per model. `block_names` is a tuple, so a model whose MoE
appears under several
class names is covered as long as they share a layout (`gpt_oss`'s
`GptOssMLP`/`GptOssMoE`).
`gemma4_text` imports its section from `gemma4` rather than restating
it.

### Testing

`tests/unit/torch/export/` — **202 passed**;
`tests/unit/torch/quantization/` — **903
passed**. The `test_export_diffusers.py` collection error and 6
`test_quant_aware_conversion.py` failures are pre-existing and reproduce
on untouched
`main`, as does the one failing pre-commit hook
(`generate-arguments-md`, no torch in its
venv).

`tests/unit/torch/models/test_model_specs.py`: registry matching (MRO /
quantized classes
/ model-type scoping); every registered MoE section's `(block_names,
expert_linear_names,
fused_expert_names, gate_up_pair)` as an exhaustive table that fails if
a spec is added
without a row; the exhaustive `grouped_expert_export` support set; the
structural
fused-expert shortcut and the spec's precedence over it; the
fused-container skip in
`get_experts_list`; the `sync_moe_gate_up_amax` fallback; and
legacy-equivalence of the
aggregated `pqs_fuse_rules` / `gate_up_pairs` /
`weight_plus_one_norm_names`.
Mutation-checked — flipping a flag, typo-ing a block name, or swapping
the precedence each
fail.

`tests/unit/torch/models/test_specs_vs_transformers.py` validates the
specs against the
*installed* transformers rather than a mirrored table: every registered
block class must
name a real class from its `min_transformers_version` on, and every
`remote_code` model
must still be absent. It names no model — which specs to check, and
which cannot be, both
come from the registry.

> The run counts above were measured before the follow-up commits in
this branch (spec
> precedence, the policy split, the `MoESpec`/`MoELayout` merge) and
have not been
> re-measured; CI on the current head is the authority.

**Local per-model-type E2E sweep** (transformers 5.3.0, tiny models from
config, run
against this branch and base `main`; not checked in):

| model_type | real block class | this branch | base `main` |
|---|---|---|---|
| `qwen3_moe` / `qwen2_moe` / `qwen3_next` | `Qwen*SparseMoeBlock` |
`gate/down/up` | same |
| `mixtral` | `MixtralSparseMoeBlock` | `w1/w2/w3` | same |
| `nemotron_h` | `NemotronHMoE` | `up/down` | same |
| `gpt_oss` | `GptOssMLP` | `gate_up/down` | ❌ `w1/w2/w3` |
| `deepseek_v3` | `DeepseekV3MoE` | raises | ❌ `w1/w2/w3` |

Five types identical to base; both differences are base bugs the
registry surfaces.
`dbrx` could not be built on transformers 5.3.0 (`DbrxAttentionConfig`
missing
`rope_theta`). `sync_moe_gate_up_amax` was separately checked against
base across
registered, unregistered, and no-config models — identical in all cases.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `get_expert_linear_names` no
longer falls back to `w1/w2/w3`; see "Behavior changes" #2.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code, no new dependency (stdlib only).
- Did you write any new necessary tests?: ✅ —
`tests/unit/torch/export/test_model_specs.py`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — happy to add an entry given the backward-incompatible item.
- Did you get Claude approval on this PR?: ✅ — run; its CRITICAL finding
on `sync_moe_gate_up_amax` is fixed above.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added model-aware export support across a broad range of MoE and dense
architectures, including newer DeepSeek, Gemma, Qwen, Nemotron, Arctic,
DBRX, GPT-OSS, Llama, and Mixtral variants.
- Improved quantization and export handling for model-specific expert
layouts, projection fusion, normalization, and gate/up synchronization.
- Added automatic architecture recognition using Hugging Face model
metadata.

- **Bug Fixes**
- Unsupported or ambiguous model layouts now produce explicit errors
instead of applying potentially incorrect defaults.

- **Documentation**
- Added documentation describing model specifications and supported
architecture metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-11 12:12:54 -07:00
realAsmaandKeval Morabia dbe28e1e05 Document the MSE calibration API (#2405)
### What does this PR do?

Type of change: documentation

Expose `mse_calibrate` through `model_calib.__all__` so Sphinx
autosummary includes the existing MSE calibration API on the generated
`model_calib` reference page.

The documentation configuration honors each module's curated `__all__`
surface. Although `mse_calibrate` was implemented and used by the
calibration dispatcher, it was missing from that surface and was
therefore filtered out during API generation.

### Usage

```python
from modelopt.torch.quantization.model_calib import mse_calibrate
```

### Testing

- `pre-commit run --files modelopt/torch/quantization/model_calib.py`
- `git diff --check -- modelopt/torch/quantization/model_calib.py`
- Generated the recursive autosummary API tree using the repository
Sphinx configuration and module template.
- Verified the generated RST contains `mse_calibrate`.
- Rendered the focused module page and verified its function-table link,
anchor, signature, and docstring.

The canonical full build was unavailable in the active environment
because the configured `shibuya` theme is not installed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing
documentation build directly exercises this declarative autosummary
contract.
- Did you update Changelog?: N/A — this is a documentation-visibility
repair for an existing API.
- Did you get Claude approval on this PR?: N/A

### Additional Information

No source implementation or calibration behavior changed.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Made the MSE calibration capability publicly available for
quantization workflows.

- **Chores**
- Increased documentation build time limits to improve reliability for
longer-running builds.
- Increased multi-version test time limits to better accommodate
extended test runs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-11 18:18:23 +00:00
noeyy-mino a1bcda4727 Fix protobuf size-check failures in ONNX deployment (#2403)
### What does this PR do?

Type of change:  Bug fix:6701737

The ONNX deployment path assumed that ModelProto.ByteSize() would always
return a valid size. With newer protobuf versions, querying the size of
a model exceeding the protobuf serialization limit can itself raise
EncodeError: Failed to serialize proto.

Replaced both direct size checks with the existing
is_model_too_large_for_protobuf() helper. This helper handles size-query
failures conservatively and checks the protobuf size limit:

Shape inference now selects the external-data/file-based path when
ByteSize() fails or the model is too large.
Metadata creation uses the same safe check instead of raising another
serialization error.
The unused TWO_GB constant was also removed.

### Usage

```
python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image
```

### Testing
the above test case pass on B100

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved ONNX model size detection during shape inference and export
processing.
* Ensured large models consistently use the appropriate external-data
handling path.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
2026-09-11 10:46:10 -04:00
Wei-Ming Chen 59d93af064 [OMNIML-5570, OMNIML-5569] 1/2 Add layer-wise KV-cache AutoQuant with forward KL (#2272)
### What does this PR do?

Type of change: new feature.

Adds standalone layer-wise KV-cache AutoQuantize through the existing
public
`mtq.auto_quantize` API:

- dispatches KV search with
`constraints={"effective_bits": ..., "cost_model": "kv_cache"}` and
forward-KL
  sensitivity;
- selects one supported K/V format for every eligible causal-attention
layer;
- supports persistent/exportable FP8 K/V, NVFP4 K/V, and FP8-K/NVFP4-V
candidates;
- solves a K/V-width- and scale-storage-aware additive recipe with the
existing
  PuLP-backed constrained solver;
- uses `BaseSearcher` lifecycle and safe checkpoint restore/save
machinery;
- preserves existing non-KV execution while isolating K/V candidate
calibration;
- returns standard AutoQuantize state that can be re-solved at another
KV budget;
- produces a complete KV-only replay config that disables every non-KV
quantizer;
- saves JSON-safe sensitivity metadata and the exact selected layer
mapping; and
- invokes the public API from `examples/hf_ptq/hf_ptq.py` through a
standalone
  calibration-free recipe.

The implementation is architecture-driven. Plain and
conditional-generation Qwen
causal attention is supported, VLM vision attention is excluded through
the existing
language-model extraction boundary, hybrid full-attention mixers are
discovered through
their paired K/V quantizers, and nonattention/Mamba modules remain
outside the search.
Ambiguous language-model roots, unsupported distributed execution,
structural
algorithms, invalid storage declarations, nonpersistent scales, and
unsupported K/V
pairs fail closed.

KV-only unified HF exports leave weight-quantization fields unset.
Uniform all-FP8 or
all-NVFP4 selections retain their legacy KV scheme while also carrying
the complete
`kv_cache_quantized_layers` map and schema version; genuinely
layer-mixed selections use
the KV-side `MIXED_PRECISION` marker plus the same map. This keeps
weight-loader metadata
accurate and prevents disabled vision attention from making uniform
language-model KV
quantization appear partially quantized.

GEMM PTQ/AutoQuantize followed by KV AutoQuantize is intentionally
excluded and proposed
separately in stacked PR #2273.

### Why KV search has a dedicated backend

The user-facing entry point remains `mtq.auto_quantize`; no separate
public KV search API
is introduced. `AutoQuantizeKVSearcher` extends `BaseSearcher` and
reuses its reset,
checkpoint load/save, and search lifecycle, along with existing Pydantic
configuration,
calibration, safe checkpoint I/O, and PuLP-backed selection utilities.

The backend remains KV-specific because a decision owns paired K/V
quantizers on one
attention layer, its cost depends on separate K/V widths and data/scale
storage, BF16 is a
scoring reference but not a deployable solver choice, and the
optimization objective is
additive isolated forward KL under a KV-storage constraint. These
contracts do not match
the weight-domain hparam grouping, parameter-count cost, or
threshold-selection behavior
of the existing weight AutoQuant searchers. Keeping the specialization
behind the shared
API avoids changing established weight-search solver and scoring
behavior.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3.8-27B \
  --recipe general/auto_quantize/kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3.8-27b-mixed-kv
```

The search checkpoint is compatible only with the same model,
eligible-layer geometry,
candidate configurations, and scoring setup. Use a distinct checkpoint
path after any of
those inputs change.

KV-cache AutoQuantize rejects `--use_fsdp2` before model loading because
its sensitivity
scoring, selection, and checkpoint writes are single-process. Existing
weight
AutoQuantize retains its previous experimental FSDP2 warning and
behavior.

### Testing

- Focused coverage exercises candidate validation/calibration, paired
K/V scoring and
storage accounting, solving, checkpoint resume, failure atomicity,
disabled layers,
fresh-model replay, Qwen/VLM/hybrid boundaries, JSON-safe reports, and
unified export.
- Uniform FP8/NVFP4 KV-only exports retain the legacy KV scheme and
complete layer map
without claiming a weight algorithm; disabled VLM vision attention is
excluded from
  causal-KV eligibility.
- The shipped recipe runs end to end on a tiny offline Qwen fixture and
preserves
  exportable scale state.
- After merging current `main`: 432 focused recipe/KV/export/hf_ptq
tests passed, with one
unrelated optional-dependency skip; changed-file pre-commit hooks
passed.

### Deployment gate

The producer schema is covered here. Runtime consumption of
`kv_cache_quantized_layers` is tracked in vLLM PR
https://github.com/vllm-project/vllm/pull/52813. Do not treat a produced
checkpoint as
runtime-supported until that consumer lands and the target K/V kernels
are available.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow
  guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅

### Additional information

- This is split from the combined ground-truth implementation in draft
PR #2211 to reduce
  review scope; composition is isolated in #2273.
- The standalone core tree contains no composed GEMM→KV recipe schema or
orchestration.
- No model-name checks, checkpoint-specific layer lists, campaign data
contracts, cluster
  launch logic, or runtime-kernel implementations are included.

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-10 22:48:34 -07:00