100 Commits
Author SHA1 Message Date
Chenjie LuoandClaude Opus 5.5 ad8cd63847 Share one CUDA encoder per IQ family (#2615)
### What does this PR do?

Type of change: refactor (no behaviour change)

The five GGML IQ CUDA encoders were five copies of the same search.
IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and
launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same
structure with a different choice space. Every scaled packer also
validated its scales twice, in the `ggml.cpp` pybind wrapper and again
in the CUDA entry point.

This PR keeps **one encoder per family**, as two templates:

- **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in
shared memory, the 16 local scales are scored per group, and each vector
then takes its best entry under the chosen scale. A format supplies its
group shape, whether it stores seven sign bits and recovers the eighth
from parity, and a `store()` that writes the chosen entries, sign masks
and local scales into its layout.
- **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary
grid. Each group picks one of `kChoices` options. With `kSharedShift`
the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`);
otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now
writes FP16 scales, so both IQ1 formats take the same input.

Each format file is now one `Format` struct, holding its layout
constants and `store()`, plus its entry point: 58–100 lines each.
Validation lives once in `common.cuh`, as `check_pack_inputs` and
`check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points
directly instead of through five wrappers. **The kernel sources shrink
from 1,536 to 1,241 lines** (+665 / −960).

This is the first of two PRs. #2604 builds on it: it adds CUDA decoders
as a `decode()` next to each format's `store()`, and makes export reuse
fake quant's packed payloads.

### Testing

**Nothing changes in the output.** Before the refactor I hashed 40
outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode
and decode, on a weight with zero, tiny, oversized and non-finite
blocks. All 40 hash the same afterwards.

**Encode speed is unchanged.** Old and new were timed alternately for
four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048
weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms,
IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 /
15.4.

- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed**
- **Validation reports the same errors in the same order.** Over 5
formats × 8 combinations of bad arguments (devices, dtype, width, grid
shape, scales dtype, length and sign), every first error matches main's.
- `tests/gpu/_extensions/test_torch_extensions.py`: the
validation-message tests pass. #2515's two Q8_0 tests fail identically
on a clean `main` on this GPU.
- IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`,
`test_convert_hf_config.py`, `test_presets.py`,
`test_export_weight.py`): **173 passed**

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Same bindings, messages and
bytes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: N/A. A refactor with no
behaviour change, verified by the hashes above and the existing GPU
tests.
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: **this** → #2604 (pack each IQ weight once and decode on
CUDA).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Quantization now checks that inputs and grids are CUDA tensors on the
same device, with compatible shapes. Scaled formats also validate scale
type, shape, and finite, non-negative values.
* **Improvements**
* IQ1 and IQ2 formats share common encoding paths while retaining their
format-specific output layouts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:14:42 -07:00
Chenjie LuoandClaude Opus 5.5 7d9e07d14b Count only added lines toward the PR size budget in AGENTS.md (#2616)
### What does this PR do?

Type of change: documentation

Changes the PR sizing rule in `AGENTS.md` to count only **added source**
lines toward the ~500-line budget, instead of total changed lines.
Deletions are cheap to review, so a PR that mostly removes code
shouldn't be pushed into a split. Tests and docs are excluded too, since
every sub-PR has to carry its own tests. The check uses the insertions
count from `git diff --shortstat` with a pathspec that excludes `tests/`
and `docs/`.

### Usage

```bash
git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs'
# N files changed, X insertions(+), Y deletions(-)  -> compare X against ~500
```

### Testing

- `pre-commit run --files AGENTS.md` (markdownlint passes).
- Ran the pathspec against recent commits (#2595, #2513) to confirm it
drops test and doc lines from the count.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Follow-up to #2494, which introduced the sizing guidance.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated review guidance to measure pull request size by added source
lines, excluding deletions, tests, and documentation. The guidance
retains the recommendation to check the size before opening a review.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 07:03:48 +00:00
Chenjie LuoandClaude Opus 5.5 aa89722d38 [6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed
the PyTorch codec; this PR adds its **CUDA encoder** and makes the
format reachable. With it, ModelOpt supports all five GGML IQ formats at
one and two bits.

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq1_m`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

### The kernel

In the kernel the delta shift is free per group, so it sits above the
entry index in the sort key: a tie still prefers the lower shift and
then the lower entry, as the reference encoder does. The 2048-entry grid
IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory
limit, so both kernels read it from global memory and rely on the cache.

| 5632×2048 weight | torch | CUDA | |
|---|---|---|---|
| IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** |

### Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and
apply the ±1/8 shift the same way. So those three steps move into
`common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and
IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA
output on a 5632×2048 weight hashes the same before and after, and so
does IQ1_M's, compared against the pre-split version of this change.
IQ1_S encodes at the same speed (306 M elem/s).

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ1_M-specific test code: backend dispatch, weight caching, the
`num_bits` guard, `convert_hf_config` metadata, Megatron export and the
`TensorQuantizer` tests in the shared battery. The shared CUDA battery
gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **166 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **221 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO
6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch
encoder parity
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** (9 tests × 5 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed**
- reconstruction error falls monotonically across all five formats,
pinned by a test
- `general/ptq` now holds 31 recipes; `ptq.md` is updated.

Rebased onto `main` after #2513 merged. The resulting tree is identical
to the one the runs above tested, and the unit set was rerun on it: 166
passed.

On this GPU, two of #2515's Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M
codec), all merged → **this**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M weight-only quantization at 1.75 bits per weight, with
CUDA acceleration and a 256-value block size.
* Added an IQ1_M post-training quantization recipe for eligible linear
layers; calibration data is not required.
* Added IQ1_M to the supported GGML-compatible formats and recipe
listings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:23:48 -07:00
Chenjie LuoandClaude Opus 5.5 e5b63320ab [5/6] Add the IQ1_M codec (#2513)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 09:45:48 -07:00
Chenjie LuoandClaude Opus 5.5 3091b8ff69 [4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512
landed the PyTorch codec; this PR adds its **CUDA encoder** and makes
the format reachable:

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq2_s`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B
parameters**.

### The kernel

IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its
search the most expensive in the family. The codebook and its norms take
36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB
static limit, so they are declared statically like the IQ2_XS and
IQ2_XXS kernels.

That cost is why the kernel matters more here than anywhere else:

| | torch | CUDA | |
|---|---|---|---|
| IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×**
|
| extrapolated to a 27B model | ~9.8 hours | **~37 s** | |

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ2_S-specific test code: backend dispatch and weight caching, the
`num_bits` guard, `convert_hf_config` metadata (uniform and mixed
precision), all 9 Megatron export tests, and the two `TensorQuantizer`
tests in the shared battery. The shared CUDA battery gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **134 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **192 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO
6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder
parity, determinism, reconstruction at scale, zero and non-finite
policy, float64 input and the fallback path.
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **36 passed** (9 tests × 4 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**.
TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`,
`block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)`
uint8.
- `general/ptq` now holds 30 recipes.
- The shared-memory change in `b7739d5d0` leaves the packed bytes
identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M
elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun:
42 passed.

All of the above was rerun after rebasing onto `main` at `c2aaa44f6`.
That base adds a Q8_0 packer to the same GGML extension (#2515), and
changes the hf_ptq example and the export code this format goes through.
The packed IQ2_S bytes still hash the same. On this RTX PRO 6000
(sm_120), two of #2515's own Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
#2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ2_S weight-only quantization for eligible linear layers, at
2.5625 bits per weight.
* Added a PTQ recipe that requires no calibration data. Weights must
meet the existing 256-value block-size constraint.
* Added CUDA-accelerated packing for CUDA weights, with a Python
fallback when the CUDA extension is unavailable.
* **Documentation**
  * Updated the PTQ recipe catalog and IQ-format size tradeoffs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 20:35:38 +00:00
Chenjie LuoandClaude Opus 5.5 767ef5533e [3/5] Add the IQ2_S codec (#2512)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at
one and two bits (2.5625 bits per weight). This one lands the **PyTorch
codec**: the encoder, the decoder and the 1024-entry codebook. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2565 adds the CUDA encoder,
registers the format and adds its recipe.

### What's distinctive about it

**IQ2_S is the one format llama.cpp's own tooling gives no head start
on**, so the search is written against the GGML layout directly.

The interesting difference from IQ2_XS and IQ2_XXS is sign handling.
IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit
parity-coded index. The encoder therefore takes the input signs as they
are instead of flipping the weakest element to fix parity, and the
search compares magnitudes directly, which is simpler than its siblings.

### Why the codec lands before the kernel

The CUDA encoder's tests use this codec as their reference. They compare
against the PyTorch encoder byte for byte and draw the grid and scale
predictor from it. So the kernel cannot be tested before the codec
exists, and it follows in #2565 together with the registration. Every
registered format therefore keeps a CUDA encoder.

### Test changes that make the split possible

A codec can now land before it is registered, so two test contracts in
`test_iq_formats.py` are stated precisely:

- The two tests that go through `TensorQuantizer` (pass-through
gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`.
Every other battery test calls the codec directly and covers IQ2_S here.
- The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <=
set(FORMATS)` instead of equality. That is what its docstring already
said: a registered format must be listed, or it escapes the contract.
- `test_registry_lists_every_exported_encoder` is unchanged, and it is
why this PR leaves the package exports alone: an exported encoder must
be registered.

The error-by-bit-width failure message also labels errors by the order
they were measured in; it previously zipped them with alphabetical
names.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry.
Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so
CI keeps checking bytes we did not produce.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S
codec cases, including the llama.cpp conformance check
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by
this PR

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new
codebook is a GGML table, carried in `codebooks.py` with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
**this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added GGML-compatible IQ2_S quantization and dequantization support,
including access to its magnitude grid.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 14:03:01 -07:00
Chenjie LuoandClaude Opus 5.5 80e04b8816 docs(eval): align Terminal-Bench 2.1 / SWE-bench / MRCR with upstream configs (#2479)
### What does this PR do?

Type of change: documentation (agent skill) + template bug fixes

Aligns the `evaluation` skill's three upstream-tracked benchmarks
(Terminal-Bench 2.1,
SWE-bench Verified, MRCR) with the current
`nvidia-eval-factory-benchmarking` configs, and
fixes guidance that turned out to be wrong when a full three-benchmark
campaign was run with
the skill end to end. Rebased on #2499. The first commit is the
alignment. The rest address
review and a fresh config-generation test: MRCR serving scoped per
variant, `limit_samples`
canary guidance, the template's serve command passing
`--gpu-memory-utilization` (replacing
`command:` had silently dropped it, so the 1M golden's 0.95 never
reached vLLM), upstream's
128K values, and regression tests for the `++limit` gate and the serve
command. The skill
text keeps only the rules; the evidence behind them is below.

GDPVal is out of scope. It was removed from this skill in #2470, and
this PR replaces #2464.

**Alignment with upstream**

- Sandbox region via `HARBOR_ECS_REGION`. TB2.1's ECR repo name tracks
the region;
  SWE-bench's stays in us-west-2.
- One interceptor order for both harbor benchmarks, with
`http_pairs_dump` **last**. Upstream
is split on its position, which changes only what the dump records,
never the score.
Interceptor lists replace wholesale on merge, so a leaf must restate the
whole chain.
- `capture_request_body` goes on the service. A shared block injects an
alias-only entry with
  no `type`.
- MLflow tags gain `task_name` and `nemo-evaluator-next-version`.
- TB2.1 `max_concurrent`: 50 for nano-class models, 15 for larger ones
(all upstream non-nano
  leaves override it to 15).
- `proxy.request_timeout` must always be set explicitly. Otherwise it
inherits the model
fragment's serving value, which ranges from 3600 to 36000 upstream, and
TB2.1 has no
  benchmark key for it.
- MRCR:
- `parallelism` is 512, deliberately above server capacity, so
`--max-num-seqs` must no
    longer be derived from it.
  - `limit_samples` now reaches the gym through a gated `++limit`.
  - Observability capture is on.
  - 128K is its own upstream benchmark on the condensed gym schema.

**Corrections found by running it**

| what the skill said | what actually happens |
|---|---|
| `username: ${oc.env:USER}` | nel-next only expands `${VAR}` /
`${VAR:-default}`. The config passes `--dry-run` and fails at `--submit`
with *"remote username contains invalid characters"*. |
| MRCR canary via `++limit` edited into `collect_rollout_params` |
`limit_samples` is now gated through. Under the 0.2.6 launcher the `-o`
path is
`++evaluation.nemo_evaluator_config.config.params.limit_samples`.
`++config.params…` creates a bogus top-level key. |
| the condensed gym schema's bootstrap is in the runtime image | It
lives in upstream `configs/models/gym_eval_command.yaml`, which is
composed in. A standalone config must carry the `command:` block. |
| `mean/prefix_matched ~0.55 is healthy` | That value is calibrated to
the 1M golden. A 128K run at `pass@1` ≈ 95 measured ≈ 1.0. The signal is
a collapse toward 0. |
| NVFP4 MoE `VLLM_*` env vars as reliable knobs | They are
build-dependent: one vLLM build logged them as unknown and ignored them.
Check the server log once per image. |

**Rules the skill lacked**

- **MRCR variant.** Pick the largest variant within the checkpoint's
trained context. On a
262K-context model, 1M needs `VLLM_ALLOW_LONG_MAX_MODEL_LEN` and
measures extrapolation, which
a quantization comparison would then entangle with quantization damage.
The serving setup
follows the variant: 128K serves at the trained context without the
override.
- **SWE-bench `reasoning_effort`.** openhands-sdk sends
`reasoning_effort: high` on every call,
and canonical `bench.yaml` doesn't strip it. A server whose accepted set
excludes `high`
returns HTTP 400 on the first call of every trial, so `pass@1` is 0. The
fix is to overwrite
it with the server default via `proxy.extra_body`. terminus-2 (TB2.1)
and Gym's
`simple_agent` (MRCR) never send the key (46/46 and 110/110 requests
checked).
- **Reasoning toggles** go in `extra_body.chat_template_kwargs`. Don't
write out no-op sampling
  defaults; `top_k` is *not* one (vLLM's default is `-1`).
- **Upstream model fragment.** Consult `configs/models/<model>/` when it
exists, for serving
  flags, the thinking toggle and `reasoning_replay.mode`.

### Usage

No API change. Regenerating a config from the skill now yields the
aligned values:

```yaml
# recipes/examples/example_eval_next.yaml
services:
  model:
    proxy:
      request_timeout: 3600                  # always explicit
benchmarks:
  - max_concurrent: 50                       # nano-class; larger models 15
    sandbox:
      region: ${HARBOR_ECS_REGION:-us-east-1}
cluster:
  username: ${USER}                          # NOT ${oc.env:USER}
```

### Testing

- `python -m pytest plugins/modelopt/skills/ -o addopts=""`: 7/7 pass,
including the new
`tests/test_example_mrcr.py`. It checks that `example_mrcr.yaml` emits
`++limit=N` only when
`limit_samples` is set, and that the folded vLLM serve command
shell-parses with every flag
intact, including `--gpu-memory-utilization`. Each check fails when its
defect is reintroduced.
- `markdownlint-cli2` on the changed Markdown: 0 errors. Both example
YAMLs parse.
The full `pre-commit` suite was not run after the squash, because the
sandbox could not fetch
  hook repos. The pre-squash commits passed it.
- **Exercised end to end.** Configs built from this skill ran a BF16
campaign for a 262K-context
  MoE reasoning model on an internal cluster to completion:
  - MRCR-128K: 1470/1470 rollouts
  - Terminal-Bench 2.1: 712/712 trials
  - SWE-bench Verified: 2500/2500 trials

  Each correction above is a defect that campaign surfaced.
- **Differential check.** Terminal-Bench configs generated from the pre-
and post-change skill,
  from the same brief and in isolation, differ on:
  - interceptor chain
  - concurrency
  - `top_k`
  - region interpolation
  - MLflow tags

Not run: a scored evaluation of this PR itself.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Docs + a template fix; no
ModelOpt API surface touched. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- Documentation and
config-template values. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- Agent skill docs, not a user-facing ModelOpt feature/breaking
change/deprecation. -->
- Did you get Claude approval on this PR?: ❌ <!-- Not run. -->

### Additional Information

The companion internal `eval-config` change now carries only the
internal values: the ECR URLs,
the region default and the cluster image notes. It points here for the
generic rules.

**Known gaps not fixed here**, worth a follow-up:

- `references/nel-next.md:136` and `references/launcher-workflow.md:207`
derive `--max-num-seqs`
from `parallelism / DP`. nel-next has no `parallelism` field (its
analogue is
  `max_concurrent`), and MRCR's 512 is deliberately not a server cap.
- `references/launcher-workflow.md:200-201` makes
`--max-num-batched-tokens` and
`--enable-chunked-prefill` always-include defaults, but
`example_eval_next.yaml` omits both.
- `references/launcher-workflow.md:27` says `sbatch_comment` belongs
under `execution:` and is
otherwise inert, yet all three shipped examples put it under `cluster:`.
- MoE detection (`--enable-expert-parallel`) is unresolvable from the
facts the skill asks for
  when the model handle has no `-A*B` suffix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:07:35 +00:00
Chenjie LuoandClaude Opus 5.5 400498d82d [2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do?

Type of change: refactor (no behaviour change)

Addresses review feedback on #2511. Backend dispatch and export each
kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the
backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in
export. All four listed the same formats. Adding a format meant a row in
each, and the lists could drift apart. That had already happened twice
in #2511: `convert_hf_config.py` kept its own upper-case spelling of the
family and dropped IQ2_XXS metadata, and the Megatron export tests were
hard-wired to two formats.

Each format module now declares **one `IQFormat` record** beside its
encoder and decoder: name, block geometry, `quantize`, `dequantize`, and
its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists
them.

- Backend dispatch looks formats up in the registry.
- Both exporters take the packer and block geometry from it.
- Export's `IQ_FORMATS` is derived from it instead of being written out
again.
- `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed.
- The per-format fake-quant wrappers collapse into one
`IQFormat.fake_quant`, which does the `num_bits` check and calls the
existing cache helper.

Codebooks, searches, payload layouts and CUDA encoders stay in each
format's module.

#### Series and merge order

This is one slice of the IQ format series. It targets `main` so unit CI
runs, and **its diff includes #2511's commits until #2511 merges**.

1. #2511 — IQ2_XXS format
2. **this PR** — one registration per format
3. #2512 — IQ2_S format
4. #2513 — IQ1_M format

After this lands, #2512 and #2513 are restacked onto it, so each adds a
format module and a single registry entry instead of rows in four
tables.

#### Design choices

- **An explicit list, not self-registration at import.** If formats
registered themselves when their module was imported, the registry's
contents would depend on import order.
- **Backward compatible, with one behaviour change.** `iq1_s_fake_quant`
and `iq2_xs_fake_quant` are public on main, so each format keeps its
`<fmt>_fake_quant` name as an alias of its record's method. The three
removed tables were introduced by #2511 and never released. The
behaviour change: on main, the alias looked the encoder up at call time,
so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record
captures the encoder and decoder when it's built, so patching those
module functions reaches neither dispatch nor the alias. Substitute
through `IQ_FORMAT_REGISTRY` instead.
- **Registering a format declares it exportable, and that's intended.**
Export's `IQ_FORMATS` is derived from the registry, so a format
registered for dispatch is also claimed by both exporters and
`convert_hf_config`. That can't be wrong for an IQ format: fake quant is
`dequantize(quantize(w))`, so a format can't be dispatched without the
packer and block geometry, and those are all export reads. A QAT-only IQ
format can't exist. If one ever needs to land ahead of its export path,
an `exportable` flag on the record is a one-line addition.
- **The registry is the substitution seam.** Dispatch now reads the
registry, so tests that swap an encoder or decoder swap the registry
entry. Patching the format module's function would no longer reach
dispatch.
- **Test expectations stay independent of the registry.** Tests take the
*list* of formats from the registry, but their expected values come from
each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A
mis-wired registry entry therefore can't make both sides of an assertion
agree.

#### What it does not unify

The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list,
codebook sizes in `common.cuh`) and the recipes and docs remain per
format. "One registration" holds for the Python side, which is where all
four tables lived.

### Usage

Adding a format after this PR (for example IQ2_S in #2512) needs its
module and one line in the registry:

```python
# modelopt/torch/quantization/ggml/iq2_s.py
IQ2_S_FORMAT = IQFormat(
    name="iq2_s",
    block_size=IQ2_S_BLOCK_SIZE,
    block_bytes=IQ2_S_BLOCK_BYTES,
    quantize=quantize_iq2_s,
    dequantize=dequantize_iq2_s,
    block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE,
    decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE,
)

# modelopt/torch/quantization/ggml/registry.py
IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)}
```

Looking up a format:

```python
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY

fmt = IQ_FORMAT_REGISTRY["iq2_xxs"]
packed, shape = fmt.quantize(weight)          # GGML blocks
fmt.block_bytes, fmt.effective_bits           # 66, 2.0625
```

### Testing

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py` — **94 passed**
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed**
(RTX PRO 6000)
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses
for that suite
- broader sweep of IQ, export and recipe unit tests — **153 passed**,
none failed

**New guards on the registry itself:**
- every encoder the package exports is registered
- each record points at its own format's codec, geometry and chunk
defaults
- the public `<fmt>_fake_quant` alias is the registered record's method
- export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the
registry
- a format's `fake_quant` refuses a quantizer configured for another
format. Dispatch picks the record by `num_bits`, so it never reaches
this guard; the test covers direct callers of a record or alias. The
three per-format guards it replaced were untested on main.
- every registered format is listed in the shared test batteries

Checked by mutation: leaving IQ2_XXS out of the registry, or registering
it with the IQ2_XS encoder, each fails the guard written for that case.

**Coverage gap closed along the way:** `test_ggml_backend.py` was
hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or
packed-once coverage. Those tests now run over the registry.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — public per-format fake-quant
names are kept as aliases; the removed tables were never released.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — internal refactor with no
user-visible change
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`,
`IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently
enumerate the same formats. A common pack/dequantize/fake_quant
interface would let backend dispatch and export consume one
registration."*

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* IQ quantization formats are available through a shared format
registry, keeping format details and quantization behavior consistent
across supported workflows.
* IQ-format model exports use registered format information for
quantization metadata and weight packing.
* **Tests**
* Expanded checks to cover registered IQ formats and verify consistent
format support across quantization and export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 23:32:11 +00:00
Chenjie LuoandClaude Opus 5 a21411adde Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do?

Type of change: new feature

llama.cpp defines five GGML IQ formats at one and two bits; we ship two.
This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and
IQ2_XS, and is the **first of three**.

On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`,
`Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84
B parameters — 10.6% of the file**, which a reader limited to
IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats
account for 17.3%.

| format | bpw | bytes/256 | codebook | |
|---|---|---|---|---|
| `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing |
| **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this
PR** |
| `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing |

The encoder follows the existing single-pass grid search at a fixed
anchored super-block scale, and the CUDA kernel the existing per-block
structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a
4-bit sub-block scale into the same 32-bit word as four 7-bit sign
indices, and its 256-entry grid needs no high index bits.

### Groundwork the next two reuse

Two things land here because IQ2_XXS is the first format to need them:

- **Export registry.** The IQ family was spelled as a two-element tuple
at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and
`unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset
plus per-format packer and block-geometry tables, so a format is a row
rather than a sweep through the exporters.
- **Shared test contract.** The per-format test files had drifted apart
— each of `iq1_s` and `iq2_xs` tested things the other did not. They
become one parametrized module per layer (unit and CUDA), so every
format is held to the same contract and a new one inherits it.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs
```

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared
against `dequantize_row_iq2_xxs` from `ggml-quants.c`:

```
IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry, as
does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint
ship as conformance vectors so CI keeps checking bytes we did not
produce; mutation testing confirms they catch a wrong sign-field width.

The CUDA encoder is byte-identical to the PyTorch reference on a fixed
input and runs at **1047.9 M elem/s against the torch search's 10.7** on
a 5632×2048 weight.

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99
passed
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7
checks × 3 formats)
- `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds
29 recipes, `ptq.md` updated
- reconstruction error decreases monotonically with bit width, pinned by
a test

Pre-existing failures in `tests/unit/torch/export/` and
`test_autoquant.py` are `transformers`/`torchvision` import problems in
my environment — identical counts with and without this change.

### A finding about already-merged code

Checking the new kernel against its PyTorch reference at 4096 blocks
showed that **CUDA and torch encoders disagree on roughly 1 block in
6000 — including the already-merged `iq2_xs`**, at 0.0163% against
IQ2_XXS's 0.0000%.

Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA
fuses it with `fmaf` while torch uses separate ops; where two local
scales fall within a float32 ULP the roundings pick different sides.
Adjudicated against float64, neither path is better (5 to 6). Worst-case
cost is **1.48e-08** relative reconstruction error, and run-to-run
determinism on a given device holds.

This is pre-existing, not introduced here — `test_iq2_xs_cuda.py`
asserts exact byte parity but on a 16-block weight where ties
essentially never arise. I have **not** changed that test; rewording a
guarantee on merged code belongs in its own change. The new shared GPU
tests assert exact parity on a small fixed input and compare
reconstruction error at scale.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new
codebook is a GGML table, carried in `codebooks.py` beside the existing
ones so the MIT-licensed surface stays in that one file, with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

First of three; **IQ2_S** and **IQ1_M** follow and build on this branch.
Replaces #2505, which carried all three at once. Follows #2446 / #2447 /
#2448 / #2449, which landed IQ1_S and IQ2_XS.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added IQ2_XXS weight-only quantization, including CUDA acceleration
and support for Hugging Face and Megatron exports.
- Added the `general/ptq/iq2_xxs` recipe. It requires no calibration
data and supports eligible layers with a weight dimension divisible by
256.
- Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at
approximately 1.56, 2.06, and 2.31 bits per weight, respectively.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 18:04:37 +00:00
Chenjie LuoandClaude Opus 5.5 25d8c91762 Add guidance on keeping skill updates concise to AGENTS.md (#2523)
### What does this PR do?

Type of change: documentation

Adds an `## Updating skills` section to `AGENTS.md` (symlinked as
`CLAUDE.md`). Skills are loaded into agent context, so each extra line
costs tokens every time the skill runs. The new guidance tells the agent
to:

- Keep skill edits concise: add only what changes agent behavior, and
tighten existing text instead of appending more.
- Do a final compression pass over the skill diff before opening a PR:
drop unnecessary explanations and examples, cut redundancy, and merge
overlapping guidance.

### Usage

N/A — no API or flag change.

### Testing

`pre-commit run --files AGENTS.md` (markdownlint and the other
applicable hooks pass).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- documentation-only
change -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- agent instructions only, not user-facing -->
- Did you get Claude approval on this PR?: ❌ <!-- will run /claude
review if reviewers want it -->

### Additional Information

Follows #2494, which added the PR sizing guidance to the same file.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added guidance for keeping skill updates focused on behavior changes
and reviewing edits for unnecessary detail.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 10:07:22 -07:00
Chenjie LuoandClaude Opus 5 7a35cada39 Add guidance on sizing and splitting PRs to AGENTS.md (#2494)
### What does this PR do?

Type of change: documentation

Adds a `## Sizing and splitting PRs` section to `AGENTS.md` (symlinked
as `CLAUDE.md`) so AI-assisted work stops producing one giant PR that
nobody wants to review.

The new guidance tells the agent to:

- Keep each PR that goes up for review under ~500 changed lines of
source, and check the size before opening.
- Propose the split *before* opening an oversized PR rather than after.
- Split on file/directory/module boundaries first and fall back to
feature boundaries (enabling refactor first, then one PR per behavior it
unlocks).
- Keep the series acyclic and linearly ordered — no circular
dependencies between sub-PRs — and state the merge order.
- Prefix sub-PR titles with `[x/N]` so reviewers know the PR is one
slice of a planned split, and link the siblings.
- Make every sub-PR stand on its own: it builds, carries unit tests for
the code it introduces, and passes CI without the later PRs.
- Optionally submit the whole change as a reference-only draft PR for
the big picture, cross-linked with the sub-PRs.

### Usage

N/A — no API or flag change.

### Testing

`pre-commit run --files AGENTS.md` (markdownlint and the other
applicable hooks pass).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- documentation-only
change -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- agent instructions only, not user-facing -->
- Did you get Claude approval on this PR?: ❌ <!-- will run /claude
review if reviewers want it -->

### Additional Information

This PR is itself well under the new budget (26 added lines in one
file), so no split applies.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated pull request sizing guidance to allow oversized changes when
they cannot be meaningfully split.
- Clarified that draft aggregate pull requests are optional and intended
for reference only when splitting would reduce clarity.
- Added examples covering self-contained changes and new models or
backends without a functional intermediate state.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 22:52:53 -07:00
Chenjie LuoandClaude Opus 5 a7166965e3 Drop GDPVal support from the evaluation skill (#2470)
### What does this PR do?

Type of change: deprecation (agent skill)

Removes GDPVal support from the `evaluation` agent skill. Its task
recipe, example
config and Apptainer SIF build helper are deleted.

GDPVal was not cleanly separable, so this is not just a delete:

- **MRCR depended on GDPVal's infrastructure.** `scripts/nel-gdpval.sh`
was a
generic pinned-0.2.6 `nel` launcher that only happened to be
GDPVal-named, and
MRCR ran through it; `references/gym-gdpval.md` documented the gym
bootstrap
machinery (prepare/reap, `install_on_the_fly` pin↔container coupling,
the trust
env vars) that both examples share. So the shared parts are kept and
renamed
rather than dropped: `scripts/nel-gym.sh` and `references/gym.md`. The
launcher
  pin itself is unchanged (0.2.6), as are its env-override semantics.
- **GDPVal is an AA-suite member**, so the "AA rule" in `SKILL.md` and
`references/quantization-benchmarks.md` had to change. An "AA" /
"Artificial
Analysis" request now generates the `aa/` tasks as one multi-task config
and
nothing else. Both files now state that the resulting set omits GDPVal
and is
therefore not directly comparable to a published AA Index — report
per-task
scores rather than an aggregate. `SKILL.md` also tells the agent to say
GDPVal is
unsupported rather than reconstruct a config from an older copy of the
skill.

Everything GDPVal-specific is gone from `references/gym.md`: the SIF
sandbox and
its silent-unsandboxed-exec failure mode, the 3-member judge panel,
rubric vs.
comparison scoring, the deliverables/MLflow `*cache*` trap as a GDPVal
concern, and
Stirrup-agent deploy sizing. `recipes/env.example` loses
`GDPVAL_SIF_DIR`,
`GDPVAL_MAX_TURNS` and `TAVILY_API_KEY`, and gains a documented
`NEMO_EVALUATOR_TRUST_UNLISTED_TASKS` (required by every gym task,
previously only
mentioned in prose).

One correction carried along: the old reference said the bootstrap pins
`ray==2.49.2`, but `example_mrcr.yaml` actually pins to whatever ray
version the
image already carries. `references/gym.md` now describes what the
template does.

### Usage

Not an API change. The skill-facing surface that moved:

```text
scripts/nel-gdpval.sh      -> scripts/nel-gym.sh      (NEL_GDPVAL_* -> NEL_GYM_*)
references/gym-gdpval.md   -> references/gym.md
tests/test_nel_gdpval.py   -> tests/test_nel_gym.py
```

### Testing

- `pytest plugins/modelopt/skills/evaluation/tests/` — 1 passed. The
launcher test
was renamed rather than deleted: it is the only coverage for the pinned
launcher,
  which MRCR still depends on.
- `pre-commit run --files <changed>` — clean. `sync-claude-skills` fails
in my
working copy because `.claude/skills/` is a read-only harness mount
there; it is
unrelated to this change (it trips on `speculative-decoding`) and was
skipped for
  the commit.
- Grepped the repo for `gdpval` (case-insensitive): the only remaining
hits are the
  deliberate "no longer supported" notes in `SKILL.md` and
`references/quantization-benchmarks.md`. No dangling pointers to the
deleted
  files, and no other skill referenced GDPVal.
- Checked `tests/evals.json` for both `evaluation` and `day0-release`:
no eval case
expected a GDPVal companion config, so no expectations needed updating.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — anyone with a saved GDPVal
config keeps
  it, but the skill no longer generates one and the SIF helper is gone.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — existing launcher test
renamed and kept passing.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under **Deprecations**.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

A follow-up is needed in the modelopt-internal repo:
`modelopttools:eval-config`
Step 3c is the GDPVal SIF / comparison-mode conversion checklist, and
Step 3d names
the gym image. Step 3c is now dead and the pointers this skill used to
make into it
are gone.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Evaluation Updates**
* Added standalone NeMo Gym support for MRCR tasks with pinned launcher
and Gym configuration requirements.
* Added shared guidance for Gym setup, validation, execution, and
recovery.
* AA requests now generate only `aa/` tasks and report per-task scores.

* **Removed Support**
* Removed GDPVal evaluation recipes, documentation, task guidance,
Apptainer helper tooling, and AA-suite inclusion.
* Updated default quantized-checkpoint validation recommendations to
exclude GDPVal.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:09:15 -07:00
Chenjie LuoandClaude Opus 5 b16356776b Build the GGML IQ packing kernels as a single CUDA extension (#2462)
### What does this PR do?

Type of change: Code refactoring

`#2448` added the GGML IQ packing kernels as **two** torch extensions,
`modelopt_cuda_ext_iq1_s` and `modelopt_cuda_ext_iq2_xs`. This merges
them into one,
`modelopt_cuda_ext_ggml`.

The existing per-extension split in `extensions.py` exists for reasons
that don't apply
to the IQ formats: `get_cuda_ext` gates on CUDA `>=11` while
`_fp8`/`_mx` gate on
`>=11.8`, and `_mx` needs `--use_fast_math`, which must not reach the
base `tensor_quant`
kernels. `get_cuda_ext_iq1_s` and `get_cuda_ext_iq2_xs` differed in none
of that —
same `>=11.8` gate, same `-O3` flags, same `common.cuh` — so the split
only compiled the
shared header twice, ran nvcc twice, and grew the loader, `__getattr__`,
and
`precompile()` once per format. With IQ2_XXS / IQ3_S / IQ4_NL plausibly
following, that
scales badly.

Changes:

- New `ggml/ggml.cpp` holds both host-side validation wrappers and the
single
`PYBIND11_MODULE`, binding `iq1_s_pack` and `iq2_xs_pack` (previously
each module
exported a bare `pack`). Deletes `ggml/iq1_s.cpp` and `ggml/iq2_xs.cpp`;
the
  validation logic and docstrings carry over unchanged.
- `get_cuda_ext_iq1_s` + `get_cuda_ext_iq2_xs` → `get_cuda_ext_ggml`,
which builds
`ggml.cpp`, `iq1_s.cu`, and `iq2_xs.cu` together. The
retry-on-`raise_if_failed`
  semantics of the old getters are preserved.
- Each format keeps its kernels in its own translation unit, so adding a
format is a new
`.cu` plus one `module.def` — no new extension, loader, or
`precompile()` line.

No caller outside `extensions.py` and its tests referenced the old
getters on `main`, so
nothing else changes. **Note for the follow-up PRs in the `#2448` series
(`#2446`/`#2447`/`#2449`): the codec layer should call
`get_cuda_ext_ggml().iq1_s_pack(...)` / `.iq2_xs_pack(...)` instead of
`get_cuda_ext_iq1_s().pack(...)` / `get_cuda_ext_iq2_xs().pack(...)`.**

### Usage

```python
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml

ext = get_cuda_ext_ggml(raise_if_failed=True)
iq1_s_payload = ext.iq1_s_pack(weight, iq1s_grid)            # uint8 [numel / 256, 50]
iq2_xs_payload = ext.iq2_xs_pack(weight, iq2xs_grid, scales) # uint8 [numel / 256, 74]
```

### Testing

Ran on a single H200 NVL (TRT-LLM `1.3.0rc27.dev202609170000`
container), building the
merged extension from scratch:

- `pytest tests/gpu/_extensions/test_torch_extensions.py` — **24
passed** (6:44). This
is the full existing IQ suite (zero-block layout, encode, dtype
rejection,
row-straddling rejection, invalid/negative-zero scales, byte-exact dtype
equivalence,
and the brute-force optimality round-trip) reparametrized onto the
merged module,
  plus the untouched `modelopt_cuda_ext` / `_fp8` / `_mx` load tests.
- Verified `precompile()` loads all four extensions and that the merged
module exports
  exactly `iq1_s_pack` and `iq2_xs_pack` with the expected arities.
- Off-GPU: compiled the three sources directly and linked them into one
`.so` to confirm
no duplicate-symbol collisions between the two `.cu` translation units.
- `pre-commit run --files ...` passes on all changed files (ruff, mypy,
clang-format,
  bandit, license headers).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the removed getters were
added in `#2448`
(merged today, unreleased) and have no callers outside this file's own
tests.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
code or dependencies; the moved wrappers keep their original
attribution.
- Did you write any new necessary tests?: ✅ — existing coverage
reparametrized onto the merged module; no behavior change to test.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal refactor of an unreleased, not-yet-wired-up API.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Follow-up to #2448. Merge before the remaining PRs in that series
(#2446, #2447, #2449)
land, so the codec layer is written against `get_cuda_ext_ggml` and no
rename is needed
afterwards.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
  - Added IQ1_S packing support through the GGML CUDA extension.
  - Added a unified GGML extension loader for IQ1_S and IQ2_XS packing.
- Improved extension loading reliability when a cached extension is
unavailable.

- **Changes**
  - Renamed the IQ2_XS packing binding from `pack` to `iq2_xs_pack`.
- Consolidated IQ1_S and IQ2_XS extension access under the shared GGML
loader.
- Updated GPU validation and coverage to use the unified extension
interface.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 20:46:41 +00:00
Chenjie Luo 542012d4d5 Update CODEOWNERS (#2463)
Add modelopt-torch-kernels-codeowners



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added code ownership coverage for the `modelopt/torch/kernels`
directory.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-09-17 20:08:12 +00:00
Chenjie LuoandClaude Opus 5 655f94c207 Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)
### What does this PR do?

Type of change: documentation

`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.

This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.

It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.

**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.

### Usage

No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:

```python
from modelopt.recipe import load_recipe

cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```

### Testing

Docs-only change; no code paths touched, so no new or updated tests.

- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
  including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
  and the existing `CHANGELOG.rst` entry, so the documented defaults
  (`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
  NVFP4-input-quantizer-only scope match the code.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->

### Additional Information

The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 10:56:23 -07:00
Chenjie LuoandClaude Opus 5 b80e164472 Carry a checkpoint's ModelOpt PTQ run into its eval's MLflow run (#2407)
### What does this PR do?

Type of change: documentation (agent skill)

Follow-up to #2374, which made a tracked `hf_ptq.py --mlflow` run leave
`.experiment.json` in the checkpoint it writes, naming the MLflow run
that quantized it. Nothing on the eval side read that file, so an eval
of a quantized checkpoint recorded no link back to the quantization that
produced it.

The `evaluation` skill now reads it. **Step 3** `cat`s the file as soon
as `checkpoint_path` is known; **Step 4** carries it into
`export.mlflow`:

| `.experiment.json` field | goes to |
| --- | --- |
| `experiment_name` | `export.mlflow.experiment_name`, verbatim |
| `run_name` / `run_id` / `run_url` | tags `modelopt_run_name` /
`modelopt_run_id` / `modelopt_run_url` |
| `tracking_uri`, `experiment_id` | deliberately unmapped |

The `modelopt_` prefix keeps them from reading as the eval's own run.
Values are quoted, or an all-digit `run_id` (or a `run_name` like
`20260910`) is YAML-coerced to an int or a date.

**`tracking_uri` is deliberately not inherited**, and the consequence is
documented rather than implied. Evals go to whatever
`$MLFLOW_TRACKING_URI` names. When the PTQ tracked to a different server
— the usual case, since `modelopttools:eval-config` points evals at
`mlflow.frontier-evals` while `hf_ptq --mlflow` typically writes to
`mlflow-modelopt` — the inherited name creates a *same-named, empty*
experiment on the eval server, and `modelopt_run_url` is the only route
back to the PTQ run. Evals of one checkpoint still group under a stable
name. Both cases are spelled out in Step 4 so nobody goes looking for
the PTQ run beside the eval.

Also updated: `references/nel-next.md` (nel-next configures MLflow
through its own `export_config.mlflow`) and `accessing-mlflow` (the
`tags.modelopt_run_id` query that closes the loop — without it the tag
is write-only).

### Usage

```bash
cat "$CHECKPOINT_PATH"/.experiment.json   # absent → name the experiment as usual
```

```yaml
export:
  mlflow:
    tracking_uri: ${oc.env:MLFLOW_TRACKING_URI}        # NOT the file's
    experiment_name: alice/hf_ptq/Qwen3.8-27B-NVFP4    # verbatim from .experiment.json
    tags:
      modelopt_run_name: '20260910-175422'
      modelopt_run_id: '7bec239a3a154970b062f3024a5ff20e'
      modelopt_run_url: 'https://<modelopt-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e'
```

Finding every eval of a checkpoint a given PTQ run produced:

```python
MLflow:query_runs(experiment_id, "tags.modelopt_run_id = '<ptq_run_id>'")
```

### Testing

Docs-only, so it was tested by having an agent follow the new wording
end to end and checking what NEL actually submitted.

A GPQA Diamond config (single repeat) was generated against an NVFP4
checkpoint carrying a `.experiment.json`, dry-run clean, then submitted
on a SLURM cluster. Verified in the `export_config.yml` heredoc **inside
the submitted `run.sub`** — the file the export job actually consumes,
not just the source YAML:

```yaml
experiment_name: chenjiel/hf_ptq/Qwen3.8-27B-NVFP4        # inherited verbatim
tags:
  modelopt_run_name: 20260910-000000-synthetic-test-fixture
  modelopt_run_id: 00000000fake0000fake0000fake0000
  modelopt_run_url: https://<modelopt-mlflow-server>/#/experiments/999/runs/...
tracking_uri: https://<eval-mlflow-server>/    # env var, not the file's
```

This confirmed the open question behind the change: NEL accepts
arbitrary `modelopt_*` keys in `export.mlflow.tags`. The
`.experiment.json` used was a synthetic fixture, since the checkpoints
predate #2374.

Repeated on a **second cluster with a second checkpoint** (different
account, QoS-based scheduler, FP8 rather than NVFP4) and carried through
to completion, so the result is confirmed as *stored by MLflow* rather
than merely submitted. Canary scored `gpqa pass@1 symbolic_correct:
90.0` (card: 88.01) with `no_answer: 0.0`, then the auto-export wrote:

```
experiment        chenjiel/hf_ptq/Qwen3.8-27B-FP8   (id 2106)
run               eval-<invocation>-ns_gpqa          FINISHED
modelopt_run_name 20260910-000000-synthetic-test-fixture
modelopt_run_id   00000000fake0000fake0000fake0000
modelopt_run_url  https://<modelopt-mlflow-server>/#/experiments/999/runs/...
```

Experiment 2106 did not exist on the eval server beforehand — the export
created it. That is the shadow-experiment case this PR documents,
observed rather than predicted.

The same exercise corrected the first draft of the wording: it had
claimed the rule makes a quantization and its evals "sit together on one
server", which is false whenever the two servers differ. Confirmed by
API — `chenjiel/hf_ptq/Qwen3.8-27B-NVFP4` returns
`RESOURCE_DOES_NOT_EXIST` on `mlflow.frontier-evals`. Rewritten as
described above.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — additive guidance; the
absent-file path is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — skill documentation;
validated by a real submission, above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — agent-skill guidance, not a library/example feature.
- Did you get Claude approval on this PR?: ❌

### Additional Information

Follow-up to #2374.

Three adjacent problems were found while validating and deliberately
**left out of scope** — happy to split them into their own PR:

1. `recipes/examples/example_eval.yaml` puts `sbatch_comment` under a
top-level `cluster:` key, but the executor reads
`cfg.execution.sbatch_comment` (`executors/slurm/executor.py:688`) and
SKILL.md Step 1 says the same. As shipped the template silently drops
the idle-GPU reaper exemption, so long evals get reaped.
2. Step 4's `execution.gres` bullet names the wrong key for
`internal/slurm/<cluster>` configs, which supply `gpus_per_node` and no
`gres`; following it literally emits a redundant flag.
3. Step 3's `max_new_tokens` rules can contradict each other — "take the
card's highest" can equal `max_model_len`, leaving no room for the
prompt.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Clarified MLflow evaluation configuration requirements, including
literal experiment and sampling values, CPU partition settings, and
ModelOpt provenance.
- Documented validation and fallback behavior for unresolved `${...}`
placeholders in provenance values.
- Added guidance for querying cross-server evaluations by run ID and
handling same-named local experiments.
- Clarified that provenance files are optional, independent of
quantization detection, and do not control deployment flags.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-11 13:08:56 -07:00
Chenjie LuoandClaude Opus 5 d69e93a72b Record the MLflow run that produced a checkpoint in .experiment.json (#2374)
### What does this PR do?

Type of change: new feature

A tracked `hf_ptq` run already tags itself with the checkpoint it writes
(`checkpoint_path`), so a run can be followed to its output. The reverse
was missing: given a checkpoint on disk, there was no way to find the
run that quantized it without searching the tracking server by path.

A tracked run now writes `.experiment.json` into `--export_path` naming
the experiment, the MLflow run id and the run URL, and uploads the same
bytes as the `experiment.json` artifact so a downloaded artifact set is
self-describing. `MlflowRunLogger` gains a `run_info` property carrying
that identity, with the tracking URI credential-masked the way `run_url`
already was.

Two deliberate behaviours:

- **Written from a `finally`**, so a run that crashes after export still
leaves the pointer behind.
- **Skipped when the export directory is absent** — a run that exported
nothing has nowhere to put it, and creating the directory would suggest
a checkpoint that does not exist. The artifact is still uploaded in that
case, so a failed run is traceable from the server side.

A failed local write warns and continues rather than failing the job,
consistent with the rest of the MLflow path. Only the main rank writes,
since the logger is inert on other ranks.

### Usage

```bash
python hf_ptq.py --pyt_ckpt_path Qwen/Qwen3.5-0.8B --qformat fp8 \
    --export_path /tmp/qwen35-fp8 --mlflow https://<your-mlflow-server>
```

```console
$ cat /tmp/qwen35-fp8/.experiment.json
{
  "tracking_uri": "https://<your-mlflow-server>",
  "experiment_name": "alice/hf_ptq/Qwen3.5-0.8B-fp8",
  "experiment_id": "36",
  "run_id": "7bec239a3a154970b062f3024a5ff20e",
  "run_name": "20260910-175422",
  "run_url": "https://<your-mlflow-server>/#/experiments/36/runs/7bec239a3a154970b062f3024a5ff20e"
}
```

```python
# checkpoint -> run
import json, mlflow
info = json.load(open("/tmp/qwen35-fp8/.experiment.json"))
mlflow.set_tracking_uri(info["tracking_uri"])
run = mlflow.get_run(info["run_id"])
```

### Testing

**Unit** — `tests/unit/torch/utils/test_mlflow.py` (61 passed):
`run_info` contents before/after the run opens, the defaulted run name
being reported rather than left blank, and credential masking of the
tracking URI.

**Example** — `tests/examples/hf_ptq/test_hf_ptq_args.py` (27 passed):
the file landing in the checkpoint and on the server with identical
content, the failed-run path, the no-export path, and untracked runs
writing nothing.

**Real runs**, 1x H200, `Qwen3.5-0.8B` FP8 PTQ,
`tensorrt-llm/release:1.3.0rc26`:

- Against a local MLflow server — checkpoint copy and uploaded artifact
byte-identical; artifacts on the run were `command.txt`,
`experiment.json`, `logs/hf_ptq.log`, `summary/quant_summary.txt`,
`version.txt`.
- Against the internal `mlflow-modelopt` server (experiment
`chenjiel/hf_ptq/Qwen3.5-0.8B-fp8`, run
`7bec239a3a154970b062f3024a5ff20e`) — same result, confirming artifact
upload against a real backend. Reading `.experiment.json` back and
calling `mlflow.get_run(run_id)` resolved to `FINISHED` with
`checkpoint_path` pointing at the export directory.
- Crash path exercised for real when a first attempt died on a gated
calibration dataset: no export directory created, `experiment.json`
still uploaded, run closed `FAILED`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — new entry under `*Misc*` in the open 0.48.0 section, matching where
the MLflow entries sit in 0.47.0.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Exported checkpoints now record experiment and run traceability
metadata in `.experiment.json`.
* Checkpoint metadata is uploaded with opened MLflow runs, including
runs where export fails.
* Active MLflow run details—including identifiers, resolved run name,
URL, and tracking server—are available with credentials redacted.
* **Bug Fixes**
* Improved handling of failed, untracked, and pre-existing exports to
prevent inherited metadata pointers.
* **Documentation**
* Updated MLflow integration guidance and changelog information for
checkpoint metadata.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 22:24:37 +00:00
Chenjie LuoandClaude Opus 5 a74054ab2b Let callers add MLflow tags to a fakequant serve's run (#2364)
### What does this PR do?

Type of change: new feature

The quantization run records what this library can see — the model, the
checkpoint, the vLLM and ModelOpt versions — but nothing about the
harness that launched it. A downstream tool that wants its own revision,
a sweep id, or a ticket number on the run has no way to put it there
today:

- `_run_tags()` returns a fixed dict
- `quant_config` (which becomes the run's params) is a hardcoded set of
`QUANT_*` variables
- MLflow itself has no environment variable for arbitrary tags

`MODELOPT_MLFLOW_EXTRA_TAGS` takes comma-separated `key=value` pairs and
merges them into the run's tags.

Two details worth a reviewer's attention:

**It joins `MLFLOW_ENV_VARS`.** A Ray-backed serve receives only the
variables named there, and the tracker runs in the rank-0 worker —
omitting it would make the feature silently do nothing under Ray.

**Caller tags are merged first**, so the library's own keys (`tool`,
`model`, `checkpoint_path`, `vllm_version`) are written over them and
keep describing the run truthfully whatever a caller sends.

`key=value` rather than JSON, learned from a live run: the variable
reaches the worker through a shell `export VAR="..."`, and JSON's own
double quotes terminate that quoting —

```
export MODELOPT_MLFLOW_EXTRA_TAGS_732b_DEPLOYMENT="{"internal_version": "4d8c"}"
```

arrived as `{`. A quote-free format survives verbatim and needs no
`json` import or exception handling. Splitting on the first `=` keeps
values that contain one, such as a URL with a query string.

### Usage

```bash
export MODELOPT_MLFLOW_EXTRA_TAGS="modelopt_internal_version=49fa29d5,sweep=kv-study"
python3 vllm_serve_fakequant.py "$MODEL" --mlflow https://your-mlflow-server/ ...
```

### Testing

Unit-level, over the helper: unset and empty variable, one and several
pairs, surrounding whitespace, an empty value, an entry with no `=`, a
trailing comma, and a value containing `=`. None raise; malformed
entries warn and are skipped.

End to end on a real fakequant serve (Nemotron-3-Nano-30B-A3B BF16,
`NVFP4_DEFAULT_CFG`, TP=8, Ray executor, vLLM 0.15, SLURM):

```
modelopt_internal_version  '49fa29d5'
modelopt_version           '0.47.0rc0.post32+gd38ed5ead'
git_sha                    'd38ed5ead'
quant_cfg                  'NVFP4_DEFAULT_CFG'
```

The tag was written by the `RayWorkerWrapper` process, which exercises
the whole path — env var → shell export → `--container-env` → raylet →
Ray actor → `_run_tags` — and confirms the `MLFLOW_ENV_VARS` entry is
doing its job. Also verified that the emitted payload survives a shell
export round-trip unchanged.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!--- Additive; with the
variable unset the tags are exactly as before. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <!--- Verified manually as
above; there is no existing test module for vllm_mlflow_utils. Happy to
add one if you would like it. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Small additive feature in an example; tell me if it warrants an
entry. -->
- Did you get Claude approval on this PR?: ❌

### Additional Information

Consumed by Model-Optimizer-Internal MR !141/!147, which sets the
variable so a fakequant eval records the same harness commit on both its
quantization run and its evaluation-score run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 20:29:16 +00:00
Chenjie LuoandClaude Opus 5 19de0075cb Forward kv_cache_free_gpu_memory_fraction to the lm_eval TensorRT-LLM engine (NVBug 6701763) (#2300)
### What does this PR do?

Type of change: Bug fix

`scripts/huggingface_example.sh --kv_cache_free_gpu_memory_fraction` has
no effect on the `lm_eval` task: the value is parsed by `parser.sh`,
printed, and then dropped.

lm-eval's built-in `trtllm` backend
(`lm_eval.models.trtllm_causallms.TRTLLM.__init__`, which this example
switched to in #2066) accepts `**kwargs`, but builds
`KvCacheConfig(enable_block_reuse=False)` and passes `LLM(...)` a fixed
set of keys — `kwargs` is never merged in. So an extra `--model_args`
entry is accepted by the CLI and silently discarded, and the KV cache is
sized from TensorRT-LLM's default `free_gpu_memory_fraction=0.9`. There
is no way to fix this from the caller: `--model_args` only yields
scalars, so a `KvCacheConfig` object cannot be passed in either.

On a GH200 that means ~119.6 GiB of KV cache (`119.55 / 0.9 ≈ 132.8 GiB
free`), leaving 87.8 MiB free, and `prompt_logprobs` deserialization
then OOMs asking for 2.82 GiB.

`examples/llm_eval/lm_eval_trtllm.py` already exists to patch this
backend (its `_parse_logprobs` misaligns TensorRT-LLM's
`prompt_logprobs` by one). It now also injects the fraction into the
`KvCacheConfig` the backend builds, defaulting to 0.8 — the same default
`parser.sh` declares, and below TensorRT-LLM's 0.9.
`huggingface_example.sh` passes the parsed value through in
`--model_args`.

Scoped deliberately to the `lm_eval` path: the `quant` smoke test and
`mmlu` go through `modelopt.deploy.llm.LLM` (0.7, hardcoded) and
`simple_eval`/`livecodebench` through `trtllm-serve` (0.9); those are
left as they are.

### Usage

```bash
# Via the example script (parser.sh default 0.8)
scripts/huggingface_example.sh --model $HF_PATH --quant fp8 --tp 1 \
    --tasks quant,lm_eval --lm_eval_tasks mmlu --lm_eval_limit 50 \
    --kv_cache_free_gpu_memory_fraction 0.5
```

```bash
# Standalone, via lm-eval's --model_args
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<ckpt>,tokenizer=<tok>,max_input_len=4096,kv_cache_free_gpu_memory_fraction=0.5 \
    --tasks mmlu --batch_size 8
```

### Testing

- `pytest tests/examples/llm_eval/test_lm_eval_trtllm.py` — 21 passed
(lm-eval 0.4.12, no GPU).
- The new tests instantiate the **real** upstream `TRTLLM.__init__`
through `create_from_arg_obj`, with `tensorrt_llm` and the tokenizer
stubbed, and assert the engine receives
`KvCacheConfig(enable_block_reuse=False, free_gpu_memory_fraction=0.5)`;
that an unset key still yields 0.8 rather than 0.9; and that the patch
does not outlive the constructor. Reverting the fix fails 3 of them.
- Tripwire test asserts upstream still neither declares nor forwards the
argument, so this shim gets deleted rather than silently kept once
lm-eval fixes it.
- `pre-commit run --files <changed>` clean (ruff, mypy, bandit,
markdownlint); `bash -n` on the modified script.
- Not run: the GPU end-to-end
`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8`, which
exercises `lm_eval` through the modified script — no GPU in this
environment.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the `lm_eval` KV cache goes
from TensorRT-LLM's 0.9 to 0.8, which is strictly more conservative;
`parser.sh`'s declared default is unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

NVBug 6701763. The 0.9 default on this path arrived with #2066 and was
documented as a known limitation in `examples/llm_eval/README.md` ("the
KV cache uses 90% of free GPU memory rather than 70%"); that note is
replaced by the working knob.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed the TensorRT-LLM evaluation workflow so
`kv_cache_free_gpu_memory_fraction` is correctly passed to the backend.
- The setting now defaults to `0.8`, providing more predictable GPU
memory allocation for KV-cache usage.

- **Documentation**
- Updated the TensorRT-LLM evaluation example and usage guidance to
describe the KV-cache memory setting and its default behavior.
- Updated the Hugging Face example to pass the configured KV-cache
memory fraction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-08 13:06:51 -07:00
Chenjie Luo 19ce447d62 [skill] evaluation: mandate 8 SciCode runs and report the mean (#2327)
### What does this PR do?

Type of change: documentation

SciCode's repeat count was pinned to `num_repeats: 1` in #1945, which
dropped its
effective sample count from the original `8` (set when the recipe was
written in
#1561) to `1`. #2254 later documented that a single run cannot gate on —
scored
single-shot at `temperature 1.0`, a paired comparison moved **3.92 pp
and changed
sign** once repeated, and the day-0 skill's own table records a
DeepSeek-V4-Pro
drop reading 2.96 pp (`REGRESSION`) at 1 run versus **-0.96 pp**
(`PASS`) at 8.
But #2254 left the remedy as a judgment call — *"pool until the standard
error is
below the threshold, or report the task `INDETERMINATE`"* — so a one-run
SciCode
number was still reportable.

This restores the original avg-of-8 statistics as a hard requirement,
without
reintroducing the in-run repeats that #1945 removed for a reason
(repeating
inside one run multiplies exposure to code-execution sandbox errors).
The task
keeps `num_repeats: 1` per run; the repeat budget of 8 is spent as **8
independent submissions, reported as their mean** — 8 per side for a
comparison,
`INDETERMINATE` below 8.

Changes:

- **`recipes/tasks/aa/scicode.md`** — mandatory *8 runs* section: how to
submit
the 8 inside a multi-task AA config, fresh-run requirement (no `run.sub`
replay off a warm cache), duplicate-score check, per-run validation
before
  averaging, and mean + `stdev/sqrt(8)` + run-count reporting.
- **`compare-results` / `day0-release`** — 8 runs per side is a floor,
not a
  variance-dependent choice.
- **`references/quantization-benchmarks.md`** — the repeat-count table
still
  listed SciCode at `num_repeats: 8`, stale since #1945.
- **`evaluation/SKILL.md`** — AA rule points at the 8 submissions; the
walltime
section distinguishes them from the forbidden practice of splitting a
heavy
  task across configs to dodge the 4h cap.
- **`evaluation/tests/evals.json`** — behavioral eval case for the rule.

### Usage

```bash
# Full AA suite (= SciCode run 1), then SciCode alone for runs 2-8.
nel run --config <cfg>.yaml
for _ in $(seq 7); do nel run --config <cfg>.yaml -t ns_scicode; done
# Report the mean of scicode_pass_at_1_avg-of-1_subtask_accuracy over the 8 runs,
# plus stdev/sqrt(8) and the run count.
```

### Testing

- `pre-commit run --files <changed files>` — all hooks pass
(`markdownlint-cli2`,
  `check json`, symlink sync), no hook-applied modifications.
- `json.load` on `evaluation/tests/evals.json` parses; 4 cases, existing
three
  unchanged (append-only diff, original formatting preserved).
- Cross-checked every remaining SciCode reference in the plugin
  (`grep -rn -i scicode plugins/modelopt/`) so no doc still claims
  `num_repeats: 8` or a variance-dependent pool size.

**Ran the protocol end-to-end** on Qwen3.8-27B-FP8 (gcp-nrt, 8xB200,
`temperature 1.0`),
8 independent `nel run` submissions, all `COMPLETED`, each scoring the
full 80 problems /
338 subtasks. MLflow experiment 2017:

| passed/338 | 168 | 167 | 166 | 164 | 163 | 161 | 157 | 155 |
|---|---|---|---|---|---|---|---|---|
| score | 49.70 | 49.41 | 49.11 | 48.52 | 48.22 | 47.63 | 46.45 | 45.86
|

**Result 48.11, stdev 1.39, stderr 0.49.** Pooled 1301/2704 subtasks
reproduces 48.1139 exactly.

This is the evidence for the change: **the spread is 3.85 pp and the
worst single run sits
2.26 pp from the mean**, so against a 1% gate any one of these eight,
reported alone, would
have been defensible and wrong. Previously this PR rested on the
variance figures inherited
from #2254; it now rests on a measured pool.

Also validated by that campaign: an agent given only "run a full SciCode
eval, follow the
evaluation skill" — with no mention of the protocol — read the recipe,
submitted 8 runs and
reported the mean with a standard error. And the corrected score key is
what made the numbers
harvestable at all; the `avg-of-1` name returns nothing.

Two operational findings from the run, one of which is folded into the
recipe:

- An MLflow export job timed out (`nel-export-ns_scicode.0`, elapsed
`00:30:11` against a
`00:30:00` limit) because all 8 exports hit the CPU partition at once
and each reinstalls
the launcher. Caused by the fan-out this PR mandates, so the recipe now
warns about it and
  says to re-submit `export.sbatch` rather than treat the run as lost.
- Out of scope, filed for separate fix: `.claude/agents` is a 0-byte
read-only placeholder,
so the `monitor` skill's instruction to create a session registry under
it fails with
  `Not a directory`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — added
`scicode-eight-run-average` to `evaluation/tests/evals.json`
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent-skill guidance, not a user-facing library change
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Restores the sampling behavior of #1561 while keeping the sandbox-load
fix from
#1945 and honoring the variance evidence from #2254.

Note: `origin/feature/puzzletron_v2` still carries the pre-#1945 version
of this
file (`num_repeats: 8`, single submission) and rewrites it to source the
repeat
count from `examples/llm_eval/task_contracts.yaml`. If that branch lands
it will
need to be reconciled with this protocol.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Clarified SciCode evaluation requirements: eight independent, valid
runs per comparison side with one repeat each.
- Standardized score reporting using per-run metrics, mean, standard
error, and run count.
- Added provenance checks and replacement runs for invalid or replayed
results, including timeout recovery.
  - Fewer than eight valid runs are reported as **INDETERMINATE**.
- Updated benchmarking guidance to distinguish separate submissions from
repeated runs.

- **Tests**
- Added coverage for eight-run averaging, validation, replacement runs,
and **INDETERMINATE** outcomes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-09-04 19:08:53 +00:00
Chenjie LuoandClaude Opus 5 73d7784223 docs(eval-skill): add NVFP4 model-card sampling reference (#2224)
### What does this PR do?

Type of change: documentation

Adds
`plugins/modelopt/skills/evaluation/references/nvfp4-modelcard-sampling.md`
— the published `temperature` / `top_p` / max generation length for the
**2026 NVFP4 checkpoints under
[huggingface.co/nvidia](https://huggingface.co/nvidia/models) that
disclose them** — and points `SKILL.md` Step 3 and
`model-card-research.md` at it.

**Why.** Config generation currently re-derives sampling params per run
from one model card, which is slow and silently wrong in two ways:

1. **Quickstart boilerplate reads as an eval setting.** Most NVIDIA
NVFP4 cards paste a TensorRT-LLM snippet containing
`SamplingParams(temperature=0.8, top_p=0.95)`. That string is
byte-identical across `Llama-3.1-8B`, `Llama-3.3-70B`, `Llama-4-Scout`,
`Phi-4-reasoning-plus`, `Phi-4-multimodal`, `Qwen2.5-VL-7B` and six
`Qwen3-*` repos. It is template text, not what the accuracy table was
measured with — but it is the most prominent `temperature=` in the card.
2. **A generic fallback beats an available answer.** When a card is
silent, Step 3 falls back to 65536/16384 even where a sibling in the
same family publishes an exact value.

**Scope — 25 rows.** All 69 NVFP4 checkpoints in the org were read. A
row exists only for a 2026 **target** checkpoint whose card discloses
usable settings, so absence means "read the card", not "not yet
checked". Excluded by construction: pre-2026 releases, cards that
publish nothing, and — via the curation regex `-NVFP4(-V\d+|-QAD)?$` —
`-DSpark`/`-DFlash` speculative-decoding variants (verified against the
target, so identical accuracy; they would only duplicate the base row),
`-Eagle3` draft heads, and `-MLPerf-Inference-Closed-*` snapshots.

| provenance | rows | meaning |
| --- | --- | --- |
| `eval` | 20 | card ties the values to its accuracy table —
authoritative |
| `rec` | 5 | recommended inference sampling, not tied to the eval |

A `—` in one value column means that field specifically is unpublished —
four rows give `temperature`/`top_p` but state no generation cap
(`Mistral-Medium-3.5`, `Nemotron-3.5-Lightning`,
`Nemotron-3-Super-120B`, `Nemotron-Labs-3-Elastic-30B`). Per-task
exceptions are recorded where cards state them: GLM-5.2 GPQA Diamond
`100000` vs `64000`; Qwen3.5-397B-V2 τ²-Bench Telecom `128000`; Qwen3.6
SciCode `temperature=0.6`; Kimi-K3 uncapped for Terminal-Bench.

**Posture: the card is the source of truth; the table is a reference,
not a constraint.** It is there to confirm a value you read, fill a gap
when the card is silent, and catch a misreading — never to override what
a card states. Listed and in agreement → proceed; listed and different →
the card wins, re-read, surface the discrepancy. The `max_num_tokens`
column records the card's *headline* cap, so Step 3's existing
take-the-highest rule still governs when a card names more than one.
Per-task `temperature`/`top_p` in the notes is precedent rather than
mandate — engineers do tune sampling per benchmark — so the guidance is
to follow the card and escalate only on a regime change (greedy vs
sampled), not a nudge (`0.95` vs `1.0`).

### Usage

Not an API change; the reference is consumed by the `evaluation` skill
when generating a NEL config.

```yaml
# nvidia/GLM-5.2-NVFP4 -> references/nvfp4-modelcard-sampling.md (provenance: eval)
nemo_evaluator_config:
  config:
    params:
      max_new_tokens: 100000  # card: GPQA Diamond 100000, others 64000 -> take the highest
      temperature: 1.0
      top_p: 0.95
```

### Testing

**Coverage — enumeration cross-verified two ways.** Rather than paging
`https://huggingface.co/nvidia/models?p=N` by hand, the candidate list
came from the HF API (`author=nvidia&limit=1000` → 917 repos, one page),
then verified against the website pagination: 918 unique model links
across `p=0..31`, and the 68 NVFP4-named repos were **identical in
both** (`api-only: []`, `web-only: []`). The committed curation regex
was then re-run against that full set and confirmed to select every one
of the 25 table rows plus 7 further 2026 target checkpoints whose cards
disclose nothing — i.e. the documented filter reproduces the table.

**Accuracy — every `eval` row diffed against its source sentence.** A
script re-parsed the table and printed each row beside the matching card
line. To be precise about what that covers: the diff was run when the
table was larger, and all 28 machine-checkable *Benchmarked with* rows
matched exactly (`Qwen3.6-35B-A3B` states its settings on the second
line of a blockquote and was checked by hand). Every edit since has been
a row **removal** — the scope cut to 2026 and the spec-decode removal —
each verified by re-parsing the table before and after and confirming no
surviving row's `temperature`/`top_p`/`max_num_tokens`/`provenance`
changed. Extraction was restricted to *"Benchmarked with…"* / *"…were
evaluated with…"* / *"We evaluate the model using…"* sentences,
"Recommended Sampling" rows, and accuracy-table footnotes — never the
quickstart snippets.

**Two extraction traps found while building it**, both now in the file's
refresh recipe as a second mandatory grep: some cards publish the cap
**only** as a footnote beneath the accuracy table (`*Max OSL for evals
can be as high as 64K`), and DeepSeek states sampling in its `## Input:`
usage block rather than a *Benchmarked with* sentence. Neither is
reachable from the obvious grep.

**Hygiene.** `pre-commit run --files <the 3 files>` passes, including
`markdownlint-cli2` and the `sync .claude/skills/ symlinks` hook; the
reference resolves through both `.claude/skills/evaluation/references/`
and `.agents/skills/evaluation/references/`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- documentation-only;
verification described under Testing -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- agent-skill documentation; no user-facing API or behavior
change -->
- Did you get Claude approval on this PR?: ✅ <!-- reviewed; findings
addressed or answered below -->

### Additional Information

Data collected 2026-08-20; the file is an explicitly dated snapshot and
says to trust the card for anything newer. It ends with a refresh recipe
(API enumeration → `-NVFP4(-V\d+|-QAD)?$` filter → `createdAt >=
2026-01-01` → card fetch → the grep phrasings that carry eval settings)
so the table can be regenerated as new checkpoints ship — roughly four
NVFP4 target checkpoints per month over 2026 so far.

Review findings addressed: the mandatory-match framing was softened to
reference-only, per-task sampling reframed as precedent with a
human-escalation trigger, `mkdir -p cards` added to the recipe, the HF
token no longer interpolated into a `curl` argument, and explicit
uncapped generation (`max_new_tokens: null`) distinguished from an
unpublished cap. The provenance question on the Nemotron-3.5-Lightning
rows is answered in a review reply — the differing labels were correct
(the base card says "Recommended Sampling", the sibling cards say
"Benchmarked with"), and those sibling rows have since been removed as
spec-decode duplicates.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 16:51:59 +00:00
Chenjie LuoandClaude Opus 5 d32c2c2a56 docs(eval-skill): add MRCR (NeMo Gym) benchmark (#2192)
### What does this PR do?

Type of change: documentation (agent skill)

Adds **MRCR** — OpenAI's Multi-Round Co-reference Resolution, a
long-context
retrieval benchmark — to the `evaluation` skill as a standalone NeMo Gym
task,
derived from the reviewed `nemotron_nano_v35_nvfp4_mrcr_gym` golden;
regroups the
gym tasks/examples under one `gym/` dir; and fixes several latent bugs
in the
shared gym command block that were found by running the benchmark
end-to-end.

MRCR tasks are long multi-turn conversations containing N near-identical
"needle"
responses; the model must reproduce the Nth verbatim behind a random
prefix.
Grading is deterministic (`SequenceMatcher.ratio()`, gated on the
prefix). Unlike
GDPVal it uses the `simple_agent`: no SIF, no judge, no Tavily —
`HF_TOKEN` is the
only secret, and the cost is context length (up to 1M tokens), not agent
turns.
**It is not an AA benchmark** and is never generated for an "AA"
request.

### Layout

| File | |
| --- | --- |
| `recipes/tasks/gym/mrcr.md` | new recipe — variants, 1M serving
envelope, canary, score extraction |
| `recipes/examples/gym/example_mrcr.yaml` | new self-contained SLURM +
vLLM config |
| `SKILL.md` | MRCR branch + gym index table |
| `references/quantization-benchmarks.md` | table row + comparability
notes |

`examples/gym_{gdpval,mrcr}/` → `examples/gym/example_<task>.yaml` and
`tasks/aa_gym/gdpval.md` → `tasks/gym/gdpval.md`, all tracked as git
renames.
`aa_gym` encoded "in the AA suite" in the *path*; since GDPVal is AA and
MRCR is
not, membership is now stated explicitly in both recipe headers and an
"In AA
suite?" column in the SKILL.md index. `gym/` groups by harness, not
suite.

### Bugs fixed (found by running it, not by reading it)

The first three are in the **shared** gym command block, so they also
affect
`example_gdpval.yaml` — GDPVal was broken independently of MRCR.

1. **Invalid OmegaConf interpolation.** A literal `${...}` inside a
comment in the
gym `command:` block is parsed as an interpolation and rejected, so
*every*
   `nel run --dry-run` of either gym template died with
   `hydra.errors.ConfigCompositionException`.
2. **Hardcoded `ray==2.49.2`** injected into each sub-server's
requirements. Against
an image carrying `ray[default]==2.55.1` this makes `uv` unsatisfiable
and the
gym resources server exits at startup. Now derived from the image at
runtime.
3. **`--max-num-seqs` sized against the wrong topology.** The template
shipped DP1
→ 4 replicas at 64 concurrent while citing a golden that is TP2×DP2 → 8
replicas
at 32. Following it ran double the reference's per-replica load, which
on
   1M-token prompts is what decides whether the KV cache fits.
4. **Score extraction pointed at an empty map.** `results.yml` →
`groups.nemo_gym.metrics` holds only `key_metrics/mean/*` telemetry; the
scores
are in `artifacts/evaluator_rollouts_aggregate_metrics.json` →
`[0].agent_metrics`.
Also documents that `pass@1/accuracy` is already 0-100 while
`mean/reward` is the
same number as a 0-1 fraction, and records `mean/prefix_matched ≈ 0.55`
as the
   healthy calibration.
5. **The template shipped a container its own bootstrap rejects.** After
adding the
   hard-fail on an unpinned Gym, `container:` still defaulted to public
`nemo-gym:26.05` — the image the docs say has a non-git `/opt/Gym`. Now
`???`,
so `--dry-run`'s mandatory-value check catches it instead of the job
dying at
   startup.

### Review feedback

All items from @meenchen and CodeRabbit addressed or answered inline.
Two worth
surfacing here:

- **`process_reasoning_traces` vs `use_reasoning`** — verified against
`nemo_evaluator/adapters/adapter_config.py`: both exist, `use_reasoning`
is the
**deprecated** one and they are bidirectionally aliased. Documented
rather than
  switched to the deprecated name.
- **`tiktoken` / `transformers` left unpinned** — deliberate. The n3
prepare path
uses `transformers.AutoTokenizer` to decide which samples exceed the
cap, so a
bump can shift dataset membership; but pinning would diverge from the
golden's
`pre_cmd` and therefore from the run that produced the reference number.
Risk is
now documented under "Deferred, know the risk" instead of being
implicit.

Topology, `gres` and TP/DP were aligned to the **existing** sibling
templates
rather than to a new convention: `gres` stays a comment (already the
convention in
`example_eval.yaml` and `example_gdpval.yaml`), TP is a concrete `1`
like both
siblings, and `num_nodes`/`num_instances`/`--max-num-seqs` are guidance
rather than
baked values. `--max-num-seqs` sizing follows **AA-LCR**, since MRCR is
the same
KV-bound problem at ~1M tokens vs LCR's ~120K: the formula gives a
ceiling, not a
target, because oversubscribing causes preemption and recomputing a
1M-token
prefill makes the run slower.

### Usage

```bash
cp plugins/modelopt/skills/evaluation/recipes/examples/gym/example_mrcr.yaml mrcr.yaml
# fill checkpoint_path / served_model_name / container / SLURM ??? values
export NEMO_EVALUATOR_TRUST_PRE_CMD=1 NEMO_EVALUATOR_TRUST_UNLISTED_TASKS=1
nel run --config mrcr.yaml --dry-run && nel run --config mrcr.yaml
```

### Testing

Docs/config-only; no library code touched. `pre-commit` clean on every
commit.

- Both gym YAMLs parse; asserted the folded-scalar rule (no `#` inside
`>-`), that
every string survives `OmegaConf.create` (bug 1), and that the variant
is
  identical in `data_prep_params` and `collect_rollout_params`.
- After the rename: zero stale `aa_gym`/`gym_gdpval`/`gym_mrcr` refs and
every
`recipes/**` path referenced across both skill trees resolves on disk —
this
  caught `env.example`, which an extension-filtered grep missed. The
`nemo_gym_gdpval_stirrup_agent` metric names contain the substring
`gym_gdpval`
  and were deliberately left untouched.
- **Ran the benchmark end-to-end.** A full run on gcp-nrt (B200)
completed
2363/2363 rollouts and scored; that run produced bugs 2 and 4 and the
corrected
metric paths. A second attempt on aws-cmh was preempted and later
cancelled.
`--dry-run` now passes (it did not before bug 1 was fixed), and the bare
template
  correctly fails validation on unresolved mandatory values.
- Regenerated the config from scratch with a fresh agent against the
fixed
templates as a regression check; its dry-run passed and its findings
drove the
  topology/`gres`/TP-DP alignment above.

Model-specific scores are deliberately **not** recorded in the recipe —
the
reference there stays the golden's BF16 Nano 3.5 shape, since quoting a
different
model's number beside it invites exactly the false comparison the recipe
warns
about.

### Before your PR is "*Ready for review*"

- Backward compatible?: ✅ — additive plus a doc-tree rename; all
referencing files
  updated and verified. The template fixes change only broken behaviour.
- Copied code / new PIP dependency: N/A.
- New tests: N/A — agent-skill documentation.
- Changelog: N/A — skill/docs change; consistent with prior skill-only
commits.
- Claude approval: ❌ not yet — will trigger `/claude review`.

### Additional Information

Upstream drift flagged to reviewers (**no change made to that repo**):
in
`nvidia-eval-factory-benchmarking`, `configs/benchmarks/mrcr/bench.yaml`
is still
old-style `ng_*` while `configs/models/nemotron_nano_v35/gym.yaml` moved
to the
`gym eval` CLI on 2026-07-29 — RULER migrated, MRCR did not, so the
layers no
longer compose. This template follows `ng_*`, which is what MRCR's own
bench.yaml
specifies and what the sign-off run executed.

NVIDIA-internal companion (container specifics kept out of this public
tree):
Model-Optimizer-Internal MR !116.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 22:37:30 +00:00
Chenjie Luo 195a5af430 Add modelopt-agents-codeowners as owner of plugins/ (#2212)
### What does this PR do?

Type of change: Documentation

`plugins/` holds the installable ModelOpt agent plugin — the canonical
skill tree
at `plugins/modelopt/skills/`, which `.agents/skills` and
`.claude/skills` expose
via relative symlinks. It currently has no CODEOWNERS rule, so it falls
through to
the default `* @NVIDIA/modelopt-devs` owner. This routes it to
`@NVIDIA/modelopt-agents-codeowners` instead.

```
# Agent plugin (skills, agent config)
/plugins @NVIDIA/modelopt-agents-codeowners
```

The path is root-anchored (`/plugins`) to match the style of the
`/examples` block
above it.

### Usage

N/A — no API or flag change.

### Testing

No test surface. Verified the rule resolves as intended against the
existing file:
`/plugins` is the last (and only) match for paths under `plugins/`, so
it wins over
the default `*` rule; it does not overlap any other entry.

Effective ownership of `plugins/` on this branch is
`@NVIDIA/modelopt-agents-codeowners`. GitHub validates the team handle
when the PR
is opened — if the team does not exist or lacks write access to this
repo, the
Settings > CODEOWNERS errors page will flag the line.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- Repo-internal review routing, not user-facing. -->
- Did you get Claude approval on this PR?: N/A

### Additional Information

Note for reviewers: shared agent config and scripts under `.agents/`
(everything
that is not a symlink into `plugins/`) are still owned by the default
`@NVIDIA/modelopt-devs` rule. Happy to add `/.agents` here too if the
agents team
should own that as well.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Added repository ownership rules for the plugins directory and its
agent configuration files.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-08-18 15:04:46 -07:00
Chenjie Luo b96841db3e Add optional MLflow tracking to the vLLM fake-quant server (#2120)
### What does this PR do?

Type of change: new feature

Wires `examples/vllm_serve/vllm_serve_fakequant.py` up to
`modelopt.torch.utils.mlflow` via `--mlflow <tracking-uri>`, the same
way #2023 did for `hf_ptq.py`, so a fake-quant serve records **what it
actually quantized** and an evaluation of that endpoint can be traced
back to a recipe. Without the flag, behavior is unchanged — every hook
is gated on it.

Three design points worth review:

1. **The run is recorded in the vLLM worker, not the launcher.**
`vllm_serve_fakequant.py` is the API-server frontend; the engine and its
workers are separate processes whose stdout it never sees, so a run
opened there would capture none of the calibration. The launcher instead
only settles the tracking configuration — validating the URI, naming the
experiment, recording the command the user actually typed — and
publishes it through the environment, which is how every other setting
in this example (`QUANT_CFG`, `RECIPE_PATH`, …) already reaches the
workers. Global rank 0 opens the run, so a TP-8 serve produces one run.

2. **The run covers load-through-warm-up, not the server's lifetime.**
It opens *before the weights load*, so an unreachable server or a
missing token fails in seconds rather than after a load and a full
calibration, and it closes `FINISHED` once the model is quantized and
warmed up. A run that stayed open for the serving lifetime would never
close cleanly on SIGTERM.

3. **`recipe/quant_cfg.yaml` is only written on the preset path.** With
`RECIPE_PATH`, `get_quant_config` returns the recipe's `quantize`
section unchanged and `resolved_recipe.yaml` already carries it. With
`QUANT_CFG`/`KV_QUANT_CFG` it is the *only* record of what ran: the
params carry the preset names, while the config reaching `mtq.quantize`
is those two deep-copied, merged, and — for an MLA model — extended at
runtime with `*kv_c_bmm_quantizer` / `*k_pe_bmm_quantizer` by inspecting
the loaded model.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The launcher's invocation, copy-pasteable, credentials
masked |
| `version.txt` | The ModelOpt version that ran |
| `recipe/resolved_recipe.yaml` | `RECIPE_PATH` with its `$import`s
expanded |
| `recipe/quant_cfg.yaml` | Merged `QUANT_CFG`/`KV_QUANT_CFG` + MLA
fixup (preset path only) |
| `logs/<script>.log` | The rank-0 worker's stdout/stderr, including a
crash traceback |
| `summary/quant_summary.txt` | The per-quantizer summary |

Plus the quantization *and* serving settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` / `vllm_version`
tags. The `checkpoint_path` tag matches the one `hf_ptq.py` sets, so a
checkpoint's PTQ run and every serve of it join up.

Two small library additions, both consumed by the new example module:

- `command_text(argv=None)` — records another process's invocation,
since a spawned worker's own `sys.argv` is vLLM plumbing rather than
anything a user typed.
- `MlflowRunLogger.log_text()` — uploads a value settled midway through
a run, so a crash during calibration still keeps the config that caused
it.

The example `Dockerfile` installs the `mlflow` extra; the client remains
optional and is imported only once tracking is enabled.

### Usage

```bash
RECIPE_PATH=<recipe.yaml> python vllm_serve_fakequant.py <model_path> -tp 8 \
  --host 0.0.0.0 --port 8000 \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] tracking to https://<your-mlflow-server>, experiment $USER/vllm_serve_fakequant/<model>-<recipe>
(Worker_TP0) [mlflow] run: https://<your-mlflow-server>/#/experiments/19/runs/1c6679448f25...
```

`--mlflow-experiment` / `--mlflow-run-name` override the defaults.
`$MLFLOW_TRACKING_URI` enables tracking on its own and is best-effort;
an explicit `--mlflow` overrides it and fails loudly.

> This is the **quantization** tracking server. It is unrelated to any
server an evaluation harness exports its scores to — NeMo Evaluator
Launcher has its own `export.mlflow.tracking_uri`. The README calls this
out.

### Testing

**Unit — 87 passing**
(`tests/examples/vllm_serve/test_vllm_mlflow_utils.py`, 33 new;
`tests/unit/torch/utils/test_mlflow.py`, +5). `vllm_mlflow_utils`
deliberately imports no vLLM, so the whole launcher→worker handover is
covered without a GPU, a server, or the mlflow client.

**End to end on aws-cmh** (4× GB300, `simple_evals.gpqa_diamond`,
Nemotron-3.5-Lightning-30B-A3B-BF16 fake-quantized with
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`): run `FINISHED` in 261.5 s,
opened by `Worker_TP0` only, all artifacts present and verified by
content — `command.txt` held the launcher's invocation rather than the
worker's spawn argv, and `resolved_recipe.yaml` was 6797 B against 1845
B of source. 104 quantizers enabled (92 NVFP4 dynamic block-16 expert
weight/input with calibrated amax, 12 FP8 KV bmm). The eval then ran to
completion against the served endpoint, 22/22 requests HTTP 200.

Two bugs the hardware run caught, both fixed here with regression tests:

- `--mlflow_run_name` was rejected. vLLM's
`FlexibleArgumentParser.parse_args` rewrites **every** `--foo_bar` to
`--foo-bar` before matching, so a flag registered only under the
underscored spelling is unreachable from its CLI. Both spellings are now
registered. A unit test on a plain `ArgumentParser` could not have
caught this.
- `recipe/quant_cfg.yaml` uploaded a Python `repr` blob under a `.yaml`
name: a recipe's `quantize` is a `QuantizeConfig`, `yaml.safe_dump`
raises `RepresenterError` on it, and the old JSON fallback stringified
the object. `_dump_yaml` now unwraps pydantic via
`model_dump(mode="json")` and raises otherwise, with the caller
downgrading that to a warning so a bad config cannot take down a serve.

**Known coverage gap:** the preset (`QUANT_CFG`/`KV_QUANT_CFG`) path —
the only one that now writes `recipe/quant_cfg.yaml` — is covered by
unit test but has not been exercised on hardware; the canary used
`RECIPE_PATH`. Likewise the case where `$MLFLOW_TRACKING_URI` is present
*inside* the deployment container and `--mlflow` overrides it is
unit-tested only: NeMo Evaluator Launcher forwards only declared env
vars, so the eval server's URI never entered the container in the
canary.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency. Uses the existing optional `nvidia-modelopt[mlflow]` extra
(`mlflow-skinny`, Apache-2.0) added in #2023; the example `Dockerfile`
now installs it. No code copied from other sources.
- Did you write any new necessary tests?: ✅ — 38 new tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 Misc.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run.

### Additional Information

Follows #2023, which added `MlflowRunLogger` and the `hf_ptq.py`
integration.

Note for anyone tracking from an OCI cluster:
`mlflow-modelopt.nvidia.com` is unreachable from oci-nrt and oci-hsg.
TCP 443 completes and the connection is then reset on the first
application byte, regardless of SNI or protocol, one RTT away — the PDX
PaaS ingress appears to apply a source-IP policy, and the OCI clusters
egress from Oracle-owned addresses (`155.248.190.0`, `168.110.199.1`)
rather than NVIDIA's. gcp-nrt, aws-cmh and cw-dfw all reach it. This is
an infrastructure matter, not a property of this change, but it
determines where the feature is usable today.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added optional MLflow tracking for vLLM fake-quantization serving
runs.
* Records serving, quantization, worker, and invocation metadata,
including configuration and summary artifacts.
* Supports tracking URI, credentials, environment, and command-line
configuration.
  * Added command and text artifact logging for active MLflow runs.
* **Documentation**
* Documented setup, configuration, recorded artifacts, lifecycle, and
fallback behavior.
  * Updated the example container to include MLflow support.
* **Tests**
* Added comprehensive coverage for tracking configuration, logging,
failures, and disabled tracking.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-08-13 00:00:21 +00:00
06480249fd [skill] evaluation: align nel-next TB2.1/SWE-bench with golden toolchain (#2063)
### What does this PR do?

Type of change: Documentation / tooling (agent skill)

Terminal-Bench 2.1 configs generated from the `evaluation` skill had
drifted from the
canonical eval-factory config
(`configs/benchmarks/terminal-bench-2.1/bench.yaml`). The
**scoring contract already matched** the reference configs exactly —
playbook, `repeats: 8`,
`timeout_strategy: max`, `run_timeout: 7200`, `llm_kwargs.timeout:
3600`, concurrency. What
had drifted was the toolchain and a few proxy-level defaults.

The drift was in the **skill**, not in individual configs: a config
generated fresh from the
skill reproduced every stale value, so patching configs alone would not
have held.

- **`nel-next.sh` installs from the public upstream repo**
(`github.com/NVIDIA-NeMo/Evaluator`, default branch → `0.4.0`) instead
of PyPI. PyPI
`nemo-evaluator` tops out at `0.3.0` and cannot reach the 0.4.x
toolchain the reference runs use.
`NEL_NEXT_SPEC` becomes the PyPI escape hatch and now takes precedence
when explicitly set;
`NEL_NEXT_ORIGIN` stays overridable from `.env` so internal mirrors stay
out of this repo.
- **`eval_image`**: document the pinned `0.5.0.1-harbor` (single source
of truth:
`configs/shared/nel_next_containers.yaml`) rather than `0.3.1.1-harbor`
as a floor.
- **`proxy.request_timeout` 1800 → 3600** — must be `>=` the solver's
`llm_kwargs.timeout`,
else the proxy truncates long agent turns the harness is still awaiting.
- **`drop_params`**: add `max_input_tokens_per_task`, `no_rebuild` —
sent by the 0.5.x harbor
  eval image; vLLM returns 400 unless stripped.
- **`exclude_patterns`**: add `model_traffic.jsonl` so captured request
bodies stay in the run
  dir and never reach MLflow.
- **`http_pairs_dump`** interceptor (last in chain) for HTTP
diagnostics.
- **Sharding documented**: `max_concurrent`/`sandbox.concurrency` are
*per shard*, so
`shards: N` multiplies both serving capacity and live sandboxes (`N x
concurrency`).
- **`.gitignore`**: broaden `.env` / `.env-*` to `.env*` so secret
backups such as
  `.env.bak-tb21` cannot be staged.

**This does not move the benchmark.** The TB2.1 task set is pinned by a
vendored registry
override that has not changed since 2026-06-03, and both
`0.3.1.1-harbor` and
`0.5.0.1-harbor` score 89 samples — so the image bump is a toolchain fix
and scores stay
comparable across it.

### Usage

```bash
set -a && source .env && set +a
.agents/scripts/nel-next.sh --version        # 0.4.0 (public upstream build)
.agents/scripts/nel-next.sh eval run <tb21-config>.yaml --dry-run
```

### Testing

**End-to-end parity runs.** Configs generated from these changes were
run to completion on
aws-cmh (4x GB300 aarch64, sm_103) with an NVFP4 checkpoint of
Qwen3.6-35B-A3B, and compared
against the reference BF16 results for the same base model:

| Benchmark | Reference (BF16) | This run (NVFP4) | Delta |
|---|---|---|---|
| Terminal-Bench 2.1 | 0.4438 `[0.4215, 0.4661]` | **0.4438** `[0.4233,
0.4644]` | **0.0000** |
| SWE-bench Verified | 0.7012 `[0.6920, 0.7104]` | **0.7040** `[0.6950,
0.7130]` | **+0.0028 (+0.40%)** |

Both `pass@1` over the full task sets (89 x r8 = 712 trials; 500 x r5 =
2500 trials), with
overlapping 95% CIs in both cases.

This exercises the changes in this PR directly: the `0.5.0.1-harbor`
eval image, the
`proxy.request_timeout >= llm_kwargs.timeout` fix, the four-param
`drop_params`, the
per-benchmark interceptor ordering, and the SWE-bench `system_message` /
`instruction_template` requirements. A wrong tool-call parser, sampling
preset, or
instruction template would each have moved these numbers well outside
the intervals.

Also verified:

- `nel-next.sh --version` -> `0.4.0`, built from
`Evaluator.git@4d081325` — the commit the
scored runs above actually executed, confirmed from their recorded uv
archive
(`direct_url.json` -> `commit_id
4d081325170aababd0c8f27c58bed31a81ce82ac`). This is the SHA
  `NEL_NEXT_REF` now defaults to.
- Both configs pass `eval run --dry-run` with no schema errors. These
schemas are
`extra="forbid"`, so `http_pairs_dump` and the new `drop_params` entries
would hard-fail if
  unsupported by the pinned toolchain.
- `0.5.0.1-harbor` resolves in the generated `nel_eval.sbatch`; the tag
is multi-arch
(linux/amd64 + linux/arm64) and imported successfully on aarch64 compute
nodes.
- `pre-commit run --files <changed>` - all hooks pass, no file
modifications.
- Values cross-checked against the canonical benchmark configs in
`dl/JoC/competitive_evaluation/nvidia-eval-factory-benchmarking` @
`main`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `NEL_NEXT_SPEC` restores the
previous PyPI install.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — agent-skill docs/config;
validated via `--dry-run` + pre-commit.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no library API change.
- Did you get Claude approval on this PR?: ❌ — pending.

### Additional Information

Personal run configs under `.agents/skills/evaluation/runs/` are
deliberately **not** included:
they carry internal cluster hostnames, lustre paths, account names and
an AWS account id, which
do not belong in this public repo.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated evaluation guidance for `nemo-evaluator` 0.4.x and the
`0.5.0.1-harbor` image.
* Added canonical benchmark configurations, longer proxy timeouts,
request filtering, diagnostics, MLflow exclusions, sharding, capacity
planning, concurrency, replay behavior, and deployment verification
guidance.
  * Documented separate virtual-environment setup and version tracking.

* **Configuration**
* Evaluation tooling now defaults to a pinned Git-based installation
with configurable sources and version references.
  * Environment files with any `.env`-prefixed name are now ignored.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 09:28:32 -07:00
Chenjie LuoandClaude Opus 5 9220fac053 [NVBug: 6563509] Drop Phi-3-vision / Phi-4-multimodal PTQ support (#2115)
### What does this PR do?

Type of change: Deprecation

Resolves [NVBug 6563509](https://nvbugspro.nvidia.com/bug/6563509),
where
`hf_ptq.py` on Phi-4-multimodal-instruct died with
`RuntimeError: Tensor.item() cannot be called on meta tensors`.

The crash is real but not fixable on our side, and it is not the reason
the model
is unusable. Phi-4-multimodal's bundled remote code predates
Transformers v5 and
does not load on **any** version in our supported range
(`transformers>=4.57,<5.15`):

| Blocker | Where |
|---|---|
| `peft.get_peft_model` reads `prepare_inputs_for_generation`, gone
since transformers 4.52 dropped `GenerationMixin` from `PreTrainedModel`
| `modeling_phi4mm.py:1959` |
| `_tied_weights_keys` declared as a list; Transformers 5.x calls
`.keys()` on it in `post_init` | `modeling_phi4mm.py:1937` |
| `int(torch.tensor(...))` in `__init__`, which cannot run on a meta
device — the reported crash | `speech_conformer_encoder.py:1435` |

The model card pins `transformers==4.48.2` / `peft==0.13.2`, so there is
no
overlap with our floor and nothing on our side can bridge it. The model
is
therefore dropped rather than worked around.

**Phi-3-vision is dropped alongside it because it is the older,
superseded model
in the same family** — with its successor unsupportable there is no
reason to
keep carrying the predecessor. This is a product-scope call, not a
separate
compatibility finding: Phi-3-vision shares the list-valued
`_tied_weights_keys`
defect (`modeling_phi3_v.py:1214`) and so is likewise broken on
Transformers 5.x,
but it does **not** hit the `peft` blocker, and it was not re-verified
on 4.57.
Per the 0.46 changelog we have already bumped the floor to 4.57 and
noted that
"Transformers 4.x support will be dropped in a future release", so any
remaining
window closes on its own. Same reasoning already applied to VILA / NVILA
in this
release.

**Removed**

- the support-matrix row in `examples/hf_ptq/README.md`
- `"Phi4MMForCausalLM": "phi4mm"` from `MODEL_NAME_TO_TYPE`
- the multimodal-detection heuristics that only ever matched these two —
`vision_lora`, `audio_processor`, `embd_layer.image_embd_layer`, and the
  `phi4mm` model-type check — in both `is_multimodal_model` and
  `_is_multimodal_config`
- the `Phi3Image` / `PhiImage` exclusions in `is_embedding`
- the phi4mm input-mode warning in `hf_ptq.py`
- `modelopt_recipes/huggingface/phi4mm/` and its references in
  `modelopt_recipes/ptq.md`

**Not changed:** the device-map sizing path (meta-device skeleton,
`infer_auto_device_map`, and the `--gpu_max_mem_percentage` cap) keeps
its
original behavior. That cap is wanted exactly where it already fires —
when the
model is already offloading to CPU, where it costs little and the
headroom is
required. With the affected checkpoints removed, there is no supported
model
that trips the meta-device build, so there is nothing to work around
here.

Text-only **Phi-3/Phi-4** and **Phi-3.5-MoE** are natively supported by
transformers and are untouched.

### Testing

On H200, `nvcr.io/nvidia/tensorrt-llm/release` (torch 2.12, transformers
5.5.4),
against the real checkpoint:

- **Version matrix** (vanilla transformers, no modelopt) — Phi-4-MM
loads at
4.48.2 / 4.49.0 / 4.50.0 / 4.51.3 and fails at 4.53.3 / 4.56.2 / 4.57.1
  (`AttributeError: 'Phi4MMModel' object has no attribute
'prepare_inputs_for_generation'`) and at 5.5.4 (meta-init, then
tied-keys).
  This is what establishes that no supported version works.
- `tests/examples/hf_ptq/test_example_utils.py` — 28 passed.
- **Sweep**: `tests/examples/hf_ptq` + `tests/unit/torch/export` —
failure set
identical to the pre-change tree (GPU/model-dependent `test_vlm_ptq`,
plus
`test_quant_aware_conversion` scoped-mapping tests), so none are
introduced
  here.
- `pre-commit` clean on all changed files, including recipe validation.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — PTQ for Phi-3-vision and
  Phi-4-multimodal is removed, along with the `huggingface/phi4mm/ptq/*`
recipes. Phi-4-multimodal is already unloadable on every supported
transformers
version, so no working workflow regresses; Phi-3-vision is a deliberate
scope
  removal as its superseded predecessor.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — this is a deletion; the
existing
`test_get_model_*` / `test_resolve_init_config_*` tests are unchanged
and still
  pass.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Two related references were left in place deliberately; say the word and
I'll
fold them in:

- `tests/examples/hf_ptq/test_deploy.py` still deploys the
already-published
`nvidia/Phi-4-multimodal-instruct-{NVFP4,FP8}` checkpoints. Those
artifacts
exist and serve fine; this PR only removes the ability to *produce*
them.
- `examples/torch_onnx/README.md` still lists Phi-4-multimodal-instruct.
That
is a separate ONNX pipeline that does not go through `get_model()` and
was not
  tested here.

Earlier revisions of this branch also reworked the device-map sizing so
the
meta-tensor crash could not occur. That was reverted in 701180ed6: the
guard is
correct as written, and every alternative either changed behavior for
models that
fit today or moved the guard somewhere it does not belong, for a crash
that only
ever affected the checkpoints this PR removes.

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 23:22:23 +00:00
Chenjie LuoandClaude Opus 5 9b8caf623a Use lm-eval 0.4.12's built-in trtllm backend, deprecate lm_eval_tensorrt_llm.py (#2066)
### What does this PR do?

Type of change: documentation / example update (with a behaviour fix)

lm-evaluation-harness **0.4.12** is the first release that ships a
TensorRT-LLM backend
(`lm_eval.models.trtllm_causallms`, registered as `trtllm`) — it is
absent in 0.4.10 and
0.4.11. This example no longer maintains its own, so:

- Pin `lm_eval[api,ifeval]>=0.4.12,<0.5` (the 0.5.0.dev line drops the
file) and bump
  `lm_eval_hf.py`'s version guard to match.
- **Delete** `examples/llm_eval/lm_eval_tensorrt_llm.py` (the `trt-llm`
model). Replace
`python lm_eval_tensorrt_llm.py --model trt-llm --model_args
tokenizer=<tok>,checkpoint_dir=<ckpt>`
with `python lm_eval_trtllm.py --model trtllm --model_args
model=<ckpt>,tokenizer=<tok>`.
- Add `examples/llm_eval/lm_eval_trtllm.py`, whose entire content is one
corrected
`_parse_logprobs` plus `cli_evaluate()` (see below). `lm_eval_hf.py`
stays HF-only.
- `examples/hf_ptq/scripts/huggingface_example.sh` and the docs use the
upstream backend.
`parser.sh` gains `--input` (`BUILD_MAX_INPUT_LEN`, default 4096) — it
already *echoed*
that variable but never parsed or defaulted it, so it printed empty on
every run.

#### Why `lm_eval_trtllm.py` exists: an upstream off-by-one

TensorRT-LLM aligns `prompt_logprobs` to the *next* token.
`executor/base_worker.py`:

```python
# Pass prompt_token_ids with an offset of 1 for correct mapping to the context logits
prompt_token_ids = generation_result._generation_request.prompt_token_ids[1:] + first_generation_token
```

So entry `i` is the distribution that predicted `tokens[i + 1]`, and
`_topk_logprobs`
appends that token's id when it is not in the top-k. lm-eval's
`_parse_logprobs` instead
reads `prompt_logprobs[i][tokens[i]]` and applies its own shift on top,
which raises
`KeyError` on the **first request of every loglikelihood task**
(hellaswag, mmlu, arc, ...):

```
File ".../lm_eval/models/trtllm_causallms.py", line 324, in _parse_logprobs
    current_token_logprob = prompt_logprob[tokens[i]]
KeyError: 6503
```

Probed against TRT-LLM 1.3.0rc23 with a 14-token prompt for
`prompt_logprobs` 0, 1 and 2:
`tokens[i]` is missing at **every** position, `tokens[i+1]` is present
at every position.
Only `generate_until` tasks work unpatched. **This wants an upstream
issue against
EleutherAI/lm-evaluation-harness.**

The override also fails loudly rather than quietly: it checks
`prompt_logprobs` covers
every prompt token and raises on a missing token, instead of skipping
the term and
silently inflating the reported accuracy.

#### Defaults that must be set explicitly

`TRTLLM.__init__` accepts `**kwargs` but forwards only a fixed set to
the **TensorRT-LLM
`LLM` API**, so extra `--model_args` aimed at the engine are silently
dropped. (lm-eval's
own named parameters — `max_gen_toks`, `batch_size`, `truncation_side`,
... — are honored
normally.) Two engine defaults are unsafe for few-shot eval:

- `tensor_parallel_size` defaults to **1** (the deleted wrapper used
every visible GPU).
- `max_input_len` defaults to **2048**, and longer prompts are silently
left-truncated —
  5-shot MMLU/gsm8k prompts exceed that.

### Usage

```bash
python lm_eval_trtllm.py --model trtllm \
    --model_args model=<quantized checkpoint dir>,tokenizer=<HF model folder>,tensor_parallel_size=<tp>,max_batch_size=<bs>,max_input_len=4096,max_output_len=512 \
    --tasks hellaswag,gsm8k \
    --batch_size <bs>
```

Flat arguments (no `run` subcommand) are what 0.4.12's
`HarnessCLI.parse_args` inserts
`run` for automatically (`_cli/harness.py:48-51`); this is the exact
command form used for
the results below.

### Testing

**Unit** — `tests/examples/llm_eval/test_lm_eval_trtllm.py`, no GPU and
no `tensorrt_llm`
install: stubs the response object and pins the `i-1` alignment, the
`rank != 1` →
`is_greedy` rule, the `ctxlen=0` edge, and both `RuntimeError` paths.
Mutation-checked —
dropping the `-1` shift is caught by 5/5 cases, ignoring `ctxlen` by
4/5. A sixth test is a
**tripwire**: it asserts lm-eval's own implementation is still
misaligned, so a future
0.4.x that fixes the bug fails the test and says to delete this file
rather than being
silently re-broken by the override.

**End to end** — `nvidia/Qwen3.5-122B-A10B-NVFP4` (NVFP4 MoE, 256
experts) on **4x B300**,
TRT-LLM 1.3.0rc23, lm-eval 0.4.12, `--limit 32`:

| run | hellaswag acc | hellaswag acc_norm | gsm8k flexible | gsm8k
strict |
|---|---|---|---|---|
| deleted impl (`trt-llm`), tp=4 | 0.7188 | 0.7812 | 0.8438 | 0.7812 |
| `lm_eval_trtllm.py`, tp=1 | 0.7188 | 0.7812 | 0.8438 | 0.8125 |
| `lm_eval_trtllm.py`, tp=2 | 0.7188 | 0.7812 | 0.9062 | 0.8125 |
| `lm_eval_trtllm.py`, tp=4 | 0.7188 | 0.7812 | 0.8750 | 0.8438 |

- hellaswag (the loglikelihood path this PR fixes) is **identical at
every tp and identical
to the deleted implementation** — the alignment fix is exact, not
approximate.
- gsm8k varies by 1–2 samples out of 32 (generation path: upstream uses
native `stop=`
sequences and per-request `SamplingParams`; the old wrapper used
beam-search-of-1 with
  post-hoc string truncation).
- Without the override, every hellaswag run above dies with the
`KeyError`.
- Re-verified at tp=4 after the code moved out of `lm_eval_hf.py` into
`lm_eval_trtllm.py`.

Note: NVFP4 fused-MoE has no CUTLASS tactic on Hopper (`No supported MoE
GEMM tactic
remains after replacing unsupported NO_SMEM epilogues.`), so this had to
be validated on
Blackwell.

### Feature parity notes

Gained from upstream: `loglikelihood_rolling` (was
`NotImplementedError`), pipeline
parallelism, `add_bos_token` auto-detection, prompt truncation,
per-request sampling params,
`prompt_logprobs` instead of full-vocab context logits (much lower
memory), thinking-tag
handling, `batch_size=auto`.

Not reachable through the upstream backend (were set by
`modelopt.deploy.llm.LLM`):
`enable_attention_dp` for MoE, `CudaGraphConfig`,
`enable_chunked_prefill`,
`moe_expert_parallel_size=1`, and `free_gpu_memory_fraction=0.7` with a
capped
`kv_cache.max_tokens` — upstream uses the TRT-LLM default 0.9 (observed
allocating 218 GiB
of paged KV cache on B300), so OOM risk is higher on smaller GPUs. This
is documented in
`examples/llm_eval/README.md`, and `huggingface_example.sh` honours a
preset `LM_EVAL_TP`
so users can lower the tensor-parallel size without editing the script.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `lm_eval_tensorrt_llm.py` is
removed and the CLI changes (`--model trt-llm` → `trtllm`,
`checkpoint_dir=` → `model=`). Migration command is in the README and
CHANGELOG.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependency; existing `lm_eval` pin tightened.
- Did you write any new necessary tests?: ✅ —
`tests/examples/llm_eval/test_lm_eval_trtllm.py` (6 cases, no GPU).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.47 *Deprecations*.
- Did you get Claude approval on this PR?: ✅ — reviewed, feedback
addressed in `dcedd37b4` and `622b97c26`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added TensorRT-LLM evaluation through lm-evaluation-harness’s `trtllm`
backend.
* Added configurable input/output lengths, batching, tensor parallelism,
and build input length.
* Improved prompt log-probability alignment for more accurate evaluation
results.

* **Documentation**
* Updated evaluation instructions, truncation guidance, backend
limitations, and configuration examples.

* **Deprecations**
  * Removed the legacy TensorRT-LLM evaluation script and entry point.

* **Updates**
  * lm-evaluation-harness now requires versions 0.4.12 through 0.4.x.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 22:36:55 +00:00
Chenjie LuoandClaude Opus 5 77dbeb1872 Add optional MLflow tracking to hf_ptq.py (#2023)
### What does this PR do?

Type of change: new feature

Adds `modelopt.torch.utils.mlflow.MlflowRunLogger`, a reusable helper
for recording a script run on an MLflow tracking server, and wires
`examples/hf_ptq/hf_ptq.py` up to it via `--mlflow <tracking-uri>` so a
PTQ run can be reproduced from its MLflow entry alone. Without the flag,
behavior is unchanged — every hook is gated on it.

The logger lives in the library rather than the example so other scripts
can record runs the same way: it takes a tracking URI, an experiment
name and an explicit `enabled` flag, with params, tags and artifacts
passed in. `hf_ptq.py` supplies only the PTQ-specific pieces (its
params, the resolved recipe, the quantization summaries). `mlflow` is an
optional dependency, imported only once tracking is enabled, so it is
not a new requirement for the library.

The run is opened **before the model loads**, so a bad URI or an
unreachable server fails in seconds rather than after hours of
calibration. The invocation and the recipe are uploaded at that point
too, which keeps a crashed run useful: it is still recorded, with status
`FAILED` and its log attached.

Uploaded artifacts:

| Artifact | Contents |
| --- | --- |
| `command.txt` | The full invocation, copy-pasteable |
| `version.txt` | The ModelOpt version that ran (also a searchable tag)
|
| `recipe/resolved_recipe.yaml` | The `--recipe` with `$import`s
expanded |
| `logs/hf_ptq.log` | Everything the run printed, including a crash
traceback |
| `summary/quant_summary.txt` | Per-quantizer summary (unless
`--no-verbose`) |
| `summary/moe.html` | Per-expert calibration token counts, when the run
produces them |

Plus model / format / calibration settings as searchable params, and
`user` / `hostname` / `modelopt_version` / `git_sha` tags.

Three design points worth review:

1. **The recipe is uploaded resolved, not verbatim.** A recipe may be a
directory or use `$import`s, so the source file is not self-contained.
For
`huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe`
the source is 2,230 B / 58 lines against 7,563 B / 308 lines resolved —
the raw file records under 30% of what actually ran.
2. **`hf_ptq.py` has no logging framework** (bare `print()`), so the log
is produced by teeing stdout/stderr. Handlers that libraries bound to
`sys.stderr` at import time are re-pointed at the tee for the run's
duration and handed back afterwards; without that, `transformers` /
`huggingface_hub` warnings reach the console but never the log. Native
(C-level) output is still not captured — documented in the README.
3. **The recipe upload lives in the caller, not the library.** That
keeps `modelopt.recipe` out of `modelopt.torch.utils`, which would
otherwise risk a `modelopt.torch.utils` → `modelopt.recipe` →
`modelopt.torch.quantization` → `modelopt.torch.utils` import cycle.
4. **MLflow failures never fail the quantization.** Startup validation
is fatal by design (it is before any GPU work); the end-of-run upload is
best-effort.

Only the main rank uploads, so `--use_fsdp2` runs produce a single run.

### Usage

```bash
python hf_ptq.py \
  --pyt_ckpt_path <huggingface_model_card> \
  --recipe general/ptq/nvfp4_default-kv_fp8_cast \
  --export_path <quantized_ckpt_path> \
  --mlflow https://<your-mlflow-server>/
```

```
[mlflow] experiment: $USER/hf_ptq/<checkpoint basename>-<recipe name>
[mlflow] run: https://<your-mlflow-server>/#/experiments/13/runs/c243352e...
```

`--mlflow_experiment` and `--mlflow_run_name` override the defaults
(`$USER/hf_ptq/<basename>-<recipe name or --qformat>`, and the UTC start
time). Passing `--mlflow` with no value uses `$MLFLOW_TRACKING_URI`.
Authentication uses MLflow's own env vars.

### Testing

**Unit** — 51 tests in `tests/unit/torch/utils/test_mlflow.py` for the
library, plus 13 in `tests/examples/hf_ptq/test_hf_ptq_args.py` for the
hf_ptq wiring. CPU-only, no network and no `mlflow` dependency (driven
against a stub module). Covers experiment-name derivation and
sanitization, URI accept/reject, tee pass-through, the pre-bound-handler
redirect, artifact renaming, skipping absent optional outputs, the
disabled path, and `version.txt`. 85 tests pass together with the
existing `test_hf_ptq_args.py` / `test_example_utils.py`.

**Hardware** — real PTQ runs against a live MLflow server:

| Run | Result |
| --- | --- |
| Qwen3-0.6B, NVFP4 PTQ, 1×B200 | `FINISHED`, all artifacts, sane
post-quant generations |
| Qwen3.6-35B-A3B MoE, AutoQuantize
`w4a16_nvfp4_fp8_at_6p0bits-active_moe`, 2×B200 | `FINISHED` in 63 min,
search hit `effective bits: 6.00`; 106 KB log capturing every per-layer
decision, 4.4 MB quant summary |
| Qwen3.6-35B-A3B, plain NVFP4 PTQ, 2×B200 | `FINISHED` |
| Qwen3-0.6B re-run after the library move, 1×H200 | `FINISHED`, all
five artifacts including `version.txt` |
| Run **without** `--mlflow` after the review fixes | exactly 1
`[load_recipe]` line and 0 `[mlflow]` lines, confirming the untracked
path is untouched |
| Two runs sharing one `--export_path`, second crashed early | second
run uploads **no** summary — the first run's 124 KB file on disk is
correctly not attributed to it, and its traceback is in the log |
| Crash mid-run (gated HF dataset) | `FAILED` recorded with log +
traceback attached, summaries correctly absent |
| Malformed URI | Rejected by `argparse` with a `Did you mean
https://…?` hint |
| Unreachable host | Fails in 9.9 s total, before any model load |
| No `--mlflow` | Exit 0, no MLflow output, unchanged export |

**Coverage gap, stated plainly:** `summary/moe.html` is verified only
against a synthetic file (unit test + a real upload). It could not be
produced naturally — `expert_token_count` buffers live on
`_QuantSparseSequentialMoe`, while Qwen3.5/3.6 experts take the fused
`_QuantFusedExperts` path, so no such file is written for these models
regardless of `--moe_calib_experts_ratio`. The uploader's conditional is
correct; the branch simply had no natural input available here.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — new optional flags only; no
`--mlflow` means no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — adds
`mlflow` as an optional extra in `pyproject.toml`
(`nvidia-modelopt[mlflow]`, folded into `all`) and to
`examples/hf_ptq/requirements.txt`. Apache-2.0 (permissive). Imported
lazily, so it is not required to install or import ModelOpt. No code
copied from other sources.
- Did you write any new necessary tests?: ✅ — 29 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.47 New Features.
- Did you get Claude approval on this PR?: ❌ — `/claude review` not yet
run. A self-review was done first and its six findings are fixed in the
third commit (the notable one: gathering the MLflow inputs re-read the
recipe on *every* run, including without `--mlflow`).

### Additional Information

The one deliberate coverage gap is `summary/moe.html`, described under
Testing: no model available here takes the sparse-sequential MoE path
that writes it, so it is covered by unit test and a synthetic upload
rather than a natural one. The uploader treats it as an optional output
and skips it when absent, which is exercised by test.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 11:25:58 +05:30
5e019d882f Add PR review-feedback guidance to AGENTS.md (#2057)
### What does this PR do?

Type of change: documentation

Adds a `## Responding to PR review feedback` section to `AGENTS.md`
(which `CLAUDE.md` symlinks to, so it covers both Claude Code and
Codex). Today the agent instructions stop at "open the PR" — nothing
says what to do with the review comments that come back, so agents
either apply every suggestion reflexively (including stale or wrong bot
findings) or fix things silently and leave reviewers guessing whether
their comment landed.

The new section says to:

- **Triage before acting** — check each comment against the current
code; CODEOWNERS reviewers outweigh bot reviewers (CodeRabbit, Claude),
whose findings are claims to verify rather than instructions. A reviewer
reaffirming after pushback settles it.
- **Pick one outcome per thread** — address in a commit, push back
citing the code that shows the comment is wrong, or postpone as out of
scope, and report which threads got which when asking for push approval.
- **Reply in the thread after pushing** — a sentence on what changed and
where. Those replies ride on the approval to push; pushback and postpone
replies need their own approval, since no commit backs them. The agent
never resolves threads — that stays the reviewer's call.

### Usage

N/A — agent instructions only, no code change.

### Testing

`pre-commit run --files AGENTS.md` passes (markdownlint included).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
  * Added guidance for responding to pull request review feedback.
* Clarified how to validate comments, prioritize reviewers, report
outcomes, and respond to addressed threads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-08-03 22:02:33 -07:00
Chenjie LuoandClaude Opus 5 14b20c0a12 [skill] evaluation: add GDPVal (NeMo Gym Stirrup agent) support (#2039)
### What does this PR do?

Type of change: new feature (agent skill)

Adds GDPVal support to the `evaluation` agent skill. GDPVal is an
agentic AA
benchmark: the NeMo Gym "Stirrup" agent produces office/PDF deliverables
inside a
per-task Apptainer code-exec sandbox, and a judge panel scores them. It
runs on the
0.2.6 launcher as a `nemo_gym` task, but it is **standalone** (one gym
eval per
config) and mechanically unlike the `aa/` nemo-skills tasks, so it gets
its own
branch in the skill rather than being merged into the `aa/` task list.

- `recipes/tasks/aa_gym/gdpval.md` — task recipe: standalone rule,
rubric-vs-comparison
  scoring, canary, score extraction.
- `references/gym-gdpval.md` — the machinery: Apptainer SIF sandbox, the
`_gym_prepare` venv-repair / process-group-reap workaround, deployment
sizing,
scoring modes, the MLflow deliverables trap, canary failure modes, and
the
  SIF ↔ Gym-version rebuild coupling.
- `recipes/examples/gym_gdpval/` — self-contained SLURM + vLLM template
plus the
co-located `_gym_prepare.yaml` Hydra include (it must travel with the
config).
- `scripts/gdpval-sif.sh` — build-if-absent / reuse-if-present Apptainer
SIF helper.
Builds on the target cluster only (never copies across clusters),
flock-guarded and
  atomic, driven by `$GDPVAL_SIF_DIR`.
- `SKILL.md` / `references/quantization-benchmarks.md` — GDPVal is part
of the AA
suite but a different harness, so it is generated as a companion
standalone config.
- `recipes/env.example` — `TAVILY_API_KEY` (agent web search) and
`GDPVAL_SIF_DIR`.

### Usage

```bash
# 1. Set GDPVAL_SIF_DIR in .env, then build the sandbox once on the target cluster
#    (build-if-absent, reuse-if-present):
srun -p cpu -t 01:00:00 --pty .agents/scripts/gdpval-sif.sh

# 2. Copy the whole example dir (the _gym_prepare.yaml include must travel with it),
#    fill in the ??? values, then dry-run -> canary -> full:
nel run --config gym_gdpval/example_gym_gdpval.yaml --dry-run
```

### Testing

Validated end-to-end on an aarch64 GB300 SLURM cluster with an NVFP4 MoE
checkpoint:
deploy → SIF build + sandboxed exec → gym head server → 220 rollouts +
deliverables →
judge scoring, producing a real rubric score with zero judge failures.
Several traps
found during that run are now documented in the reference (silent
unsandboxed
fallback, gym-commit/head-server hang, judge api-key value-vs-name, SIF
↔ Gym version
coupling, `limit_samples` not limiting gym rollouts).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (agent skill documentation
+ helper script; no library code)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (agent skill only, no API change)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

Docs/skill-only change under `.agents/`; no `modelopt/` source is
touched.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added standalone GDPVal evaluation support through NeMo Gym, with
Slurm, vLLM, sandbox, judge, reasoning, sampling, and logging
configuration.
- Added automated Apptainer/Singularity image setup with reuse,
validation, locking, and reliable publishing.
- Added configurable settings for judge services, web search, and shared
image caching.
- Added GDPVal preparation and execution examples, including dry-run,
canary, and full-run guidance.

- **Documentation**
- Added GDPVal setup, troubleshooting, scoring, deployment, and
quantization guidance.
- Clarified separate GDPVal configuration, mandatory thinking mode, and
repeat-count requirements.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 15:10:16 -07:00
Chenjie LuoandClaude Opus 5 7ed91540fa [NVBug: 6538278] Pass trust_remote_code to the TRT-LLM model load in deploy/eval examples (#2056)
### What does this PR do?

Type of change: Bug fix

Fixes [nvbug 6538278](https://nvbugspro.nvidia.com/bug/6538278).

`examples/hf_ptq/run_tensorrt_llm.py` loaded the **tokenizer** with
`args.trust_remote_code` but constructed `LLM()` without it, so it
defaulted to `False`. Deploying any checkpoint that ships custom
modeling code (`auto_map`) — e.g. `Llama-3.3-Nemotron-Super-49B-v1`
(DeciLM) — failed at executor init:

```
Failed to initialize executor on rank 0: The repository ... contains custom code which
must be executed to correctly load the model ... Please pass the argument trust_remote_code=True.
-> ValueError -> Executor worker returned error -> run_tensorrt_llm.py subprocess exit 1
```

The `modelopt.deploy.llm.LLM` wrapper already accepts and forwards
`trust_remote_code` (`modelopt/deploy/llm/generate.py:153`), and
`scripts/huggingface_example.sh` already forwards `--trust_remote_code`
to the script — only the call site dropped it.

The same omission exists in the sibling TRT-LLM deploy paths driven by
the same launcher and the same checkpoints, so they are fixed together:

| File | Fix |
| --- | --- |
| `examples/hf_ptq/run_tensorrt_llm.py` | the reported bug |
| `examples/llm_eval/lm_eval_tensorrt_llm.py` | `LLM()` ignored the
`trust_remote_code` that lm-eval injects into `model_args` |
| `examples/llm_eval/mmlu.py` | `LLM()` ignored it (the tokenizer
already used it) |
| `examples/hf_ptq/scripts/huggingface_example.sh` | the `mmlu` stage
never forwarded the flag at all |

All four propagate the **user-provided** flag; none hardcode
`trust_remote_code=True` (per SECURITY.md).

### Usage

No API change. Existing flag now takes effect on the model load:

```bash
scripts/huggingface_example.sh --model <Llama-3.3-Nemotron-Super-49B-v1> \
    --quant fp8 --tasks quant --trust_remote_code
```

### Testing

- New CPU regression test
`tests/examples/hf_ptq/test_run_tensorrt_llm.py` stubs the
`tensorrt_llm`-dependent import and asserts both the tokenizer **and**
the model load receive the flag, parametrized over `True`/`False`.
Confirmed it fails without the fix (`KeyError: 'trust_remote_code'`) and
passes with it.
- Verified against the real `fire` package that a bare
`--trust_remote_code` maps to `True` in `mmlu.py`'s `**kwargs`, and is
absent (defaulting to `False`) when not passed.
- Simulated the launcher's argument assembly both ways: with
`TRUST_REMOTE_CODE=false` the emitted commands are byte-identical to
before this change.
- `bash -n` on the launcher; full pre-commit clean.
- GPU deploy of the `fp8` DeciLM checkpoint was verified by the bug
reporter on GB10 with this one-line change (model loaded, TRT engine
built, generation succeeded, deploy exit 0). The `mmlu` / `lm_eval`
stages are not GPU-verified here.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- examples-only bug fix -->
- Did you get Claude approval on this PR?: ❌ <!-- not yet run -->

### Additional Information

nvbug 6538278 / OMNIML-5659. Reported against modelopt 0.46.0rc0 on GB10
/ DGX Spark (`tensorrt-llm/release:1.3.0rc22`).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Added consistent handling of the `trust_remote_code` setting across
TensorRT-LLM inference and MMLU evaluation workflows.
* Ensured tokenizer and model loading receive the configured remote-code
behavior.
* **Tests**
* Added coverage validating remote-code settings and preserving KV-cache
behavior when context logits are requested.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 13:07:14 -07:00
Chenjie LuoandClaude Opus 5 b75227ce1b [NVBug: 5987078] Fix unified HF export of compressed NVFP4 weights (--low_memory_mode) (#2038)
### What does this PR do?

Type of change: Bug fix

Fixes unified HF export of already-compressed NVFP4 weights —
`mtq.compress` and, through it, `examples/llm_ptq/hf_ptq.py
--low_memory_mode` (NVBug 5987078). Two defects, one root cause: the
NVFP4 export branch has no handling for weights that were already
real-quantized, unlike the `FP8_PB_REAL` branch which consumes
`weight_quantizer._scale`.

**1. `weight_scale` was recomputed from packed data.** After compression
the weight is a `QTensorWrapper` of packed NVFP4 nibbles, and
`QTensorWrapper.__new__` builds the Parameter from `_quantized_data`, so
`.shape` reports the *packed* shape (the logical shape survives only in
`metadata["shape"]`). The export derived the block count from
`weight.shape[-1]` and took amax over nibble-pair bytes, so it wrote a
scale of half the required size with meaningless values:

| | `weight` | `weight_scale` written | expected |
| --- | --- | --- | --- |
| TinyLlama-1.1B `q_proj` | `[2048, 1024]` U8 | `[2048, 64]` | `[2048,
128]` |
| DeepSeek-R1-Distill-Llama-70B `q_proj` | `[8192, 4096]` U8 | `[8192,
256]` | `[8192, 512]` |

**2. An internal quantizer buffer leaked into the checkpoint.**
`postprocess_state_dict` strips `weight_quantizer.<name>` for every name
in `RealQuantLinear.list_of_scale_tensors`, but that list carried
`"double_scale"` where the buffer is `_double_scale` — a missing
underscore. So `_scale` was stripped and `_double_scale` was not, and it
reached the checkpoint as `*.weight_quantizer._double_scale` (560
entries in the 70B checkpoint). Downstream loaders reject it before
loading any weight:

```
KeyError: 'layers.0.mlp.down_proj.weight_quantizer._double_scale'
RuntimeError: Engine core initialization failed.
```

This is why only the TensorRT backend appeared usable in the bug report
— its converter tolerates the stray key, then produces `!!!!!!` output
from the broken scales, while the PyTorch backend (vLLM / TensorRT-LLM)
fails to load outright.

The fix reuses the per-block scale captured at compression time,
rescaled into the exported `weight_scale_2` convention, and corrects the
typo above. The rescale matters: compression normalizes per-block FP8
scales against the global scale it captured at that moment, which is not
the post-calibration `weight_scale_2` the export writes. Exporting the
stored scale as-is loads fine but leaves every block off by a constant
factor (~1.96x measured), so the `weight_scale * weight_scale_2` product
that dequantization consumes must be preserved.

Note the typo fix also affects
`modelopt/torch/quantization/plugins/megatron.py:505,511`, which filter
on the same list — these are internal buffers so excluding them looks
correct, but calling it out since it is a behavior change outside the
export path.

**Compression-time scale layout.** `TensorQuantizer._real_quantize`
calls `NVFP4QTensor.quantize(..., try_tensorrt=True)`, so on an
FP4-capable device with TensorRT-LLM importable the stored `_scale` is
the **cutlass-swizzled 1-D uint8** scale rather than the modelopt 2-D
E4M3 layout. Confirmed on GB10 in a TRT-LLM container:

```
logical weight (512, 256)  -> modelopt scale should be (512, 16) e4m3
_scale : (8192,) torch.uint8  (ndim=1)     <- cutlass-swizzled
after cutlass_fp4_scale_to_modelopt_fp4_scale: (512, 16) torch.float8_e4m3fn
```

The export therefore normalizes it the same way
`NVFP4QTensor.dequantize` does, and raises if `tensorrt_llm` cannot be
imported to convert, rather than writing raw byte values.

### Usage

No API change. The previously broken path now works:

```bash
python hf_ptq.py --pyt_ckpt_path <local_ckpt_dir> --qformat nvfp4 \
  --low_memory_mode --export_path <out>
```

### Testing

All runs on DGX Spark (GB10, sm121, aarch64). Both compression-time
scale layouts are covered, since the layout depends on whether
TensorRT-LLM is importable in the process:

**New tests**

-
`tests/gpu/torch/export/test_export_weight_gpu.py::test_export_compressed_nvfp4_weight`
— dense E4M3 path. Asserts the per-block scale covers the logical input
dim, that `weight_scale * weight_scale_2` matches an uncompressed export
of the same model, and that `postprocess_state_dict` strips both
internal buffers.
-
`tests/gpu_trtllm/torch/export/test_export_compressed_nvfp4.py::test_export_compressed_nvfp4_weight_trtllm_scale`
— cutlass-swizzled path. Asserts as a *precondition* that the
environment really produced a 1-D uint8 scale, so it cannot silently
degrade into the dense case when TensorRT-LLM is absent.

**Results**

| suite | TensorRT-LLM 1.3.0rc17 container | vLLM 26.05 container |
| --- | --- | --- |
| both files above | 3 passed | 2 passed, 1 skipped |

**Negative controls** (each test fails without the code it guards)

- On `main` with only the test file applied:
`test_export_compressed_nvfp4_weight` fails.
- With only the un-swizzle conversion neutered, the rest of the fix
intact: `test_export_compressed_nvfp4_weight_trtllm_scale` fails.

**End-to-end, TensorRT-LLM container** (the environment from the bug
report, where `_scale` is swizzled) — TinyLlama-1.1B, `--qformat nvfp4
--low_memory_mode`, plus a normal export as control. Exported
checkpoints are structurally identical (0 stray `_double_scale` keys vs.
560 on `main`; 663 keys each; `q_proj.weight_scale` `[2048, 128]` FP8 in
both; QKV `weight_scale_2` unified in both). Loaded on the
**TensorRT-LLM PyTorch backend** — the backend reported as unusable:

```
##### trt_normal #####
GEN: France and is the most populous city in the country. It is located on the Seine River...
##### trt_lowmem #####
GEN: France and is the most popular tourist destination in the country. It is a city of art, history...
```

On `main` that same load fails with `KeyError:
'...weight_quantizer._double_scale'` before a single weight is read.

**End-to-end, vLLM container** (dense scale path) — TinyLlama-1.1B,
NVFP4 + `--low_memory_mode`:

- before: `KeyError: '...weight_quantizer._double_scale'`, engine fails
to start
- after: loads and generates coherently (`"Paris is the capital of"` →
`" France and is best known for the awe inspiring Notre-Dame C"`)
- The fixed checkpoint is structurally equivalent to a normal
(non-low-memory) export: identical key set (663 tensors), `input_scale`
identical, and `weight_scale_2` identical — including the values unified
across fused groups, so the fused-GEMM contract (one shared
`weight_scale_2` per QKV / gate-up group) still holds.
- DeepSeek-R1-Distill-Llama-70B (the model in the bug) reproduces the
same signature on `main` and is the source of the numbers in the table
above.

### Known limitation: fused groups are slightly less accurate than a
full-memory PTQ

This restores a working checkpoint, but `--low_memory_mode` NVFP4 is
**not numerically identical** to a normal PTQ, and cannot be made so at
export time.

The nibbles are packed at *load* time against the layer's own global
scale, so the effective per-block scale baked into them is `fp8_own *
ws2_own`. `preprocess_linear_fusion` later unifies `weight_scale_2`
across a fused group to the group max, and the format requires the
per-block scale be E4M3, so the best the export can write is
`round_fp8(fp8_own * ws2_own / ws2_unified)`. For the group member
owning the max amax that ratio is exactly 1 and the round trip is
bit-exact; for the others it costs one extra E4M3 rounding, bounded by a
half-ULP (6.25%). The uncompressed path never pays this because its
weights are still high precision at export, so `to_quantized_weight`
re-quantizes the nibbles *after* unification.

Measured on TinyLlama-1.1B (weight relative error vs. the source BF16
weights, 154 quantized tensors):

| | mean rel. error |
| --- | --- |
| normal PTQ | 0.090248 |
| `--low_memory_mode` (this PR) | 0.091443 |

The degradation is confined to exactly 66 of 154 tensors = 22 layers x
3, i.e. the non-max members of each `q/k/v` and `gate/up` group (worst
observed +0.0066, e.g. `layers.11.self_attn.k_proj` 0.0898 -> 0.0964).
The 88 remaining tensors — group winners plus the unfused `o_proj` /
`down_proj` — are bit-exact.

Removing this requires compressing a fusion group against one shared
scale in the compress-on-load path (`RealQuantParameterDict`), where the
group max first becomes known; that is a larger change and is left as a
follow-up.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

NVBug 5987078.

Not addressed here: at 70B scale on DGX Spark, `--low_memory_mode` can
also crash *before* export when the device map offloads, because
`QTensorWrapper.to()` cannot represent a `meta` tensor:

```
accelerate/hooks.py: set_module_tensor_to_device(module, name, "meta")
RuntimeError: Attempted to call `variable.set_data(tensor)`, but `variable` and `tensor` have incompatible tensor type.
```

It is reproducible on demand under low free GPU memory (crashes at 41.2
GB and 58.5 GB free; succeeds at 88.1 GB) and is easy to hit on Spark's
unified memory, where page cache from reading the checkpoint counts
against `torch.cuda.mem_get_info()`. That is an independent defect and
will be filed separately.

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 17:46:24 +00:00
Chenjie LuoandClaude Opus 4.8 23be82655b Add nvfp4_act_headroom activation calibration for NVFP4 (#2028)
### What does this PR do?

Type of change: new feature

Adds `nvfp4_act_headroom`, a calibration algorithm for NVFP4
**activation** global scales.

NVFP4 scales a tensor in two levels: an FP8-E4M3 scale per 16-element
block, plus one per-tensor global scale. Plain `max` calibration sets
that global scale from the largest per-block amax seen during
calibration, which leaves no room above it — any activation larger than
the calibration max saturates.

`nvfp4_act_headroom` instead anchors the global scale to a low
percentile of the per-block amax distribution:

```
amax = max(rho * anchor, floor)
```

`anchor` is the per-block amax at `anchor_percentile`; `upper` is the
per-block amax at `upper_percentile` and is the top of the range the
scale commits to representing. Multiplying the low anchor by `rho`
places the calibrated blocks in the lower part of the FP8 scale range
and leaves the rest as headroom.

`upper_percentile` defaults to 99.99 rather than the literal maximum on
purpose. Flooring at the literal max means one freak block drags the
global scale up until every other block's FP8 block scale falls below
subnormal: on a tensor with a single block seven orders of magnitude
above the rest, that flushes 99.998% of elements to zero, versus 6.7%
when the rare blocks are clipped instead. On a benign wide-range tensor
the two choices differ by 0.4% relative MSE. Set `upper_percentile=100`
to floor at the literal observed max, which guarantees no calibration
data is clipped, at that exposure. The calibrator warns when the
per-block range is too wide for `rho` to clear any headroom.

The algorithm applies only to NVFP4 dynamic-block **input** quantizers;
weight quantizers and everything else keep plain `max`.

Files:

- `calib/nvfp4_act_headroom.py` — `NVFP4ActHeadroomCalibrator` (log2
histogram of per-block amaxes, bounded memory)
- `config.py` — `NVFP4ActHeadroomCalibConfig` (`anchor_percentile`
default 1, `rho` default 16384)
- `model_calib.py` — `nvfp4_act_headroom_calibrate`
- `mode.py` — mode registration
- `modelopt_recipes/general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` —
mirrors `nvfp4_default-kv_fp8_cast` (dynamic NVFP4 W4A4 + FP8 KV-cache
cast, same module coverage) with only the calibration algorithm swapped:

```diff
-  algorithm: max
+  algorithm:
+    method: nvfp4_act_headroom
+    anchor_percentile: 1
```

### Why a dedicated collector rather than reusing `HistogramCalibrator`

`HistogramCalibrator` histograms the tensor values themselves into
linear, dynamically re-ranged bins, and its percentile mode returns an
amax that *clips* a high-tail fraction. This algorithm needs a different
statistic (per-block amaxes, a derived quantity), different bin spacing
(log2, because block amaxes span many decades and the anchor is read
from the **low** tail, where linear bins have almost no resolution), and
a different reduction (a low-percentile anchor scaled up, not an
upper-tail clip). Reusing it would mean changing its collection
semantics for every existing caller; composing it would still leave the
log-spacing and the derived statistic unaddressed. The collector here is
a fixed 512-bin int64 histogram per quantizer, so the memory argument
for a histogram (rather than retaining values) is preserved.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> \
    --recipe general/ptq/nvfp4_act_headroom-kv_fp8_cast --dataset <calib.jsonl>
```

Or directly:

```python
import modelopt.torch.quantization as mtq

NVFP4 = {"num_bits": (2, 1), "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)}}
config = {
    "quant_cfg": [
        {"quantizer_name": "*", "enable": False},
        {"quantizer_name": "*weight_quantizer", "cfg": NVFP4},
        {"quantizer_name": "*input_quantizer", "cfg": NVFP4},
    ],
    "algorithm": {"method": "nvfp4_act_headroom", "anchor_percentile": 1, "rho": 16384},
}
model = mtq.quantize(model, config, forward_loop)
```

### Testing

**Unit tests** —
`tests/unit/torch/quantization/test_nvfp4_act_headroom.py`, 26 tests
covering the anchor and range terms, monotonicity in
`anchor_percentile`, headroom above plain max, the too-wide-range
fallback and its warning, the rare-outlier case (the default clips it;
`upper_percentile=100` chases it), the `upper_percentile=100`
no-clipping guarantee across six distributions, input validation (`rho`
/ percentile bounds, NaN/Inf, all-zero, a last dim that is not a
multiple of the block size), quantizer selection, calibrator restoration
after calibration, `shared_states` / `sync_expert_weight_amax`
forwarding, and W4A4 end-to-end confirming weights fall back to plain
max while only activations get the headroom scale.

Full `tests/unit/torch/quantization/` and `tests/unit/recipe/` suites
pass with no regressions (the remaining failures in this environment are
a pre-existing broken-`torchvision` import and
`test_data_parallel_auto_quantize`, both of which fail identically on
`main`).

**End-to-end on a real model** — ran the shipped recipe on a 30B hybrid
Mamba/attention MoE (1024 calibration samples @ seq 4096) and verified
the export is a standard NVFP4 checkpoint: 6268 activation quantizers
calibrated; 19G vs the 62G BF16 source; 6260 `input_scale`, 6267
`weight_scale` / `weight_scale_2`; `quant_algo: NVFP4`,
`kv_cache_quant_algo: FP8`; 53 excluded modules; no module carrying an
`input_scale` without a `weight_scale`. 32 of 6268 quantizers (0.5%) had
a per-block range too wide for `rho` to clear and warned.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — purely additive: a new opt-in
algorithm, config class, and recipe. No existing behavior changes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code, no new dependencies.
- Did you write any new necessary tests?: ✅ — 12 new unit tests.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — entry under 0.47 → New Features → Quantization.
- Did you get Claude approval on this PR?: ❌ — not yet run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-03 17:22:23 +00:00
Chenjie LuoandClaude Opus 5 55f1880f70 [skill] evaluation: correct the stale vLLM CUDA-13 image tag rule (#2042)
### What does this PR do?

Type of change: Bug fix (agent skill documentation)

The `evaluation` skill instructed the agent to **"append `-cu130` to the
image tag"** for NVFP4 checkpoints on Blackwell B300/GB300 (sm_103).
That was correct for v0.19.x, but **vLLM inverted its tag convention at
v0.20.0**: the *unsuffixed* tag is now the CUDA-13 build, and `-cu129`
is the CUDA-12 opt-out.

Consequences of the stale rule:

- `v0.20.1-cu130` / `v0.24.0-cu130` / `v0.26.0-cu130` **do not exist** —
following the rule literally asks for a missing tag.
- The documented fallback (`cu130-nightly-<arch>`) points at ~v0.20-era
builds that are *older* than several models' documented minimum vLLM
version, so it can't serve as an escape hatch either.

This replaces the "append a suffix" instruction with a version-keyed
table plus the durable check: **select a tag whose config blob reports
`CUDA_VERSION` ≥ 13**, resolving the child manifest for the platform you
actually deploy on.

While the PR was open the default image was also bumped, and review
surfaced two follow-on corrections. Full contents:

1. **Tag-convention fix** — version-keyed table, `-cu130` fallback
removed.
2. **Default image `v0.19.1` → `v0.26.0`** (latest vLLM release)
everywhere it was pinned: `SKILL.md` Step 3 and the Step 7.5 table,
`example_eval.yaml`, `example_eval_next.yaml`. Version specifics that
the bump made stale or self-contradictory were dropped (the `e.g.
v0.20.0` bump example and the MiniMax-M2.7 `≥0.20.0` anecdote, both now
*below* the default; the failure-mode lesson is kept).
3. **Convention boundary corrected to v0.20.0** — the first draft said
`≤ v0.20.x` suffixed / `≥ ~v0.21` unsuffixed. Off by a minor release in
both rows; see Testing.
4. **Config blob resolved per deployment platform** — the check said
"arm64 child". GB300/Grace is arm64, but plenty of B300 deployments are
`linux/amd64`.
5. **Same corrections applied to the `deployment` skill**, which carried
the original append rule untouched: its NVFP4 note,
`references/support-matrix.md`, `references/benchmarking.md`, and the
`:latest` pins in `references/setup.md` (now `v0.26.0`, matching the
evaluation skill's never-`:latest` stance).

An earlier revision of this branch also carried a
`.claude/skills/benchmark-model-kernels` symlink, added automatically by
`tools/precommit/sync_claude_skills.sh` — it repairs missing symlinks
repo-wide on any touch of `.agents/skills/`, and #1980 landed that skill
without its link. It has been dropped from this branch to keep the scope
on the vLLM image guidance. Worth its own one-line PR: without the
symlink, Claude Code doesn't load that skill at all.

### Usage

```bash
# Durable check — resolve the child manifest for YOUR platform (arm64 for
# Grace/GB300, amd64 for x86) and read CUDA_VERSION from its config blob:
#   v0.19.1            -> CUDA_VERSION=12.9.1   (unsuffixed = CUDA 12, old convention)
#   v0.19.1-cu130      -> CUDA_VERSION=13.0.1   (suffixed = CUDA 13, old convention)
#   v0.20.0            -> CUDA_VERSION=13.0.2   (transition release: ships both suffixes)
#   v0.26.0            -> CUDA_VERSION=13.0.2   (unsuffixed = CUDA 13, new convention)
#   v0.26.0-cu129      -> CUDA_VERSION=12.9.1   (suffixed = CUDA 12, new convention)
```

### Testing

Verified empirically against the Docker registry API for
`vllm/vllm-openai` — resolved each tag's child manifests and read
`CUDA_VERSION` / `TORCH_CUDA_ARCH_LIST` from the config blob.

**Where the convention flips (arm64):**

| release | unsuffixed | `-cu130` | `-cu129` |
|---|---|---|---|
| v0.18.0 | 12.9.1 | 13.0.1 | absent |
| v0.19.0 / v0.19.1 | 12.9.1 | 13.0.1 | absent |
| **v0.20.0** | **13.0.2** | 13.0.2 | 12.9.1 |
| v0.20.1 | 13.0.2 | **absent** | 12.9.1 |
| v0.20.2 | 13.0.2 | **absent** | 12.9.1 |
| v0.21.0 … v0.26.0 | 13.0.2 | absent | 12.9.1 |

v0.20.0 is the transition release — it publishes both suffixes *and* its
unsuffixed tag is already CUDA 13. That duplication is what hid the
boundary: confirming `v0.20.0-cu130` exists reads as "old convention
still applies at 0.20", while `v0.20.1-cu130` and `v0.20.2-cu130` don't
exist at all.

**Why the platform matters.** `CUDA_VERSION` is identical across
children on every tag checked (v0.26.0, v0.26.0-cu129, v0.20.0, v0.19.1,
v0.19.1-cu130, kimi-k3), but `TORCH_CUDA_ARCH_LIST` is not:

```
v0.26.0  amd64  7.5 8.0 8.6 8.9 9.0 10.0 12.0
v0.26.0  arm64  8.0 8.7 8.9 9.0 10.0 11.0 12.0
v0.19.1  amd64  7.0 7.5 8.0 8.9 9.0 10.0 12.0
v0.19.1  arm64  8.7 8.9 9.0 10.0+PTX 12.0
```

`11.0` appears only on arm64, `7.5` / `8.6` only on amd64 — so the arch
check has to read the child you'll actually run.

**Default bump.** The `0.26.0` family is `{,-aarch64,-x86_64} ×
{,-cu129} × {,-ubuntu2404}` — 12 tags, `cu129` the only CUDA axis, no
`-cu130`. The new default is therefore already a CUDA-13 build, and
NVFP4 on B300/GB300 needs no suffix at all.

Docs-only change; no runtime code touched. `pre-commit` clean.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (agent skill
documentation)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

Split out of the GDPVal skill work (#2039) because it is independent of
GDPVal and applies to every NVFP4-on-Blackwell deployment the skill
generates. Now spans both the `evaluation` and `deployment` skills, all
under `.agents/`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
  * Updated deployment and evaluation guidance to use vLLM v0.26.0.
* Clarified CUDA image-tag conventions, including CUDA 13 defaults and
CUDA 12 opt-outs.
* Added validation guidance for resolved CUDA versions, platform
architecture settings, and image compatibility.
* Updated NVFP4 Blackwell B300/GB300 support notes, including required
`sm_103` kernel availability.
* Refreshed example recipes, setup commands, benchmarking notes, and
serving-image requirements.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 17:01:32 +00:00
Chenjie LuoandClaude Fable 5 7b8da80205 docs(recipes): sync ptq.md with shipped recipes and enforce via unit tests (#1970)
### What does this PR do?

**Type of change:** documentation (+ new tests)

`modelopt_recipes/ptq.md` is meant to track every shipped PTQ recipe,
but recent recipe PRs did not update it. This PR brings the doc back in
sync and adds unit tests so it can't drift again.

**Doc sync — general recipes:**
- Add `nvfp4_experts_only_input_scale1-kv_fp8_cast` (#1947) to the
shipped-recipes table (now 20 recipes) and document the `input_scale1`
variant (constant amax 2688 → exported NVFP4 `input_scale == 1.0`,
expert activations uncalibrated).
- Add a pointer that the PTQ recipes' `quantize` sections also drive
QAT/QAD, with the LSQ / Dual-LSQ QAD recipes under `general/qad/`
(#1884).

**Doc sync — model-specific recipes:**
- Add `gemma4` (algorithm override, #1690), `diffusion_gemma` (extra
`*self_conditioning*` exclusion, #1707), `vit` (FP8 with attention BMM
quantizers for Torch-TRT, #1569), and the qwen3_5 / qwen3_5_moe
`w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast` twins (#1620).
- Rewrite the checkpoint-mirror section for the
`huggingface/models/nvidia/` tier: Super recipe rename
(`super-nvfp4.yaml` → `nvfp4-mse.yaml` / `nvfp4-max-calib.yaml`),
Nano-4B GGUF Q4_K_M mirror (#1327/#1606), Ultra-550B `nvfp4-4o6`
(#1684); fix stale `models/<checkpoint>` paths.

**Enforcement — new `tests/unit/recipe/test_recipe_docs.py`:**
1. Every `general/ptq/*.yaml` stem must appear (backticked) in `ptq.md`.
2. Every recipe row in the shipped-recipes table must exist on disk
(catches renames/removals).
3. The "All N `general/ptq/` recipes" count must match the file count.
4. Every model folder under `huggingface/` containing `ptq/*.yaml` must
be mentioned in the doc.

### Usage

```bash
pytest tests/unit/recipe/test_recipe_docs.py -q
```

### Testing

- `pytest tests/unit/recipe/ -q` → 218 passed (includes the 4 new
tests).
- Verified enforcement bites: against the pre-update `ptq.md`, 3 of the
4 new tests fail with actionable messages.
- `pre-commit run --files modelopt_recipes/ptq.md
tests/unit/recipe/test_recipe_docs.py` → all hooks pass.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (docs + tests only; no runtime
code changed)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ —
`tests/unit/recipe/test_recipe_docs.py`
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (documentation/test change)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Follow-up to the recipe guide added in #1662. The doc-consistency tests
could also be wired into `tools/precommit/check_modelopt_recipes.py` for
commit-time feedback if reviewers prefer both layers.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added the `nvfp4_experts_only_input_scale1-kv_fp8_cast` recipe to the
general PTQ catalog.
* Documented the new `input_scale1` calibration behavior and related
quantization/caching details.
* Expanded and reworked PTQ documentation, including updated navigation,
“Beyond PTQ” guidance, new/updated model-specific recipes, and expanded
checkpoint mirror entries.
  * Updated the general PTQ shipped-recipe count to 20.

* **Tests**
* Added automated consistency checks between on-disk shipped recipe YAML
files and the PTQ documentation (coverage, existence, and counts).
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 09:19:31 +08:00
Chenjie LuoandClaude Opus 4.8 a05850bffa Add nel-next (0.3.x) agentic AA benchmark support to eval skill (#1861)
### What does this PR do?

Type of change: new feature (agent skill — nel-next agentic eval
support)

Adds support for the **nel-next** evaluator (`nemo-evaluator[harbor]`
0.3.x) to the `evaluation` agent skill, for the **agentic** AA
benchmarks (Terminal-Bench 2.1, SWE-bench Verified) that do not run on
the default `nemo-evaluator-launcher` 0.2.6 (different package, CLI,
override syntax, and config schema).

- `.agents/scripts/nel-next.sh` — bootstrap + run nel-next from an
**isolated venv**, leaving the 0.2.6 install untouched.
- `references/nel-next.md` — shared reference (venv,
`services`/`benchmarks`/`cluster`/`output` schema, harbor/ECS-Fargate
architecture, timeout strategy, MLflow export, run flow, gotchas).
- `recipes/tasks/aa_next/{terminal_bench_2_1,swebench_verified}.md` +
`recipes/examples/example_eval_next.yaml` — per-benchmark recipes +
self-contained template.
- Internal harbor infra (`eval_image`, ECR repos) is **not hardcoded** —
configs read `${NEL_NEXT_EVAL_IMAGE}` / `${HARBOR_*_ECR_REPOSITORY}`
from `.env`, written by the companion `modelopttools:eval-config` skill
(internal repo, separate change). `SKILL.md` + `recipes/env.example`
updated.

### Usage

```bash
.agents/scripts/nel-next.sh --setup-only                 # one-time isolated 0.3.x venv
set -a && source .env && set +a                          # HF_TOKEN, AWS_*, harbor infra vars
.agents/scripts/nel-next.sh eval run <config>.yaml --dry-run   # then --submit
```

### Testing

- Dry-run validated against `nemo-evaluator` 0.3.0 (`extra="forbid"`
schema) for both the Terminal-Bench 2.1 and SWE-bench Verified configs.
- Live runs on gcp-nrt (MiniMax-M2.7 NVFP4, 8×B200): Terminal-Bench 2.0
completed `pass@1=0.469` (712 trials, auto-resume across walltime
windows); SWE-bench Verified canary passed end-to-end (vLLM deploy →
OpenHands agent → AWS ECS Fargate sandbox → scoring/merge/report).
- `pre-commit run` clean (insert-license, markdownlint, YAML format,
symlink sync).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive — new files + one
branch section in the eval skill; the 0.2.6 path is unchanged)
- If you copied code or added a new PIP dependency, did you follow
`CONTRIBUTING.md`: N/A (the helper installs `nemo-evaluator` into a
runtime venv; no new repo dependency)
- Did you write any new necessary tests?: N/A (agent-skill docs/tooling;
validated via dry-run + live cluster runs)
- Did you update Changelog?: N/A (agent skill content, not a
Model-Optimizer feature/API)
- Did you get Claude approval on this PR?: N/A

### Additional Information

Companion change (separate internal repo, `modelopt-internal`): the
`modelopttools:eval-config` skill gains a Step 3b that writes the harbor
infra rows (`NEL_NEXT_EVAL_IMAGE`, `HARBOR_ECR_REPOSITORY`,
`HARBOR_SWEBENCH_ECR_REPOSITORY`) into `.env` and points at the
canonical per-benchmark `bench.yaml` source of truth.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `nel-next` CLI wrapper to run NeMo Evaluator 0.3.x from an
isolated venv, including setup/version helpers and safer install/refresh
behavior.
* Introduced nel-next (harbor) templates and benchmark runbooks for
Terminal-Bench 2.1 and SWE-bench Verified.
* **Documentation**
* Added a dedicated nel-next reference guide covering configuration,
required workflow, reporting, and common gotchas.
* Updated evaluation skill docs and the `.env` example with
nel-next/harbor credential and output variables, plus new example
evaluation configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-01 02:16:50 +00:00
Chenjie Luo 138564f443 Add AA-Omniscience eval recipe; harden judge/run conventions in the eval skill (#1834)
### What does this PR do?

Type of change: documentation (agent `evaluation` skill)

Updates the `evaluation` agent skill (`.agents/skills/evaluation/`) with
an AA-Omniscience recipe plus several judge/run hardening conventions
found while running the AA Index v2 suite on quantized checkpoints.

- **AA-Omniscience recipe** (`recipes/tasks/aa/omniscience.md`):
nemo-skills `ns_omniscience`, AA Index v2 params
(`++parse_reasoning=False`, `num_repeats=10`),
`gcp/google/gemini-3-flash-preview` judge; score
`omniscience_pass_at_1_avg-of-N_judge_correct`. Added to the AA Index v2
suite list in `SKILL.md`.
- **Judge model_ids hardcoded** in the HLE / AA-LCR / Tau2 recipes (with
a "swap for an equivalent on your own endpoint" note); shared judge URL
var renamed `NS_JUDGE_URL` -> `INFERENCE_JUDGE_URL`; judge API key
folded into `INFERENCE_API_KEY` in `env.example`.
- **Idle-reaper exemption** in `example_eval.yaml`:
`cluster.sbatch_comment` exempts eval jobs from the
`OccupiedIdleGPUsJobReaper`, which otherwise CANCELs jobs whose GPUs sit
idle during model load / judge calls / aggregation.
- **Resume-on-kill note** in `SKILL.md`: after a preemption /
idle-reaper CANCEL (not a walltime timeout, which NEL auto-resumes),
re-submit the job's `run.sub` to resume from the response cache
(`skip_filled`) with cumulative progress.

### Usage

\`\`\`bash
nel run --config recipes/examples/example_eval.yaml # now ships the
idle-reaper exemption
# extend evaluation.tasks with the AA-Omniscience fragment from
recipes/tasks/aa/omniscience.md
\`\`\`

### Testing

Validated end-to-end on gcp-nrt (B200): MiniMax-M2.7-NVFP4
AA-Omniscience full run (600 questions x 10 repeats) completed with
\`omniscience_pass_at_1_avg-of-10_judge_correct = 18.77\`. The
idle-reaper exemption and the \`sbatch run.sub\` resume path were both
exercised - a reaped run resumed from cache and finished all 10 repeats.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: yes (skill docs/recipes only; no
library API change)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (agent skill recipes/docs)
- Did you update Changelog?: N/A (agent skill, not a library feature)
- Did you get Claude approval on this PR?: pending (/claude review)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added an **AA-Omniscience** evaluation recipe.
* Expanded the default **AA Index v2** quantized-checkpoint validation
suite to include Omniscience.

* **Documentation**
* Clarified which judge/user-simulator **model identifiers** are fixed
in recipes vs supplied via environment variables.
* Updated HLE, LCR, and Tau2-Bench Telecom guidance for judge
configuration and API-key/endpoint usage.
* Added instructions for resuming evaluations after scheduler
preemption/cancellation.

* **Chores**
* Updated the example evaluation config to include a GPU job reaper
policy.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-06-26 18:10:35 +00:00
Chenjie LuoandClaude Opus 4.8 1c6bdb3021 Fix reduce_amax NotImplementedError on FP8 weights (NVBug 6360175) (#1824)
### What does this PR do?

Type of change: Bug fix

Fixes [NVBug 6360175](https://nvbugspro.nvidia.com/bug/6360175) /
OMNIML-5265: quantizing a model whose weights are stored natively in FP8
(e.g. DeepSeek-V3 in `float8_e4m3fn`) crashes during `mtq.quantize`
calibration with:

```
File ".../modelopt/torch/quantization/utils/core_utils.py", line 162, in reduce_amax
    max_val = torch.max(input)
NotImplementedError: "max_all_cuda" not implemented for 'Float8_e4m3fn'
```

**Root cause:** FP8 dtypes (`float8_e4m3fn` / `float8_e5m2`) implement
no full-tensor reduction kernel (`max_all_cuda`/`min_all_cuda`), nor
`amax`/`amin`, `abs`, or elementwise `maximum`. `reduce_amax` called
these directly on the FP8 weight tensor.

**Fix:** Upcast FP8 inputs to the default float dtype
(`torch.get_default_dtype()`) at the top of `reduce_amax`, before any
reduction. The upcast is **lossless** (any default float dtype
represents every FP8 value exactly) and only affects the FP8 path — the
common (fp16/bf16/fp32) path is untouched. Placing the upcast at the top
covers all branches (`torch.max`/`min`, `torch.amax`/`amin`,
`torch.abs`), not just the line in the traceback.

### Usage

No API change. Quantization of natively-FP8 checkpoints (e.g.
DeepSeek-V3 NVFP4 PTQ) now runs through calibration instead of raising.

### Testing

- New CPU regression test `test_reduce_amax_fp8` in
`tests/unit/torch/quantization/test_utils.py` covering both FP8 dtypes
(`float8_e4m3fn`, `float8_e5m2`) across all axis modes (`None`, `0`,
`1`, `(0, 1)`); asserts results equal the float reference and the output
dtype is the default float dtype. CPU reproduces the original error (no
FP8 reduction kernel there either), so the test is GPU-free.
- `pre-commit run --files ...` passes (ruff, mypy, bandit, license, rst
checks).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅ (0.45 Bug Fixes)
- Did you get Claude approval on this PR?: ❌ (not yet)

### Additional Information

NVBug 6360175 is tagged `Committed_ModelOpt_0.45.0` (regression); the
changelog entry is under 0.45 and this will be cherry-picked to
`release/0.45` after merge.

Supersedes #1823, which got a stuck head ref (frozen at the original
commit, no sync on force-push) after the repo move
`TensorRT-Model-Optimizer` → `Model-Optimizer`; it could not be
re-synced or reopened, so this PR replaces it from the same branch.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 18:51:36 +00:00
Chenjie LuoandClaude Opus 4.8 9cfd7dd3e8 Align eval skill AA benchmarks to golden NeMo configs + harden skill (#1790)
### What does this PR do?

Type of change: Documentation (agent skill under
`.agents/skills/evaluation/`)

Aligns the `evaluation` agent skill's AA benchmark recipes with NeMo
Evaluator's Nemotron-3-Ultra golden reproducibility configs, and fixes
several robustness gaps surfaced while validating the changes end-to-end
on SLURM (MiniMax-M2.7, FP8 + ModelOpt NVFP4).

**AA benchmark alignment**
- GPQA: simple-evals `gpqa_diamond_aa_v3` → nemo-skills `ns_gpqa`
(`++prompt_config=eval/aai/mcq-4choices`).
- MMLU-Pro: `mmlu_pro_aa_v3` → `ns_mmlu_pro` (`mcq-10choices-boxed`).
- Repeat counts unchanged (GPQA 16, MMLU-Pro 1); score-extraction metric
keys updated to the nemo-skills names
(`gpqa_pass_at_1_avg-of-N_symbolic_correct`,
`mmlu-pro_pass_at_1_symbolic_correct`, verified against MLflow run
data).
- HLE: add golden knobs `hle_strict_judge: true` +
`++server.enable_soft_fail=True`.
- Updated `references/{quantization-benchmarks,parallelism}.md` and the
example config's default task to match.

**Skill robustness fixes**
- Invoke `modelopttools:eval-config` at Step 1 for judge-scored runs;
create/populate `.env` before Step 5 (it was only created at Step 8) —
fixes a fresh-environment ordering gap.
- Secret-safety rule: never open `.env` with file tools (the harness
mirrors later external edits back into context, leaking keys) — interact
via shell only.
- Require fetching `recipes.vllm.ai` for the **exact model variant** and
bumping the image to its minimum vLLM — variant minimums differ (e.g.
MiniMax-M2 ≥0.11.0 vs M2.7 ≥0.20.0), and running below crashes
mid-inference with `CUDA illegal memory access`. Note that `--dry-run`
does not validate the image/version.
- Remove internal cluster names from the public skill docs.
- gitignore the local eval run-config dir.

### Usage

N/A — agent skill docs/recipes; no library/API change.

### Testing

Validated on SLURM with MiniMax-M2.7 (16-sample canaries):
- FP8: GPQA + MMLU-Pro ran through the new nemo-skills harness and
produced valid scores (e.g. MMLU-Pro `symbolic_correct=75.0`),
confirming the harness switch + metric keys.
- NVFP4: confirmed the vLLM-min-version finding — deploy crashed on
0.19.1 (`CUDA illegal memory access` mid-inference) and succeeded on
0.20.2 (the recipe minimum).
- (HLE tokenizer gap was reproduced deterministically; its fix is a
separate follow-up, not in this PR.)

### Before your PR is "*Ready for review*"

Commits are signed (`git commit -s -S`). ✅

- Is this change backward compatible?: N/A — agent skill/recipe docs.
Note: the GPQA/MMLU-Pro harness switch changes the MLflow metric keys,
so new runs are not directly comparable to prior simple-evals baselines
(re-baseline).
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (docs/recipes only)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (run `/claude review`)

### Additional Information

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Strengthened evaluation setup guidance: `.env` is now required earlier
with safer creation/sourcing rules, and endpoint placeholders are
clarified.
* Updated judge/user-simulator and vLLM deployment instructions,
including exact minimum `recipes.vllm.ai` variant/image requirements,
expected failure symptoms, and a warning about `--dry-run`.
* Refreshed AA benchmark/task recipes to the `nemo-skills` harness
(e.g., `ns_gpqa`, `ns_mmlu_pro`) with updated parameters and
score-metric guidance, plus stricter HLE judging options.
* **Bug Fixes**
* Improved score-reporting guidance for GPQA and MMLU-Pro to match
current outputs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-25 02:18:58 +05:30
Chenjie LuoandClaude Opus 4.8 106781659e Add W4A16 NVFP4-MSE Qwen3.5 dense/MoE PTQ recipes (#1620)
### What does this PR do?

Type of change: new feature (PTQ recipe)

Adds an MSE-calibrated counterpart of the existing
`w4a16_nvfp4-fp8_attn-kv_fp8_cast` PTQ recipe for the Qwen3.5 family
(dense `qwen3_5` and MoE `qwen3_5_moe`).

New files:
-
`modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`
— shared `quant_cfg` snippet
-
`modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`
— dense recipe
-
`modelopt_recipes/huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`
— MoE recipe

The only difference from the `max` variant: NVFP4 MLP / `lm_head` weight
scales come from an MSE FP8-scale sweep (`method: mse`,
`fp8_scale_sweep: true`, `nvfp4_static`) instead of max calibration. FP8
attention / linear-attention projections and the FP8 KV cast are
unchanged. The dense and MoE families share a single `quant_cfg` snippet
under `qwen3_5/ptq`, matching the existing recipe's layout.

### Usage

```bash
# Dense
python hf_ptq.py --recipe huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ...
# MoE
python hf_ptq.py --recipe huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ...
```

### Testing

- Both recipes load and resolve their `$import`s via
`modelopt.recipe.loader.load_recipe`.
- The `check-modelopt-recipes` pre-commit validator passes on all three
files.
- Verified against the source that weight-only MSE (`method: mse` +
`fp8_scale_sweep: true`) is supported for W4A16: `mse_calibrate` refines
only weight quantizers (`iter_weights_for_calibration`) and needs no
input/activation quantizers, and the `nvfp4_static` numeric satisfies
`is_nvfp4_static` so the FP8 scale sweep engages on the MLP/lm_head
weights. No code changes were required.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (config-only; covered by
existing recipe-loader validation)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (not yet)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## New Features

* Added PTQ recipe configurations for Qwen3.5 and Qwen3.5-MoE model
families
* Supports W4A16 quantization with NVFP4 static weights and MSE-based
calibration
* Enables FP8 precision for self-attention and KV-cache optimization for
improved model performance and reduced memory footprint

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-16 15:45:50 -07:00
Chenjie LuoandClaude Opus 4.8 d7df14d12a [tools/debugger] Enforce a single relay owner across hosts (#1735)
### What does this PR do?

Type of change: Bug fix

The `tools/debugger` file-based relay assumes a single server but never
enforced it.
Because the relay lives on shared NFS (the repo is often the same
checkout mounted on
multiple hosts), a forgotten `server.sh` on another host kept polling
the same
`.relay/` and could **steal commands** (executing them on the wrong
host), and killing
one server's cleanup could **wipe the active server's markers**.

This adds a `.relay/owner` ownership token (`host:pid:nanos`):

- Each server writes `owner` atomically at startup and **takes over**
instead of
refusing when a stale `server.ready` exists (the old `kill -0 <pid>`
guard was
  host-local and meaningless across hosts).
- The handshake and main loops exit cleanly if `owner` changes
(`[server] Superseded by <id> — exiting.`), so a freshly started server
**evicts**
  any stale one — even on another host.
- `cleanup()` only clears shared markers if we still own them, so a
stepping-down
  server never clobbers its successor's `server.ready`/`owner`.

Also gitignores `tools/debugger/logs/` and documents the `owner` file in
the README.

### Usage

```bash
# Inside the container; a previously-running server elsewhere that shares this
# NFS .relay/ steps down automatically once this one claims ownership:
bash tools/debugger/server.sh
# [server] Note: existing server.ready found (<host:pid:ts>); taking over.
#   (the stale server logs: "[server] Superseded by <id> — exiting.")
```

### Testing

Verified live on computelab: a forgotten `server.sh` on another host was
evicted when
a new server started, after which `client.sh run` executed on the
correct (new) host;
confirmed the stepping-down server's cleanup does not remove the
successor's
`server.ready`/`owner`. `server.sh` passes `bash -n`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- additive: new .relay/owner
file; client.sh and the wire protocol are unchanged -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- the file-based relay
tool has no test harness; behavior verified manually -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- internal dev tooling, not a shipped feature/API -->
- Did you get Claude approval on this PR?: N/A <!-- can run /claude
review -->

### Additional Information

Scope is limited to `tools/debugger/` (`server.sh`, `README.md`,
`.gitignore`).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated relay protocol documentation to clarify ownership-based server
coordination.

* **Bug Fixes**
* Improved reliability of multi-server coordination in shared relay
environments.

* **Chores**
  * Updated ignore patterns for logging files.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 18:27:52 +00:00
Chenjie LuoandClaude Opus 4.8 d26c8af002 Fix DeepSeek V3 ptq.py inference-repo path resolution (nvbug 6311147) (#1702)
### What does this PR do?

Type of change: Bug fix

Fixes nvbug **6311147** (OMNIML-5103).
`examples/deepseek/deepseek_v3/ptq.py` resolved the cloned DeepSeek-V3 /
DeepSeek-V3.2-Exp inference repos relative to its own directory
(`deepseek_v3/`) via `Path(__file__).resolve().parent`. But the
[README](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/deepseek)
clones those repos into the parent `examples/deepseek/` directory and
runs the script from there, so the lookup landed one level too deep and
raised `ValueError: DeepSeek-V3 or DeepSeek-V3.2-Exp not found` (the
error message also printed the wrong directory).

The fix resolves from `parent.parent` via a single `DEEPSEEK_DIR` base
shared by both repo paths and the error message.

### Usage

```bash
# Run from examples/deepseek/ as documented in the README, after cloning
# DeepSeek-V3 (or DeepSeek-V3.2-Exp) into that directory:
torchrun --nproc-per-node 8 --master_port=12346 deepseek_v3/ptq.py \
  --model_path $DS_CKPT \
  --config DeepSeek-V3/inference/configs/config_671B.json \
  --quant_cfg NVFP4_DEFAULT_CFG \
  --output_path $FP4_QUANT_PATH
```

### Testing

- Confirmed against the repro path: with the file at
`examples/deepseek/deepseek_v3/ptq.py` and the repos cloned into
`examples/deepseek/`, `Path(__file__).resolve().parent.parent` now
points at `examples/deepseek/` so `DeepSeek-V3/inference` resolves
correctly.
- Verified the sibling `examples/deepseek/deepseek_v4/` does not share
the bug (it takes an explicit `--dsv4_inference_dir` argument instead).
- `pre-commit` clean.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (one-line path fix in an
example script that requires the DeepSeek repos + multi-GPU checkpoint
to exercise)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (bug is in a 0.45-cycle example, not a regression from a released
version)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

nvbug 6311147 / OMNIML-5103.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved path resolution in the example script to more reliably locate
the required inference repository.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 22:25:29 +00:00
Chenjie Luo 60b1af5fb2 Fix GPT-OSS MXFP4->NVFP4 PTQ load, export, and cast (nvbug 6295279, 6295242) (#1678)
### What does this PR do?

Type of change: Bug fix

Fixes the GPT-OSS MXFP4 → NVFP4 PTQ path (`examples/llm_ptq/hf_ptq.py`
with `--cast_mxfp4_to_nvfp4`), which failed in three independent ways.
The documented command now runs end-to-end and produces a bit-exact
(100% lossless) NVFP4 checkpoint. Addresses **nvbug 6295279**
(OMNIML-5046) and **nvbug 6295242** (OMNIML-5045).

1. **nvbug 6295242 — CUDA illegal memory access on load.** GPT-OSS ships
native MXFP4 weights that Transformers dequantizes to BF16; the threaded
weight loader trips an illegal-memory access when `device_map="auto"`
shards the dequant across **multiple GPUs**. The missing optional
`kernels` package only *forces* the dequant path — it is not the root
cause. `get_model` now detects MXFP4 checkpoints and loads them with
`Mxfp4Config(dequantize=True)` on a **sequential** device map so the
dequant stays on a single device. `kernels` is no longer required.
2. **nvbug 6295279 #1 — `NotImplementedError: Mxfp4GptOssExperts` during
unified HF export.** Forcing `dequantize=True` yields plain
`GptOssExperts` (even when `kernels` is installed), which ModelOpt wraps
and exports normally.
3. **nvbug 6295279 #2 — `FileNotFoundError` in the cast step.**
`--cast_mxfp4_to_nvfp4` treated `--pyt_ckpt_path` as a local dir; a HF
Hub ID now resolves to its cached snapshot dir via
`_resolve_model_path`.

Also fixes a **static-block NVFP4 regression** (surfaced by the cast's
`force_weight_quantizers_static`, introduced by #1560's
now-unconditional `weight_only_quantize`): `_QuantGptOssExperts` /
`_QuantLlama4TextExperts` quantize their expert weights transposed in
the forward (`_transposed_quantize`), but the inherited
`iter_weights_for_calibration` fed the non-transposed weight, locking a
mismatched block-quant `_original_shape` and raising `ValueError: Input
shape has changed`. The override now calibrates on the transposed view,
matching both the forward and the export's `_amax` orientation.

### Why this regressed (it worked when the cast was added)

`get_model` never had explicit handling for a *natively pre-quantized
MXFP4* checkpoint — GPT-OSS fell through the generic
*unquantized-checkpoint* branch and relied on Transformers' **implicit**
MXFP4 behavior, which is fragile across three axes. The cast was
originally validated (#1372, 2026-05-01) in the "lucky" quadrant of
each:

- **GPU count:** `device_map="auto"` on a single GPU never shards, so
the dequant stays on one device. On multiple GPUs `auto` balances the
model and shards the MXFP4→BF16 dequant across devices → CUDA
illegal-memory crash (6295242).
- **`kernels` presence:** without `kernels`, Transformers
auto-dequantizes to BF16 `GptOssExperts` (exportable). With `kernels`
installed it keeps the packed `Mxfp4GptOssExperts` kernel path → export
`NotImplementedError` (6295279 #1).
- **Transformers version:** the kernel-backed experts wrapper and the
threaded multi-GPU weight loader are newer-Transformers behavior (env
here is 5.5.4). Earlier versions simply dequantized MXFP4 → BF16, which
is what the old generic path happened to need.

The QA env sat in the *breaking* quadrant (multi-GPU and/or `kernels`
present, newer Transformers), so the implicit path failed. The new
branch makes both decisions explicit and deterministic
(`dequantize=True` + single-device load), regardless of environment —
mirroring the existing `has_pack_quantized_config` branch for
compressed-tensors checkpoints.

The fourth issue (static-block `Input shape has changed`) is a separate
regression: it was introduced by **#1560 (2026-06-02, "Make sure all
weight quantizers have `_amax`")**, a month *after* the cast landed.
#1560 made `weight_only_quantize` unconditional in `max_calibrate`;
previously it ran only when no calibration `forward_loop` was supplied,
and the cast always supplies one — so the non-transposed
weight-quantizer call simply never happened before. The conflict only
appears at the intersection of (a) transposed-quantize experts
(GPT-OSS/Llama4), (b) static-block NVFP4 — which `--cast_mxfp4_to_nvfp4`
forces via `force_weight_quantizers_static` — and (c) #1560. CI's
GPT-OSS NVFP4 coverage uses the *dynamic*-block path, which never locks
the block shape, so #1560 looked safe.

### Usage

```bash
python hf_ptq.py \
  --pyt_ckpt_path openai/gpt-oss-20b \
  --qformat nvfp4_mlp_only \
  --cast_mxfp4_to_nvfp4 \
  --export_path ./gpt-oss-20b-nvfp4
```

### Testing

- Ran the documented command end-to-end on 2xB200
(`openai/gpt-oss-20b`): cast overrode **48/48** expert weight
quantizers, **100% lossless** layers/blocks, exported a valid
packed-NVFP4 HF checkpoint (uint8 weights + FP8 per-block `weight_scale`
+ per-tensor `weight_scale_2` + `hf_quant_config.json`).
- Verified plain `--qformat nvfp4_mlp_only` (no cast) still works
end-to-end.
- **Independently verified the export is bit-exact:** dequantized the
exported NVFP4 weights (ModelOpt's E2M1 LUT + pack layout) and compared
against Transformers' canonical MXFP4→BF16 dequant
(`Mxfp4Config(dequantize=True)`) over all 24 layers × both expert
weights — `max_abs_err = 0`, 100% bitwise-equal in bf16. So
`dequant(exported NVFP4) == dequant(original MXFP4)` exactly.
- New unit tests: `test_get_original_hf_quant_method_*` (load detection)
and `test_gpt_oss_experts_iter_weights_for_calibration_transposed` (the
transpose regression). Existing `test_cast_mxfp4_to_nvfp4.py` (8 tests)
still pass. `pre-commit` clean.

**Known limitation:** verified for gpt-oss-20b (fits one GPU).
gpt-oss-120b dequantized does not fit a single GPU, so `sequential`
would still span GPUs — that case would need a CPU-dequant-then-dispatch
path and is left as a follow-up.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (0.45 Bug Fixes)
- Did you get Claude approval on this PR?: ❌ (not yet run)

### Additional Information

nvbug 6295279, nvbug 6295242 / OMNIML-5046, OMNIML-5045.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Prevented CUDA illegal-memory access during MXFP4→NVFP4 casting.
* Fixed expert-weight calibration orientation to avoid shape mismatches.

* **New Features**
* Support loading native MXFP4 checkpoints with automatic
dequantization.
* Resolve remote model identifiers to local checkpoints when casting
MXFP4→NVFP4, improving reliability.

* **Tests**
* Added unit and GPU regression tests covering quant-method detection,
casting, and expert-weight calibration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-06-12 22:30:13 +05:30
Chenjie LuoandClaude Opus 4.8 9cb00ce43a docs: add modelopt_recipes README and PTQ recipe/scheme guide (#1662)
### What does this PR do?

Type of change: documentation

Adds two docs under `modelopt_recipes/` (no code or behavior changes):

- **`README.md`** — catalog of the recipe library: its purpose (a recipe
is the
single, version-controlled source of truth for *how* a model is
optimized), the
directory layout (`general/`, `huggingface/`, `models/`, `configs/`),
how to
load/select recipes (`load_recipe`, `--recipe`), and a high-level map of
the
  general PTQ combos, speculative-decoding, and distillation recipes.
- **`recipe.md`** — a focused guide to the PTQ schemes: the general
`general/ptq/`
body scopes (full-model FP8/NVFP4, scoped experts-only / mlp-only /
omlp-only,
weight-only), KV-cache modes (`kv_fp8_cast` / `kv_nvfp4_cast` /
`kv_fp8`),
calibration variants (max / mse / gptq / layerwise), low- vs
high-concurrency
deployment guidance, and the model-specific recipes under `huggingface/`
and
  `models/` — each compared to its general baseline.

### Usage

```python
# Documentation only. The recipes themselves load as before, e.g.:
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_experts_only-kv_fp8_cast")
```

### Testing

`pre-commit run --files modelopt_recipes/README.md
modelopt_recipes/recipe.md`
passes (markdownlint, modelopt recipe validation, license/format hooks).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A <!-- docs only -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs only -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- docs only -->
- Did you get Claude approval on this PR?: ❌ <!-- not yet -->

### Additional Information

Documentation for the `modelopt_recipes/` library; content verified
against the
recipe YAMLs and the `modelopt.recipe` / config-loader source.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive ModelOpt recipes guide describing YAML-based,
composable optimization workflows, directory/lookup layout, reuse via
imports, and how to add or share recipes.
* Added PTQ quantization guide covering recipe naming/structure,
quantization scopes and KV-cache options, calibration variant guidance,
model-specific overrides, multimodal considerations, and a
checkpoint-mirroring example.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 15:01:12 -07:00
Chenjie LuoandClaude Opus 4.8 bde162a701 feat(deepseek): add --cast_mxfp4_to_nvfp4 to deepseek_v4 quantize step (#1653)
### What does this PR do?

Type of change: new feature

Brings the GPT-OSS lossless MXFP4 → NVFP4 cast (#1372) to DeepSeek V4's
routed-expert export by adding a `--cast_mxfp4_to_nvfp4` flag to
`examples/deepseek/deepseek_v4/quantize_to_nvfp4.py`.

To avoid duplicating the closed-form math, the shared numerics —
`mxfp4_to_nvfp4_global_amax`, `mxfp4_to_nvfp4_per_block_amax`, and the
E2M1/E4M3/E8M0 constants — are **hoisted out of the GPT-OSS example cast
into the library** at
`modelopt/torch/quantization/utils/numeric_utils.py`. Both the GPT-OSS
cast (`examples/llm_ptq/cast_mxfp4_to_nvfp4.py`) and the new DeepSeek
path now import them from there.

DeepSeek V4's routed experts ship as MXFP4 (E2M1 nibbles + a
power-of-two E8M0 scale per 32-element block). By default the export
dequantizes them to BF16 and re-quantizes to NVFP4 using the calibrated
per-tensor weight amax, which re-derives per-block scales from the data
and is therefore lossy. With the flag, the cast pins `scale_2 =
2^(k_max-8)` and each per-block E4M3 scale to `2^(k_j-m)` straight from
the source E8M0 scales, so `per_block_scale * scale_2 = 2^k_j` and the
NVFP4 nibbles equal the source MXFP4 nibbles bit-for-bit (for every
block whose `k_j` lands in E4M3's representable window; rare
out-of-range blocks clamp). The one V4-specific addition is that w1/w3
share a single `scale_2` for the fused GEMM1, so `k_max` is taken over
both projections. The flag only affects routed-expert **weights** —
activation `input_scale` still comes from `--amax_path` calibration.

### Usage

```bash
python deepseek_v4/quantize_to_nvfp4.py \
    --amax_path ${AMAX} \
    --source_ckpt ${DS_V4} \
    --output_ckpt ${HF_NVFP4_PATH} \
    --cast_mxfp4_to_nvfp4
```

### Testing

- The hoisted numerics get unit tests in
`tests/unit/torch/quantization/test_numeric_utils.py` (10 cases:
per-tensor global_amax, per-block amax incl. out-of-range,
magnitude-table cache) — 10/10 pass. The example test
`tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py` keeps the
cast-specific cases (quantizer naming, `build_amax_map`,
`apply_to_model`).
- Validated on real DeepSeek-V4-Flash expert tensors (incl. the on-disk
`float8_e8m0fnu` scale dtype): 23.5M blocks, 100% lossless, 0 error.
- Generated a full NVFP4 checkpoint for DeepSeek-V4-Flash (43 layers,
256 routed experts) end-to-end: `[cast] lossless MXFP4->NVFP4 blocks:
8,657,043,456/8,657,043,456 (100.0000%)`. Output weights match an
independently-produced reference cast byte-for-byte (`weight_scale`,
`weight_scale_2`, packed nibbles modulo the harmless sign-of-zero).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (new opt-in flag; default
export behavior unchanged; hoist re-exports through the existing example
module)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ N/A (no new
deps; shared numerics moved into the library rather than duplicated)
- Did you write any new necessary tests?: ✅ (library numerics covered by
`tests/unit/torch/quantization/test_numeric_utils.py`; end-to-end
validated on a real DeepSeek-V4 checkpoint)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (will run `/claude review`)

### Additional Information

Mirrors and reuses #1372 (GPT-OSS MXFP4 → NVFP4 cast); the closed-form
numerics are now shared via
`modelopt.torch.quantization.utils.numeric_utils`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `--cast_mxfp4_to_nvfp4` flag to perform a closed-form, mostly
lossless MXFP4→NVFP4 conversion for routed-expert weights with
aggregated lossless/block statistics.

* **Documentation**
* Updated DeepSeek V4 export instructions and README to document the new
flag and clarify calibration behavior for activation scales.

* **Chores**
* Exposed shared numeric quantization utilities for MXFP4→NVFP4 casting.

* **Tests**
* Added and updated tests to validate the new numeric helpers and
conversion behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-10 10:54:34 +05:30
Chenjie LuoandClaude Opus 4.8 56c4af2333 feat(recipes): add kv_fp8_cast variants for partial-NVFP4 and weight-only PTQ recipes (#1652)
### What does this PR do?

Type of change: new feature (recipes)

Several `general/ptq` recipe families shipped a data-driven FP8 KV-cache
(`-kv_fp8`) variant but lacked the constant-amax `kv_fp8_cast` companion
that `fp8_default` and `nvfp4_default` already have. This PR adds the
missing cast variants so every KV-quantizing (and the weight-only)
family offers the calibration-free FP8 KV-cache option:

- `general/ptq/nvfp4_experts_only-kv_fp8_cast`
- `general/ptq/nvfp4_mlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_omlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_weight_only-kv_fp8_cast`

Each new recipe composes the exact same model-quant config as its
existing sibling and swaps the `kv_fp8` unit for the shared
`kv_fp8_cast` unit (constant-amax FP8 KV cache; no KV calibration
forward pass). The docs guide table/tree and the changelog are updated
to match.

### Usage

```bash
python examples/llm_ptq/hf_ptq.py \
    --pyt_ckpt_path <model> \
    --recipe general/ptq/nvfp4_mlp_only-kv_fp8_cast
```

### Testing

Extended the built-in PTQ smoke test
`tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins` with
the four new recipe paths; all four load into a valid
`ModelOptPTQRecipe` with a populated `quantize` section.

```
$ python -m pytest tests/unit/recipe/test_loader.py tests/unit/recipe/test_presets.py -q
180 passed
```

`pre-commit` (including the `validate modelopt recipes` hook) passes on
all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive — only new recipe
files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (extended the builtin recipe
smoke test)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (not yet)

### Additional Information

The two weight-only families were discussed for scope;
`nvfp4_weight_only` is included (it already names a KV mode, `kv_fp16`),
while `int4_blockwise_weight_only` is intentionally left untouched since
it carries no `-kv_` composition.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added four new NVFP4 PTQ (Post-Training Quantization) recipe variants:
experts-only, MLP-only, OMLP-only, and weight-only configurations.
* All new recipes include FP8 KV-cache cast mode support for improved
inference performance.

* **Documentation**
* Updated built-in recipes guide with new NVFP4 recipe options and
repository layout.

* **Tests**
  * Expanded recipe loader test coverage for new recipe configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:13:25 -07:00
Chenjie LuoandClaude Opus 4.8 6f08731fe6 docs(eval skill): vLLM backend env vars + SLURM HF-cache/cpu_partition guidance (#1625)
### What does this PR do?

Type of change: documentation

Hardens the **evaluation** skill with operational guidance discovered
while running AA-Index evals (NVFP4 Nemotron-3-Nano) on SLURM. Two
themes:

**1. vLLM backend env vars (original commits).** Some models need a
model-card backend toggle that is an env var, not a CLI flag — e.g.
NVFP4 MoE models like
[NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4](https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4)
need `VLLM_USE_FLASHINFER_MOE_FP4=1` +
`VLLM_FLASHINFER_MOE_BACKEND=throughput`. These go in
`deployment.env_vars` (with the `lit:` prefix), not the `vllm serve`
command.

**2. SLURM deploy/eval operational lessons (new commits).**
- **`mount_home: false` always.** Some internal cluster templates
default it `true`, which mounts the host `~/.cache`; where that is a
symlink into a shared/networked filesystem, it dangles in the container
and the vLLM `trust-remote-code` deploy dies with `FileNotFoundError:
/root/.cache/huggingface` — **invisible to `--dry-run`**.
- **HF cache:** mount the realpath of `~/.cache/huggingface` to
`/hf-cache` and set `HF_HOME: lit:/hf-cache` for both stages.
- **`execution.cpu_partition`:** on split GPU/CPU-partition clusters,
the CPU-only MLflow auto-export job is rejected by the GPU partition and
marks the whole task FAILED despite `EVAL_EXIT_CODE=0`. Set
`cpu_partition` to route it correctly.
- **Top-level `env_vars:`** for vars both stages need (`HF_TOKEN`,
`HF_HOME`); `execution.env_vars` is unsupported and hard-errors.
- **`lcr.md`:** `LOG_LEVEL=WARNING` to skip logging AA-LCR's ~120K-token
inputs.

### Usage

```yaml
execution:
  cpu_partition: <cpu-partition>   # CPU-only auto-export job
  mounts:
    mount_home: false
    deployment: { <shared-fs>/<user>/.cache/huggingface: /hf-cache }
    evaluation: { <shared-fs>/<user>/.cache/huggingface: /hf-cache }
env_vars:                          # shared by both stages
  HF_TOKEN: host:HF_TOKEN
  HF_HOME: lit:/hf-cache
deployment:
  env_vars:
    VLLM_USE_FLASHINFER_MOE_FP4: lit:1
    VLLM_FLASHINFER_MOE_BACKEND: lit:throughput
```

### Testing

Docs/skill-only. Validated end-to-end by running the full AA suite (GPQA
+ SciCode + AA-LCR) across 5 NVFP4 checkpoints on SLURM: dry-run +
canary + full runs all SUCCESS with these settings. Pre-commit (yamlfmt,
markdownlint) passed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Clarified where backend-related environment variables must be
configured (deployment-level) and not embedded in execution commands.
* Added guidance to mount the real HuggingFace cache path and set
home-mounting to false for containerized SLURM runs.
* Documented routing CPU-only jobs via CPU partitions and added a
task-level LOG_LEVEL setting to reduce verbose long-context logging.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-05 23:36:51 +00:00
Chenjie LuoandClaude Opus 4.8 a548be5da6 [Eval skill] Add reasoning adapter_config to eval template (#1610)
### What does this PR do?

Type of change: documentation

Adds a reasoning-aware adapter (interceptor) config to the `evaluation`
skill's
default eval template and documents how to set `use_reasoning` per model
type.

- **`recipes/examples/example_eval.yaml`** — adds an `adapter_config`
block under
`evaluation.nemo_evaluator_config.target.api_endpoint` with
request/response
  logging, `use_reasoning`, `log_failed_requests`, and a
`params_to_add.chat_template_kwargs` thinking block, with inline
comments on
  which lines to flip for instruct vs. reasoning vs. hybrid models.
- **`SKILL.md`** — new "Reasoning adapter config (`use_reasoning`)"
subsection at
  the end of Step 3:
- Instruct (non-reasoning) models → `use_reasoning: false` (no CoT trace
to
    strip; drop the `chat_template_kwargs` thinking block).
- Reasoning models → `use_reasoning: true`, especially when the
deployment
    command sets `--reasoning-parser`.
- Hybrid models (can run reasoning on or off) → always force thinking on
via
`chat_template_kwargs` and keep `use_reasoning: true` (highest scores).

### Usage

```bash
nel run --config recipes/examples/example_eval.yaml \
  -o deployment.checkpoint_path=/path/to/checkpoint \
  -o deployment.served_model_name=my-model
```

### Testing

Docs/template-only change. Ran `pre-commit` on both files (yamlfmt +
markdownlint
pass).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs/template only
-->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <!-- not yet run -->

### Additional Information

Note: there is a pre-existing inconsistency between Step 7 (which
references
`nemo_evaluator_config.config.target.api_endpoint.adapter_config`) and
the
template's `nemo_evaluator_config.target.api_endpoint` placement. This
PR follows
the template's existing path; reconciling Step 7 is left as a follow-up.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a guide for configuring adapter reasoning/chain-of-thought
behavior by model type (instruct, reasoning, hybrid), how reasoning
traces can be stripped before scoring, model-family template keyword
variants, and guidance to verify template kwargs and select appropriate
reasoning-effort settings.

* **Chores**
* Added an example configuration demonstrating enabling the reasoning
adapter with caching, limited request/response logging, failed-request
logging, forced “thinking” flags for hybrid models, and a
reasoning-effort option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 14:25:35 -07:00
Chenjie LuoandClaude Opus 4.8 88fd7ff958 eval skill: auto-detect predefined per-cluster execution configs (#1599)
### What does this PR do?

Type of change: documentation

Some NEL installs ship ready-made per-cluster execution configs as an
`internal/slurm/<cluster>` group (via the optional
`nemo_evaluator_launcher_internal` package). When present, the matching
one pre-fills the cluster's `hostname` / `partition` / `gres` (and
node-exclusivity), so the user only sets `account` / `output_dir` /
`walltime` instead of hand-entering the hostname.

- **`SKILL.md` Step 4** — adds a "check FIRST" step: run a discovery
snippet that lists the available `cluster → hostname` pairs **from the
installed package at runtime**; on a hostname match, use `defaults: -
execution: internal/slurm/<cluster>` (replacing `slurm/default`) and
drop the now-redundant `execution.hostname`. If the package isn't
installed or nothing matches, fall back to `slurm/default` and fill the
fields manually.
- **`example_eval.yaml`** — a short, name-free comment on the `defaults`
block pointing to that check.

**Discovery-based by design (no internal data committed):** cluster
names, hostnames, and accounts are read from the install at runtime and
are **not** hardcoded here. The only internal references in the repo are
the package name `nemo_evaluator_launcher_internal` and the generic
`internal/slurm/<cluster>` group pattern. External users without the
package degrade gracefully (import fails → `slurm/default`).

### Usage

N/A — documentation / skill guidance only.

### Testing

Ran the discovery snippet verbatim (lists `cluster → hostname` from the
installed package) and confirmed `internal/slurm/<cluster>` resolves via
a `--dry-run`. `pre-commit run` passes (markdownlint + YAML format).
Grep-verified no cluster names/hostnames are hardcoded in the changed
files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (falls back to `slurm/default`)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs)
- Did you get Claude approval on this PR?: ❌ (pending)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added Step 4 guidance for discovering and using optional internal
per-cluster SLURM execution configs: how to list available internal
configs, select the matching cluster by switching the default execution,
remove redundant hostname entries, and verify with a --dry-run.
* Clarified that slurm/default is cluster-agnostic, recommend using
internal per-cluster configs when present, and explained fallback
behavior to manual hostname/account/output_dir entry.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 11:22:06 -07:00
Chenjie LuoandClaude Opus 4.8 db9ea8f17e eval skill: parameterize external judge/user-sim endpoints via .env (#1591)
### What does this PR do?

Type of change: documentation

Several AA tasks call an external **judge / user-simulator / scoring
endpoint** whose `model_id` + `url` vary per user/site (HLE, AA-LCR,
Tau2 today — and the guidance is written as a general pattern for any
future such benchmark). Previously each recipe hardcoded `<...>`
placeholders that a user had to hand-edit in every config. This makes
them reusable **without committing any internal infrastructure**:

- **`recipes/env.example`**: add placeholders — `NS_JUDGE_URL`,
`HLE_JUDGE_MODEL_ID`, `LCR_JUDGE_MODEL_ID`, `TAU2_USER_MODEL_ID`,
`TAU2_JUDGER_MODEL_ID`, `TAU2_ENDPOINT_URL` — with the **recommended
model named** (GPT-4o / Qwen3 235B / gpt-oss-120B) but only **generic
hosts** (`https://<your-inference-host>/v1`). No internal hostnames or
gateway model-routing strings are committed; real values live in the
user's gitignored `.env`.
- **`recipes/tasks/aa/{hle,lcr,tau2_bench_telecom}.md`**: carry `<VAR>`
literal placeholders (named after the `.env` keys) that the skill
**substitutes as literal values** from the user's `.env`. These are
*config, not secrets*, so they are **not** exported — which avoids the
`${oc.env:...}` footgun (it silently fails unless the var was exported
with `set -a`). Only `api_key` (`INFERENCE_API_KEY`) stays an exported
env var read by the harness.
- **`SKILL.md`** (Step 5): instructs literal substitution from `.env`,
framed as a general pattern for any external-endpoint task.

The `/v1`-base (nemo-skills) vs full `/v1/chat/completions` (tau2-bench)
URL distinction is documented.

### Usage

N/A — documentation / skill-template only.

### Testing

`pre-commit run` passes (markdownlint). Verified no internal hostnames /
gateway model IDs are present in any committed file.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Branched off latest `main` (includes #1583). Touches `lcr.md`, which
#1583 also edited (parallelism field) — different lines, rebased
cleanly.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guidance for configuring external judge and
user-simulator endpoints in evaluation tasks
* Clarified best practices for substituting configuration values from
environment configuration files while keeping API keys as environment
variables
* Updated task documentation for HLE, LCR, and Tau2-Bench with improved
instructions for credential and endpoint handling

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 22:35:03 -07:00
Chenjie Luo 905259fbf5 Fix: use python3 in debugger server.sh (#1590)
### What does this PR do?

Type of change: Bug fix

`tools/debugger/server.sh` invokes `python -c "..."` inside
`check_modelopt_local` (and `pip install -e .[dev]` in the fallback
path). Containers that only ship `python3` (no `python` shim) cause the
check to fail with `python: command not found`, exit 127. The server
interprets this as "modelopt not editable-installed", runs `pip
install`, the second check fails the same way, and the server aborts
before it ever listens for commands.

Switches both the inline import check and the install fallback to
`python3` / `python3 -m pip`.

### Usage

```bash
bash tools/debugger/server.sh
# now starts cleanly on python3-only images
```

### Testing

Verified inside a container where `which python` returns nothing and
`python3` is `/usr/bin/python3` (Python 3.12). With the old script,
`server.sh` aborted at `check_modelopt_local`. With this fix, the check
passes against an existing editable install and the server proceeds to
wait for the client handshake.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `python3` is present on every
image that previously had `python`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — internal tooling shell
script.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — internal debugger tooling, not user-facing.
- Did you get Claude approval on this PR?: N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated debugger tooling to ensure consistent use of Python 3 for
package validation and installation processes.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-06-01 21:47:21 +00:00
Chenjie LuoandClaude Opus 4.8 38c78439f2 Add parallelism sizing reference to evaluation skill (#1583)
### What does this PR do?

Type of change: documentation

Adds `.claude/skills/evaluation/references/parallelism.md`, a reference
for sizing both the **GPU topology** and the **request concurrency** of
NEL evaluation runs, and wires the main `SKILL.md` to it. Pure docs for
the agent-facing evaluation skill — no code or runtime behavior changes.

The new reference covers:

- **GPU topology (TP / DP / PP):** decision procedure
(smallest-TP-then-max-DP), the TP-up triggers, and the TP/DP-split
tradeoff for a fixed world size (e.g. all `TP×DP=8` factorizations on an
8-GPU node).
- **Expert parallelism (EP):** `--enable-expert-parallel` is a boolean
and `EP = TP × DP` (no direct EP-size flag); the DP-attention + EP-MoE
dataflow (per-layer dispatch/combine all-to-all); when to enable vs not.
- **Concurrency (`parallelism` / `--max-num-seqs`):** the
request-count-vs-serving-capacity ceiling, KV-driven `--max-num-seqs`,
and empirical tuning from vLLM startup logs + preemption.
- **Gotcha + worked examples:** bit-width read from `config.json` (not
the model name) sets the topology; includes the FP8-vs-4-bit Kimi
example.

`SKILL.md` gains pointers to the reference from the deployment-command,
expert-parallel, evaluation-params, Step 4, and canary sections.

### Usage

N/A — documentation only (agent skill guidance).

### Testing

`pre-commit run` passes on both files (markdownlint included).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs, not a library feature)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Commits are signed (`git commit -s -S`).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded deployment guidance with clearer GPU topology, MoE/EP
semantics, and bit‑width caveats.
* Refined concurrency guidance: explicit rules for sizing top-level
parallelism and computing max concurrent sequences, plus canary-run
tuning to watch preemption and KV utilization.
* Added task-level guidance for long‑context/judge‑bound jobs with
recomputation advice and worked examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 14:33:40 -07:00
Chenjie LuoandClaude Opus 4.8 2b7668d227 Fix MLflow auto-export in evaluation skill template (#1586)
### What does this PR do?

Type of change: documentation

Fixes two issues in the evaluation skill that caused completed eval runs
to silently **not** upload to MLflow (found while running a real GPQA
eval — the run finished successfully but never appeared in MLflow):

1. **Missing auto-export trigger.** NEL only auto-exports when
`execution.auto_export.destinations` is set; the `export.mlflow` block
alone just *configures* the exporter. The skill's shortcut said to "copy
the `export.mlflow` block" without calling out the trigger, so a config
could end up with the block but no upload.

2. **Interpolated export values break at submit time.** With auto-export
enabled, NEL resolves the `export.mlflow` block at submit time in a
scope that does **not** include `deployment` / `evaluation`, so
`${deployment.served_model_name}` / `${evaluation.*}` cross-references
fail hard with `Interpolation key '...' not found`. (`example_eval.yaml`
shipped this interpolated pattern *with* auto-export, so it would hit
the error.)

Changes:

- **`example_eval.yaml`**: mark the `auto_export.destinations` trigger
as REQUIRED; switch the `export.mlflow` block to **literal** values
(`CHANGEME-served-model-name` placeholders) with a comment explaining
why cross-refs break. `${oc.env:USER}` (env interpolation) is retained
since it resolves fine.
- **`SKILL.md`**: rewrite the MLflow shortcut step to require **both**
the trigger and literal export values, and warn against
`${deployment.*}` / `${evaluation.*}` cross-references.

### Usage

N/A — documentation/skill-template only.

### Testing

`pre-commit run` passes on both files (markdownlint + YAML format).
Verified the literal-export approach uploads correctly by exporting a
real run to MLflow.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (documentation)
- Did you update Changelog?: N/A (skill docs)
- Did you get Claude approval on this PR?: ❌ (pending)

### Additional Information

Separate from #1583 (parallelism reference); both touch the evaluation
skill but are independent.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Clarified MLflow auto-export: auto-export must be explicitly enabled
for uploads and the MLflow export block is ignored otherwise.
* Example config now uses literal experiment name, description, and tags
(no cross-scope interpolation) and shows concrete sampling values.
* Warned to keep sampling params consistent between tags and runtime
params and to set the tracking URI.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-01 13:11:35 -07:00
Chenjie LuoandClaude Opus 4.7 ed0a4b175d Refactor evaluation skill: vLLM cross-check, MLflow defaults, walltime cap (#1561)
### What does this PR do?

Type of change: documentation

Refactor of the `evaluation` skill (`.claude/skills/evaluation/`) with
several substantive rule additions plus a significant compression pass.

**Skill rules added / tightened:**

- **Cross-check `recipes.vllm.ai` + HF model card before composing the
vLLM command.** Both sources matter; conflicts get surfaced to the user
instead of silently picked. WebFetch caveat triage (3 cases) for
JS-rendered variant tabs (this caveat bit the agent multiple times
during testing; the triage rules name the failure modes).
- **Single `deployment.command:` field** replaces separate
`tensor_parallel_size` / `data_parallel_size` / `extra_args` YAML
fields. NEL mounts the model at `/checkpoint`; Hydra interpolates
`${deployment.port}`.
- **vLLM defaults always included unless a recipe contradicts them**
(silence ≠ contradiction):
  - `--max-num-batched-tokens 8192`
  - `--enable-chunked-prefill`
- `--enable-expert-parallel` (MoE-only, detected via active-param suffix
or `num_experts`-like config field)
- `--max-num-seqs N` where `N = ceil(max_parallelism /
data_parallel_size)`, computed after Step 4 fills in `parallelism`.
- **Six-field evaluation params template:** `parallelism`,
`request_timeout`, `max_retries`, `max_new_tokens`, `temperature`,
`top_p`. No `top_k` / `presence_penalty` / `repetition_penalty` /
`min_p` at top level (task harnesses have their own defaults that
conflict). No per-task `max_new_tokens` overrides — one ceiling
everywhere.
- **`max_new_tokens` mandatory model-card lookup:** highest
card-recommended value wins; "card not yet checked + use generic
default" is explicitly forbidden (this was a real bug the user caught).
- **AA Index v2 (`recipes/tasks/aa/`)** is the default benchmark set for
quantized-checkpoint validation. "AA" / "Artificial Analysis" triggers
AA-only mode (no MMLU-Pro / AIME / LiveCodeBench unless asked).
- **MLflow auto-export on by default in shortcut path**, with
Hydra-interpolated `experiment_name` (`${USER}/${served_model_name}`),
`description` (embeds T / top_p / max_new_tokens), and `tags`
(string-coerced via single quotes for MLflow's tag type requirement).
Only `tracking_uri` needs user input.
- **Walltime capped at 4h** in generated configs (longer walltimes lower
scheduler priority → longer queue). Skill suggests three alternatives
for over-4h runs.
- **Example template** updated to the new conventions; `--max-model-len`
fallback bumped 32K → 131K to cover AA-LCR.
- **`tau2_bench_telecom` recipe:** `parallelism` left as `???` with
rate-limit guidance (canary 32–128, cap 512).
- **`model-card-research.md`:** output-length extraction is now a
mandatory, top-level checklist item with a cross-reference to the
SKILL.md rule.
- **`env.example`:** added `JUDGE_API_KEY` entry for AIME.

**Compression pass:** SKILL.md compressed ~670 → ~330 lines while
preserving every rule. Long prose collapsed into tables and bullets;
duplicated workflow checklist at the bottom removed.

### Usage

The skill is invoked when users ask to evaluate a model. The shortcut
path now produces a config like:

```yaml
deployment:
  command: >-
    vllm serve /checkpoint
    --host 0.0.0.0
    --port ${deployment.port}
    --tensor-parallel-size 8
    ...

evaluation:
  nemo_evaluator_config:
    config:
      params:
        parallelism: ???    # Required — ask user
        request_timeout: 3600
        max_retries: 10
        max_new_tokens: 81920   # from model card (highest)
        temperature: 1.0
        top_p: 0.95

export:
  mlflow:
    tracking_uri: ???
    experiment_name: ${oc.env:USER}/${deployment.served_model_name}
    description: '...'
    tags:
      framework: vllm
      ...
```

### Testing

Manually exercised the skill end-to-end on four models during refactor
(Qwen3.5-122B-A10B-FP8, Qwen3.6-35B-A3B-FP8, Kimi-K2.6, GLM-5.1-NVFP4)
to validate that the cross-check rules surface conflicts, the MLflow
defaults populate correctly, the AA suite excludes the right tasks, and
the walltime cap holds. Test configs are not included in this PR
(session artifacts at repo root, gitignored locally).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — the skill is
editorial/operational; no runtime API changes. Existing configs continue
to work.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — skill documentation;
manually exercised on four models.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — skill documentation only.
- Did you get Claude approval on this PR?: ❌ — happy to run `/claude
review` if maintainers want it.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
- Rewrote the evaluation workflow with a stricter end-to-end checklist,
workspace-reuse guidance, AA-only shortcut, mandatory validation steps,
iterative finalization loop, and a hard 04:00:00 walltime cap
- Enforced vLLM as a single deployment command and stricter
generation/config constraints, including exact top-level params and
model-card–derived max_new_tokens
- Updated env var guidance (JUDGE_API_KEY / INFERENCE_API_KEY), MLflow
export metadata, and registry/auth preflight flow
- Added several AA-task recipes, added new benchmark recipes, and
removed or consolidated legacy task docs

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1561?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-30 01:47:10 +05:30
Chenjie LuoandClaude Opus 4.7 a5bc6f8123 Add DATASET_COMBOS for grouped calibration datasets (#1508)
## Summary

- Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` —
single ``--dataset`` tokens that fan out to several entries in
``SUPPORTED_DATASET_CONFIG``. The per-entry ``num_samples`` is split
evenly across the members inside ``get_dataset_dataloader``.
- Two initial combos:
- ``cnn_nemotron_v2_mix`` → ``cnn_dailymail`` +
``nemotron-post-training-dataset-v2``. Replaces the hardcoded
two-element fallback list in ``hf_ptq.py`` when ``--dataset`` is
omitted.
- ``nemotron-post-training-v3`` → the seven ``nvidia/Nemotron-*`` SFT
datasets registered in #1498 (mirroring the upstream
[`nemotron-post-training-v3`
collection](https://huggingface.co/collections/nvidia/nemotron-post-training-v3)).
- ``get_supported_datasets()`` now appends combo names so they show up
in ``--dataset`` help.
- ``hf_ptq.py``'s default ``--calib_size`` bumped from ``512`` to
``1024`` so the ``cnn_nemotron_v2_mix`` combo's even split preserves the
previous total sample count (was 512 per-dataset × 2 datasets = 1024;
now 1024 split → 512 per-dataset × 2). ``--calib_size`` now denotes the
total calibration budget regardless of combo cardinality.
- Reject mixing a combo with one of its member datasets in the same
``--dataset`` list (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) — combo
would otherwise double-sample the explicit member with a smaller
per-member quota.
- Reject combo names in ``get_dataset_samples``; combos are
dataloader-only. The error message points callers to
``get_dataset_dataloader``.
- Validate ``DATASET_COMBOS`` at import time: empty member lists, name
collisions with ``SUPPORTED_DATASET_CONFIG``, and references to unknown
datasets raise ``ValueError`` up front.

## Test plan

End-to-end validated against ``/hf-local/Qwen/Qwen3.5-0.8B`` via
``get_dataset_dataloader`` on the actual streamed data, plus 5 new unit
tests in ``TestDatasetCombosExpansion`` (all 44 tests in
``test_dataset_utils.py`` pass with no regressions).

- [x] ``python -c "from modelopt.torch.utils.dataset_utils import
DATASET_COMBOS, get_supported_datasets; assert 'cnn_nemotron_v2_mix' in
get_supported_datasets() and 'nemotron-post-training-v3' in
get_supported_datasets()"``
- [x] ``hf_ptq.py`` with no ``--dataset`` flag still calibrates on
cnn_dailymail + nemotron-post-training-dataset-v2 with the same total
sample count as before.
- [x] ``--dataset nemotron-post-training-v3 --calib_size 1024``
allocates 146 per member across the seven Nemotron datasets; full
1022-sample dataloader builds without error.
- [x] ``--dataset cnn_dailymail,nemotron-post-training-v3 --calib_size
256,1024`` composes correctly: 256 from cnn_dailymail (as a plain entry)
plus the 7-way split from the combo. (The earlier
``cnn_dailymail,cnn_nemotron_v2_mix`` example is rejected by design
since ``cnn_dailymail`` is a member of that combo.)
- [x] ``--dataset cnn_dailymail,cnn_nemotron_v2_mix`` raises
``ValueError`` with a clear message.
- [x] ``get_dataset_samples("cnn_nemotron_v2_mix", ...)`` raises
``ValueError`` pointing to ``get_dataset_dataloader``.
- [x] Unit tests: ``pytest
tests/unit/torch/utils/test_dataset_utils.py`` — 44 passed.

## Post-validation fix

End-to-end testing surfaced that the original
``nemotron-sft-agentic-v2`` entry kept the two splits
(``interactive_agent``, ``tool_calling``) that pyarrow's streaming JSON
reader cannot parse, and excluded ``search`` which is the only clean
split. Failures reproduce deterministically across cache wipes with
``force_redownload``, so they are content-level defects in the published
JSONL files at the pinned revision, not local artifacts:

- ``interactive_agent`` — heterogeneous schema
(``Column(.../member_id/type) changed from string to array``) at JSONL
row 4.
- ``tool_calling`` — malformed JSON row in a later shard, fails at
sample ~885 with ``Missing a closing quotation mark in string``.
- ``search`` — streams cleanly (verified to 2500 samples).

Commit ``10f3cfd`` corrects ``nemotron-sft-agentic-v2`` to use only
``search``, with an updated comment. The CHANGELOG calls this out as a
separate bullet so the behavior change on a previously-released dataset
entry from #1498 is discoverable.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset combo support: a single dataset token can expand into multiple
registered datasets with even sample splitting; predefined combos added
(e.g., cnn_nemotron_v2_mix, nemotron-post-training-v3) and listed as
supported.

* **Updates**
  * Default dataset when none specified now uses cnn_nemotron_v2_mix.
  * Calibration size default increased from 512 to 1024.

* **Bug Fixes**
* nemotron-sft-agentic-v2 now uses only the deterministic "search" split
to avoid streaming JSON errors.

* **Tests**
* Added coverage for combo expansion, splitting, overlap validation, and
rejection behavior.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1508?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-20 03:08:43 +05:30
Chenjie Luo 5508c327fb [Quantization] MSE-calibrate every per-expert weight in fused-experts MoE (#1421)
### What does this PR do?

Type of change: Bug fix

Two-part fix for transformers 5.x fused-experts containers (Qwen3-MoE /
Qwen3.5-MoE / Mixtral / DeepSeek / Kimi-K2.x ...) where weight
quantizers live in `nn.ModuleList`s (`gate_up_proj_weight_quantizers`,
`down_proj_weight_quantizers`):

1. **Per-expert weight iteration for calibration.** Add
`_QuantFusedExperts.iter_weights_for_calibration` that yields per-expert
`(weight_slice, quantizer)` pairs for both projections. The base impl
uses singular `*_weight_quantizer` and silently skips fused-experts
modules, so weight-only calibration paths never reached per-expert
quantizers.

2. **`mse_calibrate` refactor.**
- Add `_bootstrap_uncalibrated_weight_quantizers` after `max_calibrate`
to populate `_amax` on quantizers the forward pass didn't reach (dead
MoE experts that received no calibration tokens). Runs the existing
calibrator on the weight slice surfaced by
`iter_weights_for_calibration`.
- Replace the singular-only `weight_attr_names` discovery +
`getattr`-by-name walk with an `iter_weights_for_calibration` walk done
inside each parent module's `enable_weight_access_and_writeback`
context, so MSE processes every per-expert quantizer (active and dead)
and remains FSDP-safe.

Without this, the export-time fallback in `_export_fused_experts`
derived separate gate/up amaxes from each half of the fused weight,
breaking the `gate==up` `weight_scale_2` invariant on dead experts.

Also includes:
- `_sanitize_generation_config_for_save` in `unified_export_hf` —
coerces `do_sample=True` when an upstream `generation_config.json` has
`top_k`/`top_p` set, so newer transformers' strict validate doesn't
block `save_pretrained`.
- Small companion plumbing in `moe_utils.py`, `tensor_quantizer.py`, and
`core_utils.py` to support the per-expert iteration and bootstrap path.

### Usage

```python
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_config

# Recipe `nvfp4_experts_only_mse-kv_fp8_cast` (already on main) now correctly
# MSE-calibrates every per-expert weight quantizer in fused-experts MoE models.
cfg = load_config("general/ptq/nvfp4_experts_only_mse-kv_fp8_cast")
mtq.quantize(model, cfg, forward_loop=calibration_forward_loop)
```

### Testing

**Original validation — Qwen3.5-122B-A10B with
`nvfp4_experts_only_mse-fp8_cast_kv`:**
- **Before:** 1/12288 (layer 38 expert 69) `gate \!= up`; 0 weights
MSE-calibrated.
- **After:** 0/12288 mismatches; 24576 weights MSE-calibrated; ~4.2 min.

**End-to-end pipeline validation — Qwen3.5-35B-A3B (40 layers × 256
experts × 2 projections = 20,480 per-expert weight quantizers), TRT-LLM
1.3.0rc13 + transformers 5.6 docker, single B200:**

| | Path A (4-sample calib, deliberately undercalibrated) | Path B (zero
forward-pass tokens) |
|---|---|---|
| Per-expert weight quantizers calibrated | 20,480 / 20,480 | 20,480 /
20,480 |
| Missing `_amax` | 0 | 0 |
| All-zero `_amax` | 0 | 0 |
| `mtq.quantize` time | 25–34 s | 23 s |

- **Cross-path diff:** every per-expert weight amax matches
**bit-for-bit** between the two paths (`n=20480 exact=20480 diff=0
max_rel=0`). With 8/256 experts routed per token and 4 calib samples,
almost all experts are "dead" in Path A. Bootstrap fills them from
`max(|weight|)`, MSE searches deterministically from there → identical
to Path B which bootstraps everything.
- **Export to HF NVFP4 checkpoint** succeeded (~95 s, 22 GB checkpoint).
Resulting `generation_config.json` has `do_sample: true` (upstream had
`top_k=20` + `top_p=0.95` which would have failed strict validate).
- **TRT-LLM inference loaded the checkpoint and generated text:** `"Born
in north-east France, Soyer trained as a"` → `" tailor. Demonstrating
his craft at a young age, at 20 he moved to Paris at the requests of the
noble people of Picardy."` (coherent grammar; factually wrong as
expected with 4-sample calib, but no NaN/Inf in logits, no
scale-mismatch crash). 92 GB GPU memory used.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ <\!-- relies on existing
recipe-level integration coverage; verified end-to-end on
Qwen3.5-122B-A10B and Qwen3.5-35B-A3B + TRT-LLM 1.3.0rc13 -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌ <\!-- will run \`/claude
review\` -->

### Additional Information

Follow-up to PR #1407 (MSE+FP8-cast-KV recipes). The recipe YAML files
landed there; this PR fixes the calibration codepath so the MSE recipes
actually exercise per-expert weight quantizers in fused-experts MoE
containers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed generation configuration validation for HuggingFace model
exports.
* Improved handling of quantization shape mismatches during expert
weight export.

* **New Features**
* Enhanced calibration process with automatic population of missing
expert quantizers.
* Added grouped quantizer synchronization for improved multi-expert
quantization.

* **Tests**
* Added regression tests for fused expert export and calibration
correctness.

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1421)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-12 12:07:34 -07:00
Chenjie Luo 1d796f9714 [Quantization] Fused Triton kernel for NVFP4 FP8 scale sweep search (#1387)
## Summary

- Replaces the 126-iteration Python sweep in `NVFP4MSECalibrator` with a
fused Triton kernel that loads each NVFP4 block once, evaluates all 126
valid FP8 E4M3 scale candidates in registers, and emits the per-block
``best_amax`` directly.
- Triton-backed `TritonNVFP4MSECalibrator` is the **default** for
`mse_calibrate(..., fp8_scale_sweep=True)`. Set
`MODELOPT_NVFP4_TRITON_SWEEP=0` to fall back to the reference for
debugging or numerics comparison.
- Microbenchmark: **~42x speedup** on a B300 over the reference
(`NVFP4MSECalibrator`) on a representative LLM weight (`8192x4096`, ~2M
NVFP4 blocks): `176.68 ms -> 4.23 ms`.
- End-to-end on Qwen3-8B PTQ (`hf_ptq.py --qformat nvfp4_mse`,
calib=128): **~6.7x faster mtq.quantize** with **identical global
weight-MSE** as the reference.

## End-to-end Qwen3-8B PTQ (B300)

> Note: ran on **Qwen3-8B** instead of Qwen3.5-9B because the docker's
transformers (4.57.3) doesn't yet recognize the `qwen3_5` (multimodal)
architecture and the model dir doesn't ship `modeling_*.py` for
`trust_remote_code`. Qwen3-8B is the same family / similar size and
gives a representative comparison.

Settings: `--calib_size 128 --calib_seq 512`, default `nvfp4_mse` config
(`NVFP4_W4A4_WEIGHT_MSE_FP8_SWEEP_CFG`). The first `nvfp4` run was
discarded as a warm-up for HF weight loading.

| qformat | `mtq.quantize` time | global weight MSE vs FP16 orig | notes
|
|---|---:|---:|---|
| `nvfp4` (no MSE search) | **3.46 s** | 6.363e-6 | max-calibration
baseline |
| `nvfp4_mse` (reference, slow) | **43.42 s** | 4.788e-6 |
`MODELOPT_NVFP4_TRITON_SWEEP=0` |
| `nvfp4_mse` (Triton, fast — this PR)| **6.51 s** | 4.788e-6 | default
in this PR |

- The Triton path's quantized weights produce **bit-identical global
weight MSE** to the reference (4.788e-6 vs 4.788e-6), validating the
kernel on a real model.
- Triton `nvfp4_mse` is **6.67x faster** than the reference `nvfp4_mse`,
and adds only **~3 s** over the no-MSE baseline (vs ~40 s for the
reference) — making the MSE-search option nearly free in practice.
- The MSE search itself reduces weight quantization error by **~25%** vs
plain max-calibration (6.363e-6 → 4.788e-6).

Per-layer MSE CSVs are produced by `tools/debugger/compare_mse_qwen.py`
(one row per Linear weight) for closer inspection if needed.

## Microbenchmark (B300, 8192x4096 weight, ~2M NVFP4 blocks)

```
reference NVFP4MSECalibrator:   176.68 ms
triton  TritonNVFP4MSECalibrator: 4.23 ms
speedup: 41.8x
```

## Why this works

Each candidate is constructed as `valid_fp8_e4m3_value / 448`. With
`block_amax = global_amax * candidate`, the FP8 round-trip on the
per-block scale `block_amax / 6` (using `global_amax / 6` as the FP8
amax) is the **identity** — so the kernel can compute `scale = candidate
* global_amax / 6.0` inline and skip the FP8 cast. This keeps the kernel
runnable on any CUDA + Triton (no `tl.float8e4nv` requirement).

Because every candidate's per-block scale is just a rescaling of the
same input block, all 126 candidates can be evaluated against a single
`[BLOCKS_PER_PROGRAM, BLOCK_SIZE]` tile held in registers — replacing
126 weight-bandwidth passes with 1.

Two follow-on optimizations close the gap to the compute ceiling:
- `@triton.autotune` over `(BLOCKS_PER_PROGRAM, num_warps)` — the
original hand-picked default (BPP=4, num_warps=4) left ~4x on the table;
the best B300 config is `BPP=64, num_warps=8`.
- Drop the sign-handling `tl.where`: FP4 quant preserves sign, so `(w -
w_q)^2 == (|w| - |w_q|)^2` and the kernel works on `|w|` throughout (one
fewer where + negation per element per candidate).

## Files

- `modelopt/torch/kernels/quantization/gemm/nvfp4_fp8_sweep.py` — new
kernel + `nvfp4_fp8_scale_sweep` wrapper, autotuned.
- `modelopt/torch/kernels/quantization/gemm/__init__.py` — wire-in.
- `modelopt/torch/quantization/calib/mse.py` — new
`TritonNVFP4MSECalibrator(NVFP4MSECalibrator)`.
- `modelopt/torch/quantization/model_calib.py` — opt-out env var
(`MODELOPT_NVFP4_TRITON_SWEEP=0`); Triton path is default.
- `tests/gpu/torch/quantization/test_nvfp4_fp8_sweep_kernel.py` — 16 GPU
tests covering parity (across seeds, block counts, dtypes), input
validation, output round-trip, reset, and a wall-clock speedup report.

## Numerics

Bit-identical to the reference for typical block counts (`{4, 64, 1024}`
blocks, 3 seeds, fp32/fp16/bf16 — 14/15 microbenchmark tests bit-exact).

On multi-million-block weights an occasional adjacent-candidate
tie-break can differ at the fp32-noise level (observed **2 / 2,097,152
blocks** in the speedup test, per-block MSE within ~1e-7 relative). The
reference's CUDA `fake_e4m3fy` and the Triton inline math have slightly
different op ordering, which lets nearly-tied candidates flip. The
speedup test asserts the worst per-block MSE gap is `< 1e-5` relative on
differing blocks — both choices are valid argmins; the resulting
quantized weights are equally good. The Qwen3-8B end-to-end run confirms
this: aggregate weight MSE matches the reference exactly at the
displayed precision.

## Test plan

- [x] `pytest
tests/gpu/torch/quantization/test_nvfp4_fp8_sweep_kernel.py -v` — 16/16
pass on B300
- [x] `pytest
tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py -v` —
existing NVFP4 tests still pass (9/9)
- [x] End-to-end PTQ on Qwen3-8B with `--qformat nvfp4_mse`: 6.67x
speedup, identical weight MSE to reference

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
- Added a Triton-based fused NVFP4 FP8 scale-sweep for faster per-block
scale selection.
- Exposed a device-aware FP8 scale candidate generator and a GPU sweep
API used by the calibrator.
- Calibrator now uses the Triton fast-path by default with env var
opt-out (MODELOPT_NVFP4_TRITON_SWEEP).

* **Bug Fixes / Behavior**
- One-shot fast-path enforcement in the NVFP4 calibrator; reset restores
reuse.

* **Tests**
- Added comprehensive GPU tests validating parity, input validation,
dispatch behavior, one-shot semantics, and speed benchmarks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-08 22:53:37 +00:00
Chenjie Luo e2d29c869b [NVBug 6143871] Fix awq_lite uncalibrated branch leaving input_quantizer disabled (#1410)
## Summary

`awq_lite.setup()` disables `module.input_quantizer` at the start of
search. The calibrated branch re-enables it inside `postprocess()`, but
the uncalibrated branch (no cache-pass tokens, e.g. an MoE expert that
never gets routed) never did. Worse, for experts that had cache hits but
missed the search pass, the per-channel `_amax` left over from
`max_calibrate` during cache mode tripped `preprocess_linear_fusion`'s
`numel == 1` assertion and prevented `_export_quantized_weight` from
emitting a per-tensor `input_scale`.

Result: per-expert `input_scale` was missing in the exported HF
checkpoint, and TRT-LLM `CutlassFusedMoE` crashed on load with
`KeyError: '<idx>.w1.input_scale'` for any expert that did not see
enough calibration tokens (e.g. Qwen3-30B-A3B + `nvfp4_awq` from the bug
report).

## Fix

Mirror the calibrated `postprocess()` path in
`modelopt/torch/quantization/model_calib.py`: collapse any per-channel
`_amax` to scalar (axis=None) and re-enable the `input_quantizer`.

## Test plan

- [x] Added regression test
`test_awq_lite_uncalibrated_linear_keeps_input_quantizer_enabled` using
`NVFP4_AWQ_LITE_CFG` with a two-branch model where only one branch is
exercised; verifies the uncalibrated linear's `input_quantizer` remains
enabled after `mtq.quantize`.
- [x] End-to-end pipeline test on a tiny synthetic Qwen3-MoE (8 experts,
top-1 routing) confirms all 48 expert `input_scale` keys are present in
the exported state_dict (vs. multiple missing pre-fix).
- [x] Manual repro of the bug command (Qwen3-30B-A3B, `--qformat
nvfp4_awq`) confirmed the missing `input_scale` keys before the fix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed AWQ-Lite quantization for uncalibrated modules so
export/preprocessing invariants are preserved even when
calibration/parameter updates are skipped.

* **Tests**
* Added regression tests: one verifies uncalibrated linear modules keep
their input quantizer enabled after quantization; another verifies
weight-only AWQ-Lite cases keep the input quantizer disabled as
expected.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-08 04:44:14 +00:00
Chenjie Luo 6a3b6b8329 [Recipes][LLM PTQ] Add nvfp4 MSE+FP8-cast-KV recipes (experts_only / mlp_only) + --recipe in example scripts (#1407)
## Summary

- Adds two PTQ recipes that combine **experts/MLP-only NVFP4 W4A4** with
**MSE FP8 scale-sweep weight calibration** and **FP8 KV cache with
`use_constant_amax: true`** (skips KV calibration; matches the
`nvfp4_default-fp8_cast_kv` contract):
- `modelopt_recipes/general/ptq/nvfp4_experts_only_mse-fp8_cast_kv.yaml`
— applies to `*mlp.experts*` / `*block_sparse_moe*` only.
- `modelopt_recipes/general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv.yaml` —
applies to all `*mlp*` / `*block_sparse_moe*` (dense MLP + MoE).
- Threads a new `--recipe` flag through
`examples/llm_ptq/scripts/parser.sh` and `huggingface_example.sh`.
Either `--quant` or `--recipe` is required; passing **both errors out**.
Recipe names are not validated in the script — `hf_ptq.py` is the source
of truth.
- Drops the bash-side `qformat` whitelist case-statement in
`huggingface_example.sh` for the same reason.

## Files

**New recipes (`modelopt_recipes/general/ptq/`):**
- `nvfp4_experts_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_experts_only-fp8_kv.yaml`.
- `nvfp4_mlp_only_mse-fp8_cast_kv.yaml` — same patterns as
`nvfp4_mlp_only-fp8_kv.yaml`.

Both differ from their `_kv` siblings by:
- `algorithm: max` → `{ method: mse, fp8_scale_sweep: true, layerwise:
false }`
- All targeted **weight quantizers** switch `type: dynamic` → `type:
static` (otherwise `mse_calibrate` skips them: only static block-quant
weight quantizers are recognized for the FP8 sweep — see
`model_calib.py:369-374`).
- Input quantizers stay dynamic.
- KV bmm adds `use_constant_amax: true` (the `_cast_kv` flavor).

**Scripts (`examples/llm_ptq/scripts/`):**
- `parser.sh` — adds `--recipe` long-option, default `RECIPE=""`,
validates one-of-{`--quant`, `--recipe`} and not-both.
- `huggingface_example.sh` — when `RECIPE` is set, derives `MODEL_NAME`
from the recipe basename, passes `--recipe=…` to `hf_ptq.py` instead of
`--qformat=…`, and exits after export with a TRT-LLM deployment hint
(recipes can produce arbitrary configs that the script's downstream
`run_tensorrt_llm.py` path doesn't know how to handle generically).
Drops the `qformat` whitelist; defers to `hf_ptq.py`.

## Behavior

```
# Errors with: "Cannot specify both --quant and --recipe; pick one."
bash huggingface_example.sh --model=... --quant=nvfp4 --recipe=... --tasks=quant

# Errors with usage if neither is given
bash huggingface_example.sh --model=... --tasks=quant

# Both of these are now accepted; --recipe is forwarded verbatim to hf_ptq.py
bash huggingface_example.sh --model=... --quant=nvfp4 --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_experts_only_mse-fp8_cast_kv --tasks=quant
bash huggingface_example.sh --model=... --recipe=general/ptq/nvfp4_mlp_only_mse-fp8_cast_kv  --tasks=quant
```

## Test plan

- [x] `experts_only_mse-fp8_cast_kv` loads via
`modelopt.recipe.load_recipe(...)` and produces the expected algorithm +
per-pattern `quant_cfg` (verified in a working env: `algorithm ==
{'method': 'mse', 'fp8_scale_sweep': True, 'layerwise': False}`; expert
weight quantizers `type: static`; KV bmm has `use_constant_amax: True`).
- [x] Parser sanity: 4 flag combinations (both, neither, only `--quant`,
only `--recipe`) all behave as designed.

## Note

Pre-commit hook `check-modelopt-recipes` was skipped on both commits
because the local conda env has a broken `torchvision` install
(`AttributeError: partially initialized module 'torchvision' has no
attribute 'extension'`) that prevents `from modelopt.recipe.loader
import load_recipe`. The `experts_only` recipe was validated
independently by running `tools/precommit/check_modelopt_recipes.py` in
a working environment (exits 0); the `mlp_only` one is the same shape
with a different glob.

Rebased onto `main` from #1391 (which targeted
`chenjiel/nvfp4-fp8-sweep-triton`). The diff is scoped to the recipes +
script wiring; no kernel/sweep changes are included here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added recipe-based quantization as an alternative to format-based
quantization with a new `--recipe` CLI option.
* Added two new quantization recipes for targeted layer optimization:
one for expert-layer-only quantization and one for MLP-layer-only
quantization, both featuring NVFP4 and FP8 KV-cache optimization.

* **Configuration**
* `--quant` and `--recipe` options are now mutually exclusive; specify
one to configure quantization behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-07 15:38:47 -07:00
Chenjie Luo 097293b44d [Quantization] Saturate NVFP4 export FP8 scale cast to avoid NaN (#1397)
## Summary

- Saturates `per_block_scale * 448 / per_block_scale_max` to ≤ 448
before the `to(torch.float8_e4m3fn)` cast in
`NVFP4QTensor.get_weights_scaling_factor_from_quantizer`.
- Adds a regression test that reproduces the NaN byte without the clamp.

## Why

When `_amax` contains a zero entry (e.g. an all-zero weight block left
untouched by max calibration), the existing
`per_block_scale[per_block_scale == 0] = 1.0` safety net drives the
pre-cast value to `1.0 * 448 / (global_amax / 6)`. `fp8_e4m3fn` has no
Inf — anything `≥ 480` rounds to NaN — so a 0x7F byte slips into the
exported `weight_scale`.

This was observed in a saved Kimi-K2.6-NVFP4-MSE checkpoint at
`language_model.model.layers.1.mlp.experts.21.down_proj.weight_scale[4001,
18]`. The MSE FP8 sweep itself never produces zero per-block amax (it
always emits at least `c[0] * global_amax`), but any export path where
`_amax` ends up zero — including pure max calibration — hits the bug.
With the clamp the byte saturates to `0x7E` (= 448, fp8 max finite) and
dequantization is unaffected: the FP4 nibbles for an all-zero block are
all 0, so `0 × 448 × weight_scale_2 = 0` regardless of the stored fp8
scale. For non-degenerate blocks the clamp is a no-op since
`per_block_amax ≤ global_amax` already bounds the pre-cast value at 448.

## Test plan

- [x] New regression test
`test_export_fp8_scale_no_nan_for_zero_amax_block` fails on `main`'s
export code (reproduces the 0x7F NaN byte) and passes with the clamp.
- [x] Existing tests in
`tests/gpu/torch/quantization/test_nvfp4_static_quantizer_cuda.py` still
pass (10/10).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved numerical stability in FP8 quantization scaling by preventing
overflow and NaN conditions
* Enhanced handling of edge cases in quantization processing for
zero-weight blocks

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-05-06 08:33:18 -07:00
Chenjie Luo b5df2a5a3c Update llm_ptq requirements.txt (#1394)
### What does this PR do?

Type of change: Dependency update

compressed_tensors 0.15 is not compatible with our current quant
implementation for Kimi K2.5, K2.6

### Testing
Unittest

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Broadened the `compressed-tensors` dependency constraint to allow a
wider range of compatible versions.
  * Removed the `rouge_score` dependency from the example requirements.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-05-05 17:00:57 +00:00
Chenjie Luo 1d21ab9e29 [DeepSeek] Default to top-k calibration with peer-max input amax sync (#1380)
## Summary

- DeepSeek PTQ (`examples/deepseek/ptq.py`) now defaults to native top-k
routing during MoE calibration. The previous all-tokens-to-all-experts
path (`CalibMoe`) is preserved behind a new `--calib_all_experts` flag.
- After `mtq.quantize`, `fixup_moe_expert_amax` syncs every expert's
`input_quantizer.amax` (w1/w2/w3) to the per-layer global peer max via
`dist.all_reduce(MAX)` across EP ranks. `weight_quantizer.amax` stays
per-expert; any uncalibrated expert is filled by computing amax over the
dequantized FP8 weight.
- `mtq.print_quant_summary` is now also written to
`<output_path>/.quant_summary.txt`, mirroring `llm_ptq/hf_ptq.py`.

## Why

Forcing all tokens through every expert doubled calibration time and
inflated `input_quantizer.amax` for cold-routing experts with outliers
they never see at inference. The new flow matches the inference
distribution, runs roughly 2x faster, and mirrors the
`layer_sync_moe_local_experts_amax` semantics that mtq runs
automatically for `QuantSequentialMLP`-derived MoEs.

## Validation (DeepSeek-V3.2-Exp, MP=8, NVFP4_DEFAULT_CFG)

Compared `_amax_baseline` (CalibMoe) vs `_amax_synced` (new default):
- All 44,544 expert weight amaxes bit-identical.
- Attention, shared experts, gate: identical.
- Expert `w1.input` and `w3.input` (shared MoE block input): identical.
- Expert `w2.input` (post-SiLU gated, expert-specific): synced to
layer-wide peer max — 99.3% are larger than baseline (median 11.4x)
since peer-max captures the worst-case outlier from any expert in the
layer; 0.7% are smaller. This is the same trade-off
`set_expert_quantizer_amax` makes for HF MoEs in `unified_export_hf.py`.

## Test plan

- [x] DeepSeek-V3.2-Exp MP8 PTQ with default flags — completes in ~7 min
(vs ~27 min with CalibMoe), produces `_amax_synced/` consistent with the
comparison above.
- [x] DeepSeek-V3.2-Exp MP8 PTQ with `--calib_all_experts` — produces
`_amax_baseline/` identical (other than rounding) to the prior
`CalibMoe`-default behavior.
- [x] `.quant_summary.txt` written under `output_path` on rank 0.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--calib_all_experts` option to enable an alternate PTQ
calibration mode; default remains top-k routing with a post-calibration
per-layer peer-max synchronization and a compute fallback for
uncalibrated experts.
* **Documentation**
* Clarified default and alternate calibration behaviors and added note
about generation of a `.quant_summary.txt` summary file.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-05-04 11:08:01 -07:00
Chenjie LuoandClaude Opus 4.7 50706d1750 Add closed-form MXFP4 -> NVFP4 weight cast (--cast_mxfp4_to_nvfp4) (#1372)
## Summary

- New `--cast_mxfp4_to_nvfp4` flag in `hf_ptq.py` (and
`huggingface_example.sh`) that converts an MXFP4 source checkpoint (e.g.
`openai/gpt-oss-20b`) into an NVFP4 export with **bit-exact** weight
reconstruction for the in-range blocks.
- The cast pins NVFP4's `scale_2 = 2^m` (where `m = k_max − 8`) and
`_amax = 6·2^k_j` per NVFP4 block, both read from the source `*_scales`.
The resulting per-block scale `2^(k_j − m)` is exactly representable in
E4M3, so `round_to_E2M1(value / 2^k_j)` yields the original MXFP4 nibble
verbatim. For out-of-range blocks (`k_max − k_j > 17`) the per-block
amax falls back to data-derived `max(|w_block|)`, which keeps the
post-E4M3-clamp scale close to the block's actual magnitude.

## Verification

End-to-end on `openai/gpt-oss-20b` with `--qformat=nvfp4_mlp_only
--cast_mxfp4_to_nvfp4`:

```
[cast_mxfp4_to_nvfp4] overrode 48/48 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 48/48 (100.00%)
[cast_mxfp4_to_nvfp4] lossless blocks: 597196800/597196800 (100.0000%)
```

End-to-end on `openai/gpt-oss-120b` with the same flags (4×B200,
`--use_seq_device_map --gpu_max_mem_percentage 0.5 --calib_batch_size
4`):

```
[cast_mxfp4_to_nvfp4] overrode 72/72 weight quantizers
[cast_mxfp4_to_nvfp4] lossless layers: 67/72 (93.06%)
[cast_mxfp4_to_nvfp4] lossless blocks: 3583179586/3583180800 (100.0000%)
```

Five layers fall into the OOR regime (block-spread > 17); the remaining
1,214 OOR blocks use the data-derived per-block amax fallback.
Block-level losslessness is **99.99996%** end-to-end.

Per-tensor MSE between MXFP4 source dequant and NVFP4 export dequant
(~19B elements):

| Metric | Without cast | With cast |
|---|---|---|
| Per-tensor SNR | ~26.4 dB (FP4 noise floor) | **∞ (every tensor)** |
| Total RMSE | 8.67e−02 | **0** |
| max\|err\| | up to 8.0e+1 | **0** |

## Modelopt-side enablers

- `max_calibrate` auto-promotes static-block NVFP4 weight quantizers to
`NVFP4StaticQuantizer` at the end of calibration.
- `static_blockwise_fp4_fake_quant` kernel accepts N-D inputs (was
2D-only), unblocking MoE expert weights of shape `(E, F, K)`.
- BMM-experts NVFP4 export routes through
`get_weights_scaling_factor_from_quantizer` for static-mode quantizers,
so the pinned `_amax` is actually consumed.
- `set_expert_quantizer_amax` scalar-reduces per-quantizer amax before
stacking, supporting per-block (vs scalar) static-mode amax.

## Test plan

- [x] Unit tests at `tests/examples/llm_ptq/test_cast_mxfp4_to_nvfp4.py`
(15 tests, all passing) cover: scalar/global-amax math, per-block hybrid
(in-range closed-form vs OOR data-derived), shape preservation, key
collection, and end-to-end `build_amax_map` against a synthetic
safetensors checkpoint.
- [x] End-to-end PTQ → export on `openai/gpt-oss-20b` (`nvfp4_mlp_only`
qformat) with `--cast_mxfp4_to_nvfp4` succeeds; export takes ~21 s. 100%
lossless cast (48/48 layers, 597,196,800 / 597,196,800 blocks).
- [x] End-to-end PTQ → export on `openai/gpt-oss-120b` (4×B200,
`nvfp4_mlp_only`, `--use_seq_device_map --gpu_max_mem_percentage 0.5
--calib_batch_size 4`). 67/72 layers fully lossless; 99.99996%
block-level losslessness (3,583,179,586 / 3,583,180,800).
- [x] TRT-LLM serving validation (TRT-LLM 1.3.0rc11, B200) on both
exported NVFP4 checkpoints via `examples/llm_ptq/run_tensorrt_llm.py`:
- **20b** (TP=1): 18.3 GB GPU memory; coherent generation. Sample:
*"Quantum computing is poised to revolutionize data analysis. However,
its potential is currently limited by quantum hardware constraints,
including error rates, qubit lifetimes, and lack of fault tolerance…"*
- **120b** (TP=4): 36.4 GB / GPU; coherent generation. Sample: *"Quantum
computing is poised to revolutionize data storage and processing. These
rare earth-based systems could serve as robust qubits; resistant to
environmental decoherence…"*
- [x] MSE comparison script (run separately during development) confirms
per-tensor SNR=∞ across all 48 MoE expert tensors.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a MXFP4→NVFP4 weight-format cast utility and a CLI flag to
enable it; helper scripts updated to expose the option.

* **Bug Fixes**
  * Fixed static NVFP4 export for expert weights.
* Improved collection/handling of quantizer amax values to avoid shape
issues.
  * Generalized FP4 kernel to accept flexible tensor dimensionality.
  * Ensured static-block NVFP4 promotion during calibration.

* **Tests**
* Added comprehensive tests for the conversion workflow, helpers, and
end-to-end application.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 19:04:04 +00:00
Chenjie LuoandKeval Morabia 168cd828c1 Add qwen3 moe experts only test (#1274)
## Summary
- Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG`
quantization config
- Verifies that `hf_quant_config.json` correctly reports `quant_algo:
NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in
`exclude_modules` while routed expert layers (`mlp.experts.*`) do not
- Reference:
https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json

Type of change: New tests

### Known issue
On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not
recognized by `get_quant_config`, causing `quant_algo=None` in the
exported config. This test currently **fails** on transformers 5.x and
is intended to be fixed by a follow-up change.

## Testing
- **transformers 4.57.6**: PASSED
- **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused
expert export gap)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models
with NVFP4 quantization.
* Verifies the exported checkpoint records the NVFP4 quantization
algorithm and that module exclusion patterns correctly exclude attention
and LM head components while not excluding routed expert paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-01 17:56:04 +00:00
Chenjie Luo 3ad4f4f093 [Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary
- `get_max_batch_size` halved `target_data_batch` on
`torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each
retry re-fed the same too-large tensor — the retry loop was effectively
a no-op.
- Refactor the expand logic into an `_expand_to(batch)` helper, rebuild
`target_input` after halving, and call `torch.cuda.empty_cache()`
between attempts.

## Test plan
- [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input`
mocks `torch.cuda.*` and asserts the second retry receives the halved
tensor (shapes seen: `[1, 10, 5]`, regulated result `4`).
- [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass
(skipping the network-only minipile test).
- [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on
commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced GPU memory management during batch size detection. When
out-of-memory errors occur during the initial probing phase, the system
now properly adapts input tensors to smaller batch sizes and clears GPU
cache before retry attempts, resulting in more reliable recovery and
stable batch sizing across diverse hardware environments.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-30 18:25:28 +00:00
Chenjie Luo 47a33db9b6 [NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary

- NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) —
`Qwen3-30B-A3B nvfp4_awq` quantization fails at export with
`AssertionError: Modules have different quantization formats`.
- Root cause: in `model_calib.awq_lite`, MoE experts that end up
disabled (NaN in act/weight scales, or no search-pass tokens) get
`max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format`
then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`.
`unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128
experts of each linear name (gate_proj/down_proj/up_proj) and calls
`preprocess_linear_fusion(..., resmooth_only=True)`, which asserts
uniform format → fires for any single mismatched expert.
- Fix: unify the disabled-expert paths in the awq_lite postprocess loop
so any expert with `is_enabled == False` (no cache hits, NaN scales, or
no search-pass tokens) receives `max_calibrate` + a neutral all-ones
`pre_quant_scale`, matching the existing behavior for `num_cache_steps
== 0`. Emit a warning so users notice that calibration coverage is
incomplete and accuracy may degrade.

## Test plan

- [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5
passed
- [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a
small calib set that leaves many experts uncalibrated:
- All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq`
(no mismatch)
  - `export_hf_checkpoint` succeeds with no `AssertionError`
- The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning
fires for each affected expert
- [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI
(cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to
fit budget) on B200:
- 36 "the second time did not forward data through
..experts.X.{gate,up,down}_proj" warnings — i.e. the exact
bug-triggering condition from the original NVBug log naturally
reproduces
- 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for
uncalibrated/disabled experts
  - 0 `AssertionError`s — export completes
- `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq`
and post-PTQ generation works

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-27 12:40:14 -07:00
Chenjie Luo fda0899e40 feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary

- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
  - `general/ptq/fp8_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.

## Motivation

Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).

## Test plan

- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.

* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.

* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 18:45:16 +00:00
Chenjie Luo 01788bb007 Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary

- Removes Mllama (Llama 3.2 Vision) model-type branches from the
`llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the
now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`.
- Drops the legacy `MllamaImageProcessor` path in
`modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF
ProcessorMixin path handles the remaining cases.
- Adds a CHANGELOG entry under 0.44 Backward Breaking Changes.

## Test plan

- [x] CI lint / unit tests pass
- [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh
--model <llm> --quant fp8`` (text-only path, non-mllama)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Removed Mllama (Llama 3.2 Vision) support from quantization examples.
This includes removal of dedicated image processor implementation,
specialized model handling, and related calibration logic.
* Updated VLM image-text calibration guidance to use
`--calib_with_images` flag with other supported VLMs instead of
Mllama-specific processing paths.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 23:24:42 +05:30
Chenjie Luo 04fcf24227 Fix LLM deploy test failure by defaulting expert parallelism to 1 (#1273)
### What does this PR do?

Type of change: Bug fix

Fixes TRT-LLM DeepEP kernel failures during LLM deployment on
unsupported GPUs (e.g. Blackwell SM 12.0) by defaulting expert
parallelism (`ep`) to 1 instead of auto-setting it to the GPU count for
MoE models.

Previously, when the model config contained expert-related keys, `ep`
was automatically set to `torch.cuda.device_count()`, which triggered
DeepEP kernel failures on GPUs that don't support it. Now `ep` defaults
to 1 while still enabling attention data parallelism for MoE models.
Expert parallelism can be enabled explicitly by the caller when the
environment is known to support it.

### Testing

- [x] Verified that the `llm_ptq` test passes with this fix on Blackwell
GPUs.
- [x] 2-gpu CI test triggered:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24495054531/job/71588037727

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-16 16:18:38 -07:00
Chenjie Luo d45219b390 Fix debugger server failing to detect editable-installed modelopt (#1270)
## Summary
- Removed `PYTHONPATH="" python -I` override in `check_modelopt_local()`
so the PYTHONPATH validation uses the actual environment instead of an
isolated one
- Moved the workdir log line earlier in `server.sh` for better debugging
visibility

## Test plan
- [x] Start `server.sh` inside a Docker container and verify it
correctly detects editable-installed modelopt
- [x] Confirm the workdir is logged before the modelopt check runs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Server startup now displays the configured work directory earlier in
the initialization process, providing improved visibility of the active
directory during server launch.
* Simplified the modelopt validation check during server initialization
while maintaining the same validation behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 22:53:46 -07:00
Chenjie Luo 7c8557158d Add job cancellation support to the debugger command relay (#1262)
## Summary

- Add a `cancel` subcommand to the client that terminates the currently
running command on the server
- Server now runs commands in the background with PID tracking, enabling
cancellation mid-execution
- Client-side timeouts automatically cancel the running command on the
server (previously the server process was left running)
- Hardened against race conditions through 4 rounds of adversarial
review (15 fixes total)

### Key changes

**server.sh:**
- Commands run in background with PID tracked in `$RELAY_DIR/running`
(atomic tmp+mv write)
- Cancel detection loop checks for `$RELAY_DIR/cancel` file with cmd_id
verification
- SIGTERM with 5s grace period, then SIGKILL escalation for stuck
processes
- `.exit` file written before `running` marker removed (ordering
guarantee)
- `set -e`-safe: `wait` uses `|| exit_code=$?` pattern; cleanup trap
fully guarded
- Stale cancel files cleared at command start; mismatched/empty signals
rejected
- Command file read into memory and removed before execution (eliminates
TOCTOU with client timeout)

**client.sh:**
- New `cancel` subcommand: writes target cmd_id to cancel file, waits
for server acknowledgment (30s timeout)
- `run` timeout now sends targeted cancel signal (verifies cmd_id match
to avoid killing wrong command)
- `run` timeout cleans up orphaned result files
- `status` shows currently running command
- `flush` rejects if a command is currently running (prevents state
corruption)
- Exit code validated as numeric before use

### Protocol additions

```
.relay/
├── running    # server writes cmd_id:pid while executing (atomic)
├── cancel     # client writes target cmd_id to request cancellation
```

## Test plan

- [ ] Start server in Docker, handshake from host
- [ ] Run a command (`client.sh run "sleep 30"`), cancel it (`client.sh
cancel`), verify exit code 130
- [ ] Run a command with short timeout (`--timeout 5 run "sleep 30"`),
verify auto-cancel
- [ ] Run a command that exits non-zero, verify server stays alive
- [ ] Run `status` during execution, verify it shows the running command
- [ ] Attempt `flush` during execution, verify it is rejected
- [ ] Cancel when nothing is running, verify clean message

🤖 Generated with [Claude Code](https://claude.com/claude-code)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (bash scripts for dev
tooling, tested manually)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A (internal tooling)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive debug skill guide and protocol reference with
quick CLI examples and a new “Cancelling Commands” section.

* **New Features**
* Client-side `cancel` command to terminate the currently running remote
command.
  * Status now reports active command (or `(idle)`).

* **Improvements**
* Stronger startup/validation guidance, safer shutdown/cleanup,
deterministic cancel exit semantics (130), and auto-cancel on
client-side timeout.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-15 16:49:42 +00:00
Chenjie LuoandClaude Opus 4.6 952a62bf65 Fix missing attention_mask in calibration dataloader (#1261)
## Summary
- When `include_labels=False` (the default for PTQ calibration),
`get_dataset_dataloader` was discarding the `attention_mask` produced by
the tokenizer and only returning `input_ids`.
- Without `attention_mask`, HuggingFace models create a full causal
mask, causing padding tokens to participate in attention during
calibration and skewing quantization statistics.
- This fix includes `attention_mask` alongside `input_ids` so the model
correctly ignores padding tokens during calibration forward passes.

## Details
In `modelopt/torch/utils/dataset_utils.py`, the tokenizer call at line
387 with `padding=True` produces both `input_ids` and `attention_mask`.
The `include_labels=True` path (line 406) already preserves the full
`batch_encoded` dict including `attention_mask`. However, the
`include_labels=False` path was only keeping `input_ids` "for backward
compatibility."

During the calibration forward loop (`_forward_loop` →
`_process_batch`), the batch dict is unpacked as `**kwargs` into
`model.forward()`. Without `attention_mask`, HF models default to
attending to all positions including padding, which pollutes calibration
statistics.

**Practical impact**: With `batch_size=1` there is no padding so the bug
is invisible. With larger batch sizes and variable-length samples,
shorter sequences get padded and the effect grows.

## Test plan
- [x] Existing unit tests pass
(`tests/unit/torch/utils/test_dataset_utils.py`)
- [x] Pre-commit hooks pass
- [ ] Verify PTQ accuracy with batch_size > 1 on a padded calibration
dataset (GPU required)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 00:04:27 -07:00
Chenjie Luo f7557221e3 Add file-based command relay for remote Docker testing (#1174)
## Summary
- Adds a lightweight file-based client/server relay (`tools/debugger/`)
that enables Claude Code (or any host-side automation) to execute
commands inside a remote Docker container using only a shared filesystem
— no networking setup required.
- The server auto-detects the repo root, installs modelopt (`pip install
-e .[dev]`), sets `PYTHONPATH`, and listens for commands.
- The client supports `handshake`, `run`, `status`, and `flush`
subcommands.
- Includes `README.md` (full protocol docs) and `CLAUDE.md` (quick
reference for Claude Code).

## Tested
- Ran Qwen3.5-35B-A3B MoE PTQ with `nvfp4_experts_only` quantization via
the relay:
  ```
bash examples/llm_ptq/scripts/huggingface_example.sh --model
/hf-local/Qwen/Qwen3.5-35B-A3B/ --quant nvfp4_experts_only
  ```
  - 42,140 quantizers inserted, MTP layers correctly excluded
- Quantized checkpoint exported successfully (~208s, 147.61 GB peak GPU
memory)

## Test plan
- [x] Start `server.sh` inside a Docker container with the repo mounted
- [x] Run `client.sh handshake` from the host
- [x] Run `client.sh run "echo hello"` and verify output
- [x] Run `client.sh flush` and verify `.relay/` is cleared
- [x] Run a real PTQ workload end-to-end

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a file-based command relay system with host↔container client and
server CLIs, supporting handshake, run, status and flush workflows to
execute commands inside containers.

* **Documentation**
* Added guides describing the relay protocol, usage examples, CLI
options, lifecycle, and operational notes (workdir, timeouts, sequential
execution).

* **Chores**
  * Updated ignore rules to exclude ephemeral relay artifacts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-05 22:44:33 -07:00
Chenjie Luo 18ce04f1ce Update the hf_ptq.yaml (#1175)
### What does this PR do?

Type of change: Bug fix

Fix the CLI override example comment in
`tools/launcher/examples/Qwen/Qwen3-8B/hf_ptq.yaml` by adding a missing
`--` (double-dash separator) to the `task_0.args` override.

Without the `--` separator, the `--quant` flag would be parsed as an
argument to the launcher/download script rather than being passed
through to the PTQ script (`huggingface_example.sh`). This aligns the
comment example with the actual `task_0.args` definition in the YAML
(line 42), which already correctly includes the `--` separator.

### Testing

Verified the comment now matches the actual args format used in the YAML
config.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example configuration to correct command-line parameter
formatting in commented usage example.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-03 19:04:33 +00:00
Chenjie Luo 87ea8babe1 Add HuggingFace PTQ pipeline to launcher (#1100)
### What does this PR do?

Type of change: New feature

Adds a HuggingFace PTQ pipeline to the launcher, replacing the old
`hf_ptq.sh`/`hf_ptq_local.yaml` approach with a cleaner wrapper around
`huggingface_example.sh`.

**Key changes:**

- **New `common/hf/ptq.sh`** — wrapper script that downloads the model
via `huggingface-cli` if needed, then delegates to
`examples/llm_ptq/scripts/huggingface_example.sh`
- **New `examples/Qwen/Qwen3-8B/hf_ptq.yaml`** — example config for
Qwen3-8B nvfp4 quantization, supports both Slurm and local Docker
- **Removed `common/hf_ptq/hf_ptq.sh`** and
**`examples/Qwen/Qwen3-8B/hf_ptq_local.yaml`** — replaced by the new
unified pipeline
- **Configurable Slurm time limit** — `SlurmConfig.time` field replaces
the hardcoded `"04:00:00"` in `build_slurm_executor`
- **Configurable Slurm partition** — `slurm_factory` now reads
`SLURM_PARTITION` env var (default: `batch`)
- **`--clean` flag** — new `launch.py` option to `git clean -xdf` the
examples directory before job submission
- **Package `modelopt_recipes/`** — added to the nemo_run packager
include list

### Testing

- Tested HF PTQ pipeline on Slurm with Qwen3-8B

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Hugging Face PTQ workflow configuration for Qwen model
quantization.
* Added `clean` parameter to launcher for clearing directories before
job execution.

* **Improvements**
* Made Slurm execution time configurable per job instead of hardcoded
values.
  * Slurm partition configuration now respects environment variables.

* **Deprecated**
* Removed legacy PTQ launcher scripts, replaced with unified wrapper for
improved maintainability.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-02 18:55:38 +00:00
Chenjie Luo fcb09bf11d [NVBug: 6038899] Fix MoE export crash on meta tensors with CPU offload (#1155)
## Summary
Fixes `NotImplementedError` in `sync_moe_gate_up_amax` when quantizing
MoE models (e.g. Qwen3-30B-A3B) on a single GPU with insufficient VRAM.

When GPU memory is insufficient, ModelOpt enables CPU offload via
accelerate, leaving uncalibrated expert parameters on the `meta` device.
During export, `sync_moe_gate_up_amax` calls `torch.equal()` on these
meta tensors, which raises `NotImplementedError` because `aten::equal`
does not support meta tensors — even though calibration itself completed
successfully.

## Changes
- Add a guard in `sync_moe_gate_up_amax` to skip amax sync for meta
tensors (which have no real data to sync) and emit a warning explaining
the root cause.

Bug: https://nvbugspro.nvidia.com/bug/6038899

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Added warning messages for unsupported tensor configurations in
quantization workflows.
* Improved edge case detection to gracefully skip processing in
incompatible scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-02 06:59:57 +00:00
Chenjie LuoandClaude Opus 4.6 ada1e26ba0 [NVBug: 6000530] Fix AWQ crash for uncalibrated MoE experts (#1142)
## Summary
- Fixes NVBugs 6000530: `AttributeError: 'float' object has no attribute
'pow'` when running AWQ lite with `moe_calib_experts_ratio < 1.0` on MoE
models (e.g. Qwen3-30B-A3B).
- **Root cause**: When `moe_calib_experts_ratio=0.5`, some MoE experts
receive zero tokens during the AWQ cache phase, leaving `act_scale` as a
Python float `0.0` instead of a tensor. This causes two failures:
1. **Search phase crash**: Uncalibrated experts crash in `get_scale()`
because `float.pow()` doesn't exist.
2. **Export crash**: Calibrated experts have `pre_quant_scale` but
uncalibrated ones don't, causing `torch.stack()` to fail on mixed
`None`/tensor values in `preprocess_linear_fusion()`.
- **Fix**: Handle uncalibrated experts (`num_cache_steps == 0`) in two
stages:
1. **Before search**: Disable AWQ search (`is_enabled = False`) to
prevent `get_scale()` crash on float `act_scale`.
2. **During postprocessing**: Max calibrate weights and apply a neutral
(all-ones) `pre_quant_scale` so export can stack scaling factors
consistently across all experts. The `pre_quant_scale` buffer must be
registered outside `enable_weight_access_and_writeback` because HF
accelerate's `post_forward` hook drops newly-registered submodule
buffers.

## Test plan
- [x] Reproduce with `Qwen/Qwen3-30B-A3B`, `--qformat int4_awq`,
`--moe_calib_experts_ratio 0.5` — verify no crash during calibration and
export

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 13:55:21 -07:00
Chenjie LuoandKeval Morabia 16203d676a [NVBug 6007314] Deprecate MT-Bench support, remove openai pin, and add NeMo Evaluator reference (#1116)
### What does this PR do?

Type of change: Deprecation, Bug fix, Documentation

Removes MT-Bench (FastChat) evaluation support from `examples/llm_eval`
and `examples/llm_ptq`. Also removes the stale `openai>=0.28.1` pin from
`requirements.txt` that caused dependency conflicts with TRT-LLM (see
[NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314)). Adds a NeMo
Evaluator section to the llm_eval README as the recommended evaluation
workflow for quantized checkpoints.

**Changes:**
- Delete `examples/llm_eval/run_fastchat.sh` and
`examples/llm_eval/gen_model_answer.py`
- Remove `mtbench` task from `examples/llm_ptq/scripts/parser.sh` and
`huggingface_example.sh`
- Remove `openai` dependency from `examples/llm_eval/requirements.txt`
- Add NeMo Evaluator section to `examples/llm_eval/README.md` as the
recommended way to evaluate quantized checkpoints from llm_ptq via
TensorRT-LLM, vLLM, or SGLang
- Update README docs in both `llm_eval` and `llm_ptq`
- Add deprecation note to CHANGELOG.rst for 0.43

### Usage

N/A — this is a removal and documentation update.

### Testing

N/A — removed code paths; no new functionality introduced.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — MT-Bench evaluation via
`--tasks mtbench` is no longer supported.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information

Related: [NVBug 6007314](https://nvbugspro.nvidia.com/bug/6007314) —
openai dependency conflict caused by FastChat's `llm_judge` extra
pinning `openai<1`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Deprecations**
* Removed MT-Bench (FastChat) evaluation support. NeMo Evaluator is now
the recommended approach for evaluating quantized model checkpoints
across multiple benchmarks.

* **Documentation**
* Updated evaluation guides to reflect NeMo Evaluator as the primary
evaluation method, with support for TensorRT-LLM, vLLM, and SGLang
serving backends.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-25 10:22:32 +00:00
Chenjie Luo 07bc4852a8 Default limit max hf quantzed safetensors file to 10GB (#1087)
### What does this PR do?

Type of change: New feature

Adds a `max_shard_size` parameter to `export_hf_checkpoint()` (and the
internal `_export_diffusers_checkpoint()`) that controls the maximum
size of each exported safetensors shard file. Defaults to `"10GB"`.

Previously, the shard size was not explicitly controlled, which could
result in very large single-file checkpoints [E.g. Qwen3.5] that are
difficult to handle (e.g., slow uploads, memory issues, or exceeding
file size limits on model hubs). This change ensures exported
checkpoints are automatically sharded into ≤10GB files by default, while
allowing users to customize the threshold.

### Usage

```python
import modelopt.torch.quantization as mtq
from modelopt.torch.export import export_hf_checkpoint

# Default: shards capped at 10GB
export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output")

# Custom shard size
export_hf_checkpoint(model, dtype=torch.float16, export_dir="./output", max_shard_size="5GB")
```

### Testing

Verified that the `max_shard_size` parameter is correctly propagated to
both the HF `save_pretrained()` (for transformers models) and diffusers
component `save_pretrained()` calls.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (new optional parameter with
default matching previous behavior of HF's `save_pretrained`)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information

The 10GB default was chosen to keep exported safetensors files within
common file size limits while minimizing unnecessary sharding for most
models.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added configurable maximum shard size parameter for Hugging Face
checkpoint exports. When exporting both transformer and diffusion
models, users can now specify the maximum size of each safetensors
shard, controlling how checkpoint files are divided. The default shard
size is set to 10GB. This feature works with both quantized and
non-quantized models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-20 16:11:32 -07:00
Chenjie Luo acce79ffa3 Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?

Type of change: New feature

Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.

Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config

### Usage

```python
import modelopt.torch.quantization as mtq

model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```

Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe

recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```

### Testing

- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->

### Additional Information

The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.

* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-03-20 05:36:37 +00:00
Chenjie Luo 1dc890d971 Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0

## Summary

- **Remove `moe_count_expert_calib_tokens`** config field and the
`_moe_count_expert_calib_tokens` internal flag. Token counting is now
implicitly enabled when `moe_calib_experts_ratio` is set, removing a
redundant knob.
- **Change `--moe_calib_experts_ratio` default to `None`** in
`hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by
default; now the feature is opt-in and non-MoE models are unaffected
without any flag.
- **Disable `layer_sync_moe_local_experts_amax`** when
`moe_calib_experts_ratio` is set, since each expert is calibrated
independently with sufficient token coverage in that mode.
- **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks
on `_moe_calib_experts_ratio` inside the branch that already assumes it
is set.

## Changed files

| File | Change |
|------|--------|
| `modelopt/torch/quantization/config.py` | Remove
`moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio`
description to document amax sync behavior |
| `modelopt/torch/quantization/mode.py` | Remove
`moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` |
| `modelopt/torch/quantization/plugins/huggingface.py` | Remove
`_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify
`forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set |
| `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to
`None`; guard validation |
| `tests/unit/.../test_sparse_moe.py` | Update tests to use
`_moe_calib_experts_ratio` instead of removed flag |

## Test plan

- [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio`
(non-MoE model, default `None`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Configuration Changes**
* moe_calib_experts_ratio now defaults to None (disabled) instead of
1.0; validation only occurs when a value is provided.

* **Refactor**
* Simplified MoE calibration flow and token-counting behavior; removed a
deprecated expert-calibration configuration field.

* **Documentation**
* Changelog and docstrings updated to reflect the new default and
calibration behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-18 10:09:24 -07:00
Chenjie Luo 42482b1b0f Add nvfp4_omlp_only config and simplify the config.py (#973)
### What does this PR do?

Type of change: ? new feature

1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant
2) Add block sparse MOE to mlp only config
3) Simplfiy config.py
4) Update readme in llm_ptq mention these two configs for better
accuracy.

### Usage

huggingface_script.sh ... --quant nvfp4_omlp_only


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added nvfp4_omlp_only quantization format for NVFP4, enabling
selective quantization of MLP and output projection layers while
preserving attention QKV projection accuracy.

* **Changed**
* pass_through_bwd now defaults to True; set to False if using STE with
zeroed outlier gradients for better QAT accuracy.

* **Documentation**
* Updated post-training quantization guidance with NVFP4-specific
configuration recommendations and usage examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-06 01:03:25 +00:00
Chenjie Luo a4fde491cc Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do?

Type of change: Bug fix

Add moe expert calib ratio in huggingface_script.sh
Also fix minimax2.5 MOE detection which does not follow other HF MOE
layer convention

### Usage

scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4
--moe_calib_experts_ratio 1.0 --trust_remote_code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Configure MOE calibration experts ratio for quantization via an
environment/option, enabling finer control over calibration.

* **Bug Fixes**
* Improved detection of sparse MOE blocks to handle varying
expert/topology layouts, inferring expert counts when needed for more
reliable processing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-03 21:30:35 +00:00
03a1899dda Support force tokens to % of total experts during calibration (#910)
## What does this PR do?

**Type of change:** New feature

**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.

**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`

## Usage

Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5

Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq

quant_cfg = {
    "quant_cfg": { ... },
    "algorithm": {
        "method": "awq_lite",
        "moe_calib_experts_ratio": 0.25,  # calibrate 1/4 of experts
    },
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)

## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-02-24 12:35:19 -08:00
Chenjie Luo 9975ba1065 Fix DeepSeek PTQ script (#912)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?

Fix two bugs in the PTQ script

## Testing

Run DeepseekV3.2 PTQ and export


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced data type handling in quantization examples for bf16
operations
* Updated internal dependencies for quantization utilities to improve
modularity

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-20 19:45:37 +00:00
Chenjie Luo ac7c985d96 [NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?

**Type of change:** New feature, new tests

**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.

Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints

## Usage

Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:

import modelopt.torch.quantization as mtq

# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")

mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)

# During export, an .moe.html report with per-expert token counts is
saved automatically

## Testing
unittest, also test exporting qwen MOE

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.

* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-19 19:46:58 +00:00
Chenjie Luo 3801923e9d Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ?

Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ.

## Usage
scripts/huggingface_example.sh --model <minimax checkpoint> --quant
nvfp4 --trust_remote_code


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added MiniMax M2.1 model quantization support with nvfp4 format.
* Extended FP8 quantization capabilities with configurable dtype
parameter for enhanced precision control.

* **Improvements**
  * Enhanced detection of quantized linear module variants.
  * Improved weight unpacking for FP8-based linear modules.

* **Documentation**
  * Updated supported models table to include MiniMax M2.1.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-02-17 18:47:22 +00:00
Chenjie LuoandZhiyu 5e43b2a5f5 Support Qwen3 Next MTP load and export (#860)
## What does this PR do?

Fix MTP export for Qwen3 Next

**Overview:** ?

For Qwen3 next, the MTP weights are not stored separately in
safetensors. So we use "mtp" weights key to decide if the weights are
for MTP or not.


## Testing
Qwen3 Next PTQ and check if MTP is in the exported checkpoint.

scripts/huggingface_example.sh --model
<Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Optimized Multi-Token Prediction weight loading with improved layer
detection and handling.

* **Chores**
* Simplified status reporting to display total loaded weights and
detected layers.
  * Removed verbose per-file warnings for cleaner console output.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Zhiyu <zhiyuc@nvidia.com>
2026-02-09 22:48:15 +00:00
Chenjie Luo 615f99e746 Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support KIMI K2 Thinking PTQ from the original int4 checkpoint.
Tested with transformers  4.57.1, compressed-tensors 0.12.0

The model weights are dequantized on the fly to save GPU memory

## Usage
scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant
nvfp4_mlp_only --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Support for nvfp4_mlp_only quantization format, enabling new
layer-wise quantization options
  * Quantization support for CompressedLinear layers in quantized models

* **Improvements**
* Enhanced quantization for DeepSeek models with improved attention
configuration handling
* Optimized model loading with automatic precision configuration and
weight unpacking
* Better memory management during model export with automatic cache
cleanup
  * Conditional sample generation output controlled via verbose mode

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-21 09:32:33 +00:00
Chenjie Luo b0e7d9fd96 Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do?

**Overview:** ?

Unified the FP8 and NVFP4 kv cache scaling factor definition so the same
checkpoint can be used for both FP8 and NVFP4 kv cache quantization
deployment

## Testing
Unit test

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization,
simplifying configuration logic.

* **Chores**
  * Removed internal constants from public exports.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-20 08:34:32 +00:00
Chenjie Luo 9c24e2c08e Fix Deepseek transformers model loading (#740)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?
For Deepseek, let's force the user to apply trust_remote_code and use
AutoModelForCausalLM for loading the model.

## Testing
python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4
--export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64
--trust_remote_code --dataset cnn_dailymail

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-08 11:12:53 -08:00
Chenjie Luo d541324e84 Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do?

**Type of change:** ? Recipe improvement

**Overview:** ?

Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3
Next recipe for accuracy recovery

## Testing
Model accuracy benchmarking

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-02 11:06:25 -08:00
Chenjie Luo b1b9321877 Update llm_ptq doc (#685)
## What does this PR do?

**Type of change:** ? documentation

**Overview:** 
Update example doc

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-15 22:13:02 +00:00
Chenjie Luo bfec1b86fd Support Qwen3Next NVFP4 quantization (#681)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support Qwen3Next NVFP4 quantization. The QKV layers are not quantized
in PTQ to retain checkpoint accuracy across benchmarks.

## Testing
Checkpoint export and benchmarking using nemo-evaluator

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-15 21:14:04 +00:00
Chenjie Luo e4c5a68bd8 [NVBug: 5707914] Update DS V3 repo commit id (#635)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** 

DS V3 repo commit id update so the DS R1/V3.1 models can be correctly
loaded.

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-02 14:32:08 -08:00
Chenjie Luo 5ade7b03f8 [NNBUG: 5701866] Update DS V3.2 PTQ code (#630)
## What does this PR do?

**Type of change:** ?  Bug fix

**Overview:** 

1) Update the DS V3.2 repo code reference to the latest version
2) The new DS V3.2 model now includes fp32 layers. We cast it down to
match the checkpoint format during loading
3) Fix get_quant_config API change.

## Testing
Generate the deepseek-ai/DeepSeek-V3.2 checkpoint

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-12-02 09:36:55 -08:00
Chenjie Luo e20d218b43 [OMNIML-2857] [Experimental] Support the DeepSeek V3.2 model (#435)
## What does this PR do?

**Type of change:** ? New model support

**Overview:** ?

## Usage
Please see examples/deepseek/README.md


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* New Features
* Support for DeepSeek V3.2 quantization and automatic detection of
available DeepSeek versions.
* Triton-backed weight dequantization utility and MoE-aware calibration
mode to improve calibration fidelity.

* Documentation
* DeepSeek examples README expanded with setup, conversion, calibration,
and FP8→FP4 quantization workflows for R1, V3, and V3.2.

* Bug Fixes
* More robust, failure-tolerant copying of auxiliary files/assets during
quantization.

* Chores
  * Updated changelog and lint/ignore rules for example artifacts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-17 16:55:16 +00:00
Chenjie Luo ce8ce2229c [NVBUG: 5617733] Update LLM generate API for modelopt LLM eval (#498)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?

1) Remove kv_cache_config in the generate API. It's no longer used in
the code as well. We just estimate KV cache usage from other parameters
2) Add max_seq_len in the generate API to better estimate the real KV
cache usage.
3) Assume default lm_eval max input sequence length to be 4096

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-11-04 11:08:03 -08:00
Chenjie Luo 9a85e4922b [NVBUG: 5612606] Clear GPU cache for large models layer quantization during export (#497)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** ?

For large models like llama4 maverick, the stacked weights to fp8
conversion might hit OOM. This change aim to fix that.

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-04 18:46:58 +00:00
Chenjie Luo 5f0ef3b310 [NVBUG: 5608888] Update link in vlm_ptq README for support matrix details (#499)
## What does this PR do?

**Type of change:** ? documentation

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2025-11-04 22:52:06 +05:30