1310 Commits
Author SHA1 Message Date
Chenjie LuoandClaude Opus 5.5 ad8cd63847 Share one CUDA encoder per IQ family (#2615)
### What does this PR do?

Type of change: refactor (no behaviour change)

The five GGML IQ CUDA encoders were five copies of the same search.
IQ2_XS and IQ2_XXS shared 139 of their roughly 150 lines of encoder and
launcher code, and IQ2_S 113 of them. IQ1_S and IQ1_M had the same
structure with a different choice space. Every scaled packer also
validated its scales twice, in the `ggml.cpp` pybind wrapper and again
in the CUDA entry point.

This PR keeps **one encoder per family**, as two templates:

- **`iq2_family.cuh`** for IQ2_XS, IQ2_XXS and IQ2_S. The grid sits in
shared memory, the 16 local scales are scored per group, and each vector
then takes its best entry under the chosen scale. A format supplies its
group shape, whether it stores seven sign bits and recovers the eighth
from parity, and a `store()` that writes the chosen entries, sign masks
and local scales into its layout.
- **`iq1_family.cuh`** for IQ1_S and IQ1_M, over the shared ternary
grid. Each group picks one of `kChoices` options. With `kSharedShift`
the option also fixes the ±1/8 delta (IQ1_S: `shift * 8 + local`);
otherwise each vector picks its own (IQ1_M). IQ1_S's scale kernel now
writes FP16 scales, so both IQ1 formats take the same input.

Each format file is now one `Format` struct, holding its layout
constants and `store()`, plus its entry point: 58–100 lines each.
Validation lives once in `common.cuh`, as `check_pack_inputs` and
`check_scaled_pack_inputs`. `ggml.cpp` binds the CUDA entry points
directly instead of through five wrappers. **The kernel sources shrink
from 1,536 to 1,241 lines** (+665 / −960).

This is the first of two PRs. #2604 builds on it: it adds CUDA decoders
as a `decode()` next to each format's `store()`, and makes export reuse
fake quant's packed payloads.

### Testing

**Nothing changes in the output.** Before the refactor I hashed 40
outputs: 5 formats × float32/bfloat16/float16/float64 inputs × encode
and decode, on a weight with zero, tiny, oversized and non-finite
blocks. All 40 hash the same afterwards.

**Encode speed is unchanged.** Old and new were timed alternately for
four rounds, in both orders, on an idle RTX PRO 6000 with a 5632×2048
weight. They were within 1% for every format: IQ1_S 37.6 / 37.6 ms,
IQ1_M 37.0 / 37.0, IQ2_XXS 10.9 / 10.9, IQ2_XS 11.9 / 12.0, IQ2_S 15.5 /
15.4.

- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed**
- **Validation reports the same errors in the same order.** Over 5
formats × 8 combinations of bad arguments (devices, dtype, width, grid
shape, scales dtype, length and sign), every first error matches main's.
- `tests/gpu/_extensions/test_torch_extensions.py`: the
validation-message tests pass. #2515's two Q8_0 tests fail identically
on a clean `main` on this GPU.
- IQ unit tests (`test_ggml_backend.py`, `test_iq_formats.py`,
`test_convert_hf_config.py`, `test_presets.py`,
`test_export_weight.py`): **173 passed**

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ Same bindings, messages and
bytes.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: N/A. A refactor with no
behaviour change, verified by the hashes above and the existing GPU
tests.
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: **this** → #2604 (pack each IQ weight once and decode on
CUDA).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Quantization now checks that inputs and grids are CUDA tensors on the
same device, with compatible shapes. Scaled formats also validate scale
type, shape, and finite, non-negative values.
* **Improvements**
* IQ1 and IQ2 formats share common encoding paths while retaining their
format-specific output layouts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:14:42 -07:00
Daniel Korzekwa fadbf74d31 Make the QAT/QAD guide the central place for concepts, background, and framework selection (#2590)
### What does this PR do?

Make the QAT/QAD guide the central place for concepts, background, and
framework selection. Have the Hugging Face and Megatron Bridge tutorials
link back to it instead of repeating explanations of QAT and QAD,
keeping the tutorials focused on setup and execution. In main QAT/QAD
guide make links to all relevant blogposts.

Note: MBridge example doc is out of scope for this MR.

### Testing
Doc changes only, manual check.

### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?: N/A docs changes only

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded the QAT/QAD guide with workflows, use cases, and a
comparison, including QAD’s use of a frozen BF16 teacher and logit-level
loss to recover accuracy after quantization.
* Updated README and quick-start navigation to link to the combined
QAT/QAD guide; the previous standalone QAT guide now redirects readers
there.
* Reorganized the LLM QAT tutorial: recipe guidance is now part of the
end-to-end example, while trainer examples and Python
quantize-and-fine-tune guidance are in Advanced Topics. The tutorial
also notes Triton accelerated kernels.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
2026-10-01 20:56:57 +02:00
Ajinkya Rasane 4f62d418f4 Use TensorRT optimization level 0 in Torch ONNX example tests (#2619)
### What does this PR do?

Type of change: Bug fix

The Torch ONNX example tests can exceed their 300-second deadline while
building a ResNet50 INT8 TensorRT engine at optimization level 4. Add
`--trt_builder_optimization_level` to the vision example and select
level 0 in the existing integration tests. The example and helper retain
level 4 by default. Quantization, ONNX export, residual Q/DQ assertions,
engine execution, and test timeout limits are unchanged.

Document the build-time versus inference-performance tradeoff in the
example README.

### Usage

```bash
cd examples/torch_onnx
python torch_quant_to_onnx.py \
    --timm_model_name resnet50 \
    --recipe timm/resnet/ptq/int8 \
    --onnx_save_path resnet50.int8.onnx \
    --calibration_data_size 1 --no_pretrained \
    --trt_build --trt_builder_optimization_level 0
```

### Testing

Validation used the TensorRT 26.05 container, TensorRT 10.16.1.11, and
PyTorch 2.13.0, with the existing 300-second per-test deadline.

- RTX 6000 Ada: **8 passed**, covering FP8 and INT8 on ViT, Swin,
SwinV2, and ResNet50. ResNet50 INT8 passed in 96.65 seconds.
- RTX PRO 6000 Blackwell Max-Q: **20 passed, 3 existing skips**,
covering the complete test file. ResNet50 INT8 passed in 81.27 seconds;
the baseline timed out at 300 seconds.

```bash
# RTX 6000 Ada: supported FP8/INT8 cases
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py \
    -k '(fp8 or int8) and not mxfp8' --cov

# RTX PRO 6000 Blackwell: complete example test file
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py --cov
```

On each GPU, levels 4 and 0 used the same exported ResNet50 INT8 graph
with TensorRT 10.16.1.11:

- RTX 6000 Ada: TensorRT-reported engine build time decreased from 122.6
seconds at level 4 to 27.7 seconds at level 0. Both builds and inference
runs succeeded.
- RTX PRO 6000 Blackwell Max-Q: the original level-4 test hit its
300-second deadline during the engine build; level 0 built that saved
graph in 10.9 seconds and completed inference successfully.

All pre-commit checks passed for the changed files. The existing
integration tests exercise the real engine build; no redundant mocked
tests were added.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing integration
tests were updated to exercise level 0; no new test cases were needed.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — minor example/CI fix; library behavior and example defaults are
preserved.
- Did you get Claude approval on this PR?: N/A — not requested for this
focused change.

### Additional Information

Example timeout: [ResNet50 INT8 CI
failure](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36560176214/job/109381161169).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a configurable TensorRT builder optimization level for engine
builds, with a default of 4 and support for values from 0 to 5.
* Documented that level 0 can speed up builds, while lower optimization
levels may reduce inference performance.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-10-01 18:32:18 +00:00
Keval MorabiaandClaude Opus 5 9c1cf80f1b test(megatron_bridge): cover context parallelism in the VLM QAD test (#2592)
### What does this PR do?

Type of change: new tests

Runs the VLM case of `test_qad` under context parallelism, so QAD on a
Qwen3-VL model is covered on the path that until now could not run at
all.

`Qwen3VLMultimodalRotaryEmbedding` CP-shards its own embedding, so the
batch has to hand it full-length `position_ids`. Megatron-Bridge's
`get_batch` was sharding them too, leaving the rotary embedding at `seq
/ cp**2` against hidden states at `seq / cp`. The fix is upstream in
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243);
this PR is the coverage that would have caught it.

The VLM case moves from tensor to context parallelism and the PTQ step
is sized to the same TP, so QAD still loads a matching checkpoint. The
LLM case is unchanged (`tp_size=num_gpus, cp_size=1`), and
`test_distill_vlm` still covers TP for a VLM, so nothing loses coverage.

### Usage

```bash
# Unchanged: --cp_size is already a distill.py flag. On a container carrying Megatron-Bridge#6243
# it now works for VLMs, where it previously died in the rotary embedding.
python examples/megatron_bridge/distill.py --cp_size 2 --tp_size 1 ...
```

### Testing

On 2x RTX 6000 Ada, in `nemo:26.08` with Megatron-Bridge#6243 on
`PYTHONPATH`:

- `test_qad[qwen3_5_moe_vl]` at `--tp_size 1 --cp_size 2` — FP8 PTQ, QAD
across 2 CP ranks, export; quantizers survive and the vision tower is
byte-identical. **1 passed (183 s).** Without the upstream fix the same
run dies with `AttributeError: 'NoneType' object has no attribute
'ndim'` in `rope.py:175`.
- `test_qad[qwen3]`, the unchanged LLM path — **1 passed (194 s).**
- Gate check: on today's `nemo:26.08` (no #6243) the probe resolves
`False` and the VLM case stays at `cp_size=1`, byte-identical to current
CI; with #6243 it resolves `True` and runs at `cp_size=num_gpus`. On a
1-GPU runner it degenerates to today's config either way.
- `pre-commit run --files ...` clean (ruff check, ruff format, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — existing tests extended
rather than new ones added.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test coverage and one doc line; no feature, break, deprecation, or
fix for a released bug.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

- Depends on
[NVIDIA-NeMo/Megatron-Bridge#6243](https://github.com/NVIDIA-NeMo/Megatron-Bridge/pull/6243).
Safe to merge before it lands: the gate keeps the VLM case at
`cp_size=1` until a container ships the fix, at which point the coverage
switches on by itself.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the Qwen3.6 QAD instructions to keep tensor and pipeline
parallelism set to 1, while allowing context parallelism to increase for
longer sequences with the `nemo:26.10` container.
* **Tests**
* QAD validation now selects parallelism settings based on whether the
Megatron-Bridge context-parallel fix is available, and reports when
multi-GPU VLM coverage is reduced.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-10-01 23:05:55 +05:30
Chenjie LuoandClaude Opus 5.5 7d9e07d14b Count only added lines toward the PR size budget in AGENTS.md (#2616)
### What does this PR do?

Type of change: documentation

Changes the PR sizing rule in `AGENTS.md` to count only **added source**
lines toward the ~500-line budget, instead of total changed lines.
Deletions are cheap to review, so a PR that mostly removes code
shouldn't be pushed into a split. Tests and docs are excluded too, since
every sub-PR has to carry its own tests. The check uses the insertions
count from `git diff --shortstat` with a pathspec that excludes `tests/`
and `docs/`.

### Usage

```bash
git diff --shortstat origin/main...HEAD -- . ':!tests' ':!docs'
# N files changed, X insertions(+), Y deletions(-)  -> compare X against ~500
```

### Testing

- `pre-commit run --files AGENTS.md` (markdownlint passes).
- Ran the pathspec against recent commits (#2595, #2513) to confirm it
drops test and doc lines from the count.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Follow-up to #2494, which introduced the sizing guidance.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated review guidance to measure pull request size by added source
lines, excluding deletions, tests, and documentation. The guidance
retains the recommendation to check the size before opening a review.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 07:03:48 +00:00
Chenjie LuoandClaude Opus 5.5 aa89722d38 [6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed
the PyTorch codec; this PR adds its **CUDA encoder** and makes the
format reachable. With it, ModelOpt supports all five GGML IQ formats at
one and two bits.

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq1_m`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

### The kernel

In the kernel the delta shift is free per group, so it sits above the
entry index in the sort key: a tie still prefers the lower shift and
then the lower entry, as the reference encoder does. The 2048-entry grid
IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory
limit, so both kernels read it from global memory and rely on the cache.

| 5632×2048 weight | torch | CUDA | |
|---|---|---|---|
| IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** |

### Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and
apply the ±1/8 shift the same way. So those three steps move into
`common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and
IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA
output on a 5632×2048 weight hashes the same before and after, and so
does IQ1_M's, compared against the pre-split version of this change.
IQ1_S encodes at the same speed (306 M elem/s).

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ1_M-specific test code: backend dispatch, weight caching, the
`num_bits` guard, `convert_hf_config` metadata, Megatron export and the
`TensorQuantizer` tests in the shared battery. The shared CUDA battery
gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **166 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **221 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO
6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch
encoder parity
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** (9 tests × 5 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed**
- reconstruction error falls monotonically across all five formats,
pinned by a test
- `general/ptq` now holds 31 recipes; `ptq.md` is updated.

Rebased onto `main` after #2513 merged. The resulting tree is identical
to the one the runs above tested, and the unit set was rerun on it: 166
passed.

On this GPU, two of #2515's Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M
codec), all merged → **this**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M weight-only quantization at 1.75 bits per weight, with
CUDA acceleration and a 256-value block size.
* Added an IQ1_M post-training quantization recipe for eligible linear
layers; calibration data is not required.
* Added IQ1_M to the supported GGML-compatible formats and recipe
listings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:23:48 -07:00
sychen52 333ace1bc9 Fix multi-turn synthetic generation and mark incomplete outputs (#2571)
### What does this PR do?

  Type of change: Bug fix

Fixes synthetic conversation generation that assumed alternating user
and assistant messages. That assumption skipped user turns in prompt
skeletons and mishandled
  leading system messages.

• Preserve conversation history. Regenerate every user turn while
retaining system messages and generated reasoning for subsequent
requests.
• Expose generation controls. Support model-specific request parameters,
configurable timeouts, and server-managed response budgets.
• Handle failures explicitly. Reject empty final answers and unsupported
tool calls. Failed conversations remain retryable without duplicating
saved output.
• Identify incomplete outputs. Mark length- and repetition-stopped
conversations as truncated, preserve stop metadata, and stop generating
follow-up turns.

  ### Usage

Run from the repository root against a compatible Qwen server with
reasoning parsing enabled:

  python examples/speculative_decoding/scripts/server_generate.py \
      --data_path input_conversations/train.jsonl \
      --output_path synthetic/train.jsonl \
      --url http://localhost:8000/v1 \
      --model model \
      --max_tokens 0 \
      --request_timeout 3600 \
--extra_body
'{"chat_template_kwargs":{"enable_thinking":true,"preserve_thinking":true}}'

The model name must match the server’s configured name. Filter truncated
conversations before training.

  ### Testing

  Focused regression tests: 15 passed.

The tests execute the command-line entry point using the real OpenAI
client library with mocked HTTP transport.

Coverage includes multi-turn generation, system prompts, reasoning
preservation, request parameters, failure recovery, resume
deduplication, truncation, and invalid
  responses.

  python -m pytest \
      --confcutdir=tests/examples/speculative_decoding \
      tests/examples/speculative_decoding/test_server_generate.py -q

The isolated test configuration avoids an unrelated parent configuration
import failure. All applicable pre-commit checks passed for the
generator, tests, and
  documentation.

  ### Before your PR is "Ready for review"

• Is this change backward compatible?: ✅ Existing valid inputs,
defaults, conversation output structure, and resume behavior remain
supported. Invalid inputs and
    failed requests now raise errors instead of being silently accepted.

• Copied code or new PIP dependencies?: N/A. No new third-party code or
dependencies were added.
• Did you write any new necessary tests?: ✅ Added focused command-line
regression tests.
• Did you update Changelog?: N/A. These are example-script correctness
fixes, not critical released library fixes.
  • Did you get Claude approval on this PR?: ❌ Not yet obtained.

  ### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Data generation supports `conversations` and `messages` inputs,
preserves reasoning content, and accepts additional chat settings and
configurable request timeouts.
* Failed conversations are recorded separately, with options to retry
failures or exit when errors occur. Resume behavior distinguishes
retryable failures from rejected inputs.
* Outputs identify conversations truncated by length or repetition
limits.

* **Documentation**
* Updated data preparation guides with generation setup, input formats,
failure handling, resuming, and training guidance.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-10-01 00:37:08 +00:00
Trenton Starkey - Product @ NVIDIA bc5d2c3610 Update roadmap link in README.md (#1700)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated the Roadmap link in the README to point to the current roadmap
reference.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Trenton Starkey - Product @ NVIDIA <trenton.starkey@outlook.com>
2026-10-01 00:06:47 +00:00
Chenjie LuoandClaude Opus 5.5 e5b63320ab [5/6] Add the IQ1_M codec (#2513)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 09:45:48 -07:00
kinjalpatel27 ad9ea97a4b Fix vLLM compilation guard for models without marker (#2518)
### What does this PR do?

Type of change: Bug fix

Makes the vLLM `disable_compilation` context manager support inner model
implementations that do not predefine a `do_not_compile` attribute,
including GLM-5.3. The context manager now installs the marker
temporarily and removes it afterward, while preserving and restoring
existing marker values for other vLLM models.

Adds regression coverage for both supported wrapper layouts:
`model.model` and `model.language_model.model`.

### Usage

```python
with disable_compilation(model):
    mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```

No caller changes are required.

### Testing

- Ran `tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py`:
24 passed with vLLM 0.28.
- Ran pre-commit on both changed files: all applicable hooks passed.
- Installed this branch into `vllm/vllm-openai:glm53-flash` on OCI-JHB
and served the GLM-5.3-Flash BF16 checkpoint with
`QUANT_CFG=NVFP4_DEFAULT_CFG`, TP=4, eager mode, and BF16 KV cache.
- GLM passed the previous `do_not_compile` failure point, inserted 1,700
quantizers, enabled 456 weight quantizers, reached a healthy API server,
and returned a relevant manual prompt response.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — integration compatibility fix; no user-facing API change.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Validated against GLM-5.3-Flash using ModelOpt commit
`869b64fcee0b20be323663449b00e8c52940a289`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Compilation settings are now handled across supported nested model
configurations and restored after calibration, including when errors
occur.
- Calibration inputs correctly exclude padding when an attention mask is
provided and reject empty sequences.
- vLLM warmup reserves the required cache space for supported tail-cache
configurations.
- Serving startup supports an alternate vLLM launcher import path when
the OpenAI entrypoint is unavailable.
- **Compatibility**
- The vLLM serving example now defaults to vLLM 0.30.0 and documents
tested support for Nemotron 3 Nano hybrid attention/Mamba serving on
vLLM 0.26.0 and 0.30.0.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-09-29 14:28:32 -07:00
Chenjie LuoandClaude Opus 5.5 3091b8ff69 [4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512
landed the PyTorch codec; this PR adds its **CUDA encoder** and makes
the format reachable:

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq2_s`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B
parameters**.

### The kernel

IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its
search the most expensive in the family. The codebook and its norms take
36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB
static limit, so they are declared statically like the IQ2_XS and
IQ2_XXS kernels.

That cost is why the kernel matters more here than anywhere else:

| | torch | CUDA | |
|---|---|---|---|
| IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×**
|
| extrapolated to a 27B model | ~9.8 hours | **~37 s** | |

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ2_S-specific test code: backend dispatch and weight caching, the
`num_bits` guard, `convert_hf_config` metadata (uniform and mixed
precision), all 9 Megatron export tests, and the two `TensorQuantizer`
tests in the shared battery. The shared CUDA battery gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **134 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **192 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO
6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder
parity, determinism, reconstruction at scale, zero and non-finite
policy, float64 input and the fallback path.
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **36 passed** (9 tests × 4 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**.
TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`,
`block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)`
uint8.
- `general/ptq` now holds 30 recipes.
- The shared-memory change in `b7739d5d0` leaves the packed bytes
identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M
elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun:
42 passed.

All of the above was rerun after rebasing onto `main` at `c2aaa44f6`.
That base adds a Q8_0 packer to the same GGML extension (#2515), and
changes the hf_ptq example and the export code this format goes through.
The packed IQ2_S bytes still hash the same. On this RTX PRO 6000
(sm_120), two of #2515's own Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
#2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ2_S weight-only quantization for eligible linear layers, at
2.5625 bits per weight.
* Added a PTQ recipe that requires no calibration data. Weights must
meet the existing 256-value block-size constraint.
* Added CUDA-accelerated packing for CUDA weights, with a Python
fallback when the CUDA extension is unavailable.
* **Documentation**
  * Updated the PTQ recipe catalog and IQ-format size tradeoffs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 20:35:38 +00:00
Ajinkya Rasane 1576f7d4ad Increase diffusers example-test timeout to 60 minutes (#2591)
### What does this PR do?

Type of change: Bug fix

The diffusers example job can exhaust its 45-minute job budget while
tests are still progressing. On the same commit, an [initial attempt
timed
out](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109049102711),
while a [retry passed all 47 tests in
44m50s](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36456149060/job/109115221419),
leaving only 10 seconds of headroom.

Increase the diffusers timeout to 60 minutes in
`.github/workflows/example_tests.yml` to accommodate the workload and
observed runtime variation. The other ONNX matrix entries retain their
45-minute timeout. This applies to both PR and nightly diffusers jobs.

### Usage

N/A — CI configuration change.

### Testing

- `pre-commit run --files .github/workflows/example_tests.yml` — passed
all applicable hooks.
- Parsed the caller and reusable workflow with `yaml.safe_load` and
inspected the timeout input and consumer.
- `git diff --check` — passed; reviewed the one-line diff.
- GPU tests were not rerun locally. CI validation of the increased
timeout is pending.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code or new dependencies.
- Did you write any new necessary tests?: N/A — one-line CI
configuration change; validation described above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — CI-only change.
- Did you get Claude approval on this PR?: ❌ Not run; opening as a
draft.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated automated checks to allow ONNX example tests up to 60 minutes.
Other example tests retain their existing 45-minute limit. This change
affects test execution time limits only; it does not change application
features or behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-09-29 19:15:05 +00:00
sychen52andClaude Opus 5.5 834c90d7a1 Fix NVFP4 fake quant zeroing blocks with small scales (#2549)
### What does this PR do?

  Type of change: Bug fix

The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute
>= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block
scale below 1e-5 with
1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static
Triton kernel, the CUDA extension fallback and NVFP4 export have no such
floor; the floor was
only guarding division by zero. The Triton kernel had its own copy of
the scale/round code instead of the shared `nvfp4_scalar_quant`.

- Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block
scale).
- Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule.
- Both: a zero, inf or NaN global amax uses a unit block scale, like the
CUDA extension (the conv kernels returned NaN for inf/NaN before).

  Blocks with scale >= 1e-5 are unchanged.

- New tests: power-of-two scaling of input and global amax (2^-10,
2^-20) scales the output by the same factor (Triton, standalone conv
FP4, fused conv3d); invalid
global amax gives unit-scale rounding. The conv test's Python reference
drops the floor.
- B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization`
NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ).
- Negative control on `main`: the small-input tests fail on both
kernels; conv also fails for inf/NaN global amax.

  ### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (only blocks with scale < 1e-5
or an inf/NaN global amax change, from zeros/NaN to correct values)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
  - Did you write any new necessary tests?: ✅
  - Did you update Changelog?: N/A
  - Did you get Claude approval on this PR?: ❌




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* FP4 quantization now preserves proportional output when inputs and
valid global scales are reduced together, including at small scales.
* Zero, infinite, or NaN global scales use a safe fallback, preserving
inputs already representable in FP4.
* Small positive block scales are no longer discarded by an absolute
scale threshold; subnormal scale handling is covered across quantization
paths.
* **Tests**
* Added coverage for scale consistency across input types and block
sizes, invalid global scales, and subnormal FP8 block scales.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:31:31 -07:00
Daniel KorzekwaandKeval Morabia 94d7272e8b Improve technique-specific documentation links in the main README (OM… (#2584)
### What does this PR do?

Improve technique-specific documentation links in the main README
(OMNIML-5944):

The Techniques table in ./README.md contains links that do not lead
directly to the relevant documentation:

- The Docs link for Quantization Aware Training / Distillation points to
the general quantization documentation. It should point to
./guides/quantization_aware_training.html.

- The Megatron Bridge links for Post Training Quantization, Quantization
Aware Training / Distillation, Pruning, and Distillation all point to
the same folder. Users must then find the relevant section themselves.
Update the Megatron Bridge links to target the corresponding sections in
./examples/megatron_bridge/README.md:

- In ./README.md’s Techniques table, rename Docs to Getting started and
reverse the order of the two link columns: currently Examples → Docs,
proposed Getting started → Examples.

### Usage

just see the table of techniques in the main readme.me

### Testing

tested manually

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ 
- Did you write any new necessary tests?: ❌ , only manual testing, only
docs changes

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the Techniques table with “Getting started” and “Examples”
columns, replacing the former “Examples” and “Docs” columns.
* Added technique-specific guide links, including a link to the README
for pruning.
* Updated some example destinations and anchors to point to relevant
workflow sections.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Daniel Korzekwa <dkorzekwa@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-29 12:42:42 +02:00
Keval Morabia 8990897c56 Increase Unit Test timeout
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-09-29 15:00:11 +05:30
h-guo18andClaude Opus 5 c2aaa44f60 [Speculative Decoding] DFlash2 draft variant (grouped sublayer convolution + candidate selector) (#2216)
### What does this PR do?

Type of change: new feature

Adds **DFlash2** ([blog](https://inco.ai/blog/dflash2/)) as a draft
variant of the existing DFlash mode, selected with
`dflash_architecture_config.projector_type="dflash2"` alongside
`domino`, `dspark` and `lilicorr`.

DFlash2 keeps DFlash's one-pass parallel backbone and adds two
components that recover the acceptance a purely parallel draft loses:

- **Grouped dynamic depthwise convolution** around every attention and
MLP sublayer, giving each block position a view of its predecessors
*inside* the block. Taps do not cross the block boundary, so the draft
stays one forward pass.
- **Low-rank candidate selector** scoring transitions between adjacent
positions' top-k candidates, so serving walks one coherent path instead
of taking an independent argmax per position.

Both start as exact no-ops — the convolution's `base_kernel` is an
identity and `kernel_projection` is zeroed; the selector's
`successor_codebook` is zeroed — so a freshly built DFlash2 draft *is*
its DFlash backbone, and enabling the variant is an extension rather
than a perturbation. This matches the reference implementation
([SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772),
merged) and the way `modeling_lilicorr` already installs this same
convolution class.

**This unblocks a recipe already shipped on `main`.**
`modeling_lilicorr._install_sublayer_convs` imports `DFlashGroupedConv`
from `modeling_dflash2`, so
`modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml`
raises at model build today and its CHANGELOG entry documents a feature
that cannot run. Landing this makes it runnable.

Module and parameter names match the SGLang/vLLM `DFlash2DraftModel`
loaders. Verified against the released `z-lab/Qwen3.8-27B-DFlash2`
checkpoint: **81 tensors, 21 name patterns, zero difference in either
direction**. The serving side,
[vllm-project/vllm#52816](https://github.com/vllm-project/vllm/pull/52816),
has since merged (`b389ac294`) with no change to the checkpoint
contract.

### Usage

```bash
python examples/speculative_decoding/main.py \
  --config modelopt_recipes/general/speculative_decoding/dflash2.yaml \
  model.model_name_or_path=Qwen/Qwen3-8B \
  data.data_path=<corpus>.jsonl \
  training.output_dir=<out>
```

```yaml
# modelopt_recipes/general/speculative_decoding/dflash2.yaml
dflash:
  dflash_selector_loss_alpha: 1.0      # weight of the candidate-selector CE term
  dflash_architecture_config:
    projector_type: dflash2
    conv_kernel_size: 2                # taps; must not exceed the block size
    conv_group_size: 16                # must divide hidden_size
    selector_rank: 256
    selector_top_k: 16
```

### Testing
<img width="2000" height="1320" alt="image"
src="https://github.com/user-attachments/assets/869004d1-92c6-41a4-a03d-fd824a06255c"
/>

**Unit** — 25 CPU tests in
`tests/unit/torch/speculative/plugins/test_hf_dflash2.py`; the full
`tests/unit/torch/speculative/` suite passes with no regressions. The
ones worth keeping pin invariants that a decreasing loss does not catch:

- the convolution is an exact identity on the **default** construction,
and its taps stay inside the block while a position still sees its
predecessors;
- the block-offset contract shared by the training objective and
`CandidateSelector.greedy_path` — a misaligned objective still
converges;
- the export fields the vLLM loader requires, including the top-level
`block_size` that `DFlash2Exporter` derives the nested copy from;
- which selector factors receive gradient on the first step.
`successor_codebook` starts at zero, so `predecessor_codebook` and
`hidden_projection` take one step to begin moving. That is a warm start,
not a dead branch, and both sides are asserted.

**End-to-end** — trained on Qwen3-8B against a plain DFlash control with
every other argument identical (plot above). Monotonic convergence, no
NaN/divergence, no DDP unused-parameter issues. Note the losses are
**not comparable** across arms: DFlash2's includes the selector CE term.

**Serving (vLLM)** — the exported drafter loads and drafts under the
merged DFlash2 path (`RESOLVED draft architectures:
['DFlash2DraftModel']`). Two notes for anyone reproducing: vLLM sizes
the convolution from `1 + num_speculative_tokens` at runtime rather than
from the checkpoint, so a `block_size=16` drafter is only correct at
`num_speculative_tokens=15`; and at that value the upstream path
currently hits an illegal memory access in `_cache_draft_logits`
([vllm#55279](https://github.com/vllm-project/vllm/issues/55279)),
independent of which checkpoint is used.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — additive. New
`projector_type`, its own registry and exporter, one new config field;
DFlash / Domino / DSpark / LiLiCorr numerics and `state_dict` contents
are untouched.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ —
`modeling_dflash2.py` is adapted from
[SpecForge#772](https://github.com/sgl-project/SpecForge/pull/772) and
carries its MIT notice. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under `0.48.0`.
- Did you get Claude approval on this PR?: ✅ — run on 2026-08-20; all
review threads addressed and resolved.

### Additional Information

Rebased onto current `main`. Two commits from the original branch were
dropped because
[#2342](https://github.com/NVIDIA/Model-Optimizer/pull/2342) landed them
first, with authorship preserved: the no-op sublayer seam in
`modeling_dflash.py`, and the `rope_theta`/`rope_parameters` fix —
`main`'s version of the latter is stricter, so this PR no longer touches
`hf_dflash.py` at all.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DFlash2 speculative decoding with grouped dynamic convolutions
and low-rank candidate selection.
* Added configurable selector-loss weighting, including an option to
disable it.
  * Added DFlash2 model conversion and export support.
  * Added checkpoints compatible with SGLang and vLLM DFlash2 serving.

* **Documentation**
* Added training recipes and a Qwen3-8B online DFlash2 training
configuration.

* **Tests**
* Added coverage for conversion, training, metrics, gradients, and
export compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 11:43:02 +08:00
sychen52 be74001256 Add concise AgentX benchmark skill (#2573)
### What does this PR do?

  Type of change: Documentation.

Adds a concise AgentX skill covering harness installation, automatic
dataset downloads, benchmark execution, and result reporting. Reuses the
existing deployment skill
  and adds a Claude discovery link.

  ### Usage

Use run-agentx to benchmark my deployed model with a concurrency sweep.

  ### Testing

  • Skill structure and metadata validation passed.
  • Shell syntax checks passed.
• Benchmark arguments parsed and produced a valid configuration using
the pinned harness.
  • All applicable pre-commit checks passed.
  • No GPU benchmark was launched.

  ### Before your PR is "Ready for review"

Contributor guidelines and security practices were reviewed. The commit
is signed and signed off.

  • Is this change backward compatible?: ✅
• If you copied code from other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md?: N/A. No copied
implementation or project dependency
    changes.

• Did you write any new necessary tests?: N/A. Documentation changes
were validated as described above.
  • Did you update Changelog?: N/A. Skill documentation only.
  • Did you get Claude approval on this PR?: ❌ Not run.

  ### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added guidance for configuring and running SemiAnalysis AgentX serving
benchmarks, including endpoint, model, tokenizer, context limits,
caching, and dataset setup.
* Documented using a pinned benchmark harness in a separate client
environment and running each concurrency level in a fresh artifact
directory with a fixed seed.
* Expanded reporting guidance to cover overlapping requests, cache and
preemption metrics, errors, unfinished requests, warmup failures, and
submission validity.
* Clarified that missing server-reported usage makes cache-hit data
unknown, invalid or missing submission validity should be flagged, and
smoke runs are not benchmark results.
* Directed AgentX agentic workloads from the optional AIPerf guidance to
the AgentX benchmark instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-09-29 00:33:53 +00:00
Wei-Ming Chen 1d392999b4 [OMNIML-5570] 2/2 Compose GEMM and KV-cache AutoQuant workflows (#2273)
### What does this PR do?

Type of change: new feature.

Follow-up to merged #2272. Adds composition of existing GEMM
quantization with KV-cache AutoQuantize:

- fixed FP8 GEMM PTQ followed by mixed-KV AutoQuantize;
- gradient-based NVFP4/FP8 GEMM AutoQuantize followed by independent
mixed-KV AutoQuantize;
- an optional `kv_auto_quantize` recipe stage with independent method,
constraints, candidates, and checkpoint path;
- ordered `hf_ptq.py` orchestration that keeps selected
weight/activation QDQ active while its calibration state remains frozen
during KV candidate calibration;
- fail-closed validation when a preceding stage leaves actual K/V
quantizers enabled; and
- unified export of a uniform-weight or mixed-weight checkpoint together
with the selected per-layer KV map.

The KV search still uses the public `mtq.auto_quantize(...,
constraints={"cost_model": "kv_cache", ...})` API from #2272. On a
converted model, the API preserves existing non-KV quantizers and
requires K/V to be disabled before search. Fresh-model behavior is
unchanged and starts from a deny-all quantizer baseline.

#### Why a follow-up field instead of a generic stage list?

This PR deliberately supports the two composition forms required by
`hf_ptq.py` without replacing the stable recipe schema. Existing recipes
already express a fixed `quantize` baseline plus one primary
`auto_quantize` search. A generic ordered `stages` list would require a
broader recipe/API migration, indexed checkpoint semantics, and
compatibility rules for arbitrary stage sequences. There is not yet a
demonstrated third search stage that justifies that surface-area change.

The two searches are not combined inside `mtq.auto_quantize`: each
invocation owns one search domain, constraint model, scoring method, and
resumable checkpoint. Their ordering and independent checkpoint paths
are orchestration concerns, while candidate calibration, scoring,
selection, and state application remain in the shared public API. A
general stage pipeline can be considered separately if more than this
one optional KV follow-up is needed.

Both solvers and scoring protocols are unchanged. The KV checkpoint
compatibility signature additionally fingerprints the preceding
quantizer configuration and calibrated state. Unsupported uniform-weight
plus mixed-KV exports record `kv_cache_deployment_supported: false` in
both ModelOpt and converted HF metadata.

### Usage

Fixed FP8 GEMM PTQ followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/fp8_ptq_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-fp8-and-mixed-kv
```

Weight AutoQuantize followed by KV AutoQuantize:

```bash
python examples/hf_ptq/hf_ptq.py \
  --pyt_ckpt_path Qwen/Qwen3-8B \
  --recipe general/auto_quantize/nvfp4_fp8_gradient_then_kv_fp8_nvfp4_cast_kl_div_at_5p4bits \
  --auto_quantize_checkpoint /path/to/weight_autoquant.pth \
  --kv_auto_quantize_checkpoint /path/to/kv_autoquant.pth \
  --export_path /path/to/qwen3-8b-autoquant-and-mixed-kv
```

KV checkpoint resume requires identical preceding non-K/V quantizer
configuration and calibrated state. If rerunning the preceding stage
changes that state, use a new KV checkpoint path to recompute
sensitivities; configuration identity alone is insufficient to reuse the
scores safely.

### Testing

- Latest changed-area validation: 126 tests passed across `hf_ptq.py`
orchestration, KV checkpoint compatibility, export metadata, and HF
configuration conversion.
- A broader local run had 604 passes, one skip, and six failures: two
socket-binding failures under the sandbox and four local Transformers
API incompatibilities. This is not a full-suite pass.
- The fixed-PTQ→KV recipe executes end to end on a tiny offline Qwen
fixture.
- Public API coverage verifies that composed KV search preserves
preceding weight quantization and rejects enabled K/V state.
- Changed-file pre-commit hooks passed; the isolated recipe validator
also passed after dependency bootstrap.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅ (0.48.0 composition feature and KV
checkpoint flag deprecation)
- Did you get Claude approval on this PR?: ❌

### Additional Information

- This follow-up targets `main`, which contains merged #2272.
- `--auto_quantize_checkpoint` and `--kv_auto_quantize_checkpoint` are
intentionally separate because KV sensitivities depend on the preceding
GEMM state.
- Uniform-weight plus mixed-KV exports are for artifact inspection until
the runtime's uniform-weight ModelOpt configuration consumes
`kv_cache_quantized_layers`. Export emits an actionable warning and
records `kv_cache_deployment_supported: false`; this marker does not
itself add runtime support.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added staged post-training quantization workflows for weights and KV
caches, including dedicated KV-cache checkpoints.
- Added FP8/NVFP4 recipes with configurable bit constraints and scoring.
  - KV-cache quantization now supports pre-quantized models.

- **Bug Fixes**
- Mixed weight and KV-cache quantization now exports with a warning
instead of failing.
- Improved validation and checkpoint compatibility for staged
configurations.
- Added safeguards for configurations without enabled weight quantizers.

- **Documentation**
- Clarified staged KV-cache workflows, checkpoint options, configuration
behavior, and unsupported deployment combinations.
- Documented deprecated legacy quantization options and their
replacement behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-09-29 00:15:40 +00:00
hychiang de2e810219 [OMNIML-5899] Add Q8_0 CUDA packing kernel (#2515)
### What does this PR do?

Type of change: new feature

Adds the Q8_0 CUDA packing layer for the three-PR Q8_0 series:

- packs 32-value blocks into the 34-byte GGML-compatible payload;
- exposes `q8_0_pack` through the shared GGML extension;
- validates device, dtype, row alignment, and CUDA launch bounds;
- tests byte layout, accepted dtypes, non-finite handling, float64
narrowing, FP16 scale boundaries, reconstruction error, and invalid
inputs.

This PR contains only the kernel and extension boundary. The
codec/backend and export/recipe layers remain in the later PRs.

### Usage

```python
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml

extension = get_cuda_ext_ggml(raise_if_failed=True)
packed = extension.q8_0_pack(weight)
```

### Testing

- Combined #2515 -> #2516 -> #2517 stack: 108 focused CPU codec,
backend, export, and recipe tests passed.
- Ruff, formatting, and whitespace checks passed for the changed Python
test.
- The shared-extension Q8_0 and existing IQ tests require CUDA CI; the
latest run is pending on the current PR head.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: yes
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: yes; no
dependency was added and no implementation code was copied
- Did you write any new necessary tests?: yes
- Did you update `CHANGELOG.rst`?: N/A; the user-facing entry is in
#2517
- Did you get Claude approval on this PR?: pending

### Provenance

The CUDA encoder was independently written for ModelOpt. It implements
the packed-format contract and scalar quantization formula documented by
the pinned llama.cpp definitions:

- [Q8_0 packed
structure](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h)
- [Q8_0 scalar reference
formula](https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c)

No llama.cpp implementation code is incorporated into this CUDA source.
Human code-owner confirmation of this provenance and attribution is
requested before merge.

### Related PRs

Merge order:

1. **Kernel - this PR**
2. [#2516 - Q8_0 quantization codec and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2516)
3. [#2517 - Q8_0 checkpoint export and
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2517)

All three PRs target `main`.

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
2026-09-28 15:22:25 -07:00
Keval MorabiaandClaude Opus 5 0058a15537 [2/2] Track every Megatron-Bridge script with MLflow (#2514)
### What does this PR do?

Type of change: new feature

**[2/2] of a split. Based on #2544 — merge that first; this PR's diff is
only the Megatron-Bridge half.**

#2477 added MLflow tracking to `examples/megatron_bridge/quantize.py`.
It was one of five scripts in that directory that write a checkpoint;
the other four recorded nothing, so the provenance chain stopped at the
PTQ checkpoint and a deployed model could not be traced back to the run
that produced it.

All five now take the same `--mlflow` / `--mlflow_experiment` /
`--mlflow_run_name` flags, and **each declares what it records as a
`Tool` beside its own flags** — the shared `mlflow_utils.py` knows none
of them:

| Script | Records |
| --- | --- |
| `prune_minitron.py` | command, arguments, log, `prune_score` metric,
pointer |
| `quantize.py` (#2477, moved onto the shared `Tool` in #2544) | +
resolved recipe, quantizer summary |
| `distill.py` | + Megatron-Bridge's per-iteration metrics and resolved
config |
| `export_quantized_megatron_to_hf.py` | command, arguments, log,
pointer |
| `export_distilled_megatron_to_hf.py` | same, one pointer per exported
checkpoint |

Each writes `.experiment.json` into the checkpoint it produced, and each
tags what it consumed, so `prune → quantize → distill → export` is
walkable both from disk and by tag query on the server.

**`distill.py` opens the run and Megatron-Bridge joins it.** Its
`LoggerConfig` records per-iteration metrics and the full resolved
config — which a wrapper around `main()` cannot see — but nothing of
`distill.py`'s own arguments and no invocation. Megatron-Bridge takes
`mlflow.active_run()` when one exists, applies the tags and logs into
it, so `distill_run()` opens the run on the rank Megatron-Bridge looks
at (the **last** one) and the two share it. Its early exit is handled
explicitly: `train()` leaves through `sys.exit(0)` on `--exit_interval`,
which a blanket handler would record as `FAILED`.

**The library pieces that exist for that shared run land here with their
first caller**, rather than in [1/2] where they would have none:
`split_tracking_credentials`, so a URI handed to something which
*records* it carries no credential; `log_active_run_experiment_json`,
for pointing a checkpoint at a run this process did not open; and
`MlflowRunLogger._reattach`, because a co-owner can end the run first —
Megatron-Bridge does, as `KILLED`, when SIGTERM arrives mid-training.

Two of Megatron-Bridge's defaults are deliberately not inherited:
**checkpoint artifact upload stays off** unless
`--mlflow_log_checkpoints` (it pushes the whole checkpoint over HTTP
after every save), and **an untracked run passes no `mlflow_*` fields at
all**, since they landed in Megatron-Bridge 0.6 and sending them
unconditionally would break an untracked run on an older one.

### Usage

```bash
# Any of the five, same flags:
torchrun --nproc_per_node 8 prune_minitron.py  ... --mlflow https://<server>/
torchrun --nproc_per_node 8 quantize.py        ... --mlflow https://<server>/
torchrun --nproc_per_node 8 distill.py         ... --mlflow https://<server>/
torchrun --nproc_per_node 8 export_quantized_megatron_to_hf.py ... --mlflow https://<server>/

# Each checkpoint names the run that wrote it:
cat /output/qad/checkpoints/.experiment.json
```

Experiments default to
`$USER/megatron_bridge_{prune,quantize,distill,export,distill_export}/<model
basename>-<variant>`.

### Testing

- Real runs on a toy Qwen3 in one MLflow experiment covering all five
Megatron-Bridge scripts and `hf_ptq` — prune, quantize, QAD
distillation, quantized export, BF16 distillation, distilled export, HF
PTQ — each closing `FINISHED` with the invocation, its arguments as
params, its log, and a matching `.experiment.json` on disk. The chain
tags line up: each stage's `source_checkpoint_path` is the previous
stage's `checkpoint_path`.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08`, the
only lane that runs it: **76 passed**. Plus the three suites from #2544:
**195 pass**.
- `pre-commit run --files <changed>`: all hooks pass.
- Each fix from the review rounds has a test that fails with the fix
reverted: the resumed run, the foreign active run, the percent-decoded
credential, the credential that cannot be moved, the rank-dependent
`LoggerConfig`, the exit-callback guard, and the `iter_*` join.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split from a single ~1150-line PR at review's request; #2544 carries the
library consolidation this builds on, and this branch is based on it.
Earlier review threads here show as outdated after the rebases — they
are all resolved and their fixes are in this branch.

One known gap, stated in the README rather than implied: `distill.py
--hf_export_path` writes a second HuggingFace checkpoint from rank 0,
which is not the rank that owns the run, so it carries no pointer yet.
For the same reason the uploaded `logs/distill.log` holds the last
rank's output — `print_rank_0` keeps the script's own lines on rank 0 —
which the README now says outright; carrying rank 0's log into a run
owned by another rank needs cross-rank upload and is a follow-up.

Two defects found on shared-run paths during review, both verified
against the installed Megatron-Bridge 0.6 rather than its docs.
Megatron-Bridge ends the run it shares with `distill.py` as `KILLED`
from its SIGTERM handler (`train.py:1413`) and then leaves through
`sys.exit()` (`train.py:805`), i.e. before `distill_run`'s `finally` —
and MLflow's fluent calls resolve their target by *opening* a run when
none is active, so a preempted distillation's log and metrics went to a
second, empty run and its `KILLED` status was overwritten. Separately,
an unreachable server disabled our logger but `logger_kwargs` still
handed Megatron-Bridge the same URI, and `state.py` calls
`set_experiment` unguarded from inside the training loop — so a
best-effort `$MLFLOW_TRACKING_URI` aborted the training instead of
degrading to untracked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 21:29:25 +00:00
Chenjie LuoandClaude Opus 5.5 767ef5533e [3/5] Add the IQ2_S codec (#2512)
### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ2_S**, the widest of the GGML IQ formats at
one and two bits (2.5625 bits per weight). This one lands the **PyTorch
codec**: the encoder, the decoder and the 1024-entry codebook. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2565 adds the CUDA encoder,
registers the format and adds its recipe.

### What's distinctive about it

**IQ2_S is the one format llama.cpp's own tooling gives no head start
on**, so the search is written against the GGML layout directly.

The interesting difference from IQ2_XS and IQ2_XXS is sign handling.
IQ2_S stores a **full 8-bit sign mask** per group rather than a 7-bit
parity-coded index. The encoder therefore takes the input signs as they
are instead of flipping the weakest element to fix parity, and the
search compares magnitudes directly, which is simpler than its siblings.

### Why the codec lands before the kernel

The CUDA encoder's tests use this codec as their reference. They compare
against the PyTorch encoder byte for byte and draw the grid and scale
predictor from it. So the kernel cannot be tested before the codec
exists, and it follows in #2565 together with the registration. Every
registered format therefore keeps a CUDA encoder.

### Test changes that make the split possible

A codec can now land before it is registered, so two test contracts in
`test_iq_formats.py` are stated precisely:

- The two tests that go through `TensorQuantizer` (pass-through
gradient, error falls with bit width) iterate `IQ_FORMAT_REGISTRY`.
Every other battery test calls the codec directly and covers IQ2_S here.
- The coverage check now asserts `set(IQ_FORMAT_REGISTRY) <=
set(FORMATS)` instead of equality. That is what its docstring already
said: a registered format must be listed, or it escapes the contract.
- `test_registry_lists_every_exported_encoder` is unchanged, and it is
why this PR leaves the package exports alone: an exported encoder must
be registered.

The error-by-bit-width failure message also labels errors by the order
they were measured in; it previously zipped them with alphabetical
names.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ2_S: 9 tensors, 2,355,200 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry.
Blocks from `unsloth/Qwen3.8-27B-GGUF` ship as conformance vectors, so
CI keeps checking bytes we did not produce.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **121 passed**, 14 of them IQ2_S
codec cases, including the llama.cpp conformance check
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **35 passed**, unchanged by
this PR

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ The new
codebook is a GGML table, carried in `codebooks.py` with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2565
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
**this** → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added GGML-compatible IQ2_S quantization and dequantization support,
including access to its magnitude grid.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 14:03:01 -07:00
Edwardssss 57f929e358 Bound the stale-capture warning to once per capture and name its cause (#2492)
### What does this PR do?

Type of change: Bug fix

The activation-capture hooks latch `_intermediate_output` on every
forward, and only
`DistillationModel.compute_kd_loss()` ever clears it. Any training loop
that does not call
`compute_kd_loss()` therefore warns on every forward after the first,
once per hooked module:

```
plain transformers Trainer, 16 micro-batches
  before: 30 warnings (Teacher 15 + Student 15)
  after:   2 warnings (Teacher  1 + Student  1)
```

The count scales with forwards, not optimizer steps —
`gradient_accumulation_steps` of 1, 4 and
16 all produced 30 warnings for the same 16 forwards — so the
accumulator the report blamed is
not involved. A plain `transformers.Trainer` reproduces it because its
default `compute_loss`
reads the student's own CE from `outputs.loss` and never calls
`compute_kd_loss()`; the
warning is a symptom of that, but neither message said so. The student
message attributed the
re-forward to Activation Checkpointing only, and the teacher message
called the situation
"expected" while still raising `UserWarning`.

This PR:
- warns once per captured activation via
`_warn_once_about_stale_output`, re-armed wherever a
capture is cleared, so a consumed capture that is followed by another
unconsumed one is still
reported (Activation Checkpointing re-runs forwards and relies on that);
- states the cause and the consequence in both messages:
`compute_kd_loss()` did not run since
  the previous forward, so no KD loss is applied;
- moves the three `_intermediate_output = None` reset sites onto one
`_clear_captured_output`
helper so the capture and its warning state cannot drift apart, and has
the layerwise teacher
hook share both helpers rather than keeping a second copy of the
warning.

### Usage

No API change. A loop that applies KD should call `compute_kd_loss()`
once per forward:

```python
class KDLossTrainer(Trainer):
    def compute_loss(self, model, inputs, return_outputs=False, **kwargs):
        outputs = model(**inputs)
        loss = model.compute_kd_loss(student_loss=outputs.loss)  # consumes the captures
        return (loss, outputs) if return_outputs else loss
```

### Testing

`tests/unit/torch/distill/test_distill.py`:
- `test_duplicate_fwd_hook_call` now pins the bound under
`warnings.simplefilter("always")`
(3 forwards -> exactly 2 warnings) instead of relying on the
interpreter's per-message
  deduplication, and asserts the message names `compute_kd_loss`;
- `test_stale_output_warning_rearms_after_consuming` covers the re-arm:
warn, consume, warn
again -> 4 warnings. This one passes before and after the change; it
guards against a future
"warn once ever" simplification silently muting the Activation
Checkpointing case.

`tests/unit/torch/distill/test_layerwise.py`:
- `test_layerwise_stale_output_warning_is_bounded` covers the layerwise
hooks (3 forwards ->
  exactly 2 warnings).

The first and third tests fail on `main` (`assert 4 == 2`), verified in
a worktree of `main`
carrying these test files.

Measured with a `transformers.Trainer` over a 16-micro-batch run,
counting
`"already has an intermediate output stored"`:

| setup | before | after |
|---|---|---|
| plain `Trainer`, `gradient_accumulation_steps=1` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=4` | 30 | 2 |
| plain `Trainer`, `gradient_accumulation_steps=16` | 30 | 2 |
| `Trainer` calling `compute_kd_loss()`, accum=4 | 0 | 0 |

```
$ python -m pytest tests/unit/torch/distill -q
34 passed

$ python -m pytest tests/unit/torch -q --ignore=tests/unit/torch/deploy
1 failed, 2673 passed, 18 skipped in 217.17s
```

The single failure is
`tests/unit/torch/quantization/plugins/test_huggingface.py::test_dbrx`,
which fails identically on `main` in this environment because
`transformers 5.17` is outside the
`transformers>=4.57,<5.15` range pinned in `pyproject.toml`. It is
unrelated to this change.

`pre-commit run --files <changed files>` passes every hook (ruff check,
ruff format, mypy,
bandit, insert-license, large files, line endings).

Note on severity, since it affects how the issue reads: Python already
deduplicates a warning
by (message, location), so with the default filters this shows up twice
rather than 30 times.
The repetition is user-visible under `-W always`,
`PYTHONWARNINGS=always`, pytest, or one
warning registry per DDP rank, and the count is what scales with epoch
length. The behavioural
fix that matters for all filters is the latch.

### Before your PR is "*Ready for review*"

Is this change backward compatible?: ✅
If you copied code from any other sources or added a new PIP dependency,
did you follow guidance in CONTRIBUTING.md: N/A
Did you write any new necessary tests?: ✅
Did you update Changelog?: N/A
Did you get Claude approval on this PR?: N/A

### Additional Information

Fixes #2487.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved handling of captured activations during distillation,
including clearing stale outputs and resetting warning behavior after
outputs are consumed.
- Stale-output warnings now appear once per affected activation and
clarify when knowledge-distillation loss was not applied.
- Warnings distinguish expected cases involving activation checkpointing
or teacher evaluation.

- **Tests**
- Expanded coverage for warning counts, warning reset behavior, and
stale-output handling in standard and layerwise distillation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Edwardssss <ed_129@qq.com>
2026-09-28 22:01:07 +02:00
Keval MorabiaandClaude Opus 5 4eb86524f0 [1/2] One MLflow tracking core behind a Tool record (#2544)
### What does this PR do?

Type of change: refactor (no functional change)

**[1/2] of a split. Merge this first; #2514 is [2/2] and is based on
this branch.**

Three example scripts had each reimplemented the same MLflow wiring: the
flags, the `$USER/<tool>/<model>-<variant>` experiment convention, the
params/tags/artifacts a run uploads, and the open/close dance with its
status. The copies had already drifted — only `hf_ptq` wrote a
provenance pointer, only `vllm_serve` republished the resolved URI — and
every new tracked script meant another copy.

What a script records is now one declarative `Tool` record, **declared
in the script itself, beside the flags it reads**:

```python
# examples/megatron_bridge/quantize.py
QUANTIZE = Tool(
    name="megatron_bridge_quantize",
    tracks="Track this run on an MLflow server, uploading the command, the resolved recipe, ...",
    variant_help="recipe name, or --quant_cfg if no --recipe",
    variant=lambda args: Path(args.recipe).stem if args.recipe else (args.quant_cfg or "none"),
    model=lambda args: args.hf_model_name_or_path,
    checkpoint=lambda args: args.export_megatron_path,
    texts=lambda args: resolved_recipe_texts(args.recipe),
    outputs=lambda args: {"summary/quant_summary.txt": Path(args.export_megatron_path) / ".quant_summary.txt"},
)
```

`tracked_run` takes that record and runs the whole thing, so a script
adds tracking in three lines: `add_mlflow_args(parser, TOOL)`,
`resolve_mlflow_args(args, parser, TOOL)`, and `with mlflow_run(args,
TOOL):`. The shared module knows no script's flags.

`examples/hf_ptq`, `examples/vllm_serve` and
`examples/megatron_bridge/quantize.py` move onto it. Three helpers fall
away as redundant (`track_run`, `checkpoint_run_tags`, and `hf_ptq`'s
two flag pass-throughs).

### Usage

No user-facing change. The flags, their spellings and the experiment
naming are exactly as before; a script author now writes a `Tool`
instead of four functions.

### Testing

- `tests/unit/torch/utils/test_mlflow.py`,
`tests/examples/hf_ptq/test_hf_ptq_args.py`,
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` — **179 pass**.
- `tests/examples/megatron_bridge` in `nvcr.io/nvidia/nemo:26.08` (the
only lane that runs it), which drives `quantize.py` for real: **34
passed**, locally and in this PR's `megatron` lane.
- `pre-commit run --files <changed>`: all hooks pass.
- The four suites shared four copies of a stand-in for the `mlflow`
module, which had drifted — one recorded artifacts as a list, another as
a dict, a third made `log_artifact` a no-op, so a test asserting on an
upload asserted nothing. They now share one
`tests/_test_utils/mlflow.py`, which also emulates the fluent API's
habit of opening a run when none is active.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `track_run` and
`checkpoint_run_tags` are removed, but neither shipped in a release
(0.47.0's `__all__` is `MlflowRunLogger`, `command_text`,
`current_user`, `default_experiment_name`, `validate_tracking_uri`, all
unchanged here). Three deliberate behaviour changes, each in shared code
and each tested:
- The `source_checkpoint_path` tag resolves to an absolute path where it
recorded the raw argument, which a chain of runs needs to join on the
pair. `run_tags` is shared, so this applies to every script that writes
the tag — `hf_ptq` **and** `megatron_bridge/quantize.py`, for a local
`--hf_model_name_or_path`. A source that names no directory, such as a
Hub `org/name` id, is still recorded as given.
- `MlflowRunLogger.track()` — which *did* ship in 0.47.0 — records a
block ending in `SystemExit(0)` as `FINISHED` where it recorded
`FAILED`, since a script that ends by calling `sys.exit()` rather than
returning has still finished.
- `.experiment.json`'s `tracking_uri` and the `run_url` built from it
drop a trailing `/` from the tracking URI, so the link is
`https://host/#/...` rather than `https://host//#/...`. Only reachable
by constructing `MlflowRunLogger` directly; every CLI path already
stripped the slash in `resolve_tracking_uri`.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-visible change; the entry is in [2/2].
- Did you get Claude approval on this PR?: several rounds; re-requested
on this head.

### Additional Information

Split out of #2514. This half is the enabling refactor with no behaviour
change; #2514 is the feature it unlocks and is based on this branch. At
~605 changed lines of core logic it is over the ~500 guideline; the
owner accepted a two-PR split rather than three, and everything #2514
alone consumes — `split_tracking_credentials`,
`log_active_run_experiment_json`, `MlflowRunLogger._reattach` — lands
there rather than here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-28 19:57:03 +00:00
Chenjie LuoandClaude Opus 5.5 80e04b8816 docs(eval): align Terminal-Bench 2.1 / SWE-bench / MRCR with upstream configs (#2479)
### What does this PR do?

Type of change: documentation (agent skill) + template bug fixes

Aligns the `evaluation` skill's three upstream-tracked benchmarks
(Terminal-Bench 2.1,
SWE-bench Verified, MRCR) with the current
`nvidia-eval-factory-benchmarking` configs, and
fixes guidance that turned out to be wrong when a full three-benchmark
campaign was run with
the skill end to end. Rebased on #2499. The first commit is the
alignment. The rest address
review and a fresh config-generation test: MRCR serving scoped per
variant, `limit_samples`
canary guidance, the template's serve command passing
`--gpu-memory-utilization` (replacing
`command:` had silently dropped it, so the 1M golden's 0.95 never
reached vLLM), upstream's
128K values, and regression tests for the `++limit` gate and the serve
command. The skill
text keeps only the rules; the evidence behind them is below.

GDPVal is out of scope. It was removed from this skill in #2470, and
this PR replaces #2464.

**Alignment with upstream**

- Sandbox region via `HARBOR_ECS_REGION`. TB2.1's ECR repo name tracks
the region;
  SWE-bench's stays in us-west-2.
- One interceptor order for both harbor benchmarks, with
`http_pairs_dump` **last**. Upstream
is split on its position, which changes only what the dump records,
never the score.
Interceptor lists replace wholesale on merge, so a leaf must restate the
whole chain.
- `capture_request_body` goes on the service. A shared block injects an
alias-only entry with
  no `type`.
- MLflow tags gain `task_name` and `nemo-evaluator-next-version`.
- TB2.1 `max_concurrent`: 50 for nano-class models, 15 for larger ones
(all upstream non-nano
  leaves override it to 15).
- `proxy.request_timeout` must always be set explicitly. Otherwise it
inherits the model
fragment's serving value, which ranges from 3600 to 36000 upstream, and
TB2.1 has no
  benchmark key for it.
- MRCR:
- `parallelism` is 512, deliberately above server capacity, so
`--max-num-seqs` must no
    longer be derived from it.
  - `limit_samples` now reaches the gym through a gated `++limit`.
  - Observability capture is on.
  - 128K is its own upstream benchmark on the condensed gym schema.

**Corrections found by running it**

| what the skill said | what actually happens |
|---|---|
| `username: ${oc.env:USER}` | nel-next only expands `${VAR}` /
`${VAR:-default}`. The config passes `--dry-run` and fails at `--submit`
with *"remote username contains invalid characters"*. |
| MRCR canary via `++limit` edited into `collect_rollout_params` |
`limit_samples` is now gated through. Under the 0.2.6 launcher the `-o`
path is
`++evaluation.nemo_evaluator_config.config.params.limit_samples`.
`++config.params…` creates a bogus top-level key. |
| the condensed gym schema's bootstrap is in the runtime image | It
lives in upstream `configs/models/gym_eval_command.yaml`, which is
composed in. A standalone config must carry the `command:` block. |
| `mean/prefix_matched ~0.55 is healthy` | That value is calibrated to
the 1M golden. A 128K run at `pass@1` ≈ 95 measured ≈ 1.0. The signal is
a collapse toward 0. |
| NVFP4 MoE `VLLM_*` env vars as reliable knobs | They are
build-dependent: one vLLM build logged them as unknown and ignored them.
Check the server log once per image. |

**Rules the skill lacked**

- **MRCR variant.** Pick the largest variant within the checkpoint's
trained context. On a
262K-context model, 1M needs `VLLM_ALLOW_LONG_MAX_MODEL_LEN` and
measures extrapolation, which
a quantization comparison would then entangle with quantization damage.
The serving setup
follows the variant: 128K serves at the trained context without the
override.
- **SWE-bench `reasoning_effort`.** openhands-sdk sends
`reasoning_effort: high` on every call,
and canonical `bench.yaml` doesn't strip it. A server whose accepted set
excludes `high`
returns HTTP 400 on the first call of every trial, so `pass@1` is 0. The
fix is to overwrite
it with the server default via `proxy.extra_body`. terminus-2 (TB2.1)
and Gym's
`simple_agent` (MRCR) never send the key (46/46 and 110/110 requests
checked).
- **Reasoning toggles** go in `extra_body.chat_template_kwargs`. Don't
write out no-op sampling
  defaults; `top_k` is *not* one (vLLM's default is `-1`).
- **Upstream model fragment.** Consult `configs/models/<model>/` when it
exists, for serving
  flags, the thinking toggle and `reasoning_replay.mode`.

### Usage

No API change. Regenerating a config from the skill now yields the
aligned values:

```yaml
# recipes/examples/example_eval_next.yaml
services:
  model:
    proxy:
      request_timeout: 3600                  # always explicit
benchmarks:
  - max_concurrent: 50                       # nano-class; larger models 15
    sandbox:
      region: ${HARBOR_ECS_REGION:-us-east-1}
cluster:
  username: ${USER}                          # NOT ${oc.env:USER}
```

### Testing

- `python -m pytest plugins/modelopt/skills/ -o addopts=""`: 7/7 pass,
including the new
`tests/test_example_mrcr.py`. It checks that `example_mrcr.yaml` emits
`++limit=N` only when
`limit_samples` is set, and that the folded vLLM serve command
shell-parses with every flag
intact, including `--gpu-memory-utilization`. Each check fails when its
defect is reintroduced.
- `markdownlint-cli2` on the changed Markdown: 0 errors. Both example
YAMLs parse.
The full `pre-commit` suite was not run after the squash, because the
sandbox could not fetch
  hook repos. The pre-squash commits passed it.
- **Exercised end to end.** Configs built from this skill ran a BF16
campaign for a 262K-context
  MoE reasoning model on an internal cluster to completion:
  - MRCR-128K: 1470/1470 rollouts
  - Terminal-Bench 2.1: 712/712 trials
  - SWE-bench Verified: 2500/2500 trials

  Each correction above is a defect that campaign surfaced.
- **Differential check.** Terminal-Bench configs generated from the pre-
and post-change skill,
  from the same brief and in isolation, differ on:
  - interceptor chain
  - concurrency
  - `top_k`
  - region interpolation
  - MLflow tags

Not run: a scored evaluation of this PR itself.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- Docs + a template fix; no
ModelOpt API surface touched. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- Documentation and
config-template values. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- Agent skill docs, not a user-facing ModelOpt feature/breaking
change/deprecation. -->
- Did you get Claude approval on this PR?: ❌ <!-- Not run. -->

### Additional Information

The companion internal `eval-config` change now carries only the
internal values: the ECR URLs,
the region default and the cluster image notes. It points here for the
generic rules.

**Known gaps not fixed here**, worth a follow-up:

- `references/nel-next.md:136` and `references/launcher-workflow.md:207`
derive `--max-num-seqs`
from `parallelism / DP`. nel-next has no `parallelism` field (its
analogue is
  `max_concurrent`), and MRCR's 512 is deliberately not a server cap.
- `references/launcher-workflow.md:200-201` makes
`--max-num-batched-tokens` and
`--enable-chunked-prefill` always-include defaults, but
`example_eval_next.yaml` omits both.
- `references/launcher-workflow.md:27` says `sbatch_comment` belongs
under `execution:` and is
otherwise inert, yet all three shipped examples put it under `cluster:`.
- MoE detection (`--enable-expert-parallel`) is unresolvable from the
facts the skill asks for
  when the model handle has no `-A*B` suffix.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:07:35 +00:00
h-guo18andClaude Opus 5 23355eda90 fix(deps): declare httpx, unbreaking partial-install (torch) for every PR (#2547)
## Summary

`partial-install (torch)` has been failing on **every** PR since
2026-09-24 — including PRs whose branches predate the breakage — and
because it is a *collection* error rather than a test failure, it aborts
the entire run:

```
ImportError while importing test module '.../tests/unit/torch/speculative/plugins/test_hf_streaming_dataset.py'
E   ModuleNotFoundError: No module named 'httpx'
collected 2243 items / 1 error / 45 skipped
!!!!!!!!!!!!!!!!!!!! Interrupted: 1 error during collection !!!!!!!!!!!!!!!!!!!!
```

`unit-pr-required-check` aggregates it, so nothing currently merges on a
fresh run.

## What happened

**No code changed.**
`modelopt/torch/speculative/plugins/hf_streaming_dataset.py` has
imported `httpx` at module scope since #1509 (2026-06-02), and `httpx`
has never appeared in `pyproject.toml`. It arrived only transitively:
`dev-test` → `timm` → `huggingface_hub` → `httpx`.

**huggingface_hub 2.0.0**, published **2026-09-24T12:01:21Z**, replaced
`httpx<1,>=0.23.0` with the separate **`httpx2<3,>=2.0.0`**
distribution. Different package name, so `httpx` stopped being installed
and the chain disappeared.

The boundary is exact — every run *created* before that timestamp
passes, every one after fails:

| PR | run created | result |
|---|---|---|
| #2536 / #2535 | 09-23 22:02 | pass |
| #2500 | 09-23 23:45 | pass — **merged 09-24 20:01 on this stale-green
result** |
| *hub 1.33.0 (still requires httpx)* | *09-24 09:49* | |
| **hub 2.0.0 published** | **09-24 12:01** | ← |
| #2539 | 09-24 16:57 | fail |
| #2544 | 09-24 18:32 | fail |
| #2216 | 09-25 11:58 | fail |

#2500 merging afterwards is not a counterexample: GitHub does not re-run
checks at merge time, so it merged on a result from ~20 hours earlier.
That is also why this went unnoticed.

## The changes

### 1. Declare `httpx` in the `hf` extra

`httpx` is not incidental to streaming — it is the only transport:

- every fetch is HTTP: `POST /v1/completions` to the vLLM serve plus
`GET /meta` and `/desc` against the connector's sidecar, all through
`httpx.Client`;
- there is no non-HTTP path — the base `StreamingDataset._fetch` is an
abstract seam and `EagleVllmStreamingDataset._fetch` is its only
implementation;
- no other HTTP library appears in the module (`requests` / `urllib` /
`aiohttp`: zero hits, and `requests` is not declared either);
- even the retry predicate is built from it: `_TRANSIENT_FETCH_ERRORS =
(httpx.HTTPError, OSError)`.

It belongs in `hf` rather than in the core `dependencies`: the same
module needs `transformers.trainer_pt_utils` at module scope, so one
extra already gates the whole file, and a core install has no use for an
HTTP client. The bound matches the 0.x API the code uses — `httpx` has
no 1.0 release, and 2.x is a different distribution.

This is the part that stops it recurring. `[hf]` currently gets `httpx`
only because `datasets` happens to require it — the same accident with a
different supplier, one release away from repeating.

### 2. Acquire `httpx` in the test through the existing skip guard

The test file already intends to skip where the extra is absent — it has
`pytest.importorskip("transformers")` and a comment explaining why, and
`transformers` is absent in this job too. It broke only because `import
httpx` sat **five lines above** that guard, where a missing module ends
collection instead of skipping one file.

## Verification

- With everything installed: **18 passed**, no behaviour change.
- The import-order property is checked with an AST walk over the
module's top-level statements: no `hf`-extra-only import precedes the
first `importorskip` (which is now line 41).
- A faithful local reproduction was attempted and abandoned honestly:
hiding `httpx` locally also breaks `huggingface_hub` 1.28, which
`modelopt.torch.opt.plugins.huggingface` imports, so the local failure
is not the CI one. CI is the oracle for that half — this PR's own
`partial-install (torch)` run is the check that matters.

## Scope

Two files, five lines of declaration and four of test import order.
Deliberately not folded into any feature PR: it blocks the whole repo,
and burying a repo-wide fix inside unrelated work is how these stay
invisible.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Optional Hugging Face installations now include `httpx`, supporting
features that require HTTP communication without requiring it for all
installations.
* **Tests**
* Hugging Face streaming dataset tests now skip when `httpx` is
unavailable, allowing the remaining test suite to be collected and run
without it.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-26 04:27:24 +05:30
hychiang ed7e87953c Fix grouped expert quantizer checkpoint replicas (#2500)
### What does this PR do?

Type of change: Bug fix

Fix distributed checkpoint saving for quantized Transformer Engine
grouped MoE experts when tensor parallelism and expert parallelism are
both greater than one.

#### Observed error

With NeMo 26.08 and TP=2, EP=4, ETP=1, quantization and calibration
complete, but the run crashes while saving the distributed checkpoint:

```text
megatron.core.dist_checkpointing.core.CheckpointingException:
Invalid sharding pattern validation.
Invalid access pattern for ShardedTensor(
    key='decoder.layers.1.mlp.experts.experts.32.linear_fc1.weight_quantizer._amax',
    ...
)
```

#### Root cause

`_QuantMegatronTEGroupedLinear.sharded_state_dict` created each expert
quantizer's sharded tensors without passing the expert tensor-parallel
and expert data-parallel process groups to
`make_sharded_tensors_for_checkpoint`. It then manually replaced only
the expert-data-parallel component of `replica_id`.

That manual rewrite is insufficient when tensor and expert parallelism
are both enabled: replica ownership is derived using the wrong
process-group topology, so ranks can publish an inconsistent access
pattern for the same globally indexed expert quantizer key.
Megatron-Core correctly rejects that checkpoint during sharding
validation.

#### Fix

Resolve the expert model-, tensor-, and data-parallel groups from one
`_pg_collection`-aware helper, then pass the expert TP/DP groups to
`make_sharded_tensors_for_checkpoint` for per-expert quantizer buffers.
This avoids mixing process-group sources when models use a non-global
`ProcessGroupCollection` or grouped child modules convert without their
parent MLP, and removes the manual `replica_id` rewrite. Shared,
whole-linear quantizer buffers deliberately retain the default dense
TP/DP replica groups because their keys and offsets carry no expert
identity; using only expert TP/DP groups would make replica IDs collide
across EP ranks.

#### Relationship to PR #2319

[PR #2319](https://github.com/NVIDIA/Model-Optimizer/pull/2319) fixed
two earlier TEGroupedMLP checkpoint problems:

1. Per-expert quantizer state did not retain the same globally unique
expert identity as the grouped-expert weights, so it could not be
redistributed reliably when the EP layout changed.
2. After ModelOpt extra-state restoration, an expert that moved to a
different rank could be missing the `_amax` or `_global_amax`
destination buffer required by the subsequent distributed checkpoint
load.

That PR therefore fixed resharding and restore correctness across
topology changes. Its regression matrix changed one parallel dimension
at a time: EP changed while TP=1, or TP changed while EP=1.

The remaining issue was on the save path when TP and EP were
simultaneously greater than one. Although the expert keys and restore
buffers were correct after #2319, the per-expert quantizer shards were
still constructed without the expert TP/DP process groups and then had
only part of their `replica_id` rewritten manually. Under TP=2, EP=4,
ETP=1, this produced the invalid access pattern rejected by
Megatron-Core before the checkpoint could be saved.

This PR complements #2319 by fixing that process-group/replica mapping
and adding combined TP+EP coverage.

A four-rank save-and-restore regression test covers combined TP=2, EP=2,
and ETP=1.

### Usage

N/A. This fixes checkpoint behavior without changing the public API.

### Testing

- [x] Focused Ruff check and format validation
- [x] Focused mypy validation
- [x] `git diff --check origin/main..HEAD`
- [x] Four-rank grouped-expert checkpoint save/restore test with TP=2,
EP=2, ETP=1: 1 passed on the original fix
- [x] End-to-end NeMo 26.08 quantization and checkpoint save on 8 B200
GPUs with TP=2, EP=4, ETP=1 on the original fix
- [x] Review-response validation on `dbfa1cdb5`: Ruff 0.15.20
format/check, `git diff --check`, and Python compilation
- [x] Final four-rank grouped-expert checkpoint regression on
`c2be0f2c6`: TP=2, EP=2, ETP=1 with both ordinary and overridden
ModelOpt TP state; 2 passed in 459.38s on 4 B200 GPUs with NeMo 26.08

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

The end-to-end validation produced a complete eight-shard distributed
checkpoint and exited successfully.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Fixed checkpoint saving for quantized grouped MoE experts when tensor
and expert parallelism are enabled.
- Improved handling of shared and per-expert quantizer state across
parallel configurations to support more reliable checkpoint save and
restore.

- **Tests**
- Added NVFP4 grouped expert checkpoint round-trip coverage with tensor
and expert parallelism, including configurations with the
tensor-parallel group override enabled and disabled.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
2026-09-24 20:01:56 +00:00
JoshuaandCursor 63c4b660bd Add Aumann-Shapley sensitivity scoring method to auto_quantize (#2183)
[Paper](https://arxiv.org/abs/2607.12266) ·
[Overview](https://x.com/waterloo_intern/status/2076460984475263401) ·
[Implementation
thread](https://x.com/the_joshua_hill/status/2076427869388255635)

Depends on #2231, which provides the shared AutoQuantize
backward-scoring infrastructure. Until that PR merges, the focused PR B
diff is available
[here](https://github.com/joshua-hill/Model-Optimizer/compare/fix/autoquant-scoring-infrastructure...feat/aumann-shapley-autoquant).

### What does this PR do?

Type of change: new feature

This PR adds `method="aumann_shapley"` to `mtq.auto_quantize`. It is a
label-free scoring method that measures how each candidate quantization
format affects the model across the path from full precision to
quantized.

The new method is opt-in; existing `gradient` and `kl_div` behavior is
unchanged.

For each calibration batch, the method:

1. Runs the baseline model and saves its next-token distribution.
2. Measures each candidate format at a configurable number of points
along the quantization path.
3. Uses the KL-divergence gradients at those points to assign a damage
contribution to every runtime group and candidate format.
4. Measures the most aggressive candidate configuration once and uses
that value to calibrate the per-group damage model.

The resulting scores use the existing AutoQuantize linear-program
solver. The search can either:

- choose the least damaging configuration that meets an `effective_bits`
target; or
- choose the smallest configuration whose predicted damage stays below
`max_predicted_damage`.

The selected recipe records `predicted_damage` in mean per-token KL
units together with its validity and fit diagnostics.

### Public API

`auto_quantize` gains an optional `method_options` dictionary. For
`method="aumann_shapley"`, it accepts:

| Option | Default | Meaning |
|---|---:|---|
| `num_path_nodes` | `2` | Number of points used to average gradients
along the quantization path. |
| `damage_link` | `"coverage"` | How per-group scores combine.
`"coverage"` uses `damage = c * (1 - exp(-sum(b)))`; `"additive"` sums
the path contributions. |
| `max_predicted_damage` | `None` | Replaces the bit target with a
maximum predicted mean per-token KL. |

Method options are validated before the model is modified. Unknown
options and incompatible targets fail early.

### Usage

Select a configuration for a target effective bit width:

```python
import modelopt.torch.quantization as mtq

model, search_state = mtq.auto_quantize(
    model,
    constraints={"effective_bits": 4.8},
    quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"],
    data_loader=calib_loader,
    forward_step=lambda model, batch: model(**batch),
    method="aumann_shapley",
)

print(search_state["best"]["predicted_damage"])
print(search_state["best"]["predicted_damage_valid"])
```

Or let the search choose the bit width for a predicted-damage target:

```python
model, search_state = mtq.auto_quantize(
    model,
    constraints={},
    quantization_formats=["NVFP4_DEFAULT_CFG", "FP8_DEFAULT_CFG"],
    data_loader=calib_loader,
    forward_step=forward_step,
    method="aumann_shapley",
    method_options={"max_predicted_damage": 0.05},
)
```

### Implementation

- Reuses the candidate-replay and backward-scoring lifecycle introduced
in #2231.
- Reuses the existing AutoQuantize linear-program solver for both search
directions.
- Numerically integrates the coverage path when converting measured
contributions into per-group damage costs.
- Preserves deterministic runtime-group and candidate ordering.
- Keeps raw measurements, solver scores, and damage-model diagnostics
distinct in the search state.
- Rejects incompatible checkpoint resumes while allowing the same scores
to be re-solved for a new bit budget.
- Retains the shared MoE score-module rules so routed experts are scored
at their enclosing block.

Recipe integration will follow separately.

### Testing

Focused tests:

```text
pytest -q \
  tests/unit/torch/quantization/test_autoquant.py::test_backward_scoring_session_restores_partial_setup \
  tests/unit/torch/quantization/test_autoquant_shapley.py
```

Result: **51 passed**.

The tests cover:

- end-to-end scoring and configuration generation;
- agreement between path contributions and measured quantization damage;
- exact allocation checks against exhaustive search;
- effective-bits and predicted-damage search modes;
- checkpoint resume and offline re-solving;
- custom formats and heterogeneous candidate ladders;
- distributed reductions and nested MoE score modules;
- reused score modules and model-specific backward support;
- non-finite measurements and invalid-fit reporting; and
- input validation before model conversion.

End-to-end checks with NVFP4 and FP8 candidates at a 6.0-bit target:

- `Qwen/Qwen2.5-0.5B-Instruct` reaches 5.998 effective bits, with the
summed path contributions reproducing 98% of the directly measured
lowest-precision KL.
- `Qwen/Qwen3-30B-A3B` reaches 6.000 effective bits and reproduces 99%,
with all 48 MoE layers scored once at the sparse-MoE block rather than
per expert.

### Production use

We use this method in production for NVFP4 checkpoints of Kimi-K3,
MiniMax-M3, and GLM-5.2.

### Before your PR is "Ready for review"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow the contributing guidance?: ✅ No copied code
and no new dependencies.
- Did you write the necessary tests?: ✅
- Did you update `CHANGELOG.rst`?: ✅
- Are the commits signed and signed off?: ✅


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
  - Added label-free Aumann–Shapley scoring for automatic quantization.
- Added configurable path sampling, damage modeling, effective-bit
targets, and predicted-damage bounds.
- Added temporary weight-folding support with automatic state
restoration.
- Added method-specific search options, checkpoint resumption, and
distributed scoring.

- **Bug Fixes**
- Improved cleanup and restoration of quantizer state, gradients, hooks,
and forward behavior after scoring or failures.
- Added validation and clearer handling for unsupported configurations
and invalid measurements.

- **Tests**
- Expanded coverage for scoring, solver behavior, distributed execution,
checkpointing, and custom quantization formats.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Joshua Hill <joshua.hill@baseten.co>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-23 20:51:55 -07:00
Chenjie LuoandClaude Opus 5.5 400498d82d [2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do?

Type of change: refactor (no behaviour change)

Addresses review feedback on #2511. Backend dispatch and export each
kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the
backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in
export. All four listed the same formats. Adding a format meant a row in
each, and the lists could drift apart. That had already happened twice
in #2511: `convert_hf_config.py` kept its own upper-case spelling of the
family and dropped IQ2_XXS metadata, and the Megatron export tests were
hard-wired to two formats.

Each format module now declares **one `IQFormat` record** beside its
encoder and decoder: name, block geometry, `quantize`, `dequantize`, and
its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists
them.

- Backend dispatch looks formats up in the registry.
- Both exporters take the packer and block geometry from it.
- Export's `IQ_FORMATS` is derived from it instead of being written out
again.
- `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed.
- The per-format fake-quant wrappers collapse into one
`IQFormat.fake_quant`, which does the `num_bits` check and calls the
existing cache helper.

Codebooks, searches, payload layouts and CUDA encoders stay in each
format's module.

#### Series and merge order

This is one slice of the IQ format series. It targets `main` so unit CI
runs, and **its diff includes #2511's commits until #2511 merges**.

1. #2511 — IQ2_XXS format
2. **this PR** — one registration per format
3. #2512 — IQ2_S format
4. #2513 — IQ1_M format

After this lands, #2512 and #2513 are restacked onto it, so each adds a
format module and a single registry entry instead of rows in four
tables.

#### Design choices

- **An explicit list, not self-registration at import.** If formats
registered themselves when their module was imported, the registry's
contents would depend on import order.
- **Backward compatible, with one behaviour change.** `iq1_s_fake_quant`
and `iq2_xs_fake_quant` are public on main, so each format keeps its
`<fmt>_fake_quant` name as an alias of its record's method. The three
removed tables were introduced by #2511 and never released. The
behaviour change: on main, the alias looked the encoder up at call time,
so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record
captures the encoder and decoder when it's built, so patching those
module functions reaches neither dispatch nor the alias. Substitute
through `IQ_FORMAT_REGISTRY` instead.
- **Registering a format declares it exportable, and that's intended.**
Export's `IQ_FORMATS` is derived from the registry, so a format
registered for dispatch is also claimed by both exporters and
`convert_hf_config`. That can't be wrong for an IQ format: fake quant is
`dequantize(quantize(w))`, so a format can't be dispatched without the
packer and block geometry, and those are all export reads. A QAT-only IQ
format can't exist. If one ever needs to land ahead of its export path,
an `exportable` flag on the record is a one-line addition.
- **The registry is the substitution seam.** Dispatch now reads the
registry, so tests that swap an encoder or decoder swap the registry
entry. Patching the format module's function would no longer reach
dispatch.
- **Test expectations stay independent of the registry.** Tests take the
*list* of formats from the registry, but their expected values come from
each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A
mis-wired registry entry therefore can't make both sides of an assertion
agree.

#### What it does not unify

The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list,
codebook sizes in `common.cuh`) and the recipes and docs remain per
format. "One registration" holds for the Python side, which is where all
four tables lived.

### Usage

Adding a format after this PR (for example IQ2_S in #2512) needs its
module and one line in the registry:

```python
# modelopt/torch/quantization/ggml/iq2_s.py
IQ2_S_FORMAT = IQFormat(
    name="iq2_s",
    block_size=IQ2_S_BLOCK_SIZE,
    block_bytes=IQ2_S_BLOCK_BYTES,
    quantize=quantize_iq2_s,
    dequantize=dequantize_iq2_s,
    block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE,
    decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE,
)

# modelopt/torch/quantization/ggml/registry.py
IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)}
```

Looking up a format:

```python
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY

fmt = IQ_FORMAT_REGISTRY["iq2_xxs"]
packed, shape = fmt.quantize(weight)          # GGML blocks
fmt.block_bytes, fmt.effective_bits           # 66, 2.0625
```

### Testing

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py` — **94 passed**
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed**
(RTX PRO 6000)
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses
for that suite
- broader sweep of IQ, export and recipe unit tests — **153 passed**,
none failed

**New guards on the registry itself:**
- every encoder the package exports is registered
- each record points at its own format's codec, geometry and chunk
defaults
- the public `<fmt>_fake_quant` alias is the registered record's method
- export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the
registry
- a format's `fake_quant` refuses a quantizer configured for another
format. Dispatch picks the record by `num_bits`, so it never reaches
this guard; the test covers direct callers of a record or alias. The
three per-format guards it replaced were untested on main.
- every registered format is listed in the shared test batteries

Checked by mutation: leaving IQ2_XXS out of the registry, or registering
it with the IQ2_XS encoder, each fails the guard written for that case.

**Coverage gap closed along the way:** `test_ggml_backend.py` was
hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or
packed-once coverage. Those tests now run over the registry.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — public per-format fake-quant
names are kept as aliases; the removed tables were never released.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — internal refactor with no
user-visible change
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`,
`IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently
enumerate the same formats. A common pack/dequantize/fake_quant
interface would let backend dispatch and export consume one
registration."*

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* IQ quantization formats are available through a shared format
registry, keeping format details and quantization behavior consistent
across supported workflows.
* IQ-format model exports use registered format information for
quantization metadata and weight packing.
* **Tests**
* Expanded checks to cover registered IQ formats and verify consistent
format support across quantization and export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 23:32:11 +00:00
Frida HouandClaude Opus 5 1c4cde7788 Fix: Release per-layer expert weights in layerwise export under offload (#2466)
### What does this PR do?

Type of change: Bug fix

Layerwise export leaks one layer's worth of quantized tensors per layer
on offloaded models, and dies of OOM partway through a large MoE. Two
things accumulate, both because the offload window cannot reclaim what
the export pass adds inside it.

**1. Per-expert holder modules.** `_export_fused_experts` splits a fused
MoE experts module into per-expert holders and attaches them to the live
model:

```python
proj = nn.Module();  proj.weight = wrapper.weight   # packed U8
expert.add_module(proj_name, proj)
module.add_module(str(idx), expert)
```

They are plain `nn.Module`s built *inside* the weight-access window, so
they carry no accelerate `_hf_hook`.
`weight_access_and_writeback_context` closes by iterating the modules it
collected at entry and calling `hook.post_forward()` through each one's
own offload hook — the holders satisfy neither condition, so nothing
returns them to meta.

**2. Scale buffers.** `_export_quantized_weight` registers
`weight_scale` / `weight_scale_2` / `input_scale` on the layer's
pre-existing, hooked sub-modules, and `AlignDevicesHook.post_forward`
runs with `offload_buffers=False`:

```
after post_forward: {'weight': 'meta', 'weight_scale': 'cpu'}
```

The packed weight goes back to meta; the scales do not. For
`_QuantMoELinear` models this half is the whole leak on its own —
`_reconstruct_fused_moe_linear` restacks every expert's scales into one
`register_buffer` on the hooked wrapper.

A whole-model export never notices: one pass, write the state dict,
exit. Layerwise runs the same pass once per decoder layer, so every
finished layer stays resident.

### Measurements

Through unmodified `examples/hf_ptq/hf_ptq.py` on a Qwen3.5-MoE-shaped
model (10 layers, 64 experts), printing `torch.cuda.memory_allocated()`
after each exported layer:

| placement | per layer | over 9 layers |
| --- | --- | --- |
| offload | +0.052 GiB | 0.579 → 1.047 GiB |
| resident | -0.135 GiB | falls, as designed |

+0.052 GiB is exactly one layer's quantized experts: packed U8 100.7M/2
= 0.047 GiB plus FP8 block scales 100.7M/16 = 0.006 GiB.

At scale it is fatal rather than wasteful. Qwen/Qwen3.8-2.4T-A95B (92
layers, 512 experts) leaks 12.9 GB packed + 1.6 GB scales per layer, so
92 layers want 1.33 TB that no budget on a 283 GB card or 952 GB host
absorbs. The run died of CUDA OOM at layer 16/92 with
`--max_gpu_memory_gb 240`, and at `--max_gpu_memory_gb 30` leaked the
same 15 GB/layer onto the host instead.

### The fix

`_release_exported_tensors` (`model_utils.py`) is a context manager that
snapshots each sub-module's child-module and buffer names on entry, and
on exit drops whatever appeared. Persisting happens *inside* the block,
so "release only once it is on disk" is structural rather than a
comment.

Both packing sites use it: `LayerwiseExporter.export_layer` and the
offload decoder loop in `_export_transformers_checkpoint_streaming`.

The streaming writer had solved the same leak inline with a heuristic —
null every CUDA buffer, and every CUDA parameter on a hook-less module —
and that block is deleted in favour of the shared helper. Keying on
*what the pass added* rather than on device and hook presence drops two
assumptions that only held for a terminal, offloaded export: it no
longer nulls buffers the layer already had, nor parameters of
sub-modules accelerate simply did not hook. That is also what makes it
safe for the layerwise path, where resident models are supported and the
model outlives the export.

Deliberately out of scope: the FSDP2 per-unit loop in
`collect_export_tensors` keeps its per-unit holders. That predates this
PR, this PR does not touch that loop, and closing it needs its own
change and its own test.

### Usage

No API change. Existing layerwise export under offload simply stops
growing:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3.8-2.4T-A95B \
    --qformat nvfp4 \
    --export_path /path/to/export \
    --max_gpu_memory_gb 240
```

### Testing

- `tests/unit/torch/export` and `tests/unit/torch/quantization` — 1246
passed
- `tests/gpu/torch/export/test_layerwise_export.py` +
`test_offload_export.py` — 35 passed
- `cuda_alloc` over 10 offloaded layers: 0.526 → 0.518 GiB (-0.008), was
+0.468
- Exported checkpoint byte-identical to the unfixed run: 7841 tensors, 0
mismatches, max abs diff 0.0; `hf_quant_config.json` / `config.json` /
index identical
- Full Qwen3.8-2.4T-A95B PTQ then completed all 92 layers: peak GPU 117
GB of 283, peak host RSS 71 GB of 952, flat across 50 consecutive layers
at 23-26 s/layer

Coverage added: the existing `test_export_creates_per_expert_submodules`
now runs the export inside the context manager and asserts the holders
are gone on exit, and a new test in `test_offload_export.py` pins the
`offload_buffers=False` behaviour the buffer half exists for —
export-registered scales dropped, pre-existing buffers untouched.

`tests/gpu/torch/export/test_fsdp2_export.py` reports 34 failures in my
environment. They are **pre-existing and unrelated**: the same 34 fail
identically on this branch and on the merge-base (`2b1f33d0ef`), with
byte-identical failure sets and runtimes within 3 s. All 34 are `Failed:
Timeout (>120.0s)` from hung NCCL collectives, with zero assertion
failures.

The GPU export suites and the whole-model measurements above were run at
`2999d7cfb7`. The two commits since — swapping the holder marker for a
child-name diff, and moving the helper to `model_utils` — are covered by
the unit suites; a GPU re-run before merge is worthwhile.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — one new offload test for
the buffer half, and the existing fused-experts export test now covers
holder release
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — `layerwise.export_dir` is new in the unreleased 0.48.0, so this
bug was introduced and fixed within the same cycle
- Did you get Claude approval on this PR?: ✅

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 14:43:12 -07:00
Asha AnooshehandClaude Sonnet 5 d16dad1c20 docs(llm_distill): document KDTrainer-based example flow (#2524)
## Summary
- Update `examples/llm_distill/README.md` to describe the current
`main.py` flow, which uses `KDTrainer` (from
`modelopt.torch.distill.plugins.huggingface`) instead of `mtd.convert()`
/ `DistillationModel` wrapping.
- Document that `KDTrainer` only supports logit-level distillation
today, and that hidden-state/intermediate-layer KD still requires
`mtd.convert()` + `DistillationModel` until `KDTrainer` gains that
support.

## Test plan
- [x] Reviewed rendered README diff for accuracy against
`modelopt/torch/distill/plugins/huggingface.py` and
`examples/llm_distill/main.py`
- [x] `pre-commit` hooks (markdownlint-cli2, etc.) passed on commit

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated the Hugging Face getting-started example to use `KDTrainer`
with the standard training and model-saving workflow, including an
example of combining it with `SFTTrainer`.
* Clarified that `KDTrainer` supports logit-level distillation;
hidden-state distillation uses a separate approach.
* Explained that KD loss and evaluation cross-entropy are reported
separately, with weighted CE/KD loss combination unsupported.
* Added guidance on distributed training options, including the FSDP2
requirement and alternatives to default DataParallel.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-23 23:06:45 +02:00
Chad Voegele 0fdda7937b Date 0.47.0 changelog for official release (#2529)
Set the 0.47.0 changelog date to 2026-09-23. No code changes.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated the 0.47.0 changelog entry to show September 23, 2026, as the
release date. No user-facing product behavior changes are included in
this update.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-09-23 18:54:41 +00:00
Chenjie LuoandClaude Opus 5 a21411adde Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do?

Type of change: new feature

llama.cpp defines five GGML IQ formats at one and two bits; we ship two.
This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and
IQ2_XS, and is the **first of three**.

On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`,
`Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84
B parameters — 10.6% of the file**, which a reader limited to
IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats
account for 17.3%.

| format | bpw | bytes/256 | codebook | |
|---|---|---|---|---|
| `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing |
| **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this
PR** |
| `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing |

The encoder follows the existing single-pass grid search at a fixed
anchored super-block scale, and the CUDA kernel the existing per-block
structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a
4-bit sub-block scale into the same 32-bit word as four 7-bit sign
indices, and its 256-entry grid needs no high index bits.

### Groundwork the next two reuse

Two things land here because IQ2_XXS is the first format to need them:

- **Export registry.** The IQ family was spelled as a two-element tuple
at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and
`unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset
plus per-format packer and block-geometry tables, so a format is a row
rather than a sweep through the exporters.
- **Shared test contract.** The per-format test files had drifted apart
— each of `iq1_s` and `iq2_xs` tested things the other did not. They
become one parametrized module per layer (unit and CUDA), so every
format is held to the same contract and a new one inherits it.

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs
```

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared
against `dequantize_row_iq2_xxs` from `ggml-quants.c`:

```
IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0
```

The new codebook matches the `ggml-common.h` table entry for entry, as
does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint
ship as conformance vectors so CI keeps checking bytes we did not
produce; mutation testing confirms they catch a wrong sign-field width.

The CUDA encoder is byte-identical to the PyTorch reference on a fixed
input and runs at **1047.9 M elem/s against the torch search's 10.7** on
a 5632×2048 weight.

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99
passed
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7
checks × 3 formats)
- `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds
29 recipes, `ptq.md` updated
- reconstruction error decreases monotonically with bit width, pinned by
a test

Pre-existing failures in `tests/unit/torch/export/` and
`test_autoquant.py` are `transformers`/`torchvision` import problems in
my environment — identical counts with and without this change.

### A finding about already-merged code

Checking the new kernel against its PyTorch reference at 4096 blocks
showed that **CUDA and torch encoders disagree on roughly 1 block in
6000 — including the already-merged `iq2_xs`**, at 0.0163% against
IQ2_XXS's 0.0000%.

Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA
fuses it with `fmaf` while torch uses separate ops; where two local
scales fall within a float32 ULP the roundings pick different sides.
Adjudicated against float64, neither path is better (5 to 6). Worst-case
cost is **1.48e-08** relative reconstruction error, and run-to-run
determinism on a given device holds.

This is pre-existing, not introduced here — `test_iq2_xs_cuda.py`
asserts exact byte parity but on a 16-block weight where ties
essentially never arise. I have **not** changed that test; rewording a
guarantee on merged code belongs in its own change. The new shared GPU
tests assert exact parity on a small fixed input and compare
reconstruction error at scale.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new
codebook is a GGML table, carried in `codebooks.py` beside the existing
ones so the MIT-licensed surface stays in that one file, with the source
revision recorded. No new dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

First of three; **IQ2_S** and **IQ1_M** follow and build on this branch.
Replaces #2505, which carried all three at once. Follows #2446 / #2447 /
#2448 / #2449, which landed IQ1_S and IQ2_XS.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added IQ2_XXS weight-only quantization, including CUDA acceleration
and support for Hugging Face and Megatron exports.
- Added the `general/ptq/iq2_xxs` recipe. It requires no calibration
data and supports eligible layers with a weight dimension divisible by
256.
- Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at
approximately 1.56, 2.06, and 2.31 bits per weight, respectively.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 18:04:37 +00:00
Keval MorabiaandClaude Opus 5 f2f0d6958e Add MLflow tracking flags to megatron_bridge quantize.py (#2477)
### What does this PR do?

Type of change: new feature

`examples/megatron_bridge/quantize.py` gains the MLflow tracking flags
`examples/hf_ptq/hf_ptq.py` already has: `--mlflow <tracking-uri>`
(MLflow's own `$MLFLOW_TRACKING_URI` is honoured too),
`--mlflow_experiment` and `--mlflow_run_name`. Only the master rank
opens a run, so a `torchrun` launch produces one run carrying the
invocation, every command-line argument as a searchable param, the
resolved `--recipe` (with `$import`s expanded), that rank's log and the
quantizer summary. Once `bridge.save_megatron_model` returns,
`.experiment.json` is written into `--export_megatron_path`, so a
Megatron checkpoint found on disk names the run that produced it; a run
that fails is still recorded as `FAILED` with its traceback.

Rather than copy the wiring a third time, the part `hf_ptq` and
`vllm_serve` had each duplicated moves into
`modelopt.torch.utils.mlflow`:

- `add_mlflow_args(parser, tool, tracks=, variant_help=)` — the three
flags, registered under both the `--mlflow_x` and `--mlflow-x` spellings
(vLLM's `FlexibleArgumentParser` only matches the dashed one).
- `resolve_tracking_uri(uri, parser)` → `(uri, required)` — the flag
overrides the environment and is fatal when the URI is unusable; a URI
inferred from `$MLFLOW_TRACKING_URI` warns and continues untracked,
since that variable is commonly exported for unrelated tooling.
- `resolve_mlflow_args(args, parser, tool, model, variant)` — the same,
settled onto `args`, plus the default experiment name.
- `EXPERIMENT_JSON`, `MlflowRunLogger.log_experiment_json()` and
`drop_experiment_json()` — the checkpoint→run provenance pointer,
previously private to `hf_ptq`.

Both existing callers now delegate to those, keeping their own help
wording and variant naming, so the three scripts share one convention
instead of three copies (`example_utils.py` and `vllm_mlflow_utils.py`
each lose ~60 lines). Their flags and defaults are unchanged; the only
user-visible difference is that `hf_ptq`'s ignored-URI warning gains the
`$` the vLLM one already had (`Ignoring $MLFLOW_TRACKING_URI, continuing
untracked`), so one shared message serves both.

One behaviour change reaches `hf_ptq` through the shared helper, and it
is a fix: when tracking was inferred from `$MLFLOW_TRACKING_URI` and the
run never opened (unreachable server, or `mlflow` not installed), it
used to leave the previous run's `.experiment.json` beside a freshly
exported checkpoint. `log_experiment_json` now drops the pointer when it
has no run to record, so after a completed export the file is this run's
or absent.

The new example-side code lives in
`examples/megatron_bridge/mlflow_utils.py`, which deliberately imports
no Megatron, so the whole flag-to-artifact path is testable without the
Megatron container (the same split
`examples/vllm_serve/vllm_mlflow_utils.py` uses).

### Usage

```bash
torchrun --nproc_per_node 2 quantize.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --recipe general/ptq/nvfp4_default-kv_fp8 \
    --tp_size 2 \
    --export_megatron_path /tmp/Qwen3-8B-NVFP4-megatron \
    --mlflow https://<your-mlflow-server>/

# The checkpoint then names the run that produced it:
cat /tmp/Qwen3-8B-NVFP4-megatron/.experiment.json
```

The experiment defaults to `$USER/megatron_bridge_quantize/<model
basename>-<recipe name, or --quant_cfg>`.

### Testing

- `tests/examples/megatron_bridge/test_mlflow_utils.py` — 20 new tests
covering the flags (both spellings, env-vs-flag precedence, the
fatal/best-effort split), the params/tags/artifacts a run records, rank
gating, and the `.experiment.json` lifecycle. The last one guards the
seam with `quantize.py` as text, since that script needs Megatron to
import.
- `tests/unit/torch/utils/test_mlflow.py` — 13 new tests for the
extracted library API; suite at **75 passed**.
- Full `tests/examples/megatron_bridge` suite in
`nvcr.io/nvidia/nemo:26.08` on an RTX 6000 Ada: **37 passed (26m)**,
including the three `test_quantize_export` cases that drive the real
`quantize.py`, plus QAD, distill and prune.
- Regression proof for the refactor:
`tests/examples/hf_ptq/test_hf_ptq_args.py` **47 passed** and
`tests/examples/vllm_serve/test_vllm_mlflow_utils.py` **32 passed**,
unchanged apart from one renamed constant reference.
- Both new guards were shown to fire: mutating the `checkpoint_exported`
gate and removing `with mlflow_run(args):` each failed exactly one test.
- End-to-end tracked run in `nvcr.io/nvidia/nemo:26.08` (tiny Qwen3-MoE,
`general/ptq/fp8_default-kv_fp8`, 1 GPU) against an internal MLflow
server: run `47d4ccd7cd9e48269e7248868347ccd0` under experiment
`$USER/megatron_bridge_quantize/mbridge-ptq-validation` closed
`FINISHED` carrying `command.txt`, `version.txt`, `experiment.json`,
`recipe/resolved_recipe.yaml`, `logs/quantize.log` and
`summary/quant_summary.txt`; all 19 CLI arguments plus `world_size`
logged as params with no `mlflow_*` leakage, the
`model`/`checkpoint_path`/`source_checkpoint_path` tags set, and
`.experiment.json` written into the Megatron checkpoint beside
`iter_0000000/`.
- `pre-commit run --files <changed>`: all hooks pass (ruff, mypy,
bandit, markdownlint).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under *Megatron Framework (M-LM / M-Bridge)*.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

`mlflow` stays an optional dependency, imported only once tracking is
enabled, so an untracked run behaves exactly as before.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added optional MLflow tracking for Megatron-Bridge quantization runs.
- Configure tracking with `--mlflow` or `MLFLOW_TRACKING_URI`, with
customizable experiment and run names.
- Records searchable parameters, resolved recipes, quantization
summaries, logs, and checkpoint provenance.
- Captures successful and failed runs and cleans up stale checkpoint
metadata when appropriate.

- **Documentation**
- Added setup instructions and usage examples covering artifacts,
naming, checkpoint metadata, validation, and authentication.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:25:30 -07:00
Chenjie LuoandClaude Opus 5.5 25d8c91762 Add guidance on keeping skill updates concise to AGENTS.md (#2523)
### What does this PR do?

Type of change: documentation

Adds an `## Updating skills` section to `AGENTS.md` (symlinked as
`CLAUDE.md`). Skills are loaded into agent context, so each extra line
costs tokens every time the skill runs. The new guidance tells the agent
to:

- Keep skill edits concise: add only what changes agent behavior, and
tighten existing text instead of appending more.
- Do a final compression pass over the skill diff before opening a PR:
drop unnecessary explanations and examples, cut redundancy, and merge
overlapping guidance.

### Usage

N/A — no API or flag change.

### Testing

`pre-commit run --files AGENTS.md` (markdownlint and the other
applicable hooks pass).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- documentation-only
change -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- agent instructions only, not user-facing -->
- Did you get Claude approval on this PR?: ❌ <!-- will run /claude
review if reviewers want it -->

### Additional Information

Follows #2494, which added the PR sizing guidance to the same file.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added guidance for keeping skill updates focused on behavior changes
and reviewing edits for unnecessary detail.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 10:07:22 -07:00
h-guo18andClaude Opus 5 87f7d1432f fix(speculative): hold the DFlash draft's fp32 master weights in the optimizer (#2483)
### What does this PR do?

Type of change: Bug fix

**Follow-up to #2342**, which split this out on review (commit
`c67784d9`), and a rethink of how
the flag is implemented.

`dflash_fp32_master_weights` exists because the DFlash draft is cast to
the frozen bf16 target's
dtype, so AdamW allocates its moments in bf16 — and bf16 is too coarse
to hold them. At
`beta2=0.999` a single step changes `v` by at most **0.100%**, while the
smallest change bf16 can
represent near `v` is **0.164% mean / 0.388% max** (measured): every
decrease rounds away, `v` only
grows, and the effective step size decays on its own from step 1.

#2342 fixed that by **promoting the draft model to fp32**. Everything
else followed from giving the
model a dtype the rest of it does not have — a bf16 autocast at every
entry point, two transformers
loader hints so `from_pretrained(dtype="auto")` would not round the
draft away, a post-condition
check because those hints fail silently, and a doubled DDP gradient
all-reduce.

**This PR puts the fp32 in the optimizer instead**, where Megatron-LM,
DeepSpeed and apex put it.
`MasterWeightAdamW` holds an fp32 master copy of each non-fp32 parameter
plus fp32 moments in
`self.state[p]`, steps on the master, and copies back at the parameter's
dtype. The model is never
anything but the base dtype, so every one of those follow-on pieces is
deleted, gradients stay
bf16, and the exported drafter is unchanged. What the placement costs is
that wiring the optimizer
becomes the training loop's job:
`EagleTrainerWithAccLog.create_optimizer` builds it, and
`VerifyMasterWeightsCallback` raises at the end of step 1 if the moments
are not fp32.

**The default flips to `True`** — the flag now changes optimizer memory
and optimizer arithmetic
and nothing else. Flipping it on the model-promoted implementation turns
**25 of 259** unit tests
red; flipping it here is **259 passed**. Set it to `False` to reclaim
the memory, about 12 bytes per
draft parameter instead of 4.

<details>
<summary>Three drive-by fixes, independent of the above</summary>

- `_place_draft` is folded back into `modify()` — it fused the draft's
dtype, its device and an
  eager rotary buffer behind one meta guard.
- The module docstring's claim that `DFlashModule` has an `_apply`
meta-buffer fix is removed
  (`grep "def _apply"` matches nothing, and never did).
- #2342's field description no longer lists `evaluation` as a broken
path — `forward`
short-circuits to the base model when `not self.training`, so the draft
never runs there.

</details>

### Usage

No API change. `dflash_fp32_master_weights` now means the *optimizer*
holds fp32 master weights
rather than the draft model being fp32.

### Testing

**1 · The refactor is arithmetically a no-op.** Both implementations run
AdamW on an fp32 tensor,
so given the same starting values and the same gradients the
trajectories are identical — 1000
steps, `weight_decay=0.01`:

```
old fp32 parameter  vs  new fp32 master : bitwise equal = True  (max |diff| 0.0e+00)
exp_avg / exp_avg_sq                    : bitwise equal = True
optimizer state dtypes                  : ['torch.float32']
model parameter dtype                   : torch.bfloat16
```

Initial values have to be matched at bf16 first, or the bf16 arm's
one-time rounding of the draw
shows up as a 2e-4 "difference" that is not arithmetic. With that
controlled, the two
implementations differ only in their *inputs*: gradient precision (fp32
vs bf16 — torch 2.10
requires `grad.dtype == param.dtype`) and that one-time rounding.

**2 · End to end on GPU: the effect survives the refactor.** Qwen3-1.7B
base, real corpus, one GPU
per arm, three arms — pure bf16 (flag off), the #2342 implementation,
and this one — on two
algorithms trained independently, sharing seed, data order and
initialisation within an algorithm.

<img width="2925" height="960" alt="image"
src="https://github.com/user-attachments/assets/40d37059-ab8a-4924-b049-85f76b70b156"
/>


The two fp32 arms sit on top of each other for the whole run while bf16
stays above both, and the
old-vs-new gap is 10–23× smaller than the fp32-vs-bf16 effect it has to
be compared against.

**Acceptance length says the same thing, and settles what the loss could
not.** All six drafters at
the end of those curves were exported and served under vLLM against the
same base, and measured on
MT-Bench (80 prompts, 8 categories, greedy, one request at a time,
`num_speculative_tokens` =
trained `block_size` − 1, every knob but the drafter held fixed):

| | pure bf16 | fp32 in model (#2342) | fp32 in optimizer (this PR) |
new − old | fp32 − bf16 |
|---|---|---|---|---|---|
| `dflash` | 1.3068 | 1.3708 | **1.3666** | −0.0042 &nbsp;`t=−0.90` |
+0.0619 &nbsp;`t=+13.8` |
| `lilicorr` | 1.2536 | 1.2814 | **1.2882** | +0.0068 &nbsp;`t=+1.42` |
+0.0312 &nbsp;`t=+9.2` |

Paired by prompt, n=80. On both algorithms the new-vs-old 95% CI
straddles zero (`dflash`
[−0.0134, +0.0051], `lilicorr` [−0.0028, +0.0164]) while fp32-vs-bf16
does not come close to it,
and the sign of new-vs-old **flips between the two algorithms** — what a
rounding difference looks
like, not a bias. This is also the comparison the training loss could
not give: all three arms are
**exported and served in bf16**, so the old implementation's fp32 draft
weights are rounded at
export exactly as they would be for deployment, and the "its loss was
computed on a more precise
forward" caveat below does not apply. `lilicorr` needs

[vllm-project/vllm#57934](https://github.com/vllm-project/vllm/pull/57934),
applied as an overlay so
that both algorithms are measured on one engine build.

The right panel is the mechanism, and the one signal that depends on
neither the seed nor the
choice of loss statistic: Adam's updates to the draft's RMSNorm gains
are smaller than the bf16 ULP
at 1.0 (0.0078), so in the bf16 arm every one of them rounds away and
the gains never move — not
one of `dflash`'s 14 in 30000 steps, and two of `lilicorr`'s 20 by
3e-06. Both fp32 arms move
all of them, by the same amount.

Two results behind the figure rather than in it. **fp32-vs-bf16 grows
with the horizon** while
old-vs-new does not — on `dflash` −0.129 at 1500 steps → −0.262 at 15000
→ −0.341 at 30000, and on
`lilicorr` −0.191 → −0.220 → −0.285, against an old-vs-new difference
that stays near 0.02 at every
horizon and changes sign between them (−0.026 → +0.028 on `lilicorr`).
That is what a compounding
bias and a rounding difference respectively should look like, and it is
the reason the longer runs
were worth doing. And
**across seeds**, the paired old-vs-new difference at 1500 steps is
+0.0003 (n=6) on `dflash` and
+0.0643 (n=10) on `lilicorr`, both with a 95% CI straddling zero.

<details>
<summary>Limits of the above, stated rather than smoothed over</summary>

At 5 seeds the `lilicorr` paired difference read +0.1610 ± 0.0557
(t=+2.89, 4/5 seeds in the same
direction) — nominally significant, suggesting the new implementation
was genuinely worse there.
Four further `lilicorr` seeds were run against that pre-declared
question; two came back strongly
negative and the estimate settled at +0.0643 (95% CI [−0.086, +0.214]).
The earlier reading was
small-sample noise.

At 1500 steps on `lilicorr` that CI is *not* narrower than the
fp32-vs-bf16 effect it is being
compared against, so the 1500-step sweep alone cannot certify
equivalence there — `lilicorr` is
still at loss 9.3 and deep in its early transient, and it is the long
runs that resolve it. On
`dflash` the 1500-step CI (±0.031) is already 4× tighter than the effect
(−0.129).

One asymmetry the loss comparison cannot separate: the old
implementation held the draft weights in
fp32 *at forward time*, so its training loss was computed on a more
precise forward, while both
implementations export bf16. Any residual advantage it appears to have
is therefore an upper bound.

</details>

**3 · Unit tests.** `tests/unit/torch/speculative/` — **259 passed** on
**transformers 5.0.0** and
**5.3.0**, both ends of the supported `>=5.0,<5.13` (CPU, torch 2.10).
`TestDFlashFp32MasterWeights`
is rewritten for the new mechanism; the two that would have caught the
traps in this design are
`test_resume_does_not_round_the_master_back_down`
(`Optimizer.load_state_dict` casts float state to
its parameter's dtype, so a naive subclass rounds the master and both
moments to bf16 on *every*
resume, silently, with the loss still falling) and
`test_the_callback_refuses_a_loop_that_forgot_the_optimizer`. The rest
cover the draft's dtype with
the flag either way, that no forward path needs an autocast any more,
that plain AdamW really does
leave the moments in bf16, and that an fp32 model allocates no redundant
master. A sharded FSDP2
`DTensor` keeps an fp32 master and fp32 moments through a step.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ for artifacts, with one
intentional default change.
The draft's stored dtype goes back to matching the base, as it was
before #2342; existing
checkpoints load unchanged and the exported drafter is unaffected. The
flag now defaults to
**`True`** — the measurements above are the reason, and the cost is fp32
master + fp32 moments for
the draft only. A training loop that builds its own optimizer instead of
using the shipped
`create_optimizer` gets plain AdamW and none of this;
`VerifyMasterWeightsCallback` makes that
  fail loudly at step 1 rather than skip the feature quietly.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — draft; will run `/claude
review` before marking ready.

### Additional Information

**On the 7–14% acceptance-length gain quoted in #2342:** that was
measured on the fp32-model
arithmetic and is not re-derived here. What is measured above is the
like-for-like comparison this
PR has to answer — same corpus, same horizon, same serving path, one
implementation swapped.

**History:** commits 1–3 restore the autocast design as it was split
out; commits 4–6 replace it.
Happy to squash before review.

<details>
<summary>Alternatives measured and rejected, so they do not get
re-proposed</summary>

- **Swapping `p.data` to the master and calling `super().step()`**
(reuses all of AdamW, ~20 lines
instead of ~50): bit-identical on ordinary parameters over 25 steps, but
silently wrong under
FSDP2 — assigning `.data` on a `DTensor` parameter updates the wrapper's
reported dtype while the
local shard keeps the model's, so `p.dtype` reads fp32, `p.data.dtype`
reads bf16, and
`zeros_like(p)` allocates the moments in bf16 anyway. CPU tests pass
either way.
- **Narrowing the autocast from `__call__` to `forward`** (while it
still existed): turns 10
Domino/DSpark tests red — the variants apply their heads in their own
`forward` overrides,
  outside `DFlashModule.forward`.
- **Building the rotary buffer on meta and letting the loader
materialise it**: makes RoPE
correctness depend on transformers selecting a branch by class-name
substring
(`"RotaryEmbedding" in module.__class__.__name__`), and the `if not
hasattr` guard is then
permanently satisfied, so a later `to_empty()` leaves garbage forever —
measured `4.56e-41`, i.e.
  cos=1 / sin=0, no positional encoding at all.
- **Building it eagerly in `DFlashModule.__init__`**: lands before the
dtype cast, so `Module.to`
rounds the RoPE frequencies to bf16 on the default path — measured
`0.8659643530845642` →
  `0.8671875`, loss `3.47230935097` → `3.47114777565`.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- DFlash now uses FP32 optimizer master weights and Adam moments by
default while keeping draft parameters in the base model’s dtype.
- Master-weight training preserves optimizer precision when restoring
checkpoints.
  - The feature can be disabled to reduce optimizer memory usage.
  - Draft models consistently follow the base model’s dtype and device.

- **Bug Fixes**
  - DFlash workflows now support operation without autocast.
- Added validation for compatible AdamW-family optimizers and
master-weight precision, including resumed training runs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 15:39:33 +08:00
Zhiyu f2ee751089 Validate eval limit accounting and raise TB 2.1 timeout (#2499)
### What does this PR do?

Type of change: documentation

Successful evaluations can still contain timed-out trials or responses
stopped by output limits. Existing skills scan for errors but do not
require counts or rates, allowing serving-speed effects to be mistaken
for quantization accuracy changes.

Require per-task timeout and output-limit accounting before scores are
treated as validated, with explicit denominators, telemetry coverage,
effective limits, and artifact evidence. Separate request retries,
terminal trial failures, and resumable SLURM walltime events. Carry that
evidence into baseline/candidate comparisons; unknown accounting or
unresolved infrastructure effects prevent an acceptable verdict.
Preserve benchmark-defined failures and label diagnostic protocol
changes explicitly. Allow valid-with-warnings results for small
fractions of limit-hit responses/trials when coverage, protocol, and
other checks pass, without automatically retrying. Clarify that a BF16
acceptance gate cannot be satisfied with an FP8/INT4 baseline.

Correct Terminal-Bench 2.1 guidance that sharding cannot affect scores
and require accounting across the full trial set. Raise the ModelOpt TB
2.1 agent budget from 7200 to 14400 seconds in the recipe and set it
explicitly in the template. With timeout_strategy=max, this is a minimum
budget, not a hard ceiling. Set sandbox lifetime to six hours and
example SLURM walltime to eight hours to allow setup and verification;
partition limits and effective task budgets must still be checked.
Baseline and candidate must use the same policy, and old two-hour
results require remeasurement for a matched comparison. This changes
skill defaults, not the upstream benchmark protocol or harness
instrumentation. The upstream-vendored launching-evals skill is
unchanged per repository policy; its standalone workflow still needs an
upstream update.

Reduce the evaluation entrypoint from 6,156 to 888 words (86%) by moving
launcher Steps 1–8 into an on-demand reference without changing their
instructions. Preserve step headings/links for existing callers. Shorten
timeout accounting from 550 to 352 words (36%) while retaining the
checks and warning policy. This reduces initial context; full launcher
workflows still load the relevant detailed sections.

### Usage

TB 2.1 recipe/template defaults now include:

```yaml
solver:
  timeout_strategy: max
  run_timeout: 14400
sandbox:
  max_task_lifetime_sec: 21600
```

The example uses `cluster.walltime: "08:00:00"` where the partition
permits it. Request timeout remains 3600 seconds. For leaderboard
comparisons, use the benchmark protocol rather than assuming the
ModelOpt override is comparable.

### Testing

- Pre-commit passed on all six changed Markdown/YAML files.
- skill-creator quick_validate.py passed for evaluation and
compare-results.
- git diff --check passed.
- Validated recipe and template solver/sandbox settings against NEL
schemas with the pinned TB playbook; checked max and task timeout
resolution.
- Verified extracted launcher Steps 1–8 match the original text, apart
from a trailing blank line.
- Manually reviewed handling of recovered retries, resumed artifacts,
incomplete telemetry, benchmark-defined limits, and timeout-sensitive
comparisons. No evaluation jobs were launched; increased timeout
defaults have not been benchmarked.

### Before your PR is "*Ready for review*"

Contributor guidelines and security coding practices reviewed.

- Is this change backward compatible?: ✅ Existing configs remain valid.
New TB configs use longer budgets and may consume more runtime;
comparison requires matched timeout policies.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
- Did you write any new necessary tests?: N/A — skill/config changes
validated as listed above.
- Did you update Changelog?: N/A — internal skill guidance.
- Did you get Claude approval on this PR?: ❌ Not requested yet.

### Additional Information

A nonzero benchmark-defined timeout or output-cap rate does not
automatically invalidate a score. The change requires evidence and
protocol-aware interpretation, without choosing a universal acceptable
rate or silently excluding affected trials.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Expanded evaluation guidance with a centralized launcher workflow
covering setup, configuration, execution, monitoring, authentication,
and failure handling.
* Required timeout and output-limit accounting for every task, including
successful and non-reasoning runs, with rates, denominators, telemetry
coverage, effective limits, and recovery status.
* Clarified that unknown telemetry, mismatched limits, or unresolved
infrastructure effects prevent an acceptable verdict.
* Updated comparison guidance to require matched reruns, aligned
precision baselines, and limit-hit rate comparisons.
* Clarified distributed benchmark timeout handling and score extraction;
GDPVal support is no longer documented.
  * Updated example evaluation time limits and SLURM walltime.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-09-23 00:32:06 +00:00
github-actions[bot]andgithub-actions[bot] 7159c01d9d [chore]: weekly bump of uv.lock on main (2026-09-21) (#2490)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

<details>
<summary>uv lock --upgrade output</summary>

```
Using CPython 3.12.14 interpreter at: /opt/hostedtoolcache/Python/3.12.14/x64/bin/python3
Resolved 206 packages in 9.16s
Updated cachetools v7.1.8 -> v7.2.0
Updated cuda-pathfinder v1.8.1 -> v1.8.2
Updated databricks-sdk v0.139.0 -> v0.140.0
Updated deepspeed v0.19.6 -> v0.19.7
Updated filelock v3.32.6 -> v3.32.7
Updated huggingface-hub v1.31.0 -> v1.32.0
Updated hydra-core v1.3.6 -> v1.3.7
Updated idna v3.19 -> v3.20
Updated mlflow-skinny v3.16.0 -> v3.16.1
Updated multidict v6.8.0 -> v6.9.0
Updated pandas v2.3.3, v3.0.5 -> v2.3.3, v3.0.6
Updated peft v0.20.0 -> v0.21.0
Updated platformdirs v4.11.8 -> v4.11.11
Updated propcache v0.5.2 -> v0.5.4
Updated pyparsing v3.3.2 -> v3.3.3
Updated python-discovery v1.6.0 -> v1.6.1
Updated urllib3 v2.7.0 -> v2.8.0
Updated uv v0.12.13 -> v0.12.17
Updated virtualenv v21.7.9 -> v21.9.0
Updated watchfiles v1.2.0 -> v1.3.0
Updated yarl v1.24.5 -> v1.25.1
```

</details>

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 08:24:27 +00:00
1b4e7dfb14 [OMNIML-5899] Add IQ post-training quantization recipes (#2449)
## Summary

- add numerics configs for IQ1_S and IQ2_XS
- add model presets and general PTQ recipes
- document the supported weight-shape requirement
- reuse the packed IQ weight across forwards instead of re-encoding it
every time
- add end-to-end `hf_ptq` coverage for both formats
- add the release-note entry

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

Sixteen files. The PR started as eight recipe config, documentation and
changelog files; the end-to-end test added for them surfaced two
performance
bugs in the already-merged codec, and fixing those pulled in the codec
files
and their tests.

Beyond the original recipe set it now touches four files owned by #2446
—
`ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and
`ggml/backend.py` —
plus three test files. It still contains no kernel or export files.

Keeping those fixes here rather than moving them to #2446 is deliberate
and
confirmed with the stack owner: #2446 is already merged, and both bugs
are only
observable through the end-to-end test this PR adds, so splitting them
would
separate each fix from the test that demonstrates it.

## Packed-weight cache fix

Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S
`hf_ptq`
run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing
the same
154 weights 15400 times — 100 times each.

The repacking is not calibration. These recipes set `algorithm: null`
and
`hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100
passes are
the sample `generate()` calls `hf_ptq` makes before and after
quantization: one
decode step re-runs weight fake-quant on every linear, and the packed
payload
was thrown away each time.

`_PackedWeightCache` was already there to prevent exactly this, and it
never
hit. It keyed on the identity of the tensor the backend was handed, but
`TensorQuantizer` passes a fresh *view* of the weight on every forward,
so the
identity check never matched twice.

The fix anchors the entry to `inputs._base` — the parameter the view is
taken
from — held as a weakref. The parameter is stable across forwards, so
the cache
hits; the reference stays weak, so the payload is released with the
weight and
offloaded/meta-device flows are unaffected. (A strong reference does
make the
cache hit, but pins full-precision storage for the life of the
quantizer, which
is the opposite of what those flows need.)

Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and
after:

| | packer calls | packing time | wall clock |
|---|---|---|---|
| before | 15400 | 499.0 s | 8m57s |
| after | 154 (one per weight) | 6.2 s | 3m21s |

`test_ggml_weight_is_packed_once_across_forwards` pins this: it counts
encoder
calls across five forwards under `torch.inference_mode()` (what
`generate()`
runs under) and asserts exactly one.

## Decode chunk fix

With packing cached, the end-to-end cost moved entirely into the decode,
and
IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case).
Instrumenting
both showed packing was no longer the cost at all — IQ2_XS packs
*faster*:

| | pack calls | packing time | wall clock |
|---|---|---|---|
| IQ1_S | 154 | 6.2 s | 3m21s |
| IQ2_XS | 154 | 1.6 s | 17m59s |

The cause was one constant serving two loops with opposite
characteristics.
`_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which
holds the
large codebook-search temporaries and runs once per weight; IQ2_XS sets
it to
256 rather than IQ1_S's 1024 because its search sweeps sixteen local
scales per
grid tile. But the *decode* shared it — and the decode has tiny
temporaries,
runs on every forward, and is never cached, so a small chunk only
multiplies
kernel launches.

Decoding a 2048x5632 weight:

| chunk | IQ2_XS decode | transient peak |
|---|---|---|
| 256 (was) | 91.4 ms | +24 MiB |
| 1024 | 22.9 ms | +31 MiB |
| 4096 (now) | 5.8 ms | +56 MiB |

The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded
through
`fake_quantize_with_cache`. The encode bounds are untouched, so the
memory
ceiling stays where it was aimed. The seven-iteration sign-parity loop
in
`dequantize_iq2_xs` is also folded into three XOR steps, off the same
per-forward path.

Per case in `tests/examples/hf_ptq` on 2xH100:

| | before | after |
|---|---|---|
| IQ1_S | 259.65 s | 94.64 s |
| IQ2_XS | 1081.05 s | 102.77 s |

Both now fit the 300s `tests/examples` default, so the cases carry no
explicit
timeout.

## Integration contracts

- `block_sizes: {-1: 256}` records the native packed-block contract; it
does not drive the GGML fake-quant scale search. The export path reads
`TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`,
validates it against the format block size, and records `group_size:
256` in checkpoint metadata.
- The model presets intentionally expose `--qformat iq1_s` and
`--qformat iq2_xs` in `hf_ptq`.

## Testing

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed,
1 skipped
- `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load
with the expected backend and no unused search option
- `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2
passed on 2×H100 (TinyLlama, both formats end to end through export),
3m17s for the pair
- pre-commit hooks pass on all changed files
- larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers)
quantizes
and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself
192s

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:46:50 +00:00
yingguo-trt 5bb7343592 [https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect disabled quantization during calibration (#2434)
### What does this PR do?

Type of change: Bug fix

During max calibration, `enable_stats_collection()` calls
`disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ
attention paths bypass `TensorQuantizer.forward()` and previously
selected the Triton/Kitchen path from the configured enabled state
alone, so P quant-dequant could still execute while quantization was
inactive.

This change:

- enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is
true;
- bypasses all fused P-QDQ paths when the P quantizer is disabled or
quantization is inactive;
- adds focused parameterized regression tests covering Triton and
Kitchen dispatch across enabled, `disable()`, and `disable_quant()`
states.

### Usage

N/A. This restores the existing `disable_quant()` contract and does not
introduce a new API.

### Testing

- `pytest
tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14
passed
- Targeted pre-commit checks passed
- Kitchen coverage verifies the full `{disable(), disable_quant()} ×
{lazy, already initialized}` dispatch matrix remains bypassed while
inactive and resumes after re-enabling
- Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides
on B300: [module regression build
#121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/)
- Controlled B300 isolation passed when only the P-BMM quantizers were
hard-disabled during calibration, and also passed when the existing
single Triton attention configuration was forced
- Fixed-code B300 validation passed in three independent runs: [Jenkins
build
123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/),
[Jenkins build
124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/),
and [Jenkins build
125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/).
Jenkins build 123 compared all 178 module outputs byte-identically and
found 0/377 `amax` and 0/377 `scale` changes.

The confirmed impact is incorrect calibration behavior plus unstable
quantizer state and module outputs; downstream benchmark accuracy impact
has not been established.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

The failure was isolated to fused P-QDQ runtime-state dispatch.
Quantizer topology and configuration were identical in the failing
same-wheel comparison.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Corrected attention dispatch so disabled or inactive quantization uses
the original attention implementation.
- Preserved optimized quantized attention when quantization is enabled.
- Ensured attention masks remain unchanged when using the original
attention implementation.
- Improved fallback behavior for fused attention paths, including
correct initialization and reuse when quantization is re-enabled.

- **Tests**
- Added coverage for enabled and disabled quantization states,
attention-mask handling, fallback selection, and result consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com>
2026-09-22 06:51:23 +00:00
Chenjie LuoandClaude Opus 5 7a35cada39 Add guidance on sizing and splitting PRs to AGENTS.md (#2494)
### What does this PR do?

Type of change: documentation

Adds a `## Sizing and splitting PRs` section to `AGENTS.md` (symlinked
as `CLAUDE.md`) so AI-assisted work stops producing one giant PR that
nobody wants to review.

The new guidance tells the agent to:

- Keep each PR that goes up for review under ~500 changed lines of
source, and check the size before opening.
- Propose the split *before* opening an oversized PR rather than after.
- Split on file/directory/module boundaries first and fall back to
feature boundaries (enabling refactor first, then one PR per behavior it
unlocks).
- Keep the series acyclic and linearly ordered — no circular
dependencies between sub-PRs — and state the merge order.
- Prefix sub-PR titles with `[x/N]` so reviewers know the PR is one
slice of a planned split, and link the siblings.
- Make every sub-PR stand on its own: it builds, carries unit tests for
the code it introduces, and passes CI without the later PRs.
- Optionally submit the whole change as a reference-only draft PR for
the big picture, cross-linked with the sub-PRs.

### Usage

N/A — no API or flag change.

### Testing

`pre-commit run --files AGENTS.md` (markdownlint and the other
applicable hooks pass).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- documentation-only
change -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- agent instructions only, not user-facing -->
- Did you get Claude approval on this PR?: ❌ <!-- will run /claude
review if reviewers want it -->

### Additional Information

This PR is itself well under the new budget (26 added lines in one
file), so no split applies.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated pull request sizing guidance to allow oversized changes when
they cannot be meaningfully split.
- Clarified that draft aggregate pull requests are optional and intended
for reference only when splitting would reduce clarity.
- Added examples covering self-contained changes and new models or
backends without a functional intermediate state.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 22:52:53 -07:00
noeyy-mino ee5c256204 Noeyy/fix bug 6701777 (#2402)
### What does this PR do?

Type of change: Bug fix: 6701777

Regression source: "[OMNIML-3349] Add FP8 MHA
quantization support for HuggingFace ViT" (#1289), merged into
0.44.0rc3 via the batch cherry-pick #1350. This PR:
  1. Registers nn.LayerNorm as a QuantModule for the first time
     (modelopt/torch/quantization/nn/modules/quant_layernorm.py),
     intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output
     quantizer rules apply to ViT.
  2. Removes the prior forced Cast-alignment logic in export_onnx.py
that used to normalize Q/DQ node dtypes to trt_high_precision_dtype.

Root Cause:
Once nn.LayerNorm became a registered QuantModule, these wildcards
started unintentionally matching norm1.norm inside FLUX's
AdaLayerNormZero block — an elementwise_affine=False LayerNorm with
no learnable weight/bias. Its input got routed through NVFP4 Q/DQ
(emitted as Float32) while its synthesized affine scale remained
native BFloat16, producing the dtype mismatch. 

Chosen fix:
Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets
rather than touching the global QuantModuleRegistry (which ViT FP8
MHA still needs). Add, in both
modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and
nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules
(list order matters — later entries override earlier ones):

  - parent_class: 'nn.LayerNorm'
    quantizer_name: '*'
    enable: false

### Usage

```
python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16

trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1
```

### Testing
 The above test commands.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A 

### Additional Information
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by
excluding LayerNorm modules from quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-09-22 00:55:33 +00:00
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00
Zhehao-Hu 051d6adb20 Require PTQ recipe guidance before selecting quantization candidates (#2478)
### What does this PR do?

Type of change: documentation

Updates `quant-recipe-search` to require reading
`modelopt_recipes/ptq.md` before selecting initial or subsequent
quantization candidates.

The skill uses the guide to inform quantization scope, KV-cache scheme,
and calibration method. It requires inspecting candidate YAMLs and
supporting configs, citing relevant guide sections, and explaining
deviations. Existing coverage, runtime compatibility, and evaluation
checks remain required.

### Usage

Invoke `quant-recipe-search` as usual. The skill reads the guide from
the ModelOpt source checkout.

### Testing

- Passed a local Codex skill smoke test recommending an initial NVFP4
candidate for Qwen/Qwen3-8B targeting high-concurrency inference on
Blackwell.
- The test prompt named the skill without explicitly requesting that the
agent read ptq.md.
- Verified successful reads of the updated skill, ptq.md, the selected
recipe YAML, and supporting configs before the final recommendation.
- The agent selected general/ptq/nvfp4_mlp_only-kv_fp8_cast and cited
relevant guide sections. It explicitly left coverage, accuracy, and
throughput pending validation.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅

### Additional Information
This change places the reading requirement in the shared recipe-search
skill so consumers receive it directly. Consumers that pin the ModelOpt
plugin to a commit must update their pin to adopt it.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the quantization recipe search workflow to require reviewing
the PTQ reference documentation before selecting a recipe.
- Clarified that available quantization schemes and model-specific
exceptions should be considered during recipe evaluation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhehao Hu <zhehaoh@nvidia.com>
2026-09-21 14:58:38 -07:00
Shengliang Xu f1abc75626 Carry unplaced checkpoint weights using the loader's accounting, replacing MTP name-matching (#2427)
### What does this PR do?

Type of change: bug fix + new tests

Replaces `hf_ptq`'s name-based MTP detection with the Transformers
loader's own accounting of what it could not place.

#### The problem

`load_mtp_weights` found MTP weights by name — `"mtp" in key` plus
config-derived layer indices — backed by a support matrix of three
storage conventions (GLM-5.1 inlined, GLM-4.7 standalone file,
Qwen3-Next tail shard). Every architecture that spells it differently is
a silent miss, and a miss means a checkpoint exported without a
component its own config still advertises. The failure is quiet on both
sides: Transformers drops unexpected keys, and vLLM's weight loading is
pull-based, so a missing MTP produces no warning at all.

#### The fix

`from_pretrained(..., output_loading_info=True)` reports
`unexpected_keys` — *"keys that are found in the checkpoints, but not
expected in the model's architecture"* — which is exactly the carry-over
set, derived structurally rather than by naming, and already accounting
for on-the-fly key conversion that a set re-derived afterwards would
have to replay. Record those keys at load; carry them at export. Weights
the loader *did* place go through the normal export path unchanged.

#### Two mechanisms, disjoint by construction

Measured against a real `from_pretrained`, an off-index file reports
**nothing**: the loader opens only shards named in
`model.safetensors.index.json`, so it never saw those tensors to call
them unexpected. Those are sidecars, not weights — untouched by
quantization and absent from the export — so they are **copied
verbatim**, which costs no host memory, preserves the bytes and file
layout, and leaves the filename a consumer looks for where it was.

| on-disk layout | mechanism |
|---|---|
| inlined layer past `num_hidden_layers` (GLM-5.1, DeepSeek-V3) |
`extra_state_dict` |
| indexed `mtp.*` tail shard (Qwen3-Next) | `extra_state_dict` |
| standalone off-index file (GLM-4.7) | **copied whole** |

A tensor is never both copied and carried; that would export it twice,
and a test asserts it.

#### Removed as redundant

`load_mtp_weights`, `mtp_layer_prefixes_from_checkpoint`,
`get_inlined_mtp_prefixes`, `_load_tensors_matching`,
`_apply_to_model_state_dict`, `_keys_to_prefixes`, `_add_mtp_exclusions`
and its three call sites, the pre-quantization `enable: False` entries
`hf_ptq` appended to the recipe's `quant_cfg`, and the dead
`_mtp_layer_prefixes` fallback in `_get_num_nextn_predict_layers`.

#### Two deliberate behavioural changes

**MTP now follows the recipe** instead of being force-excluded by the
script — which is what `examples/megatron_bridge` already does; it has
no MTP-specific code at all. Recipes importing
`configs/ptq/units/default_disabled_quantizers` still disable `mtp.*`,
so their behaviour is unchanged; a recipe omitting that unit will now
quantize an MTP the model actually built.

**`quantization_config.ignore` can no longer claim a layer is
unquantized that the export in fact quantized.** That contradiction came
from `_add_mtp_exclusions` firing off a model attribute with no
cross-check against quantizer state.

### Usage

No API change for callers of `export_hf_checkpoint`. Within
`examples/hf_ptq`, model loading now goes through a wrapper that records
the loader's accounting:

```python
model, loading_info = auto_class.from_pretrained(ckpt_path, output_loading_info=True, **kwargs)
record_unplaced_source_keys(model, ckpt_path, loading_info.get("unexpected_keys"))
```

### Testing

`tests/examples/hf_ptq/test_carry_over_layouts.py` — 8 tests driving a
**real** `from_pretrained` against a tiny model, covering each of the
three conventions above plus an auxiliary (non-MTP) tower, two layouts
at once, a checkpoint with nothing stray, and that indexed shards are
never copied. CPU-only: the mechanism is bookkeeping during load, so a
GPU adds nothing; the export side already has GPU coverage in
`tests/gpu/torch/export/test_export_carry_over.py`.

The six `load_mtp_weights` tests are replaced by three on the recording
path, and the `get_model` test doubles now model `output_loading_info`
the way Transformers does.

All passing: 8 layout tests, 82 in the surrounding `examples/hf_ptq`
suite. `ruff` findings at parity with `main` on every changed file.

Files named like a main weight shard are excluded from the off-index set
whatever the index says — a fixture with an empty `weight_map` would
otherwise have made the source weights look like sidecars and copied
them into an export beside the quantized ones.

### The algorithm: which weights get carried, and how

Two disjoint sets of source weights reach the export without passing
through quantization. They are distinguished by **what the loader did
with the file**, and that difference decides both how each is found and
how each is moved.

**Set 1 — unplaced weights.** The loader opened the file and read the
tensor, but the model had no parameter for it, so Transformers reports
it in `unexpected_keys`. An MTP head the recipe did not quantize is the
common case. Moved as **tensors**: located in whichever shard holds
them, read, and merged into the exporter's `extra_state_dict`.

**Set 2 — off-index sidecars.** The index never names the file, so the
loader never opened it and never had the chance to call anything
unexpected. GLM-4.7 keeps its MTP head in a standalone `mtp.safetensors`
exactly this way. Moved as **files**: copied byte for byte, so no host
memory is spent re-serialising tensors the export does not otherwise
touch.

#### The index is not an inventory of the checkpoint

This is the part that is easy to get wrong, and it cost a silent
data-loss bug during review.

`model.safetensors.index.json` selects which **files** the loader opens
— not which **tensors** it sees. Within a file it opens, Transformers
enumerates every tensor present and reports the unexpected ones.
Verified by experiment against transformers 5.3.0:

| case | reported in `unexpected_keys`? |
|---|---|
| key absent from the index, in a shard the index names for *other*
tensors | **yes** |
| key in a file the index never names (`mtp.safetensors`) | **no** — the
file is never opened |

So a tensor missing from `weight_map` but sitting inside a main shard is
**set 1, not set 2**. An MTP head stored that way is reported, recorded
— and was then silently dropped, because resolution went through
`weight_map`, which by construction has no entry for it. The
`--vllm_fakequant_export` guard shared that lookup, so the check written
to refuse exports that drop weights stayed silent in exactly the case it
existed for.

`locate_source_keys` now resolves through the index first (free for
everything it lists) and header-scans the shards only for the leftovers
— names, never tensor data — warning when a key is in no file at all.
The carry and the guard share it, so they cannot disagree again.

#### Flow

1. **At load.** `record_unplaced_source_keys` stores Transformers' own
`unexpected_keys` on the model (`_modelopt_unplaced_source_keys`) plus
the resolved local checkpoint path. The question asked is "does the
model have a parameter for this key", never "is this an MTP head" — so
the mechanism is architecture-agnostic.
2. **At export, before dispatch.** `read_unplaced_weights` resolves each
recorded key to its shard and reads the tensors, merging them into
`extra_state_dict`. An explicitly passed `extra_state_dict` wins on a
name clash: a caller naming a tensor is more specific than our
inference.
3. **Rank behaviour.** Only the rank that writes `extra_state_dict`
reads the bytes — the FSDP2 writer emits it from rank 0 alone, so a full
read on every rank would be host memory spent and discarded (a
DeepSeek-V3-class MTP head is 10 GB+ in bf16). The **key list** is still
resolved on every rank, because `get_quant_config` runs per rank and the
configs must agree.
4. **Recording what was written.** `export_hf_checkpoint` records
`_modelopt_carried_over_names` — the union of carried tensors and the
off-index sidecars' tensor names — before `get_quant_config` runs,
because that is the first point that knows what was *written* rather
than what was merely unplaced.
5. **Exclusions.** Both sets must reach `quantization_config.ignore`, or
a deployment framework reads the top-level `quant_algo` and tries to
load an original-precision weight as a quantized one (the NVBug 5718750
class). `seed_carried_over_exclusions` is the single path for this,
called by `get_quant_config` after its per-layer pass and again by the
layerwise exporter from `finalize()` — which snapshots its config during
`bind()`, while calibration is still running, so it cannot see the
carried set any earlier.

#### What is deliberately excluded

- **Files that re-ship indexed weights.** Mistral's
`consolidated.safetensors` is a second full copy of the model, and
PEFT's `adapter_model.safetensors` is an adapter. Both are off-index,
and copying either would put unquantized weights beside the quantized
ones — vLLM's mistral load-format looks for `consolidated.safetensors`
by name, so it is not inert. Caught by name for the known conventions
and by tensor-name overlap for the rest.
- **Files named like a main weight shard**, whatever the index says, so
a broken or partial index cannot make the real weights look like
sidecars.
- **Symlinks are *not* excluded.** A Hugging Face snapshot stores every
file as a symlink into a sibling `blobs/`, so refusing links would drop
the sidecar of every hub-downloaded checkpoint.
`resolve_checkpoint_file` checks where the link *lands* — regular file,
inside the checkpoint dir or its blob root — rather than whether it is a
link.
- **Buffers Transformers recomputes.** `*.inv_freq` is skipped by the
fake-quant guard even when a shard provides it: older
Llama/Mistral-lineage conversions do list it in the index, and refusing
an export over it would reject checkpoints that export correctly today.

### How a carried, never-quantized MTP head reaches
`quantization_config.ignore`

Raised in review: `_add_mtp_exclusions` is gone, and an unplaced weight
has no module, so
`get_quant_config` walks right past it. That was a real gap, not just a
documentation one —
fixed here.

The export writes a carried MTP head in its original precision. If it is
absent from
`exclude_modules`, a deployment framework reads the top-level
`quant_algo` and tries to load
`eh_proj` as an FP8/NVFP4 weight — the same class of failure as NVBug
5718750, where a
`transformers>=5.0` MoE router was written in BF16 but never excluded.

The two cases now differ only in *why* the module is invisible to the
quantizer walk:

| | why invisible | handled by |
|---|---|---|
| MoE router (tf≥5.0) | module exists, never gets a quantizer |
`_get_unquantized_moe_router_names` |
| carried weight | no module at all in the live model |
`_get_carried_over_module_names` |

`_get_carried_over_module_names` reads the keys the loader recorded as
unplaced
(`_modelopt_unplaced_source_keys`), strips the trailing parameter name —
a state-dict key is
`<module path>.<parameter>` — and dedupes. Those names are seeded into
`layer_config_dict` as
`QUANTIZATION_NONE`, exactly as the router pass does, so they flow
through
`process_layer_quant_config` into `exclude_modules`, which
`convert_hf_config` emits as
`quantization_config.ignore`.

An MTP head the recipe *does* quantize is unaffected: it has a module,
is loaded normally, is
never in the unplaced set, and is reported as quantized. The
backward-breaking note above still
holds — MTP follows the recipe instead of being force-excluded — but a
head that ends up carried
rather than quantized is no longer silently missing from `ignore`.

Covered by `test_carried_over_weights_are_excluded_from_quantization`
and
`test_carried_over_module_names_strip_parameter_and_dedup`.

`--vllm_fakequant_export` does not carry unplaced weights; it now raises
rather than writing a
checkpoint quietly missing them.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — MTP layers now follow the
recipe rather than being force-excluded, and
`quantization_config.ignore` no longer lists layers the export may have
quantized (carried, never-quantized weights are still listed -- see
below). Shipped recipes are unaffected; see the Changelog entry.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under 0.48.0 Backward Breaking Changes.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

Draft: the behavioural change to MTP quantization is the part most worth
a second opinion — it aligns `hf_ptq` with `megatron_bridge`, which
special-cases nothing.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Quantization supports enabled operators outside transformer layers,
including language-model heads.
* Exports preserve checkpoint weights not loaded into the quantized
model.
* Additional safetensors sidecar files are copied unchanged into
exported checkpoints.
* Unquantized auxiliary components, such as vision layers, remain
available in exported models.

* **Behavior Changes**
  * Local recipe files take precedence over built-in recipes.
* Legacy architecture-specific recipe paths remain supported with
warnings.
  * MTP-specific export exclusions are no longer applied.

* **Bug Fixes**
  * Preserved weights are no longer incorrectly reported as unquantized.
* Incomplete exports are rejected when source weights cannot be placed
or preserved.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-21 14:00:34 -07:00
yueshen2016andClaude Opus 5 fc4c40fcbe Fix HF export crash when a dynamic-block quantizer has zero amax (#2438)
### What does this PR do?

Type of change: Bug fix

`TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized
for dynamic-block
quantizers, while the static path immediately below it has always
substituted `maxbound` for
zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a
recipe that applies it to
an *activation* quantizer — e.g.
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets
`*mlp*input_quantizer` — feeds a raw `0.0` into
`NVFP4QTensor.get_activation_scaling_factor`,
whose assert aborts the entire export:

```
AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj'
(type=QuantLinear):  activation scaling factor 0.0 not positive.
```

Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE
expert — saw only zeros, so
one dead layer costs the whole run at the final export step.

This factors the substitution into `_sanitize_export_amax()` and calls
it from both branches. Two
details beyond de-duplication:

- **Branch-free, so it survives a meta `amax`.** `torch.where` +
`nan_to_num` both have meta
kernels; `bool()` on a meta tensor raises. The layerwise and streaming
export flows carry meta
`amax` — `validate_attr` short-circuits on `is_meta` for exactly that
reason — so only the
  warning is gated on a materialized tensor.
- **No longer mutates calibrated state.** The old in-place `amax[amax ==
0] = ...` wrote through a
view of `self._amax`; `torch.where` returns a fresh tensor, so that
hazard disappears.
- **Warns, with a count.** The fix turns a loud failure into a silent
one, and a zero amax means
calibration never activated that layer — worth surfacing rather than
papering over. The message
reports how many entries were substituted, since per-location dedup
otherwise collapses many
dead experts into one uninformative message. A healthy model emits none.

Scope: only the activation path is data-dependent and reachable this
way. Weight-side `_amax` uses
are left alone, since a weight amax of 0 would require an all-zero
weight matrix.

**Knowingly left as follow-up:**
`export/quant_utils.py::get_scaling_factor` discards the sanitized
`amax` when `num_bits == (2, 1)` and recomputes via
`get_weights_scaling_factor_2_from_quantizer`,
which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input*
quantizer on a module whose
*weight* quantizer is a different format (or disabled) can still trip
`assert torch.all(scaling_factor > 0)`. Format dispatch is
weight-driven, so the reported recipe
does not reach that branch; fixing it properly changes a signature
shared with the weight-side
callers and is out of scope here.

Not a regression. The dynamic early return, the `type: dynamic` numerics
unit, and the recipe that
combines them all ship in released 0.46.0 / 0.46.1.

### Usage

No new or changed API. Exports that previously aborted now complete and
warn:

```python
# Recipe applies dynamic NVFP4 to *mlp*input_quantizer; layer 37 never activated during calibration.
mtq.quantize(model, quant_cfg, forward_loop)
export_hf_checkpoint(model, export_dir=out)   # before: AssertionError; now: exports + UserWarning
```

### Testing

- New `test_amax_export_unusable_amax`, parametrized over zero and NaN,
covering the
dynamic-NVFP4 and static per-tensor configs; asserts the exported scale
is positive and that
export leaves the calibrated `amax` untouched. Plus
`test_amax_export_meta_amax`, pinning that
a meta `amax` survives export rather than raising. Both run on CPU and
CUDA via the shared tester.
- `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40
passed.
`tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed
(GB300).
- End-to-end repro on GB300, small Llama with one MLP fed all-zero
activations under
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()`
`0.0` → `6.0`, live layer
unchanged at `3.921875`, and `export_hf_checkpoint` goes from the
`AssertionError` above to
  writing `model.safetensors`.
- Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on
a healthy model
(Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal
path is unaffected.
- `pre-commit run` clean on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; its
one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs
addressed or answered in 251f2e3d

### Additional Information

Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter
also notes it passed on
0.47.0rc0; that is not explained by code — `git diff
0.47.0rc0..0.47.0rc1` touches
`export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new
`clamp_fp8_scales` argument
whose default preserves the old behaviour) and the INT4-AWQ packing
path, neither of which is on
the dense-HF NVFP4 activation-scale path. Whether `amax` lands on
exactly 0 is
calibration/model-state dependent, which is what makes it look
version-flaky.

Worth flagging separately: in the reported log the **pre-PTQ** sample
output is already
gibberish, so that BF16 checkpoint looks broken independently of
quantization. This change stops
the crash, but such a run will now export a valid-but-garbage checkpoint
— the new warning is the
signal to investigate.

Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing
release.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed Hugging Face checkpoint export when dynamic-block quantizers
have zero or invalid calibration scales.
* Exports now use a positive fallback scale and issue a warning instead
of failing when applicable.
* Export operations no longer modify the original calibrated quantizer
state.
* Meta-device exports remain non-erroring and preserve device placement.
* **Tests**
* Added coverage for zero- and invalid-scale exports across dynamic and
static quantization modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Yue <yueshen@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:47:34 +00:00
Keval MorabiaandClaude Opus 5 b311c054de Fix hybrid stack spec serialization in Megatron-Bridge checkpoints (#2452)
### What does this PR do?

Type of change: Bug fix

Hybrid (e.g. Nemotron-H) checkpoints saved by the
`examples/megatron_bridge` scripts could not be reloaded or exported to
HuggingFace:

TypeError: MLPSubmodules.__init__() missing 2 required positional
arguments: 'linear_fc1' and 'linear_fc2'

**Root cause.** Megatron-LM's YAML writer represents a
`functools.partial` via `_partial_representer`, which passes each
keyword value through `represent_data`. A dataclass instance has no
representer, so it falls through to `_safe_object_representer`, which
emits only `{_target_, _call_}` and drops every field. The default
hybrid stack spec builds its dense-MLP and MoE layers as exactly such
partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec`
on the provider — which is serialized into every checkpoint's
`run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written
with no fields at all. Both export paths hit it: `convert.sh` via
`from_auto_config`, and `export_distilled_megatron_to_hf.py` via
`export_ckpt → load_megatron_model`. Every hybrid provider is affected,
dense or MoE.

**Fix.** `set_moe_expert_layout()` stores a named, zero-argument
*factory* instead. The provider already calls a callable spec at build
time (`_resolve_hybrid_stack_spec`), so model construction is unchanged
— only the serialized form differs:

```yaml
  hybrid_stack_spec:
    _call_: false
    _target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec
```

The grouped-GEMM factory is Megatron-Bridge's own
`transformer_engine_hybrid_stack_spec`, so stock tooling
(`scripts/conversion/convert.sh`) resolves it without importing
ModelOpt.

**Known limitation — the two layouts are not symmetric.** The
SequentialMLP layout has no bridge-side equivalent (the upstream TE
hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt
target, which `instantiate` only accepts in a process that has imported
`mbridge.py` and thereby run `register_allowed_target_prefix`. A
SequentialMLP hybrid checkpoint therefore converts through the ModelOpt
entrypoints but not through stock `convert.sh`, where it fails on the
disallowed prefix instead of on `MLPSubmodules` — no regression, but
that path stays broken for this one layout. The reach is narrow:
`use_moe_grouped_gemm()` returns True for any architecture with a
grouped-expert export rule, NemotronH included, so SequentialMLP
requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs
an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge.

Both spec builders also move out of `nas/plugins/megatron.py`, which
never used them, into a new `utils/plugins/megatron_layer_specs.py`
beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that
module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is
reached by 16 test files through
`tests/_test_utils/torch/megatron/models.py`, which is bridge-free.

The underlying defect is upstream in
`megatron/training/config/yaml_utils.py`; this only stops ModelOpt from
stepping on it, so it is worth a separate Megatron-LM issue.

### Usage

No API change — hybrid checkpoints saved after this fix convert with the
existing commands:

```bash
torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \
    --student_hf_path <student_hf_model_or_path> \
    --megatron_path   <distill_out>/checkpoints \
    --hf_export_path  <hf_out> \
    --export_iterations all
```

### Testing

Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B
Nemotron-3.5-Lightning pruned+distilled run:

- **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a
real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml`
(the writer used for `run_config.yaml`), reloaded via `instantiate`,
resolved. `moe_grouped_gemm=True` →
`TELayerNormColumnParallelLinear`/`TERowParallelLinear` +
`TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays
callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a
saved config cannot regress.
- Applying the equivalent `run_config.yaml` fix to 32 iteration
checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104
experts) with populated `MLPSubmodules` / `MoESubmodules`.
- **End-to-end exports**, 6 iterations, all rc 0, each producing exactly
the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards /
41.5 GiB, all weights finite, drift from the base rising monotonically
with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU,
`convert.sh` GPU (4×GB300, TP=4), and
`export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU
123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still
builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side
conversion.
- `ruff check` / `ruff format --check` passed on the source files before
the module move.

**Not yet run:**
`tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) —
the GPU allocation expired. It asserts the round-trip property verified
manually above, but its `HybridModelProvider(num_layers=2,
hidden_size=64, num_attention_heads=4)` construction is unverified. The
module move is verified only by reference grep and syntax check, so
please also run one test that uses
`tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run
either (unavailable in the environment used).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ behavior; note
`get_te_hybrid_stack_spec` moved module
(`modelopt.torch.nas.plugins.megatron` →
`modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint
from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the
changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (added, not yet executed —
see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
before marking ready.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Hybrid checkpoints now record complete layer specifications in
`run_config.yaml`.
  * Recorded specifications support conversion to Hugging Face format.
* Hybrid MoE configurations support grouped-GEMM and sequential-MLP
modes.
* Configuration-based reconstruction preserves the selected MoE layout.

* **Compatibility**
* Checkpoints from earlier releases may require manually setting the
hybrid layer specification before conversion.

* **Tests**
* Added coverage confirming hybrid specifications survive configuration
serialization and can be recreated successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 13:45:21 +05:30
ZhiyuandClaude Sonnet 5 0a8e70804a Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failure (#2480)
### What does this PR do?

Type of change: Bug fix

`post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional
post-quantization sanity-check `full_model.generate()` unguarded,
directly before `export_quantized()`. Any exception raised there aborted
the whole run and discarded a completed calibration without exporting a
checkpoint.

Root cause (traced from [NVBug
6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 /
DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with
`device_map="auto"`, relying on `accelerate`'s
`infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU
placement. On DGX Spark's unified-memory single-GPU host, that memory
probe under-reports GPU capacity, so part of even an 8B model can land
on CPU — and the existing fallback shrinks the GPU budget further (`*
gpu_mem_percentage`), compounding it. Calibration survives this because
it never invokes the real fake-quant kernel, but the post-PTQ sanity
`generate()` does, and NVFP4's dynamic-block-quantize op
(`modelopt/torch/quantization/tensor_quant.py`) hard-asserts
`amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes
there — after ~5.8 hours of calibration, before export.

This PR does not attempt to fix the underlying
`device_map`/memory-probing behavior (unverified without the actual
hardware/logs, which weren't reachable from this environment). Instead
it makes the failure mode safe: a failure in the optional sanity check
now only skips that check and warns, and export always proceeds,
regardless of why `generate()` failed.

### Usage

No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now
completes export even if the post-quantization sanity `generate()` call
raises.

### Testing

- Added
`tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`,
which drives `post_quantize()` with a `full_model.generate()` that
raises and asserts `export_quantized()` still runs.
- Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed).
- Ran `pre-commit` on the changed files (`ruff-format` reformatted line
wrapping only).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude
review` -->

### Additional Information

Fixes NVBug 6752977. Linked JIRA: OMNIML-5932.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantized checkpoint export now continues when the optional
post-quantization generation check fails.
  - A warning is shown when the generation check cannot complete.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-19 01:21:07 +00:00
ed5c5ed369 [OMNIML-5899] Export IQ checkpoints from HF and Megatron (#2447)
## Summary

- add IQ format metadata and packed-weight export
- support Hugging Face and TP=1 Megatron export paths
- reject fused-MoE IQ export until a deployment loader owns its packed
layout
- document the shaped `uint8` weight contract and the fused-expert
boundary
- add Hugging Face, Megatron, metadata, and fused-expert export tests

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

This PR owns only export code, deployment documentation, and export
tests. It targets `main` and should merge after #2448 and #2446. It does
not contain kernel, codec/backend, or recipe files.

## Deployment consumer boundary

Dense weights and individually named expert weights use the documented
shaped `uint8` contract. Megatron fused-MoE IQ export is intentionally
rejected with `NotImplementedError`: its payload would have shape
`[num_experts, out_features, in_features // 256, payload_bytes]`, and no
deployment loader in this stack currently owns that layout. Support
should be enabled only with a loader integration test.

## Test coverage

- [Hugging Face packed-weight
export](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_export_weight.py)
- [quantization
metadata](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/unit/torch/export/test_get_quantization.py)
- [Megatron unified export and fused-MoE
rejection](https://github.com/NVIDIA/Model-Optimizer/blob/11cd58d907465933f5a552bc1a8065f84c9ba3b1/tests/gpu_megatron/torch/export/test_unified_export_megatron.py)

## Validation

- all pre-commit hooks pass for the changed files
- 89 focused Hugging Face export, metadata, and fused-expert tests pass
locally
- direct checks cover both fused-MoE export entry points for IQ1_S and
IQ2_XS
- Megatron GPU execution remains delegated to GPU CI
- restricted-term scan passes


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for IQ1_S and IQ2_XS GGML quantization formats in
unified Hugging Face and Megatron exports.
* Added quantization metadata, tensor-shape recovery, packing details,
and IQ2_XS size documentation.
* Added validation for required block sizes and tensor parallelism
settings.

* **Limitations**
  * Fused-MoE and GPT-OSS IQ expert packing are not supported.
* IQ exports require standard `weight` attributes in Hugging Face
models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-19 00:00:33 +00:00
Chenjie LuoandClaude Opus 5 a7166965e3 Drop GDPVal support from the evaluation skill (#2470)
### What does this PR do?

Type of change: deprecation (agent skill)

Removes GDPVal support from the `evaluation` agent skill. Its task
recipe, example
config and Apptainer SIF build helper are deleted.

GDPVal was not cleanly separable, so this is not just a delete:

- **MRCR depended on GDPVal's infrastructure.** `scripts/nel-gdpval.sh`
was a
generic pinned-0.2.6 `nel` launcher that only happened to be
GDPVal-named, and
MRCR ran through it; `references/gym-gdpval.md` documented the gym
bootstrap
machinery (prepare/reap, `install_on_the_fly` pin↔container coupling,
the trust
env vars) that both examples share. So the shared parts are kept and
renamed
rather than dropped: `scripts/nel-gym.sh` and `references/gym.md`. The
launcher
  pin itself is unchanged (0.2.6), as are its env-override semantics.
- **GDPVal is an AA-suite member**, so the "AA rule" in `SKILL.md` and
`references/quantization-benchmarks.md` had to change. An "AA" /
"Artificial
Analysis" request now generates the `aa/` tasks as one multi-task config
and
nothing else. Both files now state that the resulting set omits GDPVal
and is
therefore not directly comparable to a published AA Index — report
per-task
scores rather than an aggregate. `SKILL.md` also tells the agent to say
GDPVal is
unsupported rather than reconstruct a config from an older copy of the
skill.

Everything GDPVal-specific is gone from `references/gym.md`: the SIF
sandbox and
its silent-unsandboxed-exec failure mode, the 3-member judge panel,
rubric vs.
comparison scoring, the deliverables/MLflow `*cache*` trap as a GDPVal
concern, and
Stirrup-agent deploy sizing. `recipes/env.example` loses
`GDPVAL_SIF_DIR`,
`GDPVAL_MAX_TURNS` and `TAVILY_API_KEY`, and gains a documented
`NEMO_EVALUATOR_TRUST_UNLISTED_TASKS` (required by every gym task,
previously only
mentioned in prose).

One correction carried along: the old reference said the bootstrap pins
`ray==2.49.2`, but `example_mrcr.yaml` actually pins to whatever ray
version the
image already carries. `references/gym.md` now describes what the
template does.

### Usage

Not an API change. The skill-facing surface that moved:

```text
scripts/nel-gdpval.sh      -> scripts/nel-gym.sh      (NEL_GDPVAL_* -> NEL_GYM_*)
references/gym-gdpval.md   -> references/gym.md
tests/test_nel_gdpval.py   -> tests/test_nel_gym.py
```

### Testing

- `pytest plugins/modelopt/skills/evaluation/tests/` — 1 passed. The
launcher test
was renamed rather than deleted: it is the only coverage for the pinned
launcher,
  which MRCR still depends on.
- `pre-commit run --files <changed>` — clean. `sync-claude-skills` fails
in my
working copy because `.claude/skills/` is a read-only harness mount
there; it is
unrelated to this change (it trips on `speculative-decoding`) and was
skipped for
  the commit.
- Grepped the repo for `gdpval` (case-insensitive): the only remaining
hits are the
  deliberate "no longer supported" notes in `SKILL.md` and
`references/quantization-benchmarks.md`. No dangling pointers to the
deleted
  files, and no other skill referenced GDPVal.
- Checked `tests/evals.json` for both `evaluation` and `day0-release`:
no eval case
expected a GDPVal companion config, so no expectations needed updating.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — anyone with a saved GDPVal
config keeps
  it, but the skill no longer generates one and the SIF helper is gone.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — existing launcher test
renamed and kept passing.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — under **Deprecations**.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

A follow-up is needed in the modelopt-internal repo:
`modelopttools:eval-config`
Step 3c is the GDPVal SIF / comparison-mode conversion checklist, and
Step 3d names
the gym image. Step 3c is now dead and the pointers this skill used to
make into it
are gone.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Evaluation Updates**
* Added standalone NeMo Gym support for MRCR tasks with pinned launcher
and Gym configuration requirements.
* Added shared guidance for Gym setup, validation, execution, and
recovery.
* AA requests now generate only `aa/` tasks and report per-task scores.

* **Removed Support**
* Removed GDPVal evaluation recipes, documentation, task guidance,
Apptainer helper tooling, and AA-suite inclusion.
* Updated default quantized-checkpoint validation recommendations to
exclude GDPVal.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 16:09:15 -07:00
didi 76c04dfd99 docs: clarify canonical pruning documentation source (#1871) (#2469)
Fixes #1871

The pruning documentation is split between
`docs/source/guides/3_pruning.rst`
and `examples/pruning/README.md`, causing confusion about which is
authoritative.

The examples/pruning/README.md is the comprehensive, up-to-date
reference
covering Minitron, Puzzletron, FastNAS, support matrix, guidelines, and
distillation hyperparameters.

This PR adds a note to the RST guide making clear:
- The README is canonical for Minitron and Puzzletron (LLM/VLM pruning)
- The guide covers FastNAS for Computer Vision models

Signed-off-by: Diya <diyaismahil7@gmail.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the pruning guide’s introductory content for clearer
separation of general guidance and the related Minitron/Puzzletron note.
- Clarified that the guide focuses on FastNAS pruning for computer
vision models.
- Added references to the Pruning README for Minitron and Puzzletron API
examples, support information, guidelines, and distillation
hyperparameters.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: didi <diyaismahil7@gmail.com>
2026-09-18 19:32:06 +00:00