Files
Model-Optimizer/modelopt
1b4e7dfb14 [OMNIML-5899] Add IQ post-training quantization recipes (#2449)
## Summary

- add numerics configs for IQ1_S and IQ2_XS
- add model presets and general PTQ recipes
- document the supported weight-shape requirement
- reuse the packed IQ weight across forwards instead of re-encoding it
every time
- add end-to-end `hf_ptq` coverage for both formats
- add the release-note entry

## PR split

This work is split into four focused PRs. Each PR targets `main` and
owns a disjoint file set:

1. **Kernel** — [#2448: Add CUDA kernels for IQ
packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448)
2. **Quantization** — [#2446: Add IQ quantization codecs and
backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446)
3. **Export** — [#2447: Export IQ checkpoints from HF and
Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447)
4. **Recipes** — [#2449: Add IQ post-training quantization
recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449)

The required merge order is #2448, #2446, #2447, then #2449.

## Scope

Sixteen files. The PR started as eight recipe config, documentation and
changelog files; the end-to-end test added for them surfaced two
performance
bugs in the already-merged codec, and fixing those pulled in the codec
files
and their tests.

Beyond the original recipe set it now touches four files owned by #2446
—
`ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and
`ggml/backend.py` —
plus three test files. It still contains no kernel or export files.

Keeping those fixes here rather than moving them to #2446 is deliberate
and
confirmed with the stack owner: #2446 is already merged, and both bugs
are only
observable through the end-to-end test this PR adds, so splitting them
would
separate each fix from the test that demonstrates it.

## Packed-weight cache fix

Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S
`hf_ptq`
run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing
the same
154 weights 15400 times — 100 times each.

The repacking is not calibration. These recipes set `algorithm: null`
and
`hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100
passes are
the sample `generate()` calls `hf_ptq` makes before and after
quantization: one
decode step re-runs weight fake-quant on every linear, and the packed
payload
was thrown away each time.

`_PackedWeightCache` was already there to prevent exactly this, and it
never
hit. It keyed on the identity of the tensor the backend was handed, but
`TensorQuantizer` passes a fresh *view* of the weight on every forward,
so the
identity check never matched twice.

The fix anchors the entry to `inputs._base` — the parameter the view is
taken
from — held as a weakref. The parameter is stable across forwards, so
the cache
hits; the reference stays weak, so the payload is released with the
weight and
offloaded/meta-device flows are unaffected. (A strong reference does
make the
cache hit, but pins full-precision storage for the life of the
quantizer, which
is the opposite of what those flows need.)

Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and
after:

| | packer calls | packing time | wall clock |
|---|---|---|---|
| before | 15400 | 499.0 s | 8m57s |
| after | 154 (one per weight) | 6.2 s | 3m21s |

`test_ggml_weight_is_packed_once_across_forwards` pins this: it counts
encoder
calls across five forwards under `torch.inference_mode()` (what
`generate()`
runs under) and asserts exactly one.

## Decode chunk fix

With packing cached, the end-to-end cost moved entirely into the decode,
and
IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case).
Instrumenting
both showed packing was no longer the cost at all — IQ2_XS packs
*faster*:

| | pack calls | packing time | wall clock |
|---|---|---|---|
| IQ1_S | 154 | 6.2 s | 3m21s |
| IQ2_XS | 154 | 1.6 s | 17m59s |

The cause was one constant serving two loops with opposite
characteristics.
`_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which
holds the
large codebook-search temporaries and runs once per weight; IQ2_XS sets
it to
256 rather than IQ1_S's 1024 because its search sweeps sixteen local
scales per
grid tile. But the *decode* shared it — and the decode has tiny
temporaries,
runs on every forward, and is never cached, so a small chunk only
multiplies
kernel launches.

Decoding a 2048x5632 weight:

| chunk | IQ2_XS decode | transient peak |
|---|---|---|
| 256 (was) | 91.4 ms | +24 MiB |
| 1024 | 22.9 ms | +31 MiB |
| 4096 (now) | 5.8 ms | +56 MiB |

The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded
through
`fake_quantize_with_cache`. The encode bounds are untouched, so the
memory
ceiling stays where it was aimed. The seven-iteration sign-parity loop
in
`dequantize_iq2_xs` is also folded into three XOR steps, off the same
per-forward path.

Per case in `tests/examples/hf_ptq` on 2xH100:

| | before | after |
|---|---|---|
| IQ1_S | 259.65 s | 94.64 s |
| IQ2_XS | 1081.05 s | 102.77 s |

Both now fit the 300s `tests/examples` default, so the cases carry no
explicit
timeout.

## Integration contracts

- `block_sizes: {-1: 256}` records the native packed-block contract; it
does not drive the GGML fake-quant scale search. The export path reads
`TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`,
validates it against the format block size, and records `group_size:
256` in checkpoint metadata.
- The model presets intentionally expose `--qformat iq1_s` and
`--qformat iq2_xs` in `hf_ptq`.

## Testing

- `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed,
1 skipped
- `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load
with the expected backend and no unused search option
- `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2
passed on 2×H100 (TinyLlama, both formats end to end through export),
3m17s for the pair
- pre-commit hooks pass on all changed files
- larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers)
quantizes
and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself
192s

---------

Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 07:46:50 +00:00
..
…