mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary - add numerics configs for IQ1_S and IQ2_XS - add model presets and general PTQ recipes - document the supported weight-shape requirement - reuse the packed IQ weight across forwards instead of re-encoding it every time - add end-to-end `hf_ptq` coverage for both formats - add the release-note entry ## PR split This work is split into four focused PRs. Each PR targets `main` and owns a disjoint file set: 1. **Kernel** — [#2448: Add CUDA kernels for IQ packing](https://github.com/NVIDIA/Model-Optimizer/pull/2448) 2. **Quantization** — [#2446: Add IQ quantization codecs and backend](https://github.com/NVIDIA/Model-Optimizer/pull/2446) 3. **Export** — [#2447: Export IQ checkpoints from HF and Megatron](https://github.com/NVIDIA/Model-Optimizer/pull/2447) 4. **Recipes** — [#2449: Add IQ post-training quantization recipes](https://github.com/NVIDIA/Model-Optimizer/pull/2449) The required merge order is #2448, #2446, #2447, then #2449. ## Scope Sixteen files. The PR started as eight recipe config, documentation and changelog files; the end-to-end test added for them surfaced two performance bugs in the already-merged codec, and fixing those pulled in the codec files and their tests. Beyond the original recipe set it now touches four files owned by #2446 — `ggml/common.py`, `ggml/iq1_s.py`, `ggml/iq2_xs.py` and `ggml/backend.py` — plus three test files. It still contains no kernel or export files. Keeping those fixes here rather than moving them to #2446 is deliberate and confirmed with the stack owner: #2446 is already merged, and both bugs are only observable through the end-to-end test this PR adds, so splitting them would separate each fix from the test that demonstrates it. ## Packed-weight cache fix Adding the end-to-end test made the cost visible: a TinyLlama IQ1_S `hf_ptq` run spent **499 of its 537 seconds inside the IQ1_S encoder**, packing the same 154 weights 15400 times — 100 times each. The repacking is not calibration. These recipes set `algorithm: null` and `hf_ptq` logs `Dynamic quantization. Calibration skipped.`. The 100 passes are the sample `generate()` calls `hf_ptq` makes before and after quantization: one decode step re-runs weight fake-quant on every linear, and the packed payload was thrown away each time. `_PackedWeightCache` was already there to prevent exactly this, and it never hit. It keyed on the identity of the tensor the backend was handed, but `TensorQuantizer` passes a fresh *view* of the weight on every forward, so the identity check never matched twice. The fix anchors the entry to `inputs._base` — the parameter the view is taken from — held as a weakref. The parameter is stable across forwards, so the cache hits; the reference stays weak, so the payload is released with the weight and offloaded/meta-device flows are unaffected. (A strong reference does make the cache hit, but pins full-precision storage for the life of the quantizer, which is the opposite of what those flows need.) Measured on TinyLlama IQ1_S `hf_ptq`, 2×H100, same command before and after: | | packer calls | packing time | wall clock | |---|---|---|---| | before | 15400 | 499.0 s | 8m57s | | after | 154 (one per weight) | 6.2 s | 3m21s | `test_ggml_weight_is_packed_once_across_forwards` pins this: it counts encoder calls across five forwards under `torch.inference_mode()` (what `generate()` runs under) and asserts exactly one. ## Decode chunk fix With packing cached, the end-to-end cost moved entirely into the decode, and IQ2_XS was still 4x slower than IQ1_S (1081s vs 260s per case). Instrumenting both showed packing was no longer the cost at all — IQ2_XS packs *faster*: | | pack calls | packing time | wall clock | |---|---|---|---| | IQ1_S | 154 | 6.2 s | 3m21s | | IQ2_XS | 154 | 1.6 s | 17m59s | The cause was one constant serving two loops with opposite characteristics. `_DEFAULT_BLOCK_CHUNK_SIZE` bounds the torch encode fallback, which holds the large codebook-search temporaries and runs once per weight; IQ2_XS sets it to 256 rather than IQ1_S's 1024 because its search sweeps sixteen local scales per grid tile. But the *decode* shared it — and the decode has tiny temporaries, runs on every forward, and is never cached, so a small chunk only multiplies kernel launches. Decoding a 2048x5632 weight: | chunk | IQ2_XS decode | transient peak | |---|---|---| | 256 (was) | 91.4 ms | +24 MiB | | 1024 | 22.9 ms | +31 MiB | | 4096 (now) | 5.8 ms | +56 MiB | The decode now takes its own `_DEFAULT_DECODE_CHUNK_SIZE`, threaded through `fake_quantize_with_cache`. The encode bounds are untouched, so the memory ceiling stays where it was aimed. The seven-iteration sign-parity loop in `dequantize_iq2_xs` is also folded into three XOR steps, off the same per-forward path. Per case in `tests/examples/hf_ptq` on 2xH100: | | before | after | |---|---|---| | IQ1_S | 259.65 s | 94.64 s | | IQ2_XS | 1081.05 s | 102.77 s | Both now fit the 300s `tests/examples` default, so the cases carry no explicit timeout. ## Integration contracts - `block_sizes: {-1: 256}` records the native packed-block contract; it does not drive the GGML fake-quant scale search. The export path reads `TensorQuantizer.block_sizes[-1]` through `get_weight_block_size`, validates it against the format block size, and records `group_size: 256` in checkpoint metadata. - The model presets intentionally expose `--qformat iq1_s` and `--qformat iq2_xs` in `hf_ptq`. ## Testing - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2'` — 52 passed, 1 skipped - `tests/unit/recipe/test_presets.py` — both shipped IQ recipes load with the expected backend and no unused search option - `tests/examples/hf_ptq/test_llm_ptq.py -k 'iq1_s or iq2_xs'` — 2 passed on 2×H100 (TinyLlama, both formats end to end through export), 3m17s for the pair - pre-commit hooks pass on all changed files - larger-model sanity check outside CI: Qwen3.8-27B (2256 quantizers) quantizes and exports with `general/ptq/iq1_s` on 2xH100 in 66m46s, export itself 192s --------- Signed-off-by: Hung-Yueh Chiang <hungyuehc@nvidia.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>