mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature llama.cpp defines five GGML IQ formats at one and two bits; we ship two. This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and IQ2_XS, and is the **first of three**. On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`, `Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84 B parameters — 10.6% of the file**, which a reader limited to IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats account for 17.3%. | format | bpw | bytes/256 | codebook | | |---|---|---|---|---| | `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing | | **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this PR** | | `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing | The encoder follows the existing single-pass grid search at a fixed anchored super-block scale, and the CUDA kernel the existing per-block structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a 4-bit sub-block scale into the same 32-bit word as four 7-bit sign indices, and its 256-entry grid needs no high index bits. ### Groundwork the next two reuse Two things land here because IQ2_XXS is the first format to need them: - **Export registry.** The IQ family was spelled as a two-element tuple at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and `unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset plus per-format packer and block-geometry tables, so a format is a row rather than a sweep through the exporters. - **Shared test contract.** The per-format test files had drifted apart — each of `iq1_s` and `iq2_xs` tested things the other did not. They become one parametrized module per layer (unit and CUDA), so every format is held to the same contract and a new one inherits it. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs ``` ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared against `dequantize_row_iq2_xxs` from `ggml-quants.c`: ``` IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry, as does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint ship as conformance vectors so CI keeps checking bytes we did not produce; mutation testing confirms they catch a wrong sign-field width. The CUDA encoder is byte-identical to the PyTorch reference on a fixed input and runs at **1047.9 M elem/s against the torch search's 10.7** on a 5632×2048 weight. - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99 passed - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7 checks × 3 formats) - `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds 29 recipes, `ptq.md` updated - reconstruction error decreases monotonically with bit width, pinned by a test Pre-existing failures in `tests/unit/torch/export/` and `test_autoquant.py` are `transformers`/`torchvision` import problems in my environment — identical counts with and without this change. ### A finding about already-merged code Checking the new kernel against its PyTorch reference at 4096 blocks showed that **CUDA and torch encoders disagree on roughly 1 block in 6000 — including the already-merged `iq2_xs`**, at 0.0163% against IQ2_XXS's 0.0000%. Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA fuses it with `fmaf` while torch uses separate ops; where two local scales fall within a float32 ULP the roundings pick different sides. Adjudicated against float64, neither path is better (5 to 6). Worst-case cost is **1.48e-08** relative reconstruction error, and run-to-run determinism on a given device holds. This is pre-existing, not introduced here — `test_iq2_xs_cuda.py` asserts exact byte parity but on a 16-block weight where ties essentially never arise. I have **not** changed that test; rewording a guarantee on merged code belongs in its own change. The new shared GPU tests assert exact parity on a small fixed input and compare reconstruction error at scale. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new codebook is a GGML table, carried in `codebooks.py` beside the existing ones so the MIT-licensed surface stays in that one file, with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information First of three; **IQ2_S** and **IQ1_M** follow and build on this branch. Replaces #2505, which carried all three at once. Follows #2446 / #2447 / #2448 / #2449, which landed IQ1_S and IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ2_XXS weight-only quantization, including CUDA acceleration and support for Hugging Face and Megatron exports. - Added the `general/ptq/iq2_xxs` recipe. It requires no calibration data and supports eligible layers with a weight dimension divisible by 256. - Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at approximately 1.56, 2.06, and 2.31 bits per weight, respectively. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>