mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: refactor (no behaviour change) Addresses review feedback on #2511. Backend dispatch and export each kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in export. All four listed the same formats. Adding a format meant a row in each, and the lists could drift apart. That had already happened twice in #2511: `convert_hf_config.py` kept its own upper-case spelling of the family and dropped IQ2_XXS metadata, and the Megatron export tests were hard-wired to two formats. Each format module now declares **one `IQFormat` record** beside its encoder and decoder: name, block geometry, `quantize`, `dequantize`, and its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists them. - Backend dispatch looks formats up in the registry. - Both exporters take the packer and block geometry from it. - Export's `IQ_FORMATS` is derived from it instead of being written out again. - `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed. - The per-format fake-quant wrappers collapse into one `IQFormat.fake_quant`, which does the `num_bits` check and calls the existing cache helper. Codebooks, searches, payload layouts and CUDA encoders stay in each format's module. #### Series and merge order This is one slice of the IQ format series. It targets `main` so unit CI runs, and **its diff includes #2511's commits until #2511 merges**. 1. #2511 — IQ2_XXS format 2. **this PR** — one registration per format 3. #2512 — IQ2_S format 4. #2513 — IQ1_M format After this lands, #2512 and #2513 are restacked onto it, so each adds a format module and a single registry entry instead of rows in four tables. #### Design choices - **An explicit list, not self-registration at import.** If formats registered themselves when their module was imported, the registry's contents would depend on import order. - **Backward compatible, with one behaviour change.** `iq1_s_fake_quant` and `iq2_xs_fake_quant` are public on main, so each format keeps its `<fmt>_fake_quant` name as an alias of its record's method. The three removed tables were introduced by #2511 and never released. The behaviour change: on main, the alias looked the encoder up at call time, so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record captures the encoder and decoder when it's built, so patching those module functions reaches neither dispatch nor the alias. Substitute through `IQ_FORMAT_REGISTRY` instead. - **Registering a format declares it exportable, and that's intended.** Export's `IQ_FORMATS` is derived from the registry, so a format registered for dispatch is also claimed by both exporters and `convert_hf_config`. That can't be wrong for an IQ format: fake quant is `dequantize(quantize(w))`, so a format can't be dispatched without the packer and block geometry, and those are all export reads. A QAT-only IQ format can't exist. If one ever needs to land ahead of its export path, an `exportable` flag on the record is a one-line addition. - **The registry is the substitution seam.** Dispatch now reads the registry, so tests that swap an encoder or decoder swap the registry entry. Patching the format module's function would no longer reach dispatch. - **Test expectations stay independent of the registry.** Tests take the *list* of formats from the registry, but their expected values come from each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A mis-wired registry entry therefore can't make both sides of an assertion agree. #### What it does not unify The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list, codebook sizes in `common.cuh`) and the recipes and docs remain per format. "One registration" holds for the Python side, which is where all four tables lived. ### Usage Adding a format after this PR (for example IQ2_S in #2512) needs its module and one line in the registry: ```python # modelopt/torch/quantization/ggml/iq2_s.py IQ2_S_FORMAT = IQFormat( name="iq2_s", block_size=IQ2_S_BLOCK_SIZE, block_bytes=IQ2_S_BLOCK_BYTES, quantize=quantize_iq2_s, dequantize=dequantize_iq2_s, block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE, decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE, ) # modelopt/torch/quantization/ggml/registry.py IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)} ``` Looking up a format: ```python from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY fmt = IQ_FORMAT_REGISTRY["iq2_xxs"] packed, shape = fmt.quantize(weight) # GGML blocks fmt.block_bytes, fmt.effective_bits # 66, 2.0625 ``` ### Testing - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py` — **94 passed** - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed** (RTX PRO 6000) - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses for that suite - broader sweep of IQ, export and recipe unit tests — **153 passed**, none failed **New guards on the registry itself:** - every encoder the package exports is registered - each record points at its own format's codec, geometry and chunk defaults - the public `<fmt>_fake_quant` alias is the registered record's method - export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the registry - a format's `fake_quant` refuses a quantizer configured for another format. Dispatch picks the record by `num_bits`, so it never reaches this guard; the test covers direct callers of a record or alias. The three per-format guards it replaced were untested on main. - every registered format is listed in the shared test batteries Checked by mutation: leaving IQ2_XXS out of the registry, or registering it with the IQ2_XS encoder, each fails the guard written for that case. **Coverage gap closed along the way:** `test_ggml_backend.py` was hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or packed-once coverage. Those tests now run over the registry. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — public per-format fake-quant names are kept as aliases; the removed tables were never released. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A — internal refactor with no user-visible change - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`, `IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently enumerate the same formats. A common pack/dequantize/fake_quant interface would let backend dispatch and export consume one registration."* 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * IQ quantization formats are available through a shared format registry, keeping format details and quantization behavior consistent across supported workflows. * IQ-format model exports use registered format information for quantization metadata and weight packing. * **Tests** * Expanded checks to cover registered IQ formats and verify consistent format support across quantization and export. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>