Files
Model-Optimizer/tests/gpu
Chenjie LuoandClaude Opus 5.5 400498d82d [2/4] Register each GGML IQ format once for dispatch and export (#2525)
### What does this PR do?

Type of change: refactor (no behaviour change)

Addresses review feedback on #2511. Backend dispatch and export each
kept their own list of the GGML IQ formats: `_FAKE_QUANTS` in the
backend, and `IQ_FORMATS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` in
export. All four listed the same formats. Adding a format meant a row in
each, and the lists could drift apart. That had already happened twice
in #2511: `convert_hf_config.py` kept its own upper-case spelling of the
family and dropped IQ2_XXS metadata, and the Megatron export tests were
hard-wired to two formats.

Each format module now declares **one `IQFormat` record** beside its
encoder and decoder: name, block geometry, `quantize`, `dequantize`, and
its encode and decode chunk defaults. **`IQ_FORMAT_REGISTRY`** lists
them.

- Backend dispatch looks formats up in the registry.
- Both exporters take the packer and block geometry from it.
- Export's `IQ_FORMATS` is derived from it instead of being written out
again.
- `_FAKE_QUANTS`, `IQ_BLOCK_METADATA` and `IQ_PACKERS` are removed.
- The per-format fake-quant wrappers collapse into one
`IQFormat.fake_quant`, which does the `num_bits` check and calls the
existing cache helper.

Codebooks, searches, payload layouts and CUDA encoders stay in each
format's module.

#### Series and merge order

This is one slice of the IQ format series. It targets `main` so unit CI
runs, and **its diff includes #2511's commits until #2511 merges**.

1. #2511 — IQ2_XXS format
2. **this PR** — one registration per format
3. #2512 — IQ2_S format
4. #2513 — IQ1_M format

After this lands, #2512 and #2513 are restacked onto it, so each adds a
format module and a single registry entry instead of rows in four
tables.

#### Design choices

- **An explicit list, not self-registration at import.** If formats
registered themselves when their module was imported, the registry's
contents would depend on import order.
- **Backward compatible, with one behaviour change.** `iq1_s_fake_quant`
and `iq2_xs_fake_quant` are public on main, so each format keeps its
`<fmt>_fake_quant` name as an alias of its record's method. The three
removed tables were introduced by #2511 and never released. The
behaviour change: on main, the alias looked the encoder up at call time,
so patching `iq1_s.quantize_iq1_s` changed what it ran. Now the record
captures the encoder and decoder when it's built, so patching those
module functions reaches neither dispatch nor the alias. Substitute
through `IQ_FORMAT_REGISTRY` instead.
- **Registering a format declares it exportable, and that's intended.**
Export's `IQ_FORMATS` is derived from the registry, so a format
registered for dispatch is also claimed by both exporters and
`convert_hf_config`. That can't be wrong for an IQ format: fake quant is
`dequantize(quantize(w))`, so a format can't be dispatched without the
packer and block geometry, and those are all export reads. A QAT-only IQ
format can't exist. If one ever needs to land ahead of its export path,
an `exportable` flag on the record is a one-line addition.
- **The registry is the substitution seam.** Dispatch now reads the
registry, so tests that swap an encoder or decoder swap the registry
entry. Patching the format module's function would no longer reach
dispatch.
- **Test expectations stay independent of the registry.** Tests take the
*list* of formats from the registry, but their expected values come from
each format's own module (`quantize_<fmt>`, `<FMT>_BLOCK_BYTES`, …). A
mis-wired registry entry therefore can't make both sides of an assertion
agree.

#### What it does not unify

The CUDA side (`ggml.cpp` bindings, the `extensions.py` source list,
codebook sizes in `common.cuh`) and the recipes and docs remain per
format. "One registration" holds for the Python side, which is where all
four tables lived.

### Usage

Adding a format after this PR (for example IQ2_S in #2512) needs its
module and one line in the registry:

```python
# modelopt/torch/quantization/ggml/iq2_s.py
IQ2_S_FORMAT = IQFormat(
    name="iq2_s",
    block_size=IQ2_S_BLOCK_SIZE,
    block_bytes=IQ2_S_BLOCK_BYTES,
    quantize=quantize_iq2_s,
    dequantize=dequantize_iq2_s,
    block_chunk_size=_DEFAULT_BLOCK_CHUNK_SIZE,
    decode_chunk_size=_DEFAULT_DECODE_CHUNK_SIZE,
)

# modelopt/torch/quantization/ggml/registry.py
IQ_FORMAT_REGISTRY = {fmt.name: fmt for fmt in (IQ1_S_FORMAT, IQ2_XXS_FORMAT, IQ2_XS_FORMAT, IQ2_S_FORMAT)}
```

Looking up a format:

```python
from modelopt.torch.quantization.ggml import IQ_FORMAT_REGISTRY

fmt = IQ_FORMAT_REGISTRY["iq2_xxs"]
packed, shape = fmt.quantize(weight)          # GGML blocks
fmt.block_bytes, fmt.effective_bits           # 66, 2.0625
```

### Testing

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py` — **94 passed**
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — **22 passed**
(RTX PRO 6000)
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
iq` — **27 passed** in `nvcr.io/nvidia/nemo:26.08`, the image CI uses
for that suite
- broader sweep of IQ, export and recipe unit tests — **153 passed**,
none failed

**New guards on the registry itself:**
- every encoder the package exports is registered
- each record points at its own format's codec, geometry and chunk
defaults
- the public `<fmt>_fake_quant` alias is the registered record's method
- export's `IQ_FORMATS` and `QUANTIZATION_IQ*` constants match the
registry
- a format's `fake_quant` refuses a quantizer configured for another
format. Dispatch picks the record by `num_bits`, so it never reaches
this guard; the test covers direct callers of a record or alias. The
three per-format guards it replaced were untested on main.
- every registered format is listed in the shared test batteries

Checked by mutation: leaving IQ2_XXS out of the registry, or registering
it with the IQ2_XS encoder, each fails the guard written for that case.

**Coverage gap closed along the way:** `test_ggml_backend.py` was
hard-wired to IQ1_S and IQ2_XS, so IQ2_XXS had no backend, cache or
packed-once coverage. Those tests now run over the registry.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — public per-format fake-quant
names are kept as aliases; the removed tables were never released.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A — internal refactor with no
user-visible change
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

Review feedback on #2511 that this addresses: *"`_FAKE_QUANTS`,
`IQ_FORMATS`, `IQ_BLOCK_METADATA`, and `IQ_PACKERS` independently
enumerate the same formats. A common pack/dequantize/fake_quant
interface would let backend dispatch and export consume one
registration."*

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* IQ quantization formats are available through a shared format
registry, keeping format details and quantization behavior consistent
across supported workflows.
* IQ-format model exports use registered format information for
quantization metadata and weight packing.
* **Tests**
* Expanded checks to cover registered IQ formats and verify consistent
format support across quantization and export.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 23:32:11 +00:00
..
…