Files
Chenjie LuoandClaude Opus 5.5 49f8d1eda8 Add the IQ1_M CUDA encoder and register the format
Second of two changes adding IQ1_M. The previous change landed the PyTorch
codec; this one adds its CUDA encoder and makes the format reachable.

In the kernel the delta shift is free per group, so it sits above the entry
index in the sort key: a tie still prefers the lower shift and then the lower
entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with
IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels
read it from global memory and rely on the cache. On a 5632x2048 weight the
encoder runs at 318 M elem/s against the torch search's 5.6, and its packed
bytes match the PyTorch encoder's.

The two IQ1 kernels load each vector, score it against a grid entry and apply
the +/-1/8 shift the same way, so those three steps move into common.cuh as
load_vector, grid_terms and shifted_error, and IQ1_S uses them too. IQ1_S's
packed bytes are unchanged.

IQ1_M gets an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend
dispatch, both exporters and convert_hf_config take it from there. The ggml
package exports it, and a general/ptq/iq1_m recipe uses it. Registering it
brings it under every registry-driven test with no IQ1_M-specific test code,
and the hf_ptq example test gains an IQ1_M case.

IQ1_M covers 25 tensors and 1.2B parameters of the mixed-precision checkpoint
IQ2_XXS measured. With it, ModelOpt supports all five GGML IQ formats at one
and two bits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-09-30 16:47:00 +00:00
..
…

This directory holds model-agnostic general recipes

ptq/*: model-agnostic general Post Training Quantization recipes.