[5/6] Add the IQ1_M codec (#2513)

### What does this PR do?

Type of change: new feature (not yet user-reachable)

**First of two PRs adding IQ1_M** at 1.75 bits per weight, just above
IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is
deliberately **not registered**, so no quantizer dispatches to it and
the `ggml` package does not export it. #2595 adds the CUDA encoder,
registers the format and adds its recipe. With both, ModelOpt supports
all five GGML IQ formats at one and two bits.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B
parameters**. With all five formats we can read 89.0% of that file; the
rest is k-quants and F32.

### What's distinctive about it

**IQ1_M is the most irregular layout of the five.** There is no leading
block scale field at all. The FP16 super-block scale is reassembled from
the top nibble of each of four scale words:

```c
scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000);
```

It is also finer grained than IQ1_S: a local scale per **two** groups
rather than four, and a delta shift chosen **per group** rather than per
sub-block. That is where its extra 0.1875 bits go.

### Shared with IQ1_S rather than copied

IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same
±1/8 delta, every (shift, local scale) choice for every 8-value vector.
It differs only in how it selects among those choices afterwards. So the
search moves out of IQ1_S's encoder into `_search_shifted_grid`, which
both call, and `iq1_m.py` keeps only its selection and packing.
**IQ1_S's encoded bytes are unchanged**, checked by hashing its output
before and after on a fixed input.

### A scale-anchor correction

IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a
block's peak-to-RMS** rather than being flat, and clamps higher. It uses
`clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat
`0.61`. Measured over 15 Qwen3.8-27B MLP weights:

| | flat 0.61 | correct anchor | |
|---|---|---|---|
| relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** |

It is consistent on every tensor, with no outliers. The anchor changes
quality without touching layout, so neither round-trip nor conformance
tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms`
now pins it, for all five formats; see Testing.

### Family parity

Two surface asymmetries close here, so the five are uniform. `IQ1_S` now
exposes `_predict_iq1_s_scales` like the other four, instead of
computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the
IQ1_S table it shares.

### Testing

**The decoder is validated against llama.cpp's own output, not just
round-tripped:**

```
IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0
```

This mattered: **my first IQ1_M decoder had a real bug.** A
`repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where
llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder
still passed, because the encoder made the matching mistake. Only
comparison against bytes we did not produce caught it. Blocks from that
checkpoint ship as conformance vectors, and mutation testing confirms
they catch a mis-set scale nibble.

The decoder unpacks every field in one vectorized pass, since fake quant
decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3).

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M
codec cases, including the llama.cpp conformance check
- `test_scale_anchor_follows_peak_to_rms` pins every format's scale
anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4,
8 and 16, reaching both clamps and two points on each slope, and
compares them against anchors written out in the test. Mutations each
fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61,
moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S
or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's
CUDA-vs-PyTorch parity still holds after its encoder refactor.
- IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash
identically to the pre-split version of this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds
no codebook; it reuses the IQ1_S table already carried in
`codebooks.py`. The new conformance vectors come from
`unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model
`Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595
carries the entry.
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration), all merged →
**this** → #2595 (IQ1_M CUDA encoder and registration).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M quantization and dequantization for compact,
GGML-compatible blocks of 256 values.
* Added access to the IQ1_M grid and configurable chunk sizes for
processing data.

* **Bug Fixes**
* Improved IQ1_S scale prediction and grid-search organization while
preserving its encoding behavior.

* **Tests**
* Added IQ1_M conformance data and included the format in shared
IQ-format test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chenjie Luo
2026-09-30 09:45:48 -07:00
committed by GitHub
co-authored by Claude Opus 5.5
parent ad9ea97a4b
commit e5b63320ab
4 changed files with 327 additions and 17 deletions
+219
View File
@@ -0,0 +1,219 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""IQ1_M fake quantization and GGML-compatible block packing.
The encoder performs a single-pass squared-error grid search at a fixed,
empirically anchored super-block scale, mirroring :mod:`.iq1_s`. Every 256
logical values become one 56-byte block_iq1_m payload:
* bytes 0..31: 32 low bytes of the grid index, four per sub-block
* bytes 32..47: 16 bytes holding, per nibble, three grid-index high bits and
one delta-shift bit -- two groups per byte
* bytes 48..55: four little-endian uint16 holding four 3-bit local scales each
in bits 0..11, and one nibble of the FP16 super-block scale in bits 12..15
IQ1_M has no dedicated ``d`` field: the FP16 super-block scale is reassembled
from the top nibble of each of the four scale words. It is also finer-grained
than IQ1_S -- a local scale covers two groups rather than four, and the delta
shift is chosen per group rather than per sub-block -- which is where its extra
0.1875 bits per weight go.
The grid is the same canonical 2048 x 8 ternary table IQ1_S uses, carried from
llama.cpp ggml-common.h revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
The matching dequantization formula is in ggml-quants.c at the same revision:
https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c#L2573-L2620
"""
import torch
from .common import (
GGML_BLOCK_SIZE,
narrow_to_float32,
validate_block_chunk_size,
validate_packed_weights,
validate_weight,
)
from .iq1_s import _search_shifted_grid, iq1_s_grid
__all__ = [
"IQ1_M_BLOCK_BYTES",
"IQ1_M_BLOCK_SIZE",
"IQ1_M_EFFECTIVE_BITS",
"dequantize_iq1_m",
"iq1_m_grid",
"quantize_iq1_m",
]
IQ1_M_BLOCK_SIZE = GGML_BLOCK_SIZE
IQ1_M_BLOCK_BYTES = 56
IQ1_M_EFFECTIVE_BITS = IQ1_M_BLOCK_BYTES * 8 / IQ1_M_BLOCK_SIZE
_IQ1_M_DELTA = 0.125
# Largest representable magnitude: (1 + delta) at local scale 7 -> 15 * 1.125.
_IQ1_M_NATIVE_MAX = 16.875
# IQ1_M anchors differently from IQ1_S: the ratio rises with a block's peak-to-RMS
# instead of being flat, and it is allowed closer to full range. Values follow the
# reference predictor these encoders are derived from.
_IQ1_M_SCALE_ANCHOR_BASE = 0.58
_IQ1_M_SCALE_ANCHOR_MIN = 0.65
_IQ1_M_SCALE_ANCHOR_MAX = 0.95
_IQ1_M_PEAK_TO_RMS_TAPER = 0.035
_IQ1_M_GROUPS = 32
_IQ1_M_SUBBLOCKS = 8
_DEFAULT_BLOCK_CHUNK_SIZE = 1024
_DEFAULT_DECODE_CHUNK_SIZE = 4096
def iq1_m_grid(device: torch.device | str | None = None) -> torch.Tensor:
"""Return the canonical IQ1_M ternary grid as float32.
IQ1_M indexes the same 2048-entry table as IQ1_S; this alias exists so every format
exposes a grid accessor under its own name.
"""
return iq1_s_grid(device)
def _predict_iq1_m_scales(blocks: torch.Tensor) -> torch.Tensor:
"""Predict one FP16 super-block scale for each flattened block."""
x = narrow_to_float32(blocks)
amax = x.abs().amax(dim=1)
rms = x.square().mean(dim=1).sqrt()
peak_to_rms = torch.where(rms > 0, amax / rms, torch.zeros_like(rms))
anchor_ratio = (_IQ1_M_SCALE_ANCHOR_BASE + _IQ1_M_PEAK_TO_RMS_TAPER * peak_to_rms).clamp(
_IQ1_M_SCALE_ANCHOR_MIN, _IQ1_M_SCALE_ANCHOR_MAX
)
return ((amax / _IQ1_M_NATIVE_MAX) * anchor_ratio).clamp(max=65504.0).to(torch.float16)
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
"""Encode a moderate-size batch of flattened 256-value blocks."""
x = narrow_to_float32(blocks)
block_count = x.shape[0]
d = _predict_iq1_m_scales(x)
d_float = d.float()
best_error, best_entry = _search_shifted_grid(
x.reshape(block_count, _IQ1_M_GROUPS, 8), d_float, grid
)
# Each group picks its own shift; only the 3-bit local scale is shared, over two groups.
per_shift_error = best_error.reshape(block_count, _IQ1_M_GROUPS, 2, 8)
per_shift_entry = best_entry.reshape(block_count, _IQ1_M_GROUPS, 2, 8)
local_error, shift_index = per_shift_error.min(dim=2)
local_entry = per_shift_entry.gather(2, shift_index.unsqueeze(2)).squeeze(2)
pair_error = local_error.reshape(block_count, 16, 2, 8).sum(dim=2)
selected_local = pair_error.argmin(dim=-1)
group_local = selected_local.repeat_interleave(2, dim=1)
selected_entry = local_entry.gather(2, group_local.unsqueeze(-1)).squeeze(-1)
selected_shift = shift_index.gather(2, group_local.unsqueeze(-1)).squeeze(-1)
packed = torch.empty((block_count, IQ1_M_BLOCK_BYTES), dtype=torch.uint8, device=x.device)
packed[:, :32] = (selected_entry & 0xFF).to(torch.uint8)
# Each group's nibble is three index-high bits and its shift bit; two groups share a byte,
# low nibble first.
nibbles = ((selected_entry >> 8) & 0x7) | (selected_shift << 3)
packed[:, 32:48] = (nibbles[:, 0::2] | (nibbles[:, 1::2] << 4)).to(torch.uint8)
# Four scale words, each carrying four 3-bit local scales (two sub-blocks, two halves each)
# in bits 0..11 and one nibble of d in bits 12..15, low nibble in the first word. The fields
# are disjoint, so summing them is the same as OR-ing them.
d_bits = d.contiguous().view(torch.int16).to(torch.int64) & 0xFFFF
local_shifts = torch.tensor([0, 3, 6, 9], dtype=torch.int64, device=x.device)
nibble_shifts = torch.tensor([0, 4, 8, 12], dtype=torch.int64, device=x.device)
words = (selected_local.reshape(block_count, 4, 4) << local_shifts).sum(dim=-1)
words |= ((d_bits.unsqueeze(-1) >> nibble_shifts) & 0xF) << 12
packed[:, 48:56:2] = (words & 0xFF).to(torch.uint8)
packed[:, 49:56:2] = ((words >> 8) & 0xFF).to(torch.uint8)
return torch.where((d_float == 0).unsqueeze(1), 0, packed)
@torch.no_grad()
def quantize_iq1_m(
weight: torch.Tensor, *, block_chunk_size: int = _DEFAULT_BLOCK_CHUNK_SIZE
) -> tuple[torch.Tensor, torch.Tensor]:
"""Pack a floating-point weight into GGML-compatible IQ1_M blocks.
Returned shapes are ``[*weight.shape[:-1], weight.shape[-1] // 256, 56]``
and ``[weight.ndim]``.
"""
validate_weight(weight, "IQ1_M")
validate_block_chunk_size(block_chunk_size)
logical_shape = torch.tensor(weight.shape, dtype=torch.int64)
blocks = weight.contiguous().reshape(-1, IQ1_M_BLOCK_SIZE)
grid = iq1_s_grid(weight.device)
packed_shape = (*weight.shape[:-1], weight.shape[-1] // IQ1_M_BLOCK_SIZE, IQ1_M_BLOCK_BYTES)
chunks = [
_encode_blocks(blocks[start : start + block_chunk_size], grid)
for start in range(0, blocks.shape[0], block_chunk_size)
]
return torch.cat(chunks).reshape(packed_shape), logical_shape
@torch.no_grad()
def dequantize_iq1_m(
packed_weights: torch.Tensor,
weight_shape: torch.Tensor,
*,
dtype: torch.dtype = torch.bfloat16,
block_chunk_size: int = _DEFAULT_DECODE_CHUNK_SIZE,
) -> torch.Tensor:
"""Decode GGML-compatible IQ1_M payload bytes."""
shape = validate_packed_weights(
packed_weights, weight_shape, block_bytes=IQ1_M_BLOCK_BYTES, format_name="IQ1_M"
)
validate_block_chunk_size(block_chunk_size)
blocks = packed_weights.contiguous().reshape(-1, IQ1_M_BLOCK_BYTES)
grid = iq1_s_grid(blocks.device)
scale_shifts = torch.tensor([0, 3, 6, 9], dtype=torch.int64, device=blocks.device)
decoded = torch.empty((blocks.shape[0], IQ1_M_BLOCK_SIZE), dtype=dtype, device=blocks.device)
for start in range(0, blocks.shape[0], block_chunk_size):
stop = min(start + block_chunk_size, blocks.shape[0])
block_chunk = blocks[start:stop]
count = block_chunk.shape[0]
low = block_chunk[:, :32].to(torch.int64).reshape(count, _IQ1_M_SUBBLOCKS, 4)
qh = block_chunk[:, 32:48].to(torch.int64).reshape(count, _IQ1_M_SUBBLOCKS, 2)
words = block_chunk[:, 48:56:2].to(torch.int64) | (
block_chunk[:, 49:56:2].to(torch.int64) << 8
)
# The FP16 super-block scale is the four words' top nibbles, low word first.
d_bits = (
(words[:, 0] >> 12)
| ((words[:, 1] >> 8) & 0x00F0)
| ((words[:, 2] >> 4) & 0x0F00)
| (words[:, 3] & 0xF000)
)
d = d_bits.to(torch.int16).view(torch.float16).float()
# Each qh byte carries two groups, low nibble first: three high index bits and a delta sign.
nibbles = torch.stack((qh & 0xF, qh >> 4), dim=-1).reshape(count, _IQ1_M_SUBBLOCKS, 4)
entries = low | ((nibbles & 0x7) << 8)
deltas = torch.where((nibbles & 0x8).bool(), -_IQ1_M_DELTA, _IQ1_M_DELTA)
# Each word packs four 3-bit local scales below its nibble of d: two sub-blocks, two
# halves each, so word w, slot k is sub-block 2 * w + k // 2, half k % 2.
local = ((words.unsqueeze(-1) >> scale_shifts) & 0x7).reshape(count, _IQ1_M_SUBBLOCKS, 2)
scales = d.unsqueeze(-1).unsqueeze(-1) * (2 * local + 1).float()
values = grid[entries] + deltas.unsqueeze(-1)
# Groups 0,1 take the first local scale and groups 2,3 the second, so the repeat
# runs along the half axis: [h0, h0, h1, h1], not [h0, h1, h0, h1].
group_scale = scales.repeat_interleave(2, dim=2)
chunk_decoded = values * group_scale.unsqueeze(-1)
decoded[start:stop] = chunk_decoded.reshape(-1, IQ1_M_BLOCK_SIZE)
return decoded.reshape(shape)
+36 -14
View File
@@ -87,22 +87,34 @@ def iq1_s_grid(device: torch.device | str | None = None) -> torch.Tensor:
return _GRID_CACHE[resolved_device]
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
"""Encode a moderate-size batch of flattened 256-value blocks."""
def _predict_iq1_s_scales(blocks: torch.Tensor) -> torch.Tensor:
"""Predict one FP16 super-block scale for each flattened block.
The fixed-scale search favors a compressed super-block scale, so this empirical anchor
puts d below the full-range value. The CUDA encoder computes the same quantity in its
own scale kernel; this is the reference the torch path uses.
"""
x = narrow_to_float32(blocks)
block_count = x.shape[0]
vectors = x.reshape(block_count, 32, 8)
amax = x.abs().amax(dim=1)
return ((amax / _IQ1_S_NATIVE_MAX) * _IQ1_S_SCALE_ANCHOR).clamp(max=65504.0).to(torch.float16)
def _search_shifted_grid(
vectors: torch.Tensor, d: torch.Tensor, grid: torch.Tensor
) -> tuple[torch.Tensor, torch.Tensor]:
"""Score every grid entry against each 8-value vector at every shift and local scale.
Returns the lowest error and the entry reaching it, both ``[blocks, 32, 16]`` and indexed
by choice ``shift * 8 + local``. IQ1_M runs the same search over the same grid and delta;
the two formats differ only in how they select among the choices afterwards.
"""
block_count = vectors.shape[0]
xnorm = vectors.square().sum(dim=-1)
xsum = vectors.sum(dim=-1)
amax = x.abs().amax(dim=1)
# The fixed-scale search favors a compressed super-block scale. This
# empirical anchor initializes d below the full-range value.
d = ((amax / _IQ1_S_NATIVE_MAX) * _IQ1_S_SCALE_ANCHOR).clamp(max=65504.0).to(torch.float16)
d_float = d.float()
best_error = torch.full((block_count, 32, 16), torch.inf, dtype=torch.float32, device=x.device)
best_entry = torch.zeros((block_count, 32, 16), dtype=torch.int64, device=x.device)
best_error = torch.full(
(block_count, 32, 16), torch.inf, dtype=torch.float32, device=vectors.device
)
best_entry = torch.zeros((block_count, 32, 16), dtype=torch.int64, device=vectors.device)
grid_norm = grid.square().sum(dim=-1)
grid_sum = grid.sum(dim=-1)
@@ -120,7 +132,7 @@ def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
shifted_norm = tile_norm + 2 * delta * tile_sum + 8 * delta * delta
for local in range(8):
choice = shift * 8 + local
scale = d_float.reshape(-1, 1, 1) * (2 * local + 1)
scale = d.reshape(-1, 1, 1) * (2 * local + 1)
error = (
xnorm.unsqueeze(-1) - 2 * scale * shifted_dot + scale.square() * shifted_norm
).clamp_min_(0)
@@ -132,6 +144,16 @@ def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
best_entry[:, :, choice] = torch.where(
replace, tile_index + entry_start, best_entry[:, :, choice]
)
return best_error, best_entry
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
"""Encode a moderate-size batch of flattened 256-value blocks."""
x = narrow_to_float32(blocks)
block_count = x.shape[0]
d = _predict_iq1_s_scales(x)
d_float = d.float()
best_error, best_entry = _search_shifted_grid(x.reshape(block_count, 32, 8), d_float, grid)
group_error = best_error.reshape(block_count, 8, 4, 16).sum(dim=2)
selected_choice = group_error.argmin(dim=-1)
@@ -18,7 +18,8 @@
Each entry holds packed block bytes lifted verbatim from
unsloth/Qwen3.8-27B-GGUF (Qwen3.8-27B-UD-IQ1_S.gguf) together with the values
llama.cpp's own dequantize_row_* produces for them, computed from
ggml-quants.c revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
ggml-quants.c revision 9b05354ec6fb58b4e665e9a39ebc40285c015638. The checkpoint
is published under Apache-2.0, as is Qwen/Qwen3.8-27B, the model it quantizes.
These pin our decoders against bytes we did not produce. A decoder that drifts
from the GGML layout -- a mis-set high bit, a swapped scale nibble, a sign
@@ -62,6 +63,40 @@ _VECTORS = {
"w85GPeSrj5LVBGpy6iev/eS19nOE/x+gaS2i"
),
},
"iq1_m": {
"source": "blk.1.attn_gate.weight",
"block_bytes": 56,
"blocks": (
"eNoBUAGv/sgI5+PIe/BwwioFcBUzxw3iyphnqNoBMNQxtkYxKIg/RpRL53G82jgiUKrRvh5d/dzUmweXtBQV+zXsYEc4"
"BhC3bF7wWvQFzy2AI93zFgUGNb4wPvSYWoJISJPERQmGNC8t1ESzwwUl0VqLjpLMmxRBBIezwirprpBVD7/r/Ig7oFIl"
"9J2LtXeKAZqOhjjGIDxTHmLssDKbvAWf9um/y85ttT0H26bzHeDrAJbuisGPnSGlUT2c0uWVv/620cLVhOoO+ztPFia9"
"WPgjYsisuSabXH64WhO5WhRt5Z1kl3MXjvFr3vd3GQWsYMSRdfH1s7aUT2yDlAT+RMHIacwBCG0GIkBFwJylQDLB9wbg"
"zU483Lojdx+p3BogUaPTo7/iBfE6P39ha5fQr7yx4Uj2l/ngDgV3qsfCM5LfckI5HeHRM+6Q5AU8LF+SvtLEksaZEln9"
"qfk="
),
"expected": (
"eNqVV01oXkUUfWBAFNIWoVKhFNFVXViRuDAzDz8QhWjpQlcWVIqCRiy4aFy0FF+kRBdStBR3FgtCNkJWEsnMK6mgDdZV"
"LbRQq9BFEd3Unyq6EOfOe2e+M/fNl+ricO49d+bO3Dv3fSTVxaOmued7L6iCzRid21IzF/H6E7LfRj7zqeSxYNECXPBt"
"xnJesOOZwYfGnPSQq89jYFNuU5170wuaT05ZMCNotpqadsGWmIctHGM6t9yd66Dz01rU27H7r0DNWY3PrEaM7L4WtmD9"
"lXtr0YQF0DgebdU7zt9z6c6p3mrp2yJG2x9ttbb+x8EaeopT//kNoCdMTRtwWisavUfk8Cb8VvrtWOvZImfi/Ly0B7OQ"
"2RePutiTwBpaD99BCx22zp/dU86gHlDNBvfjM0o2vkv9DcIf9KY70+pas5lHTPjG5bXqQm2rLYf8wBa+UPvANgL2jcuz"
"CXtCTRszLsOB8z7xgfM2YU+4g0DsjRkTseNoQvN8qAm8cNX3uhMErYXNWrP6Yl1995XZBK6pHqwFYg9A5yfI2QF8D9wt"
"6Vg77pNNLBr3T/rEvWMNfUBvANG6vnW9lD6j19zfD2+3/wfNkwfbTENufj/4wngr8PCOfvD+nIP3oVbtw9a1d3FHs+QH"
"+ad+Npth9PCxerN4aeZK85fmTc/heJ2ZtD/bs3DVZvv5W9D90d8MfNaxvrSWY3oN9v56zDGa5R+t0mSNCboXHr2/vWZf"
"xyfZWhNILj4v+G3hLtkefWbz5zu2+vyXWXDC4rM+cUCKQ8eeQt2FHiQdkLvGdU99ZgSj356uhdc/3tUqiO4EYU0rgC9o"
"2t9ts/+Mjwh29d5zRiB+iTWa6ZO2mtntYjzwwA5x9hnxTIlTDvaLe/o4cgaWtUYDeljr4Q+0jt3Ee/fn0FkeMbqv5zMH"
"vtqfcqo6s/MKPeVzOVe16yM3OnxHvX72oVYgPjQBfEa2rp8Hngs9I6zxPEXQWZwX+i3vs2NuDHPND9hcC79XcybyzRUf"
"bYHYor1wxN8Sbz9uI974YS36wtA4L+cHSnHh3m5eur+udv/jI9sPjNhF7m2sFz/aGwuuGd3ZMldbL1mxox9saMlm6N7d"
"XLER0GEjzmtEO7nsJqFZsm3mn76r1nqquwDEimvQF31PfVeOd/PQgX3Yug+dbwfviveTmO4no+u5j9hYMOE9akFmv7rT"
"xVzCV961KTd00QDWhQHEWRvnFDYRUhN8sQOneeI5028wIRZ1XS/5xfnr/TSruLeuYVyvSfOGOqQGgWh7501z3zcevD57"
"Ww1fbO3r9dWhR4wg2Bac8OUpn7TeTuspxlq2j3xA76lOHDfNW395ZsHo5fD/cW9vCnV3vm8PmxDqjQh2WOci9s670Jc2"
"7u1t9oWhiQ+N1saeMtD33nd6H/ZGm3u5Sf/026T+vfZFwvrOB1oBbObR6v6a16V9/Kacd3y+nO3QM90/njmeN+rBoGfw"
"e20wv3qWucc8z3qu+XyA+89vmO6EOaDa9JzoerP56ePo00CnuS/ixHEX5t4KlyDfhHwL+B7E5++kOrsc/h5+zESshN8W"
"truYC2wzBlYu+YwZslYgMWbY8HEeny/n4l6wEWMWfL21iOb0kXZSLIOuS9953AdT6I1JtaAHw16ZDFxLsJu/766rKz/5"
"iJklE/3AEZ0e/uZcchHBDvFWAH+QX/eR6+L6xjXaibXxnfW7jJHn51nQ81Ly8Y4MnIP7c2/1+/R9495lPUSfqIeZPb/Y"
"Ydv12Yh9cz6xjkFnINckhPvpN+Y7F3uSa+M34bcBzy/6iG3X18J9bATbXAPX0tv/ArwdXdM="
),
},
"iq2_xxs": {
"source": "blk.0.attn_gate.weight",
"block_bytes": 66,
@@ -29,6 +29,7 @@ from _test_utils.torch.quantization.iq_llama_cpp_vectors import (
packed_blocks,
)
import modelopt.torch.quantization.ggml.iq1_m as iq1_m_module
import modelopt.torch.quantization.ggml.iq1_s as iq1_s_module
import modelopt.torch.quantization.ggml.iq2_s as iq2_s_module
import modelopt.torch.quantization.ggml.iq2_xs as iq2_xs_module
@@ -40,6 +41,7 @@ from modelopt.torch.quantization.nn import TensorQuantizer
# name -> (module, packed bytes per block, codebook entries, bits per weight)
FORMATS = {
"iq1_s": (iq1_s_module, 50, 2048, 1.5625),
"iq1_m": (iq1_m_module, 56, 2048, 1.75),
"iq2_xxs": (iq2_xxs_module, 66, 256, 2.0625),
"iq2_xs": (iq2_xs_module, 74, 512, 2.3125),
"iq2_s": (iq2_s_module, 82, 1024, 2.5625),
@@ -49,7 +51,18 @@ NAMES = sorted(FORMATS)
# tests that go through TensorQuantizer iterate these rather than every codec above.
DISPATCHED = sorted(IQ_FORMAT_REGISTRY)
# IQ1 grids are ternary; IQ2 grids hold the magnitudes 8, 25 and 43.
TERNARY = {"iq1_s"}
TERNARY = {"iq1_s", "iq1_m"}
# name -> (largest representable magnitude, scale anchor at peak-to-RMS 1, 4, 8 and 16), where the
# anchor is d * native max / amax. IQ1_S is flat; IQ1_M rises with peak-to-RMS and the IQ2 formats
# fall with it, each clamped at both ends. Written out rather than read from the modules, so a
# changed constant fails here.
ANCHORS = {
"iq1_s": (15 * 1.125, (0.61, 0.61, 0.61, 0.61)),
"iq1_m": (15 * 1.125, (0.65, 0.72, 0.86, 0.95)),
"iq2_xxs": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
"iq2_xs": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
"iq2_s": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
}
def _parts(name):
@@ -83,9 +96,11 @@ def test_canonical_grid(name):
@pytest.mark.parametrize("name", NAMES)
def test_grid_normalizes_unindexed_cuda_device(monkeypatch, name):
module, _, _, grid_fn, _, _, _ = _parts(name)
# IQ1_M shares the IQ1_S cache, being the same table.
cache_owner = iq1_s_module if name == "iq1_m" else module
cached = torch.empty(0)
monkeypatch.setattr(torch.cuda, "current_device", lambda: 7)
monkeypatch.setitem(module._GRID_CACHE, torch.device("cuda", 7), cached)
monkeypatch.setitem(cache_owner._GRID_CACHE, torch.device("cuda", 7), cached)
assert grid_fn("cuda") is cached
@@ -240,6 +255,25 @@ def test_search_is_independent_of_default_dtype(name):
torch.set_default_dtype(torch.float32)
@pytest.mark.parametrize("name", NAMES)
def test_scale_anchor_follows_peak_to_rms(name):
"""The predicted block scale must follow the format's anchor, through both clamps.
The anchor is an encoder choice: it changes quality without touching the layout, so no
round-trip or conformance test would notice it drifting. k equal unit spikes among 256 zeros
have a peak-to-RMS of exactly 16 / sqrt(k).
"""
native_max, anchors = ANCHORS[name]
blocks = torch.zeros(4, 256)
for row, spikes in enumerate((256, 16, 4, 1)):
blocks[row, :spikes] = 1.0
d = getattr(FORMATS[name][0], f"_predict_{name}_scales")(blocks)
assert d.dtype == torch.float16
# rtol covers the FP16 rounding of d, under 0.1%; a flat 0.61 anchor misses IQ1_M by 6% or more.
torch.testing.assert_close(d.float() * native_max, torch.tensor(anchors), rtol=2**-10, atol=0)
@pytest.mark.parametrize("name", formats())
def test_decoder_matches_llama_cpp_on_captured_blocks(name):
"""Decode bytes we did not produce and match llama.cpp's own output exactly.