mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[5/6] Add the IQ1_M codec (#2513)
### What does this PR do? Type of change: new feature (not yet user-reachable) **First of two PRs adding IQ1_M** at 1.75 bits per weight, just above IQ1_S. This one lands the **PyTorch codec**: encoder and decoder. It is deliberately **not registered**, so no quantizer dispatches to it and the `ggml` package does not export it. #2595 adds the CUDA encoder, registers the format and adds its recipe. With both, ModelOpt supports all five GGML IQ formats at one and two bits. On the mixed-precision checkpoint #2511 measured (`unsloth/Qwen3.8-27B-GGUF`), IQ1_M covers **25 tensors and 1.2 B parameters**. With all five formats we can read 89.0% of that file; the rest is k-quants and F32. ### What's distinctive about it **IQ1_M is the most irregular layout of the five.** There is no leading block scale field at all. The FP16 super-block scale is reassembled from the top nibble of each of four scale words: ```c scale.u16 = (sc[0] >> 12) | ((sc[1] >> 8) & 0x00f0) | ((sc[2] >> 4) & 0x0f00) | (sc[3] & 0xf000); ``` It is also finer grained than IQ1_S: a local scale per **two** groups rather than four, and a delta shift chosen **per group** rather than per sub-block. That is where its extra 0.1875 bits go. ### Shared with IQ1_S rather than copied IQ1_M searches exactly as IQ1_S does: the same 2048-entry grid, the same ±1/8 delta, every (shift, local scale) choice for every 8-value vector. It differs only in how it selects among those choices afterwards. So the search moves out of IQ1_S's encoder into `_search_shifted_grid`, which both call, and `iq1_m.py` keeps only its selection and packing. **IQ1_S's encoded bytes are unchanged**, checked by hashing its output before and after on a fixed input. ### A scale-anchor correction IQ1_M anchors its scale differently from IQ1_S: the ratio **rises with a block's peak-to-RMS** rather than being flat, and clamps higher. It uses `clamp(0.58 + 0.035 * peak_to_rms, 0.65, 0.95)` against IQ1_S's flat `0.61`. Measured over 15 Qwen3.8-27B MLP weights: | | flat 0.61 | correct anchor | | |---|---|---|---| | relative reconstruction MSE | 0.17372 | **0.17291** | **−0.47%** | It is consistent on every tensor, with no outliers. The anchor changes quality without touching layout, so neither round-trip nor conformance tests would catch it drifting. `test_scale_anchor_follows_peak_to_rms` now pins it, for all five formats; see Testing. ### Family parity Two surface asymmetries close here, so the five are uniform. `IQ1_S` now exposes `_predict_iq1_s_scales` like the other four, instead of computing its anchor inline. `IQ1_M` exposes `iq1_m_grid`, aliasing the IQ1_S table it shares. ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped:** ``` IQ1_M: 25 tensors, 4,730,880 blocks → 0 mismatched, max|diff| 0.0 ``` This mattered: **my first IQ1_M decoder had a real bug.** A `repeat_interleave` on the wrong axis produced `[h0,h1,h0,h1]` where llama.cpp needs `[h0,h0,h1,h1]`. A round-trip against our own encoder still passed, because the encoder made the matching mistake. Only comparison against bytes we did not produce caught it. Blocks from that checkpoint ship as conformance vectors, and mutation testing confirms they catch a mis-set scale nibble. The decoder unpacks every field in one vectorized pass, since fake quant decodes on every forward: 5.2 ms for a 5632×2048 weight (IQ1_S: 3.3). - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **153 passed**, 15 of them IQ1_M codec cases, including the llama.cpp conformance check - `test_scale_anchor_follows_peak_to_rms` pins every format's scale anchor. It predicts scales for blocks whose peak-to-RMS is exactly 1, 4, 8 and 16, reaching both clamps and two points on each slope, and compares them against anchors written out in the test. Mutations each fail exactly the mutated format: reverting IQ1_M to IQ1_S's flat 0.61, moving either IQ1_M clamp, changing its taper by 0.001, moving an IQ2_S or IQ2_XS clamp, and changing IQ1_S's anchor to 0.62. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed**. IQ1_S's CUDA-vs-PyTorch parity still holds after its encoder refactor. - IQ1_S and IQ1_M PyTorch encoder output and IQ1_M decoder output hash identically to the pre-split version of this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ IQ1_M adds no codebook; it reuses the IQ1_S table already carried in `codebooks.py`. The new conformance vectors come from `unsloth/Qwen3.8-27B-GGUF`, which is Apache-2.0 like its base model `Qwen/Qwen3.8-27B`; the vectors' docstring now records that. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A. Nothing is user-reachable yet; #2595 carries the entry. - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration), all merged → **this** → #2595 (IQ1_M CUDA encoder and registration). 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M quantization and dequantization for compact, GGML-compatible blocks of 256 values. * Added access to the IQ1_M grid and configurable chunk sizes for processing data. * **Bug Fixes** * Improved IQ1_S scale prediction and grid-search organization while preserving its encoding behavior. * **Tests** * Added IQ1_M conformance data and included the format in shared IQ-format test coverage. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
ad9ea97a4b
commit
e5b63320ab
@@ -0,0 +1,219 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""IQ1_M fake quantization and GGML-compatible block packing.
|
||||
|
||||
The encoder performs a single-pass squared-error grid search at a fixed,
|
||||
empirically anchored super-block scale, mirroring :mod:`.iq1_s`. Every 256
|
||||
logical values become one 56-byte block_iq1_m payload:
|
||||
|
||||
* bytes 0..31: 32 low bytes of the grid index, four per sub-block
|
||||
* bytes 32..47: 16 bytes holding, per nibble, three grid-index high bits and
|
||||
one delta-shift bit -- two groups per byte
|
||||
* bytes 48..55: four little-endian uint16 holding four 3-bit local scales each
|
||||
in bits 0..11, and one nibble of the FP16 super-block scale in bits 12..15
|
||||
|
||||
IQ1_M has no dedicated ``d`` field: the FP16 super-block scale is reassembled
|
||||
from the top nibble of each of the four scale words. It is also finer-grained
|
||||
than IQ1_S -- a local scale covers two groups rather than four, and the delta
|
||||
shift is chosen per group rather than per sub-block -- which is where its extra
|
||||
0.1875 bits per weight go.
|
||||
|
||||
The grid is the same canonical 2048 x 8 ternary table IQ1_S uses, carried from
|
||||
llama.cpp ggml-common.h revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
|
||||
The matching dequantization formula is in ggml-quants.c at the same revision:
|
||||
https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c#L2573-L2620
|
||||
"""
|
||||
|
||||
import torch
|
||||
|
||||
from .common import (
|
||||
GGML_BLOCK_SIZE,
|
||||
narrow_to_float32,
|
||||
validate_block_chunk_size,
|
||||
validate_packed_weights,
|
||||
validate_weight,
|
||||
)
|
||||
from .iq1_s import _search_shifted_grid, iq1_s_grid
|
||||
|
||||
__all__ = [
|
||||
"IQ1_M_BLOCK_BYTES",
|
||||
"IQ1_M_BLOCK_SIZE",
|
||||
"IQ1_M_EFFECTIVE_BITS",
|
||||
"dequantize_iq1_m",
|
||||
"iq1_m_grid",
|
||||
"quantize_iq1_m",
|
||||
]
|
||||
|
||||
IQ1_M_BLOCK_SIZE = GGML_BLOCK_SIZE
|
||||
IQ1_M_BLOCK_BYTES = 56
|
||||
IQ1_M_EFFECTIVE_BITS = IQ1_M_BLOCK_BYTES * 8 / IQ1_M_BLOCK_SIZE
|
||||
_IQ1_M_DELTA = 0.125
|
||||
# Largest representable magnitude: (1 + delta) at local scale 7 -> 15 * 1.125.
|
||||
_IQ1_M_NATIVE_MAX = 16.875
|
||||
# IQ1_M anchors differently from IQ1_S: the ratio rises with a block's peak-to-RMS
|
||||
# instead of being flat, and it is allowed closer to full range. Values follow the
|
||||
# reference predictor these encoders are derived from.
|
||||
_IQ1_M_SCALE_ANCHOR_BASE = 0.58
|
||||
_IQ1_M_SCALE_ANCHOR_MIN = 0.65
|
||||
_IQ1_M_SCALE_ANCHOR_MAX = 0.95
|
||||
_IQ1_M_PEAK_TO_RMS_TAPER = 0.035
|
||||
_IQ1_M_GROUPS = 32
|
||||
_IQ1_M_SUBBLOCKS = 8
|
||||
_DEFAULT_BLOCK_CHUNK_SIZE = 1024
|
||||
_DEFAULT_DECODE_CHUNK_SIZE = 4096
|
||||
|
||||
|
||||
def iq1_m_grid(device: torch.device | str | None = None) -> torch.Tensor:
|
||||
"""Return the canonical IQ1_M ternary grid as float32.
|
||||
|
||||
IQ1_M indexes the same 2048-entry table as IQ1_S; this alias exists so every format
|
||||
exposes a grid accessor under its own name.
|
||||
"""
|
||||
return iq1_s_grid(device)
|
||||
|
||||
|
||||
def _predict_iq1_m_scales(blocks: torch.Tensor) -> torch.Tensor:
|
||||
"""Predict one FP16 super-block scale for each flattened block."""
|
||||
x = narrow_to_float32(blocks)
|
||||
amax = x.abs().amax(dim=1)
|
||||
rms = x.square().mean(dim=1).sqrt()
|
||||
peak_to_rms = torch.where(rms > 0, amax / rms, torch.zeros_like(rms))
|
||||
anchor_ratio = (_IQ1_M_SCALE_ANCHOR_BASE + _IQ1_M_PEAK_TO_RMS_TAPER * peak_to_rms).clamp(
|
||||
_IQ1_M_SCALE_ANCHOR_MIN, _IQ1_M_SCALE_ANCHOR_MAX
|
||||
)
|
||||
return ((amax / _IQ1_M_NATIVE_MAX) * anchor_ratio).clamp(max=65504.0).to(torch.float16)
|
||||
|
||||
|
||||
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
"""Encode a moderate-size batch of flattened 256-value blocks."""
|
||||
x = narrow_to_float32(blocks)
|
||||
block_count = x.shape[0]
|
||||
d = _predict_iq1_m_scales(x)
|
||||
d_float = d.float()
|
||||
best_error, best_entry = _search_shifted_grid(
|
||||
x.reshape(block_count, _IQ1_M_GROUPS, 8), d_float, grid
|
||||
)
|
||||
|
||||
# Each group picks its own shift; only the 3-bit local scale is shared, over two groups.
|
||||
per_shift_error = best_error.reshape(block_count, _IQ1_M_GROUPS, 2, 8)
|
||||
per_shift_entry = best_entry.reshape(block_count, _IQ1_M_GROUPS, 2, 8)
|
||||
local_error, shift_index = per_shift_error.min(dim=2)
|
||||
local_entry = per_shift_entry.gather(2, shift_index.unsqueeze(2)).squeeze(2)
|
||||
|
||||
pair_error = local_error.reshape(block_count, 16, 2, 8).sum(dim=2)
|
||||
selected_local = pair_error.argmin(dim=-1)
|
||||
group_local = selected_local.repeat_interleave(2, dim=1)
|
||||
selected_entry = local_entry.gather(2, group_local.unsqueeze(-1)).squeeze(-1)
|
||||
selected_shift = shift_index.gather(2, group_local.unsqueeze(-1)).squeeze(-1)
|
||||
|
||||
packed = torch.empty((block_count, IQ1_M_BLOCK_BYTES), dtype=torch.uint8, device=x.device)
|
||||
packed[:, :32] = (selected_entry & 0xFF).to(torch.uint8)
|
||||
|
||||
# Each group's nibble is three index-high bits and its shift bit; two groups share a byte,
|
||||
# low nibble first.
|
||||
nibbles = ((selected_entry >> 8) & 0x7) | (selected_shift << 3)
|
||||
packed[:, 32:48] = (nibbles[:, 0::2] | (nibbles[:, 1::2] << 4)).to(torch.uint8)
|
||||
|
||||
# Four scale words, each carrying four 3-bit local scales (two sub-blocks, two halves each)
|
||||
# in bits 0..11 and one nibble of d in bits 12..15, low nibble in the first word. The fields
|
||||
# are disjoint, so summing them is the same as OR-ing them.
|
||||
d_bits = d.contiguous().view(torch.int16).to(torch.int64) & 0xFFFF
|
||||
local_shifts = torch.tensor([0, 3, 6, 9], dtype=torch.int64, device=x.device)
|
||||
nibble_shifts = torch.tensor([0, 4, 8, 12], dtype=torch.int64, device=x.device)
|
||||
words = (selected_local.reshape(block_count, 4, 4) << local_shifts).sum(dim=-1)
|
||||
words |= ((d_bits.unsqueeze(-1) >> nibble_shifts) & 0xF) << 12
|
||||
packed[:, 48:56:2] = (words & 0xFF).to(torch.uint8)
|
||||
packed[:, 49:56:2] = ((words >> 8) & 0xFF).to(torch.uint8)
|
||||
return torch.where((d_float == 0).unsqueeze(1), 0, packed)
|
||||
|
||||
|
||||
@torch.no_grad()
|
||||
def quantize_iq1_m(
|
||||
weight: torch.Tensor, *, block_chunk_size: int = _DEFAULT_BLOCK_CHUNK_SIZE
|
||||
) -> tuple[torch.Tensor, torch.Tensor]:
|
||||
"""Pack a floating-point weight into GGML-compatible IQ1_M blocks.
|
||||
|
||||
Returned shapes are ``[*weight.shape[:-1], weight.shape[-1] // 256, 56]``
|
||||
and ``[weight.ndim]``.
|
||||
"""
|
||||
validate_weight(weight, "IQ1_M")
|
||||
validate_block_chunk_size(block_chunk_size)
|
||||
|
||||
logical_shape = torch.tensor(weight.shape, dtype=torch.int64)
|
||||
blocks = weight.contiguous().reshape(-1, IQ1_M_BLOCK_SIZE)
|
||||
grid = iq1_s_grid(weight.device)
|
||||
packed_shape = (*weight.shape[:-1], weight.shape[-1] // IQ1_M_BLOCK_SIZE, IQ1_M_BLOCK_BYTES)
|
||||
chunks = [
|
||||
_encode_blocks(blocks[start : start + block_chunk_size], grid)
|
||||
for start in range(0, blocks.shape[0], block_chunk_size)
|
||||
]
|
||||
return torch.cat(chunks).reshape(packed_shape), logical_shape
|
||||
|
||||
|
||||
@torch.no_grad()
|
||||
def dequantize_iq1_m(
|
||||
packed_weights: torch.Tensor,
|
||||
weight_shape: torch.Tensor,
|
||||
*,
|
||||
dtype: torch.dtype = torch.bfloat16,
|
||||
block_chunk_size: int = _DEFAULT_DECODE_CHUNK_SIZE,
|
||||
) -> torch.Tensor:
|
||||
"""Decode GGML-compatible IQ1_M payload bytes."""
|
||||
shape = validate_packed_weights(
|
||||
packed_weights, weight_shape, block_bytes=IQ1_M_BLOCK_BYTES, format_name="IQ1_M"
|
||||
)
|
||||
validate_block_chunk_size(block_chunk_size)
|
||||
|
||||
blocks = packed_weights.contiguous().reshape(-1, IQ1_M_BLOCK_BYTES)
|
||||
grid = iq1_s_grid(blocks.device)
|
||||
scale_shifts = torch.tensor([0, 3, 6, 9], dtype=torch.int64, device=blocks.device)
|
||||
decoded = torch.empty((blocks.shape[0], IQ1_M_BLOCK_SIZE), dtype=dtype, device=blocks.device)
|
||||
for start in range(0, blocks.shape[0], block_chunk_size):
|
||||
stop = min(start + block_chunk_size, blocks.shape[0])
|
||||
block_chunk = blocks[start:stop]
|
||||
count = block_chunk.shape[0]
|
||||
low = block_chunk[:, :32].to(torch.int64).reshape(count, _IQ1_M_SUBBLOCKS, 4)
|
||||
qh = block_chunk[:, 32:48].to(torch.int64).reshape(count, _IQ1_M_SUBBLOCKS, 2)
|
||||
words = block_chunk[:, 48:56:2].to(torch.int64) | (
|
||||
block_chunk[:, 49:56:2].to(torch.int64) << 8
|
||||
)
|
||||
|
||||
# The FP16 super-block scale is the four words' top nibbles, low word first.
|
||||
d_bits = (
|
||||
(words[:, 0] >> 12)
|
||||
| ((words[:, 1] >> 8) & 0x00F0)
|
||||
| ((words[:, 2] >> 4) & 0x0F00)
|
||||
| (words[:, 3] & 0xF000)
|
||||
)
|
||||
d = d_bits.to(torch.int16).view(torch.float16).float()
|
||||
|
||||
# Each qh byte carries two groups, low nibble first: three high index bits and a delta sign.
|
||||
nibbles = torch.stack((qh & 0xF, qh >> 4), dim=-1).reshape(count, _IQ1_M_SUBBLOCKS, 4)
|
||||
entries = low | ((nibbles & 0x7) << 8)
|
||||
deltas = torch.where((nibbles & 0x8).bool(), -_IQ1_M_DELTA, _IQ1_M_DELTA)
|
||||
|
||||
# Each word packs four 3-bit local scales below its nibble of d: two sub-blocks, two
|
||||
# halves each, so word w, slot k is sub-block 2 * w + k // 2, half k % 2.
|
||||
local = ((words.unsqueeze(-1) >> scale_shifts) & 0x7).reshape(count, _IQ1_M_SUBBLOCKS, 2)
|
||||
scales = d.unsqueeze(-1).unsqueeze(-1) * (2 * local + 1).float()
|
||||
|
||||
values = grid[entries] + deltas.unsqueeze(-1)
|
||||
# Groups 0,1 take the first local scale and groups 2,3 the second, so the repeat
|
||||
# runs along the half axis: [h0, h0, h1, h1], not [h0, h1, h0, h1].
|
||||
group_scale = scales.repeat_interleave(2, dim=2)
|
||||
chunk_decoded = values * group_scale.unsqueeze(-1)
|
||||
decoded[start:stop] = chunk_decoded.reshape(-1, IQ1_M_BLOCK_SIZE)
|
||||
return decoded.reshape(shape)
|
||||
@@ -87,22 +87,34 @@ def iq1_s_grid(device: torch.device | str | None = None) -> torch.Tensor:
|
||||
return _GRID_CACHE[resolved_device]
|
||||
|
||||
|
||||
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
"""Encode a moderate-size batch of flattened 256-value blocks."""
|
||||
def _predict_iq1_s_scales(blocks: torch.Tensor) -> torch.Tensor:
|
||||
"""Predict one FP16 super-block scale for each flattened block.
|
||||
|
||||
The fixed-scale search favors a compressed super-block scale, so this empirical anchor
|
||||
puts d below the full-range value. The CUDA encoder computes the same quantity in its
|
||||
own scale kernel; this is the reference the torch path uses.
|
||||
"""
|
||||
x = narrow_to_float32(blocks)
|
||||
block_count = x.shape[0]
|
||||
vectors = x.reshape(block_count, 32, 8)
|
||||
amax = x.abs().amax(dim=1)
|
||||
return ((amax / _IQ1_S_NATIVE_MAX) * _IQ1_S_SCALE_ANCHOR).clamp(max=65504.0).to(torch.float16)
|
||||
|
||||
|
||||
def _search_shifted_grid(
|
||||
vectors: torch.Tensor, d: torch.Tensor, grid: torch.Tensor
|
||||
) -> tuple[torch.Tensor, torch.Tensor]:
|
||||
"""Score every grid entry against each 8-value vector at every shift and local scale.
|
||||
|
||||
Returns the lowest error and the entry reaching it, both ``[blocks, 32, 16]`` and indexed
|
||||
by choice ``shift * 8 + local``. IQ1_M runs the same search over the same grid and delta;
|
||||
the two formats differ only in how they select among the choices afterwards.
|
||||
"""
|
||||
block_count = vectors.shape[0]
|
||||
xnorm = vectors.square().sum(dim=-1)
|
||||
xsum = vectors.sum(dim=-1)
|
||||
|
||||
amax = x.abs().amax(dim=1)
|
||||
# The fixed-scale search favors a compressed super-block scale. This
|
||||
# empirical anchor initializes d below the full-range value.
|
||||
d = ((amax / _IQ1_S_NATIVE_MAX) * _IQ1_S_SCALE_ANCHOR).clamp(max=65504.0).to(torch.float16)
|
||||
d_float = d.float()
|
||||
|
||||
best_error = torch.full((block_count, 32, 16), torch.inf, dtype=torch.float32, device=x.device)
|
||||
best_entry = torch.zeros((block_count, 32, 16), dtype=torch.int64, device=x.device)
|
||||
best_error = torch.full(
|
||||
(block_count, 32, 16), torch.inf, dtype=torch.float32, device=vectors.device
|
||||
)
|
||||
best_entry = torch.zeros((block_count, 32, 16), dtype=torch.int64, device=vectors.device)
|
||||
grid_norm = grid.square().sum(dim=-1)
|
||||
grid_sum = grid.sum(dim=-1)
|
||||
|
||||
@@ -120,7 +132,7 @@ def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
shifted_norm = tile_norm + 2 * delta * tile_sum + 8 * delta * delta
|
||||
for local in range(8):
|
||||
choice = shift * 8 + local
|
||||
scale = d_float.reshape(-1, 1, 1) * (2 * local + 1)
|
||||
scale = d.reshape(-1, 1, 1) * (2 * local + 1)
|
||||
error = (
|
||||
xnorm.unsqueeze(-1) - 2 * scale * shifted_dot + scale.square() * shifted_norm
|
||||
).clamp_min_(0)
|
||||
@@ -132,6 +144,16 @@ def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
best_entry[:, :, choice] = torch.where(
|
||||
replace, tile_index + entry_start, best_entry[:, :, choice]
|
||||
)
|
||||
return best_error, best_entry
|
||||
|
||||
|
||||
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
"""Encode a moderate-size batch of flattened 256-value blocks."""
|
||||
x = narrow_to_float32(blocks)
|
||||
block_count = x.shape[0]
|
||||
d = _predict_iq1_s_scales(x)
|
||||
d_float = d.float()
|
||||
best_error, best_entry = _search_shifted_grid(x.reshape(block_count, 32, 8), d_float, grid)
|
||||
|
||||
group_error = best_error.reshape(block_count, 8, 4, 16).sum(dim=2)
|
||||
selected_choice = group_error.argmin(dim=-1)
|
||||
|
||||
@@ -18,7 +18,8 @@
|
||||
Each entry holds packed block bytes lifted verbatim from
|
||||
unsloth/Qwen3.8-27B-GGUF (Qwen3.8-27B-UD-IQ1_S.gguf) together with the values
|
||||
llama.cpp's own dequantize_row_* produces for them, computed from
|
||||
ggml-quants.c revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
|
||||
ggml-quants.c revision 9b05354ec6fb58b4e665e9a39ebc40285c015638. The checkpoint
|
||||
is published under Apache-2.0, as is Qwen/Qwen3.8-27B, the model it quantizes.
|
||||
|
||||
These pin our decoders against bytes we did not produce. A decoder that drifts
|
||||
from the GGML layout -- a mis-set high bit, a swapped scale nibble, a sign
|
||||
@@ -62,6 +63,40 @@ _VECTORS = {
|
||||
"w85GPeSrj5LVBGpy6iev/eS19nOE/x+gaS2i"
|
||||
),
|
||||
},
|
||||
"iq1_m": {
|
||||
"source": "blk.1.attn_gate.weight",
|
||||
"block_bytes": 56,
|
||||
"blocks": (
|
||||
"eNoBUAGv/sgI5+PIe/BwwioFcBUzxw3iyphnqNoBMNQxtkYxKIg/RpRL53G82jgiUKrRvh5d/dzUmweXtBQV+zXsYEc4"
|
||||
"BhC3bF7wWvQFzy2AI93zFgUGNb4wPvSYWoJISJPERQmGNC8t1ESzwwUl0VqLjpLMmxRBBIezwirprpBVD7/r/Ig7oFIl"
|
||||
"9J2LtXeKAZqOhjjGIDxTHmLssDKbvAWf9um/y85ttT0H26bzHeDrAJbuisGPnSGlUT2c0uWVv/620cLVhOoO+ztPFia9"
|
||||
"WPgjYsisuSabXH64WhO5WhRt5Z1kl3MXjvFr3vd3GQWsYMSRdfH1s7aUT2yDlAT+RMHIacwBCG0GIkBFwJylQDLB9wbg"
|
||||
"zU483Lojdx+p3BogUaPTo7/iBfE6P39ha5fQr7yx4Uj2l/ngDgV3qsfCM5LfckI5HeHRM+6Q5AU8LF+SvtLEksaZEln9"
|
||||
"qfk="
|
||||
),
|
||||
"expected": (
|
||||
"eNqVV01oXkUUfWBAFNIWoVKhFNFVXViRuDAzDz8QhWjpQlcWVIqCRiy4aFy0FF+kRBdStBR3FgtCNkJWEsnMK6mgDdZV"
|
||||
"LbRQq9BFEd3Unyq6EOfOe2e+M/fNl+ricO49d+bO3Dv3fSTVxaOmued7L6iCzRid21IzF/H6E7LfRj7zqeSxYNECXPBt"
|
||||
"xnJesOOZwYfGnPSQq89jYFNuU5170wuaT05ZMCNotpqadsGWmIctHGM6t9yd66Dz01rU27H7r0DNWY3PrEaM7L4WtmD9"
|
||||
"lXtr0YQF0DgebdU7zt9z6c6p3mrp2yJG2x9ttbb+x8EaeopT//kNoCdMTRtwWisavUfk8Cb8VvrtWOvZImfi/Ly0B7OQ"
|
||||
"2RePutiTwBpaD99BCx22zp/dU86gHlDNBvfjM0o2vkv9DcIf9KY70+pas5lHTPjG5bXqQm2rLYf8wBa+UPvANgL2jcuz"
|
||||
"CXtCTRszLsOB8z7xgfM2YU+4g0DsjRkTseNoQvN8qAm8cNX3uhMErYXNWrP6Yl1995XZBK6pHqwFYg9A5yfI2QF8D9wt"
|
||||
"6Vg77pNNLBr3T/rEvWMNfUBvANG6vnW9lD6j19zfD2+3/wfNkwfbTENufj/4wngr8PCOfvD+nIP3oVbtw9a1d3FHs+QH"
|
||||
"+ad+Npth9PCxerN4aeZK85fmTc/heJ2ZtD/bs3DVZvv5W9D90d8MfNaxvrSWY3oN9v56zDGa5R+t0mSNCboXHr2/vWZf"
|
||||
"xyfZWhNILj4v+G3hLtkefWbz5zu2+vyXWXDC4rM+cUCKQ8eeQt2FHiQdkLvGdU99ZgSj356uhdc/3tUqiO4EYU0rgC9o"
|
||||
"2t9ts/+Mjwh29d5zRiB+iTWa6ZO2mtntYjzwwA5x9hnxTIlTDvaLe/o4cgaWtUYDeljr4Q+0jt3Ee/fn0FkeMbqv5zMH"
|
||||
"vtqfcqo6s/MKPeVzOVe16yM3OnxHvX72oVYgPjQBfEa2rp8Hngs9I6zxPEXQWZwX+i3vs2NuDHPND9hcC79XcybyzRUf"
|
||||
"bYHYor1wxN8Sbz9uI974YS36wtA4L+cHSnHh3m5eur+udv/jI9sPjNhF7m2sFz/aGwuuGd3ZMldbL1mxox9saMlm6N7d"
|
||||
"XLER0GEjzmtEO7nsJqFZsm3mn76r1nqquwDEimvQF31PfVeOd/PQgX3Yug+dbwfviveTmO4no+u5j9hYMOE9akFmv7rT"
|
||||
"xVzCV961KTd00QDWhQHEWRvnFDYRUhN8sQOneeI5028wIRZ1XS/5xfnr/TSruLeuYVyvSfOGOqQGgWh7501z3zcevD57"
|
||||
"Ww1fbO3r9dWhR4wg2Bac8OUpn7TeTuspxlq2j3xA76lOHDfNW395ZsHo5fD/cW9vCnV3vm8PmxDqjQh2WOci9s670Jc2"
|
||||
"7u1t9oWhiQ+N1saeMtD33nd6H/ZGm3u5Sf/026T+vfZFwvrOB1oBbObR6v6a16V9/Kacd3y+nO3QM90/njmeN+rBoGfw"
|
||||
"e20wv3qWucc8z3qu+XyA+89vmO6EOaDa9JzoerP56ePo00CnuS/ixHEX5t4KlyDfhHwL+B7E5++kOrsc/h5+zESshN8W"
|
||||
"truYC2wzBlYu+YwZslYgMWbY8HEeny/n4l6wEWMWfL21iOb0kXZSLIOuS9953AdT6I1JtaAHw16ZDFxLsJu/766rKz/5"
|
||||
"iJklE/3AEZ0e/uZcchHBDvFWAH+QX/eR6+L6xjXaibXxnfW7jJHn51nQ81Ly8Y4MnIP7c2/1+/R9495lPUSfqIeZPb/Y"
|
||||
"Ydv12Yh9cz6xjkFnINckhPvpN+Y7F3uSa+M34bcBzy/6iG3X18J9bATbXAPX0tv/ArwdXdM="
|
||||
),
|
||||
},
|
||||
"iq2_xxs": {
|
||||
"source": "blk.0.attn_gate.weight",
|
||||
"block_bytes": 66,
|
||||
|
||||
@@ -29,6 +29,7 @@ from _test_utils.torch.quantization.iq_llama_cpp_vectors import (
|
||||
packed_blocks,
|
||||
)
|
||||
|
||||
import modelopt.torch.quantization.ggml.iq1_m as iq1_m_module
|
||||
import modelopt.torch.quantization.ggml.iq1_s as iq1_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_s as iq2_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xs as iq2_xs_module
|
||||
@@ -40,6 +41,7 @@ from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
# name -> (module, packed bytes per block, codebook entries, bits per weight)
|
||||
FORMATS = {
|
||||
"iq1_s": (iq1_s_module, 50, 2048, 1.5625),
|
||||
"iq1_m": (iq1_m_module, 56, 2048, 1.75),
|
||||
"iq2_xxs": (iq2_xxs_module, 66, 256, 2.0625),
|
||||
"iq2_xs": (iq2_xs_module, 74, 512, 2.3125),
|
||||
"iq2_s": (iq2_s_module, 82, 1024, 2.5625),
|
||||
@@ -49,7 +51,18 @@ NAMES = sorted(FORMATS)
|
||||
# tests that go through TensorQuantizer iterate these rather than every codec above.
|
||||
DISPATCHED = sorted(IQ_FORMAT_REGISTRY)
|
||||
# IQ1 grids are ternary; IQ2 grids hold the magnitudes 8, 25 and 43.
|
||||
TERNARY = {"iq1_s"}
|
||||
TERNARY = {"iq1_s", "iq1_m"}
|
||||
# name -> (largest representable magnitude, scale anchor at peak-to-RMS 1, 4, 8 and 16), where the
|
||||
# anchor is d * native max / amax. IQ1_S is flat; IQ1_M rises with peak-to-RMS and the IQ2 formats
|
||||
# fall with it, each clamped at both ends. Written out rather than read from the modules, so a
|
||||
# changed constant fails here.
|
||||
ANCHORS = {
|
||||
"iq1_s": (15 * 1.125, (0.61, 0.61, 0.61, 0.61)),
|
||||
"iq1_m": (15 * 1.125, (0.65, 0.72, 0.86, 0.95)),
|
||||
"iq2_xxs": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
|
||||
"iq2_xs": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
|
||||
"iq2_s": (43 * 31 / 8, (0.92, 0.86, 0.72, 0.65)),
|
||||
}
|
||||
|
||||
|
||||
def _parts(name):
|
||||
@@ -83,9 +96,11 @@ def test_canonical_grid(name):
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_grid_normalizes_unindexed_cuda_device(monkeypatch, name):
|
||||
module, _, _, grid_fn, _, _, _ = _parts(name)
|
||||
# IQ1_M shares the IQ1_S cache, being the same table.
|
||||
cache_owner = iq1_s_module if name == "iq1_m" else module
|
||||
cached = torch.empty(0)
|
||||
monkeypatch.setattr(torch.cuda, "current_device", lambda: 7)
|
||||
monkeypatch.setitem(module._GRID_CACHE, torch.device("cuda", 7), cached)
|
||||
monkeypatch.setitem(cache_owner._GRID_CACHE, torch.device("cuda", 7), cached)
|
||||
|
||||
assert grid_fn("cuda") is cached
|
||||
|
||||
@@ -240,6 +255,25 @@ def test_search_is_independent_of_default_dtype(name):
|
||||
torch.set_default_dtype(torch.float32)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_scale_anchor_follows_peak_to_rms(name):
|
||||
"""The predicted block scale must follow the format's anchor, through both clamps.
|
||||
|
||||
The anchor is an encoder choice: it changes quality without touching the layout, so no
|
||||
round-trip or conformance test would notice it drifting. k equal unit spikes among 256 zeros
|
||||
have a peak-to-RMS of exactly 16 / sqrt(k).
|
||||
"""
|
||||
native_max, anchors = ANCHORS[name]
|
||||
blocks = torch.zeros(4, 256)
|
||||
for row, spikes in enumerate((256, 16, 4, 1)):
|
||||
blocks[row, :spikes] = 1.0
|
||||
d = getattr(FORMATS[name][0], f"_predict_{name}_scales")(blocks)
|
||||
|
||||
assert d.dtype == torch.float16
|
||||
# rtol covers the FP16 rounding of d, under 0.1%; a flat 0.61 anchor misses IQ1_M by 6% or more.
|
||||
torch.testing.assert_close(d.float() * native_max, torch.tensor(anchors), rtol=2**-10, atol=0)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", formats())
|
||||
def test_decoder_matches_llama_cpp_on_captured_blocks(name):
|
||||
"""Decode bytes we did not produce and match llama.cpp's own output exactly.
|
||||
|
||||
Reference in New Issue
Block a user