mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add the IQ2_XXS weight-only quantization format (#2511)
### What does this PR do? Type of change: new feature llama.cpp defines five GGML IQ formats at one and two bits; we ship two. This adds **IQ2_XXS** at 2.0625 bits per weight, between IQ1_S and IQ2_XS, and is the **first of three**. On a real mixed-precision checkpoint (`unsloth/Qwen3.8-27B-GGUF`, `Qwen3.8-27B-UD-IQ1_S.gguf`) IQ2_XXS alone covers **59 tensors and 2.84 B parameters — 10.6% of the file**, which a reader limited to IQ1_S/IQ2_XS cannot consume. Across all three PRs the missing formats account for 17.3%. | format | bpw | bytes/256 | codebook | | |---|---|---|---|---| | `iq1_s` | 1.5625 | 50 | `iq1s_grid` (2048) | existing | | **`iq2_xxs`** | **2.0625** | **66** | **`iq2xxs_grid` (256)** | **this PR** | | `iq2_xs` | 2.3125 | 74 | `iq2xs_grid` (512) | existing | The encoder follows the existing single-pass grid search at a fixed anchored super-block scale, and the CUDA kernel the existing per-block structure. IQ2_XXS reuses IQ2_XS's even-parity sign rule but packs a 4-bit sub-block scale into the same 32-bit word as four 7-bit sign indices, and its 256-entry grid needs no high index bits. ### Groundwork the next two reuse Two things land here because IQ2_XXS is the first format to need them: - **Export registry.** The IQ family was spelled as a two-element tuple at **nine** sites across `quant_utils.py`, `unified_export_hf.py` and `unified_export_megatron.py`. Those become an `IQ_FORMATS` frozenset plus per-format packer and block-geometry tables, so a format is a row rather than a sweep through the exporters. - **Shared test contract.** The per-format test files had drifted apart — each of `iq1_s` and `iq2_xs` tested things the other did not. They become one parametrized module per layer (unit and CUDA), so every format is held to the same contract and a new one inherits it. ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_xxs ``` ### Testing **The decoder is validated against llama.cpp's own output, not just round-tripped.** Every IQ2_XXS tensor in the checkpoint above, compared against `dequantize_row_iq2_xxs` from `ggml-quants.c`: ``` IQ2_XXS: 59 tensors, 11,100,160 blocks → 0 mismatched, max|diff| 0.0 ``` The new codebook matches the `ggml-common.h` table entry for entry, as does the `ksigns_iq2xs` sign table. Blocks lifted from that checkpoint ship as conformance vectors so CI keeps checking bytes we did not produce; mutation testing confirms they catch a wrong sign-field width. The CUDA encoder is byte-identical to the PyTorch reference on a fixed input and runs at **1047.9 M elem/s against the torch search's 10.7** on a 5632×2048 weight. - `tests/unit/torch/quantization/ -k 'ggml or iq1 or iq2 or iq_'` — 99 passed - `tests/gpu/torch/quantization/test_iq_formats_cuda.py` — 21 passed (7 checks × 3 formats) - `tests/unit/recipe/test_presets.py` — passing; `general/ptq` now holds 29 recipes, `ptq.md` updated - reconstruction error decreases monotonically with bit width, pinned by a test Pre-existing failures in `tests/unit/torch/export/` and `test_autoquant.py` are `transformers`/`torchvision` import problems in my environment — identical counts with and without this change. ### A finding about already-merged code Checking the new kernel against its PyTorch reference at 4096 blocks showed that **CUDA and torch encoders disagree on roughly 1 block in 6000 — including the already-merged `iq2_xs`**, at 0.0163% against IQ2_XXS's 0.0000%. Root cause: both compute `xnorm − 2·scale·dot + scale²·qnorm`, but CUDA fuses it with `fmaf` while torch uses separate ops; where two local scales fall within a float32 ULP the roundings pick different sides. Adjudicated against float64, neither path is better (5 to 6). Worst-case cost is **1.48e-08** relative reconstruction error, and run-to-run determinism on a given device holds. This is pre-existing, not introduced here — `test_iq2_xs_cuda.py` asserts exact byte parity but on a 16-block weight where ties essentially never arise. I have **not** changed that test; rewording a guarantee on merged code belongs in its own change. The new shared GPU tests assert exact parity on a small fixed input and compare reconstruction error at scale. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — the new codebook is a GGML table, carried in `codebooks.py` beside the existing ones so the MIT-licensed surface stays in that one file, with the source revision recorded. No new dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ — not yet run ### Additional Information First of three; **IQ2_S** and **IQ1_M** follow and build on this branch. Replaces #2505, which carried all three at once. Follows #2446 / #2447 / #2448 / #2449, which landed IQ1_S and IQ2_XS. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added IQ2_XXS weight-only quantization, including CUDA acceleration and support for Hugging Face and Megatron exports. - Added the `general/ptq/iq2_xxs` recipe. It requires no calibration data and supports eligible layers with a weight dimension divisible by 256. - Updated the PTQ recipe catalog to list IQ1_S, IQ2_XXS, and IQ2_XS at approximately 1.56, 2.06, and 2.31 bits per weight, respectively. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
f2f0d6958e
commit
a21411adde
@@ -14,6 +14,7 @@ Changelog
|
||||
*Quantization*
|
||||
|
||||
- Add IQ1_S and IQ2_XS weight-only quantization with GGML-compatible 256-value block encoders, built-in ``iq1_s`` / ``iq2_xs`` PTQ recipes, and unified HF and Megatron export of the packed blocks. Quantized weights must have a final dimension divisible by 256, and Megatron export requires tensor and pipeline parallel sizes of 1.
|
||||
- Add ``iq2_xxs`` weight-only quantization with a CUDA encoder and a ``general/ptq`` recipe, at 2.0625 bits per weight between ``iq1_s`` and ``iq2_xs``. The same 256-value block constraint applies.
|
||||
- A recipe can now **delegate its whole body to another recipe** with a top-level ``$import``; any top-level key given alongside it overrides the imported one. ``metadata.recipe_type`` became optional along with it: a recipe states its kind with a ``# modelopt-schema:`` comment, with ``metadata.recipe_type``, or by delegating to a recipe that does, and only a recipe that another file imports has to carry the schema comment. Whatever a recipe does state must be true: a schema comment and a ``recipe_type`` must agree, and so must a recipe and the recipe it delegates to. ``modelopt_recipes/models/`` uses this for checkpoint entries that a portable recipe already reproduces: the entry aliases that recipe instead of copying it.
|
||||
- Backfill the recipes behind NVIDIA's already-published checkpoints under ``modelopt_recipes/models/``, so a released checkpoint's quantization scheme is reachable from its own model-hub path rather than only from the general tier. For example, ``moonshotai/Kimi-K2.6`` (published as ``nvidia/Kimi-K2.6-NVFP4``) and ``Qwen/Qwen3.5-397B-A17B`` (published as ``nvidia/Qwen3.5-397B-A17B-NVFP4-V2``) each alias a portable recipe wholesale -- the general expert-only NVFP4 recipe and the ``qwen3_5_moe`` architecture recipe respectively -- rather than copying its body; other checkpoints follow in separate changes.
|
||||
- Add ``layerwise.export_dir``: layerwise calibration writes each decoder layer to its own quantized checkpoint shard as it finishes, so no separate ``export_hf_checkpoint()`` pass is needed and, with ``layerwise.checkpoint_dir``, an interrupted run resumes without redoing finished layers. Calibration writes the layer shards; ``finalize()`` on the exporter left on the model adds the tail shard, the index and the config artifacts, and the checkpoint does not load until it runs. ``examples/hf_ptq`` does this for you. Supports FP8 and NVFP4 on single-process models, resident or offloaded, including multimodal models and models with MTP layers; other formats and placements raise ``NotImplementedError`` before calibration starts.
|
||||
|
||||
@@ -19,14 +19,7 @@ import warnings
|
||||
from collections import defaultdict
|
||||
from typing import Any
|
||||
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
)
|
||||
from .quant_format import IQ_BLOCK_METADATA, IQ_FORMATS
|
||||
|
||||
|
||||
def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None) -> dict[str, Any]:
|
||||
@@ -127,15 +120,8 @@ def _quant_algo_to_group_config(quant_algo: str, group_size: int | None = None)
|
||||
},
|
||||
"weights": {"dynamic": False, "num_bits": 8, "type": "float", "group_size": gs},
|
||||
}
|
||||
elif quant_algo in ("IQ1_S", "IQ2_XS"):
|
||||
if quant_algo == "IQ1_S":
|
||||
block_size = IQ1_S_BLOCK_SIZE
|
||||
payload_bytes = IQ1_S_BLOCK_BYTES
|
||||
effective_bits = IQ1_S_EFFECTIVE_BITS
|
||||
else:
|
||||
block_size = IQ2_XS_BLOCK_SIZE
|
||||
payload_bytes = IQ2_XS_BLOCK_BYTES
|
||||
effective_bits = IQ2_XS_EFFECTIVE_BITS
|
||||
elif quant_algo.lower() in IQ_FORMATS:
|
||||
block_size, payload_bytes, effective_bits = IQ_BLOCK_METADATA[quant_algo.lower()]
|
||||
if group_size not in (None, block_size):
|
||||
raise ValueError(f"{quant_algo} requires group size {block_size}, got {group_size}")
|
||||
# IQ payloads are self-contained blocks, not compressed-tensors integer groups.
|
||||
@@ -239,7 +225,7 @@ def convert_hf_quant_config_format(input_config: dict[str, Any]) -> dict[str, An
|
||||
"targets": ["Linear"],
|
||||
}
|
||||
new_config["config_groups"] = {"group_0": config_group_details}
|
||||
elif quant_algo_value in ("IQ1_S", "IQ2_XS"):
|
||||
elif str(quant_algo_value).lower() in IQ_FORMATS:
|
||||
# Forward the caller's group size so a mismatched one is rejected rather than rewritten
|
||||
# to the format's block size.
|
||||
iq_metadata = _quant_algo_to_group_config(
|
||||
|
||||
@@ -19,6 +19,21 @@ Backend-specific names live with their backend: the TensorRT-LLM checkpoint layo
|
||||
constants, for example, are in :mod:`modelopt.torch.export.trtllm.model_config`.
|
||||
"""
|
||||
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
IQ2_XXS_BLOCK_BYTES,
|
||||
IQ2_XXS_BLOCK_SIZE,
|
||||
IQ2_XXS_EFFECTIVE_BITS,
|
||||
quantize_iq1_s,
|
||||
quantize_iq2_xs,
|
||||
quantize_iq2_xxs,
|
||||
)
|
||||
|
||||
QUANTIZATION_NONE = None
|
||||
QUANTIZATION_FP8 = "fp8"
|
||||
QUANTIZATION_INT8_SQ = "int8_sq"
|
||||
@@ -37,16 +52,45 @@ QUANTIZATION_FP8_PB_REAL = "fp8_pb_real"
|
||||
QUANTIZATION_FP8_PB_WO = "fp8_pb_wo"
|
||||
QUANTIZATION_FP8_PC_PT = "fp8_pc_pt"
|
||||
QUANTIZATION_IQ1_S = "iq1_s"
|
||||
QUANTIZATION_IQ2_XXS = "iq2_xxs"
|
||||
QUANTIZATION_IQ2_XS = "iq2_xs"
|
||||
|
||||
# Every GGML IQ format. They share the weight-only, 256-value-block, per-module-scale
|
||||
# shape, so export treats them as one family; adding a format means adding it here
|
||||
# rather than extending a tuple at each use site.
|
||||
IQ_FORMATS = frozenset(
|
||||
{
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XXS,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
}
|
||||
)
|
||||
|
||||
# Block geometry per IQ format: (block size, packed bytes per block, bits per weight). Checkpoint
|
||||
# metadata spells the algorithm in upper case, so consumers look up
|
||||
# ``IQ_BLOCK_METADATA[algo.lower()]`` rather than carrying a second spelling of the family.
|
||||
IQ_BLOCK_METADATA = {
|
||||
QUANTIZATION_IQ1_S: (IQ1_S_BLOCK_SIZE, IQ1_S_BLOCK_BYTES, IQ1_S_EFFECTIVE_BITS),
|
||||
QUANTIZATION_IQ2_XXS: (IQ2_XXS_BLOCK_SIZE, IQ2_XXS_BLOCK_BYTES, IQ2_XXS_EFFECTIVE_BITS),
|
||||
QUANTIZATION_IQ2_XS: (IQ2_XS_BLOCK_SIZE, IQ2_XS_BLOCK_BYTES, IQ2_XS_EFFECTIVE_BITS),
|
||||
}
|
||||
|
||||
|
||||
# The packer each format's checkpoint weights are written with. Both exporters resolve through
|
||||
# this one mapping so they cannot drift apart.
|
||||
IQ_PACKERS = {
|
||||
QUANTIZATION_IQ1_S: quantize_iq1_s,
|
||||
QUANTIZATION_IQ2_XXS: quantize_iq2_xxs,
|
||||
QUANTIZATION_IQ2_XS: quantize_iq2_xs,
|
||||
}
|
||||
|
||||
|
||||
# Formats whose scales are purely per-module, so export never merges them across the q/k/v
|
||||
# and gate/up groups that share an input. Every other format unifies input_amax (and, for
|
||||
# NVFP4, weight_scale_2) across such a group, which only a whole-model forward can discover.
|
||||
FUSION_FREE_FORMATS = frozenset(
|
||||
FUSION_FREE_FORMATS = IQ_FORMATS | frozenset(
|
||||
{
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
}
|
||||
|
||||
@@ -27,14 +27,6 @@ import torch.nn as nn
|
||||
|
||||
from modelopt import __version__
|
||||
from modelopt.torch.models import get_spec, list_all_possible
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
)
|
||||
from modelopt.torch.quantization.model_calib import (
|
||||
enable_stats_collection,
|
||||
finish_stats_collection,
|
||||
@@ -59,6 +51,8 @@ from modelopt.torch.utils import clear_cuda_cache
|
||||
from ..quantization.nn import NVFP4StaticQuantizer, SequentialQuantizer, TensorQuantizer
|
||||
from .model_utils import TiedWeightMap, get_language_model_from_vl
|
||||
from .quant_format import (
|
||||
IQ_BLOCK_METADATA,
|
||||
IQ_FORMATS,
|
||||
KV_CACHE_FP8,
|
||||
KV_CACHE_FP8_K_NVFP4_V,
|
||||
KV_CACHE_INT8,
|
||||
@@ -71,8 +65,6 @@ from .quant_format import (
|
||||
QUANTIZATION_INT4_AWQ,
|
||||
QUANTIZATION_INT8_SQ,
|
||||
QUANTIZATION_INT8_WO,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_MXFP4,
|
||||
QUANTIZATION_MXFP8,
|
||||
QUANTIZATION_NONE,
|
||||
@@ -474,8 +466,7 @@ def uses_iq_quantization(module) -> bool:
|
||||
if (
|
||||
weight_quantizer is not None
|
||||
and weight_quantizer.is_enabled
|
||||
and getattr(weight_quantizer, "num_bits", None)
|
||||
in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS)
|
||||
and getattr(weight_quantizer, "num_bits", None) in IQ_FORMATS
|
||||
):
|
||||
return True
|
||||
return any(uses_iq_quantization(child) for _, child in module.named_children())
|
||||
@@ -515,7 +506,7 @@ def get_quantization_format(module) -> str | None:
|
||||
return QUANTIZATION_W4A8_AWQ
|
||||
|
||||
# Handle individual num_bits cases
|
||||
if weight_quantizer.num_bits in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if weight_quantizer.num_bits in IQ_FORMATS:
|
||||
if weight_quantizer.backend != "ggml":
|
||||
raise ValueError("IQ formats require the built-in 'ggml' quantization backend")
|
||||
# Both exporters return before collecting input_scale and before the pre_quant_scale
|
||||
@@ -781,15 +772,8 @@ def process_layer_quant_config(layer_config_dict):
|
||||
"quant_algo": "MXFP8",
|
||||
"group_size": block_size_value,
|
||||
}
|
||||
elif v in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if v == QUANTIZATION_IQ1_S:
|
||||
block_size = IQ1_S_BLOCK_SIZE
|
||||
payload_bytes = IQ1_S_BLOCK_BYTES
|
||||
effective_bits = IQ1_S_EFFECTIVE_BITS
|
||||
else:
|
||||
block_size = IQ2_XS_BLOCK_SIZE
|
||||
payload_bytes = IQ2_XS_BLOCK_BYTES
|
||||
effective_bits = IQ2_XS_EFFECTIVE_BITS
|
||||
elif v in IQ_FORMATS:
|
||||
block_size, payload_bytes, effective_bits = IQ_BLOCK_METADATA[v]
|
||||
if block_size_value != block_size:
|
||||
raise ValueError(
|
||||
f"{v.upper()} requires block size {block_size}, got {block_size_value}"
|
||||
|
||||
@@ -67,7 +67,6 @@ except ImportError:
|
||||
from modelopt.torch.opt.conversion import ModeloptStateManager, modelopt_state
|
||||
from modelopt.torch.opt.plugins.huggingface import _MODELOPT_STATE_SAVE_NAME
|
||||
from modelopt.torch.quantization import set_quantizer_by_cfg_context
|
||||
from modelopt.torch.quantization.ggml import quantize_iq1_s, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn import SequentialQuantizer, TensorQuantizer
|
||||
from modelopt.torch.quantization.qtensor import MXFP8QTensor, NVFP4QTensor
|
||||
from modelopt.torch.quantization.qtensor.base_qtensor import QTensorWrapper
|
||||
@@ -101,11 +100,11 @@ from .quant_aware_conversion import (
|
||||
)
|
||||
from .quant_format import (
|
||||
FUSION_FREE_FORMATS,
|
||||
IQ_FORMATS,
|
||||
IQ_PACKERS,
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
QUANTIZATION_FP8_PC_PT,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_MXFP8,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_NVFP4,
|
||||
@@ -630,15 +629,13 @@ def _export_quantized_weight(
|
||||
"which dispatches to the streaming writer that materialises weights layer-by-layer."
|
||||
)
|
||||
|
||||
if quantization_format in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if quantization_format in IQ_FORMATS:
|
||||
if weight_name != "weight":
|
||||
raise NotImplementedError(
|
||||
"IQ unified export currently supports modules with a standard 'weight' "
|
||||
f"attribute, got {weight_name!r} on {type(sub_module).__name__}"
|
||||
)
|
||||
quantize_iq = (
|
||||
quantize_iq1_s if quantization_format == QUANTIZATION_IQ1_S else quantize_iq2_xs
|
||||
)
|
||||
quantize_iq = IQ_PACKERS[quantization_format]
|
||||
packed_weight, _ = quantize_iq(weight.to(dtype))
|
||||
setattr(sub_module, weight_name, nn.Parameter(packed_weight, requires_grad=False))
|
||||
maybe_clear_cuda_cache()
|
||||
|
||||
@@ -35,7 +35,6 @@ from safetensors import safe_open
|
||||
from safetensors.torch import save_file
|
||||
|
||||
from modelopt import __version__
|
||||
from modelopt.torch.quantization.ggml import quantize_iq1_s, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn.modules.tensor_quantizer import GroupedQuantizer
|
||||
from modelopt.torch.utils import import_plugin, warn_rank_0
|
||||
from modelopt.torch.utils.plugins.hf_checkpoint_utils import (
|
||||
@@ -57,13 +56,13 @@ from .plugins.mcore_custom import (
|
||||
)
|
||||
from .plugins.megatron_importer import GPTModelImporter, _get_mamba_conv1d
|
||||
from .quant_format import (
|
||||
IQ_FORMATS,
|
||||
IQ_PACKERS,
|
||||
KV_CACHE_FP8,
|
||||
KV_CACHE_NVFP4,
|
||||
QUANTIZATION_FP8,
|
||||
QUANTIZATION_FP8_PB_REAL,
|
||||
QUANTIZATION_FP8_PB_WO,
|
||||
QUANTIZATION_IQ1_S,
|
||||
QUANTIZATION_IQ2_XS,
|
||||
QUANTIZATION_NONE,
|
||||
QUANTIZATION_NVFP4,
|
||||
QUANTIZATION_W4A16_NVFP4,
|
||||
@@ -85,6 +84,7 @@ with import_plugin("transformers", verbose=False):
|
||||
import transformers
|
||||
from transformers import AutoProcessor
|
||||
|
||||
|
||||
has_mcore = False
|
||||
with import_plugin("megatron"):
|
||||
from megatron.core.models.gpt import GPTModel
|
||||
@@ -351,7 +351,7 @@ class GPTModelExporter:
|
||||
quantization = "NVFP4"
|
||||
elif quantization_format == QUANTIZATION_W4A16_NVFP4:
|
||||
quantization = "W4A16_NVFP4"
|
||||
elif quantization_format in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
elif quantization_format in IQ_FORMATS:
|
||||
quantization = quantization_format.upper()
|
||||
|
||||
if is_last_stage_main_rank:
|
||||
@@ -1115,7 +1115,7 @@ class GPTModelExporter:
|
||||
self._record_excluded_module(prefix)
|
||||
block_size = get_weight_block_size(module)
|
||||
|
||||
is_iq = qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS)
|
||||
is_iq = qformat in IQ_FORMATS
|
||||
name_to_value = self._get_weight_bias(
|
||||
module, dtype, name_to_value, keep_weight_device=is_iq
|
||||
)
|
||||
@@ -1185,7 +1185,7 @@ class GPTModelExporter:
|
||||
@staticmethod
|
||||
def _pack_iq_weight(weight: torch.Tensor, qformat: str) -> torch.Tensor:
|
||||
"""Pack one ``[out, in]`` weight and return its CPU payload."""
|
||||
quantize_iq = quantize_iq1_s if qformat == QUANTIZATION_IQ1_S else quantize_iq2_xs
|
||||
quantize_iq = IQ_PACKERS[qformat]
|
||||
packed_weight, _ = quantize_iq(weight)
|
||||
return packed_weight.detach().cpu()
|
||||
|
||||
@@ -1210,7 +1210,7 @@ class GPTModelExporter:
|
||||
The one gap left is a rank holding no local expert at all, which needs expert-parallel
|
||||
size to exceed the expert count. Worth revisiting if that becomes a supported topology.
|
||||
"""
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
raise NotImplementedError(
|
||||
"Fused-MoE IQ export requires a deployment loader that supports "
|
||||
"[num_experts, out_features, in_features // 256, payload_bytes]"
|
||||
@@ -1280,7 +1280,7 @@ class GPTModelExporter:
|
||||
weight = weight + 1.0
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
self._state_dict.update(self._get_iq_weight_state(prefix + "weight", weight, qformat))
|
||||
elif weight_scale is None:
|
||||
self._state_dict[prefix + "weight"] = weight
|
||||
@@ -1327,7 +1327,7 @@ class GPTModelExporter:
|
||||
gate_proj_weight = weight[:ffn_hidden_size, :]
|
||||
up_proj_weight = weight[ffn_hidden_size:, :]
|
||||
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
self._state_dict.update(
|
||||
self._get_iq_weight_state(gate_proj_prefix + "weight", gate_proj_weight, qformat)
|
||||
)
|
||||
@@ -1501,7 +1501,7 @@ class GPTModelExporter:
|
||||
seen_qformat, seen_block_size = qformat, block_size
|
||||
|
||||
weight = state_dict[weight_key].to(self.dtype)
|
||||
if qformat not in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat not in IQ_FORMATS:
|
||||
weight = weight.cpu()
|
||||
weight_scale_cpu = (
|
||||
weight_scale.detach().cpu().clone() if weight_scale is not None else None
|
||||
@@ -1533,7 +1533,7 @@ class GPTModelExporter:
|
||||
]
|
||||
|
||||
for shard_prefix, shard_weight, shard_scale in shards:
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
local_expert_state.update(
|
||||
self._get_iq_weight_state(
|
||||
shard_prefix + "weight", shard_weight, qformat
|
||||
@@ -1702,7 +1702,7 @@ class GPTModelExporter:
|
||||
proj_weights = [_take(weight, s, hidden_size, g) for s, g in zip(slices, gated)]
|
||||
proj_keys = [p + "weight" for p in prefixes]
|
||||
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
for key, weight in zip(proj_keys, proj_weights):
|
||||
self._state_dict.update(self._get_iq_weight_state(key, weight, qformat))
|
||||
elif weight_scale is None:
|
||||
@@ -1820,7 +1820,7 @@ class GPTModelExporter:
|
||||
proj_keys = [p + "weight" for p in proj_prefixes]
|
||||
weight_scale, weight_scale_2 = self._get_weight_scales(name_to_value, qformat)
|
||||
|
||||
if qformat in (QUANTIZATION_IQ1_S, QUANTIZATION_IQ2_XS):
|
||||
if qformat in IQ_FORMATS:
|
||||
for proj_prefix, proj_weight in zip(proj_prefixes, proj_weights):
|
||||
if proj_prefix in keep_bf16:
|
||||
self._state_dict[proj_prefix + "weight"] = proj_weight.cpu()
|
||||
|
||||
@@ -49,6 +49,7 @@ constexpr int kScaleBytes = 2;
|
||||
// them cannot drift apart.
|
||||
constexpr int kIq1sEntries = 2048;
|
||||
constexpr int kIq2xsEntries = 512;
|
||||
constexpr int kIq2xxsEntries = 256;
|
||||
|
||||
// One CUDA block encodes one GGML block. The reductions below fold over exactly this many warps,
|
||||
// and each kernel static_asserts that its codebook divides evenly among the threads.
|
||||
|
||||
@@ -23,6 +23,7 @@
|
||||
|
||||
at::Tensor iq1_s_pack_cuda(at::Tensor input, at::Tensor grid);
|
||||
at::Tensor iq2_xs_pack_cuda(at::Tensor input, at::Tensor grid, at::Tensor scales);
|
||||
at::Tensor iq2_xxs_pack_cuda(at::Tensor input, at::Tensor grid, at::Tensor scales);
|
||||
|
||||
namespace {
|
||||
|
||||
@@ -53,6 +54,23 @@ at::Tensor iq2_xs_pack(at::Tensor input, at::Tensor grid, at::Tensor scales) {
|
||||
return iq2_xs_pack_cuda(input.contiguous(), grid.contiguous(), scales.contiguous());
|
||||
}
|
||||
|
||||
at::Tensor iq2_xxs_pack(at::Tensor input, at::Tensor grid, at::Tensor scales) {
|
||||
TORCH_CHECK(input.is_cuda(), "IQ2_XXS packing requires a CUDA input");
|
||||
TORCH_CHECK(grid.is_cuda(), "IQ2_XXS packing requires a CUDA grid");
|
||||
TORCH_CHECK(scales.is_cuda(), "IQ2_XXS packing requires CUDA scales");
|
||||
modelopt::ggml::check_pack_inputs("IQ2_XXS", input, grid, modelopt::ggml::kIq2xxsEntries);
|
||||
const auto num_blocks = input.numel() / modelopt::ggml::kBlockSize;
|
||||
TORCH_CHECK(scales.scalar_type() == at::kHalf && scales.dim() == 1 &&
|
||||
scales.numel() == num_blocks,
|
||||
"scales must be float16 [numel / 256]");
|
||||
// Same rule as IQ2_XS: these bits become the block scale verbatim, and a negative or non-finite
|
||||
// one packs cleanly while decoding to garbage.
|
||||
TORCH_CHECK((scales.isfinite() & (scales >= 0)).all().item<bool>(),
|
||||
"scales must be finite and non-negative");
|
||||
TORCH_CHECK(input.get_device() == scales.get_device(), "input and scales must share a device");
|
||||
return iq2_xxs_pack_cuda(input.contiguous(), grid.contiguous(), scales.contiguous());
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
PYBIND11_MODULE(TORCH_EXTENSION_NAME, module) {
|
||||
@@ -69,4 +87,12 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, module) {
|
||||
"Returns uint8 [numel / 256, 74] on the input device. Non-finite input elements are "
|
||||
"treated as zero during packing, and finite elements outside the float32 range "
|
||||
"saturate.");
|
||||
module.def("iq2_xxs_pack", &iq2_xxs_pack,
|
||||
"Pack a non-empty float32, float64, float16, or bfloat16 CUDA tensor whose innermost "
|
||||
"dimension is a multiple of 256. The grid must be float32 [256, 8] holding "
|
||||
"non-negative codebook magnitudes, and scales must be finite non-negative float16 "
|
||||
"[numel / 256]. "
|
||||
"Returns uint8 [numel / 256, 66] on the input device. Non-finite input elements are "
|
||||
"treated as zero during packing, and finite elements outside the float32 range "
|
||||
"saturate.");
|
||||
}
|
||||
|
||||
@@ -0,0 +1,240 @@
|
||||
/*
|
||||
* SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
* SPDX-License-Identifier: Apache-2.0
|
||||
*
|
||||
* Licensed under the Apache License, Version 2.0 (the "License");
|
||||
* you may not use this file except in compliance with the License.
|
||||
* You may obtain a copy of the License at
|
||||
*
|
||||
* http://www.apache.org/licenses/LICENSE-2.0
|
||||
*
|
||||
* Unless required by applicable law or agreed to in writing, software
|
||||
* distributed under the License is distributed on an "AS IS" BASIS,
|
||||
* WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
* See the License for the specific language governing permissions and
|
||||
* limitations under the License.
|
||||
*/
|
||||
|
||||
#include "common.cuh"
|
||||
|
||||
namespace {
|
||||
|
||||
using namespace modelopt::ggml;
|
||||
|
||||
// The IQ2_XXS packed payload layout and format constants below follow the GGML
|
||||
// definition at:
|
||||
// https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-common.h
|
||||
constexpr int kEntries = kIq2xxsEntries;
|
||||
constexpr int kGroups = 8; // one 4-bit local scale per 32 values
|
||||
constexpr int kVectorsPerGroup = 4; // four 8-value codebook vectors per group
|
||||
constexpr int kLocalScales = 16;
|
||||
constexpr int kRecordBytes = 8; // four index bytes then one little-endian uint32
|
||||
constexpr int kCodeOffset = kScaleBytes;
|
||||
constexpr int kPayloadBytes = kCodeOffset + kGroups * kRecordBytes;
|
||||
constexpr float kLocalScaleStep = 0.125f; // Encoded scale is d * (2 * ls + 1) / 8.
|
||||
|
||||
static_assert(kEntries % kThreads == 0, "every thread must visit the same number of entries");
|
||||
static_assert((kEntries & (kEntries - 1)) == 0, "the codebook index mask assumes a power of two");
|
||||
static_assert(kPayloadBytes == 66, "IQ2_XXS blocks are 66 bytes");
|
||||
|
||||
// Dot product of |x| against one codebook vector, under the format's even-parity sign rule.
|
||||
// IQ2_XXS stores seven sign bits per vector and recovers the eighth from their parity, exactly as
|
||||
// IQ2_XS does, so an odd sign pattern must flip the coordinate with the smallest |x| * q penalty.
|
||||
__device__ __forceinline__ float even_parity_dot(const float *x, const float *q, bool odd_parity) {
|
||||
float dot = 0.0f;
|
||||
float weakest = FLT_MAX;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorSize; ++j) {
|
||||
const float term = fabsf(x[j]) * q[j];
|
||||
dot += term;
|
||||
weakest = fminf(weakest, term);
|
||||
}
|
||||
return odd_parity ? dot - 2.0f * weakest : dot;
|
||||
}
|
||||
|
||||
template <typename scalar_t>
|
||||
__global__ void encode(const scalar_t *input, int64_t num_blocks, const float *grid,
|
||||
const __half *scales, uint8_t *output) {
|
||||
__shared__ float shared_grid[kEntries * kVectorSize];
|
||||
__shared__ float grid_norm[kEntries];
|
||||
__shared__ float warp_best[kWarps * kLocalScales];
|
||||
__shared__ float group_error[kLocalScales];
|
||||
__shared__ unsigned long long warp_keys[kWarps];
|
||||
__shared__ int selected_local;
|
||||
__shared__ uint8_t locals[kGroups];
|
||||
__shared__ uint8_t entry_bytes[kGroups * kVectorsPerGroup];
|
||||
__shared__ uint8_t sign_bits[kGroups * kVectorsPerGroup];
|
||||
|
||||
const int tid = threadIdx.x;
|
||||
const int64_t block = blockIdx.x;
|
||||
if (block >= num_blocks)
|
||||
return;
|
||||
|
||||
for (int i = tid; i < kEntries * kVectorSize; i += blockDim.x)
|
||||
shared_grid[i] = grid[i];
|
||||
__syncthreads();
|
||||
for (int entry = tid; entry < kEntries; entry += blockDim.x) {
|
||||
float norm = 0.0f;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorSize; ++j) {
|
||||
const float q = shared_grid[entry * kVectorSize + j];
|
||||
norm = fmaf(q, q, norm);
|
||||
}
|
||||
grid_norm[entry] = norm;
|
||||
}
|
||||
__syncthreads();
|
||||
|
||||
const scalar_t *source = input + block * kBlockSize;
|
||||
uint8_t *payload = output + block * kPayloadBytes;
|
||||
const __half d_half = scales[block];
|
||||
const uint16_t d_bits = __half_as_ushort(d_half);
|
||||
const float d = __half2float(d_half);
|
||||
if (!store_block_scale<kPayloadBytes>(payload, d_bits))
|
||||
return;
|
||||
|
||||
#pragma unroll 1
|
||||
for (int group = 0; group < kGroups; ++group) {
|
||||
if (tid < kLocalScales)
|
||||
group_error[tid] = 0.0f;
|
||||
__syncthreads();
|
||||
|
||||
#pragma unroll
|
||||
for (int vector = 0; vector < kVectorsPerGroup; ++vector) {
|
||||
float x[kVectorSize];
|
||||
float xnorm = 0.0f;
|
||||
int negative_count = 0;
|
||||
const int offset = group * (kVectorsPerGroup * kVectorSize) + vector * kVectorSize;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorSize; ++j) {
|
||||
x[j] = load_float(source + offset + j);
|
||||
xnorm = fmaf(x[j], x[j], xnorm);
|
||||
negative_count += x[j] < 0.0f;
|
||||
}
|
||||
const bool odd_parity = (negative_count & 1) != 0;
|
||||
float local_best[kLocalScales];
|
||||
#pragma unroll
|
||||
for (int local = 0; local < kLocalScales; ++local)
|
||||
local_best[local] = FLT_MAX;
|
||||
for (int entry = tid; entry < kEntries; entry += blockDim.x) {
|
||||
const float *q = shared_grid + entry * kVectorSize;
|
||||
const float dot = even_parity_dot(x, q, odd_parity);
|
||||
#pragma unroll
|
||||
for (int local = 0; local < kLocalScales; ++local) {
|
||||
const float scale = d * (2 * local + 1) * kLocalScaleStep;
|
||||
local_best[local] =
|
||||
fminf(local_best[local], clamped_quant_error(xnorm, dot, grid_norm[entry], scale));
|
||||
}
|
||||
}
|
||||
block_min_accumulate<kLocalScales>(local_best, warp_best, group_error);
|
||||
}
|
||||
|
||||
if (tid == 0) {
|
||||
selected_local = 0;
|
||||
float best = group_error[0];
|
||||
#pragma unroll
|
||||
for (int local = 1; local < kLocalScales; ++local) {
|
||||
if (group_error[local] < best) {
|
||||
best = group_error[local];
|
||||
selected_local = local;
|
||||
}
|
||||
}
|
||||
locals[group] = static_cast<uint8_t>(selected_local);
|
||||
}
|
||||
__syncthreads();
|
||||
const float selected_scale = d * (2 * selected_local + 1) * kLocalScaleStep;
|
||||
|
||||
#pragma unroll
|
||||
for (int vector = 0; vector < kVectorsPerGroup; ++vector) {
|
||||
float x[kVectorSize];
|
||||
float xnorm = 0.0f;
|
||||
int negative_count = 0;
|
||||
const int offset = group * (kVectorsPerGroup * kVectorSize) + vector * kVectorSize;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorSize; ++j) {
|
||||
x[j] = load_float(source + offset + j);
|
||||
xnorm = fmaf(x[j], x[j], xnorm);
|
||||
negative_count += x[j] < 0.0f;
|
||||
}
|
||||
const bool odd_parity = (negative_count & 1) != 0;
|
||||
unsigned long long key = ~0ULL;
|
||||
for (int entry = tid; entry < kEntries; entry += blockDim.x) {
|
||||
const float error = clamped_quant_error(
|
||||
xnorm, even_parity_dot(x, shared_grid + entry * kVectorSize, odd_parity),
|
||||
grid_norm[entry], selected_scale);
|
||||
const unsigned long long candidate = error_key(error, entry);
|
||||
key = candidate < key ? candidate : key;
|
||||
}
|
||||
key = block_min_key(key, warp_keys);
|
||||
if (tid == 0) {
|
||||
const int entry = static_cast<int>(key & (kEntries - 1));
|
||||
const float *q = shared_grid + entry * kVectorSize;
|
||||
int flip_index = 0;
|
||||
float weakest = fabsf(x[0]) * q[0];
|
||||
#pragma unroll
|
||||
for (int j = 1; j < kVectorSize; ++j) {
|
||||
const float term = fabsf(x[j]) * q[j];
|
||||
if (term < weakest) {
|
||||
weakest = term;
|
||||
flip_index = j;
|
||||
}
|
||||
}
|
||||
int sign_mask = 0;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorSize; ++j) {
|
||||
bool is_negative = x[j] < 0.0f;
|
||||
if (odd_parity && j == flip_index)
|
||||
is_negative = !is_negative;
|
||||
sign_mask |= static_cast<int>(is_negative) << j;
|
||||
}
|
||||
const int slot = group * kVectorsPerGroup + vector;
|
||||
entry_bytes[slot] = static_cast<uint8_t>(entry);
|
||||
sign_bits[slot] = static_cast<uint8_t>(sign_mask & 0x7f);
|
||||
}
|
||||
__syncthreads();
|
||||
}
|
||||
}
|
||||
|
||||
// One 8-byte record per group: four index bytes, then a uint32 holding four 7-bit sign
|
||||
// indices in bits 0..27 and the 4-bit local scale in bits 28..31.
|
||||
if (tid < kGroups) {
|
||||
uint8_t *record = payload + kCodeOffset + tid * kRecordBytes;
|
||||
const int base = tid * kVectorsPerGroup;
|
||||
#pragma unroll
|
||||
for (int j = 0; j < kVectorsPerGroup; ++j)
|
||||
record[j] = entry_bytes[base + j];
|
||||
const uint32_t aux = static_cast<uint32_t>(sign_bits[base]) |
|
||||
(static_cast<uint32_t>(sign_bits[base + 1]) << 7) |
|
||||
(static_cast<uint32_t>(sign_bits[base + 2]) << 14) |
|
||||
(static_cast<uint32_t>(sign_bits[base + 3]) << 21) |
|
||||
(static_cast<uint32_t>(locals[tid]) << 28);
|
||||
#pragma unroll
|
||||
for (int j = 0; j < 4; ++j)
|
||||
record[kVectorsPerGroup + j] = static_cast<uint8_t>(aux >> (8 * j));
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
at::Tensor iq2_xxs_pack_cuda(at::Tensor input, at::Tensor grid, at::Tensor scales) {
|
||||
TORCH_CHECK(input.is_contiguous() && grid.is_contiguous() && scales.is_contiguous(),
|
||||
"inputs must be contiguous");
|
||||
check_pack_inputs("IQ2_XXS", input, grid, kEntries);
|
||||
const int64_t num_blocks = input.numel() / kBlockSize;
|
||||
TORCH_CHECK(scales.scalar_type() == at::kHalf && scales.dim() == 1 &&
|
||||
scales.numel() == num_blocks,
|
||||
"scales must be float16 [numel / 256]");
|
||||
TORCH_CHECK(input.get_device() == scales.get_device(), "input and scales must share a device");
|
||||
c10::cuda::CUDAGuard guard(input.device());
|
||||
auto output = at::empty({num_blocks, kPayloadBytes}, input.options().dtype(at::kByte));
|
||||
const auto stream = c10::cuda::getCurrentCUDAStream();
|
||||
|
||||
AT_DISPATCH_FLOATING_TYPES_AND2(
|
||||
at::ScalarType::Half, at::ScalarType::BFloat16, input.scalar_type(), "iq2_xxs_pack", [&] {
|
||||
encode<scalar_t><<<static_cast<int>(num_blocks), kThreads, 0, stream>>>(
|
||||
input.data_ptr<scalar_t>(), num_blocks, grid.data_ptr<float>(),
|
||||
reinterpret_cast<const __half *>(scales.data_ptr<at::Half>()),
|
||||
output.data_ptr<uint8_t>());
|
||||
C10_CUDA_KERNEL_LAUNCH_CHECK();
|
||||
});
|
||||
return output;
|
||||
}
|
||||
@@ -95,6 +95,7 @@ def get_cuda_ext_ggml(raise_if_failed: bool = False):
|
||||
kernels_ggml / "ggml.cpp",
|
||||
kernels_ggml / "iq1_s.cu",
|
||||
kernels_ggml / "iq2_xs.cu",
|
||||
kernels_ggml / "iq2_xxs.cu",
|
||||
],
|
||||
cuda_version_specifiers=">=11.8",
|
||||
fail_msg="GGML IQ CUDA packing extension is unavailable.",
|
||||
|
||||
@@ -21,5 +21,11 @@ from .iq1_s import *
|
||||
from .iq1_s import __all__ as _iq1_s_all
|
||||
from .iq2_xs import *
|
||||
from .iq2_xs import __all__ as _iq2_xs_all
|
||||
from .iq2_xxs import *
|
||||
from .iq2_xxs import __all__ as _iq2_xxs_all
|
||||
|
||||
__all__ = [*_iq1_s_all, *_iq2_xs_all] # noqa: PLE0604
|
||||
__all__ = [ # noqa: PLE0604
|
||||
*_iq1_s_all,
|
||||
*_iq2_xs_all,
|
||||
*_iq2_xxs_all,
|
||||
]
|
||||
|
||||
@@ -20,6 +20,14 @@ import torch
|
||||
from ..nn.modules.tensor_quantizer import register_quant_backend
|
||||
from .iq1_s import iq1_s_fake_quant
|
||||
from .iq2_xs import iq2_xs_fake_quant
|
||||
from .iq2_xxs import iq2_xxs_fake_quant
|
||||
|
||||
# One entry per GGML IQ format; adding a format is adding a row here.
|
||||
_FAKE_QUANTS = {
|
||||
"iq1_s": iq1_s_fake_quant,
|
||||
"iq2_xs": iq2_xs_fake_quant,
|
||||
"iq2_xxs": iq2_xxs_fake_quant,
|
||||
}
|
||||
|
||||
|
||||
def ggml_fake_quant(inputs: torch.Tensor, quantizer) -> torch.Tensor:
|
||||
@@ -29,11 +37,13 @@ def ggml_fake_quant(inputs: torch.Tensor, quantizer) -> torch.Tensor:
|
||||
unknown_args = set(extra_args) - {"block_chunk_size", "decode_chunk_size"}
|
||||
if unknown_args:
|
||||
raise ValueError(f"Unsupported ggml backend_extra_args: {sorted(unknown_args)}")
|
||||
if num_bits == "iq1_s":
|
||||
return iq1_s_fake_quant(inputs, quantizer, **extra_args)
|
||||
if num_bits == "iq2_xs":
|
||||
return iq2_xs_fake_quant(inputs, quantizer, **extra_args)
|
||||
raise ValueError("The ggml backend requires num_bits='iq1_s' or 'iq2_xs'")
|
||||
# num_bits arrives untyped from the quantizer and is a tuple for scalar formats,
|
||||
# so narrow before the lookup rather than relying on the dict to reject it.
|
||||
fake_quant = _FAKE_QUANTS.get(num_bits) if isinstance(num_bits, str) else None
|
||||
if fake_quant is None:
|
||||
supported = ", ".join(repr(name) for name in sorted(_FAKE_QUANTS))
|
||||
raise ValueError(f"The ggml backend requires num_bits in ({supported})")
|
||||
return fake_quant(inputs, quantizer, **extra_args)
|
||||
|
||||
|
||||
register_quant_backend("ggml", ggml_fake_quant)
|
||||
|
||||
@@ -166,6 +166,19 @@ _IQ2_XS_GRID_B64 = (
|
||||
)
|
||||
|
||||
|
||||
# Compact byte representation of the canonical [256, 8] IQ2_XXS grid. Same
|
||||
# 8/25/43 magnitude alphabet as IQ2_XS, so it compresses well.
|
||||
_IQ2_XXS_GRID_ZLIB_B64 = (
|
||||
"eNqFVVuS5CAM++UKOoPuf78ZLMmYLLXT1SkSYvyQZGct/egVuDfoFQtPA34M6RXLB74HqRW/N9rm5ZBZmWc5hK+vY+ATII55Oz4X"
|
||||
"c1+OMZ2PANgWtccrIBOYqYzwEdoH7cfP5EyQSZRJmNpHlbLNMZHCBghKgy/kfh/pOLPgxOgrwUkXoRKrCjo7OpsJlGOXAY+LMhvQ"
|
||||
"fIGM8/OST2CBvwAWw4JlAk5ZIEC/AZZ7Gz0A54Z5AJ4rkTr1hYuAijSIEEdlXMqryIOQ7paoHJEMMaVPr5LNJGRq90isoTxEVbGC"
|
||||
"6xCHFaiptZQ2CXVH7Jcidiuqkf2HaGMR8XTsdXUBrOaA1wwdgeAjFOca+ZsFwOu65FtxRlUllEtQfAsLlCKiobobAotS+sSnwyXK"
|
||||
"0elKQHnT7Wo7wXjB2T3XGr/hOQJmhKxuiaqLvuNeOMKh91qGAl9aAOZkESSOW4KnIS9sJL3/NQJ5NYRiUX9pGrNBgEysAb4ahXru"
|
||||
"STY/O/mcdAJRbDOSXlzd1WuOymbMktGQnRrkNQF7LBgFz106e/rDkjZmRpFm1KMx3WvRuqemoa19esbHuBgw7VKggELYclb7VH91"
|
||||
"Xf5hzoD21KEk6CpbkT/vBJbU"
|
||||
)
|
||||
|
||||
|
||||
@cache
|
||||
def iq1_s_grid_bytes() -> bytes:
|
||||
"""Decoded little-endian int8 bytes of the [2048, 8] IQ1_S ternary table."""
|
||||
@@ -176,3 +189,9 @@ def iq1_s_grid_bytes() -> bytes:
|
||||
def iq2_xs_grid_bytes() -> bytes:
|
||||
"""Decoded bytes of the [512, 8] IQ2_XS magnitude table."""
|
||||
return base64.b64decode(_IQ2_XS_GRID_B64)
|
||||
|
||||
|
||||
@cache
|
||||
def iq2_xxs_grid_bytes() -> bytes:
|
||||
"""Decoded bytes of the [256, 8] IQ2_XXS magnitude table."""
|
||||
return zlib.decompress(base64.b64decode(_IQ2_XXS_GRID_ZLIB_B64))
|
||||
|
||||
@@ -0,0 +1,286 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""IQ2_XXS fake quantization and GGML-compatible block packing.
|
||||
|
||||
The encoder performs a single-pass squared-error grid search at a fixed,
|
||||
empirically anchored super-block scale, mirroring :mod:`.iq2_xs`. Every 256
|
||||
logical values become one 66-byte block_iq2_xxs payload:
|
||||
|
||||
* bytes 0..1: little-endian FP16 super-block scale d
|
||||
* bytes 2..65: eight 8-byte sub-block records, each holding four 8-bit grid
|
||||
indices followed by a little-endian uint32 of four 7-bit sign indices
|
||||
(bits 0..27) and one 4-bit local scale (bits 28..31)
|
||||
|
||||
IQ2_XXS differs from IQ2_XS in three ways: the grid is 256 entries rather than
|
||||
512 so an index needs no high bits, one local scale covers a whole 32-value
|
||||
sub-block rather than 16 values, and the scale shares a word with the signs
|
||||
instead of living in a trailing array.
|
||||
|
||||
The canonical 256 x 8 magnitude grid lives in :mod:`.codebooks`, carried from
|
||||
llama.cpp ggml-common.h revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
|
||||
The matching dequantization formula is in ggml-quants.c at the same revision:
|
||||
https://github.com/ggml-org/llama.cpp/blob/9b05354ec6fb58b4e665e9a39ebc40285c015638/ggml/src/ggml-quants.c#L2489-L2514
|
||||
"""
|
||||
|
||||
import torch
|
||||
|
||||
from ..extensions import get_cuda_ext_ggml
|
||||
from .codebooks import iq2_xxs_grid_bytes
|
||||
from .common import (
|
||||
GGML_BLOCK_SIZE,
|
||||
fake_quantize_with_cache,
|
||||
narrow_to_float32,
|
||||
validate_block_chunk_size,
|
||||
validate_packed_weights,
|
||||
validate_weight,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"IQ2_XXS_BLOCK_BYTES",
|
||||
"IQ2_XXS_BLOCK_SIZE",
|
||||
"IQ2_XXS_EFFECTIVE_BITS",
|
||||
"dequantize_iq2_xxs",
|
||||
"iq2_xxs_fake_quant",
|
||||
"iq2_xxs_grid",
|
||||
"quantize_iq2_xxs",
|
||||
]
|
||||
|
||||
IQ2_XXS_BLOCK_SIZE = GGML_BLOCK_SIZE
|
||||
IQ2_XXS_BLOCK_BYTES = 66
|
||||
IQ2_XXS_EFFECTIVE_BITS = IQ2_XXS_BLOCK_BYTES * 8 / IQ2_XXS_BLOCK_SIZE
|
||||
_IQ2_XXS_GRID_ENTRIES = 256
|
||||
_IQ2_XXS_LOCAL_SCALES = 16
|
||||
_IQ2_XXS_GROUPS = 32
|
||||
_IQ2_XXS_GROUPS_PER_SUBBLOCK = 4
|
||||
_IQ2_XXS_SUBBLOCKS = _IQ2_XXS_GROUPS // _IQ2_XXS_GROUPS_PER_SUBBLOCK
|
||||
# Largest representable magnitude: grid entry 43 at local scale 15 -> (0.5 + 15) * 0.25.
|
||||
_IQ2_XXS_NATIVE_MAX = 43 * 31 / 8
|
||||
_IQ2_XXS_SCALE_ANCHOR_MIN = 0.65
|
||||
_IQ2_XXS_SCALE_ANCHOR_MAX = 0.92
|
||||
_IQ2_XXS_PEAK_TO_RMS_TAPER = 0.035
|
||||
# Bounds the encode search temporaries; see the note in .iq2_xs. The grid is half the size
|
||||
# of IQ2_XS's, so the same chunk holds half the search tile.
|
||||
_DEFAULT_BLOCK_CHUNK_SIZE = 512
|
||||
# The decode runs on every forward and is launch-bound, so it takes a much larger chunk.
|
||||
_DEFAULT_DECODE_CHUNK_SIZE = 4096
|
||||
_SCALE_BLOCK_CHUNK_SIZE = 4096
|
||||
|
||||
_GRID_CACHE: dict[torch.device, torch.Tensor] = {}
|
||||
|
||||
|
||||
def iq2_xxs_grid(device: torch.device | str | None = None) -> torch.Tensor:
|
||||
"""Return the canonical IQ2_XXS magnitude grid as float32."""
|
||||
resolved_device = torch.device(device or "cpu")
|
||||
if resolved_device.type == "cuda" and resolved_device.index is None:
|
||||
resolved_device = torch.device("cuda", torch.cuda.current_device())
|
||||
if resolved_device not in _GRID_CACHE:
|
||||
values = torch.tensor(list(iq2_xxs_grid_bytes()), dtype=torch.float32)
|
||||
_GRID_CACHE[resolved_device] = values.reshape(_IQ2_XXS_GRID_ENTRIES, 8).to(
|
||||
device=resolved_device
|
||||
)
|
||||
return _GRID_CACHE[resolved_device]
|
||||
|
||||
|
||||
def _predict_iq2_xxs_scales(blocks: torch.Tensor) -> torch.Tensor:
|
||||
"""Predict one FP16 super-block scale for each flattened block."""
|
||||
x = narrow_to_float32(blocks)
|
||||
amax = x.abs().amax(dim=1)
|
||||
rms = x.square().mean(dim=1).sqrt()
|
||||
peak_to_rms = torch.where(rms > 0, amax / rms, torch.zeros_like(rms))
|
||||
anchor_ratio = (1.0 - _IQ2_XXS_PEAK_TO_RMS_TAPER * peak_to_rms).clamp(
|
||||
_IQ2_XXS_SCALE_ANCHOR_MIN, _IQ2_XXS_SCALE_ANCHOR_MAX
|
||||
)
|
||||
return ((amax / _IQ2_XXS_NATIVE_MAX) * anchor_ratio).clamp(max=65504.0).to(torch.float16)
|
||||
|
||||
|
||||
def _encode_blocks(blocks: torch.Tensor, grid: torch.Tensor) -> torch.Tensor:
|
||||
"""Encode a moderate-size batch of flattened 256-value blocks."""
|
||||
x = narrow_to_float32(blocks)
|
||||
block_count = x.shape[0]
|
||||
vectors = x.reshape(block_count, _IQ2_XXS_GROUPS, 8)
|
||||
magnitudes = vectors.abs()
|
||||
negative = vectors < 0
|
||||
# Only seven sign bits are stored; the eighth is their parity, so an odd sign pattern
|
||||
# has to flip one element. Account for that cost while searching, not after.
|
||||
odd_parity = negative.sum(dim=-1).remainder(2).bool()
|
||||
|
||||
d = _predict_iq2_xxs_scales(x)
|
||||
d_float = d.float()
|
||||
|
||||
xnorm = vectors.square().sum(dim=-1)
|
||||
qnorm = grid.square().sum(dim=-1)
|
||||
shape = (block_count, _IQ2_XXS_GROUPS, _IQ2_XXS_LOCAL_SCALES)
|
||||
best_error = torch.full(shape, torch.inf, dtype=torch.float32, device=x.device)
|
||||
best_entry = torch.zeros(shape, dtype=torch.int64, device=x.device)
|
||||
# Search the codebook in tiles to cap temporary memory. Strict comparison
|
||||
# preserves the lowest grid index on equal error.
|
||||
for entry_start in range(0, _IQ2_XXS_GRID_ENTRIES, 64):
|
||||
grid_tile = grid[entry_start : entry_start + 64]
|
||||
products = magnitudes.unsqueeze(2) * grid_tile.reshape(1, 1, -1, 8)
|
||||
dot = products.sum(dim=-1)
|
||||
dot = torch.where(odd_parity.unsqueeze(-1), dot - 2.0 * products.amin(dim=-1), dot)
|
||||
tile_qnorm = qnorm[entry_start : entry_start + 64].reshape(1, 1, -1)
|
||||
|
||||
for local in range(_IQ2_XXS_LOCAL_SCALES):
|
||||
scale = d_float.reshape(-1, 1, 1) * ((2 * local + 1) / 8.0)
|
||||
error = (
|
||||
xnorm.unsqueeze(-1) - 2.0 * scale * dot + scale.square() * tile_qnorm
|
||||
).clamp_min_(0)
|
||||
tile_error, tile_index = error.min(dim=-1)
|
||||
replace = tile_error < best_error[:, :, local]
|
||||
best_error[:, :, local] = torch.where(replace, tile_error, best_error[:, :, local])
|
||||
best_entry[:, :, local] = torch.where(
|
||||
replace, tile_index + entry_start, best_entry[:, :, local]
|
||||
)
|
||||
|
||||
# One local scale covers four groups here, against two for IQ2_XS.
|
||||
subblock_error = best_error.reshape(
|
||||
block_count, _IQ2_XXS_SUBBLOCKS, _IQ2_XXS_GROUPS_PER_SUBBLOCK, _IQ2_XXS_LOCAL_SCALES
|
||||
).sum(dim=2)
|
||||
selected_local = subblock_error.argmin(dim=-1)
|
||||
group_local = selected_local.repeat_interleave(_IQ2_XXS_GROUPS_PER_SUBBLOCK, dim=1)
|
||||
selected_entry = best_entry.gather(2, group_local.unsqueeze(-1)).squeeze(-1)
|
||||
|
||||
selected_grid = grid[selected_entry]
|
||||
weakest_index = (magnitudes * selected_grid).argmin(dim=-1)
|
||||
flip = torch.nn.functional.one_hot(weakest_index, num_classes=8).bool()
|
||||
encoded_negative = negative ^ (flip & odd_parity.unsqueeze(-1))
|
||||
sign_bits = torch.arange(8, dtype=torch.int64, device=x.device)
|
||||
sign_mask = (encoded_negative.to(torch.int64) << sign_bits).sum(dim=-1) & 0x7F
|
||||
|
||||
signs = sign_mask.reshape(block_count, _IQ2_XXS_SUBBLOCKS, _IQ2_XXS_GROUPS_PER_SUBBLOCK)
|
||||
aux = (
|
||||
signs[:, :, 0]
|
||||
| (signs[:, :, 1] << 7)
|
||||
| (signs[:, :, 2] << 14)
|
||||
| (signs[:, :, 3] << 21)
|
||||
| (selected_local << 28)
|
||||
)
|
||||
|
||||
body = torch.empty((block_count, _IQ2_XXS_SUBBLOCKS, 8), dtype=torch.uint8, device=x.device)
|
||||
body[:, :, 0:4] = (
|
||||
selected_entry.reshape(block_count, _IQ2_XXS_SUBBLOCKS, _IQ2_XXS_GROUPS_PER_SUBBLOCK)
|
||||
).to(torch.uint8)
|
||||
for byte in range(4):
|
||||
body[:, :, 4 + byte] = ((aux >> (8 * byte)) & 0xFF).to(torch.uint8)
|
||||
|
||||
packed = torch.empty((block_count, IQ2_XXS_BLOCK_BYTES), dtype=torch.uint8, device=x.device)
|
||||
packed[:, :2] = d.contiguous().view(torch.uint8).reshape(block_count, 2)
|
||||
packed[:, 2:] = body.reshape(block_count, -1)
|
||||
return torch.where((d_float == 0).unsqueeze(1), 0, packed)
|
||||
|
||||
|
||||
@torch.no_grad()
|
||||
def quantize_iq2_xxs(
|
||||
weight: torch.Tensor, *, block_chunk_size: int = _DEFAULT_BLOCK_CHUNK_SIZE
|
||||
) -> tuple[torch.Tensor, torch.Tensor]:
|
||||
"""Pack a floating-point weight into GGML-compatible IQ2_XXS blocks.
|
||||
|
||||
Returned shapes are ``[*weight.shape[:-1], weight.shape[-1] // 256, 66]``
|
||||
and ``[weight.ndim]``. The packed payload remains on the weight's device;
|
||||
the logical-shape metadata is kept on CPU. Non-finite input elements are
|
||||
treated as zero during packing.
|
||||
"""
|
||||
validate_weight(weight, "IQ2_XXS")
|
||||
validate_block_chunk_size(block_chunk_size)
|
||||
|
||||
logical_shape = torch.tensor(weight.shape, dtype=torch.int64)
|
||||
blocks = weight.contiguous().reshape(-1, IQ2_XXS_BLOCK_SIZE)
|
||||
grid = iq2_xxs_grid(weight.device)
|
||||
packed_shape = (
|
||||
*weight.shape[:-1],
|
||||
weight.shape[-1] // IQ2_XXS_BLOCK_SIZE,
|
||||
IQ2_XXS_BLOCK_BYTES,
|
||||
)
|
||||
if weight.is_cuda:
|
||||
extension = get_cuda_ext_ggml()
|
||||
if extension is not None:
|
||||
scale_chunks = [
|
||||
_predict_iq2_xxs_scales(blocks[start : start + _SCALE_BLOCK_CHUNK_SIZE])
|
||||
for start in range(0, blocks.shape[0], _SCALE_BLOCK_CHUNK_SIZE)
|
||||
]
|
||||
packed = extension.iq2_xxs_pack(blocks, grid, torch.cat(scale_chunks))
|
||||
return packed.reshape(packed_shape), logical_shape
|
||||
|
||||
chunks = [
|
||||
_encode_blocks(blocks[start : start + block_chunk_size], grid)
|
||||
for start in range(0, blocks.shape[0], block_chunk_size)
|
||||
]
|
||||
return torch.cat(chunks).reshape(packed_shape), logical_shape
|
||||
|
||||
|
||||
@torch.no_grad()
|
||||
def dequantize_iq2_xxs(
|
||||
packed_weights: torch.Tensor,
|
||||
weight_shape: torch.Tensor,
|
||||
*,
|
||||
dtype: torch.dtype = torch.bfloat16,
|
||||
block_chunk_size: int = _DEFAULT_DECODE_CHUNK_SIZE,
|
||||
) -> torch.Tensor:
|
||||
"""Decode GGML-compatible IQ2_XXS payload bytes."""
|
||||
shape = validate_packed_weights(
|
||||
packed_weights, weight_shape, block_bytes=IQ2_XXS_BLOCK_BYTES, format_name="IQ2_XXS"
|
||||
)
|
||||
validate_block_chunk_size(block_chunk_size)
|
||||
|
||||
blocks = packed_weights.contiguous().reshape(-1, IQ2_XXS_BLOCK_BYTES)
|
||||
bit_positions = torch.arange(8, dtype=torch.int64, device=blocks.device)
|
||||
sign_shifts = 7 * torch.arange(
|
||||
_IQ2_XXS_GROUPS_PER_SUBBLOCK, dtype=torch.int64, device=blocks.device
|
||||
)
|
||||
grid = iq2_xxs_grid(blocks.device)
|
||||
decoded = torch.empty((blocks.shape[0], IQ2_XXS_BLOCK_SIZE), dtype=dtype, device=blocks.device)
|
||||
for start in range(0, blocks.shape[0], block_chunk_size):
|
||||
stop = min(start + block_chunk_size, blocks.shape[0])
|
||||
block_chunk = blocks[start:stop]
|
||||
count = block_chunk.shape[0]
|
||||
d = block_chunk[:, :2].contiguous().view(torch.float16).reshape(-1).float()
|
||||
body = block_chunk[:, 2:].reshape(count, _IQ2_XXS_SUBBLOCKS, 8).to(torch.int64)
|
||||
entries = body[:, :, 0:4]
|
||||
aux = body[:, :, 4] | (body[:, :, 5] << 8) | (body[:, :, 6] << 16) | (body[:, :, 7] << 24)
|
||||
# Top nibble is the sub-block scale; the low 28 bits are four 7-bit sign indices.
|
||||
scales = d.unsqueeze(-1) * (0.5 + ((aux >> 28) & 0xF).float()) * 0.25
|
||||
sign_index = (aux.unsqueeze(-1) >> sign_shifts) & 0x7F
|
||||
folded = sign_index ^ (sign_index >> 4)
|
||||
folded ^= folded >> 2
|
||||
folded ^= folded >> 1
|
||||
sign_mask = sign_index | ((folded & 1) << 7)
|
||||
signs = 1.0 - 2.0 * ((sign_mask.unsqueeze(-1) >> bit_positions) & 1).float()
|
||||
values = grid[entries] * signs
|
||||
chunk_decoded = values * scales.unsqueeze(-1).unsqueeze(-1)
|
||||
decoded[start:stop] = chunk_decoded.reshape(-1, IQ2_XXS_BLOCK_SIZE)
|
||||
return decoded.reshape(shape)
|
||||
|
||||
|
||||
def iq2_xxs_fake_quant(
|
||||
inputs: torch.Tensor,
|
||||
quantizer,
|
||||
*,
|
||||
block_chunk_size: int = _DEFAULT_BLOCK_CHUNK_SIZE,
|
||||
decode_chunk_size: int = _DEFAULT_DECODE_CHUNK_SIZE,
|
||||
) -> torch.Tensor:
|
||||
"""IQ2_XXS weight backend for TensorQuantizer, with pass-through backward."""
|
||||
if getattr(quantizer, "num_bits", None) != "iq2_xxs":
|
||||
raise ValueError("The ggml IQ2_XXS backend requires num_bits='iq2_xxs'")
|
||||
return fake_quantize_with_cache(
|
||||
inputs,
|
||||
quantizer,
|
||||
format_name="iq2_xxs",
|
||||
block_chunk_size=block_chunk_size,
|
||||
decode_chunk_size=decode_chunk_size,
|
||||
quantize=quantize_iq2_xxs,
|
||||
dequantize=dequantize_iq2_xxs,
|
||||
)
|
||||
@@ -0,0 +1,26 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# IQ2_XXS weight quantizer using the built-in fixed-scale codebook search.
|
||||
|
||||
# modelopt-schema: modelopt.torch.quantization.config.QuantizerAttributeConfig
|
||||
num_bits: iq2_xxs
|
||||
# Cost metadata for AutoQuantize's compression estimate only; it drives no packing or
|
||||
# numerics. num_bits is the string "iq2_xxs", so the generic estimator cannot derive the
|
||||
# storage cost: 66 packed bytes * 8 / 256 weights. Keep in sync with IQ2_XXS_BLOCK_BYTES.
|
||||
effective_bits: 2.0625
|
||||
block_sizes:
|
||||
-1: 256
|
||||
backend: ggml
|
||||
@@ -0,0 +1,32 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# QuantizeConfig preset for IQ2_XXS weight-only quantization.
|
||||
|
||||
# modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
iq2_xxs: configs/numerics/iq2_xxs
|
||||
|
||||
algorithm:
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- quantizer_name: '*weight_quantizer'
|
||||
cfg:
|
||||
$import: iq2_xxs
|
||||
- quantizer_name: '*input_quantizer'
|
||||
enable: false
|
||||
- $import: default_disabled_quantizers
|
||||
@@ -0,0 +1,27 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# IQ2_XXS weight-only PTQ.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
preset: configs/ptq/presets/model/iq2_xxs
|
||||
|
||||
metadata:
|
||||
description: >-
|
||||
Applies uniform GGML-compatible IQ2_XXS weight-only quantization to eligible linear layers.
|
||||
This is not a mixed per-tensor precision preset. No calibration data is required.
|
||||
quantize:
|
||||
$import: preset
|
||||
@@ -28,7 +28,7 @@ supported combinations.
|
||||
### The shipped recipes
|
||||
|
||||
<details>
|
||||
<summary>All 28 <code>general/ptq/</code> recipes (click to expand)</summary>
|
||||
<summary>All 29 <code>general/ptq/</code> recipes (click to expand)</summary>
|
||||
|
||||
| Recipe | Model body | KV cache | Calibration |
|
||||
|--------|-----------|----------|-------------|
|
||||
@@ -59,7 +59,8 @@ supported combinations.
|
||||
| `nvfp4_mlp_weight_only` | NVFP4 W4A16 (block 32), MLP + MoE weights only | none | max |
|
||||
| `mxfp4_mlp_weight_only` | MXFP4 W4A16, MLP + MoE weights only | none | none (no calibration) |
|
||||
| `iq1_s` | IQ1_S W1A16, eligible linears | none | none (no calibration) |
|
||||
| `iq2_xs` | IQ2_XS W2A16, eligible linears | none | none (no calibration) |
|
||||
| `iq2_xxs` | IQ2_XXS W2A16 (2.06 bpw), eligible linears | none | none (no calibration) |
|
||||
| `iq2_xs` | IQ2_XS W2A16 (2.31 bpw), eligible linears | none | none (no calibration) |
|
||||
|
||||
</details>
|
||||
|
||||
@@ -138,9 +139,10 @@ activations and tensor-core math are what deliver the throughput.
|
||||
- **`mxfp4_mlp_weight_only`** — MXFP4 weights on MLP/MoE layers only, BF16
|
||||
activations. Needs no calibration forward pass; the QAT starting point for the
|
||||
GPT-OSS family (see `examples/gpt-oss`).
|
||||
- **`iq1_s` / `iq2_xs`** — GGML-compatible IQ1_S or IQ2_XS weights on the eligible
|
||||
linear layers, with BF16 activations; `lm_head`, MoE routers, `conv1d` and the
|
||||
vision branch stay in BF16 like every other preset. No calibration data is
|
||||
- **`iq1_s` / `iq2_xxs` / `iq2_xs`** — GGML-compatible IQ weights
|
||||
on the eligible linear layers, with BF16 activations; `lm_head`, MoE routers,
|
||||
`conv1d` and the vision branch stay in BF16 like every other preset. The formats
|
||||
trade size against accuracy in order: 1.56, 2.06 and 2.31 bits per weight. No calibration data is
|
||||
required. Quantized weights must have a final dimension divisible by 256.
|
||||
Unified HF export writes the packed GGML blocks; Megatron export additionally
|
||||
requires tensor and pipeline parallel sizes of 1, and does not support
|
||||
|
||||
@@ -0,0 +1,162 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""Conformance vectors captured from a real llama.cpp-quantized checkpoint.
|
||||
|
||||
Each entry holds packed block bytes lifted verbatim from
|
||||
unsloth/Qwen3.8-27B-GGUF (Qwen3.8-27B-UD-IQ1_S.gguf) together with the values
|
||||
llama.cpp's own dequantize_row_* produces for them, computed from
|
||||
ggml-quants.c revision 9b05354ec6fb58b4e665e9a39ebc40285c015638.
|
||||
|
||||
These pin our decoders against bytes we did not produce. A decoder that drifts
|
||||
from the GGML layout -- a mis-set high bit, a swapped scale nibble, a sign
|
||||
parity mistake -- fails here even though a round-trip test against our own
|
||||
encoder would still pass.
|
||||
"""
|
||||
|
||||
import base64
|
||||
import zlib
|
||||
|
||||
import numpy as np
|
||||
|
||||
_VECTORS = {
|
||||
"iq1_s": {
|
||||
"source": "blk.0.attn_qkv.weight",
|
||||
"block_bytes": 50,
|
||||
"blocks": (
|
||||
"eNoBLAHT/sgYPIuVHR5unIa25z3nTMBrrDsowEZpEmhZYzdCh6ybWSOReiY2MD9xWx/BZF4J3Z/LTxnt0w2T+DYRwT0N"
|
||||
"IIwJoKkC4mejy6P3sqbJWgv0FgfWRjO1zs0h+ZpZymO6MyQ9Ms9aF+DLPrjwY54mJRLDZo2emJfnx3B7rQMfrk8bWdoJ"
|
||||
"Hc7n5ue6duVmzL9SXOzWw0SmcOsXFEu0ESDUFxuXCM0YGhrzGpYf6sJ0hTc5bGk4hFUG9NtWPKTaovj92oteeEPl9Fc9"
|
||||
"8BYIR2j/8Hv38eeyZ9fzxt60f8/DEfda/DyI7XWEZ6CXKNxq2nbhUvHVwm4FckftMeVoGA2YAjcSngU6wBLrdmQNavPg"
|
||||
"js7fDIkqa+Domf9cZoy23tsEc05IMcBUzWTS8th0Vso1lo8="
|
||||
),
|
||||
"expected": (
|
||||
"eNqFVzGIHVUUHTCRgGCaLQJugoFUQYLoFuK8T3YtxDKIBASRDRZiRLCwsRtCYrGsWBjEwuAvAiJJJTEE5wW+pAsJNmJC"
|
||||
"QNlOkG3EJiiic+7OmT3v/jv/F4d77znn3vf+vPdgtzrwT64+/KwO4TXWiKpFPvVGnrePTyy+dj5XX32RDMhZL+N3DtYF"
|
||||
"6EGuPeQYfc41xtb0Pl1zzBftTXldc+VWNmy+XlfXd5NFQjXWqsPPHkSds2gWc3iefa4dcPXH/d+PWnPW9CBqr86gx/cQ"
|
||||
"9AD+d3Nf0d49H30zP4t9GpnveVvD9d1sceVWyZFnVIBTv+/3oKZr+W+v39p/Z+qei7DsnIZZK23z0jdZowJcc/hC8rxp"
|
||||
"wo96+n6dbdz+mqlaPd0SWjNv/vslA+rzfOFl7me7Gaa/8nU7hvWTL95mPnvzqckiL/Qxj+e1Xn/niUl1+b2asTl3L7Nm"
|
||||
"ztp7fO7ryD+HJ7dCNFt/5oifvfrMxHsiL3zqJQcvwLy7AzVh54F4+ALuTNbc0OlzUJ/Uxbye1xnao371RXvyPdEcj2Kf"
|
||||
"fr3+vkZ3Uzn1Rve88DpfMYO/n/WhB7k60ekdmtMd3+daw0MMNTXxsFYtrHXNv2e52n6rBpqjj9OQ93yhCwcv/Vb3nNZz"
|
||||
"fehhL7lDD7C32kP5bs/ZMOKln76C7/vIz3mOriXDw/s/DIjqz9/NA88Ijh7q5HSu5k5vLn6aq3y8bv76PiES5KmhZlTd"
|
||||
"uB5+TjEjyPt56M2IHdda3eesmVvdeSNedc4Mda/he7xxoGZsfrucgIEXrYiEr5XX/ojDeu6eF3eYuYsew/0feyvCz72t"
|
||||
"G7++3Py7sRfPf5wAcqwLKN/3RX7lOLuI1F441QKz979MAGsFtfXHdzJqRPWS9z2LfAO/9cEcplcO3gZYzx7+lCKf6gT7"
|
||||
"tJ95OEf2HO5Pfgt1avwurKPvsGx+xyHWjECnZc1Zw+M18tqvmu/X9Wze2WmqPnm6RmQOdHcD9yOzVg/4KNd+nct5Omvg"
|
||||
"5Pz0nPSsNIfHa9pLzp+5+vR+8T0U70jekmpzb1XeV+Qr3lrgtXpzra527+7FZaAP8ecr2XJGArVyWus64FCv3uz+PzlT"
|
||||
"G078kYacvOoRxx5E5RfNVNy4lgzbl7KBuUafA9Mj9QBomvs68nDm9Ejb5dmwfSkZmCNCVx41AY45Z6hfdeXU68+G5+XP"
|
||||
"NTpP9fBORPfBz9d7FJ0TzyriNKfH3we9J9GZ64zdu9hzKqDc5lo7RAI6o9foJx95yY316jrarzW9vh6by171oN7ZaAtc"
|
||||
"/TZZXH2ULQeQs6bm+9hLD2ud4zXE5y92fw/9ng03VyZWK8D7HD72eH+EMR/4e1vdOzhWGzQHPvouGVTzfuiLIn0E5+6v"
|
||||
"AT5ZnB5rLUcEyCmoax+i9lFnv/d43q9NnlzUp7nfH3X2Rr+NNe+S3o1liLy8o57Tu+x7VY/m+Xfg1/Wzovehb0i1/f1k"
|
||||
"w85GPeSrj5LVBGpy6iev/eS19nOE/x+gaS2i"
|
||||
),
|
||||
},
|
||||
"iq2_xxs": {
|
||||
"source": "blk.0.attn_gate.weight",
|
||||
"block_bytes": 66,
|
||||
"blocks": (
|
||||
"eNoBjAFz/h4NhxlMgsOHqpTEeFmmMeZw4V0WDG0g4Cb0pTZTgvOM3r2oaw9MQEyjfPeuF/uVJwiZksqBaESGaYExuAFT"
|
||||
"KMEBk0MNCEkDIWIGt8+8CRrWkLA2d7ZC/0m3vOOCDi0WUT6Mgv8VisTSWq625wLzKca0Lpy+i2qXT7x15YoTItH5Rzf0"
|
||||
"hmMMII18mY4im+1dZIo77zVm6MUboI8D79T31BB1QJuOy5qNKSG6TWwbwyDmyNH/Kx6iY4bTDpD+o9HE4ACHp1TGvQwM"
|
||||
"yGBr0/KEtaKcEJDJ4GDO/xjDFkyPmje7B5AlhYYbWP8QBfPTejik75531FEJNlH+mGkNxp5IHM3KzeC8Sr+o1pkM9A8t"
|
||||
"sqJ02ZIZDS8LwnHA040JCXza+pedDYUCN94HQvK5kLGPRoUCkZtRJ4dVbQm0J+VwazuGl3caymqF/qftjjYMkfIiNBP2"
|
||||
"4c7htG3cwS21qXiRkebCWsy4lJiFhE8AP6HXO76mrJpLsAYAURyv5J/4toWGKjGmD6mo9CS+VbBguEDGxLA="
|
||||
),
|
||||
"expected": (
|
||||
"eNp1VE2IVmUUvhhhCMlQQi7Shv60gppVi865NauiRSBRKUg0i4wkSBfiJnJRhn9BhgsJUobQfikGixx9z4XRLHSiEnUR"
|
||||
"ljpEmJnJV0EOktY5r+fceeb1tng45zznec77c9/vq1btT72dM5rhs1u5WrWfAEk5sRiwukNH3ifXh5fBx9bTdWrr23qe"
|
||||
"p+q5OTJ06hfWGBBD3/JHauUtZ++LI9fab4AT8Oe+e6Xw5Wizo+4/cYFH7nqyrt6Zz5Z7FM8F+SKXgOnDF56Azm6wDm+s"
|
||||
"WT37WzL0n95tOVk+sGB+45x4n7yfe6gFT42ewisGn0sWY41q7VuEGDy5goce/ZY1T5qL1cFbrT0DAy8WjY8ckBzk8xjy"
|
||||
"K1p/O/Ym4g3Z27A341wqQKGD90bxLuHdte8w3h28y6k3+M/HZBge/1UsTpxbx86lyLXHmls/WQ5cqw1/zHJdArQzwZvi"
|
||||
"rHAm7jhXe444C/ymmvJuyt+ln7eG33R7Xwf3D9TVpmvEIzsEwAMbvkGOgcuIGYWmjYXf6lY3vGM1V8e3UaB38zExziHK"
|
||||
"JeXYALoEINeFV0CTOZyNPpvZt6wn1eldyTDyxSZWCObeJ+yhx3vTZmg/ELOiT+DhXM+7jSfvXdosOnFZDFoHGGJAXIcc"
|
||||
"uy/6XHgYZrVrROy9cJ6rrXNFkaPWbe6Rix5qc6/w5PzMrMeacnaBrDl8U8OGatafSUEe23x0+62N5R5zzz1inHvJNbX3"
|
||||
"pZhD4Qu+9eo3sG9VfCf8Rsm/ef6e/n2lQys+S/w7l+9CYh3smT7eifcJ30nk+IZwLnpiv7hXXL/rPL0vv+LqxcsJQM4R"
|
||||
"cmf+uLsBPve0FqydY+NN7746OOfr/9G280LbsQ7uRaIOlByuizrcQ3VpphgWvaT/L5dmtphM99fBad5oLhFdk33God/6"
|
||||
"nk/jUOvI/MC/pzKqXe+SIjksJ+XZ6j56hU2jUbAPCF8Kj8fcD5/NQU3uP/VDUtDBp29sBpZ+JJY7ktZsCM40wbsvddTk"
|
||||
"c3Ktnto5Rn9gZe8NqR4cTYgubtGWC6I8RwxOIwFaL+g5+uDLvGk39+3h6qdzBEjfL7ildj45SGuBOsM0qm0Kvp1lc2Ce"
|
||||
"+Ixpa1Xjh9LhndvEooIcaXTLzFp5Dt5r01FE5ZrS11FnXbHG1LqTH5AieaSxQ2uNS4PLfhbN2aLxGtl6xgVcbzFFz/Sh"
|
||||
"dV+5RrtW7p25SBNbFtYWA4PXHjA+aeSIpvE6dAmjaxqvW84hCOubNvdWvCmBauHepJE1EiCN3XHRuGS15mIwXWjD634J"
|
||||
"H4AKHbU6P4PvnTqQ/A7E7wnPJ8W9Md5fcZ8E3nZmNXqWxp7ZLcOfzqstap0QyrFrciz75jM+YngMyjWGLl/rhX3GHcA5"
|
||||
"GL75tHOH1u4E7mXavcG7Sh1v68qdLBmn/r/0d7RkPI3N2S6eZ15rNs4icKapgUugZa8DhDBveGydXK8/TkNHb2gG937I"
|
||||
"mifNa+DEeIvWC1gNHgpf1K4n5yQ0HXO5+vwTGlizXjy2teWbD5zl6GGuSA4qo/lcF3zuqV86dQ9PkiJZHPl7HxuQ85g6"
|
||||
"em3/uu/urD2/aobnMUPAdwVT+yv3nfw+cq/Y/zRYz/td90Loi3O389bNZoX0dp3MsDy4APDRy1C+BXgYgLOy7vnXHmpA"
|
||||
"z3BW2xPjt8Hv7e9B8Eyuy3zcVdnDd9Xhl2rN0aQgjCsf2MkQy555KKJx78+dXaMeexatD7NaPuvOb0iHF++xSBOvLs61"
|
||||
"R7LcIwWvWtbIxkGeZ4TfezGLY7b7xbWZrx5/m/oHV4nGNDD7iETuoJLXnI0HrtVYzzjvcfRgJsHslDWvH0kOMgy9vEO6"
|
||||
"ctQZb+hbfX0dfOHj8DifQh/eds0fR5KCAhP3bGSD8cPv/R5RrOcxuU4MrsE5qdCmmAfeqfmf7aPhJ7aKxYDW3Lt9RuN1"
|
||||
"wujayBn61NGPucnm+dw6fHmNqX3lc/reyPcd5+bYN5w3zkHF+RKcW2JGcUdT/unnSHgm37MU50xwNwR3kwofu4+K+7J+"
|
||||
"3d7PxhkyfOHrDMsVXMTIuSO/CjGnmNd6Vy6/r8Y63ha8PYk3E3dXvp/y+0CeineG/niTOK/zffn9xLegjiigQw+hrvwu"
|
||||
"AX+DecZ/P39gQw=="
|
||||
),
|
||||
},
|
||||
"iq2_xs": {
|
||||
"source": "blk.2.ssm_out.weight",
|
||||
"block_bytes": 74,
|
||||
"blocks": (
|
||||
"eNoBvAFD/ooQA6AAiCN6lmuU9VeqmJ0A5LvON22xpJNoA/QAWAGcDZaWQKyDoBePECMoCX4R2nrRr2MDSLwulHU1HAI0"
|
||||
"CJIA7qhpU211f1h7XxAAzCTiWUXSqgJfCmAn5Z2fsULmOlXie4P+2Hs6K4JXIkNMR9BMRi0ZH7TGIaAzha65Ev8svrLU"
|
||||
"1ifnCjik2AR9VzVCZmaZ9mj0FNMeyOaKjppuTNaS4wbCEv+Gcmfqfj5EnRP95Ooudk/83EyYhB7dAMgAhGrkfgw9mb+0"
|
||||
"iVCEeyeZATYADOqTF9s0MyQ0hTxVjwANNM9T0JRdV5xbUJAyTRtpVWhT61gGwqYlAPxNcKRXB3CIfKqRCfyjf2kPFAEX"
|
||||
"LmoX/nqif6ITk2ALpVceb5N4Tbpntlv4epaLABIJsJWJzbQAtif3Y3IAVhfoXfpC8GUkXYI35fkDxVr/SmoyAAJJMjgb"
|
||||
"H0hKjGpgJVmrJd94HzoDLeSgUQN6OSomyKRERK+Iw6ljEmOmdWwqnHp6sDpbIynDWhCVbdJGRFvGr1757kJ49dKyb1ZJ"
|
||||
"vCaNS9cwmkYajOIXLCaNWoLNhkaeOkv/ABP8eMYiEiIh58/YfcDzwlE="
|
||||
),
|
||||
"expected": (
|
||||
"eNp9Vl+IncUVHymxoUVYIZYktXpRCLQNuiw0LX5n9BZDjSS1C81Dgwveh6ASKElpK5GK3hppI0RYTRM01nYfUtiaNmyb"
|
||||
"kj/OmeU++CeBPIS2lH0odNOq9UFwzUsjVew5M+fMd+bbmz78+J3zO7+Zb+bMuey6nRthefAd74gF2GH4v/nxu5EAvZV1"
|
||||
"kWNl1oRTPPmFYRyj+eXR9d49sRmVLebefCRKDMrki8YL/YfPglv7fmBQjMSNIAyO3OI1trpB4yafDMMj30JmQkMxcN7R"
|
||||
"sH/+VdRYask7+uwvwZ19Kwg3GhPQ6pyr1+jB3fxjuHTyeGQm4BhApwaW5Y71ncydpSc2rupzd17vFd07CKNoBVrjNW7z"
|
||||
"Why9/lckBolBYuj9d0vUmGF8ID50pw8BY/D3B6NlA/bg7LoYNRbkOu8he9tzGNTfkzqfjf00+5HnuH/7h2WuKU6zrTX5"
|
||||
"feDY30L7nmjf2bw7VO/dIvXeXZwGAiom7j+kWtIHt97uCbHjK3W+A93Fd+7X7YHtRZ3f9hkGClblC1+bjMy9dy6CMFb+"
|
||||
"naOGEOa+eTTFwikXoOioXuZ937gumjxIPVgPa+Tz4vfWL3Hjzv+IEYRT3Id7UBiYB48dR4X1cT56kfqwdEdDDMwSo8mD"
|
||||
"ahIHrvWv7Mn1mVfC6B/0W5x5pZn74V9AcmRmzXCjPvEi54PDfwC3418NIQhKLLXCEuPEmfVefA3FUWP2dfazrB40nqJZ"
|
||||
"n+RprX7LnMH6LdRTrWeN97DfMD6Yfp3edeMmJAAz5VFjZqmDQdFXph6KssaPWQ9clxztN2R91s29tZedPnb7U/fppz9n"
|
||||
"gAC/d9PpOHvuu54QVRPdG2/ypfxX6xQgwOUt76Hy9J/u86wTx3Fe8478VtHMDJq74Jg3y+efnsLhA5+PDI4JwBi+vNdL"
|
||||
"jJXO3pf3Fv9w5q7oFj6H7tQSEkNBm6OB6uy/Vq316L7ZC9396Nu+7HUNjzlbe1ar5/t379rVQDV794QLx9D13gyM4ZfW"
|
||||
"RM6JPeWNaN7WWFeNcjB7p752zyDf6p6j9dp+5l61/RTuvFHtWdnFwOG9X4kcE3vO+99/IbFoqSYaWM09Os9Axmj7P2P/"
|
||||
"+QNec4qj1NVT+RPm/wMEZQu8Vry8e9Ni0cwcpXmgXOYCVs0jQT2lD3IXc+8Ul/u19y39kH5pL9Awai+5V5090PTTq6+3"
|
||||
"84FFd2hDJHhhjS2upXvuue0r5dx/YF1qqumb1O8zQX8T3Z+DoKEcipa5SVh/gy+5XbNwCpx7lxkTu3ebwgc3epMH8UFZ"
|
||||
"w7UB/R11R5VD4sEeTBpj9kIbZz1U/t52+k1tp/jpIGgSWOMaxwsnUfRQ9LberhlXXzgJsh+O0bt3albd1/Yj9ySa2I/x"
|
||||
"12tbT1N6nDmvW/4qxTPMkGJmzt1M4y4dhoRcUz2UOvOlZ8Dte5v2eDVUYD3HTcK+t7OPOWshx19koNs2HRPXgCqfvJKh"
|
||||
"tW3TPs3Q+htiZ/7AzGFTgfUqTvOQMXsByny0CEXjWcrzBWXOdEba+WnnqJ2PUGI7M2kGFilepP78IseZm4TpT6HkzOqp"
|
||||
"fGYm7O+i/k2N+/3IrJk30pjf5dIzKADzrk31dlx3OyBh4ll6ix0ogCoe3OoJsXi1ntYc5Z6j4RzPXoirtLHxUhA0JR7N"
|
||||
"ozBI3JR678ZYa0+gAMawBf39/K03fvHIt3lf/bZ+n3U+g/2+PZ/V1ZfP113Dui/75Xq7nvdr61jd3+6t69mbEaveaV/q"
|
||||
"M4RV32vvRv7XGCDI+cpHvsrbehdY+rq8te1tjmHV2yxvjammdc55ttp5A8nbudR84tm4SrO+ei5rL89p1tDMOeSZnqc+"
|
||||
"XKaeXwYBJi3r4PoHMljvH8hw8yHV2DPYU895m4P8DnxHx1JjrTdF9d1ggEmbfCmWnHnypdrHnt5ULPPNvbT9zm9Sa7n3"
|
||||
"0f4mRkfuQ7f/qUBoGHPHT7AGVuOYfaqzx9Qa8VqPao3sX3nM+iD76ffUA919ZC/Q/dQzHH4QCI3btqGANeIg3Iyufhuk"
|
||||
"lrzGX87YPaf0Qr8XzHkae2a9i/StxJ3zBnOP8s3ky+fVswU6K2qsOZ9fzyxx0Dt13iqYnqK9l/EFe8bphZuiO3YPEntm"
|
||||
"AuwPQ9ZAwbn6Ol50Xz8HBGa8+sdZ/97lq141yqPWjK/UUv2TN5AAH69Zszj/+2OR4DlXXZlrJm49J36N55/7m5/8yeOe"
|
||||
"YuCcQXmUvGjsk7jVN9H/5pt2MePCb77sD77zvFeN8shsNGAPQ9aA+8HvUAAWS2f+HU0Nt539meZAsbc1WYOyzosXDPvu"
|
||||
"/qrzW/DbKMq7yFtqTG/oGeZ9U/1/TYe49w=="
|
||||
),
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def packed_blocks(name: str) -> np.ndarray:
|
||||
"""Packed bytes for ``name`` as ``(n_blocks, block_bytes)`` uint8."""
|
||||
entry = _VECTORS[name]
|
||||
raw = zlib.decompress(base64.b64decode(entry["blocks"]))
|
||||
return np.frombuffer(raw, dtype=np.uint8).reshape(-1, entry["block_bytes"]).copy()
|
||||
|
||||
|
||||
def expected_values(name: str) -> np.ndarray:
|
||||
"""llama.cpp's dequantized output for ``packed_blocks(name)``, as ``(n, 256)``."""
|
||||
raw = zlib.decompress(base64.b64decode(_VECTORS[name]["expected"]))
|
||||
return np.frombuffer(raw, dtype=np.float32).reshape(-1, 256).copy()
|
||||
|
||||
|
||||
def formats() -> list[str]:
|
||||
"""Every format with captured vectors."""
|
||||
return sorted(_VECTORS)
|
||||
@@ -76,12 +76,14 @@ def test_ptq_whisper(command):
|
||||
PTQCommand(quant="int8_weight_only", kv_cache_quant="none"),
|
||||
PTQCommand(quant="int4_awq", kv_cache_quant="none"),
|
||||
PTQCommand(quant="w4a8_awq_beta", kv_cache_quant="none"),
|
||||
# GGML IQ weight-only, recipe-driven. These encoders require every weight's input
|
||||
# dimension to be a multiple of 256; TinyLlama's 2048 and 5632 both are. Neither
|
||||
# recipe calibrates -- both set algorithm: null -- so the only IQ-specific cost is
|
||||
# packing each weight once and decoding it on each forward. 95s and 103s on 2xH100,
|
||||
# inside the 300s tests/examples default.
|
||||
# GGML IQ weight-only, recipe-driven: three formats between 1.56 and 2.31 bits
|
||||
# per weight. These encoders require every weight's input dimension to be a
|
||||
# multiple of 256; TinyLlama's 2048 and 5632 both are. None of them calibrates --
|
||||
# every recipe sets algorithm: null -- so the only IQ-specific cost is packing each
|
||||
# weight once on a CUDA encoder and decoding it on each forward, which fits the
|
||||
# 300s tests/examples default.
|
||||
PTQCommand(recipe="general/ptq/iq1_s", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq2_xxs", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq2_xs", kv_cache_quant="none"),
|
||||
PTQCommand(quant="nvfp4"),
|
||||
PTQCommand(quant="nvfp4_awq_lite"),
|
||||
|
||||
@@ -0,0 +1,179 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""CUDA encoders for every GGML IQ format.
|
||||
|
||||
Parametrized over the family rather than written per format, so a behaviour asserted
|
||||
for one is asserted for all. IQ1_S is the odd one out only in its entry point: it
|
||||
derives the block scale in its own kernel instead of taking a precomputed one.
|
||||
"""
|
||||
|
||||
import pytest
|
||||
import torch
|
||||
|
||||
import modelopt.torch.quantization.ggml.iq1_s as iq1_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xs as iq2_xs_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xxs as iq2_xxs_module
|
||||
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
IQ2_XXS_BLOCK_BYTES,
|
||||
)
|
||||
|
||||
# module, packer name, per-block payload size, whether the packer takes precomputed scales
|
||||
FORMATS = {
|
||||
"iq1_s": (iq1_s_module, "iq1_s_pack", IQ1_S_BLOCK_BYTES, False),
|
||||
"iq2_xxs": (iq2_xxs_module, "iq2_xxs_pack", IQ2_XXS_BLOCK_BYTES, True),
|
||||
"iq2_xs": (iq2_xs_module, "iq2_xs_pack", IQ2_XS_BLOCK_BYTES, True),
|
||||
}
|
||||
|
||||
|
||||
def _extension():
|
||||
extension = get_cuda_ext_ggml(raise_if_failed=True)
|
||||
assert extension is not None
|
||||
return extension
|
||||
|
||||
|
||||
def _pack(name, weight):
|
||||
module, packer, _, takes_scales = FORMATS[name]
|
||||
grid = getattr(module, f"{name}_grid")("cuda")
|
||||
if not takes_scales:
|
||||
return getattr(_extension(), packer)(weight, grid)
|
||||
blocks = weight.contiguous().reshape(-1, 256)
|
||||
scales = getattr(module, f"_predict_{name}_scales")(blocks)
|
||||
return getattr(_extension(), packer)(weight, grid, scales)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_pack_matches_pytorch_encoder_and_is_decodable(monkeypatch, name):
|
||||
"""Byte parity with the reference encoder on a fixed small weight.
|
||||
|
||||
Exact parity is asserted on this input, not in general: the two encoders evaluate the same
|
||||
squared error with different floating-point fusion, so where two local scales fall within a
|
||||
float32 ULP they can round to different sides. That happens in roughly one block in several
|
||||
thousand, costs under 1e-8 of relative reconstruction error, and favours neither encoder --
|
||||
see ``test_cuda_pack_reconstruction_matches_pytorch_at_scale``.
|
||||
"""
|
||||
module, _, block_bytes, _ = FORMATS[name]
|
||||
generator = torch.Generator(device="cuda").manual_seed(1234)
|
||||
weight = torch.randn((8, 512), generator=generator, device="cuda", dtype=torch.bfloat16)
|
||||
|
||||
packed = _pack(name, weight).reshape(8, 2, block_bytes)
|
||||
packed_again = _pack(name, weight).reshape(8, 2, block_bytes)
|
||||
monkeypatch.setattr(module, "get_cuda_ext_ggml", lambda: None)
|
||||
reference, shape = getattr(module, f"quantize_{name}")(weight)
|
||||
reconstructed = getattr(module, f"dequantize_{name}")(packed, shape)
|
||||
|
||||
assert packed.shape == (8, 2, block_bytes)
|
||||
assert torch.equal(packed, packed_again)
|
||||
assert torch.equal(packed, reference)
|
||||
assert shape.device.type == "cpu"
|
||||
normalized_mse = (
|
||||
reconstructed.float() - weight.float()
|
||||
).square().mean() / weight.float().square().mean()
|
||||
assert normalized_mse < 0.25
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_pack_is_deterministic_on_one_device(name):
|
||||
"""The contract reproducibility actually needs: same machine, same bytes."""
|
||||
generator = torch.Generator(device="cuda").manual_seed(7)
|
||||
weight = torch.randn((64, 1024), generator=generator, device="cuda", dtype=torch.bfloat16)
|
||||
first = _pack(name, weight)
|
||||
for _ in range(3):
|
||||
assert torch.equal(_pack(name, weight), first)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_pack_reconstruction_matches_pytorch_at_scale(monkeypatch, name):
|
||||
"""Over many blocks the encoders may disagree on a near-tied scale, but not on quality."""
|
||||
module, _, block_bytes, _ = FORMATS[name]
|
||||
generator = torch.Generator(device="cuda").manual_seed(11)
|
||||
weight = torch.randn((128, 2048), generator=generator, device="cuda", dtype=torch.bfloat16)
|
||||
blocks = weight.shape[0] * weight.shape[1] // 256
|
||||
|
||||
packed = _pack(name, weight).reshape(weight.shape[0], weight.shape[1] // 256, block_bytes)
|
||||
monkeypatch.setattr(module, "get_cuda_ext_ggml", lambda: None)
|
||||
reference, shape = getattr(module, f"quantize_{name}")(weight)
|
||||
|
||||
dequantize = getattr(module, f"dequantize_{name}")
|
||||
target = weight.float()
|
||||
denominator = target.square().sum()
|
||||
cuda_error = (
|
||||
dequantize(packed, shape, dtype=torch.float32) - target
|
||||
).square().sum() / denominator
|
||||
torch_error = (
|
||||
dequantize(reference, shape, dtype=torch.float32) - target
|
||||
).square().sum() / denominator
|
||||
|
||||
differing = int((packed != reference).any(dim=-1).sum())
|
||||
assert differing <= blocks // 1000, f"{differing} of {blocks} blocks differ"
|
||||
assert torch.isclose(cuda_error, torch_error, rtol=1e-5)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_zero_encoding_matches_ggml_block_layout(name):
|
||||
module, _, block_bytes, _ = FORMATS[name]
|
||||
weight = torch.zeros((1, 256), device="cuda", dtype=torch.bfloat16)
|
||||
packed = _pack(name, weight).reshape(1, 1, block_bytes)
|
||||
shape = torch.tensor(weight.shape, device="cuda")
|
||||
|
||||
assert not packed.any()
|
||||
assert torch.equal(getattr(module, f"dequantize_{name}")(packed, shape), weight)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_nonfinite_policy_matches_pytorch_encoder(monkeypatch, name):
|
||||
module, _, _, _ = FORMATS[name]
|
||||
weight = torch.zeros((1, 256), device="cuda", dtype=torch.float32)
|
||||
weight[0, 0] = float("nan")
|
||||
weight[0, 1] = float("inf")
|
||||
weight[0, 2] = float("-inf")
|
||||
|
||||
packed, _ = getattr(module, f"quantize_{name}")(weight)
|
||||
monkeypatch.setattr(module, "get_cuda_ext_ggml", lambda: None)
|
||||
reference, _ = getattr(module, f"quantize_{name}")(weight)
|
||||
assert torch.equal(packed, reference)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_falls_back_to_pytorch_encoder(monkeypatch, name):
|
||||
"""Without the extension the format still packs, through the torch search."""
|
||||
module, _, block_bytes, _ = FORMATS[name]
|
||||
monkeypatch.setattr(module, "get_cuda_ext_ggml", lambda: None)
|
||||
generator = torch.Generator(device="cuda").manual_seed(5)
|
||||
weight = torch.randn((2, 256), generator=generator, device="cuda", dtype=torch.bfloat16)
|
||||
|
||||
packed, shape = getattr(module, f"quantize_{name}")(weight)
|
||||
reconstructed = getattr(module, f"dequantize_{name}")(packed, shape)
|
||||
normalized_mse = (
|
||||
reconstructed.float() - weight.float()
|
||||
).square().mean() / weight.float().square().mean()
|
||||
|
||||
assert packed.shape == (2, 1, block_bytes)
|
||||
assert normalized_mse < 0.25
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", sorted(FORMATS))
|
||||
def test_cuda_float64_matches_pytorch_encoder(monkeypatch, name):
|
||||
module, _, _, _ = FORMATS[name]
|
||||
generator = torch.Generator(device="cuda").manual_seed(99)
|
||||
weight = torch.randn((2, 256), generator=generator, device="cuda", dtype=torch.float64)
|
||||
|
||||
packed, _ = getattr(module, f"quantize_{name}")(weight)
|
||||
monkeypatch.setattr(module, "get_cuda_ext_ggml", lambda: None)
|
||||
reference, _ = getattr(module, f"quantize_{name}")(weight)
|
||||
assert torch.equal(reference, packed)
|
||||
@@ -43,11 +43,12 @@ from transformers.models.qwen3_vl.modeling_qwen3_vl import Qwen3VLForConditional
|
||||
|
||||
import modelopt.torch.export.unified_export_megatron as uem
|
||||
import modelopt.torch.quantization as mtq
|
||||
import modelopt.torch.quantization.ggml as ggml
|
||||
import modelopt.torch.speculative as mtsp
|
||||
from modelopt.torch.export import KV_CACHE_FP8, export_mcore_gpt_to_hf, import_mcore_gpt_from_hf
|
||||
from modelopt.torch.export.quant_format import IQ_FORMATS
|
||||
from modelopt.torch.export.unified_export_megatron import GPTModelExporter
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig
|
||||
from modelopt.torch.quantization.ggml import dequantize_iq1_s, dequantize_iq2_xs, quantize_iq2_xs
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
from modelopt.torch.speculative.eagle.default_config import default_eagle_config
|
||||
from modelopt.torch.speculative.plugins.megatron_eagle import _DynamicEagleGPTModel
|
||||
@@ -89,12 +90,17 @@ def _verify_model_quant_config(
|
||||
assert quant_config_dict["kv_cache_quant_algo"] == KV_CACHE_FP8
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("qformat", "payload_bytes", "dequantize"),
|
||||
[("iq1_s", 50, dequantize_iq1_s), ("iq2_xs", 74, dequantize_iq2_xs)],
|
||||
)
|
||||
def test_megatron_name_remapping_exports_iq_payload(qformat, payload_bytes, dequantize):
|
||||
# Every IQ format the exporter accepts. Only the list of formats comes from the export
|
||||
# tables; each test resolves what it expects from the codec module itself, so a wrong entry
|
||||
# in IQ_PACKERS or IQ_BLOCK_METADATA cannot make both sides of an assertion agree.
|
||||
IQ_FORMAT_NAMES = sorted(IQ_FORMATS)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_name_remapping_exports_iq_payload(qformat):
|
||||
"""Megatron export writes the same scale-free IQ representation as HF export."""
|
||||
payload_bytes = getattr(ggml, f"{qformat.upper()}_BLOCK_BYTES")
|
||||
dequantize = getattr(ggml, f"dequantize_{qformat}")
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
@@ -112,20 +118,23 @@ def test_megatron_name_remapping_exports_iq_payload(qformat, payload_bytes, dequ
|
||||
exporter._name_remapping(linear, "model.layers.0.mlp.down_proj.")
|
||||
|
||||
packed_key = "model.layers.0.mlp.down_proj.weight"
|
||||
assert exporter._state_dict[packed_key].shape == (2, 1, payload_bytes)
|
||||
assert exporter._state_dict[packed_key].dtype == torch.uint8
|
||||
logical_shape = torch.tensor(
|
||||
[
|
||||
*exporter._state_dict[packed_key].shape[:-2],
|
||||
exporter._state_dict[packed_key].shape[-2] * 256,
|
||||
]
|
||||
packed = exporter._state_dict[packed_key]
|
||||
assert packed.shape == (2, 1, payload_bytes)
|
||||
assert packed.dtype == torch.uint8
|
||||
# Exact bytes against the format's own packer, as the slicing tests below check.
|
||||
_assert_iq_payload_matches(qformat, packed, linear.weight)
|
||||
# And the payload decodes to exactly what the fake quantizer reconstructs. Compare with the
|
||||
# decoded reference, not the fake-quant forward: that returns the straight-through form
|
||||
# a + (r - a), which in bf16 differs from r by up to one ULP of a -- enough to fail a
|
||||
# relative tolerance wherever r is small next to a, as IQ1_S's grid near zero often is.
|
||||
logical_shape = torch.tensor([*packed.shape[:-2], packed.shape[-2] * 256])
|
||||
reference, _ = getattr(ggml, f"quantize_{qformat}")(linear.weight)
|
||||
torch.testing.assert_close(
|
||||
dequantize(packed, logical_shape, dtype=torch.bfloat16),
|
||||
dequantize(reference, logical_shape, dtype=torch.bfloat16),
|
||||
rtol=0,
|
||||
atol=0,
|
||||
)
|
||||
reconstructed = dequantize(
|
||||
exporter._state_dict[packed_key],
|
||||
logical_shape,
|
||||
dtype=torch.bfloat16,
|
||||
)
|
||||
torch.testing.assert_close(reconstructed, linear.weight_quantizer(linear.weight))
|
||||
assert exporter.layer_config_dict == {
|
||||
"model.layers.0.mlp.down_proj.quantization": qformat,
|
||||
"model.layers.0.mlp.down_proj.awq_block_size": 256,
|
||||
@@ -167,28 +176,30 @@ def _make_iq_weight(rows):
|
||||
return torch.linspace(-1, 1, rows * 256, dtype=torch.float32).reshape(rows, 256).bfloat16()
|
||||
|
||||
|
||||
def _assert_iq2_payload_matches(packed, logical_weight):
|
||||
expected, _ = quantize_iq2_xs(logical_weight)
|
||||
def _assert_iq_payload_matches(qformat, packed, logical_weight):
|
||||
expected, _ = getattr(ggml, f"quantize_{qformat}")(logical_weight)
|
||||
torch.testing.assert_close(packed, expected.cpu(), rtol=0, atol=0)
|
||||
|
||||
|
||||
def test_megatron_gated_mlp_slicing_exports_iq_payloads():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_gated_mlp_slicing_exports_iq_payloads(qformat):
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(config=SimpleNamespace(ffn_hidden_size=4))
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, qformat, 256)
|
||||
|
||||
exporter._gated_mlp_slicing(module, "model.layers.0.mlp.")
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.gate_proj.weight"], weight[:4]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mlp.gate_proj.weight"], weight[:4]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.up_proj.weight"], weight[4:]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mlp.up_proj.weight"], weight[4:]
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_grouped_mlp_slicing_exports_iq_payloads():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_grouped_mlp_slicing_exports_iq_payloads(qformat):
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(
|
||||
num_gemms=1,
|
||||
@@ -199,7 +210,7 @@ def test_megatron_grouped_mlp_slicing_exports_iq_payloads():
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: (
|
||||
{"weight": module.weight},
|
||||
"iq2_xs",
|
||||
qformat,
|
||||
256,
|
||||
)
|
||||
|
||||
@@ -210,15 +221,16 @@ def test_megatron_grouped_mlp_slicing_exports_iq_payloads():
|
||||
up_proj_name="up_proj",
|
||||
)
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.experts.0.gate_proj.weight"], weight[:4]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mlp.experts.0.gate_proj.weight"], weight[:4]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mlp.experts.0.up_proj.weight"], weight[4:]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mlp.experts.0.up_proj.weight"], weight[4:]
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_qkv_slicing_exports_iq_payloads():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_qkv_slicing_exports_iq_payloads(qformat):
|
||||
weight = _make_iq_weight(8)
|
||||
module = SimpleNamespace(
|
||||
config=SimpleNamespace(
|
||||
@@ -230,7 +242,7 @@ def test_megatron_qkv_slicing_exports_iq_payloads():
|
||||
)
|
||||
)
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, qformat, 256)
|
||||
|
||||
exporter._qkv_slicing(module, "model.layers.0.self_attn.")
|
||||
|
||||
@@ -241,13 +253,15 @@ def test_megatron_qkv_slicing_exports_iq_payloads():
|
||||
"v_proj": reshaped[3].reshape(2, 256),
|
||||
}
|
||||
for projection, logical_weight in expected.items():
|
||||
_assert_iq2_payload_matches(
|
||||
_assert_iq_payload_matches(
|
||||
qformat,
|
||||
exporter._state_dict[f"model.layers.0.self_attn.{projection}.weight"],
|
||||
logical_weight,
|
||||
)
|
||||
|
||||
|
||||
def test_megatron_gated_delta_net_slicing_exports_iq_payloads():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_gated_delta_net_slicing_exports_iq_payloads(qformat):
|
||||
weight = _make_iq_weight(12)
|
||||
module = SimpleNamespace(
|
||||
in_proj=object(),
|
||||
@@ -255,15 +269,15 @@ def test_megatron_gated_delta_net_slicing_exports_iq_payloads():
|
||||
in_proj_split_sections=(2, 2, 2, 2, 2, 2),
|
||||
)
|
||||
exporter = _make_iq_exporter()
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, "iq2_xs", 256)
|
||||
exporter._get_quantized_state = lambda *a, **k: ({"weight": weight}, qformat, 256)
|
||||
|
||||
exporter._gated_delta_net_slicing(module, "model.layers.0.mixer.")
|
||||
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_qkv.weight"], weight[:6]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mixer.in_proj_qkv.weight"], weight[:6]
|
||||
)
|
||||
_assert_iq2_payload_matches(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_z.weight"], weight[6:8]
|
||||
_assert_iq_payload_matches(
|
||||
qformat, exporter._state_dict["model.layers.0.mixer.in_proj_z.weight"], weight[6:8]
|
||||
)
|
||||
torch.testing.assert_close(
|
||||
exporter._state_dict["model.layers.0.mixer.in_proj_b.weight"], weight[8:10]
|
||||
@@ -273,7 +287,7 @@ def test_megatron_gated_delta_net_slicing_exports_iq_payloads():
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("qformat", ["iq1_s", "iq2_xs"])
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_packed_experts_reject_iq_without_deployment_loader(qformat):
|
||||
experts = _make_iq_experts(qformat, "linear_fc2")
|
||||
exporter = _make_iq_exporter()
|
||||
@@ -287,8 +301,9 @@ def test_megatron_packed_experts_reject_iq_without_deployment_loader(qformat):
|
||||
assert exporter._state_dict == {}
|
||||
|
||||
|
||||
def test_megatron_gpt_oss_packed_experts_reject_iq_without_deployment_loader():
|
||||
experts = _make_iq_experts("iq2_xs", "linear_fc1", bias=True)
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_gpt_oss_packed_experts_reject_iq_without_deployment_loader(qformat):
|
||||
experts = _make_iq_experts(qformat, "linear_fc1", bias=True)
|
||||
exporter = _make_iq_exporter()
|
||||
|
||||
with pytest.raises(NotImplementedError, match="Fused-MoE IQ export requires"):
|
||||
@@ -300,12 +315,13 @@ def test_megatron_gpt_oss_packed_experts_reject_iq_without_deployment_loader():
|
||||
assert exporter._state_dict == {}
|
||||
|
||||
|
||||
def test_megatron_iq_export_rejects_tensor_parallelism():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_iq_export_rejects_tensor_parallelism(qformat):
|
||||
"""IQ packing is intentionally limited to complete TP=1 weights."""
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits="iq2_xs",
|
||||
num_bits=qformat,
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
@@ -324,7 +340,8 @@ def test_megatron_iq_export_rejects_tensor_parallelism():
|
||||
exporter.save_pretrained("unused", "unused")
|
||||
|
||||
|
||||
def test_megatron_iq_export_rejects_pipeline_parallelism():
|
||||
@pytest.mark.parametrize("qformat", IQ_FORMAT_NAMES)
|
||||
def test_megatron_iq_export_rejects_pipeline_parallelism(qformat):
|
||||
"""IQ packing requires PP=1 so the fused-MoE rejection reaches every rank.
|
||||
|
||||
The rejection raises from inside the per-expert loops, so a stage owning no expert would
|
||||
@@ -334,7 +351,7 @@ def test_megatron_iq_export_rejects_pipeline_parallelism():
|
||||
linear = torch.nn.Linear(256, 2, bias=False, dtype=torch.bfloat16)
|
||||
linear.weight_quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(
|
||||
num_bits="iq2_xs",
|
||||
num_bits=qformat,
|
||||
block_sizes={-1: 256},
|
||||
backend="ggml",
|
||||
)
|
||||
|
||||
@@ -38,6 +38,8 @@ from modelopt.torch.quantization.ggml import (
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_XS_BLOCK_SIZE,
|
||||
IQ2_XS_EFFECTIVE_BITS,
|
||||
IQ2_XXS_BLOCK_SIZE,
|
||||
IQ2_XXS_EFFECTIVE_BITS,
|
||||
)
|
||||
|
||||
|
||||
@@ -135,6 +137,7 @@ def test_mlp_weight_only_recipe_matches_its_mtq_cfg(recipe_name, cfg_name):
|
||||
("qformat", "block_size", "effective_bits"),
|
||||
[
|
||||
("iq1_s", IQ1_S_BLOCK_SIZE, IQ1_S_EFFECTIVE_BITS),
|
||||
("iq2_xxs", IQ2_XXS_BLOCK_SIZE, IQ2_XXS_EFFECTIVE_BITS),
|
||||
("iq2_xs", IQ2_XS_BLOCK_SIZE, IQ2_XS_EFFECTIVE_BITS),
|
||||
],
|
||||
)
|
||||
|
||||
@@ -15,7 +15,9 @@
|
||||
|
||||
import pytest
|
||||
|
||||
import modelopt.torch.quantization.ggml as ggml
|
||||
from modelopt.torch.export.convert_hf_config import convert_hf_quant_config_format
|
||||
from modelopt.torch.export.quant_format import IQ_BLOCK_METADATA, IQ_FORMATS
|
||||
from modelopt.torch.export.unified_export_hf import _revert_hf_quant_config_names
|
||||
|
||||
|
||||
@@ -110,3 +112,104 @@ def test_reverse_quant_config_name_mapping_is_atomic():
|
||||
"kv_cache_quantized_layers": {"model.bad": {"quant_algo": "FP8"}},
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fmt", sorted(IQ_FORMATS))
|
||||
def test_iq_config_carries_block_metadata(fmt):
|
||||
"""Every IQ format must describe its packed block, not just name itself.
|
||||
|
||||
A consumer reads group_size and block_payload_bytes to walk the payload, so a format
|
||||
that falls through to the generic branch produces a checkpoint that cannot be decoded.
|
||||
"""
|
||||
block_size, payload_bytes, effective_bits = IQ_BLOCK_METADATA[fmt]
|
||||
converted = convert_hf_quant_config_format(
|
||||
{
|
||||
"producer": {"name": "modelopt", "version": "test"},
|
||||
"quantization": {"quant_algo": fmt.upper()},
|
||||
}
|
||||
)
|
||||
|
||||
assert converted["quant_algo"] == fmt.upper()
|
||||
assert converted["group_size"] == block_size
|
||||
assert converted["block_payload_bytes"] == payload_bytes
|
||||
assert converted["effective_bits"] == pytest.approx(effective_bits)
|
||||
assert converted["packing"] == "ggml"
|
||||
# IQ payloads are self-contained blocks, not compressed-tensors integer groups.
|
||||
assert "config_groups" not in converted
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fmt", sorted(IQ_FORMATS))
|
||||
def test_iq_config_rejects_mismatched_group_size(fmt):
|
||||
"""A caller's group size is rejected rather than silently rewritten to the block size."""
|
||||
block_size, _, _ = IQ_BLOCK_METADATA[fmt]
|
||||
with pytest.raises(ValueError, match=f"requires group size {block_size}"):
|
||||
convert_hf_quant_config_format(
|
||||
{
|
||||
"producer": {"name": "modelopt", "version": "test"},
|
||||
"quantization": {"quant_algo": fmt.upper(), "group_size": block_size // 2},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fmt", sorted(IQ_FORMATS))
|
||||
def test_iq_block_metadata_matches_the_codec(fmt):
|
||||
"""The exported geometry is the codec's own, so a checkpoint cannot claim a wrong layout."""
|
||||
block_size, payload_bytes, effective_bits = IQ_BLOCK_METADATA[fmt]
|
||||
upper = fmt.upper()
|
||||
assert block_size == getattr(ggml, f"{upper}_BLOCK_SIZE")
|
||||
assert payload_bytes == getattr(ggml, f"{upper}_BLOCK_BYTES")
|
||||
assert effective_bits == pytest.approx(getattr(ggml, f"{upper}_EFFECTIVE_BITS"))
|
||||
assert effective_bits == pytest.approx(payload_bytes * 8 / block_size)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fmt", sorted(IQ_FORMATS))
|
||||
def test_iq_mixed_precision_config_group_carries_block_metadata(fmt):
|
||||
"""A per-layer IQ config must describe its block too, not only a uniform one.
|
||||
|
||||
Mixed exports route each distinct layer config through the same helper, so a format
|
||||
missing there loses its geometry for exactly the layers that use it.
|
||||
"""
|
||||
block_size, payload_bytes, effective_bits = IQ_BLOCK_METADATA[fmt]
|
||||
converted = convert_hf_quant_config_format(
|
||||
{
|
||||
"producer": {"name": "modelopt", "version": "test"},
|
||||
"quantization": {
|
||||
"quant_algo": "MIXED_PRECISION",
|
||||
"quantized_layers": {
|
||||
"model.layers.0.mlp.gate_proj": {"quant_algo": fmt.upper()},
|
||||
"model.layers.1.mlp.up_proj": {"quant_algo": "FP8"},
|
||||
},
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
groups = converted["config_groups"].values()
|
||||
iq_group = next(g for g in groups if g.get("quant_algo") == fmt.upper())
|
||||
assert iq_group["group_size"] == block_size
|
||||
assert iq_group["block_payload_bytes"] == payload_bytes
|
||||
assert iq_group["effective_bits"] == pytest.approx(effective_bits)
|
||||
assert iq_group["packing"] == "ggml"
|
||||
assert iq_group["targets"] == ["model.layers.0.mlp.gate_proj"]
|
||||
# The FP8 layer keeps its own compressed-tensors scheme.
|
||||
assert any("weights" in g for g in groups)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("fmt", sorted(IQ_FORMATS))
|
||||
def test_iq_mixed_precision_rejects_bad_per_layer_group_size(fmt):
|
||||
"""A per-layer group size is validated, not silently rewritten to the block size."""
|
||||
block_size, _, _ = IQ_BLOCK_METADATA[fmt]
|
||||
with pytest.raises(ValueError, match=f"requires group size {block_size}"):
|
||||
convert_hf_quant_config_format(
|
||||
{
|
||||
"producer": {"name": "modelopt", "version": "test"},
|
||||
"quantization": {
|
||||
"quant_algo": "MIXED_PRECISION",
|
||||
"quantized_layers": {
|
||||
"model.layers.0.mlp.gate_proj": {
|
||||
"quant_algo": fmt.upper(),
|
||||
"group_size": block_size // 2,
|
||||
}
|
||||
},
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
@@ -82,7 +82,8 @@ def test_ggml_backend_forwards_chunk_sizes(monkeypatch, extra_args):
|
||||
received.update(kwargs)
|
||||
return inputs
|
||||
|
||||
monkeypatch.setattr(backend_module, "iq1_s_fake_quant", fake_quant)
|
||||
# The dispatcher resolves through its registry, so that is the seam to patch.
|
||||
monkeypatch.setitem(backend_module._FAKE_QUANTS, "iq1_s", fake_quant)
|
||||
inputs = torch.ones(1, 256)
|
||||
quantizer = SimpleNamespace(num_bits="iq1_s", backend_extra_args=extra_args)
|
||||
|
||||
|
||||
@@ -0,0 +1,274 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""Behaviour every GGML IQ format shares, exercised identically for each of them.
|
||||
|
||||
The formats differ only in codebook size, payload layout and bits per weight. Anything that
|
||||
should hold for one should hold for all, so the contract lives here once and is parametrized
|
||||
rather than duplicated per format -- a new format is a row in ``FORMATS``.
|
||||
"""
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
import torch
|
||||
from _test_utils.torch.quantization.iq_llama_cpp_vectors import (
|
||||
expected_values,
|
||||
formats,
|
||||
packed_blocks,
|
||||
)
|
||||
|
||||
import modelopt.torch.quantization.ggml.iq1_s as iq1_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xs as iq2_xs_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xxs as iq2_xxs_module
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
|
||||
# name -> (module, packed bytes per block, codebook entries, bits per weight)
|
||||
FORMATS = {
|
||||
"iq1_s": (iq1_s_module, 50, 2048, 1.5625),
|
||||
"iq2_xxs": (iq2_xxs_module, 66, 256, 2.0625),
|
||||
"iq2_xs": (iq2_xs_module, 74, 512, 2.3125),
|
||||
}
|
||||
NAMES = sorted(FORMATS)
|
||||
# IQ1 grids are ternary; IQ2 grids hold the magnitudes 8, 25 and 43.
|
||||
TERNARY = {"iq1_s"}
|
||||
|
||||
|
||||
def _parts(name):
|
||||
module, block_bytes, entries, bits = FORMATS[name]
|
||||
return (
|
||||
module,
|
||||
getattr(module, f"quantize_{name}"),
|
||||
getattr(module, f"dequantize_{name}"),
|
||||
getattr(module, f"{name}_grid"),
|
||||
block_bytes,
|
||||
entries,
|
||||
bits,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_canonical_grid(name):
|
||||
_, _, _, grid_fn, _, entries, _ = _parts(name)
|
||||
grid = grid_fn()
|
||||
|
||||
assert grid.shape == (entries, 8)
|
||||
assert grid.dtype == torch.float32
|
||||
if name in TERNARY:
|
||||
assert set(grid.unique().tolist()) == {-1.0, 0.0, 1.0}
|
||||
assert grid[0].tolist() == [-1.0] * 8
|
||||
else:
|
||||
assert set(grid.unique().tolist()) <= {8.0, 25.0, 43.0}
|
||||
assert grid[0].tolist() == [8.0] * 8
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_grid_normalizes_unindexed_cuda_device(monkeypatch, name):
|
||||
module, _, _, grid_fn, _, _, _ = _parts(name)
|
||||
cached = torch.empty(0)
|
||||
monkeypatch.setattr(torch.cuda, "current_device", lambda: 7)
|
||||
monkeypatch.setitem(module._GRID_CACHE, torch.device("cuda", 7), cached)
|
||||
|
||||
assert grid_fn("cuda") is cached
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_effective_bits_matches_the_payload(name):
|
||||
module, _, _, _, block_bytes, _, bits = _parts(name)
|
||||
assert getattr(module, f"{name.upper()}_BLOCK_BYTES") == block_bytes
|
||||
assert getattr(module, f"{name.upper()}_EFFECTIVE_BITS") == pytest.approx(bits)
|
||||
assert getattr(module, f"{name.upper()}_BLOCK_SIZE") == 256
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_round_trip_and_payload_fields(name):
|
||||
_, quantize, dequantize, _, block_bytes, _, _ = _parts(name)
|
||||
generator = torch.Generator().manual_seed(1234)
|
||||
weight = torch.randn((2, 512), generator=generator, dtype=torch.bfloat16)
|
||||
|
||||
packed, shape = quantize(weight)
|
||||
reconstructed = dequantize(packed, shape)
|
||||
chunked = dequantize(packed, shape, block_chunk_size=1)
|
||||
|
||||
assert packed.shape == (2, 2, block_bytes)
|
||||
assert packed.dtype == torch.uint8
|
||||
assert reconstructed.shape == weight.shape
|
||||
assert reconstructed.dtype == torch.bfloat16
|
||||
assert torch.equal(reconstructed, chunked)
|
||||
normalized_mse = (
|
||||
reconstructed.float() - weight.float()
|
||||
).square().mean() / weight.float().square().mean()
|
||||
assert normalized_mse < 0.25
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_decode_is_invariant_to_chunk_size(name):
|
||||
"""Chunking the decode is a memory bound, not a numerical choice."""
|
||||
_, quantize, dequantize, _, _, _, _ = _parts(name)
|
||||
torch.manual_seed(0)
|
||||
weight = torch.randn(3, 1024, dtype=torch.bfloat16)
|
||||
packed, shape = quantize(weight)
|
||||
|
||||
reference = dequantize(packed, shape, dtype=weight.dtype, block_chunk_size=1)
|
||||
for chunk in (2, 7, 4096):
|
||||
assert torch.equal(
|
||||
dequantize(packed, shape, dtype=weight.dtype, block_chunk_size=chunk), reference
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_zero_block_has_canonical_zero_encoding(name):
|
||||
_, quantize, dequantize, _, _, _, _ = _parts(name)
|
||||
weight = torch.zeros((2, 256), dtype=torch.bfloat16)
|
||||
packed, shape = quantize(weight)
|
||||
|
||||
assert not packed.any()
|
||||
assert torch.equal(dequantize(packed, shape), weight)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_underflowed_scale_has_canonical_zero_encoding(name):
|
||||
"""A block whose scale rounds to zero in FP16 packs as all zero bytes.
|
||||
|
||||
The magnitude has to clear every format's threshold at once: the IQ1 formats divide by a
|
||||
native max of 16.875 against the IQ2 formats' 166.6, so a weight that underflows an IQ2
|
||||
scale still lands on an FP16 subnormal for IQ1.
|
||||
"""
|
||||
_, quantize, dequantize, _, _, _, _ = _parts(name)
|
||||
weight = torch.full((1, 256), -1e-8, dtype=torch.bfloat16)
|
||||
packed, shape = quantize(weight)
|
||||
|
||||
assert not packed.any()
|
||||
assert torch.equal(dequantize(packed, shape), torch.zeros_like(weight))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_requires_complete_last_dimension_blocks(name):
|
||||
_, quantize, _, _, _, _, _ = _parts(name)
|
||||
with pytest.raises(ValueError, match="last weight dimension"):
|
||||
quantize(torch.ones(2, 257))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_treats_nonfinite_values_as_zero(name):
|
||||
_, quantize, _, _, _, _, _ = _parts(name)
|
||||
weight = torch.zeros(1, 256)
|
||||
weight[0, 0] = float("nan")
|
||||
weight[0, 1] = float("inf")
|
||||
weight[0, 2] = float("-inf")
|
||||
|
||||
packed, _ = quantize(weight)
|
||||
assert not packed.any()
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_saturates_finite_values_above_the_float32_range(name):
|
||||
"""float64 weights are accepted, so a finite value too large for float32 must saturate.
|
||||
|
||||
Converting before sanitizing would turn it into infinity and then zero, which silently
|
||||
encodes a large weight as nothing and diverges from the CUDA ``load_float`` policy.
|
||||
"""
|
||||
_, quantize, _, _, _, _, _ = _parts(name)
|
||||
torch.manual_seed(0)
|
||||
weight = torch.randn(1, 256, dtype=torch.float64)
|
||||
weight[0, 7] = 1e100
|
||||
saturated = weight.clone()
|
||||
saturated[0, 7] = torch.finfo(torch.float32).max
|
||||
zeroed = weight.clone()
|
||||
zeroed[0, 7] = 0.0
|
||||
|
||||
packed, _ = quantize(weight)
|
||||
assert torch.equal(packed, quantize(saturated)[0])
|
||||
assert not torch.equal(packed, quantize(zeroed)[0])
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_rejects_invalid_shape_metadata(name):
|
||||
_, _, dequantize, _, block_bytes, _, _ = _parts(name)
|
||||
packed = torch.zeros((1, 1, block_bytes), dtype=torch.uint8)
|
||||
for weight_shape in (torch.tensor([[1, 256]]), torch.tensor([1.0, 256.0]), torch.tensor([257])):
|
||||
with pytest.raises(ValueError, match=r"weight_shape|logical weight shape"):
|
||||
dequantize(packed, weight_shape)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_rejects_scalar_packed_payload(name):
|
||||
_, _, dequantize, _, _, _, _ = _parts(name)
|
||||
with pytest.raises(ValueError, match="packed_weights"):
|
||||
dequantize(torch.tensor(0, dtype=torch.uint8), torch.tensor([1, 256]))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_fake_quant_has_pass_through_gradient(name):
|
||||
quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(num_bits=name, block_sizes={-1: 256}, backend="ggml")
|
||||
)
|
||||
weight = torch.randn(2, 256, requires_grad=True)
|
||||
quantizer(weight).sum().backward()
|
||||
assert torch.equal(weight.grad, torch.ones_like(weight))
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", NAMES)
|
||||
def test_search_is_independent_of_default_dtype(name):
|
||||
"""The encoder must not inherit a global default dtype; it works in float32 throughout."""
|
||||
_, quantize, _, _, _, _, _ = _parts(name)
|
||||
torch.manual_seed(0)
|
||||
weight = torch.randn(2, 256)
|
||||
expected, _ = quantize(weight)
|
||||
try:
|
||||
torch.set_default_dtype(torch.float64)
|
||||
assert torch.equal(quantize(weight)[0], expected)
|
||||
finally:
|
||||
torch.set_default_dtype(torch.float32)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("name", formats())
|
||||
def test_decoder_matches_llama_cpp_on_captured_blocks(name):
|
||||
"""Decode bytes we did not produce and match llama.cpp's own output exactly.
|
||||
|
||||
A round-trip against our own encoder cannot catch a layout error that the encoder makes
|
||||
symmetrically; these blocks come from a real checkpoint, so they can.
|
||||
"""
|
||||
_, _, dequantize, _, block_bytes, _, _ = _parts(name)
|
||||
blocks = packed_blocks(name)
|
||||
assert blocks.shape[1] == block_bytes
|
||||
|
||||
count = blocks.shape[0]
|
||||
decoded = dequantize(
|
||||
torch.from_numpy(blocks).reshape(count, 1, block_bytes),
|
||||
torch.tensor([count, 256]),
|
||||
dtype=torch.float32,
|
||||
).reshape(count, 256)
|
||||
|
||||
assert np.array_equal(decoded.numpy(), expected_values(name))
|
||||
|
||||
|
||||
def test_every_format_has_conformance_vectors():
|
||||
"""A new format must arrive with blocks captured from a real llama.cpp checkpoint."""
|
||||
assert sorted(formats()) == NAMES
|
||||
|
||||
|
||||
def test_error_decreases_with_bit_width():
|
||||
"""More bits must buy less error, or a format's scale handling is wrong."""
|
||||
generator = torch.Generator().manual_seed(7)
|
||||
weight = torch.randn((4, 1024), generator=generator)
|
||||
errors = []
|
||||
for name in sorted(NAMES, key=lambda n: FORMATS[n][3]):
|
||||
quantizer = TensorQuantizer(
|
||||
QuantizerAttributeConfig(num_bits=name, block_sizes={-1: 256}, backend="ggml")
|
||||
)
|
||||
errors.append(float((quantizer(weight) - weight).square().mean()))
|
||||
|
||||
assert errors == sorted(errors, reverse=True), dict(zip(sorted(NAMES), errors))
|
||||
Reference in New Issue
Block a user