mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do? Type of change: new feature **Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed the PyTorch codec; this PR adds its **CUDA encoder** and makes the format reachable. With it, ModelOpt supports all five GGML IQ formats at one and two bits. - the CUDA encoder, its binding and extension build wiring, plus the CUDA path in `quantize_iq1_m` - an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so backend dispatch, both exporters and `convert_hf_config` take it from there - the `ggml` package export - the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG entry The kernel lands with the registration so every registered format keeps a CUDA encoder. ### The kernel In the kernel the delta shift is free per group, so it sits above the entry index in the sort key: a tie still prefers the lower shift and then the lower entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels read it from global memory and rely on the cache. | 5632×2048 weight | torch | CUDA | | |---|---|---|---| | IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** | ### Shared with IQ1_S rather than copied The two IQ1 kernels load each vector, score it against a grid entry and apply the ±1/8 shift the same way. So those three steps move into `common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA output on a 5632×2048 weight hashes the same before and after, and so does IQ1_M's, compared against the pre-split version of this change. IQ1_S encodes at the same speed (306 M elem/s). ### Usage ```bash python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m ``` ### Testing Registering the format brings it under every registry-driven test with no IQ1_M-specific test code: backend dispatch, weight caching, the `num_bits` guard, `convert_hf_config` metadata, Megatron export and the `TensorQuantizer` tests in the shared battery. The shared CUDA battery gains one row. - `tests/unit/torch/quantization/test_ggml_backend.py`, `test_iq_formats.py`, `tests/unit/torch/export/test_convert_hf_config.py`, `tests/unit/recipe/test_presets.py`: **166 passed** - broader unit sweep (`-k 'ggml or iq or gguf or registry'` over quantization, export and recipe tests): **221 passed**. The one failure, `test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`, is a `torchvision` import error in my environment, unrelated to IQ. - `tests/gpu/torch/quantization/test_iq_formats_cuda.py`, `test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO 6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch encoder parity - `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k 'iq or ggml'`: **45 passed** (9 tests × 5 formats) in `nvcr.io/nvidia/nemo:26.08` - `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed** - reconstruction error falls monotonically across all five formats, pinned by a test - `general/ptq` now holds 31 recipes; `ptq.md` is updated. Rebased onto `main` after #2513 merged. The resulting tree is identical to the one the runs above tested, and the unit set was rerun on it: 166 passed. On this GPU, two of #2515's Q8_0 tests in `tests/gpu/_extensions/test_torch_extensions.py` fail: `test_cuda_ext_q8_0_zero_and_roundf_layout` and `test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically on a clean `main` checkout, so they are not from this PR. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code sources or dependencies. - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ✅ - Did you get Claude approval on this PR?: ❌ Not yet run. ### Additional Information Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M codec), all merged → **this**. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added IQ1_M weight-only quantization at 1.75 bits per weight, with CUDA acceleration and a 256-value block size. * Added an IQ1_M post-training quantization recipe for eligible linear layers; calibration data is not required. * Added IQ1_M to the supported GGML-compatible formats and recipe listings. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5.5
parent
333ace1bc9
commit
aa89722d38
@@ -76,13 +76,14 @@ def test_ptq_whisper(command):
|
||||
PTQCommand(quant="int8_weight_only", kv_cache_quant="none"),
|
||||
PTQCommand(quant="int4_awq", kv_cache_quant="none"),
|
||||
PTQCommand(quant="w4a8_awq_beta", kv_cache_quant="none"),
|
||||
# GGML IQ weight-only, recipe-driven: four formats between 1.56 and 2.56 bits
|
||||
# per weight. These encoders require every weight's input dimension to be a multiple of
|
||||
# GGML IQ weight-only, recipe-driven: the five formats between 1.56 and 2.56 bits per
|
||||
# weight. These encoders require every weight's input dimension to be a multiple of
|
||||
# 256; TinyLlama's 2048 and 5632 both are. None of them calibrates -- every recipe
|
||||
# sets algorithm: null -- so the only IQ-specific cost is packing each weight once on
|
||||
# a CUDA encoder and decoding it on each forward, which fits the 300s
|
||||
# tests/examples default.
|
||||
PTQCommand(recipe="general/ptq/iq1_s", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq1_m", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq2_xxs", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq2_xs", kv_cache_quant="none"),
|
||||
PTQCommand(recipe="general/ptq/iq2_s", kv_cache_quant="none"),
|
||||
|
||||
@@ -23,12 +23,14 @@ derives the block scale in its own kernel instead of taking a precomputed one.
|
||||
import pytest
|
||||
import torch
|
||||
|
||||
import modelopt.torch.quantization.ggml.iq1_m as iq1_m_module
|
||||
import modelopt.torch.quantization.ggml.iq1_s as iq1_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_s as iq2_s_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xs as iq2_xs_module
|
||||
import modelopt.torch.quantization.ggml.iq2_xxs as iq2_xxs_module
|
||||
from modelopt.torch.quantization.extensions import get_cuda_ext_ggml
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_M_BLOCK_BYTES,
|
||||
IQ1_S_BLOCK_BYTES,
|
||||
IQ2_S_BLOCK_BYTES,
|
||||
IQ2_XS_BLOCK_BYTES,
|
||||
@@ -39,6 +41,7 @@ from modelopt.torch.quantization.ggml import (
|
||||
# module, packer name, per-block payload size, whether the packer takes precomputed scales
|
||||
FORMATS = {
|
||||
"iq1_s": (iq1_s_module, "iq1_s_pack", IQ1_S_BLOCK_BYTES, False),
|
||||
"iq1_m": (iq1_m_module, "iq1_m_pack", IQ1_M_BLOCK_BYTES, True),
|
||||
"iq2_xxs": (iq2_xxs_module, "iq2_xxs_pack", IQ2_XXS_BLOCK_BYTES, True),
|
||||
"iq2_xs": (iq2_xs_module, "iq2_xs_pack", IQ2_XS_BLOCK_BYTES, True),
|
||||
"iq2_s": (iq2_s_module, "iq2_s_pack", IQ2_S_BLOCK_BYTES, True),
|
||||
|
||||
@@ -34,6 +34,8 @@ from modelopt.recipe.presets import RecipeSupersededAction
|
||||
from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT
|
||||
from modelopt.torch.quantization.config import LocalHessianCalibConfig, QuantizeConfig
|
||||
from modelopt.torch.quantization.ggml import (
|
||||
IQ1_M_BLOCK_SIZE,
|
||||
IQ1_M_EFFECTIVE_BITS,
|
||||
IQ1_S_BLOCK_SIZE,
|
||||
IQ1_S_EFFECTIVE_BITS,
|
||||
IQ2_S_BLOCK_SIZE,
|
||||
@@ -139,6 +141,7 @@ def test_mlp_weight_only_recipe_matches_its_mtq_cfg(recipe_name, cfg_name):
|
||||
("qformat", "block_size", "effective_bits"),
|
||||
[
|
||||
("iq1_s", IQ1_S_BLOCK_SIZE, IQ1_S_EFFECTIVE_BITS),
|
||||
("iq1_m", IQ1_M_BLOCK_SIZE, IQ1_M_EFFECTIVE_BITS),
|
||||
("iq2_xxs", IQ2_XXS_BLOCK_SIZE, IQ2_XXS_EFFECTIVE_BITS),
|
||||
("iq2_xs", IQ2_XS_BLOCK_SIZE, IQ2_XS_EFFECTIVE_BITS),
|
||||
("iq2_s", IQ2_S_BLOCK_SIZE, IQ2_S_EFFECTIVE_BITS),
|
||||
|
||||
Reference in New Issue
Block a user