Files
Chenjie LuoandClaude Opus 5.5 aa89722d38 [6/6] Add the IQ1_M CUDA encoder and register the format (#2595)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ1_M** (1.75 bits per weight). #2513 landed
the PyTorch codec; this PR adds its **CUDA encoder** and makes the
format reachable. With it, ModelOpt supports all five GGML IQ formats at
one and two bits.

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq1_m`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq1_m` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

### The kernel

In the kernel the delta shift is free per group, so it sits above the
entry index in the sort key: a tie still prefers the lower shift and
then the lower entry, as the reference encoder does. The 2048-entry grid
IQ1_M shares with IQ1_S is 64 KiB, past the 48 KiB static shared-memory
limit, so both kernels read it from global memory and rely on the cache.

| 5632×2048 weight | torch | CUDA | |
|---|---|---|---|
| IQ1_M encode | 5.6 M elem/s | **318 M elem/s** | **57×** |

### Shared with IQ1_S rather than copied

The two IQ1 kernels load each vector, score it against a grid entry and
apply the ±1/8 shift the same way. So those three steps move into
`common.cuh` as `load_vector`, `grid_terms` and `shifted_error`, and
IQ1_S uses them too. **IQ1_S's packed bytes are unchanged**: its CUDA
output on a 5632×2048 weight hashes the same before and after, and so
does IQ1_M's, compared against the pre-split version of this change.
IQ1_S encodes at the same speed (306 M elem/s).

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq1_m
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ1_M-specific test code: backend dispatch, weight caching, the
`num_bits` guard, `convert_hf_config` metadata, Megatron export and the
`TensorQuantizer` tests in the shared battery. The shared CUDA battery
gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **166 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **221 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **49 passed** on RTX PRO
6000 Blackwell (sm_120), 7 of them IQ1_M, including CUDA-vs-PyTorch
encoder parity
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **45 passed** (9 tests × 5 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq1_m`: **passed**
- reconstruction error falls monotonically across all five formats,
pinned by a test
- `general/ptq` now holds 31 recipes; `ptq.md` is updated.

Rebased onto `main` after #2513 merged. The resulting tree is identical
to the one the runs above tested, and the unit set was rerun on it: 166
passed.

On this GPU, two of #2515's Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS) → #2525 (format registry) → #2512 (IQ2_S
codec) → #2565 (IQ2_S CUDA encoder and registration) → #2513 (IQ1_M
codec), all merged → **this**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ1_M weight-only quantization at 1.75 bits per weight, with
CUDA acceleration and a 256-value block size.
* Added an IQ1_M post-training quantization recipe for eligible linear
layers; calibration data is not required.
* Added IQ1_M to the supported GGML-compatible formats and recipe
listings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:23:48 -07:00
..
…

ModelOpt Recipes

This folder is the library of ModelOpt optimization recipes — declarative YAML files that describe a complete model-optimization workflow (post-training quantization, speculative-decoding training, diffusion distillation).

Purpose: a recipe is the single, version-controlled source of truth for how a model is optimized — algorithm, per-layer numeric formats, and calibration — expressed as data instead of code. That makes an optimization run reproducible, diffable, and shareable without hand-writing Python config, and lets a tuned configuration be looked up by name. The same YAML drives the Python API (load_recipe), the example CLIs (--recipe), and — for the presets under configs/ — the built-in *_CFG constants.

Recipes are composed from small, reusable building blocks via an $import system, then loaded by path relative to this folder, e.g.:

# PTQ recipe -> mtq.quantize()
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_default-kv_fp8_cast")

# distillation recipe -> DMDConfig
from modelopt.torch.fastgen import load_dmd_config
cfg = load_dmd_config("general/distillation/dmd2_qwen_image")

or selected from a script/CLI flag, e.g. hf_ptq.py --recipe model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.

📖 Must-read for PTQ recipe tuning → ptq.md. It is the guide to every PTQ scheme — body scopes (NVFP4/FP8, experts-only / mlp-only / weight-only), KV-cache modes, and calibration variants — with concrete guidance on choosing and tuning a recipe for your model and deployment. Start there before picking a recipe.

This README is the catalog across all recipe families; ptq.md is the how-to for PTQ.

Layout

Directory What lives here
general/ Model-agnostic recipes — a good starting point for any model. PTQ combos, speculative-decoding training, and distillation.
model_type/<model_type>/ Architecture-specific recipes keyed by a HF model_type; one recipe covers every checkpoint of that architecture.
timm/<architecture>/ Architecture-specific recipes for timm models.
models/<org>/<model_id>/ Checkpoint-specific recipes that mirror a particular published checkpoint, keyed by its model-hub path (e.g. nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16).
configs/ Shared building blocks (numerics/, ptq/units/, ptq/presets/) that recipes compose from via $import. Not run directly.

ℹ️ model_type/ was previously named huggingface/. Old huggingface/<model_type>/... recipe paths still resolve for backward compatibility (a source-tree symlink plus a loader alias), but model_type/ is the canonical location — please use it in new recipes, configs, and --recipe flags.

Choosing where to look: check models/<org>/<model_id>/ for your exact checkpoint first, then model_type/<model_type>/ for its architecture or timm/<architecture>/ for a timm model; if none has an entry, fall back to general/. The presence of a model folder signals a recommended, tuned recipe.


General recipes

The model-agnostic recipes live under general/. For PTQ, recipes are mix-and-match combinations of formats, scope, KV-cache mode, and calibration — ptq.md is the guide; read it to understand the schemes and choose one.

Other general recipe families are documented inside their own folders: general/speculative_decoding/ (EAGLE3 / DFlash draft-head training) and general/distillation/ (diffusion distillation, e.g. DMD2).


model_type/ — architecture-specific recipes

Each lives under its HF model_type. The point of a model folder is to capture what differs from the generic preset — usually an algorithm tweak or a disabled-quantizer pattern for non-text branches. The numerics and standard exclusions are still inherited from configs/. Browse model_type/ for the available model_types; each <task>/ folder has a README.md describing the exact delta. See ptq.md for how the model-specific recipes compare to the general ones and why they deviate.

models/ — checkpoint-specific recipes

These mirror a single published checkpoint's quantization config exactly — a per-component mixed-precision scheme tuned to match a specific release. Each is keyed by the checkpoint's model-hub path <org>/<model_id> (as on the Hugging Face Hub, ModelScope, etc.). Browse models/ for the available checkpoints; see models/README.md for the naming convention.


Adding a recipe

  • New combo for any model → add to general/ptq/ by composing existing configs/ units; follow the <formats-scope>-<kv-mode>[-<algorithm>] naming.
  • Tuned for a HF architecture → model_type/<model_type>/<task>/, with a README.md documenting the delta from the generic preset. Verify the exact model_type against the checkpoint's config.json before placing it.
  • Tuned for a timm architecture → timm/<architecture>/<task>/.
  • Mirrors a specific released checkpoint → models/<org>/<model_id>/ (its model-hub path).
  • Share reused bodies via a # modelopt-schema:-tagged snippet and $import it; keep recipe wrappers thin.