Files
Chenjie LuoandClaude Opus 5.5 3091b8ff69 [4/5] Add the IQ2_S CUDA encoder and register the format (#2565)
### What does this PR do?

Type of change: new feature

**Second of two PRs adding IQ2_S** (2.5625 bits per weight). #2512
landed the PyTorch codec; this PR adds its **CUDA encoder** and makes
the format reachable:

- the CUDA encoder, its binding and extension build wiring, plus the
CUDA path in `quantize_iq2_s`
- an `IQFormat` record and **one `IQ_FORMAT_REGISTRY` entry**, so
backend dispatch, both exporters and `convert_hf_config` take it from
there
- the `ggml` package export
- the `general/ptq/iq2_s` recipe, its presets, `ptq.md` and a CHANGELOG
entry

The kernel lands with the registration so every registered format keeps
a CUDA encoder.

On the mixed-precision checkpoint #2511 measured
(`unsloth/Qwen3.8-27B-GGUF`), IQ2_S covers **9 tensors and 0.6 B
parameters**.

### The kernel

IQ2_S's **1024-entry codebook is twice IQ2_XS's**, which makes its
search the most expensive in the family. The codebook and its norms take
36 KiB of shared memory, the most of any IQ kernel but inside the 48 KiB
static limit, so they are declared statically like the IQ2_XS and
IQ2_XXS kernels.

That cost is why the kernel matters more here than anywhere else:

| | torch | CUDA | |
|---|---|---|---|
| IQ2_S, 5632×2048 weight | 0.8 M elem/s | **725.7 M elem/s** | **907×**
|
| extrapolated to a 27B model | ~9.8 hours | **~37 s** | |

### Usage

```bash
python examples/hf_ptq/hf_ptq.py --pyt_ckpt_path <model> --recipe general/ptq/iq2_s
```

### Testing

Registering the format brings it under every registry-driven test with
no IQ2_S-specific test code: backend dispatch and weight caching, the
`num_bits` guard, `convert_hf_config` metadata (uniform and mixed
precision), all 9 Megatron export tests, and the two `TensorQuantizer`
tests in the shared battery. The shared CUDA battery gains one row.

- `tests/unit/torch/quantization/test_ggml_backend.py`,
`test_iq_formats.py`,
`tests/unit/torch/export/test_convert_hf_config.py`,
`tests/unit/recipe/test_presets.py`: **134 passed**
- broader unit sweep (`-k 'ggml or iq or gguf or registry'` over
quantization, export and recipe tests): **192 passed**. The one failure,
`test_export_registry.py::test_builtin_dispatch_covers_all_handler_shapes`,
is a `torchvision` import error in my environment, unrelated to IQ.
- `tests/gpu/torch/quantization/test_iq_formats_cuda.py`,
`test_iq1_s_cuda.py`, `test_iq2_xs_cuda.py`: **42 passed** on RTX PRO
6000 Blackwell (sm_120). 7 of them are IQ2_S: CUDA-vs-PyTorch encoder
parity, determinism, reconstruction at scale, zero and non-finite
policy, float64 input and the fallback path.
- `tests/gpu_megatron/torch/export/test_unified_export_megatron.py -k
'iq or ggml'`: **36 passed** (9 tests × 4 formats) in
`nvcr.io/nvidia/nemo:26.08`
- `tests/examples/hf_ptq/test_llm_ptq.py -k iq2_s`: **passed**.
TinyLlama PTQ through unified HF export writes `quant_algo: IQ2_S`,
`block_payload_bytes: 82`, and `down_proj` packed as `(2048, 22, 82)`
uint8.
- `general/ptq` now holds 30 recipes.
- The shared-memory change in `b7739d5d0` leaves the packed bytes
identical (same hash on a 5632×2048 weight), and packing runs at 849.1 M
elem/s against 825.7 before on RTX PRO 6000. The GPU battery was rerun:
42 passed.

All of the above was rerun after rebasing onto `main` at `c2aaa44f6`.
That base adds a Q8_0 packer to the same GGML extension (#2515), and
changes the hf_ptq example and the export code this format goes through.
The packed IQ2_S bytes still hash the same. On this RTX PRO 6000
(sm_120), two of #2515's own Q8_0 tests in
`tests/gpu/_extensions/test_torch_extensions.py` fail:
`test_cuda_ext_q8_0_zero_and_roundf_layout` and
`test_cuda_ext_q8_0_dequantizes_with_small_error`. They fail identically
on a clean `main` checkout, so they are not from this PR.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ No new code
sources or dependencies.
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
- Did you get Claude approval on this PR?: ❌ Not yet run.

### Additional Information

Merge order: #2511 (IQ2_XXS, merged) → #2525 (format registry, merged) →
#2512 (IQ2_S codec, merged) → **this** → #2513 (IQ1_M).

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added IQ2_S weight-only quantization for eligible linear layers, at
2.5625 bits per weight.
* Added a PTQ recipe that requires no calibration data. Weights must
meet the existing 256-value block-size constraint.
* Added CUDA-accelerated packing for CUDA weights, with a Python
fallback when the CUDA extension is unavailable.
* **Documentation**
  * Updated the PTQ recipe catalog and IQ-format size tradeoffs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 20:35:38 +00:00
..

ModelOpt Recipes

This folder is the library of ModelOpt optimization recipes — declarative YAML files that describe a complete model-optimization workflow (post-training quantization, speculative-decoding training, diffusion distillation).

Purpose: a recipe is the single, version-controlled source of truth for how a model is optimized — algorithm, per-layer numeric formats, and calibration — expressed as data instead of code. That makes an optimization run reproducible, diffable, and shareable without hand-writing Python config, and lets a tuned configuration be looked up by name. The same YAML drives the Python API (load_recipe), the example CLIs (--recipe), and — for the presets under configs/ — the built-in *_CFG constants.

Recipes are composed from small, reusable building blocks via an $import system, then loaded by path relative to this folder, e.g.:

# PTQ recipe -> mtq.quantize()
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_default-kv_fp8_cast")

# distillation recipe -> DMDConfig
from modelopt.torch.fastgen import load_dmd_config
cfg = load_dmd_config("general/distillation/dmd2_qwen_image")

or selected from a script/CLI flag, e.g. hf_ptq.py --recipe model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.

📖 Must-read for PTQ recipe tuning → ptq.md. It is the guide to every PTQ scheme — body scopes (NVFP4/FP8, experts-only / mlp-only / weight-only), KV-cache modes, and calibration variants — with concrete guidance on choosing and tuning a recipe for your model and deployment. Start there before picking a recipe.

This README is the catalog across all recipe families; ptq.md is the how-to for PTQ.

Layout

Directory What lives here
general/ Model-agnostic recipes — a good starting point for any model. PTQ combos, speculative-decoding training, and distillation.
model_type/<model_type>/ Architecture-specific recipes keyed by a HF model_type; one recipe covers every checkpoint of that architecture.
timm/<architecture>/ Architecture-specific recipes for timm models.
models/<org>/<model_id>/ Checkpoint-specific recipes that mirror a particular published checkpoint, keyed by its model-hub path (e.g. nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16).
configs/ Shared building blocks (numerics/, ptq/units/, ptq/presets/) that recipes compose from via $import. Not run directly.

ℹ️ model_type/ was previously named huggingface/. Old huggingface/<model_type>/... recipe paths still resolve for backward compatibility (a source-tree symlink plus a loader alias), but model_type/ is the canonical location — please use it in new recipes, configs, and --recipe flags.

Choosing where to look: check models/<org>/<model_id>/ for your exact checkpoint first, then model_type/<model_type>/ for its architecture or timm/<architecture>/ for a timm model; if none has an entry, fall back to general/. The presence of a model folder signals a recommended, tuned recipe.


General recipes

The model-agnostic recipes live under general/. For PTQ, recipes are mix-and-match combinations of formats, scope, KV-cache mode, and calibration — ptq.md is the guide; read it to understand the schemes and choose one.

Other general recipe families are documented inside their own folders: general/speculative_decoding/ (EAGLE3 / DFlash draft-head training) and general/distillation/ (diffusion distillation, e.g. DMD2).


model_type/ — architecture-specific recipes

Each lives under its HF model_type. The point of a model folder is to capture what differs from the generic preset — usually an algorithm tweak or a disabled-quantizer pattern for non-text branches. The numerics and standard exclusions are still inherited from configs/. Browse model_type/ for the available model_types; each <task>/ folder has a README.md describing the exact delta. See ptq.md for how the model-specific recipes compare to the general ones and why they deviate.

models/ — checkpoint-specific recipes

These mirror a single published checkpoint's quantization config exactly — a per-component mixed-precision scheme tuned to match a specific release. Each is keyed by the checkpoint's model-hub path <org>/<model_id> (as on the Hugging Face Hub, ModelScope, etc.). Browse models/ for the available checkpoints; see models/README.md for the naming convention.


Adding a recipe

  • New combo for any model → add to general/ptq/ by composing existing configs/ units; follow the <formats-scope>-<kv-mode>[-<algorithm>] naming.
  • Tuned for a HF architecture → model_type/<model_type>/<task>/, with a README.md documenting the delta from the generic preset. Verify the exact model_type against the checkpoint's config.json before placing it.
  • Tuned for a timm architecture → timm/<architecture>/<task>/.
  • Mirrors a specific released checkpoint → models/<org>/<model_id>/ (its model-hub path).
  • Share reused bodies via a # modelopt-schema:-tagged snippet and $import it; keep recipe wrappers thin.