Files
Chenjie LuoandClaude Opus 5.5 49f8d1eda8 Add the IQ1_M CUDA encoder and register the format
Second of two changes adding IQ1_M. The previous change landed the PyTorch
codec; this one adds its CUDA encoder and makes the format reachable.

In the kernel the delta shift is free per group, so it sits above the entry
index in the sort key: a tie still prefers the lower shift and then the lower
entry, as the reference encoder does. The 2048-entry grid IQ1_M shares with
IQ1_S is 64 KiB, past the 48 KiB static shared-memory limit, so both kernels
read it from global memory and rely on the cache. On a 5632x2048 weight the
encoder runs at 318 M elem/s against the torch search's 5.6, and its packed
bytes match the PyTorch encoder's.

The two IQ1 kernels load each vector, score it against a grid entry and apply
the +/-1/8 shift the same way, so those three steps move into common.cuh as
load_vector, grid_terms and shifted_error, and IQ1_S uses them too. IQ1_S's
packed bytes are unchanged.

IQ1_M gets an IQFormat record and one IQ_FORMAT_REGISTRY entry, so backend
dispatch, both exporters and convert_hf_config take it from there. The ggml
package exports it, and a general/ptq/iq1_m recipe uses it. Registering it
brings it under every registry-driven test with no IQ1_M-specific test code,
and the hf_ptq example test gains an IQ1_M case.

IQ1_M covers 25 tensors and 1.2B parameters of the mixed-precision checkpoint
IQ2_XXS measured. With it, ModelOpt supports all five GGML IQ formats at one
and two bits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-09-30 16:47:00 +00:00
..

ModelOpt Recipes

This folder is the library of ModelOpt optimization recipes — declarative YAML files that describe a complete model-optimization workflow (post-training quantization, speculative-decoding training, diffusion distillation).

Purpose: a recipe is the single, version-controlled source of truth for how a model is optimized — algorithm, per-layer numeric formats, and calibration — expressed as data instead of code. That makes an optimization run reproducible, diffable, and shareable without hand-writing Python config, and lets a tuned configuration be looked up by name. The same YAML drives the Python API (load_recipe), the example CLIs (--recipe), and — for the presets under configs/ — the built-in *_CFG constants.

Recipes are composed from small, reusable building blocks via an $import system, then loaded by path relative to this folder, e.g.:

# PTQ recipe -> mtq.quantize()
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_default-kv_fp8_cast")

# distillation recipe -> DMDConfig
from modelopt.torch.fastgen import load_dmd_config
cfg = load_dmd_config("general/distillation/dmd2_qwen_image")

or selected from a script/CLI flag, e.g. hf_ptq.py --recipe model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.

📖 Must-read for PTQ recipe tuning → ptq.md. It is the guide to every PTQ scheme — body scopes (NVFP4/FP8, experts-only / mlp-only / weight-only), KV-cache modes, and calibration variants — with concrete guidance on choosing and tuning a recipe for your model and deployment. Start there before picking a recipe.

This README is the catalog across all recipe families; ptq.md is the how-to for PTQ.

Layout

Directory What lives here
general/ Model-agnostic recipes — a good starting point for any model. PTQ combos, speculative-decoding training, and distillation.
model_type/<model_type>/ Architecture-specific recipes keyed by a HF model_type; one recipe covers every checkpoint of that architecture.
timm/<architecture>/ Architecture-specific recipes for timm models.
models/<org>/<model_id>/ Checkpoint-specific recipes that mirror a particular published checkpoint, keyed by its model-hub path (e.g. nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16).
configs/ Shared building blocks (numerics/, ptq/units/, ptq/presets/) that recipes compose from via $import. Not run directly.

ℹ️ model_type/ was previously named huggingface/. Old huggingface/<model_type>/... recipe paths still resolve for backward compatibility (a source-tree symlink plus a loader alias), but model_type/ is the canonical location — please use it in new recipes, configs, and --recipe flags.

Choosing where to look: check models/<org>/<model_id>/ for your exact checkpoint first, then model_type/<model_type>/ for its architecture or timm/<architecture>/ for a timm model; if none has an entry, fall back to general/. The presence of a model folder signals a recommended, tuned recipe.


General recipes

The model-agnostic recipes live under general/. For PTQ, recipes are mix-and-match combinations of formats, scope, KV-cache mode, and calibration — ptq.md is the guide; read it to understand the schemes and choose one.

Other general recipe families are documented inside their own folders: general/speculative_decoding/ (EAGLE3 / DFlash draft-head training) and general/distillation/ (diffusion distillation, e.g. DMD2).


model_type/ — architecture-specific recipes

Each lives under its HF model_type. The point of a model folder is to capture what differs from the generic preset — usually an algorithm tweak or a disabled-quantizer pattern for non-text branches. The numerics and standard exclusions are still inherited from configs/. Browse model_type/ for the available model_types; each <task>/ folder has a README.md describing the exact delta. See ptq.md for how the model-specific recipes compare to the general ones and why they deviate.

models/ — checkpoint-specific recipes

These mirror a single published checkpoint's quantization config exactly — a per-component mixed-precision scheme tuned to match a specific release. Each is keyed by the checkpoint's model-hub path <org>/<model_id> (as on the Hugging Face Hub, ModelScope, etc.). Browse models/ for the available checkpoints; see models/README.md for the naming convention.


Adding a recipe

  • New combo for any model → add to general/ptq/ by composing existing configs/ units; follow the <formats-scope>-<kv-mode>[-<algorithm>] naming.
  • Tuned for a HF architecture → model_type/<model_type>/<task>/, with a README.md documenting the delta from the generic preset. Verify the exact model_type against the checkpoint's config.json before placing it.
  • Tuned for a timm architecture → timm/<architecture>/<task>/.
  • Mirrors a specific released checkpoint → models/<org>/<model_id>/ (its model-hub path).
  • Share reused bodies via a # modelopt-schema:-tagged snippet and $import it; keep recipe wrappers thin.