### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
28 KiB
PTQ Recipes & Schemes
This doc walks through the PTQ quantization schemes in two parts: the
model-agnostic recipes under general/ptq/ (the recommended
starting point for any model), and then the
model-specific recipes — per-model_type folders
under model_type/ plus the checkpoint-mirror models/<org>/<checkpoint>/
tier — comparing each to its general baseline and explaining why it deviates.
General recipes
The general recipes are model-agnostic. Each file name combines a formats + scope (what gets quantized, and to what format) with a KV-cache mode, optionally with an algorithm (calibration variant):
<formats-scope>-<kv-mode>[-<algorithm>].yaml
nvfp4_experts_only - kv_fp8_cast
Pick one model-body scheme + one KV-cache scheme; the shipped files are the supported combinations.
The shipped recipes
All 25 general/ptq/ recipes (click to expand)
| Recipe | Model body | KV cache | Calibration |
|---|---|---|---|
fp8_default-kv_fp8 |
FP8 W8A8, all linears | FP8 (calibrated) | max |
fp8_default-kv_fp8_cast |
FP8 W8A8, all linears | FP8 (constant amax) | max |
nvfp4_default-kv_fp8 |
NVFP4 W4A4, all linears | FP8 (calibrated) | max |
nvfp4_default-kv_fp8_cast |
NVFP4 W4A4, all linears | FP8 (constant amax) | max |
nvfp4_act_headroom-kv_fp8_cast |
NVFP4 W4A4, all linears | FP8 (constant amax) | nvfp4_act_headroom |
nvfp4_default-kv_nvfp4_cast |
NVFP4 W4A4, all linears | NVFP4 (constant amax) | max |
nvfp4_default-kv_none-gptq |
NVFP4 W4A4 (static W), all linears | none | GPTQ (layerwise) |
nvfp4_mlp_only-kv_fp8 |
NVFP4 W4A4, MLP + MoE experts | FP8 (calibrated) | max |
nvfp4_mlp_only-novit-kv_fp8 |
NVFP4 W4A4, MLP + MoE experts (VL vision tower excluded) | FP8 (calibrated) | max |
nvfp4_mlp_only-kv_fp8_cast |
NVFP4 W4A4, MLP + MoE experts | FP8 (constant amax) | max |
nvfp4_mlp_only_mse-kv_fp8_cast |
NVFP4 W4A4, MLP + MoE experts | FP8 (constant amax) | MSE + FP8 sweep |
nvfp4_experts_only-kv_fp8 |
NVFP4 W4A4, MoE experts only | FP8 (calibrated) | max |
nvfp4_experts_only-kv_fp8_cast |
NVFP4 W4A4, MoE experts only | FP8 (constant amax) | max |
nvfp4_experts_only-kv_fp8_layerwise |
NVFP4 W4A4, MoE experts only | FP8 (calibrated) | max, layerwise |
nvfp4_experts_only-kv_fp8_layerwise_offload |
NVFP4 W4A4, MoE experts only | FP8 (calibrated) | max, layerwise (non-mutating, for disk offload) |
nvfp4_experts_only-kv_fp8_layerwise_export |
NVFP4 W4A4, MoE experts only | FP8 (calibrated) | max, layerwise (exports each layer as it is calibrated) |
nvfp4_experts_only_mse-kv_fp8_cast |
NVFP4 W4A4, MoE experts only | FP8 (constant amax) | MSE + FP8 sweep |
nvfp4_experts_only_input_scale1-kv_fp8_cast |
NVFP4 W4A4, MoE experts only, expert input_scale pinned to 1.0 |
FP8 (constant amax) | max (weights); expert activations uncalibrated |
nvfp4_omlp_only-kv_fp8 |
NVFP4 W4A4, o_proj + MLP/MoE | FP8 (calibrated) | max |
nvfp4_omlp_only-kv_fp8_cast |
NVFP4 W4A4, o_proj + MLP/MoE | FP8 (constant amax) | max |
nvfp4_weight_only-kv_fp16 |
NVFP4 W4A16, weights only | none (BF16/FP16) | max |
nvfp4_weight_only-kv_fp8_cast |
NVFP4 W4A16, weights only | FP8 (constant amax) | max |
int4_blockwise_weight_only |
INT4 W4A16, block 128, weights only | none | max |
nvfp4_mlp_weight_only |
NVFP4 W4A16 (block 32), MLP + MoE weights only | none | max |
mxfp4_mlp_weight_only |
MXFP4 W4A16, MLP + MoE weights only | none | none (no calibration) |
Model-body schemes
The body scheme is the main lever: it trades accuracy against memory/throughput by choosing which parts of the model drop to low precision and whether activations are quantized too (W4A4/W8A8 vs weight-only W4A16).
Full-model schemes (quantize everything)
fp8_default— per-tensor FP8 E4M3 W8A8 on every linear (attention q/k/v/o + MLP/MoE) — one scale per weight/activation tensor. The safest aggressive option: FP8 has a wide dynamic range, so accuracy loss is usually negligible. Needs Hopper+ for FP8 kernels. Good default when the target hardware is FP8-class and you want the broadest speedup.nvfp4_default— NVFP4 (E2M1, block-16, FP8 block scales) W4A4 on every linear. The most aggressive scheme — 4-bit weights and activations everywhere — for maximum memory/throughput on Blackwell+. Highest risk of accuracy loss; if it regresses, fall back to one of the scoped schemes below rather than abandoning NVFP4.
Scoped schemes (quantize part of the model)
nvfp4_experts_only— NVFP4 W4A4 on MoE routed experts only (*.experts.*,*block_sparse_moe*). Dense layers, shared experts, and attention stay BF16. The most recommended NVFP4 recipe for MoE models: it's the narrowest, most accuracy-preserving NVFP4 scope, so it recovers the most accuracy — while still compressing well, because the routed experts are usually the largest share of the model's total weights.nvfp4_mlp_only— NVFP4 W4A4 on all MLP/FFN compute: dense MLP layers, MoE routed experts, andblock_sparse_moeblocks. Attention stays BF16. Recommended for dense models: most FLOPs/params live in the MLP, so this captures most of the win while leaving the sensitive attention path untouched for accuracy.nvfp4_omlp_only— NVFP4 W4A4 on MLP/MoE plus the attention output projection (o_proj), but not q/k/v. A middle ground betweenmlp_onlyanddefault: adds the o_proj GEMM (often safe) without quantizing the more sensitive q/k/v projections.
Scope vs. compression. These schemes keep the accuracy-sensitive attention path (or the whole dense path) at the model's original precision — BF16 for most checkpoints — and quantize only the FFN/expert weights, which dominate params and compute. That's good for accuracy, but the left-out layers stay uncompressed. How much that matters is model-dependent: on MoE models the routed experts are the vast majority of the weights, so leaving attention in BF16 costs almost nothing on disk; on some dense models the attention projections are large enough that they noticeably bound the checkpoint size. If the attention weights are large and you want to compress them, we recommend adding an FP8 rule for the attention projections (keep NVFP4 on the MLP/experts) rather than leaving them BF16 — FP8 keeps that sensitive path at a safer precision than NVFP4 while still halving those weights vs. BF16.
Weight-only schemes (W4A16 — activations stay BF16)
Quantize weights only; activations run in BF16. This shrinks the model (memory-bound decode win) with much lower accuracy risk than W4A4, and needs no calibration forward pass.
These are usually recommended for low-concurrency deployments — edge and on-device/client use cases — where the workload is memory-bandwidth-bound and shrinking the weights is the main win. For high-concurrency data-center serving, prefer a scheme that also quantizes activations (the W4A4/W8A8 body schemes above): at large batch sizes the GEMMs become compute-bound, so low-bit activations and tensor-core math are what deliver the throughput.
nvfp4_weight_only— NVFP4 weights, BF16 activations. Memory savings of 4-bit weights without the activation-quantization risk.int4_blockwise_weight_only— INT4 weights, block size 128, BF16 activations. Classic W4A16 weight compression; works without NVFP4-class hardware.nvfp4_mlp_weight_only— NVFP4 (block size 32) weights on MLP/MoE layers only, BF16 activations.mxfp4_mlp_weight_only— MXFP4 weights on MLP/MoE layers only, BF16 activations. Needs no calibration forward pass; the QAT starting point for the GPT-OSS family (seeexamples/gpt-oss).
KV-cache schemes
The kv_* suffix controls how the attention KV cache is quantized — independent
of the body scheme. Quantizing the KV cache reduces memory at long context.
kv_fp8_cast— FP8 KV with a constant amax (cast mode): skips KV calibration entirely. Cheaper to produce and the safe default for KV. For most models it is as accurate as the calibratedkv_fp8below, so prefer it unless you have a specific reason to calibrate KV scales. Hopper+.kv_nvfp4_cast— NVFP4 KV cache with constant amax. More aggressive KV compression (4-bit); combines with any body scheme. Blackwell+.kv_fp8— FP8 E4M3 KV cache with calibrated per-tensor amax. The KV scales are measured during the calibration pass. Hopper+.
kv_fp8_castvskv_fp8: both produce an FP8 KV cache._castuses a fixed scale and skips the KV calibration step (faster, no extra data dependence); plainkv_fp8calibrates the scale from data. The cast version usually matches calibrated accuracy, so start withkv_fp8_cast.
Calibration variants
How the quantization scales are searched. The default (no suffix) is max.
max(default) — amax/max calibration. Fast, one calibration pass; the baseline choice.mse(e.g.nvfp4_mlp_only_mse,nvfp4_experts_only_mse) — MSE search for static NVFP4 weight scales, with an FP8-scale sweep over the e4m3 scale values. The MSE search applies to the weights; activations are still max (amax) calibrated as in the default recipes. Costs more calibration time but recovers accuracy NVFP4 W4A4 can lose under plain max. Reach for it when amaxrecipe regresses.input_scale1(nvfp4_experts_only_input_scale1-kv_fp8_cast) — pins the expert activation per-tensor amax to a constant2688.0(= E2M1_MAX × E4M3_MAX = 6 × 448) viaconstant_amax, so the exported NVFP4input_scaleis exactly 1.0 and those quantizers skip activation calibration entirely (no forward statistics collected). Weights are still max-calibrated (computed directly from the weight tensors, no data needed), and the per-block E4M3 activation scales remain dynamic. Because nothing in the recipe needs a calibration forward pass — expert activations are pinned, weight amax comes from the weights, and the KV cast uses a constant amax — it may work out of the box for very large LLMs (hundreds of billions of parameters and up), where running calibration PTQ is difficult on a resource-limited setup. Also reach for it when the deployment stack expects a unit expertinput_scale(e.g. NVFP4 expert kernels that assumeinput_scale == 1.0) or to take expert activation calibration out of the picture.gptq(nvfp4_default-kv_none-gptq) — GPTQ layerwise calibration of the weight scales; writes layerwise checkpoints. GPTQ is best established for INT4 weight-only quantization; its effectiveness on NVFP4 weight quantization varies model by model — it tends to help most when the other recipes show a larger accuracy loss. Applying GPTQ to MoE models is still an open research topic and needs extra recipe tuning.layerwise(nvfp4_experts_only-kv_fp8_layerwise) — max calibration done one decoder layer at a time to lower peak memory; same numerics as the non-layerwise variant.layerwise export(nvfp4_experts_only-kv_fp8_layerwise_export) — the same calibration, additionally writing each decoder layer to the export checkpoint as soon as it is calibrated. A run interrupted part-way resumes without redoing finished layers, and no separate export pass is needed. Same numerics again.
These can also be stacked when a single method isn't enough — e.g. mse +
gptq combines an MSE-searched weight scale with GPTQ's layerwise update.
Choosing a general recipe
- Match the format to your hardware/target. FP8 (
fp8_default) on Hopper+; NVFP4 (nvfp4_*) on Blackwell+; weight-only (*_weight_only) when you want compression with minimal risk or lack NVFP4-class kernels. - Start from the most accurate scope, then quantize more toward your
memory/performance target. For low-concurrency deployments (edge,
on-device/client), start from a weight-only recipe (
nvfp4_weight_only/int4_blockwise_weight_only) — shrinking weights is the main win there. For higher-concurrency serving, begin with the narrowest activation-quantized scope —nvfp4_experts_onlyfor MoE,nvfp4_mlp_onlyfor dense — then widen (mlp_only→omlp_only→default) only as far as your memory/throughput target requires, checking accuracy as you go. - Recover accuracy via calibration before backing off the scope. If a
wider-scope recipe regresses, switch its
maxto themsevariant before retreating to a narrower scope. - Pick KV by deployment.
kv_fp8_castis the safe default (usually as accurate as calibratedkv_fp8); usekv_nvfp4_castfor maximum KV compression.
Beyond PTQ: these recipes'
quantizesections are also reused as the quantization config for QAT/QAD training flows. Dedicated scale-learning (LSQ / Dual-LSQ) QAD recipes live separately undergeneral/qad/.
Model-specific recipes
The general recipes above are model-agnostic: they select layers by wildcard
(*mlp*, *self_attn*, *[kv]_bmm_quantizer) and lean on the shared
default_disabled_quantizers exclusions, so the same file works on any
architecture whose module names follow the usual conventions. A recipe only
earns a place under model_type/<model_type>/ or
models/<org>/<checkpoint>/ when a model has to deviate from
that baseline. The deviations come in four kinds:
ℹ️
model_type/was previously namedhuggingface/; oldhuggingface/<model_type>/...--recipepaths still resolve for backward compatibility, but usemodel_type/going forward.
| Kind | What changes vs. the general recipe | Examples |
|---|---|---|
Architecture-aware quant_cfg |
Per-sub-module format choices a single wildcard scheme can't express | minimax_m3_vl, qwen3_vl, qwen3_5, qwen3_5_moe, vit, nemotron_llama |
| Algorithm override | Same numerics & scope, but the calibration algorithm is tweaked because the default breaks or regresses | gemma, gemma4, mpt |
| Extra exclusions | Adds disabled-quantizer patterns so non-language branches stay full precision | nemotron_vl, diffusion_gemma |
| Checkpoint mirror | A mixed-precision map reproducing one published checkpoint exactly | models/nvidia/NVIDIA-Nemotron-3-*, models/mistralai/Mistral-Medium-3.5-128B |
The numerics and standard exclusions are still inherited from configs/
wherever possible — the model folder captures only the delta. Each <task>/
folder may carry a README.md spelling out that delta.
Architecture-aware quant_cfg — minimax_m3_vl, qwen3_vl, qwen3_5, qwen3_5_moe, vit, nemotron_llama
minimax_m3_vl/ptq/mxfp8_nvfp4_experts applies MXFP8 to the language-model
linear layers and MSE-calibrated NVFP4 to routed experts, with expert
input_scale fixed to 1.0. The vision branch, routers, lm_head, and KV cache
remain unquantized.
qwen3_vl/ptq and qwen3_5/ptq provide FP8 recipes for the visual branch of
validated Qwen3-VL and dense Qwen3.5 checkpoints. fp8_vision-kv_none enables only Vision
Encoder nn.Linear weights and inputs, including the primary merger and any deepstack mergers.
fp8_vision_lm-kv_fp8_cast combines that visual configuration with the standard W8A8 FP8 model
and FP8 KV-cache-cast units. Both keep patch embedding and vision-attention BMM operands in high
precision. The shared visual snippet lives under qwen3_vl; thin wrappers remain discoverable
under each exact Hugging Face model_type.
model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast (and its MoE twin,
which shares the same quant_cfg snippet) is a mixed scheme no single general
body covers: NVFP4 W4A16 on MLP / expert projection weights and lm_head,
FP8 on self-attention and the large linear-attention projections
(in_proj_qkv, in_proj_z, out_proj), plus FP8 KV cast. It also disables
architecture-specific submodules that aren't in the reference recipe
(linear_attn.in_proj_a/b, conv1d, and any visual/mtp siblings).
Why special: these are hybrid linear-attention + softmax-attention models. A general scheme would apply one format per wildcard class; this architecture needs FP8 for the big linear-attention projections but NVFP4 for MLP weights, and needs the linear-attention conv/gate submodules left alone. The dense and MoE families share the identical wildcard rules, so one snippet drives both.
On Qwen3.5 / Qwen3.6 this W4A16 recipe usually does not regress accuracy
versus the official checkpoint, and it is designed for best performance in
low-concurrency use cases (weight-only on the MLP keeps the memory-bound decode
path fast without quantizing activations). Both the dense and MoE folders also
ship an MSE twin, w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast — the identical
layout with the NVFP4 weight scales chosen by an MSE FP8-scale sweep instead of
max calibration (the mse variant applied to this
architecture-specific scheme) — for when the max-calibrated recipe regresses.
vit/ptq/fp8 covers ViT image classifiers (the FP8 + Torch-TRT example):
FP8 W8A8 on every linear, like fp8_default, but it additionally enables the
attention BMM quantizers — q/k/v BMM plus the softmax-P quantizer — in FP8
and disables the output quantizers. Why special: the general LLM recipes never
quantize the attention BMM inputs; for ViT the whole attention block runs in FP8
so Torch-TRT can compile it end-to-end.
nemotron_llama/ptq/{nvfp4,fp8}_output_quant_proj covers the Llama-Nemotron
embedding/reranking encoders (e.g. llama-nemotron-embed-1b-v2): same numerics
as the general nvfp4/fp8 presets, plus output quantizers on the projection
Linears (*_proj.output_quantizer) so TensorRT engines carry inter-layer
activations in the low-precision format instead of FP16 — roughly half the
engine activation memory on these models. The sequence-classification score
head stays unquantized, like lm_head. In the NVFP4 recipe, its [1, hidden]
weight cannot be packed by the NVFP4 exporter. Why special: the
general recipes never enable output quantizers, and the pattern must stay scoped
to GEMM outputs — a DynamicQuantize on non-GEMM outputs (embedding lookup,
pooling) fails to compile in TensorRT.
A lighter case: models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only is close to
general/ptq/nvfp4_mlp_only (NVFP4 on MoE/MLP weights+inputs, FP8 KV) but pinned
to one released checkpoint and carrying instance-specific disables
(share_expert, moe.gate, the conv1d branches).
step3p7/ptq/{nvfp4_experts_only-kv_fp8_cast,nvfp4_mlp_only-kv_fp8} are the
Step-3.7 equivalents, and the reason they exist is module naming: Step calls
the MoE block moe and the dense sibling share_expert, so the general
recipes' *.experts.*, *block_sparse_moe* and *mlp* patterns match nothing
on the routed experts — the general recipe would quantize nothing and export a
checkpoint with quant_algo: null. These select *moe* instead and disable the
router (moe.gate) and share_expert on top. Use them, not the general
recipes, for Step-3.7 checkpoints; Step-3.5 has its own recipe above.
Algorithm overrides — gemma, gemma4, mpt
These quantize the same layers as the general recipes; only the
quantize.algorithm block differs, to work around model-specific numerics:
gemma/ptq/w4a8_awq-kv_fp8_cast(INT4 block weights + FP8 inputs + FP8 KV cast), its multimodal siblinggemma4/ptq/w4a8_awq-kv_fp8_cast(the Gemma 4gemma4model type, whose vision branch stays BF16 via the standard exclusions), andmpt/ptq/w4a8_awq-kv_fp8_castuseawq_litewithalpha_step: 1instead of the default AWQ search. The default search overflows the TRT-LLM kernels on these models; the coarser sweep avoids it without measurably hurting accuracy.gemma/ptq/int8_sq-kv_fp8_cast(INT8 per-channel weights + INT8 inputs + FP8 KV cast) sets SmoothQuantalpha: 0.5instead of the default1.0— Gemma 7B regresses at1.0, and0.5recovers it.
Why special: identical scope/numerics to a general scheme, but a general recipe's default algorithm would overflow or regress here.
Extra exclusions — nemotron_vl, diffusion_gemma
Each of these is numerically identical to a general recipe. What makes them
special is a model-local disabled_quantizers.yaml unit that extends the
standard exclusions so a model-specific branch stays in full precision:
nemotron_vl(vision-language, incl. Nemotron-Parse) — generalnvfp4_default-kv_fp8_castnumerics, adding*vision*,*image*,*radio*,*visual*,*encoder*,*model_encoder*so only the language decoder is quantized.diffusion_gemma(block-diffusion encoder-decoder text LLM on a Gemma4 MoE backbone) — generalnvfp4_experts_only-kv_fp8_castnumerics, adding*self_conditioning*: the self-conditioning network is text-only and never exercised by standard PTQ calibration data, so its quantizers collect no amax and export crashes; the exclusion keeps it in BF16.
Why special: a general recipe would happily quantize the vision/audio encoders (or the never-calibrated self-conditioning branch), regressing those modalities or crashing export. The extra patterns keep them in full precision; everything else matches the general recipe.
Checkpoint mirrors — models/<org>/<checkpoint>
The models/ tier reproduces a single published (or planned)
checkpoint's quant config verbatim:
-
models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attentionmirrorsnvidia/Kimi-K3-NVFP4: the source MXFP4 routed experts are cast to NVFP4, with activationinput_scale=1.0, while KDA and MLA projection weights use 128x128 block FP8. Attention activations are dynamic; shared and latent experts, routers, convolutions, norms, the vision tower,lm_head, and KV cache remain BF16. Because the 2.8T source uses packed MXFP4 expert tensors, use the calibration-free streaming converter inexamples/kimi/rather than the in-memoryhf_ptq.pyflow. -
models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_onlymirrorsnvidia/DeepSeek-V4-Pro-0813-NVFP4: the source MXFP4 routed experts are cast to NVFP4 (weights and activations, block 16), while shared experts, attention, router gates,lm_headand the MTP/DSpark speculative-decoding block stay in their source format. Because the 1.65T source ships its experts packed as MXFP4 in DeepSeek's native layout, calibration runs throughexamples/deepseek/deepseek_v4/ptq.pyand the cast throughquantize_to_nvfp4.py --cast_mxfp4_to_nvfp4, rather than the in-memoryhf_ptq.pyflow. -
models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calibmirrorsnvidia/Mistral-Medium-3.5-128B-NVFP4: decoder MLP layers 4–86 use NVFP4 W4A4, edge MLP layers 0–3 and 87 use FP8 W8A8, and all attention projections and the KV cache use FP8. It uses max calibration. -
models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-msemirrorsnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4exactly — a hybrid Mamba-MoE with a hand-mapped, per-component precision scheme:- MoE routed experts → NVFP4 W4A4,
group_size 16, static weight scales - shared experts and Mamba
in/out_proj→ FP8 per-tensor - KV cache → FP8
- attention q/k/v, MTP head,
lm_head, latent-MoE, Mamba conv1d → BF16
nvfp4-mse.yamluses MSE calibration with an FP8-scale sweep (matches the release);nvfp4-max-calib.yamlis the identical layer map under plainmaxcalibration, kept for comparison. - MoE routed experts → NVFP4 W4A4,
-
models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6follows the same Super-style component map (routed experts NVFP4 W4A4 block-16; shared experts + Mambain/out_proj+ KV cache FP8; everything else BF16), but the routed-expert weights use Four-over-Six (4/6) NVFP4: an MSE search picks each weight's amax multiplier from[1.0, 1.5](M=6 vs. M=4). Activations stay dynamic NVFP4 (not MSE-calibrated). -
models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6applies Four-over-Six NVFP4 W4A16 to routed experts, shared experts, and the language model head; Mambain/out_projweights and inputs plus the KV cache use FP8, while attention remains BF16. -
models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16mirrors the GGUF Q4_K_M bit allocation of the Nemotron-H hybrid, mapped onto NVFP4/FP8 per layer: Q4_K/Q5_0 linears → NVFP4 W4A4 (attention q/k/v/o kept uniform so export can fuse them), the Q6_K MLPdown_projlayers → FP8 W8A8, embeddings → NVFP4 W4A16,lm_head→ FP8 W8A16, and the F32 tensors (conv1d, norms) → BF16. -
models/Qwen/Qwen3.8-2.4T-A95B/ptq/nvfp4_experts_mse-fp8_self_attn-fp8_linear_attn-kv_fp8_castmirrorsnvidia/Qwen3.8-2.4T-A95B-NVFP4: aqwen3_5_moe_textMoE with hybrid attention — gated-delta (linear-attention) layers interleaved with full-attention layers. Routed experts → NVFP4 (MSE-searched static weight scales, dynamic input scales); both self-attention and the full gated-delta linear-attention path (conv1d,in_proj_qkv/z/a/b,out_proj) → FP8 W8A8; KV cache → FP8 cast; everything else, including the MTP block, stays BF16. The gated-delta norms stay BF16, but theconv1dis FP8 like the projections:nn.Conv1dis a registered quant module, so the recipe's broad*linear_attn*wildcards reachlinear_attn.conv1dtoo — matching the published checkpoint, whosehf_quant_config.jsonlistslinear_attn.conv1das FP8 on every gated-delta layer (the full-attention layers have noconv1d). The source ships as native block-FP8 (weight_block_size [128, 128]); the loader dequantizes it to BF16 before quantizers are inserted, so the scales are calibrated against BF16 weights, not the shipped FP8. -
models/zai-org/GLM-5.3-Flash/ptq/nvfp4_experts_dense_mlp-kv_fp8_castis the NVFP4 config forzai-org/GLM-5.3-Flash, aglm5_nextVLM MoE with hybrid attention — KDA (linear-attention) layers interleaved with NoPE sparse-MLA layers. Routed experts and the dense MLP → NVFP4 W4A4; KV cache → FP8 cast; everything else stays BF16 (shared experts, router gate, both attention families, the vision tower, embeddings andlm_head).mlp_layer_typesmarks only layers 0-2dense, so the dense-MLP scope adds just 9 modules (mlp.gate_proj/up_proj/down_proj) on top of the routed experts. The vision tower reuses those same leaf names, so a single*visual*disable is appended last to keepmodel.visual.*in BF16. It pinslayerwise.enable=false, which this VLM requires because its decoder layers nest undermodel.language_model.layers. The MTP layer is not built by the HF class atnum_hidden_layers: 45, so it is neither quantized nor exported. (For plain experts-only NVFP4 on this model, use the generalgeneral/ptq/nvfp4_experts_only-kv_fp8_cast— the model-specific delta here is the dense-MLP scope plus the vision-tower exclusion.)
Why special: unlike any general recipe, each is pinned to one checkpoint and
captures a model-specific deviation a portable general recipe can't express. Most
mix FP8 and NVFP4 across different component types — or individual layers —
and hardcode the precise published layout (for Super, matched on both HF and
Megatron-Core module names) rather than a portable wildcard scheme. GLM-5.3-Flash
is the exception: its deviation is a model-specific scope — a wildcard scheme
plus a load-bearing vision-tower exclusion and the VLM-required
layerwise.enable=false — rather than a per-component precision map.
For the full catalog and how to pick a starting recipe for a given model, see
README.md.