Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)

### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-15 12:16:12 -07:00
committed by GitHub
parent 30f89908f0
commit c7ed23a103
75 changed files with 426 additions and 178 deletions
+4 -3
View File
@@ -18,11 +18,12 @@ Changelog
**Backward Breaking Changes**
- Unified HuggingFace export now fails with ``NotImplementedError`` when it meets an MoE block whose expert projection names it does not know, instead of assuming Mixtral's ``w1``/``w2``/``w3``. If you hit this, register a ``ModelSpec`` for the model under ``modelopt/torch/models/``. Every MoE architecture ModelOpt exported correctly before this change is registered, so no supported model regresses.
- ``--recipe`` (and ``modelopt.recipe.load_recipe``) now resolve a recipe path **filesystem-first**: a recipe of the same relative path in the current working directory takes precedence over the shipped built-in of that name, matching how recipe ``$import`` paths already resolve. Previously the built-in won.
**Deprecations**
- The single-format quantization CLI flags are deprecated in favour of ``--recipe`` and will be removed in a future release; passing one now emits a ``FutureWarning``. ``examples/hf_ptq``: ``--qformat`` and ``--kv_cache_qformat``. ``examples/megatron_bridge/quantize.py``: ``--quant_cfg``, ``--kv_cache_quant`` and ``--weight_only``. ``examples/torch_onnx/torch_quant_to_onnx.py``: ``--qformat``. A recipe carries the quantization config, the calibration algorithm and the KV-cache setting in one file, so they cannot drift apart the way separate flags can -- and ``--recipe`` already took precedence over all six, silently on ``hf_ptq`` and with a warning on ``megatron_bridge`` -- with one gap the recipe closes rather than inherits: a weight AutoQuantize recipe that omits ``kv_cache`` still falls back to ``--kv_cache_qformat``, so set ``kv_cache`` in the recipe when migrating. Use a recipe from ``modelopt_recipes/general/ptq/`` or a model-specific one under ``modelopt_recipes/models/``. The warning fires only when a flag is passed explicitly: ``--qformat`` defaults to ``fp8`` and ``--kv_cache_qformat`` to ``fp8_cast``, so warning on the defaults would fire on every run, including runs that correctly use ``--recipe``. ``examples/speculative_decoding/scripts/quantize_drafter.py`` keeps ``--qformat`` undeprecated: it has no ``--recipe`` alternative yet.
- Rename the architecture-specific recipe tier from ``modelopt_recipes/huggingface/`` to ``modelopt_recipes/model_type/`` to clarify that it holds recipes shared across every checkpoint of a Hugging Face ``model_type``. Saved ``--recipe huggingface/<model_type>/...`` paths still resolve via a backward-compatibility alias but now emit a ``FutureWarning``, so update them to ``model_type/<model_type>/...`` as the ``huggingface/`` prefix is deprecated.
- The single-format quantization CLI flags are deprecated in favour of ``--recipe`` and will be removed in a future release; passing one now emits a ``FutureWarning``. ``examples/hf_ptq``: ``--qformat`` and ``--kv_cache_qformat``. ``examples/megatron_bridge/quantize.py``: ``--quant_cfg``, ``--kv_cache_quant`` and ``--weight_only``. ``examples/torch_onnx/torch_quant_to_onnx.py``: ``--qformat``. A recipe carries the quantization config, the calibration algorithm and the KV-cache setting in one file, so they cannot drift apart the way separate flags can -- and ``--recipe`` already took precedence over all six, silently on ``hf_ptq`` and with a warning on ``megatron_bridge`` -- with one gap the recipe closes rather than inherits: a weight AutoQuantize recipe that omits ``kv_cache`` still falls back to ``--kv_cache_qformat``, so set ``kv_cache`` in the recipe when migrating. Use a recipe from ``modelopt_recipes/general/ptq/``, an architecture-specific one under ``modelopt_recipes/model_type/<model_type>/``, or a checkpoint-specific one under ``modelopt_recipes/models/``. The warning fires only when a flag is passed explicitly: ``--qformat`` defaults to ``fp8`` and ``--kv_cache_qformat`` to ``fp8_cast``, so warning on the defaults would fire on every run, including runs that correctly use ``--recipe``. ``examples/speculative_decoding/scripts/quantize_drafter.py`` keeps ``--qformat`` undeprecated: it has no ``--recipe`` alternative yet.
- The TensorRT-LLM checkpoint export format is deprecated and will be removed in 0.49.0: ``export_tensorrt_llm_checkpoint`` and ``torch_to_tensorrt_llm_checkpoint`` now emit a ``DeprecationWarning`` on use. Use ``export_hf_checkpoint``, which exports a unified Hugging Face checkpoint deployable on TensorRT-LLM, vLLM and SGLang. Its implementation moved to ``modelopt.torch.export.trtllm``, so import those two functions from there and the ``ModelConfig`` dataclasses from ``modelopt.torch.export.trtllm.model_config``; both functions remain importable from ``modelopt.torch.export`` for this release only.
**Bug Fixes**
@@ -59,7 +60,7 @@ Changelog
- Fix ``training.gradient_checkpointing`` to reach the DFlash draft. Previously the flag applied only to the frozen target model, saving no activations; the draft now honours it in its decoder-layer loop.
- Add optional **grouped sublayer convolutions for LiLiCorr**, reusing DFlash2's ``DFlashGroupedConv``; enabled by ``conv_kernel_size`` and ``conv_group_size`` in ``dflash_architecture_config``. Requires the DFlash2 branch. Recipe at ``modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml``.
- Add PTQ support for Step-3.7 (``stepfun-ai/Step-3.7-Flash``), whose routed experts were previously left unquantized. Quantize with the new ``huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast`` or ``huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`` recipes rather than the general ones, which select experts by module names Step does not use.
- Add PTQ support for Step-3.7 (``stepfun-ai/Step-3.7-Flash``), whose routed experts were previously left unquantized. Quantize with the new ``model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast`` or ``model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8`` recipes rather than the general ones, which select experts by module names Step does not use.
*Megatron Framework (M-LM / M-Bridge)*
+9 -2
View File
@@ -1,2 +1,9 @@
exclude modelopt_recipes/huggingface/models
prune modelopt_recipes/huggingface/models
# Backward-compat symlinks in the recipe library (huggingface -> model_type and the
# nested model_type/models -> ../models). Walking them would ship the recipe library
# multiple times; the sdist ships each recipe once from its real path. Old
# huggingface/... --recipe paths keep working via the loader alias in
# modelopt/recipe/loader.py.
exclude modelopt_recipes/huggingface
prune modelopt_recipes/huggingface
exclude modelopt_recipes/model_type/models
prune modelopt_recipes/model_type/models
+12 -5
View File
@@ -520,14 +520,21 @@ Model-specific recipes
----------------------
Model-specific recipes come in two tiers: architecture recipes keyed by a
Hugging Face ``model_type`` under ``huggingface/<model_type>/<task>/``, and
Hugging Face ``model_type`` under ``model_type/<model_type>/<task>/``, and
checkpoint mirrors keyed by a model-hub path under
``models/<org>/<model_id>/<task>/``. See
`modelopt_recipes/huggingface/README.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/huggingface/README.md>`_
`modelopt_recipes/model_type/README.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/model_type/README.md>`_
and
`modelopt_recipes/models/README.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/models/README.md>`_
for the layout conventions and recipe-lookup order.
.. note::
``model_type/`` was previously named ``huggingface/``. Old
``huggingface/<model_type>/...`` recipe paths still resolve for backward
compatibility, but ``model_type/`` is the canonical location — prefer it in
new ``--recipe`` flags and ``load_recipe`` calls.
.. list-table::
:header-rows: 1
:widths: 40 60
@@ -536,7 +543,7 @@ for the layout conventions and recipe-lookup order.
- Description
* - ``models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only``
- NVFP4 MLP-only for Step 3.5 Flash MoE model
* - ``huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts``
* - ``model_type/minimax_m3_vl/ptq/mxfp8_nvfp4_experts``
- MXFP8 language-model base with MSE-calibrated NVFP4 routed experts for MiniMax-M3
@@ -689,8 +696,8 @@ The ``modelopt_recipes/`` package is organized as follows:
| +-- nvfp4_omlp_only-kv_fp8_cast.yaml
| +-- nvfp4_omlp_only-kv_fp8.yaml
| +-- nvfp4_weight_only-kv_fp8_cast.yaml
+-- huggingface/ # Architecture-specific recipes (by model_type)
| +-- <model_type>/ # see modelopt_recipes/huggingface/README.md
+-- model_type/ # Architecture-specific recipes (by model_type)
| +-- <model_type>/ # see modelopt_recipes/model_type/README.md
| +-- <task>/
| +-- <recipe>.yaml
+-- models/ # Checkpoint-specific recipes (by model-hub path)
+5 -5
View File
@@ -201,7 +201,7 @@ python hf_ptq.py \
--export_path <quantized_ckpt_path>
```
Built-in recipes are located in `modelopt_recipes/general/ptq/` for model-agnostic recipes and in `modelopt_recipes/huggingface/<model_type>/ptq/` for recipes tuned to a specific Hugging Face `model_type` (see [`modelopt_recipes/huggingface/README.md`](../../modelopt_recipes/huggingface/README.md)). You can also provide a path to your own custom YAML recipe file or directory. See the [recipe documentation](https://nvidia.github.io/Model-Optimizer) for details on the YAML schema and available recipes.
Built-in recipes are located in `modelopt_recipes/general/ptq/` for model-agnostic recipes and in `modelopt_recipes/model_type/<model_type>/ptq/` for recipes tuned to a specific Hugging Face `model_type` (see [`modelopt_recipes/model_type/README.md`](../../modelopt_recipes/model_type/README.md)). You can also provide a path to your own custom YAML recipe file or directory. See the [recipe documentation](https://nvidia.github.io/Model-Optimizer) for details on the YAML schema and available recipes.
> *When `--recipe` is specified, `--qformat` is ignored. KV cache handling depends on the recipe type: a **PTQ** recipe bakes KV cache into its config and ignores `--kv_cache_qformat`; an **AutoQuantize** recipe falls back to `--kv_cache_qformat` unless it sets an explicit `kv_cache` field.*
@@ -287,7 +287,7 @@ Use the recipe directory matching the checkpoint's `model_type`: `qwen3_vl` or `
# Vision encoder only: FP8 vision Linears and merger, BF16 LLM and KV cache.
python hf_ptq.py \
--pyt_ckpt_path <Qwen3-VL-or-Qwen3.5-checkpoint> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none \
--calib_with_images \
--calib_size 512 \
--skip_generate \
@@ -296,7 +296,7 @@ python hf_ptq.py \
# Joint vision encoder + language model FP8 with FP8 KV-cache cast.
python hf_ptq.py \
--pyt_ckpt_path <Qwen3-VL-or-Qwen3.5-checkpoint> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \
--recipe model_type/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \
--calib_with_images \
--calib_size 512 \
--skip_generate \
@@ -410,7 +410,7 @@ search-disabled layers, and cost-excluded layers — see
[`AutoQuantizeConfig`](../../modelopt/recipe/config.py). Shipped recipes live in
[`modelopt_recipes/general/auto_quantize/`](../../modelopt_recipes/general/auto_quantize); model-specific
recipes (carrying architecture-specific disabled layers — e.g. VL vision towers) live under
`modelopt_recipes/huggingface/<model>/auto_quantize/`.
`modelopt_recipes/model_type/<model>/auto_quantize/`.
[Script](./scripts/huggingface_example.sh)
@@ -469,7 +469,7 @@ not actually searched.
The fixed baseline may also reuse a model-specific PTQ configuration. For example, the Qwen3.6 MoE
AutoQuantize recipe imports the same model-specific `quant_cfg` used by
`huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`, reproduces that recipe's `quantize`
`model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`, reproduces that recipe's `quantize`
section, and lists only shared experts, attention, and `lm_head` under `module_search_spaces`. A
loader test asserts that the inherited fixed baseline remains equal to the original PTQ recipe while
leaving the original recipe unchanged.
+2
View File
@@ -2,5 +2,7 @@
accelerate>=1.0.0,<1.14
flash-attn>=2.6.0
liger-kernel>=0.5.0; platform_system != 'Darwin' and platform_system != 'Windows'
# peft 0.21 stops writing adapter_model.safetensors where export.py expects it.
peft<0.21
py7zr
tensorboard
+1 -1
View File
@@ -27,7 +27,7 @@ so it never loads either complete model.
python examples/minimax_m3/hf_ptq_mixed_mxfp8_nvfp4.py \
--mxfp8_ckpt /models/minimax-m3-mxfp8 \
--bf16_ckpt /models/minimax-m3-bf16 \
--recipe huggingface/minimax_m3_vl/ptq/nvfp4_experts_only \
--recipe model_type/minimax_m3_vl/ptq/nvfp4_experts_only \
--output_ckpt /models/minimax-m3-mxfp8-nvfp4 \
--device cuda
```
@@ -24,7 +24,7 @@ Usage:
python hf_ptq_mixed_mxfp8_nvfp4.py \\
--mxfp8_ckpt /models/minimax-m3-mxfp8 \\
--bf16_ckpt /models/minimax-m3-bf16 \\
--recipe huggingface/minimax_m3_vl/ptq/nvfp4_experts_only \\
--recipe model_type/minimax_m3_vl/ptq/nvfp4_experts_only \\
--output_ckpt /workspace/quant/minimax-m3-mxfp8-nvfp4-mixed \\
--device cuda
"""
+2 -2
View File
@@ -122,7 +122,7 @@ mean pooling and L2 normalization on top of the encoder; reranking
graphs take `input_ids` and `attention_mask` with dynamic batch/sequence axes.
The default recipe
(`modelopt_recipes/huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj.yaml`)
(`modelopt_recipes/model_type/nemotron_llama/ptq/nvfp4_output_quant_proj.yaml`)
quantizes weights and activations to NVFP4 and additionally quantizes the
projection-Linear outputs. Without output-side quantization, quantized GEMMs
emit FP16 activations, so FP8/FP4 engines can use as much or more activation
@@ -143,7 +143,7 @@ engines, 5 dynamic-shape profiles up to 32x512), engine activation memory:
python hf_embedding_quant_to_onnx.py \
--model_path=nvidia/llama-nemotron-embed-1b-v2 \
--trust_remote_code \
--recipe=huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj \
--recipe=model_type/nemotron_llama/ptq/nvfp4_output_quant_proj \
--onnx_save_path=llama_nemotron_embed_nvfp4.onnx
# Reranking variant (auto-detected from the model architecture)
@@ -46,7 +46,7 @@ __all__ = [
"register_bidirectional_sdpa",
]
DEFAULT_RECIPE = "huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj"
DEFAULT_RECIPE = "model_type/nemotron_llama/ptq/nvfp4_output_quant_proj"
# TODO: Add an accuracy evaluation pipeline for the embedding and reranking models.
CALIBRATION_TEXTS = [
+6 -6
View File
@@ -68,7 +68,7 @@ from modelopt.recipe import load_recipe
from modelopt.torch.quantization.utils import export_torch_mode
# 1. Quantize the eager PyTorch model with a Model Optimizer PTQ recipe.
recipe = load_recipe("huggingface/vit/ptq/fp8")
recipe = load_recipe("model_type/vit/ptq/fp8")
mtq.quantize(model, recipe.quantize.model_dump(), forward_loop=calibrate)
# 2. Compile the quantized (Q/DQ) graph with Torch-TensorRT.
@@ -138,7 +138,7 @@ This is the recipe the CLI selects by default when `--model_id` points at a HF V
| `--recipe` value | Calibration | What it quantizes |
| :---: | :---: | :--- |
| `huggingface/vit/ptq/fp8` (default) | `max` | Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the `*weight_quantizer` / `*input_quantizer` globs — encoder Linears, the patch-embed `nn.Conv2d` projection, and the `classifier` head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled. |
| `model_type/vit/ptq/fp8` (default) | `max` | Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the `*weight_quantizer` / `*input_quantizer` globs — encoder Linears, the patch-embed `nn.Conv2d` projection, and the `classifier` head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled. |
</div>
@@ -153,7 +153,7 @@ This is the recipe the CLI selects by default when `--model_id` points at a HF V
| Flag | Default | Description |
| :---: | :---: | :--- |
| `--model_id` | `google/vit-large-patch16-224` | HuggingFace model id of the ViT classifier to quantize. |
| `--recipe` | `huggingface/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--recipe` | `model_type/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--calib_samples` | `1024` | Number of tiny-imagenet samples to use for calibration. |
| `--batch_size` | `128` | Batch size for calibration / TRT compile. |
| `--save_dir` | `./modelopt_quantized` | Directory the quantized Model Optimizer state-dict (FP16 weights + Q/DQ metadata) is always saved to, as `vit_modelopt_state.pt` — re-usable across runs without recalibration. |
@@ -182,7 +182,7 @@ python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt
| Flag | Default | Description |
| :---: | :---: | :--- |
| `--model_id` | `google/vit-large-patch16-224` | HuggingFace model id of the ViT classifier to quantize and score. |
| `--recipe` | `huggingface/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--recipe` | `model_type/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--calib_samples` | `1024` | Number of tiny-imagenet samples to use for calibration. |
| `--batch_size` | `128` | Calibration / compile / eval batch size. The Torch-TRT engine is dynamic (`min=1`, `opt=max(--batch_size, 2)`, `max=1024`) and handles any batch including the trailing partial batch. |
| `--eval_data_size` | full 50k | Number of ImageNet validation images to score. |
@@ -199,7 +199,7 @@ python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt
```bash
python torch_tensorrt_accuracy.py \
--recipe huggingface/vit/ptq/fp8 \
--recipe model_type/vit/ptq/fp8 \
--batch_size 128 \
--baseline \
--eval_data_size 5000 \
@@ -215,7 +215,7 @@ python torch_tensorrt_accuracy.py \
## Custom Recipes
Use `--recipe <path>` to plug in a different recipe — either a path relative to `modelopt_recipes/` (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via `modelopt.recipe.load_recipe`, must declare `metadata.recipe_type: ptq` and a `quantize:` section, and its `quantize` config is passed straight to `mtq.quantize`. See the existing [`modelopt_recipes/huggingface/vit/ptq/*.yaml`](../../modelopt_recipes/huggingface/vit/ptq/) for the patterns used here.
Use `--recipe <path>` to plug in a different recipe — either a path relative to `modelopt_recipes/` (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via `modelopt.recipe.load_recipe`, must declare `metadata.recipe_type: ptq` and a `quantize:` section, and its `quantize` config is passed straight to `mtq.quantize`. See the existing [`modelopt_recipes/model_type/vit/ptq/*.yaml`](../../modelopt_recipes/model_type/vit/ptq/) for the patterns used here.
### Resuming From a Saved Checkpoint
+3 -3
View File
@@ -21,7 +21,7 @@ Pipeline:
2. Build a calibration loader from `zh-plus/tiny-imagenet` so the recipe runs
end-to-end without ImageNet access.
3. Run ``mtq.quantize`` with the ViT-specific FP8 recipe under
`modelopt_recipes/huggingface/vit/ptq/`.
`modelopt_recipes/model_type/vit/ptq/`.
4. Compile the quantized model with ``torch_tensorrt.compile(ir="dynamo",
min_block_size=1)`` and verify the compiled-model argmax matches the
fake-quant argmax on a sample input.
@@ -44,10 +44,10 @@ import modelopt.torch.quantization as mtq
from modelopt.recipe import ModelOptPTQRecipe, load_recipe
from modelopt.torch.quantization.utils import export_torch_mode
# Default ViT PTQ recipe under `modelopt_recipes/huggingface/vit/ptq/`. The
# Default ViT PTQ recipe under `modelopt_recipes/model_type/vit/ptq/`. The
# recipe loader resolves this relative path against the built-in recipe library;
# pass `--recipe` for a different one.
DEFAULT_RECIPE = "huggingface/vit/ptq/fp8"
DEFAULT_RECIPE = "model_type/vit/ptq/fp8"
def load_model_and_processor(model_id: str, device: torch.device, dtype: torch.dtype):
+40 -18
View File
@@ -15,6 +15,8 @@
"""Recipe loading utilities."""
import warnings
try:
from importlib.resources.abc import Traversable
except ImportError: # Python < 3.11
@@ -24,7 +26,7 @@ from pathlib import Path
from omegaconf import OmegaConf
from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT as BUILTIN_RECIPES_LIB
from modelopt.torch.opt.config_loader import load_config
from modelopt.torch.opt.config_loader import _alias_builtin_recipe_prefix, load_config
from modelopt.torch.quantization.config import QuantizeConfig
from .config import (
@@ -50,7 +52,7 @@ _REQUIRED_SECTION_PER_RECIPE_TYPE: dict[RecipeType, str] = {
def _resolve_recipe_path(recipe_path: str | Path | Traversable) -> Path | Traversable:
"""Resolve a recipe path, checking the built-in library first then the filesystem.
"""Resolve a recipe path, checking the filesystem first then the built-in library.
Returns the resolved path (file or directory).
"""
@@ -58,24 +60,44 @@ def _resolve_recipe_path(recipe_path: str | Path | Traversable) -> Path | Traver
isinstance(recipe_path, Path) and recipe_path.is_absolute()
):
rp_str = str(recipe_path)
# Backward-compat alias: checkpoint-mirror recipes moved from the old
# ``huggingface/models/<org>/<model_id>/`` layout to the top-level ``models/``
# tier. A source checkout also keeps a ``huggingface/models`` -> ``../models``
# symlink, but symlinks don't survive into built wheels, so rewrite the old
# prefix here too — that keeps saved ``--recipe huggingface/models/...`` paths
# working for pip-installed users, not just source checkouts.
_bc_prefix = "huggingface/models/"
if rp_str.replace("\\", "/").startswith(_bc_prefix):
rp_str = "models/" + rp_str.replace("\\", "/")[len(_bc_prefix) :]
suffixes = [""] if rp_str.endswith((".yml", ".yaml")) else ["", ".yml", ".yaml"]
for suffix in suffixes:
candidate = BUILTIN_RECIPES_LIB.joinpath(rp_str + suffix)
# Backward-compat aliases for the recipe-library restructure. A source checkout keeps
# ``huggingface`` -> ``model_type`` (and the nested ``model_type/models`` -> ``../models``)
# symlinks, but symlinks don't survive into built wheels, so the deprecated tier
# prefixes are rewritten (see ``_alias_builtin_recipe_prefix``) for the built-in lookup,
# keeping saved ``--recipe huggingface/...`` paths working for pip-installed users.
aliased = _alias_builtin_recipe_prefix(rp_str)
def _suffixes(s: str) -> list[str]:
return [""] if s.endswith((".yml", ".yaml")) else ["", ".yml", ".yaml"]
# Filesystem first, probing the path exactly as given (then its aliased form), so a
# user's own local recipe tree overrides a built-in of the same name -- the same
# precedence as ``config_loader._resolve_config_path``. A local ``huggingface/<type>/...``
# file therefore still wins even when ``<type>`` collides with a shipped ``model_type``.
for probe in dict.fromkeys((rp_str, aliased)):
for suffix in _suffixes(probe):
fs_candidate = Path(probe + suffix)
if fs_candidate.is_file() or fs_candidate.is_dir():
return fs_candidate
# Then the built-in library, using the aliased (renamed-tier) form so old
# ``huggingface/...`` paths resolve from wheels where the compat symlink is gone.
# Resolving via a rewritten prefix means the user passed a deprecated tier name, so
# nudge them off it (only here, not when a local file above already won by its own name).
for suffix in _suffixes(aliased):
candidate = BUILTIN_RECIPES_LIB.joinpath(aliased + suffix)
if candidate.is_file() or candidate.is_dir():
if aliased != rp_str:
warnings.warn(
f"Recipe path {rp_str!r} uses a deprecated recipe-tier prefix; it "
f"resolved to the built-in {aliased!r}. Update saved ``--recipe`` paths "
"to the new prefix -- the deprecated one will be removed in a future "
"release.",
FutureWarning,
# Point at the caller of the public ``load_recipe`` (one internal frame
# above this one), the common entry point through which a path arrives.
stacklevel=3,
)
return candidate
for suffix in suffixes:
fs_candidate = Path(rp_str + suffix)
if fs_candidate.is_file() or fs_candidate.is_dir():
return fs_candidate
return Path(rp_str)
return recipe_path
@@ -329,11 +329,14 @@ def _scope_prefixes(rev) -> tuple[str, ...]:
"""Candidate key prefixes a scoped sub-model transform may apply under.
transformers tags a conversion collected from a sub-model with ``scope_prefix`` (the
sub-module path) and ``base_model_prefix``, then matches keys against
``base_model_prefix.scope_prefix.`` first and ``scope_prefix.`` second (see
``WeightTransform._scoped_match``). Returned in that same priority order, each with a
trailing dot. Empty tuple when the transform is unscoped (owned by the root model),
in which case its patterns already address the full key space.
sub-module path). Older versions also tagged a ``base_model_prefix`` and matched keys
against ``base_model_prefix.scope_prefix.`` first and ``scope_prefix.`` second;
transformers>=5.9 dropped ``base_model_prefix`` and ``WeightTransform._scoped_match``
now keys off ``scope_prefix`` alone. The ``getattr`` fallback below covers both: an
absent ``base_model_prefix`` collapses to just the ``scope_prefix.`` candidate.
Returned in priority order, each with a trailing dot. Empty tuple when the transform is
unscoped (owned by the root model), in which case its patterns already address the full
key space.
"""
scope = getattr(rev, "scope_prefix", None)
if scope is None:
+39 -6
View File
@@ -80,6 +80,34 @@ class _ResolvedImport:
# Root to all built-in configs and recipes.
BUILTIN_CONFIG_ROOT = files("modelopt_recipes")
# Deprecated ``modelopt_recipes/`` tier prefixes mapped to their current location. The recipe
# library was restructured and keeps source-tree symlinks (``huggingface`` -> ``model_type``
# and ``model_type/models`` -> ``../models``) for backward compatibility, but symlinks do not
# survive into built wheels. Every relative path handed to the built-in library -- a top-level
# ``--recipe`` / ``load_recipe`` input *or* an ``$import`` inside a recipe -- is rewritten
# through :func:`_alias_builtin_recipe_prefix` so old paths keep resolving for pip-installed
# users too. Ordered longest-prefix first so ``.../models/`` wins over the bare rename.
_DEPRECATED_RECIPE_PREFIXES: tuple[tuple[str, str], ...] = (
("huggingface/models/", "models/"),
("model_type/models/", "models/"),
("huggingface/", "model_type/"),
)
def _alias_builtin_recipe_prefix(config_path: str) -> str:
"""Rewrite a deprecated ``modelopt_recipes/`` tier prefix to its current location.
Returns *config_path* unchanged when it does not start with a deprecated prefix. Only
the built-in-library candidates should use the rewritten form; filesystem probes keep
the original path so a user's local ``huggingface/`` recipe tree still loads by that name.
"""
norm = config_path.replace("\\", "/")
for old, new in _DEPRECATED_RECIPE_PREFIXES:
if norm.startswith(old):
return new + norm[len(old) :]
return config_path
_EXMY_RE = re.compile(r"^[Ee](\d+)[Mm](\d+)$")
_EXMY_KEYS = frozenset({"num_bits", "scale_bits"})
_MODELOPT_SCHEMA_RE = re.compile(r"^\s*#\s*modelopt-schema:\s*(\S+)\s*$")
@@ -120,27 +148,32 @@ def _resolve_config_path(config_file: str | Path | Traversable) -> Path | Traver
"""
# Probe order: filesystem first, then built-in library.
# This lets users override built-in configs by placing a file locally.
# Built-in candidates use the deprecated-tier alias (huggingface/ -> model_type/,
# .../models/ -> models/) so old ``$import`` paths resolve from wheels; filesystem
# candidates keep the original path so a local override tree loads by its own name.
paths_to_check: list[Path | Traversable] = []
if isinstance(config_file, str):
builtin = _alias_builtin_recipe_prefix(config_file)
if not config_file.endswith(".yml") and not config_file.endswith(".yaml"):
paths_to_check.append(Path(f"{config_file}.yml"))
paths_to_check.append(Path(f"{config_file}.yaml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{config_file}.yml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{config_file}.yaml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{builtin}.yml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{builtin}.yaml"))
else:
paths_to_check.append(Path(config_file))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(config_file))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(builtin))
elif isinstance(config_file, Path):
builtin = _alias_builtin_recipe_prefix(str(config_file))
if config_file.suffix in (".yml", ".yaml"):
paths_to_check.append(config_file)
if not config_file.is_absolute():
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(str(config_file)))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(builtin))
else:
paths_to_check.append(Path(f"{config_file}.yml"))
paths_to_check.append(Path(f"{config_file}.yaml"))
if not config_file.is_absolute():
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{config_file}.yml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{config_file}.yaml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{builtin}.yml"))
paths_to_check.append(BUILTIN_CONFIG_ROOT.joinpath(f"{builtin}.yaml"))
elif isinstance(config_file, Traversable):
paths_to_check.append(config_file)
else:
+1 -1
View File
@@ -214,7 +214,7 @@ def _check_weight_quantization_took_effect(model: nn.Module, config: QuantizeCon
"The quantization config asks for weight quantization but no weight quantizer is "
f"enabled, so nothing would be quantized. These patterns asked for it:\n {patterns}\n"
"Either the patterns do not match this architecture's module names (check the "
"model-specific recipes under modelopt_recipes/huggingface/<model_type>/), or the "
"model-specific recipes under modelopt_recipes/model_type/<model_type>/), or the "
"modules holding the weights were never converted to quantized modules (an "
"unsupported custom module, e.g. a trust_remote_code MoE layout).\n"
"Under pipeline parallelism, a rank whose local stage genuinely has none of the "
+12 -6
View File
@@ -26,7 +26,7 @@ cfg = load_dmd_config("general/distillation/dmd2_qwen_image")
```
or selected from a script/CLI flag, e.g. `hf_ptq.py --recipe
huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`.
model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`.
> 📖 **Must-read for PTQ recipe tuning → [`ptq.md`](ptq.md).** It is the
> guide to every PTQ scheme — body scopes (NVFP4/FP8, experts-only / mlp-only /
@@ -42,13 +42,19 @@ huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`.
| Directory | What lives here |
|-----------|-----------------|
| `general/` | **Model-agnostic** recipes — a good starting point for any model. PTQ combos, speculative-decoding training, and distillation. |
| `huggingface/<model_type>/` | **Architecture-specific** recipes keyed by a HF `model_type`; one recipe covers every checkpoint of that architecture. |
| `model_type/<model_type>/` | **Architecture-specific** recipes keyed by a HF `model_type`; one recipe covers every checkpoint of that architecture. |
| `timm/<architecture>/` | **Architecture-specific** recipes for timm models. |
| `models/<org>/<model_id>/` | **Checkpoint-specific** recipes that mirror a particular published checkpoint, keyed by its model-hub path (e.g. `nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16`). |
| `configs/` | Shared building blocks (`numerics/`, `ptq/units/`, `ptq/presets/`) that recipes compose from via `$import`. Not run directly. |
> ℹ️ **`model_type/` was previously named `huggingface/`.** Old
> `huggingface/<model_type>/...` recipe paths still resolve for backward
> compatibility (a source-tree symlink plus a loader alias), but `model_type/` is
> the canonical location — please use it in new recipes, configs, and `--recipe`
> flags.
**Choosing where to look:** check `models/<org>/<model_id>/` for your exact
checkpoint first, then `huggingface/<model_type>/` for its architecture or
checkpoint first, then `model_type/<model_type>/` for its architecture or
`timm/<architecture>/` for a timm model; if none has an entry, fall back to
`general/`. The presence of a model folder signals a recommended, tuned recipe.
@@ -67,13 +73,13 @@ Other general recipe families are documented inside their own folders:
---
## `huggingface/` — architecture-specific recipes
## `model_type/` — architecture-specific recipes
Each lives under its HF `model_type`. The point of a model folder is to capture
**what differs from the generic preset** — usually an algorithm tweak or a
disabled-quantizer pattern for non-text branches. The numerics and standard
exclusions are still inherited from `configs/`. Browse
[`huggingface/`](huggingface/) for the available `model_type`s; each `<task>/`
[`model_type/`](model_type/) for the available `model_type`s; each `<task>/`
folder has a `README.md` describing the exact delta. See [`ptq.md`](ptq.md) for
how the model-specific recipes compare to the general ones and why they deviate.
@@ -92,7 +98,7 @@ convention.
- **New combo for any model** → add to `general/ptq/` by composing existing
`configs/` units; follow the `<formats-scope>-<kv-mode>[-<algorithm>]` naming.
- **Tuned for a HF architecture** → `huggingface/<model_type>/<task>/`, with a
- **Tuned for a HF architecture** → `model_type/<model_type>/<task>/`, with a
`README.md` documenting the delta from the generic preset. Verify the exact
`model_type` against the checkpoint's `config.json` before placing it.
- **Tuned for a timm architecture** → `timm/<architecture>/<task>/`.
@@ -38,7 +38,7 @@ auto_quantize:
score_size: 128
# Base (model-agnostic) non-quantizable layers, spliced from the shared unit. Arch-specific
# models use a recipe under huggingface/<model>/auto_quantize/ that appends to this set.
# models use a recipe under model_type/<model>/auto_quantize/ that appends to this set.
disabled_layers:
- $import: base_disabled_layers
@@ -40,7 +40,7 @@ auto_quantize:
score_size: 128
# Base (model-agnostic) non-quantizable layers, spliced from the shared unit. Arch-specific
# models use a recipe under huggingface/<model>/auto_quantize/ that appends to this set.
# models use a recipe under model_type/<model>/auto_quantize/ that appends to this set.
disabled_layers:
- $import: base_disabled_layers
@@ -38,7 +38,7 @@ auto_quantize:
score_size: 128
# Base (model-agnostic) non-quantizable layers, spliced from the shared unit. Arch-specific
# models use a recipe under huggingface/<model>/auto_quantize/ that appends to this set.
# models use a recipe under model_type/<model>/auto_quantize/ that appends to this set.
disabled_layers:
- $import: base_disabled_layers
@@ -47,7 +47,7 @@ auto_quantize:
# kv_cache omitted -> falls back to --kv_cache_qformat (none in the reference command).
# Base (model-agnostic) non-quantizable layers. Arch-specific models use a recipe under
# huggingface/<model>/auto_quantize/ that extends this set.
# model_type/<model>/auto_quantize/ that extends this set.
disabled_layers:
- $import: base_disabled_layers
@@ -38,7 +38,7 @@ auto_quantize:
score_size: 128
# Base (model-agnostic) non-quantizable layers, spliced from the shared unit. Arch-specific
# models use a recipe under huggingface/<model>/auto_quantize/ that appends to this set.
# models use a recipe under model_type/<model>/auto_quantize/ that appends to this set.
disabled_layers:
- $import: base_disabled_layers
+1
View File
@@ -0,0 +1 @@
model_type
@@ -13,7 +13,7 @@ Built-in recipes live in three tiers — pick the most specific that applies:
1. **[`../models/<org>/<model_id>/`](../models/)** first, if there is an entry
for your **exact** published checkpoint (keyed by its model-hub path). It
mirrors a validated, per-checkpoint scheme.
2. **`huggingface/<model_type>/`** for the target model's Hugging Face
2. **`model_type/<model_type>/`** for the target model's Hugging Face
`model_type` — an architecture-level recipe that applies to every checkpoint
of that `model_type`. The presence of a folder here signals a recommended
recipe for that architecture.
@@ -30,7 +30,7 @@ value of the top-level `model_type` field in the model's `config.json`
language model). Use the exact `model_type` as the directory name:
```text
modelopt_recipes/huggingface/
modelopt_recipes/model_type/
<model_type>/
<task>/
<recipe>.yaml
@@ -43,7 +43,7 @@ modelopt_recipes/huggingface/
Selecting a recipe at runtime uses the path relative to
`modelopt_recipes/`, e.g.
`--recipe huggingface/<model_type>/<task>/<recipe>`.
`--recipe model_type/<model_type>/<task>/<recipe>`.
### Verifying a model's `model_type`
@@ -25,7 +25,7 @@ imports:
base_disable_all: configs/ptq/units/base_disable_all
experts_nvfp4: configs/ptq/units/experts_nvfp4
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
disabled_quantizers: huggingface/diffusion_gemma/ptq/disabled_quantizers
disabled_quantizers: model_type/diffusion_gemma/ptq/disabled_quantizers
metadata:
recipe_type: ptq
@@ -2,7 +2,7 @@
| Recipe | What's model-specific |
|--------|-----------------------|
| `w4a8_awq-kv_fp8_cast.yaml` | Uses `awq_lite` with `alpha_step: 1` instead of the default AWQ search. The default search overflows in TRT-LLM kernels on MPT; the coarser sweep avoids it. Numerics: INT4 block weights + FP8 inputs + FP8 KV-cache cast (constant amax, no KV calibration). Same algorithm override applied to Gemma — see `huggingface/gemma/ptq/`. |
| `w4a8_awq-kv_fp8_cast.yaml` | Uses `awq_lite` with `alpha_step: 1` instead of the default AWQ search. The default search overflows in TRT-LLM kernels on MPT; the coarser sweep avoids it. Numerics: INT4 block weights + FP8 inputs + FP8 KV-cache cast (constant amax, no KV calibration). Same algorithm override applied to Gemma — see `model_type/gemma/ptq/`. |
The base numerics units and the standard disabled-quantizer list are inherited
from the shared `configs/`; only the AWQ algorithm fields are model-specific.
@@ -20,7 +20,7 @@
imports:
base_disable_all: configs/ptq/units/base_disable_all
w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
disabled_quantizers: huggingface/nemotron_vl/ptq/disabled_quantizers
disabled_quantizers: model_type/nemotron_vl/ptq/disabled_quantizers
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
metadata:
@@ -17,7 +17,7 @@
imports:
base_disable_all: configs/ptq/units/base_disable_all
vision_fp8: huggingface/qwen3_vl/ptq/vision_fp8.quant_cfg
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
metadata:
recipe_type: ptq
@@ -19,7 +19,7 @@ imports:
base_disable_all: configs/ptq/units/base_disable_all
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
vision_fp8: huggingface/qwen3_vl/ptq/vision_fp8.quant_cfg
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
metadata:
@@ -15,8 +15,8 @@
# Shared `quant_cfg` snippet for the Qwen3.5 family's
# `w4a16_nvfp4-fp8_attn-kv_fp8_cast` recipe. Imported by both
# `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` (dense `qwen3_5`)
# and `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` (MoE
# `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` (dense `qwen3_5`)
# and `model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml` (MoE
# `qwen3_5_moe`); the two families share the hybrid linear-attention +
# softmax-attention architecture, so the wildcard rules apply identically.
# MoE-only patterns inside `default_disabled_quantizers`
@@ -17,12 +17,12 @@
# HuggingFace `qwen3_5` (dense) models. Covers Qwen3.5 and Qwen3.6 dense
# releases, which share the `qwen3_5` model_type and hybrid linear-attention +
# softmax-attention architecture. Shares its `quant_cfg` with the MoE
# counterpart at `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`;
# counterpart at `model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`;
# the snippet lives under
# `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
# `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
imports:
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
metadata:
recipe_type: ptq
@@ -15,8 +15,8 @@
# Shared `quant_cfg` snippet for the Qwen3.5 family's
# `w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast` recipe. Imported by both
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (dense `qwen3_5`)
# and `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (MoE
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (dense `qwen3_5`)
# and `model_type/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (MoE
# `qwen3_5_moe`); the two families share the hybrid linear-attention +
# softmax-attention architecture, so the wildcard rules apply identically.
# MoE-only patterns inside `default_disabled_quantizers`
@@ -20,12 +20,12 @@
# `w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`: NVFP4 weight scales come from an MSE
# FP8-scale sweep instead of max calibration. Shares its `quant_cfg` with the
# MoE counterpart at
# `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`;
# `model_type/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`;
# the snippet lives under
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
imports:
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
metadata:
recipe_type: ptq
@@ -18,11 +18,11 @@
# releases, which share the `qwen3_5_moe` model_type and hybrid
# linear-attention + softmax-attention MoE architecture. Shares its
# `quant_cfg` with the dense counterpart at
# `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`; the snippet lives
# under `huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
# `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`; the snippet lives
# under `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
imports:
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
metadata:
recipe_type: ptq
@@ -20,12 +20,12 @@
# `w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`: NVFP4 weight scales come from an MSE
# FP8-scale sweep instead of max calibration. Shares its `quant_cfg` with the
# dense counterpart at
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`; the
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`; the
# snippet lives under
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
imports:
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
metadata:
recipe_type: ptq
@@ -10,7 +10,7 @@ imports:
base_disabled_layers: configs/auto_quantize/units/base_disabled_layers
base_cost_excluded_layers: configs/auto_quantize/units/base_cost_excluded_layers
fp8: configs/ptq/presets/model/fp8
model_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
model_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
w4a16_nvfp4: configs/ptq/presets/model/w4a16_nvfp4
metadata:
@@ -17,7 +17,7 @@
imports:
base_disable_all: configs/ptq/units/base_disable_all
vision_fp8: huggingface/qwen3_vl/ptq/vision_fp8.quant_cfg
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
metadata:
recipe_type: ptq
@@ -19,7 +19,7 @@ imports:
base_disable_all: configs/ptq/units/base_disable_all
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
vision_fp8: huggingface/qwen3_vl/ptq/vision_fp8.quant_cfg
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
metadata:
+3 -3
View File
@@ -4,7 +4,7 @@ This folder holds model-optimization recipes (e.g. PTQ recipes) tuned for a
**specific published model instance** — one checkpoint released on a model hub
such as the [Hugging Face Hub](https://huggingface.co/),
[ModelScope](https://modelscope.cn/), or similar. Unlike
[`../huggingface/`](../huggingface/), which keys recipes by a transformers
[`../model_type/`](../model_type/), which keys recipes by a transformers
`model_type` (an architecture shared by many checkpoints), a recipe here mirrors
**one checkpoint's** quantization scheme verbatim.
@@ -48,7 +48,7 @@ Prefer the most specific entry that applies to your model:
1. **`models/<org>/<model_id>/`** — if there is an entry for your **exact**
checkpoint. It reproduces a validated, often per-component mixed-precision
scheme for that release; use it to match a published quantized checkpoint.
2. **[`huggingface/<model_type>/`](../huggingface/)** — an architecture-level
2. **[`model_type/<model_type>/`](../model_type/)** — an architecture-level
recipe that applies to every checkpoint of that `model_type`.
3. **[`general/`](../general/)** — model-agnostic recipes; a good starting point
for any model without a more specific entry.
@@ -75,7 +75,7 @@ A recipe earns a place here only when it mirrors **one specific released (or
planned) checkpoint** — a hand-mapped, usually per-layer or per-component
precision scheme tuned to match that exact release. If the tuning generalizes to
every checkpoint of an architecture, it belongs under
[`../huggingface/<model_type>/`](../huggingface/) instead; if it is
[`../model_type/<model_type>/`](../model_type/) instead; if it is
model-agnostic, it belongs under [`../general/`](../general/). See
[`../ptq.md`](../ptq.md) for what each checkpoint mirror does and how it compares
to its general baseline.
+7 -3
View File
@@ -4,7 +4,7 @@ This doc walks through the **PTQ quantization schemes** in two parts: the
model-agnostic recipes under [`general/ptq/`](general/ptq/) (the recommended
starting point for any model), and then the
[model-specific recipes](#model-specific-recipes) — per-`model_type` folders
under `huggingface/` plus the checkpoint-mirror `models/<org>/<checkpoint>/`
under `model_type/` plus the checkpoint-mirror `models/<org>/<checkpoint>/`
tier — comparing each to its general baseline and explaining why it deviates.
---
@@ -237,10 +237,14 @@ The general recipes above are **model-agnostic**: they select layers by wildcard
(`*mlp*`, `*self_attn*`, `*[kv]_bmm_quantizer`) and lean on the shared
`default_disabled_quantizers` exclusions, so the same file works on any
architecture whose module names follow the usual conventions. A recipe only
earns a place under `huggingface/<model_type>/` or
earns a place under `model_type/<model_type>/` or
`models/<org>/<checkpoint>/` when a model has to **deviate** from
that baseline. The deviations come in four kinds:
> ℹ️ `model_type/` was previously named `huggingface/`; old
> `huggingface/<model_type>/...` `--recipe` paths still resolve for backward
> compatibility, but use `model_type/` going forward.
| Kind | What changes vs. the general recipe | Examples |
|------|-------------------------------------|----------|
| **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama` |
@@ -267,7 +271,7 @@ and FP8 KV-cache-cast units. Both keep patch embedding and vision-attention BMM
precision. The shared visual snippet lives under `qwen3_vl`; thin wrappers remain discoverable
under each exact Hugging Face `model_type`.
`huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast` (and its MoE twin,
`model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast` (and its MoE twin,
which shares the same `quant_cfg` snippet) is a **mixed scheme no single general
body covers**: NVFP4 **W4A16** on MLP / expert projection weights and `lm_head`,
**FP8** on self-attention *and* the large linear-attention projections
+2 -2
View File
@@ -58,7 +58,7 @@ If extra deps are needed:
```bash
ls modelopt_recipes/models/ 2>/dev/null
ls modelopt_recipes/huggingface/<model_type>/ptq/ 2>/dev/null # per-arch; <model_type> from local config.json (Hub ID: AutoConfig.from_pretrained)
ls modelopt_recipes/model_type/<model_type>/ptq/ 2>/dev/null # per-arch; <model_type> from local config.json (Hub ID: AutoConfig.from_pretrained)
```
If a model-specific recipe exists, prefer `--recipe <path>` — but **inspect its include/exclude patterns** rather than assuming (e.g. for VLMs, confirm the vision tower is actually excluded).
@@ -72,7 +72,7 @@ Use `--qformat <name>` (e.g., `--qformat nvfp4`). Format definitions: `modelopt/
Before running PTQ, sanity-check the selected qformat/recipe against the model structure. Inspect the recipe's include/exclude patterns and summarize which layer groups will be quantized and approximately how many modules/layers match (attention projections, MLP projections, experts, etc.). If the match count is 0, or far smaller than expected for the model, stop and fix the recipe or ask the user before launching calibration.
**VLMs:** generic `*mlp*`/`*experts*` recipes also match the vision tower (`model.visual.*`); quantizing the ViT silently breaks image benchmarks. Use the `huggingface/<model_type>/ptq/` recipe or add `*visual*`/`*vision_tower*` excludes, then verify in Step 5 — see `references/checkpoint-validation.md`.
**VLMs:** generic `*mlp*`/`*experts*` recipes also match the vision tower (`model.visual.*`); quantizing the ViT silently breaks image benchmarks. Use the `model_type/<model_type>/ptq/` recipe or add `*visual*`/`*vision_tower*` excludes, then verify in Step 5 — see `references/checkpoint-validation.md`.
If the source checkpoint is already quantized and the requested recipe/config reduces quantization coverage, confirm that intent with the user before running. For example, if an FP8 checkpoint is used as input and the recipe excludes some layers so they would fall back to BF16 instead of staying quantized, call out the affected layer groups and ask whether that FP8-to-BF16 fallback is intended.
@@ -61,7 +61,7 @@ for k in vis_q[:8]: print(' ', k)
"
```
Nonzero → the ViT was quantized; re-quantize with the `huggingface/<model_type>/ptq/` recipe or add `*visual*`/`*vision_tower*` exclusions.
Nonzero → the ViT was quantized; re-quantize with the `model_type/<model_type>/ptq/` recipe or add `*visual*`/`*vision_tower*` exclusions.
## Expected quantization patterns by recipe
@@ -226,7 +226,7 @@ paths and validation gates are authoritative.
When ModelOpt is available, start from `modelopt_recipes`:
1. Check model-specific recipes first, for example
`modelopt_recipes/huggingface/<model_family>/ptq/`.
`modelopt_recipes/model_type/<model_family>/ptq/`.
2. Check general PTQ recipes and presets.
3. Use recipe fragments to build controlled manual variants.
4. Summarize include/exclude coverage before calibration. If a pattern misses the
+16 -7
View File
@@ -153,16 +153,25 @@ modelopt = ["**/*.h", "**/*.cpp", "**/*.cu"]
modelopt_recipes = ["**/*.yml", "**/*.yaml"]
[tool.setuptools.exclude-package-data]
# huggingface/models is a backward-compat symlink to ../models. The recursive
# package-data glob above follows it, so drop the aliased copies here to avoid
# shipping every checkpoint recipe twice; setuptools' exclude glob is non-recursive,
# hence the explicit org/model/task depths. MANIFEST.in prunes the symlink dir entry
# itself (build_py can't copy a symlink-to-dir). Old huggingface/models/... --recipe
# paths keep working via the loader alias in modelopt/recipe/loader.py.
# The recipe library keeps two backward-compat symlinks: the top-level
# ``huggingface`` -> ``model_type`` rename alias, and the nested
# ``model_type/models`` -> ``../models`` alias. The recursive package-data glob above
# follows both, so drop the aliased copies here to avoid shipping every recipe twice;
# setuptools' exclude glob is non-recursive, hence the explicit per-depth entries.
# MANIFEST.in prunes the symlink dir entries themselves (build_py can't copy a
# symlink-to-dir). Old huggingface/... --recipe paths keep working via the loader
# alias in modelopt/recipe/loader.py.
modelopt_recipes = [
"huggingface/models",
"huggingface",
"model_type/models",
# Architecture recipes duplicated via ``huggingface`` -> ``model_type``.
"huggingface/*/*/*.yaml", "huggingface/*/*/*.yml",
# Checkpoint mirrors duplicated via ``huggingface/models`` and
# ``model_type/models`` (both resolve to the real top-level ``models/`` tier).
"huggingface/models/*/*/*/*.yaml", "huggingface/models/*/*/*/*.yml",
"huggingface/models/*/*/*/*/*.yaml", "huggingface/models/*/*/*/*/*.yml",
"model_type/models/*/*/*/*.yaml", "model_type/models/*/*/*/*.yml",
"model_type/models/*/*/*/*/*.yaml", "model_type/models/*/*/*/*/*.yml",
]
[tool.uv]
+3 -3
View File
@@ -292,7 +292,7 @@ def test_autoquant_recipe_cost_excluded_layers_map_into_cost(monkeypatch):
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
)
aq = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
).auto_quantize
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(aq, args)
@@ -313,12 +313,12 @@ def test_autoquant_recipe_maps_module_search_spaces(monkeypatch):
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
)
recipe = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
)
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(
recipe.auto_quantize, args, fixed_quantize_config=recipe.quantize
)
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
assert inputs["quantization_formats"] == []
assert inputs["fixed_quantization_config"] == model_ptq.quantize.model_dump()
@@ -33,7 +33,7 @@ def hf_ptq(monkeypatch):
("recipe", "extracts_language_model"),
[
(None, True),
("huggingface/qwen3_vl/ptq/fp8_vision-kv_none", False),
("model_type/qwen3_vl/ptq/fp8_vision-kv_none", False),
],
)
def test_image_calibration_model_target_follows_recipe(
@@ -160,10 +160,10 @@ def test_image_calibration_uses_full_vlm_forward(hf_ptq, monkeypatch):
@pytest.mark.parametrize(
"recipe",
[
"huggingface/qwen3_vl/ptq/fp8_vision-kv_none",
"huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
"huggingface/qwen3_5/ptq/fp8_vision-kv_none",
"huggingface/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
"model_type/qwen3_vl/ptq/fp8_vision-kv_none",
"model_type/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
"model_type/qwen3_5/ptq/fp8_vision-kv_none",
"model_type/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
],
)
def test_vision_recipe_requires_image_calibration(hf_ptq, recipe):
@@ -45,12 +45,12 @@ _TINY_CONFIG = {
[
(
"embedding",
"huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj",
"model_type/nemotron_llama/ptq/nvfp4_output_quant_proj",
"TRT_FP4DynamicQuantize",
),
(
"reranking",
"huggingface/nemotron_llama/ptq/fp8_output_quant_proj",
"model_type/nemotron_llama/ptq/fp8_output_quant_proj",
"QuantizeLinear",
),
],
@@ -19,7 +19,7 @@ from _test_utils.torch.transformers_models import create_tiny_vit_dir
# Recipe variants the example ships.
_RECIPES = [
"huggingface/vit/ptq/fp8",
"model_type/vit/ptq/fp8",
]
@@ -60,7 +60,7 @@ def test_qwen_vision_recipe_calibrates_and_exports(
):
model = _get_tiny_qwen_vlm(model_type).to("cuda").eval()
model.config.architectures = [architecture]
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
vision_config = model.config.vision_config
pixel_width = (
vision_config.in_channels * vision_config.temporal_patch_size * vision_config.patch_size**2
+132 -13
View File
@@ -37,8 +37,13 @@ from modelopt.recipe.config import (
ModelOptPTQRecipe,
RecipeType,
)
from modelopt.recipe.loader import _apply_dotlist, load_config, load_recipe
from modelopt.torch.opt.config_loader import _load_raw_config, _schema_type
from modelopt.recipe.loader import _apply_dotlist, _resolve_recipe_path, load_config, load_recipe
from modelopt.torch.opt.config_loader import (
_alias_builtin_recipe_prefix,
_load_raw_config,
_resolve_config_path,
_schema_type,
)
from modelopt.torch.quantization.config import QuantizerAttributeConfig, normalize_quant_cfg_list
from modelopt.torch.quantization.mode import CalibrateModeRegistry, get_modelike_from_algo_cfg
@@ -160,28 +165,142 @@ def test_load_recipe_builtin_description():
assert len(recipe.description) > 0
def _first_builtin_ptq_recipe(root: Path, glob_pattern: str) -> Path:
"""Deterministically pick the first built-in *PTQ* recipe matching *glob_pattern*.
``glob`` order is filesystem-dependent (NTFS returns entries sorted, ext4 does not), and
recipe directories also hold non-recipe ``$import`` fragments (e.g. ``*.quant_cfg.yaml``,
``disabled_quantizers.yaml``) that are not loadable on their own. Sorting makes the pick
stable across platforms; skipping anything that does not load as a PTQ recipe keeps those
fragments from being mistaken for one.
"""
for path in sorted(root.glob(glob_pattern)):
rel = str(path.relative_to(root).with_suffix(""))
try:
if load_recipe(rel).recipe_type == RecipeType.PTQ:
return path
except Exception:
continue
raise AssertionError(f"no built-in PTQ recipe matched {glob_pattern!r} under {root}")
def test_load_recipe_huggingface_arch_backward_compat_alias():
"""Old ``huggingface/<model_type>/...`` recipe paths resolve to the renamed
``model_type/`` tier.
``huggingface/`` was renamed to ``model_type/``. A source checkout keeps a
``huggingface`` -> ``model_type`` symlink, but symlinks don't survive into built
wheels, so the loader rewrites the prefix directly. This guards that saved
``--recipe huggingface/<model_type>/...`` paths keep working for pip-installed
users, not just source checkouts.
"""
root = Path(str(files("modelopt_recipes")))
sample = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
new_path = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
old_path = "huggingface/" + new_path[len("model_type/") :] # huggingface/<arch>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(old_path)
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_huggingface_models_backward_compat_alias():
"""Old ``huggingface/models/<org>/<model_id>/...`` recipe paths resolve to the
top-level ``models/`` tier.
The restructure keeps a ``huggingface/models`` -> ``../models`` source symlink, but
symlinks don't survive into built wheels, so the loader rewrites the prefix directly.
This guards that saved ``--recipe huggingface/models/...`` paths keep working for
pip-installed users, not just source checkouts.
The ``huggingface/models/`` prefix is more specific than the ``huggingface/`` ->
``model_type/`` rename and must win: checkpoint mirrors moved all the way out to
the top-level ``models/`` tier. A source checkout keeps the
``huggingface`` -> ``model_type`` -> ``models`` symlink chain, but symlinks don't
survive into built wheels, so the loader rewrites the prefix directly. This guards
that saved ``--recipe huggingface/models/...`` paths keep working for pip-installed
users, not just source checkouts.
"""
from modelopt.recipe.loader import _resolve_recipe_path
root = Path(str(files("modelopt_recipes")))
sample = next(root.glob("models/*/*/ptq/*.yaml"))
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
new_path = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
old_path = "huggingface/" + new_path # huggingface/models/<org>/<model>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(old_path)
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
recipe = load_recipe(old_path)
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_model_type_models_alias_resolves_like_wheel():
"""``model_type/models/<org>/<model_id>/...`` resolves to the top-level ``models/`` tier.
``model_type/models`` is a source-only ``../models`` symlink that packaging prunes, so
without the loader alias the path would resolve in a checkout but 404 from a built wheel.
The alias rewrites the prefix to ``models/`` so both behave identically.
"""
root = Path(str(files("modelopt_recipes")))
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
canonical = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
aliased = "model_type/" + canonical # model_type/models/<org>/<model>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(aliased)
assert str(_resolve_recipe_path(aliased)) == str(_resolve_recipe_path(canonical))
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_local_tree_overrides_builtin_even_on_name_collision(tmp_path, monkeypatch):
"""A local recipe tree overrides a built-in of the same name — even when the name
collides with a shipped ``model_type``.
``_resolve_recipe_path`` probes the filesystem before the built-in library (matching
``config_loader._resolve_config_path``), so a user who keeps their own recipe tree on disk
is never silently shadowed by the deprecated-tier alias. This uses a *shipped* recipe's
exact relative path, spelled with the old ``huggingface/`` prefix that aliases to it, to
prove the local file wins over the built-in — the collision case a non-shipped name misses.
"""
root = Path(str(files("modelopt_recipes")))
shipped = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
# The path a user would keep locally: the shipped recipe's own relative path, but under the
# deprecated ``huggingface/`` tier that the alias rewrites to ``model_type/``.
old_rel = Path("huggingface") / shipped.relative_to(root).relative_to("model_type")
local = tmp_path / old_rel
local.parent.mkdir(parents=True)
local.write_text(
"metadata:\n recipe_type: ptq\nquantize:\n quant_cfg: {}\n algorithm: max\n"
)
monkeypatch.chdir(tmp_path)
resolved = _resolve_recipe_path(str(old_rel.with_suffix("")))
assert Path(resolved).resolve() == local.resolve()
def test_import_resolution_honors_huggingface_alias():
"""``$import`` resolution rewrites deprecated tier prefixes just like ``load_recipe``.
``$import`` paths go through ``config_loader._resolve_config_path`` (not the recipe-path
alias), so a custom recipe that imports a shipped snippet by its old ``huggingface/...``
path must still resolve from a wheel where the ``huggingface`` symlink is gone.
"""
# Prefix-rewrite mapping: architecture rename plus both checkpoint-mirror aliases.
assert _alias_builtin_recipe_prefix("huggingface/qwen3_vl/ptq/x") == "model_type/qwen3_vl/ptq/x"
assert (
_alias_builtin_recipe_prefix("huggingface/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
)
assert (
_alias_builtin_recipe_prefix("model_type/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
)
assert _alias_builtin_recipe_prefix("general/ptq/x") == "general/ptq/x" # untouched
root = Path(str(files("modelopt_recipes")))
sample = next(root.glob("model_type/*/ptq/*.yaml"))
canonical = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
old = "huggingface/" + canonical[len("model_type/") :] # huggingface/<arch>/ptq/<file>
assert str(_resolve_config_path(old)) == str(_resolve_config_path(canonical))
def _all_shipped_ptq_recipe_paths():
"""Every shipped PTQ recipe, discovered from disk rather than a hardcoded list."""
root = files("modelopt_recipes")
@@ -201,7 +320,7 @@ def _all_shipped_ptq_recipe_paths():
# Discovered from disk (not hardcoded) so the smoke tests cover every shipped PTQ
# recipe — general/, huggingface/<model_type>/, and models/<org>/<model_id>/ — and
# recipe — general/, model_type/<model_type>/, and models/<org>/<model_id>/ — and
# never drift as recipes are added, moved, or removed.
_BUILTIN_PTQ_RECIPES = _all_shipped_ptq_recipe_paths()
@@ -1905,10 +2024,10 @@ def test_load_recipe_autoquantize_builtin_active_moe():
def test_load_recipe_autoquantize_module_search_spaces():
"""Qwen recipe separates its fixed PTQ baseline from explicit search spaces."""
recipe = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
)
aq = recipe.auto_quantize
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
assert recipe.quantize is not None
assert recipe.quantize == model_ptq.quantize
assert aq.candidate_formats == []
+1 -1
View File
@@ -61,7 +61,7 @@ class _MiniMaxModel(nn.Module):
def test_mxfp8_nvfp4_experts_recipe_quantizer_precedence():
model = _MiniMaxModel()
register_fused_experts_on_the_fly(model)
recipe = load_recipe("huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
recipe = load_recipe("model_type/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
config = recipe.quantize.model_dump()
assert config["algorithm"]["layerwise"]["enable"] is True
config["algorithm"] = None
+1 -1
View File
@@ -45,7 +45,7 @@ def test_qwen_vision_recipes_select_expected_quantizers(
):
model = _get_tiny_qwen_vlm(model_type)
assert model.config.model_type == model_type
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
mtq.quantize(model, quant_cfg, forward_loop=None)
modules = dict(model.named_modules())
+41 -27
View File
@@ -89,22 +89,22 @@ def test_general_ptq_recipe_count_in_ptq_md():
def test_every_model_specific_ptq_dir_is_mentioned():
"""Every model-specific PTQ recipe must be identifiable in ptq.md.
``huggingface/<model_type>/ptq/`` recipes are checked by their ``model_type``
``model_type/<model_type>/ptq/`` recipes are checked by their ``model_type``
(e.g. ``gemma4``); ``models/<org>/<model_id>/ptq/`` recipes are checked by their
full ``<org>/<model_id>`` hub path (e.g. ``nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16``), so
the org — the whole point of the top-level tier — is verified too and an org
re-key (e.g. ``step3p5`` → ``stepfun-ai``) can't silently drift from the doc.
"""
doc = _ptq_md_text()
# model_type recipes: huggingface/<model_type>/ptq/<recipe>.yaml -> <model_type>
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "huggingface").glob("**/ptq/*.yaml")}
# model_type recipes: model_type/<model_type>/ptq/<recipe>.yaml -> <model_type>
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "model_type").glob("**/ptq/*.yaml")}
# checkpoint recipes: models/<org>/<model_id>/ptq/<recipe>.yaml -> <org>/<model_id>
model_ids = {
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
for p in (RECIPES_DIR / "models").glob("**/ptq/*.yaml")
}
identifiers = sorted(hf_ids | model_ids)
assert identifiers, "No model-specific PTQ recipes found under huggingface/ or models/"
assert identifiers, "No model-specific PTQ recipes found under model_type/ or models/"
missing = [name for name in identifiers if name not in doc]
assert not missing, (
f"Model-specific PTQ recipe folders are missing from "
@@ -114,40 +114,54 @@ def test_every_model_specific_ptq_dir_is_mentioned():
def test_checkpoint_recipes_live_in_the_top_level_models_tier():
"""Lock in the model_type-vs-checkpoint split.
"""Lock in the model_type-vs-checkpoint split and the backward-compat symlinks.
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``huggingface/``
holds only per-``model_type`` recipes. ``huggingface/models`` is kept as a
backward-compatibility **symlink** to the top-level ``models/`` tier, so the old
``--recipe huggingface/models/<org>/<model_id>/...`` paths still resolve; it must
stay a symlink that points at ``../models`` and never become a real directory that
holds recipes. A checkpoint recipe nested under a ``model_type`` (e.g.
``huggingface/<model_type>/<checkpoint>/<task>/``) still fails loudly here instead
of silently shipping both tiers — e.g. on a bad merge that re-adds the old layout.
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``model_type/``
(formerly ``huggingface/``) holds only per-``model_type`` recipes. Two
backward-compatibility **symlinks** are kept so old ``--recipe`` paths still
resolve: the top-level ``huggingface`` -> ``model_type`` rename alias, and the
nested ``model_type/models`` -> ``../models`` alias for the old
``huggingface/models/<org>/<model_id>/...`` checkpoint paths. Both must stay
symlinks and never become real directories that hold recipes. A checkpoint recipe
nested under a ``model_type`` (e.g. ``model_type/<model_type>/<checkpoint>/<task>/``)
still fails loudly here instead of silently shipping both tiers — e.g. on a bad
merge that re-adds the old layout.
"""
hf = RECIPES_DIR / "huggingface"
model_type = RECIPES_DIR / "model_type"
models = RECIPES_DIR / "models"
hf_models = hf / "models"
assert hf_models.is_symlink(), (
"huggingface/models must be a symlink to the top-level modelopt_recipes/models/ "
"tier (a backward-compat alias for the old --recipe paths), not a real directory."
hf_alias = RECIPES_DIR / "huggingface"
mt_models = model_type / "models"
# Top-level huggingface -> model_type rename alias.
assert hf_alias.is_symlink(), (
"modelopt_recipes/huggingface must be a backward-compat symlink to model_type/ "
"(the rename alias), not a real directory."
)
assert hf_models.resolve() == models.resolve(), (
f"huggingface/models must resolve to the top-level models/ tier; resolves to "
f"{hf_models.resolve()} instead of {models.resolve()}."
assert hf_alias.resolve() == model_type.resolve(), (
f"huggingface must resolve to the model_type/ tier; resolves to "
f"{hf_alias.resolve()} instead of {model_type.resolve()}."
)
# Every recipe under huggingface/ must be <model_type>/<task>/<file> (3 parts);
# Nested model_type/models -> ../models alias for the old huggingface/models/... paths.
assert mt_models.is_symlink(), (
"model_type/models must be a symlink to the top-level modelopt_recipes/models/ "
"tier (a backward-compat alias for the old --recipe huggingface/models/... paths), "
"not a real directory."
)
assert mt_models.resolve() == models.resolve(), (
f"model_type/models must resolve to the top-level models/ tier; resolves to "
f"{mt_models.resolve()} instead of {models.resolve()}."
)
# Every recipe under model_type/ must be <model_type>/<task>/<file> (3 parts);
# anything deeper is a checkpoint nested under a model_type and belongs in models/.
# Skip the huggingface/models symlink so the models/ recipes it aliases (4 parts)
# Skip the model_type/models symlink so the models/ recipes it aliases (4 parts)
# aren't miscounted as nested here.
nested = sorted(
str(p.relative_to(RECIPES_DIR))
for ext in ("*.yaml", "*.yml")
for p in hf.glob(f"**/{ext}")
if hf_models not in p.parents and len(p.relative_to(hf).parts) != 3
for p in model_type.glob(f"**/{ext}")
if mt_models not in p.parents and len(p.relative_to(model_type).parts) != 3
)
assert not nested, (
f"Recipes under huggingface/ must be <model_type>/<task>/<file>; found nested "
f"Recipes under model_type/ must be <model_type>/<task>/<file>; found nested "
f"paths (a checkpoint recipe belongs under models/<org>/<model_id>/): {nested}"
)
# Every recipe under models/ must be <org>/<model_id>/<task>/<file> (4 parts) so the
@@ -181,7 +195,7 @@ def test_launcher_yaml_recipe_paths_resolve():
# ``--recipe <p>`` / ``QUANT_CFG: <p>`` are modelopt_recipes-relative — only tier-prefixed
# values are recipe paths; bare names like ``auto`` or ``FP8_DEFAULT_CFG`` are not. The
# ``modelopt_recipes/<p>.yaml`` form (e.g. ``--config``) embeds the path directly.
tier = r"(?:general|huggingface|models|configs)/[A-Za-z0-9._/-]+"
tier = r"(?:general|model_type|models|configs)/[A-Za-z0-9._/-]+"
rel_re = re.compile(rf"(?:--recipe\s+|QUANT_CFG:\s*)({tier})")
abs_re = re.compile(rf"modelopt_recipes/({tier}\.ya?ml)")
+6 -6
View File
@@ -121,8 +121,8 @@ def _enabled(module, quantizer="weight_quantizer"):
@pytest.mark.parametrize(
"recipe_name",
[
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
],
)
def test_routed_experts_are_quantized(recipe_name):
@@ -144,8 +144,8 @@ def test_routed_experts_are_quantized(recipe_name):
@pytest.mark.parametrize(
"recipe_name",
[
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
],
)
def test_router_shared_expert_and_head_stay_bf16(recipe_name):
@@ -160,7 +160,7 @@ def test_router_shared_expert_and_head_stay_bf16(recipe_name):
def test_experts_only_leaves_dense_mlp_bf16():
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
dense_mlp = model.model.language_model.layers[1].mlp
for proj in ("gate_proj", "up_proj", "down_proj"):
@@ -168,7 +168,7 @@ def test_experts_only_leaves_dense_mlp_bf16():
def test_mlp_only_also_quantizes_dense_mlp():
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
dense_mlp = model.model.language_model.layers[1].mlp
for proj in ("gate_proj", "up_proj", "down_proj"):
@@ -40,6 +40,26 @@ from modelopt.torch.export.quant_aware_conversion import (
BLOCK = 16
def _set_scope_attr(transform, name, value):
"""Set an optional scoped-match attribute that only some transformers versions expose.
transformers>=5.9 dropped ``base_model_prefix`` from ``WeightTransform``'s ``__slots__``
(scoped matching now keys off ``scope_prefix`` alone); older supported versions still
carry it. Production ``_scope_prefixes`` reads it via ``getattr(..., None)``, so skipping
the assignment where the slot is absent is equivalent — and lets these tests run across
the whole supported transformers range instead of ``AttributeError``-ing on the setattr.
The suppression is scoped to that one known version-dependent slot: a setattr failure for
any other name (a typo or a future rename) still raises instead of silently no-op-ing.
"""
try:
setattr(transform, name, value)
except AttributeError:
if name != "base_model_prefix":
raise
# Tiny Mixtral shaped to match the synthetic expert tensors built by ``_nvfp4_linear`` below.
_MIXTRAL_KWARGS = {
"hidden_size": 32,
@@ -341,7 +361,7 @@ def test_scoped_submodel_prefix_change_does_not_capture_siblings():
prefix_change = PrefixChange(prefix_to_remove="vision_model")
prefix_change.scope_prefix = "model.vision_tower"
prefix_change.base_model_prefix = "model"
_set_scope_attr(prefix_change, "base_model_prefix", "model")
model._weight_conversions = [prefix_change]
state_dict = {
@@ -381,7 +401,7 @@ def test_scoped_rule_maps_config_module_names_consistently():
prefix_change = PrefixChange(prefix_to_remove="vision_model")
prefix_change.scope_prefix = "model.vision_tower"
prefix_change.base_model_prefix = "model"
_set_scope_attr(prefix_change, "base_model_prefix", "model")
model._weight_conversions = [prefix_change]
mapper = build_reverse_name_mapper(model)
@@ -420,7 +440,7 @@ def test_root_scoped_rule_still_faces_shadowing_guard():
)
# Root scope: reaches every key, exactly like an unscoped rule.
renaming.scope_prefix = ""
renaming.base_model_prefix = ""
_set_scope_attr(renaming, "base_model_prefix", "")
model._weight_conversions = [renaming]
state_dict = {
@@ -454,7 +474,7 @@ def test_scoped_weight_converter_is_refused():
operations=[Chunk(dim=0)],
)
conv.scope_prefix = "model.language_model"
conv.base_model_prefix = "model"
_set_scope_attr(conv, "base_model_prefix", "model")
model = types.SimpleNamespace(_weight_conversions=[conv])
sd = _nvfp4_linear("model.language_model.layers.0.mlp.gate_up_proj", 8, 16)