Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)

### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-15 12:16:12 -07:00
committed by GitHub
parent 30f89908f0
commit c7ed23a103
75 changed files with 426 additions and 178 deletions
+5 -5
View File
@@ -201,7 +201,7 @@ python hf_ptq.py \
--export_path <quantized_ckpt_path>
```
Built-in recipes are located in `modelopt_recipes/general/ptq/` for model-agnostic recipes and in `modelopt_recipes/huggingface/<model_type>/ptq/` for recipes tuned to a specific Hugging Face `model_type` (see [`modelopt_recipes/huggingface/README.md`](../../modelopt_recipes/huggingface/README.md)). You can also provide a path to your own custom YAML recipe file or directory. See the [recipe documentation](https://nvidia.github.io/Model-Optimizer) for details on the YAML schema and available recipes.
Built-in recipes are located in `modelopt_recipes/general/ptq/` for model-agnostic recipes and in `modelopt_recipes/model_type/<model_type>/ptq/` for recipes tuned to a specific Hugging Face `model_type` (see [`modelopt_recipes/model_type/README.md`](../../modelopt_recipes/model_type/README.md)). You can also provide a path to your own custom YAML recipe file or directory. See the [recipe documentation](https://nvidia.github.io/Model-Optimizer) for details on the YAML schema and available recipes.
> *When `--recipe` is specified, `--qformat` is ignored. KV cache handling depends on the recipe type: a **PTQ** recipe bakes KV cache into its config and ignores `--kv_cache_qformat`; an **AutoQuantize** recipe falls back to `--kv_cache_qformat` unless it sets an explicit `kv_cache` field.*
@@ -287,7 +287,7 @@ Use the recipe directory matching the checkpoint's `model_type`: `qwen3_vl` or `
# Vision encoder only: FP8 vision Linears and merger, BF16 LLM and KV cache.
python hf_ptq.py \
--pyt_ckpt_path <Qwen3-VL-or-Qwen3.5-checkpoint> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none \
--calib_with_images \
--calib_size 512 \
--skip_generate \
@@ -296,7 +296,7 @@ python hf_ptq.py \
# Joint vision encoder + language model FP8 with FP8 KV-cache cast.
python hf_ptq.py \
--pyt_ckpt_path <Qwen3-VL-or-Qwen3.5-checkpoint> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \
--recipe model_type/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast \
--calib_with_images \
--calib_size 512 \
--skip_generate \
@@ -410,7 +410,7 @@ search-disabled layers, and cost-excluded layers — see
[`AutoQuantizeConfig`](../../modelopt/recipe/config.py). Shipped recipes live in
[`modelopt_recipes/general/auto_quantize/`](../../modelopt_recipes/general/auto_quantize); model-specific
recipes (carrying architecture-specific disabled layers — e.g. VL vision towers) live under
`modelopt_recipes/huggingface/<model>/auto_quantize/`.
`modelopt_recipes/model_type/<model>/auto_quantize/`.
[Script](./scripts/huggingface_example.sh)
@@ -469,7 +469,7 @@ not actually searched.
The fixed baseline may also reuse a model-specific PTQ configuration. For example, the Qwen3.6 MoE
AutoQuantize recipe imports the same model-specific `quant_cfg` used by
`huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`, reproduces that recipe's `quantize`
`model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`, reproduces that recipe's `quantize`
section, and lists only shared experts, attention, and `lm_head` under `module_search_spaces`. A
loader test asserts that the inherited fixed baseline remains equal to the original PTQ recipe while
leaving the original recipe unchanged.
+2
View File
@@ -2,5 +2,7 @@
accelerate>=1.0.0,<1.14
flash-attn>=2.6.0
liger-kernel>=0.5.0; platform_system != 'Darwin' and platform_system != 'Windows'
# peft 0.21 stops writing adapter_model.safetensors where export.py expects it.
peft<0.21
py7zr
tensorboard
+1 -1
View File
@@ -27,7 +27,7 @@ so it never loads either complete model.
python examples/minimax_m3/hf_ptq_mixed_mxfp8_nvfp4.py \
--mxfp8_ckpt /models/minimax-m3-mxfp8 \
--bf16_ckpt /models/minimax-m3-bf16 \
--recipe huggingface/minimax_m3_vl/ptq/nvfp4_experts_only \
--recipe model_type/minimax_m3_vl/ptq/nvfp4_experts_only \
--output_ckpt /models/minimax-m3-mxfp8-nvfp4 \
--device cuda
```
@@ -24,7 +24,7 @@ Usage:
python hf_ptq_mixed_mxfp8_nvfp4.py \\
--mxfp8_ckpt /models/minimax-m3-mxfp8 \\
--bf16_ckpt /models/minimax-m3-bf16 \\
--recipe huggingface/minimax_m3_vl/ptq/nvfp4_experts_only \\
--recipe model_type/minimax_m3_vl/ptq/nvfp4_experts_only \\
--output_ckpt /workspace/quant/minimax-m3-mxfp8-nvfp4-mixed \\
--device cuda
"""
+2 -2
View File
@@ -122,7 +122,7 @@ mean pooling and L2 normalization on top of the encoder; reranking
graphs take `input_ids` and `attention_mask` with dynamic batch/sequence axes.
The default recipe
(`modelopt_recipes/huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj.yaml`)
(`modelopt_recipes/model_type/nemotron_llama/ptq/nvfp4_output_quant_proj.yaml`)
quantizes weights and activations to NVFP4 and additionally quantizes the
projection-Linear outputs. Without output-side quantization, quantized GEMMs
emit FP16 activations, so FP8/FP4 engines can use as much or more activation
@@ -143,7 +143,7 @@ engines, 5 dynamic-shape profiles up to 32x512), engine activation memory:
python hf_embedding_quant_to_onnx.py \
--model_path=nvidia/llama-nemotron-embed-1b-v2 \
--trust_remote_code \
--recipe=huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj \
--recipe=model_type/nemotron_llama/ptq/nvfp4_output_quant_proj \
--onnx_save_path=llama_nemotron_embed_nvfp4.onnx
# Reranking variant (auto-detected from the model architecture)
@@ -46,7 +46,7 @@ __all__ = [
"register_bidirectional_sdpa",
]
DEFAULT_RECIPE = "huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj"
DEFAULT_RECIPE = "model_type/nemotron_llama/ptq/nvfp4_output_quant_proj"
# TODO: Add an accuracy evaluation pipeline for the embedding and reranking models.
CALIBRATION_TEXTS = [
+6 -6
View File
@@ -68,7 +68,7 @@ from modelopt.recipe import load_recipe
from modelopt.torch.quantization.utils import export_torch_mode
# 1. Quantize the eager PyTorch model with a Model Optimizer PTQ recipe.
recipe = load_recipe("huggingface/vit/ptq/fp8")
recipe = load_recipe("model_type/vit/ptq/fp8")
mtq.quantize(model, recipe.quantize.model_dump(), forward_loop=calibrate)
# 2. Compile the quantized (Q/DQ) graph with Torch-TensorRT.
@@ -138,7 +138,7 @@ This is the recipe the CLI selects by default when `--model_id` points at a HF V
| `--recipe` value | Calibration | What it quantizes |
| :---: | :---: | :--- |
| `huggingface/vit/ptq/fp8` (default) | `max` | Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the `*weight_quantizer` / `*input_quantizer` globs — encoder Linears, the patch-embed `nn.Conv2d` projection, and the `classifier` head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled. |
| `model_type/vit/ptq/fp8` (default) | `max` | Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the `*weight_quantizer` / `*input_quantizer` globs — encoder Linears, the patch-embed `nn.Conv2d` projection, and the `classifier` head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled. |
</div>
@@ -153,7 +153,7 @@ This is the recipe the CLI selects by default when `--model_id` points at a HF V
| Flag | Default | Description |
| :---: | :---: | :--- |
| `--model_id` | `google/vit-large-patch16-224` | HuggingFace model id of the ViT classifier to quantize. |
| `--recipe` | `huggingface/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--recipe` | `model_type/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--calib_samples` | `1024` | Number of tiny-imagenet samples to use for calibration. |
| `--batch_size` | `128` | Batch size for calibration / TRT compile. |
| `--save_dir` | `./modelopt_quantized` | Directory the quantized Model Optimizer state-dict (FP16 weights + Q/DQ metadata) is always saved to, as `vit_modelopt_state.pt` — re-usable across runs without recalibration. |
@@ -182,7 +182,7 @@ python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt
| Flag | Default | Description |
| :---: | :---: | :--- |
| `--model_id` | `google/vit-large-patch16-224` | HuggingFace model id of the ViT classifier to quantize and score. |
| `--recipe` | `huggingface/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--recipe` | `model_type/vit/ptq/fp8` | Recipe path (relative to `modelopt_recipes/` or an absolute YAML). |
| `--calib_samples` | `1024` | Number of tiny-imagenet samples to use for calibration. |
| `--batch_size` | `128` | Calibration / compile / eval batch size. The Torch-TRT engine is dynamic (`min=1`, `opt=max(--batch_size, 2)`, `max=1024`) and handles any batch including the trailing partial batch. |
| `--eval_data_size` | full 50k | Number of ImageNet validation images to score. |
@@ -199,7 +199,7 @@ python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt
```bash
python torch_tensorrt_accuracy.py \
--recipe huggingface/vit/ptq/fp8 \
--recipe model_type/vit/ptq/fp8 \
--batch_size 128 \
--baseline \
--eval_data_size 5000 \
@@ -215,7 +215,7 @@ python torch_tensorrt_accuracy.py \
## Custom Recipes
Use `--recipe <path>` to plug in a different recipe — either a path relative to `modelopt_recipes/` (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via `modelopt.recipe.load_recipe`, must declare `metadata.recipe_type: ptq` and a `quantize:` section, and its `quantize` config is passed straight to `mtq.quantize`. See the existing [`modelopt_recipes/huggingface/vit/ptq/*.yaml`](../../modelopt_recipes/huggingface/vit/ptq/) for the patterns used here.
Use `--recipe <path>` to plug in a different recipe — either a path relative to `modelopt_recipes/` (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via `modelopt.recipe.load_recipe`, must declare `metadata.recipe_type: ptq` and a `quantize:` section, and its `quantize` config is passed straight to `mtq.quantize`. See the existing [`modelopt_recipes/model_type/vit/ptq/*.yaml`](../../modelopt_recipes/model_type/vit/ptq/) for the patterns used here.
### Resuming From a Saved Checkpoint
+3 -3
View File
@@ -21,7 +21,7 @@ Pipeline:
2. Build a calibration loader from `zh-plus/tiny-imagenet` so the recipe runs
end-to-end without ImageNet access.
3. Run ``mtq.quantize`` with the ViT-specific FP8 recipe under
`modelopt_recipes/huggingface/vit/ptq/`.
`modelopt_recipes/model_type/vit/ptq/`.
4. Compile the quantized model with ``torch_tensorrt.compile(ir="dynamo",
min_block_size=1)`` and verify the compiled-model argmax matches the
fake-quant argmax on a sample input.
@@ -44,10 +44,10 @@ import modelopt.torch.quantization as mtq
from modelopt.recipe import ModelOptPTQRecipe, load_recipe
from modelopt.torch.quantization.utils import export_torch_mode
# Default ViT PTQ recipe under `modelopt_recipes/huggingface/vit/ptq/`. The
# Default ViT PTQ recipe under `modelopt_recipes/model_type/vit/ptq/`. The
# recipe loader resolves this relative path against the built-in recipe library;
# pass `--recipe` for a different one.
DEFAULT_RECIPE = "huggingface/vit/ptq/fp8"
DEFAULT_RECIPE = "model_type/vit/ptq/fp8"
def load_model_and_processor(model_id: str, device: torch.device, dtype: torch.dtype):