### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
14 KiB
Torch-TensorRT Quantization
Torch-TensorRT compiles a PyTorch model into an optimized TensorRT engine with no separate export or runtime. This example quantizes a PyTorch / HuggingFace model with NVIDIA Model Optimizer and then compiles the quantized graph in-framework with Torch-TensorRT for deployment.
Quantization is an effective model optimization technique that compresses your models. Model Optimizer inserts Q/DQ nodes into the eager PyTorch graph; torch_tensorrt.compile(ir="dynamo") then converts those Q/DQ nodes into native TensorRT FP8 precision layers, following the Torch-TensorRT quantization guide.
This section focuses on the in-framework Torch-TensorRT path: a PyTorch front end (mtq.quantize) feeding a Dynamo-compiled TensorRT engine, demonstrated end-to-end on a HuggingFace ViT image classifier. If you instead want a portable ONNX → TensorRT artifact, or you start from an ONNX model, see the sibling torch_onnx and onnx_ptq examples (compared in the Support Matrix).
| Section | Description | Link | Docs |
|---|---|---|---|
| Pre-Requisites | Required packages and installation | [Link] | |
| Getting Started | Quantize and compile a ViT in a few lines | [Link] | [docs] |
| Support Matrix | How this path compares to the ONNX examples | [Link] | |
| ViT Recipes | The FP8 recipe shipped with the example | [Link] | |
| Usage | CLI flags for the quantize and accuracy scripts | [Link] | |
| Evaluate Accuracy | Measure ImageNet top-1 / top-5 accuracy | [Link] | |
| Custom Recipes | Plug in your own recipe / model | [Link] | |
| Resources | Roadmap, docs, benchmarks, and support | [Link] |
Pre-Requisites
Docker
Please use the TensorRT docker image (e.g., nvcr.io/nvidia/tensorrt:26.02-py3) or visit our installation docs for more information.
docker run --gpus all -it --rm -v $(pwd):/workspace -w /workspace nvcr.io/nvidia/tensorrt:26.02-py3 bash
Also follow the installation steps below to upgrade to the latest version of Model Optimizer and install example-specific dependencies.
Local Installation
pip install -U "nvidia-modelopt[hf]"
pip install -r requirements.txt
Hardware Requirements
The low-precision kernels Torch-TensorRT emits need a GPU that supports the target format:
| Recipe | Minimum GPU |
|---|---|
fp8 |
Ada / Hopper — compute capability 8.9+ |
Note
Older GPUs still let
mtq.quantizesucceed — it emits fake-quant nodes in PyTorch — buttorch_tensorrt.compilewill not find a real low-precision kernel for an unsupported format.
Getting Started
Quantize a HuggingFace ViT, then compile the Q/DQ graph with Torch-TensorRT into a single torch.nn.Module you call from PyTorch:
import torch
import torch_tensorrt
import modelopt.torch.quantization as mtq
from modelopt.recipe import load_recipe
from modelopt.torch.quantization.utils import export_torch_mode
# 1. Quantize the eager PyTorch model with a Model Optimizer PTQ recipe.
recipe = load_recipe("model_type/vit/ptq/fp8")
mtq.quantize(model, recipe.quantize.model_dump(), forward_loop=calibrate)
# 2. Compile the quantized (Q/DQ) graph with Torch-TensorRT.
# export_torch_mode() makes Model Optimizer emit Q/DQ in the TRT-friendly form,
# and min_block_size=1 lets single-node Q/DQ + matmul subgraphs become TRT
# precision layers (per the Torch-TensorRT quantization guide).
with export_torch_mode():
trt_model = torch_tensorrt.compile(
model,
ir="dynamo",
min_block_size=1,
truncate_double=True,
inputs=[torch_tensorrt.Input(
min_shape=(1, 3, 224, 224),
opt_shape=(128, 3, 224, 224),
max_shape=(1024, 3, 224, 224),
dtype=torch.float16,
)],
)
logits = trt_model(pixel_values) # call it like any nn.Module
The runnable script torch_tensorrt_ptq.py wraps this flow end-to-end. It:
- Loads a HuggingFace ViT classifier (default
google/vit-large-patch16-224). - Builds a tiny calibration loader from
zh-plus/tiny-imagenet(avoids the gatedILSVRC/imagenet-1krepo, so the example runs unauthenticated). - Runs
mtq.quantizewith one of the recipes undermodelopt_recipes/(see ViT Recipes). - Saves the quantized Model Optimizer state (FP16 weights + Q/DQ metadata) to
<save_dir>/vit_modelopt_state.ptfor reuse without recalibration (see Custom Recipes). - Compiles the quantized model with
torch_tensorrt.compileand verifies that the compiled-model argmax matches the fake-quant argmax on a sample input.
# Default model is google/vit-large-patch16-224, default recipe is the ViT FP8 recipe.
python torch_tensorrt_ptq.py --calib_samples 1024 --batch_size 128
# Quantize but don't TRT-compile (handy on a non-TRT host).
python torch_tensorrt_ptq.py --skip_trt
Note
Both
torch_tensorrt_ptq.pyand the accuracy script (torch_tensorrt_accuracy.py) run the model infloat16.
Support Matrix
All three of these examples reach the same destination — a low-precision TensorRT engine — but quantize at a different point in the pipeline and emit a different artifact, so they suit different deployment stacks:
| Torch-TensorRT (this example) | torch_onnx |
onnx_ptq |
|
|---|---|---|---|
| Starting point | a PyTorch / HF model | a PyTorch / timm model | an already-exported ONNX model |
| Quantize on | the eager PyTorch graph (mtq.quantize) |
the eager PyTorch graph (mtq.quantize) |
the ONNX graph directly (ONNX PTQ) |
| Export step | none — the FX/Dynamo graph stays in-process | torch.onnx.export of the Q/DQ graph, postprocessed for TRT |
none — Q/DQ inserted straight into the ONNX graph |
| Intermediate artifact | none | a Q/DQ ONNX file | a Q/DQ ONNX file |
| Compiler + runtime | torch_tensorrt.compile(ir="dynamo") → a torch.nn.Module you call from PyTorch |
TensorRT builds a standalone engine from the ONNX | TensorRT builds a standalone engine from the ONNX |
| Best when | PyTorch-native serving; you want a drop-in compiled module | you quantize in PyTorch but deploy via a portable ONNX → TRT engine | you only have an ONNX model and never touch PyTorch |
This example and torch_onnx share the same PyTorch front end (mtq.quantize), so the numerics are identical — they differ only in the back end: this one keeps the graph in-process and hands it to Torch-TensorRT, while torch_onnx exports a portable ONNX artifact for the standalone TensorRT runtime. onnx_ptq instead quantizes the ONNX graph directly, for when you start from an ONNX model rather than PyTorch. Pick this example when your serving stack is PyTorch-native and you'd rather avoid an ONNX export step.
ViT Recipes
This is the recipe the CLI selects by default when --model_id points at a HF ViT classifier. It is tuned for the HF ViT module layout and is composed from the shared $import building blocks under modelopt_recipes/configs/ (ptq/units/{w8a8_fp8_fp8,attention_qkv_fp8}) rather than spelling out each quant_cfg entry.
--recipe value |
Calibration | What it quantizes |
|---|---|---|
model_type/vit/ptq/fp8 (default) |
max |
Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the *weight_quantizer / *input_quantizer globs — encoder Linears, the patch-embed nn.Conv2d projection, and the classifier head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled. |
Usage
torch_tensorrt_ptq.py
Script — quantize and (optionally) Torch-TensorRT-compile a ViT.
| Flag | Default | Description |
|---|---|---|
--model_id |
google/vit-large-patch16-224 |
HuggingFace model id of the ViT classifier to quantize. |
--recipe |
model_type/vit/ptq/fp8 |
Recipe path (relative to modelopt_recipes/ or an absolute YAML). |
--calib_samples |
1024 |
Number of tiny-imagenet samples to use for calibration. |
--batch_size |
128 |
Batch size for calibration / TRT compile. |
--save_dir |
./modelopt_quantized |
Directory the quantized Model Optimizer state-dict (FP16 weights + Q/DQ metadata) is always saved to, as vit_modelopt_state.pt — re-usable across runs without recalibration. |
--skip_trt |
off | Quantize + run the fake-quant model only; skip torch_tensorrt.compile. Useful for environments without Torch-TensorRT installed. |
--layer_info_path |
unset | If set, write the compiled TRT engine's per-layer info (get_layer_info()) to this file. |
# Custom model + custom recipe, saving the quantized state elsewhere.
python torch_tensorrt_ptq.py \
--model_id <huggingface/model-id> \
--recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml> \
--save_dir ./my_quantized
# Dump the compiled engine's per-layer info to inspect FP8 fusion.
python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt
torch_tensorrt_accuracy.py
Script — quantize, compile, and score on ImageNet (see Evaluate Accuracy).
| Flag | Default | Description |
|---|---|---|
--model_id |
google/vit-large-patch16-224 |
HuggingFace model id of the ViT classifier to quantize and score. |
--recipe |
model_type/vit/ptq/fp8 |
Recipe path (relative to modelopt_recipes/ or an absolute YAML). |
--calib_samples |
1024 |
Number of tiny-imagenet samples to use for calibration. |
--batch_size |
128 |
Calibration / compile / eval batch size. The Torch-TRT engine is dynamic (min=1, opt=max(--batch_size, 2), max=1024) and handles any batch including the trailing partial batch. |
--eval_data_size |
full 50k | Number of ImageNet validation images to score. |
--imagenet_path |
ILSVRC/imagenet-1k |
HF dataset card or local path to the ImageNet validation set (gated). |
--baseline |
off | Also score the unquantized model as a reference. It is Torch-TensorRT-compiled like the quantized model (or run eager under --skip_trt) so the comparison is apples-to-apples. |
--skip_trt |
off | Score the fake-quant (Model Optimizer) model; skip torch_tensorrt.compile. Useful for environments without Torch-TensorRT installed. |
--results_path |
unset | If set, write the accuracy results to this CSV path. |
Evaluate Accuracy
torch_tensorrt_accuracy.py reuses the quantize → compile pipeline above and reports ImageNet-1k top-1 / top-5 accuracy via the onnx_ptq example's evaluate() harness (examples/onnx_ptq/evaluation.py):
python torch_tensorrt_accuracy.py \
--recipe model_type/vit/ptq/fp8 \
--batch_size 128 \
--baseline \
--eval_data_size 5000 \
--results_path results.csv
--baselinealso scores the unquantized model. It is Torch-TensorRT-compiled the same way as the quantized model, so every reported number comes from the same TRT runtime (pass--skip_trtto score the eager / fake-quant models instead).- The eval uses a dynamic engine (default
--batch_size 128) for both precisions, so it serves the trailing partial batch at any batch size. --results_path results.csvwrites the metrics table (Metric,Top1 (%),Top5 (%)) to CSV.
Note
Validation uses the gated
ILSVRC/imagenet-1ksplit: accept its license / setHF_TOKEN, or point--imagenet_pathat a local copy.evaluate()shuffles the split, so a partial--eval_data_sizedraws a different random subset each run — omit it (full 50k set) for a stable, comparable score.
Custom Recipes
Use --recipe <path> to plug in a different recipe — either a path relative to modelopt_recipes/ (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via modelopt.recipe.load_recipe, must declare metadata.recipe_type: ptq and a quantize: section, and its quantize config is passed straight to mtq.quantize. See the existing modelopt_recipes/model_type/vit/ptq/*.yaml for the patterns used here.
Resuming From a Saved Checkpoint
torch_tensorrt_ptq.py always saves the quantized Model Optimizer state to <save_dir>/vit_modelopt_state.pt (default --save_dir ./modelopt_quantized) via mto.save. To reload it without recalibrating, restore it onto a freshly-loaded model before the TRT compile step:
import modelopt.torch.opt as mto
mto.restore(model, "./modelopt_quantized/vit_modelopt_state.pt")
Note
See the save / restore guide for the full
mto.save/mto.restoreworkflow.