Files
Model-Optimizer/examples/torch_trt/README.md
T
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00

14 KiB

Torch-TensorRT Quantization

Torch-TensorRT compiles a PyTorch model into an optimized TensorRT engine with no separate export or runtime. This example quantizes a PyTorch / HuggingFace model with NVIDIA Model Optimizer and then compiles the quantized graph in-framework with Torch-TensorRT for deployment.

Quantization is an effective model optimization technique that compresses your models. Model Optimizer inserts Q/DQ nodes into the eager PyTorch graph; torch_tensorrt.compile(ir="dynamo") then converts those Q/DQ nodes into native TensorRT FP8 precision layers, following the Torch-TensorRT quantization guide.

This section focuses on the in-framework Torch-TensorRT path: a PyTorch front end (mtq.quantize) feeding a Dynamo-compiled TensorRT engine, demonstrated end-to-end on a HuggingFace ViT image classifier. If you instead want a portable ONNX → TensorRT artifact, or you start from an ONNX model, see the sibling torch_onnx and onnx_ptq examples (compared in the Support Matrix).

Section Description Link Docs
Pre-Requisites Required packages and installation [Link]
Getting Started Quantize and compile a ViT in a few lines [Link] [docs]
Support Matrix How this path compares to the ONNX examples [Link]
ViT Recipes The FP8 recipe shipped with the example [Link]
Usage CLI flags for the quantize and accuracy scripts [Link]
Evaluate Accuracy Measure ImageNet top-1 / top-5 accuracy [Link]
Custom Recipes Plug in your own recipe / model [Link]
Resources Roadmap, docs, benchmarks, and support [Link]

Pre-Requisites

Docker

Please use the TensorRT docker image (e.g., nvcr.io/nvidia/tensorrt:26.02-py3) or visit our installation docs for more information.

docker run --gpus all -it --rm -v $(pwd):/workspace -w /workspace nvcr.io/nvidia/tensorrt:26.02-py3 bash

Also follow the installation steps below to upgrade to the latest version of Model Optimizer and install example-specific dependencies.

Local Installation

pip install -U "nvidia-modelopt[hf]"
pip install -r requirements.txt

Hardware Requirements

The low-precision kernels Torch-TensorRT emits need a GPU that supports the target format:

Recipe Minimum GPU
fp8 Ada / Hopper — compute capability 8.9+

Note

Older GPUs still let mtq.quantize succeed — it emits fake-quant nodes in PyTorch — but torch_tensorrt.compile will not find a real low-precision kernel for an unsupported format.

Getting Started

Quantize a HuggingFace ViT, then compile the Q/DQ graph with Torch-TensorRT into a single torch.nn.Module you call from PyTorch:

import torch
import torch_tensorrt

import modelopt.torch.quantization as mtq
from modelopt.recipe import load_recipe
from modelopt.torch.quantization.utils import export_torch_mode

# 1. Quantize the eager PyTorch model with a Model Optimizer PTQ recipe.
recipe = load_recipe("model_type/vit/ptq/fp8")
mtq.quantize(model, recipe.quantize.model_dump(), forward_loop=calibrate)

# 2. Compile the quantized (Q/DQ) graph with Torch-TensorRT.
#    export_torch_mode() makes Model Optimizer emit Q/DQ in the TRT-friendly form,
#    and min_block_size=1 lets single-node Q/DQ + matmul subgraphs become TRT
#    precision layers (per the Torch-TensorRT quantization guide).
with export_torch_mode():
    trt_model = torch_tensorrt.compile(
        model,
        ir="dynamo",
        min_block_size=1,
        truncate_double=True,
        inputs=[torch_tensorrt.Input(
            min_shape=(1, 3, 224, 224),
            opt_shape=(128, 3, 224, 224),
            max_shape=(1024, 3, 224, 224),
            dtype=torch.float16,
        )],
    )

logits = trt_model(pixel_values)  # call it like any nn.Module

The runnable script torch_tensorrt_ptq.py wraps this flow end-to-end. It:

  1. Loads a HuggingFace ViT classifier (default google/vit-large-patch16-224).
  2. Builds a tiny calibration loader from zh-plus/tiny-imagenet (avoids the gated ILSVRC/imagenet-1k repo, so the example runs unauthenticated).
  3. Runs mtq.quantize with one of the recipes under modelopt_recipes/ (see ViT Recipes).
  4. Saves the quantized Model Optimizer state (FP16 weights + Q/DQ metadata) to <save_dir>/vit_modelopt_state.pt for reuse without recalibration (see Custom Recipes).
  5. Compiles the quantized model with torch_tensorrt.compile and verifies that the compiled-model argmax matches the fake-quant argmax on a sample input.
# Default model is google/vit-large-patch16-224, default recipe is the ViT FP8 recipe.
python torch_tensorrt_ptq.py --calib_samples 1024 --batch_size 128

# Quantize but don't TRT-compile (handy on a non-TRT host).
python torch_tensorrt_ptq.py --skip_trt

Note

Both torch_tensorrt_ptq.py and the accuracy script (torch_tensorrt_accuracy.py) run the model in float16.

Support Matrix

All three of these examples reach the same destination — a low-precision TensorRT engine — but quantize at a different point in the pipeline and emit a different artifact, so they suit different deployment stacks:

Torch-TensorRT (this example) torch_onnx onnx_ptq
Starting point a PyTorch / HF model a PyTorch / timm model an already-exported ONNX model
Quantize on the eager PyTorch graph (mtq.quantize) the eager PyTorch graph (mtq.quantize) the ONNX graph directly (ONNX PTQ)
Export step none — the FX/Dynamo graph stays in-process torch.onnx.export of the Q/DQ graph, postprocessed for TRT none — Q/DQ inserted straight into the ONNX graph
Intermediate artifact none a Q/DQ ONNX file a Q/DQ ONNX file
Compiler + runtime torch_tensorrt.compile(ir="dynamo") → a torch.nn.Module you call from PyTorch TensorRT builds a standalone engine from the ONNX TensorRT builds a standalone engine from the ONNX
Best when PyTorch-native serving; you want a drop-in compiled module you quantize in PyTorch but deploy via a portable ONNX → TRT engine you only have an ONNX model and never touch PyTorch

This example and torch_onnx share the same PyTorch front end (mtq.quantize), so the numerics are identical — they differ only in the back end: this one keeps the graph in-process and hands it to Torch-TensorRT, while torch_onnx exports a portable ONNX artifact for the standalone TensorRT runtime. onnx_ptq instead quantizes the ONNX graph directly, for when you start from an ONNX model rather than PyTorch. Pick this example when your serving stack is PyTorch-native and you'd rather avoid an ONNX export step.

ViT Recipes

This is the recipe the CLI selects by default when --model_id points at a HF ViT classifier. It is tuned for the HF ViT module layout and is composed from the shared $import building blocks under modelopt_recipes/configs/ (ptq/units/{w8a8_fp8_fp8,attention_qkv_fp8}) rather than spelling out each quant_cfg entry.

--recipe value Calibration What it quantizes
model_type/vit/ptq/fp8 (default) max Per-tensor FP8 (E4M3) on every weight + input quantizer matched by the *weight_quantizer / *input_quantizer globs — encoder Linears, the patch-embed nn.Conv2d projection, and the classifier head — plus FP8 on the attention Q/K/V BMMs and softmax. All output quantizers disabled.

Usage

torch_tensorrt_ptq.py

Script — quantize and (optionally) Torch-TensorRT-compile a ViT.

Flag Default Description
--model_id google/vit-large-patch16-224 HuggingFace model id of the ViT classifier to quantize.
--recipe model_type/vit/ptq/fp8 Recipe path (relative to modelopt_recipes/ or an absolute YAML).
--calib_samples 1024 Number of tiny-imagenet samples to use for calibration.
--batch_size 128 Batch size for calibration / TRT compile.
--save_dir ./modelopt_quantized Directory the quantized Model Optimizer state-dict (FP16 weights + Q/DQ metadata) is always saved to, as vit_modelopt_state.pt — re-usable across runs without recalibration.
--skip_trt off Quantize + run the fake-quant model only; skip torch_tensorrt.compile. Useful for environments without Torch-TensorRT installed.
--layer_info_path unset If set, write the compiled TRT engine's per-layer info (get_layer_info()) to this file.
# Custom model + custom recipe, saving the quantized state elsewhere.
python torch_tensorrt_ptq.py \
    --model_id <huggingface/model-id> \
    --recipe <recipe-path-relative-to-modelopt_recipes-or-absolute-yaml> \
    --save_dir ./my_quantized

# Dump the compiled engine's per-layer info to inspect FP8 fusion.
python torch_tensorrt_ptq.py --layer_info_path ./vit_fp8_layers.txt

torch_tensorrt_accuracy.py

Script — quantize, compile, and score on ImageNet (see Evaluate Accuracy).

Flag Default Description
--model_id google/vit-large-patch16-224 HuggingFace model id of the ViT classifier to quantize and score.
--recipe model_type/vit/ptq/fp8 Recipe path (relative to modelopt_recipes/ or an absolute YAML).
--calib_samples 1024 Number of tiny-imagenet samples to use for calibration.
--batch_size 128 Calibration / compile / eval batch size. The Torch-TRT engine is dynamic (min=1, opt=max(--batch_size, 2), max=1024) and handles any batch including the trailing partial batch.
--eval_data_size full 50k Number of ImageNet validation images to score.
--imagenet_path ILSVRC/imagenet-1k HF dataset card or local path to the ImageNet validation set (gated).
--baseline off Also score the unquantized model as a reference. It is Torch-TensorRT-compiled like the quantized model (or run eager under --skip_trt) so the comparison is apples-to-apples.
--skip_trt off Score the fake-quant (Model Optimizer) model; skip torch_tensorrt.compile. Useful for environments without Torch-TensorRT installed.
--results_path unset If set, write the accuracy results to this CSV path.

Evaluate Accuracy

torch_tensorrt_accuracy.py reuses the quantize → compile pipeline above and reports ImageNet-1k top-1 / top-5 accuracy via the onnx_ptq example's evaluate() harness (examples/onnx_ptq/evaluation.py):

python torch_tensorrt_accuracy.py \
    --recipe model_type/vit/ptq/fp8 \
    --batch_size 128 \
    --baseline \
    --eval_data_size 5000 \
    --results_path results.csv
  • --baseline also scores the unquantized model. It is Torch-TensorRT-compiled the same way as the quantized model, so every reported number comes from the same TRT runtime (pass --skip_trt to score the eager / fake-quant models instead).
  • The eval uses a dynamic engine (default --batch_size 128) for both precisions, so it serves the trailing partial batch at any batch size.
  • --results_path results.csv writes the metrics table (Metric, Top1 (%), Top5 (%)) to CSV.

Note

Validation uses the gated ILSVRC/imagenet-1k split: accept its license / set HF_TOKEN, or point --imagenet_path at a local copy. evaluate() shuffles the split, so a partial --eval_data_size draws a different random subset each run — omit it (full 50k set) for a stable, comparable score.

Custom Recipes

Use --recipe <path> to plug in a different recipe — either a path relative to modelopt_recipes/ (resolved against the built-in recipe library) or an absolute filesystem path to a YAML file. The recipe is loaded via modelopt.recipe.load_recipe, must declare metadata.recipe_type: ptq and a quantize: section, and its quantize config is passed straight to mtq.quantize. See the existing modelopt_recipes/model_type/vit/ptq/*.yaml for the patterns used here.

Resuming From a Saved Checkpoint

torch_tensorrt_ptq.py always saves the quantized Model Optimizer state to <save_dir>/vit_modelopt_state.pt (default --save_dir ./modelopt_quantized) via mto.save. To reload it without recalibrating, restore it onto a freshly-loaded model before the TRT compile step:

import modelopt.torch.opt as mto

mto.restore(model, "./modelopt_quantized/vit_modelopt_state.pt")

Note

See the save / restore guide for the full mto.save / mto.restore workflow.

Resources