Files
Model-Optimizer/examples/onnx_ptq
Ajinkya RasaneandCodex 5c123ce183 [OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do?

Type of change: new example, example simplification, and
backward-breaking example migration

Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one
shared workflow:

- quantizes the shared VoVNet image backbone/encoder to INT8 or FP8;
- runs both the selected historical and current PETRv2 six-camera sweeps
through the same precision-matched TensorRT backbone engine using
distinct execution contexts during accuracy evaluation;
- keeps the PETR head and FAR3D decoder in their exported mixed
FP16/FP32 precision;
- reuses one NPZ calibration format, VoVNet exclusion helper,
quantization entry point, and TensorRT runner;
- does not change generic Model Optimizer calibration behavior or its
public CLI.

### Container boundary

Both examples use two targets from one Dockerfile, with no virtual
environments:

- `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base
with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup,
metadata generation, ONNX export, direct PyTorch calibration capture,
and final accuracy evaluation;
- `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base
for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization,
and TensorRT engine builds.

Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated
on the same GPU architecture. Final metrics remain in the evaluator
because they import the legacy model-framework postprocessing and
dataset code; only artifacts cross the container boundary through the
shared workspace.

PETR is used without patches. FAR3D applies only the official
`patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR
carries no patch files.

### Evaluator dependencies

The dependencies intentionally installed without transitive dependencies
are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on
runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator
base.

`lyft-dataset-sdk` is required only by mmdet3d's eager dataset import;
neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main
evaluator requirements because its compiled installation uses the
evaluator build step rather than the intentionally dependency-free
legacy package step.

Fresh setup and dependency approval is requested for the final reduced
dependency set.

### Reproducible PETR metadata

The documented workflow mounts raw nuScenes read-only and creates a
writable dataset view using symlinks. It then runs the pinned
mmdetection3d converter and a temporary, untracked copy of PETR's pinned
sweep generator configured only for the validation prefix and writable
dataset root.

A clean run generated both metadata files with 6,019 validation records.
The referenced camera, lidar, and sweep paths are absolute and
resolvable through the writable dataset view.

### Example-local utilities

The per-batch NPZ streaming and TensorRT runtime utilities remain
example-local because they execute in the legacy evaluator, where Model
Optimizer is not installed. The core `CalibrationDataProvider` consumes
one in-memory mapping of stacked arrays and does not provide this
streamed per-file workflow.

### Validation

- Focused CPU tests: 10 passed.
- Broader ONNX quantization CPU tests: 326 passed.
- All applicable pre-commit and documentation checks, plus `git diff
--check`, passed.
- Rebuilt both Docker targets and verified their exact dependency
versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization,
and absence of virtual environments.
- Generated both PETR metadata files from a clean writable dataset view
and verified 6,019 validation records plus resolvable data paths.
- PETRv1 passed a one-sample TensorRT regression smoke.
- PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full
6,019-sample validation. Both the selected historical and current sweeps
are computed by the matching backbone engine; accuracy evaluation no
longer extracts image features with PyTorch.
- FAR3D passed a recurrent two-frame TensorRT smoke covering plugin
loading and recurrent state.

TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after
correcting its temporal feature path; the PETRv1 and FAR3D numerical
paths are unchanged.

| Pipeline | FP16 | INT8 | FP8 |
| --- | ---: | ---: | ---: |
| PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 |
0.3707 | 0.3756 |
| PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head |
0.4102 | 0.3982 | 0.4084 |
| FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235
| 0.239 |

Normalized engine-only performance improvement over each matching FP16
pipeline:

| Pipeline | INT8 speedup | FP8 speedup |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
| FAR3D | 1.69x | 1.40x |

Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX
6000 Ada Generation GPU using five interleaved trials per engine
component. Each component uses the median `trtexec`-reported GPU Compute
Time with data transfers disabled and CUDA Graphs enabled. Component
times are summed before normalization: PETRv1 uses one backbone pass
plus its fixed head, PETRv2 uses two serial backbone passes plus its
fixed head with no temporal cache assumed, and FAR3D uses one encoder
pass plus its fixed decoder. Absolute latency values are intentionally
not published.

Adapted files retain exact public-source references and upstream
notices, and the top-level license attribution is updated.

- Is this change backward compatible?: ❌
- Did you write the necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-08 17:32:56 +00:00
..

ONNX Post-training quantization (PTQ)

This ONNX PTQ Toolkit provides a comprehensive suite of tools designed to optimize ONNX (Open Neural Network Exchange) models through quantization. Our toolkit is aimed at developers looking to enhance performance, reduce model size, and accelerate inference times without compromising the accuracy of their neural networks when deployed with TensorRT.

Quantization is an effective model optimization technique that compresses your models. Quantization with Model Optimizer can compress model size by 2x-4x, speeding up inference while preserving model quality.

Model Optimizer enables highly performant quantization formats including NVFP4, FP8, INT8, INT4 and supports advanced algorithms such as AWQ and Double Quantization with easy-to-use Python APIs.

Section Description Link Docs
Pre-Requisites Required & optional packages to use this technique Link
Getting Started Learn how to optimize your models using PTQ to reduce precision and improve inference efficiency Link docs
PyTorch to ONNX Example scripts demonstrating how to quantize with PyTorch and then convert to ONNX Link
Advanced Features Examples demonstrating use advanced ONNX quantization features Link
Resources Extra links to relevant resources Link

Pre-Requisites

Docker

Please use the TensorRT docker image (e.g., nvcr.io/nvidia/tensorrt:26.02-py3) or visit our installation docs for more information.

Note: If you are using onnxruntime-gpu, we recommend using nvcr.io/nvidia/tensorrt:25.06-py3 as it is built with CUDA 12, which is required by the stable onnxruntime-gpu package.

PETR and FAR3D containers

PETR and FAR3D share two targets from one Dockerfile. The evaluator target contains the legacy OpenMMLab stack used for data preparation, ONNX export, calibration, and final accuracy evaluation. The modelopt target uses the PyTorch 26.07 container for Model Optimizer, ONNX Runtime CUDA, and TensorRT 11.1 engine builds. Neither target creates a virtual environment.

From the Model Optimizer repository root:

docker build --target evaluator -f examples/onnx_ptq/Dockerfile -t modelopt-onnx-evaluator .
docker build --target modelopt -f examples/onnx_ptq/Dockerfile -t modelopt-onnx-trt11 .

Mount the same workspace into both containers to hand off ONNX models, calibration batches, and TensorRT engines:

docker run --rm -it --gpus=all --ipc=host \
  --user "$(id -u):$(id -g)" -e HOME=/tmp \
  -e USER="$(id -un)" -e LOGNAME="$(id -un)" \
  -v /path/to/workspace:/workspace \
  modelopt-onnx-evaluator

docker run --rm -it --gpus=all --ipc=host \
  --user "$(id -u):$(id -g)" -e HOME=/tmp \
  -e USER="$(id -un)" -e LOGNAME="$(id -un)" \
  -v /path/to/workspace:/workspace \
  modelopt-onnx-trt11

TensorRT engines must be built and evaluated with TensorRT 11.1.0.106 on the same GPU architecture. See the PETR and FAR3D guides for their source and dataset mounts.

Set the following environment variables inside the TensorRT docker.

export CUDNN_LIB_DIR=/usr/lib/x86_64-linux-gnu/
export LD_LIBRARY_PATH="${CUDNN_LIB_DIR}:${LD_LIBRARY_PATH}"

Also follow the installation steps below to upgrade to the latest version of Model Optimizer and install example-specific dependencies.

Local Installation

Install Model Optimizer with onnx dependencies using pip from PyPI and install the requirements for the example:

pip install -U nvidia-modelopt[onnx]
pip install -r requirements.txt

For TensorRT Compiler framework workloads:

Install the latest TensorRT from here.

Getting Started

Prepare the example model

Most of the examples in this doc use vit_base_patch16_224.onnx as the input model. The model can be downloaded with the following script:

python download_example_onnx.py \
    --timm_model_name=vit_base_patch16_224 \
    --onnx_save_path=vit_base_patch16_224.onnx \
    --fp16 # <Optional, if the desired output ONNX precision is FP16>

Prepare calibration data

Calibration data is a representative subset of your training or validation dataset used during quantization to determine the optimal scale factors for converting floating-point values to lower precision formats (INT8, FP8, INT4). This data helps maintain model accuracy after quantization by analyzing the distribution of activations throughout the network.

First, prepare some calibration data. TensorRT recommends calibration data size to be at least 500 for CNN and ViT models. The following command picks up 500 images from the tiny-imagenet dataset and converts them to a numpy-format calibration array. Reduce the calibration data size for resource constrained environments.

python image_prep.py \
    --calibration_data_size=500 \
    --output_path=calib.npy \
    --fp16 # <Optional, if the input ONNX is in FP16 precision>

For Int4 quantization, it is recommended to set --calibration_data_size=64.

Quantize ONNX Model to FP8, INT8 or INT4

The model can be quantized as an FP8, INT8 or INT4 model using either the CLI or Python API. For FP8 and INT8 quantization, you have a choice between max and entropy calibration algorithms. For INT4 quantization, awq_clip or rtn_dq algorithms can be chosen.

For NVFP4 and MXFP8 ONNX, see the PyTorch to ONNX example.

Minimum opset requirements: int8 (13+), fp8 (21+), int4 (21+). ModelOpt will automatically upgrade lower opset versions to meet these requirements.

Option 1: Command-line interface

python -m modelopt.onnx.quantization \
    --onnx_path=vit_base_patch16_224.onnx \
    --quantize_mode=<fp8|int8|int4> \
    --calibration_data=calib.npy \
    --calibration_method=<max|entropy|awq_clip|rtn_dq> \
    --output_path=vit_base_patch16_224.quant.onnx

Option 2: Python API

from modelopt.onnx.quantization import quantize

quantize(
    onnx_path="vit_base_patch16_224.onnx",
    quantize_mode="int8",       # fp8, int8, int4 etc.
    calibration_data="calib.npy",
    calibration_method="max",   # max, entropy, awq_clip, rtn_dq etc.
    output_path="vit_base_patch16_224.quant.onnx",
)

Evaluate the quantized ONNX model

The evaluation script automatically downloads and uses the ILSVRC/imagenet-1k dataset from Hugging Face. This gated repository requires authentication via Hugging Face access token. See https://huggingface.co/docs/hub/en/security-tokens for details. The quantized ONNX ViT model can be evaluated on the ImageNet dataset as follows:

python evaluate.py \
    --onnx_path=<path to classification model> \
    --imagenet_path=<HF dataset card or local path to the ImageNet dataset> \
    --engine_precision=stronglyTyped \
    --model_name=vit_base_patch16_224

This script converts the quantized ONNX model to a TensorRT engine and does the evaluation with that engine. Finally, the evaluation result will be reported as follows:

The top1 accuracy of the model is <accuracy score between 0-100%>
The top5 accuracy of the model is <accuracy score between 0-100%>
Inference latency of the model is <X> ms

FAR3D 3D object detection

The FAR3D example exports and quantizes the FAR3D ONNX image encoder, builds TensorRT engines, and evaluates 3D object detection mAP on the Argoverse 2 validation set.

PETR 3D object detection

The PETR example exports and quantizes the PETRv1 and PETRv2 ONNX backbones, builds TensorRT engines, and evaluates 3D object detection mAP on the nuScenes validation set.

Advanced Features

Per node calibration of ONNX models

Per node calibration is a memory optimization feature designed to reduce memory consumption during quantization of large ONNX models. Instead of running inference over the entire network at once, this feature processes the model node-by-node, which can significantly reduce peak memory usage and prevent out-of-memory (OOM) errors.

How it works

When per node calibration is enabled, the quantization process:

  1. Decomposes the model: Splits the original ONNX model into multiple single-node sub-models
  2. Manages dependencies: Tracks input/output dependencies between nodes to ensure correct execution order
  3. Processes sequentially: Runs calibration on each node individually using a topological processing order
  4. Manages memory: Automatically cleans up intermediate results and manages reference counting to minimize memory usage
  5. Aggregates results: Combines calibration data from all nodes to produce the final quantized model

When to use per node calibration

Per node calibration is particularly beneficial for:

  • Large models that cause OOM errors during standard calibration
  • Memory-constrained environments where GPU memory is limited
  • Models with complex architectures that have high intermediate memory requirements

Usage

To enable per node calibration, add the --calibrate_per_node flag to your quantization command:

python -m modelopt.onnx.quantization \
    --onnx_path=vit_base_patch16_224.onnx \
    --quantize_mode=<int8/fp8> \
    --calibration_data=calib.npy \
    --calibrate_per_node \
    --output_path=vit_base_patch16_224.quant.onnx

Note

: Per node calibration is not available for INT4 quantization methods (awq_clip, rtn_dq)

Quantize an ONNX model with custom op

This feature requires TensorRT 10+ and ORT>=1.20. For proper usage, please make sure that the paths to libcudnn*.so and TensorRT lib/ are in the LD_LIBRARY_PATH env variable and that the tensorrt python package is installed.

A self-contained example is provided in the custom_op_plugin/ subfolder, based on leimao/TensorRT-Custom-Plugin-Example. Please see the steps below.

Step 1: Build the TensorRT plugin and create the sample ONNX model.

  1.1. Compile the TensorRT plugin:

cmake -S custom_op_plugin/plugin -B /tmp/plugin_build
cmake --build /tmp/plugin_build --config Release --parallel

This generates /tmp/plugin_build/libidentity_conv_plugin.so.

  1.2. Create the ONNX model with a custom IdentityConv operator:

python custom_op_plugin/create_identity_neural_network.py \
    --output_path=/tmp/identity_neural_network.onnx

Step 2: Quantize the ONNX model using the compiled plugin.

python -m modelopt.onnx.quantization \
    --onnx_path=/tmp/identity_neural_network.onnx \
    --trt_plugins=/tmp/plugin_build/libidentity_conv_plugin.so

Step 3: Deploy the quantized model with TensorRT.

trtexec --onnx=/tmp/identity_neural_network.quant.onnx \
    --staticPlugins=/tmp/plugin_build/libidentity_conv_plugin.so

Optimize Q/DQ node placement with Autotune

This feature automates Q/DQ (Quantize/Dequantize) node placement optimization for ONNX models using TensorRT performance measurements.

To access this feature in the ONNX quantization workflow, simply add --autotune in your CLI:

python -m modelopt.onnx.quantization \
    --onnx_path=vit_base_patch16_224.onnx \
    --quantize_mode=<fp8|int8|int4> \
    --calibration_data=calib.npy \
    --calibration_method=<max|entropy|awq_clip|rtn_dq> \
    --output_path=vit_base_patch16_224.quant.onnx \
    --autotune=<quick,default,extensive>

For more fine-tuned Autotune flags, please refer to the API guide and the Autotune guide.

Resources

Technical Resources

There are many quantization schemes supported in the example scripts:

  1. The FP8 format is available on the Hopper and Ada GPUs with CUDA compute capability greater than or equal to 8.9.

  2. The INT4 AWQ is an INT4 weight only quantization and calibration method. INT4 AWQ is particularly effective for low batch inference where inference latency is dominated by weight loading time rather than the computation time itself. For low batch inference, INT4 AWQ could give lower latency than FP8/INT8 and lower accuracy degradation than INT8.

  3. The NVFP4 is one of the new FP4 formats supported by NVIDIA Blackwell GPU and demonstrates good accuracy compared with other 4-bit alternatives. NVFP4 can be applied to both model weights as well as activations, providing the potential for both a significant increase in math throughput and reductions in memory footprint and memory bandwidth usage compared to the FP8 data format on Blackwell.