[OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)

### What does this PR do?

Type of change: new example, example simplification, and
backward-breaking example migration

Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one
shared workflow:

- quantizes the shared VoVNet image backbone/encoder to INT8 or FP8;
- runs both the selected historical and current PETRv2 six-camera sweeps
through the same precision-matched TensorRT backbone engine using
distinct execution contexts during accuracy evaluation;
- keeps the PETR head and FAR3D decoder in their exported mixed
FP16/FP32 precision;
- reuses one NPZ calibration format, VoVNet exclusion helper,
quantization entry point, and TensorRT runner;
- does not change generic Model Optimizer calibration behavior or its
public CLI.

### Container boundary

Both examples use two targets from one Dockerfile, with no virtual
environments:

- `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base
with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup,
metadata generation, ONNX export, direct PyTorch calibration capture,
and final accuracy evaluation;
- `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base
for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization,
and TensorRT engine builds.

Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated
on the same GPU architecture. Final metrics remain in the evaluator
because they import the legacy model-framework postprocessing and
dataset code; only artifacts cross the container boundary through the
shared workspace.

PETR is used without patches. FAR3D applies only the official
`patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR
carries no patch files.

### Evaluator dependencies

The dependencies intentionally installed without transitive dependencies
are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on
runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator
base.

`lyft-dataset-sdk` is required only by mmdet3d's eager dataset import;
neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main
evaluator requirements because its compiled installation uses the
evaluator build step rather than the intentionally dependency-free
legacy package step.

Fresh setup and dependency approval is requested for the final reduced
dependency set.

### Reproducible PETR metadata

The documented workflow mounts raw nuScenes read-only and creates a
writable dataset view using symlinks. It then runs the pinned
mmdetection3d converter and a temporary, untracked copy of PETR's pinned
sweep generator configured only for the validation prefix and writable
dataset root.

A clean run generated both metadata files with 6,019 validation records.
The referenced camera, lidar, and sweep paths are absolute and
resolvable through the writable dataset view.

### Example-local utilities

The per-batch NPZ streaming and TensorRT runtime utilities remain
example-local because they execute in the legacy evaluator, where Model
Optimizer is not installed. The core `CalibrationDataProvider` consumes
one in-memory mapping of stacked arrays and does not provide this
streamed per-file workflow.

### Validation

- Focused CPU tests: 10 passed.
- Broader ONNX quantization CPU tests: 326 passed.
- All applicable pre-commit and documentation checks, plus `git diff
--check`, passed.
- Rebuilt both Docker targets and verified their exact dependency
versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization,
and absence of virtual environments.
- Generated both PETR metadata files from a clean writable dataset view
and verified 6,019 validation records plus resolvable data paths.
- PETRv1 passed a one-sample TensorRT regression smoke.
- PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full
6,019-sample validation. Both the selected historical and current sweeps
are computed by the matching backbone engine; accuracy evaluation no
longer extracts image features with PyTorch.
- FAR3D passed a recurrent two-frame TensorRT smoke covering plugin
loading and recurrent state.

TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after
correcting its temporal feature path; the PETRv1 and FAR3D numerical
paths are unchanged.

| Pipeline | FP16 | INT8 | FP8 |
| --- | ---: | ---: | ---: |
| PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 |
0.3707 | 0.3756 |
| PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head |
0.4102 | 0.3982 | 0.4084 |
| FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235
| 0.239 |

Normalized engine-only performance improvement over each matching FP16
pipeline:

| Pipeline | INT8 speedup | FP8 speedup |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
| FAR3D | 1.69x | 1.40x |

Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX
6000 Ada Generation GPU using five interleaved trials per engine
component. Each component uses the median `trtexec`-reported GPU Compute
Time with data transfers disabled and CUDA Graphs enabled. Component
times are summed before normalization: PETRv1 uses one backbone pass
plus its fixed head, PETRv2 uses two serial backbone passes plus its
fixed head with no temporal cache assumed, and FAR3D uses one encoder
pass plus its fixed decoder. Absolute latency values are intentionally
not published.

Adapted files retain exact public-source references and upstream
notices, and the top-level license attribution is updated.

- Is this change backward compatible?: ❌
- Did you write the necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
This commit is contained in:
Ajinkya Rasane
2026-09-08 17:32:56 +00:00
committed by GitHub
co-authored by Codex
parent 5cae394040
commit 5c123ce183
26 changed files with 1419 additions and 574 deletions
+1
View File
@@ -1,4 +1,5 @@
docker
.git
examples/**/.git
examples/**/saved_models*
**/experimental
+2
View File
@@ -127,6 +127,8 @@ repos:
examples/llm_eval/mmlu.py|
examples/llm_eval/modeling.py|
examples/onnx_ptq/far3d/evaluate.py|
examples/onnx_ptq/petr/evaluate.py|
examples/onnx_ptq/trt_runner.py|
examples/llm_qat/train.py|
examples/llm_sparsity/weight_sparsity/finetune.py|
examples/specdec_bench/specdec_bench/models/specbench_medusa.py|
+2
View File
@@ -25,6 +25,7 @@ Changelog
- Add a Muse Glimmer AutoQuantize recipe that searches language-model MLP projections, self-attention projections, and ``lm_head`` over W4A16 NVFP4 Four-Over-Six, FP8, and BF16 fallback at 5.5 effective bits while leaving the vision tower unquantized.
- Add ``examples/alpamayo/qad.py``, which runs quantization-aware distillation on the quantized Alpamayo checkpoint produced by ``examples/alpamayo/quantize.py``. It distills the quantized VLM against the original FP16 VLM with ``QADTrainer``, supports FSDP2 for multi-GPU runs, and ``--export`` reassembles the trained VLM into a full AlpamayoR1 checkpoint that ``AlpamayoR1.from_pretrained`` can reload.
- Add a calibration-free streaming Kimi-K3 converter and checkpoint-mirror recipe for NVFP4 routed experts with ``input_scale=1.0`` and 128x128 block-FP8 KDA/MLA attention weights. The converter operates shard-by-shard on the source checkpoint's packed MXFP4 experts instead of loading the 2.8T model through the in-memory ``hf_ptq.py`` path.
- Add end-to-end PETRv1 and PETRv2 ONNX PTQ examples covering calibration, INT8 and FP8 VoVNet backbone quantization, TensorRT deployment, and accuracy evaluation.
- Add opt-in FP8 Vision Encoder recipes under the ``qwen3_vl`` and ``qwen3_5`` model types. The vision-only recipe keeps the language model and KV cache in high precision; the joint recipe quantizes Vision Encoder and language-model Linears and uses FP8 KV-cache cast. Both quantize primary and deepstack merger Linears where present, while leaving patch embedding and vision-attention BMMs in high precision. Exported checkpoints require an inference runtime that supports quantized Vision Encoder Linears.
- Add ``mtq.temporarily_fold_weights`` for repeated frozen-weight inference and ``mtq.preserve_quantizer_attributes_context`` for restoring temporary quantizer property and type changes. Temporary folding snapshots affected fake-quant weights on a configurable device and restores them with their quantizer state; retained pre-quant scales are inactive, while shared weights, shared quantizers, and ``SequentialQuantizer`` weights are unsupported.
- Add the ``nvfp4_act_headroom`` calibration algorithm for NVFP4 **activation** global scales. Instead of setting the global scale from the largest per-block amax seen during calibration (plain ``max``, which leaves no room above it so any larger activation saturates), it anchors the scale to a low percentile of the per-block amax distribution, leaving the rest of the FP8 block-scale range as headroom: ``amax = max(rho * anchor, upper)``, where ``anchor`` and ``upper`` are the per-block amaxes at ``anchor_percentile`` (default 1) and ``upper_percentile`` (default 99.99; set to 100 to never clip calibration data), and ``rho`` (default 16384) is the headroom factor. Applies only to NVFP4 dynamic-block input quantizers; ``SequentialQuantizer`` activation quantizers raise. Weight scales are an orthogonal axis selected by a nested ``weight_scale_algorithm`` (``max`` by default, or ``mse`` / ``local_hessian``), so one recipe can combine a weight calibration with this activation policy in a single pass. Ships ``modelopt_recipes/general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml``, which mirrors ``nvfp4_default-kv_fp8_cast`` with only the calibration algorithm swapped and exports a standard NVFP4 checkpoint.
@@ -48,6 +49,7 @@ Changelog
**Backward Breaking Changes**
- Migrate the FAR3D ONNX PTQ example to the shared evaluator and ModelOpt containers and ``quantize_vovnet.py``. Only the encoder supports INT8 and FP8; decoder calibration, quantization, and related CLI flags are removed, and the decoder remains in its exported mixed FP16/FP32 precision.
- Image-text calibration with ``--calib_with_images`` now forwards multimodal batches through the complete VLM for all VLM families, so existing non-Nemotron commands may produce different language-model activation ranges and output scales. Recipe-based VLM PTQ also targets the complete VLM: vision modules stay in high precision by default and are quantized only when a model-specific recipe enables them, so custom recipes must explicitly exclude vision modules when required.
- Move the checkpoint-mirror recipe tier from ``huggingface/models/<org>/<checkpoint>/`` to the top-level ``models/<org>/<model_id>/``, keyed by each recipe's canonical Hugging Face Hub id — so the Step 3.5 Flash recipe moves to ``models/stepfun-ai/Step-3.5-Flash/ptq/`` and the NVIDIA Nemotron recipes gain the ``NVIDIA-`` prefix (e.g. ``models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse``). Update any saved ``--recipe`` paths for these checkpoint recipes accordingly; the per-``model_type`` recipes under ``huggingface/`` are unchanged.
- Move the Mistral Medium 3.5 checkpoint-mirror recipe from ``huggingface/models/nvidia/Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib`` to ``models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib``, keying it by the canonical Hugging Face base model. Update any saved ``--recipe`` paths to the new location.
+1
View File
@@ -224,6 +224,7 @@ the following copyright holders, licensed under the Apache License, Version 2.0
Copyright (c) 2024 Heming Xia
Copyright 2025 The Qwen team, Alibaba Group and the HuggingFace Inc. team
Copyright (c) OpenMMLab. All rights reserved.
Copyright (c) 2022 megvii-model. All Rights Reserved.
Licensed under the Apache License, Version 2.0 (the "License"); you may not
use these files except in compliance with the License. You may obtain a copy
+63
View File
@@ -0,0 +1,63 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
FROM nvcr.io/nvidia/pytorch:22.06-py3@sha256:6f9a1fdfcbc1d1aa6f28791ed7dc41d651d7c47d634c6a11b4f0692c50c8a664 AS evaluator
RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \
libgl1-mesa-glx \
libglib2.0-0 \
libsm6 \
libxext6 && \
rm -rf /var/lib/apt/lists/*
RUN env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip install --no-cache-dir \
--extra-index-url https://download.pytorch.org/whl/cu117 \
torch==1.13.1+cu117 \
torchvision==0.14.1+cu117
COPY examples/onnx_ptq/requirements-evaluator*.txt /tmp/
# These legacy projects rely on runtime packages supplied by the digest-pinned evaluator base.
# lyft-dataset-sdk only satisfies mmdet3d's eager dataset import;
# neither example uses Lyft data.
RUN env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null FLASH_ATTENTION_FORCE_BUILD=TRUE MAX_JOBS=8 \
python -m pip install --no-cache-dir --no-build-isolation \
-r /tmp/requirements-evaluator.txt && \
env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip install \
--no-cache-dir --no-build-isolation --no-deps \
-r /tmp/requirements-evaluator-nodeps.txt
# Both TensorRT packages own the Python namespace, so install the CUDA wrapper last.
RUN env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip uninstall -y tensorrt && \
env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip install \
--no-cache-dir --force-reinstall \
tensorrt==11.1.0.106 && \
env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip install \
--no-cache-dir --force-reinstall --no-deps \
tensorrt-cu13==11.1.0.106 && \
python -c "import tensorrt as trt; assert trt.__version__ == '11.1.0.106'"
COPY examples/onnx_ptq /opt/Model-Optimizer/examples/onnx_ptq
WORKDIR /workspace
FROM nvcr.io/nvidia/pytorch:26.07-py3@sha256:2140e699b3beaf7f96a0081fd9c9406bc3832b435cdb60dfa2d261f7d2f34a1c AS modelopt
COPY . /opt/Model-Optimizer
RUN env -u PIP_CONSTRAINT PIP_CONFIG_FILE=/dev/null python -m pip install --no-cache-dir \
"/opt/Model-Optimizer[onnx]" \
"onnxruntime-gpu[cuda,cudnn]~=1.24.2" && \
python -c "import modelopt, tensorrt as trt; assert trt.__version__ == '11.1.0.106'"
WORKDIR /workspace
+34 -1
View File
@@ -26,6 +26,35 @@ Please use the TensorRT docker image (e.g., `nvcr.io/nvidia/tensorrt:26.02-py3`)
> **Note:** If you are using `onnxruntime-gpu`, we recommend using `nvcr.io/nvidia/tensorrt:25.06-py3` as it is built with CUDA 12, which is required by the stable `onnxruntime-gpu` package.
#### PETR and FAR3D containers
PETR and FAR3D share two targets from one Dockerfile. The `evaluator` target contains the legacy OpenMMLab stack used for data preparation, ONNX export, calibration, and final accuracy evaluation. The `modelopt` target uses the PyTorch 26.07 container for Model Optimizer, ONNX Runtime CUDA, and TensorRT 11.1 engine builds. Neither target creates a virtual environment.
From the Model Optimizer repository root:
```bash
docker build --target evaluator -f examples/onnx_ptq/Dockerfile -t modelopt-onnx-evaluator .
docker build --target modelopt -f examples/onnx_ptq/Dockerfile -t modelopt-onnx-trt11 .
```
Mount the same workspace into both containers to hand off ONNX models, calibration batches, and TensorRT engines:
```bash
docker run --rm -it --gpus=all --ipc=host \
--user "$(id -u):$(id -g)" -e HOME=/tmp \
-e USER="$(id -un)" -e LOGNAME="$(id -un)" \
-v /path/to/workspace:/workspace \
modelopt-onnx-evaluator
docker run --rm -it --gpus=all --ipc=host \
--user "$(id -u):$(id -g)" -e HOME=/tmp \
-e USER="$(id -un)" -e LOGNAME="$(id -un)" \
-v /path/to/workspace:/workspace \
modelopt-onnx-trt11
```
TensorRT engines must be built and evaluated with TensorRT 11.1.0.106 on the same GPU architecture. See the [PETR](./petr/) and [FAR3D](./far3d/) guides for their source and dataset mounts.
Set the following environment variables inside the TensorRT docker.
```bash
@@ -131,7 +160,11 @@ Inference latency of the model is <X> ms
### FAR3D 3D object detection
The [FAR3D example](./far3d/) demonstrates an end-to-end workflow that exports and quantizes the FAR3D ONNX image encoder, builds TensorRT engines, and evaluates 3D object detection mAP on the Argoverse 2 validation set.
The [FAR3D example](./far3d/) exports and quantizes the FAR3D ONNX image encoder, builds TensorRT engines, and evaluates 3D object detection mAP on the Argoverse 2 validation set.
### PETR 3D object detection
The [PETR example](./petr/) exports and quantizes the PETRv1 and PETRv2 ONNX backbones, builds TensorRT engines, and evaluates 3D object detection mAP on the nuScenes validation set.
## Advanced Features
-56
View File
@@ -1,56 +0,0 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# TensorRT 11.1 from the base image builds and runs the FAR3D engines.
FROM nvcr.io/nvidia/pytorch:26.07-py3
ENV LD_LIBRARY_PATH=/usr/local/cuda/compat/lib:/usr/local/nvidia/lib:/usr/local/nvidia/lib64
RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \
libgl1 \
libglib2.0-0 && \
rm -rf /var/lib/apt/lists/*
ENV UV_PYTHON_INSTALL_DIR=/opt/python
# FAR3D requires the legacy PyTorch 1.13/MMCV stack in Python 3.8. ModelOpt is installed
# separately below in the base image's Python 3.12 environment.
RUN python -m pip install --no-cache-dir uv && \
uv python install 3.8 && \
uv venv --seed --python 3.8 /opt/far3d
COPY examples/onnx_ptq/far3d/requirements*.txt /tmp/far3d-requirements/
RUN env -u PIP_CONSTRAINT /opt/far3d/bin/python -m pip install --no-cache-dir \
-r /tmp/far3d-requirements/requirements-torch.txt && \
env -u PIP_CONSTRAINT /opt/far3d/bin/python -m pip install --no-cache-dir \
-r /tmp/far3d-requirements/requirements.txt && \
env -u PIP_CONSTRAINT /opt/far3d/bin/python -m pip install --no-cache-dir \
--no-build-isolation \
-r /tmp/far3d-requirements/requirements-mmdet3d.txt && \
mkdir -p /opt/far3d/lib/python3.8/site-packages/tensorrt && \
cp /opt/far3d/lib/python3.8/site-packages/tensorrt_bindings/__init__.py \
/opt/far3d/lib/python3.8/site-packages/tensorrt/__init__.py && \
cp /opt/far3d/lib/python3.8/site-packages/tensorrt_bindings/tensorrt.so \
/opt/far3d/lib/python3.8/site-packages/tensorrt/tensorrt.so
COPY . /opt/Model-Optimizer
RUN cd /opt/Model-Optimizer && \
env -u PIP_CONSTRAINT python -m pip install --no-cache-dir \
-e ".[onnx]" \
"onnxruntime-gpu[cuda,cudnn]~=1.24.2" \
"tensorrt-cu12-libs==10.11.0.33"
# The TensorRT EP in ONNX Runtime 1.24 requires TensorRT 10 during decoder quantization.
ENV ORT_TRT10_LIB_PATH=/usr/local/lib/python3.12/dist-packages/tensorrt_libs
+83 -117
View File
@@ -1,163 +1,129 @@
# FAR3D ONNX PTQ and Argoverse 2 evaluation
This example quantizes the FAR3D image encoder and decoder to INT8 or FP8 with Model Optimizer and evaluates the complete pipeline on the Argoverse 2 validation set. It follows the [NVIDIA DL4AGX FAR3D workflow](https://github.com/NVIDIA/DL4AGX/tree/master/AV-Solutions/far3d-trt).
This example quantizes the FAR3D VoVNet image encoder to INT8 or FP8, keeps the decoder in its exported mixed FP16/FP32 precision, and evaluates TensorRT 11.1 engines on the Argoverse 2 validation set. It follows the [NVIDIA DL4AGX FAR3D workflow](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/far3d-trt).
FAR3D uses a legacy PyTorch/MMCV environment that is incompatible with the current Model Optimizer Python dependencies. The provided image uses `nvcr.io/nvidia/pytorch:26.07-py3` with TensorRT 11.1 for engine build and evaluation, and isolates the legacy FAR3D packages in a Python 3.8 virtual environment. The TensorRT EP in ONNX Runtime 1.24 requires CUDA 12 and TensorRT 10.11 compatibility libraries during decoder quantization; these libraries are not used to build or run the TensorRT 11.1 engines.
Build the shared `evaluator` and `modelopt` images as described in the [parent guide](../README.md#petr-and-far3d-containers). Use the evaluator for source setup, metadata, export, calibration, and accuracy evaluation. Use the ModelOpt image for AutoCast, quantization, and TensorRT engine builds.
## 1. Prepare FAR3D and Argoverse 2
Clone DL4AGX, initialize its submodules, and apply its FAR3D patch:
Download the [Argoverse 2 sensor validation set](https://www.argoverse.org/av2.html) on the host, then start the evaluator with the workspace and dataset mounted:
```bash
git clone https://github.com/NVIDIA/DL4AGX.git
cd DL4AGX
git submodule update --init --recursive
cd AV-Solutions/far3d-trt/dependencies/Far3D
git apply ../../patch/far3d.patch
git apply /path/to/Model-Optimizer/examples/onnx_ptq/far3d/far3d_optional_flash_attn.patch
cd ../..
docker run --rm -it --gpus=all --ipc=host \
--user "$(id -u):$(id -g)" -e HOME=/tmp \
-e USER="$(id -un)" -e LOGNAME="$(id -un)" \
-v /path/to/workspace:/workspace \
-v /path/to/av2_sensor:/data/av2:ro \
modelopt-onnx-evaluator
```
The second patch makes the unused CUDA 11-only FlashAttention implementation optional; the reference configuration uses MMCV `MultiheadAttention`.
Clone the pinned DL4AGX tree and apply its official FAR3D export patch. This is the only source patch in the workflow.
Download the [Argoverse 2 sensor validation set](https://www.argoverse.org/av2.html), the [reference FAR3D checkpoint](https://github.com/NVIDIA/DL4AGX/tree/master/AV-Solutions/far3d-trt#pytorch-model-to-onnx), and its configuration. The remaining commands assume:
```bash
git clone https://github.com/NVIDIA/DL4AGX.git /workspace/DL4AGX
git -C /workspace/DL4AGX checkout 9f7b29104c253d5bc68334e7b83b3eecb72d4572
git -C /workspace/DL4AGX submodule update --init \
AV-Solutions/far3d-trt/dependencies/Far3D \
AV-Solutions/far3d-trt/dependencies/mmdetection3d
git -C /workspace/DL4AGX/AV-Solutions/far3d-trt/dependencies/Far3D \
apply ../../patch/far3d.patch
```
Download the [FAR3D checkpoint](https://github.com/megvii-research/Far3D/releases/download/v1.0/iter_82548.pth). Keep the raw dataset read-only and store generated metadata in the workspace:
```text
far3d-trt/
├── data/av2/val/
├── dependencies/Far3D/projects/configs/far3d.py
/workspace/DL4AGX/AV-Solutions/far3d-trt/
├── data/av2/
│ └── val -> /data/av2/val
└── weights/iter_82548.pth
```
Build the example image from the Model Optimizer checkout:
```bash
docker build \
-f /path/to/Model-Optimizer/examples/onnx_ptq/far3d/Dockerfile \
-t far3d-modelopt \
/path/to/Model-Optimizer
cd /workspace/DL4AGX/AV-Solutions/far3d-trt
# Replace DL4AGX's dataset-root symlink so generated metadata stays in the workspace.
unlink data/av2
mkdir -p data/av2 weights
ln -s /data/av2/val data/av2/val
```
Start the image and mount the FAR3D checkout:
## 2. Export and calibrate in the evaluator
```bash
docker run --rm -it --network=host --gpus=all --shm-size=80G --privileged \
-v /data/av2:/data/av2 \
-v /path/to/far3d-trt:/workspace/far3d-trt \
far3d-modelopt
```
cd /workspace/DL4AGX/AV-Solutions/far3d-trt
export PYTHONPATH=$PWD/dependencies/Far3D
Use `/opt/far3d/bin/python` for data preparation, export, and evaluation. It selects the isolated legacy FAR3D environment:
```bash
export PYTHONPATH=/workspace/far3d-trt/dependencies/Far3D
cd /workspace/far3d-trt
/opt/far3d/bin/python /opt/Model-Optimizer/examples/onnx_ptq/far3d/prepare_metadata.py data/av2
```
## 2. Export the ONNX models
```bash
/opt/far3d/bin/python tools/export_onnx.py \
python /opt/Model-Optimizer/examples/onnx_ptq/far3d/prepare_metadata.py data/av2
python tools/export_onnx.py \
dependencies/Far3D/projects/configs/far3d.py \
weights/iter_82548.pth
```
This produces `far3d.encoder.onnx` and `far3d.decoder.onnx`.
## 3. Prepare calibration batches
Build temporary engines from the exported models. They run the reference pipeline while collecting representative encoder and decoder inputs:
```bash
trtexec \
--onnx=far3d.encoder.onnx \
--saveEngine=far3d.encoder.fp16.engine \
--fp16 \
--skipInference
trtexec \
--onnx=far3d.decoder.onnx \
--saveEngine=far3d.decoder.fp16.engine \
--stronglyTyped \
--skipInference
```
Extract 512 batches sampled every 20 frames from the Argoverse 2 validation loader:
```bash
/opt/far3d/bin/python /opt/Model-Optimizer/examples/onnx_ptq/far3d/prepare_calibration.py \
python /opt/Model-Optimizer/examples/onnx_ptq/far3d/prepare_calibration.py \
dependencies/Far3D/projects/configs/far3d.py \
data/far3d_calibration \
--encoder-engine far3d.encoder.fp16.engine \
--decoder-engine far3d.decoder.fp16.engine \
--num-samples 512 \
--sample-skip-interval 20
far3d.encoder.onnx calibration/encoder
```
The calibration directory contains separate `encoder/` and `decoder/` batches. Decoder batches include the image features, camera geometry, and temporal state seen by the reference decoder.
The calibration command writes 512 NPZ batches directly from the data loader. No temporary TensorRT engine or decoder calibration data is needed.
## 4. Quantize the models
## 3. Optimize and build in the ModelOpt container
Use the base Python environment for Model Optimizer:
Restart the workspace with `modelopt-onnx-trt11`, then run:
```bash
LD_LIBRARY_PATH="${ORT_TRT10_LIB_PATH}:${LD_LIBRARY_PATH}" \
python /opt/Model-Optimizer/examples/onnx_ptq/far3d/quantize.py \
--encoder-onnx far3d.encoder.onnx \
--decoder-onnx far3d.decoder.onnx \
--calibration-dir data/far3d_calibration
cd /workspace/DL4AGX/AV-Solutions/far3d-trt
python -m modelopt.onnx.autocast \
--onnx_path far3d.encoder.onnx \
--output_path far3d.encoder.fp16.onnx \
--calibration_data calibration/encoder/batch_0000.npz \
--low_precision_type fp16 --keep_io_types --providers cuda:0 cpu
for precision in int8 fp8; do
python /opt/Model-Optimizer/examples/onnx_ptq/quantize_vovnet.py \
far3d.encoder.onnx calibration/encoder \
--precision "$precision" --output "far3d.encoder.${precision}.onnx"
done
for precision in fp16 int8 fp8; do
trtexec --onnx="far3d.encoder.${precision}.onnx" \
--saveEngine="far3d.encoder.${precision}.engine" --skipInference
done
trtexec --onnx=far3d.decoder.onnx \
--saveEngine=far3d.decoder.mixed.engine --skipInference
```
Both models use max calibration. INT8 is the default; use `--quantization-mode fp8` to produce `far3d.encoder.fp8.onnx` and `far3d.decoder.fp8.onnx` instead. FP8 deployment requires an FP8-capable GPU.
TensorRT 11.1 uses typed ONNX graphs; neither `--fp16` nor `--stronglyTyped` is needed. Serialized engines are not portable across TensorRT versions or GPU architectures.
The quantizer preserves the accuracy-sensitive exclusions used by the DL4AGX reference: the `OSA4_5` block and nodes downstream of `lateral_convs` remain in high precision.
## 4. Evaluate in the evaluator
To keep the decoder in its original mixed FP16/FP32 precision, add `--fp16-decoder`; decoder calibration batches are not required in that mode. This flag can be combined with either quantization mode.
Build both engines in the same container. Serialized TensorRT engines are not portable across TensorRT versions or GPU architectures.
Set the precision to the quantization mode used above:
Restart `modelopt-onnx-evaluator` with the same mounts:
```bash
precision=int8 # Use fp8 for FP8 models.
trtexec \
--onnx=far3d.encoder.${precision}.onnx \
--saveEngine=far3d.encoder.${precision}.engine \
--stronglyTyped \
--skipInference
trtexec \
--onnx=far3d.decoder.${precision}.onnx \
--saveEngine=far3d.decoder.${precision}.engine \
--stronglyTyped \
--skipInference
cd /workspace/DL4AGX/AV-Solutions/far3d-trt
export PYTHONPATH=$PWD/dependencies/Far3D
for precision in fp16 int8 fp8; do
python /opt/Model-Optimizer/examples/onnx_ptq/far3d/evaluate.py \
dependencies/Far3D/projects/configs/far3d.py \
"far3d.encoder.${precision}.engine" far3d.decoder.mixed.engine
done
```
When using `--fp16-decoder`, build `far3d.decoder.onnx` as `far3d.decoder.fp16.engine` instead.
Add `--max-samples 2` for a smoke test that also exercises recurrent decoder state. Full validation contains 23,522 frames.
## 5. Evaluate accuracy
## Reference accuracy and performance
```bash
precision=int8 # Use fp8 for FP8 models.
/opt/far3d/bin/python /opt/Model-Optimizer/examples/onnx_ptq/far3d/evaluate.py \
dependencies/Far3D/projects/configs/far3d.py \
far3d.encoder.${precision}.engine \
far3d.decoder.${precision}.engine
```
Accuracy was measured with TensorRT 11.1.0.106 on an NVIDIA RTX 6000 Ada Generation GPU using 512 calibration batches.
Use `--max-samples N` for an inference smoke test. Dataset metrics are skipped when only part of the validation set is processed.
| Encoder | Decoder | mAP |
| --- | --- | ---: |
| FP16 | Mixed FP16/FP32 | 0.241 |
| INT8 | Mixed FP16/FP32 | 0.235 |
| FP8 | Mixed FP16/FP32 | 0.239 |
## Results on Argoverse 2 validation set
Engine-only performance is normalized to the FP16 pipeline. Each comparison comprises one encoder pass plus the same exported mixed FP16/FP32 decoder; only the encoder precision changes.
The following historical results use TensorRT 10.11.0.33 on an NVIDIA RTX 6000 Ada Generation GPU. Model quantization uses PyTorch 2.8.0a0 from the 25.06 PyTorch container, while the FAR3D export and evaluation environment uses PyTorch 1.13.1. Accuracy is measured over all 23,522 validation frames after calibration with 512 batches sampled every 20 frames. These numbers are not directly reproducible with the current 26.07/TensorRT 11.1 image; rerun the workflow to measure the current toolchain.
Measurements use TensorRT 11.1.0.106 on an NVIDIA RTX 6000 Ada Generation GPU with five interleaved trials per engine component. Each component uses the median `trtexec`-reported GPU Compute Time with data transfers disabled and CUDA Graphs enabled. Component times are summed before normalization. Only speedups are reported.
| Encoder precision | Decoder precision | Framework | GPU compute time (ms) | Accuracy (mAP) |
| --- | --- | --- | ---: | ---: |
| FP32 | FP32 | TensorRT 10.11 | 92.5 | 0.241 |
| FP16 | FP32 | TensorRT 10.11 | 47.8 | 0.241 |
| FP16 | FP16 | TensorRT 10.11 | 45.0 | 0.241 |
| INT8 | FP16 | TensorRT 10.11 | 24.6 | 0.236 |
| FP8 | FP16 | TensorRT 10.11 | 31.5 | 0.241 |
Quantizing the decoder to INT8 or FP8 produced severe accuracy degradation in this evaluation and is not recommended. Keep the decoder in its original mixed FP16/FP32 precision.
GPU compute time is the sum of the encoder and decoder median times reported by `trtexec`, with host-to-device and device-to-host transfers disabled. Results depend on the TensorRT version and GPU architecture and are not directly comparable with the DRIVE Orin-X measurements in the [DL4AGX reference](https://github.com/NVIDIA/DL4AGX/tree/master/AV-Solutions/far3d-trt#results-on-argoverse2-validation-set).
| Pipeline | INT8 speedup vs. FP16 | FP8 speedup vs. FP16 |
| --- | ---: | ---: |
| FAR3D | 1.69x | 1.40x |
+18 -128
View File
@@ -21,9 +21,10 @@
import argparse
import importlib
import os
import sys
import warnings
from pathlib import Path
import tensorrt as trt
import torch
from mmcv import Config, DictAction
from mmcv.utils import import_modules_from_strings
@@ -33,115 +34,9 @@ from mmdet3d.datasets import build_dataset
from projects.mmdet3d_plugin.datasets.builder import build_dataloader
from tqdm import tqdm
TRT_TO_TORCH = {
trt.DataType.FLOAT: torch.float32,
trt.DataType.HALF: torch.float16,
trt.DataType.INT8: torch.int8,
trt.DataType.INT32: torch.int32,
trt.DataType.BOOL: torch.bool,
trt.DataType.UINT8: torch.uint8,
}
if int(trt.__version__.split(".")[0]) >= 10:
TRT_TO_TORCH[trt.DataType.INT64] = torch.int64
TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
trt.init_libnvinfer_plugins(TRT_LOGGER, "")
def aligned_tensor(shape, dtype, device, alignment=256):
element_size = torch.empty((), dtype=dtype).element_size()
element_count = int(torch.tensor(shape).prod().item())
storage = torch.empty(element_count + alignment // element_size, dtype=dtype, device=device)
offset_bytes = (-storage.data_ptr()) % alignment
offset = offset_bytes // element_size
return storage[offset : offset + element_count].view(shape)
class TensorRTRunner:
def __init__(self, engine_path, state_names=()):
with open(engine_path, "rb") as engine_file:
engine_bytes = engine_file.read()
self.engine = trt.Runtime(TRT_LOGGER).deserialize_cuda_engine(engine_bytes)
if self.engine is None:
raise RuntimeError(f"Failed to deserialize {engine_path}")
self.context = self.engine.create_execution_context()
if self.context is None:
raise RuntimeError(f"Failed to create an execution context for {engine_path}")
self.tensor_names = [
self.engine.get_tensor_name(index) for index in range(self.engine.num_io_tensors)
]
self.input_shapes = {}
self.output_shapes = {}
self.tensor_dtypes = {}
for name in self.tensor_names:
shape = tuple(self.engine.get_tensor_shape(name))
dtype = TRT_TO_TORCH[self.engine.get_tensor_dtype(name)]
self.tensor_dtypes[name] = dtype
if self.engine.get_tensor_mode(name) == trt.TensorIOMode.INPUT:
self.input_shapes[name] = shape
else:
self.output_shapes[name] = shape
self.state = {}
for base_name in state_names:
name = self.resolve_name(base_name)
if name in self.input_shapes:
tensor = aligned_tensor(self.input_shapes[name], self.tensor_dtypes[name], "cuda")
tensor.zero_()
self.state[name] = tensor
self.context.set_tensor_address(name, tensor.data_ptr())
if self.state:
torch.cuda.synchronize()
def resolve_name(self, base_name):
if base_name in self.tensor_names:
return base_name
suffixed_name = f"{base_name}.1"
return suffixed_name if suffixed_name in self.tensor_names else base_name
def reset_state(self):
for tensor in self.state.values():
tensor.zero_()
def prepare_input(self, name, inputs):
shape = self.input_shapes[name]
base_name = name.rsplit(".1", maxsplit=1)[0] if name.endswith(".1") else name
if base_name not in inputs:
raise KeyError(f"Missing TensorRT input {base_name}")
value = inputs[base_name].to(device="cuda", dtype=self.tensor_dtypes[name])
if tuple(value.shape) != shape:
if tuple(value.shape[1:]) == shape:
value = value.squeeze(0)
elif tuple(shape[1:]) == tuple(value.shape):
value = value.unsqueeze(0)
else:
raise ValueError(
f"Input {base_name} has shape {tuple(value.shape)}, expected {shape}"
)
return value
def __call__(self, stream, **inputs):
input_buffers = {}
for name, shape in self.input_shapes.items():
if name in self.state:
continue
value = self.prepare_input(name, inputs)
buffer = aligned_tensor(shape, value.dtype, value.device)
buffer.copy_(value)
input_buffers[name] = buffer
self.context.set_tensor_address(name, buffer.data_ptr())
outputs = {}
for name, shape in self.output_shapes.items():
output = aligned_tensor(shape, self.tensor_dtypes[name], "cuda")
outputs[name] = output
self.context.set_tensor_address(name, output.data_ptr())
if not self.context.execute_async_v3(stream.cuda_stream):
raise RuntimeError("TensorRT execution failed")
stream.synchronize()
return outputs
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.trt_runner import TensorRTRunner
STATE_NAMES = (
"memory_embedding",
@@ -152,10 +47,13 @@ STATE_NAMES = (
)
def get_image_input(data):
return data["img"][0].data[0].flip(2).permute(0, 1, 3, 4, 2).contiguous()
class Far3DDecoderRunner(TensorRTRunner):
def __init__(self, engine_path, input_callback=None):
super().__init__(engine_path, STATE_NAMES)
self.input_callback = input_callback
def __init__(self, engine_path):
super().__init__(engine_path, state_names=STATE_NAMES)
self.scene_token = None
self.timestamp_offset = None
@@ -175,16 +73,6 @@ class Far3DDecoderRunner(TensorRTRunner):
device="cuda",
)
inputs["timestamp"] = (timestamp - self.timestamp_offset).float()
if self.input_callback:
calibration_inputs = {}
for name in self.input_shapes:
base_name = name.rsplit(".1", maxsplit=1)[0] if name.endswith(".1") else name
if name in self.state:
value = self.state[name]
else:
value = self.prepare_input(name, inputs)
calibration_inputs[base_name] = value
self.input_callback(calibration_inputs)
outputs = super().__call__(stream, **inputs)
for base_name in STATE_NAMES:
input_name = self.resolve_name(base_name)
@@ -196,15 +84,15 @@ class Far3DDecoderRunner(TensorRTRunner):
class Far3DPipeline:
def __init__(self, encoder_engine, decoder_engine, decoder_input_callback=None):
def __init__(self, encoder_engine, decoder_engine):
self.encoder = TensorRTRunner(encoder_engine)
self.decoder = Far3DDecoderRunner(decoder_engine, decoder_input_callback)
self.decoder = Far3DDecoderRunner(decoder_engine)
@staticmethod
def unpack(data):
lidar2img = data["lidar2img"][0].data[0][0].unsqueeze(0).cuda()
return {
"img": data["img"][0].data[0].flip(2).permute(0, 1, 3, 4, 2).contiguous().cuda(),
"img": get_image_input(data).cuda(),
"intrinsics": data["intrinsics"][0].data[0][0].unsqueeze(0).cuda(),
"extrinsics": data["extrinsics"][0].data[0][0].unsqueeze(0).cuda(),
"lidar2img": lidar2img,
@@ -246,8 +134,9 @@ def parse_args():
def import_plugin(cfg):
plugin_dir = os.path.dirname(cfg.plugin_dir).split("/")
importlib.import_module(".".join(plugin_dir))
plugin_dir = cfg.get("plugin_dir")
if cfg.get("plugin") and plugin_dir:
importlib.import_module(".".join(os.path.dirname(plugin_dir).split("/")))
def main():
@@ -277,6 +166,7 @@ def main():
outputs = []
for data in tqdm(data_loader):
result = pipeline(stream, data)
torch.cuda.current_stream().wait_stream(stream)
boxes = LiDARInstance3DBoxes(result["bboxes"].cpu())
outputs.append(
{
@@ -287,7 +177,7 @@ def main():
}
}
)
if args.max_samples is not None and len(outputs) == args.max_samples:
if args.max_samples is not None and len(outputs) >= args.max_samples:
break
if len(outputs) < len(dataset):
@@ -1,17 +0,0 @@
--- a/projects/mmdet3d_plugin/models/utils/petr_transformer.py
+++ b/projects/mmdet3d_plugin/models/utils/petr_transformer.py
@@ -17,3 +17,6 @@
from torch.nn import ModuleList
-from .attention import FlashMHA
+try:
+ from .attention import FlashMHA
+except ImportError:
+ FlashMHA = None
import torch.utils.checkpoint as cp
@@ -65,3 +68,5 @@
self.batch_first = True
-
+ if FlashMHA is None:
+ raise ImportError("flash-attn is required for PETRMultiheadFlashAttention")
+
self.attn = FlashMHA(embed_dims, num_heads, attn_drop, dtype=torch.float16, device='cuda',
+18 -69
View File
@@ -14,24 +14,26 @@
# limitations under the License.
import argparse
import sys
from pathlib import Path
import numpy as np
import torch
from evaluate import Far3DPipeline
from evaluate import get_image_input
from mmcv import Config
from mmdet.datasets import replace_ImageToTensor
from mmdet3d.datasets import build_dataset
from projects.mmdet3d_plugin.datasets.builder import build_dataloader
from torch.utils.data import Subset
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationWriter
def parse_args():
parser = argparse.ArgumentParser(description="Prepare FAR3D calibration batches")
parser = argparse.ArgumentParser(description="Prepare FAR3D encoder calibration batches")
parser.add_argument("config", help="Path to the FAR3D configuration file")
parser.add_argument("encoder_onnx")
parser.add_argument("output_dir", type=Path)
parser.add_argument("--encoder-engine")
parser.add_argument("--decoder-engine")
parser.add_argument("--num-samples", type=int, default=512)
parser.add_argument("--sample-skip-interval", type=int, default=20)
return parser.parse_args()
@@ -61,9 +63,8 @@ def build_validation_loader(config_path, num_samples, sample_skip_interval):
min(len(dataset), num_samples * sample_skip_interval),
sample_skip_interval,
)
dataset = Subset(dataset, sample_indices)
return build_dataloader(
dataset,
Subset(dataset, sample_indices),
samples_per_gpu=samples_per_gpu,
workers_per_gpu=cfg.data.workers_per_gpu,
dist=False,
@@ -72,71 +73,19 @@ def build_validation_loader(config_path, num_samples, sample_skip_interval):
)
class DecoderCalibrationWriter:
def __init__(self, output_dir):
self.output_dir = output_dir
self.saved = 0
def __call__(self, inputs):
batch = {name: value.detach().cpu().numpy() for name, value in inputs.items()}
np.savez(self.output_dir / f"batch_{self.saved:04d}.npz", **batch)
self.saved += 1
def main():
args = parse_args()
if args.num_samples < 1:
raise ValueError("--num-samples must be positive")
if args.sample_skip_interval < 1:
raise ValueError("--sample-skip-interval must be positive")
if bool(args.encoder_engine) != bool(args.decoder_engine):
raise ValueError("--encoder-engine and --decoder-engine must be specified together")
if args.num_samples < 1 or args.sample_skip_interval < 1:
raise ValueError("Sample count and skip interval must be positive")
encoder_dir = args.output_dir / "encoder"
encoder_dir.mkdir(parents=True, exist_ok=True)
if any(encoder_dir.glob("*.npy")):
raise FileExistsError(
f"{encoder_dir} already contains calibration batches; use an empty directory"
)
writer = NpzCalibrationWriter(args.output_dir, args.encoder_onnx)
loader = build_validation_loader(args.config, args.num_samples, args.sample_skip_interval)
for data in loader:
writer.write({"img": get_image_input(data)})
decoder_writer = pipeline = None
if args.encoder_engine:
decoder_dir = args.output_dir / "decoder"
decoder_dir.mkdir(parents=True, exist_ok=True)
if any(decoder_dir.glob("*.npz")):
raise FileExistsError(
f"{decoder_dir} already contains calibration batches; use an empty directory"
)
decoder_writer = DecoderCalibrationWriter(decoder_dir)
pipeline = Far3DPipeline(
args.encoder_engine,
args.decoder_engine,
decoder_input_callback=decoder_writer,
)
stream = torch.cuda.Stream()
saved = 0
data_loader = build_validation_loader(args.config, args.num_samples, args.sample_skip_interval)
for data in data_loader:
images = data["img"][0].data[0].cpu().permute(0, 1, 3, 4, 2).numpy()
np.save(encoder_dir / f"batch_{saved:04d}.npy", images)
if pipeline:
pipeline(stream, data)
saved += 1
if saved == args.num_samples:
break
if saved < args.num_samples:
raise RuntimeError(
f"Only prepared {saved} of {args.num_samples} requested calibration batches"
)
if decoder_writer and decoder_writer.saved != saved:
raise RuntimeError(f"Prepared {saved} encoder and {decoder_writer.saved} decoder batches")
print(f"Saved {saved} encoder calibration batches to {encoder_dir}")
if decoder_writer:
print(
f"Saved {decoder_writer.saved} decoder calibration batches to {decoder_writer.output_dir}"
)
if writer.count != args.num_samples:
raise RuntimeError(f"Prepared {writer.count} batches; expected {args.num_samples}")
print(f"Saved {writer.count} calibration batches to {args.output_dir}")
if __name__ == "__main__":
-158
View File
@@ -1,158 +0,0 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import re
from pathlib import Path
import numpy as np
import onnx
from onnxruntime.quantization.calibrate import CalibrationDataReader
from modelopt.onnx.quantization import quantize
from modelopt.onnx.utils import topologically_sort_graph_nodes
class FileCalibrationReader(CalibrationDataReader):
def __init__(self, calibration_dir, pattern):
self.batch_paths = sorted(Path(calibration_dir).glob(pattern))
if not self.batch_paths:
raise ValueError(f"No {pattern} calibration batches found in {calibration_dir}")
self.rewind()
def get_next(self):
batch_path = next(self._iterator, None)
return None if batch_path is None else self.load(batch_path)
def get_first(self):
return self.load(self.batch_paths[0])
def rewind(self):
self._iterator = iter(self.batch_paths)
def load(self, batch_path):
raise NotImplementedError
class EncoderCalibrationReader(FileCalibrationReader):
def __init__(self, calibration_dir):
super().__init__(calibration_dir, "*.npy")
def load(self, batch_path):
return {"img": np.load(batch_path)}
class DecoderCalibrationReader(FileCalibrationReader):
def __init__(self, calibration_dir, onnx_path):
graph = onnx.load(onnx_path, load_external_data=False).graph
self.input_dtypes = {
value.name: onnx.helper.tensor_dtype_to_np_dtype(value.type.tensor_type.elem_type)
for value in graph.input
}
super().__init__(calibration_dir, "*.npz")
def load(self, batch_path):
with np.load(batch_path) as batch:
missing = self.input_dtypes.keys() - batch.files
if missing:
raise ValueError(f"{batch_path} is missing decoder inputs: {sorted(missing)}")
return {
name: batch[name].astype(dtype, copy=False)
for name, dtype in self.input_dtypes.items()
}
def find_encoder_nodes_to_exclude(onnx_path):
graph = onnx.load(onnx_path, load_external_data=False).graph
topologically_sort_graph_nodes(graph)
excluded = set()
downstream_tensors = set()
for node in graph.node:
is_osa = "OSA4_5" in node.name
is_downstream = any(name in downstream_tensors for name in node.input)
if is_osa or is_downstream:
excluded.add(node.name)
if "lateral_convs" in node.name or (is_downstream and not is_osa):
downstream_tensors.update(node.output)
return sorted(excluded)
def parse_args():
parser = argparse.ArgumentParser(description="Quantize the FAR3D ONNX models")
parser.add_argument("--encoder-onnx", required=True, help="Path to far3d.encoder.onnx")
parser.add_argument("--decoder-onnx", required=True, help="Path to far3d.decoder.onnx")
parser.add_argument(
"--calibration-dir", required=True, help="Directory created by prepare_calibration.py"
)
parser.add_argument("--quantization-mode", choices=("int8", "fp8"), default="int8")
parser.add_argument("--encoder-output")
parser.add_argument("--decoder-output")
parser.add_argument(
"--fp16-decoder",
action="store_true",
help="Skip decoder quantization and use the original mixed-precision decoder",
)
return parser.parse_args()
def quantize_encoder(args):
encoder_dir = Path(args.calibration_dir)
if (encoder_dir / "encoder").is_dir():
encoder_dir /= "encoder"
excluded_nodes = [
rf"^{re.escape(name)}$" for name in find_encoder_nodes_to_exclude(args.encoder_onnx)
]
print(f"Excluding {len(excluded_nodes)} accuracy-sensitive nodes from quantization")
quantize(
onnx_path=args.encoder_onnx,
quantize_mode=args.quantization_mode,
calibration_data_reader=EncoderCalibrationReader(encoder_dir),
calibration_method="max",
calibration_eps=["cuda:0", "cpu"],
nodes_to_exclude=excluded_nodes,
high_precision_dtype="fp16",
output_path=args.encoder_output,
)
def quantize_decoder(args):
decoder_dir = Path(args.calibration_dir) / "decoder"
quantize(
onnx_path=args.decoder_onnx,
quantize_mode=args.quantization_mode,
calibration_data_reader=DecoderCalibrationReader(decoder_dir, args.decoder_onnx),
calibration_method="max",
calibration_eps=["cuda:0", "cpu"],
high_precision_dtype="fp16" if args.quantization_mode == "fp8" else "fp32",
output_path=args.decoder_output,
)
def main():
args = parse_args()
if args.encoder_output is None:
args.encoder_output = f"far3d.encoder.{args.quantization_mode}.onnx"
if args.decoder_output is None:
args.decoder_output = f"far3d.decoder.{args.quantization_mode}.onnx"
quantize_encoder(args)
if args.fp16_decoder:
print("Skipping decoder quantization; use the original mixed-precision decoder ONNX")
else:
quantize_decoder(args)
if __name__ == "__main__":
main()
@@ -1 +0,0 @@
mmdet3d==1.0.0rc6
@@ -1,4 +0,0 @@
--extra-index-url https://download.pytorch.org/whl/cu117
torch==1.13.1+cu117
torchvision==0.14.1+cu117
-23
View File
@@ -1,23 +0,0 @@
# Dependencies for the isolated FAR3D Python 3.8 environment. ModelOpt and its ONNX
# dependencies are installed separately in the base Python environment.
--extra-index-url https://pypi.nvidia.com
--find-links https://download.openmmlab.com/mmcv/dist/cu117/torch1.13.0/index.html
av2==0.2.1
einops
ipython<9
kornia==0.6.12
mmcv-full==1.7.0
mmdet==2.28.2
mmsegmentation==0.30.0
numpy<1.24
onnx
onnx-graphsurgeon==0.6.1
onnxruntime
onnxsim
opencv-python==4.5.5.64
refile
setuptools<81
tensorrt-cu13-bindings==11.1.0.106
yapf==0.32.0
+218
View File
@@ -0,0 +1,218 @@
# PETR ONNX PTQ and nuScenes evaluation
This example quantizes the VoVNet image backbone in PETRv1 and PETRv2 to INT8 or FP8, keeps the detection head in mixed FP16/FP32, and evaluates TensorRT 11.1 engines on the nuScenes validation set. It follows the [NVIDIA DL4AGX PETR workflow](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/petr-trt).
Build the shared `evaluator` and `modelopt` images as described in the [parent guide](../README.md#petr-and-far3d-containers). Use the evaluator for source setup, export, calibration, and accuracy evaluation. Use the ModelOpt image for AutoCast, quantization, and TensorRT engine builds.
## 1. Prepare PETR and nuScenes
Place the nuScenes data in the host dataset directory. Then start the evaluator with the workspace and raw dataset mounted:
```bash
docker run --rm -it --gpus=all --ipc=host \
--user "$(id -u):$(id -g)" -e HOME=/tmp \
-e USER="$(id -un)" -e LOGNAME="$(id -un)" \
-v /path/to/workspace:/workspace \
-v /path/to/nuscenes:/data/nuscenes:ro \
modelopt-onnx-evaluator
```
The final workspace layout is:
```text
/workspace/
├── DL4AGX/
├── PETR/
│ └── ckpts/
│ ├── PETR-vov-p4-800x320_e24.pth
│ └── PETRv2-vov-p4-800x320_e24.pth
└── nuscenes/
├── nuscenes_infos_val.pkl
└── mmdet3d_nuscenes_30f_infos_val.pkl
```
Pin the source repositories. PETR does not need a patch; dataset paths are passed through its existing configuration overrides.
The mmdetection3d checkout supplies the base configuration files that PETR references by relative path.
```bash
git clone https://github.com/NVIDIA/DL4AGX.git /workspace/DL4AGX
git -C /workspace/DL4AGX checkout 9f7b29104c253d5bc68334e7b83b3eecb72d4572
git clone https://github.com/megvii-research/PETR.git /workspace/PETR
git -C /workspace/PETR checkout f7525f93467a33707ef401c587a52d5e7b34de74
git clone https://github.com/open-mmlab/mmdetection3d.git /workspace/PETR/mmdetection3d
git -C /workspace/PETR/mmdetection3d checkout f1107977dfd26155fc1f83779ee6535d2468f449
mkdir -p /workspace/PETR/ckpts
```
Download the checkpoints linked from the DL4AGX guide. Use nuScenes only under its [terms of use](https://www.nuscenes.org/terms-of-use).
Create a writable view of the read-only dataset, then generate the standard nuScenes metadata with the pinned [mmdetection3d data converter](https://github.com/open-mmlab/mmdetection3d/blob/f1107977dfd26155fc1f83779ee6535d2468f449/tools/data_converter/nuscenes_converter.py). Run it from `/tmp` so the converter keeps the `/workspace/nuscenes` paths absolute.
```bash
mkdir -p /workspace/nuscenes
for name in lidarseg maps panoptic samples sweeps v1.0-trainval; do
ln -sfn "/data/nuscenes/$name" "/workspace/nuscenes/$name"
done
(
cd /tmp
PYTHONPATH=/workspace/PETR/mmdetection3d/tools python -c \
'from data_converter.nuscenes_converter import create_nuscenes_infos; create_nuscenes_infos("/workspace/nuscenes", "nuscenes", version="v1.0-trainval", max_sweeps=10)'
)
```
The pinned [PETR sweep generator](https://github.com/megvii-research/PETR/blob/f7525f93467a33707ef401c587a52d5e7b34de74/tools/generate_sweep_pkl.py) uses fixed training paths. Run a temporary validation configuration without modifying the PETR checkout:
```bash
sed \
-e "s/^info_prefix = 'train'$/info_prefix = 'val'/" \
-e 's#^data_root = "/data/Dataset/nuScenes/"$#data_root = "/workspace/nuscenes/"#' \
/workspace/PETR/tools/generate_sweep_pkl.py \
> /tmp/generate_sweep_pkl_val.py
python /tmp/generate_sweep_pkl_val.py
test -s /workspace/nuscenes/nuscenes_infos_val.pkl
test -s /workspace/nuscenes/mmdet3d_nuscenes_30f_infos_val.pkl
```
## 2. Export and calibrate in the evaluator
```bash
cd /workspace/DL4AGX/AV-Solutions/petr-trt/export_eval
export PYTHONPATH=/workspace/PETR:$PWD
mkdir -p onnx_files engines calibration
DATA_ROOT=/workspace/nuscenes
V1_CONFIG=/workspace/PETR/projects/configs/petr/petr_vovnet_gridmask_p4_800x320.py
V2_CONFIG=/workspace/PETR/projects/configs/petrv2/petrv2_vovnet_gridmask_p4_800x320.py
V1_CHECKPOINT=/workspace/PETR/ckpts/PETR-vov-p4-800x320_e24.pth
V2_CHECKPOINT=/workspace/PETR/ckpts/PETRv2-vov-p4-800x320_e24.pth
V1_INFO="$DATA_ROOT/nuscenes_infos_val.pkl"
V2_INFO="$DATA_ROOT/mmdet3d_nuscenes_30f_infos_val.pkl"
```
Export both models without modifying the PETR checkout:
```bash
python v1/v1_export_to_onnx.py "$V1_CONFIG" "$V1_CHECKPOINT" --eval bbox \
--cfg-options \
data.val.data_root="$DATA_ROOT/" data.val.ann_file="$V1_INFO" \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V1_INFO"
python v2/v2_export_to_onnx.py "$V2_CONFIG" "$V2_CHECKPOINT" --eval bbox \
--cfg-options \
data.val.data_root="$DATA_ROOT/" data.val.ann_file="$V2_INFO" \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V2_INFO"
for model in PETRv1 PETRv2; do
python -m onnxsim "onnx_files/${model}.extract_feat.onnx" \
"onnx_files/${model}.backbone.onnx"
python -m onnxsim "onnx_files/${model}.pts_bbox_head.forward.onnx" \
"onnx_files/${model}.head.onnx"
done
```
Collect 512 backbone batches and one representative head batch directly from PyTorch. No temporary TensorRT engines are needed.
For PETRv2 calibration, PyTorch supplies the historical feature inputs because the TensorRT engines have not been built yet. Accuracy evaluation does not reuse that path: it computes both sweeps with the selected TensorRT backbone engine.
```bash
python /opt/Model-Optimizer/examples/onnx_ptq/petr/prepare_calibration.py \
v1 "$V1_CONFIG" "$V1_CHECKPOINT" \
onnx_files/PETRv1.backbone.onnx onnx_files/PETRv1.head.onnx calibration/PETRv1 \
--cfg-options \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V1_INFO"
python /opt/Model-Optimizer/examples/onnx_ptq/petr/prepare_calibration.py \
v2 "$V2_CONFIG" "$V2_CHECKPOINT" \
onnx_files/PETRv2.backbone.onnx onnx_files/PETRv2.head.onnx calibration/PETRv2 \
--cfg-options \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V2_INFO"
```
## 3. Optimize and build in the ModelOpt container
Restart the workspace with `modelopt-onnx-trt11`, then return to the export directory:
```bash
cd /workspace/DL4AGX/AV-Solutions/petr-trt/export_eval
for model in PETRv1 PETRv2; do
python -m modelopt.onnx.autocast \
--onnx_path "onnx_files/${model}.backbone.onnx" \
--output_path "onnx_files/${model}.backbone.fp16.onnx" \
--calibration_data "calibration/${model}/backbone/batch_0000.npz" \
--low_precision_type fp16 --keep_io_types --providers cuda:0 cpu
python -m modelopt.onnx.autocast \
--onnx_path "onnx_files/${model}.head.onnx" \
--output_path "onnx_files/${model}.head.fp16.onnx" \
--calibration_data "calibration/${model}/head/batch_0000.npz" \
--low_precision_type fp16 --keep_io_types --providers cuda:0 cpu
for precision in int8 fp8; do
python /opt/Model-Optimizer/examples/onnx_ptq/quantize_vovnet.py \
"onnx_files/${model}.backbone.onnx" "calibration/${model}/backbone" \
--precision "$precision" \
--output "onnx_files/${model}.backbone.${precision}.onnx"
done
trtexec --onnx="onnx_files/${model}.head.fp16.onnx" \
--saveEngine="engines/${model}.head.fp16.engine" --skipInference
for precision in fp16 int8 fp8; do
trtexec --onnx="onnx_files/${model}.backbone.${precision}.onnx" \
--saveEngine="engines/${model}.backbone.${precision}.engine" --skipInference
done
done
```
TensorRT 11.1 uses typed ONNX graphs; the removed `--fp16` builder flag is not used.
## 4. Evaluate in the evaluator
Restart `modelopt-onnx-evaluator` with the same mounts, restore the variables from step 2, and run each backbone precision with the shared typed mixed FP16/FP32 head:
```bash
cd /workspace/DL4AGX/AV-Solutions/petr-trt/export_eval
export PYTHONPATH=/workspace/PETR:$PWD
for precision in fp16 int8 fp8; do
python /opt/Model-Optimizer/examples/onnx_ptq/petr/evaluate.py \
v1 "$V1_CONFIG" "$V1_CHECKPOINT" \
"engines/PETRv1.backbone.${precision}.engine" engines/PETRv1.head.fp16.engine \
--cfg-options \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V1_INFO"
python /opt/Model-Optimizer/examples/onnx_ptq/petr/evaluate.py \
v2 "$V2_CONFIG" "$V2_CHECKPOINT" \
"engines/PETRv2.backbone.${precision}.engine" engines/PETRv2.head.fp16.engine \
--cfg-options \
data.test.data_root="$DATA_ROOT/" data.test.ann_file="$V2_INFO"
done
```
Add `--max-samples 1` for an export-to-inference smoke test. Full validation contains 6,019 samples.
PETRv2 evaluates the selected historical six-camera sweep and then the current six-camera sweep through the same backbone engine at the selected precision. The first pass supplies temporal features to the second; no PyTorch image feature extraction is used during accuracy evaluation.
## Reference accuracy and performance
Accuracy was measured with TensorRT 11.1.0.106 on an NVIDIA RTX 6000 Ada Generation GPU using 512 calibration batches and the full 6,019-sample validation set.
| Model | Backbone | Head | mAP |
| --- | --- | --- | ---: |
| PETRv1 | FP16 | Mixed FP16/FP32 | 0.3778 |
| PETRv1 | INT8 | Mixed FP16/FP32 | 0.3707 |
| PETRv1 | FP8 | Mixed FP16/FP32 | 0.3756 |
| PETRv2 | FP16 | Mixed FP16/FP32 | 0.4102 |
| PETRv2 | INT8 | Mixed FP16/FP32 | 0.3982 |
| PETRv2 | FP8 | Mixed FP16/FP32 | 0.4084 |
Engine-only performance is normalized to each model's FP16 pipeline. PETRv1 comprises one backbone pass plus its fixed typed mixed FP16/FP32 head; PETRv2 comprises two serial backbone passes plus its fixed typed mixed FP16/FP32 head, with no temporal cache assumed.
Measurements use TensorRT 11.1.0.106 on an NVIDIA RTX 6000 Ada Generation GPU with five interleaved trials per engine component. Each component uses the median `trtexec`-reported GPU Compute Time with data transfers disabled and CUDA Graphs enabled. Component times are summed before normalization. Only speedups are reported.
| Model | INT8 speedup vs. FP16 | FP8 speedup vs. FP16 |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
+179
View File
@@ -0,0 +1,179 @@
# Adapted from https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/petr-trt/export_eval/v1/v1_evaluate_trt.py
# and https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/petr-trt/export_eval/v2/v2_evaluate_trt.py.
# Copyright (c) OpenMMLab. All rights reserved.
#
# SPDX-FileCopyrightText: Copyright (c) 2023-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import importlib
import os
import sys
from pathlib import Path
import torch
import torch.nn.functional as F
from mmcv import Config, DictAction
from mmcv.runner import load_checkpoint, wrap_fp16_model
from mmcv.utils import import_modules_from_strings
from mmdet.apis import set_random_seed
from mmdet3d.core import bbox3d2result
from mmdet3d.datasets import build_dataloader, build_dataset
from mmdet3d.models import build_model
from tqdm import tqdm
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.petr.petr_utils import run_backbone
from examples.onnx_ptq.trt_runner import TensorRTRunner
__all__ = ["PETRPipeline", "build_runtime", "get_head_inputs"]
def import_plugin(cfg):
if cfg.get("custom_imports"):
import_modules_from_strings(**cfg.custom_imports)
plugin_dir = cfg.get("plugin_dir")
if cfg.get("plugin") and plugin_dir:
importlib.import_module(".".join(os.path.dirname(plugin_dir).split("/")))
def build_runtime(config_path, checkpoint_path, cfg_options=None):
cfg = Config.fromfile(config_path)
if cfg_options:
cfg.merge_from_dict(cfg_options)
import_plugin(cfg)
cfg.model.pretrained = None
cfg.model.train_cfg = None
cfg.data.test.test_mode = True
dataset = build_dataset(cfg.data.test)
loader = build_dataloader(
dataset,
samples_per_gpu=1,
workers_per_gpu=cfg.data.workers_per_gpu,
dist=False,
shuffle=False,
)
model = build_model(cfg.model, test_cfg=cfg.get("test_cfg"))
if cfg.get("fp16"):
wrap_fp16_model(model)
checkpoint = load_checkpoint(model, checkpoint_path, map_location="cpu")
model.CLASSES = checkpoint.get("meta", {}).get("CLASSES", dataset.CLASSES)
if hasattr(dataset, "PALETTE"):
model.PALETTE = checkpoint.get("meta", {}).get("PALETTE", dataset.PALETTE)
model = model.cuda().eval()
return cfg, dataset, loader, model
def get_head_inputs(version, model, features, img_metas):
batch_size, num_cams = features[0].shape[:2]
input_h, input_w, _ = img_metas[0]["pad_shape"][0]
masks = features[0].new_ones((batch_size, num_cams, input_h, input_w))
for image_id in range(batch_size):
for camera_id in range(num_cams):
image_h, image_w, _ = img_metas[image_id]["img_shape"][camera_id]
masks[image_id, camera_id, :image_h, :image_w] = 0
masks = F.interpolate(masks, size=features[0].shape[-2:]).to(torch.bool)
coords, _ = model.pts_bbox_head.position_embeding(features, img_metas, masks)
inputs = {
"mlvl_feats.0": features[0],
"img_metas.0[coords_position_embeding]": coords,
}
if version == "v2":
timestamps = features[0].new_tensor([meta["timestamp"] for meta in img_metas])
timestamps = timestamps.view(1, -1, 6)
inputs["img_metas.0[mean_time_stamp]"] = (timestamps[:, 1] - timestamps[:, 0]).mean(-1)
return inputs
class PETRPipeline:
def __init__(self, version, model, backbone_engine, head_engine):
self.version = version
self.model = model
self.backbone = TensorRTRunner(backbone_engine)
self.history_backbone = (
self.backbone.new_context(state_names=("prev.0", "prev.1")) if version == "v2" else None
)
self.head = TensorRTRunner(head_engine)
def __call__(self, stream, data):
images = data["img"][0].data[0].cuda()
img_metas = data["img_metas"][0].data[0]
with torch.cuda.stream(stream), torch.no_grad():
feature_outputs = run_backbone(
self.version, self.backbone, self.history_backbone, stream, images
)
camera_count = 6 if self.version == "v1" else 12
features = [
feature_outputs[name].reshape(1, camera_count, *feature_outputs[name].shape[-3:])
for name in ("out.0", "out.1")
]
outputs = self.head(
stream, **get_head_inputs(self.version, self.model, features, img_metas)
)
head_outputs = {
"all_cls_scores": outputs["out.all_cls_scores"].float(),
"all_bbox_preds": outputs["out.all_bbox_preds"].float(),
"enc_cls_scores": None,
"enc_bbox_preds": None,
}
boxes = self.model.pts_bbox_head.get_bboxes(head_outputs, img_metas, rescale=True)
torch.cuda.current_stream().wait_stream(stream)
return [
{"pts_bbox": bbox3d2result(boxes_3d, scores_3d, labels_3d)}
for boxes_3d, scores_3d, labels_3d in boxes
]
def parse_args():
parser = argparse.ArgumentParser(description="Evaluate PETR TensorRT engines on nuScenes")
parser.add_argument("version", choices=("v1", "v2"))
parser.add_argument("config")
parser.add_argument("checkpoint")
parser.add_argument("backbone_engine")
parser.add_argument("head_engine")
parser.add_argument("--cfg-options", nargs="+", action=DictAction)
parser.add_argument("--eval-options", nargs="+", action=DictAction)
parser.add_argument("--max-samples", type=int)
args = parser.parse_args()
if args.max_samples is not None and args.max_samples < 1:
raise ValueError("--max-samples must be positive")
return args
def main():
args = parse_args()
set_random_seed(0, deterministic=False)
cfg, dataset, loader, model = build_runtime(args.config, args.checkpoint, args.cfg_options)
pipeline = PETRPipeline(args.version, model, args.backbone_engine, args.head_engine)
stream = torch.cuda.Stream()
outputs = []
for data in tqdm(loader):
outputs.extend(pipeline(stream, data))
if args.max_samples is not None and len(outputs) >= args.max_samples:
break
if len(outputs) < len(dataset):
print(f"Processed {len(outputs)} samples; skipping dataset metrics")
return
eval_kwargs = cfg.get("evaluation", {}).copy()
for key in ("interval", "tmpdir", "start", "gpu_collect", "save_best", "rule"):
eval_kwargs.pop(key, None)
if args.eval_options:
eval_kwargs.update(args.eval_options)
print(dataset.evaluate(outputs, **eval_kwargs))
if __name__ == "__main__":
main()
+38
View File
@@ -0,0 +1,38 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
__all__ = ["run_backbone"]
_CAMERAS_PER_SWEEP = 6
_FEATURE_OUTPUT_NAMES = ("out.0", "out.1")
def run_backbone(version, backbone, history_backbone, stream, images):
if version == "v1":
return backbone(stream, img=images.squeeze(0))
current = images[:, :_CAMERAS_PER_SWEEP].contiguous()
history = images[:, _CAMERAS_PER_SWEEP : 2 * _CAMERAS_PER_SWEEP].contiguous()
history_outputs = history_backbone(stream, img=history.squeeze(0))
history_features = {
f"prev.{index}": history_outputs[name][:, :_CAMERAS_PER_SWEEP]
for index, name in enumerate(_FEATURE_OUTPUT_NAMES)
}
return backbone(
stream,
img=current.squeeze(0),
**history_features,
)
@@ -0,0 +1,95 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import sys
from pathlib import Path
import torch
from evaluate import build_runtime, get_head_inputs
from mmcv import DictAction
from mmdet3d.datasets import build_dataloader
from torch.utils.data import Subset
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationWriter
def get_backbone_inputs(version, model, images, img_metas):
if version == "v1":
return {"img": images.squeeze(0)}
current = images[:, :6].contiguous()
previous = images[:, 6:12].contiguous()
previous_features = model.extract_img_feat(previous, img_metas)
return {
"img": current.squeeze(0),
**{f"prev.{index}": value for index, value in enumerate(previous_features)},
}
def parse_args():
parser = argparse.ArgumentParser(description="Prepare PETR calibration batches")
parser.add_argument("version", choices=("v1", "v2"))
parser.add_argument("config")
parser.add_argument("checkpoint")
parser.add_argument("backbone_onnx")
parser.add_argument("head_onnx")
parser.add_argument("output_dir", type=Path)
parser.add_argument("--num-samples", type=int, default=512)
parser.add_argument("--sample-skip-interval", type=int, default=10)
parser.add_argument("--cfg-options", nargs="+", action=DictAction)
return parser.parse_args()
def main():
args = parse_args()
if args.num_samples < 1 or args.sample_skip_interval < 1:
raise ValueError("Sample count and skip interval must be positive")
cfg, dataset, _, model = build_runtime(args.config, args.checkpoint, args.cfg_options)
stop = min(len(dataset), args.num_samples * args.sample_skip_interval)
subset = Subset(dataset, range(args.sample_skip_interval - 1, stop, args.sample_skip_interval))
loader = build_dataloader(
subset,
samples_per_gpu=1,
workers_per_gpu=cfg.data.workers_per_gpu,
dist=False,
shuffle=False,
)
backbone_writer = NpzCalibrationWriter(args.output_dir / "backbone", args.backbone_onnx)
head_writer = NpzCalibrationWriter(args.output_dir / "head", args.head_onnx)
with torch.no_grad():
for data in loader:
images = data["img"][0].data[0].cuda()
img_metas = data["img_metas"][0].data[0]
backbone_writer.write(get_backbone_inputs(args.version, model, images, img_metas))
if head_writer.count == 0:
features = model.extract_img_feat(images.clone(), img_metas)
head_writer.write(get_head_inputs(args.version, model, features, img_metas))
if backbone_writer.count != args.num_samples or head_writer.count != 1:
raise RuntimeError(
f"Prepared {backbone_writer.count} backbone and {head_writer.count} head batches; "
f"expected {args.num_samples} and 1"
)
print(
f"Saved {backbone_writer.count} backbone and one head calibration batch "
f"to {args.output_dir}"
)
if __name__ == "__main__":
main()
+133
View File
@@ -0,0 +1,133 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import re
from pathlib import Path
import numpy as np
import onnx
from onnxruntime.quantization.calibrate import CalibrationDataReader
__all__ = ["NpzCalibrationReader", "NpzCalibrationWriter", "find_vovnet_nodes_to_exclude"]
def _onnx_input_specs(onnx_path):
graph = onnx.load(onnx_path, load_external_data=False).graph
initializer_names = {initializer.name for initializer in graph.initializer}
input_specs = {}
for value in graph.input:
if value.name in initializer_names:
continue
tensor_type = value.type.tensor_type
shape = None
if tensor_type.HasField("shape"):
shape = tuple(
dimension.dim_value if dimension.HasField("dim_value") else None
for dimension in tensor_type.shape.dim
)
input_specs[value.name] = (
np.dtype(onnx.helper.tensor_dtype_to_np_dtype(tensor_type.elem_type)),
shape,
)
return input_specs
class NpzCalibrationWriter:
"""Write calibration batches that match an ONNX model's inputs."""
def __init__(self, output_dir, onnx_path):
self.output_dir = Path(output_dir)
self.output_dir.mkdir(parents=True, exist_ok=True)
if any(self.output_dir.glob("batch_*.npz")):
raise FileExistsError(f"{self.output_dir} already contains calibration batches")
self.input_specs = _onnx_input_specs(onnx_path)
self.count = 0
def write(self, values):
missing = self.input_specs.keys() - values.keys()
unexpected = values.keys() - self.input_specs.keys()
if missing or unexpected:
raise ValueError(
f"Calibration input mismatch; missing={sorted(missing)}, "
f"unexpected={sorted(unexpected)}"
)
batch = {}
for name, (dtype, expected_shape) in self.input_specs.items():
value = values[name]
if hasattr(value, "detach"):
value = value.detach().cpu().numpy()
value = np.asarray(value)
if expected_shape is not None and (
value.ndim != len(expected_shape)
or any(
expected is not None and actual != expected
for actual, expected in zip(value.shape, expected_shape)
)
):
raise ValueError(
f"Calibration input {name!r} has shape {value.shape}; expected {expected_shape}"
)
batch[name] = value.astype(dtype, copy=False)
np.savez(self.output_dir / f"batch_{self.count:04d}.npz", **batch)
self.count += 1
class NpzCalibrationReader(CalibrationDataReader):
"""Stream example-generated NPZ calibration batches."""
def __init__(self, calibration_dir):
self.batch_paths = sorted(Path(calibration_dir).glob("batch_*.npz"))
if not self.batch_paths:
raise ValueError(f"No calibration batches found in {calibration_dir}")
self.rewind()
@staticmethod
def load(batch_path):
with np.load(batch_path, allow_pickle=False) as batch:
return {name: batch[name] for name in batch.files}
def get_next(self):
batch_path = next(self._iterator, None)
return None if batch_path is None else self.load(batch_path)
def get_first(self):
return self.load(self.batch_paths[0])
def rewind(self):
self._iterator = iter(self.batch_paths)
def find_vovnet_nodes_to_exclude(onnx_path):
"""Find the VoVNet OSA4_5 stage and nodes downstream of FPN lateral_convs."""
# The evaluator image uses the calibration writer without installing ModelOpt.
from modelopt.onnx.utils import topologically_sort_graph_nodes
graph = onnx.load(onnx_path, load_external_data=False).graph
topologically_sort_graph_nodes(graph)
excluded = set()
downstream_tensors = set()
for node in graph.node:
is_osa = "OSA4_5" in node.name
is_downstream = any(name in downstream_tensors for name in node.input)
if is_osa or is_downstream:
excluded.add(node.name)
if "lateral_convs" in node.name or (is_downstream and not is_osa):
downstream_tensors.update(node.output)
if not excluded:
raise ValueError(f"No accuracy-sensitive VoVNet nodes found in {onnx_path}")
return [rf"^{re.escape(name)}$" for name in sorted(excluded)]
+72
View File
@@ -0,0 +1,72 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import shutil
import sys
import tempfile
from pathlib import Path
from modelopt.onnx.quantization import quantize
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationReader, find_vovnet_nodes_to_exclude
def parse_args():
parser = argparse.ArgumentParser(description="Quantize a VoVNet ONNX image encoder")
parser.add_argument("onnx_path")
parser.add_argument("calibration_dir", type=Path)
parser.add_argument("--precision", choices=("int8", "fp8"), default="int8")
parser.add_argument("--output")
return parser.parse_args()
def main():
args = parse_args()
onnx_path = Path(args.onnx_path)
output_path = args.output or onnx_path.with_name(
f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}"
)
excluded_nodes = find_vovnet_nodes_to_exclude(onnx_path)
print(f"Excluding {len(excluded_nodes)} accuracy-sensitive VoVNet nodes")
# Shape inference updates its input in place; a sibling copy preserves external-data paths.
temporary_file = tempfile.NamedTemporaryFile(
dir=onnx_path.parent,
prefix=f".{onnx_path.stem}.",
suffix=onnx_path.suffix,
delete=False,
)
temporary_onnx = Path(temporary_file.name)
temporary_file.close()
try:
shutil.copyfile(onnx_path, temporary_onnx)
quantize(
onnx_path=str(temporary_onnx),
quantize_mode=args.precision,
calibration_data_reader=NpzCalibrationReader(args.calibration_dir),
calibration_method="max",
calibration_eps=["cuda:0", "cpu"],
nodes_to_exclude=excluded_nodes,
high_precision_dtype="fp16",
output_path=str(output_path),
)
finally:
temporary_onnx.unlink(missing_ok=True)
if __name__ == "__main__":
main()
@@ -0,0 +1,6 @@
# Training and development dependencies are intentionally omitted for these legacy projects.
av2==0.2.1
lyft-dataset-sdk==0.0.8
mmdet3d==1.0.0rc6
nuscenes-devkit==1.1.11
refile==0.4.1
@@ -0,0 +1,19 @@
--find-links https://download.openmmlab.com/mmcv/dist/cu117/torch1.13.0/index.html
descartes==1.1.0
einops==0.8.1
flash-attn==0.2.8
kornia==0.6.12
mmcv-full==1.7.0
mmdet==2.28.2
mmsegmentation==0.30.0
numpy==1.23.5
onnx==1.17.0
onnx-graphsurgeon==0.6.1
onnxruntime==1.19.2
onnxsim==0.5.0
opencv-python==4.5.5.64
pyquaternion==0.9.9
Shapely==1.8.5.post1
trimesh==2.35.39
yapf==0.32.0
+151
View File
@@ -0,0 +1,151 @@
# Adapted from https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/far3d-trt/tools/test_tensorrt.py
# which was modified from https://github.com/megvii-research/Far3D/blob/5efb9d73a246c39fac79b3cf8c20a8e059611c3f/tools/test.py.
# Copyright (c) OpenMMLab. All rights reserved.
# Modified by Zhiqi Li.
#
# SPDX-FileCopyrightText: Copyright (c) 2023-2024, 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import math
import tensorrt as trt
import torch
__all__ = ["TensorRTRunner"]
TRT_TO_TORCH = {
trt.DataType.FLOAT: torch.float32,
trt.DataType.HALF: torch.float16,
trt.DataType.INT8: torch.int8,
trt.DataType.INT32: torch.int32,
trt.DataType.BOOL: torch.bool,
trt.DataType.UINT8: torch.uint8,
}
if int(trt.__version__.split(".")[0]) >= 10:
TRT_TO_TORCH[trt.DataType.INT64] = torch.int64
TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
trt.init_libnvinfer_plugins(TRT_LOGGER, "")
def aligned_tensor(shape, dtype, device, alignment=256):
element_size = torch.empty((), dtype=dtype).element_size()
element_count = math.prod(shape)
storage = torch.empty(element_count + alignment // element_size, dtype=dtype, device=device)
offset_bytes = (-storage.data_ptr()) % alignment
offset = offset_bytes // element_size
return storage[offset : offset + element_count].view(shape)
def _base_tensor_name(name):
return name.rsplit(".1", maxsplit=1)[0] if name.endswith(".1") else name
class TensorRTRunner:
def __init__(self, engine_path, state_names=()):
with open(engine_path, "rb") as engine_file:
engine_bytes = engine_file.read()
self.engine = trt.Runtime(TRT_LOGGER).deserialize_cuda_engine(engine_bytes)
if self.engine is None:
raise RuntimeError(f"Failed to deserialize {engine_path}")
self.tensor_names = [
self.engine.get_tensor_name(index) for index in range(self.engine.num_io_tensors)
]
self.input_shapes = {}
self.output_shapes = {}
self.tensor_dtypes = {}
for name in self.tensor_names:
shape = tuple(self.engine.get_tensor_shape(name))
dtype = TRT_TO_TORCH[self.engine.get_tensor_dtype(name)]
self.tensor_dtypes[name] = dtype
if self.engine.get_tensor_mode(name) == trt.TensorIOMode.INPUT:
self.input_shapes[name] = shape
else:
self.output_shapes[name] = shape
self._create_context(state_names)
def new_context(self, state_names=()):
runner = object.__new__(type(self))
runner.engine = self.engine
runner.tensor_names = self.tensor_names
runner.input_shapes = self.input_shapes
runner.output_shapes = self.output_shapes
runner.tensor_dtypes = self.tensor_dtypes
runner._create_context(state_names)
return runner
def _create_context(self, state_names):
self.context = self.engine.create_execution_context()
if self.context is None:
raise RuntimeError("Failed to create a TensorRT execution context")
self.state = {}
for base_name in state_names:
name = self.resolve_name(base_name)
if name in self.input_shapes:
tensor = aligned_tensor(self.input_shapes[name], self.tensor_dtypes[name], "cuda")
tensor.zero_()
self.state[name] = tensor
self.context.set_tensor_address(name, tensor.data_ptr())
if self.state:
torch.cuda.synchronize()
def resolve_name(self, base_name):
if base_name in self.tensor_names:
return base_name
suffixed_name = f"{base_name}.1"
return suffixed_name if suffixed_name in self.tensor_names else base_name
def reset_state(self):
for tensor in self.state.values():
tensor.zero_()
def prepare_input(self, name, inputs):
input_key = name if name in inputs else _base_tensor_name(name)
if input_key not in inputs:
raise KeyError(f"Missing TensorRT input {name}")
shape = self.input_shapes[name]
value = inputs[input_key].to(device="cuda", dtype=self.tensor_dtypes[name])
if tuple(value.shape) != shape:
if tuple(value.shape[1:]) == shape:
value = value.squeeze(0)
elif tuple(shape[1:]) == tuple(value.shape):
value = value.unsqueeze(0)
else:
raise ValueError(
f"Input {input_key} has shape {tuple(value.shape)}, expected {shape}"
)
return value
def __call__(self, stream, **inputs):
input_buffers = []
for name, shape in self.input_shapes.items():
if name in self.state:
continue
value = self.prepare_input(name, inputs)
buffer = aligned_tensor(shape, value.dtype, value.device)
buffer.copy_(value)
input_buffers.append(buffer)
self.context.set_tensor_address(name, buffer.data_ptr())
outputs = {}
for name, shape in self.output_shapes.items():
output = aligned_tensor(shape, self.tensor_dtypes[name], "cuda")
outputs[name] = output
self.context.set_tensor_address(name, output.data_ptr())
if not self.context.execute_async_v3(stream.cuda_stream):
raise RuntimeError("TensorRT execution failed")
return outputs
@@ -0,0 +1,70 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import torch
from examples.onnx_ptq.petr.petr_utils import run_backbone
class FakeBackbone:
def __init__(self, name, events, outputs):
self.name = name
self.events = events
self.outputs = outputs
def __call__(self, stream, **inputs):
self.events.append((self.name, inputs))
return self.outputs
def test_run_backbone_v1_uses_one_context():
events = []
outputs = {"out.0": torch.ones(1)}
backbone = FakeBackbone("current", events, outputs)
images = torch.arange(6).reshape(1, 6, 1)
result = run_backbone("v1", backbone, None, None, images)
assert result is outputs
assert [name for name, _ in events] == ["current"]
assert set(events[0][1]) == {"img"}
torch.testing.assert_close(events[0][1]["img"], images.squeeze(0))
def test_run_backbone_v2_uses_history_features_from_matching_context():
events = []
history_outputs = {
"out.0": torch.arange(12).reshape(1, 12, 1),
"out.1": torch.arange(100, 112).reshape(1, 12, 1),
}
current_outputs = {
"out.0": torch.full((1, 12, 1), 200),
"out.1": torch.full((1, 12, 1), 300),
}
history_backbone = FakeBackbone("history", events, history_outputs)
backbone = FakeBackbone("current", events, current_outputs)
images = torch.arange(12).reshape(1, 12, 1)
result = run_backbone("v2", backbone, history_backbone, None, images)
assert result is current_outputs
assert [name for name, _ in events] == ["history", "current"]
assert set(events[0][1]) == {"img"}
torch.testing.assert_close(events[0][1]["img"], images[:, 6:12].squeeze(0))
torch.testing.assert_close(events[1][1]["img"], images[:, :6].squeeze(0))
for index, name in enumerate(("out.0", "out.1")):
previous = events[1][1][f"prev.{index}"]
torch.testing.assert_close(previous, history_outputs[name][:, :6])
assert previous.is_contiguous()
@@ -0,0 +1,216 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
from pathlib import Path
from types import SimpleNamespace
import numpy as np
import onnx
import pytest
from onnx import TensorProto, helper, numpy_helper
from examples.onnx_ptq import quantize_vovnet
from examples.onnx_ptq.quantization_utils import (
NpzCalibrationReader,
NpzCalibrationWriter,
find_vovnet_nodes_to_exclude,
)
def make_calibration_model(tmp_path, image_shape=(1, 2)):
model_path = tmp_path / "model.onnx"
inputs = [
helper.make_tensor_value_info("image", TensorProto.FLOAT, image_shape),
helper.make_tensor_value_info("index", TensorProto.INT64, (1,)),
]
outputs = [
helper.make_tensor_value_info("image_out", TensorProto.FLOAT, image_shape),
helper.make_tensor_value_info("index_out", TensorProto.INT64, (1,)),
]
graph = helper.make_graph(
[
helper.make_node("Identity", ["image"], ["image_out"]),
helper.make_node("Identity", ["index"], ["index_out"]),
],
"calibration",
inputs,
outputs,
)
onnx.save(helper.make_model(graph), model_path)
return model_path
def test_npz_calibration_round_trip_and_rewind(tmp_path):
model_path = make_calibration_model(tmp_path)
writer = NpzCalibrationWriter(tmp_path / "batches", model_path)
writer.write(
{
"image": np.array([[1, 2]], dtype=np.float64),
"index": np.array([3], dtype=np.int32),
}
)
writer.write(
{
"image": np.array([[4, 5]], dtype=np.float64),
"index": np.array([6], dtype=np.int32),
}
)
assert writer.count == 2
assert [path.name for path in sorted((tmp_path / "batches").glob("*.npz"))] == [
"batch_0000.npz",
"batch_0001.npz",
]
reader = NpzCalibrationReader(tmp_path / "batches")
first = reader.get_first()
assert first["image"].dtype == np.float32
assert first["index"].dtype == np.int64
np.testing.assert_array_equal(reader.get_next()["image"], [[1, 2]])
np.testing.assert_array_equal(reader.get_next()["image"], [[4, 5]])
assert reader.get_next() is None
reader.rewind()
np.testing.assert_array_equal(reader.get_next()["image"], [[1, 2]])
with pytest.raises(FileExistsError, match="already contains"):
NpzCalibrationWriter(tmp_path / "batches", model_path)
@pytest.mark.parametrize(
"values",
[
{"image": np.ones((1, 2))},
{
"image": np.ones((1, 2)),
"index": np.ones((1,)),
"unexpected": np.ones((1,)),
},
],
)
def test_npz_writer_rejects_wrong_input_names(tmp_path, values):
writer = NpzCalibrationWriter(tmp_path / "batches", make_calibration_model(tmp_path))
with pytest.raises(ValueError, match="Calibration input mismatch"):
writer.write(values)
@pytest.mark.parametrize(
"image_shape",
[(2,), (1, 3)],
ids=("wrong-rank", "wrong-static-dimension"),
)
def test_npz_writer_rejects_wrong_input_shape(tmp_path, image_shape):
writer = NpzCalibrationWriter(tmp_path / "batches", make_calibration_model(tmp_path))
with pytest.raises(ValueError, match="Calibration input 'image' has shape"):
writer.write(
{
"image": np.ones(image_shape),
"index": np.ones((1,)),
}
)
def test_npz_writer_accepts_dynamic_input_shape(tmp_path):
writer = NpzCalibrationWriter(
tmp_path / "batches", make_calibration_model(tmp_path, image_shape=("batch", 2))
)
writer.write(
{
"image": np.ones((3, 2)),
"index": np.ones((1,)),
}
)
assert NpzCalibrationReader(tmp_path / "batches").get_first()["image"].shape == (3, 2)
def test_find_vovnet_nodes_to_exclude(tmp_path):
model_path = tmp_path / "vovnet.onnx"
nodes = [
helper.make_node("Identity", ["branch_out"], ["tail_out"], name="tail"),
helper.make_node("Identity", ["osa_out"], ["osa_tail_out"], name="osa_tail"),
helper.make_node("Identity", ["input"], ["osa_out"], name="backbone.OSA4_5"),
helper.make_node("Identity", ["lateral_out"], ["branch_out"], name="branch"),
helper.make_node("Identity", ["input"], ["lateral_out"], name="neck.lateral_convs.0"),
]
graph = helper.make_graph(
nodes,
"vovnet",
[helper.make_tensor_value_info("input", TensorProto.FLOAT, (1,))],
[
helper.make_tensor_value_info("tail_out", TensorProto.FLOAT, (1,)),
helper.make_tensor_value_info("osa_tail_out", TensorProto.FLOAT, (1,)),
],
)
onnx.save(helper.make_model(graph), model_path)
assert find_vovnet_nodes_to_exclude(model_path) == [
r"^backbone\.OSA4_5$",
r"^branch$",
r"^tail$",
]
def test_quantize_vovnet_preserves_source_model(tmp_path, monkeypatch):
model_path = tmp_path / "model.onnx"
graph = helper.make_graph(
[helper.make_node("Add", ["input", "weight"], ["output"])],
"external_data",
[helper.make_tensor_value_info("input", TensorProto.FLOAT, (1,))],
[helper.make_tensor_value_info("output", TensorProto.FLOAT, (1,))],
[numpy_helper.from_array(np.ones(1, dtype=np.float32), name="weight")],
)
onnx.save_model(
helper.make_model(graph),
model_path,
save_as_external_data=True,
all_tensors_to_one_file=True,
location="weights.bin",
size_threshold=0,
)
weights_path = tmp_path / "weights.bin"
source_bytes = model_path.read_bytes()
weight_bytes = weights_path.read_bytes()
temporary_paths = []
def fake_quantize(**kwargs):
temporary_path = Path(kwargs["onnx_path"])
onnx.load(temporary_path, load_external_data=True)
temporary_paths.append(temporary_path)
temporary_path.write_bytes(b"mutated")
Path(kwargs["output_path"]).write_bytes(b"quantized")
monkeypatch.chdir(tmp_path)
monkeypatch.setattr(
quantize_vovnet,
"parse_args",
lambda: SimpleNamespace(
onnx_path=str(model_path),
calibration_dir=tmp_path,
precision="int8",
output="quantized.onnx",
),
)
monkeypatch.setattr(quantize_vovnet, "find_vovnet_nodes_to_exclude", lambda _: [])
monkeypatch.setattr(quantize_vovnet, "NpzCalibrationReader", lambda _: object())
monkeypatch.setattr(quantize_vovnet, "quantize", fake_quantize)
quantize_vovnet.main()
assert model_path.read_bytes() == source_bytes
assert weights_path.read_bytes() == weight_bytes
assert not temporary_paths[0].exists()
assert (tmp_path / "quantized.onnx").read_bytes() == b"quantized"