[OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)

### What does this PR do?

Type of change: new example

Adds an end-to-end BEVFormer-tiny ONNX PTQ example under
`examples/onnx_ptq/bevformer` with:

- The [BEVFormer
Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile)
from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second
container definition in this repository.
- A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10
source patch, and TensorRT plugin compilation at container runtime.
- Ordered temporal calibration-data generation that propagates
`prev_bev`, resets state at scene boundaries, computes CAN bus deltas,
writes the exact requested sample count, and publishes output
atomically.
- INT8 and FP8 quantization with BEVFormer-specific calibration
defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback
policy.
- Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus
nuScenes evaluation instructions.
- Shared temporary-ONNX-copy handling for BEVFormer and VoVNet
quantization so shape inference cannot mutate the source model.
- Focused CPU-only tests for temporal state, exact-count cleanup,
quantization defaults, plugin configuration, and source-model
preservation.

The ONNX PTQ index continues to document the shared PETR/FAR3D
containers separately and links to the BEVFormer guide.

### Usage

Clone the pinned DL4AGX revision and build its BEVFormer image:

```bash
git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX
git -C /path/to/DL4AGX checkout --detach \
  9f7b29104c253d5bc68334e7b83b3eecb72d4572
docker build \
  --build-arg TORCH_CUDA_ARCH_LIST=8.9 \
  --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \
  --tag modelopt-onnx-bevformer \
  /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq
```

The example command targets compute capability 8.9; use the deployment
GPU's compute capability for another architecture.

After exporting the model and generating temporal calibration data,
quantize it with:

```bash
python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \
  --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \
  --calibration-dir=/artifacts/calibration \
  --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \
  --quantization-mode=fp8 \
  --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx
```

See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin
compilation, export, FP16 feedback-engine creation, temporal
calibration, INT8/FP8 engine builds, and evaluation.

### Validation

#### Current revision: three-sample smoke validation

The Dockerfile at the pinned DL4AGX revision built successfully, and the
current Model Optimizer checkout was mounted into it for the workflow
below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch
2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage.

A fresh three-frame A/B/B smoke passed without calculating partial NDS
or mAP:

- Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and
FP8 engine builds passed.
- Temporal calibration published exactly three batches with
`use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent
feedback on the third frame, and the expected CAN bus deltas.
- The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs
with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero
points.
- Every engine produced finite outputs and 300 non-empty decoded
detections for each frame; maximum scores ranged from 0.9489 to 0.9608.

CPU-only validation passed the 16-test focused example set and the full
ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the
warning-as-error documentation build also passed.

#### Historical full accuracy reference

The following results were collected previously with TensorRT 10.14.1.48
on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration
samples and all 6,019 nuScenes validation samples. They are reference
results, not a completed full evaluation of the current revision.

| Precision | NDS | mAP |
| :-- | --: | --: |
| FP16 | 0.3546 | 0.2515 |
| INT8/FP16 | 0.3512 | 0.2505 |
| FP8/FP16 | 0.3526 | 0.2489 |

#### Performance

Following the PETR and FAR3D convention, performance is reported only as
speedup normalized to FP16. The archived full-validation engines were
benchmarked with five interleaved TensorRT 10.14 trials on the same
NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU
Compute Time median with data transfers disabled, CUDA Graphs enabled,
spin-wait, a 1-second warmup, a 10-second measurement window, and one
inference stream.

| Precision | Speedup vs. FP16 |
| :-- | --: |
| FP16 | 1.00x |
| INT8/FP16 | 1.87x |
| FP8/FP16 | 1.18x |

TODO: Investigate why FP8 delivers less speedup than INT8 for
BEVFormer-tiny.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from another source or added a new PIP dependency,
did you follow the guidance in `CONTRIBUTING.md`?: ✅
- Did you write the necessary tests?: ✅
- Did you update the
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional information

Reference workflow: [NVIDIA DL4AGX BEVFormer INT8
example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq).

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end BEVFormer 3D detection workflow for ONNX
post-training quantization, including temporal calibration, INT8/FP8
quantization, TensorRT engine generation, and nuScenes evaluation.
* Added command-line tools for preparing calibration data and quantizing
BEVFormer models.
* **Bug Fixes**
* Prevented source ONNX models from being overwritten during
quantization and improved temporary model cleanup.
  * Fixed FP8 export for BF16 models during real-weight compression.
* **Documentation**
* Added setup, usage, compatibility, and performance guidance for the
BEVFormer workflow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
This commit is contained in:
Ajinkya Rasane
2026-09-11 16:15:59 -04:00
committed by GitHub
co-authored by Codex
parent b80e164472
commit bd90a5ed51
10 changed files with 606 additions and 19 deletions
+4
View File
@@ -162,6 +162,10 @@ Inference latency of the model is <X> ms
The [FAR3D example](./far3d/) exports and quantizes the FAR3D ONNX image encoder, builds TensorRT engines, and evaluates 3D object detection mAP on the Argoverse 2 validation set.
### BEVFormer 3D object detection
The [BEVFormer example](./bevformer/) exports BEVFormer-tiny to ONNX, generates temporal calibration data, quantizes the model to INT8 or FP8, builds TensorRT engines, and evaluates NDS and mAP on the nuScenes validation set.
### PETR 3D object detection
The [PETR example](./petr/) exports and quantizes the PETRv1 and PETRv2 ONNX backbones, builds TensorRT engines, and evaluates 3D object detection mAP on the nuScenes validation set.
+176
View File
@@ -0,0 +1,176 @@
# BEVFormer ONNX PTQ and nuScenes evaluation
This example extends the [DL4AGX BEVFormer workflow at source revision `9f7b291`](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq) with temporal calibration and FP8 quantization. Build the TensorRT 10.14 environment with its [Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile). The TensorRT plugins are compiled after the container starts so they target the GPU used for deployment.
## 1. Build and run the container
Clone the pinned DL4AGX revision and build its BEVFormer image. Set `MODELOPT_ROOT` to this Model Optimizer checkout so the container can use the example scripts and current Model Optimizer source.
```bash
export MODELOPT_ROOT=$(git rev-parse --show-toplevel)
export DL4AGX_ROOT=/path/to/DL4AGX
git clone https://github.com/NVIDIA/DL4AGX.git ${DL4AGX_ROOT}
git -C ${DL4AGX_ROOT} checkout --detach \
9f7b29104c253d5bc68334e7b83b3eecb72d4572
export DL4AGX_EXAMPLE=${DL4AGX_ROOT}/AV-Solutions/bevformer-int8-eq
docker build \
--build-arg TORCH_CUDA_ARCH_LIST=8.9 \
--file ${DL4AGX_EXAMPLE}/docker/tensorrt.Dockerfile \
--tag modelopt-onnx-bevformer \
${DL4AGX_EXAMPLE}
```
This build targets compute capability 8.9, matching the NVIDIA RTX 6000 Ada Generation validation system. Set `TORCH_CUDA_ARCH_LIST` to the compute capability of the deployment GPU when using another architecture.
Download nuScenes v1.0 trainval and CAN bus expansion data under the [nuScenes terms of use](https://www.nuscenes.org/terms-of-use). Create a host directory outside the Model Optimizer checkout for generated artifacts, then start the container with the dataset and artifact directories mounted:
```bash
export BEVFORMER_ARTIFACTS=/path/to/bevformer_artifacts
mkdir -p "${BEVFORMER_ARTIFACTS}"
docker run --rm -it --gpus=all --shm-size=20g \
-e PYTHONPATH=/opt/Model-Optimizer \
-v "${MODELOPT_ROOT}:/opt/Model-Optimizer:ro" \
-v "${DL4AGX_EXAMPLE}:/mnt/dl4agx:ro" \
-v "${BEVFORMER_ARTIFACTS}:/artifacts" \
-v /path/to/nuscenes:/workspace/BEVFormer_tensorrt/data/nuscenes \
-v /path/to/can_bus:/workspace/BEVFormer_tensorrt/data/can_bus \
modelopt-onnx-bevformer
```
The DL4AGX image provides and owns the BEVFormer dependency stack. The read-only Model Optimizer mount and `PYTHONPATH` ensure the commands below use this checkout rather than the Model Optimizer release installed by the image.
Inside the container, set the paths used by the remaining steps:
```bash
export BEVFORMER_ROOT=/workspace/BEVFormer_tensorrt
export MODELOPT_ROOT=/opt/Model-Optimizer
export ARTIFACTS=/artifacts
export CONFIG=${BEVFORMER_ROOT}/configs/bevformer/plugin/bevformer_tiny_trt_p2.py
export ONNX_PATH=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.onnx
export FP16_ONNX=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.fp16.onnx
export PLUGIN_PATH=${BEVFORMER_ROOT}/TensorRT/lib/libtensorrt_ops.so
export FP16_ENGINE=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.fp16.engine
```
Treat `${CONFIG}`, `${PLUGIN_PATH}`, and `${ARTIFACTS}/calibration` as trusted local inputs. MMCV executes the Python configuration, TensorRT loads the native plugin, and quantization reads every matching calibration batch. Use only the configuration and locally compiled plugin from the checked-out source revision and the calibration data produced by this workflow.
Compile the TensorRT plugins for the deployment GPU and prepare the nuScenes metadata:
```bash
cd ${BEVFORMER_ROOT}
git apply --check /mnt/dl4agx/bevformer_trt10.patch
git apply /mnt/dl4agx/bevformer_trt10.patch
cmake -S ${BEVFORMER_ROOT}/TensorRT -B ${BEVFORMER_ROOT}/TensorRT/build \
-DCMAKE_TENSORRT_PATH=/usr
cmake --build ${BEVFORMER_ROOT}/TensorRT/build --parallel
cmake --install ${BEVFORMER_ROOT}/TensorRT/build
bash samples/bevformer/create_data.sh
```
The source repositories, patches, checkpoint, and dataset retain their upstream licenses and terms.
## 2. Export and build the FP16 engine
Download and verify the BEVFormer-tiny checkpoint, then export the original ONNX model:
```bash
wget --continue --directory-prefix="${ARTIFACTS}" \
https://github.com/zhiqi-li/storage/releases/download/v1.0/bevformer_tiny_epoch_24.pth
echo "7305046dbaa4fe8b1fa6d6acb9e0e3d605a70a3c473f763e936103428d2b2f12 ${ARTIFACTS}/bevformer_tiny_epoch_24.pth" \
| sha256sum --check
cd ${BEVFORMER_ROOT}
python tools/pth2onnx.py ${CONFIG} ${ARTIFACTS}/bevformer_tiny_epoch_24.pth \
--opset_version=13 --cuda --flag=cp2_op13
cp checkpoints/onnx/bevformer_tiny_epoch_24_cp2_op13.onnx ${ONNX_PATH}
```
Use ModelOpt AutoCast to create a typed mixed FP16/FP32 graph. Build its strongly typed TensorRT engine for temporal feedback and the FP16 accuracy baseline:
```bash
python -m modelopt.onnx.autocast \
--onnx_path=${ONNX_PATH} \
--output_path=${FP16_ONNX} \
--low_precision_type=fp16 \
--keep_io_types \
--providers trt cuda:0 cpu \
--trt_plugins ${PLUGIN_PATH}
trtexec --onnx=${FP16_ONNX} \
--saveEngine=${FP16_ENGINE} \
--staticPlugins=${PLUGIN_PATH} \
--stronglyTyped \
--skipInference
```
## 3. Generate temporal calibration data
Generate exactly 600 ordered training samples. Each saved input name, shape, and dtype is checked against the original ONNX model, while the FP16 engine propagates `prev_bev`. The generator resets `prev_bev` and CAN bus deltas at scene boundaries and publishes the calibration directory only after all requested samples complete.
```bash
cd ${BEVFORMER_ROOT}
PYTHONPATH=${MODELOPT_ROOT}:${BEVFORMER_ROOT} python \
${MODELOPT_ROOT}/examples/onnx_ptq/bevformer/prepare_calibration.py \
${CONFIG} \
--onnx=${ONNX_PATH} \
--engine=${FP16_ENGINE} \
--trt-plugin=${PLUGIN_PATH} \
--output-dir=${ARTIFACTS}/calibration \
--num-samples=600
```
Use a smaller `--num-samples` only for a smoke test. Accuracy results should use all 600 calibration samples.
## 4. Quantize and build INT8 and FP8 engines
INT8 uses entropy calibration by default; FP8 uses max calibration. Both modes preserve the source ONNX model and leave custom plugin operations and `MatMul` nodes in FP16.
```bash
for precision in int8 fp8; do
python ${MODELOPT_ROOT}/examples/onnx_ptq/bevformer/quantize.py \
--onnx=${ONNX_PATH} \
--calibration-dir=${ARTIFACTS}/calibration \
--trt-plugins=${PLUGIN_PATH} \
--quantization-mode=${precision} \
--output=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.onnx
trtexec --onnx=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.onnx \
--saveEngine=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.engine \
--staticPlugins=${PLUGIN_PATH} \
--stronglyTyped \
--skipInference
done
```
## 5. Evaluate
Evaluate each engine inside the same image and on the same GPU architecture used for the build:
```bash
cd ${BEVFORMER_ROOT}
for precision in fp16 int8 fp8; do
python tools/bevformer/evaluate_trt.py ${CONFIG} \
${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.engine \
--trt_plugins=${PLUGIN_PATH} \
| tee ${ARTIFACTS}/evaluate_${precision}.log
done
```
The following historical reference results were collected by the earlier TensorRT 10.14 workflow on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are retained as full-validation baselines; the DL4AGX-container integration was validated with three samples and did not rerun NDS or mAP.
| Precision | NDS | mAP |
| :-- | --: | --: |
| FP16 | 0.3546 | 0.2515 |
| INT8/FP16 | 0.3512 | 0.2505 |
| FP8/FP16 | 0.3526 | 0.2489 |
The accuracy run must complete all 6,019 validation samples. Serialized TensorRT engines are specific to the TensorRT version and GPU architecture used to build them.
## Performance
Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. These informational reference values use the archived full-validation engines and the median of five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream.
| Precision | Speedup vs. FP16 |
| :-- | --: |
| FP16 | 1.00x |
| INT8/FP16 | 1.87x |
| FP8/FP16 | 1.18x |
TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny.
@@ -0,0 +1,144 @@
# Adapted from https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/tools/calib_data_prep.py.
#
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import ctypes
import sys
from pathlib import Path
from tempfile import TemporaryDirectory
import numpy as np
import torch
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationWriter
def parse_args(arguments=None):
parser = argparse.ArgumentParser(description="Prepare BEVFormer calibration batches")
parser.add_argument("config", help="Path to the BEVFormer TensorRT configuration")
parser.add_argument("--onnx", required=True, type=Path, help="Original ONNX model")
parser.add_argument("--engine", required=True, type=Path, help="FP16 TensorRT engine")
parser.add_argument("--trt-plugin", required=True, type=Path)
parser.add_argument("--output-dir", required=True, type=Path)
parser.add_argument("--num-samples", type=int, default=600)
parser.add_argument("--workers", type=int, default=6)
return parser.parse_args(arguments)
def build_inputs(data, prev_bev, previous_frame):
image = data["img"][0].data[0].numpy().astype(np.float32, copy=False)
metadata = data["img_metas"][0].data[0][0]
can_bus = metadata["can_bus"]
same_scene = metadata["scene_token"] == previous_frame["scene_token"]
position, angle = can_bus[:3].copy(), can_bus[-1].copy()
if same_scene:
can_bus[:3] -= previous_frame["position"]
can_bus[-1] -= previous_frame["angle"]
else:
can_bus[:3] = 0
can_bus[-1] = 0
prev_bev = np.zeros_like(prev_bev)
previous_frame.update(scene_token=metadata["scene_token"], position=position, angle=angle)
return {
"image": image,
"prev_bev": prev_bev,
"use_prev_bev": np.array([same_scene], dtype=np.float32),
"can_bus": can_bus.astype(np.float32),
"lidar2img": np.stack(metadata["lidar2img"])[None].astype(np.float32),
}
def run_feedback(runner, stream, inputs, output_name) -> np.ndarray:
with torch.cuda.stream(stream), torch.no_grad():
device_inputs = {
name: torch.from_numpy(value).cuda(non_blocking=True) for name, value in inputs.items()
}
outputs = runner(stream, **device_inputs)
stream.synchronize()
return outputs[output_name].detach().cpu().numpy()
def prepare_batches(loader, runner, onnx_path, prev_bev_shape, output_dir, num_samples, stream):
output_dir = Path(output_dir)
output_dir.parent.mkdir(parents=True, exist_ok=True)
if output_dir.exists():
if not output_dir.is_dir() or any(output_dir.iterdir()):
raise FileExistsError(f"{output_dir} must be empty")
output_dir.rmdir()
with TemporaryDirectory(
dir=output_dir.parent, prefix=f".{output_dir.name}.staging-"
) as staging_dir:
writer = NpzCalibrationWriter(staging_dir, onnx_path)
prev_bev = np.zeros(prev_bev_shape, dtype=np.float32)
previous_frame = {"scene_token": None, "position": 0, "angle": 0}
output_name = runner.resolve_name("bev_embed")
for data in loader:
inputs = build_inputs(data, prev_bev, previous_frame)
prev_bev = run_feedback(runner, stream, inputs, output_name)
writer.write(inputs)
if writer.count == num_samples:
break
if writer.count != num_samples:
raise RuntimeError(f"Prepared {writer.count} of {num_samples} requested samples")
Path(staging_dir).replace(output_dir)
return writer.count
def main(arguments=None):
args = parse_args(arguments)
if args.num_samples < 1 or args.workers < 0:
raise ValueError("Sample count must be positive and workers must be non-negative")
for path in (args.onnx, args.engine, args.trt_plugin):
if not path.is_file():
raise FileNotFoundError(path)
# Keep optional BEVFormer dependencies out of CPU-only module imports.
from mmcv import Config
from third_party.bev_mmdet3d.datasets.builder import build_dataloader, build_dataset
config = Config.fromfile(args.config)
loader = build_dataloader(
build_dataset(cfg=config.data.quant),
samples_per_gpu=1,
workers_per_gpu=args.workers,
shuffle=False,
dist=False,
)
_plugin = ctypes.CDLL(str(args.trt_plugin), mode=ctypes.RTLD_GLOBAL)
from examples.onnx_ptq.trt_runner import TensorRTRunner
runner = TensorRTRunner(args.engine)
saved = prepare_batches(
loader,
runner,
args.onnx,
(config.bev_h_ * config.bev_w_, 1, config._dim_),
args.output_dir,
args.num_samples,
torch.cuda.Stream(),
)
print(f"Saved {saved} calibration batches to {args.output_dir}")
if __name__ == "__main__":
main()
+68
View File
@@ -0,0 +1,68 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationReader, temporary_onnx_copy
from modelopt.onnx.quantization import quantize
def parse_args(arguments=None):
parser = argparse.ArgumentParser(description="Quantize the BEVFormer ONNX model")
parser.add_argument("--onnx", required=True, type=Path)
parser.add_argument("--calibration-dir", required=True, type=Path)
parser.add_argument("--trt-plugins", required=True, nargs="+", type=Path)
parser.add_argument("--quantization-mode", choices=("int8", "fp8"), default="int8")
parser.add_argument("--calibration-method", choices=("entropy", "max"))
parser.add_argument("--output", type=Path)
return parser.parse_args(arguments)
def main(arguments=None):
args = parse_args(arguments)
if not args.onnx.is_file():
raise FileNotFoundError(args.onnx)
for plugin in args.trt_plugins:
if not plugin.is_file():
raise FileNotFoundError(plugin)
if args.output is None:
args.output = args.onnx.with_name(f"{args.onnx.stem}.{args.quantization_mode}.onnx")
if args.output.resolve() == args.onnx.resolve():
raise ValueError("Output path must differ from the source ONNX path")
calibration_method = args.calibration_method or (
"entropy" if args.quantization_mode == "int8" else "max"
)
with temporary_onnx_copy(args.onnx) as temporary_onnx:
quantize(
onnx_path=str(temporary_onnx),
quantize_mode=args.quantization_mode,
calibration_data_reader=NpzCalibrationReader(args.calibration_dir),
calibration_method=calibration_method,
calibration_eps=["trt", "cuda:0", "cpu"],
op_types_to_exclude=["MatMul"],
disable_mha_qdq=True,
trt_plugins=[str(plugin) for plugin in args.trt_plugins],
high_precision_dtype="fp16",
output_path=str(args.output),
)
if __name__ == "__main__":
main()
+28 -1
View File
@@ -13,14 +13,41 @@
# See the License for the specific language governing permissions and
# limitations under the License.
import contextlib
import re
import shutil
import tempfile
from pathlib import Path
import numpy as np
import onnx
from onnxruntime.quantization.calibrate import CalibrationDataReader
__all__ = ["NpzCalibrationReader", "NpzCalibrationWriter", "find_vovnet_nodes_to_exclude"]
__all__ = [
"NpzCalibrationReader",
"NpzCalibrationWriter",
"find_vovnet_nodes_to_exclude",
"temporary_onnx_copy",
]
@contextlib.contextmanager
def temporary_onnx_copy(onnx_path):
"""Yield a sibling copy so relative external-data paths remain valid."""
onnx_path = Path(onnx_path)
temporary_file = tempfile.NamedTemporaryFile(
dir=onnx_path.parent,
prefix=f".{onnx_path.stem}.",
suffix=onnx_path.suffix,
delete=False,
)
temporary_path = Path(temporary_file.name)
temporary_file.close()
try:
shutil.copyfile(onnx_path, temporary_path)
yield temporary_path
finally:
temporary_path.unlink(missing_ok=True)
def _onnx_input_specs(onnx_path):
+12 -18
View File
@@ -14,16 +14,18 @@
# limitations under the License.
import argparse
import shutil
import sys
import tempfile
from pathlib import Path
from modelopt.onnx.quantization import quantize
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
from examples.onnx_ptq.quantization_utils import NpzCalibrationReader, find_vovnet_nodes_to_exclude
from examples.onnx_ptq.quantization_utils import (
NpzCalibrationReader,
find_vovnet_nodes_to_exclude,
temporary_onnx_copy,
)
def parse_args():
@@ -38,22 +40,16 @@ def parse_args():
def main():
args = parse_args()
onnx_path = Path(args.onnx_path)
output_path = args.output or onnx_path.with_name(
f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}"
output_path = (
Path(args.output)
if args.output
else onnx_path.with_name(f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}")
)
if output_path.resolve() == onnx_path.resolve():
raise ValueError("Output path must differ from the source ONNX path")
excluded_nodes = find_vovnet_nodes_to_exclude(onnx_path)
print(f"Excluding {len(excluded_nodes)} accuracy-sensitive VoVNet nodes")
# Shape inference updates its input in place; a sibling copy preserves external-data paths.
temporary_file = tempfile.NamedTemporaryFile(
dir=onnx_path.parent,
prefix=f".{onnx_path.stem}.",
suffix=onnx_path.suffix,
delete=False,
)
temporary_onnx = Path(temporary_file.name)
temporary_file.close()
try:
shutil.copyfile(onnx_path, temporary_onnx)
with temporary_onnx_copy(onnx_path) as temporary_onnx:
quantize(
onnx_path=str(temporary_onnx),
quantize_mode=args.precision,
@@ -64,8 +60,6 @@ def main():
high_precision_dtype="fp16",
output_path=str(output_path),
)
finally:
temporary_onnx.unlink(missing_ok=True)
if __name__ == "__main__":