mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do? Type of change: new example Adds an end-to-end BEVFormer-tiny ONNX PTQ example under `examples/onnx_ptq/bevformer` with: - The [BEVFormer Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile) from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second container definition in this repository. - A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10 source patch, and TensorRT plugin compilation at container runtime. - Ordered temporal calibration-data generation that propagates `prev_bev`, resets state at scene boundaries, computes CAN bus deltas, writes the exact requested sample count, and publishes output atomically. - INT8 and FP8 quantization with BEVFormer-specific calibration defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback policy. - Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus nuScenes evaluation instructions. - Shared temporary-ONNX-copy handling for BEVFormer and VoVNet quantization so shape inference cannot mutate the source model. - Focused CPU-only tests for temporal state, exact-count cleanup, quantization defaults, plugin configuration, and source-model preservation. The ONNX PTQ index continues to document the shared PETR/FAR3D containers separately and links to the BEVFormer guide. ### Usage Clone the pinned DL4AGX revision and build its BEVFormer image: ```bash git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX git -C /path/to/DL4AGX checkout --detach \ 9f7b29104c253d5bc68334e7b83b3eecb72d4572 docker build \ --build-arg TORCH_CUDA_ARCH_LIST=8.9 \ --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \ --tag modelopt-onnx-bevformer \ /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq ``` The example command targets compute capability 8.9; use the deployment GPU's compute capability for another architecture. After exporting the model and generating temporal calibration data, quantize it with: ```bash python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \ --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \ --calibration-dir=/artifacts/calibration \ --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \ --quantization-mode=fp8 \ --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx ``` See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin compilation, export, FP16 feedback-engine creation, temporal calibration, INT8/FP8 engine builds, and evaluation. ### Validation #### Current revision: three-sample smoke validation The Dockerfile at the pinned DL4AGX revision built successfully, and the current Model Optimizer checkout was mounted into it for the workflow below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch 2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage. A fresh three-frame A/B/B smoke passed without calculating partial NDS or mAP: - Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and FP8 engine builds passed. - Temporal calibration published exactly three batches with `use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent feedback on the third frame, and the expected CAN bus deltas. - The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero points. - Every engine produced finite outputs and 300 non-empty decoded detections for each frame; maximum scores ranged from 0.9489 to 0.9608. CPU-only validation passed the 16-test focused example set and the full ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the warning-as-error documentation build also passed. #### Historical full accuracy reference The following results were collected previously with TensorRT 10.14.1.48 on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are reference results, not a completed full evaluation of the current revision. | Precision | NDS | mAP | | :-- | --: | --: | | FP16 | 0.3546 | 0.2515 | | INT8/FP16 | 0.3512 | 0.2505 | | FP8/FP16 | 0.3526 | 0.2489 | #### Performance Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. The archived full-validation engines were benchmarked with five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream. | Precision | Speedup vs. FP16 | | :-- | --: | | FP16 | 1.00x | | INT8/FP16 | 1.87x | | FP8/FP16 | 1.18x | TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from another source or added a new PIP dependency, did you follow the guidance in `CONTRIBUTING.md`?: ✅ - Did you write the necessary tests?: ✅ - Did you update the [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional information Reference workflow: [NVIDIA DL4AGX BEVFormer INT8 example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq). > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end BEVFormer 3D detection workflow for ONNX post-training quantization, including temporal calibration, INT8/FP8 quantization, TensorRT engine generation, and nuScenes evaluation. * Added command-line tools for preparing calibration data and quantizing BEVFormer models. * **Bug Fixes** * Prevented source ONNX models from being overwritten during quantization and improved temporary model cleanup. * Fixed FP8 export for BF16 models during real-weight compression. * **Documentation** * Added setup, usage, compatibility, and performance guidance for the BEVFormer workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
This commit is contained in:
@@ -162,6 +162,10 @@ Inference latency of the model is <X> ms
|
||||
|
||||
The [FAR3D example](./far3d/) exports and quantizes the FAR3D ONNX image encoder, builds TensorRT engines, and evaluates 3D object detection mAP on the Argoverse 2 validation set.
|
||||
|
||||
### BEVFormer 3D object detection
|
||||
|
||||
The [BEVFormer example](./bevformer/) exports BEVFormer-tiny to ONNX, generates temporal calibration data, quantizes the model to INT8 or FP8, builds TensorRT engines, and evaluates NDS and mAP on the nuScenes validation set.
|
||||
|
||||
### PETR 3D object detection
|
||||
|
||||
The [PETR example](./petr/) exports and quantizes the PETRv1 and PETRv2 ONNX backbones, builds TensorRT engines, and evaluates 3D object detection mAP on the nuScenes validation set.
|
||||
|
||||
@@ -0,0 +1,176 @@
|
||||
# BEVFormer ONNX PTQ and nuScenes evaluation
|
||||
|
||||
This example extends the [DL4AGX BEVFormer workflow at source revision `9f7b291`](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq) with temporal calibration and FP8 quantization. Build the TensorRT 10.14 environment with its [Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile). The TensorRT plugins are compiled after the container starts so they target the GPU used for deployment.
|
||||
|
||||
## 1. Build and run the container
|
||||
|
||||
Clone the pinned DL4AGX revision and build its BEVFormer image. Set `MODELOPT_ROOT` to this Model Optimizer checkout so the container can use the example scripts and current Model Optimizer source.
|
||||
|
||||
```bash
|
||||
export MODELOPT_ROOT=$(git rev-parse --show-toplevel)
|
||||
export DL4AGX_ROOT=/path/to/DL4AGX
|
||||
git clone https://github.com/NVIDIA/DL4AGX.git ${DL4AGX_ROOT}
|
||||
git -C ${DL4AGX_ROOT} checkout --detach \
|
||||
9f7b29104c253d5bc68334e7b83b3eecb72d4572
|
||||
export DL4AGX_EXAMPLE=${DL4AGX_ROOT}/AV-Solutions/bevformer-int8-eq
|
||||
docker build \
|
||||
--build-arg TORCH_CUDA_ARCH_LIST=8.9 \
|
||||
--file ${DL4AGX_EXAMPLE}/docker/tensorrt.Dockerfile \
|
||||
--tag modelopt-onnx-bevformer \
|
||||
${DL4AGX_EXAMPLE}
|
||||
```
|
||||
|
||||
This build targets compute capability 8.9, matching the NVIDIA RTX 6000 Ada Generation validation system. Set `TORCH_CUDA_ARCH_LIST` to the compute capability of the deployment GPU when using another architecture.
|
||||
|
||||
Download nuScenes v1.0 trainval and CAN bus expansion data under the [nuScenes terms of use](https://www.nuscenes.org/terms-of-use). Create a host directory outside the Model Optimizer checkout for generated artifacts, then start the container with the dataset and artifact directories mounted:
|
||||
|
||||
```bash
|
||||
export BEVFORMER_ARTIFACTS=/path/to/bevformer_artifacts
|
||||
mkdir -p "${BEVFORMER_ARTIFACTS}"
|
||||
docker run --rm -it --gpus=all --shm-size=20g \
|
||||
-e PYTHONPATH=/opt/Model-Optimizer \
|
||||
-v "${MODELOPT_ROOT}:/opt/Model-Optimizer:ro" \
|
||||
-v "${DL4AGX_EXAMPLE}:/mnt/dl4agx:ro" \
|
||||
-v "${BEVFORMER_ARTIFACTS}:/artifacts" \
|
||||
-v /path/to/nuscenes:/workspace/BEVFormer_tensorrt/data/nuscenes \
|
||||
-v /path/to/can_bus:/workspace/BEVFormer_tensorrt/data/can_bus \
|
||||
modelopt-onnx-bevformer
|
||||
```
|
||||
|
||||
The DL4AGX image provides and owns the BEVFormer dependency stack. The read-only Model Optimizer mount and `PYTHONPATH` ensure the commands below use this checkout rather than the Model Optimizer release installed by the image.
|
||||
|
||||
Inside the container, set the paths used by the remaining steps:
|
||||
|
||||
```bash
|
||||
export BEVFORMER_ROOT=/workspace/BEVFormer_tensorrt
|
||||
export MODELOPT_ROOT=/opt/Model-Optimizer
|
||||
export ARTIFACTS=/artifacts
|
||||
export CONFIG=${BEVFORMER_ROOT}/configs/bevformer/plugin/bevformer_tiny_trt_p2.py
|
||||
export ONNX_PATH=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.onnx
|
||||
export FP16_ONNX=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.fp16.onnx
|
||||
export PLUGIN_PATH=${BEVFORMER_ROOT}/TensorRT/lib/libtensorrt_ops.so
|
||||
export FP16_ENGINE=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.fp16.engine
|
||||
```
|
||||
|
||||
Treat `${CONFIG}`, `${PLUGIN_PATH}`, and `${ARTIFACTS}/calibration` as trusted local inputs. MMCV executes the Python configuration, TensorRT loads the native plugin, and quantization reads every matching calibration batch. Use only the configuration and locally compiled plugin from the checked-out source revision and the calibration data produced by this workflow.
|
||||
|
||||
Compile the TensorRT plugins for the deployment GPU and prepare the nuScenes metadata:
|
||||
|
||||
```bash
|
||||
cd ${BEVFORMER_ROOT}
|
||||
git apply --check /mnt/dl4agx/bevformer_trt10.patch
|
||||
git apply /mnt/dl4agx/bevformer_trt10.patch
|
||||
cmake -S ${BEVFORMER_ROOT}/TensorRT -B ${BEVFORMER_ROOT}/TensorRT/build \
|
||||
-DCMAKE_TENSORRT_PATH=/usr
|
||||
cmake --build ${BEVFORMER_ROOT}/TensorRT/build --parallel
|
||||
cmake --install ${BEVFORMER_ROOT}/TensorRT/build
|
||||
bash samples/bevformer/create_data.sh
|
||||
```
|
||||
|
||||
The source repositories, patches, checkpoint, and dataset retain their upstream licenses and terms.
|
||||
|
||||
## 2. Export and build the FP16 engine
|
||||
|
||||
Download and verify the BEVFormer-tiny checkpoint, then export the original ONNX model:
|
||||
|
||||
```bash
|
||||
wget --continue --directory-prefix="${ARTIFACTS}" \
|
||||
https://github.com/zhiqi-li/storage/releases/download/v1.0/bevformer_tiny_epoch_24.pth
|
||||
echo "7305046dbaa4fe8b1fa6d6acb9e0e3d605a70a3c473f763e936103428d2b2f12 ${ARTIFACTS}/bevformer_tiny_epoch_24.pth" \
|
||||
| sha256sum --check
|
||||
cd ${BEVFORMER_ROOT}
|
||||
python tools/pth2onnx.py ${CONFIG} ${ARTIFACTS}/bevformer_tiny_epoch_24.pth \
|
||||
--opset_version=13 --cuda --flag=cp2_op13
|
||||
cp checkpoints/onnx/bevformer_tiny_epoch_24_cp2_op13.onnx ${ONNX_PATH}
|
||||
```
|
||||
|
||||
Use ModelOpt AutoCast to create a typed mixed FP16/FP32 graph. Build its strongly typed TensorRT engine for temporal feedback and the FP16 accuracy baseline:
|
||||
|
||||
```bash
|
||||
python -m modelopt.onnx.autocast \
|
||||
--onnx_path=${ONNX_PATH} \
|
||||
--output_path=${FP16_ONNX} \
|
||||
--low_precision_type=fp16 \
|
||||
--keep_io_types \
|
||||
--providers trt cuda:0 cpu \
|
||||
--trt_plugins ${PLUGIN_PATH}
|
||||
trtexec --onnx=${FP16_ONNX} \
|
||||
--saveEngine=${FP16_ENGINE} \
|
||||
--staticPlugins=${PLUGIN_PATH} \
|
||||
--stronglyTyped \
|
||||
--skipInference
|
||||
```
|
||||
|
||||
## 3. Generate temporal calibration data
|
||||
|
||||
Generate exactly 600 ordered training samples. Each saved input name, shape, and dtype is checked against the original ONNX model, while the FP16 engine propagates `prev_bev`. The generator resets `prev_bev` and CAN bus deltas at scene boundaries and publishes the calibration directory only after all requested samples complete.
|
||||
|
||||
```bash
|
||||
cd ${BEVFORMER_ROOT}
|
||||
PYTHONPATH=${MODELOPT_ROOT}:${BEVFORMER_ROOT} python \
|
||||
${MODELOPT_ROOT}/examples/onnx_ptq/bevformer/prepare_calibration.py \
|
||||
${CONFIG} \
|
||||
--onnx=${ONNX_PATH} \
|
||||
--engine=${FP16_ENGINE} \
|
||||
--trt-plugin=${PLUGIN_PATH} \
|
||||
--output-dir=${ARTIFACTS}/calibration \
|
||||
--num-samples=600
|
||||
```
|
||||
|
||||
Use a smaller `--num-samples` only for a smoke test. Accuracy results should use all 600 calibration samples.
|
||||
|
||||
## 4. Quantize and build INT8 and FP8 engines
|
||||
|
||||
INT8 uses entropy calibration by default; FP8 uses max calibration. Both modes preserve the source ONNX model and leave custom plugin operations and `MatMul` nodes in FP16.
|
||||
|
||||
```bash
|
||||
for precision in int8 fp8; do
|
||||
python ${MODELOPT_ROOT}/examples/onnx_ptq/bevformer/quantize.py \
|
||||
--onnx=${ONNX_PATH} \
|
||||
--calibration-dir=${ARTIFACTS}/calibration \
|
||||
--trt-plugins=${PLUGIN_PATH} \
|
||||
--quantization-mode=${precision} \
|
||||
--output=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.onnx
|
||||
trtexec --onnx=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.onnx \
|
||||
--saveEngine=${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.engine \
|
||||
--staticPlugins=${PLUGIN_PATH} \
|
||||
--stronglyTyped \
|
||||
--skipInference
|
||||
done
|
||||
```
|
||||
|
||||
## 5. Evaluate
|
||||
|
||||
Evaluate each engine inside the same image and on the same GPU architecture used for the build:
|
||||
|
||||
```bash
|
||||
cd ${BEVFORMER_ROOT}
|
||||
for precision in fp16 int8 fp8; do
|
||||
python tools/bevformer/evaluate_trt.py ${CONFIG} \
|
||||
${ARTIFACTS}/bevformer_tiny_epoch_24_cp2_op13.${precision}.engine \
|
||||
--trt_plugins=${PLUGIN_PATH} \
|
||||
| tee ${ARTIFACTS}/evaluate_${precision}.log
|
||||
done
|
||||
```
|
||||
|
||||
The following historical reference results were collected by the earlier TensorRT 10.14 workflow on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are retained as full-validation baselines; the DL4AGX-container integration was validated with three samples and did not rerun NDS or mAP.
|
||||
|
||||
| Precision | NDS | mAP |
|
||||
| :-- | --: | --: |
|
||||
| FP16 | 0.3546 | 0.2515 |
|
||||
| INT8/FP16 | 0.3512 | 0.2505 |
|
||||
| FP8/FP16 | 0.3526 | 0.2489 |
|
||||
|
||||
The accuracy run must complete all 6,019 validation samples. Serialized TensorRT engines are specific to the TensorRT version and GPU architecture used to build them.
|
||||
|
||||
## Performance
|
||||
|
||||
Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. These informational reference values use the archived full-validation engines and the median of five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream.
|
||||
|
||||
| Precision | Speedup vs. FP16 |
|
||||
| :-- | --: |
|
||||
| FP16 | 1.00x |
|
||||
| INT8/FP16 | 1.87x |
|
||||
| FP8/FP16 | 1.18x |
|
||||
|
||||
TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny.
|
||||
@@ -0,0 +1,144 @@
|
||||
# Adapted from https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/tools/calib_data_prep.py.
|
||||
#
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
import argparse
|
||||
import ctypes
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from tempfile import TemporaryDirectory
|
||||
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
|
||||
|
||||
from examples.onnx_ptq.quantization_utils import NpzCalibrationWriter
|
||||
|
||||
|
||||
def parse_args(arguments=None):
|
||||
parser = argparse.ArgumentParser(description="Prepare BEVFormer calibration batches")
|
||||
parser.add_argument("config", help="Path to the BEVFormer TensorRT configuration")
|
||||
parser.add_argument("--onnx", required=True, type=Path, help="Original ONNX model")
|
||||
parser.add_argument("--engine", required=True, type=Path, help="FP16 TensorRT engine")
|
||||
parser.add_argument("--trt-plugin", required=True, type=Path)
|
||||
parser.add_argument("--output-dir", required=True, type=Path)
|
||||
parser.add_argument("--num-samples", type=int, default=600)
|
||||
parser.add_argument("--workers", type=int, default=6)
|
||||
return parser.parse_args(arguments)
|
||||
|
||||
|
||||
def build_inputs(data, prev_bev, previous_frame):
|
||||
image = data["img"][0].data[0].numpy().astype(np.float32, copy=False)
|
||||
metadata = data["img_metas"][0].data[0][0]
|
||||
can_bus = metadata["can_bus"]
|
||||
same_scene = metadata["scene_token"] == previous_frame["scene_token"]
|
||||
position, angle = can_bus[:3].copy(), can_bus[-1].copy()
|
||||
if same_scene:
|
||||
can_bus[:3] -= previous_frame["position"]
|
||||
can_bus[-1] -= previous_frame["angle"]
|
||||
else:
|
||||
can_bus[:3] = 0
|
||||
can_bus[-1] = 0
|
||||
prev_bev = np.zeros_like(prev_bev)
|
||||
|
||||
previous_frame.update(scene_token=metadata["scene_token"], position=position, angle=angle)
|
||||
return {
|
||||
"image": image,
|
||||
"prev_bev": prev_bev,
|
||||
"use_prev_bev": np.array([same_scene], dtype=np.float32),
|
||||
"can_bus": can_bus.astype(np.float32),
|
||||
"lidar2img": np.stack(metadata["lidar2img"])[None].astype(np.float32),
|
||||
}
|
||||
|
||||
|
||||
def run_feedback(runner, stream, inputs, output_name) -> np.ndarray:
|
||||
with torch.cuda.stream(stream), torch.no_grad():
|
||||
device_inputs = {
|
||||
name: torch.from_numpy(value).cuda(non_blocking=True) for name, value in inputs.items()
|
||||
}
|
||||
outputs = runner(stream, **device_inputs)
|
||||
stream.synchronize()
|
||||
return outputs[output_name].detach().cpu().numpy()
|
||||
|
||||
|
||||
def prepare_batches(loader, runner, onnx_path, prev_bev_shape, output_dir, num_samples, stream):
|
||||
output_dir = Path(output_dir)
|
||||
output_dir.parent.mkdir(parents=True, exist_ok=True)
|
||||
if output_dir.exists():
|
||||
if not output_dir.is_dir() or any(output_dir.iterdir()):
|
||||
raise FileExistsError(f"{output_dir} must be empty")
|
||||
output_dir.rmdir()
|
||||
|
||||
with TemporaryDirectory(
|
||||
dir=output_dir.parent, prefix=f".{output_dir.name}.staging-"
|
||||
) as staging_dir:
|
||||
writer = NpzCalibrationWriter(staging_dir, onnx_path)
|
||||
prev_bev = np.zeros(prev_bev_shape, dtype=np.float32)
|
||||
previous_frame = {"scene_token": None, "position": 0, "angle": 0}
|
||||
output_name = runner.resolve_name("bev_embed")
|
||||
for data in loader:
|
||||
inputs = build_inputs(data, prev_bev, previous_frame)
|
||||
prev_bev = run_feedback(runner, stream, inputs, output_name)
|
||||
writer.write(inputs)
|
||||
if writer.count == num_samples:
|
||||
break
|
||||
|
||||
if writer.count != num_samples:
|
||||
raise RuntimeError(f"Prepared {writer.count} of {num_samples} requested samples")
|
||||
Path(staging_dir).replace(output_dir)
|
||||
return writer.count
|
||||
|
||||
|
||||
def main(arguments=None):
|
||||
args = parse_args(arguments)
|
||||
if args.num_samples < 1 or args.workers < 0:
|
||||
raise ValueError("Sample count must be positive and workers must be non-negative")
|
||||
for path in (args.onnx, args.engine, args.trt_plugin):
|
||||
if not path.is_file():
|
||||
raise FileNotFoundError(path)
|
||||
|
||||
# Keep optional BEVFormer dependencies out of CPU-only module imports.
|
||||
from mmcv import Config
|
||||
from third_party.bev_mmdet3d.datasets.builder import build_dataloader, build_dataset
|
||||
|
||||
config = Config.fromfile(args.config)
|
||||
loader = build_dataloader(
|
||||
build_dataset(cfg=config.data.quant),
|
||||
samples_per_gpu=1,
|
||||
workers_per_gpu=args.workers,
|
||||
shuffle=False,
|
||||
dist=False,
|
||||
)
|
||||
|
||||
_plugin = ctypes.CDLL(str(args.trt_plugin), mode=ctypes.RTLD_GLOBAL)
|
||||
from examples.onnx_ptq.trt_runner import TensorRTRunner
|
||||
|
||||
runner = TensorRTRunner(args.engine)
|
||||
saved = prepare_batches(
|
||||
loader,
|
||||
runner,
|
||||
args.onnx,
|
||||
(config.bev_h_ * config.bev_w_, 1, config._dim_),
|
||||
args.output_dir,
|
||||
args.num_samples,
|
||||
torch.cuda.Stream(),
|
||||
)
|
||||
print(f"Saved {saved} calibration batches to {args.output_dir}")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,68 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
|
||||
|
||||
from examples.onnx_ptq.quantization_utils import NpzCalibrationReader, temporary_onnx_copy
|
||||
from modelopt.onnx.quantization import quantize
|
||||
|
||||
|
||||
def parse_args(arguments=None):
|
||||
parser = argparse.ArgumentParser(description="Quantize the BEVFormer ONNX model")
|
||||
parser.add_argument("--onnx", required=True, type=Path)
|
||||
parser.add_argument("--calibration-dir", required=True, type=Path)
|
||||
parser.add_argument("--trt-plugins", required=True, nargs="+", type=Path)
|
||||
parser.add_argument("--quantization-mode", choices=("int8", "fp8"), default="int8")
|
||||
parser.add_argument("--calibration-method", choices=("entropy", "max"))
|
||||
parser.add_argument("--output", type=Path)
|
||||
return parser.parse_args(arguments)
|
||||
|
||||
|
||||
def main(arguments=None):
|
||||
args = parse_args(arguments)
|
||||
if not args.onnx.is_file():
|
||||
raise FileNotFoundError(args.onnx)
|
||||
for plugin in args.trt_plugins:
|
||||
if not plugin.is_file():
|
||||
raise FileNotFoundError(plugin)
|
||||
if args.output is None:
|
||||
args.output = args.onnx.with_name(f"{args.onnx.stem}.{args.quantization_mode}.onnx")
|
||||
if args.output.resolve() == args.onnx.resolve():
|
||||
raise ValueError("Output path must differ from the source ONNX path")
|
||||
|
||||
calibration_method = args.calibration_method or (
|
||||
"entropy" if args.quantization_mode == "int8" else "max"
|
||||
)
|
||||
with temporary_onnx_copy(args.onnx) as temporary_onnx:
|
||||
quantize(
|
||||
onnx_path=str(temporary_onnx),
|
||||
quantize_mode=args.quantization_mode,
|
||||
calibration_data_reader=NpzCalibrationReader(args.calibration_dir),
|
||||
calibration_method=calibration_method,
|
||||
calibration_eps=["trt", "cuda:0", "cpu"],
|
||||
op_types_to_exclude=["MatMul"],
|
||||
disable_mha_qdq=True,
|
||||
trt_plugins=[str(plugin) for plugin in args.trt_plugins],
|
||||
high_precision_dtype="fp16",
|
||||
output_path=str(args.output),
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -13,14 +13,41 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
import contextlib
|
||||
import re
|
||||
import shutil
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import onnx
|
||||
from onnxruntime.quantization.calibrate import CalibrationDataReader
|
||||
|
||||
__all__ = ["NpzCalibrationReader", "NpzCalibrationWriter", "find_vovnet_nodes_to_exclude"]
|
||||
__all__ = [
|
||||
"NpzCalibrationReader",
|
||||
"NpzCalibrationWriter",
|
||||
"find_vovnet_nodes_to_exclude",
|
||||
"temporary_onnx_copy",
|
||||
]
|
||||
|
||||
|
||||
@contextlib.contextmanager
|
||||
def temporary_onnx_copy(onnx_path):
|
||||
"""Yield a sibling copy so relative external-data paths remain valid."""
|
||||
onnx_path = Path(onnx_path)
|
||||
temporary_file = tempfile.NamedTemporaryFile(
|
||||
dir=onnx_path.parent,
|
||||
prefix=f".{onnx_path.stem}.",
|
||||
suffix=onnx_path.suffix,
|
||||
delete=False,
|
||||
)
|
||||
temporary_path = Path(temporary_file.name)
|
||||
temporary_file.close()
|
||||
try:
|
||||
shutil.copyfile(onnx_path, temporary_path)
|
||||
yield temporary_path
|
||||
finally:
|
||||
temporary_path.unlink(missing_ok=True)
|
||||
|
||||
|
||||
def _onnx_input_specs(onnx_path):
|
||||
|
||||
@@ -14,16 +14,18 @@
|
||||
# limitations under the License.
|
||||
|
||||
import argparse
|
||||
import shutil
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
|
||||
from modelopt.onnx.quantization import quantize
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
|
||||
|
||||
from examples.onnx_ptq.quantization_utils import NpzCalibrationReader, find_vovnet_nodes_to_exclude
|
||||
from examples.onnx_ptq.quantization_utils import (
|
||||
NpzCalibrationReader,
|
||||
find_vovnet_nodes_to_exclude,
|
||||
temporary_onnx_copy,
|
||||
)
|
||||
|
||||
|
||||
def parse_args():
|
||||
@@ -38,22 +40,16 @@ def parse_args():
|
||||
def main():
|
||||
args = parse_args()
|
||||
onnx_path = Path(args.onnx_path)
|
||||
output_path = args.output or onnx_path.with_name(
|
||||
f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}"
|
||||
output_path = (
|
||||
Path(args.output)
|
||||
if args.output
|
||||
else onnx_path.with_name(f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}")
|
||||
)
|
||||
if output_path.resolve() == onnx_path.resolve():
|
||||
raise ValueError("Output path must differ from the source ONNX path")
|
||||
excluded_nodes = find_vovnet_nodes_to_exclude(onnx_path)
|
||||
print(f"Excluding {len(excluded_nodes)} accuracy-sensitive VoVNet nodes")
|
||||
# Shape inference updates its input in place; a sibling copy preserves external-data paths.
|
||||
temporary_file = tempfile.NamedTemporaryFile(
|
||||
dir=onnx_path.parent,
|
||||
prefix=f".{onnx_path.stem}.",
|
||||
suffix=onnx_path.suffix,
|
||||
delete=False,
|
||||
)
|
||||
temporary_onnx = Path(temporary_file.name)
|
||||
temporary_file.close()
|
||||
try:
|
||||
shutil.copyfile(onnx_path, temporary_onnx)
|
||||
with temporary_onnx_copy(onnx_path) as temporary_onnx:
|
||||
quantize(
|
||||
onnx_path=str(temporary_onnx),
|
||||
quantize_mode=args.precision,
|
||||
@@ -64,8 +60,6 @@ def main():
|
||||
high_precision_dtype="fp16",
|
||||
output_path=str(output_path),
|
||||
)
|
||||
finally:
|
||||
temporary_onnx.unlink(missing_ok=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
|
||||
Reference in New Issue
Block a user