Files
Ajinkya RasaneandCodex bd90a5ed51 [OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do?

Type of change: new example

Adds an end-to-end BEVFormer-tiny ONNX PTQ example under
`examples/onnx_ptq/bevformer` with:

- The [BEVFormer
Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile)
from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second
container definition in this repository.
- A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10
source patch, and TensorRT plugin compilation at container runtime.
- Ordered temporal calibration-data generation that propagates
`prev_bev`, resets state at scene boundaries, computes CAN bus deltas,
writes the exact requested sample count, and publishes output
atomically.
- INT8 and FP8 quantization with BEVFormer-specific calibration
defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback
policy.
- Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus
nuScenes evaluation instructions.
- Shared temporary-ONNX-copy handling for BEVFormer and VoVNet
quantization so shape inference cannot mutate the source model.
- Focused CPU-only tests for temporal state, exact-count cleanup,
quantization defaults, plugin configuration, and source-model
preservation.

The ONNX PTQ index continues to document the shared PETR/FAR3D
containers separately and links to the BEVFormer guide.

### Usage

Clone the pinned DL4AGX revision and build its BEVFormer image:

```bash
git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX
git -C /path/to/DL4AGX checkout --detach \
  9f7b29104c253d5bc68334e7b83b3eecb72d4572
docker build \
  --build-arg TORCH_CUDA_ARCH_LIST=8.9 \
  --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \
  --tag modelopt-onnx-bevformer \
  /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq
```

The example command targets compute capability 8.9; use the deployment
GPU's compute capability for another architecture.

After exporting the model and generating temporal calibration data,
quantize it with:

```bash
python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \
  --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \
  --calibration-dir=/artifacts/calibration \
  --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \
  --quantization-mode=fp8 \
  --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx
```

See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin
compilation, export, FP16 feedback-engine creation, temporal
calibration, INT8/FP8 engine builds, and evaluation.

### Validation

#### Current revision: three-sample smoke validation

The Dockerfile at the pinned DL4AGX revision built successfully, and the
current Model Optimizer checkout was mounted into it for the workflow
below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch
2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage.

A fresh three-frame A/B/B smoke passed without calculating partial NDS
or mAP:

- Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and
FP8 engine builds passed.
- Temporal calibration published exactly three batches with
`use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent
feedback on the third frame, and the expected CAN bus deltas.
- The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs
with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero
points.
- Every engine produced finite outputs and 300 non-empty decoded
detections for each frame; maximum scores ranged from 0.9489 to 0.9608.

CPU-only validation passed the 16-test focused example set and the full
ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the
warning-as-error documentation build also passed.

#### Historical full accuracy reference

The following results were collected previously with TensorRT 10.14.1.48
on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration
samples and all 6,019 nuScenes validation samples. They are reference
results, not a completed full evaluation of the current revision.

| Precision | NDS | mAP |
| :-- | --: | --: |
| FP16 | 0.3546 | 0.2515 |
| INT8/FP16 | 0.3512 | 0.2505 |
| FP8/FP16 | 0.3526 | 0.2489 |

#### Performance

Following the PETR and FAR3D convention, performance is reported only as
speedup normalized to FP16. The archived full-validation engines were
benchmarked with five interleaved TensorRT 10.14 trials on the same
NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU
Compute Time median with data transfers disabled, CUDA Graphs enabled,
spin-wait, a 1-second warmup, a 10-second measurement window, and one
inference stream.

| Precision | Speedup vs. FP16 |
| :-- | --: |
| FP16 | 1.00x |
| INT8/FP16 | 1.87x |
| FP8/FP16 | 1.18x |

TODO: Investigate why FP8 delivers less speedup than INT8 for
BEVFormer-tiny.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from another source or added a new PIP dependency,
did you follow the guidance in `CONTRIBUTING.md`?: ✅
- Did you write the necessary tests?: ✅
- Did you update the
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional information

Reference workflow: [NVIDIA DL4AGX BEVFormer INT8
example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq).

> 🤖 _Generated by Codex (AI agent)._


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an end-to-end BEVFormer 3D detection workflow for ONNX
post-training quantization, including temporal calibration, INT8/FP8
quantization, TensorRT engine generation, and nuScenes evaluation.
* Added command-line tools for preparing calibration data and quantizing
BEVFormer models.
* **Bug Fixes**
* Prevented source ONNX models from being overwritten during
quantization and improved temporary model cleanup.
  * Fixed FP8 export for BF16 models during real-weight compression.
* **Documentation**
* Added setup, usage, compatibility, and performance guidance for the
BEVFormer workflow.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-11 16:15:59 -04:00

161 lines
5.7 KiB
Python

# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import contextlib
import re
import shutil
import tempfile
from pathlib import Path
import numpy as np
import onnx
from onnxruntime.quantization.calibrate import CalibrationDataReader
__all__ = [
"NpzCalibrationReader",
"NpzCalibrationWriter",
"find_vovnet_nodes_to_exclude",
"temporary_onnx_copy",
]
@contextlib.contextmanager
def temporary_onnx_copy(onnx_path):
"""Yield a sibling copy so relative external-data paths remain valid."""
onnx_path = Path(onnx_path)
temporary_file = tempfile.NamedTemporaryFile(
dir=onnx_path.parent,
prefix=f".{onnx_path.stem}.",
suffix=onnx_path.suffix,
delete=False,
)
temporary_path = Path(temporary_file.name)
temporary_file.close()
try:
shutil.copyfile(onnx_path, temporary_path)
yield temporary_path
finally:
temporary_path.unlink(missing_ok=True)
def _onnx_input_specs(onnx_path):
graph = onnx.load(onnx_path, load_external_data=False).graph
initializer_names = {initializer.name for initializer in graph.initializer}
input_specs = {}
for value in graph.input:
if value.name in initializer_names:
continue
tensor_type = value.type.tensor_type
shape = None
if tensor_type.HasField("shape"):
shape = tuple(
dimension.dim_value if dimension.HasField("dim_value") else None
for dimension in tensor_type.shape.dim
)
input_specs[value.name] = (
np.dtype(onnx.helper.tensor_dtype_to_np_dtype(tensor_type.elem_type)),
shape,
)
return input_specs
class NpzCalibrationWriter:
"""Write calibration batches that match an ONNX model's inputs."""
def __init__(self, output_dir, onnx_path):
self.output_dir = Path(output_dir)
self.output_dir.mkdir(parents=True, exist_ok=True)
if any(self.output_dir.glob("batch_*.npz")):
raise FileExistsError(f"{self.output_dir} already contains calibration batches")
self.input_specs = _onnx_input_specs(onnx_path)
self.count = 0
def write(self, values):
missing = self.input_specs.keys() - values.keys()
unexpected = values.keys() - self.input_specs.keys()
if missing or unexpected:
raise ValueError(
f"Calibration input mismatch; missing={sorted(missing)}, "
f"unexpected={sorted(unexpected)}"
)
batch = {}
for name, (dtype, expected_shape) in self.input_specs.items():
value = values[name]
if hasattr(value, "detach"):
value = value.detach().cpu().numpy()
value = np.asarray(value)
if expected_shape is not None and (
value.ndim != len(expected_shape)
or any(
expected is not None and actual != expected
for actual, expected in zip(value.shape, expected_shape)
)
):
raise ValueError(
f"Calibration input {name!r} has shape {value.shape}; expected {expected_shape}"
)
batch[name] = value.astype(dtype, copy=False)
np.savez(self.output_dir / f"batch_{self.count:04d}.npz", **batch)
self.count += 1
class NpzCalibrationReader(CalibrationDataReader):
"""Stream example-generated NPZ calibration batches."""
def __init__(self, calibration_dir):
self.batch_paths = sorted(Path(calibration_dir).glob("batch_*.npz"))
if not self.batch_paths:
raise ValueError(f"No calibration batches found in {calibration_dir}")
self.rewind()
@staticmethod
def load(batch_path):
with np.load(batch_path, allow_pickle=False) as batch:
return {name: batch[name] for name in batch.files}
def get_next(self):
batch_path = next(self._iterator, None)
return None if batch_path is None else self.load(batch_path)
def get_first(self):
return self.load(self.batch_paths[0])
def rewind(self):
self._iterator = iter(self.batch_paths)
def find_vovnet_nodes_to_exclude(onnx_path):
"""Find the VoVNet OSA4_5 stage and nodes downstream of FPN lateral_convs."""
# The evaluator image uses the calibration writer without installing ModelOpt.
from modelopt.onnx.utils import topologically_sort_graph_nodes
graph = onnx.load(onnx_path, load_external_data=False).graph
topologically_sort_graph_nodes(graph)
excluded = set()
downstream_tensors = set()
for node in graph.node:
is_osa = "OSA4_5" in node.name
is_downstream = any(name in downstream_tensors for name in node.input)
if is_osa or is_downstream:
excluded.add(node.name)
if "lateral_convs" in node.name or (is_downstream and not is_osa):
downstream_tensors.update(node.output)
if not excluded:
raise ValueError(f"No accuracy-sensitive VoVNet nodes found in {onnx_path}")
return [rf"^{re.escape(name)}$" for name in sorted(excluded)]