mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new example Adds an end-to-end BEVFormer-tiny ONNX PTQ example under `examples/onnx_ptq/bevformer` with: - The [BEVFormer Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile) from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second container definition in this repository. - A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10 source patch, and TensorRT plugin compilation at container runtime. - Ordered temporal calibration-data generation that propagates `prev_bev`, resets state at scene boundaries, computes CAN bus deltas, writes the exact requested sample count, and publishes output atomically. - INT8 and FP8 quantization with BEVFormer-specific calibration defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback policy. - Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus nuScenes evaluation instructions. - Shared temporary-ONNX-copy handling for BEVFormer and VoVNet quantization so shape inference cannot mutate the source model. - Focused CPU-only tests for temporal state, exact-count cleanup, quantization defaults, plugin configuration, and source-model preservation. The ONNX PTQ index continues to document the shared PETR/FAR3D containers separately and links to the BEVFormer guide. ### Usage Clone the pinned DL4AGX revision and build its BEVFormer image: ```bash git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX git -C /path/to/DL4AGX checkout --detach \ 9f7b29104c253d5bc68334e7b83b3eecb72d4572 docker build \ --build-arg TORCH_CUDA_ARCH_LIST=8.9 \ --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \ --tag modelopt-onnx-bevformer \ /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq ``` The example command targets compute capability 8.9; use the deployment GPU's compute capability for another architecture. After exporting the model and generating temporal calibration data, quantize it with: ```bash python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \ --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \ --calibration-dir=/artifacts/calibration \ --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \ --quantization-mode=fp8 \ --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx ``` See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin compilation, export, FP16 feedback-engine creation, temporal calibration, INT8/FP8 engine builds, and evaluation. ### Validation #### Current revision: three-sample smoke validation The Dockerfile at the pinned DL4AGX revision built successfully, and the current Model Optimizer checkout was mounted into it for the workflow below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch 2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage. A fresh three-frame A/B/B smoke passed without calculating partial NDS or mAP: - Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and FP8 engine builds passed. - Temporal calibration published exactly three batches with `use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent feedback on the third frame, and the expected CAN bus deltas. - The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero points. - Every engine produced finite outputs and 300 non-empty decoded detections for each frame; maximum scores ranged from 0.9489 to 0.9608. CPU-only validation passed the 16-test focused example set and the full ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the warning-as-error documentation build also passed. #### Historical full accuracy reference The following results were collected previously with TensorRT 10.14.1.48 on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are reference results, not a completed full evaluation of the current revision. | Precision | NDS | mAP | | :-- | --: | --: | | FP16 | 0.3546 | 0.2515 | | INT8/FP16 | 0.3512 | 0.2505 | | FP8/FP16 | 0.3526 | 0.2489 | #### Performance Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. The archived full-validation engines were benchmarked with five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream. | Precision | Speedup vs. FP16 | | :-- | --: | | FP16 | 1.00x | | INT8/FP16 | 1.87x | | FP8/FP16 | 1.18x | TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from another source or added a new PIP dependency, did you follow the guidance in `CONTRIBUTING.md`?: ✅ - Did you write the necessary tests?: ✅ - Did you update the [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional information Reference workflow: [NVIDIA DL4AGX BEVFormer INT8 example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq). > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end BEVFormer 3D detection workflow for ONNX post-training quantization, including temporal calibration, INT8/FP8 quantization, TensorRT engine generation, and nuScenes evaluation. * Added command-line tools for preparing calibration data and quantizing BEVFormer models. * **Bug Fixes** * Prevented source ONNX models from being overwritten during quantization and improved temporary model cleanup. * Fixed FP8 export for BF16 models during real-weight compression. * **Documentation** * Added setup, usage, compatibility, and performance guidance for the BEVFormer workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com>
67 lines
2.3 KiB
Python
67 lines
2.3 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
import argparse
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
from modelopt.onnx.quantization import quantize
|
|
|
|
sys.path.insert(0, str(Path(__file__).resolve().parents[2]))
|
|
|
|
from examples.onnx_ptq.quantization_utils import (
|
|
NpzCalibrationReader,
|
|
find_vovnet_nodes_to_exclude,
|
|
temporary_onnx_copy,
|
|
)
|
|
|
|
|
|
def parse_args():
|
|
parser = argparse.ArgumentParser(description="Quantize a VoVNet ONNX image encoder")
|
|
parser.add_argument("onnx_path")
|
|
parser.add_argument("calibration_dir", type=Path)
|
|
parser.add_argument("--precision", choices=("int8", "fp8"), default="int8")
|
|
parser.add_argument("--output")
|
|
return parser.parse_args()
|
|
|
|
|
|
def main():
|
|
args = parse_args()
|
|
onnx_path = Path(args.onnx_path)
|
|
output_path = (
|
|
Path(args.output)
|
|
if args.output
|
|
else onnx_path.with_name(f"{onnx_path.stem}.{args.precision}{onnx_path.suffix}")
|
|
)
|
|
if output_path.resolve() == onnx_path.resolve():
|
|
raise ValueError("Output path must differ from the source ONNX path")
|
|
excluded_nodes = find_vovnet_nodes_to_exclude(onnx_path)
|
|
print(f"Excluding {len(excluded_nodes)} accuracy-sensitive VoVNet nodes")
|
|
with temporary_onnx_copy(onnx_path) as temporary_onnx:
|
|
quantize(
|
|
onnx_path=str(temporary_onnx),
|
|
quantize_mode=args.precision,
|
|
calibration_data_reader=NpzCalibrationReader(args.calibration_dir),
|
|
calibration_method="max",
|
|
calibration_eps=["cuda:0", "cpu"],
|
|
nodes_to_exclude=excluded_nodes,
|
|
high_precision_dtype="fp16",
|
|
output_path=str(output_path),
|
|
)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|