mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
bd90a5ed51 |
[OMNIML-5774] Add BEVFormer ONNX PTQ and evaluation example (#2208)
### What does this PR do? Type of change: new example Adds an end-to-end BEVFormer-tiny ONNX PTQ example under `examples/onnx_ptq/bevformer` with: - The [BEVFormer Dockerfile](https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile) from pinned NVIDIA DL4AGX commit `9f7b291`, rather than a second container definition in this repository. - A mounted Model Optimizer checkout, the official DL4AGX TensorRT 10 source patch, and TensorRT plugin compilation at container runtime. - Ordered temporal calibration-data generation that propagates `prev_bev`, resets state at scene boundaries, computes CAN bus deltas, writes the exact requested sample count, and publishes output atomically. - INT8 and FP8 quantization with BEVFormer-specific calibration defaults, custom-plugin handling, `MatMul` exclusions, and FP16 fallback policy. - Strongly typed FP16, INT8, and FP8 TensorRT engine generation plus nuScenes evaluation instructions. - Shared temporary-ONNX-copy handling for BEVFormer and VoVNet quantization so shape inference cannot mutate the source model. - Focused CPU-only tests for temporal state, exact-count cleanup, quantization defaults, plugin configuration, and source-model preservation. The ONNX PTQ index continues to document the shared PETR/FAR3D containers separately and links to the BEVFormer guide. ### Usage Clone the pinned DL4AGX revision and build its BEVFormer image: ```bash git clone https://github.com/NVIDIA/DL4AGX.git /path/to/DL4AGX git -C /path/to/DL4AGX checkout --detach \ 9f7b29104c253d5bc68334e7b83b3eecb72d4572 docker build \ --build-arg TORCH_CUDA_ARCH_LIST=8.9 \ --file /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq/docker/tensorrt.Dockerfile \ --tag modelopt-onnx-bevformer \ /path/to/DL4AGX/AV-Solutions/bevformer-int8-eq ``` The example command targets compute capability 8.9; use the deployment GPU's compute capability for another architecture. After exporting the model and generating temporal calibration data, quantize it with: ```bash python /opt/Model-Optimizer/examples/onnx_ptq/bevformer/quantize.py \ --onnx=/artifacts/bevformer_tiny_epoch_24_cp2_op13.onnx \ --calibration-dir=/artifacts/calibration \ --trt-plugins=/workspace/BEVFormer_tensorrt/TensorRT/lib/libtensorrt_ops.so \ --quantization-mode=fp8 \ --output=/artifacts/bevformer_tiny_epoch_24_cp2_op13.fp8.onnx ``` See `examples/onnx_ptq/bevformer/README.md` for dataset setup, plugin compilation, export, FP16 feedback-engine creation, temporal calibration, INT8/FP8 engine builds, and evaluation. ### Validation #### Current revision: three-sample smoke validation The Dockerfile at the pinned DL4AGX revision built successfully, and the current Model Optimizer checkout was mounted into it for the workflow below. The runtime audit confirmed TensorRT 10.14.1.48, CUDA 13.1, Torch 2.9, the TensorRT/CUDA/CPU execution providers, and plugin linkage. A fresh three-frame A/B/B smoke passed without calculating partial NDS or mAP: - Fresh Torch 2.9 export, AutoCast, and strongly typed FP16, INT8, and FP8 engine builds passed. - Temporal calibration published exactly three batches with `use_prev_bev=[0, 0, 1]`, zero state at both scene starts, recurrent feedback on the third frame, and the expected CAN bus deltas. - The source ONNX hash remained unchanged. INT8 contained 136 Q/DQ pairs with INT8 zero points; FP8 contained 127 Q/DQ pairs with FP8 zero points. - Every engine produced finite outputs and 300 non-empty decoded detections for each frame; maximum scores ranged from 0.9489 to 0.9608. CPU-only validation passed the 16-test focused example set and the full ONNX session (`649 passed, 1 skipped`). Full-diff pre-commit and the warning-as-error documentation build also passed. #### Historical full accuracy reference The following results were collected previously with TensorRT 10.14.1.48 on an NVIDIA RTX 6000 Ada Generation GPU, using 600 ordered calibration samples and all 6,019 nuScenes validation samples. They are reference results, not a completed full evaluation of the current revision. | Precision | NDS | mAP | | :-- | --: | --: | | FP16 | 0.3546 | 0.2515 | | INT8/FP16 | 0.3512 | 0.2505 | | FP8/FP16 | 0.3526 | 0.2489 | #### Performance Following the PETR and FAR3D convention, performance is reported only as speedup normalized to FP16. The archived full-validation engines were benchmarked with five interleaved TensorRT 10.14 trials on the same NVIDIA RTX 6000 Ada Generation GPU. Each trial used the `trtexec` GPU Compute Time median with data transfers disabled, CUDA Graphs enabled, spin-wait, a 1-second warmup, a 10-second measurement window, and one inference stream. | Precision | Speedup vs. FP16 | | :-- | --: | | FP16 | 1.00x | | INT8/FP16 | 1.87x | | FP8/FP16 | 1.18x | TODO: Investigate why FP8 delivers less speedup than INT8 for BEVFormer-tiny. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ - If you copied code from another source or added a new PIP dependency, did you follow the guidance in `CONTRIBUTING.md`?: ✅ - Did you write the necessary tests?: ✅ - Did you update the [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ ### Additional information Reference workflow: [NVIDIA DL4AGX BEVFormer INT8 example](https://github.com/NVIDIA/DL4AGX/tree/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/bevformer-int8-eq). > 🤖 _Generated by Codex (AI agent)._ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added an end-to-end BEVFormer 3D detection workflow for ONNX post-training quantization, including temporal calibration, INT8/FP8 quantization, TensorRT engine generation, and nuScenes evaluation. * Added command-line tools for preparing calibration data and quantizing BEVFormer models. * **Bug Fixes** * Prevented source ONNX models from being overwritten during quantization and improved temporary model cleanup. * Fixed FP8 export for BF16 models during real-weight compression. * **Documentation** * Added setup, usage, compatibility, and performance guidance for the BEVFormer workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |
||
|
|
5c123ce183 |
[OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do? Type of change: new example, example simplification, and backward-breaking example migration Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one shared workflow: - quantizes the shared VoVNet image backbone/encoder to INT8 or FP8; - runs both the selected historical and current PETRv2 six-camera sweeps through the same precision-matched TensorRT backbone engine using distinct execution contexts during accuracy evaluation; - keeps the PETR head and FAR3D decoder in their exported mixed FP16/FP32 precision; - reuses one NPZ calibration format, VoVNet exclusion helper, quantization entry point, and TensorRT runner; - does not change generic Model Optimizer calibration behavior or its public CLI. ### Container boundary Both examples use two targets from one Dockerfile, with no virtual environments: - `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup, metadata generation, ONNX export, direct PyTorch calibration capture, and final accuracy evaluation; - `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization, and TensorRT engine builds. Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated on the same GPU architecture. Final metrics remain in the evaluator because they import the legacy model-framework postprocessing and dataset code; only artifacts cross the container boundary through the shared workspace. PETR is used without patches. FAR3D applies only the official `patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR carries no patch files. ### Evaluator dependencies The dependencies intentionally installed without transitive dependencies are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator base. `lyft-dataset-sdk` is required only by mmdet3d's eager dataset import; neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main evaluator requirements because its compiled installation uses the evaluator build step rather than the intentionally dependency-free legacy package step. Fresh setup and dependency approval is requested for the final reduced dependency set. ### Reproducible PETR metadata The documented workflow mounts raw nuScenes read-only and creates a writable dataset view using symlinks. It then runs the pinned mmdetection3d converter and a temporary, untracked copy of PETR's pinned sweep generator configured only for the validation prefix and writable dataset root. A clean run generated both metadata files with 6,019 validation records. The referenced camera, lidar, and sweep paths are absolute and resolvable through the writable dataset view. ### Example-local utilities The per-batch NPZ streaming and TensorRT runtime utilities remain example-local because they execute in the legacy evaluator, where Model Optimizer is not installed. The core `CalibrationDataProvider` consumes one in-memory mapping of stacked arrays and does not provide this streamed per-file workflow. ### Validation - Focused CPU tests: 10 passed. - Broader ONNX quantization CPU tests: 326 passed. - All applicable pre-commit and documentation checks, plus `git diff --check`, passed. - Rebuilt both Docker targets and verified their exact dependency versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization, and absence of virtual environments. - Generated both PETR metadata files from a clean writable dataset view and verified 6,019 validation records plus resolvable data paths. - PETRv1 passed a one-sample TensorRT regression smoke. - PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full 6,019-sample validation. Both the selected historical and current sweeps are computed by the matching backbone engine; accuracy evaluation no longer extracts image features with PyTorch. - FAR3D passed a recurrent two-frame TensorRT smoke covering plugin loading and recurrent state. TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after correcting its temporal feature path; the PETRv1 and FAR3D numerical paths are unchanged. | Pipeline | FP16 | INT8 | FP8 | | --- | ---: | ---: | ---: | | PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 | 0.3707 | 0.3756 | | PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head | 0.4102 | 0.3982 | 0.4084 | | FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235 | 0.239 | Normalized engine-only performance improvement over each matching FP16 pipeline: | Pipeline | INT8 speedup | FP8 speedup | | --- | ---: | ---: | | PETRv1 | 1.49x | 1.29x | | PETRv2 | 1.51x | 1.30x | | FAR3D | 1.69x | 1.40x | Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX 6000 Ada Generation GPU using five interleaved trials per engine component. Each component uses the median `trtexec`-reported GPU Compute Time with data transfers disabled and CUDA Graphs enabled. Component times are summed before normalization: PETRv1 uses one backbone pass plus its fixed head, PETRv2 uses two serial backbone passes plus its fixed head with no temporal cache assumed, and FAR3D uses one encoder pass plus its fixed decoder. Absolute latency values are intentionally not published. Adapted files retain exact public-source references and upstream notices, and the top-level license attribution is updated. - Is this change backward compatible?: ❌ - Did you write the necessary tests?: ✅ - Did you update the changelog?: ✅ > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> |