Files
Model-Optimizer/examples/onnx_ptq/petr/evaluate.py
T
Ajinkya RasaneandCodex 5c123ce183 [OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do?

Type of change: new example, example simplification, and
backward-breaking example migration

Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one
shared workflow:

- quantizes the shared VoVNet image backbone/encoder to INT8 or FP8;
- runs both the selected historical and current PETRv2 six-camera sweeps
through the same precision-matched TensorRT backbone engine using
distinct execution contexts during accuracy evaluation;
- keeps the PETR head and FAR3D decoder in their exported mixed
FP16/FP32 precision;
- reuses one NPZ calibration format, VoVNet exclusion helper,
quantization entry point, and TensorRT runner;
- does not change generic Model Optimizer calibration behavior or its
public CLI.

### Container boundary

Both examples use two targets from one Dockerfile, with no virtual
environments:

- `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base
with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup,
metadata generation, ONNX export, direct PyTorch calibration capture,
and final accuracy evaluation;
- `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base
for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization,
and TensorRT engine builds.

Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated
on the same GPU architecture. Final metrics remain in the evaluator
because they import the legacy model-framework postprocessing and
dataset code; only artifacts cross the container boundary through the
shared workspace.

PETR is used without patches. FAR3D applies only the official
`patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR
carries no patch files.

### Evaluator dependencies

The dependencies intentionally installed without transitive dependencies
are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on
runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator
base.

`lyft-dataset-sdk` is required only by mmdet3d's eager dataset import;
neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main
evaluator requirements because its compiled installation uses the
evaluator build step rather than the intentionally dependency-free
legacy package step.

Fresh setup and dependency approval is requested for the final reduced
dependency set.

### Reproducible PETR metadata

The documented workflow mounts raw nuScenes read-only and creates a
writable dataset view using symlinks. It then runs the pinned
mmdetection3d converter and a temporary, untracked copy of PETR's pinned
sweep generator configured only for the validation prefix and writable
dataset root.

A clean run generated both metadata files with 6,019 validation records.
The referenced camera, lidar, and sweep paths are absolute and
resolvable through the writable dataset view.

### Example-local utilities

The per-batch NPZ streaming and TensorRT runtime utilities remain
example-local because they execute in the legacy evaluator, where Model
Optimizer is not installed. The core `CalibrationDataProvider` consumes
one in-memory mapping of stacked arrays and does not provide this
streamed per-file workflow.

### Validation

- Focused CPU tests: 10 passed.
- Broader ONNX quantization CPU tests: 326 passed.
- All applicable pre-commit and documentation checks, plus `git diff
--check`, passed.
- Rebuilt both Docker targets and verified their exact dependency
versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization,
and absence of virtual environments.
- Generated both PETR metadata files from a clean writable dataset view
and verified 6,019 validation records plus resolvable data paths.
- PETRv1 passed a one-sample TensorRT regression smoke.
- PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full
6,019-sample validation. Both the selected historical and current sweeps
are computed by the matching backbone engine; accuracy evaluation no
longer extracts image features with PyTorch.
- FAR3D passed a recurrent two-frame TensorRT smoke covering plugin
loading and recurrent state.

TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after
correcting its temporal feature path; the PETRv1 and FAR3D numerical
paths are unchanged.

| Pipeline | FP16 | INT8 | FP8 |
| --- | ---: | ---: | ---: |
| PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 |
0.3707 | 0.3756 |
| PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head |
0.4102 | 0.3982 | 0.4084 |
| FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235
| 0.239 |

Normalized engine-only performance improvement over each matching FP16
pipeline:

| Pipeline | INT8 speedup | FP8 speedup |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
| FAR3D | 1.69x | 1.40x |

Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX
6000 Ada Generation GPU using five interleaved trials per engine
component. Each component uses the median `trtexec`-reported GPU Compute
Time with data transfers disabled and CUDA Graphs enabled. Component
times are summed before normalization: PETRv1 uses one backbone pass
plus its fixed head, PETRv2 uses two serial backbone passes plus its
fixed head with no temporal cache assumed, and FAR3D uses one encoder
pass plus its fixed decoder. Absolute latency values are intentionally
not published.

Adapted files retain exact public-source references and upstream
notices, and the top-level license attribution is updated.

- Is this change backward compatible?: ❌
- Did you write the necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-08 17:32:56 +00:00

180 lines
7.2 KiB
Python

# Adapted from https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/petr-trt/export_eval/v1/v1_evaluate_trt.py
# and https://github.com/NVIDIA/DL4AGX/blob/9f7b29104c253d5bc68334e7b83b3eecb72d4572/AV-Solutions/petr-trt/export_eval/v2/v2_evaluate_trt.py.
# Copyright (c) OpenMMLab. All rights reserved.
#
# SPDX-FileCopyrightText: Copyright (c) 2023-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
import argparse
import importlib
import os
import sys
from pathlib import Path
import torch
import torch.nn.functional as F
from mmcv import Config, DictAction
from mmcv.runner import load_checkpoint, wrap_fp16_model
from mmcv.utils import import_modules_from_strings
from mmdet.apis import set_random_seed
from mmdet3d.core import bbox3d2result
from mmdet3d.datasets import build_dataloader, build_dataset
from mmdet3d.models import build_model
from tqdm import tqdm
sys.path.insert(0, str(Path(__file__).resolve().parents[3]))
from examples.onnx_ptq.petr.petr_utils import run_backbone
from examples.onnx_ptq.trt_runner import TensorRTRunner
__all__ = ["PETRPipeline", "build_runtime", "get_head_inputs"]
def import_plugin(cfg):
if cfg.get("custom_imports"):
import_modules_from_strings(**cfg.custom_imports)
plugin_dir = cfg.get("plugin_dir")
if cfg.get("plugin") and plugin_dir:
importlib.import_module(".".join(os.path.dirname(plugin_dir).split("/")))
def build_runtime(config_path, checkpoint_path, cfg_options=None):
cfg = Config.fromfile(config_path)
if cfg_options:
cfg.merge_from_dict(cfg_options)
import_plugin(cfg)
cfg.model.pretrained = None
cfg.model.train_cfg = None
cfg.data.test.test_mode = True
dataset = build_dataset(cfg.data.test)
loader = build_dataloader(
dataset,
samples_per_gpu=1,
workers_per_gpu=cfg.data.workers_per_gpu,
dist=False,
shuffle=False,
)
model = build_model(cfg.model, test_cfg=cfg.get("test_cfg"))
if cfg.get("fp16"):
wrap_fp16_model(model)
checkpoint = load_checkpoint(model, checkpoint_path, map_location="cpu")
model.CLASSES = checkpoint.get("meta", {}).get("CLASSES", dataset.CLASSES)
if hasattr(dataset, "PALETTE"):
model.PALETTE = checkpoint.get("meta", {}).get("PALETTE", dataset.PALETTE)
model = model.cuda().eval()
return cfg, dataset, loader, model
def get_head_inputs(version, model, features, img_metas):
batch_size, num_cams = features[0].shape[:2]
input_h, input_w, _ = img_metas[0]["pad_shape"][0]
masks = features[0].new_ones((batch_size, num_cams, input_h, input_w))
for image_id in range(batch_size):
for camera_id in range(num_cams):
image_h, image_w, _ = img_metas[image_id]["img_shape"][camera_id]
masks[image_id, camera_id, :image_h, :image_w] = 0
masks = F.interpolate(masks, size=features[0].shape[-2:]).to(torch.bool)
coords, _ = model.pts_bbox_head.position_embeding(features, img_metas, masks)
inputs = {
"mlvl_feats.0": features[0],
"img_metas.0[coords_position_embeding]": coords,
}
if version == "v2":
timestamps = features[0].new_tensor([meta["timestamp"] for meta in img_metas])
timestamps = timestamps.view(1, -1, 6)
inputs["img_metas.0[mean_time_stamp]"] = (timestamps[:, 1] - timestamps[:, 0]).mean(-1)
return inputs
class PETRPipeline:
def __init__(self, version, model, backbone_engine, head_engine):
self.version = version
self.model = model
self.backbone = TensorRTRunner(backbone_engine)
self.history_backbone = (
self.backbone.new_context(state_names=("prev.0", "prev.1")) if version == "v2" else None
)
self.head = TensorRTRunner(head_engine)
def __call__(self, stream, data):
images = data["img"][0].data[0].cuda()
img_metas = data["img_metas"][0].data[0]
with torch.cuda.stream(stream), torch.no_grad():
feature_outputs = run_backbone(
self.version, self.backbone, self.history_backbone, stream, images
)
camera_count = 6 if self.version == "v1" else 12
features = [
feature_outputs[name].reshape(1, camera_count, *feature_outputs[name].shape[-3:])
for name in ("out.0", "out.1")
]
outputs = self.head(
stream, **get_head_inputs(self.version, self.model, features, img_metas)
)
head_outputs = {
"all_cls_scores": outputs["out.all_cls_scores"].float(),
"all_bbox_preds": outputs["out.all_bbox_preds"].float(),
"enc_cls_scores": None,
"enc_bbox_preds": None,
}
boxes = self.model.pts_bbox_head.get_bboxes(head_outputs, img_metas, rescale=True)
torch.cuda.current_stream().wait_stream(stream)
return [
{"pts_bbox": bbox3d2result(boxes_3d, scores_3d, labels_3d)}
for boxes_3d, scores_3d, labels_3d in boxes
]
def parse_args():
parser = argparse.ArgumentParser(description="Evaluate PETR TensorRT engines on nuScenes")
parser.add_argument("version", choices=("v1", "v2"))
parser.add_argument("config")
parser.add_argument("checkpoint")
parser.add_argument("backbone_engine")
parser.add_argument("head_engine")
parser.add_argument("--cfg-options", nargs="+", action=DictAction)
parser.add_argument("--eval-options", nargs="+", action=DictAction)
parser.add_argument("--max-samples", type=int)
args = parser.parse_args()
if args.max_samples is not None and args.max_samples < 1:
raise ValueError("--max-samples must be positive")
return args
def main():
args = parse_args()
set_random_seed(0, deterministic=False)
cfg, dataset, loader, model = build_runtime(args.config, args.checkpoint, args.cfg_options)
pipeline = PETRPipeline(args.version, model, args.backbone_engine, args.head_engine)
stream = torch.cuda.Stream()
outputs = []
for data in tqdm(loader):
outputs.extend(pipeline(stream, data))
if args.max_samples is not None and len(outputs) >= args.max_samples:
break
if len(outputs) < len(dataset):
print(f"Processed {len(outputs)} samples; skipping dataset metrics")
return
eval_kwargs = cfg.get("evaluation", {}).copy()
for key in ("interval", "tmpdir", "start", "gpu_collect", "save_best", "rule"):
eval_kwargs.pop(key, None)
if args.eval_options:
eval_kwargs.update(args.eval_options)
print(dataset.evaluate(outputs, **eval_kwargs))
if __name__ == "__main__":
main()