Files
Model-Optimizer/examples/onnx_ptq/requirements-evaluator-nodeps.txt
Ajinkya RasaneandCodex 5c123ce183 [OMNIML-5563] Add PETR ONNX PTQ and accuracy evaluation example (#2180)
### What does this PR do?

Type of change: new example, example simplification, and
backward-breaking example migration

Adds end-to-end PETRv1/PETRv2 ONNX PTQ and reduces PETR/FAR3D to one
shared workflow:

- quantizes the shared VoVNet image backbone/encoder to INT8 or FP8;
- runs both the selected historical and current PETRv2 six-camera sweeps
through the same precision-matched TensorRT backbone engine using
distinct execution contexts during accuracy evaluation;
- keeps the PETR head and FAR3D decoder in their exported mixed
FP16/FP32 precision;
- reuses one NPZ calibration format, VoVNet exclusion helper,
quantization entry point, and TensorRT runner;
- does not change generic Model Optimizer calibration behavior or its
public CLI.

### Container boundary

Both examples use two targets from one Dockerfile, with no virtual
environments:

- `evaluator`: a digest-pinned `nvcr.io/nvidia/pytorch:22.06-py3` base
with the legacy PyTorch 1.13.1/OpenMMLab stack for source setup,
metadata generation, ONNX export, direct PyTorch calibration capture,
and final accuracy evaluation;
- `modelopt`: a digest-pinned `nvcr.io/nvidia/pytorch:26.07-py3` base
for Model Optimizer, ONNX Runtime CUDA, AutoCast, INT8/FP8 quantization,
and TensorRT engine builds.

Both targets use TensorRT `11.1.0.106`. Engines are built and evaluated
on the same GPU architecture. Final metrics remain in the evaluator
because they import the legacy model-framework postprocessing and
dataset code; only artifacts cross the container boundary through the
shared workspace.

PETR is used without patches. FAR3D applies only the official
`patch/far3d.patch` from the pinned NVIDIA DL4AGX revision. This PR
carries no patch files.

### Evaluator dependencies

The dependencies intentionally installed without transitive dependencies
are listed in `requirements-evaluator-nodeps.txt`. Their pins rely on
runtime packages supplied by the digest-pinned PyTorch 22.06 evaluator
base.

`lyft-dataset-sdk` is required only by mmdet3d's eager dataset import;
neither PETR nor FAR3D uses Lyft data. `flash-attn` remains in the main
evaluator requirements because its compiled installation uses the
evaluator build step rather than the intentionally dependency-free
legacy package step.

Fresh setup and dependency approval is requested for the final reduced
dependency set.

### Reproducible PETR metadata

The documented workflow mounts raw nuScenes read-only and creates a
writable dataset view using symlinks. It then runs the pinned
mmdetection3d converter and a temporary, untracked copy of PETR's pinned
sweep generator configured only for the validation prefix and writable
dataset root.

A clean run generated both metadata files with 6,019 validation records.
The referenced camera, lidar, and sweep paths are absolute and
resolvable through the writable dataset view.

### Example-local utilities

The per-batch NPZ streaming and TensorRT runtime utilities remain
example-local because they execute in the legacy evaluator, where Model
Optimizer is not installed. The core `CalibrationDataProvider` consumes
one in-memory mapping of stacked arrays and does not provide this
streamed per-file workflow.

### Validation

- Focused CPU tests: 10 passed.
- Broader ONNX quantization CPU tests: 326 passed.
- All applicable pre-commit and documentation checks, plus `git diff
--check`, passed.
- Rebuilt both Docker targets and verified their exact dependency
versions, imports, TensorRT `11.1.0.106`, GPU runtime initialization,
and absence of virtual environments.
- Generated both PETR metadata files from a clean writable dataset view
and verified 6,019 validation records plus resolvable data paths.
- PETRv1 passed a one-sample TensorRT regression smoke.
- PETRv2 passed FP16, INT8, and FP8 TensorRT smokes and full
6,019-sample validation. Both the selected historical and current sweeps
are computed by the matching backbone engine; accuracy evaluation no
longer extracts image features with PyTorch.
- FAR3D passed a recurrent two-frame TensorRT smoke covering plugin
loading and recurrent state.

TensorRT `11.1.0.106` mAP follows. PETRv2 was remeasured after
correcting its temporal feature path; the PETRv1 and FAR3D numerical
paths are unchanged.

| Pipeline | FP16 | INT8 | FP8 |
| --- | ---: | ---: | ---: |
| PETRv1: 1 backbone pass + fixed typed mixed FP16/FP32 head | 0.3778 |
0.3707 | 0.3756 |
| PETRv2: 2 serial backbone passes + fixed typed mixed FP16/FP32 head |
0.4102 | 0.3982 | 0.4084 |
| FAR3D: 1 encoder pass + fixed mixed FP16/FP32 decoder | 0.241 | 0.235
| 0.239 |

Normalized engine-only performance improvement over each matching FP16
pipeline:

| Pipeline | INT8 speedup | FP8 speedup |
| --- | ---: | ---: |
| PETRv1 | 1.49x | 1.29x |
| PETRv2 | 1.51x | 1.30x |
| FAR3D | 1.69x | 1.40x |

Performance was measured with TensorRT `11.1.0.106` on an NVIDIA RTX
6000 Ada Generation GPU using five interleaved trials per engine
component. Each component uses the median `trtexec`-reported GPU Compute
Time with data transfers disabled and CUDA Graphs enabled. Component
times are summed before normalization: PETRv1 uses one backbone pass
plus its fixed head, PETRv2 uses two serial backbone passes plus its
fixed head with no temporal cache assumed, and FAR3D uses one encoder
pass plus its fixed decoder. Absolute latency values are intentionally
not published.

Adapted files retain exact public-source references and upstream
notices, and the top-level license attribution is updated.

- Is this change backward compatible?: ❌
- Did you write the necessary tests?: ✅
- Did you update the changelog?: ✅

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-08 17:32:56 +00:00

7 lines
184 B
Plaintext

# Training and development dependencies are intentionally omitted for these legacy projects.
av2==0.2.1
lyft-dataset-sdk==0.0.8
mmdet3d==1.0.0rc6
nuscenes-devkit==1.1.11
refile==0.4.1