Files
Ajinkya Rasane 4f62d418f4 Use TensorRT optimization level 0 in Torch ONNX example tests (#2619)
### What does this PR do?

Type of change: Bug fix

The Torch ONNX example tests can exceed their 300-second deadline while
building a ResNet50 INT8 TensorRT engine at optimization level 4. Add
`--trt_builder_optimization_level` to the vision example and select
level 0 in the existing integration tests. The example and helper retain
level 4 by default. Quantization, ONNX export, residual Q/DQ assertions,
engine execution, and test timeout limits are unchanged.

Document the build-time versus inference-performance tradeoff in the
example README.

### Usage

```bash
cd examples/torch_onnx
python torch_quant_to_onnx.py \
    --timm_model_name resnet50 \
    --recipe timm/resnet/ptq/int8 \
    --onnx_save_path resnet50.int8.onnx \
    --calibration_data_size 1 --no_pretrained \
    --trt_build --trt_builder_optimization_level 0
```

### Testing

Validation used the TensorRT 26.05 container, TensorRT 10.16.1.11, and
PyTorch 2.13.0, with the existing 300-second per-test deadline.

- RTX 6000 Ada: **8 passed**, covering FP8 and INT8 on ViT, Swin,
SwinV2, and ResNet50. ResNet50 INT8 passed in 96.65 seconds.
- RTX PRO 6000 Blackwell Max-Q: **20 passed, 3 existing skips**,
covering the complete test file. ResNet50 INT8 passed in 81.27 seconds;
the baseline timed out at 300 seconds.

```bash
# RTX 6000 Ada: supported FP8/INT8 cases
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py \
    -k '(fp8 or int8) and not mxfp8' --cov

# RTX PRO 6000 Blackwell: complete example test file
python -m pytest tests/examples/torch_onnx/test_torch_quant_to_onnx.py --cov
```

On each GPU, levels 4 and 0 used the same exported ResNet50 INT8 graph
with TensorRT 10.16.1.11:

- RTX 6000 Ada: TensorRT-reported engine build time decreased from 122.6
seconds at level 4 to 27.7 seconds at level 0. Both builds and inference
runs succeeded.
- RTX PRO 6000 Blackwell Max-Q: the original level-4 test hit its
300-second deadline during the engine build; level 0 built that saved
graph in 10.9 seconds and completed inference successfully.

All pre-commit checks passed for the changed files. The existing
integration tests exercise the real engine build; no redundant mocked
tests were added.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing integration
tests were updated to exercise level 0; no new test cases were needed.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — minor example/CI fix; library behavior and example defaults are
preserved.
- Did you get Claude approval on this PR?: N/A — not requested for this
focused change.

### Additional Information

Example timeout: [ResNet50 INT8 CI
failure](https://github.com/NVIDIA/Model-Optimizer/actions/runs/36560176214/job/109381161169).


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added a configurable TensorRT builder optimization level for engine
builds, with a default of 4 and support for values from 0 to 5.
* Documented that level 0 can speed up builds, while lower optimization
levels may reduce inference performance.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-10-01 18:32:18 +00:00
..

Model Optimizer End-to-End Tests

This directory contains end-to-end tests for the Model Optimizer's examples.

Adding new tests

To add a test for a new example, create a new example directory in tests/examples following the existing examples as guidance. Make sure to use as small models and less data as possible to keep the tests fast. Unless needed, have tests finish under 15 minutes.

Running the tests

To run a test, start from the recommended docker image from our installation docs. Then mount your local modelopt directory to /workspace/Model-Optimizer and run this from the root of the repository.

cd /workspace/Model-Optimizer
pip install -e ".[all,dev-test]"
pytest tests/examples/$TEST

Environment variables

The following environment variables can be set to control the behavior of the tests:

  • MODELOPT_LOCAL_MODEL_ROOT: If set, the tests will use the local model directory instead of downloading the model from the internet. Default is not set, which means the model will be downloaded.