[6008361][ONNX][Quantization] Clarify autotune guidance (#1989)

### What does this PR do?

Type of change: documentation

This PR clarifies when users should use ONNX quantization with Autotune
enabled versus the direct Autotune entry point.

- Adds a warning to the Autotune guide explaining that direct Autotune
is a lower-level Q/DQ placement tool and does not replace calibrated
ONNX PTQ.
- Updates the ONNX quantization guide to show `autotune=True` in the
Python API and explain that it uses default Autotune settings.
- Updates ONNX PTQ example documentation to prefer `python -m
modelopt.onnx.quantization ... --autotune=<mode>` for accuracy-sensitive
PTQ from an unquantized model.
- Updates the direct Autotune CLI help text to point users back to the
full ONNX quantization workflow when calibration data and accuracy
validation are required.

### Usage

```python
N/A — documentation/help text change.
```

### Testing

- Ran `python -m py_compile
modelopt/onnx/quantization/autotune/__main__.py`.
- Built the Sphinx documentation with `python -m sphinx -b html
docs/source docs/build/html`; build succeeded. Remaining warnings are
from optional documentation imports and existing cross-reference labels.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified that Direct Autotune is an advanced tool for Q/DQ placement
experiments, not a replacement for full calibrated ONNX quantization.
* Added guidance on when to use Direct Autotune versus the end-to-end
ONNX PTQ workflow.
* Documented the optional `autotune=True` setting, expected
calibration-time impact, and representative calibration data
requirements.
* Expanded links and guidance across ONNX PTQ examples and updated CLI
help text with the recommended workflow.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gwenaelle Cunha Sergio <gcunhasergio@nvidia.com>
This commit is contained in:
Gwena Cunha
2026-07-23 21:17:49 +00:00
committed by GitHub
parent 3edd137d22
commit d984de3795
5 changed files with 31 additions and 6 deletions
+12 -3
View File
@@ -2,6 +2,12 @@
Autotune (ONNX)
===============================================
.. warning::
Direct Autotune is an advanced ONNX quantization subtool for optimizing Q/DQ placement using TensorRT latency measurements. It does not replace the full calibrated ONNX quantization workflow.
To quantize an ONNX model taking into consideration accuracy, please use ``python -m modelopt.onnx.quantization ... --autotune=<quick|default|extensive>`` with representative calibration data. See the `ONNX quantization Autotune options <_onnx_quantization.html#python-m-modelopt.onnx.quantization-autotune-only-applicable-when-autotune-is-set>`_.
.. contents:: Table of Contents
:local:
:depth: 2
@@ -22,9 +28,10 @@ The ``modelopt.onnx.quantization.autotune`` module automates Q/DQ (Quantize/Dequ
**When to Use This Tool:**
* Quantizing an ONNX model for TensorRT deployment
* Optimizing Q/DQ placement for best performance
* The model has repeating structures (e.g., transformer blocks, ResNet layers)
* Debugging or developing the Q/DQ placement autotuning algorithm
* Running the lower-level workflow without invoking the full quantization CLI
* Programmatic experiments with direct Autotune classes and workflow functions
* Expert workflows that intentionally start from already-quantized or pre-patterned Q/DQ models
Quick Start
===========
@@ -54,6 +61,8 @@ The command will:
4. Select the best scheme based on TensorRT latency measurements
5. Export an optimized ONNX model with Q/DQ nodes
Autotune searches for Q/DQ placement schemes that improve TensorRT runtime. It does not by itself define the full calibration and quantization policy for an accuracy-sensitive deployment. For end-to-end ONNX PTQ that starts from an unquantized model, run ONNX quantization with calibration data and enable ``--autotune`` there. See the `ONNX quantization Autotune options <_onnx_quantization.html#python-m-modelopt.onnx.quantization-autotune-only-applicable-when-autotune-is-set>`_.
**Output Files:**
Files are written under the output directory (default ``./autotuner_output``, or the path given by ``--output_dir``):
+10
View File
@@ -74,6 +74,16 @@ Call PTQ function
quantize_mode="int8",
)
Optionally enable Autotune for more optimized Q/DQ placement. Note that this will likely increase the time required to calibrate the model.
.. code-block:: python
moq.quantize(
...
# Default Autotune settings, can be tuned with the autotune_* arguments below.
autotune=True,
)
Alternatively, you can call PTQ function in command line:
.. argparse::
+1 -2
View File
@@ -210,7 +210,6 @@ trtexec --onnx=/tmp/identity_neural_network.quant.onnx \
### Optimize Q/DQ node placement with Autotune
This feature automates Q/DQ (Quantize/Dequantize) node placement optimization for ONNX models using TensorRT performance measurements.
For more information on the standalone toolkit, please refer to [autotune](./autotune).
To access this feature in the ONNX quantization workflow, simply add `--autotune` in your CLI:
@@ -224,7 +223,7 @@ python -m modelopt.onnx.quantization \
--autotune=<quick,default,extensive>
```
For more fine-tuned Autotune flags, please refer to the [API guide](https://nvidia.github.io/Model-Optimizer/guides/_onnx_quantization.html).
For more fine-tuned Autotune flags, please refer to the [API guide](https://nvidia.github.io/Model-Optimizer/guides/_onnx_quantization.html) and the [Autotune guide](https://nvidia.github.io/Model-Optimizer/guides/9_autotune.html).
## Resources
+2
View File
@@ -2,6 +2,8 @@
This example demonstrates automated Q/DQ (Quantize/Dequantize) node placement optimization for ONNX models using TensorRT performance measurements.
> **Warning:** This example uses the direct Autotune entry point for lower-level Q/DQ placement experiments. If you are starting from an unquantized ONNX model and care about **accuracy**, please use the ONNX PTQ workflow with `--autotune` enabled and with representative calibration data. See [../README#optimize-qdq-node-placement-with-autotune](../README#optimize-qdq-node-placement-with-autotune).
## Table of Contents
<div align="center">
@@ -154,7 +154,12 @@ def get_parser() -> argparse.ArgumentParser:
"""
parser = argparse.ArgumentParser(
prog="modelopt.onnx.quantization.autotune",
description="ONNX Q/DQ Autotuning with TensorRT",
description=(
"ONNX Q/DQ placement autotuning with TensorRT latency measurements. "
"For end-to-end ONNX quantization taking into consideration accuracy, use "
"`python -m modelopt.onnx.quantization ... --autotune=<mode>` "
"with representative calibration data."
),
formatter_class=argparse.RawDescriptionHelpFormatter,
epilog="""
Examples: