mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
[6008361][ONNX][Quantization] Clarify autotune guidance (#1989)
### What does this PR do? Type of change: documentation This PR clarifies when users should use ONNX quantization with Autotune enabled versus the direct Autotune entry point. - Adds a warning to the Autotune guide explaining that direct Autotune is a lower-level Q/DQ placement tool and does not replace calibrated ONNX PTQ. - Updates the ONNX quantization guide to show `autotune=True` in the Python API and explain that it uses default Autotune settings. - Updates ONNX PTQ example documentation to prefer `python -m modelopt.onnx.quantization ... --autotune=<mode>` for accuracy-sensitive PTQ from an unquantized model. - Updates the direct Autotune CLI help text to point users back to the full ONNX quantization workflow when calibration data and accuracy validation are required. ### Usage ```python N/A — documentation/help text change. ``` ### Testing - Ran `python -m py_compile modelopt/onnx/quantization/autotune/__main__.py`. - Built the Sphinx documentation with `python -m sphinx -b html docs/source docs/build/html`; build succeeded. Remaining warnings are from optional documentation imports and existing cross-reference labels. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified that Direct Autotune is an advanced tool for Q/DQ placement experiments, not a replacement for full calibrated ONNX quantization. * Added guidance on when to use Direct Autotune versus the end-to-end ONNX PTQ workflow. * Documented the optional `autotune=True` setting, expected calibration-time impact, and representative calibration data requirements. * Expanded links and guidance across ONNX PTQ examples and updated CLI help text with the recommended workflow. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gwenaelle Cunha Sergio <gcunhasergio@nvidia.com>
This commit is contained in:
@@ -2,6 +2,12 @@
|
||||
Autotune (ONNX)
|
||||
===============================================
|
||||
|
||||
.. warning::
|
||||
|
||||
Direct Autotune is an advanced ONNX quantization subtool for optimizing Q/DQ placement using TensorRT latency measurements. It does not replace the full calibrated ONNX quantization workflow.
|
||||
|
||||
To quantize an ONNX model taking into consideration accuracy, please use ``python -m modelopt.onnx.quantization ... --autotune=<quick|default|extensive>`` with representative calibration data. See the `ONNX quantization Autotune options <_onnx_quantization.html#python-m-modelopt.onnx.quantization-autotune-only-applicable-when-autotune-is-set>`_.
|
||||
|
||||
.. contents:: Table of Contents
|
||||
:local:
|
||||
:depth: 2
|
||||
@@ -22,9 +28,10 @@ The ``modelopt.onnx.quantization.autotune`` module automates Q/DQ (Quantize/Dequ
|
||||
|
||||
**When to Use This Tool:**
|
||||
|
||||
* Quantizing an ONNX model for TensorRT deployment
|
||||
* Optimizing Q/DQ placement for best performance
|
||||
* The model has repeating structures (e.g., transformer blocks, ResNet layers)
|
||||
* Debugging or developing the Q/DQ placement autotuning algorithm
|
||||
* Running the lower-level workflow without invoking the full quantization CLI
|
||||
* Programmatic experiments with direct Autotune classes and workflow functions
|
||||
* Expert workflows that intentionally start from already-quantized or pre-patterned Q/DQ models
|
||||
|
||||
Quick Start
|
||||
===========
|
||||
@@ -54,6 +61,8 @@ The command will:
|
||||
4. Select the best scheme based on TensorRT latency measurements
|
||||
5. Export an optimized ONNX model with Q/DQ nodes
|
||||
|
||||
Autotune searches for Q/DQ placement schemes that improve TensorRT runtime. It does not by itself define the full calibration and quantization policy for an accuracy-sensitive deployment. For end-to-end ONNX PTQ that starts from an unquantized model, run ONNX quantization with calibration data and enable ``--autotune`` there. See the `ONNX quantization Autotune options <_onnx_quantization.html#python-m-modelopt.onnx.quantization-autotune-only-applicable-when-autotune-is-set>`_.
|
||||
|
||||
**Output Files:**
|
||||
|
||||
Files are written under the output directory (default ``./autotuner_output``, or the path given by ``--output_dir``):
|
||||
|
||||
@@ -74,6 +74,16 @@ Call PTQ function
|
||||
quantize_mode="int8",
|
||||
)
|
||||
|
||||
Optionally enable Autotune for more optimized Q/DQ placement. Note that this will likely increase the time required to calibrate the model.
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
moq.quantize(
|
||||
...
|
||||
# Default Autotune settings, can be tuned with the autotune_* arguments below.
|
||||
autotune=True,
|
||||
)
|
||||
|
||||
Alternatively, you can call PTQ function in command line:
|
||||
|
||||
.. argparse::
|
||||
|
||||
@@ -210,7 +210,6 @@ trtexec --onnx=/tmp/identity_neural_network.quant.onnx \
|
||||
### Optimize Q/DQ node placement with Autotune
|
||||
|
||||
This feature automates Q/DQ (Quantize/Dequantize) node placement optimization for ONNX models using TensorRT performance measurements.
|
||||
For more information on the standalone toolkit, please refer to [autotune](./autotune).
|
||||
|
||||
To access this feature in the ONNX quantization workflow, simply add `--autotune` in your CLI:
|
||||
|
||||
@@ -224,7 +223,7 @@ python -m modelopt.onnx.quantization \
|
||||
--autotune=<quick,default,extensive>
|
||||
```
|
||||
|
||||
For more fine-tuned Autotune flags, please refer to the [API guide](https://nvidia.github.io/Model-Optimizer/guides/_onnx_quantization.html).
|
||||
For more fine-tuned Autotune flags, please refer to the [API guide](https://nvidia.github.io/Model-Optimizer/guides/_onnx_quantization.html) and the [Autotune guide](https://nvidia.github.io/Model-Optimizer/guides/9_autotune.html).
|
||||
|
||||
## Resources
|
||||
|
||||
|
||||
@@ -2,6 +2,8 @@
|
||||
|
||||
This example demonstrates automated Q/DQ (Quantize/Dequantize) node placement optimization for ONNX models using TensorRT performance measurements.
|
||||
|
||||
> **Warning:** This example uses the direct Autotune entry point for lower-level Q/DQ placement experiments. If you are starting from an unquantized ONNX model and care about **accuracy**, please use the ONNX PTQ workflow with `--autotune` enabled and with representative calibration data. See [../README#optimize-qdq-node-placement-with-autotune](../README#optimize-qdq-node-placement-with-autotune).
|
||||
|
||||
## Table of Contents
|
||||
|
||||
<div align="center">
|
||||
|
||||
@@ -154,7 +154,12 @@ def get_parser() -> argparse.ArgumentParser:
|
||||
"""
|
||||
parser = argparse.ArgumentParser(
|
||||
prog="modelopt.onnx.quantization.autotune",
|
||||
description="ONNX Q/DQ Autotuning with TensorRT",
|
||||
description=(
|
||||
"ONNX Q/DQ placement autotuning with TensorRT latency measurements. "
|
||||
"For end-to-end ONNX quantization taking into consideration accuracy, use "
|
||||
"`python -m modelopt.onnx.quantization ... --autotune=<mode>` "
|
||||
"with representative calibration data."
|
||||
),
|
||||
formatter_class=argparse.RawDescriptionHelpFormatter,
|
||||
epilog="""
|
||||
Examples:
|
||||
|
||||
Reference in New Issue
Block a user