Files
Model-Optimizer/examples
Gwena Cunha 0593df0eb3 [6241485] Add support for ONNX Q/DQ node placement for DLA (#1661)
### What does this PR do?

**Type of change**: New feature

On DLA, the whole DLA-eligible region is compiled as one node, which
runs in INT8 or FP16, and it expects scales to be present throughout. A
tensor without a usable scale typically forces either that region to run
in FP16 or a GPU fallback (if enabled) — otherwise the build fails.

With IQ (implicit quantization) being deprecated in TensorRT, users are
migrating to ModelOpt for quantization/calibration. However, this breaks
the DLA workflow since DLA still only supports IQ. The suggested
workflow is then to:
1. Use ModelOpt to obtain the EQ (explicitly quantized) model;
2. Use [NVIDIA's Q/DQ Translator
Toolkit](https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/tree/main/tools/qdq-translator)
to obtain the `calib.cache` and `layer_arg.txt` files, which can be used
with the non-quantized model to generate a DLA loadable.

A [study on
Yolov5](https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes)
has shown that EQ can achieve perf parity with IQ on DLA if Q/DQ nodes
are inserted at every layer, making sure all tensors have INT8 scales.
From the study: _"With this option, all layers’ scales can be obtained
during model fine-tuning. However, this method may potentially disrupt
TensorRT fusion strategy with Q/DQ layers when running inference on GPU
and lead to higher latency on the GPU. For DLA, on the other hand, the
rule of thumb with PTQ scales is, “The more available scales, the lower
the latency.” "_

This PR aims to enable a quantization path targeting DLA.

### Usage

```python
$ python -m modelopt.onnx.quantization --onnx=model.onnx --target_dla
```

### Testing
- Two new parametrized tests (target_dla=False/True) cover both the
Conv/Mul quantization expansion and the GEMV (MatMul m=1) exclusion
bypass, with dedicated model builders.
- Internal test: 6241485@10

I ran the following experiments on various `timm` models:

| Exp | ModelOpt flag | QDQ-Translator flag |
|-------|-----------------------|------------------------------|
| 1      | `--high_precision_dtype=fp32` | default |
| 2      | `--high_precision_dtype=fp32 --target_dla` | default |
| 3 | `--high_precision_dtype=fp32` |
`--addtl_ops_to_infer_adjacent_scales` [1] |
| 4 | `--high_precision_dtype=fp32 --target_dla` |
`--addtl_ops_to_infer_adjacent_scales` [1] |
> [1] See https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/pull/35

Results (DOS Orin Linux with TRT 10.15.3.2):

| Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 |
|-------|------|----|------|----|
| resnet50 | 5.09 | 1.25 | 1.26 | 1.22 |
| mobilenetv2_100 | 4.07 | 3.78 | 0.80 | 0.77 |
| efficientnet_lite0 | 6.10 | 5.65 | 1.07 | 1.07 |
| inception_v3 | 11.43 | 1.57 | 1.56 | 1.57 |
| res2net50_14w_8s | 17.39 | 3.06 | 3.82 | 3.02 |

Observations: 
1. Exp 1 vs 2: `--target_dla` is essential to recover performance.
2. Exp 3 vs 4: `--target_dla` is necessary for perf parity or improved
perf compared to the default ModelOpt behavior. This is demonstrated in
the `res2net50_14w_8s`, which benefits from this new flag due to its
architecture containing 8 Convs operating on 14-channel tensors (below
the 16-channel minimum check in `int8.py /
find_nodes_from_convs_to_exclude()`.

Accuracy evaluation also shows no degradation for any of the
experiments. Top-1 with 1,000 ImageNet samples (%):

| Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 |
|-------|------|----|------|----|
| resnet50 | 75.6 | 76.0 | 75.1 | 76.0 |
| mobilenetv2_100 | 72.1 | 72.0 | 72.1 | 72.3 |
| efficientnet_lite0 | 75.1 | 75.2 | 75.1 | 75.2 |
| inception_v3 | 76.4 | 75.5 | 76.2 | 75.5 |
| res2net50_14w_8s | 75.5 | 75.6 | 75.8 | 75.6 |

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->
- Did you get Claude approval on this PR?: ❌ <!--- Run `/claude review`.
NVIDIA org members can self-trigger for complex changes; orthogonal to
CodeRabbit. -->

### Additional info
Related blogpost:
https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a `--target_dla` option for INT8 quantization to enable
optimized Q/DQ placement for DLA.

* **Behavior Changes**
* Adjusts quantization pre-processing rules when DLA targeting (or
autotune) is enabled, and defaults to quantizing all op types when none
are specified.

* **Examples**
* Added deterministic `--seed` for evaluation; enhanced ImageNet dataset
and calibration image loading/preprocessing (local or dataset-based).

* **Tests**
* Added coverage to verify Q/DQ placement differences for Conv and
MatMul with `target_dla`.

* **Documentation**
  * Updated the changelog to highlight the new DLA targeting option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-07-09 15:40:07 -04:00
..
2026-06-27 01:00:25 +05:30
2026-06-27 01:00:25 +05:30
2026-06-27 01:00:25 +05:30
2026-06-27 01:00:25 +05:30