mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? **Type of change**: New feature On DLA, the whole DLA-eligible region is compiled as one node, which runs in INT8 or FP16, and it expects scales to be present throughout. A tensor without a usable scale typically forces either that region to run in FP16 or a GPU fallback (if enabled) — otherwise the build fails. With IQ (implicit quantization) being deprecated in TensorRT, users are migrating to ModelOpt for quantization/calibration. However, this breaks the DLA workflow since DLA still only supports IQ. The suggested workflow is then to: 1. Use ModelOpt to obtain the EQ (explicitly quantized) model; 2. Use [NVIDIA's Q/DQ Translator Toolkit](https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/tree/main/tools/qdq-translator) to obtain the `calib.cache` and `layer_arg.txt` files, which can be used with the non-quantized model to generate a DLA loadable. A [study on Yolov5](https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes) has shown that EQ can achieve perf parity with IQ on DLA if Q/DQ nodes are inserted at every layer, making sure all tensors have INT8 scales. From the study: _"With this option, all layers’ scales can be obtained during model fine-tuning. However, this method may potentially disrupt TensorRT fusion strategy with Q/DQ layers when running inference on GPU and lead to higher latency on the GPU. For DLA, on the other hand, the rule of thumb with PTQ scales is, “The more available scales, the lower the latency.” "_ This PR aims to enable a quantization path targeting DLA. ### Usage ```python $ python -m modelopt.onnx.quantization --onnx=model.onnx --target_dla ``` ### Testing - Two new parametrized tests (target_dla=False/True) cover both the Conv/Mul quantization expansion and the GEMV (MatMul m=1) exclusion bypass, with dedicated model builders. - Internal test: 6241485@10 I ran the following experiments on various `timm` models: | Exp | ModelOpt flag | QDQ-Translator flag | |-------|-----------------------|------------------------------| | 1 | `--high_precision_dtype=fp32` | default | | 2 | `--high_precision_dtype=fp32 --target_dla` | default | | 3 | `--high_precision_dtype=fp32` | `--addtl_ops_to_infer_adjacent_scales` [1] | | 4 | `--high_precision_dtype=fp32 --target_dla` | `--addtl_ops_to_infer_adjacent_scales` [1] | > [1] See https://github.com/NVIDIA/Deep-Learning-Accelerator-SW/pull/35 Results (DOS Orin Linux with TRT 10.15.3.2): | Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 | |-------|------|----|------|----| | resnet50 | 5.09 | 1.25 | 1.26 | 1.22 | | mobilenetv2_100 | 4.07 | 3.78 | 0.80 | 0.77 | | efficientnet_lite0 | 6.10 | 5.65 | 1.07 | 1.07 | | inception_v3 | 11.43 | 1.57 | 1.56 | 1.57 | | res2net50_14w_8s | 17.39 | 3.06 | 3.82 | 3.02 | Observations: 1. Exp 1 vs 2: `--target_dla` is essential to recover performance. 2. Exp 3 vs 4: `--target_dla` is necessary for perf parity or improved perf compared to the default ModelOpt behavior. This is demonstrated in the `res2net50_14w_8s`, which benefits from this new flag due to its architecture containing 8 Convs operating on 14-channel tensors (below the 16-channel minimum check in `int8.py / find_nodes_from_convs_to_exclude()`. Accuracy evaluation also shows no degradation for any of the experiments. Top-1 with 1,000 ImageNet samples (%): | Model | Exp 1 | Exp 2 | Exp 3 | Exp 4 | |-------|------|----|------|----| | resnet50 | 75.6 | 76.0 | 75.1 | 76.0 | | mobilenetv2_100 | 72.1 | 72.0 | 72.1 | 72.3 | | efficientnet_lite0 | 75.1 | 75.2 | 75.1 | 75.2 | | inception_v3 | 76.4 | 75.5 | 76.2 | 75.5 | | res2net50_14w_8s | 75.5 | 75.6 | 75.8 | 75.6 | ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ❌ <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional info Related blogpost: https://developer.nvidia.com/blog/deploying-yolov5-on-nvidia-jetson-orin-with-cudla-quantization-aware-training-to-inference/#adding_qdq_nodes <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a `--target_dla` option for INT8 quantization to enable optimized Q/DQ placement for DLA. * **Behavior Changes** * Adjusts quantization pre-processing rules when DLA targeting (or autotune) is enabled, and defaults to quantizing all op types when none are specified. * **Examples** * Added deterministic `--seed` for evaluation; enhanced ImageNet dataset and calibration image loading/preprocessing (local or dataset-based). * **Tests** * Added coverage to verify Q/DQ placement differences for Conv and MatMul with `target_dla`. * **Documentation** * Updated the changelog to highlight the new DLA targeting option. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>