mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature Adds opt-in support for using the standalone TensorRT-RTX ABI Execution Provider during ModelOpt ONNX quantization. Users select the ABI backend with: `--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi` When selected, ModelOpt imports and registers the installed TensorRT-RTX ABI provider before creating the ONNX Runtime inference session. The backend selection is propagated through INT8, FP8, and INT4 AWQ calibration paths, including the Windows GenAI LLM quantization example. The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward compatible. The `legacy` backend is still the default and continues to use TensorRT-RTX libraries supplied through `PATH`. For Windows x64 with Python 3.11 or newer, the ONNX dependencies now include: - `onnxruntime-gpu~=1.26.0` - `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0` Keeping `onnxruntime-gpu` allows users to select either CUDA EP or TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build instructions are intentionally out of scope and will be documented separately. ### Usage ```powershell python -m modelopt.onnx.quantization ` --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" ` --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" ` --quantize_mode=int8 ` --output_path="C:\path\to\int8_abi\model.onnx" ` --calibration_eps=NvTensorRtRtx ` --trt_rtx_backend=abi ` --use_external_data_format ` --high_precision_dtype=fp32 ` --log_level=INFO ### Testing unit test have been added ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: pending <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64. - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default. - Added validation for unsupported backends and incompatible TensorRT plugin configurations. - Updated Windows ARM64 installation support and platform-specific package configuration. - **Documentation** - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps. - Documented the new TensorRT-RTX backend command-line option. - **Tests** - Added coverage for ABI provider registration, backend validation, and compatibility checks. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>