Files
haoxiz-nvidia acdf330414 Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do?

Type of change: new feature

Adds opt-in support for using the standalone TensorRT-RTX ABI Execution
Provider during ModelOpt ONNX quantization.

Users select the ABI backend with:

`--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi`

When selected, ModelOpt imports and registers the installed TensorRT-RTX
ABI provider before creating the ONNX Runtime inference session. The
backend selection is propagated through INT8, FP8, and INT4 AWQ
calibration paths, including the Windows GenAI LLM quantization example.

The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward
compatible. The `legacy` backend is still the default and continues to
use TensorRT-RTX libraries supplied through `PATH`.

For Windows x64 with Python 3.11 or newer, the ONNX dependencies now
include:

- `onnxruntime-gpu~=1.26.0`
- `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0`

Keeping `onnxruntime-gpu` allows users to select either CUDA EP or
TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build
instructions are intentionally out of scope and will be documented
separately.

### Usage

```powershell
python -m modelopt.onnx.quantization `
  --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" `
  --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" `
  --quantize_mode=int8 `
  --output_path="C:\path\to\int8_abi\model.onnx" `
  --calibration_eps=NvTensorRtRtx `
  --trt_rtx_backend=abi `
  --use_external_data_format `
  --high_precision_dtype=fp32 `
  --log_level=INFO

### Testing
unit test have been added

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ 
- Did you get Claude approval on this PR?: pending



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

- **New Features**
  - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64.
  - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default.
  - Added validation for unsupported backends and incompatible TensorRT plugin configurations.
  - Updated Windows ARM64 installation support and platform-specific package configuration.

- **Documentation**
  - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps.
  - Documented the new TensorRT-RTX backend command-line option.

- **Tests**
  - Added coverage for ABI provider registration, backend validation, and compatibility checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-09 04:04:46 +00:00
..