12 Commits
Author SHA1 Message Date
vishalpandya1990 87c9f8cf83 Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host (#2022)
### What does this PR do?

Type of change: Documentation update

- Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host -
mention about compatible onnxruntim-gpu and cupy-cuda13x packages.

### Testing

- Windows's onnx_ptq\genai_llm INT4 PTQ example with a 1B genai-cuda-ep
ONNX model + local doc building

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified Windows CUDA prerequisites for calibration and
GPU-accelerated quantization.
* Added setup guidance for CUDA 12 and CUDA 13.x, including compatible
packages and cuDNN requirements.
* Expanded installation verification steps for CUDA, ONNX Runtime, and
CuPy.
* Updated the GenAI LLM example with CUDA version compatibility
guidance.

* **Enhancements**
* Added runtime logging of detected CUDA environment paths and version
details during quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-07-28 11:42:47 +05:30
vishalpandya1990 769ea5f51c Pass USE_CUDA to compilation of cuda-ext to avoid failure on Windows (#1761)
### What does this PR do?

Type of change: Bug Fix

- Fixes a Windows-only failure when compiling ModelOpt CUDA extensions
(e.g. modelopt_cuda_ext_mx for NVFP4/MX quantization) with torch 2.9 +
CUDA 12.9:
`error C2872: 'std': ambiguous symbol
(torch/csrc/dynamo/compiled_autograd.h)
RuntimeError: Error building extension 'modelopt_cuda_ext_mx'`
- Compiled_autograd.h's IValuePacker<T>::packed_type() gates its
std-heavy branch on a preprocessor guard that changed between 2.8 and
2.9:
```cpp
// torch 2.8 — blanket Windows guard
#ifdef _WIN32
    TORCH_CHECK_NOT_IMPLEMENTED(false, "torch.compile not supported on Windows");
#else
    ... else if constexpr (::std::is_same_v<T, ::std::string>) ...   // never compiled on Windows
#endif

// torch 2.9 — now also requires USE_CUDA/USE_ROCM
#if defined(_WIN32) && (defined(USE_CUDA) || defined(USE_ROCM))
    TORCH_CHECK_NOT_IMPLEMENTED(false, "torch.compile not supported on Windows");
#else
    ... else if constexpr (::std::is_same_v<T, ::std::string>) ...   // line 1134: C2872 under nvcc+MSVC
#endif
```
- `load_cpp_extension` now adds -DUSE_CUDA=1 on Windows (os.name ==
"nt") so the header compiles. No op on Linux.
- Open PyTorch issue -
[pytorch/pytorch#148317](https://github.com/pytorch/pytorch/issues/148317)

### Usage

```python
# On Windows + torch 2.9/CUDA 12.9 this previously failed to build; now succeeds:
from modelopt.torch.quantization import extensions as e
print(e.get_cuda_ext_mx(raise_if_failed=True))
```

### Testing
- Reproduced and verified on Windows (RTX 5090 / sm_120, MSVC 2022
14.44, CUDA 12.9, torch 2.9.0+cu129)
- Before: get_cuda_ext_mx /
[quantize.py](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/diffusers/quantization/quantize.py)
--format fp4 (SD3.5) fail with C2872.
- After: extension compiles/loads and FP4 quantization proceeds and ONNX
export completes.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Bug Fixes**
* Fixed CUDA compilation issues on Windows systems. The extension
loading mechanism now automatically applies Windows-specific compiler
definitions, resolving previous build failures. Windows users no longer
need manual configuration workarounds, ensuring consistent, seamless
functionality across all supported platforms when using CUDA-enabled
extensions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-06-17 14:46:35 +05:30
vishalpandya1990 555be6c203 Add unit test for checking any leak of temporary augmented onnx files, on exception during ONNX INT4 AWQ quantization (#1383)
### What does this PR do?

Type of change: New unit test

- Adding a unit test for checking whether ONNX INT4 AWQ quantization, on
exception, leaks temporary augmented ONNX files it created during
quantization process.

- Monkey patches create-inference-session to raise / simulate exception.
Monkey patches temp-filep-creation utility to track temp files created
during quantization process.

- It is a follow up to fix made in to
https://github.com/NVIDIA/Model-Optimizer/pull/1359 -
[here](https://github.com/NVIDIA/Model-Optimizer/pull/1359#issuecomment-4349540073).

### Testing

- Local test run.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
* Added regression test to verify temporary files are properly cleaned
up during ONNX quantization operations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-05-07 14:39:08 +05:30
vishalpandya1990 a492fa9a14 Ensure removal of temp files on error in ONNX INT4 quantization (#1359)
### What does this PR do?

Type of change: Minor bug fix


- Put quantization steps inside try-finally to ensure removal of temp
files on error in ONNX INT4 quantization.
- To avoid redundancy between awq_lite() and awq_clip() methods, created
a utility _remove_augmented_onnx() for exception-handling based removal
of augmented onnx file and its data file.

### Testing

- Locally performed ONNX INT4 awq-lite and awq-clip quantization with
Llama 1B model.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Improved reliability of the quantization pipeline by ensuring
temporary conversion artifacts are always removed, making cleanup more
robust.
* Consolidated handling of external-data companions and added safer
deletion behavior that logs failures instead of raising errors.
* Ensured consistent session teardown and forced memory collection to
reduce resource leakage and intermittent errors during model conversion.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-04-30 09:22:36 +05:30
vishalpandya1990 20a46e04a7 Update ModelOpt-with-Olive documentation to mention CUDA EP commands (#1099)
### What does this PR do?

Type of change: Minor documentation update

- Update documentation (olive installation instructions) to mention
install commands for CUDA EP packages for ORT / ORT-genai.

### Testing

- Locally checked the readme, doc.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated Windows Olive installation guidance to recommend CUDA-based
ONNX Runtime packages instead of DirectML and added a link to ONNX
Runtime’s Execution-Provider docs for alternative EPs and requirements.
* Simplified Windows examples by removing explicit package install
commands and pointing users to the consolidated Olive setup
instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-03-24 17:54:15 +05:30
vishalpandya1990 4a848c4f52 Modelopt-windows documentation update (#812)
## What does this PR do?

Documentation

**Overview:**

- Update support matrix, changelog, deployment page, example readmes as
per recent feature and model support on Windows side.

## Testing
- No testing, its just documentation change

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ONNX Mixed Precision Weight-only quantization (INT4/INT8)
support.
  * Introduced diffusion-model quantization on Windows.
  * Added new accuracy benchmarks (Perplexity and KL-Divergence).
* Expanded deployment with multiple ONNX Runtime Execution Providers
(CUDA, DirectML, TensorRT-RTX).

* **Bug Fixes**
* Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ
quantization.

* **Documentation**
* Updated installation guides with system requirements and multiple
backend options.
* Reorganized deployment documentation with comprehensive execution
provider guidance.
* Expanded example workflows with improved setup instructions and
support matrices.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-01-28 10:48:59 +05:30
vishalpandya1990 0a66b37772 Add diffusion quantization guide for windows (#705)
## What does this PR do?

**Type of change:** New example for Windows

**Overview:** 

- Add a Windows' guide for quantization, ONNX export, and ORT-TRTRTX EP
run of diffusion models.

## Testing

- FP4 quantization and ORT-TRTRTX EP inference of SD3.5 Medium model is
done on Windows RTX 5090.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2025-12-18 14:31:39 +05:30
vishalpandya1990 78dd40918f Use ONNX DQ node instead of DQ custom-op for activation dequantization in nvfp4 (#536)
## What does this PR do?

**Type of change:** Minor update

**Overview:** 

- We are currently using trt::DequantizeLinear custom op for
activation's dequantization path in NVFP4 model.
- With this change, substituting trt::DequantizeLinear custom op with
DequantizeLinear (ONNX node) since official ONNX DQ node should already
support it.

## Testing

_System details:_ Windows 11 22621, RTX 5090, TRT 10.10.0.31
_Model:_ SD3.5-Medium's FP4 transformer ONNX model (with quant-mha)
_Command:_
`trtexec --builderOptimizationLevel=4 --onnx=<.onnx file path>
--minShapes=hidden_states:2x16x64x64,timestep:2,encoder_hidden_states:2x77x4096,pooled_projections:2x2048
--optShapes=hidden_states:16x16x64x64,timestep:16,encoder_hidden_states:16x154x4096,pooled_projections:16x2048
--maxShapes=hidden_states:16x16x128x128,timestep:16,encoder_hidden_states:16x333x4096,pooled_projections:16x2048
--stronglyTyped`

_With change (i.e. with ONNX DQ node for activation):_ Throughput:
12.2501 qps, GPU Compute Time: min = 65.9408 ms, max = 71.7161 ms, mean
= 67.0051 ms, median = 66.4348 ms, percentile(90%) = 69.2656 ms,
percentile(95%) = 70.3727 ms, percentile(99%) = 71.7161 ms

_Without change (i.e. with custom DQ node for activation):_ Throughput:
12.0729 qps, GPU Compute Time: min = 66.5288 ms, max = 80.5145 ms, mean
= 68.2194 ms, median = 67.043 ms, percentile(90%) = 71.2693 ms,
percentile(95%) = 72.0937 ms, percentile(99%) = 80.5145 ms

Attached the trtexec log for reference (both with and without change)

[trtexec_log_fp4_custom_op_A_removal_with_without_change.txt](https://github.com/user-attachments/files/23475935/trtexec_log_fp4_custom_op_A_removal_with_without_change.txt)


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: vipandya <vipandya@nvidia.com>
2025-11-19 11:29:21 +05:30
vishalpandya1990 9cd0824d42 Make all tensors on same device for svdquant with cpu-offloading (#550)
## What does this PR do?

**Type of change:** Bug Fix

**Overview:** ?

While running SVDQuant with cpu-offloading enabled using diffuser-ptq
example (sd3.5-medium model), error about "not all tensors on same
device" were observed at following steps:

1. awq-scale computation - get_scale() using x_max and w_max
2. loss update for each alpha - update_loss()
3. _apply_weight_pre_quant_scale() - while multiplying with
pre-quant-scale
4. apply_pre_quant_scale_and_smooth() - while multiplying with
pre-quant-scale

These errors should also be seen with flux model - with SVDQuant and
cpu-offloading enabled.

So, in this change, updating above places to ensure that concerned
tensors are on same device. Using ".to(device)" for this effect.

## Testing

- Tried SVDQuant with cpu-offloading enabled - with sd3.5-medium, on RTX
5090, Windows 11 22621. With this change, final ONNX model (transformer)
was produced without any error.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2025-11-17 16:02:05 +05:30
vishalpandya1990 1dfee8b12d Fix DQ1 output type error in DQ1->DQ2 for FP4 weights in NVFP4 model (#513)
## What does this PR do?

**Type of change:** Bug Fix

**Overview:**

- In post-processing after NVFP4 PTQ and ONNX Export, we convert FP4-QDQ
into DQ1->DQ2 for FP4 weights of the MatMuls. The output of DQ1 is of
the original weight-type (FP16 for FP16 base model) but its scale is in
FP32. There is a cast-to-fp16 after DQ2.
- In above setting, with FP16 base model weights, DQ1 has x_scale in
FP32 but its output type is set to FP16. This hybrid precision mode is
not allowed up to opset-21, and thereby it leads to error when run with
Onnxruntime.
- Note that such hybrid precision mode is allowed in opset-23+ but they
are not fully supported with onnxruntime EPs today, and even in future
we would want to support opset < 23 too.
- So, in this change, setting output of DQ1 to FP32 since its scale is
in FP32. There is already a cast-to-fp16 after DQ2 (before Gemm).

## Testing

- Checked with trtexec binary and onnxruntime-trt-rtx ep - using
sd3.5-medium model, on Windows RTX 5090.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2025-11-11 14:05:11 +05:30
vishalpandya1990 99c76fff26 Add SD3.5-medium quantization support in ModelOpt Diffusers example (#444)
Signed-off-by: vipandya <vipandya@nvidia.com>
2025-10-22 23:28:15 -07:00
vishalpandya1990 ee19a7ebd1 Add command-line options in Windows' llm-ptq example for Gather nodes' INT4 ONNX quantization (#418)
Signed-off-by: vipandya <vipandya@nvidia.com>
2025-10-10 11:15:05 +05:30