mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
main
12
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
87c9f8cf83 |
Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host (#2022)
### What does this PR do? Type of change: Documentation update - Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host - mention about compatible onnxruntim-gpu and cupy-cuda13x packages. ### Testing - Windows's onnx_ptq\genai_llm INT4 PTQ example with a 1B genai-cuda-ep ONNX model + local doc building ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Clarified Windows CUDA prerequisites for calibration and GPU-accelerated quantization. * Added setup guidance for CUDA 12 and CUDA 13.x, including compatible packages and cuDNN requirements. * Expanded installation verification steps for CUDA, ONNX Runtime, and CuPy. * Updated the GenAI LLM example with CUDA version compatibility guidance. * **Enhancements** * Added runtime logging of detected CUDA environment paths and version details during quantization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
769ea5f51c |
Pass USE_CUDA to compilation of cuda-ext to avoid failure on Windows (#1761)
### What does this PR do?
Type of change: Bug Fix
- Fixes a Windows-only failure when compiling ModelOpt CUDA extensions
(e.g. modelopt_cuda_ext_mx for NVFP4/MX quantization) with torch 2.9 +
CUDA 12.9:
`error C2872: 'std': ambiguous symbol
(torch/csrc/dynamo/compiled_autograd.h)
RuntimeError: Error building extension 'modelopt_cuda_ext_mx'`
- Compiled_autograd.h's IValuePacker<T>::packed_type() gates its
std-heavy branch on a preprocessor guard that changed between 2.8 and
2.9:
```cpp
// torch 2.8 — blanket Windows guard
#ifdef _WIN32
TORCH_CHECK_NOT_IMPLEMENTED(false, "torch.compile not supported on Windows");
#else
... else if constexpr (::std::is_same_v<T, ::std::string>) ... // never compiled on Windows
#endif
// torch 2.9 — now also requires USE_CUDA/USE_ROCM
#if defined(_WIN32) && (defined(USE_CUDA) || defined(USE_ROCM))
TORCH_CHECK_NOT_IMPLEMENTED(false, "torch.compile not supported on Windows");
#else
... else if constexpr (::std::is_same_v<T, ::std::string>) ... // line 1134: C2872 under nvcc+MSVC
#endif
```
- `load_cpp_extension` now adds -DUSE_CUDA=1 on Windows (os.name ==
"nt") so the header compiles. No op on Linux.
- Open PyTorch issue -
[pytorch/pytorch#148317](https://github.com/pytorch/pytorch/issues/148317)
### Usage
```python
# On Windows + torch 2.9/CUDA 12.9 this previously failed to build; now succeeds:
from modelopt.torch.quantization import extensions as e
print(e.get_cuda_ext_mx(raise_if_failed=True))
```
### Testing
- Reproduced and verified on Windows (RTX 5090 / sm_120, MSVC 2022
14.44, CUDA 12.9, torch 2.9.0+cu129)
- Before: get_cuda_ext_mx /
[quantize.py](https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/diffusers/quantization/quantize.py)
--format fp4 (SD3.5) fail with C2872.
- After: extension compiles/loads and FP4 quantization proceeds and ONNX
export completes.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Bug Fixes**
* Fixed CUDA compilation issues on Windows systems. The extension
loading mechanism now automatically applies Windows-specific compiler
definitions, resolving previous build failures. Windows users no longer
need manual configuration workarounds, ensuring consistent, seamless
functionality across all supported platforms when using CUDA-enabled
extensions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: vipandya <vipandya@nvidia.com>
|
||
|
|
555be6c203 |
Add unit test for checking any leak of temporary augmented onnx files, on exception during ONNX INT4 AWQ quantization (#1383)
### What does this PR do? Type of change: New unit test - Adding a unit test for checking whether ONNX INT4 AWQ quantization, on exception, leaks temporary augmented ONNX files it created during quantization process. - Monkey patches create-inference-session to raise / simulate exception. Monkey patches temp-filep-creation utility to track temp files created during quantization process. - It is a follow up to fix made in to https://github.com/NVIDIA/Model-Optimizer/pull/1359 - [here](https://github.com/NVIDIA/Model-Optimizer/pull/1359#issuecomment-4349540073). ### Testing - Local test run. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Tests** * Added regression test to verify temporary files are properly cleaned up during ONNX quantization operations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
a492fa9a14 |
Ensure removal of temp files on error in ONNX INT4 quantization (#1359)
### What does this PR do? Type of change: Minor bug fix - Put quantization steps inside try-finally to ensure removal of temp files on error in ONNX INT4 quantization. - To avoid redundancy between awq_lite() and awq_clip() methods, created a utility _remove_augmented_onnx() for exception-handling based removal of augmented onnx file and its data file. ### Testing - Locally performed ONNX INT4 awq-lite and awq-clip quantization with Llama 1B model. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Improved reliability of the quantization pipeline by ensuring temporary conversion artifacts are always removed, making cleanup more robust. * Consolidated handling of external-data companions and added safer deletion behavior that logs failures instead of raising errors. * Ensured consistent session teardown and forced memory collection to reduce resource leakage and intermittent errors during model conversion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
20a46e04a7 |
Update ModelOpt-with-Olive documentation to mention CUDA EP commands (#1099)
### What does this PR do? Type of change: Minor documentation update - Update documentation (olive installation instructions) to mention install commands for CUDA EP packages for ORT / ORT-genai. ### Testing - Locally checked the readme, doc. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated Windows Olive installation guidance to recommend CUDA-based ONNX Runtime packages instead of DirectML and added a link to ONNX Runtime’s Execution-Provider docs for alternative EPs and requirements. * Simplified Windows examples by removing explicit package install commands and pointing users to the consolidated Olive setup instructions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
4a848c4f52 |
Modelopt-windows documentation update (#812)
## What does this PR do? Documentation **Overview:** - Update support matrix, changelog, deployment page, example readmes as per recent feature and model support on Windows side. ## Testing - No testing, its just documentation change ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ONNX Mixed Precision Weight-only quantization (INT4/INT8) support. * Introduced diffusion-model quantization on Windows. * Added new accuracy benchmarks (Perplexity and KL-Divergence). * Expanded deployment with multiple ONNX Runtime Execution Providers (CUDA, DirectML, TensorRT-RTX). * **Bug Fixes** * Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ quantization. * **Documentation** * Updated installation guides with system requirements and multiple backend options. * Reorganized deployment documentation with comprehensive execution provider guidance. * Expanded example workflows with improved setup instructions and support matrices. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
0a66b37772 |
Add diffusion quantization guide for windows (#705)
## What does this PR do? **Type of change:** New example for Windows **Overview:** - Add a Windows' guide for quantization, ONNX export, and ORT-TRTRTX EP run of diffusion models. ## Testing - FP4 quantization and ORT-TRTRTX EP inference of SD3.5 Medium model is done on Windows RTX 5090. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
78dd40918f |
Use ONNX DQ node instead of DQ custom-op for activation dequantization in nvfp4 (#536)
## What does this PR do? **Type of change:** Minor update **Overview:** - We are currently using trt::DequantizeLinear custom op for activation's dequantization path in NVFP4 model. - With this change, substituting trt::DequantizeLinear custom op with DequantizeLinear (ONNX node) since official ONNX DQ node should already support it. ## Testing _System details:_ Windows 11 22621, RTX 5090, TRT 10.10.0.31 _Model:_ SD3.5-Medium's FP4 transformer ONNX model (with quant-mha) _Command:_ `trtexec --builderOptimizationLevel=4 --onnx=<.onnx file path> --minShapes=hidden_states:2x16x64x64,timestep:2,encoder_hidden_states:2x77x4096,pooled_projections:2x2048 --optShapes=hidden_states:16x16x64x64,timestep:16,encoder_hidden_states:16x154x4096,pooled_projections:16x2048 --maxShapes=hidden_states:16x16x128x128,timestep:16,encoder_hidden_states:16x333x4096,pooled_projections:16x2048 --stronglyTyped` _With change (i.e. with ONNX DQ node for activation):_ Throughput: 12.2501 qps, GPU Compute Time: min = 65.9408 ms, max = 71.7161 ms, mean = 67.0051 ms, median = 66.4348 ms, percentile(90%) = 69.2656 ms, percentile(95%) = 70.3727 ms, percentile(99%) = 71.7161 ms _Without change (i.e. with custom DQ node for activation):_ Throughput: 12.0729 qps, GPU Compute Time: min = 66.5288 ms, max = 80.5145 ms, mean = 68.2194 ms, median = 67.043 ms, percentile(90%) = 71.2693 ms, percentile(95%) = 72.0937 ms, percentile(99%) = 80.5145 ms Attached the trtexec log for reference (both with and without change) [trtexec_log_fp4_custom_op_A_removal_with_without_change.txt](https://github.com/user-attachments/files/23475935/trtexec_log_fp4_custom_op_A_removal_with_without_change.txt) ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
9cd0824d42 |
Make all tensors on same device for svdquant with cpu-offloading (#550)
## What does this PR do? **Type of change:** Bug Fix **Overview:** ? While running SVDQuant with cpu-offloading enabled using diffuser-ptq example (sd3.5-medium model), error about "not all tensors on same device" were observed at following steps: 1. awq-scale computation - get_scale() using x_max and w_max 2. loss update for each alpha - update_loss() 3. _apply_weight_pre_quant_scale() - while multiplying with pre-quant-scale 4. apply_pre_quant_scale_and_smooth() - while multiplying with pre-quant-scale These errors should also be seen with flux model - with SVDQuant and cpu-offloading enabled. So, in this change, updating above places to ensure that concerned tensors are on same device. Using ".to(device)" for this effect. ## Testing - Tried SVDQuant with cpu-offloading enabled - with sd3.5-medium, on RTX 5090, Windows 11 22621. With this change, final ONNX model (transformer) was produced without any error. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
1dfee8b12d |
Fix DQ1 output type error in DQ1->DQ2 for FP4 weights in NVFP4 model (#513)
## What does this PR do? **Type of change:** Bug Fix **Overview:** - In post-processing after NVFP4 PTQ and ONNX Export, we convert FP4-QDQ into DQ1->DQ2 for FP4 weights of the MatMuls. The output of DQ1 is of the original weight-type (FP16 for FP16 base model) but its scale is in FP32. There is a cast-to-fp16 after DQ2. - In above setting, with FP16 base model weights, DQ1 has x_scale in FP32 but its output type is set to FP16. This hybrid precision mode is not allowed up to opset-21, and thereby it leads to error when run with Onnxruntime. - Note that such hybrid precision mode is allowed in opset-23+ but they are not fully supported with onnxruntime EPs today, and even in future we would want to support opset < 23 too. - So, in this change, setting output of DQ1 to FP32 since its scale is in FP32. There is already a cast-to-fp16 after DQ2 (before Gemm). ## Testing - Checked with trtexec binary and onnxruntime-trt-rtx ep - using sd3.5-medium model, on Windows RTX 5090. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
99c76fff26 |
Add SD3.5-medium quantization support in ModelOpt Diffusers example (#444)
Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
ee19a7ebd1 |
Add command-line options in Windows' llm-ptq example for Gather nodes' INT4 ONNX quantization (#418)
Signed-off-by: vipandya <vipandya@nvidia.com> |