44 Commits
Author SHA1 Message Date
haoxiz-nvidia acdf330414 Add TensorRT-RTX ABI EP support for ONNX quantization (#2262)
### What does this PR do?

Type of change: new feature

Adds opt-in support for using the standalone TensorRT-RTX ABI Execution
Provider during ModelOpt ONNX quantization.

Users select the ABI backend with:

`--calibration_eps=NvTensorRtRtx --trt_rtx_backend=abi`

When selected, ModelOpt imports and registers the installed TensorRT-RTX
ABI provider before creating the ONNX Runtime inference session. The
backend selection is propagated through INT8, FP8, and INT4 AWQ
calibration paths, including the Windows GenAI LLM quantization example.

The existing `--calibration_eps=NvTensorRtRtx` behavior remains backward
compatible. The `legacy` backend is still the default and continues to
use TensorRT-RTX libraries supplied through `PATH`.

For Windows x64 with Python 3.11 or newer, the ONNX dependencies now
include:

- `onnxruntime-gpu~=1.26.0`
- `onnxruntime-ep-nv-tensorrt-rtx-cu13==0.4.0`

Keeping `onnxruntime-gpu` allows users to select either CUDA EP or
TensorRT-RTX ABI EP for calibration. Windows-on-Arm source-build
instructions are intentionally out of scope and will be documented
separately.

### Usage

```powershell
python -m modelopt.onnx.quantization `
  --onnx_path="C:\path\to\Llama-3.2-3B-Instruct\model.onnx" `
  --model_id="C:\path\to\Llama-3.2-3B-Instruct\config.json" `
  --quantize_mode=int8 `
  --output_path="C:\path\to\int8_abi\model.onnx" `
  --calibration_eps=NvTensorRtRtx `
  --trt_rtx_backend=abi `
  --use_external_data_format `
  --high_precision_dtype=fp32 `
  --log_level=INFO

### Testing
unit test have been added

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ 
- Did you get Claude approval on this PR?: pending



<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

- **New Features**
  - Added optional TensorRT-RTX ABI backend support for ONNX calibration on Windows ARM64.
  - Added `legacy` and `abi` backend selection to quantization APIs and command-line tools; `legacy` remains the default.
  - Added validation for unsupported backends and incompatible TensorRT plugin configurations.
  - Updated Windows ARM64 installation support and platform-specific package configuration.

- **Documentation**
  - Updated Windows installation guidance, Python compatibility requirements, ARM64 setup, and verification steps.
  - Documented the new TensorRT-RTX backend command-line option.

- **Tests**
  - Added coverage for ABI provider registration, backend validation, and compatibility checks.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-09 04:04:46 +00:00
haoxiz-nvidia 4773f72f8a Docs: Add WOA documentation (#2264)
### What does this PR do?

Add WoA env setup guide. Includes build instruction of pyarrow, which
used by datatsets

### Usage

N/A

### Testing
N/A

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: N/A 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Added comprehensive Windows on Arm installation guidance, including
prerequisites, environment setup, dependency installation, verification,
and troubleshooting.
- Documented experimental Windows ARM64 support, supported quantization
formats, native dependency requirements, and Support Matrix details.
  - Expanded supported Windows Python versions through 3.13.
- Expanded TensorRT-RTX guidance for calibration, deployment, provider
setup, and standalone plugin usage.
- Clarified PyArrow requirements and local build instructions for
Windows ARM64.
- Added links to dedicated Windows on Arm installation resources and
shared TensorRT-RTX documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Haoxi Zhang <haoxiz@nvidia.com>
2026-09-08 15:04:04 +05:30
Keval MorabiaandClaude Opus 5 6a2ae5a25b Fix the llm_eval timeout: reachable MMLU mirror + no pipe deadlock (#2270)
### What does this PR do?

Type of change: Bug fix

`tests/examples/llm_eval/test_llm_eval.py::test_qwen3_eval_fp8` has been
failing with `Failed: Timeout (>900.0s) from pytest-timeout` on
unrelated branches (runs 33027056246 and the one for `ad83a428`, while
the 2026-08-24 nightly passed). It is not the test being slow — it is
the harness deadlocking, and the deadlock also destroys the diagnostics
that would explain the underlying kill.

**Mechanism.** The traceback shows `self = <Popen: returncode: -9 args:
['scripts/huggingface_example.sh', ...]>` while still blocked in
`stdout.read()`. The launcher was SIGKILLed (nothing in pytest sends
SIGKILL — pytest-timeout raises in the main thread, and the test's
`finally` `pkill` sends SIGTERM and only runs afterwards — so an OOM
kill is the likely source). But `subprocess.run(..., stdout=PIPE,
stderr=STDOUT)` waits for **EOF on the pipe**, not for the process, and
a surviving grandchild (the TRT-LLM serve/build worker) still holds the
write end. EOF never arrives, so the test blocks until the 900 s alarm.
Because the pipe is never drained, **every line of child output is
discarded**, which is why the CI log says nothing about what the script
was doing when it died.

Reduced to a self-contained reproducer:

```python
script = "sleep 300 & echo 'launcher output'; sleep 0.3; kill -9 $$"
subprocess.run(["bash", "-c", script], stdout=PIPE, stderr=STDOUT, text=True, timeout=20)
# -> TimeoutExpired: still blocked in communicate() after 20.0s, output lost
```

**Fix.** `_run_capturing` now starts the command in its own session,
drains its output on a reader thread (so logs stream as they arrive
instead of being buffered until the end), waits on the *process*, and
kills the process group if descendants still hold the pipe after a 30 s
grace period. A killed launcher now fails in seconds with its logs
intact instead of silently burning the test's whole timeout.

This does not fix whatever kills the script; it makes it diagnosable.
Worth noting separately: `test_qwen3_eval_fp8` took **749.10 s against
its 900 s mark** on the last green nightly, so it is fragile regardless
and may want its work trimmed or its budget raised once the logs show
where the time goes.

### Usage

```python
# unchanged public API
run_example_command(cmd_parts, example_path="llm_eval")
```

### Testing

Verified against the reproducer above and on the normal paths:

| scenario | before | after |
| --- | --- | --- |
| launcher SIGKILLed, survivor holds the pipe | blocks indefinitely (900
s in CI) | `rc=-9` in 3.5 s, `'launcher output'` captured |
| the surviving descendant | keeps running | killed with the process
group (stopped ticking, 20 -> 20 bytes) |
| normal exit | ok | `rc=0`, stdout and stderr interleaved in order |
| non-zero exit | ok | `rc=3`, output captured |

The example-test suites that use this helper run through the same code
path; `tests/examples/megatron_bridge` (16 passed, 1 skipped) exercised
it on nemo:26.08 in the branch this was extracted from.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `_run_capturing` keeps its
`(returncode, output)` contract; only the buffering strategy changed.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌ — this is test
infrastructure; the scenario needs a process that outlives a SIGKILLed
parent, which is awkward to assert in CI. Verified manually with the
reproducer above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — test-infrastructure fix.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved command execution reliability with real-time output capture.
* Ensured lingering child processes are cleaned up after commands exit
or are interrupted.
  * Added warnings when forced cleanup may truncate output.
  * Prevented hangs when descendant processes keep output streams open.

* **Documentation**
* Updated MMLU setup instructions to use the Hugging Face dataset
repository.
  * Improved Windows instructions by explicitly using `curl.exe`.

* **Examples**
* Improved MMLU downloads with retries, separate timeouts, resume
support, and automatic temporary-file cleanup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---

### Update: the second half of the failure

With the streaming fix in place, the next CI run showed the *actual*
cause, which the old code had been hiding. That run blocked in
`process.wait()` with `<Popen: returncode: None ...>` — child alive and
working, not the previous dead-child pipe deadlock — and the now-visible
output was:

```
--2026-08-27 18:09:20--  (try: 4)  https://people.eecs.berkeley.edu/~hendrycks/data.tar
Connecting to people.eecs.berkeley.edu ...|128.32.139.28|:443... failed: Connection timed out.
Retrying.
--2026-08-27 18:11:40--  (try: 5)  ...
```

`huggingface_example.sh` downloads the MMLU tarball from
`people.eecs.berkeley.edu`, that host stopped answering around
2026-08-25, and wget's default retry policy (20 tries, ~2 min per
connect timeout) consumed the whole 900 s budget. Not runner-specific:
the URL also times out from a developer workstation, and the nightlies
flipped 08-24 ✅ / 08-25 ✅ / **08-26 ❌ / 08-27 ❌**, matching the outage.

So this PR now carries both halves of the same failure:

1. the harness no longer deadlocks and no longer swallows the logs
(`985809cc2d`), and
2. the MMLU data comes from HuggingFace's copy of the same tarball, with
bounded retries (`40f1d89154`).

The mirror is byte-for-byte the same dataset in the same layout the
script already expects — verified by running the exact download/extract
commands:

```
https://huggingface.co/datasets/cais/mmlu/resolve/main/data.tar  ->  HTTP 200, 166 MB
data/mmlu/{dev,test,val}/  ->  57 subject CSVs each, plus auxiliary_train/
```

`wget --timeout=20 --tries=3` plus an explicit error means the next
dataset-host outage fails in about a minute with "Could not download the
MMLU test data. Set MMLU_DATA_PATH to a local copy." instead of silently
eating a test's timeout. The same URL is updated in
`examples/llm_eval/README.md` so a manual run does not hit the dead host
either.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 21:38:25 +05:30
vishalpandya1990 87c9f8cf83 Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host (#2022)
### What does this PR do?

Type of change: Documentation update

- Update documentation guide for ONNX INT4 PTQ on Windows cuda13 host -
mention about compatible onnxruntim-gpu and cupy-cuda13x packages.

### Testing

- Windows's onnx_ptq\genai_llm INT4 PTQ example with a 1B genai-cuda-ep
ONNX model + local doc building

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Clarified Windows CUDA prerequisites for calibration and
GPU-accelerated quantization.
* Added setup guidance for CUDA 12 and CUDA 13.x, including compatible
packages and cuDNN requirements.
* Expanded installation verification steps for CUDA, ONNX Runtime, and
CuPy.
* Updated the GenAI LLM example with CUDA version compatibility
guidance.

* **Enhancements**
* Added runtime logging of detected CUDA environment paths and version
details during quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-07-28 11:42:47 +05:30
haoxiz-nvidia cba8a5c62a Add: support input_shape_profile for trt-rtx ep (#1782)
### What does this PR do?

Add support for onnx quantization and support model_id as input, which
fix missing input_shpae_profile problem for some version of trt-rtx

### Usage

```python
python -m modelopt.onnx.quantization --onnx_path="path\to\model.onnx" --quantize_mode=int8 --output_path="path\to\output\model.onnx" --calibration_eps=NvTensorRtRtx --use_external_data_format --high_precision_dtype=fp32 --model_id="huggingface_model_id"
```

### Testing
Tested on 4 popular llm models on all popular quantization method(int4,
fp8, int8)

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌
- Did you get Claude approval on this PR?:  N/A



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added `model_id` and `trust_remote_code` support to the ONNX PTQ
CLI/API to enable automatic `input_shapes_profile` generation.
* Added `input_shapes_profile` input parsing from inline JSON or a JSON
file, with validation.

* **Enhancements**
* Threaded `input_shapes_profile` through INT8/FP8 quantization and
MatMul/MHA exclusion logic, including realignment after
calibration-provider changes.
* Updated ORT execution-provider wiring to apply per-provider shape
profiles/options during session setup.

* **Bug Fixes**
  * Improved Windows behavior for TensorRT provider setup.

* **Tests**
* Added unit and CLI integration coverage for parsing, forwarding, EP
filtering, and realignment behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: haoxiz <haoxiz@nvidia.com>
Signed-off-by: haoxiz-nvidia <45587794+haoxiz-nvidia@users.noreply.github.com>
2026-07-16 07:05:57 +00:00
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Hrishith Thadicherla fe8c5178c7 Removed version fixes for torch transformers in windows ptq example requirements (#1275)
### What does this PR do?

Type of change: Bug fix

Removed version fixes for torch and transformers



### Testing
Tested quantization with a couple of models . Working as expected.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Relaxed dependency specs: removed strict pin for torch to allow latest
compatible installs, and constrained transformers to <5.0.0 for broader
compatibility and easier updates.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
Signed-off-by: Hrishith Thadicherla <99313418+hthadicherla@users.noreply.github.com>
2026-04-17 09:16:48 +05:30
Keval MorabiaandClaude Sonnet 4.6 80a77d1cc0 Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary

LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party
dependency developed and provided by Lightricks. It is governed by the
[LTX Community License
Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE),
**not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per
legal guidance, all integration points must clearly surface this to
users.

- Add `[!WARNING]` license notice blocks at the top of all LTX-2-related
READMEs (`examples/diffusers`, `examples/diffusers/distillation`,
`examples/windows/diffusers/qad_example`)
- Add `warnings.warn(UserWarning)` at every LTX package import site in
Python files, covering both top-level and lazy imports:
  - `examples/diffusers/distillation/distillation_trainer.py`
  - `examples/diffusers/quantization/calibration.py`
  - `examples/diffusers/quantization/pipeline_manager.py`
-
`examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py`
  - `modelopt/torch/export/diffusers_utils.py`
  - `modelopt/torch/quantization/plugins/diffusion/ltx2.py`
- Add license notice comment to `requirements.txt` files that list LTX
packages, so the obligation is visible at install time
- Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering
variants like `requirements-dev.txt`) are owned by
`@NVIDIA/modelopt-setup-codeowners` regardless of location, via a
last-match-wins rule

**Design notes:**
- For library files (`diffusers_utils.py`, `ltx2.py`), the warning is
placed at the lazy import site inside functions — it fires only when
LTX-2 code paths are actually invoked, not at module import time, to
avoid polluting non-LTX users
- For example entry-point scripts that are LTX-2-only, the warning fires
at module load time (after all imports, to satisfy ruff E402)

## Test plan

- [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy,
markdownlint, bandit all clean)
- [ ] Verify warning appears at runtime when running an LTX-2
quantization or distillation example
- [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the
warning

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added third-party license notices across documentation and
requirements files clarifying LTX-2 packages are governed by the LTX
Community License Agreement rather than NVIDIA Model Optimizer's Apache
2.0 license.

* **Chores**
  * Updated code ownership configuration for requirements files.
* Added runtime warnings to notify when LTX-2 dependencies are accessed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 11:04:29 +05:30
Ajinkya RasaneandClaude Opus 4.6 3baa2da62e Upgrade ONNX from 1.19 to 1.21 (#1207)
### What does this PR do?

Type of change: new feature

Upgrade ONNX dependency from `~=1.19.0` to `~=1.21.0`. ONNX 1.20+
removed several deprecated
helper functions (`float32_to_bfloat16`, `float32_to_float8e4m3`,
`pack_float32_to_4bit`) that
`onnx_graphsurgeon` 0.5.x still references at import time. This PR adds
a compatibility shim
(`modelopt/onnx/_onnx_compat.py`) that restores these functions using
`ml_dtypes` before any
`onnx_graphsurgeon` import occurs. This supersedes the partial inline
fix from #1204 by also
handling `float32_to_float8e4m3`.

Changes:
- Bump `onnx~=1.19.0` to `onnx~=1.21.0` in `pyproject.toml`
- Add `modelopt/onnx/_onnx_compat.py` compatibility shim for removed
ONNX APIs
- Import shim in `modelopt/onnx/__init__.py` and
`tests/unit/onnx/conftest.py`
- Remove usage of removed `onnx.helper.pack_float32_to_4bit` in
`test_quant_utils.py`
- Update example requirements (`genai_llm`, `whisper`) to `onnx==1.21.0`

**TensorRT Compatibility:** TRT 10.16-GA supports opsets 9–24. ModelOpt
quantization modes
use opsets 19–23, all within range. ONNX 1.21 does not force opset 26.

### Usage

```python
# No API changes — the upgrade is transparent to users.
# The compatibility shim is applied automatically on import.
import modelopt.onnx
```

### Testing

- 469/470 ONNX unit tests pass inside
`nvcr.io/nvidia/tensorrt:25.06-py3` (1 pre-existing ORT
`CopyTensorAsync` EP issue, not ONNX-related)
- 6/6 `torch_onnx` integration tests pass (fp8, int8, nvfp4, mxfp8,
int4_awq, auto)
- ViT FP8 quantization via `torch_onnx` → TRT engine build → ImageNet
eval: **85.3% top-1, 97.8% top-5**
- ViT FP8 quantization via `onnx_ptq` → TRT engine build succeeds
- All pre-commit hooks pass (ruff, mypy, bandit, license headers)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅ (updated existing tests,
added conftest.py for compat shim)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ (dependency upgrade, no API change)

### Additional Information

Related: #1204 (partial fix for `float32_to_bfloat16` only — this PR
supersedes it with full coverage)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Dependencies**
* Removed unpinned ONNX from example requirement files and updated the
ONNX optional dependency to ~=1.21.0.

* **Refactor**
* Centralized an ONNX compatibility shim to restore missing helper APIs
when needed.

* **Tests**
* Added tests for the compatibility shim, adjusted quantization tests to
remove reliance on removed ONNX helpers, and ensured shim runs before
related tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 11:55:56 +05:30
Keval Morabia 04cd596d79 Add experimental support for transformers>=5.0 + min torch 2.8 (#975)
### What does this PR do?

- Add experimental support for transformers >=5.0 and remove deprecated
usages:
https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- ⚠️ For accelerate examples that used `--warmup-ratio: float`
(deprecated in 5.x), we now change it to `--warmup-steps: float | int`
which works as ratio if float but only for 5.x. For 4.x, it will error
out if float and prompt user to change back to `--warmup-ratio` or pass
an int absolute step count.
- ⚠️ Unified Hugging Face checkpoint export for quantized checkpoints
may not work for some models with transformers>=5.0 yet as it requires a
lot of fixes (e.g. change in how MoE experts are organized)
- ~Add Workaround for TRT-LLM's import of deprecated transformers
functions so trt-llm based gpu unit tests work fine. Still deployment
for models needs proper fixes directly in TRT-LLM hence llm/vlm ptq
example tests still run with transformers 4.57~
- Everything except PTQ and Export (mainly MoE) should work fine with
transformers>=5.0
- Bump min torch to 2.8 and enable 2.11 cicd testing
- NOTE: Upcoming Nemo:26.04 container comes with transformers 5.3

### Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] CI/CD tests passing
- [x] Manually tested unit tests, gpu tests with transformers 4.56 and
5.4
- [x] Manually tested example tests (except trt-llm container tests)
with transformers 4.56 and 5.4
- [x] 2-gpu nightly CICD tests manually triggered and passing: [gpu
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867257540),
[example
tests](https://github.com/NVIDIA/Model-Optimizer/actions/runs/23867260643)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Make remote-code usage opt-in via a configurable --trust_remote_code
flag across examples and tools.

* **Bug Fixes**
* Improve checkpoint/resume detection and related training guidance to
avoid erroneous errors.

* **Refactor**
* Consolidate dtype/config naming, switch warmup settings from ratio →
steps, and unify tokenizer invocation patterns.

* **Documentation**
  * Simplify changelog title and add misc notes for release 0.44.

* **Chores**
* Remove scheduled PR-branch cleanup workflow and relax/remove several
transformers version pins.

* **Tests**
* Adjust test gates, skips, and structures to align with updated deps
and behaviors.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-09 09:59:37 +05:30
ynankani-nv 4a70040ddd Ynankani/fvd benchmark (#1182)
### What does this PR do?

Type of change: ? New benchmark addition

<!-- Details about the change. -->
This PR adds the following to examples/windows/:
1) FVD (Fréchet Video Distance) evaluation tool — a standalone script
for computing FVD between two sets of videos using a pre-trained I3D
model (Kinetics-400, 1024-dim pooled features).
2) Directory reorganization — moved torch_onnx/ and qad_example/ under a
new diffusers/ folder for clearer grouping of diffusion model examples.
3) Minor fixes: Updated WikiText-2 dataset link in
perplexity_metrics/README.md (old wikitext → Salesforce/wikitext)
### Usage

```python
python compute_fvd.py \
  --ref-dir /path/to/reference/videos \
  --gen-dir /path/to/generated/videos \
  --output results.json \
  --device cuda
```

### Testing
Evaluated FVD tool on LTX-2.3 outputs comparing QAD vs PTQ checkpoints
(both NVFp4) against a BF16 baseline across 11 VBench dimensions.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added FVD (Fréchet Video Distance) evaluation: command‑line tool,
I3D‑based feature extraction, PCA option, and quantized comparisons (PTQ
vs QAD) with per‑category and overall results showing QAD's lower
average FVD.

* **Documentation**
* Expanded benchmarks TOC with “Additional Metrics”, added FVD
evaluation guide, usage examples, troubleshooting, and updated
perplexity dataset reference.

* **Chores**
  * Added Python requirements for the FVD benchmark.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: Yash Nankani <ynankani@nvidia.com>
2026-04-08 19:19:13 +00:00
Hrishith Thadicherla abf45580da Added workaround to onnx.helper function depreciated in onnx 1.21 (#1204)
### What does this PR do?

Type of change: Bug Fix

With onnx 1.21 and onnx graphsurgeon there is a incompatibility issue
where onnx discontinued the float32 to bfloat16 helper , thereby causing
an import failure while importing onnx graph surgeon which has that
function dependency.

```
Traceback (most recent call last):
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\Model-Optimizer\examples\windows\onnx_ptq\genai_llm\quantize.py", line 29, in <module>
    from modelopt.onnx.quantization.int4 import quantize as quantize_int4
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\Model-Optimizer\modelopt\onnx\__init__.py", line 36, in <module>
    from . import quantization
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\Model-Optimizer\modelopt\onnx\quantization\__init__.py", line 19, in <module>
    from .int4 import quantize as quantize_int4
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\Model-Optimizer\modelopt\onnx\quantization\int4.py", line 30, in <module>
    import onnx_graphsurgeon as gs
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\venv\Lib\site-packages\onnx_graphsurgeon\__init__.py", line 1, in <module>
    from onnx_graphsurgeon.exporters.onnx_exporter import export_onnx
  File "C:\Users\local-hthadicherla\Downloads\ModelOpt_Bench\venv\Lib\site-packages\onnx_graphsurgeon\exporters\onnx_exporter.py", line 134, in <module>
    np.uint16, onnx.helper.float32_to_bfloat16
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
AttributeError: module 'onnx.helper' has no attribute 'float32_to_bfloat16'
```

Added mldtypes converter and overrided the helper to use this workaround
function as onnx.helper.float32_to_bfloat16.

Long term fix is an AI for TensorRT team to fix this within onnx graph
surgeon i guess.

### Testing
Tested quantization works after fix with onnx 1.21.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Relaxed ONNX version constraints in example configurations to support
a wider range of compatible versions.

* **Bug Fixes**
* Added compatibility support for ONNX to ensure proper functionality
across different library versions.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-04-08 18:36:42 +05:30
Keval MorabiaandRinZ27 5dc17dfd15 [Security] Enable torch.load(weights_only=True) for secure checkpoint loading + trust_remote_code fix (#1181)
### What does this PR do?

- Add secure checkpoint loading support using
`torch.serialization.add_safe_globals([cls])`. This also removes 1
existing pickle usage.
- Remove hard-coded `trust_remote_code=True`
- Replaces https://github.com/NVIDIA/Model-Optimizer/pull/1056 by
@RinZ27

### Testing
<!-- Mention how have you tested your change if applicable. -->

CICD tests ran

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information

NVBug: 5999336

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added safe checkpoint save/load helpers and a --trust_remote_code CLI
flag in examples to control remote-code loading.

* **Bug Fixes**
* Checkpoint loading now defaults to safer, weights-only semantics to
reduce arbitrary-code exposure.

* **Documentation**
* CHANGELOG updated with security guidance and opt-in procedure for
unsafe checkpoint loading.

* **Tests**
  * New unit tests validating the safe-load behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: RinZ27 <222222878+RinZ27@users.noreply.github.com>
2026-04-08 00:36:28 +05:30
Shengliang Xu 1cceb950d6 [OMNIML-3689] PTQ quant_cfg semantic correction. Design in doc _quant_cfg.rst (#1094)
### What does this PR do?

#### Summary

Redesigns the `quant_cfg` configuration format in ModelOpt's PyTorch
quantization stack, replacing the previous dict-based format with an
**ordered list of typed `QuantizerCfgEntry` dicts**.

##### Motivation

The old `quant_cfg` dict had several pain points:
- **Ambiguous precedence**: no explicit way to reason about which entry
wins when multiple keys match a quantizer
- **Mixed key namespaces**: wildcard paths and PyTorch class names lived
in the same dict level, requiring ad-hoc dispatch
- **Magic `"default"` key**: an implicit, undocumented catch-all that
was easy to misuse
- **Poor composability**: merging two configs required dict updates that
silently discarded keys
- **No YAML round-trip fidelity**: the nested structure couldn't be
expressed cleanly in YAML

##### New format

`quant_cfg` is now an ordered list of `QuantizerCfgEntry` TypedDicts.
Each entry has:
- `quantizer_name` *(required)*: `fnmatch` wildcard matched against
quantizer module names
- `cfg` *(optional)*: dict (or list of dicts) of
`QuantizerAttributeConfig` fields
- `enable` *(optional)*: toggles quantizer on/off independently of `cfg`
- `parent_class` *(optional)*: restricts match to quantizers whose
parent module is of the given PyTorch class (e.g. `"nn.BatchNorm2d"`)

Entries are applied in list order; later entries override earlier ones.
The canonical pattern is deny-all first (`_base_disable_all`), then
selectively re-enable and configure, then apply standard exclusions
(`_default_disabled_quantizer_cfg`).

##### Changes

**Core library (`modelopt/torch/quantization/`)**

- **`config.py`**:
- Added `QuantizerCfgEntry` TypedDict (line 163) and
`find_quant_cfg_entry_by_path()` helper for exact-match lookup of
entries by path.
- Added `normalize_quant_cfg_list()` (line 1539) that converts legacy
formats (flat dict, single-key dicts, `nn.*`-scoped dicts, `"default"`
key) to canonical `QuantizerCfgEntry` lists. After normalization every
entry is guaranteed to have explicit `quantizer_name`, `enable`, and
`cfg` keys.
- Converted `_default_disabled_quantizer_cfg` and
`_mamba_moe_disabled_quantizer_cfg` from dicts to lists of
`QuantizerCfgEntry`.
- Added `_base_disable_all` (line 205): canonical deny-all entry
(`[{"quantizer_name": "*", "enable": False}]`).
- Converted all ~30 built-in config constants (`INT8_DEFAULT_CFG`,
`FP8_DEFAULT_CFG`, `NVFP4_DEFAULT_CFG`, etc.) to list format using
`*_base_disable_all` and `*_default_disabled_quantizer_cfg` unpacking.
- KV-cache configs (`FP8_KV_CFG`, `NVFP4_KV_CFG`, etc.) are now minimal
lists designed to be concatenated with a primary config — they
intentionally omit `_base_disable_all` and `"algorithm"`.
- Added two `QuantizeConfig` Pydantic field validators: a
`mode="before"` validator that calls `normalize_quant_cfg_list()`, and a
`mode="after"` validator that validates `cfg` dicts against
`QuantizerAttributeConfig`.
- Updated `need_calibration()` to iterate the normalized list instead of
the old dict.
- Changed `QuantizeQuantCfgType` alias from `dict[str | Callable, ...]`
to `list[QuantizerCfgEntry]`.

- **`conversion.py`**:
- Rewrote `set_quantizer_by_cfg()` (line 217) to iterate the list
directly. Each entry's `parent_class` is resolved via
`QuantModuleRegistry[parent_class_name]` (the existing `_DMRegistryCls`
registry).
- Added `set_quantizer_attributes_full()` (line 314): full replacement
of quantizer attributes from a `QuantizerAttributeConfig`. Unspecified
fields revert to defaults, enforcing entry atomicity. Can also upgrade
`TensorQuantizer` → `SequentialQuantizer` or downgrade the reverse.
- Added `set_quantizer_attributes_partial()` (line 384): merges a
partial `dict` of attributes into existing quantizer state. Does NOT
change quantizer structure. Used for enable-only entries.
- Added `set_quantizer_by_cfg_context()` context manager (line 447) that
temporarily applies a `quant_cfg` list and restores original quantizer
state on exit.
- Deprecated `set_quantizer_attribute()` (line 525) with a
`DeprecationWarning` pointing to the new functions.

- **`tensor_quantizer.py`**:
- `TensorQuantizer.set_from_attribute_config()`: narrowed type hint from
`dict` to `dict[str, Any]`.
- Added `_axis_setter` and `_block_sizes_setter` custom setters so that
`axis` and `block_sizes` changes properly propagate to the calibrator
and maintain mutual exclusivity.
- `SequentialQuantizer.set_from_attribute_config()`: narrowed signature
to `list[QuantizerAttributeConfig] | list[dict[str, Any]]` (removed the
old union with single values).

- **`algorithms.py`**:
- Updated `_match_quantizer_cfg()` to iterate the list and return
`(matched_cfg, matched_enable)` tuple with last-match-wins.
- Updated `_cfg_to_dict()`, `estimate_quant_compression()`, and
`QuantRecipe` to work with the list-based format.
- Updated `get_auto_quantize_config()` to emit list-format `quant_cfg`.

- **`model_quant.py`**: `disable_quantizer()` / `enable_quantizer()` now
call `set_quantizer_attributes_partial()` directly instead of the
deprecated `set_quantizer_attribute()`. Updated docstrings and code
examples to show the list format.

- **`utils/core_utils.py`**: `disable_lora_quantizers_in_config()` and
`update_quant_cfg_with_kv_cache_quant()` updated to append
`QuantizerCfgEntry` dicts to the list.

- **Other**: minor updates to `backends/fp8_per_tensor_gemm.py`,
`backends/nvfp4_gemm.py`, `compress.py`, `model_calib.py`,
`export/unified_export_hf.py`, and
`sparsity/attention_sparsity/conversion.py` to use the list format.

- **`onnx/llm_export_utils/quantization_utils.py`**: Updated
quantization config construction to use list format.

**YAML recipes (`modelopt_recipes/`)**

- Converted all 5 general PTQ recipes to the new list format:
  - `general/ptq/fp8_default-fp8_kv.yml`
  - `general/ptq/nvfp4_default-fp8_kv.yml`
  - `general/ptq/nvfp4_experts_only-fp8_kv.yml`
  - `general/ptq/nvfp4_mlp_only-fp8_kv.yml`
  - `general/ptq/nvfp4_omlp_only-fp8_kv.yml`
- Converted model-specific recipe:
`models/Step3.5-Flash/nvfp4-mlp-only.yaml`

**Documentation (`docs/`)**

- New guide: `docs/source/guides/_quant_cfg.rst` — comprehensive
reference covering entry format, ordering semantics, entry atomicity,
`enable` vs `cfg` independence, `parent_class` filtering, and common
patterns (deny-all-then-enable, customizing a built-in config, building
from scratch).
- Updated `_pytorch_quantization.rst` code examples to show the list
format with `copy.deepcopy` and `.append()`.
- Added `_quant_cfg.rst` to the quantization guide table of contents.

**Examples**

- Updated all quantization examples to use the list format:
`deepseek/ptq.py`, `diffusers/quantization/config.py`,
`llm_ptq/hf_ptq.py`, `llm_qat/main.py`, `vllm_serve/vllm_ptq_utils.py`,
`llm_autodeploy/run_auto_quantize.py`, `llm_eval/quantization_utils.py`,
`llm_ptq/example_utils.py`,
`windows/torch_onnx/diffusers/qad_example/sample_example_qad_diffusers.py`,
and 2 notebooks.

**Tests**

- New test file:
`tests/unit/torch/quantization/test_config_validation.py` — unit tests
for `need_calibration()`, `normalize_quant_cfg_list()` (new format,
legacy format conversions, error cases),
`find_quant_cfg_entry_by_path()`, `_match_quantizer_cfg()`, and
`QuantizeConfig` Pydantic validators.
- Extended `tests/unit/torch/quantization/test_quantize_cpu.py` with
tests for `set_quantizer_attributes_full()` (atomicity, parent_class
filtering, SequentialQuantizer creation), list ordering, enable-only
entry behavior, and end-to-end legacy dict format.
- Updated 20+ existing test files across `tests/unit/`, `tests/gpu/`,
`tests/gpu_megatron/`, and `tests/_test_utils/` to use the list format.

##### Backward compatibility

`normalize_quant_cfg_list()` is called automatically by the
`QuantizeConfig` Pydantic `mode="before"` validator, so existing code
passing the old dict-based format (flat dict like `{"*weight_quantizer":
{"num_bits": 8}}`, single-key dict lists, or `nn.*`-scoped dicts with
`parent_class` semantics) continues to work without modification. The
legacy `"default"` key is converted to `quantizer_name: "*"`.

`set_quantizer_attribute()` is preserved as a deprecated wrapper around
`set_quantizer_attributes_partial()`.

#### Test coverage

- **Unit tests**: new `test_config_validation.py` with tests for
normalization, validation, path lookup, and cfg matching. Extended
`test_quantize_cpu.py` with tests for full/partial attribute setting,
ordering, atomicity, and legacy backward compatibility.
- **System testing**:

```
python examples/llm_ptq/hf_ptq.py \
      --model Qwen/Qwen3-8B  \
      --recipe general/ptq/fp8_default-fp8_kv \
      --export_path=build/fp8_default-fp8_kv42  \
      --calib_size=16 \
      --batch_size=0 \
      --trust_remote_code \
      --export_fmt=hf
```

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-06 15:38:44 -07:00
Keval Morabia d6c8e9d2b6 Remove internal lustre paths (#1148)
Remove internal /lusre paths

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Configuration Updates**
* Set a concrete default container image to an NVIDIA PyTorch container
for templates.
* Removed built-in default values in scripts; key environment/config
variables must be provided externally.
* Replaced absolute example paths with placeholder relative paths and
TODO notes requiring user customization.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-01 17:53:20 +05:30
vishalpandya1990 20a46e04a7 Update ModelOpt-with-Olive documentation to mention CUDA EP commands (#1099)
### What does this PR do?

Type of change: Minor documentation update

- Update documentation (olive installation instructions) to mention
install commands for CUDA EP packages for ORT / ORT-genai.

### Testing

- Locally checked the readme, doc.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated Windows Olive installation guidance to recommend CUDA-based
ONNX Runtime packages instead of DirectML and added a link to ONNX
Runtime’s Execution-Provider docs for alternative EPs and requirements.
* Simplified Windows examples by removing explicit package install
commands and pointing users to the consolidated Olive setup
instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-03-24 17:54:15 +05:30
ynankani-nv 296a865c70 sample QAD example script (#933)
## What does this PR do?
sample QAD example script

**Type of change:** ?  new example
Example script for QAD on diffusion model like ltx-2

**Overview:** ?
1) Model loading 
2) NvFP4 fake quant PTQ using mtq.quantize
3) Distillation class wrapping using mtd.convert 
4) Using ltx-2 trainer code for training 
5) Checkpoint save in bf16 . post process bf16 model using Comfy-kitchen
to produce real quantized model for ComyUI inference

## Usage
<!-- You can potentially add a usage example below. -->

```python
accelerate launch --config_file fsdp_custom.yaml sample_example_qad_diffusers.py train  --config ltx2_qad.yaml 
```

## Testing
1) Tested improvement in Vbench score for PTQ and QAD checkpoint.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: NA <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive README documenting the new Windows QAD example:
setup, usage, project layout, and workflow.

* **New Features**
* Added a complete Quantization-Aware Distillation training example with
distributed training config, PTQ calibration, teacher-student
distillation, CLI for training/inference, and an inference-checkpoint
creation utility.
  * Added requirements file listing needed Python packages and tooling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-03-06 04:15:54 +00:00
Keval Morabia 860d0b4a70 Merge Linux and Windows Changelog (#954)
Since we no longer have any compiled packages, all releases are for all
platforms so we dont have separate windows releases hence merging
changelog as well.

Going forward, windows can test on latest linux version and if fixes
needed, they can go in next usual monthly linux release

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Reorganized changelog structure to consolidate Windows and Linux
release information in a unified view.
* Expanded Windows Support documentation across recent release versions.

* **Chores**
* Updated Windows example release badge to dynamically reflect the
latest PyPI release version.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 11:31:59 +05:30
Keval Morabia 82f1d216d1 Add Security and IP related contributing guide and configure coderabbit to catch such issues (#935)
### What does this PR do?

- Add Security related coding practices in `SECURITY.md` and merge with
`2_security.rst`
- Update `CONTRIBUTING.md` for instructions to follow if copying code
from other repositories
- Update PR template
- Cleanup dependency files
- New API `mto.load_modelopt_state` doing the insecure `torch.load(f,
weights_only=False)` instead of doing it separately everywhere. This
also allows us to later improve the input validation for
`modelopt_state_path` or use safer alternatives to `torch.load`

### Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=True)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ <!--- Mandatory -->
- Did you write any new necessary tests?: NA <!--- Mandatory for new
features or examples. -->
- Did you add or update any necessary documentation and update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
NA <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded and reorganized security guidance and contributor procedures;
updated PR template and several READMEs with clearer security,
submission, and installation instructions
* Replaced an older security document with an enhanced, centralized
security guidance

* **Chores**
* Adjusted example dependency lists and optional extras (adds, removals,
and version constraints)
* Enabled automated incremental reviews, added pre-merge security
checks, and introduced a knowledge-base of coding/security guidelines
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 03:49:15 +05:30
Hrishith Thadicherla fc6a211b4a Added column-major storage of weights and scales in INT4 quantization for model load time improvement in TRT-RTX (#811)
## What does this PR do?

**Type of change:** ? New feature

**Overview:** 
TensorRT-RTX requires the weights and scales in the ONNX models to be in
column-major format. So whenever the model loads TRT-RTX JIT transposes
the weights and scales during load time, causing increased load time.

Proposed feature is after quantization, transpose the weights and scales
in DQ node and add a transpose node right after i.e,
A × B = A × ((Bᵀ)ᵀ)

The transformation is post processing step and is disabled by default.
It can be enabled by quantizing with --use_column_major

## Usage
```
python -m modelopt.onnx.quantization --onnx_path "model.onnx" --output_path "model_quant.onnx" --quantize_mode int4 --calibration_method awq_lite --use_column_major --skip_shared_constants_duplication
```

## Testing
Tested a few LLM's and their MMLU scores with and without this
transformation. No degradations were observed.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added --use_column_major CLI flag to enable column-major weight
storage optimization (applies to DQ-only quantization paths).

* **Documentation**
  * CLI docs updated to describe the new flag and its applicability.

* **Tests**
* New unit tests validating column-major transformation behavior and
output equivalence.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-02-02 20:24:47 +05:30
vishalpandya1990 4a848c4f52 Modelopt-windows documentation update (#812)
## What does this PR do?

Documentation

**Overview:**

- Update support matrix, changelog, deployment page, example readmes as
per recent feature and model support on Windows side.

## Testing
- No testing, its just documentation change

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ONNX Mixed Precision Weight-only quantization (INT4/INT8)
support.
  * Introduced diffusion-model quantization on Windows.
  * Added new accuracy benchmarks (Perplexity and KL-Divergence).
* Expanded deployment with multiple ONNX Runtime Execution Providers
(CUDA, DirectML, TensorRT-RTX).

* **Bug Fixes**
* Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ
quantization.

* **Documentation**
* Updated installation guides with system requirements and multiple
backend options.
* Reorganized deployment documentation with comprehensive execution
provider guidance.
* Expanded example workflows with improved setup instructions and
support matrices.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-01-28 10:48:59 +05:30
ynankani-nv 5cc2a54519 Ynankani/update windows benchmark md (#762)
## What does this PR do? 
**Type of change:** ? documentation

**Overview:** Md update to add perplexity and kl divergence benchmark
info.




## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: NA
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Expanded accuracy comparison section with three detailed benchmark
metrics: MMLU scores, Perplexity (PPL), and KL-divergence.
* Added comprehensive tables showing results across models and
quantization configurations.
  * Included evaluation guides and references for each metric.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: unknown <ynankani@nvidia.com>
2026-01-25 06:23:25 +00:00
Hrishith Thadicherla 03dc3860af Update onnxruntime-gpu (#697)
## What does this PR do?

**Type of change:** Bug fix 

**Overview:** Updated setup.py to use only onnxruntime-gpu and removed
onnxruntime-directml as dependency.
Also changed onnxruntime-gpu version in examples.


## Testing
Tested int4 quantization and MMLU benchmark with updated onnxruntime-gpu
, working as expected

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2025-12-22 08:35:58 +00:00
vishalpandya1990 0a66b37772 Add diffusion quantization guide for windows (#705)
## What does this PR do?

**Type of change:** New example for Windows

**Overview:** 

- Add a Windows' guide for quantization, ONNX export, and ORT-TRTRTX EP
run of diffusion models.

## Testing

- FP4 quantization and ORT-TRTRTX EP inference of SD3.5 Medium model is
done on Windows RTX 5090.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2025-12-18 14:31:39 +05:30
Hrishith Thadicherla 01b0bc889e Updated dependencies in Windows examples (#622)
## What does this PR do?

**Type of change:** ? Bug Fix

**Overview:** Updated torch and transformers to latest versions for
normal quantization examples.
For whisper quantization, updated README with steps to install and
enable latest torch and torchaudio.

## Testing
Tested quantization and MMLU benchmarks with updated torch and
transformers version. It was working as expected.

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2025-12-15 15:52:32 +05:30
Keval Morabia 53a2ddebab Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer
(OMNIML-3033)
- [x] Mention in Latest News section with date on the date of merging
this PR (12/08)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-07 12:33:21 +05:30
ynankani-nv 2e19b5afe4 [4975376][5541172]perplexity and kl-divergence benchmark metrics (#411)
Signed-off-by: unknown <ynankani@nvidia.com>
2025-10-30 13:46:05 -07:00
Keval Morabia 90b1e68fcf Update HF nvidia collection links in docs (#475)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-28 22:23:01 +05:30
vishalpandya1990 ee19a7ebd1 Add command-line options in Windows' llm-ptq example for Gather nodes' INT4 ONNX quantization (#418)
Signed-off-by: vipandya <vipandya@nvidia.com>
2025-10-10 11:15:05 +05:30
Chenjie Luo 340eb7a756 Remove Qwen tokenizer modification (#390)
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2025-10-04 17:41:08 +05:30
ynankani-nv 8b1cedf2b7 [5506930]Add support in ModelOpt for generating mixed-precision (INT4… (#310)
Signed-off-by: unknown <ynankani@nvidia.com>
2025-09-23 12:09:32 +00:00
ajrasane 0adb1b6772 Upgrade to ONNX 1.19.0 (#289)
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-09-10 00:40:39 +00:00
Keval Morabia 1ef1d72a1b Code quality improvements - typos, formatting, etc.
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-02 19:58:33 +05:30
Keval Morabia 4d1eb0caf5 Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-30 10:54:28 +05:30
Keval Morabia dcdc2df084 Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-26 01:12:26 +05:30
Keval Morabia 37085354c9 Update files on Github (#258)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-17 02:44:16 +05:30
Keval Morabia 4c611e47a6 Update files on GitHub
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-07-31 22:48:09 +05:30
Keval Morabia cafa7f60ed Update for 0.33.0 release 2025-07-14 23:08:26 +05:30
Keval Morabia 7af33d29ce Update for 0.31.0 release 2025-06-05 13:24:07 -07:00
Keval Morabia 7a047435ae Update for 0.29.0 release 2025-05-08 23:43:51 +05:30
Keval Morabia e048fb239d Add changes for 0.27 Windows release 2025-04-29 22:11:52 +05:30
Keval Morabia 92f430f6ab Add files for 0.27.0 release 2025-04-03 10:31:41 +05:30
Keval Morabia 2017cd9063 Update for 0.25.0 release 2025-03-03 22:54:22 +05:30
Keval Morabia 73d6af785f Update 0.23.0 - OSS release 2025-01-29 01:45:47 +05:30