13 Commits
Author SHA1 Message Date
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Keval Morabia 1ccd945a51 Remove unused diffusers/cache_diffusion/pipeline and cuda-python dependency (#996)
`cuda-python` has mixed license and needs EStaff approval for usage. And
till 0.42, it was only used in
`examples/diffusers/cache_diffusion/pipeline` which has not been updated
in 9 months and not used anymore hence removing.

Also cherry-picked to `release/0.42.0` branch:
https://github.com/NVIDIA/Model-Optimizer/pull/984

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Removed TensorRT/ONNX deployment and inference tooling, related model
export/configuration, and runtime helpers from the cache-optimized
diffusion examples; removed the cuda-python example dependency.
* **Tests**
* Removed the example benchmarking script and its associated benchmark
test.
* **Documentation**
* Strengthened dependency-review, security, and PR guidance; updated PR
template and contributing documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-07 00:37:19 +05:30
Shengliang XuandKeval Morabia e0a6efbe70 fix trt engine building of the diffusers pipelines (#637)
## What does this PR do?

**Type of change:**

Bug fix

**Overview:**

1. The diffusion_trt.py needs the dynamic_shapes when running trtexec
for engine building. A previous change altered the format of
dynamic_shapes, fix it here.

2. the dynamic_shapes logic gets cleaned up. The existing logic is very
confusing

3. recover min-batch_size config for some pipelines. Previously some
pipelines set the min batch_size to be > 1, which was odd, so a previous
change sets them to be 1, but it turns out the oddity has a reason, the
trt engine building fails with the altered batch_size min/opt, thus
recover them.


## Testing

pytest tests/examples/diffusers

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-03 10:34:15 +00:00
jingyu-ml fa84955283 Fixed the cache diffusion ci/cd (#620)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Fixed the cache diffusion CI/CD issue related to Torch
2.9.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
pytest tests/examples/diffusers/test_cache_diffusion.py::test_sdxl_benchmarks -v -s
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2025-11-30 16:57:19 +05:30
jingyu-mlandKeval Morabia 46a9e49ab3 Fixed cache diffusion ci/cd (#426)
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-10-13 16:31:55 +05:30
Keval Morabia c0590b0255 Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-23 01:43:25 +05:30
Keval Morabia 1ef1d72a1b Code quality improvements - typos, formatting, etc.
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-02 19:58:33 +05:30
Keval Morabia 4d1eb0caf5 Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-30 10:54:28 +05:30
Keval Morabia cafa7f60ed Update for 0.33.0 release 2025-07-14 23:08:26 +05:30
Keval Morabia 7af33d29ce Update for 0.31.0 release 2025-06-05 13:24:07 -07:00
Keval Morabia 7a047435ae Update for 0.29.0 release 2025-05-08 23:43:51 +05:30
Keval Morabia 2017cd9063 Update for 0.25.0 release 2025-03-03 22:54:22 +05:30
Keval Morabia 73d6af785f Update 0.23.0 - OSS release 2025-01-29 01:45:47 +05:30