20 Commits
Author SHA1 Message Date
ZhiyuandClaude Opus 5 6261f854aa docs: rebuild the unified HF deployment support matrix from the deploy test suite (NVBug 6550792) (#2087)
### What does this PR do?

Type of change: documentation

Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) /
OMNIML-5693.

The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9
model families and **no VLMs**, while
`tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80
checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`,
`Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not
use the doc to scope testing, and users could not tell what is actually
covered.

Filing also surfaced that the matrix lived in **three places that had
drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed
Qwen3.5 MoE, and the skill reference had neither.

#### Changes

1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`**
from `test_deploy.py`, split into language models,
vision-language/multimodal, speculative decoding drafters, and
diffusion.

2. **Stated plainly what the matrix is and is not.** Review established
that the original "CI-validated" framing claimed more than the suite
substantiates, so a *What this matrix is based on* section now leads
with two limits:
- The suite is marked `release` and collects only under `--run-release`,
which **no workflow passes** — these are declared cases, not PR-gated
coverage.
- Each case is a **load-and-generate smoke check on the text path**: no
accuracy, no image/audio input, no diffusion output, no verification
that speculative decoding engages.

The legend follows from that: ✅ = declared in the suite, ⚠ = expected to
work but not a suite entry (or an entry that does not exercise the
feature the row names), `-` = not in the suite. Sections that would
otherwise over-read carry their own qualifiers — VLM rows are labelled
text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason
stated.

3. **Removed the two duplicate copies**, replacing them with links, so
there is one table to maintain.

4. **Fixed stale prose**: the deployment tabs still claimed FP8-only
support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025,
both contradicting the version table above them. The TRT-LLM floor moves
to v1.2.0, qualified as the oldest version stated rather than the oldest
that works.

5. **Dropped the Phi series** from the deployment matrix, following
#2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated.

### Usage

N/A — documentation only.

### Testing

- `docutils` parse of the modified `.rst`: no warnings or errors from
the new content; all 5 tables parse with every cell in the correct
column.
- Cell contents cross-checked against `test_deploy.py` by AST-parsing
the `ModelDeployerList(...)` calls rather than by eye; the scope caveats
were each verified against `tests/_test_utils/deploy_utils.py`.
- `pre-commit run --files …` passes; `build-docs` green.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation only
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

**Two known follow-ups, neither in scope here:**

1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.**
Consolidating to one copy removes the three-way drift but not the
doc-vs-test drift; a generator plus a CI check would close it.
2. **The release deployment suite does not run in CI.** Wiring it into
per-backend release CI is what would let ✅ mean "verified to pass"
rather than "declared". That needs GPU capacity across three backends
and should be tracked on its own.

**For the filer (@Kenny Kang):** the ✅ cells are the scope the release
deploy suite declares, and `test_deploy.py` carries the checkpoint, TP
size, and minimum SM version per entry — but please read the legend
first, since those cases are not currently executed by CI.

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 01:37:07 +05:30
Keval Morabia b6bf6b7997 Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-23 21:42:37 +05:30
jingyu-mlandClaude Opus 4.8 c458ad36f1 [2/2] Remove examples/diffusers/eval image-quality evaluation example (#1694)
> **Part 2 of 2** — removal for the 0.46 release. Depends on **Part 1**:
#1798 (the 0.45 deprecation).

### What does this PR do?

Type of change: Backward breaking change (removal of a deprecated
example)

Removes the `examples/diffusers/eval` image-quality evaluation example
(ImageReward / CLIP-IQA / CLIP metrics) and its references in
`examples/diffusers/README.md`. The example was **deprecated in 0.45**
(see #1798) and is removed here for **0.46**, per the [Deprecation
Policy](https://github.com/NVIDIA/Model-Optimizer#deprecation-policy)
(1-release migration before removal).

Scope is limited to the diffusers example; the unrelated
`examples/llm_ptq` "Evaluate Accuracy" section is untouched.

### Usage

N/A — removes example scripts; no library API changes.

### Testing

N/A — deletion of example scripts plus documentation cleanup. Verified
`examples/diffusers/README.md` has no remaining references to the
deleted `eval/` directory.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — removes the (previously
deprecated) `examples/diffusers/eval` example. No public `modelopt` API
is affected.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — removal only.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.46 Backward Breaking Changes.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Paired with #1798 (the 0.45 deprecation). This PR targets **0.46** and
should **not** carry the `cherry-pick-0.45.0` label (that belongs on
#1798).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Documentation**
* Updated the Diffusers “Model Optimizations” README by shortening the
overview, removing the “Evaluate Accuracy” guidance (links, inputs,
commands, and example results), and refining notes about per-subsection
`requirements.txt` usage.

* **Breaking Changes / Deprecations**
* Updated the 0.46 changelog to reflect that the Diffusers image-quality
evaluation example (ImageReward/CLIP-IQA/CLIP) is no longer maintained.

* **Chores**
* Removed the Diffusers evaluation example, including its evaluation
entrypoint, metrics, and shared utilities.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 08:44:59 +05:30
c78e654744 Skip Softmax diffusion export (#1269)
### What does this PR do?

Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.

- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.

### Usage

```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 --calib-size 4 \
    --initial-disabled-steps 5 \
    --export-dir ./wan22_skip_softmax_ckpt
```

Resulting layout — a `config.json` per component, **no `sparse.yaml`**:

```
wan22_skip_softmax_ckpt/
├── transformer/config.json        # carries sparse_attention_config
├── transformer_2/config.json      # carries sparse_attention_config
├── vae/ …  text_encoder/ …  tokenizer/ …  scheduler/ …
└── model_index.json
```

A representative `config.json` entry for a diffusion transformer:

```json
"sparse_attention_config": {
  "config_groups": {
    "group_0": {
      "algorithm": "skip_softmax",
      "targets": ["WanAttention"],
      "ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
      "initial_disabled_steps": 5,
      "threshold_scale_factor": {
        "formula": "a * exp(b * target_sparsity)",
        "prefill": {"a": 1443.49, "b": 4.30}
      },
      "target_sparsity": {"prefill": 0.5}
    }
  },
  "producer": {"name": "modelopt", "version": "0.45.0..."}
}
```

The N:M variant adds a second group:

```json
"group_1": {
  "algorithm": "sparse_softmax",
  "targets": ["WanAttention"],
  "sparsity_n": 2, "sparsity_m": 4,
  "dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```

### Testing

- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:02:41 -07:00
3ff15ccef3 Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do?

Type of change: ?  new feature

<!-- Details about the change. -->
Adds post-processing support for exported diffusion model checkpoints to
enable NVFP4 block scale swizzling and configurable padding strategies.
This allows exported quantized checkpoints to be directly consumed by
inference runtimes (e.g., ComfyUI with comfy_kitchen) that require
cuBLAS 2-D block-scaling-factors layout.
Changes:
1) Unified post-processing step (_postprocess_safetensors): Loads saved
safetensors files and applies merge, padding, swizzle, and quantization
metadata injection in a single pass.
2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales
from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled
layout per the cuBLAS specification.
3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale
tensors to multiples of 16, with "row" (rows only) or "row_col" (both
dimensions) strategies.
4) Standalone quantization metadata (build_layerwise_quant_metadata):
Extracted from merge_diffusion_checkpoint so _quantization_metadata can
be injected independently of merging — works for both merged (LTX-2) and
standalone (Flux2) exports.
5) Bug fix (conversion.py): Wrapped yield in try/finally in
set_quantizer_by_cfg_context so quantizer states are always restored,
fixing an issue when yield fails.
### Usage

```python
# LTX-2 export with merge + swizzle + padding
export_hf_checkpoint(
  pipeline,
  export_dir="./output",
  merged_base_safetensor_path="./ltx-2-22b-dev.safetensors",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
  enable_layerwise_quant_metadata=True,
)
# Flux2 standalone export with swizzle + padding (no merge needed)
export_hf_checkpoint(
  transformer,
  export_dir="./output",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
)
# Via quantize.py CLI
python quantize.py \
  --model ltx-2 --format fp4 \
  --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \
  --extra-param enable_swizzle_layout=true \
  --extra-param padding_strategy=row_col \
  --hf-ckpt-dir ./output
```

### Testing
1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base
checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn
scales in swizzled layout, and _quantization_metadata . Ran the
checkpoint with ComfyUI
2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has
correct uint8 weights, float8_e4m3fn scales in swizzled layout, and
_quantization_metadata . Ran the checkpoint with ComfyUI

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Diffusers export: optional NVFP4 support — swizzle layout, row/row_col
padding, and optional per-layer quantization metadata; exports are now
post-processed to apply these options.
* Export flow accepts new flags to enable swizzle, padding strategy, and
layerwise metadata.

* **Bug Fixes**
* Quantizer context manager now always restores state, including on
exceptions.

* **Tests**
* Added unit tests for NVFP4 padding, swizzling, metadata injection, and
post-processing.

* **Documentation**
  * README example updated to show swizzle and padding flags.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Signed-off-by: ynankani-nv <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
2026-05-22 10:36:59 +00:00
jingyu-ml c7966119eb Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do?

Type of change: re-org code <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ We changed the import path
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅
<!--- Only for new features, API changes, critical bug fixes or backward
incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Calibration support for skip-softmax multi-threshold measurement in
sparse attention.
  * N:M sparse softmax masking and helpers for sparsity-aware attention.

* **Chores**
* Reorganized and consolidated kernel/backends for quantization and
sparsity to a unified kernels layout, updating tests and examples to
match.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-22 23:34:56 +00:00
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-19 12:20:14 +05:30
jingyu-ml feec81ad2b Add the Skip softmax for diffusion (#1166)
### What does this PR do?

Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

## Summary

- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)

## Changes

### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.

### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.

### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.

### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.

### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.

### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.

## Threshold modes

| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |

## Usage

```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --raw-threshold -0.7 \
    --prompt "A cat playing piano" --output out.mp4

# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 \
    --prompt "A cat playing piano" --output out.mp4

# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --baseline \
    --prompt "A cat playing piano" --output baseline.mp4
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization

* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples

* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-18 06:54:19 +00:00
Keval MorabiaandClaude Sonnet 4.6 80a77d1cc0 Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary

LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party
dependency developed and provided by Lightricks. It is governed by the
[LTX Community License
Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE),
**not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per
legal guidance, all integration points must clearly surface this to
users.

- Add `[!WARNING]` license notice blocks at the top of all LTX-2-related
READMEs (`examples/diffusers`, `examples/diffusers/distillation`,
`examples/windows/diffusers/qad_example`)
- Add `warnings.warn(UserWarning)` at every LTX package import site in
Python files, covering both top-level and lazy imports:
  - `examples/diffusers/distillation/distillation_trainer.py`
  - `examples/diffusers/quantization/calibration.py`
  - `examples/diffusers/quantization/pipeline_manager.py`
-
`examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py`
  - `modelopt/torch/export/diffusers_utils.py`
  - `modelopt/torch/quantization/plugins/diffusion/ltx2.py`
- Add license notice comment to `requirements.txt` files that list LTX
packages, so the obligation is visible at install time
- Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering
variants like `requirements-dev.txt`) are owned by
`@NVIDIA/modelopt-setup-codeowners` regardless of location, via a
last-match-wins rule

**Design notes:**
- For library files (`diffusers_utils.py`, `ltx2.py`), the warning is
placed at the lazy import site inside functions — it fires only when
LTX-2 code paths are actually invoked, not at module import time, to
avoid polluting non-LTX users
- For example entry-point scripts that are LTX-2-only, the warning fires
at module load time (after all imports, to satisfy ruff E402)

## Test plan

- [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy,
markdownlint, bandit all clean)
- [ ] Verify warning appears at runtime when running an LTX-2
quantization or distillation example
- [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the
warning

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added third-party license notices across documentation and
requirements files clarifying LTX-2 packages are governed by the LTX
Community License Agreement rather than NVIDIA Model Optimizer's Apache
2.0 license.

* **Chores**
  * Updated code ownership configuration for requirements files.
* Added runtime warnings to notify when LTX-2 dependencies are accessed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 11:04:29 +05:30
Ajinkya Rasane 8f4c11aa44 Upgrade the TensorRT container version (#1112)
### What does this PR do?

Type of change: Container version update

- Upgraded the TensorRT container version to 26.02
- This supports TensorRT 10.15.1

### Testing
Unit and integrations tests pass

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example guides to reference the latest TensorRT Docker image
version (26.02)
  * Added compatibility note for onnxruntime-gpu usage

* **Chores**
* Updated CI/CD workflows to use the latest TensorRT Docker image
version

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-03-25 04:00:52 +05:30
jingyu-ml 37d3f10cbd To support LTX2 ComfyUI format (#972)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

- Added a flag merged_base_safetensor_path to the example code so that
user can export the ComfyUI style ckpt.

### Usage

```bash
python quantize.py \
    --model ltx-2 --format fp4 --batch-size 1 --calib-size 32 --n-steps 40 \
    --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors \
    --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors \
    --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors \
    --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized \
    --extra-param fp8transformer=true \
    --quantized-torch-ckpt-save-path ./ltx-2-transformer.pt \
    --hf-ckpt-dir ./LTX2-NVFP4/ \
    --extra-param merged_base_safetensor_path=./ltx-2-19b-dev-fp8.safetensors
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added new command-line parameters documentation for LTX-2 FP4
quantization examples (--hf-ckpt-dir and merged_base_safetensor_path
configuration options)

* **Improvements**
* Enhanced quantization pipeline to support conditional export behavior
based on model type
* Expanded LTX-Video model filtering patterns for more comprehensive
block detection

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-06 05:06:36 +00:00
jingyu-ml 10efcb65f6 [3.1/4] Diffusion Quantized ckpt export - WAN 2.2 14B (#855)
## What does this PR do?

**Type of change:** documentation <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

**Overview:** 

1. Added multi‑backbone support for quantization: --backbone now accepts
space- or comma-separated lists and resolves to a list of backbone
modules.
2. Introduced PipelineManager.iter_backbones() to iterate named backbone
modules and updated get_backbone() to return a single module or a
ModuleList for multi‑backbone.
3. Updated ExportManager to save/restore per‑backbone checkpoints when a
directory is provided, with {backbone_name}.pt files, and to create
target directories when missing.
4. Simplified save_checkpoint() calls to rely on the registered
pipeline_manager by default.


**Usage: **

```bash
python quantize.py --model wan2.2-t2v-14b --format fp4 --batch-size 1 --calib-size 32 \
    --n-steps 30 --backbone transformer transformer_2 --model-dtype BFloat16 \
    --quantized-torch-ckpt-save-path ./wan22_mo_ckpts \
    --hf-ckpt-dir ./wan2.2-t2v-14b 
```

Plans

- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml 
- [x] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml 

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Unified Hugging Face export support for diffusers pipelines and
components
  * LTX-2 and Wan2.2 (T2V) support in diffusers quantization workflow
* Comprehensive ONNX export and TensorRT engine build documentation for
diffusion models

* **Documentation**
* Updated to clarify support for both transformers and diffusers models
in unified export API
* Expanded diffusers examples with LoRA fusion guidance and additional
model options (Flux, SD3, SDXL variants)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-10 22:34:12 -08:00
Asha Anoosheh 510451322c Streamline KD & QAD transformers Trainers (#708)
## What does this PR do?

**Type of change:** ? Refactor and stabilization

**Overview:** 
* Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so
that we can remove multiple restrictions

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-01-10 01:28:56 +00:00
Keval Morabia 53a2ddebab Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer
(OMNIML-3033)
- [x] Mention in Latest News section with date on the date of merging
this PR (12/08)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-07 12:33:21 +05:30
Shengliang XuandKeval Morabia e0a6efbe70 fix trt engine building of the diffusers pipelines (#637)
## What does this PR do?

**Type of change:**

Bug fix

**Overview:**

1. The diffusion_trt.py needs the dynamic_shapes when running trtexec
for engine building. A previous change altered the format of
dynamic_shapes, fix it here.

2. the dynamic_shapes logic gets cleaned up. The existing logic is very
confusing

3. recover min-batch_size config for some pipelines. Previously some
pipelines set the min batch_size to be > 1, which was odd, so a previous
change sets them to be 1, but it turns out the oddity has a reason, the
trt engine building fails with the altered batch_size min/opt, thus
recover them.


## Testing

pytest tests/examples/diffusers

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-03 10:34:15 +00:00
ajrasane 229053323a [NVBUG: 5619158] Enforce high precision model dtype for diffusion trt (#526)
## What does this PR do?

**Type of change:** 
Minor code change

**Overview:** 
- Select the high precision dtype directly based on model type - FP16
for Stable Diffusion models, BF16 for Flux


## Testing
```python
python diffusion_trt.py --model flux-dev --benchmark
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No (No option to specify
dtype while loading pipeline)
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-11-06 23:44:37 -08:00
Keval Morabia c0590b0255 Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-23 01:43:25 +05:30
Keval Morabia 1ef1d72a1b Code quality improvements - typos, formatting, etc.
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-09-02 19:58:33 +05:30
Keval Morabia 4d1eb0caf5 Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-08-30 10:54:28 +05:30
Keval Morabia 2017cd9063 Update for 0.25.0 release 2025-03-03 22:54:22 +05:30