67 Commits
Author SHA1 Message Date
Ajinkya RasaneandCodex cf1f48fa0f [5565357] Fix SDXL NVFP4 export and performance (#2336)
### What does this PR do?

Type of change: Bug fix

Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:

- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.

For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.

This PR also changes shared exporter behavior:

- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.

Other model recipe configurations remain unchanged.

### Usage

```bash
python quantize.py \
    --model sdxl-1.0 \
    --model-dtype Half \
    --trt-high-precision-dtype Half \
    --format fp4 \
    --block-size 16 \
    --batch-size 2 \
    --calib-size 128 \
    --n-steps 20 \
    --quantized-torch-ckpt-save-path ./sdxl-fp4 \
    --onnx-dir ./onnx-sdxl-fp4
```

### Testing

- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
  - 302 native block-scaled NVFP4 GEMM tactics;
  - 38 native FP8 Conv tactics;
  - no FP4 Q/K/V projections;
  - all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [5565357]

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.

- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
2026-09-18 17:22:28 +00:00
Keval MorabiaandClaude Opus 5 449a39922b Pin nemo_automodel below 0.6 for the fastgen example (#2260)
### What does this PR do?

Type of change: Bug fix

`nemo_automodel` 0.6.0 removed
`nemo_automodel.recipes.diffusion.train.is_main_process` without a
replacement (it was a three-line rank-zero predicate in 0.5.0, and 0.6.0
defines no equivalent anywhere in the package).
`examples/diffusers/fastgen/dmd2_recipe.py` imports it, so the example's
import guard fires and **every** test in `tests/examples/diffusers/`
errors at collection:

```
ImportError: cannot import name 'is_main_process' from 'nemo_automodel.recipes.diffusion.train'
  tests/examples/diffusers/fastgen/test_resume_dataloader.py
E   ImportError: The DMD2 fastgen example requires `nemo_automodel`. ...
collected 42 items / 1 error
```

The requirement was `>=0.4.0,<1.0`, so CI picked 0.6.0 as soon as it was
published and the `onnx (diffusers)` job started failing on every PR
(e.g. runs 33020467654, 33019418460, 33010815298, 33007613265,
33006944292 — all unrelated branches). Capping at `<0.6` restores the
tested range.

Every other `nemo_automodel` symbol the example imports still exists in
0.6.0 (`_diffusers.auto_diffusion_pipeline.NeMoAutoDiffusionPipeline`,
`recipes.diffusion.train.TrainDiffusionRecipe`, and the four
`components.datasets.diffusion.*` helpers), so `is_main_process` is the
only blocker; the alternative is defining that predicate locally and
widening the cap again, which is worth doing separately if the example
is meant to track 0.6.

### Usage

```bash
pip install -r examples/diffusers/fastgen/requirements.txt
```

### Testing

Reproduced the break by diffing the published wheels: `is_main_process`
is defined at `nemo_automodel/recipes/diffusion/train.py:692` in 0.5.0
and absent from 0.6.0 (`grep -rn "def is_main_process"` over the
unpacked 0.6.0 wheel returns nothing). Confirmed the remaining imported
symbols are all still present in 0.6.0.

CI on this PR exercises the fix directly: the `onnx (diffusers)` job
installs from this requirements file and is the job that has been
failing.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — existing
dependency, tightened bound.
- Did you write any new necessary tests?: N/A — the existing
`tests/examples/diffusers/` suite is what this unblocks.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — dependency-pin fix for a break introduced and fixed within the
same unreleased cycle.
- Did you get Claude approval on this PR?: ❌

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed dependency compatibility for the FastGen diffusion example.
* Prevented installation of versions that could cause the example to
fail at startup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 05:59:02 +05:30
ZhiyuandClaude Opus 5 6261f854aa docs: rebuild the unified HF deployment support matrix from the deploy test suite (NVBug 6550792) (#2087)
### What does this PR do?

Type of change: documentation

Fixes [NVBug 6550792](https://nvbugspro.nvidia.com/bug/6550792) /
OMNIML-5693.

The **Unified HF Checkpoint Deployment Model Support Matrix** listed 9
model families and **no VLMs**, while
`tests/examples/hf_ptq/test_deploy.py` declares deployment cases for ~80
checkpoints across TRT-LLM, vLLM, and SGLang — including `Qwen2.5-VL`,
`Qwen3-VL-235B`, and `Nemotron-3-Nano-Omni`. QA (the filer) could not
use the doc to scope testing, and users could not tell what is actually
covered.

Filing also surfaced that the matrix lived in **three places that had
drifted apart**: only the `.rst` listed Qwen3-VL, only the README listed
Qwen3.5 MoE, and the skill reference had neither.

#### Changes

1. **Rebuilt the matrix in `docs/source/deployment/3_unified_hf.rst`**
from `test_deploy.py`, split into language models,
vision-language/multimodal, speculative decoding drafters, and
diffusion.

2. **Stated plainly what the matrix is and is not.** Review established
that the original "CI-validated" framing claimed more than the suite
substantiates, so a *What this matrix is based on* section now leads
with two limits:
- The suite is marked `release` and collects only under `--run-release`,
which **no workflow passes** — these are declared cases, not PR-gated
coverage.
- Each case is a **load-and-generate smoke check on the text path**: no
accuracy, no image/audio input, no diffusion output, no verification
that speculative decoding engages.

The legend follows from that: ✅ = declared in the suite, ⚠ = expected to
work but not a suite entry (or an entry that does not exercise the
feature the row names), `-` = not in the suite. Sections that would
otherwise over-read carry their own qualifiers — VLM rows are labelled
text-only smoke coverage, and Medusa and Wan 2.2 are ⚠ with the reason
stated.

3. **Removed the two duplicate copies**, replacing them with links, so
there is one table to maintain.

4. **Fixed stale prose**: the deployment tabs still claimed FP8-only
support on vLLM v0.6.5 and a source build of SGLang main from Jan 2025,
both contradicting the version table above them. The TRT-LLM floor moves
to v1.2.0, qualified as the oldest version stated rather than the oldest
that works.

5. **Dropped the Phi series** from the deployment matrix, following
#2115 (NVBug 6563509) and confirmation that Phi-4 is being deprecated.

### Usage

N/A — documentation only.

### Testing

- `docutils` parse of the modified `.rst`: no warnings or errors from
the new content; all 5 tables parse with every cell in the correct
column.
- Cell contents cross-checked against `test_deploy.py` by AST-parsing
the `ModelDeployerList(...)` calls rather than by eye; the scope caveats
were each verified against `tests/_test_utils/deploy_utils.py`.
- `pre-commit run --files …` passes; `build-docs` green.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — documentation only
- Did you get Claude approval on this PR?: ❌ — not yet run

### Additional Information

**Two known follow-ups, neither in scope here:**

1. **Nothing enforces that the doc matrix tracks `test_deploy.py`.**
Consolidating to one copy removes the three-way drift but not the
doc-vs-test drift; a generator plus a CI check would close it.
2. **The release deployment suite does not run in CI.** Wiring it into
per-backend release CI is what would let ✅ mean "verified to pass"
rather than "declared". That needs GPU capacity across three backends
and should be tracked on its own.

**For the filer (@Kenny Kang):** the ✅ cells are the scope the release
deploy suite declares, and `test_deploy.py` carries the checkpoint, TP
size, and minimum SM version per entry — but please read the legend
first, since those cases are not currently executed by CI.

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 01:37:07 +05:30
jingyu-mlandClaude Opus 4.8 6b4ad85849 Qwen-Image diffusers PTQ: FP8 / NVFP4 / NVFP4-SVDQuant HF checkpoints (#1706)
### What does this PR do?

Type of change: New feature

Adds **Qwen-Image** (`Qwen/Qwen-Image`, `QwenImageTransformer2DModel`)
to the diffusers quantization example and exports HuggingFace
checkpoints in three precisions — **FP8**, **NVFP4**, and **NVFP4 +
SVDQuant** — through the unified HF export.

- Registers `--model qwen-image` (lazy diffusers import; no
`trust_remote_code`).
- Transformer-block-range recipe: quantizes only the linears under
`transformer_blocks`, keeping the **first 2 / last 2** blocks (and
everything outside `transformer_blocks`) in original precision. Applied
**before** calibration so SVDQuant never mutates the excluded blocks.
Expressed with the top-level `enable` `QuantizerCfgEntry` field
(disable-all → re-enable `transformer_blocks` → disable first/last-N).
- SVDQuant export (AWQ-style): promotes quantizer-owned tensors to clean
module-level safetensors keys at export time —
`weight_quantizer.svdquant_lora_a/b → <module>.svdquant_lora_a/b` and
`input_quantizer._pre_quant_scale → <module>.pre_quant_scale` — with a
documented `NVFP4_SVD` `quantization_config` (`group_size`,
`has_zero_point: false`, `pre_quant_scale: true`, `lora_rank`). **Core
SVDQuant quantization code (`modelopt/torch/quantization`) is
unchanged.**
- Shared export-path change — **intentionally global** (applies to all
diffusers exports — SDXL / Flux / Wan, not just Qwen; the full export
suite was verified green on GB200): `hide_quantizers_from_state_dict`
now strips quantizer state from *all* modules (not just quant-linears)
so calibrated norm-layer input quantizers no longer leak
`input_quantizer._amax`. (An earlier `max_shard_size` workaround was
dropped after merging `main`: #1794 makes the ComfyUI layerwise-metadata
post-processing a no-op unless explicitly opted in, so a default sharded
export no longer hits the unsupported-sharded path.)

### Usage

```bash
python examples/diffusers/quantization/quantize.py \
    --model qwen-image --override-model-path <Qwen-Image> --model-dtype BFloat16 \
    --format fp4 --quant-algo svdquant --lowrank 32 \
    --calib-size 64 --n-steps 20 \
    --hf-ckpt-dir <out>
# FP8:   --format fp8 --quant-algo max
# NVFP4: --format fp4 --quant-algo max
```

### Testing

- Focused unit + example tests pass on GB200 (sm_100): block-range
recipe, `NVFP4_SVD` config schema, SVDQuant forward/fold (LoRA stays on
`weight_quantizer`), Qwen dummy-input / strict-QKV-fusion / promotion,
pipeline loading, and the diffusers HF-export test for Qwen FP8 / NVFP4
/ SVDQuant.
- Full `tests/examples/diffusers/test_export_diffusers_hf_ckpt.py` is
green (SDXL, Flux, Qwen, Wan2.2) — confirms the shared export changes do
not regress other models.
- End-to-end on the real `Qwen/Qwen-Image` (~20B): all three formats
export valid HF checkpoints — only `transformer_blocks` 2..57 quantized,
nothing outside, no quantizer/`_amax` leak, correct
`weight_scale`(`_2`)/`input_scale`, promoted SVDQuant keys
(rank-consistent shapes), and the expected `quantization_config`.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ <!-- live-model LoRA storage
unchanged; existing exports unaffected -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ❌ <!-- new-example feature; add a
CHANGELOG.rst entry if required -->
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

All changes are confined to the diffusers example
(`examples/diffusers/quantization`) plus the shared export path
(`modelopt/torch/export`); the core quantization library is untouched.

### Follow-up (next step): fused-QKV SVDQuant for sglang / Nunchaku

This export keeps attention `q/k/v` (and `add_q/k/v_proj`) as
**separate** projections — the diffusers-native layout. That matches
sglang's bf16 / FP8 / plain-NVFP4 paths (which also keep QKV separate)
and ModelOpt/TRT-LLM consumers, so those load 1:1.

sglang's **NVFP4-SVDQuant (Nunchaku)** path, however, builds a **fused**
`to_qkv` with a *single* fused rank-r LoRA in Nunchaku-native format
(`proj_down`/`proj_up`, `smooth_factor`, `wscales`/`wtscale`). Our
per-projection tensors (`svdquant_lora_a/b` + `pre_quant_scale`; three
independent rank-r decompositions) are not directly loadable there — and
cannot be fused at load time, because the fp16 weight residual needed to
derive a single fused rank-r is not preserved after export.

**Planned next step:** an opt-in fused-QKV SVDQuant export mode that
fuses q/k/v **before** SVDQuant calibration (yielding one rank-r over
the fused weight) and emits a Nunchaku-compatible layout, enabling
lower-latency fused-QKV inference in sglang. Tracked as a separate
follow-up.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Qwen-Image (`QWEN_IMAGE`) model quantization and Diffusers
export support
  * Added NVFP4_SVD (SVDQuant) export configuration support
* Added transformer block-range quantization recipes (exclude first/last
blocks)
* **Bug Fixes**
  * Improved missing-pipeline error messaging for Qwen-Image
* Prevented quantizer-related tensor/buffer leakage by promoting and
cleaning quantizer outputs during export
* **Tests**
  * Added Qwen-Image HF checkpoint export tests and offline fixtures
  * Added unit coverage for SVDQuant promotion/clean state-dict keys
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-07 17:21:31 -07:00
Keval MorabiaandClaude Sonnet 4.6 33bfa8b1fe CI/Dev env bump (#1818)
### What does this PR do?

Type of change: chore

Bumps CI/dev tooling and test containers.

**Container bumps**
- NeMo test containers → 26.06
- TRT-LLM container → 1.3.0rc19
- transformers max version → 5.12

**Dev tooling bumps**
- ruff bump 0.12.11 → 0.15.18
- mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`,
`strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4
modules (rather than blanket-suppressing them); remove 2 stale `# type:
ignore` comments
- pre-commit 4.3.0 → 4.6.0
- sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings
= ["ref.python"]` to fix cross-reference ambiguity error new in sphinx
9.x
- trl fix for newly released 1.7 version

**Bug fixes surfaced by the bumps**
- sparsity (weight): make the weight mask DTensor-aware under FSDP. The
transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save
through torch's DTensor-based `get_optimizer_state_dict`, which
triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the
dynamic `weight` getter. The mask is now distributed to the weight's
mesh/placements before masking, cached, and rebuilt only when the
sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity`
example test.

### Testing

- `pre-commit run --all-files` ✅ (including mypy 2.1.0)
- `nox -s docs` ✅
- `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅
- `llm_sparsity` GPU example test (FSDP path) verified in CI

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **Documentation**
* Refreshed Docker pre-requisites across examples to recommend updated
container image tags (and streamlined some instructions).
* **Bug Fixes**
* Improved sparse weight mask handling for DTensor/FSDP by aligning and
caching distributed masks.
  * Made TensorRT engine byte retrieval return immutable `bytes`.
* Reduced Sphinx cross-reference warnings and tuned Transformers
compatibility warning thresholds.
* **Tests**
  * Increased default unit test timeout on Windows runners.
* **Chores**
* Updated CI workflow container tags and refreshed linting/typing/docs
version pins, plus related mypy configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-06-27 01:00:25 +05:30
Keval Morabia b6bf6b7997 Update Roadmap Issue link
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-23 21:42:37 +05:30
jingyu-mlandClaude Opus 4.8 c458ad36f1 [2/2] Remove examples/diffusers/eval image-quality evaluation example (#1694)
> **Part 2 of 2** — removal for the 0.46 release. Depends on **Part 1**:
#1798 (the 0.45 deprecation).

### What does this PR do?

Type of change: Backward breaking change (removal of a deprecated
example)

Removes the `examples/diffusers/eval` image-quality evaluation example
(ImageReward / CLIP-IQA / CLIP metrics) and its references in
`examples/diffusers/README.md`. The example was **deprecated in 0.45**
(see #1798) and is removed here for **0.46**, per the [Deprecation
Policy](https://github.com/NVIDIA/Model-Optimizer#deprecation-policy)
(1-release migration before removal).

Scope is limited to the diffusers example; the unrelated
`examples/llm_ptq` "Evaluate Accuracy" section is untouched.

### Usage

N/A — removes example scripts; no library API changes.

### Testing

N/A — deletion of example scripts plus documentation cleanup. Verified
`examples/diffusers/README.md` has no remaining references to the
deleted `eval/` directory.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — removes the (previously
deprecated) `examples/diffusers/eval` example. No public `modelopt` API
is affected.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — removal only.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — 0.46 Backward Breaking Changes.
- Did you get Claude approval on this PR?: N/A

### Additional Information

Paired with #1798 (the 0.45 deprecation). This PR targets **0.46** and
should **not** carry the `cherry-pick-0.45.0` label (that belongs on
#1798).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Documentation**
* Updated the Diffusers “Model Optimizations” README by shortening the
overview, removing the “Evaluate Accuracy” guidance (links, inputs,
commands, and example results), and refining notes about per-subsection
`requirements.txt` usage.

* **Breaking Changes / Deprecations**
* Updated the 0.46 changelog to reflect that the Diffusers image-quality
evaluation example (ImageReward/CLIP-IQA/CLIP) is no longer maintained.

* **Chores**
* Removed the Diffusers evaluation example, including its evaluation
entrypoint, metrics, and shared utilities.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 08:44:59 +05:30
090b1c5114 fastgen DMD2: make the Qwen-Image example self-contained on stock nemo_automodel (#1688)
### What does this PR do?

**Type of change:** new example / refactor (example self-containment)

The published `examples/diffusers/fastgen` DMD2 Qwen-Image distillation
example previously relied on
**local, unpublished modifications to the sibling `nemo_automodel`
package** (data path, collate, a
partial-load checkpointer, and Qwen-Image preprocessing). An external
user running the published
Model-Optimizer example against **official stock `nemo_automodel`**
would hit import/attribute errors.

This PR makes the example **self-contained on stock
`nemo_automodel>=0.4.0`** — with the changes kept
**as small as possible**: the team's actual AutoModel delta is only ~430
lines, so wherever the
upstream module is importable, the delta is expressed as a thin subclass
/ small reimplementation rather than a
full-file copy.

- **`fastgen_data/`** — the DMD2 data path:
- `collate_fns.py` — **reuses the upstream `SequentialBucketSampler` and
reimplements the collate**:
it builds the DMD2 batch directly from the vendored dataset's per-item
output (`image_latents` /
`text_embeddings` / `text_embeddings_mask` + an optional broadcast
`negative_text_embeddings` for
CFG) and deliberately does **not** call upstream
`collate_fn_production`, which stacks
model-specific token keys (`clip_tokens` / `t5_tokens`) absent from the
Qwen-Image cache. The
    builder loads an optional `negative_prompt_embedding_path`.
- `text_to_image_dataset.py` — a **faithful vendored copy** of the
upstream reader (its
`prompt_embeds_mask` emission is interleaved with cache loading, so
wrapping it would force a
    redundant per-item `torch.load`; carried verbatim instead).
- **`fastgen_checkpoint.py`** — `PartialLoadCheckpointer(Checkpointer)`
that overrides only
`load_optimizer` (FSDP2 `DefaultLoadPlanner(allow_partial_load=True)`)
so optimizer resume works
without patching upstream. Injected via an in-place re-bless of
`self.checkpointer` in the
  recipe's `load_checkpoint` (model-state load stays strict).
- **`preprocess/`** — Qwen-Image preprocessing
(`preprocessing_multiprocess.py` + `processors/`),
trimmed to the image path (drops the flux/wan/hunyuan processors and the
video base class).
It lives in AutoModel's top-level `tools/` tree, which is **not**
shipped in the pip package, so
it cannot be wrapped and is vendored; `MultiTierBucketCalculator` is
imported from stock upstream.
- **`make_negative_prompt_embedding.py`** — generates the optional CFG
negative-prompt embedding.
- All `configs/*.yaml` target `fastgen_data.build_*` (a test enumerates
every config).
- Licensing: the AutoModel-copied files are NVIDIA-authored Apache-2.0,
so they carry only the
standard NVIDIA SPDX header (managed by the `insert-license` hook) — no
per-file provenance note,
no duplicated license, no pre-commit exclusion, and no separate
`LICENSE` note.
  `nemo_automodel[diffusion]` version bound in `requirements.txt`.

The DMD2 math in `modelopt/torch/fastgen/` is **unchanged** — only
example/training-time glue moved.

### Usage

```bash
# Install example deps (stock nemo_automodel) from a source checkout
pip install -r examples/diffusers/fastgen/requirements.txt

# Build the training cache from raw images (Qwen-Image VAE latents + text embeddings)
python examples/diffusers/fastgen/preprocess_qwen_image.py image \
    --image_dir <raw images> --output_dir <cache dir> --processor qwen_image \
    --caption_format meta_json

# Generate the CFG negative-prompt embedding once
python examples/diffusers/fastgen/make_negative_prompt_embedding.py \
    --output <cache dir>/negative_prompt_embedding.pt

# Point the config's data.dataloader.cache_dir + negative_prompt_embedding_path at the cache, then train.
```

See `examples/diffusers/fastgen/README.md` → "Requirements &
self-contained data path".

### Testing

- New `tests/examples/diffusers/fastgen/test_vendored_migration.py`:
environment-independent
invariants (every config targets a vendored builder; no `tools.*`
imports; each former AutoModel
patch is vendored / wrapped / a documented exclusion; the
former-vendored files carry the standard NVIDIA SPDX header, no
provenance note or duplicate license) plus
dependency-guarded structural tests (the collate emits the batch
contract + broadcasts the
negative embedding; the builder accepts
`negative_prompt_embedding_path`; the checkpointer
overrides only `load_optimizer`; the Qwen-Image processor
self-registers).
- **Validated 9/9 against a pure stock `nemo_automodel` 0.4.0 worktree**
(none of the local patches
  present) via SLURM — re-run after this slim-down.
- The migrated code path is exercised by a live multi-GPU DMD2 run that
resumed from a checkpoint
  through the vendored `PartialLoadCheckpointer`.
- `ruff check` + `ruff format --check` clean on all changed files.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (additive; bundled configs
target the vendored builders)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
(NeMo-AutoModel @ `e42584e3`, Apache-2.0; per review these
NVIDIA-authored files carry the standard NVIDIA SPDX header, no separate
provenance / `LICENSE` note; `nemo_automodel` was already a dependency)
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: N/A (example-only change)
- Did you get Claude approval on this PR?: ❌ (pending — opened as draft)

### Additional Information

Opened as a **draft** pending: OSRB review of the vendored
NeMo-AutoModel (Apache-2.0) code, CI green,
and (optional) a multi-GPU smoke-train + resume on stock upstream.
Vendored from
NVIDIA-NeMo/Automodel at commit `e42584e3`.

### Update (post-review)

- **Mid-run resume data-correctness fix** (`6ffbc52c9`): on resume the
`StatefulDataLoader`'s restored state did not advance past the resume
point, so each window re-served the same data slice and multi-window
(SLURM-windowed) runs under-covered the dataset. Fixed by rebuilding a
fresh loader and skipping the deterministic sampler to the position
implied by `global_step`; added a SLURM-free, GPU-free CPU regression
test (`tests/examples/diffusers/fastgen/test_resume_dataloader.py`).
- **Licensing review** (`be832ae95`): the AutoModel-copied files are
NVIDIA-authored Apache-2.0, so they now carry only the standard NVIDIA
SPDX header — dropped the per-file provenance note, the duplicated
original-license block, the `insert-license` pre-commit exclusion, and
the `LICENSE` note.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added vendored Qwen‑Image preprocessing (multi-process) with processor
registry support.
* Updated DMD2 data loading with a dedicated dataset/collation pipeline,
including negative-prompt embedding/mask handling.
* Improved training resume behavior by rebuilding dataloader state and
making checkpoint optimizer restore tolerant of partial FSDP2 optimizer
shards.
* **Documentation**
* Refreshed the fastgen README and config notes for real-data training;
removed the prior mock-data smoke workflow.
* **Tests**
* Added regression and migration tests covering vendored wiring,
collate/dataloader contracts, processor registration, and
resume/checkpoint behavior.
* **Chores**
* Updated licenses/attribution, vendoring/tooling guards, requirements
pinning, linting configuration, and repository ownership rules.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-22 16:22:38 -07:00
Keval Morabia a21197c277 Remove unsafe torch.load from examples/diffusers/fastgen (#1740)
Follow-up to #1326 - Remove unsafe `torch.load(..., weights_only=False)`
in `examples/diffusers/fastgen`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated checkpoint loading mechanisms in FastGen examples to improve
compatibility and reliability during model restoration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-16 08:44:13 +05:30
jingyu-mlandClaude Opus 4.8 aec72ffa68 Add DMD2 distillation for Qwen-Image (fastgen) (#1326)
### What does this PR do?

**Type of change:** New example + new `modelopt.torch.fastgen` library
module.

Adds **DMD2 (Distribution Matching Distillation) for Qwen-Image** —
distilling the base model into a few-step (1–4) generator. Includes the
framework-agnostic `modelopt.torch.fastgen` loss library (DMD pipeline,
EMA, optional GAN discriminator) and a NeMo AutoModel–based training
example with a mock-data smoke config, a real-data config, and inference
/ export scripts.

**Noted**: the example script will be migrated to AutoModel repo

### Usage

```bash
# Mock-data wiring smoke — runs end-to-end with no dataset to prepare
torchrun --nproc-per-node=8 \
    examples/diffusers/fastgen/dmd2_finetune.py \
    --config examples/diffusers/fastgen/configs/dmd2_qwen_image_smoke.yaml
```

See `examples/diffusers/fastgen/README.md` for real-data training and
inference.

### Testing

Unit tests under `tests/unit/torch/fastgen/`; `pre-commit` /
code-quality clean.

### Before your PR is "*Ready for review*"

- Backward compatible?: ✅ (new, additive module)
- Followed `CONTRIBUTING.md` for any copied code / new deps: ✅
- New tests added?: ✅
- Updated Changelog?: N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Adds a FastGen-based distillation framework (DMD2) with
student/fake-score training, EMA support, GAN discriminator branch,
inference pipeline, and export utilities.
* Qwen-Image integration with latent packing and feature-capture for
plugin-enabled pipelines.

* **Documentation**
* New README, example configs, and runnable example scripts for
Qwen-Image distillation and inference.

* **Tests**
* Comprehensive unit tests covering math parity, gradient routing,
plugins, hooks, EMA, and recipe setup.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-08 16:03:35 -07:00
c78e654744 Skip Softmax diffusion export (#1269)
### What does this PR do?

Type of change: New Feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Adds HuggingFace `config.json` export of skip-softmax sparse-attention
calibration for diffusion pipelines (e.g. Wan 2.2), on top of the base
skip-softmax work.

- **`_export_diffusers_checkpoint`** walks every `nn.Module` component
of a diffusers pipeline, calls `export_sparse_attention_config`, and
writes the result into that component's `config.json` under the
`sparse_attention_config` key. The sparse config lives **only** in
`config.json` — there is no standalone `sparse.yaml`.
- **`export_sparse_attention_config`** emits a `config_groups` schema
where each algorithm's parameters are nested inside its own group; only
`config_groups` and `producer` are top-level:
- skip-softmax group → `algorithm: "skip_softmax"`, `targets`, `ignore`
(layers kept dense — e.g. cross-attention + first/last blocks),
`initial_disabled_steps` (opt-in, user-set; emitted only when `> 0`),
`threshold_scale_factor` (`a * exp(b * target_sparsity)`), and
`target_sparsity`.
- N:M group → `algorithm: "sparse_softmax"` with
`sparsity_n`/`sparsity_m`, `dense_sink_tokens`, `dense_recent_tokens`
flattened into the group.
- **Deploy reader**
(`modelopt/torch/sparsity/attention_sparsity/plugins/sparse_attn_config.py`)
reads these per-group params back, keeping the export↔load round-trip
consistent.
- **Example wiring**:
`examples/diffusers/sparsity/wan22_skip_softmax.py` gains
`--export-dir`, `--skip-softmax-threshold`, and
`--initial-disabled-steps`. `--export-dir` runs
`export_hf_checkpoint(pipe, export_dir=...)` after calibration.
- Updated `CHANGELOG.rst`.

### Usage

```bash
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path Wan-AI/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 --calib-size 4 \
    --initial-disabled-steps 5 \
    --export-dir ./wan22_skip_softmax_ckpt
```

Resulting layout — a `config.json` per component, **no `sparse.yaml`**:

```
wan22_skip_softmax_ckpt/
├── transformer/config.json        # carries sparse_attention_config
├── transformer_2/config.json      # carries sparse_attention_config
├── vae/ …  text_encoder/ …  tokenizer/ …  scheduler/ …
└── model_index.json
```

A representative `config.json` entry for a diffusion transformer:

```json
"sparse_attention_config": {
  "config_groups": {
    "group_0": {
      "algorithm": "skip_softmax",
      "targets": ["WanAttention"],
      "ignore": ["blocks.0.attn1", "blocks.0.attn2", "…"],
      "initial_disabled_steps": 5,
      "threshold_scale_factor": {
        "formula": "a * exp(b * target_sparsity)",
        "prefill": {"a": 1443.49, "b": 4.30}
      },
      "target_sparsity": {"prefill": 0.5}
    }
  },
  "producer": {"name": "modelopt", "version": "0.45.0..."}
}
```

The N:M variant adds a second group:

```json
"group_1": {
  "algorithm": "sparse_softmax",
  "targets": ["WanAttention"],
  "sparsity_n": 2, "sparsity_m": 4,
  "dense_sink_tokens": 0, "dense_recent_tokens": 64
}
```

### Testing

- `tests/examples/diffusers_sparsity/test_sparsity.py`: baseline /
triton-baseline / fixed-threshold runs of the Wan 2.2 example, plus a
Python-API calibrate → **export** test asserting the nested
`sparse_attention_config` (`threshold_scale_factor`, `target_sparsity`,
`ignore`, `initial_disabled_steps`) and the absence of any
`sparse.yaml`.
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
and `test_sparse_attn_config.py`: unit coverage of the per-group export
schema and the deploy-reader round-trip (writer nests → reader reads
from groups → internal mtsa config unchanged).
- Validated end-to-end on Wan 2.2 T2V-A14B: full 4-prompt / 40-step /
81-frame calibration; the exported checkpoint carries the nested schema
in both `transformer` and `transformer_2` `config.json`, and runtime
measurement shows ~47–49% tile sparsity at a 0.5 target.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ The exported
`sparse_attention_config` schema was renamed and nested per-group during
0.45.x development, and the loader reads only the new layout —
checkpoints exported by earlier 0.45.x builds must be re-exported. No
released version is affected. <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 16:02:41 -07:00
Keval MorabiaandClaude Opus 4.8 0081473861 Speed up slow unit/gpu/example tests (#1616)
### What does this PR do?

Type of change: test infrastructure / test speedups + CI stabilization

Make the test suite faster, `tests/unit` hermetic, and the CI lanes
stable, without losing coverage. Most changes are mechanical test/infra
edits; the buckets below cover the diff broadly.

**Unit tests — hermetic (no HF Hub):** toy local datasets/configs + the
local tiny tokenizer (with a checked-in chat template) replace Hub
assets; `tests/unit/conftest.py` enforces offline mode. Genuinely-HF
tests moved to `tests/gpu*` (e.g. the new
`tests/gpu/torch/utils/test_dataset_utils.py`). `CONTRIBUTING.md`
documents the hermetic-unit-test expectation.

**Unit-test speedups (no coverage loss):** speculative (disable CPU
torch.compile), calibrator (fewer histogram bins), ONNX conv/dynamo
(smaller shapes + representative subset), Ruler/sparse-attention (local
tokenizer), data-parallel autoquant (world size 4→2). Shared
`tiny_tokenizer` fixture. The distributed test helper now uses a private
`spawn` context instead of mutating the global start method (avoids
cross-test contamination).

**Rarely-used autonas/fastnas tests:** heavy parametrize cases marked
`@pytest.mark.manual`, one representative kept per test (fastnas
preferred); lighter sibling tests still cover core behavior. The legacy
FSDP1 NAS distributed test is also dropped: FastNAS/AutoNAS aren't used
with either FSDP1 or FSDP2, and FSDP1 is superseded by the newer FSDP2
API — so we keep a single FSDP2 case as a sanity check and drop FSDP1,
leaving the suite leaner.

**gpu_megatron:** deduplicate distributed worker pools by world_size
within a module (saves a redundant pool spin-up in multi-pool files;
module-scoped, no cross-module reuse).

**Example tests:** reduce per-test work via args that default to current
behavior (tests pass the fast values) — torch_onnx TRT optimization
level, diffusers calibration/inference steps, eagle `sample_size`,
megatron_bridge iters/calib, llm_sparsity data slice, export
safetensors-structure `calib_size`. Also enable the recently added
`gpt-oss` example tests in CI.

**Per-test timeouts:** `pytest-timeout` with a default per-directory
timeout (60s unit / 300s gpu+example) enforced in `tests/conftest.py`
(`timeout_func_only` in `pyproject.toml`), so a new test cannot silently
exceed the budget — an unmapped test dir crashes collection. A few
inherently slow tests carry explicit higher per-test overrides
(CUDA-compile, autotune, dflash).

**CUDA kernel pre-compilation:** a dedicated `tests/gpu/_extensions`
test JIT-builds the conv3d implicit-GEMM kernel up front (collected
before the functional tests in the same process) so the one-time build
cost no longer lands on — and time out — the first functional test that
uses it. Mirrored into the `llm_ptq`/`vlm_ptq` example lanes.

**Test relocation & optional-dependency guards:** vLLM sparsity plugin
test moved to `tests/gpu_vllm` (drops the in-test `importorskip`);
diffusers-dependent unit test guarded with `importorskip("diffusers")`
for partial-install lanes; `gpt_oss` example test dir renamed to
`gpt-oss` to match the CI matrix.

**Diffusers test models:** shared model-path constants in
`tests/_test_utils/examples/models.py` consolidated/renamed and point at
tiny `hf-internal-testing` test pipes (SDXL/SD3/FLUX) so
cachify/quantize/export tests run on toy weights; `local_id`s
normalized.

**Shared dataset utils:** `examples/llm_sparsity/.../hf_pts.py` now uses
`get_dataset_dataloader` (drops the bespoke cnn_dailymail-only
`get_calib_dataloader`; supports any registered/HF/JSONL dataset,
includes attention_mask); `data_prep.py` gains `--max_samples`.

**CI workflows:** container image bumps (pytorch 26.04→26.05, TRT-LLM
rc16→rc17) and tightened lane timeouts (unit 30→15 min, gpu lanes
trimmed, onnx example lane 45 min).

**Imports at top of file:** in-function imports across the test suite
are moved to module top per the coding guideline, conservatively —
optional deps stay guarded (in-function or behind a module-level
`importorskip`) in `tests/unit` since the partial-install lane runs
without them, and build/hardware-availability imports (apex, triton,
megatron/transformer_engine, tensorrt_llm) plus `_test_utils` lazy
guards are left in place.

**Kernel warning filters:** the repeated `filterwarnings` blanket-ignore
in six `tests/gpu/torch/kernels/**` modules is consolidated into a
scoped hook in `tests/gpu/torch/kernels/conftest.py` (kernel tests only
— the rest of the suite keeps surfacing warnings).

**Eagle example speedups:** `torch.compile` (eagle recipe default) added
~2 min to every eagle training test; it's now disabled in the eagle
example tests except one smoke (`test_llama_eagle3[1-False]`), and the
downstream resume / AR-validate / export tests point at the compile-free
checkpoint. Measured: `test_ar_validate` 139s→17s, offline training
142s→22s, streaming 140s→23s — the compile path is still smoke-tested
once.

**Example lanes install editable (`-e`):** so example scripts launched
as subprocesses resolve `modelopt` to the same source path as the test
process and reuse the pre-compiled CUDA-extension cache instead of
recompiling (~2 min/test); verified in the TRT-LLM container.

**Tiny test tokenizer:** `get_tiny_tokenizer` defaults to left padding
(what decoder-LM calibration expects) and ships a terse
generation-tagged chat template — replacing a verbose ChatML one that
inflated tokenized length on the 128-vocab tokenizer and broke the
offline-PTQ example tests' `max-seq-len` filter.

**Restored Hub-download coverage:** the live (ungated) HF dataset
round-trips exercising `get_dataset_samples`' download branch now live
in `tests/gpu/torch/utils/test_dataset_utils.py` (they had been dropped
from the hermetic unit file without a counterpart).

Individual file changes not explicitly called out above fall under this
general test/CI cleanup.

### Testing

Unit + the touched gpu_megatron files validated locally; example/GPU
lanes validated in CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (tests + example CLI args
default to prior behavior)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (optimizes/relocates
existing tests)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Chores**
* Updated container image versions for PyTorch (26.04→26.05),
TensorRT-LLM (1.3.0rc16→1.3.0rc17), and ONNX/TensorRT (26.04→26.05).

* **Tests**
* Enhanced test isolation: unit tests now run hermetically without
HuggingFace Hub access.
* Optimized test runtime via smaller model/dataset parameters and
parallel test caching.
* Added CUDA extension availability tests and extended dataset utility
coverage.

* **Documentation**
* Updated testing guidelines in `CONTRIBUTING.md` to emphasize offline
test design.

* **Chores**
* Added pytest timeout configuration and improved CI/CD workflow
efficiency with editable installs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 17:13:10 +00:00
3ff15ccef3 Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do?

Type of change: ?  new feature

<!-- Details about the change. -->
Adds post-processing support for exported diffusion model checkpoints to
enable NVFP4 block scale swizzling and configurable padding strategies.
This allows exported quantized checkpoints to be directly consumed by
inference runtimes (e.g., ComfyUI with comfy_kitchen) that require
cuBLAS 2-D block-scaling-factors layout.
Changes:
1) Unified post-processing step (_postprocess_safetensors): Loads saved
safetensors files and applies merge, padding, swizzle, and quantization
metadata injection in a single pass.
2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales
from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled
layout per the cuBLAS specification.
3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale
tensors to multiples of 16, with "row" (rows only) or "row_col" (both
dimensions) strategies.
4) Standalone quantization metadata (build_layerwise_quant_metadata):
Extracted from merge_diffusion_checkpoint so _quantization_metadata can
be injected independently of merging — works for both merged (LTX-2) and
standalone (Flux2) exports.
5) Bug fix (conversion.py): Wrapped yield in try/finally in
set_quantizer_by_cfg_context so quantizer states are always restored,
fixing an issue when yield fails.
### Usage

```python
# LTX-2 export with merge + swizzle + padding
export_hf_checkpoint(
  pipeline,
  export_dir="./output",
  merged_base_safetensor_path="./ltx-2-22b-dev.safetensors",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
  enable_layerwise_quant_metadata=True,
)
# Flux2 standalone export with swizzle + padding (no merge needed)
export_hf_checkpoint(
  transformer,
  export_dir="./output",
  enable_swizzle_layout=True,
  padding_strategy="row_col",
)
# Via quantize.py CLI
python quantize.py \
  --model ltx-2 --format fp4 \
  --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \
  --extra-param enable_swizzle_layout=true \
  --extra-param padding_strategy=row_col \
  --hf-ckpt-dir ./output
```

### Testing
1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base
checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn
scales in swizzled layout, and _quantization_metadata . Ran the
checkpoint with ComfyUI
2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has
correct uint8 weights, float8_e4m3fn scales in swizzled layout, and
_quantization_metadata . Ran the checkpoint with ComfyUI

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Diffusers export: optional NVFP4 support — swizzle layout, row/row_col
padding, and optional per-layer quantization metadata; exports are now
post-processed to apply these options.
* Export flow accepts new flags to enable swizzle, padding strategy, and
layerwise metadata.

* **Bug Fixes**
* Quantizer context manager now always restores state, including on
exceptions.

* **Tests**
* Added unit tests for NVFP4 padding, swizzling, metadata injection, and
post-processing.

* **Documentation**
  * README example updated to show swizzle and padding flags.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Signed-off-by: ynankani-nv <ynankani@nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com>
Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com>
2026-05-22 10:36:59 +00:00
kaix-nv c9098b63fb [4/n] Add vLLM integration for modelopt sparse attention (#1127)
### What does this PR do?

Type of change: New feature, new example, new tests, documentation.

Adds vLLM integration for ModelOpt sparse attention with paged KV cache
support.

This PR extends the ModelOpt Triton flash attention path so K/V can be
read directly from vLLM's paged KV cache through `block_table` lookup.
This avoids gather-to-contiguous copies when serving exported
sparse-attention checkpoints with vLLM.

The vLLM integration swaps vLLM's `FlashAttentionImpl` with
`ModelOptSparseAttentionImpl` after model load. The sparse configuration
is read from the exported checkpoint's `config.json`
`sparse_attention_config` block, written by
`examples/llm_sparsity/attention_sparsity/hf_sa.py`.

The restored checkpoint metadata supports:

- calibrated skip-softmax metadata (`threshold_scale_factor`,
`target_sparse_ratio`)
- N:M sparse-softmax metadata (`sparsity_n`, `sparsity_m`)
- dense token preservation metadata (`dense_sink_tokens`,
`dense_recent_tokens`)

The vLLM path uses ModelOpt Triton for sparse prefill launches.
Decode-only launches, cascade/prefix-cache metadata, and launches
without active sparse work delegate back to vLLM FlashAttention.

### Limitations

- Sparse attention is enabled for sparse prefill only.
- Decode-only launches currently fall back to vLLM FlashAttention.
- Attention sinks from vLLM FlashAttention are rejected until the
ModelOpt Triton path supports them.
- CUDA graph capture is not validated with this sparse attention path
yet; use `--enforce-eager`.
- Quant-only serving remains covered by `vllm_serve_fakequant.py`.
- Combined sparse attention + quantization serving is not handled by
this launcher in this PR and is planned as follow-up work.

### Usage

Export a checkpoint with calibrated skip-softmax and sparse24 metadata:

```bash
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
    --pyt_ckpt_path /path/to/hf-model \
    --sparse_attn skip_softmax_calib_sparse24 \
    --target_sparse_ratio 0.5 \
    --calib_samples 64 \
    --calib_max_seqlen 16384 \
    --calib_chunk_size 4096 \
    --seq_len 2048 \
    --export_dir /path/to/modelopt-skipsoftmax-sparse24-export
```

Serve the exported checkpoint with the vLLM sparse-attention launcher:

```bash
PYTHONPATH=$PWD python examples/vllm_serve/vllm_serve_sparse_attn.py \
    /path/to/modelopt-skipsoftmax-sparse24-export \
    --tensor-parallel-size 8 \
    --host 0.0.0.0 \
    --port 8000 \
    --trust-remote-code \
    --enforce-eager
```

Send a request through the OpenAI-compatible endpoint:

```bash
curl http://localhost:8000/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
      "model": "/path/to/modelopt-skipsoftmax-sparse24-export",
      "messages": [{"role": "user", "content": "Explain sparse attention in one paragraph."}],
      "max_tokens": 128
    }'
```

### Testing

GitHub CI on the latest commit is green:

- DCO
- code-quality
- docs build / deploy preview
- unit tests, including Linux, Windows, multi-version, partial-install,
and launcher jobs
- example tests
- GPU tests, including required GPU gate
- regression tests, including required regression gate
- `codecov/project`

Focused test coverage added/updated for this PR includes:

-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_sparse_attention_conversion.py`
-
`tests/unit/torch/sparsity/attention_sparsity/test_triton_skip_softmax.py`
- `tests/gpu/torch/sparsity/attention_sparsity/test_vllm_plugin.py`
- `tests/gpu/torch/kernels/common/attention/test_triton_fa_paged.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_sparse_nm.py`
-
`tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py`

Manual / NEL eval validation:

- Served a ModelOpt exported sparse-attention checkpoint through
`examples/vllm_serve/vllm_serve_sparse_attn.py`.
- Launched RULER64K NEL evals on DFW with `coreai_nvfm_llm`.
- Current partial RULER64K prediction scores, before final `results.yml`
is written:
  - `skipsoftmax-only`: 98.59% over 4500 flushed samples
  - `skipsoftmax-r0.7`: 99.70% over 1000 flushed samples
  - `skipsoftmax-r0.9`: 99.70% over 1000 flushed samples

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - no new
PIP dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A - documentation, examples, and tests are updated for this vLLM
integration path.

### Additional Information

Follow-up work:

- Validate and enable CUDA graph capture for the sparse vLLM path.
- Add combined sparse attention + quantization serving once the combined
path is tested.
- Investigate whether skip-softmax should also be enabled during decode.

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-05-20 14:11:25 -07:00
Shengliang Xu e4dc0205d1 [OMNIML-4775] Move built-in PTQ quantization configs to YAML (#1423)
### What does this PR do?

Type of change: refactor

This PR moves the built-in PTQ quantization config definitions out of
hard-coded Python dictionaries and into schema-backed YAML config files,
and factors shared blocks into reusable composable snippets.

- Adds reusable numeric config snippets under
`modelopt_recipes/configs/numerics/`.
- Adds YAML presets for the built-in model PTQ configs under
`modelopt_recipes/configs/ptq/presets/model/`.
- Adds YAML presets for KV-cache quantization configs under
`modelopt_recipes/configs/ptq/presets/kv/`.
- Adds YAML presets for the Diffusers-specific PTQ configs under
`modelopt_recipes/configs/ptq/presets/diffusers/` and re-points
`examples/diffusers/quantization/config.py` constants at them via
`load_config`.
- Adds reusable KV quantization units (`kv_fp8_affine`, `kv_nvfp4`,
`kv_nvfp4_affine`, `kv_nvfp4_rotate`, `kv_*_cast` variants) under
`modelopt_recipes/configs/ptq/units/`.
- Adds reusable model-side units following the
`component_numerics[_type]` convention:
- `attention_qkv_fp8` — FP8 E4M3 on attention q/k/v bmm and softmax
quantizers; shared by `model/` and `diffusers/` `nvfp4_fp8_mha` presets.
- `block_sparse_moe_nvfp4` — NVFP4 W4A4 on `*block_sparse_moe*`
weight/input quantizers; shared by `nvfp4_mlp_only`,
`nvfp4_experts_only`, `nvfp4_omlp_only`.
- `experts_nvfp4` — NVFP4 W4A4 on `*.experts.*` weight/input quantizers;
shared by `nvfp4_mlp_only` and `nvfp4_experts_only`.
- Switches the existing 5 NVFP4 presets (default + awq lite/clip/full +
svdquant) and 4 mamba_moe presets to `$import` the existing
`w4a4_nvfp4_nvfp4` / `w8a8_fp8_fp8` units instead of re-inlining the
same weight+input quantizer pairs.
- Moves the recently-added `W4A16_NVFP4_CFG` to YAML
(`presets/model/w4a16_nvfp4.yaml`) composed from the existing
`units/w4_nvfp4` snippet.
- Updates `modelopt.torch.quantization.config` built-in config constants
to load `QuantizeConfig` objects from YAML with `load_config(...,
schema_type=QuantizeConfig).model_dump(exclude_unset=True)` via a new
`_load_quantize_config_dict` helper; the constants remain plain
`dict[str, Any]` for backwards compatibility with consumers that do
mapping-style mutation (e.g. `entry["cfg"]` assignment).
- Simplifies the cfg-list loader (`_load_quantizer_cfg_dict_list`) down
to a 4-line list/single normalization now that the three call sites all
load schema-typed YAMLs.
- Adds/updates recipe loader coverage for built-in schema-backed config
snippets.

### Latent-bug fixes surfaced by the refactor

Two small correctness fixes are included alongside the mechanical
refactor; flagging them explicitly:

- **`examples/diffusers/quantization/quantize.py`** — adds an explicit
`base_cfg = copy.deepcopy(base_cfg)` before applying runtime overrides.
The existing `# Build a fresh config dict so we never mutate the global
constants` comment had been aspirational only; in practice
`reset_set_int8_config` accumulated `PercentileCalibrator` entries into
`mtq.INT8_SMOOTHQUANT_CFG`/`INT8_DEFAULT_CONFIG` across repeated calls,
and `set_quant_config_attr` added `trt_high_precision_dtype` keys into
globally-shared cfg dicts. The deepcopy makes the code match the
comment.
- **`choices` set in `modelopt/torch/quantization/config.py`** — adds
`MXFP6_DEFAULT_CFG` and `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` to the
documented public set of valid `mtq.*_CFG` names. Both constants exist
on main but were missing from `choices`, so CLIs that gate on
`mtq.config.choices` (e.g., `hf_ptq.py --qformat`) couldn't reach them
even though the configs themselves were fully supported.

### Usage

Existing Python imports continue to work:

```python
import modelopt.torch.quantization as mtq

cfg = mtq.FP8_DEFAULT_CFG
model = mtq.quantize(model, cfg, forward_loop)
```

The built-in constants are plain `dict[str, Any]` (sparse — only
explicitly-set fields are present), but their definitions now come from
YAML snippets and presets composed through the existing `$import`
system.

Reusable YAML snippets can be composed through `$import`, for example:

```yaml
# modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig
imports:
  base_disable_all: configs/ptq/units/base_disable_all
  w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
  default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers

algorithm: max
quant_cfg:
  - $import: base_disable_all
  - $import: w4a4_nvfp4_nvfp4
  - $import: default_disabled_quantizers
```

### Testing

Local checks run:

- `nox -s "unit-3.10(torch_211, tf_latest)"` — 2329 passed, 12 skipped.
- `nox -s pre_commit_all` — all hooks pass (ruff check / ruff format /
mypy / YAML format / license / bandit / markdownlint).
- YAML parse + `$import` resolution sanity check across all changed
config files.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ Existing built-in Python config
constants keep the same public names and dict semantics.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ Adds/updates recipe loader
coverage for schema-backed built-in snippets.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ❌

### Additional Information

This PR was previously stacked on #1405, which has since merged to
`main`. The branch has been rebased onto `main` and no longer depends on
any other open PR.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Many new quantization numeric configs and PTQ presets added
(INT4/INT8/MXFP4/MXFP6/MXFP8/MXINT8/NVFP4), plus Diffusers, KV-cache
(affine/cast/rotate) and MLP/MoE-targeted presets.

* **Refactor**
* Presets and shared snippets migrated to schema-backed YAML sources and
centralized loading; INT8 percentile calibration avoids mutating shared
base configs.

* **Tests**
* Tests now discover packaged config snippets at runtime and validate
import/append behaviors.

* **Documentation**
  * Presets README and numerous header descriptions updated.

* **Chores**
  * Minor typing and script improvements.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1423?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-05-20 08:21:25 -07:00
jingyu-ml c7966119eb Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do?

Type of change: re-org code <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ We changed the import path
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅
<!--- Only for new features, API changes, critical bug fixes or backward
incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Calibration support for skip-softmax multi-threshold measurement in
sparse attention.
  * N:M sparse softmax masking and helpers for sparsity-aware attention.

* **Chores**
* Reorganized and consolidated kernel/backends for quantization and
sparsity to a unified kernels layout, updating tests and examples to
match.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-22 23:34:56 +00:00
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-19 12:20:14 +05:30
jingyu-ml feec81ad2b Add the Skip softmax for diffusion (#1166)
### What does this PR do?

Type of change: new feature, new example <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

## Summary

- Add skip-softmax sparse attention (BLASST) for diffusion models via
dedicated Triton kernels — an inference kernel with tile skipping and a
calibration kernel with vectorized multi-threshold sparsity measurement
- Add `triton_skip_softmax` method with exponential model calibration
(`scale_factor = a * exp(b * sparsity)`) and log-space fitting for
diffusion models
- Add Triton kernel backends for diffusers and LTX attention dispatch
- Fix calibration to skip RULER dataset generation when user provides
their own `forward_loop` (required for non-LLM models)

## Changes

### Triton kernels (`modelopt/torch/kernels/triton_fa.py`)
- **`_attn_fwd`**: Forward kernel with optional tile skipping — tiles
whose max attention score is far below the running softmax max are
skipped entirely (no V load, no softmax, no accumulation). Runtime
sparsity measurement via atomic counters.
- **`_attn_fwd_calibrate`**: Calibration kernel that computes full
attention while measuring how many tiles would be skipped at each of N
thresholds simultaneously. Uses per-program output buffers (zero atomic
contention) and vectorized multi-threshold comparison.
- **`attention()`** / **`attention_calibrate()`**: Python wrappers for
inference and calibration kernels.

### Kernel backends
(`modelopt/torch/sparsity/attention_sparsity/kernels/`)
- **`diffusers_triton_attention.py`**: Registers `modelopt_triton`
backend in diffusers' attention dispatch. Handles [B, S, H, D] → varlen
layout conversion, calibration/inference mode switching, thread-local
configuration, and counter accumulation.
- **`ltx_triton_attention.py`**: Patches `ltx_core.Attention` modules
for Triton dispatch with the same calibration/inference modes.

### Method
(`modelopt/torch/sparsity/attention_sparsity/methods/triton_skip_softmax.py`)
- `TritonSkipSoftmaxMethod`: Context managers for calibration (→
calibration kernel) and inference (→ forward kernel with tile skipping).
Three threshold priority levels: raw threshold > calibrated scale_factor
> static threshold.

### Calibration
(`modelopt/torch/sparsity/attention_sparsity/calibration/`)
- **`calibrator.py`**: `DynamicThresholdCalibrator` with `fit_logspace`
option — fits exponential model in log space (minimizes relative error)
for diffusion models where scale_factors span many orders of magnitude.
Records observed sparsity range for extrapolation warnings.
- **`calibrate.py`**: Skips RULER dataset when `forward_loop` is
provided; passes `fit_logspace` through from config.

### Config & conversion
- **`config.py`**: `CalibrationConfig.fit_logspace` field (default
False, recommended True for diffusion models).
`skip_softmax_raw_threshold` field for direct threshold mode.
- **`conversion.py`**: Auto-registers diffusers/LTX Triton backends on
`sparsify()`. Updated summary display.

### Example
- **`wan22_skip_softmax.py`**: End-to-end example for WAN 2.2 5B/14B
with baseline, raw-threshold, and calibrated modes. Supports runtime
sparsity reporting.

## Threshold modes

| Mode | How it works | Use case |
|------|-------------|----------|
| **Raw threshold** (`--raw-threshold -0.7`) | Passed directly to kernel
as `skip_threshold_log2` | Quick testing, sweeps |
| **Calibrated** (`--calibrate --target-sparsity 0.5`) | `scale_factor =
a * exp(b * target)`, then `threshold = scale_factor / seq_k` at runtime
| Production use with seqlen adaptation |
| **Static** (default `skip_softmax_threshold=0.1`) | `log2(lambda) *
sm_scale` | Fallback |

## Usage

```bash
# Fixed raw threshold (no calibration)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --raw-threshold -0.7 \
    --prompt "A cat playing piano" --output out.mp4

# With calibration (log-space fit for diffusion models)
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --calibrate --target-sparsity 0.5 \
    --prompt "A cat playing piano" --output out.mp4

# Dense baseline for comparison
python examples/diffusers/sparsity/wan22_skip_softmax.py \
    --model-path /path/to/Wan2.2-T2V-A14B-Diffusers \
    --baseline \
    --prompt "A cat playing piano" --output baseline.mp4
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added skip-softmax sparse attention support for Diffusers models,
enabling efficient video generation
* Added support for both eager and Triton attention backends for sparse
attention
* Added new example script for Wan 2.2 text-to-video generation with
sparse attention optimization

* **Documentation**
* Updated documentation with sparse attention configuration guide and
usage examples

* **Tests**
* Added comprehensive unit tests for kernel backend registration and
skip-softmax functionality
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-18 06:54:19 +00:00
Keval MorabiaandClaude Sonnet 4.6 80a77d1cc0 Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary

LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party
dependency developed and provided by Lightricks. It is governed by the
[LTX Community License
Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE),
**not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per
legal guidance, all integration points must clearly surface this to
users.

- Add `[!WARNING]` license notice blocks at the top of all LTX-2-related
READMEs (`examples/diffusers`, `examples/diffusers/distillation`,
`examples/windows/diffusers/qad_example`)
- Add `warnings.warn(UserWarning)` at every LTX package import site in
Python files, covering both top-level and lazy imports:
  - `examples/diffusers/distillation/distillation_trainer.py`
  - `examples/diffusers/quantization/calibration.py`
  - `examples/diffusers/quantization/pipeline_manager.py`
-
`examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py`
  - `modelopt/torch/export/diffusers_utils.py`
  - `modelopt/torch/quantization/plugins/diffusion/ltx2.py`
- Add license notice comment to `requirements.txt` files that list LTX
packages, so the obligation is visible at install time
- Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering
variants like `requirements-dev.txt`) are owned by
`@NVIDIA/modelopt-setup-codeowners` regardless of location, via a
last-match-wins rule

**Design notes:**
- For library files (`diffusers_utils.py`, `ltx2.py`), the warning is
placed at the lazy import site inside functions — it fires only when
LTX-2 code paths are actually invoked, not at module import time, to
avoid polluting non-LTX users
- For example entry-point scripts that are LTX-2-only, the warning fires
at module load time (after all imports, to satisfy ruff E402)

## Test plan

- [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy,
markdownlint, bandit all clean)
- [ ] Verify warning appears at runtime when running an LTX-2
quantization or distillation example
- [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the
warning

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added third-party license notices across documentation and
requirements files clarifying LTX-2 packages are governed by the LTX
Community License Agreement rather than NVIDIA Model Optimizer's Apache
2.0 license.

* **Chores**
  * Updated code ownership configuration for requirements files.
* Added runtime warnings to notify when LTX-2 dependencies are accessed.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 11:04:29 +05:30
Shengliang Xu 1cceb950d6 [OMNIML-3689] PTQ quant_cfg semantic correction. Design in doc _quant_cfg.rst (#1094)
### What does this PR do?

#### Summary

Redesigns the `quant_cfg` configuration format in ModelOpt's PyTorch
quantization stack, replacing the previous dict-based format with an
**ordered list of typed `QuantizerCfgEntry` dicts**.

##### Motivation

The old `quant_cfg` dict had several pain points:
- **Ambiguous precedence**: no explicit way to reason about which entry
wins when multiple keys match a quantizer
- **Mixed key namespaces**: wildcard paths and PyTorch class names lived
in the same dict level, requiring ad-hoc dispatch
- **Magic `"default"` key**: an implicit, undocumented catch-all that
was easy to misuse
- **Poor composability**: merging two configs required dict updates that
silently discarded keys
- **No YAML round-trip fidelity**: the nested structure couldn't be
expressed cleanly in YAML

##### New format

`quant_cfg` is now an ordered list of `QuantizerCfgEntry` TypedDicts.
Each entry has:
- `quantizer_name` *(required)*: `fnmatch` wildcard matched against
quantizer module names
- `cfg` *(optional)*: dict (or list of dicts) of
`QuantizerAttributeConfig` fields
- `enable` *(optional)*: toggles quantizer on/off independently of `cfg`
- `parent_class` *(optional)*: restricts match to quantizers whose
parent module is of the given PyTorch class (e.g. `"nn.BatchNorm2d"`)

Entries are applied in list order; later entries override earlier ones.
The canonical pattern is deny-all first (`_base_disable_all`), then
selectively re-enable and configure, then apply standard exclusions
(`_default_disabled_quantizer_cfg`).

##### Changes

**Core library (`modelopt/torch/quantization/`)**

- **`config.py`**:
- Added `QuantizerCfgEntry` TypedDict (line 163) and
`find_quant_cfg_entry_by_path()` helper for exact-match lookup of
entries by path.
- Added `normalize_quant_cfg_list()` (line 1539) that converts legacy
formats (flat dict, single-key dicts, `nn.*`-scoped dicts, `"default"`
key) to canonical `QuantizerCfgEntry` lists. After normalization every
entry is guaranteed to have explicit `quantizer_name`, `enable`, and
`cfg` keys.
- Converted `_default_disabled_quantizer_cfg` and
`_mamba_moe_disabled_quantizer_cfg` from dicts to lists of
`QuantizerCfgEntry`.
- Added `_base_disable_all` (line 205): canonical deny-all entry
(`[{"quantizer_name": "*", "enable": False}]`).
- Converted all ~30 built-in config constants (`INT8_DEFAULT_CFG`,
`FP8_DEFAULT_CFG`, `NVFP4_DEFAULT_CFG`, etc.) to list format using
`*_base_disable_all` and `*_default_disabled_quantizer_cfg` unpacking.
- KV-cache configs (`FP8_KV_CFG`, `NVFP4_KV_CFG`, etc.) are now minimal
lists designed to be concatenated with a primary config — they
intentionally omit `_base_disable_all` and `"algorithm"`.
- Added two `QuantizeConfig` Pydantic field validators: a
`mode="before"` validator that calls `normalize_quant_cfg_list()`, and a
`mode="after"` validator that validates `cfg` dicts against
`QuantizerAttributeConfig`.
- Updated `need_calibration()` to iterate the normalized list instead of
the old dict.
- Changed `QuantizeQuantCfgType` alias from `dict[str | Callable, ...]`
to `list[QuantizerCfgEntry]`.

- **`conversion.py`**:
- Rewrote `set_quantizer_by_cfg()` (line 217) to iterate the list
directly. Each entry's `parent_class` is resolved via
`QuantModuleRegistry[parent_class_name]` (the existing `_DMRegistryCls`
registry).
- Added `set_quantizer_attributes_full()` (line 314): full replacement
of quantizer attributes from a `QuantizerAttributeConfig`. Unspecified
fields revert to defaults, enforcing entry atomicity. Can also upgrade
`TensorQuantizer` → `SequentialQuantizer` or downgrade the reverse.
- Added `set_quantizer_attributes_partial()` (line 384): merges a
partial `dict` of attributes into existing quantizer state. Does NOT
change quantizer structure. Used for enable-only entries.
- Added `set_quantizer_by_cfg_context()` context manager (line 447) that
temporarily applies a `quant_cfg` list and restores original quantizer
state on exit.
- Deprecated `set_quantizer_attribute()` (line 525) with a
`DeprecationWarning` pointing to the new functions.

- **`tensor_quantizer.py`**:
- `TensorQuantizer.set_from_attribute_config()`: narrowed type hint from
`dict` to `dict[str, Any]`.
- Added `_axis_setter` and `_block_sizes_setter` custom setters so that
`axis` and `block_sizes` changes properly propagate to the calibrator
and maintain mutual exclusivity.
- `SequentialQuantizer.set_from_attribute_config()`: narrowed signature
to `list[QuantizerAttributeConfig] | list[dict[str, Any]]` (removed the
old union with single values).

- **`algorithms.py`**:
- Updated `_match_quantizer_cfg()` to iterate the list and return
`(matched_cfg, matched_enable)` tuple with last-match-wins.
- Updated `_cfg_to_dict()`, `estimate_quant_compression()`, and
`QuantRecipe` to work with the list-based format.
- Updated `get_auto_quantize_config()` to emit list-format `quant_cfg`.

- **`model_quant.py`**: `disable_quantizer()` / `enable_quantizer()` now
call `set_quantizer_attributes_partial()` directly instead of the
deprecated `set_quantizer_attribute()`. Updated docstrings and code
examples to show the list format.

- **`utils/core_utils.py`**: `disable_lora_quantizers_in_config()` and
`update_quant_cfg_with_kv_cache_quant()` updated to append
`QuantizerCfgEntry` dicts to the list.

- **Other**: minor updates to `backends/fp8_per_tensor_gemm.py`,
`backends/nvfp4_gemm.py`, `compress.py`, `model_calib.py`,
`export/unified_export_hf.py`, and
`sparsity/attention_sparsity/conversion.py` to use the list format.

- **`onnx/llm_export_utils/quantization_utils.py`**: Updated
quantization config construction to use list format.

**YAML recipes (`modelopt_recipes/`)**

- Converted all 5 general PTQ recipes to the new list format:
  - `general/ptq/fp8_default-fp8_kv.yml`
  - `general/ptq/nvfp4_default-fp8_kv.yml`
  - `general/ptq/nvfp4_experts_only-fp8_kv.yml`
  - `general/ptq/nvfp4_mlp_only-fp8_kv.yml`
  - `general/ptq/nvfp4_omlp_only-fp8_kv.yml`
- Converted model-specific recipe:
`models/Step3.5-Flash/nvfp4-mlp-only.yaml`

**Documentation (`docs/`)**

- New guide: `docs/source/guides/_quant_cfg.rst` — comprehensive
reference covering entry format, ordering semantics, entry atomicity,
`enable` vs `cfg` independence, `parent_class` filtering, and common
patterns (deny-all-then-enable, customizing a built-in config, building
from scratch).
- Updated `_pytorch_quantization.rst` code examples to show the list
format with `copy.deepcopy` and `.append()`.
- Added `_quant_cfg.rst` to the quantization guide table of contents.

**Examples**

- Updated all quantization examples to use the list format:
`deepseek/ptq.py`, `diffusers/quantization/config.py`,
`llm_ptq/hf_ptq.py`, `llm_qat/main.py`, `vllm_serve/vllm_ptq_utils.py`,
`llm_autodeploy/run_auto_quantize.py`, `llm_eval/quantization_utils.py`,
`llm_ptq/example_utils.py`,
`windows/torch_onnx/diffusers/qad_example/sample_example_qad_diffusers.py`,
and 2 notebooks.

**Tests**

- New test file:
`tests/unit/torch/quantization/test_config_validation.py` — unit tests
for `need_calibration()`, `normalize_quant_cfg_list()` (new format,
legacy format conversions, error cases),
`find_quant_cfg_entry_by_path()`, `_match_quantizer_cfg()`, and
`QuantizeConfig` Pydantic validators.
- Extended `tests/unit/torch/quantization/test_quantize_cpu.py` with
tests for `set_quantizer_attributes_full()` (atomicity, parent_class
filtering, SequentialQuantizer creation), list ordering, enable-only
entry behavior, and end-to-end legacy dict format.
- Updated 20+ existing test files across `tests/unit/`, `tests/gpu/`,
`tests/gpu_megatron/`, and `tests/_test_utils/` to use the list format.

##### Backward compatibility

`normalize_quant_cfg_list()` is called automatically by the
`QuantizeConfig` Pydantic `mode="before"` validator, so existing code
passing the old dict-based format (flat dict like `{"*weight_quantizer":
{"num_bits": 8}}`, single-key dict lists, or `nn.*`-scoped dicts with
`parent_class` semantics) continues to work without modification. The
legacy `"default"` key is converted to `quantizer_name: "*"`.

`set_quantizer_attribute()` is preserved as a deprecated wrapper around
`set_quantizer_attributes_partial()`.

#### Test coverage

- **Unit tests**: new `test_config_validation.py` with tests for
normalization, validation, path lookup, and cfg matching. Extended
`test_quantize_cpu.py` with tests for full/partial attribute setting,
ordering, atomicity, and legacy backward compatibility.
- **System testing**:

```
python examples/llm_ptq/hf_ptq.py \
      --model Qwen/Qwen3-8B  \
      --recipe general/ptq/fp8_default-fp8_kv \
      --export_path=build/fp8_default-fp8_kv42  \
      --calib_size=16 \
      --batch_size=0 \
      --trust_remote_code \
      --export_fmt=hf
```

### Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-04-06 15:38:44 -07:00
jingyu-ml 4a5ef01acc Fixed the calib size bug (#1178)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed quantization configuration calibration batch counting to
properly account for partial final batches during the calibration
process.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-06 13:46:51 +00:00
Ajinkya Rasane 8f4c11aa44 Upgrade the TensorRT container version (#1112)
### What does this PR do?

Type of change: Container version update

- Upgraded the TensorRT container version to 26.02
- This supports TensorRT 10.15.1

### Testing
Unit and integrations tests pass

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated example guides to reference the latest TensorRT Docker image
version (26.02)
  * Added compatibility note for onnxruntime-gpu usage

* **Chores**
* Updated CI/CD workflows to use the latest TensorRT Docker image
version

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-03-25 04:00:52 +05:30
jingyu-ml 1070d895dc Flux2-Dev Quantization (#947)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Register Flux2Attention and Flux2ParallelSelfAttention in the
quantization plugin so bmm quantizers are patched (enables
--quantize-mha).
- Add Flux2-specific dummy input generation for HF checkpoint export.
- Guard check_conv_and_mha with hasattr for bmm quantizer attributes 

## Usage
<!-- You can potentially add a usage example below. -->

```bash
python quantize.py \
    --model flux2-dev \
    --model-dtype BFloat16 \
    --format fp4 --batch-size 2 --calib-size 1 \
    --n-steps 20 --quantized-torch-ckpt-save-path ./flux2-dev-fp4.pt --collect-method default \
    --hf-ckpt-dir ./flux2-dev-fp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes<!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Flux2-dev model support with Flux2-compatible dummy input
generation and default inference params (768×1024, guidance scale 4.0).

* **Refactor**
* Made attention quantization disabling more robust by iterating
available quantizers before disabling.

* **Infrastructure**
* Flux2 attention components are now optional and registered only when
present to avoid import issues.

* **Tests**
* Added Flux2 test helpers and coverage validating Flux2 dummy input
shapes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-14 11:42:02 -05:00
jingyu-ml 812e8c60a2 Minor update on the LTX2 NVFP4 recipe (#1010)
### What does this PR do?

Type of change: minor code change <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

1. Update the default calibration dataset for LTX_VIDEO_DEV and LTX2
from Gustavosta/Stable-Diffusion-Prompts to nkp37/OpenVid-1M, which
provides video-specific captions better suited for video model
calibration.
2. update the default recipe for ltx2: first 3 and last 3 layers stays
at higher precision.

### Usage

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated default dataset configuration for LTX-Video and LTX2 models.
* Refined model filtering pattern for LTX-Video to support additional
model components.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-13 22:42:48 -05:00
Keval Morabia 1ccd945a51 Remove unused diffusers/cache_diffusion/pipeline and cuda-python dependency (#996)
`cuda-python` has mixed license and needs EStaff approval for usage. And
till 0.42, it was only used in
`examples/diffusers/cache_diffusion/pipeline` which has not been updated
in 9 months and not used anymore hence removing.

Also cherry-picked to `release/0.42.0` branch:
https://github.com/NVIDIA/Model-Optimizer/pull/984

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Removed TensorRT/ONNX deployment and inference tooling, related model
export/configuration, and runtime helpers from the cache-optimized
diffusion examples; removed the cuda-python example dependency.
* **Tests**
* Removed the example benchmarking script and its associated benchmark
test.
* **Documentation**
* Strengthened dependency-review, security, and PR guidance; updated PR
template and contributing documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-07 00:37:19 +05:30
jingyu-ml 37d3f10cbd To support LTX2 ComfyUI format (#972)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

- Added a flag merged_base_safetensor_path to the example code so that
user can export the ComfyUI style ckpt.

### Usage

```bash
python quantize.py \
    --model ltx-2 --format fp4 --batch-size 1 --calib-size 32 --n-steps 40 \
    --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors \
    --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors \
    --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors \
    --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized \
    --extra-param fp8transformer=true \
    --quantized-torch-ckpt-save-path ./ltx-2-transformer.pt \
    --hf-ckpt-dir ./LTX2-NVFP4/ \
    --extra-param merged_base_safetensor_path=./ltx-2-19b-dev-fp8.safetensors
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added new command-line parameters documentation for LTX-2 FP4
quantization examples (--hf-ckpt-dir and merged_base_safetensor_path
configuration options)

* **Improvements**
* Enhanced quantization pipeline to support conditional export behavior
based on model type
* Expanded LTX-Video model filtering patterns for more comprehensive
block detection

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-06 05:06:36 +00:00
Keval Morabia 82f1d216d1 Add Security and IP related contributing guide and configure coderabbit to catch such issues (#935)
### What does this PR do?

- Add Security related coding practices in `SECURITY.md` and merge with
`2_security.rst`
- Update `CONTRIBUTING.md` for instructions to follow if copying code
from other repositories
- Update PR template
- Cleanup dependency files
- New API `mto.load_modelopt_state` doing the insecure `torch.load(f,
weights_only=False)` instead of doing it separately everywhere. This
also allows us to later improve the input validation for
`modelopt_state_path` or use safer alternatives to `torch.load`

### Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=True)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ <!--- Mandatory -->
- Did you write any new necessary tests?: NA <!--- Mandatory for new
features or examples. -->
- Did you add or update any necessary documentation and update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
NA <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded and reorganized security guidance and contributor procedures;
updated PR template and several READMEs with clearer security,
submission, and installation instructions
* Replaced an older security document with an enhanced, centralized
security guidance

* **Chores**
* Adjusted example dependency lists and optional extras (adds, removals,
and version constraints)
* Enabled automated incremental reviews, added pre-merge security
checks, and introduced a knowledge-base of coding/security guidelines
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 03:49:15 +05:30
jingyu-ml 2905cb0f2e Updated the diffusion config issue and more test cases (#937)
## What does this PR do?

**Type of change:** new tests, Bug fix <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

**Overview:** 

- **Fixed the INT8 config issue**

- **Add HF checkpoint export test coverage**

1. The `--hf-ckpt-dir` export path had zero test coverage. This MR adds
tests at two levels:
2. Unit tests (tests/unit/torch/export/test_export_diffusers.py):
- Extended test_export_diffusers_real_quantized to parametrize over
INT8, INT8 SmoothQuant, FP8, and FP4 configs
- (previously only FP8). This gives 3 models x 4 configs = 12 test
cases.

3. GPU integration tests
(tests/gpu/torch/export/test_export_diffusers_hf_ckpt.py)
- New file testing the full quantize.py --hf-ckpt-dir pipeline via
subprocess with 4 combos:
- SDXL INT8 smoothquant min-mean (the exact scenario that triggered the
bug)
    - Flux INT8 smoothquant min-mean
    - SDXL FP8
    - Flux FP4

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**:No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Tests**
* Added test coverage for exporting Diffusers models with Hugging Face
checkpoints across multiple quantization formats (INT8, FP8, FP4)
* Extended quantization export testing to validate multiple
configuration scenarios

* **Chores**
* Refined INT8 quantization configuration with improved calibrator
support for convolution layers

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-28 23:54:02 +05:30
jingyu-ml d78797b466 Update the LTX2 API calls during the calibration (#926)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Update LTX-2 integration to match latest upstream API

1. The LTX-2 codebase removed/replaced several APIs. This MR updates all
affected files:
2. Replace cfg_guidance_scale with MultiModalGuiderParams: The pipeline
__call__ no longer accepts a single cfg_guidance_scale float. It now
requires two MultiModalGuiderParams objects (video_guider_params and
audio_guider_params) that control CFG, STG, rescale, cross-modality
guidance, and skip-step settings. Updated in ltx-2.py, ltx-2-fp8.py,
ltx-2-onestage.py, calibration.py, and models_utils.py.
3. Replace fp8transformer with QuantizationPolicy: The
TI2VidTwoStagesPipeline constructor no longer accepts the fp8transformer
boolean flag. FP8 quantization is now configured via
quantization=QuantizationPolicy.fp8_cast(). Updated in ltx-2-fp8.py and
pipeline_manager.py (with backwards-compatible support for the old
--extra-param fp8transformer=true CLI flag).
4. Remove DEFAULT_CFG_GUIDANCE_SCALE constant: Replaced by
DEFAULT_VIDEO_GUIDER_PARAMS and DEFAULT_AUDIO_GUIDER_PARAMS in all
import sites.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
python quantize.py --model ltx-2 --format fp4 --batch-size 1 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Updates**
  * Default resolution for LTX2 models adjusted to 768x1280
* Guidance parameter configuration updated for video and audio pipelines
  * FP8 quantization parameter handling refined

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-24 14:23:28 -08:00
mxinO ca1f9687bd [OMNIML-3505] LTX-2 Distillation Trainer (#892)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 
Adding LTX-2 distillation trainer.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
accelerate launch \
    --config_file configs/accelerate/fsdp.yaml \
    --num_processes 8 \
    distillation_trainer.py --config configs/distillation_example.yaml
```

See readme for more details.

## Testing
Run training with single/multiple nodes.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## New Features
* Added distillation training support for LTX-2 models with quantization
integration.
* Introduced comprehensive documentation and example configurations for
distillation workflows.
* Includes multi-GPU and multi-node training setup with distributed
training support and customizable configuration templates.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Meng Xin <mxin@nvidia.com>
2026-02-14 13:28:21 -08:00
jingyu-ml 10efcb65f6 [3.1/4] Diffusion Quantized ckpt export - WAN 2.2 14B (#855)
## What does this PR do?

**Type of change:** documentation <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

**Overview:** 

1. Added multi‑backbone support for quantization: --backbone now accepts
space- or comma-separated lists and resolves to a list of backbone
modules.
2. Introduced PipelineManager.iter_backbones() to iterate named backbone
modules and updated get_backbone() to return a single module or a
ModuleList for multi‑backbone.
3. Updated ExportManager to save/restore per‑backbone checkpoints when a
directory is provided, with {backbone_name}.pt files, and to create
target directories when missing.
4. Simplified save_checkpoint() calls to rely on the registered
pipeline_manager by default.


**Usage: **

```bash
python quantize.py --model wan2.2-t2v-14b --format fp4 --batch-size 1 --calib-size 32 \
    --n-steps 30 --backbone transformer transformer_2 --model-dtype BFloat16 \
    --quantized-torch-ckpt-save-path ./wan22_mo_ckpts \
    --hf-ckpt-dir ./wan2.2-t2v-14b 
```

Plans

- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml 
- [x] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml 

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Unified Hugging Face export support for diffusers pipelines and
components
  * LTX-2 and Wan2.2 (T2V) support in diffusers quantization workflow
* Comprehensive ONNX export and TensorRT engine build documentation for
diffusion models

* **Documentation**
* Updated to clarify support for both transformers and diffusers models
in unified export API
* Expanded diffusers examples with LoRA fusion guidance and additional
model options (Flux, SD3, SDXL variants)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-10 22:34:12 -08:00
jingyu-ml 2e43c80609 [2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This MR adds HuggingFace checkpoint export support for LTX‑2 by treating
TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the
stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling
back to writing model.safetensors when save_pretrained isn’t available.
It also preserves the original forward in DynamicModule patching
(_forward_pre_dm) so downstream callers can still invoke the pre-patched
forward implementation.

**Changes**

1. Added the calibration & quantization support of the LTX2, even with
FP8 precision.
2. Preserve original forward before `DynamicModule` patching: when
patching forward, we now stash the pre-patched implementation in
`self._forward_pre_dm` (once) so downstream code can still call the
original forward, then re-bind forward to the class implementation. This
is needed for the LTX2 FP8 calibration.
3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats
ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a
“diffusion-like” object and routes it through
_export_diffusers_checkpoint() (import guarded; no hard dependency).
4. Generalized component discovery: introduced
get_diffusion_components() (aliasing the old get_diffusers_components)
to support non-diffusers pipelines; for LTX‑2 it returns only
stage_1_transformer.
5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy
forward generator (generate_diffusion_dummy_forward_fn) that builds
minimal LTX Modality inputs (including correct timesteps broadcasting)
so shared-input hooks can run and fuse QKV when applicable.
6. Export fallback for non-save_pretrained modules: when a component
lacks save_pretrained (LTX‑2 transformer), export now writes
model.safetensors + minimal config.json instead of pytorch_model.bin.

Plans

- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml 
- [ ] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml 

## Usage
<!-- You can potentially add a usage example below. -->
```bash
python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added LTX-2 video model support with complete quantization and export
pipeline integration
* Introduced `--extra-param` CLI option for flexible model configuration
and parameter passing
* Enhanced export capabilities with broader diffusion model
compatibility

* **Chores**
* Changed default model data type from Half to BFloat16 for improved
numerical stability

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-04 10:41:33 +00:00
jingyu-ml 668b8a19e8 [1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This PR adds support for exporting quantized diffusers models (DiT,
Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling
deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM.

**Changes**

New file: `diffusers_utils.py`
- Dummy input generation for various diffusion models
- Pipeline component extraction helpers
- QKV projection detection and grouping
- `hide_quantizers_from_state_dict()` context manager for clean saves

Refactored: `unified_export_hf.py`
- New `_fuse_qkv_linears_diffusion()` for QKV amax fusion
- `_export_diffusers_checkpoint()` to export full pipelines (models +
tokenizers + schedulers etc.)

Plans

- [x] [1/3] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml 
- [ ] [3/3] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml

## Usage
<!-- You can potentially add a usage example below. -->
```
mtq.quantize(pipe, quant_config, forward_call)
export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
* Added HuggingFace checkpoint export support for quantized diffusion
models with configurable output directory
* Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint
export destination
* Extended export functionality to support selective component exports
from diffusion pipelines
* Enhanced quantized model export with improved component handling and
multi-stage checkpoint generation

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:11:09 +00:00
jingyu-ml e6e4efd61e [0.5/3] Diffusion ckpt export for NVFP4 & FP8 (#783)
See https://github.com/NVIDIA/Model-Optimizer/pull/781

This is the MR that only includes the refactoring of the llm export,
please ignore the change on quantize.py from the diffusion example.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added `--hf-ckpt-dir` CLI option to save checkpoints in HuggingFace
format
  * Enabled support for exporting Diffusers-based pipelines
* Unified export system now handles both transformer and diffusion model
architectures

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-15 13:49:22 -06:00
Asha Anoosheh 510451322c Streamline KD & QAD transformers Trainers (#708)
## What does this PR do?

**Type of change:** ? Refactor and stabilization

**Overview:** 
* Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so
that we can remove multiple restrictions

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-01-10 01:28:56 +00:00
Shengliang Xu fd66be2bec [NVBug 5659126] dummy inputs are kwargs maps (#676)
## What does this PR do?

bug fix

**Overview:**

[NVBug 5659126] dummy inputs are kwargs maps, and rename the
generate_... function to more explicitly reflect the return value type


## Testing

`python diffusion_trt.py --model flux-dev --override-model-path
/models/FLUX.1-dev --torch --benchmark --skip-image
`

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2025-12-12 00:57:36 +05:30
Keval Morabia 53a2ddebab Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer
(OMNIML-3033)
- [x] Mention in Latest News section with date on the date of merging
this PR (12/08)

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-07 12:33:21 +05:30
Ajinkya Rasane c6c9905929 [OMNIML-2244] Create the nvfp4 quant exporter (#636)
## What does this PR do?

**Type of change:**
New feature

**Overview:** 
- Implemented the NVFP4QuantExporter
- Deprecated fp4qdq_to_2dq
- Updated tests

## Usage

```python
python torch_quant_to_onnx.py --quantize_mode=nvfp4 \
	--onnx_save_path=vit_base_patch16_224.nvfp4.onnx \
	--calibration_data_size 64 \
	--batch_size 128
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

```
python evaluate.py --onnx_path=vit_base_patch16_224.nvfp4.onnx \
	--model_name=vit_base_patch16_224 \
	--results_path=./results.txt \
	--batch_size 128
```

Results:
```
The top1 accuracy of the model is 84.39%
The top5 accuracy of the model is 97.312%
Inference latency of the model is 7.22412 ms
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No
- Deprecated fp4qdq_to_2dq
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-12-04 22:56:07 +00:00
Shengliang Xu cfb247a1f3 [NVBug 5659126] The same workaround for RMSNorm exporting for diffusers >=0.35.0 (#642)
## What does this PR do?

**Type of change:** ?

Bug fix

**Overview:** ?


For the trt_diffusions script


## Testing

python diffusion_trt.py --model flux-dev --benchmark --skip-image

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2025-12-04 05:16:36 +00:00
Shengliang Xu 8844a2b2bc Support attention quantization for diffusers >= 0.35.0 (#608)
## What does this PR do?

**Type of change:**

new feature

**Overview:** ?

Attention mechanism has changed from diffusers 0.35.

Many model attentions are now subclass of a new Mixin class:
AttentionModuleMixin, which is not a sub class of Attention

To fix it, patch the mixin class by forcing to use native attention
impl so the existing function monkey patch still work.


## Testing

manual quant of Wan, Flux

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2025-12-03 11:20:48 -08:00
Shengliang XuandKeval Morabia e0a6efbe70 fix trt engine building of the diffusers pipelines (#637)
## What does this PR do?

**Type of change:**

Bug fix

**Overview:**

1. The diffusion_trt.py needs the dynamic_shapes when running trtexec
for engine building. A previous change altered the format of
dynamic_shapes, fix it here.

2. the dynamic_shapes logic gets cleaned up. The existing logic is very
confusing

3. recover min-batch_size config for some pipelines. Previously some
pipelines set the min batch_size to be > 1, which was odd, so a previous
change sets them to be 1, but it turns out the oddity has a reason, the
trt engine building fails with the altered batch_size min/opt, thus
recover them.


## Testing

pytest tests/examples/diffusers

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2025-12-03 10:34:15 +00:00
Shengliang Xu d0b0c0fd46 Fix extra args and --component-dtype default value (#605)
## What does this PR do?

**Type of change:** ? 

Bug fix

**Overview:** ?

1. We are passing incorrect extra args to pipeline inference.
2. Need a default empty list for the list argument component-dtype

## Testing

Passed SDXL pipeline without component dtype

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2025-12-01 09:35:53 -08:00
jingyu-ml fa84955283 Fixed the cache diffusion ci/cd (#620)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Fixed the cache diffusion CI/CD issue related to Torch
2.9.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
pytest tests/examples/diffusers/test_cache_diffusion.py::test_sdxl_benchmarks -v -s
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2025-11-30 16:57:19 +05:30
Shengliang Xu e35db173b1 Support Wan2.2 t2v diffusers quantization (#556)
## What does this PR do?

**Type of change:**

new feature

**Overview:**

Support Wan2.2 t2v diffusers quantization
1. fix torch2.9 support

2. add Wan2.2 t2v diffusers pipeline quantization

   Main difference of the Wan2.2 pipeline comparing to exisiting
   pipelines is that there are 2 backbone models for denoising. For the
   quantization therefore we need to quantize both of them.

   However, it turns out our base library does not well support
   quantization of multiple models in the same time. Therefore, the
   change here just stick to quantize a single model each time, and then
   run the quantization multiple times.

   So, we need to allow users to pick which backbone to quantize,
   therefore adding a new argment for it

3. add a workaround for the exporting ONNX issue when we upgrade
   diffusers to >= 0.35.0. The issue lies is the exporting of the
   torch.nn.RMSNorm. Some pipelines in the diffusers > 0.35.0 use the
   torch version RMSNorm while before that they use the diffusers' own
   version of RMSNorm. It turns out they are directly replacable so the
   workaround is to simply replace the torch RMSNorm usages with
   diffusers RMSNorm. But we need to fix it properly soon by porting our
   ONNX export to be based on torch dynamo instead of torchscript. Issue
reported from external user:
https://github.com/NVIDIA/TensorRT-Model-Optimizer/issues/262

4. allow use of a prompts file, which is simply a text file with a list
   of prompts, one prompt each line

5. allow each component of a pipeline to have different dtype accuracy.
   added a new list stype command line arg --component-dtype for this.
   example: --component-dtype vae:Float

6. print the summary of the quantized model so users can capture issues
from
   log

## Usage

python quantize.py \
    --model wan2.2-t2v-14b \
    --format fp8 \
    --batch-size 4 \
    --calib-size 64 \
    --n-steps 20 \
    --backbone transformer \
    --model-dtype BFloat16 \
    --component-dtype vae:Float \
    --trt-high-precision-dtype BFloat16 \
    --quantized-torch-ckpt-save-path ./wan_transformer.pt \
    --onnx-dir wan-transformer-onnx \
    --prompts-file wan-prompts.txt



## Testing

Tested SDXL_BASE, LTX_VIDEO_DEV, WAN22_T2V

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information

https://github.com/NVIDIA/TensorRT-Model-Optimizer/issues/262

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2025-11-19 17:59:04 -08:00
Keval Morabiaandajrasane c333d36a85 Enable torch 2.9 tests in CICD + Diffusers fixes (#561)
As title

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-11-19 17:11:09 +05:30
ajrasane 8188a01b38 [NVBUG: 5619158] Optimize memory usage for diffusion_trt.py (#547)
## What does this PR do?

**Type of change:** 
Minor code change

**Overview:**
- Delete backbone after Device Model creation
- Add assertion for torch compile
- Update dummy input generation function

## Testing
```
python diffusion_trt.py --model flux-dev --benchmark --skip-image 
python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp8_autodeploy_fake.pt
python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp4_autodeploy_fake.pt
python diffusion_trt.py --model flux-dev --benchmark --skip-image --torch 
python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp8_autodeploy_fake.pt --torch 
python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp4_autodeploy_fake.pt --torch 
python diffusion_trt.py --model flux-dev --benchmark --skip-image --torch --torch-compile 
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-11-12 17:31:39 -05:00
ajrasane e74a46890d Update benchmarking for diffusers (#487)
## What does this PR do?

**Type of change:** 
Example update

**Overview:** 
- Optimize the benchmarking function in the diffusers example

```python
python diffusion_trt.py --model flux-dev --benchmark --model-dtype BFloat16 --skip-image --torch
```

## Testing
```
Backbone-only inference latency (BFloat16):
  Average: 139.48 ms
  P50: 139.36 ms
  P95: 141.13 ms
  P99: 141.35 ms
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-11-10 13:06:35 +00:00
ajrasane 229053323a [NVBUG: 5619158] Enforce high precision model dtype for diffusion trt (#526)
## What does this PR do?

**Type of change:** 
Minor code change

**Overview:** 
- Select the high precision dtype directly based on model type - FP16
for Stable Diffusion models, BF16 for Flux


## Testing
```python
python diffusion_trt.py --model flux-dev --benchmark
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No (No option to specify
dtype while loading pipeline)
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-11-06 23:44:37 -08:00
ajrasane 41f2bf4941 Add option to real quantize the model (#473)
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2025-10-28 13:29:19 -07:00
vishalpandya1990 99c76fff26 Add SD3.5-medium quantization support in ModelOpt Diffusers example (#444)
Signed-off-by: vipandya <vipandya@nvidia.com>
2025-10-22 23:28:15 -07:00