mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-04 12:21:41 +08:00
pull-request/2597
48
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cf1f48fa0f |
[5565357] Fix SDXL NVFP4 export and performance (#2336)
### What does this PR do?
Type of change: Bug fix
Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:
- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.
For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.
This PR also changes shared exporter behavior:
- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.
Other model recipe configurations remain unchanged.
### Usage
```bash
python quantize.py \
--model sdxl-1.0 \
--model-dtype Half \
--trt-high-precision-dtype Half \
--format fp4 \
--block-size 16 \
--batch-size 2 \
--calib-size 128 \
--n-steps 20 \
--quantized-torch-ckpt-save-path ./sdxl-fp4 \
--onnx-dir ./onnx-sdxl-fp4
```
### Testing
- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
- 302 native block-scaled NVFP4 GEMM tactics;
- 38 native FP8 Conv tactics;
- no FP4 Q/K/V projections;
- all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [5565357]
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.
- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
|
||
|
|
6b4ad85849 |
Qwen-Image diffusers PTQ: FP8 / NVFP4 / NVFP4-SVDQuant HF checkpoints (#1706)
### What does this PR do? Type of change: New feature Adds **Qwen-Image** (`Qwen/Qwen-Image`, `QwenImageTransformer2DModel`) to the diffusers quantization example and exports HuggingFace checkpoints in three precisions — **FP8**, **NVFP4**, and **NVFP4 + SVDQuant** — through the unified HF export. - Registers `--model qwen-image` (lazy diffusers import; no `trust_remote_code`). - Transformer-block-range recipe: quantizes only the linears under `transformer_blocks`, keeping the **first 2 / last 2** blocks (and everything outside `transformer_blocks`) in original precision. Applied **before** calibration so SVDQuant never mutates the excluded blocks. Expressed with the top-level `enable` `QuantizerCfgEntry` field (disable-all → re-enable `transformer_blocks` → disable first/last-N). - SVDQuant export (AWQ-style): promotes quantizer-owned tensors to clean module-level safetensors keys at export time — `weight_quantizer.svdquant_lora_a/b → <module>.svdquant_lora_a/b` and `input_quantizer._pre_quant_scale → <module>.pre_quant_scale` — with a documented `NVFP4_SVD` `quantization_config` (`group_size`, `has_zero_point: false`, `pre_quant_scale: true`, `lora_rank`). **Core SVDQuant quantization code (`modelopt/torch/quantization`) is unchanged.** - Shared export-path change — **intentionally global** (applies to all diffusers exports — SDXL / Flux / Wan, not just Qwen; the full export suite was verified green on GB200): `hide_quantizers_from_state_dict` now strips quantizer state from *all* modules (not just quant-linears) so calibrated norm-layer input quantizers no longer leak `input_quantizer._amax`. (An earlier `max_shard_size` workaround was dropped after merging `main`: #1794 makes the ComfyUI layerwise-metadata post-processing a no-op unless explicitly opted in, so a default sharded export no longer hits the unsupported-sharded path.) ### Usage ```bash python examples/diffusers/quantization/quantize.py \ --model qwen-image --override-model-path <Qwen-Image> --model-dtype BFloat16 \ --format fp4 --quant-algo svdquant --lowrank 32 \ --calib-size 64 --n-steps 20 \ --hf-ckpt-dir <out> # FP8: --format fp8 --quant-algo max # NVFP4: --format fp4 --quant-algo max ``` ### Testing - Focused unit + example tests pass on GB200 (sm_100): block-range recipe, `NVFP4_SVD` config schema, SVDQuant forward/fold (LoRA stays on `weight_quantizer`), Qwen dummy-input / strict-QKV-fusion / promotion, pipeline loading, and the diffusers HF-export test for Qwen FP8 / NVFP4 / SVDQuant. - Full `tests/examples/diffusers/test_export_diffusers_hf_ckpt.py` is green (SDXL, Flux, Qwen, Wan2.2) — confirms the shared export changes do not regress other models. - End-to-end on the real `Qwen/Qwen-Image` (~20B): all three formats export valid HF checkpoints — only `transformer_blocks` 2..57 quantized, nothing outside, no quantizer/`_amax` leak, correct `weight_scale`(`_2`)/`input_scale`, promoted SVDQuant keys (rank-consistent shapes), and the expected `quantization_config`. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ <!-- live-model LoRA storage unchanged; existing exports unaffected --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: ❌ <!-- new-example feature; add a CHANGELOG.rst entry if required --> - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information All changes are confined to the diffusers example (`examples/diffusers/quantization`) plus the shared export path (`modelopt/torch/export`); the core quantization library is untouched. ### Follow-up (next step): fused-QKV SVDQuant for sglang / Nunchaku This export keeps attention `q/k/v` (and `add_q/k/v_proj`) as **separate** projections — the diffusers-native layout. That matches sglang's bf16 / FP8 / plain-NVFP4 paths (which also keep QKV separate) and ModelOpt/TRT-LLM consumers, so those load 1:1. sglang's **NVFP4-SVDQuant (Nunchaku)** path, however, builds a **fused** `to_qkv` with a *single* fused rank-r LoRA in Nunchaku-native format (`proj_down`/`proj_up`, `smooth_factor`, `wscales`/`wtscale`). Our per-projection tensors (`svdquant_lora_a/b` + `pre_quant_scale`; three independent rank-r decompositions) are not directly loadable there — and cannot be fused at load time, because the fp16 weight residual needed to derive a single fused rank-r is not preserved after export. **Planned next step:** an opt-in fused-QKV SVDQuant export mode that fuses q/k/v **before** SVDQuant calibration (yielding one rank-r over the fused weight) and emits a Nunchaku-compatible layout, enabling lower-latency fused-QKV inference in sglang. Tracked as a separate follow-up. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Qwen-Image (`QWEN_IMAGE`) model quantization and Diffusers export support * Added NVFP4_SVD (SVDQuant) export configuration support * Added transformer block-range quantization recipes (exclude first/last blocks) * **Bug Fixes** * Improved missing-pipeline error messaging for Qwen-Image * Prevented quantizer-related tensor/buffer leakage by promoting and cleaning quantizer outputs during export * **Tests** * Added Qwen-Image HF checkpoint export tests and offline fixtures * Added unit coverage for SVDQuant promotion/clean state-dict keys <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
33bfa8b1fe |
CI/Dev env bump (#1818)
### What does this PR do? Type of change: chore Bumps CI/dev tooling and test containers. **Container bumps** - NeMo test containers → 26.06 - TRT-LLM container → 1.3.0rc19 - transformers max version → 5.12 **Dev tooling bumps** - ruff bump 0.12.11 → 0.15.18 - mypy 1.17.1 → 2.1.0: enable new defaults (`local_partial_types`, `strict_bytes`); fix/narrow the errors newly surfaced by mypy 2.0 in 4 modules (rather than blanket-suppressing them); remove 2 stale `# type: ignore` comments - pre-commit 4.3.0 → 4.6.0 - sphinx 8.1 → 9.1 + sphinx-rtd-theme 3.0 → 3.1: add `suppress_warnings = ["ref.python"]` to fix cross-reference ambiguity error new in sphinx 9.x - trl fix for newly released 1.7 version **Bug fixes surfaced by the bumps** - sparsity (weight): make the weight mask DTensor-aware under FSDP. The transformers→5.12 bump routes the HF Trainer FSDP optimizer-state save through torch's DTensor-based `get_optimizer_state_dict`, which triggered `aten.mul.Tensor got mixed torch.Tensor and DTensor` in the dynamic `weight` getter. The mask is now distributed to the weight's mesh/placements before masking, cached, and rebuilt only when the sharding changes (invalidated on `set_mask`). Fixes the `llm_sparsity` example test. ### Testing - `pre-commit run --all-files` ✅ (including mypy 2.1.0) - `nox -s docs` ✅ - `tests/unit/torch/sparsity` + `tests/unit/torch/nas` ✅ - `llm_sparsity` GPU example test (FSDP path) verified in CI ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ✅ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Summary * **Documentation** * Refreshed Docker pre-requisites across examples to recommend updated container image tags (and streamlined some instructions). * **Bug Fixes** * Improved sparse weight mask handling for DTensor/FSDP by aligning and caching distributed masks. * Made TensorRT engine byte retrieval return immutable `bytes`. * Reduced Sphinx cross-reference warnings and tuned Transformers compatibility warning thresholds. * **Tests** * Increased default unit test timeout on Windows runners. * **Chores** * Updated CI workflow container tags and refreshed linting/typing/docs version pins, plus related mypy configuration. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
0081473861 |
Speed up slow unit/gpu/example tests (#1616)
### What does this PR do?
Type of change: test infrastructure / test speedups + CI stabilization
Make the test suite faster, `tests/unit` hermetic, and the CI lanes
stable, without losing coverage. Most changes are mechanical test/infra
edits; the buckets below cover the diff broadly.
**Unit tests — hermetic (no HF Hub):** toy local datasets/configs + the
local tiny tokenizer (with a checked-in chat template) replace Hub
assets; `tests/unit/conftest.py` enforces offline mode. Genuinely-HF
tests moved to `tests/gpu*` (e.g. the new
`tests/gpu/torch/utils/test_dataset_utils.py`). `CONTRIBUTING.md`
documents the hermetic-unit-test expectation.
**Unit-test speedups (no coverage loss):** speculative (disable CPU
torch.compile), calibrator (fewer histogram bins), ONNX conv/dynamo
(smaller shapes + representative subset), Ruler/sparse-attention (local
tokenizer), data-parallel autoquant (world size 4→2). Shared
`tiny_tokenizer` fixture. The distributed test helper now uses a private
`spawn` context instead of mutating the global start method (avoids
cross-test contamination).
**Rarely-used autonas/fastnas tests:** heavy parametrize cases marked
`@pytest.mark.manual`, one representative kept per test (fastnas
preferred); lighter sibling tests still cover core behavior. The legacy
FSDP1 NAS distributed test is also dropped: FastNAS/AutoNAS aren't used
with either FSDP1 or FSDP2, and FSDP1 is superseded by the newer FSDP2
API — so we keep a single FSDP2 case as a sanity check and drop FSDP1,
leaving the suite leaner.
**gpu_megatron:** deduplicate distributed worker pools by world_size
within a module (saves a redundant pool spin-up in multi-pool files;
module-scoped, no cross-module reuse).
**Example tests:** reduce per-test work via args that default to current
behavior (tests pass the fast values) — torch_onnx TRT optimization
level, diffusers calibration/inference steps, eagle `sample_size`,
megatron_bridge iters/calib, llm_sparsity data slice, export
safetensors-structure `calib_size`. Also enable the recently added
`gpt-oss` example tests in CI.
**Per-test timeouts:** `pytest-timeout` with a default per-directory
timeout (60s unit / 300s gpu+example) enforced in `tests/conftest.py`
(`timeout_func_only` in `pyproject.toml`), so a new test cannot silently
exceed the budget — an unmapped test dir crashes collection. A few
inherently slow tests carry explicit higher per-test overrides
(CUDA-compile, autotune, dflash).
**CUDA kernel pre-compilation:** a dedicated `tests/gpu/_extensions`
test JIT-builds the conv3d implicit-GEMM kernel up front (collected
before the functional tests in the same process) so the one-time build
cost no longer lands on — and time out — the first functional test that
uses it. Mirrored into the `llm_ptq`/`vlm_ptq` example lanes.
**Test relocation & optional-dependency guards:** vLLM sparsity plugin
test moved to `tests/gpu_vllm` (drops the in-test `importorskip`);
diffusers-dependent unit test guarded with `importorskip("diffusers")`
for partial-install lanes; `gpt_oss` example test dir renamed to
`gpt-oss` to match the CI matrix.
**Diffusers test models:** shared model-path constants in
`tests/_test_utils/examples/models.py` consolidated/renamed and point at
tiny `hf-internal-testing` test pipes (SDXL/SD3/FLUX) so
cachify/quantize/export tests run on toy weights; `local_id`s
normalized.
**Shared dataset utils:** `examples/llm_sparsity/.../hf_pts.py` now uses
`get_dataset_dataloader` (drops the bespoke cnn_dailymail-only
`get_calib_dataloader`; supports any registered/HF/JSONL dataset,
includes attention_mask); `data_prep.py` gains `--max_samples`.
**CI workflows:** container image bumps (pytorch 26.04→26.05, TRT-LLM
rc16→rc17) and tightened lane timeouts (unit 30→15 min, gpu lanes
trimmed, onnx example lane 45 min).
**Imports at top of file:** in-function imports across the test suite
are moved to module top per the coding guideline, conservatively —
optional deps stay guarded (in-function or behind a module-level
`importorskip`) in `tests/unit` since the partial-install lane runs
without them, and build/hardware-availability imports (apex, triton,
megatron/transformer_engine, tensorrt_llm) plus `_test_utils` lazy
guards are left in place.
**Kernel warning filters:** the repeated `filterwarnings` blanket-ignore
in six `tests/gpu/torch/kernels/**` modules is consolidated into a
scoped hook in `tests/gpu/torch/kernels/conftest.py` (kernel tests only
— the rest of the suite keeps surfacing warnings).
**Eagle example speedups:** `torch.compile` (eagle recipe default) added
~2 min to every eagle training test; it's now disabled in the eagle
example tests except one smoke (`test_llama_eagle3[1-False]`), and the
downstream resume / AR-validate / export tests point at the compile-free
checkpoint. Measured: `test_ar_validate` 139s→17s, offline training
142s→22s, streaming 140s→23s — the compile path is still smoke-tested
once.
**Example lanes install editable (`-e`):** so example scripts launched
as subprocesses resolve `modelopt` to the same source path as the test
process and reuse the pre-compiled CUDA-extension cache instead of
recompiling (~2 min/test); verified in the TRT-LLM container.
**Tiny test tokenizer:** `get_tiny_tokenizer` defaults to left padding
(what decoder-LM calibration expects) and ships a terse
generation-tagged chat template — replacing a verbose ChatML one that
inflated tokenized length on the 128-vocab tokenizer and broke the
offline-PTQ example tests' `max-seq-len` filter.
**Restored Hub-download coverage:** the live (ungated) HF dataset
round-trips exercising `get_dataset_samples`' download branch now live
in `tests/gpu/torch/utils/test_dataset_utils.py` (they had been dropped
from the hermetic unit file without a counterpart).
Individual file changes not explicitly called out above fall under this
general test/CI cleanup.
### Testing
Unit + the touched gpu_megatron files validated locally; example/GPU
lanes validated in CI.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (tests + example CLI args
default to prior behavior)
- If you copied code from any other sources or added a new PIP
dependency: N/A
- Did you write any new necessary tests?: N/A (optimizes/relocates
existing tests)
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: ❌ (pending)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Chores**
* Updated container image versions for PyTorch (26.04→26.05),
TensorRT-LLM (1.3.0rc16→1.3.0rc17), and ONNX/TensorRT (26.04→26.05).
* **Tests**
* Enhanced test isolation: unit tests now run hermetically without
HuggingFace Hub access.
* Optimized test runtime via smaller model/dataset parameters and
parallel test caching.
* Added CUDA extension availability tests and extended dataset utility
coverage.
* **Documentation**
* Updated testing guidelines in `CONTRIBUTING.md` to emphasize offline
test design.
* **Chores**
* Added pytest timeout configuration and improved CI/CD workflow
efficiency with editable installs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
3ff15ccef3 |
Add support for postprocess exported model for block scale swizzling and support for different padding strategy (#1195)
### What does this PR do? Type of change: ? new feature <!-- Details about the change. --> Adds post-processing support for exported diffusion model checkpoints to enable NVFP4 block scale swizzling and configurable padding strategies. This allows exported quantized checkpoints to be directly consumed by inference runtimes (e.g., ComfyUI with comfy_kitchen) that require cuBLAS 2-D block-scaling-factors layout. Changes: 1) Unified post-processing step (_postprocess_safetensors): Loads saved safetensors files and applies merge, padding, swizzle, and quantization metadata injection in a single pass. 2) NVFP4 scale swizzle (swizzle_nvfp4_scales): Rearranges block scales from ModelOpt's flat [rows, cols // 16] layout to cuBLAS 2-D tiled layout per the cuBLAS specification. 3) Configurable padding (pad_nvfp4_weights): Pads NVFP4 weight and scale tensors to multiples of 16, with "row" (rows only) or "row_col" (both dimensions) strategies. 4) Standalone quantization metadata (build_layerwise_quant_metadata): Extracted from merge_diffusion_checkpoint so _quantization_metadata can be injected independently of merging — works for both merged (LTX-2) and standalone (Flux2) exports. 5) Bug fix (conversion.py): Wrapped yield in try/finally in set_quantizer_by_cfg_context so quantizer states are always restored, fixing an issue when yield fails. ### Usage ```python # LTX-2 export with merge + swizzle + padding export_hf_checkpoint( pipeline, export_dir="./output", merged_base_safetensor_path="./ltx-2-22b-dev.safetensors", enable_swizzle_layout=True, padding_strategy="row_col", enable_layerwise_quant_metadata=True, ) # Flux2 standalone export with swizzle + padding (no merge needed) export_hf_checkpoint( transformer, export_dir="./output", enable_swizzle_layout=True, padding_strategy="row_col", ) # Via quantize.py CLI python quantize.py \ --model ltx-2 --format fp4 \ --extra-param merged_base_safetensor_path=./ltx-2-22b-dev.safetensors \ --extra-param enable_swizzle_layout=true \ --extra-param padding_strategy=row_col \ --hf-ckpt-dir ./output ``` ### Testing 1) Exported LTX-2.3 NVFP4 with swizzle + padding + merged base checkpoint. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI 2) Exported Flux2 NVFP4 with swizzle + padding. Verified checkpoint has correct uint8 weights, float8_e4m3fn scales in swizzled layout, and _quantization_metadata . Ran the checkpoint with ComfyUI ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Diffusers export: optional NVFP4 support — swizzle layout, row/row_col padding, and optional per-layer quantization metadata; exports are now post-processed to apply these options. * Export flow accepts new flags to enable swizzle, padding strategy, and layerwise metadata. * **Bug Fixes** * Quantizer context manager now always restores state, including on exceptions. * **Tests** * Added unit tests for NVFP4 padding, swizzling, metadata injection, and post-processing. * **Documentation** * README example updated to show swizzle and padding flags. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ynankani <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Signed-off-by: ynankani-nv <ynankani@nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Signed-off-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: YASH Nankani <ynankani@2u1g-x570-0073.ipp2a1.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-1979.ipp2a2.colossus.nvidia.com> Co-authored-by: YASH Nankani <ynankani@dl325g11-0771.ipp4a1.colossus.nvidia.com> |
||
|
|
e4dc0205d1 |
[OMNIML-4775] Move built-in PTQ quantization configs to YAML (#1423)
### What does this PR do? Type of change: refactor This PR moves the built-in PTQ quantization config definitions out of hard-coded Python dictionaries and into schema-backed YAML config files, and factors shared blocks into reusable composable snippets. - Adds reusable numeric config snippets under `modelopt_recipes/configs/numerics/`. - Adds YAML presets for the built-in model PTQ configs under `modelopt_recipes/configs/ptq/presets/model/`. - Adds YAML presets for KV-cache quantization configs under `modelopt_recipes/configs/ptq/presets/kv/`. - Adds YAML presets for the Diffusers-specific PTQ configs under `modelopt_recipes/configs/ptq/presets/diffusers/` and re-points `examples/diffusers/quantization/config.py` constants at them via `load_config`. - Adds reusable KV quantization units (`kv_fp8_affine`, `kv_nvfp4`, `kv_nvfp4_affine`, `kv_nvfp4_rotate`, `kv_*_cast` variants) under `modelopt_recipes/configs/ptq/units/`. - Adds reusable model-side units following the `component_numerics[_type]` convention: - `attention_qkv_fp8` — FP8 E4M3 on attention q/k/v bmm and softmax quantizers; shared by `model/` and `diffusers/` `nvfp4_fp8_mha` presets. - `block_sparse_moe_nvfp4` — NVFP4 W4A4 on `*block_sparse_moe*` weight/input quantizers; shared by `nvfp4_mlp_only`, `nvfp4_experts_only`, `nvfp4_omlp_only`. - `experts_nvfp4` — NVFP4 W4A4 on `*.experts.*` weight/input quantizers; shared by `nvfp4_mlp_only` and `nvfp4_experts_only`. - Switches the existing 5 NVFP4 presets (default + awq lite/clip/full + svdquant) and 4 mamba_moe presets to `$import` the existing `w4a4_nvfp4_nvfp4` / `w8a8_fp8_fp8` units instead of re-inlining the same weight+input quantizer pairs. - Moves the recently-added `W4A16_NVFP4_CFG` to YAML (`presets/model/w4a16_nvfp4.yaml`) composed from the existing `units/w4_nvfp4` snippet. - Updates `modelopt.torch.quantization.config` built-in config constants to load `QuantizeConfig` objects from YAML with `load_config(..., schema_type=QuantizeConfig).model_dump(exclude_unset=True)` via a new `_load_quantize_config_dict` helper; the constants remain plain `dict[str, Any]` for backwards compatibility with consumers that do mapping-style mutation (e.g. `entry["cfg"]` assignment). - Simplifies the cfg-list loader (`_load_quantizer_cfg_dict_list`) down to a 4-line list/single normalization now that the three call sites all load schema-typed YAMLs. - Adds/updates recipe loader coverage for built-in schema-backed config snippets. ### Latent-bug fixes surfaced by the refactor Two small correctness fixes are included alongside the mechanical refactor; flagging them explicitly: - **`examples/diffusers/quantization/quantize.py`** — adds an explicit `base_cfg = copy.deepcopy(base_cfg)` before applying runtime overrides. The existing `# Build a fresh config dict so we never mutate the global constants` comment had been aspirational only; in practice `reset_set_int8_config` accumulated `PercentileCalibrator` entries into `mtq.INT8_SMOOTHQUANT_CFG`/`INT8_DEFAULT_CONFIG` across repeated calls, and `set_quant_config_attr` added `trt_high_precision_dtype` keys into globally-shared cfg dicts. The deepcopy makes the code match the comment. - **`choices` set in `modelopt/torch/quantization/config.py`** — adds `MXFP6_DEFAULT_CFG` and `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` to the documented public set of valid `mtq.*_CFG` names. Both constants exist on main but were missing from `choices`, so CLIs that gate on `mtq.config.choices` (e.g., `hf_ptq.py --qformat`) couldn't reach them even though the configs themselves were fully supported. ### Usage Existing Python imports continue to work: ```python import modelopt.torch.quantization as mtq cfg = mtq.FP8_DEFAULT_CFG model = mtq.quantize(model, cfg, forward_loop) ``` The built-in constants are plain `dict[str, Any]` (sparse — only explicitly-set fields are present), but their definitions now come from YAML snippets and presets composed through the existing `$import` system. Reusable YAML snippets can be composed through `$import`, for example: ```yaml # modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig imports: base_disable_all: configs/ptq/units/base_disable_all w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4 default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers algorithm: max quant_cfg: - $import: base_disable_all - $import: w4a4_nvfp4_nvfp4 - $import: default_disabled_quantizers ``` ### Testing Local checks run: - `nox -s "unit-3.10(torch_211, tf_latest)"` — 2329 passed, 12 skipped. - `nox -s pre_commit_all` — all hooks pass (ruff check / ruff format / mypy / YAML format / license / bandit / markdownlint). - YAML parse + `$import` resolution sanity check across all changed config files. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ Existing built-in Python config constants keep the same public names and dict semantics. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ Adds/updates recipe loader coverage for schema-backed built-in snippets. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: ❌ ### Additional Information This PR was previously stacked on #1405, which has since merged to `main`. The branch has been rebased onto `main` and no longer depends on any other open PR. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Many new quantization numeric configs and PTQ presets added (INT4/INT8/MXFP4/MXFP6/MXFP8/MXINT8/NVFP4), plus Diffusers, KV-cache (affine/cast/rotate) and MLP/MoE-targeted presets. * **Refactor** * Presets and shared snippets migrated to schema-backed YAML sources and centralized loading; INT8 percentile calibration avoids mutating shared base configs. * **Tests** * Tests now discover packaged config snippets at runtime and validate import/append behaviors. * **Documentation** * Presets README and numerous header descriptions updated. * **Chores** * Minor typing and script improvements. <!-- review_stack_entry_start --> [](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1423?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack) <!-- review_stack_entry_end --> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
26ae8da517 |
[2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do? Type of change: new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> - Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused NVFP4 activation quantization for video diffusion VAE layers - Integrate into _QuantConv3d via QuantModuleRegistry — automatically dispatched when NVFP4 quantization is applied to nn.Conv3d - Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`; move tests to `tests/gpu/torch/quantization/kernels/` ### Testing <!-- Mention how have you tested your change if applicable. --> - Added test cases to measure the difference between cuDNN and our CUDA implicit GEMM kernel - Added an NVFP4 fake quantization test using CUDA code ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Per-backbone quantization/export in a single run with per-backbone checkpoints and backbone-aware quant filters * Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D inference path and Wan 2.2 quantization support * **Bug Fixes** * Video-model calibration now respects extra params and forces video decoding during calibration * **Documentation** * Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed experimental Conv3D prototype docs/benchmark * **Tests** * New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel test coverage <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
80a77d1cc0 |
Add LTX-2 third-party license notices for legal compliance (#1226)
## Summary LTX-2 (`ltx-core`, `ltx-pipelines`, `ltx-trainer`) is a third-party dependency developed and provided by Lightricks. It is governed by the [LTX Community License Agreement](https://github.com/Lightricks/LTX-2/blob/main/LICENSE), **not** the Apache 2.0 license that covers NVIDIA Model Optimizer. Per legal guidance, all integration points must clearly surface this to users. - Add `[!WARNING]` license notice blocks at the top of all LTX-2-related READMEs (`examples/diffusers`, `examples/diffusers/distillation`, `examples/windows/diffusers/qad_example`) - Add `warnings.warn(UserWarning)` at every LTX package import site in Python files, covering both top-level and lazy imports: - `examples/diffusers/distillation/distillation_trainer.py` - `examples/diffusers/quantization/calibration.py` - `examples/diffusers/quantization/pipeline_manager.py` - `examples/windows/diffusers/qad_example/sample_example_qad_diffusers.py` - `modelopt/torch/export/diffusers_utils.py` - `modelopt/torch/quantization/plugins/diffusion/ltx2.py` - Add license notice comment to `requirements.txt` files that list LTX packages, so the obligation is visible at install time - Update `.github/CODEOWNERS` so all `requirements*.txt` files (covering variants like `requirements-dev.txt`) are owned by `@NVIDIA/modelopt-setup-codeowners` regardless of location, via a last-match-wins rule **Design notes:** - For library files (`diffusers_utils.py`, `ltx2.py`), the warning is placed at the lazy import site inside functions — it fires only when LTX-2 code paths are actually invoked, not at module import time, to avoid polluting non-LTX users - For example entry-point scripts that are LTX-2-only, the warning fires at module load time (after all imports, to satisfy ruff E402) ## Test plan - [ ] Confirm `pre-commit run --all-files` passes (ruff, mypy, markdownlint, bandit all clean) - [ ] Verify warning appears at runtime when running an LTX-2 quantization or distillation example - [ ] Confirm non-LTX code paths (FLUX, SDXL, SD3) do not emit the warning 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added third-party license notices across documentation and requirements files clarifying LTX-2 packages are governed by the LTX Community License Agreement rather than NVIDIA Model Optimizer's Apache 2.0 license. * **Chores** * Updated code ownership configuration for requirements files. * Added runtime warnings to notify when LTX-2 dependencies are accessed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
1cceb950d6 |
[OMNIML-3689] PTQ quant_cfg semantic correction. Design in doc _quant_cfg.rst (#1094)
### What does this PR do?
#### Summary
Redesigns the `quant_cfg` configuration format in ModelOpt's PyTorch
quantization stack, replacing the previous dict-based format with an
**ordered list of typed `QuantizerCfgEntry` dicts**.
##### Motivation
The old `quant_cfg` dict had several pain points:
- **Ambiguous precedence**: no explicit way to reason about which entry
wins when multiple keys match a quantizer
- **Mixed key namespaces**: wildcard paths and PyTorch class names lived
in the same dict level, requiring ad-hoc dispatch
- **Magic `"default"` key**: an implicit, undocumented catch-all that
was easy to misuse
- **Poor composability**: merging two configs required dict updates that
silently discarded keys
- **No YAML round-trip fidelity**: the nested structure couldn't be
expressed cleanly in YAML
##### New format
`quant_cfg` is now an ordered list of `QuantizerCfgEntry` TypedDicts.
Each entry has:
- `quantizer_name` *(required)*: `fnmatch` wildcard matched against
quantizer module names
- `cfg` *(optional)*: dict (or list of dicts) of
`QuantizerAttributeConfig` fields
- `enable` *(optional)*: toggles quantizer on/off independently of `cfg`
- `parent_class` *(optional)*: restricts match to quantizers whose
parent module is of the given PyTorch class (e.g. `"nn.BatchNorm2d"`)
Entries are applied in list order; later entries override earlier ones.
The canonical pattern is deny-all first (`_base_disable_all`), then
selectively re-enable and configure, then apply standard exclusions
(`_default_disabled_quantizer_cfg`).
##### Changes
**Core library (`modelopt/torch/quantization/`)**
- **`config.py`**:
- Added `QuantizerCfgEntry` TypedDict (line 163) and
`find_quant_cfg_entry_by_path()` helper for exact-match lookup of
entries by path.
- Added `normalize_quant_cfg_list()` (line 1539) that converts legacy
formats (flat dict, single-key dicts, `nn.*`-scoped dicts, `"default"`
key) to canonical `QuantizerCfgEntry` lists. After normalization every
entry is guaranteed to have explicit `quantizer_name`, `enable`, and
`cfg` keys.
- Converted `_default_disabled_quantizer_cfg` and
`_mamba_moe_disabled_quantizer_cfg` from dicts to lists of
`QuantizerCfgEntry`.
- Added `_base_disable_all` (line 205): canonical deny-all entry
(`[{"quantizer_name": "*", "enable": False}]`).
- Converted all ~30 built-in config constants (`INT8_DEFAULT_CFG`,
`FP8_DEFAULT_CFG`, `NVFP4_DEFAULT_CFG`, etc.) to list format using
`*_base_disable_all` and `*_default_disabled_quantizer_cfg` unpacking.
- KV-cache configs (`FP8_KV_CFG`, `NVFP4_KV_CFG`, etc.) are now minimal
lists designed to be concatenated with a primary config — they
intentionally omit `_base_disable_all` and `"algorithm"`.
- Added two `QuantizeConfig` Pydantic field validators: a
`mode="before"` validator that calls `normalize_quant_cfg_list()`, and a
`mode="after"` validator that validates `cfg` dicts against
`QuantizerAttributeConfig`.
- Updated `need_calibration()` to iterate the normalized list instead of
the old dict.
- Changed `QuantizeQuantCfgType` alias from `dict[str | Callable, ...]`
to `list[QuantizerCfgEntry]`.
- **`conversion.py`**:
- Rewrote `set_quantizer_by_cfg()` (line 217) to iterate the list
directly. Each entry's `parent_class` is resolved via
`QuantModuleRegistry[parent_class_name]` (the existing `_DMRegistryCls`
registry).
- Added `set_quantizer_attributes_full()` (line 314): full replacement
of quantizer attributes from a `QuantizerAttributeConfig`. Unspecified
fields revert to defaults, enforcing entry atomicity. Can also upgrade
`TensorQuantizer` → `SequentialQuantizer` or downgrade the reverse.
- Added `set_quantizer_attributes_partial()` (line 384): merges a
partial `dict` of attributes into existing quantizer state. Does NOT
change quantizer structure. Used for enable-only entries.
- Added `set_quantizer_by_cfg_context()` context manager (line 447) that
temporarily applies a `quant_cfg` list and restores original quantizer
state on exit.
- Deprecated `set_quantizer_attribute()` (line 525) with a
`DeprecationWarning` pointing to the new functions.
- **`tensor_quantizer.py`**:
- `TensorQuantizer.set_from_attribute_config()`: narrowed type hint from
`dict` to `dict[str, Any]`.
- Added `_axis_setter` and `_block_sizes_setter` custom setters so that
`axis` and `block_sizes` changes properly propagate to the calibrator
and maintain mutual exclusivity.
- `SequentialQuantizer.set_from_attribute_config()`: narrowed signature
to `list[QuantizerAttributeConfig] | list[dict[str, Any]]` (removed the
old union with single values).
- **`algorithms.py`**:
- Updated `_match_quantizer_cfg()` to iterate the list and return
`(matched_cfg, matched_enable)` tuple with last-match-wins.
- Updated `_cfg_to_dict()`, `estimate_quant_compression()`, and
`QuantRecipe` to work with the list-based format.
- Updated `get_auto_quantize_config()` to emit list-format `quant_cfg`.
- **`model_quant.py`**: `disable_quantizer()` / `enable_quantizer()` now
call `set_quantizer_attributes_partial()` directly instead of the
deprecated `set_quantizer_attribute()`. Updated docstrings and code
examples to show the list format.
- **`utils/core_utils.py`**: `disable_lora_quantizers_in_config()` and
`update_quant_cfg_with_kv_cache_quant()` updated to append
`QuantizerCfgEntry` dicts to the list.
- **Other**: minor updates to `backends/fp8_per_tensor_gemm.py`,
`backends/nvfp4_gemm.py`, `compress.py`, `model_calib.py`,
`export/unified_export_hf.py`, and
`sparsity/attention_sparsity/conversion.py` to use the list format.
- **`onnx/llm_export_utils/quantization_utils.py`**: Updated
quantization config construction to use list format.
**YAML recipes (`modelopt_recipes/`)**
- Converted all 5 general PTQ recipes to the new list format:
- `general/ptq/fp8_default-fp8_kv.yml`
- `general/ptq/nvfp4_default-fp8_kv.yml`
- `general/ptq/nvfp4_experts_only-fp8_kv.yml`
- `general/ptq/nvfp4_mlp_only-fp8_kv.yml`
- `general/ptq/nvfp4_omlp_only-fp8_kv.yml`
- Converted model-specific recipe:
`models/Step3.5-Flash/nvfp4-mlp-only.yaml`
**Documentation (`docs/`)**
- New guide: `docs/source/guides/_quant_cfg.rst` — comprehensive
reference covering entry format, ordering semantics, entry atomicity,
`enable` vs `cfg` independence, `parent_class` filtering, and common
patterns (deny-all-then-enable, customizing a built-in config, building
from scratch).
- Updated `_pytorch_quantization.rst` code examples to show the list
format with `copy.deepcopy` and `.append()`.
- Added `_quant_cfg.rst` to the quantization guide table of contents.
**Examples**
- Updated all quantization examples to use the list format:
`deepseek/ptq.py`, `diffusers/quantization/config.py`,
`llm_ptq/hf_ptq.py`, `llm_qat/main.py`, `vllm_serve/vllm_ptq_utils.py`,
`llm_autodeploy/run_auto_quantize.py`, `llm_eval/quantization_utils.py`,
`llm_ptq/example_utils.py`,
`windows/torch_onnx/diffusers/qad_example/sample_example_qad_diffusers.py`,
and 2 notebooks.
**Tests**
- New test file:
`tests/unit/torch/quantization/test_config_validation.py` — unit tests
for `need_calibration()`, `normalize_quant_cfg_list()` (new format,
legacy format conversions, error cases),
`find_quant_cfg_entry_by_path()`, `_match_quantizer_cfg()`, and
`QuantizeConfig` Pydantic validators.
- Extended `tests/unit/torch/quantization/test_quantize_cpu.py` with
tests for `set_quantizer_attributes_full()` (atomicity, parent_class
filtering, SequentialQuantizer creation), list ordering, enable-only
entry behavior, and end-to-end legacy dict format.
- Updated 20+ existing test files across `tests/unit/`, `tests/gpu/`,
`tests/gpu_megatron/`, and `tests/_test_utils/` to use the list format.
##### Backward compatibility
`normalize_quant_cfg_list()` is called automatically by the
`QuantizeConfig` Pydantic `mode="before"` validator, so existing code
passing the old dict-based format (flat dict like `{"*weight_quantizer":
{"num_bits": 8}}`, single-key dict lists, or `nn.*`-scoped dicts with
`parent_class` semantics) continues to work without modification. The
legacy `"default"` key is converted to `quantizer_name: "*"`.
`set_quantizer_attribute()` is preserved as a deprecated wrapper around
`set_quantizer_attributes_partial()`.
#### Test coverage
- **Unit tests**: new `test_config_validation.py` with tests for
normalization, validation, path lookup, and cfg matching. Extended
`test_quantize_cpu.py` with tests for full/partial attribute setting,
ordering, atomicity, and legacy backward compatibility.
- **System testing**:
```
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
--export_path=build/fp8_default-fp8_kv42 \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=hf
```
### Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
|
||
|
|
4a5ef01acc |
Fixed the calib size bug (#1178)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed quantization configuration calibration batch counting to properly account for partial final batches during the calibration process. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
1070d895dc |
Flux2-Dev Quantization (#947)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Register Flux2Attention and Flux2ParallelSelfAttention in the
quantization plugin so bmm quantizers are patched (enables
--quantize-mha).
- Add Flux2-specific dummy input generation for HF checkpoint export.
- Guard check_conv_and_mha with hasattr for bmm quantizer attributes
## Usage
<!-- You can potentially add a usage example below. -->
```bash
python quantize.py \
--model flux2-dev \
--model-dtype BFloat16 \
--format fp4 --batch-size 2 --calib-size 1 \
--n-steps 20 --quantized-torch-ckpt-save-path ./flux2-dev-fp4.pt --collect-method default \
--hf-ckpt-dir ./flux2-dev-fp4
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes<!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Flux2-dev model support with Flux2-compatible dummy input
generation and default inference params (768×1024, guidance scale 4.0).
* **Refactor**
* Made attention quantization disabling more robust by iterating
available quantizers before disabling.
* **Infrastructure**
* Flux2 attention components are now optional and registered only when
present to avoid import issues.
* **Tests**
* Added Flux2 test helpers and coverage validating Flux2 dummy input
shapes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
812e8c60a2 |
Minor update on the LTX2 NVFP4 recipe (#1010)
### What does this PR do? Type of change: minor code change <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> 1. Update the default calibration dataset for LTX_VIDEO_DEV and LTX2 from Gustavosta/Stable-Diffusion-Prompts to nkp37/OpenVid-1M, which provides video-specific captions better suited for video model calibration. 2. update the default recipe for ltx2: first 3 and last 3 layers stays at higher precision. ### Usage ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated default dataset configuration for LTX-Video and LTX2 models. * Refined model filtering pattern for LTX-Video to support additional model components. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
37d3f10cbd |
To support LTX2 ComfyUI format (#972)
### What does this PR do?
Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
- Added a flag merged_base_safetensor_path to the example code so that
user can export the ComfyUI style ckpt.
### Usage
```bash
python quantize.py \
--model ltx-2 --format fp4 --batch-size 1 --calib-size 32 --n-steps 40 \
--extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors \
--extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors \
--extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors \
--extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized \
--extra-param fp8transformer=true \
--quantized-torch-ckpt-save-path ./ltx-2-transformer.pt \
--hf-ckpt-dir ./LTX2-NVFP4/ \
--extra-param merged_base_safetensor_path=./ltx-2-19b-dev-fp8.safetensors
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added new command-line parameters documentation for LTX-2 FP4
quantization examples (--hf-ckpt-dir and merged_base_safetensor_path
configuration options)
* **Improvements**
* Enhanced quantization pipeline to support conditional export behavior
based on model type
* Expanded LTX-Video model filtering patterns for more comprehensive
block detection
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
2905cb0f2e |
Updated the diffusion config issue and more test cases (#937)
## What does this PR do?
**Type of change:** new tests, Bug fix <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
**Overview:**
- **Fixed the INT8 config issue**
- **Add HF checkpoint export test coverage**
1. The `--hf-ckpt-dir` export path had zero test coverage. This MR adds
tests at two levels:
2. Unit tests (tests/unit/torch/export/test_export_diffusers.py):
- Extended test_export_diffusers_real_quantized to parametrize over
INT8, INT8 SmoothQuant, FP8, and FP4 configs
- (previously only FP8). This gives 3 models x 4 configs = 12 test
cases.
3. GPU integration tests
(tests/gpu/torch/export/test_export_diffusers_hf_ckpt.py)
- New file testing the full quantize.py --hf-ckpt-dir pipeline via
subprocess with 4 combos:
- SDXL INT8 smoothquant min-mean (the exact scenario that triggered the
bug)
- Flux INT8 smoothquant min-mean
- SDXL FP8
- Flux FP4
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**:No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Tests**
* Added test coverage for exporting Diffusers models with Hugging Face
checkpoints across multiple quantization formats (INT8, FP8, FP4)
* Extended quantization export testing to validate multiple
configuration scenarios
* **Chores**
* Refined INT8 quantization configuration with improved calibrator
support for convolution layers
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
d78797b466 |
Update the LTX2 API calls during the calibration (#926)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Update LTX-2 integration to match latest upstream API 1. The LTX-2 codebase removed/replaced several APIs. This MR updates all affected files: 2. Replace cfg_guidance_scale with MultiModalGuiderParams: The pipeline __call__ no longer accepts a single cfg_guidance_scale float. It now requires two MultiModalGuiderParams objects (video_guider_params and audio_guider_params) that control CFG, STG, rescale, cross-modality guidance, and skip-step settings. Updated in ltx-2.py, ltx-2-fp8.py, ltx-2-onestage.py, calibration.py, and models_utils.py. 3. Replace fp8transformer with QuantizationPolicy: The TI2VidTwoStagesPipeline constructor no longer accepts the fp8transformer boolean flag. FP8 quantization is now configured via quantization=QuantizationPolicy.fp8_cast(). Updated in ltx-2-fp8.py and pipeline_manager.py (with backwards-compatible support for the old --extra-param fp8transformer=true CLI flag). 4. Remove DEFAULT_CFG_GUIDANCE_SCALE constant: Replaced by DEFAULT_VIDEO_GUIDER_PARAMS and DEFAULT_AUDIO_GUIDER_PARAMS in all import sites. ## Usage <!-- You can potentially add a usage example below. --> ```bash python quantize.py --model ltx-2 --format fp4 --batch-size 1 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Updates** * Default resolution for LTX2 models adjusted to 768x1280 * Guidance parameter configuration updated for video and audio pipelines * FP8 quantization parameter handling refined <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
10efcb65f6 |
[3.1/4] Diffusion Quantized ckpt export - WAN 2.2 14B (#855)
## What does this PR do?
**Type of change:** documentation <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->
**Overview:**
1. Added multi‑backbone support for quantization: --backbone now accepts
space- or comma-separated lists and resolves to a list of backbone
modules.
2. Introduced PipelineManager.iter_backbones() to iterate named backbone
modules and updated get_backbone() to return a single module or a
ModuleList for multi‑backbone.
3. Updated ExportManager to save/restore per‑backbone checkpoints when a
directory is provided, with {backbone_name}.pt files, and to create
target directories when missing.
4. Simplified save_checkpoint() calls to rely on the registered
pipeline_manager by default.
**Usage: **
```bash
python quantize.py --model wan2.2-t2v-14b --format fp4 --batch-size 1 --calib-size 32 \
--n-steps 30 --backbone transformer transformer_2 --model-dtype BFloat16 \
--quantized-torch-ckpt-save-path ./wan22_mo_ckpts \
--hf-ckpt-dir ./wan2.2-t2v-14b
```
Plans
- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml
- [x] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Unified Hugging Face export support for diffusers pipelines and
components
* LTX-2 and Wan2.2 (T2V) support in diffusers quantization workflow
* Comprehensive ONNX export and TensorRT engine build documentation for
diffusion models
* **Documentation**
* Updated to clarify support for both transformers and diffusers models
in unified export API
* Expanded diffusers examples with LoRA fusion guidance and additional
model options (Flux, SD3, SDXL variants)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
2e43c80609 |
[2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This MR adds HuggingFace checkpoint export support for LTX‑2 by treating TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling back to writing model.safetensors when save_pretrained isn’t available. It also preserves the original forward in DynamicModule patching (_forward_pre_dm) so downstream callers can still invoke the pre-patched forward implementation. **Changes** 1. Added the calibration & quantization support of the LTX2, even with FP8 precision. 2. Preserve original forward before `DynamicModule` patching: when patching forward, we now stash the pre-patched implementation in `self._forward_pre_dm` (once) so downstream code can still call the original forward, then re-bind forward to the class implementation. This is needed for the LTX2 FP8 calibration. 3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a “diffusion-like” object and routes it through _export_diffusers_checkpoint() (import guarded; no hard dependency). 4. Generalized component discovery: introduced get_diffusion_components() (aliasing the old get_diffusers_components) to support non-diffusers pipelines; for LTX‑2 it returns only stage_1_transformer. 5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy forward generator (generate_diffusion_dummy_forward_fn) that builds minimal LTX Modality inputs (including correct timesteps broadcasting) so shared-input hooks can run and fuse QKV when applicable. 6. Export fallback for non-save_pretrained modules: when a component lacks save_pretrained (LTX‑2 transformer), export now writes model.safetensors + minimal config.json instead of pytorch_model.bin. Plans - [x] [1/4] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml - [ ] [3/4] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml - [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ```bash python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added LTX-2 video model support with complete quantization and export pipeline integration * Introduced `--extra-param` CLI option for flexible model configuration and parameter passing * Enhanced export capabilities with broader diffusion model compatibility * **Chores** * Changed default model data type from Half to BFloat16 for improved numerical stability <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
668b8a19e8 |
[1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This PR adds support for exporting quantized diffusers models (DiT, Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM. **Changes** New file: `diffusers_utils.py` - Dummy input generation for various diffusion models - Pipeline component extraction helpers - QKV projection detection and grouping - `hide_quantizers_from_state_dict()` context manager for clean saves Refactored: `unified_export_hf.py` - New `_fuse_qkv_linears_diffusion()` for QKV amax fusion - `_export_diffusers_checkpoint()` to export full pipelines (models + tokenizers + schedulers etc.) Plans - [x] [1/3] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml - [ ] [3/3] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ``` mtq.quantize(pipe, quant_config, forward_call) export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir) ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added HuggingFace checkpoint export support for quantized diffusion models with configurable output directory * Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint export destination * Extended export functionality to support selective component exports from diffusion pipelines * Enhanced quantized model export with improved component handling and multi-stage checkpoint generation <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
e6e4efd61e |
[0.5/3] Diffusion ckpt export for NVFP4 & FP8 (#783)
See https://github.com/NVIDIA/Model-Optimizer/pull/781 This is the MR that only includes the refactoring of the llm export, please ignore the change on quantize.py from the diffusion example. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added `--hf-ckpt-dir` CLI option to save checkpoints in HuggingFace format * Enabled support for exporting Diffusers-based pipelines * Unified export system now handles both transformer and diffusion model architectures <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
fd66be2bec |
[NVBug 5659126] dummy inputs are kwargs maps (#676)
## What does this PR do? bug fix **Overview:** [NVBug 5659126] dummy inputs are kwargs maps, and rename the generate_... function to more explicitly reflect the return value type ## Testing `python diffusion_trt.py --model flux-dev --override-model-path /models/FLUX.1-dev --torch --benchmark --skip-image ` Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
53a2ddebab |
Product Rename: TensorRT Model Optimizer to Model Optimizer (#583)
- [x] Product Rename: TensorRT Model Optimizer to Model Optimizer (OMNIML-3033) - [x] Mention in Latest News section with date on the date of merging this PR (12/08) Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c6c9905929 |
[OMNIML-2244] Create the nvfp4 quant exporter (#636)
## What does this PR do? **Type of change:** New feature **Overview:** - Implemented the NVFP4QuantExporter - Deprecated fp4qdq_to_2dq - Updated tests ## Usage ```python python torch_quant_to_onnx.py --quantize_mode=nvfp4 \ --onnx_save_path=vit_base_patch16_224.nvfp4.onnx \ --calibration_data_size 64 \ --batch_size 128 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ``` python evaluate.py --onnx_path=vit_base_patch16_224.nvfp4.onnx \ --model_name=vit_base_patch16_224 \ --results_path=./results.txt \ --batch_size 128 ``` Results: ``` The top1 accuracy of the model is 84.39% The top5 accuracy of the model is 97.312% Inference latency of the model is 7.22412 ms ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No - Deprecated fp4qdq_to_2dq - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
cfb247a1f3 |
[NVBug 5659126] The same workaround for RMSNorm exporting for diffusers >=0.35.0 (#642)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? For the trt_diffusions script ## Testing python diffusion_trt.py --model flux-dev --benchmark --skip-image Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
8844a2b2bc |
Support attention quantization for diffusers >= 0.35.0 (#608)
## What does this PR do? **Type of change:** new feature **Overview:** ? Attention mechanism has changed from diffusers 0.35. Many model attentions are now subclass of a new Mixin class: AttentionModuleMixin, which is not a sub class of Attention To fix it, patch the mixin class by forcing to use native attention impl so the existing function monkey patch still work. ## Testing manual quant of Wan, Flux --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
e0a6efbe70 |
fix trt engine building of the diffusers pipelines (#637)
## What does this PR do? **Type of change:** Bug fix **Overview:** 1. The diffusion_trt.py needs the dynamic_shapes when running trtexec for engine building. A previous change altered the format of dynamic_shapes, fix it here. 2. the dynamic_shapes logic gets cleaned up. The existing logic is very confusing 3. recover min-batch_size config for some pipelines. Previously some pipelines set the min batch_size to be > 1, which was odd, so a previous change sets them to be 1, but it turns out the oddity has a reason, the trt engine building fails with the altered batch_size min/opt, thus recover them. ## Testing pytest tests/examples/diffusers --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d0b0c0fd46 |
Fix extra args and --component-dtype default value (#605)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? 1. We are passing incorrect extra args to pipeline inference. 2. Need a default empty list for the list argument component-dtype ## Testing Passed SDXL pipeline without component dtype --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
e35db173b1 |
Support Wan2.2 t2v diffusers quantization (#556)
## What does this PR do? **Type of change:** new feature **Overview:** Support Wan2.2 t2v diffusers quantization 1. fix torch2.9 support 2. add Wan2.2 t2v diffusers pipeline quantization Main difference of the Wan2.2 pipeline comparing to exisiting pipelines is that there are 2 backbone models for denoising. For the quantization therefore we need to quantize both of them. However, it turns out our base library does not well support quantization of multiple models in the same time. Therefore, the change here just stick to quantize a single model each time, and then run the quantization multiple times. So, we need to allow users to pick which backbone to quantize, therefore adding a new argment for it 3. add a workaround for the exporting ONNX issue when we upgrade diffusers to >= 0.35.0. The issue lies is the exporting of the torch.nn.RMSNorm. Some pipelines in the diffusers > 0.35.0 use the torch version RMSNorm while before that they use the diffusers' own version of RMSNorm. It turns out they are directly replacable so the workaround is to simply replace the torch RMSNorm usages with diffusers RMSNorm. But we need to fix it properly soon by porting our ONNX export to be based on torch dynamo instead of torchscript. Issue reported from external user: https://github.com/NVIDIA/TensorRT-Model-Optimizer/issues/262 4. allow use of a prompts file, which is simply a text file with a list of prompts, one prompt each line 5. allow each component of a pipeline to have different dtype accuracy. added a new list stype command line arg --component-dtype for this. example: --component-dtype vae:Float 6. print the summary of the quantized model so users can capture issues from log ## Usage python quantize.py \ --model wan2.2-t2v-14b \ --format fp8 \ --batch-size 4 \ --calib-size 64 \ --n-steps 20 \ --backbone transformer \ --model-dtype BFloat16 \ --component-dtype vae:Float \ --trt-high-precision-dtype BFloat16 \ --quantized-torch-ckpt-save-path ./wan_transformer.pt \ --onnx-dir wan-transformer-onnx \ --prompts-file wan-prompts.txt ## Testing Tested SDXL_BASE, LTX_VIDEO_DEV, WAN22_T2V ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information https://github.com/NVIDIA/TensorRT-Model-Optimizer/issues/262 --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com> |
||
|
|
c333d36a85 |
Enable torch 2.9 tests in CICD + Diffusers fixes (#561)
As title --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
8188a01b38 |
[NVBUG: 5619158] Optimize memory usage for diffusion_trt.py (#547)
## What does this PR do? **Type of change:** Minor code change **Overview:** - Delete backbone after Device Model creation - Add assertion for torch compile - Update dummy input generation function ## Testing ``` python diffusion_trt.py --model flux-dev --benchmark --skip-image python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp8_autodeploy_fake.pt python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp4_autodeploy_fake.pt python diffusion_trt.py --model flux-dev --benchmark --skip-image --torch python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp8_autodeploy_fake.pt --torch python diffusion_trt.py --model flux-dev --benchmark --skip-image --restore-from ./flux_dev_fp4_autodeploy_fake.pt --torch python diffusion_trt.py --model flux-dev --benchmark --skip-image --torch --torch-compile ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
e74a46890d |
Update benchmarking for diffusers (#487)
## What does this PR do? **Type of change:** Example update **Overview:** - Optimize the benchmarking function in the diffusers example ```python python diffusion_trt.py --model flux-dev --benchmark --model-dtype BFloat16 --skip-image --torch ``` ## Testing ``` Backbone-only inference latency (BFloat16): Average: 139.48 ms P50: 139.36 ms P95: 141.13 ms P99: 141.35 ms ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
229053323a |
[NVBUG: 5619158] Enforce high precision model dtype for diffusion trt (#526)
## What does this PR do? **Type of change:** Minor code change **Overview:** - Select the high precision dtype directly based on model type - FP16 for Stable Diffusion models, BF16 for Flux ## Testing ```python python diffusion_trt.py --model flux-dev --benchmark ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No (No option to specify dtype while loading pipeline) - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
41f2bf4941 |
Add option to real quantize the model (#473)
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
99c76fff26 |
Add SD3.5-medium quantization support in ModelOpt Diffusers example (#444)
Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
2fd67cc711 |
Add option to benchmark pipeline in diffusion_trt.py (#457)
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
e6e0d2cc8f |
Fix failing CICD nightly tests (#445)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
c0590b0255 |
Deprecate ModelOpt custom docker and directly use TRT-LLM / PyTorch / TRT docker (#346)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
358b0c61b4 |
Fixed the CICD for diffusers (#312)
Signed-off-by: jingyu <jingyu@omniml.ai> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
4d1eb0caf5 |
Major improvement of READMEs and documentation
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
d09c82779c |
Push latest changes and bug fixes (#274)
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
dcdc2df084 |
Push latest changes
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
4c611e47a6 |
Update files on GitHub
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
cafa7f60ed | Update for 0.33.0 release | ||
|
|
7af33d29ce | Update for 0.31.0 release | ||
|
|
7a047435ae | Update for 0.29.0 release | ||
|
|
9cb36cb43c | Update files for 0.27.1 release | ||
|
|
92f430f6ab | Add files for 0.27.0 release | ||
|
|
2017cd9063 | Update for 0.25.0 release | ||
|
|
73d6af785f | Update 0.23.0 - OSS release |