Commit Graph
4 Commits
Author SHA1 Message Date
jingyu-ml 1070d895dc Flux2-Dev Quantization (#947)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Register Flux2Attention and Flux2ParallelSelfAttention in the
quantization plugin so bmm quantizers are patched (enables
--quantize-mha).
- Add Flux2-specific dummy input generation for HF checkpoint export.
- Guard check_conv_and_mha with hasattr for bmm quantizer attributes 

## Usage
<!-- You can potentially add a usage example below. -->

```bash
python quantize.py \
    --model flux2-dev \
    --model-dtype BFloat16 \
    --format fp4 --batch-size 2 --calib-size 1 \
    --n-steps 20 --quantized-torch-ckpt-save-path ./flux2-dev-fp4.pt --collect-method default \
    --hf-ckpt-dir ./flux2-dev-fp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes<!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Flux2-dev model support with Flux2-compatible dummy input
generation and default inference params (768×1024, guidance scale 4.0).

* **Refactor**
* Made attention quantization disabling more robust by iterating
available quantizers before disabling.

* **Infrastructure**
* Flux2 attention components are now optional and registered only when
present to avoid import issues.

* **Tests**
* Added Flux2 test helpers and coverage validating Flux2 dummy input
shapes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-14 11:42:02 -05:00
ynankani-nv 110a44c81e Add support for export ComfyUI compatible checkpoint for diffusion model(e.g., LTX-2) (#911)
## What does this PR do
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)

**Type of change:** <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

**Overview:** 
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)
1) Added a a parameter for merging the base vae, vocoder, connectors in
the quantized checkpoint
2) storing quantization metadata and export tool as modelopt , required
for ComfyUI compatibility.
3) Internally updating the transformer block prefixes to match the
expectation of ComfyUI
 
## Usage
<!-- You can potentially add a usage example below. -->

```python
    export_hf_checkpoint(
        pipeline,
        export_dir=EXPORT_DIR,
        merged_base_safetensor_path=BASE_CKPT,  # merge VAE/vocoder from base
    )
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
1) Tested with ltx-2 model 
       a) initializing a twoStagePipeline object 
       b) calling mtq.quantize on transformer with NVFP4_DEFAULT_CFG 
c) then exporting with export_hf_checkpoint passing the param
merged_base_safetensor_path to generate merged
           checkpoint
2) Ran the generated checkpoint with step1 on ComfyUI to validate
3) Ran step1 without merged_base_safetensor_path to check backward
compatibility.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for exporting LTX-2 diffusion models with merged base
checkpoint integration
* Enhanced export functionality to preserve and attach quantization
metadata during model export
* Extended model export capabilities with automatic model type detection
for improved export handling

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-02-28 07:02:34 +00:00
jingyu-ml 2e43c80609 [2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This MR adds HuggingFace checkpoint export support for LTX‑2 by treating
TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the
stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling
back to writing model.safetensors when save_pretrained isn’t available.
It also preserves the original forward in DynamicModule patching
(_forward_pre_dm) so downstream callers can still invoke the pre-patched
forward implementation.

**Changes**

1. Added the calibration & quantization support of the LTX2, even with
FP8 precision.
2. Preserve original forward before `DynamicModule` patching: when
patching forward, we now stash the pre-patched implementation in
`self._forward_pre_dm` (once) so downstream code can still call the
original forward, then re-bind forward to the class implementation. This
is needed for the LTX2 FP8 calibration.
3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats
ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a
“diffusion-like” object and routes it through
_export_diffusers_checkpoint() (import guarded; no hard dependency).
4. Generalized component discovery: introduced
get_diffusion_components() (aliasing the old get_diffusers_components)
to support non-diffusers pipelines; for LTX‑2 it returns only
stage_1_transformer.
5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy
forward generator (generate_diffusion_dummy_forward_fn) that builds
minimal LTX Modality inputs (including correct timesteps broadcasting)
so shared-input hooks can run and fuse QKV when applicable.
6. Export fallback for non-save_pretrained modules: when a component
lacks save_pretrained (LTX‑2 transformer), export now writes
model.safetensors + minimal config.json instead of pytorch_model.bin.

Plans

- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml 
- [ ] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml 

## Usage
<!-- You can potentially add a usage example below. -->
```bash
python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added LTX-2 video model support with complete quantization and export
pipeline integration
* Introduced `--extra-param` CLI option for flexible model configuration
and parameter passing
* Enhanced export capabilities with broader diffusion model
compatibility

* **Chores**
* Changed default model data type from Half to BFloat16 for improved
numerical stability

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-04 10:41:33 +00:00
jingyu-ml 668b8a19e8 [1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This PR adds support for exporting quantized diffusers models (DiT,
Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling
deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM.

**Changes**

New file: `diffusers_utils.py`
- Dummy input generation for various diffusion models
- Pipeline component extraction helpers
- QKV projection detection and grouping
- `hide_quantizers_from_state_dict()` context manager for clean saves

Refactored: `unified_export_hf.py`
- New `_fuse_qkv_linears_diffusion()` for QKV amax fusion
- `_export_diffusers_checkpoint()` to export full pipelines (models +
tokenizers + schedulers etc.)

Plans

- [x] [1/3] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml 
- [ ] [3/3] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml

## Usage
<!-- You can potentially add a usage example below. -->
```
mtq.quantize(pipe, quant_config, forward_call)
export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
* Added HuggingFace checkpoint export support for quantized diffusion
models with configurable output directory
* Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint
export destination
* Extended export functionality to support selective component exports
from diffusion pipelines
* Enhanced quantized model export with improved component handling and
multi-stage checkpoint generation

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:11:09 +00:00