Commit Graph
449 Commits
Author SHA1 Message Date
Chenjie LuoandZhiyu 5e43b2a5f5 Support Qwen3 Next MTP load and export (#860)
## What does this PR do?

Fix MTP export for Qwen3 Next

**Overview:** ?

For Qwen3 next, the MTP weights are not stored separately in
safetensors. So we use "mtp" weights key to decide if the weights are
for MTP or not.


## Testing
Qwen3 Next PTQ and check if MTP is in the exported checkpoint.

scripts/huggingface_example.sh --model
<Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Optimized Multi-Token Prediction weight loading with improved layer
detection and handling.

* **Chores**
* Simplified status reporting to display total loaded weights and
detected layers.
  * Removed verbose per-file warnings for cleaner console output.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Zhiyu <zhiyuc@nvidia.com>
2026-02-09 22:48:15 +00:00
yeyu-nvidia a8f5314c93 fix the path change in torch v2.10 for spec dec (#863)
## What does this PR do?

**Type of change:** 
bug fix

**Overview:** 
torch v2.10 changes the path for _SDPAMerger. will need to use the new
path for import

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated internal import references to reflect organizational changes
in dependencies.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-02-09 17:56:53 +00:00
Gwena Cunha 24e358789d [5868890][ONNX][Autocast] Fix: failure when checking input shape with unknown dimension (#859)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Skip unknown dimensions when comparing input shape in
model vs calibration data.

## Usage

```python
$ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```

## Testing
See bug 5868890.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes
* Enhanced input shape validation to properly handle dynamic tensor
dimensions, allowing more flexible dimension checking while maintaining
validation accuracy.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

## Aditional info
Regression introduced in
https://github.com/NVIDIA/Model-Optimizer/pull/652.

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-02-09 09:48:13 -05:00
kaix-nv e53ca61b71 Add contribution guidelines for experimental features (#867)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for experimental optimization technique
development, including recommended structure, testing conventions,
licensing requirements, and graduation path to production.

* **New Features**
* Introduced experimental package with templates and utilities for
implementing research-stage optimization techniques. Includes
configuration framework and example code patterns. Emits stability
warnings to indicate unstable APIs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-02-07 00:31:30 +00:00
Chenhan D. Yu 62c2799be2 Fix Sequential MLP amax sync deadlock (#862)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Bug fix

**Overview:** ?

After `QuantMoELayer`, we rely on `layer_sync_moe_local_experts_amax` to
first perform local sync. This is supposed to create
`input_quantizer.amax` for all experts but the current logic will only
update experts that already have `amax`. This results in some experts
are still missing `amax`.

With the fact above, `sync_quantizer_amax_across_dp_ep` will actually
deadlock seems the collective is called based on whether
`quantizer._amax is None`. Any expert with `None` amax will not call
collective hence will never arrive the collective and cause a deadlock.

We fix `layer_sync_moe_local_experts_amax` such that even if an expert
does not have `amax`, we will overwrite it with a clone of the global
amax. The post condition should be all experts have `amax` and the pre
condition of `sync_quantizer_amax_across_dp_ep` should be the same.

**Note:** we found that `_check_moe_calibration_complete` actually
didn't raise any error even some experts have no amax. Didn't look into
this problem.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved synchronization of quantization parameters for Mixture of
Experts (MoE) models with more flexible configuration support.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-02-06 14:53:52 -08:00
realAsma ac30686c82 Track global_amax for weight FP4 MSE sweep; Refactor to NVFP4StaticQantizer, NVFP4MSECalibrator (#849)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added NVFP4StaticQuantizer for improved 4-bit quantization with
enhanced precision control
* Introduced NVFP4MSECalibrator with flexible candidate generation for
calibration optimization

* **Improvements**
* Optimized GPU kernels for Hopper+ graphics cards with better
performance
  * Extended Triton support to broader GPU compatibility
* Enhanced backward compatibility for restoring previously quantized
models

* **Tests**
* Added comprehensive test coverage for new quantizers and calibration
methods

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-02-06 19:47:36 +00:00
yueshen2016 3393e981e6 Fix TEGroupedLinear quantization for expert parallelism (EP > 1) (#833)
## What does this PR do?

**Type of change:** Bug fix / Compatibility update

**Overview:**

Fix `te_grouped_quantized_linear_fn` argument parsing for
TEGroupedLinear quantization when parallelism configuration results in
fewer local experts per GPU.

### Problem
TransformerEngine changed the _GroupedLinear.forward signature in PR
#2377 (released in TE 2.10):
Old signature (TE < 2.10): forward(ctx, inp, m_splits: List[int],
use_bias, is_first_microbatch, ...)
New signature (TE >= 2.10): forward(ctx, inp, non_tensor_args: Tuple,
*weights_and_biases) where non_tensor_args = (m_splits, use_bias,
is_first_microbatch, ...)
Without this fix, ModelOpt's quantization code fails with newer TE
versions because it tries to access m_splits directly from args[idx +
1], but in TE >= 2.10, that position contains the non_tensor_args tuple
instead.


### Root Cause
The code assumed m_splits was always directly accessible at args[idx +
1], but TransformerEngine PR #2377 changed the signature to pack all
non-tensor arguments into a tuple.
Taking Qwen3-30B-A3B (with `num_gemms=21`, threshold=44) as an example:

### Solution
Added version checking to handle both signatures:
```python
if Version("2.10") <= _TE_VERSION:
    # New signature: non_tensor_args is a tuple, m_splits is the first element
    num_gemms = len(args[idx + 1][0])
else:
    # Old signature: m_splits is directly args[idx + 1]
    num_gemms = len(args[idx + 1])
```

## Usage
<!-- You can potentially add a usage example below. -->
Works seamlessly with any TransformerEngine version:

```python
# High EP quantization - previously failed, now works
torchrun --nproc_per_node 8 examples/quantization/quantize.py \
  --hf-model-id /models/Qwen3-30B-A3B \
  --export-quant-cfg fp8 \
  --megatron-save-path /models/Qwen3-30B-A3B_fp8_mlm \
  --tp 8 \
  --ep 8

# High EP inference - previously failed, now works  
torchrun --nproc_per_node 8 examples/quantization/ptq_generate.py \
  --megatron-load-path /models/Qwen3-30B-A3B_fp8_mlm \
  --hf-model-id /models/Qwen3-30B-A3B \
  --tp 8 \
  --ep 8
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
```python
# High EP quantization - previously failed, now works
torchrun --nproc_per_node 8 examples/quantization/quantize.py \
  --hf-model-id /models/Qwen3-30B-A3B \
  --export-quant-cfg fp8 \
  --megatron-save-path /models/Qwen3-30B-A3B_fp8_mlm \
  --tp 8 \
  --ep 8

# High EP inference - previously failed, now works  
torchrun --nproc_per_node 8 examples/quantization/ptq_generate.py \
  --megatron-load-path /models/Qwen3-30B-A3B_fp8_mlm \
  --hf-model-id /models/Qwen3-30B-A3B \
  --tp 8 \
  --ep 8
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced Mixture of Experts (MoE) calibration validation and
synchronization to ensure consistency across distributed training
setups.
* Improved grouped linear quantization robustness to handle varying
input patterns and tensor dimensions.

* **Improvements**
* Better error handling for incomplete MoE expert calibration detection.
  * More flexible argument parsing for quantization operations.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-02-06 19:43:37 +00:00
willg-nv 0225e5273f Integrate Automated QDQ placement tool - part 2.1 (#844)
## What does this PR do?

This PR implements RegionPattern class. RegionPattern describes local
topology structure of a Region. Regions with same Pattern could be
autotune together. Best insertion points of a given pattern could also
be saved to accelerate the next QDQ autotuning.

**Overview:** ?

## Usage

```python
python -m modelopt.onnx.quantization.autotune.region_search --model model.onnx --verbose
```
```
    ├─ Region 212 (Level 0, Type: COMPOSITE)
    │  ├─ Direct nodes: 0
    │  ├─ Total nodes (recursive): 9
    │  ├─ Children: 1
    │  ├─ Inputs: 3 tensors
    │  │    - xxx
    │  │    - xxx
    │  │    - xxx
    │  └─ Outputs: 1 tensors
    │       - xxx
    │
    │  Child regions:
    │
      ├─ Region 209 (Level 2, Type: LEAF) 
      │  ├─ Direct nodes: 9
      │  ├─ Total nodes (recursive): 9
      │  ├─ Children: 0
      │  ├─ Inputs: 11 tensors
      │  │    - xxx
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No, document
update is in Part 4
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
CHANGELOG update could be done after all changes are ready.

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Enhanced ONNX quantization analysis with improved region pattern
matching and comparison capabilities.
* Added utility to identify quantized tensors in models for better
analysis.

* **Tests**
* Comprehensive test coverage for region pattern functionality and
quantization utilities.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-02-06 10:00:46 -05:00
Zhiyu 452c5a09b0 GLM-4.7 MTP support (#792)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Enable GLM-4.7 PTQ workflow, including loading the
standalone MTP modules and export as-is.

## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/GLM-4.7 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/GLM-4.7-NVFP4-0203 --trust_remote_code
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added quantization support for GLM-4.7 model with automatic handling
of specialized layer architecture.
* Added image-text data calibration capabilities for Nemotron VL model
quantization.

* **Documentation**
* Updated support matrix to reflect newly supported models and
quantization features.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-02-04 15:43:24 -08:00
Keval Morabia 944dd1a284 Move parallel_state init and warnings to Quant DynamicModule + MBridge pruning doc update (#854)
## What does this PR do?

**Type of change:** Minor improvement <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

Only quantization DynamicModules use the parallel_state attribute so for
all other model opt methods, we see a parallel state not initialized
warning which could be confusing hence moving it to QuantModule class
instead

Minor update to MBridge pruning docs

## Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Distributed parallel state support is now available in quantization
workflows for multi-GPU training.

* **Bug Fixes**
* Improved resource cleanup in distributed training to ensure proper
environment finalization.

* **Documentation**
* Updated example paths and added new manual pruning configuration
examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-04 22:23:21 +00:00
Jenny Chen 8a8c250d12 Latent MOE & Repeated MTP support for NemotronH; fix KV cache quant export (#830)
## What does this PR do?

**Type of change:** New feature and bug fix

**Overview:** 

Support Latent MOE and Repeated MTP for NemotronH models
- Enable latent MOE modules during megatron import/export
- Fix KV cache quantization export: remove old
`qkv_layer.output_quantizer` export & replace with proper
`k/v_bmm_quantizer` logic
- Improvements to EP amax sync
- Support repeated MTP import/export for NemotronH models (only BF16
export for MTP for now)

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for grouped MLP and self-attention scaling operations in
model export workflows
* Enhanced model parallel training capabilities with improved component
mapping
* Expanded quantization configuration handling with dynamic module
exclusion across distributed ranks
  * Improved support for additional transformer engine components

* **Refactor**
* Reorganized internal export and import logic for improved
maintainability and specialist model architecture support

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: jenchen13 <jennifchen@nvidia.com>
Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
2026-02-04 20:18:00 +00:00
jingyu-ml 2e43c80609 [2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This MR adds HuggingFace checkpoint export support for LTX‑2 by treating
TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the
stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling
back to writing model.safetensors when save_pretrained isn’t available.
It also preserves the original forward in DynamicModule patching
(_forward_pre_dm) so downstream callers can still invoke the pre-patched
forward implementation.

**Changes**

1. Added the calibration & quantization support of the LTX2, even with
FP8 precision.
2. Preserve original forward before `DynamicModule` patching: when
patching forward, we now stash the pre-patched implementation in
`self._forward_pre_dm` (once) so downstream code can still call the
original forward, then re-bind forward to the class implementation. This
is needed for the LTX2 FP8 calibration.
3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats
ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a
“diffusion-like” object and routes it through
_export_diffusers_checkpoint() (import guarded; no hard dependency).
4. Generalized component discovery: introduced
get_diffusion_components() (aliasing the old get_diffusers_components)
to support non-diffusers pipelines; for LTX‑2 it returns only
stage_1_transformer.
5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy
forward generator (generate_diffusion_dummy_forward_fn) that builds
minimal LTX Modality inputs (including correct timesteps broadcasting)
so shared-input hooks can run and fuse QKV when applicable.
6. Export fallback for non-save_pretrained modules: when a component
lacks save_pretrained (LTX‑2 transformer), export now writes
model.safetensors + minimal config.json instead of pytorch_model.bin.

Plans

- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml 
- [ ] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml 

## Usage
<!-- You can potentially add a usage example below. -->
```bash
python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added LTX-2 video model support with complete quantization and export
pipeline integration
* Introduced `--extra-param` CLI option for flexible model configuration
and parameter passing
* Enhanced export capabilities with broader diffusion model
compatibility

* **Chores**
* Changed default model data type from Half to BFloat16 for improved
numerical stability

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-04 10:41:33 +00:00
jingyu-ml 87237e7dd1 Update on the QuantModule & DynamicModule to accept external forward (#824)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This MR improves robustness when `forward()` is monkey‑patched (replaced
at runtime) on modules that later get wrapped/converted by ModelOpt
(DynamicModule + quant wrappers).

It addresses two concrete failure modes introduced/exposed by supporting
“patched forward” modules:

1. Forward “leakage” after export: a dynamic wrapper forward could
remain bound on an instance even after export() restores the original
(non‑dynamic) class, causing runtime errors in unrelated codepaths (e.g.
KD export/save/restore chains).

1. Infinite recursion in quant wrappers: _forward_pre_dm can sometimes
point to a wrapper forward that already participates in the class chain,
causing a recursion loop when quant wrappers call _forward_pre_dm
directly.

## Usage
<!-- You can potentially add a usage example below. -->

```python
lin = torch.nn.Linear(4, 4)

def upcast_forward(x):
    # external closure: NOT part of any class MRO
    return torch.nn.functional.linear(x, lin.weight.to(x.dtype), lin.bias.to(x.dtype))

lin.forward = upcast_forward  # framework/user patches forward

# Later, ModelOpt converts/wraps the module.
# It stashes the patched function as `_forward_pre_dm` and binds the wrapper forward on the class.

# During quantization, QuantInputBase.forward sees `_forward_pre_dm` is NOT in MRO -> calls it.
```

```
# Imagine a module already wrapped by quant classes:
# QuantLinearConvBase.forward -> super().forward -> QuantInputBase.forward -> ...

# If `_forward_pre_dm` accidentally points to QuantLinearConvBase.forward (which IS in MRO),
# and QuantInputBase.forward calls it directly, you get:
# QuantInputBase.forward -> _forward_pre_dm (QuantLinearConvBase.forward)
# -> super().forward -> QuantInputBase.forward -> ...
# infinite recursion

# The fix: if `_forward_pre_dm` is a forward already in MRO, ignore it and use super().forward.
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Bug Fixes**
* Improved forward method restoration during module export to prevent
state leakage
  * Enhanced quantization behavior when using chained optimization modes

* **Tests**
* Added regression tests for quantization with runtime forward patching
* Added validation tests for sparse quantization combined with
distillation workflows

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
0.42.0rc0
2026-02-03 21:25:12 -08:00
noeyy-mino 9de9b8f20c Noeyy/add test cases for the newly added checkpoints on HF (#827)
## What does this PR do?

**Type of change:** new tests

**Overview:** Add new test cases for the newly added checkpoints on
HuggingFace.

## Usage
pytest test_deploy.py --run-release

```python
None
```

## Testing
None

## Before your PR is "*Ready for review*"


- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
None


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for NVFP4 model variants across multiple model families
(DeepSeek, Llama, Qwen, and others).

* **Improvements**
* Enhanced backend availability detection to automatically identify and
manage supported deployment backends at runtime.

* **Tests**
* Improved test infrastructure for better reproducibility and backend
compatibility handling.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2026-02-04 09:27:49 +08:00
realAsma e247f5d0e4 Fixes for Megatron Expert Parallel, GroupedMLP and SequentialMLP (#831)
## What does this PR do?

**Type of change:** Bug fix / Improvement

**Overview:** 

Fix MoE quantization calibration sync by removing the force-routing
workaround and adding explicit validation for incomplete calibration.

**Problem:** During MoE calibration, some experts may not receive tokens
(router doesn't select them). This causes amax=None on some ranks while
others have valid values, leading to hangs or failures during
distributed amax sync.

Previous workaround: Force all tokens through all experts during
calibration. This was however causing the following error:

```
File "/opt/TensorRT-Model-Optimizer/modelopt/torch/quantization/plugins/transformer_engine.py", line 152, in te_grouped_quantized_linear_fn
    quantized_inputs = self.input_quantizer(inp)

File "/opt/TensorRT-Model-Optimizer/modelopt/torch/quantization/calib/max.py", line 69, in collect
    assert torch.all(local_amax >= 0), (

torch.AcceleratorError: CUDA error: an illegal memory access was encountered
```
This is probably because forcing all tokens through all experts, the
inputs become garbage and are possibly inf/nan causing calibration to
fail.


Solution:
Remove _QuantMoELayer force-routing workaround
Add validation before sync: detect if some ranks have amax=None while
others have values
Raise clear error: "MoE calibration incomplete: increase --calib-size" -
This is a cleaner solution. In case of under calibration we just raise a
clear error.


## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes (Verified backward
compatibility by loading a MoE model saved before this change)
- **Did you write any new necessary tests?**: No - existing MoE tests
cover the sync behavior
- **Did you add or update any necessary documentation?**: Not needed,
low level change
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No needed, low level change

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Improvements**
* Enhanced Mixture of Experts (MoE) quantization with comprehensive
calibration validation to ensure consistent synchronization across
distributed experts.

* **Refactor**
* Streamlined MoE quantization architecture by consolidating internal
handling mechanisms.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-02-03 16:44:22 -08:00
Keval Morabia e02409773b Increase nighytly gpu test timeout to 150mins
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-03 01:17:57 -08:00
Keval Morabia fb5923c89e Add Megatron-Bridge pruning example scripts (#800)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

Megatron-Bridge pruning example scripts (HF input, HF / Megatron
output). Also defined some utility functions we can reuse for adding
examples for quantization or other optimizations:
- `modelopt.torch.utils.plugins.mbridge.load_mbridge_model_from_hf`:
Load HF to MBridge with ModelOpt spec in desired TP/PP/etc configuration
-
`modelopt.torch.utils.plugins.mbridge.get_hf_mbridge_calibration_loop`:
Create `forward_loop` for calibration on a HF dataset
- Supports all datasets available in
`modelopt.torch.utils.dataset_utils` (`cnn_dailymail`,
`nemotron-post-training-dataset-v2`, etc)
  - Support applying chat template for chat-based data

## Usage
<!-- You can potentially add a usage example below. -->

From `nvcr.io/nvidian/nemo:26.02.rc1` container (mount latest code to
`/opt/Megatron-Bridge` and `/opt/Model-Optimizer`)

```python
torchrun --nproc_per_node 2 /opt/Model-Optimizer/examples/megatron_bridge/prune_minitron.py \
    --hf_model_name_or_path Qwen/Qwen3-8B \
    --prune_target_params 6e9 \
    --hparams_to_skip num_attention_heads \
    --output_hf_path /tmp/Qwen3-8B-Pruned-6B
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] Manually ran pruning script in nemo:25.11 container (plus modelopt
and mbridge mounted to latest) for Qwen3-8B and Nemotron-Nano-9B-v2 with
PP=8 and PP=4
- [ ] Added per-PR CI/CD test for example script

Results when pruning Qwen3 8B -> 6B (10 different configurations) with
and without chat template on the dataset samples. Perhaps MMLU is not
the right metric to look at.

| Layers | Hidden Size | FFN Hidden Size | Params | MMLU (Concatenated
messages) | MMLU (Applied Chat Template) |

|--------|-------------|-----------------|--------|------------------|----------------------|
| 34 | 3328 | 11264 | 5.99B | 0.401 | 0.393 |
| 30 | 3584 | 11776 | 5.99B | 0.588 | 0.576 |
| 36 | 3840 | 8192 | 5.98B | 0.507 | 0.518 |
| 36 | 3584 | 9216 | 5.98B | 0.477 | 0.469 |
| 36 | 3072 | 11776 | 5.97B | 0.255 | 0.249 |
| 32 | 3584 | 10752 | 5.96B | 0.554 | 0.549 |
| 28 | 4096 | 10240 | 5.94B | 0.400 | 0.438 |
| 36 | 4096 | 7168 | 5.93B | 0.461 | 0.438 |
| 36 | 3328 | 10240 | 5.92B | 0.362 | 0.359 |
| 34 | 3840 | 8704 | 5.91B | 0.515 | 0.546 |

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: ‼️ TODO
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

**New Features**
* Added new Megatron-Bridge pruning example demonstrating Minitron-based
model optimization with advanced pruning configurations.

**Documentation**
* Updated core project documentation to highlight Megatron-Bridge as a
supported optimization framework.
* Added comprehensive example documentation for Megatron-Bridge
workflows including pruning, distillation, and quantization.
* Updated pruning guides with Megatron-Bridge integration examples and
best practices.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-03 04:30:05 +00:00
willg-nv 23278a44db Integrate Automated QDQ placement tool - Part 1 (#701)
## What does this PR do?

**Type of change:** new feature

**Overview:** This PR integrates an automatical QDQ placment tool into
ModelOpt.

This PR is the 1/4 parts of the change, it contains the following
changes:
1. Defines common types: Region, RegionType, Error types
2. Defines InsertionPoints (the logical localtion to place QDQ pairs),
InsertionScheme (a set of insertion points)
3. Unit tests for new types

Part 1: https://github.com/NVIDIA/Model-Optimizer/pull/701
Part 2: https://github.com/NVIDIA/Model-Optimizer/pull/702
Part 3: https://github.com/NVIDIA/Model-Optimizer/pull/703
Part 4: https://github.com/NVIDIA/Model-Optimizer/pull/704

## Usage

```python
        # Region type usage:
        region = Region(region_id=1, level=0, region_type=RegionType.LEAF)
        assert region.get_id() == 1
        assert region.get_level() == 0
        region.add_node(1) # 1 is the index of ONNX graph node
        ...

        point = NodeInputInsertionPoint(node_index=0, input_index=2)
        assert point.node_index == 0 # relative node index in region
        assert point.input_index == 2 # relative input tensor index in specific node
        resolved = point.resolve(region, graph)
        ...
```

## Testing
Implement unit tests, all tests could get passed.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No, document
change will be included in part 4.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No, this could be done when all parts of the change are merged.

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added foundational autotuner infrastructure for quantization
optimization, including region hierarchies and insertion scheme
management.
* Introduced insertion point system for managing quantize/dequantize
operation placement across ONNX graph regions.
* Added utility functions for tensor consumer mapping and boolean
operation identification.

* **Tests**
* Added comprehensive test coverage for autotuner components, insertion
points, and region management.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-02-02 19:36:05 -08:00
Asha Anoosheh 2a467531f2 Layerwise KD mode (#802)
## What does this PR do?

**Type of change:** new feature

**Overview:** Add a subclass of `DistillationModel` which implements
slightly different hooks to inject teacher tensors into corresponding
student layers for module replacement purposes, as opposed to logits
distillation.

## Usage

```python
mtd.convert(model, mode=[("layerwise_kd", config)])
```

## Testing
New units

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Introduced bypass-enabled knowledge distillation mode with layer-level
loss mapping for fine-grained model optimization control.
* Added model export functionality with automatic cleanup of
intermediate activation capturing mechanisms.

* **API Changes**
* New bypass_kd mode configuration option available for advanced
knowledge distillation workflows.
  * Updated model export interface for improved lifecycle management.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-02-02 18:58:42 +01:00
Hrishith Thadicherla fc6a211b4a Added column-major storage of weights and scales in INT4 quantization for model load time improvement in TRT-RTX (#811)
## What does this PR do?

**Type of change:** ? New feature

**Overview:** 
TensorRT-RTX requires the weights and scales in the ONNX models to be in
column-major format. So whenever the model loads TRT-RTX JIT transposes
the weights and scales during load time, causing increased load time.

Proposed feature is after quantization, transpose the weights and scales
in DQ node and add a transpose node right after i.e,
A × B = A × ((Bᵀ)ᵀ)

The transformation is post processing step and is disabled by default.
It can be enabled by quantizing with --use_column_major

## Usage
```
python -m modelopt.onnx.quantization --onnx_path "model.onnx" --output_path "model_quant.onnx" --quantize_mode int4 --calibration_method awq_lite --use_column_major --skip_shared_constants_duplication
```

## Testing
Tested a few LLM's and their MMLU scores with and without this
transformation. No degradations were observed.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added --use_column_major CLI flag to enable column-major weight
storage optimization (applies to DQ-only quantization paths).

* **Documentation**
  * CLI docs updated to describe the new flag and its applicability.

* **Tests**
* New unit tests validating column-major transformation behavior and
output equivalence.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-02-02 20:24:47 +05:30
sugunav14 02c5f292f0 GPTQ Lite implementation (#555)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** Adds support for GPTQ algorithm. This PR implements a
modified version of the official GPTQ algorithm; the key difference is
that updated activations from each layer are not used for hessian
computation

## Usage
<!-- You can potentially add a usage example below. -->
Modify "algorithm" field in quant_cfg to "gptq_lite".

Note: Does not currently work with AWQ

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- [x] Added unit tests to test helper functions + e2e flow
- [x] Perplexity and GPQA results

| Model       | Qformat               | Perplexity wikitext2 | GPQA |
|-------------|------------------------|------------------------|------|
| Qwen3-8B | INT4 weight only (modelopt + amax/7) no GPTQ | 10.75 | n/a
|
| Qwen3-8B | INT4 weight only (modelopt + amax/7) | **10.56** | 0.388 |
| Qwen3-8B | INT4 weight only + FP-Quant hessians + amax/7.5 | 10.25 |
**0.449** |
| Qwen3-8B | INT4 weight only (FP-Quant) | 10.24 | 0.46 |
| Qwen3-8B | NVFP4 static weight only | 10.25 | n/a |
| Qwen3-8B | NVFP4 static weight only no GPTQ | 10.25 | n/a |
| Qwen3-0.6B | NVFP4 static weight only | **22.75** | n/a |
| Qwen3-0.6B    | NVFP4 dynamic weight only |  23.50          | n/a   |
| Qwen3-0.6B | NVFP4 static weight only with FP-Quant hessians | 22.0 |
n/a |
| Qwen3-0.6B | NVFP4 static weight only no GPTQ | 24.25 | n/a |

Conclusions from results
- Perplexity remains the same or shows improvement with Modelopt
implementation. The magnitude of improvement is lesser in modelopt when
compared to FP-Quant
- GPQA shows no improvement with modelopt, but shows improvement with
FP-Quant



## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* GPTQ Lite quantization mode now available for efficient model
calibration
  * GPU memory usage monitoring utility added
* Quantization configuration extended to support complex nested
structures and lists

* **Tests**
  * Comprehensive test coverage added for GPTQ quantization workflows

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-01-30 02:15:17 +00:00
yeyu-nvidia 81b67ddf06 Context parallelism for Megatron core models (#818)
## What does this PR do?

New feature

**Overview:** 
This PR implements the context manager which injects attn_mask as
attn_bias to TEDotProductAttention so that we can enable EAGLE training
with arbitrary mask.

## Usage
set CP>1 in
https://github.com/NVIDIA/Megatron-LM/blob/main/examples/post_training/modelopt/finetune.sh

```python
# Add a code snippet demonstrating how to use this
```

## Testing
Tested on DSR1 Llama 8B.
CP1->CP2
38854MB->28050MB
MTbench AL 2.26->2.31

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## New Features
* Context parallelism support added for Eagle speculative decoding with
HuggingFace and Megatron Core models.
* Model checkpoint loading enhanced to enable remote code execution
capabilities when required.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-01-29 19:47:22 +00:00
Asha Anoosheh 770962b32c Rename MLM teacher arg (#829)
## What does this PR do?

**Type of change:** Refactor

**Overview:** MLM arg changed from `--teacher-model-config` to
`--export-kd-teacher-model-config` for consistency

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated export flag naming for knowledge distillation teacher model
configuration.
  * Adjusted default top-k parameter for logits selection to 1024.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-01-29 17:39:50 +01:00
binghanc 58abdc2f38 Support MLA nvfp4 quant for Deepseek for max perf (#582)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** support for newer checkpoints

## Usage
<!-- You can potentially add a usage example below. -->

```python
torchrun --nproc-per-node=8 ptq.py --mla_quant nvfp4_wq_a_wkv_a_wq_b_wo_fp8_wkv_b --batch_size 4 --model_path $DS_CKPT --config DeepSeek-V3/inference/configs/config_671B.json --quant_cfg NVFP4_DEFAULT_CFG --output_path $AMAX_PATH
```



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
  * Added NVFP4 quantization option for MLA model quantization workflow.
* Expanded quantization configuration choices to include "nvfp4"
alongside existing per_tensor_fp8 option.
* Introduced new CLI parameter to specify MLA quantization type during
post-training quantization.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: binghanc <176802681+binghanc@users.noreply.github.com>
2026-01-28 23:21:12 -08:00
Chenhan D. Yu 9857e0a48c Nenotrom Nano PTQ fix where MoELayer forward has additional named arguments (#823)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Bug fix

**Overview:** ?

Our `QuantMoELayer` dynamic module override the forward. In the latest
version of `megatron.core` `MoELayer.forward` has additional argument ;
hence resulting in failure. Since we are passing through all the
arguments, here we change to use *args and **kwargs to avoid future
issue.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Updated argument propagation in the mixture-of-experts quantization
layer to flexibly support additional calibration-related and padding
parameters while maintaining backward compatibility.
* Improved routing behavior with conditional top-k adjustments that
activate when group routing configurations are in use.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-28 20:56:54 +00:00
danisereb 4227bb7366 Add support for MXFP8 PTQ (#736)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** Add support for MXFP8 PTQ, enabling MXFP8 hardware
acceleration during inference on Blackwell GPUs.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
export MODEL_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B
export OUTPUT_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B-MXFP8
mkdir -p $OUTPUT_PATH

python examples/llm_ptq/hf_ptq.py \
--export_fmt hf \
--dataset cnn_dailymail \
--pyt_ckpt_path $MODEL_PATH \
--export_path $OUTPUT_PATH \
--qformat mxfp8
```

The `hf_quant_config.json` of the output checkpoint:
```json
{
    "producer": {
        "name": "modelopt",
        "version": "0.41.0.dev50+g7a796a875"
    },
    "quantization": {
        "quant_algo": "MXFP8",
        "kv_cache_quant_algo": "FP8",
        "group_size": 32,
        "exclude_modules": [
            "lm_head"
        ]
    }
}
```

And `config.json` (only the `quantization_config`):
```json
...
    "quantization_config": {
        "ignore": [
            "lm_head"
        ],
        "quant_algo": "MXFP8",
        "kv_cache_scheme": {
            "dynamic": false,
            "num_bits": 8,
            "type": "float"
        },
        "producer": {
            "name": "modelopt",
            "version": "0.41.0.dev50+g7a796a875"
        },
        "quant_method": "modelopt"
    }
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
Used `hf_ptq.py` to quantize the model `nvidia/OpenMath2-Llama3.1-8B`
([available in
hugging-face](https://huggingface.co/nvidia/OpenMath2-Llama3.1-8B)), see
the example command above.

Checked that the generated MXFP8 checkpoint can be loaded with vLLM
(required changes in vLLM, not merged to main).

Added tests for `MXFP8QTensor` in
`tests/gpu/torch/quantization/test_qtensor_cuda.py`.
Added "mxfp8" in `‎tests/examples/llm_ptq/test_llm_ptq.py`

#### Support for Nemotron Models

Verify that Nemotron Nano V3 BF16 can be converted to MXFP8 using
`hf_ptq.py`:
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added MXFP8 quantization format support with new scaling mechanisms
and quantization utilities.
* Updated configuration options, example scripts, and utilities to
recognize and process MXFP8 quantization workflows.
* Extended quantization export pipelines to handle MXFP8 quantized
models.

* **Tests**
* Expanded test coverage for MXFP8 quantization across various tensor
shapes, data types, and device configurations.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
2026-01-28 11:55:19 -08:00
vishalpandya1990 4a848c4f52 Modelopt-windows documentation update (#812)
## What does this PR do?

Documentation

**Overview:**

- Update support matrix, changelog, deployment page, example readmes as
per recent feature and model support on Windows side.

## Testing
- No testing, its just documentation change

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ONNX Mixed Precision Weight-only quantization (INT4/INT8)
support.
  * Introduced diffusion-model quantization on Windows.
  * Added new accuracy benchmarks (Perplexity and KL-Divergence).
* Expanded deployment with multiple ONNX Runtime Execution Providers
(CUDA, DirectML, TensorRT-RTX).

* **Bug Fixes**
* Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ
quantization.

* **Documentation**
* Updated installation guides with system requirements and multiple
backend options.
* Reorganized deployment documentation with comprehensive execution
provider guidance.
* Expanded example workflows with improved setup instructions and
support matrices.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-01-28 10:48:59 +05:30
Gal Hubara-AgamandGwena Cunha 2c73de0405 [5525939] Allow user to select target opset in MOQ (#809)
## What does this PR do?

**Type of change:** new feature 

**Overview:** 
- Allow user to select the target opset
- Minimum opset will be defined according to quantization mode
- Add tests in tests/unit/onnx/test_quantize_api.py

## Testing
Added unit tests

tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[int8]
PASSED [ 11%]

tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[fp8]
PASSED [ 22%]

tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[int4]
PASSED [ 33%]

tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[int8]
PASSED [ 44%]

tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[fp8]
PASSED [ 55%]

tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[int4]
PASSED [ 66%]
tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[int8]
PASSED [ 77%]
tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[fp8]
PASSED [ 88%]
tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[int4]
PASSED [100%]


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes - auto
update according to argparser help
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
Requested as a WAR for a Windows-onnxruntime issue in 5525939, but
regardless, it's a useful feature to have

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `--opset` CLI option enabling users to specify target ONNX opset
version when quantizing models.
* Automatic validation ensures the opset version is compatible with
quantization requirements, with warnings when adjustments are made.

* **Tests**
* Added comprehensive test coverage for opset version handling across
quantization workflows.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
Signed-off-by: Gal Hubara-Agam <96368689+galagam@users.noreply.github.com>
Co-authored-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
2026-01-27 17:11:12 +00:00
Keval Morabia 1c7a928df8 Change cnn_dailymail to abisee/cnn_dailymail (#819)
cnn_dailymail is changed to abisee/cnn_dailymail long ago. Perhaps older
name no longer works

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated dataset configuration path for improved dataset access.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-27 20:05:15 +05:30
Frida Hou aafd388394 add FP8 sweep option for static NVFP4 MSE (#758)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature

**Overview:** ?
Adds fp8_scale_sweep mode to MSE calibrator for optimizing FP8-quantized
per-block scales in NVFP4 format.


## Usage
<!-- You can potentially add a usage example below. -->

Tested with this config
```python
NVFP4_WEIGHT_MSE_FP8_SWEEP_CFG = {
    "quant_cfg": {
        "*weight_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        "*input_quantizer": {
            "enable": False,
        },
        **_default_disabled_quantizer_cfg,
    },
    "algorithm": {
        "method": "mse",
        "fp8_scale_sweep": True,
    },
}
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

**New Features**
- Added FP8 scale sweep option for quantization calibration, enabling
optimized scale value sweeping for NVFP4 per-block quantization.
- Introduced new NVFP4_WEIGHT_MSE_CFG configuration preset for improved
weight quantization workflows.

**Tests**
- Added test coverage validating FP8 scale sweep functionality and reset
behavior.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-01-26 14:39:55 -08:00
Gal Hubara-Agam 38403095c4 [5725362] AutoCast Fixes for models with external data (#731)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Fix AutoCast ReferenceRunner to handle large models.
Models above 2GB cannot be serialized to string, which is what
polygraphy is doing under the hood. Use a temporary file instead to save
the modified onnx with all tensors marked as outputs.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Improvements**
* Enhanced model processing to better support large ONNX models during
validation and runtime execution
* Added diagnostic logging of model sizes at key processing stages for
improved debugging and performance monitoring

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
2026-01-26 22:52:24 +02:00
Zhiyu 0ebcd70878 Support VLM calibration with image-text data (#755)
## What does this PR do?

**Type of change:** New feature

**Overview:** 

The primary goal of this PR is to allow the model optimizer to use
image-text pair data during the calibration phase of quantization, which
is likely help improve accuracy of quantized VLMs like Nemotron VL on
visual understanding tasks particularly, compared to text-only
calibration data.

- New Feature: Adds support for VLM calibration specifically using
image-text data.
- Dataset Integration: Introduces support for sampling from the
`Nemotron-VLM-Dataset-v2`.
- Refactoring: Created a separate utility for VLM datasets to keep the
main Hugging Face PTQ script (`hf_ptq.py`) clean.
- Simplified logic for handling multimodal inputs.
- Addressed specific issues encountered when calibrating the
`Nemotron-Nano-VL-12B-V2` model with image data.
- Documentation: Updated the README to include instructions and examples
for VLM calibration.

This PR complements https://github.com/NVIDIA/Model-Optimizer/pull/347
and we will consolidate llm_ptq and vlm_ptq examples in follow-up PRs.

## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 hf_ptq.py   --pyt_ckpt_path /home/scratch.omniml_data_2/models/Nemotron-Nano-VL-12B-V2   --qformat nvfp4   --export_path /home/omniml_data_3/zhiyuc/checkpoints/Nemotron-Nano-VL-12B-V2-NVFP4-doccalib   --trust_remote_code   --kv_cache_qformat none --calib_with_images   --vlm_dataset nemotron_vlm_dataset_v2   --vlm_subsets sparsetables,plotqa_cot   --calib_size 512
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Not yet <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Vision-Language Model (VLM) calibration support with image-text
pair data, specifically for Nemotron VL models.
* Added new `--calib_with_images` CLI flag to enable image-based
calibration workflows.
* Integrated Nemotron VLM dataset v2 for streaming multimodal
calibration data.

* **Documentation**
* Added VLM calibration guidance in the PTQ README with usage examples
and dataset information.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-01-26 09:46:26 -08:00
Gwena Cunha 8c36f5a2f7 [5676209][ONNX][Autocast] Add support for single npz file with multiple samples (#815)
## What does this PR do?

**Type of change:** New feature

**Overview:** Currently, Autocast only supports calibration data with
shape matching the model's input. This PR adds support for calibration
data with shape that is a multiple of the model's input. It does so by
re-arranging the data as such that it contains multiple samples with
shape matching the model's input.

Simplified example:
- ONNX input: `[1, 3, 224, 224]`
- Calibration data: `[10, 3, 224, 224]`
- Calibration data with multiple samples: `[1, 3, 224, 224] * 10`

## Usage
Single `npz` file with multiple samples:
```sh
$ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```

## Testing
See bug 5676209.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
Equivalent support is already included in the quantization workflow:
https://github.com/NVIDIA/Model-Optimizer/blob/1cc8e6bf3917f61500e81d4ded0af5d5a00e2e25/modelopt/onnx/quantization/calib_utils.py#L50


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* CalibrationDataProvider now accepts both file paths and pre-loaded
ONNX models as input.
* NPZ calibration files can now provide multiple batches for calibration
workflows.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-26 10:20:14 -05:00
ynankani-nv 5cc2a54519 Ynankani/update windows benchmark md (#762)
## What does this PR do? 
**Type of change:** ? documentation

**Overview:** Md update to add perplexity and kl divergence benchmark
info.




## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: NA
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Expanded accuracy comparison section with three detailed benchmark
metrics: MMLU scores, Perplexity (PPL), and KL-divergence.
* Added comprehensive tables showing results across models and
quantization configurations.
  * Included evaluation guides and references for each metric.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: unknown <ynankani@nvidia.com>
2026-01-25 06:23:25 +00:00
h-guo18 3036a9ea9f Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?

**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp: 
  - torch2.8.0
  - transformers5.0.0
  - accelrate1.12.0 
 - Move to FSDP2 
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
 - Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`

## Usage
<!-- You can potentially add a usage example below. -->

```bash
./launch_train.sh --model $MODEL \
            --output_dir $OUTPUT_DIR \  
            --data $DATA \
            --num_epochs 0.1 \
            --train_bs 1 \
            --eagle_config eagle_config.json \
            --training_seq_len 1024 \
            --cp_size 2   #newly added
```

## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_out vs out: 0.001953125
  Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
  Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
  Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
  Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
  Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```

- E2E Training Acc
  (Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>

- Peak Mem Reserved
   (llama3.1-8B, 8xH100, train_length=4k)

    | cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
    |----|--------------------------|--------------------------|
    | 1  |         65040.20              |79018.00
    | 2  |           50409.17             |73098.00
    | 4  |              45120.92            |72052.00
    | 8  |              38882.12            |66484.00

- Max Training Length test
  (llama3.1-8B, H100)

  | cp_size               | 6k  | 12k | 24k  | 48k  |
  |--------------------|-----|-----|-----|-----|
  | 1         | ✅ | OOM | OOM | OOM  |
  |2      | ✅  | ✅ | OOM | OOM |
  | 4         | ✅ | ✅ | ✅  | OOM  |
  | 8         | ✅ | ✅ | ✅  | ✅  |

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.

* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.

* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
  * Added FSDP configuration file for distributed training setup.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-01-24 02:45:50 +00:00
realAsma 04165ace59 [Minor] Force 'fuse_wgrad_accumulation' to false for TE GroupedLinear (#814)
## What does this PR do?

**Type of change:** ? Minor

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Automatically disables fuse_wgrad_accumulation when using ModelOpt
quantization with Transformer Engine-based quantization paths. A warning
is now displayed to notify users when this adjustment occurs.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-01-23 23:59:01 +00:00
yueshen2016 044c4bc4eb Support megatron generate for vlm (#773)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ? This PR adds feature of VLM generation for
megatron_generate

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Vision Language Model support to text generation pipeline,
enabling simultaneous processing of image and text inputs during both
generation and prefill operations.

* **Improvements**
* Enhanced data flow to properly route multimodal inputs (images and
text tokens) through generation paths with automatic detection and
handling of vision-enabled model architectures.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-01-23 15:02:59 -08:00
Zhiyu 4f4558adbb Fix a nvfp4 weight amax attribute issue during export (#785)
## What does this PR do?

**Type of change:** Bugfix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Fix a nvfp4 weight amax attribute issue during export,
especially when calibration size is small. Context:
https://github.com/sgl-project/sglang/issues/14677#issuecomment-3712750444

## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 hf_ptq.py --pyt_ckpt_path /home/scratch.jingyux_coreai/kimi-k2/models/Kimi-K2-Thinking-BF16 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/Kimi-K2-Thinking-NVFP4 --kv_cache_qformat none --calib_size 20 --trust_remote_code --dataset cnn_dailymail
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes
* Improved weight quantizer calibration to ensure quantizers are
properly initialized with calibration statistics before computing
scaling factors.
* Enhanced reliability and consistency of quantized model exports.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-01-23 08:58:04 +00:00
Chenhan D. Yu 2a08622059 Fix moe amax remedy for dsr1 and remove global barrier in quantization megatron plugins (#808)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug

**Overview:** ?

This PR fix 2 bugs which impact DeepSeek calibration as well as PP
forward of MoE models.

1. The WAR in `MoELayer` that change the `topk` to `num_experts` only
works if no group-topk (a.k.a group routing) is used. Only changing topk
will lead to out-of-range error since `topk` can never be `num_experts`
when `group_topk != None`. Currently only `DeepSeek-V3` uses
`group_topk` and DeepSeek-V3 does not have difficulty to calibrate all
experts. As a result, we disable the WAR when detecting `group_topk`.

2. A previous PR inserted global barrier in
`quantization.plugin.megatron`
https://github.com/NVIDIA/Model-Optimizer/commit/6ef9954db1e73b8c4a86e5bfd31c954cfa21db61#diff-0fa2ba4ecc36c5ff031be9f9a5af080e7aa3afa331c438f02f501b9432ec6d6aL228-R515
This leads to dead lock when using PP since PP rank will never be able
to sync during pipeline forward. For MoE, this can be even worse if the
barrier is only visited by some EP/PP rank. Using collective
communication over the global world (a.k.a global comm) in megatron
plugin should be prohibited. Using collective on sub communication group
should avoid using `megatron.core.parallel_state` (a.k.a `mpu`) in the
future. Instead, use the local `pg_collection` from each module. Any
usage of collective communication must be inspected carefully with test
as PP, TP, and EP.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-23 02:11:19 +00:00
sychen52 b44c60ad34 Svdquant huggingface checkpoint export support (#754)
## What does this PR do?

**Type of change:** new feature

**Overview:** 

## Usage

```bash
cd ./examples/llm_ptq/
python hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3-4B \
    --export_path /home/scratch.shiychen_coreai/quantized_models/Qwen3-4B-svdq \
    --qformat nvfp4_awq_svdquant --kv_cache_qformat none --sparsity_fmt dense --calib_size 8
```

## Testing
exported checkpoint and loaded.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added nvfp4_svdquant as a new quantization format option for LLM model
quantization workflows.

* **Limitations**
* Multi-GPU export configurations using tensor or pipeline parallelism
are not supported with nvfp4_svdquant quantization.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-01-22 11:19:33 -08:00
jingyu-ml 945ee02f8d [1/3] Add the fastvideo support (#804)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

FastVideo is a new diffusion-focused framework that we plan to integrate
with. In this work, we added initial support for WAN 2.2 5B in
FastVideo, targeting the text-to-video use case.

For the Conv layer type, we currently use a straightforward direct
convolution call. Implicit GEMM quantization is intentionally omitted in
this first MR and will be addressed in a follow-up MR.

- [x] [1/3] Added support for the WAN 2.2 DIT + VAE layer type.
- [ ] [2/3] Added calibration support for them in the example script,
add test cases and README, doc.
- [ ] [3/3] Submitted an MR to fastvideo to enable quantization-aware
training.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added FastVideo plugin support to the quantization framework. Users
can now apply quantization to FastVideo-specific layers with specialized
weight quantization handling, optimized input processing, and caching
features for enhanced inference performance.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:17:59 +00:00
jingyu-ml 668b8a19e8 [1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This PR adds support for exporting quantized diffusers models (DiT,
Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling
deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM.

**Changes**

New file: `diffusers_utils.py`
- Dummy input generation for various diffusion models
- Pipeline component extraction helpers
- QKV projection detection and grouping
- `hide_quantizers_from_state_dict()` context manager for clean saves

Refactored: `unified_export_hf.py`
- New `_fuse_qkv_linears_diffusion()` for QKV amax fusion
- `_export_diffusers_checkpoint()` to export full pipelines (models +
tokenizers + schedulers etc.)

Plans

- [x] [1/3] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml 
- [ ] [3/3] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml

## Usage
<!-- You can potentially add a usage example below. -->
```
mtq.quantize(pipe, quant_config, forward_call)
export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
* Added HuggingFace checkpoint export support for quantized diffusion
models with configurable output directory
* Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint
export destination
* Extended export functionality to support selective component exports
from diffusion pipelines
* Enhanced quantized model export with improved component handling and
multi-stage checkpoint generation

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:11:09 +00:00
Keval Morabia 563a1e09c6 Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?

**Type of change:** New feature 

- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
  - Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.

## Usage
<!-- You can potentially add a usage example below. -->

Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
    model,
    mode="mcore_minitron",
    constraints=constraints,
    dummy_input=None,  # Not used
    config=config,
)
```

1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```

2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}

# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
    return megatron_mmlu(m, tokenizer, percentage=0.05)  # 5% sampled data for faster eval
config["score_func"] = score_func

# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune

mtp.prune(...)
```

To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
    hidden_size_divisor=256,
    ffn_hidden_size_divisor=512,
    mamba_head_dim_divisor=8,
    num_moe_experts_divisor=8,
    num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```

## Testing

**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
     0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
     0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
     0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
     0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
     0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
     0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
     0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
     0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
     0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```

**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
     0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads':  88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
     0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
     0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads':  96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
     0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
     0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
     0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
     0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
     0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads':  80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
     0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```

## Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information

OMNIML-3043

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
  * Introduced parameter counting capabilities for architecture search

* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
  * Updated configuration documentation with granular divisor parameters

* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-21 22:34:39 +00:00
Keval Morabia 792806f00e Add Security considerations in docs (#803)
Add security considerations to docs suggested by Nvidia Security team

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* Added comprehensive security considerations documentation for ModelOpt
covering multiple risk areas including untrusted input handling,
deserialization safety, input validation, resource exhaustion
prevention, data protection for transit and storage, logging and
observability practices, supply chain security, and structured
mitigation approaches with practical implementation examples.


<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-21 12:52:44 -08:00
Hrishith Thadicherla 67576d20ba Revert onnxruntime-gpu version to 1.22.0 for Windows (#801)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** 
Reverted the windows ort version in setup.py to 1.22, since I later
observed some regressions with vision models.
Keeping ORT version as 1.23 in windows examples since i didn't face any
issues with LLM's.

## Testing
Attaching results observed while testing these models.

[int8_trt_comparison.csv](https://github.com/user-attachments/files/24763088/int8_trt_comparison.csv)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated GPU runtime dependencies for the optional ONNX package.
Simplified platform-specific version constraints while maintaining
compatibility across supported platforms.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-01-22 02:08:18 +05:30
Chenjie Luo 615f99e746 Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support KIMI K2 Thinking PTQ from the original int4 checkpoint.
Tested with transformers  4.57.1, compressed-tensors 0.12.0

The model weights are dequantized on the fly to save GPU memory

## Usage
scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant
nvfp4_mlp_only --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Support for nvfp4_mlp_only quantization format, enabling new
layer-wise quantization options
  * Quantization support for CompressedLinear layers in quantized models

* **Improvements**
* Enhanced quantization for DeepSeek models with improved attention
configuration handling
* Optimized model loading with automatic precision configuration and
weight unpacking
* Better memory management during model export with automatic cache
cleanup
  * Conditional sample generation output controlled via verbose mode

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-21 09:32:33 +00:00
21a4010348 Add Quantizers for Qwen3VLMoeTextDecoderLayer (#666)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ? huggingface transformers library implements Qwen3VL Moe
layer as a monolithic module, instead of assembling it using Linear
layers, which cannot be recognized by modelopt's quantizer now. This PR
introduces a conversion from hf's qwen3vl_moe MoE layers to qewn3_moe
MoE layers which consist of a set of Linear layers.

## Testing
Tested with
```python
python hf_ptq.py --pyt_ckpt_path=Qwen/Qwen3-VL-30B-A3B-Instruct --qformat=nvfp4 --dataset wikipedia
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added quantization support for Qwen3VL models with sparse
mixture-of-experts (MoE) architecture, enabling efficient model
compression for this model type.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Qidong Su <qidongs@nvidia.com>
Signed-off-by: Qidong Su <soodoshll@gmail.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu <bestczy317@gmail.com>
2026-01-20 15:05:14 -08:00
Chenjie Luo b0e7d9fd96 Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do?

**Overview:** ?

Unified the FP8 and NVFP4 kv cache scaling factor definition so the same
checkpoint can be used for both FP8 and NVFP4 kv cache quantization
deployment

## Testing
Unit test

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization,
simplifying configuration logic.

* **Chores**
  * Removed internal constants from public exports.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-20 08:34:32 +00:00
Gwena Cunha 1cc8e6bf39 [5676209] Fix duplicated calib data (#794)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** This PR fixes an issue with calibration data with multiple
samples. Previously, calibration data with multiple samples was
generating a data loader with the same sample copied X times instead of
generating data with X different samples.

## Usage
```python
$ python -m modelopt.onnx --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```

## Testing
Use calibration data from 5676209 and observe the output of
`calibration_data_reader` in `quantize.py`:
```python
calibration_data_reader = CalibrationDataProvider(
    onnx_path, calibration_data, calibration_shapes
)
```

Each calibration sample in the list should be different.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed calibration data generation in ONNX workflow to properly handle
multiple samples during processing.

* **Documentation**
* Updated changelog with version 0.42 entry documenting bug fixes and
new features.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-19 23:39:49 +00:00
Gal Hubara-Agam 391f6cb004 [5750013][5591945][5360813]: AutoCast standalone implementation for type inference (#719)
## What does this PR do?

**Type of change:** New feature

**Overview:** 
AutoCast runs full type inference to get the new types after adding
casts. ONNX doesn't have a separate function for type inference, and it
is done as part of shape inference. Shape inference is a much more
complex task than type inference, especially when dynamic shapes are
involved. We're seeing some shape inference related bugs in AutoCast.
Typically we can WAR, but it's cumbersome. A local implementation might
allow users to WAR shape inference related issues. This is opt-in and
marked as experimental.

## Usage
python -m modelopt.onnx.autocast --onnx_path /path/to/input.onnx
[options] --use_standalone_type_inference

## Testing
Added use_standalone_type_inference=True to all existing
PrecisionConverter tests.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
A more permanent fix would be to decouple type and shape inference in
ONNX, we should invest in that when we have the resources - see
https://github.com/onnx/onnx/issues/7100
. This is a quick fix, which is also why it is opt-in and not the
default mode.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `--use_standalone_type_inference` flag to ONNX AutoCast,
enabling type-only inference as an alternative to standard shape
inference. Useful as a workaround when shape inference fails or to
reduce computational overhead.

* **Documentation**
* Added "Type Inference Control" section with usage examples and caveats
for the new standalone type inference option.

* **Tests**
* Extended test coverage to validate both standard and standalone type
inference paths.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
2026-01-18 09:14:54 +00:00