Commit Graph
511 Commits
Author SHA1 Message Date
Chenhan D. Yu 2a08622059 Fix moe amax remedy for dsr1 and remove global barrier in quantization megatron plugins (#808)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug

**Overview:** ?

This PR fix 2 bugs which impact DeepSeek calibration as well as PP
forward of MoE models.

1. The WAR in `MoELayer` that change the `topk` to `num_experts` only
works if no group-topk (a.k.a group routing) is used. Only changing topk
will lead to out-of-range error since `topk` can never be `num_experts`
when `group_topk != None`. Currently only `DeepSeek-V3` uses
`group_topk` and DeepSeek-V3 does not have difficulty to calibrate all
experts. As a result, we disable the WAR when detecting `group_topk`.

2. A previous PR inserted global barrier in
`quantization.plugin.megatron`
https://github.com/NVIDIA/Model-Optimizer/commit/6ef9954db1e73b8c4a86e5bfd31c954cfa21db61#diff-0fa2ba4ecc36c5ff031be9f9a5af080e7aa3afa331c438f02f501b9432ec6d6aL228-R515
This leads to dead lock when using PP since PP rank will never be able
to sync during pipeline forward. For MoE, this can be even worse if the
barrier is only visited by some EP/PP rank. Using collective
communication over the global world (a.k.a global comm) in megatron
plugin should be prohibited. Using collective on sub communication group
should avoid using `megatron.core.parallel_state` (a.k.a `mpu`) in the
future. Instead, use the local `pg_collection` from each module. Any
usage of collective communication must be inspected carefully with test
as PP, TP, and EP.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-23 02:11:19 +00:00
sychen52 b44c60ad34 Svdquant huggingface checkpoint export support (#754)
## What does this PR do?

**Type of change:** new feature

**Overview:** 

## Usage

```bash
cd ./examples/llm_ptq/
python hf_ptq.py \
    --pyt_ckpt_path Qwen/Qwen3-4B \
    --export_path /home/scratch.shiychen_coreai/quantized_models/Qwen3-4B-svdq \
    --qformat nvfp4_awq_svdquant --kv_cache_qformat none --sparsity_fmt dense --calib_size 8
```

## Testing
exported checkpoint and loaded.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added nvfp4_svdquant as a new quantization format option for LLM model
quantization workflows.

* **Limitations**
* Multi-GPU export configurations using tensor or pipeline parallelism
are not supported with nvfp4_svdquant quantization.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-01-22 11:19:33 -08:00
jingyu-ml 945ee02f8d [1/3] Add the fastvideo support (#804)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

FastVideo is a new diffusion-focused framework that we plan to integrate
with. In this work, we added initial support for WAN 2.2 5B in
FastVideo, targeting the text-to-video use case.

For the Conv layer type, we currently use a straightforward direct
convolution call. Implicit GEMM quantization is intentionally omitted in
this first MR and will be addressed in a follow-up MR.

- [x] [1/3] Added support for the WAN 2.2 DIT + VAE layer type.
- [ ] [2/3] Added calibration support for them in the example script,
add test cases and README, doc.
- [ ] [3/3] Submitted an MR to fastvideo to enable quantization-aware
training.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added FastVideo plugin support to the quantization framework. Users
can now apply quantization to FastVideo-specific layers with specialized
weight quantization handling, optimized input processing, and caching
features for enhanced inference performance.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:17:59 +00:00
jingyu-ml 668b8a19e8 [1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

This PR adds support for exporting quantized diffusers models (DiT,
Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling
deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM.

**Changes**

New file: `diffusers_utils.py`
- Dummy input generation for various diffusion models
- Pipeline component extraction helpers
- QKV projection detection and grouping
- `hide_quantizers_from_state_dict()` context manager for clean saves

Refactored: `unified_export_hf.py`
- New `_fuse_qkv_linears_diffusion()` for QKV amax fusion
- `_export_diffusers_checkpoint()` to export full pipelines (models +
tokenizers + schedulers etc.)

Plans

- [x] [1/3] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml 
- [ ] [3/3] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml

## Usage
<!-- You can potentially add a usage example below. -->
```
mtq.quantize(pipe, quant_config, forward_call)
export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## New Features
* Added HuggingFace checkpoint export support for quantized diffusion
models with configurable output directory
* Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint
export destination
* Extended export functionality to support selective component exports
from diffusion pipelines
* Enhanced quantized model export with improved component handling and
multi-stage checkpoint generation

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-21 23:11:09 +00:00
Keval Morabia 563a1e09c6 Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?

**Type of change:** New feature 

- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
  - Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.

## Usage
<!-- You can potentially add a usage example below. -->

Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
    model,
    mode="mcore_minitron",
    constraints=constraints,
    dummy_input=None,  # Not used
    config=config,
)
```

1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```

2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}

# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
    return megatron_mmlu(m, tokenizer, percentage=0.05)  # 5% sampled data for faster eval
config["score_func"] = score_func

# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune

mtp.prune(...)
```

To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
    hidden_size_divisor=256,
    ffn_hidden_size_divisor=512,
    mamba_head_dim_divisor=8,
    num_moe_experts_divisor=8,
    num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```

## Testing

**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
     0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
     0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
     0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
     0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
     0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
     0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
     0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
     0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
     0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```

**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
     0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads':  88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
     0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
     0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads':  96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
     0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
     0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
     0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
     0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
     0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads':  80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
     0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```

## Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information

OMNIML-3043

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
  * Introduced parameter counting capabilities for architecture search

* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
  * Updated configuration documentation with granular divisor parameters

* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-21 22:34:39 +00:00
Keval Morabia 792806f00e Add Security considerations in docs (#803)
Add security considerations to docs suggested by Nvidia Security team

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* Added comprehensive security considerations documentation for ModelOpt
covering multiple risk areas including untrusted input handling,
deserialization safety, input validation, resource exhaustion
prevention, data protection for transit and storage, logging and
observability practices, supply chain security, and structured
mitigation approaches with practical implementation examples.


<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-21 12:52:44 -08:00
Hrishith Thadicherla 67576d20ba Revert onnxruntime-gpu version to 1.22.0 for Windows (#801)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** 
Reverted the windows ort version in setup.py to 1.22, since I later
observed some regressions with vision models.
Keeping ORT version as 1.23 in windows examples since i didn't face any
issues with LLM's.

## Testing
Attaching results observed while testing these models.

[int8_trt_comparison.csv](https://github.com/user-attachments/files/24763088/int8_trt_comparison.csv)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated GPU runtime dependencies for the optional ONNX package.
Simplified platform-specific version constraints while maintaining
compatibility across supported platforms.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-01-22 02:08:18 +05:30
Chenjie Luo 615f99e746 Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** 

Support KIMI K2 Thinking PTQ from the original int4 checkpoint.
Tested with transformers  4.57.1, compressed-tensors 0.12.0

The model weights are dequantized on the fly to save GPU memory

## Usage
scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant
nvfp4_mlp_only --trust_remote_code

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Support for nvfp4_mlp_only quantization format, enabling new
layer-wise quantization options
  * Quantization support for CompressedLinear layers in quantized models

* **Improvements**
* Enhanced quantization for DeepSeek models with improved attention
configuration handling
* Optimized model loading with automatic precision configuration and
weight unpacking
* Better memory management during model export with automatic cache
cleanup
  * Conditional sample generation output controlled via verbose mode

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-21 09:32:33 +00:00
21a4010348 Add Quantizers for Qwen3VLMoeTextDecoderLayer (#666)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ? huggingface transformers library implements Qwen3VL Moe
layer as a monolithic module, instead of assembling it using Linear
layers, which cannot be recognized by modelopt's quantizer now. This PR
introduces a conversion from hf's qwen3vl_moe MoE layers to qewn3_moe
MoE layers which consist of a set of Linear layers.

## Testing
Tested with
```python
python hf_ptq.py --pyt_ckpt_path=Qwen/Qwen3-VL-30B-A3B-Instruct --qformat=nvfp4 --dataset wikipedia
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added quantization support for Qwen3VL models with sparse
mixture-of-experts (MoE) architecture, enabling efficient model
compression for this model type.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Qidong Su <qidongs@nvidia.com>
Signed-off-by: Qidong Su <soodoshll@gmail.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu <bestczy317@gmail.com>
2026-01-20 15:05:14 -08:00
Chenjie Luo b0e7d9fd96 Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do?

**Overview:** ?

Unified the FP8 and NVFP4 kv cache scaling factor definition so the same
checkpoint can be used for both FP8 and NVFP4 kv cache quantization
deployment

## Testing
Unit test

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization,
simplifying configuration logic.

* **Chores**
  * Removed internal constants from public exports.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-20 08:34:32 +00:00
Gwena Cunha 1cc8e6bf39 [5676209] Fix duplicated calib data (#794)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** This PR fixes an issue with calibration data with multiple
samples. Previously, calibration data with multiple samples was
generating a data loader with the same sample copied X times instead of
generating data with X different samples.

## Usage
```python
$ python -m modelopt.onnx --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```

## Testing
Use calibration data from 5676209 and observe the output of
`calibration_data_reader` in `quantize.py`:
```python
calibration_data_reader = CalibrationDataProvider(
    onnx_path, calibration_data, calibration_shapes
)
```

Each calibration sample in the list should be different.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed calibration data generation in ONNX workflow to properly handle
multiple samples during processing.

* **Documentation**
* Updated changelog with version 0.42 entry documenting bug fixes and
new features.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-19 23:39:49 +00:00
Gal Hubara-Agam 391f6cb004 [5750013][5591945][5360813]: AutoCast standalone implementation for type inference (#719)
## What does this PR do?

**Type of change:** New feature

**Overview:** 
AutoCast runs full type inference to get the new types after adding
casts. ONNX doesn't have a separate function for type inference, and it
is done as part of shape inference. Shape inference is a much more
complex task than type inference, especially when dynamic shapes are
involved. We're seeing some shape inference related bugs in AutoCast.
Typically we can WAR, but it's cumbersome. A local implementation might
allow users to WAR shape inference related issues. This is opt-in and
marked as experimental.

## Usage
python -m modelopt.onnx.autocast --onnx_path /path/to/input.onnx
[options] --use_standalone_type_inference

## Testing
Added use_standalone_type_inference=True to all existing
PrecisionConverter tests.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
A more permanent fix would be to decouple type and shape inference in
ONNX, we should invest in that when we have the resources - see
https://github.com/onnx/onnx/issues/7100
. This is a quick fix, which is also why it is opt-in and not the
default mode.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `--use_standalone_type_inference` flag to ONNX AutoCast,
enabling type-only inference as an alternative to standard shape
inference. Useful as a workaround when shape inference fails or to
reduce computational overhead.

* **Documentation**
* Added "Type Inference Control" section with usage examples and caveats
for the new standalone type inference option.

* **Tests**
* Extended test coverage to validate both standard and standalone type
inference paths.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
2026-01-18 09:14:54 +00:00
Wei-Ming Chen 38fb12037d [NVBug 5702186] Fix awq model export for Gemma3 (#793)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** norms laers in Gemma that use (1 + weight) in forward, we
will fold pre_quant_scale into the effective weight. That is to find
folded w' subject to: `1 + w' = (1 + w) * s` => `w' = (1 + w) * s -1`

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

./scripts/huggingface_example.sh --model google/gemma-3-1b-it --quant
int4_awq

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Improvements**
* Enhanced quantization utilities to better handle various LayerNorm
variants and normalization patterns, including support for weight-offset
variants and zero-centered gamma configurations.
* Optimized pre-quantization layer normalization fusion to apply
conditional weight scaling strategies based on normalization type.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-01-17 20:48:32 -08:00
Gwena Cunha c1956b8e2b [5763424][ONNX][Autocast] Fix ConstantOfShape layer output precision (#789)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Fixed the output precision of ConstantOfShape layers in
models with custom ops.

## Usage
<!-- You can potentially add a usage example below. -->

```python
$ python -m modelopt.onnx.quantization --onnx_path=$MODEL_NAME.onnx
```

## Testing
See bug 5763424.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
This issue only affects models with custom ops.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved type propagation handling for ConstantOfShape operations in
ONNX autocast, ensuring correct precision type conversion across related
operations.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-16 14:02:36 -08:00
jingyu-ml e6e4efd61e [0.5/3] Diffusion ckpt export for NVFP4 & FP8 (#783)
See https://github.com/NVIDIA/Model-Optimizer/pull/781

This is the MR that only includes the refactoring of the llm export,
please ignore the change on quantize.py from the diffusion example.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added `--hf-ckpt-dir` CLI option to save checkpoints in HuggingFace
format
  * Enabled support for exporting Diffusers-based pipelines
* Unified export system now handles both transformer and diffusion model
architectures

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-01-15 13:49:22 -06:00
Chenhan D. Yu 849a3501c5 Change trust_remote_code default to False for security reason (#787)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug fix

**Overview:** ?

Change `trust_remote_code` default to `False` for security reason

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Updated model loader security settings: remote code is no longer
trusted by default when loading model configurations. Users requiring
remote code execution must now explicitly enable this option.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-15 18:59:51 +00:00
Zhiyu 0f05d676d3 Remove quantization_config in config.json from original deepseek models (#753)
## What does this PR do?

**Type of change:**  Bug fix

**Overview:** DeepSeek original checkpoints may include a
`quantization_config` field in `config.json`
(describing the source checkpoint's quantization). When we export
ModelOpt quantization
configs to `hf_quant_config.json`, leaving the original
`quantization_config` in place can
    be confusing. Add a function to remove it.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
Resolve nvbug https://nvbugspro.nvidia.com/bug/5736665

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-01-14 18:14:09 -08:00
Chenhan D. Yu 406c18ce3b chg: passing through trust_remote_code (#778)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug fix

**Overview:** 

Passing `trust_remote_code` all the way through during export and
import. This is needed since `DeepSeek` will error out if
`trust_remote_code=True` but `Nemotron-H` will error out if
`trust_remote_code=False`

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Updated default `trust_remote_code` parameter from `True` to `False`
in GPT model export and import functionality.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-14 14:20:16 -08:00
Wei-Ming Chen 6038451779 Fix Qwen3 recipe and update autoquant example cmd (#749)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-01-14 11:57:56 -08:00
Wei-Ming Chen db76b1edeb Fix AWQ export when quantization of some layers are disabled (#721)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Fix AWQ export when quantization of some layers are disabled

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-01-14 10:04:19 -08:00
Gwena Cunha 951c6aad5d [5763448][ONNX][Autocast] Fix Resize input type mismatch error (#757)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** This PR fixes an input type mismatch in Resize layers when
being converted to FP16.

## Usage

```python
$ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx
```

## Testing
Added unittest.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
This issue is also fixed by using the standalone type inference logic
from https://github.com/NVIDIA/Model-Optimizer/pull/719.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Improvements**
* Enhanced the graph sanitization process to automatically duplicate
shared constants during optimization, ensuring improved model handling
and consistency.

* **Tests**
* Added test coverage for mixed precision conversion of Conv-Resize
model architectures.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-14 09:26:33 +05:30
Frida Hou 18d9b1eea4 Add static per block MSE for NVFP4 weight (#613)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature

**Overview:** ?
Support static block-wise MSE for NVFP4 weight quantization.
Add a FP4 triton kernel that take in scales for each block. It also
quantizes the scales to FP8.

This PR does the following:
1. Enable static NVFP4 implementation, i.e. block scales for weights are
calculated during calibration and feed into fake quant kernels
2.Extend mse_calibrate to support static NVFP4 with block scales
searching by MSE and global scale set as MAX
3.Refinements: calibrate weight quantizers only once during MSE
calibration



## Usage
<!-- You can potentially add a usage example below. -->


Example config:
```python
NVFP4_WEIGHT_MSE_CFG = {
    "quant_cfg": {
        "*weight_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        "*input_quantizer": {
            "enable": False,
        },
        **_default_disabled_quantizer_cfg,
    },
    "algorithm": {
        "method": "mse",
        "step_size": 0.25,
        "start_multiplier": 0.25,
        "stop_multiplier": 2.0,
    },
}

NVFP4_WEIGHT_ACT_MSE_CFG = {
    "quant_cfg": {
        "*weight_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        "*input_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        **_default_disabled_quantizer_cfg,
    },
    "algorithm": {
        "method": "mse",
        "step_size": 0.25,
        "start_multiplier": 0.25,
        "stop_multiplier": 2.0,
    },
}

```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-01-13 14:43:57 -08:00
Ajinkya Rasane b4c77c0d9b [NVBUG 5801937] Disable dq_only by default (#777)
## What does this PR do?

**Type of change:** 
Bug fix

**Overview:**
Disable dq_only flag by default in modelopt onnx quantization

## Testing
Able to build and run model with modelopt onnx Python CLI

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No
- dq_only is set to False by default
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Chores**
* Updated quantization default behavior: Q/DQ (Quantize/Dequantize)
nodes are now added by default instead of only Dequantize nodes.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-01-13 22:20:06 +00:00
Keval Morabia 9de4877d27 [1/2] Address security concerns in code (#626)
- [x] Address feedback on Threat and Vuln Analysis (TAVA) doc by ProdSec
team
- [x] Add note on safe usage of pickle deserialization of
modelopt-generated state files

**TODO: [Separate PR]** Replace pickle usage in
`modelopt/torch/opt/plugins/megatron.py` - Needs fix on
TransformerEngine first as we copy from
https://github.com/NVIDIA/TransformerEngine/blob/3ff0b8d4/transformer_engine/pytorch/module/base.py#L863

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added `--trust_calibration_data` CLI flag for secure ONNX quantization
with pickle data files.

* **Improvements**
  * Enhanced security validation for generated quantization code.
* Simplified data loading by removing pickle-based caching—data is now
always loaded fresh.
  * Added security guidance throughout model state loading operations.

* **Documentation**
* Updated guides with security best practices for model state handling.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-13 20:30:07 +00:00
yeyu-nvidia 90fa48ce14 remove duplicated RMSNorm and use LlamaRMSNorm from transformers (#774)
## What does this PR do?
Code cleanup

**Overview:** 
Remove RMSNorm which is identical to LlamaRMSNorm from transformers.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Updated the normalization implementation in the Eagle speculative
module.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-01-13 11:26:52 -08:00
Gwena Cunha 5e0d36551a [5796745][ONNX][Autocast] Fix opset check for model with custom ops (#767)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** This PR fixes the opset being incorrectly detected as
being `1` in models with custom ops. That happens because the
'trt.plugins' domain version is detected rather than the actual model's
opset version.

## Usage
<!-- You can potentially add a usage example below. -->

```python
$ python -m modelopt.onnx.autocast --onnx_path=${MODEL_NAME}.onnx
```

## Testing
See bug 5796745.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Optimized the opset version detection logic for more efficient model
conversion handling.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-01-13 11:54:58 -05:00
Asha Anoosheh b813ab548b Top-K KL Divergence loss (#747)
## What does this PR do?

**Type of change:** New feature

**Overview:** Writes a new KLDiv Logits loss which only uses top-k vocab
values

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Top-K logit filtering capability for knowledge distillation
workflows, enabling selective focus on high-probability tokens.

* **Improvements**
* Enhanced distributed tensor model-parallel operations with improved
awareness for gradient computation and reduction.
  * Simplified legacy distributed operation constructs.

* **Tests**
* Introduced comprehensive test coverage for Megatron-based
distillation, validating both standard and Top-K filtering variants.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-01-13 16:09:31 +01:00
Chenhan D. Yu 7836065f3e Set trust_remote_code default to False (#769)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

DeepSeek later has official `transformers` support and use
`trust_remote_code=True` will encounter error. Set default to `False`.

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-13 05:54:47 +00:00
Wei-Ming Chen 6ae96b5162 [NVBug 5784940] Fix autodeploy example (#764)
## What does this PR do?

**Type of change:** bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:**  update the example with new API


## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
python examples/llm_autodeploy/api_server.py --ckpt_path
TinyLlama/TinyLlama-1.1B-Chat-v1.0 --world_size 1

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* AutoDeployConfig parameter handling has been updated to improve how
parameters are prepared for language model initialization. The
configuration method now uses optimized keyword argument formatting to
ensure consistency and clarity across the deployment system.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-01-12 15:36:15 -08:00
Izzy Putterman 727da95a91 SpecDec Bench: PostProcess flag (#759)
## What does this PR do?

**Type of change:** ? Bug Fix: https://nvbugspro.nvidia.com/bug/5795144

**Overview:** ? Pass postprocess flag to handle slicing message.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Introduced --postprocess command-line option to select postprocessing
strategy. Users can choose "base" (default, preserves existing behavior)
or "gptoss" (new alternative method) with validation to reject invalid
selections.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Izzy Putterman <iputterman@nvidia.com>
2026-01-13 02:24:41 +05:30
Keval Morabia b484efb84e [CI] Cleanup ubuntu-runner disk storage before installing deps (#765)
We started seeing this issue in GitHub's free ubuntu-latest runners:
`ERROR: Could not install packages due to an OSError: [Errno 28] No
space left on device`.

Suggested by other GH runners users to remove unnecessary android /
dotnet files to avoid the storage issue.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Consolidated GitHub Actions workflow setup into a reusable custom
action for improved maintainability and consistency across CI/CD
pipelines.
* Enhanced release workflow with automated unit testing and artifact
upload capabilities.
* Streamlined runner initialization by reducing redundant configuration
steps.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-13 02:18:41 +05:30
Asha Anoosheh 510451322c Streamline KD & QAD transformers Trainers (#708)
## What does this PR do?

**Type of change:** ? Refactor and stabilization

**Overview:** 
* Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so
that we can remove multiple restrictions

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-01-10 01:28:56 +00:00
Keval Morabia 5b9261f084 Add .coderabbit.yaml for auto PR reviews (#756)
Automatically review PRs by Coderabbit

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Added configuration settings for code review automation.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-09 23:37:24 +05:30
Gal Hubara-Agam 7971fff058 [5694695][AutoCast] Preserve outer scope variable types in subgraphs (#717)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** 
When clearing type information for shape inference, preserve value_info
for outer scope variables in subgraphs. Previously, all value_info
entries were cleared indiscriminately, causing shape inference failures
when subgraph nodes referenced outer scope variables.


## Testing
pytest
tests/unit/onnx/autocast/test_precisionconverter.py::test_if_subgraph_outer_scope_type_preservation

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: N/A
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
2026-01-09 16:14:37 +02:00
sychen52 ecda7b0bfa Use kitchen FA in huggingface plugin (#674)
## What does this PR do?  new feature

**Overview:** use kitchen FA in huggingface plugin

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-01-08 21:23:18 -08:00
Chenhan D. Yu 307fe7183b Fix QuantSequentialMLP sharded_state_dict (#742)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug

**Overview:** ?

These fixes are needed for Megatron-LM `main` branch due to some changes
in `sharded_state_dict`. Qwen3-30B-A3B PTQ and resume fails while EP=4
cannot load a checkpoint generated with PP=4.

`singleton_local_shards` must be added to the metadata; otherwise, all
experts `amax` are packed to gather and currently the TP `replica_id`
for `linear_fc1` is incorrect.

**Other Finding:** This limits TP=ETP=1 when EP>1. Otherwise, there will
be `sharded_state_dict` access error. There is a potential blind spot of
using the default TP group in `ColumnParallelLinear` and
`RowParallelLinear` since it can be part of the MoE where the tensor
parallelism is controlled by ETP instead. Will need a different PR to
fix the parallel_state.

**Results:** If calibrate with EP=1, mmlu = 0.80. This can be resumed
with EP=4, TP=1, ETP=1 (TP>1 does not work as mentioned above). However
if calibrated with EP=4, then mmlu = 0.71 which shows there are some
issues with max sync in EP.


## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-01-08 16:06:44 -08:00
Wei-Ming Chen 6f18490b83 Improve AWQ init speed (#748)
## What does this PR do?

**Type of change:** ?Improvement<!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

Improve speed of accessing weight through
enable_weight_access_and_writeback in AWQ helper init. This change
reduces the time complexity from O(num_modules^2) to O(num_modules) and
the runtime from ~1hour to 30 seconds.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
python hf_ptq.py --pyt_ckpt_path
/home/scratch.omniml_data_1/models/qwen/Qwen3-30B-A3B-Instruct-2507
--qformat int4_awq

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-01-08 12:08:17 -08:00
Chenjie Luo 9c24e2c08e Fix Deepseek transformers model loading (#740)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?
For Deepseek, let's force the user to apply trust_remote_code and use
AutoModelForCausalLM for loading the model.

## Testing
python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4
--export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64
--trust_remote_code --dataset cnn_dailymail

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-08 11:12:53 -08:00
Keval Morabia 68d604dd69 Move new puzzle dist utils from feature/compress to main (#746)
- Move new `modelopt.torch.utils.distributed` from `feature/compress` to
`main` branch so they can be used via modelopt in puzzletron gitlab

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-08 19:45:29 +05:30
Keval Morabia 9a3b986f9c Fix TRT-LLM 2-gpu CI test shm issue (#744)
- As suggested by NVGHA runners team to increase SHM size to avoid issue
on 2-gpu nightly tests

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-01-08 15:31:48 +05:30
realAsmaandAsma Thekkumpate 81c509c643 Fixes & Simplifications for MCore KVCache QAT/QAD; Unittests; Distributed Sync of KVCache Quantizer params (#727)
## What does this PR do?

**Type of change:** Fix MCore KV Cache Quantization: Amax Device
Placement Bug; Code clean up; Distributed Sync of KVCache Quantizer
params; unittest expansion to hybrid models

**Overview:** Fixes bugs preventing MCore KV Cache quantization from
working during checkpoint restore.

### Bug Chain

**Bug 1:** `is_enabled = self.weight_quantizer.is_enabled if
hasattr(self, "weight_quantizer") else False`

No `weight_quantizer` for KV-cache-only quant → `is_enabled=False` →
metadata not saved → `modelopt_post_restore()` never called. *(Thanks to
@jenchen13 )*

**Bug 2:** After fixing Bug 1, `_amax` restored on CPU (via
`_reset_pytorch_state_from_metadata`). Fallback
`_calibrate_quantizers()` never called because `_amax` exists.

**Bug 3:** Even if called, `_calibrate_quantizers()` fails —
`core_attention` has no parameters → can't determine device/dtype.

### The Fix

1. Remove `is_enabled` check entirely — disabled modules may still need
metadata restore. Explicitly skip `output_layer` from extra state
callbacks (never quantized)
2. Set `dtype`/`device` on `core_attention` from parent Attention
module, `modelopt_post_restore()` calls `self.to(device, dtype)`
3. Remove dead `_calibrate_quantizers()` code (will bring back similar
logic for KV cache affine quantization)

### Previous Unit Test Was Wrong

`model_test` was `mtq.quantize()`'d, not `mto.restore()`'d. Never tested
actual restore path.

### Additional Fixes

- Amax sync across DP/TP for KV cache quantizers
- `flash_decode` auto-disabled

### Code Cleanup

Removed ~100 lines of dead code.

## Testing

1. MCore KV Cache QAD with Nano V3 + Context Parallel works
2. Unit tests: hybrid models, KV+GEMM configs, correct restore workflow,
backward pass validation

## Before your PR is "*Ready for review*"

- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update Changelog?**: Yes

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Asma Thekkumpate <akuriparambi@cw-dfw-cs-001-vscode-02.cm.cluster>
2026-01-06 16:08:01 -08:00
kinjalpatel27 fe52b2a46e Bias running average computation in float (#738)
## What does this PR do?

**Type of change:**  Bug fix

**Overview:** ?
Computing Bias running average with bf16 creates incorrect estimations.
Impact on accuracy for Qwen2.5-7B model:

With BF16 running average:

NVFP4_AFFINE_KV | 59.11%
-- | --

With running average in Float:
NVFP4_AFFINE_KV | 71.81%
-- | --


## Usage
Use examples/lm_eval/mmlu.py with batchsize of 1
Note: the issue is masked with larger batch sizes

## Testing
- Ran mmlu benchmark with mmlu.py and nv-eval
- also ploted bf16 and float running average for different layers, one
of the example for layer 0 in Qwen2.5-7B:
 
<img width="2100" height="600" alt="image"
src="https://github.com/user-attachments/assets/715059c5-34a4-495e-b6f1-0b57cf0c08af"
/>

Note: for the larger value bf16 shows smaller value compared to float

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-01-05 14:35:34 -08:00
kinjalpatel27 4eb1835df5 Support for KV cache quantization for MLA Attention vLLM fakequant (#714)
## What does this PR do?

**Type of change:**  Feature extention

**Overview:** 
Added support to quantize KV cache in vLLM fakequant by adding
quantization support for
[MLAAttention](https://github.com/vllm-project/vllm/blob/v0.11.1/vllm/attention/layer.py#L641)

## Usage
Please refer to
[Readme](https://github.com/NVIDIA/Model-Optimizer/tree/kinjal/vllm_att_quant/examples/vllm_serve#calibrate-and-serve-fake-quant-model-in-vllm)

```shell
KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CFG=NVFP4_DEFAULT_CFG python vllm_serve_fakequant.py deepseek-ai/DeepSeek-V2 --served-model-name deepseek-ai/DeepSeek-V2 --host 0.0.0.0 --port 8001 --trust-remote-code --enforce-eager --gpu-memory-utilization 0.8  
```

## Testing
Locally tested KV Cache quantization
```
(rotary_emb): DeepseekScalingRotaryEmbedding()
(mla_attn): MultiHeadLatentAttentionWrapper(
  (fused_qkv_a_proj): QuantMergedColumnParallelLinear(
    in_features=5120, output_features=2112, bias=False, tp_size=1, gather_output=False
    (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=141.0000 calibrator=MaxCalibrator quant)
    (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.4297 calibrator=MaxCalibrator quant)
    (output_quantizer): TensorQuantizer(disabled)
  )
  (q_a_layernorm): RMSNorm(hidden_size=1536, eps=1e-06)
  (q_b_proj): QuantColumnParallelLinear(
    in_features=1536, output_features=3072, bias=False, tp_size=8, gather_output=False
    (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=32.0000 calibrator=MaxCalibrator quant)
    (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.1670 calibrator=MaxCalibrator quant)
    (output_quantizer): TensorQuantizer(disabled)
  )
  (kv_a_layernorm): RMSNorm(hidden_size=512, eps=1e-06)
  (kv_b_proj): QuantColumnParallelLinear(
    in_features=512, output_features=4096, bias=False, tp_size=8, gather_output=False
    (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant)
    (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.2773 calibrator=MaxCalibrator quant)
    (output_quantizer): TensorQuantizer(disabled)
  )
  (rotary_emb): DeepseekScalingRotaryEmbedding()
  (o_proj): QuantRowParallelLinear(
    in_features=2048, output_features=5120, bias=False, tp_size=8, reduce_results=True
    (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.7188 calibrator=MaxCalibrator quant)
    (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.4336 calibrator=MaxCalibrator quant)
    (output_quantizer): TensorQuantizer(disabled)
  )
  (mla_attn): QuantMLAAttention(
    (q_bmm_quantizer): TensorQuantizer(disabled)
    (kv_c_bmm_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant)
  )
)
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**:NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-01-05 13:03:24 -08:00
Keval Morabia 8426c363bd Skip unit tests in release workflow avoid storage issues in runner
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
0.42.0dev 0.41.0rc1
2026-01-05 19:18:19 +05:30
Chenjie Luo d541324e84 Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do?

**Type of change:** ? Recipe improvement

**Overview:** ?

Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3
Next recipe for accuracy recovery

## Testing
Model accuracy benchmarking

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-01-02 11:06:25 -08:00
Wei-Ming Chen b655321d87 [Issue 543] [Bug fix] Fix dynamic input quant for AWQ (#726)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Dynamic input quantizers, e.g., MXFP4, are not restored
after AWQ. This PR fix the issue.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested with MXFP4, NVFP4, int4

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2025-12-30 12:06:40 -08:00
noeyy-mino 883c8731aa Noeyy/add new cases for newly added checkpoints on HF (#728)
## What does this PR do?

**Type of change:** Add TRT LLM/vLLM/SGLang functional test cases for
newly added checkpoints on HF

**Overview:** 
1.Since the speculative draft model only supports loading from a local
path, we should set the MODELOPT_LOCAL_MODEL_ROOT environment variable.
If we don't set it, these test cases will be skipped.
2. Newly added checkpoints:

- nvidia/gpt-oss-120b-Eagle3-short-context
- nvidia/gpt-oss-120b-Eagle3-throughput
- nvidia/EAGLE3-NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8

## Usage


```python
pytest tests/examples/llm_ptq/test_deploy.py --run-release
```

## Testing
Run release testing

## Before your PR is "*Ready for review*"
Ready for review

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information
N/A

---------

Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com>
2025-12-30 10:52:11 +05:30
Wei-Ming Chen 3350b0a45b [OMNIML-3017] MLM QAD example (#682)
## What does this PR do?

**Type of change:** New example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

### Add QAD Training example for Megatron-LM

- File Structure
  - qad.sh / sbatch_qad.sh - Training and SLURM submission scripts
  - data_utils/ - Dataset download and preprocessing utilities
- configs/ - Configuration templates for Qwen3-30B-A3B (MoE) and
Qwen3-8B (Dense)

- Key Features
  - One-button dataset generation (OpenScience + Nemotron-v2)
- Config-based training scripts, keep all tunable knobs into a single
config file

## Usage
<!-- You can potentially add a usage example below. -->
1. Generate dataset
```bash
bash data_utils/generate_dataset.sh \
    --output-dir /path/to/datasets \
    --mlm-path /path/to/Megatron-LM \
    --tokenizer Qwen/Qwen3-30B-A3B-Instruct-2507
```
2. Create a config based on templates
3. Kick off training with Slurm:
```bash
sbatch sbatch_qad.sh --config configs/my-experiment.conf
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
QAD with Qwen3-30B-A3B-instruct-2507 NVFP4 (all layers quantized)
- GPQA:
BF16: 0.549
NVFP4 (PTQ): 0.4949
NVFP4 (QAD): 0.5202

- Livecodebench:
BF16: 0.3987
NVFP4 (PTQ): 0.37
NVFP4 (QAD): 0.3855

- Scicode:
BF16: 0.325
NVFP4 (PTQ): 0.276
NVFP4 (QAD): 0.3146

- AIME
BF16: 0.6049 
NVFP4 (PTQ): 0.55
NVFP4 (QAD): 0.5431
 

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Wei-Ming Chen <weimingc@login-eos01.eos.clusters.nvidia.com>
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2025-12-22 11:22:31 -08:00
Hrishith Thadicherla 03dc3860af Update onnxruntime-gpu (#697)
## What does this PR do?

**Type of change:** Bug fix 

**Overview:** Updated setup.py to use only onnxruntime-gpu and removed
onnxruntime-directml as dependency.
Also changed onnxruntime-gpu version in examples.


## Testing
Tested int4 quantization and MMLU benchmark with updated onnxruntime-gpu
, working as expected

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2025-12-22 08:35:58 +00:00
realAsma cb343352ca Registry interface for custom quantization functional backend (#683)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Add registry interface for custom quantization functional backend

## Usage
<!-- You can potentially add a usage example below. -->

see `tests/unit/torch/quantization/test_custom_backend.py` for usage
example.

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2025-12-20 04:17:32 +00:00