mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
a4fde491cc6c3c7c747c7c946416ae09d2c6fa5d
511
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2a08622059 |
Fix moe amax remedy for dsr1 and remove global barrier in quantization megatron plugins (#808)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug **Overview:** ? This PR fix 2 bugs which impact DeepSeek calibration as well as PP forward of MoE models. 1. The WAR in `MoELayer` that change the `topk` to `num_experts` only works if no group-topk (a.k.a group routing) is used. Only changing topk will lead to out-of-range error since `topk` can never be `num_experts` when `group_topk != None`. Currently only `DeepSeek-V3` uses `group_topk` and DeepSeek-V3 does not have difficulty to calibrate all experts. As a result, we disable the WAR when detecting `group_topk`. 2. A previous PR inserted global barrier in `quantization.plugin.megatron` https://github.com/NVIDIA/Model-Optimizer/commit/6ef9954db1e73b8c4a86e5bfd31c954cfa21db61#diff-0fa2ba4ecc36c5ff031be9f9a5af080e7aa3afa331c438f02f501b9432ec6d6aL228-R515 This leads to dead lock when using PP since PP rank will never be able to sync during pipeline forward. For MoE, this can be even worse if the barrier is only visited by some EP/PP rank. Using collective communication over the global world (a.k.a global comm) in megatron plugin should be prohibited. Using collective on sub communication group should avoid using `megatron.core.parallel_state` (a.k.a `mpu`) in the future. Instead, use the local `pg_collection` from each module. Any usage of collective communication must be inspected carefully with test as PP, TP, and EP. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
b44c60ad34 |
Svdquant huggingface checkpoint export support (#754)
## What does this PR do?
**Type of change:** new feature
**Overview:**
## Usage
```bash
cd ./examples/llm_ptq/
python hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--export_path /home/scratch.shiychen_coreai/quantized_models/Qwen3-4B-svdq \
--qformat nvfp4_awq_svdquant --kv_cache_qformat none --sparsity_fmt dense --calib_size 8
```
## Testing
exported checkpoint and loaded.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added nvfp4_svdquant as a new quantization format option for LLM model
quantization workflows.
* **Limitations**
* Multi-GPU export configurations using tensor or pipeline parallelism
are not supported with nvfp4_svdquant quantization.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
945ee02f8d |
[1/3] Add the fastvideo support (#804)
## What does this PR do? **Type of change:** new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** FastVideo is a new diffusion-focused framework that we plan to integrate with. In this work, we added initial support for WAN 2.2 5B in FastVideo, targeting the text-to-video use case. For the Conv layer type, we currently use a straightforward direct convolution call. Implicit GEMM quantization is intentionally omitted in this first MR and will be addressed in a follow-up MR. - [x] [1/3] Added support for the WAN 2.2 DIT + VAE layer type. - [ ] [2/3] Added calibration support for them in the example script, add test cases and README, doc. - [ ] [3/3] Submitted an MR to fastvideo to enable quantization-aware training. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added FastVideo plugin support to the quantization framework. Users can now apply quantization to FastVideo-specific layers with specialized weight quantization handling, optimized input processing, and caching features for enhanced inference performance. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
668b8a19e8 |
[1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This PR adds support for exporting quantized diffusers models (DiT, Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM. **Changes** New file: `diffusers_utils.py` - Dummy input generation for various diffusion models - Pipeline component extraction helpers - QKV projection detection and grouping - `hide_quantizers_from_state_dict()` context manager for clean saves Refactored: `unified_export_hf.py` - New `_fuse_qkv_linears_diffusion()` for QKV amax fusion - `_export_diffusers_checkpoint()` to export full pipelines (models + tokenizers + schedulers etc.) Plans - [x] [1/3] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml - [ ] [3/3] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ``` mtq.quantize(pipe, quant_config, forward_call) export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir) ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added HuggingFace checkpoint export support for quantized diffusion models with configurable output directory * Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint export destination * Extended export functionality to support selective component exports from diffusion pipelines * Enhanced quantized model export with improved component handling and multi-stage checkpoint generation <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
563a1e09c6 |
Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?
**Type of change:** New feature
- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
- Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.
## Usage
<!-- You can potentially add a usage example below. -->
Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
model,
mode="mcore_minitron",
constraints=constraints,
dummy_input=None, # Not used
config=config,
)
```
1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```
2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
return megatron_mmlu(m, tokenizer, percentage=0.05) # 5% sampled data for faster eval
config["score_func"] = score_func
# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune
mtp.prune(...)
```
To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
hidden_size_divisor=256,
ffn_hidden_size_divisor=512,
mamba_head_dim_divisor=8,
num_moe_experts_divisor=8,
num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```
## Testing
**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```
**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads': 88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads': 96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads': 80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```
## Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
OMNIML-3043
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
* Introduced parameter counting capabilities for architecture search
* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
* Updated configuration documentation with granular divisor parameters
* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
792806f00e |
Add Security considerations in docs (#803)
Add security considerations to docs suggested by Nvidia Security team <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * Added comprehensive security considerations documentation for ModelOpt covering multiple risk areas including untrusted input handling, deserialization safety, input validation, resource exhaustion prevention, data protection for transit and storage, logging and observability practices, supply chain security, and structured mitigation approaches with practical implementation examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
67576d20ba |
Revert onnxruntime-gpu version to 1.22.0 for Windows (#801)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** Reverted the windows ort version in setup.py to 1.22, since I later observed some regressions with vision models. Keeping ORT version as 1.23 in windows examples since i didn't face any issues with LLM's. ## Testing Attaching results observed while testing these models. [int8_trt_comparison.csv](https://github.com/user-attachments/files/24763088/int8_trt_comparison.csv) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated GPU runtime dependencies for the optional ONNX package. Simplified platform-specific version constraints while maintaining compatibility across supported platforms. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
615f99e746 |
Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support KIMI K2 Thinking PTQ from the original int4 checkpoint. Tested with transformers 4.57.1, compressed-tensors 0.12.0 The model weights are dequantized on the fly to save GPU memory ## Usage scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant nvfp4_mlp_only --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support for nvfp4_mlp_only quantization format, enabling new layer-wise quantization options * Quantization support for CompressedLinear layers in quantized models * **Improvements** * Enhanced quantization for DeepSeek models with improved attention configuration handling * Optimized model loading with automatic precision configuration and weight unpacking * Better memory management during model export with automatic cache cleanup * Conditional sample generation output controlled via verbose mode <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
21a4010348 |
Add Quantizers for Qwen3VLMoeTextDecoderLayer (#666)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? huggingface transformers library implements Qwen3VL Moe layer as a monolithic module, instead of assembling it using Linear layers, which cannot be recognized by modelopt's quantizer now. This PR introduces a conversion from hf's qwen3vl_moe MoE layers to qewn3_moe MoE layers which consist of a set of Linear layers. ## Testing Tested with ```python python hf_ptq.py --pyt_ckpt_path=Qwen/Qwen3-VL-30B-A3B-Instruct --qformat=nvfp4 --dataset wikipedia ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added quantization support for Qwen3VL models with sparse mixture-of-experts (MoE) architecture, enabling efficient model compression for this model type. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Qidong Su <qidongs@nvidia.com> Signed-off-by: Qidong Su <soodoshll@gmail.com> Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> Co-authored-by: Zhiyu <bestczy317@gmail.com> |
||
|
|
b0e7d9fd96 |
Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do? **Overview:** ? Unified the FP8 and NVFP4 kv cache scaling factor definition so the same checkpoint can be used for both FP8 and NVFP4 kv cache quantization deployment ## Testing Unit test ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization, simplifying configuration logic. * **Chores** * Removed internal constants from public exports. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
1cc8e6bf39 |
[5676209] Fix duplicated calib data (#794)
## What does this PR do?
**Type of change:** Bug fix
**Overview:** This PR fixes an issue with calibration data with multiple
samples. Previously, calibration data with multiple samples was
generating a data loader with the same sample copied X times instead of
generating data with X different samples.
## Usage
```python
$ python -m modelopt.onnx --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```
## Testing
Use calibration data from 5676209 and observe the output of
`calibration_data_reader` in `quantize.py`:
```python
calibration_data_reader = CalibrationDataProvider(
onnx_path, calibration_data, calibration_shapes
)
```
Each calibration sample in the list should be different.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed calibration data generation in ONNX workflow to properly handle
multiple samples during processing.
* **Documentation**
* Updated changelog with version 0.42 entry documenting bug fixes and
new features.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
|
||
|
|
391f6cb004 |
[5750013][5591945][5360813]: AutoCast standalone implementation for type inference (#719)
## What does this PR do? **Type of change:** New feature **Overview:** AutoCast runs full type inference to get the new types after adding casts. ONNX doesn't have a separate function for type inference, and it is done as part of shape inference. Shape inference is a much more complex task than type inference, especially when dynamic shapes are involved. We're seeing some shape inference related bugs in AutoCast. Typically we can WAR, but it's cumbersome. A local implementation might allow users to WAR shape inference related issues. This is opt-in and marked as experimental. ## Usage python -m modelopt.onnx.autocast --onnx_path /path/to/input.onnx [options] --use_standalone_type_inference ## Testing Added use_standalone_type_inference=True to all existing PrecisionConverter tests. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information A more permanent fix would be to decouple type and shape inference in ONNX, we should invest in that when we have the resources - see https://github.com/onnx/onnx/issues/7100 . This is a quick fix, which is also why it is opt-in and not the default mode. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--use_standalone_type_inference` flag to ONNX AutoCast, enabling type-only inference as an alternative to standard shape inference. Useful as a workaround when shape inference fails or to reduce computational overhead. * **Documentation** * Added "Type Inference Control" section with usage examples and caveats for the new standalone type inference option. * **Tests** * Extended test coverage to validate both standard and standalone type inference paths. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> |
||
|
|
38fb12037d |
[NVBug 5702186] Fix awq model export for Gemma3 (#793)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** norms laers in Gemma that use (1 + weight) in forward, we will fold pre_quant_scale into the effective weight. That is to find folded w' subject to: `1 + w' = (1 + w) * s` => `w' = (1 + w) * s -1` ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ./scripts/huggingface_example.sh --model google/gemma-3-1b-it --quant int4_awq ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Enhanced quantization utilities to better handle various LayerNorm variants and normalization patterns, including support for weight-offset variants and zero-centered gamma configurations. * Optimized pre-quantization layer normalization fusion to apply conditional weight scaling strategies based on normalization type. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
c1956b8e2b |
[5763424][ONNX][Autocast] Fix ConstantOfShape layer output precision (#789)
## What does this PR do? **Type of change:** Bug fix **Overview:** Fixed the output precision of ConstantOfShape layers in models with custom ops. ## Usage <!-- You can potentially add a usage example below. --> ```python $ python -m modelopt.onnx.quantization --onnx_path=$MODEL_NAME.onnx ``` ## Testing See bug 5763424. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information This issue only affects models with custom ops. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved type propagation handling for ConstantOfShape operations in ONNX autocast, ensuring correct precision type conversion across related operations. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
e6e4efd61e |
[0.5/3] Diffusion ckpt export for NVFP4 & FP8 (#783)
See https://github.com/NVIDIA/Model-Optimizer/pull/781 This is the MR that only includes the refactoring of the llm export, please ignore the change on quantize.py from the diffusion example. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added `--hf-ckpt-dir` CLI option to save checkpoints in HuggingFace format * Enabled support for exporting Diffusers-based pipelines * Unified export system now handles both transformer and diffusion model architectures <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
849a3501c5 |
Change trust_remote_code default to False for security reason (#787)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? Change `trust_remote_code` default to `False` for security reason ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Updated model loader security settings: remote code is no longer trusted by default when loading model configurations. Users requiring remote code execution must now explicitly enable this option. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
0f05d676d3 |
Remove quantization_config in config.json from original deepseek models (#753)
## What does this PR do?
**Type of change:** Bug fix
**Overview:** DeepSeek original checkpoints may include a
`quantization_config` field in `config.json`
(describing the source checkpoint's quantization). When we export
ModelOpt quantization
configs to `hf_quant_config.json`, leaving the original
`quantization_config` in place can
be confusing. Add a function to remove it.
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
Resolve nvbug https://nvbugspro.nvidia.com/bug/5736665
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
||
|
|
406c18ce3b |
chg: passing through trust_remote_code (#778)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** Passing `trust_remote_code` all the way through during export and import. This is needed since `DeepSeek` will error out if `trust_remote_code=True` but `Nemotron-H` will error out if `trust_remote_code=False` ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Updated default `trust_remote_code` parameter from `True` to `False` in GPT model export and import functionality. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
6038451779 |
Fix Qwen3 recipe and update autoquant example cmd (#749)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
db76b1edeb |
Fix AWQ export when quantization of some layers are disabled (#721)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Fix AWQ export when quantization of some layers are disabled ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
951c6aad5d |
[5763448][ONNX][Autocast] Fix Resize input type mismatch error (#757)
## What does this PR do? **Type of change:** Bug fix **Overview:** This PR fixes an input type mismatch in Resize layers when being converted to FP16. ## Usage ```python $ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx ``` ## Testing Added unittest. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information This issue is also fixed by using the standalone type inference logic from https://github.com/NVIDIA/Model-Optimizer/pull/719. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Improvements** * Enhanced the graph sanitization process to automatically duplicate shared constants during optimization, ensuring improved model handling and consistency. * **Tests** * Added test coverage for mixed precision conversion of Conv-Resize model architectures. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
18d9b1eea4 |
Add static per block MSE for NVFP4 weight (#613)
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature
**Overview:** ?
Support static block-wise MSE for NVFP4 weight quantization.
Add a FP4 triton kernel that take in scales for each block. It also
quantizes the scales to FP8.
This PR does the following:
1. Enable static NVFP4 implementation, i.e. block scales for weights are
calculated during calibration and feed into fake quant kernels
2.Extend mse_calibrate to support static NVFP4 with block scales
searching by MSE and global scale set as MAX
3.Refinements: calibrate weight quantizers only once during MSE
calibration
## Usage
<!-- You can potentially add a usage example below. -->
Example config:
```python
NVFP4_WEIGHT_MSE_CFG = {
"quant_cfg": {
"*weight_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
"*input_quantizer": {
"enable": False,
},
**_default_disabled_quantizer_cfg,
},
"algorithm": {
"method": "mse",
"step_size": 0.25,
"start_multiplier": 0.25,
"stop_multiplier": 2.0,
},
}
NVFP4_WEIGHT_ACT_MSE_CFG = {
"quant_cfg": {
"*weight_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
"*input_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
**_default_disabled_quantizer_cfg,
},
"algorithm": {
"method": "mse",
"step_size": 0.25,
"start_multiplier": 0.25,
"stop_multiplier": 2.0,
},
}
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
|
||
|
|
b4c77c0d9b |
[NVBUG 5801937] Disable dq_only by default (#777)
## What does this PR do? **Type of change:** Bug fix **Overview:** Disable dq_only flag by default in modelopt onnx quantization ## Testing Able to build and run model with modelopt onnx Python CLI - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No - dq_only is set to False by default - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Chores** * Updated quantization default behavior: Q/DQ (Quantize/Dequantize) nodes are now added by default instead of only Dequantize nodes. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> |
||
|
|
9de4877d27 |
[1/2] Address security concerns in code (#626)
- [x] Address feedback on Threat and Vuln Analysis (TAVA) doc by ProdSec team - [x] Add note on safe usage of pickle deserialization of modelopt-generated state files **TODO: [Separate PR]** Replace pickle usage in `modelopt/torch/opt/plugins/megatron.py` - Needs fix on TransformerEngine first as we copy from https://github.com/NVIDIA/TransformerEngine/blob/3ff0b8d4/transformer_engine/pytorch/module/base.py#L863 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--trust_calibration_data` CLI flag for secure ONNX quantization with pickle data files. * **Improvements** * Enhanced security validation for generated quantization code. * Simplified data loading by removing pickle-based caching—data is now always loaded fresh. * Added security guidance throughout model state loading operations. * **Documentation** * Updated guides with security best practices for model state handling. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
90fa48ce14 |
remove duplicated RMSNorm and use LlamaRMSNorm from transformers (#774)
## What does this PR do? Code cleanup **Overview:** Remove RMSNorm which is identical to LlamaRMSNorm from transformers. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Updated the normalization implementation in the Eagle speculative module. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
5e0d36551a |
[5796745][ONNX][Autocast] Fix opset check for model with custom ops (#767)
## What does this PR do?
**Type of change:** Bug fix
**Overview:** This PR fixes the opset being incorrectly detected as
being `1` in models with custom ops. That happens because the
'trt.plugins' domain version is detected rather than the actual model's
opset version.
## Usage
<!-- You can potentially add a usage example below. -->
```python
$ python -m modelopt.onnx.autocast --onnx_path=${MODEL_NAME}.onnx
```
## Testing
See bug 5796745.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Refactor**
* Optimized the opset version detection logic for more efficient model
conversion handling.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
|
||
|
|
b813ab548b |
Top-K KL Divergence loss (#747)
## What does this PR do? **Type of change:** New feature **Overview:** Writes a new KLDiv Logits loss which only uses top-k vocab values ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Top-K logit filtering capability for knowledge distillation workflows, enabling selective focus on high-probability tokens. * **Improvements** * Enhanced distributed tensor model-parallel operations with improved awareness for gradient computation and reduction. * Simplified legacy distributed operation constructs. * **Tests** * Introduced comprehensive test coverage for Megatron-based distillation, validating both standard and Top-K filtering variants. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
7836065f3e |
Set trust_remote_code default to False (#769)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> DeepSeek later has official `transformers` support and use `trust_remote_code=True` will encounter error. Set default to `False`. **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
6ae96b5162 |
[NVBug 5784940] Fix autodeploy example (#764)
## What does this PR do? **Type of change:** bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** update the example with new API ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> python examples/llm_autodeploy/api_server.py --ckpt_path TinyLlama/TinyLlama-1.1B-Chat-v1.0 --world_size 1 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * AutoDeployConfig parameter handling has been updated to improve how parameters are prepared for language model initialization. The configuration method now uses optimized keyword argument formatting to ensure consistency and clarity across the deployment system. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
727da95a91 |
SpecDec Bench: PostProcess flag (#759)
## What does this PR do? **Type of change:** ? Bug Fix: https://nvbugspro.nvidia.com/bug/5795144 **Overview:** ? Pass postprocess flag to handle slicing message. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Introduced --postprocess command-line option to select postprocessing strategy. Users can choose "base" (default, preserves existing behavior) or "gptoss" (new alternative method) with validation to reject invalid selections. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Izzy Putterman <iputterman@nvidia.com> |
||
|
|
b484efb84e |
[CI] Cleanup ubuntu-runner disk storage before installing deps (#765)
We started seeing this issue in GitHub's free ubuntu-latest runners: `ERROR: Could not install packages due to an OSError: [Errno 28] No space left on device`. Suggested by other GH runners users to remove unnecessary android / dotnet files to avoid the storage issue. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Consolidated GitHub Actions workflow setup into a reusable custom action for improved maintainability and consistency across CI/CD pipelines. * Enhanced release workflow with automated unit testing and artifact upload capabilities. * Streamlined runner initialization by reducing redundant configuration steps. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
510451322c |
Streamline KD & QAD transformers Trainers (#708)
## What does this PR do? **Type of change:** ? Refactor and stabilization **Overview:** * Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so that we can remove multiple restrictions ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
5b9261f084 |
Add .coderabbit.yaml for auto PR reviews (#756)
Automatically review PRs by Coderabbit <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Added configuration settings for code review automation. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
7971fff058 |
[5694695][AutoCast] Preserve outer scope variable types in subgraphs (#717)
## What does this PR do? **Type of change:** Bug fix **Overview:** When clearing type information for shape inference, preserve value_info for outer scope variables in subgraphs. Previously, all value_info entries were cleared indiscriminately, causing shape inference failures when subgraph nodes referenced outer scope variables. ## Testing pytest tests/unit/onnx/autocast/test_precisionconverter.py::test_if_subgraph_outer_scope_type_preservation ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: N/A - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> |
||
|
|
ecda7b0bfa |
Use kitchen FA in huggingface plugin (#674)
## What does this PR do? new feature **Overview:** use kitchen FA in huggingface plugin ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> |
||
|
|
307fe7183b |
Fix QuantSequentialMLP sharded_state_dict (#742)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug **Overview:** ? These fixes are needed for Megatron-LM `main` branch due to some changes in `sharded_state_dict`. Qwen3-30B-A3B PTQ and resume fails while EP=4 cannot load a checkpoint generated with PP=4. `singleton_local_shards` must be added to the metadata; otherwise, all experts `amax` are packed to gather and currently the TP `replica_id` for `linear_fc1` is incorrect. **Other Finding:** This limits TP=ETP=1 when EP>1. Otherwise, there will be `sharded_state_dict` access error. There is a potential blind spot of using the default TP group in `ColumnParallelLinear` and `RowParallelLinear` since it can be part of the MoE where the tensor parallelism is controlled by ETP instead. Will need a different PR to fix the parallel_state. **Results:** If calibrate with EP=1, mmlu = 0.80. This can be resumed with EP=4, TP=1, ETP=1 (TP>1 does not work as mentioned above). However if calibrated with EP=4, then mmlu = 0.71 which shows there are some issues with max sync in EP. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
6f18490b83 |
Improve AWQ init speed (#748)
## What does this PR do? **Type of change:** ?Improvement<!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Improve speed of accessing weight through enable_weight_access_and_writeback in AWQ helper init. This change reduces the time complexity from O(num_modules^2) to O(num_modules) and the runtime from ~1hour to 30 seconds. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> python hf_ptq.py --pyt_ckpt_path /home/scratch.omniml_data_1/models/qwen/Qwen3-30B-A3B-Instruct-2507 --qformat int4_awq ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
9c24e2c08e |
Fix Deepseek transformers model loading (#740)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? For Deepseek, let's force the user to apply trust_remote_code and use AutoModelForCausalLM for loading the model. ## Testing python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4 --export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64 --trust_remote_code --dataset cnn_dailymail ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
68d604dd69 |
Move new puzzle dist utils from feature/compress to main (#746)
- Move new `modelopt.torch.utils.distributed` from `feature/compress` to `main` branch so they can be used via modelopt in puzzletron gitlab Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
9a3b986f9c |
Fix TRT-LLM 2-gpu CI test shm issue (#744)
- As suggested by NVGHA runners team to increase SHM size to avoid issue on 2-gpu nightly tests Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
81c509c643 |
Fixes & Simplifications for MCore KVCache QAT/QAD; Unittests; Distributed Sync of KVCache Quantizer params (#727)
## What does this PR do? **Type of change:** Fix MCore KV Cache Quantization: Amax Device Placement Bug; Code clean up; Distributed Sync of KVCache Quantizer params; unittest expansion to hybrid models **Overview:** Fixes bugs preventing MCore KV Cache quantization from working during checkpoint restore. ### Bug Chain **Bug 1:** `is_enabled = self.weight_quantizer.is_enabled if hasattr(self, "weight_quantizer") else False` No `weight_quantizer` for KV-cache-only quant → `is_enabled=False` → metadata not saved → `modelopt_post_restore()` never called. *(Thanks to @jenchen13 )* **Bug 2:** After fixing Bug 1, `_amax` restored on CPU (via `_reset_pytorch_state_from_metadata`). Fallback `_calibrate_quantizers()` never called because `_amax` exists. **Bug 3:** Even if called, `_calibrate_quantizers()` fails — `core_attention` has no parameters → can't determine device/dtype. ### The Fix 1. Remove `is_enabled` check entirely — disabled modules may still need metadata restore. Explicitly skip `output_layer` from extra state callbacks (never quantized) 2. Set `dtype`/`device` on `core_attention` from parent Attention module, `modelopt_post_restore()` calls `self.to(device, dtype)` 3. Remove dead `_calibrate_quantizers()` code (will bring back similar logic for KV cache affine quantization) ### Previous Unit Test Was Wrong `model_test` was `mtq.quantize()`'d, not `mto.restore()`'d. Never tested actual restore path. ### Additional Fixes - Amax sync across DP/TP for KV cache quantizers - `flash_decode` auto-disabled ### Code Cleanup Removed ~100 lines of dead code. ## Testing 1. MCore KV Cache QAD with Nano V3 + Context Parallel works 2. Unit tests: hybrid models, KV+GEMM configs, correct restore workflow, backward pass validation ## Before your PR is "*Ready for review*" - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update Changelog?**: Yes --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Co-authored-by: Asma Thekkumpate <akuriparambi@cw-dfw-cs-001-vscode-02.cm.cluster> |
||
|
|
fe52b2a46e |
Bias running average computation in float (#738)
## What does this PR do? **Type of change:** Bug fix **Overview:** ? Computing Bias running average with bf16 creates incorrect estimations. Impact on accuracy for Qwen2.5-7B model: With BF16 running average: NVFP4_AFFINE_KV | 59.11% -- | -- With running average in Float: NVFP4_AFFINE_KV | 71.81% -- | -- ## Usage Use examples/lm_eval/mmlu.py with batchsize of 1 Note: the issue is masked with larger batch sizes ## Testing - Ran mmlu benchmark with mmlu.py and nv-eval - also ploted bf16 and float running average for different layers, one of the example for layer 0 in Qwen2.5-7B: <img width="2100" height="600" alt="image" src="https://github.com/user-attachments/assets/715059c5-34a4-495e-b6f1-0b57cf0c08af" /> Note: for the larger value bf16 shows smaller value compared to float ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: NA - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
4eb1835df5 |
Support for KV cache quantization for MLA Attention vLLM fakequant (#714)
## What does this PR do? **Type of change:** Feature extention **Overview:** Added support to quantize KV cache in vLLM fakequant by adding quantization support for [MLAAttention](https://github.com/vllm-project/vllm/blob/v0.11.1/vllm/attention/layer.py#L641) ## Usage Please refer to [Readme](https://github.com/NVIDIA/Model-Optimizer/tree/kinjal/vllm_att_quant/examples/vllm_serve#calibrate-and-serve-fake-quant-model-in-vllm) ```shell KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CFG=NVFP4_DEFAULT_CFG python vllm_serve_fakequant.py deepseek-ai/DeepSeek-V2 --served-model-name deepseek-ai/DeepSeek-V2 --host 0.0.0.0 --port 8001 --trust-remote-code --enforce-eager --gpu-memory-utilization 0.8 ``` ## Testing Locally tested KV Cache quantization ``` (rotary_emb): DeepseekScalingRotaryEmbedding() (mla_attn): MultiHeadLatentAttentionWrapper( (fused_qkv_a_proj): QuantMergedColumnParallelLinear( in_features=5120, output_features=2112, bias=False, tp_size=1, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=141.0000 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.4297 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (q_a_layernorm): RMSNorm(hidden_size=1536, eps=1e-06) (q_b_proj): QuantColumnParallelLinear( in_features=1536, output_features=3072, bias=False, tp_size=8, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=32.0000 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.1670 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (kv_a_layernorm): RMSNorm(hidden_size=512, eps=1e-06) (kv_b_proj): QuantColumnParallelLinear( in_features=512, output_features=4096, bias=False, tp_size=8, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.2773 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (rotary_emb): DeepseekScalingRotaryEmbedding() (o_proj): QuantRowParallelLinear( in_features=2048, output_features=5120, bias=False, tp_size=8, reduce_results=True (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.7188 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.4336 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (mla_attn): QuantMLAAttention( (q_bmm_quantizer): TensorQuantizer(disabled) (kv_c_bmm_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant) ) ) ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:NA - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
8426c363bd |
Skip unit tests in release workflow avoid storage issues in runner
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>0.42.0dev 0.41.0rc1 |
||
|
|
d541324e84 |
Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do? **Type of change:** ? Recipe improvement **Overview:** ? Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3 Next recipe for accuracy recovery ## Testing Model accuracy benchmarking Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
b655321d87 |
[Issue 543] [Bug fix] Fix dynamic input quant for AWQ (#726)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Dynamic input quantizers, e.g., MXFP4, are not restored after AWQ. This PR fix the issue. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> Tested with MXFP4, NVFP4, int4 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
883c8731aa |
Noeyy/add new cases for newly added checkpoints on HF (#728)
## What does this PR do? **Type of change:** Add TRT LLM/vLLM/SGLang functional test cases for newly added checkpoints on HF **Overview:** 1.Since the speculative draft model only supports loading from a local path, we should set the MODELOPT_LOCAL_MODEL_ROOT environment variable. If we don't set it, these test cases will be skipped. 2. Newly added checkpoints: - nvidia/gpt-oss-120b-Eagle3-short-context - nvidia/gpt-oss-120b-Eagle3-throughput - nvidia/EAGLE3-NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-FP8 ## Usage ```python pytest tests/examples/llm_ptq/test_deploy.py --run-release ``` ## Testing Run release testing ## Before your PR is "*Ready for review*" Ready for review - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information N/A --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
3350b0a45b |
[OMNIML-3017] MLM QAD example (#682)
## What does this PR do?
**Type of change:** New example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
### Add QAD Training example for Megatron-LM
- File Structure
- qad.sh / sbatch_qad.sh - Training and SLURM submission scripts
- data_utils/ - Dataset download and preprocessing utilities
- configs/ - Configuration templates for Qwen3-30B-A3B (MoE) and
Qwen3-8B (Dense)
- Key Features
- One-button dataset generation (OpenScience + Nemotron-v2)
- Config-based training scripts, keep all tunable knobs into a single
config file
## Usage
<!-- You can potentially add a usage example below. -->
1. Generate dataset
```bash
bash data_utils/generate_dataset.sh \
--output-dir /path/to/datasets \
--mlm-path /path/to/Megatron-LM \
--tokenizer Qwen/Qwen3-30B-A3B-Instruct-2507
```
2. Create a config based on templates
3. Kick off training with Slurm:
```bash
sbatch sbatch_qad.sh --config configs/my-experiment.conf
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
QAD with Qwen3-30B-A3B-instruct-2507 NVFP4 (all layers quantized)
- GPQA:
BF16: 0.549
NVFP4 (PTQ): 0.4949
NVFP4 (QAD): 0.5202
- Livecodebench:
BF16: 0.3987
NVFP4 (PTQ): 0.37
NVFP4 (QAD): 0.3855
- Scicode:
BF16: 0.325
NVFP4 (PTQ): 0.276
NVFP4 (QAD): 0.3146
- AIME
BF16: 0.6049
NVFP4 (PTQ): 0.55
NVFP4 (QAD): 0.5431
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Wei-Ming Chen <weimingc@login-eos01.eos.clusters.nvidia.com>
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
03dc3860af |
Update onnxruntime-gpu (#697)
## What does this PR do? **Type of change:** Bug fix **Overview:** Updated setup.py to use only onnxruntime-gpu and removed onnxruntime-directml as dependency. Also changed onnxruntime-gpu version in examples. ## Testing Tested int4 quantization and MMLU benchmark with updated onnxruntime-gpu , working as expected --------- Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
cb343352ca |
Registry interface for custom quantization functional backend (#683)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Add registry interface for custom quantization functional backend ## Usage <!-- You can potentially add a usage example below. --> see `tests/unit/torch/quantization/test_custom_backend.py` for usage example. ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |