mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
5e43b2a5f586c2b34824c344a8b681e6c50378c4
449
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5e43b2a5f5 |
Support Qwen3 Next MTP load and export (#860)
## What does this PR do? Fix MTP export for Qwen3 Next **Overview:** ? For Qwen3 next, the MTP weights are not stored separately in safetensors. So we use "mtp" weights key to decide if the weights are for MTP or not. ## Testing Qwen3 Next PTQ and check if MTP is in the exported checkpoint. scripts/huggingface_example.sh --model <Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Optimized Multi-Token Prediction weight loading with improved layer detection and handling. * **Chores** * Simplified status reporting to display total loaded weights and detected layers. * Removed verbose per-file warnings for cleaner console output. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Zhiyu <zhiyuc@nvidia.com> |
||
|
|
a8f5314c93 |
fix the path change in torch v2.10 for spec dec (#863)
## What does this PR do? **Type of change:** bug fix **Overview:** torch v2.10 changes the path for _SDPAMerger. will need to use the new path for import ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated internal import references to reflect organizational changes in dependencies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
24e358789d |
[5868890][ONNX][Autocast] Fix: failure when checking input shape with unknown dimension (#859)
## What does this PR do? **Type of change:** Bug fix **Overview:** Skip unknown dimensions when comparing input shape in model vs calibration data. ## Usage ```python $ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz ``` ## Testing See bug 5868890. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Bug Fixes * Enhanced input shape validation to properly handle dynamic tensor dimensions, allowing more flexible dimension checking while maintaining validation accuracy. <!-- end of auto-generated comment: release notes by coderabbit.ai --> ## Aditional info Regression introduced in https://github.com/NVIDIA/Model-Optimizer/pull/652. --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
e53ca61b71 |
Add contribution guidelines for experimental features (#867)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive guide for experimental optimization technique development, including recommended structure, testing conventions, licensing requirements, and graduation path to production. * **New Features** * Introduced experimental package with templates and utilities for implementing research-stage optimization techniques. Includes configuration framework and example code patterns. Emits stability warnings to indicate unstable APIs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Kai Xu <kaix@nvidia.com> |
||
|
|
62c2799be2 |
Fix Sequential MLP amax sync deadlock (#862)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? After `QuantMoELayer`, we rely on `layer_sync_moe_local_experts_amax` to first perform local sync. This is supposed to create `input_quantizer.amax` for all experts but the current logic will only update experts that already have `amax`. This results in some experts are still missing `amax`. With the fact above, `sync_quantizer_amax_across_dp_ep` will actually deadlock seems the collective is called based on whether `quantizer._amax is None`. Any expert with `None` amax will not call collective hence will never arrive the collective and cause a deadlock. We fix `layer_sync_moe_local_experts_amax` such that even if an expert does not have `amax`, we will overwrite it with a clone of the global amax. The post condition should be all experts have `amax` and the pre condition of `sync_quantizer_amax_across_dp_ep` should be the same. **Note:** we found that `_check_moe_calibration_complete` actually didn't raise any error even some experts have no amax. Didn't look into this problem. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved synchronization of quantization parameters for Mixture of Experts (MoE) models with more flexible configuration support. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
ac30686c82 |
Track global_amax for weight FP4 MSE sweep; Refactor to NVFP4StaticQantizer, NVFP4MSECalibrator (#849)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added NVFP4StaticQuantizer for improved 4-bit quantization with enhanced precision control * Introduced NVFP4MSECalibrator with flexible candidate generation for calibration optimization * **Improvements** * Optimized GPU kernels for Hopper+ graphics cards with better performance * Extended Triton support to broader GPU compatibility * Enhanced backward compatibility for restoring previously quantized models * **Tests** * Added comprehensive test coverage for new quantizers and calibration methods <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
3393e981e6 |
Fix TEGroupedLinear quantization for expert parallelism (EP > 1) (#833)
## What does this PR do? **Type of change:** Bug fix / Compatibility update **Overview:** Fix `te_grouped_quantized_linear_fn` argument parsing for TEGroupedLinear quantization when parallelism configuration results in fewer local experts per GPU. ### Problem TransformerEngine changed the _GroupedLinear.forward signature in PR #2377 (released in TE 2.10): Old signature (TE < 2.10): forward(ctx, inp, m_splits: List[int], use_bias, is_first_microbatch, ...) New signature (TE >= 2.10): forward(ctx, inp, non_tensor_args: Tuple, *weights_and_biases) where non_tensor_args = (m_splits, use_bias, is_first_microbatch, ...) Without this fix, ModelOpt's quantization code fails with newer TE versions because it tries to access m_splits directly from args[idx + 1], but in TE >= 2.10, that position contains the non_tensor_args tuple instead. ### Root Cause The code assumed m_splits was always directly accessible at args[idx + 1], but TransformerEngine PR #2377 changed the signature to pack all non-tensor arguments into a tuple. Taking Qwen3-30B-A3B (with `num_gemms=21`, threshold=44) as an example: ### Solution Added version checking to handle both signatures: ```python if Version("2.10") <= _TE_VERSION: # New signature: non_tensor_args is a tuple, m_splits is the first element num_gemms = len(args[idx + 1][0]) else: # Old signature: m_splits is directly args[idx + 1] num_gemms = len(args[idx + 1]) ``` ## Usage <!-- You can potentially add a usage example below. --> Works seamlessly with any TransformerEngine version: ```python # High EP quantization - previously failed, now works torchrun --nproc_per_node 8 examples/quantization/quantize.py \ --hf-model-id /models/Qwen3-30B-A3B \ --export-quant-cfg fp8 \ --megatron-save-path /models/Qwen3-30B-A3B_fp8_mlm \ --tp 8 \ --ep 8 # High EP inference - previously failed, now works torchrun --nproc_per_node 8 examples/quantization/ptq_generate.py \ --megatron-load-path /models/Qwen3-30B-A3B_fp8_mlm \ --hf-model-id /models/Qwen3-30B-A3B \ --tp 8 \ --ep 8 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ```python # High EP quantization - previously failed, now works torchrun --nproc_per_node 8 examples/quantization/quantize.py \ --hf-model-id /models/Qwen3-30B-A3B \ --export-quant-cfg fp8 \ --megatron-save-path /models/Qwen3-30B-A3B_fp8_mlm \ --tp 8 \ --ep 8 # High EP inference - previously failed, now works torchrun --nproc_per_node 8 examples/quantization/ptq_generate.py \ --megatron-load-path /models/Qwen3-30B-A3B_fp8_mlm \ --hf-model-id /models/Qwen3-30B-A3B \ --tp 8 \ --ep 8 ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Enhanced Mixture of Experts (MoE) calibration validation and synchronization to ensure consistency across distributed training setups. * Improved grouped linear quantization robustness to handle varying input patterns and tensor dimensions. * **Improvements** * Better error handling for incomplete MoE expert calibration detection. * More flexible argument parsing for quantization operations. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: James Shen <yueshen@nvidia.com> |
||
|
|
0225e5273f |
Integrate Automated QDQ placement tool - part 2.1 (#844)
## What does this PR do?
This PR implements RegionPattern class. RegionPattern describes local
topology structure of a Region. Regions with same Pattern could be
autotune together. Best insertion points of a given pattern could also
be saved to accelerate the next QDQ autotuning.
**Overview:** ?
## Usage
```python
python -m modelopt.onnx.quantization.autotune.region_search --model model.onnx --verbose
```
```
├─ Region 212 (Level 0, Type: COMPOSITE)
│ ├─ Direct nodes: 0
│ ├─ Total nodes (recursive): 9
│ ├─ Children: 1
│ ├─ Inputs: 3 tensors
│ │ - xxx
│ │ - xxx
│ │ - xxx
│ └─ Outputs: 1 tensors
│ - xxx
│
│ Child regions:
│
├─ Region 209 (Level 2, Type: LEAF)
│ ├─ Direct nodes: 9
│ ├─ Total nodes (recursive): 9
│ ├─ Children: 0
│ ├─ Inputs: 11 tensors
│ │ - xxx
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No, document
update is in Part 4
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
CHANGELOG update could be done after all changes are ready.
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Enhanced ONNX quantization analysis with improved region pattern
matching and comparison capabilities.
* Added utility to identify quantized tensors in models for better
analysis.
* **Tests**
* Comprehensive test coverage for region pattern functionality and
quantization utilities.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Will Guo <willg@nvidia.com>
|
||
|
|
452c5a09b0 |
GLM-4.7 MTP support (#792)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Enable GLM-4.7 PTQ workflow, including loading the standalone MTP modules and export as-is. ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/GLM-4.7 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/GLM-4.7-NVFP4-0203 --trust_remote_code ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added quantization support for GLM-4.7 model with automatic handling of specialized layer architecture. * Added image-text data calibration capabilities for Nemotron VL model quantization. * **Documentation** * Updated support matrix to reflect newly supported models and quantization features. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
944dd1a284 |
Move parallel_state init and warnings to Quant DynamicModule + MBridge pruning doc update (#854)
## What does this PR do? **Type of change:** Minor improvement <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Only quantization DynamicModules use the parallel_state attribute so for all other model opt methods, we see a parallel state not initialized warning which could be confusing hence moving it to QuantModule class instead Minor update to MBridge pruning docs ## Testing <!-- Mention how have you tested your change if applicable. --> N/A ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Distributed parallel state support is now available in quantization workflows for multi-GPU training. * **Bug Fixes** * Improved resource cleanup in distributed training to ensure proper environment finalization. * **Documentation** * Updated example paths and added new manual pruning configuration examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
8a8c250d12 |
Latent MOE & Repeated MTP support for NemotronH; fix KV cache quant export (#830)
## What does this PR do? **Type of change:** New feature and bug fix **Overview:** Support Latent MOE and Repeated MTP for NemotronH models - Enable latent MOE modules during megatron import/export - Fix KV cache quantization export: remove old `qkv_layer.output_quantizer` export & replace with proper `k/v_bmm_quantizer` logic - Improvements to EP amax sync - Support repeated MTP import/export for NemotronH models (only BF16 export for MTP for now) ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added support for grouped MLP and self-attention scaling operations in model export workflows * Enhanced model parallel training capabilities with improved component mapping * Expanded quantization configuration handling with dynamic module exclusion across distributed ranks * Improved support for additional transformer engine components * **Refactor** * Reorganized internal export and import logic for improved maintainability and specialist model architecture support <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: jenchen13 <jennifchen@nvidia.com> Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> |
||
|
|
2e43c80609 |
[2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This MR adds HuggingFace checkpoint export support for LTX‑2 by treating TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling back to writing model.safetensors when save_pretrained isn’t available. It also preserves the original forward in DynamicModule patching (_forward_pre_dm) so downstream callers can still invoke the pre-patched forward implementation. **Changes** 1. Added the calibration & quantization support of the LTX2, even with FP8 precision. 2. Preserve original forward before `DynamicModule` patching: when patching forward, we now stash the pre-patched implementation in `self._forward_pre_dm` (once) so downstream code can still call the original forward, then re-bind forward to the class implementation. This is needed for the LTX2 FP8 calibration. 3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a “diffusion-like” object and routes it through _export_diffusers_checkpoint() (import guarded; no hard dependency). 4. Generalized component discovery: introduced get_diffusion_components() (aliasing the old get_diffusers_components) to support non-diffusers pipelines; for LTX‑2 it returns only stage_1_transformer. 5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy forward generator (generate_diffusion_dummy_forward_fn) that builds minimal LTX Modality inputs (including correct timesteps broadcasting) so shared-input hooks can run and fuse QKV when applicable. 6. Export fallback for non-save_pretrained modules: when a component lacks save_pretrained (LTX‑2 transformer), export now writes model.safetensors + minimal config.json instead of pytorch_model.bin. Plans - [x] [1/4] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml - [ ] [3/4] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml - [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ```bash python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added LTX-2 video model support with complete quantization and export pipeline integration * Introduced `--extra-param` CLI option for flexible model configuration and parameter passing * Enhanced export capabilities with broader diffusion model compatibility * **Chores** * Changed default model data type from Half to BFloat16 for improved numerical stability <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
87237e7dd1 |
Update on the QuantModule & DynamicModule to accept external forward (#824)
## What does this PR do?
**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
This MR improves robustness when `forward()` is monkey‑patched (replaced
at runtime) on modules that later get wrapped/converted by ModelOpt
(DynamicModule + quant wrappers).
It addresses two concrete failure modes introduced/exposed by supporting
“patched forward” modules:
1. Forward “leakage” after export: a dynamic wrapper forward could
remain bound on an instance even after export() restores the original
(non‑dynamic) class, causing runtime errors in unrelated codepaths (e.g.
KD export/save/restore chains).
1. Infinite recursion in quant wrappers: _forward_pre_dm can sometimes
point to a wrapper forward that already participates in the class chain,
causing a recursion loop when quant wrappers call _forward_pre_dm
directly.
## Usage
<!-- You can potentially add a usage example below. -->
```python
lin = torch.nn.Linear(4, 4)
def upcast_forward(x):
# external closure: NOT part of any class MRO
return torch.nn.functional.linear(x, lin.weight.to(x.dtype), lin.bias.to(x.dtype))
lin.forward = upcast_forward # framework/user patches forward
# Later, ModelOpt converts/wraps the module.
# It stashes the patched function as `_forward_pre_dm` and binds the wrapper forward on the class.
# During quantization, QuantInputBase.forward sees `_forward_pre_dm` is NOT in MRO -> calls it.
```
```
# Imagine a module already wrapped by quant classes:
# QuantLinearConvBase.forward -> super().forward -> QuantInputBase.forward -> ...
# If `_forward_pre_dm` accidentally points to QuantLinearConvBase.forward (which IS in MRO),
# and QuantInputBase.forward calls it directly, you get:
# QuantInputBase.forward -> _forward_pre_dm (QuantLinearConvBase.forward)
# -> super().forward -> QuantInputBase.forward -> ...
# infinite recursion
# The fix: if `_forward_pre_dm` is a forward already in MRO, ignore it and use super().forward.
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Bug Fixes**
* Improved forward method restoration during module export to prevent
state leakage
* Enhanced quantization behavior when using chained optimization modes
* **Tests**
* Added regression tests for quantization with runtime forward patching
* Added validation tests for sparse quantization combined with
distillation workflows
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
0.42.0rc0
|
||
|
|
9de9b8f20c |
Noeyy/add test cases for the newly added checkpoints on HF (#827)
## What does this PR do? **Type of change:** new tests **Overview:** Add new test cases for the newly added checkpoints on HuggingFace. ## Usage pytest test_deploy.py --run-release ```python None ``` ## Testing None ## Before your PR is "*Ready for review*" - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No ## Additional Information None <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for NVFP4 model variants across multiple model families (DeepSeek, Llama, Qwen, and others). * **Improvements** * Enhanced backend availability detection to automatically identify and manage supported deployment backends at runtime. * **Tests** * Improved test infrastructure for better reproducibility and backend compatibility handling. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: noeyy-mino <174223378+noeyy-mino@users.noreply.github.com> |
||
|
|
e247f5d0e4 |
Fixes for Megatron Expert Parallel, GroupedMLP and SequentialMLP (#831)
## What does this PR do?
**Type of change:** Bug fix / Improvement
**Overview:**
Fix MoE quantization calibration sync by removing the force-routing
workaround and adding explicit validation for incomplete calibration.
**Problem:** During MoE calibration, some experts may not receive tokens
(router doesn't select them). This causes amax=None on some ranks while
others have valid values, leading to hangs or failures during
distributed amax sync.
Previous workaround: Force all tokens through all experts during
calibration. This was however causing the following error:
```
File "/opt/TensorRT-Model-Optimizer/modelopt/torch/quantization/plugins/transformer_engine.py", line 152, in te_grouped_quantized_linear_fn
quantized_inputs = self.input_quantizer(inp)
File "/opt/TensorRT-Model-Optimizer/modelopt/torch/quantization/calib/max.py", line 69, in collect
assert torch.all(local_amax >= 0), (
torch.AcceleratorError: CUDA error: an illegal memory access was encountered
```
This is probably because forcing all tokens through all experts, the
inputs become garbage and are possibly inf/nan causing calibration to
fail.
Solution:
Remove _QuantMoELayer force-routing workaround
Add validation before sync: detect if some ranks have amax=None while
others have values
Raise clear error: "MoE calibration incomplete: increase --calib-size" -
This is a cleaner solution. In case of under calibration we just raise a
clear error.
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes (Verified backward
compatibility by loading a MoE model saved before this change)
- **Did you write any new necessary tests?**: No - existing MoE tests
cover the sync behavior
- **Did you add or update any necessary documentation?**: Not needed,
low level change
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No needed, low level change
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Improvements**
* Enhanced Mixture of Experts (MoE) quantization with comprehensive
calibration validation to ensure consistent synchronization across
distributed experts.
* **Refactor**
* Streamlined MoE quantization architecture by consolidating internal
handling mechanisms.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: realAsma <akuriparambi@nvidia.com>
|
||
|
|
e02409773b |
Increase nighytly gpu test timeout to 150mins
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
fb5923c89e |
Add Megatron-Bridge pruning example scripts (#800)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
Megatron-Bridge pruning example scripts (HF input, HF / Megatron
output). Also defined some utility functions we can reuse for adding
examples for quantization or other optimizations:
- `modelopt.torch.utils.plugins.mbridge.load_mbridge_model_from_hf`:
Load HF to MBridge with ModelOpt spec in desired TP/PP/etc configuration
-
`modelopt.torch.utils.plugins.mbridge.get_hf_mbridge_calibration_loop`:
Create `forward_loop` for calibration on a HF dataset
- Supports all datasets available in
`modelopt.torch.utils.dataset_utils` (`cnn_dailymail`,
`nemotron-post-training-dataset-v2`, etc)
- Support applying chat template for chat-based data
## Usage
<!-- You can potentially add a usage example below. -->
From `nvcr.io/nvidian/nemo:26.02.rc1` container (mount latest code to
`/opt/Megatron-Bridge` and `/opt/Model-Optimizer`)
```python
torchrun --nproc_per_node 2 /opt/Model-Optimizer/examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_target_params 6e9 \
--hparams_to_skip num_attention_heads \
--output_hf_path /tmp/Qwen3-8B-Pruned-6B
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- [x] Manually ran pruning script in nemo:25.11 container (plus modelopt
and mbridge mounted to latest) for Qwen3-8B and Nemotron-Nano-9B-v2 with
PP=8 and PP=4
- [ ] Added per-PR CI/CD test for example script
Results when pruning Qwen3 8B -> 6B (10 different configurations) with
and without chat template on the dataset samples. Perhaps MMLU is not
the right metric to look at.
| Layers | Hidden Size | FFN Hidden Size | Params | MMLU (Concatenated
messages) | MMLU (Applied Chat Template) |
|--------|-------------|-----------------|--------|------------------|----------------------|
| 34 | 3328 | 11264 | 5.99B | 0.401 | 0.393 |
| 30 | 3584 | 11776 | 5.99B | 0.588 | 0.576 |
| 36 | 3840 | 8192 | 5.98B | 0.507 | 0.518 |
| 36 | 3584 | 9216 | 5.98B | 0.477 | 0.469 |
| 36 | 3072 | 11776 | 5.97B | 0.255 | 0.249 |
| 32 | 3584 | 10752 | 5.96B | 0.554 | 0.549 |
| 28 | 4096 | 10240 | 5.94B | 0.400 | 0.438 |
| 36 | 4096 | 7168 | 5.93B | 0.461 | 0.438 |
| 36 | 3328 | 10240 | 5.92B | 0.362 | 0.359 |
| 34 | 3840 | 8704 | 5.91B | 0.515 | 0.546 |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: ‼️ TODO
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
**New Features**
* Added new Megatron-Bridge pruning example demonstrating Minitron-based
model optimization with advanced pruning configurations.
**Documentation**
* Updated core project documentation to highlight Megatron-Bridge as a
supported optimization framework.
* Added comprehensive example documentation for Megatron-Bridge
workflows including pruning, distillation, and quantization.
* Updated pruning guides with Megatron-Bridge integration examples and
best practices.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
23278a44db |
Integrate Automated QDQ placement tool - Part 1 (#701)
## What does this PR do? **Type of change:** new feature **Overview:** This PR integrates an automatical QDQ placment tool into ModelOpt. This PR is the 1/4 parts of the change, it contains the following changes: 1. Defines common types: Region, RegionType, Error types 2. Defines InsertionPoints (the logical localtion to place QDQ pairs), InsertionScheme (a set of insertion points) 3. Unit tests for new types Part 1: https://github.com/NVIDIA/Model-Optimizer/pull/701 Part 2: https://github.com/NVIDIA/Model-Optimizer/pull/702 Part 3: https://github.com/NVIDIA/Model-Optimizer/pull/703 Part 4: https://github.com/NVIDIA/Model-Optimizer/pull/704 ## Usage ```python # Region type usage: region = Region(region_id=1, level=0, region_type=RegionType.LEAF) assert region.get_id() == 1 assert region.get_level() == 0 region.add_node(1) # 1 is the index of ONNX graph node ... point = NodeInputInsertionPoint(node_index=0, input_index=2) assert point.node_index == 0 # relative node index in region assert point.input_index == 2 # relative input tensor index in specific node resolved = point.resolve(region, graph) ... ``` ## Testing Implement unit tests, all tests could get passed. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No, document change will be included in part 4. - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No, this could be done when all parts of the change are merged. ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added foundational autotuner infrastructure for quantization optimization, including region hierarchies and insertion scheme management. * Introduced insertion point system for managing quantize/dequantize operation placement across ONNX graph regions. * Added utility functions for tensor consumer mapping and boolean operation identification. * **Tests** * Added comprehensive test coverage for autotuner components, insertion points, and region management. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Will Guo <willg@nvidia.com> |
||
|
|
2a467531f2 |
Layerwise KD mode (#802)
## What does this PR do?
**Type of change:** new feature
**Overview:** Add a subclass of `DistillationModel` which implements
slightly different hooks to inject teacher tensors into corresponding
student layers for module replacement purposes, as opposed to logits
distillation.
## Usage
```python
mtd.convert(model, mode=[("layerwise_kd", config)])
```
## Testing
New units
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Introduced bypass-enabled knowledge distillation mode with layer-level
loss mapping for fine-grained model optimization control.
* Added model export functionality with automatic cleanup of
intermediate activation capturing mechanisms.
* **API Changes**
* New bypass_kd mode configuration option available for advanced
knowledge distillation workflows.
* Updated model export interface for improved lifecycle management.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
|
||
|
|
fc6a211b4a |
Added column-major storage of weights and scales in INT4 quantization for model load time improvement in TRT-RTX (#811)
## What does this PR do? **Type of change:** ? New feature **Overview:** TensorRT-RTX requires the weights and scales in the ONNX models to be in column-major format. So whenever the model loads TRT-RTX JIT transposes the weights and scales during load time, causing increased load time. Proposed feature is after quantization, transpose the weights and scales in DQ node and add a transpose node right after i.e, A × B = A × ((Bᵀ)ᵀ) The transformation is post processing step and is disabled by default. It can be enabled by quantizing with --use_column_major ## Usage ``` python -m modelopt.onnx.quantization --onnx_path "model.onnx" --output_path "model_quant.onnx" --quantize_mode int4 --calibration_method awq_lite --use_column_major --skip_shared_constants_duplication ``` ## Testing Tested a few LLM's and their MMLU scores with and without this transformation. No degradations were observed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added --use_column_major CLI flag to enable column-major weight storage optimization (applies to DQ-only quantization paths). * **Documentation** * CLI docs updated to describe the new flag and its applicability. * **Tests** * New unit tests validating column-major transformation behavior and output equivalence. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
02c5f292f0 |
GPTQ Lite implementation (#555)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Adds support for GPTQ algorithm. This PR implements a modified version of the official GPTQ algorithm; the key difference is that updated activations from each layer are not used for hessian computation ## Usage <!-- You can potentially add a usage example below. --> Modify "algorithm" field in quant_cfg to "gptq_lite". Note: Does not currently work with AWQ ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> - [x] Added unit tests to test helper functions + e2e flow - [x] Perplexity and GPQA results | Model | Qformat | Perplexity wikitext2 | GPQA | |-------------|------------------------|------------------------|------| | Qwen3-8B | INT4 weight only (modelopt + amax/7) no GPTQ | 10.75 | n/a | | Qwen3-8B | INT4 weight only (modelopt + amax/7) | **10.56** | 0.388 | | Qwen3-8B | INT4 weight only + FP-Quant hessians + amax/7.5 | 10.25 | **0.449** | | Qwen3-8B | INT4 weight only (FP-Quant) | 10.24 | 0.46 | | Qwen3-8B | NVFP4 static weight only | 10.25 | n/a | | Qwen3-8B | NVFP4 static weight only no GPTQ | 10.25 | n/a | | Qwen3-0.6B | NVFP4 static weight only | **22.75** | n/a | | Qwen3-0.6B | NVFP4 dynamic weight only | 23.50 | n/a | | Qwen3-0.6B | NVFP4 static weight only with FP-Quant hessians | 22.0 | n/a | | Qwen3-0.6B | NVFP4 static weight only no GPTQ | 24.25 | n/a | Conclusions from results - Perplexity remains the same or shows improvement with Modelopt implementation. The magnitude of improvement is lesser in modelopt when compared to FP-Quant - GPQA shows no improvement with modelopt, but shows improvement with FP-Quant ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * GPTQ Lite quantization mode now available for efficient model calibration * GPU memory usage monitoring utility added * Quantization configuration extended to support complex nested structures and lists * **Tests** * Comprehensive test coverage added for GPTQ quantization workflows <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> |
||
|
|
81b67ddf06 |
Context parallelism for Megatron core models (#818)
## What does this PR do? New feature **Overview:** This PR implements the context manager which injects attn_mask as attn_bias to TEDotProductAttention so that we can enable EAGLE training with arbitrary mask. ## Usage set CP>1 in https://github.com/NVIDIA/Megatron-LM/blob/main/examples/post_training/modelopt/finetune.sh ```python # Add a code snippet demonstrating how to use this ``` ## Testing Tested on DSR1 Llama 8B. CP1->CP2 38854MB->28050MB MTbench AL 2.26->2.31 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Context parallelism support added for Eagle speculative decoding with HuggingFace and Megatron Core models. * Model checkpoint loading enhanced to enable remote code execution capabilities when required. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
770962b32c |
Rename MLM teacher arg (#829)
## What does this PR do? **Type of change:** Refactor **Overview:** MLM arg changed from `--teacher-model-config` to `--export-kd-teacher-model-config` for consistency ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated export flag naming for knowledge distillation teacher model configuration. * Adjusted default top-k parameter for logits selection to 1024. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
58abdc2f38 |
Support MLA nvfp4 quant for Deepseek for max perf (#582)
## What does this PR do? **Type of change:** new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** support for newer checkpoints ## Usage <!-- You can potentially add a usage example below. --> ```python torchrun --nproc-per-node=8 ptq.py --mla_quant nvfp4_wq_a_wkv_a_wq_b_wo_fp8_wkv_b --batch_size 4 --model_path $DS_CKPT --config DeepSeek-V3/inference/configs/config_671B.json --quant_cfg NVFP4_DEFAULT_CFG --output_path $AMAX_PATH ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added NVFP4 quantization option for MLA model quantization workflow. * Expanded quantization configuration choices to include "nvfp4" alongside existing per_tensor_fp8 option. * Introduced new CLI parameter to specify MLA quantization type during post-training quantization. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: binghanc <176802681+binghanc@users.noreply.github.com> |
||
|
|
9857e0a48c |
Nenotrom Nano PTQ fix where MoELayer forward has additional named arguments (#823)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug fix **Overview:** ? Our `QuantMoELayer` dynamic module override the forward. In the latest version of `megatron.core` `MoELayer.forward` has additional argument ; hence resulting in failure. Since we are passing through all the arguments, here we change to use *args and **kwargs to avoid future issue. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Updated argument propagation in the mixture-of-experts quantization layer to flexibly support additional calibration-related and padding parameters while maintaining backward compatibility. * Improved routing behavior with conditional top-k adjustments that activate when group routing configurations are in use. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
4227bb7366 |
Add support for MXFP8 PTQ (#736)
## What does this PR do?
**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:** Add support for MXFP8 PTQ, enabling MXFP8 hardware
acceleration during inference on Blackwell GPUs.
## Usage
<!-- You can potentially add a usage example below. -->
```bash
export MODEL_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B
export OUTPUT_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B-MXFP8
mkdir -p $OUTPUT_PATH
python examples/llm_ptq/hf_ptq.py \
--export_fmt hf \
--dataset cnn_dailymail \
--pyt_ckpt_path $MODEL_PATH \
--export_path $OUTPUT_PATH \
--qformat mxfp8
```
The `hf_quant_config.json` of the output checkpoint:
```json
{
"producer": {
"name": "modelopt",
"version": "0.41.0.dev50+g7a796a875"
},
"quantization": {
"quant_algo": "MXFP8",
"kv_cache_quant_algo": "FP8",
"group_size": 32,
"exclude_modules": [
"lm_head"
]
}
}
```
And `config.json` (only the `quantization_config`):
```json
...
"quantization_config": {
"ignore": [
"lm_head"
],
"quant_algo": "MXFP8",
"kv_cache_scheme": {
"dynamic": false,
"num_bits": 8,
"type": "float"
},
"producer": {
"name": "modelopt",
"version": "0.41.0.dev50+g7a796a875"
},
"quant_method": "modelopt"
}
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
Used `hf_ptq.py` to quantize the model `nvidia/OpenMath2-Llama3.1-8B`
([available in
hugging-face](https://huggingface.co/nvidia/OpenMath2-Llama3.1-8B)), see
the example command above.
Checked that the generated MXFP8 checkpoint can be loaded with vLLM
(required changes in vLLM, not merged to main).
Added tests for `MXFP8QTensor` in
`tests/gpu/torch/quantization/test_qtensor_cuda.py`.
Added "mxfp8" in `tests/examples/llm_ptq/test_llm_ptq.py`
#### Support for Nemotron Models
Verify that Nemotron Nano V3 BF16 can be converted to MXFP8 using
`hf_ptq.py`:
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added MXFP8 quantization format support with new scaling mechanisms
and quantization utilities.
* Updated configuration options, example scripts, and utilities to
recognize and process MXFP8 quantization workflows.
* Extended quantization export pipelines to handle MXFP8 quantized
models.
* **Tests**
* Expanded test coverage for MXFP8 quantization across various tensor
shapes, data types, and device configurations.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
|
||
|
|
4a848c4f52 |
Modelopt-windows documentation update (#812)
## What does this PR do? Documentation **Overview:** - Update support matrix, changelog, deployment page, example readmes as per recent feature and model support on Windows side. ## Testing - No testing, its just documentation change ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ONNX Mixed Precision Weight-only quantization (INT4/INT8) support. * Introduced diffusion-model quantization on Windows. * Added new accuracy benchmarks (Perplexity and KL-Divergence). * Expanded deployment with multiple ONNX Runtime Execution Providers (CUDA, DirectML, TensorRT-RTX). * **Bug Fixes** * Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ quantization. * **Documentation** * Updated installation guides with system requirements and multiple backend options. * Reorganized deployment documentation with comprehensive execution provider guidance. * Expanded example workflows with improved setup instructions and support matrices. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
2c73de0405 |
[5525939] Allow user to select target opset in MOQ (#809)
## What does this PR do? **Type of change:** new feature **Overview:** - Allow user to select the target opset - Minimum opset will be defined according to quantization mode - Add tests in tests/unit/onnx/test_quantize_api.py ## Testing Added unit tests tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[int8] PASSED [ 11%] tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[fp8] PASSED [ 22%] tests/unit/onnx/test_quantize_api.py::test_opset_below_minimum_upgrades_to_minimum[int4] PASSED [ 33%] tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[int8] PASSED [ 44%] tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[fp8] PASSED [ 55%] tests/unit/onnx/test_quantize_api.py::test_opset_below_original_uses_original[int4] PASSED [ 66%] tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[int8] PASSED [ 77%] tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[fp8] PASSED [ 88%] tests/unit/onnx/test_quantize_api.py::test_opset_above_minimum[int4] PASSED [100%] ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - auto update according to argparser help - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information Requested as a WAR for a Windows-onnxruntime issue in 5525939, but regardless, it's a useful feature to have <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--opset` CLI option enabling users to specify target ONNX opset version when quantizing models. * Automatic validation ensures the opset version is compatible with quantization requirements, with warnings when adjustments are made. * **Tests** * Added comprehensive test coverage for opset version handling across quantization workflows. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> Signed-off-by: Gal Hubara-Agam <96368689+galagam@users.noreply.github.com> Co-authored-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com> |
||
|
|
1c7a928df8 |
Change cnn_dailymail to abisee/cnn_dailymail (#819)
cnn_dailymail is changed to abisee/cnn_dailymail long ago. Perhaps older name no longer works <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated dataset configuration path for improved dataset access. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
aafd388394 |
add FP8 sweep option for static NVFP4 MSE (#758)
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature
**Overview:** ?
Adds fp8_scale_sweep mode to MSE calibrator for optimizing FP8-quantized
per-block scales in NVFP4 format.
## Usage
<!-- You can potentially add a usage example below. -->
Tested with this config
```python
NVFP4_WEIGHT_MSE_FP8_SWEEP_CFG = {
"quant_cfg": {
"*weight_quantizer": {
"num_bits": (2, 1),
"block_sizes": {-1: 16, "type": "static", "scale_bits": (4, 3)},
"axis": None,
"enable": True,
},
"*input_quantizer": {
"enable": False,
},
**_default_disabled_quantizer_cfg,
},
"algorithm": {
"method": "mse",
"fp8_scale_sweep": True,
},
}
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
**New Features**
- Added FP8 scale sweep option for quantization calibration, enabling
optimized scale value sweeping for NVFP4 per-block quantization.
- Introduced new NVFP4_WEIGHT_MSE_CFG configuration preset for improved
weight quantization workflows.
**Tests**
- Added test coverage validating FP8 scale sweep functionality and reset
behavior.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
|
||
|
|
38403095c4 |
[5725362] AutoCast Fixes for models with external data (#731)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Fix AutoCast ReferenceRunner to handle large models. Models above 2GB cannot be serialized to string, which is what polygraphy is doing under the hood. Use a temporary file instead to save the modified onnx with all tensors marked as outputs. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Enhanced model processing to better support large ONNX models during validation and runtime execution * Added diagnostic logging of model sizes at key processing stages for improved debugging and performance monitoring <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> |
||
|
|
0ebcd70878 |
Support VLM calibration with image-text data (#755)
## What does this PR do? **Type of change:** New feature **Overview:** The primary goal of this PR is to allow the model optimizer to use image-text pair data during the calibration phase of quantization, which is likely help improve accuracy of quantized VLMs like Nemotron VL on visual understanding tasks particularly, compared to text-only calibration data. - New Feature: Adds support for VLM calibration specifically using image-text data. - Dataset Integration: Introduces support for sampling from the `Nemotron-VLM-Dataset-v2`. - Refactoring: Created a separate utility for VLM datasets to keep the main Hugging Face PTQ script (`hf_ptq.py`) clean. - Simplified logic for handling multimodal inputs. - Addressed specific issues encountered when calibrating the `Nemotron-Nano-VL-12B-V2` model with image data. - Documentation: Updated the README to include instructions and examples for VLM calibration. This PR complements https://github.com/NVIDIA/Model-Optimizer/pull/347 and we will consolidate llm_ptq and vlm_ptq examples in follow-up PRs. ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/scratch.omniml_data_2/models/Nemotron-Nano-VL-12B-V2 --qformat nvfp4 --export_path /home/omniml_data_3/zhiyuc/checkpoints/Nemotron-Nano-VL-12B-V2-NVFP4-doccalib --trust_remote_code --kv_cache_qformat none --calib_with_images --vlm_dataset nemotron_vlm_dataset_v2 --vlm_subsets sparsetables,plotqa_cot --calib_size 512 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Not yet <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Vision-Language Model (VLM) calibration support with image-text pair data, specifically for Nemotron VL models. * Added new `--calib_with_images` CLI flag to enable image-based calibration workflows. * Integrated Nemotron VLM dataset v2 for streaming multimodal calibration data. * **Documentation** * Added VLM calibration guidance in the PTQ README with usage examples and dataset information. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
8c36f5a2f7 |
[5676209][ONNX][Autocast] Add support for single npz file with multiple samples (#815)
## What does this PR do? **Type of change:** New feature **Overview:** Currently, Autocast only supports calibration data with shape matching the model's input. This PR adds support for calibration data with shape that is a multiple of the model's input. It does so by re-arranging the data as such that it contains multiple samples with shape matching the model's input. Simplified example: - ONNX input: `[1, 3, 224, 224]` - Calibration data: `[10, 3, 224, 224]` - Calibration data with multiple samples: `[1, 3, 224, 224] * 10` ## Usage Single `npz` file with multiple samples: ```sh $ python -m modelopt.onnx.autocast --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz ``` ## Testing See bug 5676209. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information Equivalent support is already included in the quantization workflow: https://github.com/NVIDIA/Model-Optimizer/blob/1cc8e6bf3917f61500e81d4ded0af5d5a00e2e25/modelopt/onnx/quantization/calib_utils.py#L50 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * CalibrationDataProvider now accepts both file paths and pre-loaded ONNX models as input. * NPZ calibration files can now provide multiple batches for calibration workflows. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
5cc2a54519 |
Ynankani/update windows benchmark md (#762)
## What does this PR do? **Type of change:** ? documentation **Overview:** Md update to add perplexity and kl divergence benchmark info. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: NA - **Did you write any new necessary tests?**: NA - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded accuracy comparison section with three detailed benchmark metrics: MMLU scores, Perplexity (PPL), and KL-divergence. * Added comprehensive tables showing results across models and quantization configurations. * Included evaluation guides and references for each metric. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: unknown <ynankani@nvidia.com> |
||
|
|
3036a9ea9f |
Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp:
- torch2.8.0
- transformers5.0.0
- accelrate1.12.0
- Move to FSDP2
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
- Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`
## Usage
<!-- You can potentially add a usage example below. -->
```bash
./launch_train.sh --model $MODEL \
--output_dir $OUTPUT_DIR \
--data $DATA \
--num_epochs 0.1 \
--train_bs 1 \
--eagle_config eagle_config.json \
--training_seq_len 1024 \
--cp_size 2 #newly added
```
## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
Absolute diff (adiff) cp_out vs out: 0.001953125
Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```
- E2E Training Acc
(Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>
- Peak Mem Reserved
(llama3.1-8B, 8xH100, train_length=4k)
| cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
|----|--------------------------|--------------------------|
| 1 | 65040.20 |79018.00
| 2 | 50409.17 |73098.00
| 4 | 45120.92 |72052.00
| 8 | 38882.12 |66484.00
- Max Training Length test
(llama3.1-8B, H100)
| cp_size | 6k | 12k | 24k | 48k |
|--------------------|-----|-----|-----|-----|
| 1 | ✅ | OOM | OOM | OOM |
|2 | ✅ | ✅ | OOM | OOM |
| 4 | ✅ | ✅ | ✅ | OOM |
| 8 | ✅ | ✅ | ✅ | ✅ |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.
* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.
* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
* Added FSDP configuration file for distributed training setup.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
04165ace59 |
[Minor] Force 'fuse_wgrad_accumulation' to false for TE GroupedLinear (#814)
## What does this PR do? **Type of change:** ? Minor **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Automatically disables fuse_wgrad_accumulation when using ModelOpt quantization with Transformer Engine-based quantization paths. A warning is now displayed to notify users when this adjustment occurs. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: realAsma <akuriparambi@nvidia.com> |
||
|
|
044c4bc4eb |
Support megatron generate for vlm (#773)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? This PR adds feature of VLM generation for megatron_generate ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Vision Language Model support to text generation pipeline, enabling simultaneous processing of image and text inputs during both generation and prefill operations. * **Improvements** * Enhanced data flow to properly route multimodal inputs (images and text tokens) through generation paths with automatic detection and handling of vision-enabled model architectures. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: James Shen <yueshen@nvidia.com> |
||
|
|
4f4558adbb |
Fix a nvfp4 weight amax attribute issue during export (#785)
## What does this PR do? **Type of change:** Bugfix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Fix a nvfp4 weight amax attribute issue during export, especially when calibration size is small. Context: https://github.com/sgl-project/sglang/issues/14677#issuecomment-3712750444 ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/scratch.jingyux_coreai/kimi-k2/models/Kimi-K2-Thinking-BF16 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/Kimi-K2-Thinking-NVFP4 --kv_cache_qformat none --calib_size 20 --trust_remote_code --dataset cnn_dailymail ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Bug Fixes * Improved weight quantizer calibration to ensure quantizers are properly initialized with calibration statistics before computing scaling factors. * Enhanced reliability and consistency of quantized model exports. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
2a08622059 |
Fix moe amax remedy for dsr1 and remove global barrier in quantization megatron plugins (#808)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Bug **Overview:** ? This PR fix 2 bugs which impact DeepSeek calibration as well as PP forward of MoE models. 1. The WAR in `MoELayer` that change the `topk` to `num_experts` only works if no group-topk (a.k.a group routing) is used. Only changing topk will lead to out-of-range error since `topk` can never be `num_experts` when `group_topk != None`. Currently only `DeepSeek-V3` uses `group_topk` and DeepSeek-V3 does not have difficulty to calibrate all experts. As a result, we disable the WAR when detecting `group_topk`. 2. A previous PR inserted global barrier in `quantization.plugin.megatron` https://github.com/NVIDIA/Model-Optimizer/commit/6ef9954db1e73b8c4a86e5bfd31c954cfa21db61#diff-0fa2ba4ecc36c5ff031be9f9a5af080e7aa3afa331c438f02f501b9432ec6d6aL228-R515 This leads to dead lock when using PP since PP rank will never be able to sync during pipeline forward. For MoE, this can be even worse if the barrier is only visited by some EP/PP rank. Using collective communication over the global world (a.k.a global comm) in megatron plugin should be prohibited. Using collective on sub communication group should avoid using `megatron.core.parallel_state` (a.k.a `mpu`) in the future. Instead, use the local `pg_collection` from each module. Any usage of collective communication must be inspected carefully with test as PP, TP, and EP. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> |
||
|
|
b44c60ad34 |
Svdquant huggingface checkpoint export support (#754)
## What does this PR do?
**Type of change:** new feature
**Overview:**
## Usage
```bash
cd ./examples/llm_ptq/
python hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--export_path /home/scratch.shiychen_coreai/quantized_models/Qwen3-4B-svdq \
--qformat nvfp4_awq_svdquant --kv_cache_qformat none --sparsity_fmt dense --calib_size 8
```
## Testing
exported checkpoint and loaded.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added nvfp4_svdquant as a new quantization format option for LLM model
quantization workflows.
* **Limitations**
* Multi-GPU export configurations using tensor or pipeline parallelism
are not supported with nvfp4_svdquant quantization.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
945ee02f8d |
[1/3] Add the fastvideo support (#804)
## What does this PR do? **Type of change:** new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** FastVideo is a new diffusion-focused framework that we plan to integrate with. In this work, we added initial support for WAN 2.2 5B in FastVideo, targeting the text-to-video use case. For the Conv layer type, we currently use a straightforward direct convolution call. Implicit GEMM quantization is intentionally omitted in this first MR and will be addressed in a follow-up MR. - [x] [1/3] Added support for the WAN 2.2 DIT + VAE layer type. - [ ] [2/3] Added calibration support for them in the example script, add test cases and README, doc. - [ ] [3/3] Submitted an MR to fastvideo to enable quantization-aware training. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added FastVideo plugin support to the quantization framework. Users can now apply quantization to FastVideo-specific layers with specialized weight quantization handling, optimized input processing, and caching features for enhanced inference performance. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
668b8a19e8 |
[1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This PR adds support for exporting quantized diffusers models (DiT, Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM. **Changes** New file: `diffusers_utils.py` - Dummy input generation for various diffusion models - Pipeline component extraction helpers - QKV projection detection and grouping - `hide_quantizers_from_state_dict()` context manager for clean saves Refactored: `unified_export_hf.py` - New `_fuse_qkv_linears_diffusion()` for QKV amax fusion - `_export_diffusers_checkpoint()` to export full pipelines (models + tokenizers + schedulers etc.) Plans - [x] [1/3] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml - [ ] [3/3] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ``` mtq.quantize(pipe, quant_config, forward_call) export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir) ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added HuggingFace checkpoint export support for quantized diffusion models with configurable output directory * Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint export destination * Extended export functionality to support selective component exports from diffusion pipelines * Enhanced quantized model export with improved component handling and multi-stage checkpoint generation <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
563a1e09c6 |
Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?
**Type of change:** New feature
- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
- Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.
## Usage
<!-- You can potentially add a usage example below. -->
Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
model,
mode="mcore_minitron",
constraints=constraints,
dummy_input=None, # Not used
config=config,
)
```
1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```
2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
return megatron_mmlu(m, tokenizer, percentage=0.05) # 5% sampled data for faster eval
config["score_func"] = score_func
# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune
mtp.prune(...)
```
To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
hidden_size_divisor=256,
ffn_hidden_size_divisor=512,
mamba_head_dim_divisor=8,
num_moe_experts_divisor=8,
num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```
## Testing
**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```
**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads': 88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads': 96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads': 80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```
## Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
OMNIML-3043
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
* Introduced parameter counting capabilities for architecture search
* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
* Updated configuration documentation with granular divisor parameters
* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
792806f00e |
Add Security considerations in docs (#803)
Add security considerations to docs suggested by Nvidia Security team <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * Added comprehensive security considerations documentation for ModelOpt covering multiple risk areas including untrusted input handling, deserialization safety, input validation, resource exhaustion prevention, data protection for transit and storage, logging and observability practices, supply chain security, and structured mitigation approaches with practical implementation examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
67576d20ba |
Revert onnxruntime-gpu version to 1.22.0 for Windows (#801)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** Reverted the windows ort version in setup.py to 1.22, since I later observed some regressions with vision models. Keeping ORT version as 1.23 in windows examples since i didn't face any issues with LLM's. ## Testing Attaching results observed while testing these models. [int8_trt_comparison.csv](https://github.com/user-attachments/files/24763088/int8_trt_comparison.csv) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated GPU runtime dependencies for the optional ONNX package. Simplified platform-specific version constraints while maintaining compatibility across supported platforms. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
615f99e746 |
Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support KIMI K2 Thinking PTQ from the original int4 checkpoint. Tested with transformers 4.57.1, compressed-tensors 0.12.0 The model weights are dequantized on the fly to save GPU memory ## Usage scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant nvfp4_mlp_only --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support for nvfp4_mlp_only quantization format, enabling new layer-wise quantization options * Quantization support for CompressedLinear layers in quantized models * **Improvements** * Enhanced quantization for DeepSeek models with improved attention configuration handling * Optimized model loading with automatic precision configuration and weight unpacking * Better memory management during model export with automatic cache cleanup * Conditional sample generation output controlled via verbose mode <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
21a4010348 |
Add Quantizers for Qwen3VLMoeTextDecoderLayer (#666)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? huggingface transformers library implements Qwen3VL Moe layer as a monolithic module, instead of assembling it using Linear layers, which cannot be recognized by modelopt's quantizer now. This PR introduces a conversion from hf's qwen3vl_moe MoE layers to qewn3_moe MoE layers which consist of a set of Linear layers. ## Testing Tested with ```python python hf_ptq.py --pyt_ckpt_path=Qwen/Qwen3-VL-30B-A3B-Instruct --qformat=nvfp4 --dataset wikipedia ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added quantization support for Qwen3VL models with sparse mixture-of-experts (MoE) architecture, enabling efficient model compression for this model type. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Qidong Su <qidongs@nvidia.com> Signed-off-by: Qidong Su <soodoshll@gmail.com> Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> Co-authored-by: Zhiyu <bestczy317@gmail.com> |
||
|
|
b0e7d9fd96 |
Define kv cache scaling factor as amax / 448 (#790)
## What does this PR do? **Overview:** ? Unified the FP8 and NVFP4 kv cache scaling factor definition so the same checkpoint can be used for both FP8 and NVFP4 kv cache quantization deployment ## Testing Unit test ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **Refactor** * Fixed KV cache maximum bound to 448 for FP8 and NVFP4 quantization, simplifying configuration logic. * **Chores** * Removed internal constants from public exports. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
1cc8e6bf39 |
[5676209] Fix duplicated calib data (#794)
## What does this PR do?
**Type of change:** Bug fix
**Overview:** This PR fixes an issue with calibration data with multiple
samples. Previously, calibration data with multiple samples was
generating a data loader with the same sample copied X times instead of
generating data with X different samples.
## Usage
```python
$ python -m modelopt.onnx --onnx_path=$MODEL_NAME.onnx --calibration_data=calib_data_10.npz
```
## Testing
Use calibration data from 5676209 and observe the output of
`calibration_data_reader` in `quantize.py`:
```python
calibration_data_reader = CalibrationDataProvider(
onnx_path, calibration_data, calibration_shapes
)
```
Each calibration sample in the list should be different.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Fixed calibration data generation in ONNX workflow to properly handle
multiple samples during processing.
* **Documentation**
* Updated changelog with version 0.42 entry documenting bug fixes and
new features.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
|
||
|
|
391f6cb004 |
[5750013][5591945][5360813]: AutoCast standalone implementation for type inference (#719)
## What does this PR do? **Type of change:** New feature **Overview:** AutoCast runs full type inference to get the new types after adding casts. ONNX doesn't have a separate function for type inference, and it is done as part of shape inference. Shape inference is a much more complex task than type inference, especially when dynamic shapes are involved. We're seeing some shape inference related bugs in AutoCast. Typically we can WAR, but it's cumbersome. A local implementation might allow users to WAR shape inference related issues. This is opt-in and marked as experimental. ## Usage python -m modelopt.onnx.autocast --onnx_path /path/to/input.onnx [options] --use_standalone_type_inference ## Testing Added use_standalone_type_inference=True to all existing PrecisionConverter tests. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information A more permanent fix would be to decouple type and shape inference in ONNX, we should invest in that when we have the resources - see https://github.com/onnx/onnx/issues/7100 . This is a quick fix, which is also why it is opt-in and not the default mode. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--use_standalone_type_inference` flag to ONNX AutoCast, enabling type-only inference as an alternative to standard shape inference. Useful as a workaround when shape inference fails or to reduce computational overhead. * **Documentation** * Added "Type Inference Control" section with usage examples and caveats for the new standalone type inference option. * **Tests** * Extended test coverage to validate both standard and standalone type inference paths. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com> |