mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
03a1899dda024fa9425118d0711e4ac2b53e0fe2
196
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
03a1899dda |
Support force tokens to % of total experts during calibration (#910)
## What does this PR do?
**Type of change:** New feature
**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.
**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`
## Usage
Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5
Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq
quant_cfg = {
"quant_cfg": { ... },
"algorithm": {
"method": "awq_lite",
"moe_calib_experts_ratio": 0.25, # calibrate 1/4 of experts
},
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)
## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
|
||
|
|
2802302bf9 |
SpecDec Bench: February Update (#875)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Addition of SpecBench Dataset Addition of NVIDID SPEED-Bench dataset, preproc scripts, and custom metrics aggregator Addition of example of converting SpecBench Medusa to this FW Addition of Initial TRTLLM AutoDeploy Specdec support Updates to all frameworks for better perf (overlap/async scheduling etc) ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added SPEED-Bench dataset support with configurable throughput and qualitative configurations * Introduced SpecBench metrics with acceptance rate analysis and visualizations * Added progress bar during benchmark execution * New model implementations for auto-deployment and Medusa-style speculative decoding * Data preparation utility for benchmark datasets * Enhanced metrics with per-category analysis and performance charts * **Documentation** * Updated README with SPEED-Bench workflow and examples * New porting guide for integrating custom benchmark runners * **Refactor** * Streamlined model and runner interfaces for improved flexibility * Consolidated dataset implementations and removed deprecated base classes * **Chores** * Added required dependencies for data handling and visualizations <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Izzy Putterman <iputterman@nvidia.com> |
||
|
|
c689ea1176 |
flush print megatron tokenization stats and update readme (#927)
## What does this PR do? When running the script, I often see the print stats for tokenization (every `log_interval`) not showing up or showing up very very delayed. Hence using `print(..., flush=True)` to fix this. Also update README that the example shown for tokenization takes too long to run, split into multiple .jsonl files for efficiently running the tokenization; and try out a smaller dataset first to test the script ## Testing <!-- Mention how have you tested your change if applicable. --> Split Nemotron-pretraining-SFT-v1 dataset into multiple .jsonl splits and then tokenize them parallelly in different slurm jobs. Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
f78385e70e |
Improve megatron dataset preprocessing script and update docs (#918)
## What does this PR do?
Improve megatron dataset preprocessing script and update docs
## Usage
<!-- You can potentially add a usage example below. -->
```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
--hf_dataset nvidia/Nemotron-Pretraining-SFT-v1 \
--hf_name Nemotron-SFT-General \
--hf_split train \
--hf_max_samples_per_split 10_000_000 \
--json_keys text \
--tokenizer Qwen/Qwen3-0.6B \
--output_dir /path/to/tokenized/data/qwen3 \
--workers 32 \
--max_sequence_length 256_000
```
```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
--jsonl_paths /path/to/data1.jsonl /path/to/data2.jsonl ... \
--json_keys text \
--tokenizer Qwen/Qwen3-0.6B \
--output_dir /path/to/tokenized/data/qwen3 \
--workers 32 \
--max_sequence_length 256_000
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- Downloaded and tokenized Nemotron-Pretraining-SFT-v1 with
Nemotron-Nano-v2 tokenizer
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Updated data preparation guides with new CLI patterns and Hugging Face
Hub integration instructions.
* **New Features**
* Added batch tokenization via directory input and direct Hugging Face
dataset downloads with flexible subset/split filtering.
* **Configuration Updates**
* Optimized distillation settings: adjusted optimizer parameters and
increased checkpoint retention.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
9975ba1065 |
Fix DeepSeek PTQ script (#912)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? Fix two bugs in the PTQ script ## Testing Run DeepseekV3.2 PTQ and export <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Enhanced data type handling in quantization examples for bf16 operations * Updated internal dependencies for quantization utilities to improve modularity <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
ac7c985d96 |
[NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?
**Type of change:** New feature, new tests
**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.
Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints
## Usage
Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:
import modelopt.torch.quantization as mtq
# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")
mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)
# During export, an .moe.html report with per-expert token counts is
saved automatically
## Testing
unittest, also test exporting qwen MOE
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.
* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
|
||
|
|
eb99488da1 |
Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Patch transformers 5.x parameter loading to preserve original `requires_grad` settings. In transformers v5.x, loading a checkpoint forcibly sets parameters' requires_grad, which unintentionally unfreeze frozen parameters (e.g. Base model in eagle training). This leads to optimizer initialization error since the restored optimizer expected more parameter than the checkpoint. This monkey-patch restores the original`requires_grad` after loading parameters. Reference: https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640 ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed model parameter loading in speculative decoding to properly preserve gradient requirements for each parameter when using HuggingFace Transformers 5.x, ensuring correct behavior during checkpoint resumption and model initialization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
b8a4586702 |
Refactor: Eagle data loading (#668)
## What does this PR do? **Type of change:** Refactor <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955 Main changes : - Consolidate Eagle data loading with @ChenhanYu 's implementation of `transformers_dataset.py` - Refactor: baked the following logics from `example/main.py` to `modelopt/torch` for cleaner example entrance: - default config selecting and merging with custom config - tokenizer post-processor (chat template and pad_tok_id) - d2t loading - Implementation refactor: In HF workflow, reuse base modfel's input hidden states as input_embedding, instead of calculating from input_ids. This has two main benefits: - Easier VLM support, which has various embedding processing logics. - Training effieicy. - Deprecating eagle1 from the example. It is still available by setting custom config. - Other minor fixes and readme updates. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> Tested that training curves after changes (both online&offline) is identical with original branch: <img width="1073" height="634" alt="image" src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d" /> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added draft vocabulary cache support for EAGLE model training, enabling runtime vocabulary customization via `--draft_vocab_cache` parameter * Introduced new data loading utilities with sharding, streaming, and tokenization support for large-scale training * Added optional `--log_steps` configuration to training launcher * **Documentation** * Updated EAGLE configuration guides with draft vocabulary cache setup instructions and examples * **Refactor** * Restructured data pipeline for offline training with improved dataset handling and batching * Updated command-line arguments across training scripts (`--input-data` replaces `--input-file`) <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
9e38041d34 |
[OMNIML-2850] [3/n] Adds sparse attention calibration (#538)
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature
**Overview:** ?
- This PR adds the sparse attention calibration algorithm
- Chunked prefill to support long ctx_len
- Separated calibration for prefill and decode
## Usage
<!-- You can potentially add a usage example below. -->
```python
import modelopt.torch.sparsity.attention_sparsity as mtsa
# Apply sparse attention with calibration
model = mtsa.sparsify(model, config=SKIP_SOFTMAX_CALIB)
# Print summary - now shows actual thresholds
mtsa.print_sparse_attention_summary(model)
# Output:
# Method: flash_skip_softmax, Threshold: Dynamic (λ=437.395926)
# Or llm_eval integration
# HuggingFace sparse attention example
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--sparse_attn skip_softmax_calib
```
# The calibration method
## Calibration Algorithm
- Implemented the Inverse Power model: scale_factor = k / (1 -
sparsity)^p
- Fit model parameters (k, p) per phase using scipy.optimize.curve_fit
- At inference: threshold = k / (1 - target_sparsity)^p / seqlen
## Why Choosing the Inverse Power model?
The inverse power model better fits the relationship between sparsity
ratio and threshold_scale_factor.
<img width="2388" height="1082" alt="sparsity_model_analysis"
src="https://github.com/user-attachments/assets/4dfb45d4-8c16-4f15-a878-c8e08a9b6128"
/>
## Runtime Flexibility
- Target sparsity can be changed at inference time without recalibration
- Users can adjust module._sparse_method_instance.target_sparse_ratio
dynamically
- Threshold automatically adapts to sequence length
## Testing
<!-- Mention how have you tested your change if applicable. -->
The calibration results for `Qwen/Qwen3-30B-A3B-Thinking-2507` are shown
below and are mostly consistent with the ground-truth numbers collected
from the kernel side.
```
Prefill Calibration Results:
Model: scale_factor = k / (1 - sparsity)^p
Fitted k: 1003.3990
Fitted p: 1.2589
R-squared: 0.827549
Scale factors for different target sparsities:
Target Scale Factor
---------- ---------------
50% 2401.35
70% 4568.26
80% 7610.98
90% 18214.70
95% 43591.65
```
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
|
||
|
|
3801923e9d |
Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do? **Type of change:** ? new feature **Overview:** ? Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ. ## Usage scripts/huggingface_example.sh --model <minimax checkpoint> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added MiniMax M2.1 model quantization support with nvfp4 format. * Extended FP8 quantization capabilities with configurable dtype parameter for enhanced precision control. * **Improvements** * Enhanced detection of quantized linear module variants. * Improved weight unpacking for FP8-based linear modules. * **Documentation** * Updated supported models table to include MiniMax M2.1. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com> |
||
|
|
ca1f9687bd |
[OMNIML-3505] LTX-2 Distillation Trainer (#892)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
Adding LTX-2 distillation trainer.
## Usage
<!-- You can potentially add a usage example below. -->
```bash
accelerate launch \
--config_file configs/accelerate/fsdp.yaml \
--num_processes 8 \
distillation_trainer.py --config configs/distillation_example.yaml
```
See readme for more details.
## Testing
Run training with single/multiple nodes.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## New Features
* Added distillation training support for LTX-2 models with quantization
integration.
* Introduced comprehensive documentation and example configurations for
distillation workflows.
* Includes multi-GPU and multi-node training setup with distributed
training support and customizable configuration templates.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Meng Xin <mxin@nvidia.com>
|
||
|
|
b43ae641a4 |
MBridge pruning minor fix for saving pruned NemotronH (#887)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> ## Testing <!-- Mention how have you tested your change if applicable. --> Nemotron Nano v2 pruned can be saved <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed Hugging Face model loading to properly respect the `trust_remote_code` parameter during model instantiation. * **Improvements** * Enhanced distributed training logging with rank-0 aware warning and logging mechanisms for cleaner, non-redundant output in multi-GPU and multi-node scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
95511a0bc9 |
Add Nemotron parse PTQ support (#786)
## What does this PR do? **Type of change:** New model support <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Add PTQ support for https://huggingface.co/nvidia/NVIDIA-Nemotron-Parse-v1.1 ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/NVIDIA-Nemotron-Parse-v1.1 --qformat fp8 --export_path /home/omniml_data_3/zhiyuc/checkpoints/NVIDIA-Nemotron-Parse-v1.1-FP8 --trust_remote_code --kv_cache_qformat none --attn_implementation eager ``` By default, image-text data will be used in calibration for VLMs. ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Not yet <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added support for Nemotron-Parse multimodal models, including proper device mapping, processor loading, and generation handling. * **Improvements** * Enhanced quantization robustness with safer handling of quantization attributes and fallback logic. * Improved model loading with better device placement and encoder buffer management for vision-language models. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
8d67cb08d5 |
Add Megatron-Bridge recipe-free distillation example script (#861)
## What does this PR do?
**Type of change:** New example script <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->
- [x] M-Bridge recipe-free distillation script so its more easier to run
and can support pruned models
- [x] Fix resuming distillation run
## Usage
<!-- You can potentially add a usage example below. -->
```python
torchrun --nproc_per_node 8 distill.py \
--teacher_hf_path Qwen/Qwen3-8B \
--student_hf_path Qwen3-8B-NAS-Pruned-6B \
--tp_size 8 \
--data_paths <climbmix 25% tokenized (~90B tokens)> \
--data_path_to_cache /path/to/cache/climbmix_dataset_indices_qwen3 \
--seq_length 4096 \
--mbs 8 \
--gbs 768 \
--train_iters 28500 \
--lr 1e-4 \
--min_lr 1e-5 \
--lr_warmup_iters 100 \
--eval_interval 500 \
--eval_iters 32 \
--log_interval 10 \
--output_dir qwen3_8b_6b_mbridge_distill
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- [x] Re-ran Qwen3 8B -> 6B experiments and compare with Nemo2 results
from blog
Best subnet from NAS: `{'num_layers': 30, 'hidden_size': 3584,
'ffn_hidden_size': 11776} -> 5.99B params, 0.5718 score`
| Model | MMLU | GSM8K - flexible, strict | MBPP (coding) |
| ------- | ------ | ------- | ------- |
| Qwen3-8B | 74.9 | 87.5, 84.6 | 65.4 |
| Qwen3-8B-Pruned-6B | 57.6 | 11.6, 10.0 | 4.8 |
| Qwen3-8B-Pruned-6B (Distilled for 16k steps i.e. 50B tokens ~3k GPU
hours) | 71.6 | 78.0, 64.7 | 43.4 |
| Qwen3-8B-Pruned-6B (Distilled for 28.5k steps i.e. 90B tokens ~5.2k
GPU hours) | 71.9 | 78.1, 64.8 | 44.2 |
| Qwen3-4B | 70.0 | 81.1, 84.7 | 62.8 |
Previous Nemo2 experiments on depth pruned Qwen3 8B -> 6B (24 layers)
had MMLU ~72.0 so more or less similar. No hparam tuning done for
current M-Bridge distillation run
- [ ] (Separate PR) GitHub CI/CD test for example script with NeMo 26.02
container
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: N/A
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added complete distillation workflow and example for Megatron-Bridge
optimization.
* **Documentation**
* Enhanced setup guide with Docker workflows, data preparation steps,
and detailed distillation instructions.
* Improved usage documentation and help references.
* **Improvements**
* Better data preprocessing output with human-readable formatting for
metrics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
eb3e6edb50 |
[fix][5875912] Fix autoquant-autodeploy example (#878)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Please check Bug ticket ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing Tested with ``` ./scripts/run_auto_quant_and_deploy.sh --hf_ckpt ./models/Qwen/Qwen3-8B --save_quantized_ckpt ./qwen3_8B_autoquant --quant fp8 --effective_bits 10.0 ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Simplified LLM initialization by removing intermediate configuration layer * Updated attention backend from triton to flashinfer <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
10efcb65f6 |
[3.1/4] Diffusion Quantized ckpt export - WAN 2.2 14B (#855)
## What does this PR do?
**Type of change:** documentation <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->
**Overview:**
1. Added multi‑backbone support for quantization: --backbone now accepts
space- or comma-separated lists and resolves to a list of backbone
modules.
2. Introduced PipelineManager.iter_backbones() to iterate named backbone
modules and updated get_backbone() to return a single module or a
ModuleList for multi‑backbone.
3. Updated ExportManager to save/restore per‑backbone checkpoints when a
directory is provided, with {backbone_name}.pt files, and to create
target directories when missing.
4. Simplified save_checkpoint() calls to rely on the registered
pipeline_manager by default.
**Usage: **
```bash
python quantize.py --model wan2.2-t2v-14b --format fp4 --batch-size 1 --calib-size 32 \
--n-steps 30 --backbone transformer transformer_2 --model-dtype BFloat16 \
--quantized-torch-ckpt-save-path ./wan22_mo_ckpts \
--hf-ckpt-dir ./wan2.2-t2v-14b
```
Plans
- [x] [1/4] Add the basic functionalities to support limited image
models with NVFP4 + FP8, with some refactoring on the previous LLM code
and the diffusers example. PIC: @jingyu-ml
- [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml
- [x] [3/4] Add test cases, refactor on the doc, and all related README.
PIC: @jingyu-ml
- [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**:No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Unified Hugging Face export support for diffusers pipelines and
components
* LTX-2 and Wan2.2 (T2V) support in diffusers quantization workflow
* Comprehensive ONNX export and TensorRT engine build documentation for
diffusion models
* **Documentation**
* Updated to clarify support for both transformers and diffusers models
in unified export API
* Expanded diffusers examples with LoRA fusion guidance and additional
model options (Flux, SD3, SDXL variants)
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
5e43b2a5f5 |
Support Qwen3 Next MTP load and export (#860)
## What does this PR do? Fix MTP export for Qwen3 Next **Overview:** ? For Qwen3 next, the MTP weights are not stored separately in safetensors. So we use "mtp" weights key to decide if the weights are for MTP or not. ## Testing Qwen3 Next PTQ and check if MTP is in the exported checkpoint. scripts/huggingface_example.sh --model <Qwen3-Next-80B-A3B-Instruct/Thinking> --quant nvfp4 --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * Optimized Multi-Token Prediction weight loading with improved layer detection and handling. * **Chores** * Simplified status reporting to display total loaded weights and detected layers. * Removed verbose per-file warnings for cleaner console output. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Zhiyu <zhiyuc@nvidia.com> |
||
|
|
a8f5314c93 |
fix the path change in torch v2.10 for spec dec (#863)
## What does this PR do? **Type of change:** bug fix **Overview:** torch v2.10 changes the path for _SDPAMerger. will need to use the new path for import ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated internal import references to reflect organizational changes in dependencies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
452c5a09b0 |
GLM-4.7 MTP support (#792)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Enable GLM-4.7 PTQ workflow, including loading the standalone MTP modules and export as-is. ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/GLM-4.7 --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/GLM-4.7-NVFP4-0203 --trust_remote_code ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added quantization support for GLM-4.7 model with automatic handling of specialized layer architecture. * Added image-text data calibration capabilities for Nemotron VL model quantization. * **Documentation** * Updated support matrix to reflect newly supported models and quantization features. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
944dd1a284 |
Move parallel_state init and warnings to Quant DynamicModule + MBridge pruning doc update (#854)
## What does this PR do? **Type of change:** Minor improvement <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> Only quantization DynamicModules use the parallel_state attribute so for all other model opt methods, we see a parallel state not initialized warning which could be confusing hence moving it to QuantModule class instead Minor update to MBridge pruning docs ## Testing <!-- Mention how have you tested your change if applicable. --> N/A ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Distributed parallel state support is now available in quantization workflows for multi-GPU training. * **Bug Fixes** * Improved resource cleanup in distributed training to ensure proper environment finalization. * **Documentation** * Updated example paths and added new manual pruning configuration examples. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
2e43c80609 |
[2/4] Diffusion Quantized ckpt export (#810)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This MR adds HuggingFace checkpoint export support for LTX‑2 by treating TI2VidTwoStagesPipeline as a diffusion-like pipeline, exporting only the stage‑1 transformer (with QKV-fusion-enabled dummy inputs) and falling back to writing model.safetensors when save_pretrained isn’t available. It also preserves the original forward in DynamicModule patching (_forward_pre_dm) so downstream callers can still invoke the pre-patched forward implementation. **Changes** 1. Added the calibration & quantization support of the LTX2, even with FP8 precision. 2. Preserve original forward before `DynamicModule` patching: when patching forward, we now stash the pre-patched implementation in `self._forward_pre_dm` (once) so downstream code can still call the original forward, then re-bind forward to the class implementation. This is needed for the LTX2 FP8 calibration. 3. Added LTX‑2 HF export path: `export_hf_checkpoint()` now also treats ltx_pipelines.ti2vid_two_stages.TI2VidTwoStagesPipeline as a “diffusion-like” object and routes it through _export_diffusers_checkpoint() (import guarded; no hard dependency). 4. Generalized component discovery: introduced get_diffusion_components() (aliasing the old get_diffusers_components) to support non-diffusers pipelines; for LTX‑2 it returns only stage_1_transformer. 5. Enabled QKV fusion for LTX‑2 backbone: added a model-aware dummy forward generator (generate_diffusion_dummy_forward_fn) that builds minimal LTX Modality inputs (including correct timesteps broadcasting) so shared-input hooks can run and fuse QKV when applicable. 6. Export fallback for non-save_pretrained modules: when a component lacks save_pretrained (LTX‑2 transformer), export now writes model.safetensors + minimal config.json instead of pytorch_model.bin. Plans - [x] [1/4] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [x] [2/4] Add support to more video gen models. PIC: @jingyu-ml - [ ] [3/4] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml - [ ] [4/4] Add the final support to ComfyUI. PIC @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ```bash python quantize.py --model ltx-2 --format fp4 --batch-size 64 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=/home/scratch.omniml_data_2/jingyux/models/LTX-2/ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=/home/scratch.omniml_data_2/jingyux/models/LTX-2/gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added LTX-2 video model support with complete quantization and export pipeline integration * Introduced `--extra-param` CLI option for flexible model configuration and parameter passing * Enhanced export capabilities with broader diffusion model compatibility * **Chores** * Changed default model data type from Half to BFloat16 for improved numerical stability <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
fb5923c89e |
Add Megatron-Bridge pruning example scripts (#800)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
Megatron-Bridge pruning example scripts (HF input, HF / Megatron
output). Also defined some utility functions we can reuse for adding
examples for quantization or other optimizations:
- `modelopt.torch.utils.plugins.mbridge.load_mbridge_model_from_hf`:
Load HF to MBridge with ModelOpt spec in desired TP/PP/etc configuration
-
`modelopt.torch.utils.plugins.mbridge.get_hf_mbridge_calibration_loop`:
Create `forward_loop` for calibration on a HF dataset
- Supports all datasets available in
`modelopt.torch.utils.dataset_utils` (`cnn_dailymail`,
`nemotron-post-training-dataset-v2`, etc)
- Support applying chat template for chat-based data
## Usage
<!-- You can potentially add a usage example below. -->
From `nvcr.io/nvidian/nemo:26.02.rc1` container (mount latest code to
`/opt/Megatron-Bridge` and `/opt/Model-Optimizer`)
```python
torchrun --nproc_per_node 2 /opt/Model-Optimizer/examples/megatron_bridge/prune_minitron.py \
--hf_model_name_or_path Qwen/Qwen3-8B \
--prune_target_params 6e9 \
--hparams_to_skip num_attention_heads \
--output_hf_path /tmp/Qwen3-8B-Pruned-6B
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
- [x] Manually ran pruning script in nemo:25.11 container (plus modelopt
and mbridge mounted to latest) for Qwen3-8B and Nemotron-Nano-9B-v2 with
PP=8 and PP=4
- [ ] Added per-PR CI/CD test for example script
Results when pruning Qwen3 8B -> 6B (10 different configurations) with
and without chat template on the dataset samples. Perhaps MMLU is not
the right metric to look at.
| Layers | Hidden Size | FFN Hidden Size | Params | MMLU (Concatenated
messages) | MMLU (Applied Chat Template) |
|--------|-------------|-----------------|--------|------------------|----------------------|
| 34 | 3328 | 11264 | 5.99B | 0.401 | 0.393 |
| 30 | 3584 | 11776 | 5.99B | 0.588 | 0.576 |
| 36 | 3840 | 8192 | 5.98B | 0.507 | 0.518 |
| 36 | 3584 | 9216 | 5.98B | 0.477 | 0.469 |
| 36 | 3072 | 11776 | 5.97B | 0.255 | 0.249 |
| 32 | 3584 | 10752 | 5.96B | 0.554 | 0.549 |
| 28 | 4096 | 10240 | 5.94B | 0.400 | 0.438 |
| 36 | 4096 | 7168 | 5.93B | 0.461 | 0.438 |
| 36 | 3328 | 10240 | 5.92B | 0.362 | 0.359 |
| 34 | 3840 | 8704 | 5.91B | 0.515 | 0.546 |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: ‼️ TODO
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
**New Features**
* Added new Megatron-Bridge pruning example demonstrating Minitron-based
model optimization with advanced pruning configurations.
**Documentation**
* Updated core project documentation to highlight Megatron-Bridge as a
supported optimization framework.
* Added comprehensive example documentation for Megatron-Bridge
workflows including pruning, distillation, and quantization.
* Updated pruning guides with Megatron-Bridge integration examples and
best practices.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
fc6a211b4a |
Added column-major storage of weights and scales in INT4 quantization for model load time improvement in TRT-RTX (#811)
## What does this PR do? **Type of change:** ? New feature **Overview:** TensorRT-RTX requires the weights and scales in the ONNX models to be in column-major format. So whenever the model loads TRT-RTX JIT transposes the weights and scales during load time, causing increased load time. Proposed feature is after quantization, transpose the weights and scales in DQ node and add a transpose node right after i.e, A × B = A × ((Bᵀ)ᵀ) The transformation is post processing step and is disabled by default. It can be enabled by quantizing with --use_column_major ## Usage ``` python -m modelopt.onnx.quantization --onnx_path "model.onnx" --output_path "model_quant.onnx" --quantize_mode int4 --calibration_method awq_lite --use_column_major --skip_shared_constants_duplication ``` ## Testing Tested a few LLM's and their MMLU scores with and without this transformation. No degradations were observed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added --use_column_major CLI flag to enable column-major weight storage optimization (applies to DQ-only quantization paths). * **Documentation** * CLI docs updated to describe the new flag and its applicability. * **Tests** * New unit tests validating column-major transformation behavior and output equivalence. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
81b67ddf06 |
Context parallelism for Megatron core models (#818)
## What does this PR do? New feature **Overview:** This PR implements the context manager which injects attn_mask as attn_bias to TEDotProductAttention so that we can enable EAGLE training with arbitrary mask. ## Usage set CP>1 in https://github.com/NVIDIA/Megatron-LM/blob/main/examples/post_training/modelopt/finetune.sh ```python # Add a code snippet demonstrating how to use this ``` ## Testing Tested on DSR1 Llama 8B. CP1->CP2 38854MB->28050MB MTbench AL 2.26->2.31 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Context parallelism support added for Eagle speculative decoding with HuggingFace and Megatron Core models. * Model checkpoint loading enhanced to enable remote code execution capabilities when required. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
770962b32c |
Rename MLM teacher arg (#829)
## What does this PR do? **Type of change:** Refactor **Overview:** MLM arg changed from `--teacher-model-config` to `--export-kd-teacher-model-config` for consistency ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated export flag naming for knowledge distillation teacher model configuration. * Adjusted default top-k parameter for logits selection to 1024. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
58abdc2f38 |
Support MLA nvfp4 quant for Deepseek for max perf (#582)
## What does this PR do? **Type of change:** new feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** support for newer checkpoints ## Usage <!-- You can potentially add a usage example below. --> ```python torchrun --nproc-per-node=8 ptq.py --mla_quant nvfp4_wq_a_wkv_a_wq_b_wo_fp8_wkv_b --batch_size 4 --model_path $DS_CKPT --config DeepSeek-V3/inference/configs/config_671B.json --quant_cfg NVFP4_DEFAULT_CFG --output_path $AMAX_PATH ``` <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added NVFP4 quantization option for MLA model quantization workflow. * Expanded quantization configuration choices to include "nvfp4" alongside existing per_tensor_fp8 option. * Introduced new CLI parameter to specify MLA quantization type during post-training quantization. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: binghanc <176802681+binghanc@users.noreply.github.com> |
||
|
|
4227bb7366 |
Add support for MXFP8 PTQ (#736)
## What does this PR do?
**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:** Add support for MXFP8 PTQ, enabling MXFP8 hardware
acceleration during inference on Blackwell GPUs.
## Usage
<!-- You can potentially add a usage example below. -->
```bash
export MODEL_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B
export OUTPUT_PATH=/my_home/hf_models/nvidia/OpenMath2-Llama3.1-8B-MXFP8
mkdir -p $OUTPUT_PATH
python examples/llm_ptq/hf_ptq.py \
--export_fmt hf \
--dataset cnn_dailymail \
--pyt_ckpt_path $MODEL_PATH \
--export_path $OUTPUT_PATH \
--qformat mxfp8
```
The `hf_quant_config.json` of the output checkpoint:
```json
{
"producer": {
"name": "modelopt",
"version": "0.41.0.dev50+g7a796a875"
},
"quantization": {
"quant_algo": "MXFP8",
"kv_cache_quant_algo": "FP8",
"group_size": 32,
"exclude_modules": [
"lm_head"
]
}
}
```
And `config.json` (only the `quantization_config`):
```json
...
"quantization_config": {
"ignore": [
"lm_head"
],
"quant_algo": "MXFP8",
"kv_cache_scheme": {
"dynamic": false,
"num_bits": 8,
"type": "float"
},
"producer": {
"name": "modelopt",
"version": "0.41.0.dev50+g7a796a875"
},
"quant_method": "modelopt"
}
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
Used `hf_ptq.py` to quantize the model `nvidia/OpenMath2-Llama3.1-8B`
([available in
hugging-face](https://huggingface.co/nvidia/OpenMath2-Llama3.1-8B)), see
the example command above.
Checked that the generated MXFP8 checkpoint can be loaded with vLLM
(required changes in vLLM, not merged to main).
Added tests for `MXFP8QTensor` in
`tests/gpu/torch/quantization/test_qtensor_cuda.py`.
Added "mxfp8" in `tests/examples/llm_ptq/test_llm_ptq.py`
#### Support for Nemotron Models
Verify that Nemotron Nano V3 BF16 can be converted to MXFP8 using
`hf_ptq.py`:
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added MXFP8 quantization format support with new scaling mechanisms
and quantization utilities.
* Updated configuration options, example scripts, and utilities to
recognize and process MXFP8 quantization workflows.
* Extended quantization export pipelines to handle MXFP8 quantized
models.
* **Tests**
* Expanded test coverage for MXFP8 quantization across various tensor
shapes, data types, and device configurations.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Daniel Serebrenik <daserebrenik@nvidia.com>
|
||
|
|
4a848c4f52 |
Modelopt-windows documentation update (#812)
## What does this PR do? Documentation **Overview:** - Update support matrix, changelog, deployment page, example readmes as per recent feature and model support on Windows side. ## Testing - No testing, its just documentation change ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ONNX Mixed Precision Weight-only quantization (INT4/INT8) support. * Introduced diffusion-model quantization on Windows. * Added new accuracy benchmarks (Perplexity and KL-Divergence). * Expanded deployment with multiple ONNX Runtime Execution Providers (CUDA, DirectML, TensorRT-RTX). * **Bug Fixes** * Fixed ONNX 1.19 compatibility issue with CuPy during INT4 AWQ quantization. * **Documentation** * Updated installation guides with system requirements and multiple backend options. * Reorganized deployment documentation with comprehensive execution provider guidance. * Expanded example workflows with improved setup instructions and support matrices. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: vipandya <vipandya@nvidia.com> |
||
|
|
1c7a928df8 |
Change cnn_dailymail to abisee/cnn_dailymail (#819)
cnn_dailymail is changed to abisee/cnn_dailymail long ago. Perhaps older name no longer works <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated dataset configuration path for improved dataset access. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
0ebcd70878 |
Support VLM calibration with image-text data (#755)
## What does this PR do? **Type of change:** New feature **Overview:** The primary goal of this PR is to allow the model optimizer to use image-text pair data during the calibration phase of quantization, which is likely help improve accuracy of quantized VLMs like Nemotron VL on visual understanding tasks particularly, compared to text-only calibration data. - New Feature: Adds support for VLM calibration specifically using image-text data. - Dataset Integration: Introduces support for sampling from the `Nemotron-VLM-Dataset-v2`. - Refactoring: Created a separate utility for VLM datasets to keep the main Hugging Face PTQ script (`hf_ptq.py`) clean. - Simplified logic for handling multimodal inputs. - Addressed specific issues encountered when calibrating the `Nemotron-Nano-VL-12B-V2` model with image data. - Documentation: Updated the README to include instructions and examples for VLM calibration. This PR complements https://github.com/NVIDIA/Model-Optimizer/pull/347 and we will consolidate llm_ptq and vlm_ptq examples in follow-up PRs. ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path /home/scratch.omniml_data_2/models/Nemotron-Nano-VL-12B-V2 --qformat nvfp4 --export_path /home/omniml_data_3/zhiyuc/checkpoints/Nemotron-Nano-VL-12B-V2-NVFP4-doccalib --trust_remote_code --kv_cache_qformat none --calib_with_images --vlm_dataset nemotron_vlm_dataset_v2 --vlm_subsets sparsetables,plotqa_cot --calib_size 512 ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Not yet <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Vision-Language Model (VLM) calibration support with image-text pair data, specifically for Nemotron VL models. * Added new `--calib_with_images` CLI flag to enable image-based calibration workflows. * Integrated Nemotron VLM dataset v2 for streaming multimodal calibration data. * **Documentation** * Added VLM calibration guidance in the PTQ README with usage examples and dataset information. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
5cc2a54519 |
Ynankani/update windows benchmark md (#762)
## What does this PR do? **Type of change:** ? documentation **Overview:** Md update to add perplexity and kl divergence benchmark info. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: NA - **Did you write any new necessary tests?**: NA - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Expanded accuracy comparison section with three detailed benchmark metrics: MMLU scores, Perplexity (PPL), and KL-divergence. * Added comprehensive tables showing results across models and quantization configurations. * Included evaluation guides and references for each metric. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: unknown <ynankani@nvidia.com> |
||
|
|
3036a9ea9f |
Feat: Context Parallel for Eagle3 Training (#745)
## What does this PR do?
**Type of change:** New Feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Supported Context Parallel by patching torch ring attention;
- Require following libirary version for stable cp:
- torch2.8.0
- transformers5.0.0
- accelrate1.12.0
- Move to FSDP2
- Removed unused arguments in training script (`--multi_gpu`,
`fsdp_wrap_layer`)
- Bump CI container to `nvcr.io/nvidia/pytorch:25.08-py3`
## Usage
<!-- You can potentially add a usage example below. -->
```bash
./launch_train.sh --model $MODEL \
--output_dir $OUTPUT_DIR \
--data $DATA \
--num_epochs 0.1 \
--train_bs 1 \
--eagle_config eagle_config.json \
--training_seq_len 1024 \
--cp_size 2 #newly added
```
## Testing
- SDPA level correctness: tested TTT attention with/without CP, diff <
1%
```
=== Compare context-parallel (CP) outputs and grads with non-CP ===
Forward output comparison (CP vs Non-CP):
Absolute diff (adiff) cp_out vs out: 0.001953125
Relative diff (rdiff) cp_out vs out: 0.00182342529296875
WQ (query proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wq_grad vs wq_grad: 0.0078125
Relative diff (rdiff) cp_wq_grad vs wq_grad: 0.00347900390625
WK (key proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wk_grad vs wk_grad: 0.0078125
Relative diff (rdiff) cp_wk_grad vs wk_grad: 0.002471923828125
WV (value proj) grad comparison (CP vs Non-CP):
Absolute diff (adiff) cp_wv_grad vs wv_grad: 0.25
Relative diff (rdiff) cp_wv_grad vs wv_grad: 0.0069580078125
==============================================================
```
- E2E Training Acc
(Llama3.1-8B, Unsynthesized magpie)
<img width="911" height="630" alt="image"
src="https://github.com/user-attachments/assets/1ecacc7f-c720-494c-9c1b-b60e7ced7baa"
/>
- Peak Mem Reserved
(llama3.1-8B, 8xH100, train_length=4k)
| cp_size | max_memory_allocated(MB) |max_memory_reserved (MB) |
|----|--------------------------|--------------------------|
| 1 | 65040.20 |79018.00
| 2 | 50409.17 |73098.00
| 4 | 45120.92 |72052.00
| 8 | 38882.12 |66484.00
- Max Training Length test
(llama3.1-8B, H100)
| cp_size | 6k | 12k | 24k | 48k |
|--------------------|-----|-----|-----|-----|
| 1 | ✅ | OOM | OOM | OOM |
|2 | ✅ | ✅ | OOM | OOM |
| 4 | ✅ | ✅ | ✅ | OOM |
| 8 | ✅ | ✅ | ✅ | ✅ |
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added context parallelism (CP) and data parallelism shard size
configuration parameters to training arguments.
* **Enhancements**
* Improved TTT attention masking support for speculative decoding
workflows.
* Enhanced training launch script with improved parallelism
configuration handling.
* **Chores**
* Updated core dependencies: torch, transformers, accelerate, and wandb.
* Added FSDP configuration file for distributed training setup.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
|
||
|
|
b44c60ad34 |
Svdquant huggingface checkpoint export support (#754)
## What does this PR do?
**Type of change:** new feature
**Overview:**
## Usage
```bash
cd ./examples/llm_ptq/
python hf_ptq.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--export_path /home/scratch.shiychen_coreai/quantized_models/Qwen3-4B-svdq \
--qformat nvfp4_awq_svdquant --kv_cache_qformat none --sparsity_fmt dense --calib_size 8
```
## Testing
exported checkpoint and loaded.
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added nvfp4_svdquant as a new quantization format option for LLM model
quantization workflows.
* **Limitations**
* Multi-GPU export configurations using tensor or pipeline parallelism
are not supported with nvfp4_svdquant quantization.
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
|
||
|
|
668b8a19e8 |
[1/3] Diffusion ckpt export for NVFP4 & FP8 (#781)
## What does this PR do? **Type of change:** New feature <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** This PR adds support for exporting quantized diffusers models (DiT, Flux, SD3, UNet, etc.) to HuggingFace checkpoint format, enabling deployment to inference frameworks like SGLang, vLLM, and TensorRT-LLM. **Changes** New file: `diffusers_utils.py` - Dummy input generation for various diffusion models - Pipeline component extraction helpers - QKV projection detection and grouping - `hide_quantizers_from_state_dict()` context manager for clean saves Refactored: `unified_export_hf.py` - New `_fuse_qkv_linears_diffusion()` for QKV amax fusion - `_export_diffusers_checkpoint()` to export full pipelines (models + tokenizers + schedulers etc.) Plans - [x] [1/3] Add the basic functionalities to support limited image models with NVFP4 + FP8, with some refactoring on the previous LLM code and the diffusers example. PIC: @jingyu-ml - [ ] [2/3] Add support to more video gen modelsPIC: @jingyu-ml - [ ] [3/3] Add test cases, refactor on the doc, and all related README. PIC: @jingyu-ml ## Usage <!-- You can potentially add a usage example below. --> ``` mtq.quantize(pipe, quant_config, forward_call) export_hf_checkpoint(pipe, export_dir=hf_ckpt_dir) ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes <!--- If No, explain why. --> - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added HuggingFace checkpoint export support for quantized diffusion models with configurable output directory * Introduced new `--hf-ckpt-dir` CLI argument for specifying checkpoint export destination * Extended export functionality to support selective component exports from diffusion pipelines * Enhanced quantized model export with improved component handling and multi-stage checkpoint generation <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
563a1e09c6 |
Add NAS to Minitron pruning for parameter based auto-pruning (#720)
## What does this PR do?
**Type of change:** New feature
- So far users we didnt have the NAS step from the Minitron paper so
users had to manually find a pruned architecture which fits that params
constraints and for that they have to tweak different combinations of
width and depth pruning
- This PR adds a simplified version of the NAS from paper. We first find
all candidate subnets that fit the user's params constraint and sort
them by parameter count.
- [Paper] Then we pick top K candidates, do distillation for ~2B tokens
and then select the one with best score (LM Loss / MMLU / other metric
we care about). Note that the one among these top K with highest params
is often not the best pruned model
- [ModelOpt] Then we pick top K candidates, and select the one with best
score (LM Loss / MMLU / other metric we care about). While doing KD
gives better indication on which one to pick, skipping it makes the
pruning much faster, much less compute intense, and finish everything in
single prune API instead of first exporting top K models, doing KD and
eval for all K models separately. We do print a Note in pruning step to
let users know this so they can do KD if they want slightly better
pruned model.
- Further full KD is still needed as usual
- We also restrict the search space choices (e.g. `hidden_size` multiple
of 256, `ffn_hidden_size` multiple of 512) to make the process
efficient. Users can configure this if they want to.
## Usage
<!-- You can potentially add a usage example below. -->
Pruning API is same as before:
```python
import modelopt.torch.prune as mtp
mtp.prune(
model,
mode="mcore_minitron",
constraints=constraints,
dummy_input=None, # Not used
config=config,
)
```
1. Manual Pruning (Existing):
```python
constraints = {"export_config": {"hidden_size: 3072", "ffn_hidden_size": 9216}}
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
mtp.prune(...)
```
2. NAS-based Auto Pruning (New):
```python
constraints = {"params": 6e9}. # prune to 6B params
config = {"forward_loop": forward_loop, "checkpoint": "/path/to/cache/pruning/scores.pth"}
# define the score_func to maximize (e.g MMLU, negative val loss, etc.)
from modelopt.torch.utils.plugins.megatron_mmlu import megatron_mmlu
def score_func(m):
return megatron_mmlu(m, tokenizer, percentage=0.05) # 5% sampled data for faster eval
config["score_func"] = score_func
# overwrite search space choices (showing defaults):
config["max_width_pruning"] = 0.4
config["max_depth_pruning"] = 0.2
config["hparams_to_skip"] = [] # can be used to disable pruning some hparams e.g. ["num_attention_heads"]
config["top_k"] = 10 # might be better to use 20 at the cost of longer time to prune
mtp.prune(...)
```
To configure search space (shows defaults):
```python
ss_config = mtp.mcore_minitron.get_mcore_minitron_config(
hidden_size_divisor=256,
ffn_hidden_size_divisor=512,
mamba_head_dim_divisor=8,
num_moe_experts_divisor=8,
num_layers_divisor=2,
)
mtp.prune(model, mode=[("mcore_minitron", ss_config)], ....)
```
## Testing
**Qwen3-8B -> 6B (~2 hours on 8xA5000)**
```python
0.4350 score -> {'num_layers': 34, 'hidden_size': 3328, 'ffn_hidden_size': 11264}
BEST 0.5705 score -> {'num_layers': 30, 'hidden_size': 3584, 'ffn_hidden_size': 11776}
0.4051 score -> {'num_layers': 36, 'hidden_size': 3840, 'ffn_hidden_size': 8192}
0.4593 score -> {'num_layers': 36, 'hidden_size': 3584, 'ffn_hidden_size': 9216}
0.2737 score -> {'num_layers': 36, 'hidden_size': 3072, 'ffn_hidden_size': 11776}
0.5556 score -> {'num_layers': 32, 'hidden_size': 3584, 'ffn_hidden_size': 10752}
0.3198 score -> {'num_layers': 28, 'hidden_size': 4096, 'ffn_hidden_size': 10240}
0.4119 score -> {'num_layers': 36, 'hidden_size': 4096, 'ffn_hidden_size': 7168}
0.3808 score -> {'num_layers': 36, 'hidden_size': 3328, 'ffn_hidden_size': 10240}
0.4783 score -> {'num_layers': 34, 'hidden_size': 3840, 'ffn_hidden_size': 8704}
```
**Nemotron-Nano-9B-v2 -> 7B (~2.5 hours on 8xA5000)**
```python
0.2629 score -> {'num_layers': 54, 'hidden_size': 4352, 'mamba_num_heads': 88, 'mamba_head_dim': 72, 'ffn_hidden_size': 15360}
0.2778 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 56, 'ffn_hidden_size': 15680}
0.5041 score -> {'num_layers': 56, 'hidden_size': 4096, 'mamba_num_heads': 96, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
BEST 0.6043 score -> {'num_layers': 48, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
0.0772 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 10240}
0.3550 score -> {'num_layers': 50, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 64, 'ffn_hidden_size': 15680}
0.1016 score -> {'num_layers': 56, 'hidden_size': 4352, 'mamba_num_heads': 120, 'mamba_head_dim': 72, 'ffn_hidden_size': 10752}
0.5461 score -> {'num_layers': 46, 'hidden_size': 4480, 'mamba_num_heads': 128, 'mamba_head_dim': 72, 'ffn_hidden_size': 14848}
0.1992 score -> {'num_layers': 54, 'hidden_size': 4480, 'mamba_num_heads': 80, 'mamba_head_dim': 80, 'ffn_hidden_size': 14336}
0.5881 score -> {'num_layers': 48, 'hidden_size': 4480, 'mamba_num_heads': 112, 'mamba_head_dim': 80, 'ffn_hidden_size': 13824}
```
## Before your PR is "*Ready for review*"
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
OMNIML-3043
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added NAS-based Auto Pruning for Minitron models as an alternative to
manual pruning using parameter constraints
* Introduced parameter counting capabilities for architecture search
* **Documentation**
* Expanded pruning guides with detailed examples and workflows for both
manual and automatic pruning approaches
* Updated configuration documentation with granular divisor parameters
* **Improvements**
* Enhanced parameter counting support for models with dynamic or
mixture-of-experts modules
<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
|
||
|
|
615f99e746 |
Support KIMI K2 Thinking int4 checkpoint PTQ (#669)
## What does this PR do? **Type of change:** ? new feature **Overview:** Support KIMI K2 Thinking PTQ from the original int4 checkpoint. Tested with transformers 4.57.1, compressed-tensors 0.12.0 The model weights are dequantized on the fly to save GPU memory ## Usage scripts/huggingface_example.sh --model <Kimi-K2-Thinking ckpt> --quant nvfp4_mlp_only --trust_remote_code ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Support for nvfp4_mlp_only quantization format, enabling new layer-wise quantization options * Quantization support for CompressedLinear layers in quantized models * **Improvements** * Enhanced quantization for DeepSeek models with improved attention configuration handling * Optimized model loading with automatic precision configuration and weight unpacking * Better memory management during model export with automatic cache cleanup * Conditional sample generation output controlled via verbose mode <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
e6e4efd61e |
[0.5/3] Diffusion ckpt export for NVFP4 & FP8 (#783)
See https://github.com/NVIDIA/Model-Optimizer/pull/781 This is the MR that only includes the refactoring of the llm export, please ignore the change on quantize.py from the diffusion example. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Added `--hf-ckpt-dir` CLI option to save checkpoints in HuggingFace format * Enabled support for exporting Diffusers-based pipelines * Unified export system now handles both transformer and diffusion model architectures <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
0f05d676d3 |
Remove quantization_config in config.json from original deepseek models (#753)
## What does this PR do?
**Type of change:** Bug fix
**Overview:** DeepSeek original checkpoints may include a
`quantization_config` field in `config.json`
(describing the source checkpoint's quantization). When we export
ModelOpt quantization
configs to `hf_quant_config.json`, leaving the original
`quantization_config` in place can
be confusing. Add a function to remove it.
## Usage
<!-- You can potentially add a usage example below. -->
```python
# Add a code snippet demonstrating how to use this
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
Resolve nvbug https://nvbugspro.nvidia.com/bug/5736665
---------
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
|
||
|
|
6038451779 |
Fix Qwen3 recipe and update autoquant example cmd (#749)
## What does this PR do? **Type of change:** Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
9de4877d27 |
[1/2] Address security concerns in code (#626)
- [x] Address feedback on Threat and Vuln Analysis (TAVA) doc by ProdSec team - [x] Add note on safe usage of pickle deserialization of modelopt-generated state files **TODO: [Separate PR]** Replace pickle usage in `modelopt/torch/opt/plugins/megatron.py` - Needs fix on TransformerEngine first as we copy from https://github.com/NVIDIA/TransformerEngine/blob/3ff0b8d4/transformer_engine/pytorch/module/base.py#L863 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added `--trust_calibration_data` CLI flag for secure ONNX quantization with pickle data files. * **Improvements** * Enhanced security validation for generated quantization code. * Simplified data loading by removing pickle-based caching—data is now always loaded fresh. * Added security guidance throughout model state loading operations. * **Documentation** * Updated guides with security best practices for model state handling. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
6ae96b5162 |
[NVBug 5784940] Fix autodeploy example (#764)
## What does this PR do? **Type of change:** bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** update the example with new API ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> python examples/llm_autodeploy/api_server.py --ckpt_path TinyLlama/TinyLlama-1.1B-Chat-v1.0 --world_size 1 ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * AutoDeployConfig parameter handling has been updated to improve how parameters are prepared for language model initialization. The configuration method now uses optimized keyword argument formatting to ensure consistency and clarity across the deployment system. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> |
||
|
|
727da95a91 |
SpecDec Bench: PostProcess flag (#759)
## What does this PR do? **Type of change:** ? Bug Fix: https://nvbugspro.nvidia.com/bug/5795144 **Overview:** ? Pass postprocess flag to handle slicing message. ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Introduced --postprocess command-line option to select postprocessing strategy. Users can choose "base" (default, preserves existing behavior) or "gptoss" (new alternative method) with validation to reject invalid selections. <sub>✏️ Tip: You can customize this high-level summary in your review settings.</sub> <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Izzy Putterman <iputterman@nvidia.com> |
||
|
|
510451322c |
Streamline KD & QAD transformers Trainers (#708)
## What does this PR do? **Type of change:** ? Refactor and stabilization **Overview:** * Enforce use of FSDP-2 on KD and QAD trainers in HF plugins/examples so that we can remove multiple restrictions ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
9c24e2c08e |
Fix Deepseek transformers model loading (#740)
## What does this PR do? **Type of change:** ? Bug fix **Overview:** ? For Deepseek, let's force the user to apply trust_remote_code and use AutoModelForCausalLM for loading the model. ## Testing python hf_ptq.py --pyt_ckpt_path <Kimi-K2-Thinking_path> --qformat nvfp4 --export_path <quantized_ckpt> --kv_cache_qformat none --calib_size 64 --trust_remote_code --dataset cnn_dailymail ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
4eb1835df5 |
Support for KV cache quantization for MLA Attention vLLM fakequant (#714)
## What does this PR do? **Type of change:** Feature extention **Overview:** Added support to quantize KV cache in vLLM fakequant by adding quantization support for [MLAAttention](https://github.com/vllm-project/vllm/blob/v0.11.1/vllm/attention/layer.py#L641) ## Usage Please refer to [Readme](https://github.com/NVIDIA/Model-Optimizer/tree/kinjal/vllm_att_quant/examples/vllm_serve#calibrate-and-serve-fake-quant-model-in-vllm) ```shell KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CFG=NVFP4_DEFAULT_CFG python vllm_serve_fakequant.py deepseek-ai/DeepSeek-V2 --served-model-name deepseek-ai/DeepSeek-V2 --host 0.0.0.0 --port 8001 --trust-remote-code --enforce-eager --gpu-memory-utilization 0.8 ``` ## Testing Locally tested KV Cache quantization ``` (rotary_emb): DeepseekScalingRotaryEmbedding() (mla_attn): MultiHeadLatentAttentionWrapper( (fused_qkv_a_proj): QuantMergedColumnParallelLinear( in_features=5120, output_features=2112, bias=False, tp_size=1, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=141.0000 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.4297 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (q_a_layernorm): RMSNorm(hidden_size=1536, eps=1e-06) (q_b_proj): QuantColumnParallelLinear( in_features=1536, output_features=3072, bias=False, tp_size=8, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=32.0000 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.1670 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (kv_a_layernorm): RMSNorm(hidden_size=512, eps=1e-06) (kv_b_proj): QuantColumnParallelLinear( in_features=512, output_features=4096, bias=False, tp_size=8, gather_output=False (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.2773 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (rotary_emb): DeepseekScalingRotaryEmbedding() (o_proj): QuantRowParallelLinear( in_features=2048, output_features=5120, bias=False, tp_size=8, reduce_results=True (input_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=1.7188 calibrator=MaxCalibrator quant) (weight_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=0.4336 calibrator=MaxCalibrator quant) (output_quantizer): TensorQuantizer(disabled) ) (mla_attn): QuantMLAAttention( (q_bmm_quantizer): TensorQuantizer(disabled) (kv_c_bmm_quantizer): TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=7.5312 calibrator=MaxCalibrator quant) ) ) ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**:No - **Did you add or update any necessary documentation?**:NA - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA ## Additional Information <!-- E.g. related issue. --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
d541324e84 |
Disable QKV NVFP4 quantization for Qwen3 MOE (#735)
## What does this PR do? **Type of change:** ? Recipe improvement **Overview:** ? Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3 Next recipe for accuracy recovery ## Testing Model accuracy benchmarking Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
3350b0a45b |
[OMNIML-3017] MLM QAD example (#682)
## What does this PR do?
**Type of change:** New example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
### Add QAD Training example for Megatron-LM
- File Structure
- qad.sh / sbatch_qad.sh - Training and SLURM submission scripts
- data_utils/ - Dataset download and preprocessing utilities
- configs/ - Configuration templates for Qwen3-30B-A3B (MoE) and
Qwen3-8B (Dense)
- Key Features
- One-button dataset generation (OpenScience + Nemotron-v2)
- Config-based training scripts, keep all tunable knobs into a single
config file
## Usage
<!-- You can potentially add a usage example below. -->
1. Generate dataset
```bash
bash data_utils/generate_dataset.sh \
--output-dir /path/to/datasets \
--mlm-path /path/to/Megatron-LM \
--tokenizer Qwen/Qwen3-30B-A3B-Instruct-2507
```
2. Create a config based on templates
3. Kick off training with Slurm:
```bash
sbatch sbatch_qad.sh --config configs/my-experiment.conf
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
QAD with Qwen3-30B-A3B-instruct-2507 NVFP4 (all layers quantized)
- GPQA:
BF16: 0.549
NVFP4 (PTQ): 0.4949
NVFP4 (QAD): 0.5202
- Livecodebench:
BF16: 0.3987
NVFP4 (PTQ): 0.37
NVFP4 (QAD): 0.3855
- Scicode:
BF16: 0.325
NVFP4 (PTQ): 0.276
NVFP4 (QAD): 0.3146
- AIME
BF16: 0.6049
NVFP4 (PTQ): 0.55
NVFP4 (QAD): 0.5431
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Wei-Ming Chen <weimingc@login-eos01.eos.clusters.nvidia.com>
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
03dc3860af |
Update onnxruntime-gpu (#697)
## What does this PR do? **Type of change:** Bug fix **Overview:** Updated setup.py to use only onnxruntime-gpu and removed onnxruntime-directml as dependency. Also changed onnxruntime-gpu version in examples. ## Testing Tested int4 quantization and MMLU benchmark with updated onnxruntime-gpu , working as expected --------- Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com> |
||
|
|
3fd8b804f4 |
Fix:add vllm dump script (#710)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? Missing one file in last PR:https://github.com/NVIDIA/Model-Optimizer/pull/689/changes ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com> |
||
|
|
2a51bbd7cb |
Optimize calibrate_draft_vocab to read only required lines when calib… (#618)
Optimize calibrate_draft_vocab to read only required lines when
calibrate_size is set
## What does this PR do?
**Type of change:** Performance improvement
**Overview:**
This PR optimizes the
[calibrate_draft_vocab.py](cci:7://file:///Users/obenshoham/PycharmProjects/TensorRT-Model-Optimizer/examples/speculative_decoding/scripts/calibrate_draft_vocab.py:0:0-0:0)
script to improve memory efficiency and I/O performance when using the
`--calibrate_size` parameter. Previously, the script would read all
lines from the data file into memory before slicing to the specified
`calibrate_size`, causing unnecessary resource usage for large datasets.
The optimization uses `itertools.islice` to read only the required
number of lines when `calibrate_size` is specified.
## Usage
The script usage remains unchanged. When using `--calibrate_size`, the
script now only reads the specified number of lines instead of loading
the entire dataset:
```bash
# Only reads first 1000 lines from the dataset (optimized)
python scripts/calibrate_draft_vocab.py \
--model meta-llama/Llama-3.2-1B-Instruct \
--data input_conversations/daring-anteater.jsonl \
--draft_vocab_size 32000 \
--calibrate_size 1000 \
--save_dir draft_vocab_cache
Signed-off-by: Ofir Ben Shoham <ofir_benshoham@intuit.com>
|