mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
acce79ffa30446cd79f80d81bf745697b062ceec
566
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
acce79ffa3 |
Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?
Type of change: New feature
Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.
Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config
### Usage
```python
import modelopt.torch.quantization as mtq
model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```
Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe
recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```
### Testing
- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->
### Additional Information
The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.
* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
|
||
|
|
839fa3d658 |
add: ModelOpt Launcher for Slurm job submission (#1031)
``` # Install cd Model-Optimizer/launcher curl -LsSf https://astral.sh/uv/install.sh | sh git submodule update --init --recursive # Run locally with Docker (single GPU) uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml hf_local=/mnt/hf-local --yes # Run on Slurm cluster (no need to export the follow SLURM_XXX envs if used in sandbox) export SLURM_HOST=login-node.example.com export SLURM_ACCOUNT=my_account export SLURM_HF_LOCAL=/shared/hf-local export SLURM_JOB_DIR=/shared/experiments uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --yes # Preview config without running uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v # Override parameters uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml \ pipeline.task_0.slurm_config.nodes=2 --yes # Dump resolved config for reproducibility (single YAML for reproducibility, great for QA, Eng, and agent to triage) uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --to-yaml resolved.yaml # Run tests uv pip install -e . pytest uv run pytest -v ``` ## Summary Add `launcher/` module for submitting quantization, training, and evaluation jobs to Slurm clusters or running them locally with Docker via `nemo-run`. `nemo-run` is used in all `NVIDIA-NeMo/*` projects. It supports modern YAML factory (superset of the `OmegaConf` and `Hydra`) and it support multiple executor backends (here we use docker and slurm mainly). A sample YAML config `launcher/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`: ``` job_name: Qwen3-8B_NVFP4_DEFAULT_CFG pipeline: # hf_local: path prefix for model weights and datasets. # # This should be a self-managed directory that mirrors the HuggingFace Hub # hierarchy (e.g., /hf-local/Qwen/Qwen3-8B/, /hf-local/cais/mmlu/). Using # a dedicated folder is preferred over the HuggingFace cache (~/.cache/huggingface) # to avoid cache corruption issues with concurrent jobs. # # Override on CLI: # pipeline.global_vars.hf_local=/mnt/my-models/ # use a different path # pipeline.global_vars.hf_local="" # download from HuggingFace Hub global_vars: hf_local: /hf-local/ task_0: script: common/megatron-lm/quantize/quantize.sh args: - --calib-dataset-path-or-name <<global_vars.hf_local>>abisee/cnn_dailymail - --calib-size 32 environment: - MLM_MODEL_CFG: Qwen/Qwen3-8B - QUANT_CFG: NVFP4_DEFAULT_CFG - HF_MODEL_CKPT: <<global_vars.hf_local>>Qwen/Qwen3-8B - MMLU_DATASET: <<global_vars.hf_local>>cais/mmlu - TP: 4 slurm_config: _factory_: "slurm_factory" nodes: 1 ntasks_per_node: 4 gpus_per_node: 4 ``` ### Key features - **`launch.py`** — public entrypoint accepting `--yaml` config format - **`core.py`** — shared logic (dataclasses, executor builders, run loop) also used by nmm-sandbox's `slurm.py` - **Factory system** — env-var-driven `slurm_factory` with `register_factory()` registry - **`<<global_vars.X>>`** interpolation for sharing values across pipeline tasks - **`hf_local`** global var for configurable model/dataset storage path - **Version reporting** — git commit/branch printed at job start for reproducibility - **`--to-yaml`** — dump resolved config for bug reports and reproducibility - **Model-Optimizer symlink** — `modules/Model-Optimizer -> ../..` (auto-created, avoids recursive submodule) ### Files | Path | Description | |------|-------------| | `launcher/launch.py` | Public entrypoint | | `launcher/core.py` | Shared dataclasses, executors, run loop | | `launcher/slurm_config.py` | SlurmConfig + env-var factory | | `launcher/common/` | Shell scripts (quantize, query, eagle3, specdec_bench) | | `launcher/Qwen/Qwen3-8B/` | Example configs (PTQ, EAGLE3 pipeline) | | `launcher/tests/` | 64 unit tests | | `launcher/README.md` | User guide | | `launcher/ADVANCED.md` | Architecture, mount mechanism, Claude Code workflows | | `launcher/CLAUDE.md` | Claude Code project instructions | | `.github/workflows/unit_tests.yml` | CI job for launcher tests | ### Verified - Same YAML produces identical MMLU results via both `slurm.py` and `launch.py`: - Local Docker (TP=1): 0.719 (128/178) - OCI-HSG Slurm (TP=4): 0.730 (130/178) ## Test plan - [x] 64 unit tests (core, factory, YAML, Docker executor, Slurm executor, Docker launch) - [x] CI workflow added to `.github/workflows/unit_tests.yml` - [x] Local Docker end-to-end with `python:3.12-slim` - [x] Qwen3-8B PTQ on OCI-HSG via both launchers - [ ] Reviewer runs: `cd launcher && uv pip install -e . pytest && uv run pytest -v` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## Release Notes * **New Features** * Introduced ModelOpt Launcher for submitting quantization, training, and evaluation jobs to Slurm clusters or running locally via Docker. * Added YAML-based job configuration with multi-task pipeline support and global variable interpolation. * Included example workflows for Qwen3-8B quantization and EAGLE3 speculative decoding. * Provided configurable Slurm and execution environment defaults. * **Documentation** * Added comprehensive README with quick start, environment setup, and configuration guidance. * Added advanced guide detailing launcher architecture and integration patterns. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
52cfa4ecff |
fix: https://github.com/NVIDIA/Model-Optimizer/issues/981 (#983)
### What does this PR do? Type of change: Bug fix <!-- Details about the change. --> An issue is reported in https://github.com/NVIDIA/Model-Optimizer/issues/981 where `str(v)` on some `TransformerConfig` fields will raise `TypeError`. We remove the yaml saving logic entirely as it's unused and can cause future errors still. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved checkpoint loading stability by handling unusual configuration values more gracefully; such values no longer cause failures and are skipped with a warning instead. * Reduced risk of crashes during configuration processing when encountering non-standard or unsupported objects. * **Chores** * Checkpoints no longer include saved run configuration or tool-version metadata, yielding smaller, simpler checkpoint files. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com> Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com> |
||
|
|
7e2e85a7e4 |
[5991789][ONNX][Autotune] Add note about remote autotuning only being available in safety (#1067)
### What does this PR do?
**Type of change**: Bug fix
**Overview**: Remote autotuning via `--autotune` is only available with
`--safe` enabled.
See `trtexec --help`:
```sh
--remoteAutoTuningConfig Set the remote auto tuning config. Must be specified with --safe.
Format: protocol://username[:password]@hostname[:port]?param1=value1¶m2=value2
Example: ssh://user:pass@192.0.2.100:22?remote_exec_path=/opt/tensorrt/bin&remote_lib_path=/opt/tensorrt/lib
```
### Usage
```python
$ python -m modelopt.onnx.quantization.autotune \
--onnx_path resnet50_Opset17_bs128.onnx \
--use_trtexec \
--trtexec_benchmark_args "--remoteAutoTuningConfig=\"<remote autotuning config>\" --safe"
```
### Testing
See `examples/onnx_ptq/autotune`.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Bug Fixes**
* Remote autotuning now automatically enforces safety mode by including
the --safe flag if not already present, with informative warning
messages when the flag is automatically added.
* **Documentation**
* Updated remote autotuning guide to clarify that safety mode (--safe
flag) must be enabled through the trtexec benchmark arguments for proper
configuration.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
|
||
|
|
6ffe4a52b3 |
Add nvfp4_local_hessian to QUANT_CFG_CHOICES (#1065)
### What does this PR do? Type of change: New feature Wire up `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` (from PR #788) to the `hf_ptq.py` CLI so it can be used via `--qformat nvfp4_local_hessian`. One-line addition to `QUANT_CFG_CHOICES` dict. ### Usage ```bash python examples/llm_ptq/hf_ptq.py \ --model Qwen/Qwen3-8B \ --qformat nvfp4_local_hessian \ --kv_cache_qformat fp8 \ --export_fmt hf ``` ### Testing Tested via modelopt-quantization CI pipeline (quant_flow) on GB200 (`oci-hsg` launcher) with Qwen3-8B. PTQ stage completed successfully. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (wiring existing config to existing CLI) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information - `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` was added in PR #788 but not exposed via the CLI. - Also used in modelopt-quantization CI (`quant_flow`) for automated NVFP4 scale-setting sweeps. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a new KV-cache quantization configuration option, expanding the available quantization choices for users. This provides an additional quantization mode to select from in configuration UIs and CLIs while preserving existing behavior and compatibility. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
1dc890d971 |
Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0 ## Summary - **Remove `moe_count_expert_calib_tokens`** config field and the `_moe_count_expert_calib_tokens` internal flag. Token counting is now implicitly enabled when `moe_calib_experts_ratio` is set, removing a redundant knob. - **Change `--moe_calib_experts_ratio` default to `None`** in `hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by default; now the feature is opt-in and non-MoE models are unaffected without any flag. - **Disable `layer_sync_moe_local_experts_amax`** when `moe_calib_experts_ratio` is set, since each expert is calibrated independently with sufficient token coverage in that mode. - **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks on `_moe_calib_experts_ratio` inside the branch that already assumes it is set. ## Changed files | File | Change | |------|--------| | `modelopt/torch/quantization/config.py` | Remove `moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio` description to document amax sync behavior | | `modelopt/torch/quantization/mode.py` | Remove `moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` | | `modelopt/torch/quantization/plugins/huggingface.py` | Remove `_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify `forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set | | `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to `None`; guard validation | | `tests/unit/.../test_sparse_moe.py` | Update tests to use `_moe_calib_experts_ratio` instead of removed flag | ## Test plan - [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio` (non-MoE model, default `None`) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Configuration Changes** * moe_calib_experts_ratio now defaults to None (disabled) instead of 1.0; validation only occurs when a value is provided. * **Refactor** * Simplified MoE calibration flow and token-counting behavior; removed a deprecated expert-calibration configuration field. * **Documentation** * Changelog and docstrings updated to reflect the new default and calibration behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> |
||
|
|
c76633ac9d |
[EAGLE] Configurable number of TTT steps (#1042)
### What does this PR do? Type of change: new CLI option for existing option <!-- Details about the change. --> - Added num_ttt_steps CLI flag - Changed num_ttt_steps default from 4 to 3 for consistency. Num_spec_tokens == 3 or == 7 are most common in practice, so rounding down to 3 and allowing users to increment higher on-demand. Will also improve training efficiency for the OOTB experience. ### Usage Users can now pass `--num_ttt_steps 7` to `launch_train.sh` when training an EAGLE3 model for extended speculation lengths. ### Testing N/A ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added ability to configure train-time-test steps for speculative decoding training via command-line argument. * Updated default train-time-test steps value from 4 to 3. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Benjamin Chislett <bchislett@nvidia.com> |
||
|
|
4292505512 |
Refactor: Clean up EAGLE training dataset preparation (#684)
## What does this PR do? **Type of change:** Refactor **Overview:** - Consolidate input dataset preparation into `make_dataset.py` - Read dataset mix spec from a YAML file - - Can now specify how many samples to take from each split - - Can no longer easily split a dataset into train/test sections. I don't think this feature was really useful to begin with. Most datasets can already be separated into train/val/test at the split level, and those that can't are usually going to be splitted by the training FW anyways. - Add support for a few new dataset types, magpie 300k/500k/1M, nemotron post-training dataset v2. ## Usage See README for detailed example ## Testing Ran it locally on all dataset modes, works successfully and output looks good. Checked shuffling, conversation IDs, and output contents were all unique and usable. - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: Yes - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated speculative decoding example documentation with new dataset references and standardized file paths. * **New Features** * Introduced configuration-driven dataset preparation supporting multiple dataset sources with centralized configuration files. * **Refactor** * Simplified dataset preparation workflow with unified tooling and updated default data paths throughout the training pipeline. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Benjamin Chislett <bchislett@nvidia.com> |
||
|
|
7c33d85607 |
[1/n] Add a Triton attention kernel with HF integration (#1034)
### What does this PR do?
Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->
<!-- Details about the change. -->
- Adds a Triton flash attention kernel (triton_fa.py) with HF
integration for use in sparse attention and quantization workflows. The
kernel implements Flash Attention with varlen support, GQA, causal
masking, and forward/backward.
- Update the sparse attention to support backend="triton".
Key components:
- modelopt/torch/kernels/triton_fa.py -- Core Triton kernel
- modelopt/torch/kernels/hf_triton_attention.py -- HF adapter,
registered as attn_implementation="modelopt_triton"
- modelopt/torch/kernels/__init__.py -- Shared kernel registry
- modelopt/torch/sparsity/attention_sparsity/conversion.py -- Backend
selection (backend="triton" or "pytorch")
- modelopt/torch/sparsity/attention_sparsity/config.py -- Added "triton"
as valid backend option
- examples/llm_sparsity/attention_sparsity/hf_sa.py -- Updated example
to support --backend triton
### Usage
```python
# Direct kernel API (varlen packed format)
from modelopt.torch.kernels import attention
o = attention(
q, k, v, # [total_tokens, heads, head_dim]
b_start_loc=b_start_loc, # [batch] per-sequence start offsets
b_seq_len=b_seq_len, # [batch] per-sequence lengths
max_input_len=max_seq_len,
is_causal=True,
)
# HuggingFace integration (automatic via sparsify)
import modelopt.torch.sparsity.attention_sparsity as mtsa
config = {"sparse_cfg": {"*attn*": {"method": "flash_skip_softmax", "backend": "triton", "enable": True}}}
model = mtsa.sparsify(model, config=config)
# model now uses the Triton kernel for attention
# Or load directly with attn_implementation
model = AutoModelForCausalLM.from_pretrained(path, attn_implementation="modelopt_triton")
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
`tests/gpu/torch/sparsity/attention_sparsity/test_triton_fa.py`
#### Kernel benchmark on RTX 6000
| SEQ_LEN | ModelOpt Triton | PyTorch SDPA | Flash Attention 2 |
|--------:|----------------:|-------------:|------------------:|
| 256.0 | 34.435353 | 26.199215 | 47.927293 |
| 512.0 | 60.216998 | 47.408218 | 80.736116 |
| 1024.0 | 81.209990 | 82.673526 | 94.197181 |
| 2048.0 | 88.800973 | 89.239451 | 94.822496 |
| 4096.0 | 88.302953 | 89.192071 | 96.826178 |
| 8192.0 | 89.538177 | 89.115835 | 91.563461 |
| 16384.0 | 85.457533 | 80.509254 | 81.391092 |
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added Triton flash attention backend support for sparse attention
operations, alongside the existing PyTorch backend, enabling improved
performance for compatible hardware.
* **Documentation**
* Updated README to document both available attention backends and their
configurations.
* **Tests**
* Added comprehensive test coverage for the new Triton backend,
including forward and backward pass validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
|
||
|
|
00fa5bd790 |
ModelOpt Framework, Recipe Lib, converting subset of existing recipes 1/N (#1000)
### What does this PR do?
1. start a new config system using yaml/yml files.
2. add a new top level package: modelopt_recipes
I want it to be a top level package so we can make it clear that the
modelopt package holds the code, this new package holds recipes
3. implement some of the existing quantization recipes using the new
config system as model agnostic general recipes, but not actually in
use. these recipes sit inside modelopt_recipes/general/ptq/...
4. make sure the configs from the new config system match the exisiting
configs
5. extend the hf_ptq script to enable recipe based PTQ
8. testted hf_ptq using both builtin and extenal config file. example
script:
### Usage
```bash
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
...
```
### Testing
```bash
python examples/llm_ptq/hf_ptq.py \
--model Qwen/Qwen3-8B \
--recipe general/ptq/fp8_default-fp8_kv \
--export_path=fp8_default-fp8_kv \
--calib_size=16 \
--batch_size=0 \
--trust_remote_code \
--export_fmt=hf
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipe-driven PTQ workflows via YAML recipes and new recipe loader;
CLI gains a --recipe option and --pyt_ckpt_path renamed to --model.
* Many new PTQ recipe and config presets (FP8, INT4/INT8, NVFP4, MXFPx,
KV-cache variants) and improved runtime config loading/merging.
* **Documentation**
* Added READMEs describing recipe/config layout.
* **Tests**
* New unit tests covering config loading, inheritance and recipe
loading.
* **Chores**
* Added YAML/OmegaConf runtime support and packaging of recipe YAMLs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
0.43.0rc1
|
||
|
|
cb1ff321ee |
Add Python 3.13 support (#1048)
Fixes https://github.com/NVIDIA/Model-Optimizer/issues/217 ## Summary - Bump `requires-python` from `>=3.10,<3.13` to `>=3.10,<3.14` to formally include Python 3.13 - Add explicit Python 3.10–3.13 PyPI classifiers for better discoverability - Add `py313` to tox CPU unit test and partial-install environment matrices - Add Python 3.10–3.13 to the `multi-py` CI matrix in `unit_tests.yml` ## Background Python 3.13 was previously excluded by the `<3.13` upper bound. Testing in a related repo with `--ignore-requires-python` confirmed that the library installs and runs correctly under Python 3.13. This PR lifts the restriction and wires up CI to verify it going forward. ## Test plan - [ ] CI `multi-py` job passes on `py313-torch210-tf_latest-unit` - [ ] `tox -e py313-torch210-tf_latest-unit` passes locally (requires Python 3.13 installed) - [ ] `tox -e py313-partial-unit-torch` passes locally - [ ] No regressions on existing Python 3.10/3.11/3.12 matrix jobs 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Extended Python support: minimum remains 3.10; added official support up through 3.13 (upper bound advanced accordingly). * **Tests** * CI and test matrix expanded to include experimental Python 3.13 coverage. * **Documentation** * Installation docs and changelog updated to reflect Python 3.13 support. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ivan Basov <ibasov@nvidia.com> Signed-off-by: Ivan Basov <5455484+ivanbasov@users.noreply.github.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>0.44.0dev |
||
|
|
e4df91bf04 |
OMNIML-2663] Replace modelopt FP8 QDQ nodes with native ONNX QDQ nodes (#852)
## What does this PR do? **Type of change:** New feature **Overview:** - Updated FP8 quant exporter to replace modelopt custom QDQ nodes with native ONNX QDQ nodes - Updated get_onnx_bytes_and_metadata to make convert_float_to_float16() default instead of autocast - Created util functions to fix graph structure after conversion ## Testing ``` python torch_quant_to_onnx.py --quantize_mode=fp8 \ --onnx_save_path=<model_path> \ --calibration_data_size 64 \ --batch_size 128 python evaluate.py --onnx_path=<model_path> \ --model_name=vit_base_patch16_224 \ --results_path=./results.txt \ --batch_size 128 ``` Results: Before replacement: ``` The top1 accuracy of the model is 85.06% The top5 accuracy of the model is 97.558% Inference latency of the model is 5.27963 ms ``` After replacement: ``` The top1 accuracy of the model is 85.054% The top5 accuracy of the model is 97.542% Inference latency of the model is 5.74771 ms ``` ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: No - Replaced modelopt QDQ nodes with native ONNX qdq nodes - **Did you write any new necessary tests?**: No - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * ONNX utilities to remove redundant Casts, fold Constant→Cast patterns, and convert targeted Casts to FP16. * **Improvements** * FP8 QDQ nodes now converted to native ONNX QDQ/Dequantize nodes for improved compatibility. * Export pipeline streamlined: consistent FP16 handling, unified weight quantization, cast cleanup ordering, and added logging for better traceability. * **Tests** * Unit tests updated to use the new ONNX utilities. * **Changelog** * Entry added noting FP8 QDQ → native ONNX QDQ conversion. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>0.43.0rc0 |
||
|
|
beac6e9faa |
Sequential calibrate refactor (#982)
### What does this PR do?
Type of change: New feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
The current sequential calibration support has O(N^2) complexity for
collecting updated activations for a decoder layer. To solve this, we
adopted a modular/plugin based approach which involves hooks to capture
the updated activations by running forward on the previous decoder layer
using cached prev layer activations. This leads to an issue with nested
modules i.e. the logic in the parent module might need to be replicated
in the lower level modules to ensure equivalence. For example, in the
nemotron model, the parent module NemotronHModel has logic to create and
select appropriate mask based on the decoder layer type (mamba vs
attention).
This PR implements a more generic solution for sequential calibration,
by choosing to collect activations using model forward, thereby ensuring
that all the parent module logic is preserved. We use an attribute
"state"on the modules to indicate whether to perform recomputation/skip
the layer while running module forward. This can help us avoid redundant
computations for getting updated activations.
The overall flow is as follows
1. The user must register a get_decoder_layers() function that returns a
list of layers to be calibrated sequentially
2. LayerActivationCollector, goes through the list of layers and patches
module forward with a "state aware" module forward
3. When model.forward() is called, all the parent logic is recomputed as
expected (embeddings, residual connections, generating attention mask
etc).
4. Lets say we are currently calibrating layer N and we want to get
updated activations; we set layer N to capture and layer N-1 to run
(because this layer was processed previously and updated activations
need to be generated). Already processed layers are set to skip. When
model.forward() is called, all the previous decoder layer computations
are skipped. Layer N-1 uses the cached inputs to generate new
activations. Layer N inputs are captured using the same logic as before
and cached so that they can be used to get updated activations for Layer
N+1.
### Usage
```python
# Sequential calibrate config
NVFP4_SEQUENTIAL_CFG = {
"quant_cfg": {
"*weight_quantizer": _nvfp4_quantizer,
"*input_quantizer": _nvfp4_quantizer,
**_default_disabled_quantizer_cfg,
},
"algorithm": {"method": "max", "use_sequential": True},
}
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Public sequential, per-layer calibration API and an
activation-collection utility.
* Broader model discovery support including Nemotron-H and homogeneous
HuggingFace variants.
* **Improvements**
* Clearer validation/error messages and deterministic
patching/unpatching with guaranteed cleanup and resource handling.
* Consolidated discovery/registration flow for decoder-layer handling
and improved per-layer logging/progress.
* **Tests**
* Extensive new unit tests covering discovery, per-layer capture/replay,
inter-layer behavior, and edge cases.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: realAsma <akuriparambi@nvidia.com>
|
||
|
|
7b34de6436 |
Unify weight_scale_2 between gate_proj/up_proj (and w1/w3) in the HF export path for MOE models (#1033)
### What does this PR do? Unify `weight_scale_2` between `gate_proj/up_proj` (and `w1/w3`) in the HF export path for MOE models. Serving engines fuse these projections into a single `gate_up_proj` and require a shared scale; this takes the element-wise max of the two independent scales as a conservative choice that avoids overflow. Type of change: ? Bug fix ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Automatic synchronization of quantization scaling between Mixture-of-Experts gate and up projections during model export for non‑fused MoE setups (e.g., Qwen MoE, DeepSeek). * **Bug Fixes / Improvements** * Export now emits a brief notification when gate/up scaling values are adjusted to ensure consistent quantization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> |
||
|
|
1070d895dc |
Flux2-Dev Quantization (#947)
## What does this PR do?
**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
- Register Flux2Attention and Flux2ParallelSelfAttention in the
quantization plugin so bmm quantizers are patched (enables
--quantize-mha).
- Add Flux2-specific dummy input generation for HF checkpoint export.
- Guard check_conv_and_mha with hasattr for bmm quantizer attributes
## Usage
<!-- You can potentially add a usage example below. -->
```bash
python quantize.py \
--model flux2-dev \
--model-dtype BFloat16 \
--format fp4 --batch-size 2 --calib-size 1 \
--n-steps 20 --quantized-torch-ckpt-save-path ./flux2-dev-fp4.pt --collect-method default \
--hf-ckpt-dir ./flux2-dev-fp4
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes<!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added Flux2-dev model support with Flux2-compatible dummy input
generation and default inference params (768×1024, guidance scale 4.0).
* **Refactor**
* Made attention quantization disabling more robust by iterating
available quantizers before disabling.
* **Infrastructure**
* Flux2 attention components are now optional and registered only when
present to avoid import issues.
* **Tests**
* Added Flux2 test helpers and coverage validating Flux2 dummy input
shapes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
0f8482a982 |
[minor]: Fix AutoQuant Megatron test (#1040)
### What does this PR do? Type of change: Bug fix <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> The `quantizer_states` are on different rank, so the tests are failing, while I can move them to the same device and compare, I think it's sufficient to just compare the search outcome. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
5b417377a8 |
Support Kimi-K2.5 PTQ (#820)
## What does this PR do? **Type of change:** New model support <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** Support Kimi-K2.5 PTQ. ## Usage <!-- You can potentially add a usage example below. --> ```python python3 hf_ptq.py --pyt_ckpt_path moonshotai/Kimi-K2.5 --qformat nvfp4_mlp_only --export_path ./kimi-k2.5-nvfp4 --trust_remote_code ``` ## Testing <!-- Mention how have you tested your change if applicable. --> You may need `pip install transformers==4.57.1` and the model file here: https://huggingface.co/nvidia/Kimi-K2.5-NVFP4/blob/main/modeling_kimi_k25.py ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixes model loading so mixed-precision (BF16) and expert weights are correctly restored and stale placeholders removed. * Adds error handling to avoid failures when optional decompression components are missing. * **New Features** * Adds conditional patch/restore around pack‑quantized model loads, final unpacking of weights after load, and on‑the‑fly decompression during inference. * Improves logging for load and decompression events. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu <zhiyuc@nvidia.com> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
bc96f1ce39 |
[OMNIML-3277] Update kv cache behavior (#1012)
### What does this PR do?
Type of change: New feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
By default, FP8 KV cache quantization in hf_ptq.py now uses a constant
scale of 1.0 (amax=448.0) without a data-driven calibration pass. No KV
scales are written to the exported checkpoint — inference engines
(TRT-LLM, vLLM) use scale=1.0 when no scale is
present, so this is lossless. Pass --calibrate_kv_cache to opt into
data-driven per-tensor KV scale calibration (previous default behavior).
To support this cleanly in the quantization stack, a constant_amax field
is added to QuantizerAttributeConfig. Quantizers configured with
constant_amax skip calibration entirely (no forward pass needed), use
the fixed amax during fake-quant, and produce no
_amax buffer in the state dict.
### Usage
```python
# Quantize with default constant KV scale (no calibration pass for KV)
python hf_ptq.py --model ... --qformat fp8 --kv_cache_qformat fp8
# Opt into data-driven KV calibration
python hf_ptq.py --model ... --qformat fp8 --kv_cache_qformat fp8 --calibrate_kv_cache
# Use constant_amax in a custom quant config
quant_cfg = {
"quant_cfg": {
"*[kv]_bmm_quantizer": {"num_bits": (4, 3), "enable": True, "constant_amax": 448.0},
},
"algorithm": "max",
}
model = mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
- All 8 existing GPU HF export tests pass
(tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py)
- Two new CPU unit tests added to TensorQuantizerTester:
test_constant_amax and test_constant_amax_skips_calibration
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ (--calibrate_kv_cache flag
defaults to False; existing scripts without the flag now skip KV
calibration)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
- Did you write any new necessary tests?: ✅
- Did you update Changelog?: ✅
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* New CLI flag --calibrate_kv_cache to optionally enable data-driven KV
cache calibration; default uses a fixed KV scale and omits KV scales
from exported checkpoints.
* Added constant_amax option to set fixed quantizer scales and skip
dynamic calibration for configured quantizers.
* **Bug Fixes**
* Removed forced flooring/clamp of KV cache scales; out-of-range
activations now emit a shorter warning.
* **Tests**
* Added tests for constant_amax behavior and calibration interaction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
|
||
|
|
6f32d242fc |
Implicit Gemm NVFP4 on Conv3D (#886)
## What does this PR do?
**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->
**Overview:**
Experimental Conv3D implicit-GEMM CUDA kernel with optional NVFP4-style
(E2M1 + FP8 E4M3 scale) fake quantization for activations.
It is intended for research/prototyping and quantization-accuracy
experiments only, not production deployment.
The implementation runs as a JIT-compiled PyTorch extension, mirrors
conv3d output shape, and provides a quantized and non-quantized path to
compare numerical behavior.
There is currently no real quantized production kernel integration in
the formal ModelOpt export/compress/runtime stack; this path is kept in
experimental/ for fake-quant accuracy validation and benchmarking.
## Usage
<!-- You can potentially add a usage example below. -->
```python
import torch
from experimental.conv.implicit_gemm_cuda import conv3d_implicit_gemm_cuda
from modelopt.torch.quantization.tensor_quant import dynamic_block_quantize_op
x = torch.randn(1, 128, 21, 60, 106, device="cuda")
w = torch.randn(512, 128, 3, 3, 3, device="cuda")
block_size = 128
# Without FP4 activation quantization (drop-in-style Conv3D call)
out = conv3d_implicit_gemm_cuda(x, w, stride=(1, 1, 1), padding=(1, 1, 1))
# Optional FP4 block quantization of weights along the GEMM K dimension.
# The kernel's A-tile (activations) is quantized along K = Cin*kD*kH*kW,
# so weights must be flattened to [Cout, K] before quantizing to match.
Cout, Cin = w.shape[:2]
K = Cin * w.shape[2] * w.shape[3] * w.shape[4]
w_flat = w.reshape(Cout, K)
w_q_flat = dynamic_block_quantize_op(
w_flat,
block_size,
w_flat.abs().max().unsqueeze(0),
4, # num_bits
2, # exponent_bits
8, # scale_num_bits
4, # scale_exponent_bits
)
w_q = w_q_flat.reshape_as(w)
# With FP4 activation fake quantization
out_q = conv3d_implicit_gemm_cuda(
x,
w_q,
stride=(1, 1, 1),
padding=(1, 1, 1),
act_amax=x.abs().max().unsqueeze(0),
quant_act=True,
fp4_block_size=block_size, # 128 or 256
)
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **New Features**
* Added experimental Conv3D implementation with implicit GEMM
acceleration and optional FP4 quantization support
* Added benchmarking tool to compare 3D convolution performance across
implementations
* Enhanced quantization framework integration for Conv3D operations
* **Documentation**
* Added comprehensive guide for experimental Conv3D prototype, including
supported scenarios, API reference, and current limitations
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
812e8c60a2 |
Minor update on the LTX2 NVFP4 recipe (#1010)
### What does this PR do? Type of change: minor code change <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> 1. Update the default calibration dataset for LTX_VIDEO_DEV and LTX2 from Gustavosta/Stable-Diffusion-Prompts to nkp37/OpenVid-1M, which provides video-specific captions better suited for video model calibration. 2. update the default recipe for ltx2: first 3 and last 3 layers stays at higher precision. ### Usage ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Updated default dataset configuration for LTX-Video and LTX2 models. * Refined model filtering pattern for LTX-Video to support additional model components. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jingyu Xin <jingyux@nvidia.com> |
||
|
|
2d7d1ec345 |
Skip softmax calibration with list of thresholds (#987)
Modify skip softmax calibration to use a list of thresholds instead of a single threshold. Sparsity during inference is unchanged, but during calibration we can use the list to gather statistics about many thresholds in a single forward pass. Makes calibration 20x faster <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** - Multi-threshold sparsity configuration for fine-grained control over attention sparsity levels. * **Improvements** - Calibration efficiency: Single forward pass for collecting all threshold data instead of iterating per threshold. - Configuration format updated to support threshold lists for prefill and decode phases. * **Breaking Changes** - Configuration API: `threshold` field renamed to `thresholds` and now expects lists of values instead of scalars. - Sparsity statistics output updated to return per-threshold values. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com> |
||
|
|
58417e5014 |
Added block wise RHT (#1014)
### What does this PR do?
Added support for RHT with non-power of 2. Rotate quantization
configuration can be used to specify block_size as well.
```
NVFP4_KV_ROTATE_BLOCK_32_CFG = {
"quant_cfg": {
"*q_bmm_quantizer": {
"enable": False,
"rotate": {"enable": True, "block_size": 32},
},
"*k_bmm_quantizer": {
**_nvfp4_quantizer,
"rotate": {"enable": True, "block_size": 32},
},
"*v_bmm_quantizer": _nvfp4_quantizer,
},
"algorithm": "max",
}
```
### Testing
```
pytest tests/gpu/torch/quantization/test_hadamard.py -k test_hadamard_transform_block
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Block-granular randomized Hadamard transform (RHT) for non-power-of-2
dimensions.
* Rotation configuration expanded to accept block-size and an option to
perform rotation in FP32; rotation settings are now exposed via the
public API.
* **Tests**
* Added tests validating block-granular RHT across varied dimensions and
block sizes.
* **Documentation**
* Changelog updated to mention the new block-granular RHT capability.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
|
||
|
|
d0abca7981 |
Support megatron tokenization for post training datasets (#1018)
### What does this PR do? Update megatron_preprocess_data.py to support applying chat template for tokenizing chat based post training datasets <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> - Tokenized Nemotron-Post-Training-Dataset-v2 (~2B tokens for stem + chat + math + code splits) - Doing distillation on pruned nano v2 7B ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ❌ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Added runtime logging for data truncation operations to enhance processing visibility * Improved handling of chat-formatted conversation data in list format * Eliminated duplicate log messages during data encoding operations <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
bc8798182d |
[minor]: fix NemotronH model export in HF path (#943)
## What does this PR do? **Type of change:** ? <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> **Overview:** ? ## Usage We can use `NemotronHForCausalLM` with `MAMBA_MOE_NVFP4_CONSERVATIVE_CFG` or `MAMBA_MOE_NVFP4_AGGRESSIVE_CFG` <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved wildcard pattern matching to emit more specific patterns when applicable, refining quantization wildcard summarization in deployment configurations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com> |
||
|
|
69c0d47946 |
[OMNIML-3252][ONNX] MOQ + Autotune moq integration docs (#1026)
### What does this PR do? **Type of change**: documentation **Overview**: This PR updates the documentation and does some folder re-structuring and file re-naming related to https://github.com/NVIDIA/Model-Optimizer/pull/951. ### Usage Documentation ### Testing Documentation ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ - Did you write any new necessary tests?: N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (renamed `AutoQDQ` to `Autotune`) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Renamed AutoQDQ to Autotune across guides and changelog. * Updated Autotune guide descriptions and wording. * Added a new section on optimizing Q/DQ node placement with Autotune, including CLI usage and API links (appears twice in one README). * Applied minor grammar and capitalization corrections. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
72a5b3df6d |
Minor fix of typo on news (#1028)
### What does this PR do? Minor fix of typo on newsType of change: ? Bug fix <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> Signed-off-by: James Shen <yueshen@nvidia.com> |
||
|
|
52f8783059 |
Update news of Nemotron=3-Super is supported on Megatron-Bridge (#1025)
### What does this PR do? Type of change: Documentation Add Nemotron-3-Super launch news entries to the README "Latest News" section: A new entry highlighting that [NeMo Megatron Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge) now supports Nemotron-3-Super quantization (PTQ) and export workflows using the Model Optimizer library, with a link to the [Quantization (PTQ and QAT) guide](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md#quantization-ptq-and-qat). ### Usage ```markdown N/A — documentation-only change (README.md update). ``` ### Testing No testing required; this is a documentation-only change to README.md. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information Related links: - Megatron Bridge Nemotron 3 Super docs: https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated latest news section with announcement of March 2026 release: NeMo Megatron Bridge now provides full support for Nemotron-3-Super quantization capabilities, supporting both Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) approaches * Added detailed documentation covering export workflows via the Model Optimizer library with direct reference links to comprehensive quantization guides <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: James Shen <yueshen@nvidia.com> |
||
|
|
34a9fc7924 |
Add support matrix for Nemotron-3 (#1023)
### What does this PR do? Type of change: documentation <!-- Use one of the following: Bug fix, new feature, new example, new tests, documentation. --> <!-- Details about the change. --> Add support matrix for Nemotron-3 ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Nemotron-3 model quantization support added; Mixture-of-Experts expert support retained in auto-quantize scoring/grouping * **Documentation** * Hugging Face model support matrix updated to include Nemotron-3 and adjust referenced models * Quantization examples revised: nvfp4_mse with fp8 and effective precision ~4.75; example configs clarified <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> Signed-off-by: Wei-Ming Chen <17592131+meenchen@users.noreply.github.com> |
||
|
|
26cad678d2 |
[OMNIML-3252][ONNX] Add real Q/DQ scales in Autotune (#951)
## What does this PR do?
**Type of change:** New feature
**Overview:** ONNX Autotune (also called Auto Q/DQ) is currently and
standalone feature of ModelOpt that automatically adds Q/DQ where
relevant according to information obtained from TensorRT inference. One
issue is that the scales in those Q/DQ nodes are random.
This PR does 2 major things:
1. Integrates Auto Q/DQ into the ONNX quantization workflow; and
2. Enables calibration data to be used to obtain the correct scales for
the Q/DQ nodes.
## Usage
```python
$ python -m modelopt.onnx.quantization --onnx_path=model.onnx --autotune={quick,default,extensive}
```
> Please see `__main__.py` for other args.
## Testing
1. Added unittest for Q/DQ node placement validation:
`tests/gpu/onnx/quantization/test_autotune_quantization_integration.py`
2. Verified that accuracy was recovered by integrating MOQ with
Autotune. Results on RTX 3090 with TRT 10.12.0.36 (`--stronglyTyped`)
with ViT, as per `examples/onnx_ptq`:
| Model | Top-1 acc | Top-5 acc |
|--------------------------|---------------|----------------|
| FP32 | 85.1% | 97.5% |
| FP16 (FP32 with --fp16) | 85.1% | 97.5% |
| Quant (MOQ) | 82.4% | 96.4% |
| Quant (Autotune) | 0.1% | 0.5%|
| Quant (MOQ + Autotune) | 79.6% | 95.0% |
Notice that accuracy was mostly recovered from standalone Autotune to
MOQ + Autotune (real Q/DQ scales). The drop in accuracy between MOQ and
MOQ + Autotune is likely due to some sensitive nodes being quantized,
such as `BiasAdd` (see bug 5916898).
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No (will be
done in a different PR)
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Autotuning added to ONNX quantization: CLI flags, presets, per-region
tuning, and FP8/INT8 support; accepts in-memory models and optional
output dirs; node-filter loading and explicit-flag CLI behavior.
* Activation-operation accessor exposed and autotune helpers added to
the package API.
* **Bug Fixes**
* Safer graph rewiring to avoid corrupting quantized graphs when targets
are absent.
* **Tests**
* New integration test and model helper validating autotune quantization
consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
## Additional information
To reproduce accuracy with ViT, call `download_example_onnx.py` and
`image_prep.py` without `--fp16`.
If `--fp16` is used here, quantizing this model with `--autotune`
results in the following error:
```
[modelopt][onnx] - ERROR - Benchmark failed: Converting dtype('float16') to a ctypes type
```
This is fixed in https://github.com/NVIDIA/Model-Optimizer/pull/978.
---------
Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
|
||
|
|
fe83270139 |
Refactor HF _QuantSparseMoe: config-driven token counting, NemotronH detection (#970)
## What does this PR do? **Type of change:** New feature **Overview:** Extend `_QuantSparseMoe` to support NemotronH-style MoE blocks (which use `n_routed_experts` instead of `num_experts`) and refactor the MoE calibration features to be config-driven and lazy-initialized. Key changes: - `_is_sparse_moe_block` in `plugins/huggingface.py` now accepts `n_routed_experts` (NemotronH pattern) in addition to `num_experts` - `_QuantSparseMoe` is refactored: token counting and forced expert forwarding are now opt-in via config knobs (`moe_calib_experts_ratio`, `moe_count_expert_calib_tokens`). When both are off (default), forward is a zero-overhead pass-through. - Token counting buffer and gate hook are lazy-initialized on first use instead of eagerly in `_setup` - `_QuantSparseMoe` gets `layer_sync_moe_local_experts_amax` to sync input quantizer amax across experts (same as Megatron path) - Extract shared `sync_moe_experts_input_amax` utility into `utils.py`, also fixing missing weight amax for experts that received no tokens during calibration. Megatron's `_MegatronSequentialMLP` now calls this shared utility. - `SequentialQuantizer` delegates `amax` property ## Testing - Updated and added unit tests in `test_sparse_moe.py` covering default config, lazy init, token counting, top_k restoration, and end-to-end quantize with both features enabled. ## Before your PR is "*Ready for review*" - **Is this change backward compatible?**: Yes - **Did you write any new necessary tests?**: Yes - **Did you add or update any necessary documentation?**: No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: No Signed-off-by: realAsma <akuriparambi@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
358ee83097 |
updated bmm and matmul for GPT-OSS (#999)
### What does this PR do? This PR fixes maximum recursion bug for GPT-OSS. It replaces `torch._bmm` and `torch.matmul` with `torch.ops.aten.bmm` and `torch.ops.aten.matmul` to avoid recursion ### Usage ```shell Docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc4 [Repro Steps]: [gpt-oss] Step1: accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b --output_dir /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft Step 1 completed: SFT checkpoint at /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft Step2: accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft --quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-qat ``` ### Testing ``` python pytest tests/examples/gpt_oss/test_gpt_oss_qat.py ``` ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`:N/A - Did you write any new necessary tests?: N/A (test already exist) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed a recursion-related instability in attention quantization that could cause errors during certain matrix operations, improving reliability. * **Performance** * Improved handling of batched and matrix-multiplication operations under quantization for more consistent and efficient runtime behavior, including better support for outputs specified by callers. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> |
||
|
|
a5d46ff12b |
Auto Quantize improvements and bug fixes for large sparse MoEs (#953)
## What does this PR do?
**Type of change:** New feature + Bug fixes
**Overview:**
Enable AutoQuantize for NemotronH and large SparseMoE models, and update
the FP8 workflow split between `mtq.auto_quantize` and `mtq.quantize`.
`mtq.auto_quantize` is now positioned as the lightweight search phase
(lite calibration + scoring), while `mtq.quantize` is used for
heavier/final calibration workflows (longer calibration passes,
force-all-token style MoE calibration, and advanced recipes such as
GPTQ, MSE, etc.).
### Algorithm & feature changes
- **NemotronH / SparseMoE support**: Updated `quant_module` and
`score_module` rules (should eventually move to the proposed modeling
lib). In future, this should be the only change needed to support new
models — the bug fixes below were unearthed while enabling NemotronH
- **Config generation**: Added
`mtq.get_auto_quantize_config(search_state, constraints=None,
verbose=False)` to re-solve from `search_state` and produce plain-dict
configs (no redundant `output_quantizer`), with optional verbose summary
- **FP8 workflow split**: Use lite calibration in `mtq.auto_quantize`,
then run longer/final calibration with `mtq.quantize` using the
generated config
- **Performance**: Pass `name_to_module` to
`enable_weight_access_and_writeback` to avoid O(N^2) overhead on large
MoE models
- **Calibration caching in checkpoint**: Save/restore quantizer
calibration states (metadata + state_dict) per recipe in the
AutoQuantize checkpoint, so resuming a search skips redundant
calibration
- **Per-rank distributed checkpointing**: When `torch.distributed` is
initialized, each rank saves/loads its own checkpoint file
(`search_state{rank}.pt`), with backward-compatible fallback to the
single-file path
### API updates
- **Config API naming**: Use `mtq.get_auto_quantize_config(...)` for
exporting the searched recipe into a quantize-ready config
- **Recommended usage pattern**:
```python
# 1) Lightweight search + lite calibration
model, search_state = mtq.auto_quantize(
model,
constraints={"effective_bits": 6.0},
quantization_formats=[mtq.NVFP4_DEFAULT_CFG, mtq.FP8_DEFAULT_CFG],
data_loader=data_loader,
forward_step=forward_step,
loss_func=loss_func,
num_calib_steps=64, # lite calibration during search
num_score_steps=128,
)
# 2) Export searched config (optionally re-solve constraints)
auto_quantize_config = mtq.get_auto_quantize_config(
search_state,
constraints={"effective_bits": 6.0},
verbose=True,
)
# 3) Final / longer calibration pass with quantize
model = mtq.quantize(
model,
config=auto_quantize_config,
forward_loop=long_calibration_loop, # e.g. force-all-token style MoE calibration
)
```
### Bug fixes
- Fixed `disabled_layers` handling so fused kernels (e.g. Mamba blocks)
are properly skipped
- Fixed gradient checkpointing to keep all modules except the
checkpointed modules in eval
- Fixed FP8 fake quant NaN/inf when `amax ≈ 0`
- Fixed `SequentialQuantizer.convert_to_single_quantizer` to operate on
`module` instead of `model`, avoiding O(N^2) CPU iteration on SparseMoE
models with 1000s of submodules
- Switched to proper `F.kl_div` for KL divergence scoring
### Not yet exposed to `llm_ptq`
`mtq.get_auto_quantize_config` is not yet wired into `llm_ptq`. The
plain config records per-expert quantization settings for all MoE
experts, resulting in large JSON files. For my experiments I used a
quick workaround. A follow-up PR will add a better config representation
and expose it to `llm_ptq`.
## Testing
- Tested on NemotronH-tiny and Nemotron-Super-RL models
- Verified auto_quantize scoring + config generation end-to-end
- Unit test for checkpoint resume verifies calibration cache correctness
(metadata + tensor values)
- Existing unit tests pass
## Before your PR is "*Ready for review*"
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: AutoQuantize end-to-end
requires GPU + large MoE models; verified manually on NemotronH-tiny and
Nemotron-Super-RL. Unit test coverage to follow.
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information
Follow-up planned: expose `mtq.get_auto_quantize_config` to `llm_ptq`
with a compact config format for MoE models. AWQ support in AutoQuantize
can also be removed in a future PR to keep it lightweight.
---------
Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
||
|
|
fff65b06d6 |
Allow HF trainer to mask sequences prior to reduction (#1009)
### What does this PR do? Type of change: Bug fix Previously HF trainer did not account for loss masking ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Improvements** * Knowledge-distillation loss now properly ignores padding/special tokens and supports masked per-token averaging. * Default loss reduction behavior adjusted for finer-grained training control and clearer per-token outputs. * More robust logit handling with consistent numeric casting for improved stability and accuracy, including mixed-precision scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
695c8e8522 |
Integrate Automated QDQ placement tool - part 4.3 (#843)
## What does this PR do? This PR upload user guide of Automated QDQ placement tool. This tool automatically search QDQ insertion points with better performance. **Overview:** ? ## Usage <!-- You can potentially add a usage example below. --> ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive guide for Automated Q/DQ Placement Optimization workflow, including quick start instructions, advanced usage patterns, configuration options, best practices, and troubleshooting. * **New Features** * Exposed public API for CLI parser programmatic access. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Will Guo <willg@nvidia.com> Signed-off-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com> Co-authored-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
0214676cc2 |
Pin torchprofile==0.0.4 to fix CI
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
cbab377983 |
Fix Mcore test utils for M-LM main (#1008)
### What does this PR do? - M-LM main removed some functions we were using for running mcore inference in unit tests - replace with alternative - Future-proof `hybrid_override_pattern` -> `hybrid_layer_pattern` rename for Mamba models ### Testing Ran megatron tests with M-LM main branch and previous 0.16 release version ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Mamba model pruning to dynamically select the appropriate hybrid pattern configuration, ensuring correct handling across different model variants. * Simplified inference logic for pipeline-parallel models, improving robustness and reliability. * **Tests** * Updated test infrastructure for Mamba and GPT model inference validation with explicit hidden size parameters. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
0a84bb2caf |
Graph Surgery Framework for TRT-RTX (#992)
### What does this PR do?
Type of change: new feature
Adds the **ONNX Graph Surgery** framework — graph-level transformations
on exported ONNX models for optimized inference with ONNX Runtime.
**GQA Attention Replacement** — Replaces native attention subgraphs
(Q/K/V projections, RoPE, softmax, KV cache) with a single fused
`GroupQueryAttention` operator. Handles RoPE cache computation,
attention mask reformatting, Q/K/V weight fusion, and KV cache I/O
automatically from a HuggingFace model ID. Supports FP16/BF16, INT4/AWQ
quantized, and combined QKV models.
**DequantizeLinear Weight Transpose** — Transposes quantized weights in
`DequantizeLinear` nodes to column-major layout for providers like
NvTensorRtRtx. Handles INT4/UINT4 packed formats.
**Whisper Encoder Cross-Attention KV** — Adds cross-attention K/V
projection outputs to the Whisper encoder for the ONNX Runtime GenAI
pipeline. Loads cross-attention weights from HuggingFace and generates
`genai_config.json`.
### Usage
```bash
# GQA attention replacement
python -m modelopt.onnx.graph_surgery replace-gqa \
-i model.onnx -o model_gqa.onnx \
-m meta-llama/Llama-3.2-1B \
--max-seq-len 4096 --dtype float16
# DequantizeLinear weight transpose
python -m modelopt.onnx.graph_surgery transpose-dq \
-i model_quantized.onnx -o model_transposed.onnx
# Whisper encoder cross-attention KV
python -m modelopt.onnx.graph_surgery add-cross-kv \
-i encoder_model.onnx -o encoder_with_kv.onnx \
-m openai/whisper-large-v3-turbo
```
### Testing
Added test case for GQA graph surgery at
`tests/unit/onnx/test_gqa_graph_surgery.py`. Builds a toy attention
subgraph with native ops matching the real Optimum export pattern and
applies GQA surgery on it. Compared end outputs of both models — outputs
match exactly.
```
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_gqa_node_exists PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_gqa_attributes PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_node_count_reduced PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_rotary_emb_nodes_removed PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_position_ids_removed PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_logits_match
Original nodes: 106 -> GQA nodes: 13
Logits shape: (1, 4, 64)
Original[0,:4]: [-2906. 7704. 15248. 8488.]
GQA [0,:4]: [-2906. 7704. 15248. 8488.]
Max abs diff: 0.000000
Mean abs diff: 0.000000
PASSED
```
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* ONNX graph-surgery toolkit: new utilities to replace attention with
GQA, add encoder cross-attention KV outputs, convert FP16→BF16, and
transpose DequantizeLinear weights.
* New CLI with subcommands to run the above transformations.
* New helper utilities for graph manipulation, RoPE cache generation,
and Whisper GenAI config creation.
* **Tests**
* Added unit tests validating GQA surgery and DQ-transpose
transformations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
|
||
|
|
a56b6f356b |
Use os.path.join for quant_summary path (#1001)
### What does this PR do?
Type of change: Bug fix
I'm seeing the error
```
Error saving quant summary: [Errno 2] No such file or directory: '/home/chad/checkpoint//.quant_summary.txt'
```
when using a trailing slash in the `export_path` to `hf_ptq.py`.
### Testing
This fix will work for all cases of with and without trailing slash, and
both `str` and `Path`.
```
cvoegele@nvdilw8aur4dw1a:~>> python
>>> os.path.join("/home/chad/output", "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join("/home/chad/output/", "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join(Path("/home/chad/output"), "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join(Path("/home/chad/output/"), "a.txt")
'/home/chad/output/a.txt'
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Chores**
* Improved internal path handling for quantization summary output to
enhance cross-platform compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
|
||
|
|
1d6ec895ff |
[OMNIML-3495] Add TEGroupedMLP export support for NemotronH models (#967)
### What does this PR do?
Type of change: New feature
Add export support for `TEGroupedMLP` (fused grouped GEMM experts) in
the MCore-to-HuggingFace checkpoint exporter. Previously, the exporter
only supported `SequentialMLP` (which has `local_experts` as a
`ModuleList`). `TEGroupedMLP` stores per-expert weights as `weight0`,
`weight1`, ..., `weight{N-1}` in a single `TEGroupedLinear` module
instead. This caused an `AttributeError: 'QuantTEGroupedMLP' object has
no attribute 'local_experts'` when exporting NemotronH models.
Changes:
- Add `GroupedMLPSlicing` class in `mcore_custom.py` — the export
counterpart of `GroupedMLPMerging`
- Add `_grouped_mlp_slicing` method in `GPTModelExporter` that iterates
`TEGroupedLinear`'s per-expert weights and exports them as individual
HF-format weights with proper quantization scale handling
- Add `"experts.linear_fc1"` and `"experts.linear_fc2"` rules using
`GroupedMLPSlicing` to `nemotron_h_causal_lm_export`
- Route `TEGroupedMLP` (detected by absence of `local_experts`
attribute) to the new `"experts.linear_fc1"` rule in
`_get_transformer_layer_state_dict`
### Usage
No API change. NemotronH models using `TEGroupedMLP` can now be
exported:
```python
import modelopt.torch.export as mtex
mtex.export_mcore_gpt_to_hf(
model=megatron_model,
export_dir="/path/to/hf_export",
pretrained_model_name_or_path="/path/to/hf_model",
)
```
### Testing
Inside Model-Bridge
```
torchrun --nproc_per_node 4 examples/quantization/export.py \
--hf-model-id /models/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/ \
--megatron-load-path /models/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4-MLM \
--export-dir /models/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4-MLM_hf \
--pp 4 \
--dtype bfloat16 \
--trust-remote-code
```
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ The existing `SequentialMLP`
(`local_experts`) path is guarded by `hasattr(layer.mlp.experts,
"local_experts")` and remains unchanged. The new `TEGroupedMLP` path
only activates when `local_experts` is absent and `"experts.linear_fc1"`
is defined in the architecture's rules.
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A
- Did you write any new necessary tests?: ❌ Tested manually with
Nemotron-3-Nano-30B-A3B. Unit test coverage should be added for
`_grouped_mlp_slicing`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ New feature for a specific model architecture.
### Additional Information
- The import counterpart (`GroupedMLPMerging` / `_grouped_mlp_merging`)
was added by @jennifchen in PR #830. This PR completes the round-trip by
adding the export side.
- `_grouped_mlp_slicing` temporarily assigns `module.weight =
module.weight0` so that `_get_quantized_state` can extract
qformat/scales from the module's quantizers, then removes it afterward.
This follows the same pattern used by `_QuantTEGroupedLinear._setup()`
in the quantization plugin.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Export now supports grouped-expert MLP slicing to split fused expert
weights into per-expert tensors for downstream formats.
* Per-expert export logic enhanced with clear fallbacks between packed
and per-expert layouts, including a grouped-MLP export path.
* Nemotron H causal LM import/export mappings updated to better align
with grouped local-expert exports.
* Added fused-normalization export support and safer handling when
loading remote model code.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: James Shen <yueshen@nvidia.com>
|
||
|
|
2bb404ebd8 |
[chore]: weekly bump of uv.lock on main (2026-03-09) (#1006)
## Summary Automated weekly update of uv.lock file for nSpect Scanning: - `uv.lock` — upgraded all transitive dependencies to latest compatible versions Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
5d0e012751 |
inplement mix hidden_states for eagle3; deprecate eagle1 (#946)
## What does this PR do? new feature **Overview:** Enable mix hidden_states in eagle3 training. Deprecate eagle1 ## Usage Add --mix_hidden_states True to launch_train.sh ```python # Add a code snippet demonstrating how to use this ``` ## Testing <!-- Mention how have you tested your change if applicable. --> ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: Yes/No <!--- If No, explain why. --> - **Did you write any new necessary tests?**: Yes/No - **Did you add or update any necessary documentation?**: Yes/No - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: Yes/No <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added --mix_hidden_states option to enable optional hidden-state mixing during training. * Added eagle_ttt_steps setting to control speculative multi-step iterations. * **Chores** * Consolidated speculative decoding to EAGLE3 only; legacy Medusa/EAGLE1 paths removed. * Unified configuration handling so models and plugins accept a single config object. * **Tests** * Updated and expanded tests for hidden-state mixing and EAGLE3-only scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> |
||
|
|
0ad287ca7b |
[ONNX][Autotune] Replace CUDA memory management from CUDART to PyTorch (#998)
### What does this PR do? **Type of change**: Bug fix **Overview**: Replace CUDA memory management from CUDART to PyTorch (higher-level API). ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing 1. Added unittests. 2. Tested that this PR does not break https://github.com/NVIDIA/Model-Optimizer/pull/951 or https://github.com/NVIDIA/Model-Optimizer/pull/978 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information Summary of changes in `benchmark.py — TensorRTPyBenchmark`: | What changed | Before | After | |---|---|---| | Imports | `contextlib` + `from cuda.bindings import runtime as cudart` | `import torch` (conditional) | | Availability flag | `CUDART_AVAILABLE` | `TORCH_CUDA_AVAILABLE = torch.cuda.is_available()` | | `__init__` guard | checks `CUDART_AVAILABLE or cudart is None` | checks `TORCH_CUDA_AVAILABLE` | | `_alloc_pinned_host` | `cudaMallocHost` + ctypes address hack, returns `(ptr, arr, err)` | `torch.empty(...).pin_memory()`, returns `(tensor, tensor.numpy())` | | `_free_buffers` | `cudaFreeHost` + `cudaFree` per buffer | `bufs.clear()` — PyTorch GC handles deallocation | | `_allocate_buffers` | raw `device_ptr` integers, error-code returns | `torch.empty(..., device="cuda")`, `tensor.data_ptr()` for TRT address | | `_run_warmup` | `cudaMemcpyAsync` + `cudaStreamSynchronize` | `tensor.copy_(non_blocking=True)` inside `torch.cuda.stream()` | | `_run_timing` | same cudart pattern | same torch pattern | | `run` — stream lifecycle | `cudaStreamCreate()` / `cudaStreamDestroy()` | `torch.cuda.Stream()` / `del stream` | | `run` — stream arg to TRT | raw integer handle | `stream.cuda_stream` (integer property) | | Error handling | `cudaError_t` return codes | PyTorch raises `RuntimeError`, caught by existing `except Exception` | Related to https://github.com/NVIDIA/Model-Optimizer/pull/961 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Refactor** * TensorRT benchmarking migrated from direct CUDA runtime calls to PyTorch CUDA tensors, pinned memory, and CUDA stream primitives — simplifying buffer management, transfers, and timing semantics. * **Tests** * Expanded GPU autotune benchmark tests with broader unit and integration coverage for CUDA/TensorRT paths, pinned-host/device buffering, stream behavior, warmup/timing, and end-to-end latency scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com> |
||
|
|
d3748c2c60 |
Allow basename of dataset paths to match registered names (#997)
### What does this PR do? Allow local dataset paths to match registered dataset configs Type of change: Bug fix <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a small sample dataset entry (minipile_100_samples) and support for loading datasets from local filesystem paths with automatic detection and config override. * **Chores** * Improved local-path resolution and substring-based matching against registered dataset keys for consistent behavior. * **Tests** * Added a unit test to verify loading samples from a local dataset snapshot. * **Documentation** * Updated docs to describe local-path support, matching behavior, and updated function docstring. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chenhan Yu <chenhany@nvidia.com> Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com> Co-authored-by: Chenhan Yu <chenhany@nvidia.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
6d77ce754f |
Integrate Automated QDQ placement tool - part 4.4 (#961)
### What does this PR do? Many minor changes: 1. Add preset mode to AutoQDQ. 2. Add pattern cache tests. 3. increase batch size for stable QDQ insertion 4. update LICENSE 2024 -> 2026. 5. add cuda-python to pyproject.toml for `[onnx]` ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, using `torch.load(..., weights_only=True)`, avoiding `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other source, did you follow IP policy in [CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes or backward incompatible changes. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added mode presets (quick, default, extensive) with a new --mode option for autotuning. * Introduced AutoQDQ: automated Q/DQ placement tool for ONNX quantization with pattern caching and checkpoint/resume. * **Documentation** * Updated CLI help and examples to show mode usage and override semantics. * **Tests** * Added comprehensive tests for pattern cache and mode-presets/explicit-override behavior; re-enabled a GPU autotuning workflow test; minor test updates. * **Chores** * Added "cuda-python" to optional dependencies. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Will Guo <willg@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
22423041f7 |
API to measure MSE for target quantizers (#940)
## What does this PR do?
**Type of change:** new feature ? <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->
**Overview:** add an API to measure MSE for target quantizers given a
forward loop
## Usage
<!-- You can potentially add a usage example below. -->
```python
# 1. Quantize the model as usual
model = mtq.quantize(model, quant_cfg, forward_loop)
# 2. Compute MSE for all quantizers
mse = mtq.compute_quantization_mse(model, forward_loop)
# 3. Print the top-5 noisiest quantizers
for name, err in sorted(mse.items(), key=lambda x: -x[1])[:5]:
print(f"{name}: {err:.4e}")
```
## Testing
<!-- Mention how have you tested your change if applicable. -->
Unit test and test with HF PTQ
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an API to measure per-quantizer mean-squared error (MSE) between
original and fake-quantized tensors; supports wildcard and callable
filtering, skips disabled/non-fake-quant quantizers, and runs safely
under no-grad.
* **Tests**
* Added comprehensive tests for MSE validity, pattern and callable
filtering, union behavior, exclusion of disabled quantizers,
preservation of model state, and forward-hook cleanup.
* **Documentation**
* Updated changelog to document the new MSE measurement API.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
Signed-off-by: Wei-Ming Chen <17592131+meenchen@users.noreply.github.com>
|
||
|
|
be6dfad920 |
[5951713] Fix benchmark allocation failure (#978)
### What does this PR do?
```
[modelopt][onnx] - ERROR - Benchmark failed: Converting dtype('float16') to a ctypes type
Traceback (most recent call last):
...
raise NotImplementedError(
NotImplementedError: Converting dtype('float16') to a ctypes type
```
### Usage
```python
# Add a code snippet demonstrating how to use this
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
## Release Notes
* **Bug Fixes**
* Improved dtype handling robustness in host memory allocation to avoid
failures for uncommon numeric types.
* Added fallback support for 2-byte floating-point formats (float16,
bfloat16); clearer errors now raised when a dtype is unsupported.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Will Guo <willg@nvidia.com>
|
||
|
|
a007820afa |
Add CLAUDE.md (#956)
### What does this PR do? Add CLAUDE.md file with repo overview for AI agents <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added a comprehensive CLAUDE.md documenting the Model Optimizer: concepts, architecture, design patterns and anti-patterns, security and contribution guidelines, common commands, architecture layout, core abstractions (modes), key components overview, CI/testing and export guidance, setup and workflow tips, and links to further documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com> |
||
|
|
1ccd945a51 |
Remove unused diffusers/cache_diffusion/pipeline and cuda-python dependency (#996)
`cuda-python` has mixed license and needs EStaff approval for usage. And till 0.42, it was only used in `examples/diffusers/cache_diffusion/pipeline` which has not been updated in 9 months and not used anymore hence removing. Also cherry-picked to `release/0.42.0` branch: https://github.com/NVIDIA/Model-Optimizer/pull/984 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Chores** * Removed TensorRT/ONNX deployment and inference tooling, related model export/configuration, and runtime helpers from the cache-optimized diffusion examples; removed the cuda-python example dependency. * **Tests** * Removed the example benchmarking script and its associated benchmark test. * **Documentation** * Strengthened dependency-review, security, and PR guidance; updated PR template and contributing documentation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> |
||
|
|
37d3f10cbd |
To support LTX2 ComfyUI format (#972)
### What does this PR do?
Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
<!-- Details about the change. -->
- Added a flag merged_base_safetensor_path to the example code so that
user can export the ComfyUI style ckpt.
### Usage
```bash
python quantize.py \
--model ltx-2 --format fp4 --batch-size 1 --calib-size 32 --n-steps 40 \
--extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors \
--extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors \
--extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors \
--extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized \
--extra-param fp8transformer=true \
--quantized-torch-ckpt-save-path ./ltx-2-transformer.pt \
--hf-ckpt-dir ./LTX2-NVFP4/ \
--extra-param merged_base_safetensor_path=./ltx-2-19b-dev-fp8.safetensors
```
### Testing
<!-- Mention how have you tested your change if applicable. -->
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).
- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
### Additional Information
<!-- E.g. related issue. -->
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **Documentation**
* Added new command-line parameters documentation for LTX-2 FP4
quantization examples (--hf-ckpt-dir and merged_base_safetensor_path
configuration options)
* **Improvements**
* Enhanced quantization pipeline to support conditional export behavior
based on model type
* Expanded LTX-Video model filtering patterns for more comprehensive
block detection
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
|
||
|
|
296a865c70 |
sample QAD example script (#933)
## What does this PR do? sample QAD example script **Type of change:** ? new example Example script for QAD on diffusion model like ltx-2 **Overview:** ? 1) Model loading 2) NvFP4 fake quant PTQ using mtq.quantize 3) Distillation class wrapping using mtd.convert 4) Using ltx-2 trainer code for training 5) Checkpoint save in bf16 . post process bf16 model using Comfy-kitchen to produce real quantized model for ComyUI inference ## Usage <!-- You can potentially add a usage example below. --> ```python accelerate launch --config_file fsdp_custom.yaml sample_example_qad_diffusers.py train --config ltx2_qad.yaml ``` ## Testing 1) Tested improvement in Vbench score for PTQ and QAD checkpoint. ## Before your PR is "*Ready for review*" <!-- If you haven't finished some of the above items you can still open `Draft` PR. --> - **Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)** and your commits are signed. - **Is this change backward compatible?**: NA <!--- If No, explain why. --> - **Did you write any new necessary tests?**: NA - **Did you add or update any necessary documentation?**: NA - **Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**: NA <!--- Only for new features, API changes, critical bug fixes or bw breaking changes. --> ## Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive README documenting the new Windows QAD example: setup, usage, project layout, and workflow. * **New Features** * Added a complete Quantization-Aware Distillation training example with distributed training config, PTQ calibration, teacher-student distillation, CLI for training/inference, and an inference-checkpoint creation utility. * Added requirements file listing needed Python packages and tooling. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: ynankani <ynankani@nvidia.com> |