Commit Graph
566 Commits
Author SHA1 Message Date
Chenjie Luo acce79ffa3 Add NVFP4_EXPERTS_ONLY_CFG quantization config and YAML recipe (#1030)
### What does this PR do?

Type of change: New feature

Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.

Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config

### Usage

```python
import modelopt.torch.quantization as mtq

model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```

Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe

recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```

### Testing

- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->

### Additional Information

The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.

* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-03-20 05:36:37 +00:00
Chenhan D. YuandClaude Opus 4.6 839fa3d658 add: ModelOpt Launcher for Slurm job submission (#1031)
```
# Install                                                                                                                                                                              
cd Model-Optimizer/launcher                                                                                                                                                            
curl -LsSf https://astral.sh/uv/install.sh | sh                                                                                                                                        
git submodule update --init --recursive                                                                                                                                                
                                                                                                                                                                                         
# Run locally with Docker (single GPU)                                                                                                                                                 
uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml hf_local=/mnt/hf-local --yes                                                                                              
                                                                                                                                                                                         
# Run on Slurm cluster (no need to export the follow SLURM_XXX envs if used in sandbox)                                                                                                                                                              
export SLURM_HOST=login-node.example.com                                                                                                                                               
export SLURM_ACCOUNT=my_account                                                                                                                                                        
export SLURM_HF_LOCAL=/shared/hf-local                                                                                                                                                 
export SLURM_JOB_DIR=/shared/experiments                                                                                                                                               
uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --yes                                                                                                                       
                                                                                                                                                                                         
# Preview config without running                                                                                                                                                       
uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --dryrun --yes -v                                                                                                           
                                                                                                                                                                                         
# Override parameters                                                                                                                                                                  
uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml \                                                                                                                           
      pipeline.task_0.slurm_config.nodes=2 --yes                                                                                                                                         
                                                                                                                                                                                         
# Dump resolved config for reproducibility (single YAML for reproducibility, great for QA, Eng, and agent to triage)
uv run launch.py --yaml Qwen/Qwen3-8B/megatron_lm_ptq.yaml --to-yaml resolved.yaml                                                                                                     
                                                                                                                                                                                         
# Run tests
uv pip install -e . pytest                                                                                                                                                             
uv run pytest -v 
```
 ## Summary

Add `launcher/` module for submitting quantization, training, and
evaluation jobs to Slurm clusters or running them locally with Docker
via `nemo-run`. `nemo-run` is used in all `NVIDIA-NeMo/*` projects. It
supports modern YAML factory (superset of the `OmegaConf` and `Hydra`)
and it support multiple executor backends (here we use docker and slurm
mainly).

A sample YAML config `launcher/Qwen/Qwen3-8B/megatron_lm_ptq.yaml`:
```
job_name: Qwen3-8B_NVFP4_DEFAULT_CFG
pipeline:
  # hf_local: path prefix for model weights and datasets.
  #
  # This should be a self-managed directory that mirrors the HuggingFace Hub
  # hierarchy (e.g., /hf-local/Qwen/Qwen3-8B/, /hf-local/cais/mmlu/). Using
  # a dedicated folder is preferred over the HuggingFace cache (~/.cache/huggingface)
  # to avoid cache corruption issues with concurrent jobs.
  #
  # Override on CLI:
  #   pipeline.global_vars.hf_local=/mnt/my-models/   # use a different path
  #   pipeline.global_vars.hf_local=""                 # download from HuggingFace Hub
  global_vars:
    hf_local: /hf-local/

  task_0:
    script: common/megatron-lm/quantize/quantize.sh
    args:
      - --calib-dataset-path-or-name <<global_vars.hf_local>>abisee/cnn_dailymail
      - --calib-size 32
    environment:
      - MLM_MODEL_CFG: Qwen/Qwen3-8B
      - QUANT_CFG: NVFP4_DEFAULT_CFG
      - HF_MODEL_CKPT: <<global_vars.hf_local>>Qwen/Qwen3-8B
      - MMLU_DATASET: <<global_vars.hf_local>>cais/mmlu
      - TP: 4
    slurm_config:
      _factory_: "slurm_factory"
      nodes: 1
      ntasks_per_node: 4
      gpus_per_node: 4
```
  
### Key features
  - **`launch.py`** — public entrypoint accepting `--yaml` config format
- **`core.py`** — shared logic (dataclasses, executor builders, run
loop) also used by nmm-sandbox's `slurm.py`
- **Factory system** — env-var-driven `slurm_factory` with
`register_factory()` registry
- **`<<global_vars.X>>`** interpolation for sharing values across
pipeline tasks
- **`hf_local`** global var for configurable model/dataset storage path
- **Version reporting** — git commit/branch printed at job start for
reproducibility
- **`--to-yaml`** — dump resolved config for bug reports and
reproducibility
- **Model-Optimizer symlink** — `modules/Model-Optimizer -> ../..`
(auto-created, avoids recursive submodule)
  ### Files
| Path | Description |
  |------|-------------|
  | `launcher/launch.py` | Public entrypoint |
  | `launcher/core.py` | Shared dataclasses, executors, run loop |
| `launcher/slurm_config.py` | SlurmConfig + env-var factory |
| `launcher/common/` | Shell scripts (quantize, query, eagle3,
specdec_bench) |
| `launcher/Qwen/Qwen3-8B/` | Example configs (PTQ, EAGLE3 pipeline) |
| `launcher/tests/` | 64 unit tests |
  | `launcher/README.md` | User guide |
| `launcher/ADVANCED.md` | Architecture, mount mechanism, Claude Code
workflows |
| `launcher/CLAUDE.md` | Claude Code project instructions |
  | `.github/workflows/unit_tests.yml` | CI job for launcher tests |
### Verified

- Same YAML produces identical MMLU results via both `slurm.py` and
`launch.py`:
    - Local Docker (TP=1): 0.719 (128/178)
- OCI-HSG Slurm (TP=4): 0.730 (130/178)
  ## Test plan
- [x] 64 unit tests (core, factory, YAML, Docker executor, Slurm
executor, Docker launch)
  - [x] CI workflow added to `.github/workflows/unit_tests.yml`
- [x] Local Docker end-to-end with `python:3.12-slim`
- [x] Qwen3-8B PTQ on OCI-HSG via both launchers
- [ ] Reviewer runs: `cd launcher && uv pip install -e . pytest && uv
run pytest -v`

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Introduced ModelOpt Launcher for submitting quantization, training,
and evaluation jobs to Slurm clusters or running locally via Docker.
* Added YAML-based job configuration with multi-task pipeline support
and global variable interpolation.
* Included example workflows for Qwen3-8B quantization and EAGLE3
speculative decoding.
  * Provided configurable Slurm and execution environment defaults.

* **Documentation**
* Added comprehensive README with quick start, environment setup, and
configuration guidance.
* Added advanced guide detailing launcher architecture and integration
patterns.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-18 15:32:15 -07:00
52cfa4ecff fix: https://github.com/NVIDIA/Model-Optimizer/issues/981 (#983)
### What does this PR do?

Type of change: Bug fix

<!-- Details about the change. -->

An issue is reported in
https://github.com/NVIDIA/Model-Optimizer/issues/981 where `str(v)` on
some `TransformerConfig` fields will raise `TypeError`. 

We remove the yaml saving logic entirely as it's unused and can cause future errors still.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved checkpoint loading stability by handling unusual
configuration values more gracefully; such values no longer cause
failures and are skipped with a warning instead.
* Reduced risk of crashes during configuration processing when
encountering non-standard or unsupported objects.

* **Chores**
* Checkpoints no longer include saved run configuration or tool-version
metadata, yielding smaller, simpler checkpoint files.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-03-18 15:30:23 -07:00
Gwena Cunha 7e2e85a7e4 [5991789][ONNX][Autotune] Add note about remote autotuning only being available in safety (#1067)
### What does this PR do?

**Type of change**: Bug fix

**Overview**: Remote autotuning via `--autotune` is only available with
`--safe` enabled.

See `trtexec --help`:
```sh
  --remoteAutoTuningConfig           Set the remote auto tuning config. Must be specified with --safe.
                                     Format: protocol://username[:password]@hostname[:port]?param1=value1&param2=value2
                                     Example: ssh://user:pass@192.0.2.100:22?remote_exec_path=/opt/tensorrt/bin&remote_lib_path=/opt/tensorrt/lib
```

### Usage

```python
$ python -m modelopt.onnx.quantization.autotune \
    --onnx_path resnet50_Opset17_bs128.onnx \
    --use_trtexec \
    --trtexec_benchmark_args "--remoteAutoTuningConfig=\"<remote autotuning config>\" --safe"
```

### Testing
See `examples/onnx_ptq/autotune`.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Remote autotuning now automatically enforces safety mode by including
the --safe flag if not already present, with informative warning
messages when the flag is automatically added.

* **Documentation**
* Updated remote autotuning guide to clarify that safety mode (--safe
flag) must be enabled through the trtexec benchmark arguments for proper
configuration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-03-18 15:41:52 -04:00
sungsoo haandClaude Opus 4.6 6ffe4a52b3 Add nvfp4_local_hessian to QUANT_CFG_CHOICES (#1065)
### What does this PR do?

Type of change: New feature

Wire up `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` (from PR #788) to the
`hf_ptq.py` CLI so it can be used via `--qformat nvfp4_local_hessian`.

One-line addition to `QUANT_CFG_CHOICES` dict.

### Usage

```bash
python examples/llm_ptq/hf_ptq.py \
  --model Qwen/Qwen3-8B \
  --qformat nvfp4_local_hessian \
  --kv_cache_qformat fp8 \
  --export_fmt hf
```

### Testing

Tested via modelopt-quantization CI pipeline (quant_flow) on GB200
(`oci-hsg` launcher) with Qwen3-8B. PTQ stage completed successfully.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A (wiring existing config to
existing CLI)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

- `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` was added in PR #788 but not
exposed via the CLI.
- Also used in modelopt-quantization CI (`quant_flow`) for automated
NVFP4 scale-setting sweeps.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a new KV-cache quantization configuration option, expanding the
available quantization choices for users. This provides an additional
quantization mode to select from in configuration UIs and CLIs while
preserving existing behavior and compatibility.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Sungsoo Ha <sungsooh@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-18 11:38:26 -07:00
Chenjie Luo 1dc890d971 Remove _moe_count_expert_calib_tokens flag; tie token counting to moe_calib_experts_ratio (#1062)
Cherry-pick for 0.43.0

## Summary

- **Remove `moe_count_expert_calib_tokens`** config field and the
`_moe_count_expert_calib_tokens` internal flag. Token counting is now
implicitly enabled when `moe_calib_experts_ratio` is set, removing a
redundant knob.
- **Change `--moe_calib_experts_ratio` default to `None`** in
`hf_ptq.py` (was `1.0`). Previously all experts were force-calibrated by
default; now the feature is opt-in and non-MoE models are unaffected
without any flag.
- **Disable `layer_sync_moe_local_experts_amax`** when
`moe_calib_experts_ratio` is set, since each expert is calibrated
independently with sufficient token coverage in that mode.
- **Simplify `_QuantSparseMoe.forward`**: remove redundant truthy checks
on `_moe_calib_experts_ratio` inside the branch that already assumes it
is set.

## Changed files

| File | Change |
|------|--------|
| `modelopt/torch/quantization/config.py` | Remove
`moe_count_expert_calib_tokens` field; update `moe_calib_experts_ratio`
description to document amax sync behavior |
| `modelopt/torch/quantization/mode.py` | Remove
`moe_count_expert_calib_tokens` propagation in `wrapped_calib_func` |
| `modelopt/torch/quantization/plugins/huggingface.py` | Remove
`_moe_count_expert_calib_tokens` from `_QuantSparseMoe`; simplify
`forward`; skip `layer_sync_moe_local_experts_amax` when ratio is set |
| `examples/llm_ptq/hf_ptq.py` | Default `--moe_calib_experts_ratio` to
`None`; guard validation |
| `tests/unit/.../test_sparse_moe.py` | Update tests to use
`_moe_calib_experts_ratio` instead of removed flag |

## Test plan

- [x] Verify `hf_ptq.py` works without `--moe_calib_experts_ratio`
(non-MoE model, default `None`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Configuration Changes**
* moe_calib_experts_ratio now defaults to None (disabled) instead of
1.0; validation only occurs when a value is provided.

* **Refactor**
* Simplified MoE calibration flow and token-counting behavior; removed a
deprecated expert-calibration configuration field.

* **Documentation**
* Changelog and docstrings updated to reflect the new default and
calibration behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-18 10:09:24 -07:00
Benjamin Chislett c76633ac9d [EAGLE] Configurable number of TTT steps (#1042)
### What does this PR do?

Type of change: new CLI option for existing option

<!-- Details about the change. -->

- Added num_ttt_steps CLI flag
- Changed num_ttt_steps default from 4 to 3 for consistency.
Num_spec_tokens == 3 or == 7 are most common in practice, so rounding
down to 3 and allowing users to increment higher on-demand. Will also
improve training efficiency for the OOTB experience.

### Usage

Users can now pass `--num_ttt_steps 7` to `launch_train.sh` when
training an EAGLE3 model for extended speculation lengths.

### Testing

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ability to configure train-time-test steps for speculative
decoding training via command-line argument.
  * Updated default train-time-test steps value from 4 to 3.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
2026-03-18 09:32:17 -07:00
Benjamin Chislett 4292505512 Refactor: Clean up EAGLE training dataset preparation (#684)
## What does this PR do?

**Type of change:** Refactor

**Overview:** 
- Consolidate input dataset preparation into `make_dataset.py`
- Read dataset mix spec from a YAML file
- - Can now specify how many samples to take from each split
- - Can no longer easily split a dataset into train/test sections. I
don't think this feature was really useful to begin with. Most datasets
can already be separated into train/val/test at the split level, and
those that can't are usually going to be splitted by the training FW
anyways.
- Add support for a few new dataset types, magpie 300k/500k/1M, nemotron
post-training dataset v2.

## Usage
See README for detailed example

## Testing
Ran it locally on all dataset modes, works successfully and output looks
good. Checked shuffling, conversation IDs, and output contents were all
unique and usable.

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated speculative decoding example documentation with new dataset
references and standardized file paths.

* **New Features**
* Introduced configuration-driven dataset preparation supporting
multiple dataset sources with centralized configuration files.

* **Refactor**
* Simplified dataset preparation workflow with unified tooling and
updated default data paths throughout the training pipeline.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
2026-03-18 09:29:49 -07:00
kaix-nv 7c33d85607 [1/n] Add a Triton attention kernel with HF integration (#1034)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->
- Adds a Triton flash attention kernel (triton_fa.py) with HF
integration for use in sparse attention and quantization workflows. The
kernel implements Flash Attention with varlen support, GQA, causal
masking, and forward/backward.
- Update the sparse attention to support backend="triton".

Key components:
- modelopt/torch/kernels/triton_fa.py -- Core Triton kernel
- modelopt/torch/kernels/hf_triton_attention.py -- HF adapter,
registered as attn_implementation="modelopt_triton"
- modelopt/torch/kernels/__init__.py -- Shared kernel registry
- modelopt/torch/sparsity/attention_sparsity/conversion.py -- Backend
selection (backend="triton" or "pytorch")
- modelopt/torch/sparsity/attention_sparsity/config.py -- Added "triton"
as valid backend option
- examples/llm_sparsity/attention_sparsity/hf_sa.py -- Updated example
to support --backend triton

### Usage

```python
# Direct kernel API (varlen packed format)
from modelopt.torch.kernels import attention

o = attention(
    q, k, v,  # [total_tokens, heads, head_dim]
    b_start_loc=b_start_loc,  # [batch] per-sequence start offsets
    b_seq_len=b_seq_len,      # [batch] per-sequence lengths
    max_input_len=max_seq_len,
    is_causal=True,
)

# HuggingFace integration (automatic via sparsify)
import modelopt.torch.sparsity.attention_sparsity as mtsa

config = {"sparse_cfg": {"*attn*": {"method": "flash_skip_softmax", "backend": "triton", "enable": True}}}
model = mtsa.sparsify(model, config=config)
# model now uses the Triton kernel for attention

# Or load directly with attn_implementation
model = AutoModelForCausalLM.from_pretrained(path, attn_implementation="modelopt_triton")
```

### Testing
<!-- Mention how have you tested your change if applicable. -->
`tests/gpu/torch/sparsity/attention_sparsity/test_triton_fa.py`

#### Kernel benchmark on RTX 6000

| SEQ_LEN | ModelOpt Triton | PyTorch SDPA | Flash Attention 2 |
|--------:|----------------:|-------------:|------------------:|
|   256.0 |       34.435353 |    26.199215 |         47.927293 |
|   512.0 |       60.216998 |    47.408218 |         80.736116 |
|  1024.0 |       81.209990 |    82.673526 |         94.197181 |
|  2048.0 |       88.800973 |    89.239451 |         94.822496 |
|  4096.0 |       88.302953 |    89.192071 |         96.826178 |
|  8192.0 |       89.538177 |    89.115835 |         91.563461 |
| 16384.0 |       85.457533 |    80.509254 |         81.391092 |


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added Triton flash attention backend support for sparse attention
operations, alongside the existing PyTorch backend, enabling improved
performance for compatible hardware.

* **Documentation**
* Updated README to document both available attention backends and their
configurations.

* **Tests**
* Added comprehensive test coverage for the new Triton backend,
including forward and backward pass validation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-03-17 16:25:16 -07:00
Shengliang Xu 00fa5bd790 ModelOpt Framework, Recipe Lib, converting subset of existing recipes 1/N (#1000)
### What does this PR do?

1. start a new config system using yaml/yml files.

2. add a new top level package: modelopt_recipes

I want it to be a top level package so we can make it clear that the
modelopt package holds the code, this new package holds recipes

3. implement some of the existing quantization recipes using the new
config system as model agnostic general recipes, but not actually in
use. these recipes sit inside modelopt_recipes/general/ptq/...

4. make sure the configs from the new config system match the exisiting
configs

5. extend the hf_ptq script to enable recipe based PTQ

8. testted hf_ptq using both builtin and extenal config file. example
script:


### Usage

```bash
   python examples/llm_ptq/hf_ptq.py         \
     --model Qwen/Qwen3-8B                   \
     --recipe general/ptq/fp8_default-fp8_kv \
     ...

```

### Testing

```bash
   python examples/llm_ptq/hf_ptq.py         \
     --model Qwen/Qwen3-8B                   \
     --recipe general/ptq/fp8_default-fp8_kv \
     --export_path=fp8_default-fp8_kv        \
     --calib_size=16                         \
     --batch_size=0                          \
     --trust_remote_code                     \
     --export_fmt=hf

```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipe-driven PTQ workflows via YAML recipes and new recipe loader;
CLI gains a --recipe option and --pyt_ckpt_path renamed to --model.
* Many new PTQ recipe and config presets (FP8, INT4/INT8, NVFP4, MXFPx,
KV-cache variants) and improved runtime config loading/merging.

* **Documentation**
  * Added READMEs describing recipe/config layout.

* **Tests**
* New unit tests covering config loading, inheritance and recipe
loading.

* **Chores**
  * Added YAML/OmegaConf runtime support and packaging of recipe YAMLs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
0.43.0rc1
2026-03-16 23:11:31 -07:00
cb1ff321ee Add Python 3.13 support (#1048)
Fixes https://github.com/NVIDIA/Model-Optimizer/issues/217

## Summary

- Bump `requires-python` from `>=3.10,<3.13` to `>=3.10,<3.14` to
formally include Python 3.13
- Add explicit Python 3.10–3.13 PyPI classifiers for better
discoverability
- Add `py313` to tox CPU unit test and partial-install environment
matrices
- Add Python 3.10–3.13 to the `multi-py` CI matrix in `unit_tests.yml`

## Background

Python 3.13 was previously excluded by the `<3.13` upper bound. Testing
in a related repo with `--ignore-requires-python` confirmed that the
library installs and runs correctly under Python 3.13. This PR lifts the
restriction and wires up CI to verify it going forward.

## Test plan

- [ ] CI `multi-py` job passes on `py313-torch210-tf_latest-unit`
- [ ] `tox -e py313-torch210-tf_latest-unit` passes locally (requires
Python 3.13 installed)
- [ ] `tox -e py313-partial-unit-torch` passes locally
- [ ] No regressions on existing Python 3.10/3.11/3.12 matrix jobs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Extended Python support: minimum remains 3.10; added official support
up through 3.13 (upper bound advanced accordingly).
* **Tests**
* CI and test matrix expanded to include experimental Python 3.13
coverage.
* **Documentation**
* Installation docs and changelog updated to reflect Python 3.13
support.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ivan Basov <ibasov@nvidia.com>
Signed-off-by: Ivan Basov <5455484+ivanbasov@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
0.44.0dev
2026-03-17 11:33:48 +05:30
Ajinkya Rasane e4df91bf04 OMNIML-2663] Replace modelopt FP8 QDQ nodes with native ONNX QDQ nodes (#852)
## What does this PR do?

**Type of change:**
New feature

**Overview:** 
- Updated FP8 quant exporter to replace modelopt custom QDQ nodes with
native ONNX QDQ nodes
- Updated get_onnx_bytes_and_metadata to make convert_float_to_float16()
default instead of autocast
- Created util functions to fix graph structure after conversion

## Testing
```
python torch_quant_to_onnx.py --quantize_mode=fp8 \
	--onnx_save_path=<model_path> \
	--calibration_data_size 64 \
	--batch_size 128

python evaluate.py --onnx_path=<model_path> \
	--model_name=vit_base_patch16_224 \
	--results_path=./results.txt \
	--batch_size 128
```

Results:
Before replacement:
```
The top1 accuracy of the model is 85.06%
The top5 accuracy of the model is 97.558%
Inference latency of the model is 5.27963 ms
```
After replacement:
```
The top1 accuracy of the model is 85.054%
The top5 accuracy of the model is 97.542%
Inference latency of the model is 5.74771 ms
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: No
- Replaced modelopt QDQ nodes with native ONNX qdq nodes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* ONNX utilities to remove redundant Casts, fold Constant→Cast patterns,
and convert targeted Casts to FP16.

* **Improvements**
* FP8 QDQ nodes now converted to native ONNX QDQ/Dequantize nodes for
improved compatibility.
* Export pipeline streamlined: consistent FP16 handling, unified weight
quantization, cast cleanup ordering, and added logging for better
traceability.

* **Tests**
  * Unit tests updated to use the new ONNX utilities.

* **Changelog**
  * Entry added noting FP8 QDQ → native ONNX QDQ conversion.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
0.43.0rc0
2026-03-17 11:14:25 +05:30
sugunav14andrealAsma beac6e9faa Sequential calibrate refactor (#982)
### What does this PR do?

Type of change: New feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

The current sequential calibration support has O(N^2) complexity for
collecting updated activations for a decoder layer. To solve this, we
adopted a modular/plugin based approach which involves hooks to capture
the updated activations by running forward on the previous decoder layer
using cached prev layer activations. This leads to an issue with nested
modules i.e. the logic in the parent module might need to be replicated
in the lower level modules to ensure equivalence. For example, in the
nemotron model, the parent module NemotronHModel has logic to create and
select appropriate mask based on the decoder layer type (mamba vs
attention).

This PR implements a more generic solution for sequential calibration,
by choosing to collect activations using model forward, thereby ensuring
that all the parent module logic is preserved. We use an attribute
"state"on the modules to indicate whether to perform recomputation/skip
the layer while running module forward. This can help us avoid redundant
computations for getting updated activations.

The overall flow is as follows
1. The user must register a get_decoder_layers() function that returns a
list of layers to be calibrated sequentially
2. LayerActivationCollector, goes through the list of layers and patches
module forward with a "state aware" module forward
3. When model.forward() is called, all the parent logic is recomputed as
expected (embeddings, residual connections, generating attention mask
etc).
4. Lets say we are currently calibrating layer N and we want to get
updated activations; we set layer N to capture and layer N-1 to run
(because this layer was processed previously and updated activations
need to be generated). Already processed layers are set to skip. When
model.forward() is called, all the previous decoder layer computations
are skipped. Layer N-1 uses the cached inputs to generate new
activations. Layer N inputs are captured using the same logic as before
and cached so that they can be used to get updated activations for Layer
N+1.


### Usage

```python
# Sequential calibrate config
NVFP4_SEQUENTIAL_CFG = {
    "quant_cfg": {
        "*weight_quantizer": _nvfp4_quantizer,
        "*input_quantizer": _nvfp4_quantizer,
        **_default_disabled_quantizer_cfg,
    },
    "algorithm": {"method": "max", "use_sequential": True},
}
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Public sequential, per-layer calibration API and an
activation-collection utility.
* Broader model discovery support including Nemotron-H and homogeneous
HuggingFace variants.

* **Improvements**
* Clearer validation/error messages and deterministic
patching/unpatching with guaranteed cleanup and resource handling.
* Consolidated discovery/registration flow for decoder-layer handling
and improved per-layer logging/progress.

* **Tests**
* Extensive new unit tests covering discovery, per-layer capture/replay,
inter-layer behavior, and edge cases.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: realAsma <akuriparambi@nvidia.com>
2026-03-16 17:51:29 -07:00
Zhiyu 7b34de6436 Unify weight_scale_2 between gate_proj/up_proj (and w1/w3) in the HF export path for MOE models (#1033)
### What does this PR do?
Unify `weight_scale_2` between `gate_proj/up_proj` (and `w1/w3`) in the
HF export path for MOE models. Serving engines fuse these projections
into a single `gate_up_proj` and require a shared scale; this takes the
element-wise max of the two independent scales as a conservative choice
that avoids overflow.

Type of change: ? Bug fix


### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / N/A <!--- If ❌, explain why.
-->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Automatic synchronization of quantization scaling between
Mixture-of-Experts gate and up projections during model export for
non‑fused MoE setups (e.g., Qwen MoE, DeepSeek).
* **Bug Fixes / Improvements**
* Export now emits a brief notification when gate/up scaling values are
adjusted to ensure consistent quantization.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-03-16 15:14:45 -07:00
jingyu-ml 1070d895dc Flux2-Dev Quantization (#947)
## What does this PR do?

**Type of change:** new example <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

- Register Flux2Attention and Flux2ParallelSelfAttention in the
quantization plugin so bmm quantizers are patched (enables
--quantize-mha).
- Add Flux2-specific dummy input generation for HF checkpoint export.
- Guard check_conv_and_mha with hasattr for bmm quantizer attributes 

## Usage
<!-- You can potentially add a usage example below. -->

```bash
python quantize.py \
    --model flux2-dev \
    --model-dtype BFloat16 \
    --format fp4 --batch-size 2 --calib-size 1 \
    --n-steps 20 --quantized-torch-ckpt-save-path ./flux2-dev-fp4.pt --collect-method default \
    --hf-ckpt-dir ./flux2-dev-fp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes<!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added Flux2-dev model support with Flux2-compatible dummy input
generation and default inference params (768×1024, guidance scale 4.0).

* **Refactor**
* Made attention quantization disabling more robust by iterating
available quantizers before disabling.

* **Infrastructure**
* Flux2 attention components are now optional and registered only when
present to avoid import issues.

* **Tests**
* Added Flux2 test helpers and coverage validating Flux2 dummy input
shapes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-14 11:42:02 -05:00
Frida Hou 0f8482a982 [minor]: Fix AutoQuant Megatron test (#1040)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

The `quantizer_states` are on different rank, so the tests are failing,
while I can move them to the same device and compare, I think it's
sufficient to just compare the search outcome.


### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-03-14 16:32:51 +05:30
ZhiyuandClaude Opus 4.6 5b417377a8 Support Kimi-K2.5 PTQ (#820)
## What does this PR do?

**Type of change:** New model support <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

**Overview:** Support Kimi-K2.5 PTQ.

## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 hf_ptq.py   --pyt_ckpt_path moonshotai/Kimi-K2.5   --qformat nvfp4_mlp_only   --export_path ./kimi-k2.5-nvfp4   --trust_remote_code
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

You may need `pip install transformers==4.57.1` and the model file here:
https://huggingface.co/nvidia/Kimi-K2.5-NVFP4/blob/main/modeling_kimi_k25.py

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixes model loading so mixed-precision (BF16) and expert weights are
correctly restored and stale placeholders removed.
* Adds error handling to avoid failures when optional decompression
components are missing.

* **New Features**
* Adds conditional patch/restore around pack‑quantized model loads,
final unpacking of weights after load, and on‑the‑fly decompression
during inference.
  * Improves logging for load and decompression events.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu <zhiyuc@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-14 06:56:24 +00:00
Wei-Ming Chen bc96f1ce39 [OMNIML-3277] Update kv cache behavior (#1012)
### What does this PR do?

Type of change: New feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->
By default, FP8 KV cache quantization in hf_ptq.py now uses a constant
scale of 1.0 (amax=448.0) without a data-driven calibration pass. No KV
scales are written to the exported checkpoint — inference engines
(TRT-LLM, vLLM) use scale=1.0 when no scale is
present, so this is lossless. Pass --calibrate_kv_cache to opt into
data-driven per-tensor KV scale calibration (previous default behavior).
To support this cleanly in the quantization stack, a constant_amax field
is added to QuantizerAttributeConfig. Quantizers configured with
constant_amax skip calibration entirely (no forward pass needed), use
the fixed amax during fake-quant, and produce no
  _amax buffer in the state dict. 

### Usage

```python
  # Quantize with default constant KV scale (no calibration pass for KV)
  python hf_ptq.py --model ... --qformat fp8 --kv_cache_qformat fp8

  # Opt into data-driven KV calibration
  python hf_ptq.py --model ... --qformat fp8 --kv_cache_qformat fp8 --calibrate_kv_cache

  # Use constant_amax in a custom quant config
  quant_cfg = {
      "quant_cfg": {
          "*[kv]_bmm_quantizer": {"num_bits": (4, 3), "enable": True, "constant_amax": 448.0},
      },
      "algorithm": "max",
  }
  model = mtq.quantize(model, quant_cfg, forward_loop=calibrate_loop)
```

### Testing
<!-- Mention how have you tested your change if applicable. -->


- All 8 existing GPU HF export tests pass
(tests/gpu/torch/export/test_unified_hf_export_and_check_safetensors.py)
- Two new CPU unit tests added to TensorQuantizerTester:
test_constant_amax and test_constant_amax_skips_calibration

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ (--calibrate_kv_cache flag
defaults to False; existing scripts without the flag now skip KV
calibration)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in CONTRIBUTING.md: N/A
  - Did you write any new necessary tests?: ✅
  - Did you update Changelog?: ✅

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* New CLI flag --calibrate_kv_cache to optionally enable data-driven KV
cache calibration; default uses a fixed KV scale and omits KV scales
from exported checkpoints.
* Added constant_amax option to set fixed quantizer scales and skip
dynamic calibration for configured quantizers.

* **Bug Fixes**
* Removed forced flooring/clamp of KV cache scales; out-of-range
activations now emit a shorter warning.

* **Tests**
  * Added tests for constant_amax behavior and calibration interaction.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-03-14 05:50:04 +00:00
jingyu-ml 6f32d242fc Implicit Gemm NVFP4 on Conv3D (#886)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 

Experimental Conv3D implicit-GEMM CUDA kernel with optional NVFP4-style
(E2M1 + FP8 E4M3 scale) fake quantization for activations.

It is intended for research/prototyping and quantization-accuracy
experiments only, not production deployment.
The implementation runs as a JIT-compiled PyTorch extension, mirrors
conv3d output shape, and provides a quantized and non-quantized path to
compare numerical behavior.

There is currently no real quantized production kernel integration in
the formal ModelOpt export/compress/runtime stack; this path is kept in
experimental/ for fake-quant accuracy validation and benchmarking.

## Usage
<!-- You can potentially add a usage example below. -->

```python
import torch

from experimental.conv.implicit_gemm_cuda import conv3d_implicit_gemm_cuda
from modelopt.torch.quantization.tensor_quant import dynamic_block_quantize_op

x = torch.randn(1, 128, 21, 60, 106, device="cuda")
w = torch.randn(512, 128, 3, 3, 3, device="cuda")
block_size = 128

# Without FP4 activation quantization (drop-in-style Conv3D call)
out = conv3d_implicit_gemm_cuda(x, w, stride=(1, 1, 1), padding=(1, 1, 1))

# Optional FP4 block quantization of weights along the GEMM K dimension.
# The kernel's A-tile (activations) is quantized along K = Cin*kD*kH*kW,
# so weights must be flattened to [Cout, K] before quantizing to match.
Cout, Cin = w.shape[:2]
K = Cin * w.shape[2] * w.shape[3] * w.shape[4]
w_flat = w.reshape(Cout, K)
w_q_flat = dynamic_block_quantize_op(
    w_flat,
    block_size,
    w_flat.abs().max().unsqueeze(0),
    4,  # num_bits
    2,  # exponent_bits
    8,  # scale_num_bits
    4,  # scale_exponent_bits
)
w_q = w_q_flat.reshape_as(w)

# With FP4 activation fake quantization
out_q = conv3d_implicit_gemm_cuda(
    x,
    w_q,
    stride=(1, 1, 1),
    padding=(1, 1, 1),
    act_amax=x.abs().max().unsqueeze(0),
    quant_act=True,
    fp4_block_size=block_size,  # 128 or 256
)
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added experimental Conv3D implementation with implicit GEMM
acceleration and optional FP4 quantization support
* Added benchmarking tool to compare 3D convolution performance across
implementations
  * Enhanced quantization framework integration for Conv3D operations

* **Documentation**
* Added comprehensive guide for experimental Conv3D prototype, including
supported scenarios, API reference, and current limitations
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-14 04:01:41 +00:00
jingyu-ml 812e8c60a2 Minor update on the LTX2 NVFP4 recipe (#1010)
### What does this PR do?

Type of change: minor code change <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

1. Update the default calibration dataset for LTX_VIDEO_DEV and LTX2
from Gustavosta/Stable-Diffusion-Prompts to nkp37/OpenVid-1M, which
provides video-specific captions better suited for video model
calibration.
2. update the default recipe for ltx2: first 3 and last 3 layers stays
at higher precision.

### Usage

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated default dataset configuration for LTX-Video and LTX2 models.
* Refined model filtering pattern for LTX-Video to support additional
model components.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-13 22:42:48 -05:00
Rohan Joshi 2d7d1ec345 Skip softmax calibration with list of thresholds (#987)
Modify skip softmax calibration to use a list of thresholds instead of a
single threshold. Sparsity during inference is unchanged, but during
calibration we can use the list to gather statistics about many
thresholds in a single forward pass. Makes calibration 20x faster


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
- Multi-threshold sparsity configuration for fine-grained control over
attention sparsity levels.

* **Improvements**
- Calibration efficiency: Single forward pass for collecting all
threshold data instead of iterating per threshold.
- Configuration format updated to support threshold lists for prefill
and decode phases.

* **Breaking Changes**
- Configuration API: `threshold` field renamed to `thresholds` and now
expects lists of values instead of scalars.
  - Sparsity statistics output updated to return per-threshold values.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
2026-03-13 22:28:01 +00:00
kinjalpatel27 58417e5014 Added block wise RHT (#1014)
### What does this PR do?

Added support for RHT with non-power of 2. Rotate quantization
configuration can be used to specify block_size as well.
```
NVFP4_KV_ROTATE_BLOCK_32_CFG = {
    "quant_cfg": {
        "*q_bmm_quantizer": {
            "enable": False,
            "rotate": {"enable": True, "block_size": 32},
        },
        "*k_bmm_quantizer": {
            **_nvfp4_quantizer,
            "rotate": {"enable": True, "block_size": 32},
        },
        "*v_bmm_quantizer": _nvfp4_quantizer,
    },
    "algorithm": "max",
}
```


### Testing
```
pytest  tests/gpu/torch/quantization/test_hadamard.py -k test_hadamard_transform_block
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Block-granular randomized Hadamard transform (RHT) for non-power-of-2
dimensions.
* Rotation configuration expanded to accept block-size and an option to
perform rotation in FP32; rotation settings are now exposed via the
public API.

* **Tests**
* Added tests validating block-granular RHT across varied dimensions and
block sizes.

* **Documentation**
  * Changelog updated to mention the new block-granular RHT capability.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-03-13 22:33:19 +05:30
Keval Morabia d0abca7981 Support megatron tokenization for post training datasets (#1018)
### What does this PR do?

Update megatron_preprocess_data.py to support applying chat template for
tokenizing chat based post training datasets

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Tokenized Nemotron-Post-Training-Dataset-v2 (~2B tokens for stem +
chat + math + code splits)
- Doing distillation on pruned nano v2 7B

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Added runtime logging for data truncation operations to enhance
processing visibility
  * Improved handling of chat-formatted conversation data in list format
  * Eliminated duplicate log messages during data encoding operations

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-13 12:32:05 +00:00
Frida Hou bc8798182d [minor]: fix NemotronH model export in HF path (#943)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage

We can use `NemotronHForCausalLM` with
`MAMBA_MOE_NVFP4_CONSERVATIVE_CFG` or `MAMBA_MOE_NVFP4_AGGRESSIVE_CFG`
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved wildcard pattern matching to emit more specific patterns when
applicable, refining quantization wildcard summarization in deployment
configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
2026-03-12 21:25:03 -07:00
Gwena Cunha 69c0d47946 [OMNIML-3252][ONNX] MOQ + Autotune moq integration docs (#1026)
### What does this PR do?

**Type of change**: documentation

**Overview**: This PR updates the documentation and does some folder
re-structuring and file re-naming related to
https://github.com/NVIDIA/Model-Optimizer/pull/951.

### Usage

Documentation

### Testing

Documentation

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (renamed `AutoQDQ` to `Autotune`)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
  * Renamed AutoQDQ to Autotune across guides and changelog.
  * Updated Autotune guide descriptions and wording.
* Added a new section on optimizing Q/DQ node placement with Autotune,
including CLI usage and API links (appears twice in one README).
  * Applied minor grammar and capitalization corrections.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-03-12 10:47:07 -07:00
yueshen2016 72a5b3df6d Minor fix of typo on news (#1028)
### What does this PR do?
Minor fix of typo on newsType of change: ? Bug fix

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-03-11 14:05:26 -07:00
yueshen2016 52f8783059 Update news of Nemotron=3-Super is supported on Megatron-Bridge (#1025)
### What does this PR do?

Type of change: Documentation

Add Nemotron-3-Super launch news entries to the README "Latest News"
section:

A new entry highlighting that [NeMo Megatron
Bridge](https://github.com/NVIDIA-NeMo/Megatron-Bridge) now supports
Nemotron-3-Super quantization (PTQ) and export workflows using the Model
Optimizer library, with a link to the [Quantization (PTQ and QAT)
guide](https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md#quantization-ptq-and-qat).

### Usage

```markdown
N/A — documentation-only change (README.md update).
```

### Testing

No testing required; this is a documentation-only change to README.md.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

Related links:
- Megatron Bridge Nemotron 3 Super docs:
https://github.com/NVIDIA-NeMo/Megatron-Bridge/blob/super-v3/docs/models/llm/nemotron3-super.md

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated latest news section with announcement of March 2026 release:
NeMo Megatron Bridge now provides full support for Nemotron-3-Super
quantization capabilities, supporting both Post-Training Quantization
(PTQ) and Quantization-Aware Training (QAT) approaches
* Added detailed documentation covering export workflows via the Model
Optimizer library with direct reference links to comprehensive
quantization guides

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-03-11 19:18:15 +00:00
Wei-Ming Chen 34a9fc7924 Add support matrix for Nemotron-3 (#1023)
### What does this PR do?

Type of change: documentation <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

<!-- Details about the change. -->
Add support matrix for Nemotron-3

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Nemotron-3 model quantization support added; Mixture-of-Experts expert
support retained in auto-quantize scoring/grouping

* **Documentation**
* Hugging Face model support matrix updated to include Nemotron-3 and
adjust referenced models
* Quantization examples revised: nvfp4_mse with fp8 and effective
precision ~4.75; example configs clarified
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
Signed-off-by: Wei-Ming Chen <17592131+meenchen@users.noreply.github.com>
2026-03-11 18:12:39 +00:00
Gwena Cunha 26cad678d2 [OMNIML-3252][ONNX] Add real Q/DQ scales in Autotune (#951)
## What does this PR do?

**Type of change:** New feature

**Overview:** ONNX Autotune (also called Auto Q/DQ) is currently and
standalone feature of ModelOpt that automatically adds Q/DQ where
relevant according to information obtained from TensorRT inference. One
issue is that the scales in those Q/DQ nodes are random.

This PR does 2 major things:
1. Integrates Auto Q/DQ into the ONNX quantization workflow; and
2. Enables calibration data to be used to obtain the correct scales for
the Q/DQ nodes.

## Usage

```python
$ python -m modelopt.onnx.quantization --onnx_path=model.onnx --autotune={quick,default,extensive}
```
> Please see `__main__.py` for other args.

## Testing
1. Added unittest for Q/DQ node placement validation:
`tests/gpu/onnx/quantization/test_autotune_quantization_integration.py`

2. Verified that accuracy was recovered by integrating MOQ with
Autotune. Results on RTX 3090 with TRT 10.12.0.36 (`--stronglyTyped`)
with ViT, as per `examples/onnx_ptq`:

| Model                    | Top-1 acc | Top-5 acc |
|--------------------------|---------------|----------------|
| FP32                       | 85.1% | 97.5% |
| FP16 (FP32 with --fp16) | 85.1% | 97.5% |
| Quant (MOQ)                      | 82.4% | 96.4% |
| Quant (Autotune)              | 0.1% | 0.5%|
| Quant (MOQ + Autotune) | 79.6% | 95.0% |

Notice that accuracy was mostly recovered from standalone Autotune to
MOQ + Autotune (real Q/DQ scales). The drop in accuracy between MOQ and
MOQ + Autotune is likely due to some sensitive nodes being quantized,
such as `BiasAdd` (see bug 5916898).

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No (will be
done in a different PR)
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Autotuning added to ONNX quantization: CLI flags, presets, per-region
tuning, and FP8/INT8 support; accepts in-memory models and optional
output dirs; node-filter loading and explicit-flag CLI behavior.
* Activation-operation accessor exposed and autotune helpers added to
the package API.

* **Bug Fixes**
* Safer graph rewiring to avoid corrupting quantized graphs when targets
are absent.

* **Tests**
* New integration test and model helper validating autotune quantization
consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

## Additional information
To reproduce accuracy with ViT, call `download_example_onnx.py` and
`image_prep.py` without `--fp16`.

If `--fp16` is used here, quantizing this model with `--autotune`
results in the following error:
```
[modelopt][onnx] - ERROR - Benchmark failed: Converting dtype('float16') to a ctypes type
```
This is fixed in https://github.com/NVIDIA/Model-Optimizer/pull/978.

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-03-11 15:32:47 +00:00
realAsmaandClaude Opus 4.6 fe83270139 Refactor HF _QuantSparseMoe: config-driven token counting, NemotronH detection (#970)
## What does this PR do?

**Type of change:** New feature

**Overview:** Extend `_QuantSparseMoe` to support NemotronH-style MoE
blocks (which use `n_routed_experts` instead of `num_experts`) and
refactor the MoE calibration features to be config-driven and
lazy-initialized.

Key changes:
- `_is_sparse_moe_block` in `plugins/huggingface.py` now accepts
`n_routed_experts` (NemotronH pattern) in addition to `num_experts`
- `_QuantSparseMoe` is refactored: token counting and forced expert
forwarding are now opt-in via config knobs (`moe_calib_experts_ratio`,
`moe_count_expert_calib_tokens`). When both are off (default), forward
is a zero-overhead pass-through.
- Token counting buffer and gate hook are lazy-initialized on first use
instead of eagerly in `_setup`
- `_QuantSparseMoe` gets `layer_sync_moe_local_experts_amax` to sync
input quantizer amax across experts (same as Megatron path)
- Extract shared `sync_moe_experts_input_amax` utility into `utils.py`,
also fixing missing weight amax for experts that received no tokens
during calibration. Megatron's `_MegatronSequentialMLP` now calls this
shared utility.
- `SequentialQuantizer` delegates `amax` property

## Testing

- Updated and added unit tests in `test_sparse_moe.py` covering default
config, lazy init, token counting, top_k restoration, and end-to-end
quantize with both features enabled.

## Before your PR is "*Ready for review*"

- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-11 00:07:35 +00:00
kinjalpatel27 358ee83097 updated bmm and matmul for GPT-OSS (#999)
### What does this PR do?

This PR fixes maximum recursion bug for GPT-OSS. It replaces
`torch._bmm` and `torch.matmul` with `torch.ops.aten.bmm` and
`torch.ops.aten.matmul` to avoid recursion


### Usage

```shell
Docker image: nvcr.io/nvidia/tensorrt-llm/release:1.3.0rc4

[Repro Steps]:

[gpt-oss]
Step1:
accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path openai/gpt-oss-20b --output_dir /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft

Step 1 completed: SFT checkpoint at /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft


Step2:

accelerate launch --config_file configs/zero3.yaml sft.py --config configs/sft_full.yaml --model_name_or_path /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-sft --quant_cfg MXFP4_MLP_WEIGHT_ONLY_CFG --output_dir /tmp/pytest-of-root/pytest-0/test_gpt_oss_complete_pipeline0/gpt-oss-20b-qat
```

### Testing
``` python
pytest tests/examples/gpt_oss/test_gpt_oss_qat.py
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`:N/A
- Did you write any new necessary tests?:  N/A (test already exist)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Fixed a recursion-related instability in attention quantization that
could cause errors during certain matrix operations, improving
reliability.

* **Performance**
* Improved handling of batched and matrix-multiplication operations
under quantization for more consistent and efficient runtime behavior,
including better support for outputs specified by callers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-03-11 02:53:28 +05:30
realAsmaandClaude Opus 4.6 a5d46ff12b Auto Quantize improvements and bug fixes for large sparse MoEs (#953)
## What does this PR do?

**Type of change:** New feature + Bug fixes

**Overview:**

Enable AutoQuantize for NemotronH and large SparseMoE models, and update
the FP8 workflow split between `mtq.auto_quantize` and `mtq.quantize`.

`mtq.auto_quantize` is now positioned as the lightweight search phase
(lite calibration + scoring), while `mtq.quantize` is used for
heavier/final calibration workflows (longer calibration passes,
force-all-token style MoE calibration, and advanced recipes such as
GPTQ, MSE, etc.).

### Algorithm & feature changes

- **NemotronH / SparseMoE support**: Updated `quant_module` and
`score_module` rules (should eventually move to the proposed modeling
lib). In future, this should be the only change needed to support new
models — the bug fixes below were unearthed while enabling NemotronH
- **Config generation**: Added
`mtq.get_auto_quantize_config(search_state, constraints=None,
verbose=False)` to re-solve from `search_state` and produce plain-dict
configs (no redundant `output_quantizer`), with optional verbose summary
- **FP8 workflow split**: Use lite calibration in `mtq.auto_quantize`,
then run longer/final calibration with `mtq.quantize` using the
generated config
- **Performance**: Pass `name_to_module` to
`enable_weight_access_and_writeback` to avoid O(N^2) overhead on large
MoE models
- **Calibration caching in checkpoint**: Save/restore quantizer
calibration states (metadata + state_dict) per recipe in the
AutoQuantize checkpoint, so resuming a search skips redundant
calibration
- **Per-rank distributed checkpointing**: When `torch.distributed` is
initialized, each rank saves/loads its own checkpoint file
(`search_state{rank}.pt`), with backward-compatible fallback to the
single-file path

### API updates

- **Config API naming**: Use `mtq.get_auto_quantize_config(...)` for
exporting the searched recipe into a quantize-ready config
- **Recommended usage pattern**:

```python
# 1) Lightweight search + lite calibration
model, search_state = mtq.auto_quantize(
    model,
    constraints={"effective_bits": 6.0},
    quantization_formats=[mtq.NVFP4_DEFAULT_CFG, mtq.FP8_DEFAULT_CFG],
    data_loader=data_loader,
    forward_step=forward_step,
    loss_func=loss_func,
    num_calib_steps=64,   # lite calibration during search
    num_score_steps=128,
)

# 2) Export searched config (optionally re-solve constraints)
auto_quantize_config = mtq.get_auto_quantize_config(
    search_state,
    constraints={"effective_bits": 6.0},
    verbose=True,
)

# 3) Final / longer calibration pass with quantize
model = mtq.quantize(
    model,
    config=auto_quantize_config,
    forward_loop=long_calibration_loop,  # e.g. force-all-token style MoE calibration
)
```

### Bug fixes

- Fixed `disabled_layers` handling so fused kernels (e.g. Mamba blocks)
are properly skipped
- Fixed gradient checkpointing to keep all modules except the
checkpointed modules in eval
- Fixed FP8 fake quant NaN/inf when `amax ≈ 0`
- Fixed `SequentialQuantizer.convert_to_single_quantizer` to operate on
`module` instead of `model`, avoiding O(N^2) CPU iteration on SparseMoE
models with 1000s of submodules
- Switched to proper `F.kl_div` for KL divergence scoring

### Not yet exposed to `llm_ptq`

`mtq.get_auto_quantize_config` is not yet wired into `llm_ptq`. The
plain config records per-expert quantization settings for all MoE
experts, resulting in large JSON files. For my experiments I used a
quick workaround. A follow-up PR will add a better config representation
and expose it to `llm_ptq`.

## Testing

- Tested on NemotronH-tiny and Nemotron-Super-RL models
- Verified auto_quantize scoring + config generation end-to-end
- Unit test for checkpoint resume verifies calibration cache correctness
(metadata + tensor values)
- Existing unit tests pass

## Before your PR is "*Ready for review*"

- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: AutoQuantize end-to-end
requires GPU + large MoE models; verified manually on NemotronH-tiny and
Nemotron-Super-RL. Unit test coverage to follow.
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes
## Additional Information

Follow-up planned: expose `mtq.get_auto_quantize_config` to `llm_ptq`
with a compact config format for MoE models. AWQ support in AutoQuantize
can also be removed in a future PR to keep it lightweight.

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-10 20:59:35 +00:00
Asha AnooshehandKeval Morabia fff65b06d6 Allow HF trainer to mask sequences prior to reduction (#1009)
### What does this PR do?

Type of change: Bug fix

Previously HF trainer did not account for loss masking

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Improvements**
* Knowledge-distillation loss now properly ignores padding/special
tokens and supports masked per-token averaging.
* Default loss reduction behavior adjusted for finer-grained training
control and clearer per-token outputs.
* More robust logit handling with consistent numeric casting for
improved stability and accuracy, including mixed-precision scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-10 20:49:07 +01:00
695c8e8522 Integrate Automated QDQ placement tool - part 4.3 (#843)
## What does this PR do?

This PR upload user guide of Automated QDQ placement tool. This tool
automatically search QDQ insertion points with better performance.

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive guide for Automated Q/DQ Placement Optimization
workflow, including quick start instructions, advanced usage patterns,
configuration options, best practices, and troubleshooting.

* **New Features**
  * Exposed public API for CLI parser programmatic access.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
Signed-off-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-10 18:42:51 +00:00
Keval Morabia 0214676cc2 Pin torchprofile==0.0.4 to fix CI
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-10 23:24:46 +05:30
Keval Morabia cbab377983 Fix Mcore test utils for M-LM main (#1008)
### What does this PR do?

- M-LM main removed some functions we were using for running mcore
inference in unit tests - replace with alternative
- Future-proof `hybrid_override_pattern` -> `hybrid_layer_pattern`
rename for Mamba models

### Testing

Ran megatron tests with M-LM main branch and previous 0.16 release
version

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved Mamba model pruning to dynamically select the appropriate
hybrid pattern configuration, ensuring correct handling across different
model variants.
* Simplified inference logic for pipeline-parallel models, improving
robustness and reliability.

* **Tests**
* Updated test infrastructure for Mamba and GPT model inference
validation with explicit hidden size parameters.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-10 14:34:57 +05:30
Hrishith Thadicherla 0a84bb2caf Graph Surgery Framework for TRT-RTX (#992)
### What does this PR do?

Type of change: new feature

Adds the **ONNX Graph Surgery** framework — graph-level transformations
on exported ONNX models for optimized inference with ONNX Runtime.

**GQA Attention Replacement** — Replaces native attention subgraphs
(Q/K/V projections, RoPE, softmax, KV cache) with a single fused
`GroupQueryAttention` operator. Handles RoPE cache computation,
attention mask reformatting, Q/K/V weight fusion, and KV cache I/O
automatically from a HuggingFace model ID. Supports FP16/BF16, INT4/AWQ
quantized, and combined QKV models.

**DequantizeLinear Weight Transpose** — Transposes quantized weights in
`DequantizeLinear` nodes to column-major layout for providers like
NvTensorRtRtx. Handles INT4/UINT4 packed formats.

**Whisper Encoder Cross-Attention KV** — Adds cross-attention K/V
projection outputs to the Whisper encoder for the ONNX Runtime GenAI
pipeline. Loads cross-attention weights from HuggingFace and generates
`genai_config.json`.


### Usage

```bash
# GQA attention replacement
python -m modelopt.onnx.graph_surgery replace-gqa \
    -i model.onnx -o model_gqa.onnx \
    -m meta-llama/Llama-3.2-1B \
    --max-seq-len 4096 --dtype float16

# DequantizeLinear weight transpose
python -m modelopt.onnx.graph_surgery transpose-dq \
    -i model_quantized.onnx -o model_transposed.onnx

# Whisper encoder cross-attention KV
python -m modelopt.onnx.graph_surgery add-cross-kv \
    -i encoder_model.onnx -o encoder_with_kv.onnx \
    -m openai/whisper-large-v3-turbo
```

### Testing

Added test case for GQA graph surgery at
`tests/unit/onnx/test_gqa_graph_surgery.py`. Builds a toy attention
subgraph with native ops matching the real Optimum export pattern and
applies GQA surgery on it. Compared end outputs of both models — outputs
match exactly.

```
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_gqa_node_exists PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_gqa_attributes PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_node_count_reduced PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_rotary_emb_nodes_removed PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_position_ids_removed PASSED
tests/unit/onnx/test_gqa_graph_surgery.py::TestGQAGraphSurgery::test_logits_match
  Original nodes: 106  ->  GQA nodes: 13
  Logits shape:   (1, 4, 64)
  Original[0,:4]: [-2906.  7704. 15248.  8488.]
  GQA     [0,:4]: [-2906.  7704. 15248.  8488.]
  Max  abs diff:  0.000000
  Mean abs diff:  0.000000
PASSED
```


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* ONNX graph-surgery toolkit: new utilities to replace attention with
GQA, add encoder cross-attention KV outputs, convert FP16→BF16, and
transpose DequantizeLinear weights.
  * New CLI with subcommands to run the above transformations.
* New helper utilities for graph manipulation, RoPE cache generation,
and Whisper GenAI config creation.

* **Tests**
* Added unit tests validating GQA surgery and DQ-transpose
transformations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Hrishith Thadicherla <hthadicherla@nvidia.com>
2026-03-10 14:26:41 +05:30
Chad Voegele a56b6f356b Use os.path.join for quant_summary path (#1001)
### What does this PR do?

Type of change: Bug fix

I'm seeing the error 
```
Error saving quant summary: [Errno 2] No such file or directory: '/home/chad/checkpoint//.quant_summary.txt'
```
when using a trailing slash in the `export_path` to `hf_ptq.py`.

### Testing

This fix will work for all cases of with and without trailing slash, and
both `str` and `Path`.

```
cvoegele@nvdilw8aur4dw1a:~>> python
>>> os.path.join("/home/chad/output", "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join("/home/chad/output/", "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join(Path("/home/chad/output"), "a.txt")
'/home/chad/output/a.txt'
>>> os.path.join(Path("/home/chad/output/"), "a.txt")
'/home/chad/output/a.txt'
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- Did you write any new necessary tests?: ✅ 
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Improved internal path handling for quantization summary output to
enhance cross-platform compatibility.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
2026-03-09 15:27:04 -05:00
yueshen2016 1d6ec895ff [OMNIML-3495] Add TEGroupedMLP export support for NemotronH models (#967)
### What does this PR do?

Type of change: New feature

Add export support for `TEGroupedMLP` (fused grouped GEMM experts) in
the MCore-to-HuggingFace checkpoint exporter. Previously, the exporter
only supported `SequentialMLP` (which has `local_experts` as a
`ModuleList`). `TEGroupedMLP` stores per-expert weights as `weight0`,
`weight1`, ..., `weight{N-1}` in a single `TEGroupedLinear` module
instead. This caused an `AttributeError: 'QuantTEGroupedMLP' object has
no attribute 'local_experts'` when exporting NemotronH models.

Changes:
- Add `GroupedMLPSlicing` class in `mcore_custom.py` — the export
counterpart of `GroupedMLPMerging`
- Add `_grouped_mlp_slicing` method in `GPTModelExporter` that iterates
`TEGroupedLinear`'s per-expert weights and exports them as individual
HF-format weights with proper quantization scale handling
- Add `"experts.linear_fc1"` and `"experts.linear_fc2"` rules using
`GroupedMLPSlicing` to `nemotron_h_causal_lm_export`
- Route `TEGroupedMLP` (detected by absence of `local_experts`
attribute) to the new `"experts.linear_fc1"` rule in
`_get_transformer_layer_state_dict`

### Usage

No API change. NemotronH models using `TEGroupedMLP` can now be
exported:

```python
import modelopt.torch.export as mtex

mtex.export_mcore_gpt_to_hf(
    model=megatron_model,
    export_dir="/path/to/hf_export",
    pretrained_model_name_or_path="/path/to/hf_model",
)
```

### Testing
Inside Model-Bridge
```
torchrun --nproc_per_node 4 examples/quantization/export.py \
    --hf-model-id /models/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/ \
    --megatron-load-path /models/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4-MLM \
    --export-dir /models/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4-MLM_hf \
    --pp 4 \
    --dtype bfloat16 \
    --trust-remote-code
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ The existing `SequentialMLP`
(`local_experts`) path is guarded by `hasattr(layer.mlp.experts,
"local_experts")` and remains unchanged. The new `TEGroupedMLP` path
only activates when `local_experts` is absent and `"experts.linear_fc1"`
is defined in the architecture's rules.
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A
- Did you write any new necessary tests?: ❌ Tested manually with
Nemotron-3-Nano-30B-A3B. Unit test coverage should be added for
`_grouped_mlp_slicing`.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ New feature for a specific model architecture.

### Additional Information

- The import counterpart (`GroupedMLPMerging` / `_grouped_mlp_merging`)
was added by @jennifchen in PR #830. This PR completes the round-trip by
adding the export side.
- `_grouped_mlp_slicing` temporarily assigns `module.weight =
module.weight0` so that `_get_quantized_state` can extract
qformat/scales from the module's quantizers, then removes it afterward.
This follows the same pattern used by `_QuantTEGroupedLinear._setup()`
in the quantization plugin.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Export now supports grouped-expert MLP slicing to split fused expert
weights into per-expert tensors for downstream formats.
* Per-expert export logic enhanced with clear fallbacks between packed
and per-expert layouts, including a grouped-MLP export path.
* Nemotron H causal LM import/export mappings updated to better align
with grouped local-expert exports.
* Added fused-normalization export support and safer handling when
loading remote model code.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-03-09 12:47:12 -07:00
github-actions[bot]andKeval Morabia 2bb404ebd8 [chore]: weekly bump of uv.lock on main (2026-03-09) (#1006)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-09 19:16:45 +00:00
yeyu-nvidia 5d0e012751 inplement mix hidden_states for eagle3; deprecate eagle1 (#946)
## What does this PR do?

new feature

**Overview:** 
Enable mix hidden_states in eagle3 training. Deprecate eagle1

## Usage
Add --mix_hidden_states True to launch_train.sh

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added --mix_hidden_states option to enable optional hidden-state
mixing during training.
* Added eagle_ttt_steps setting to control speculative multi-step
iterations.

* **Chores**
* Consolidated speculative decoding to EAGLE3 only; legacy Medusa/EAGLE1
paths removed.
* Unified configuration handling so models and plugins accept a single
config object.

* **Tests**
* Updated and expanded tests for hidden-state mixing and EAGLE3-only
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-03-09 11:14:10 -07:00
Gwena Cunha 0ad287ca7b [ONNX][Autotune] Replace CUDA memory management from CUDART to PyTorch (#998)
### What does this PR do?

**Type of change**: Bug fix

**Overview**: Replace CUDA memory management from CUDART to PyTorch
(higher-level API).

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
1. Added unittests.
2. Tested that this PR does not break
https://github.com/NVIDIA/Model-Optimizer/pull/951 or
https://github.com/NVIDIA/Model-Optimizer/pull/978

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information

Summary of changes in `benchmark.py — TensorRTPyBenchmark`:
 | What changed | Before | After |
|---|---|---|
| Imports | `contextlib` + `from cuda.bindings import runtime as cudart`
| `import torch` (conditional) |
| Availability flag | `CUDART_AVAILABLE` | `TORCH_CUDA_AVAILABLE =
torch.cuda.is_available()` |
| `__init__` guard | checks `CUDART_AVAILABLE or cudart is None` |
checks `TORCH_CUDA_AVAILABLE` |
| `_alloc_pinned_host` | `cudaMallocHost` + ctypes address hack, returns
`(ptr, arr, err)` | `torch.empty(...).pin_memory()`, returns `(tensor,
tensor.numpy())` |
| `_free_buffers` | `cudaFreeHost` + `cudaFree` per buffer |
`bufs.clear()` — PyTorch GC handles deallocation |
| `_allocate_buffers` | raw `device_ptr` integers, error-code returns |
`torch.empty(..., device="cuda")`, `tensor.data_ptr()` for TRT address |
| `_run_warmup` | `cudaMemcpyAsync` + `cudaStreamSynchronize` |
`tensor.copy_(non_blocking=True)` inside `torch.cuda.stream()` |
| `_run_timing` | same cudart pattern | same torch pattern |
| `run` — stream lifecycle | `cudaStreamCreate()` /
`cudaStreamDestroy()` | `torch.cuda.Stream()` / `del stream` |
| `run` — stream arg to TRT | raw integer handle | `stream.cuda_stream`
(integer property) |
| Error handling | `cudaError_t` return codes | PyTorch raises
`RuntimeError`, caught by existing `except Exception` |

Related to https://github.com/NVIDIA/Model-Optimizer/pull/961

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* TensorRT benchmarking migrated from direct CUDA runtime calls to
PyTorch CUDA tensors, pinned memory, and CUDA stream primitives —
simplifying buffer management, transfers, and timing semantics.
* **Tests**
* Expanded GPU autotune benchmark tests with broader unit and
integration coverage for CUDA/TensorRT paths, pinned-host/device
buffering, stream behavior, warmup/timing, and end-to-end latency
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
2026-03-09 17:09:09 +00:00
d3748c2c60 Allow basename of dataset paths to match registered names (#997)
### What does this PR do?  Allow local dataset paths to match registered dataset configs

Type of change: Bug fix

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a small sample dataset entry (minipile_100_samples) and support
for loading datasets from local filesystem paths with automatic
detection and config override.

* **Chores**
* Improved local-path resolution and substring-based matching against
registered dataset keys for consistent behavior.

* **Tests**
* Added a unit test to verify loading samples from a local dataset
snapshot.

* **Documentation**
* Updated docs to describe local-path support, matching behavior, and
updated function docstring.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
Co-authored-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-09 15:24:44 +00:00
willg-nvandKeval Morabia 6d77ce754f Integrate Automated QDQ placement tool - part 4.4 (#961)
### What does this PR do?

Many minor changes:
1. Add preset mode to AutoQDQ.
2. Add pattern cache tests.
3. increase batch size for stable QDQ insertion
4. update LICENSE 2024 -> 2026.
5. add cuda-python to pyproject.toml for `[onnx]`
### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added mode presets (quick, default, extensive) with a new --mode
option for autotuning.
* Introduced AutoQDQ: automated Q/DQ placement tool for ONNX
quantization with pattern caching and checkpoint/resume.

* **Documentation**
* Updated CLI help and examples to show mode usage and override
semantics.

* **Tests**
* Added comprehensive tests for pattern cache and
mode-presets/explicit-override behavior; re-enabled a GPU autotuning
workflow test; minor test updates.

* **Chores**
  * Added "cuda-python" to optional dependencies.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-09 15:10:14 +00:00
Wei-Ming Chen 22423041f7 API to measure MSE for target quantizers (#940)
## What does this PR do?

**Type of change:** new feature ? <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

**Overview:** add an API to measure MSE for target quantizers given a
forward loop

## Usage
<!-- You can potentially add a usage example below. -->

```python
 # 1. Quantize the model as usual
model = mtq.quantize(model, quant_cfg, forward_loop)
# 2. Compute MSE for all quantizers
mse = mtq.compute_quantization_mse(model, forward_loop)
# 3. Print the top-5 noisiest quantizers
for name, err in sorted(mse.items(), key=lambda x: -x[1])[:5]:
   print(f"{name}: {err:.4e}")
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
Unit test and test with HF PTQ

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added an API to measure per-quantizer mean-squared error (MSE) between
original and fake-quantized tensors; supports wildcard and callable
filtering, skips disabled/non-fake-quant quantizers, and runs safely
under no-grad.

* **Tests**
* Added comprehensive tests for MSE validity, pattern and callable
filtering, union behavior, exclusion of disabled quantizers,
preservation of model state, and forward-hook cleanup.

* **Documentation**
  * Updated changelog to document the new MSE measurement API.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
Signed-off-by: Wei-Ming Chen <17592131+meenchen@users.noreply.github.com>
2026-03-06 22:02:49 +00:00
willg-nv be6dfad920 [5951713] Fix benchmark allocation failure (#978)
### What does this PR do?

```
[modelopt][onnx] - ERROR - Benchmark failed: Converting dtype('float16') to a ctypes type
Traceback (most recent call last):
...
    raise NotImplementedError(
NotImplementedError: Converting dtype('float16') to a ctypes type
```

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Bug Fixes**
* Improved dtype handling robustness in host memory allocation to avoid
failures for uncommon numeric types.
* Added fallback support for 2-byte floating-point formats (float16,
bfloat16); clearer errors now raised when a dtype is unsupported.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-03-06 20:35:10 +00:00
Rohan Joshi a007820afa Add CLAUDE.md (#956)
### What does this PR do?

Add CLAUDE.md file with repo overview for AI agents


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added a comprehensive CLAUDE.md documenting the Model Optimizer:
concepts, architecture, design patterns and anti-patterns, security and
contribution guidelines, common commands, architecture layout, core
abstractions (modes), key components overview, CI/testing and export
guidance, setup and workflow tips, and links to further documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
2026-03-06 19:30:42 +00:00
Keval Morabia 1ccd945a51 Remove unused diffusers/cache_diffusion/pipeline and cuda-python dependency (#996)
`cuda-python` has mixed license and needs EStaff approval for usage. And
till 0.42, it was only used in
`examples/diffusers/cache_diffusion/pipeline` which has not been updated
in 9 months and not used anymore hence removing.

Also cherry-picked to `release/0.42.0` branch:
https://github.com/NVIDIA/Model-Optimizer/pull/984

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Removed TensorRT/ONNX deployment and inference tooling, related model
export/configuration, and runtime helpers from the cache-optimized
diffusion examples; removed the cuda-python example dependency.
* **Tests**
* Removed the example benchmarking script and its associated benchmark
test.
* **Documentation**
* Strengthened dependency-review, security, and PR guidance; updated PR
template and contributing documentation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-07 00:37:19 +05:30
jingyu-ml 37d3f10cbd To support LTX2 ComfyUI format (#972)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

- Added a flag merged_base_safetensor_path to the example code so that
user can export the ComfyUI style ckpt.

### Usage

```bash
python quantize.py \
    --model ltx-2 --format fp4 --batch-size 1 --calib-size 32 --n-steps 40 \
    --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors \
    --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors \
    --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors \
    --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized \
    --extra-param fp8transformer=true \
    --quantized-torch-ckpt-save-path ./ltx-2-transformer.pt \
    --hf-ckpt-dir ./LTX2-NVFP4/ \
    --extra-param merged_base_safetensor_path=./ltx-2-19b-dev-fp8.safetensors
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Added new command-line parameters documentation for LTX-2 FP4
quantization examples (--hf-ckpt-dir and merged_base_safetensor_path
configuration options)

* **Improvements**
* Enhanced quantization pipeline to support conditional export behavior
based on model type
* Expanded LTX-Video model filtering patterns for more comprehensive
block detection

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-03-06 05:06:36 +00:00
ynankani-nv 296a865c70 sample QAD example script (#933)
## What does this PR do?
sample QAD example script

**Type of change:** ?  new example
Example script for QAD on diffusion model like ltx-2

**Overview:** ?
1) Model loading 
2) NvFP4 fake quant PTQ using mtq.quantize
3) Distillation class wrapping using mtd.convert 
4) Using ltx-2 trainer code for training 
5) Checkpoint save in bf16 . post process bf16 model using Comfy-kitchen
to produce real quantized model for ComyUI inference

## Usage
<!-- You can potentially add a usage example below. -->

```python
accelerate launch --config_file fsdp_custom.yaml sample_example_qad_diffusers.py train  --config ltx2_qad.yaml 
```

## Testing
1) Tested improvement in Vbench score for PTQ and QAD checkpoint.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: NA <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Added comprehensive README documenting the new Windows QAD example:
setup, usage, project layout, and workflow.

* **New Features**
* Added a complete Quantization-Aware Distillation training example with
distributed training config, PTQ calibration, teacher-student
distillation, CLI for training/inference, and an inference-checkpoint
creation utility.
  * Added requirements file listing needed Python packages and tooling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-03-06 04:15:54 +00:00