Commit Graph
511 Commits
Author SHA1 Message Date
Chenjie Luo a4fde491cc Update MOE block detection logic and enable in huggingface_script.sh (#962)
### What does this PR do?

Type of change: Bug fix

Add moe expert calib ratio in huggingface_script.sh
Also fix minimax2.5 MOE detection which does not follow other HF MOE
layer convention

### Usage

scripts/huggingface_example.sh --model <MiniMax-M2.5> --quant nvfp4
--moe_calib_experts_ratio 1.0 --trust_remote_code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Configure MOE calibration experts ratio for quantization via an
environment/option, enabling finer control over calibration.

* **Bug Fixes**
* Improved detection of sparse MOE blocks to handle varying
expert/topology layouts, inferring expert counts when needed for more
reliable processing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-03 21:30:35 +00:00
yueshen2016 78779f4a94 [OMNIML-3495] Fix megatron_generate to pass position_ids for MTP model support (#955)
### What does this PR do?

Type of change: Bug fix

`megatron_generate` and `megatron_prefill` set `position_ids = None` for
text-only LLM models. This crashes models with Multi-Token Prediction
(MTP) layers (e.g., Nemotron Super with `mtp_num_layers=2`) because
MTP's `roll_tensor` calls `torch.roll(position_ids, ...)` which requires
a Tensor, not None.

```
File "megatron/core/transformer/multi_token_prediction.py", line 161, in roll_tensor
    rolled_tensor = torch.roll(tensor, shifts=shifts, dims=dims)
TypeError: roll(): argument 'input' (position 1) must be Tensor, not NoneType
```

This fix generates sequential `position_ids = [0, 1, 2, ..., seq_len-1]`
per batch element.

### Usage

No API change. Models with MTP now work with `megatron_generate` without
any caller-side changes:

```python
from modelopt.torch.utils.plugins.megatron_generate import megatron_generate

# This now works for MTP models (previously crashed with TypeError)
megatron_generate(model, tokens.input_ids.cuda(), osl=1)
```

### Testing

- Verified Nemotron Super (88-layer hybrid Mamba/Attention/MoE model
with `mtp_num_layers=2`) quantization + generation works end-to-end with
`megatron_generate` on 16 GPUs (2 nodes, TP=16, EP=16).
```
torchrun --nproc_per_node 8 examples/quantization/quantize.py \
  --hf-model-id /models/nemotron-super-rl-021126/ \
  --calib-size 1 \
  --export-quant-cfg fp8 \
  --megatron-save-path /models/nemotron-super-rl-021126-FP8-MLM \
  --pp 1 \
  --tp 8 \
  --ep 8 \
  --trust-remote-code
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ All modern models (RoPE, YaRN,
Mamba) have `add_position_embedding = False`, so `position_ids` is
ignored by the embedding layer. RoPE generates identical sequential
positions internally whether `position_ids` is provided or `None`.
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A
- Did you write any new necessary tests?: ❌ Existing quantization tests
cover the code path. The fix is a one-line behavioral change (None →
Tensor) with no new branches.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ Bug fix for a previously unsupported model type.

### Additional Information

The VLM code path in `_forward_step_func` already generates
`position_ids` correctly (lines 223-227). This fix applies the same
pattern to the text-only LLM code path for consistency.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved positional ID handling in the text generation pipeline,
ensuring proper position tracking across batch processing in both
prefill and generation stages.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: James Shen <yueshen@nvidia.com>
2026-03-03 20:45:42 +00:00
Keval Morabia a076e6c218 Make GPU tests >2x faster by reusing spawn processes between tests (#958)
### What does this PR do?

Type of change: Test infra improvement <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

* Instead of spawning multiprocesses per test case (adds ~10s overhead x
150 tests), we spawn once per file (gpu_megatron) or session (gpu) to
make the tests much much faster
* Remove some unused / unnecessary tests
* Fix and enable Minitron NAS test on Hybrid models for Mcore 0.16

Test time (`pytest tests/<test_type>`) on 2x RTXPro 6000 Blackwell:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/22625069275

|            | gpu | gpu_megatron |
|------------|-----|--------------|
| **Before** | 27m | 40m          |
| **After**  | 16m | 19m          |

Full PR-merge CI/CD now only takes ~45mins

### Testing
<!-- Mention how have you tested your change if applicable. -->

Tested on 1,2,4,8 GPU setup

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Reworked GPU test infra to use a persistent distributed worker pool
with improved per-worker teardown, aggregated error reporting,
Megatron-specific cleanup, reduced GPU CI timeouts, and a small
build-system dependency change.
* **Tests**
* Converted many multi-process GPU tests to a fixture-driven distributed
runner, updating test entry points and orchestration while preserving
test logic; adjusted imports and conditional guards for several GPU
tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 20:09:18 +00:00
Keval Morabia d780fa593f Add uv.lock for experimental/dms (#964)
For nspect scanning, we need `uv.lock` with each `pyproject.toml`

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
  * Updated dependency management and lock configuration settings.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 18:39:01 +05:30
Keval Morabia 860d0b4a70 Merge Linux and Windows Changelog (#954)
Since we no longer have any compiled packages, all releases are for all
platforms so we dont have separate windows releases hence merging
changelog as well.

Going forward, windows can test on latest linux version and if fixes
needed, they can go in next usual monthly linux release

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Reorganized changelog structure to consolidate Windows and Linux
release information in a unified view.
* Expanded Windows Support documentation across recent release versions.

* **Chores**
* Updated Windows example release badge to dynamically reflect the
latest PyPI release version.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 11:31:59 +05:30
Chenhan D. Yu edde087bbd Adding a special list to handle HF models that require (#950)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. --> Bug fix

**Overview:** ?

1. Some models check `layer_types` and our offline overwrites
`num_hidden_layers=0` which will requires passing `layer_types=[]`. It
is also possible that the model will raise unexpected arguments when
passing `layer_types=[]`. So the current WAR is to use a list to handle
by checking the architectures.
2. Fix relative path issue in `launch_train.sh`



## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
  * Enhanced model loading robustness for specific layer configurations.

* **Chores**
* Improved example script path handling and GPU initialization tracking.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
2026-03-03 00:07:10 +00:00
Keval Morabia 82f1d216d1 Add Security and IP related contributing guide and configure coderabbit to catch such issues (#935)
### What does this PR do?

- Add Security related coding practices in `SECURITY.md` and merge with
`2_security.rst`
- Update `CONTRIBUTING.md` for instructions to follow if copying code
from other repositories
- Update PR template
- Cleanup dependency files
- New API `mto.load_modelopt_state` doing the insecure `torch.load(f,
weights_only=False)` instead of doing it separately everywhere. This
also allows us to later improve the input validation for
`modelopt_state_path` or use safer alternatives to `torch.load`

### Testing
<!-- Mention how have you tested your change if applicable. -->

N/A

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=True)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ <!--- Mandatory -->
- Did you write any new necessary tests?: NA <!--- Mandatory for new
features or examples. -->
- Did you add or update any necessary documentation and update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
NA <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Expanded and reorganized security guidance and contributor procedures;
updated PR template and several READMEs with clearer security,
submission, and installation instructions
* Replaced an older security document with an enhanced, centralized
security guidance

* **Chores**
* Adjusted example dependency lists and optional extras (adds, removals,
and version constraints)
* Enabled automated incremental reviews, added pre-merge security
checks, and introduced a knowledge-base of coding/security guidelines
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-03 03:49:15 +05:30
Adi Renduchintalaandcoderabbitai[bot] f26b9c3a2f skip generate option for large models and mxfp8 (#942)
## What does this PR do?

**Type of change:** New feature

**Overview:** Adds a `--skip_generate` flag to `hf_ptq.py` that skips
the pre/post-quantization generation preview calls. These calls run
`model.generate()` which crashes for very large models (500B+) that are
split across GPU and CPU via `device_map="auto"` (e.g., models with
Mamba/Triton kernels that cannot handle CPU-offloaded tensors).

## Usage
```
python examples/llm_ptq/hf_ptq.py \
    --pyt_ckpt_path /path/to/model \
    --export_path /path/to/output \
    --qformat mxfp8 \
    --trust_remote_code \
    --export_fmt hf \
    --batch_size 1 \
    --skip_generate \
    --kv_cache_qformat none
```
## Testing
Tested with a 500B parameter NemotronH hybrid Mamba/attention model on
4x GB200 GPUs. Without --skip_generate, the script crashes at
model.generate() due to Mamba Triton kernels failing on CPU-offloaded
tensors. With --skip_generate, the generation preview is skipped and
quantization proceeds normally.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
The --skip_generate flag sets generated_ids_before_ptq = None early,
which also causes the post-quantization generate to be skipped via the
existing if generated_ids_before_ptq is None: pass guard. Combined with
--batch_size 1 (to skip the get_max_batch_size forward-pass probe), this
eliminates all forward passes that can crash for device-map-split
models.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Introduced `--skip_generate` CLI option to skip pre-quantization text
and image generation, reducing processing time for very large models.
Useful when generation previews are computationally expensive.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: adithyare <adithyare@nvidia.com>
Signed-off-by: Adi Renduchintala <adithya.r@gmail.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-03-02 13:41:25 -08:00
Keval Morabia ba29ad706d Enable RTX Pro 6000 Blackwell runners for CI/CD (#944)
## What does this PR do?

**Type of change:** CI/CD Improvement <!-- Use one of the following: Bug
fix, new feature, new example, new tests, documentation. -->

Updated CI/CD test matrix (new `cuda13-gpu-trtllm` dedicated job for gpu
tests on trtllm container)

| Workflow | Trigger | Test Matrix | GPU Runner |
|---|---|---|---|
| **GPU tests** | PR | `cuda13-gpu`, `cuda13-gpu-megatron`,
`cuda13-gpu-trtllm` | 1x RTX Pro 6000 |
| **GPU tests** | Nightly | `cuda13-gpu`, `cuda13-gpu-megatron`,
`cuda13-gpu-trtllm` | 2x RTX Pro 6000 |
| **Example tests (torch)** | PR | `llm_distill`, `llm_qat`,
`llm_sparsity`, `speculative_decoding` | 1x H100 |
| **Example tests (torch)** | Nightly | `llm_distill`, `llm_qat`,
`llm_sparsity`, `speculative_decoding` | 2x RTX Pro 6000 |
| **Example tests (trtllm)** | PR | `llm_ptq`, `vlm_ptq` | 1x RTX Pro
6000 |
| **Example tests (trtllm)** | Nightly | `llm_autodeploy`, `llm_eval`,
`llm_ptq`, `vlm_ptq` | 2x RTX Pro 6000 |
| **Example tests (onnx)** | PR | `diffusers`, `torch_onnx` | 1x L4 |
| **Example tests (onnx)** | Nightly | `diffusers`, `torch_onnx` | 2x
RTX Pro 6000 |

## Testing
<!-- Mention how have you tested your change if applicable. -->

- [ ] Per PR tests pass in this PR
- [ ] Nightly tests manually triggered:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/22495679199 (GPU
tests), ? (Example tests)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Unified and simplified CI test matrices and reduced duplicate workflow
configuration.
* Updated GPU runner targets and container images for test jobs to newer
GPU types.
* **Tests**
  * Added new GPU test variants and several new test modules.
* Simplified test gating: megatron auto-skip removed; tests now use a
single dependency check for the mamba provider.
* **Tooling**
  * Added a new tox environment for an additional GPU test suite.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-02 21:00:46 +00:00
Keval Morabia 0f668a312f Replace setup.py with uv-compatible pyproject.toml and add uv.lock for nspect scanning (#945)
## What does this PR do?

**Type of change:** Project config update <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

Remove unnecessary setup.py as the same can be done with pyproject.toml.
No change for users in terms of installation command.

## Testing
<!-- Mention how have you tested your change if applicable. -->

Per PR tests passing is sufficient validation

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Migrated project configuration from setup.py to pyproject.toml,
adopting modern Python packaging standards.
* Updated CI/CD workflows and development tools to reference the new
configuration location.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-03-02 09:55:50 -08:00
willg-nv 94b2cb272d Integrate Automated QDQ placement tool - part 3.3 (#839)
## What does this PR do?

This PR implements QDQ autotuner CLI. This is the initial version of
CLI, it will be integrated to modelopt.onnx.quantization.autotune.
Usage:
```
  python -m modelopt.onnx.quantization.autotune
      --onnx_path model.onnx --schemes_per_region 50
      --pattern_cache cache.yaml --qdq_baseline baseline.onnx
      --quant_type int8 --verbose
```

PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837
PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838
PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839

**Overview:** ?

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Document will
be added in part 4.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
CHANGE log will be added in part 4.

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added a command-line interface for ONNX quantization autotuning with
configurable parameters for models, output paths, quantization
strategies, and TensorRT benchmarking.
* Introduced an automated workflow for pattern-based region optimization
with state management, baseline comparison, and benchmarking
capabilities.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-03-02 07:45:12 +00:00
sugunav14 fcdaf6519d Support decoder block-level sequential calibration (#924)
## What does this PR do?

**Type of change:** New feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** Add support for sequential calibration of layers (at
decoder level granularity) in ModelOpt.

Calibration flow
1. Get list of decoder blocks
2. For current block call get input activations (considering weight and
activation QDQ from all other previous blocks) and call specified
calibration function.

functions added
1. get_decoder_layers() -> to detect and get list of blocks to iterate
over
2. LayerActivationCollector class -> to get input activations to the
layer
3. sequential_calibrate() -> to perform the described calibration flow
4. use_sequential field in QuantizeAlgorithmConfig

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Sample config
NVFP4_DEFAULT_CFG = {
    "quant_cfg": {
        "*weight_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        "*input_quantizer": {
            "num_bits": (2, 1),
            "block_sizes": {-1: 16, "type": "dynamic", "scale_bits": (4, 3)},
            "axis": None,
            "enable": True,
        },
        **_default_disabled_quantizer_cfg,
    },
    "algorithm": {
           "method": "max",
           "use_sequential": True,
}
```
Set use_sequential=True in QUANT_CFG's "algorithm" section.

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Sequential layer-by-layer calibration: Quantization now supports
processing decoder layers sequentially to improve memory efficiency on
large models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-03-01 23:38:59 -08:00
jingyu-ml 2905cb0f2e Updated the diffusion config issue and more test cases (#937)
## What does this PR do?

**Type of change:** new tests, Bug fix <!-- Use one of the following:
Bug fix, new feature, new example, new tests, documentation. -->

**Overview:** 

- **Fixed the INT8 config issue**

- **Add HF checkpoint export test coverage**

1. The `--hf-ckpt-dir` export path had zero test coverage. This MR adds
tests at two levels:
2. Unit tests (tests/unit/torch/export/test_export_diffusers.py):
- Extended test_export_diffusers_real_quantized to parametrize over
INT8, INT8 SmoothQuant, FP8, and FP4 configs
- (previously only FP8). This gives 3 models x 4 configs = 12 test
cases.

3. GPU integration tests
(tests/gpu/torch/export/test_export_diffusers_hf_ckpt.py)
- New file testing the full quantize.py --hf-ckpt-dir pipeline via
subprocess with 4 combos:
- SDXL INT8 smoothquant min-mean (the exact scenario that triggered the
bug)
    - Flux INT8 smoothquant min-mean
    - SDXL FP8
    - Flux FP4

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**:No <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**:No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:No
<!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Tests**
* Added test coverage for exporting Diffusers models with Hugging Face
checkpoints across multiple quantization formats (INT8, FP8, FP4)
* Extended quantization export testing to validate multiple
configuration scenarios

* **Chores**
* Refined INT8 quantization configuration with improved calibrator
support for convolution layers

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-28 23:54:02 +05:30
sugunav14 56e97c8096 Bug fix 5875873 (#865)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Newer version of trl uses dtype instead of torch_dtype.
Modified code to set float32 as default for older versions of trl that
you torch_dtype.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced error handling in model training examples to safely manage
missing dtype attributes, preventing crashes during initialization when
torch_dtype is not configured.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
2026-02-28 23:52:44 +05:30
ynankani-nv 110a44c81e Add support for export ComfyUI compatible checkpoint for diffusion model(e.g., LTX-2) (#911)
## What does this PR do
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)

**Type of change:** <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

**Overview:** 
Add support for export ComfyUI compatible checkpoint for diffusion
model(e.g., LTX-2)
1) Added a a parameter for merging the base vae, vocoder, connectors in
the quantized checkpoint
2) storing quantization metadata and export tool as modelopt , required
for ComfyUI compatibility.
3) Internally updating the transformer block prefixes to match the
expectation of ComfyUI
 
## Usage
<!-- You can potentially add a usage example below. -->

```python
    export_hf_checkpoint(
        pipeline,
        export_dir=EXPORT_DIR,
        merged_base_safetensor_path=BASE_CKPT,  # merge VAE/vocoder from base
    )
```

## Testing
<!-- Mention how have you tested your change if applicable. -->
1) Tested with ltx-2 model 
       a) initializing a twoStagePipeline object 
       b) calling mtq.quantize on transformer with NVFP4_DEFAULT_CFG 
c) then exporting with export_hf_checkpoint passing the param
merged_base_safetensor_path to generate merged
           checkpoint
2) Ran the generated checkpoint with step1 on ComfyUI to validate
3) Ran step1 without merged_base_safetensor_path to check backward
compatibility.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: NA
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
NA <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for exporting LTX-2 diffusion models with merged base
checkpoint integration
* Enhanced export functionality to preserve and attach quantization
metadata during model export
* Extended model export capabilities with automatic model type detection
for improved export handling

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-02-28 07:02:34 +00:00
mxinO 6f094d7fa5 [For RL] Keep attrs after folding weight and fix empty extra state for Megatron (#779)
## What does this PR do?

**Type of change:** improvement  

**Overview:** 
- For Quantization aware reinforcement learning, after folding weight of
rollout, we want to keep the quantization attrs for next step.
- Minor fix for empty extra state
- Support getting dataloader from jsonl file, useful for using training
data as calibration data. I can separate this to another PR if
necessary.

## Usage
`mtq.fold_weight(keep_attrs=True)` will keep quantizer attrs after
folding weight,


## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: NA
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for loading dataset samples directly from JSONL/JSONL.GZ
files
* Added optional parameter to skip logits return in generation prefill
operations
* Enhanced weight folding operations to optionally preserve quantization
attributes during model optimization

* **Bug Fixes**
* Fixed handling of empty tensor states to prevent deserialization
errors in Megatron module

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Meng Xin <mxin@nvidia.com>
2026-02-28 02:53:32 +00:00
Rohan JoshiandClaude Sonnet 4.6 a538f2e256 Fix skip softmax calibration memory issue (#923)
Fix OOM issue when running skip softmax calibration

Test:

```
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
    --pyt_ckpt_path Qwen/Qwen3-30B-Instruct-A3B-2507 \
    --sparse_attn skip_softmax_calib 
```

works with >= 96GB GPU memory

---------

Signed-off-by: Rohan Joshi <rohjoshi@nvidia.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-02-27 22:48:08 +00:00
willg-nv 35e60991c7 Integrate Automated QDQ autotuner - part 3.2 (#838)
## What does this PR do?

This PR implements QDQAutotuner class. This class is used to drive the
main Autotuner workflow.

The workflow is:
1. uses RegionSearch to build regions
2. generate QDQ ONNX models and evaluate perf
3. save best model

This PR is part 2/4 of
https://github.com/NVIDIA/Model-Optimizer/pull/703.

PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837
PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838
PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839

**Overview:** ?

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: Not in this part.
- **Did you add or update any necessary documentation?**: No, document
will be updated in part 4.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No, change log will be updated when all changes are ready.

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Introduced ONNX Q/DQ autotuning framework with automatic region
discovery and pattern-based optimization.
* Added model profiling and quantization scheme generation capabilities.
* Enabled state persistence and quantization model export functionality.
* Introduced configuration management for quantization parameters and
profiling workflows.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-02-27 07:29:29 +00:00
Zhiyu a415667992 Enable Qwen3.5-MoE PTQ (#897)
## What does this PR do?

**Type of change:**  New model support

**Overview:** Add ModelOpt PTQ support for
https://huggingface.co/Qwen/Qwen3.5-397B-A17B

## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 hf_ptq.py --pyt_ckpt_path /home/omniml_data_3/models/Qwen3.5-397B-A17B --qformat nvfp4_mlp_only --export_path /home/omniml_data_3/zhiyuc/checkpoints/Qwen3.5-397B-A17B-NVFP4 --trust_remote_code
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Not yet <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Qwen3.5 Mixture-of-Experts model support in quantization
workflows.

* **Bug Fixes**
* Enhanced error diagnostics during model export with detailed module
information.
* Improved dataset tokenizer processing with proper truncation and
length handling.
  * Fixed model export stability issue related to framework integration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-02-27 00:06:07 +00:00
Zhiyu a6cbcbad0a Support mixed-precision per layer quant config in config.json (#929)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** Support mixed-precision per layer quant config in
config.json, since it's the first-class source of truth in deployment
FWs.


## Usage
<!-- You can potentially add a usage example below. -->

```python
python3 -c "
from modelopt.torch.export.convert_hf_config import convert_hf_quant_config_format
import json

# Test 1: Existing FP8 case still works
fp8_config = {
    'producer': {'name': 'modelopt', 'version': '0.29.0'},
    'quantization': {
        'quant_algo': 'FP8',
        'kv_cache_quant_algo': 'FP8',
        'exclude_modules': ['lm_head'],
    },
}
result = convert_hf_quant_config_format(fp8_config)
print('=== FP8 (existing) ===')
print(json.dumps(result, indent=2))
assert result['quant_algo'] == 'FP8'
assert 'group_0' in result['config_groups']
assert result['ignore'] == ['lm_head']

# Test 2: Mixed precision
mixed_config = {
    'producer': {'name': 'modelopt', 'version': '0.29.0'},
    'quantization': {
        'quant_algo': 'MIXED_PRECISION',
        'kv_cache_quant_algo': 'FP8',
        'quantized_layers': {
            'model.layers.0.self_attn.q_proj': {'quant_algo': 'FP8'},
            'model.layers.0.self_attn.k_proj': {'quant_algo': 'FP8'},
            'model.layers.0.mlp.gate_proj': {'quant_algo': 'NVFP4', 'group_size': 16},
            'model.layers.0.mlp.up_proj': {'quant_algo': 'NVFP4', 'group_size': 16},
            'model.layers.1.self_attn.q_proj': {'quant_algo': 'FP8'},
        },
    },
}
result = convert_hf_quant_config_format(mixed_config)
print()
print('=== MIXED_PRECISION ===')
print(json.dumps(result, indent=2))

assert result['quant_algo'] == 'MIXED_PRECISION'
assert 'config_groups' in result
assert 'quantized_layers' in result
# Should have 2 groups: one for FP8 layers, one for NVFP4 layers
assert len(result['config_groups']) == 2

# Verify per-layer detail is preserved
assert 'model.layers.0.self_attn.q_proj' in result['quantized_layers']
assert result['quantized_layers']['model.layers.0.mlp.gate_proj']['quant_algo'] == 'NVFP4'

# Check that FP8 group has correct targets
for gname, gcfg in result['config_groups'].items():
    if gcfg.get('weights', {}).get('num_bits') == 8:
        assert len(gcfg['targets']) == 3  # 3 FP8 layers
    elif gcfg.get('weights', {}).get('num_bits') == 4:
        assert len(gcfg['targets']) == 2  # 2 NVFP4 layers

print()
print('All tests passed!')
"

```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for MIXED_PRECISION quantization configurations in model
export, enabling automatic aggregation of layers by their individual
quantization settings.
* Enhanced quantization config handling to dynamically manage multiple
configuration groups for complex quantization scenarios.
* Maintained backward compatibility with existing quantization
algorithms.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-02-26 21:31:38 +00:00
Frida Hou dfe705aab5 support static NVFP4 HF export (#858)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ?
Supports export `NVFP4StaticQantizer` in unified huggingface checkpoint,
as a deployment path for PTQ algorithms such as MSE

## Usage
<!-- You can potentially add a usage example below. -->

```python
# checkpoint generation
python examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-8B  --qformat nvfp4_mse --export_path test-Qwen3-8B-Instruct-MSE-FP8-sweep-FP4 --kv_cache_qformat none --trust_remote_code 

```

## Testing
Tested generated Qwen3 8B checkpoint with trtllm serve and nv_eval
example in `Model-Optimizer-Internal/examples/nv_eval`.

NV eval results:
```
| Groups |Version|Filter|n-shot|  Metric   |   |Value |   |Stderr|
|--------|-------|------|------|-----------|---|-----:|---|-----:|
|mmlu_str|       |none  |      |exact_match|↑  |0.7186|±  |0.0036|
```


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for static NVFP4 quantizers that utilize pre-computed
calibration scales.
* Introduced new NVFP4 W4A4 quantization configuration with optional FP8
scale sweep.

* **Performance Improvements**
* Static quantizers now skip unnecessary dynamic scaling factor
recalculation.
* Unified quantization handling for improved consistency and efficiency.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
2026-02-26 10:41:29 -08:00
Gwena CunhaandKeval Morabia 3fe7e65b70 [5615343,5597780,5371126] Upgrade ORT to 1.24 (#928)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Upgrade ORT to 1.24.x to fix various bugs (5615343,
5597780, 5371126).

TODO: Verify no regressions by bumping ORT once @ajrasane is back

## Usage
See each bug.

## Testing
See each bug.

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes
* Upgraded ONNX Runtime to version 1.24.2, addressing multiple reported
issues
* Updated system requirements and installation documentation to reflect
the new ONNX Runtime version

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: gcunhase <4861122+gcunhase@users.noreply.github.com>
Signed-off-by: Gwena Cunha <4861122+gcunhase@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-27 00:03:13 +05:30
yeyu-nvidia 4a11486f6c Enable multinode training for HF speculative decoding (#922)
## What does this PR do?

**Type of change:** 
New example

**Overview:** 
Modify launch_train.sh script to enable multi-node training.
Provide a slurm template script.

## Usage
Add required fields in slurm.sh.
Then use the command below to submit multi-node job:

```bash
bash slurm.sh
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added multi-node training support with new configuration options for
distributed setups.
* Introduced Slurm batch script for streamlined job submission to
cluster environments.

* **Improvements**
* Enhanced GPU resource management for multi-GPU and distributed
training configurations.
  * Updated speculative decoding model support and validation handling.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
2026-02-26 18:23:43 +00:00
Michael Feil ef5a2dfc5d feat: Baseten contrib third-party-dataset support (#851)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?
Background: the outage on `cnn_dailymail` has a couple of issues with
out platform, with 20+ customers asking about breaking support for own
datasets.
- We still require trt-engines to support, which means e.g. the modelopt
version cannot be upgraded to latest easily.
- Applying the chat template on short context or out-of-distribution
material leads to a stronger degradation on harder or generation tasks.
This is visible when e.g. running bfcl benchmark, where the model is
generating a lot of tokens, the degration with modelopt is large, e.g.
5%+ tokens.

Going forward, we are looking to get stable support for customer-brought
datasets into modelopt, because the existing options do not suffice. A
hard validation is not desirable, many baseten users will see things
like `abisee/cnn_dailymail` is not in supported datasets, because we use
modelopt 0.35.x.

## Usage
<!-- You can potentially add a usage example below. -->
`llm_ptq --dataset baseten/quant_calibration_dataset_v1`

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Extended dataset loading to support third‑party and custom datasets,
including message-style, prompt, and text formats.
* Tokenizer-aware processing and an option to apply a chat template for
message-style datasets.
* Data-loading pipeline now propagates tokenizer context through sample
retrieval.

* **Bug Fixes / Improvements**
* Improved warnings and validation for datasets that lack expected
structures or chat-template support.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: michaelfeil <63565275+michaelfeil@users.noreply.github.com>
2026-02-26 18:34:08 +05:30
Jenny Chen 6b0ea4d43d Fix MOE layer sync test (#936)
## What does this PR do?

**Type of change:** Bug Fix

**Overview:** ?
Fix MOE layer sync test by initializing weights in MOE layer differently
[Link to
bug](https://github.com/NVIDIA/Model-Optimizer/actions/runs/22288772586/job/64472124733#step:7:851)
## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

https://github.com/NVIDIA/Model-Optimizer/actions/runs/22414958311

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Tests**
* Enhanced quantization testing with improved expert weight
initialization patterns for expert-based models.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-02-26 12:44:42 +05:30
Keval Morabia 4eacb0da72 Add trust_remote_code cli option for mbridge distillation (#934)
## What does this PR do?

Nemotron2/3 need `AutoBridge(..., trust_remote_code=True)` which was
missing previously

## Testing
<!-- Mention how have you tested your change if applicable. -->

Nemotron-nano-v2 can be distilled using tokenized
Nemotron-Pretraining-SFT-v1 data

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-25 22:30:53 +01:00
willg-nv e589ac80f6 Integrate Automated QDQ benchmark - part 3.1 (#837)
## What does this PR do?

This PR integrates benchmark module to QDQ autotunner. This benchamrk
module is used to evaluate ONNX model perf.
This PR is 1/3 of https://github.com/NVIDIA/Model-Optimizer/pull/703.
Once all small PRs are merged #703 could be closed.

PR 3.1: https://github.com/NVIDIA/Model-Optimizer/pull/837
PR 3.2 https://github.com/NVIDIA/Model-Optimizer/pull/838
PR 3.3: https://github.com/NVIDIA/Model-Optimizer/pull/839

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No, document
will be added in part 4.
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No, change log will be updated when all changes are merged.
## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added ONNX quantization autotuning capabilities with a consolidated
module providing streamlined import paths for core components.
* Introduced unified benchmarking framework supporting TensorRT-based
model evaluation with both command-line and Python API implementations.
* Added support for timing cache persistence, custom plugin libraries,
shape validation, and dynamic input shape configuration for flexible
model testing and optimization.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Will Guo <willg@nvidia.com>
2026-02-25 08:18:30 +00:00
jingyu-ml d78797b466 Update the LTX2 API calls during the calibration (#926)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Update LTX-2 integration to match latest upstream API

1. The LTX-2 codebase removed/replaced several APIs. This MR updates all
affected files:
2. Replace cfg_guidance_scale with MultiModalGuiderParams: The pipeline
__call__ no longer accepts a single cfg_guidance_scale float. It now
requires two MultiModalGuiderParams objects (video_guider_params and
audio_guider_params) that control CFG, STG, rescale, cross-modality
guidance, and skip-step settings. Updated in ltx-2.py, ltx-2-fp8.py,
ltx-2-onestage.py, calibration.py, and models_utils.py.
3. Replace fp8transformer with QuantizationPolicy: The
TI2VidTwoStagesPipeline constructor no longer accepts the fp8transformer
boolean flag. FP8 quantization is now configured via
quantization=QuantizationPolicy.fp8_cast(). Updated in ltx-2-fp8.py and
pipeline_manager.py (with backwards-compatible support for the old
--extra-param fp8transformer=true CLI flag).
4. Remove DEFAULT_CFG_GUIDANCE_SCALE constant: Replaced by
DEFAULT_VIDEO_GUIDER_PARAMS and DEFAULT_AUDIO_GUIDER_PARAMS in all
import sites.

## Usage
<!-- You can potentially add a usage example below. -->

```bash
python quantize.py --model ltx-2 --format fp4 --batch-size 1 --calib-size 1 --n-steps 40 --extra-param checkpoint_path=./ltx-2-19b-dev-fp8.safetensors --extra-param distilled_lora_path=./ltx-2-19b-distilled-lora-384.safetensors --extra-param spatial_upsampler_path=./ltx-2-spatial-upscaler-x2-1.0.safetensors --extra-param gemma_root=./gemma-3-12b-it-qat-q4_0-unquantized --extra-param fp8transformer=true --hf-ckpt-dir ./ltx2-nvfp4
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Updates**
  * Default resolution for LTX2 models adjusted to 768x1280
* Guidance parameter configuration updated for video and audio pipelines
  * FP8 quantization parameter handling refined

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-24 14:23:28 -08:00
03a1899dda Support force tokens to % of total experts during calibration (#910)
## What does this PR do?

**Type of change:** New feature

**Overview:** Adds a configurable `moe_calib_experts_ratio` parameter
that controls the percentage of experts to calibrate during the forward
pass in MoE (Mixture of Experts) models. Previously, the calibration
forward always routed tokens to **all** experts, which is expensive.
This PR allows the user to specify a ratio (default: still all experts
so no behavior change) to improve expert calibration coverage without
the cost of a full-expert forward. The token counting for the expert
coverage table now tracks the calibration routing and runs on CUDA for
efficiency.

**Changes include:**
- New `moe_calib_experts_ratio` field in `QuantizeAlgorithmConfig`
(`config.py`)
- Propagation of the ratio from the algorithm config to MoE modules
during calibration (`mode.py`)
- Updated `_QuantSparseMoe.forward` to use the configurable ratio
instead of hard-coding all experts (`huggingface.py`)
- New `--moe_calib_experts_ratio` CLI flag in `hf_ptq.py` (default
`0.25`)
- Moved `expert_token_count` tensor to CUDA and updated the HTML table
title in `moe_utils.py`

## Usage

Via hf_ptq.py CLI — calibrate 50% of experts during MoE calibration
python hf_ptq.py --model <model> --qformat int4_awq
--moe_calib_experts_ratio 0.5

Via Python API — pass the ratio through the algorithm config
import modelopt.torch.quantization as mtq

quant_cfg = {
    "quant_cfg": { ... },
    "algorithm": {
        "method": "awq_lite",
        "moe_calib_experts_ratio": 0.25,  # calibrate 1/4 of experts
    },
}
mtq.quantize(model, quant_cfg, forward_loop=calib_loop)

## Testing
Test with Qwen3 30B A3B calibration and check the tokens per expert.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added support for configurable expert calibration during Mixture of
Experts (MOE) model quantization. Users can now specify the percentage
of experts to include during calibration, enabling better expert
coverage and improved quantization accuracy for MOE models. Default: 25%
of all experts.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: realAsma <86726418+realAsma@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-02-24 12:35:19 -08:00
h-guo18 75b5da9b83 Fix: quant config error on quantized offline eagle (#925)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **Refactor**
* Enhanced quantization configuration handling for transformer models
through improved type validation, ensuring more robust processing of
quantized model configurations.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-24 09:08:34 -08:00
Izzy Putterman 2802302bf9 SpecDec Bench: February Update (#875)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 
Addition of SpecBench Dataset
Addition of NVIDID SPEED-Bench dataset, preproc scripts, and custom
metrics aggregator
Addition of example of converting SpecBench Medusa to this FW
Addition of Initial TRTLLM AutoDeploy Specdec support

Updates to all frameworks for better perf (overlap/async scheduling etc)

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added SPEED-Bench dataset support with configurable throughput and
qualitative configurations
* Introduced SpecBench metrics with acceptance rate analysis and
visualizations
  * Added progress bar during benchmark execution
* New model implementations for auto-deployment and Medusa-style
speculative decoding
  * Data preparation utility for benchmark datasets
  * Enhanced metrics with per-category analysis and performance charts

* **Documentation**
  * Updated README with SPEED-Bench workflow and examples
  * New porting guide for integrating custom benchmark runners

* **Refactor**
  * Streamlined model and runner interfaces for improved flexibility
* Consolidated dataset implementations and removed deprecated base
classes

* **Chores**
  * Added required dependencies for data handling and visualizations

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Izzy Putterman <iputterman@nvidia.com>
2026-02-24 22:24:13 +05:30
Keval Morabia c689ea1176 flush print megatron tokenization stats and update readme (#927)
## What does this PR do?

When running the script, I often see the print stats for tokenization
(every `log_interval`) not showing up or showing up very very delayed.
Hence using `print(..., flush=True)` to fix this.

Also update README that the example shown for tokenization takes too
long to run, split into multiple .jsonl files for efficiently running
the tokenization; and try out a smaller dataset first to test the script

## Testing
<!-- Mention how have you tested your change if applicable. -->

Split Nemotron-pretraining-SFT-v1 dataset into multiple .jsonl splits
and then tokenize them parallelly in different slurm jobs.

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-24 17:41:38 +01:00
Keval Morabia 52e662dd1d Fix test_transformers_tp for torch 2.10 env (#915)
After bumping CICD dev containers to latest (with torch 2.10),
`test_transformers_tp.py` is failing (was skpipped in PR-merge CICD as
it requires 2-gpu)

Failing test:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/22258743173/job/64393623736#step:7:617

Passing test after this fix:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/22259791793/job/64396179609

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization calibration by converting outputs to local
tensor representations and adding a normalization step before loss
computation, ensuring more reliable and accurate model calibration
results.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-24 02:22:11 +05:30
Keval Morabia f78385e70e Improve megatron dataset preprocessing script and update docs (#918)
## What does this PR do?

Improve megatron dataset preprocessing script and update docs

## Usage
<!-- You can potentially add a usage example below. -->

```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
    --hf_dataset nvidia/Nemotron-Pretraining-SFT-v1 \
    --hf_name Nemotron-SFT-General \
    --hf_split train \
    --hf_max_samples_per_split 10_000_000 \
    --json_keys text \
    --tokenizer Qwen/Qwen3-0.6B \
    --output_dir /path/to/tokenized/data/qwen3 \
    --workers 32 \
    --max_sequence_length 256_000
```

```python
python -m modelopt.torch.utils.plugins.megatron_preprocess_data \
    --jsonl_paths /path/to/data1.jsonl /path/to/data2.jsonl ... \
    --json_keys text \
    --tokenizer Qwen/Qwen3-0.6B \
    --output_dir /path/to/tokenized/data/qwen3 \
    --workers 32 \
    --max_sequence_length 256_000
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

- Downloaded and tokenized Nemotron-Pretraining-SFT-v1 with
Nemotron-Nano-v2 tokenizer

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated data preparation guides with new CLI patterns and Hugging Face
Hub integration instructions.

* **New Features**
* Added batch tokenization via directory input and direct Hugging Face
dataset downloads with flexible subset/split filtering.

* **Configuration Updates**
* Optimized distillation settings: adjusted optimizer parameters and
increased checkpoint retention.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-24 01:12:55 +05:30
kinjalpatel27 df47e816a7 Added support to rotate in fp32 (optional) (#885)
## What does this PR do?

**Type of change:**  New Feature

**Overview:** 
This MR adds support to perform rotation for RHT in float32 if enabled
by quantization configuration. It also makes rotate argument in
quantization configuration of type bool (for backward compatibility) or
dict (added option for float32 rotation)

## Usage
```
python hf_ptq.py --pyt_ckpt_path meta-llama/Llama-3.2-3B-Instruct --qformat nvfp4 --export_fmt hf --dataset cnn_dailymail --export_path test --trust_remote_code --inference_pipeline_parallel 1 --batch_size 1 --calib_size 4 --kv_cache_qformat nvfp4_rotate
```

Updated `NVFP4_KV_ROTATE_CFG` locally with `"rotate": {"enable": True,
"rotate_fp32": True}`

```
...
model.layers.27.self_attn.k_bmm_quantizer                                        TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=8.3750 rotated (fp32) calibrator
=MaxCalibrator quant) 
...
```

Updated `NVFP4_KV_ROTATE_CFG` locally with `"rotate": {"enable": True,
"rotate_fp32": False}`

```
model.layers.27.self_attn.k_bmm_quantizer                                        TensorQuantizer((2, 1) bit fake block_sizes={-1: 16, 'type': 'dynamic', 'scale_bits': (4, 3)}, amax=8.3750 rotated calibrator=MaxCalibrator quant)
```


## Testing
Updated unit test in `tests/gpu/torch/quantization/test_hadamard.py`

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No (updated existing test)
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added rotational input capability prior to quantization for RHT
(Rotated Hyperplane Transform).
* Introduced granular rotation configuration options enabling FP32
casting for improved numerical stability during transforms.

* **Tests**
* Expanded test coverage for rotation functionality with parameterized
FP32 casting scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-02-24 01:10:13 +05:30
Jenny Chen 02fa3623d8 Sync MOE layer input quantizer only (#903)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** in MOE layer we currently sync both the weight and input
quantizers so that all experts have the same weight amaxes and
activation amaxes.

VLLM/TRTLLM actually support non-uniform weight amaxes in MOE so we only
need to sync the activation amaxes.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved input quantizer synchronization for Mixture of Experts models
to ensure correct amax value handling across local experts.

* **Documentation**
  * Fixed typos and clarified wording in quantization documentation.

* **Tests**
* Added test coverage for Mixture of Experts quantizer synchronization
functionality.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-02-21 22:04:17 +00:00
kaix-nv 70ffb6f87c Remove test_llama_eval_sparse_attention (#914)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
Bug fix

**Overview:** ?
Fix the example test.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Tests**
  * Removed a sparse attention test variant to streamline test coverage.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-02-21 15:06:50 +05:30
Keval Morabia 9e23c6c312 Upgrade Dev containers for CICD to latest (#891)
## What does this PR do?

- Upgrade CICD test containers to latest
- Enable torch 2.10 testing in CICD

## Testing
<!-- Mention how have you tested your change if applicable. -->

CI/CD in this PR should pass


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * Added support for mixed-precision gradient handling with FSDP2.

* **Documentation**
* Updated Linux installation guide with CUDA 13.x support and cupy
dependency guidance.

* **Chores**
* Updated CI/CD workflows and test infrastructure to support PyTorch
2.10 and CUDA 13.
* Updated container image versions and test environment configurations.
  * Updated TensorRT-LLM version requirements.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-02-21 09:00:18 +05:30
Chenjie Luo 9975ba1065 Fix DeepSeek PTQ script (#912)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?

Fix two bugs in the PTQ script

## Testing

Run DeepseekV3.2 PTQ and export


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced data type handling in quantization examples for bf16
operations
* Updated internal dependencies for quantization utilities to improve
modularity

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-20 19:45:37 +00:00
Tony@NV 7c4c9fdbc9 Support multiple-batch input for autocast calibration. (#760)
## What does this PR do?

Add multi-batch calibration data support for autocast precision
conversion. This enhancement allows users to provide multiple batches of
calibration data (via a directory of NPZ files or Polygraphy JSON with
multiple batches) to aggregate tensor statistics across batches,
resulting in more robust precision conversion decisions.

## Usage

### Single NPZ file (existing behavior)
```
python -m modelopt.onnx.autocast --onnx_path model.onnx --calibration_data calibration_data.npz --output_path model_fp16.onnx
```
### Directory containing multiple NPZ files for multi-batch calibration
(new)
```
python -m modelopt.onnx.autocast --onnx_path model.onnx --calibration_data calibration_data_dir/ --output_path model_fp16.onnx
```
## Testing
- Tested with single NPZ file to ensure backward compatibility
- Tested with directory containing multiple NPZ files for multi-batch
calibration
- Verified that aggregated statistics (absmax, min, max) are correctly
computed across batches

## Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: Yes
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information

Key changes:
- Added `TensorStats` dataclass to store aggregated tensor statistics
(absmax, min_val, max_val, shape)
- Updated `ReferenceRunner` to:
  - Load multiple NPZ files from a directory (`_load_inputs_from_npz`)
  - Aggregate statistics across batches (`_aggregate_tensor_stats`)
  - Process multi-batch inference in `run()` method
- Updated `IORangeRule` and `DepthOfReductionRule` to handle both raw
numpy arrays and `TensorStats` objects
- Enhanced `--calibration_data` CLI help text to document multi-batch
support

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added multi-batch calibration support via directories of NPZ files or
Polygraphy JSON files.
* Implemented cross-batch statistics aggregation for more robust
precision conversion decisions.

* **Documentation**
* Expanded calibration_data CLI option guidance with detailed support
for multiple input formats and batch processing benefits.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Tony Yin <toyin@nvidia.com>
2026-02-20 08:42:28 +02:00
Frida Hou adcce614cb add local hessian calibration (#788)
## What does this PR do?

**Type of change:** new feature <!-- Use one of the following: Bug fix,
new feature, new example, new tests, documentation. -->

**Overview:** 
Add a new calibration method for weight scale search. It considers
activation information by weighing scale candidates with local hessian
matrix. Initial experiments with Qwen3 8B NVFP4 shows improvements.

## Usage
<!-- You can potentially add a usage example below. -->

Use `NVFP4_W4A4_WEIGHT_LOCAL_HESSIAN_CFG` quantization config for
quantization and evaluation.

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added local Hessian-weighted MSE calibration pathway for NVFP4
per-block quantization with configurable amax search parameters and FP8
scale sweep support.

* **Tests**
* Added test coverage for the new local Hessian weight-only quantization
configuration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Signed-off-by: Frida Hou <201670829+Fridah-nv@users.noreply.github.com>
2026-02-20 00:24:07 +00:00
Chenjie Luo ac7c985d96 [NVBUG: 5804406] Auto detect MOE layers (#900)
## What does this PR do?

**Type of change:** New feature, new tests

**Overview:** Replace hardcoded per-model MoE class registrations
(Mixtral, Qwen2Moe, Qwen3Moe, Qwen3Next, Llama4TextMoe, Qwen3VLMoe,
MiniMaxM2, etc.) with a single generic auto-detection mechanism
(`register_sparse_moe_on_the_fly`) that walks the model tree and
identifies MoE blocks by their structural attributes (`gate` + `experts`
with `top_k`/`num_experts`). This makes MoE quantization
forward-compatible with new HuggingFace MoE architectures without
requiring explicit registration for each model family.

Additionally, this PR:
- Tracks per-expert token routing counts during calibration via a gate
forward hook, enabling visibility into expert utilization.
- Saves an HTML report of expert token counts during export
(`save_expert_token_count_table`), highlighting under-utilized experts.
- Fixes the `topk` -> `top_k` attribute name for transformers >= 5.0
compatibility.
- Also move the ptq summary prints to a file in hf_ptq.py to reduce the
prints

## Usage

Auto-detection is transparent -- no user-facing API changes are needed.
Any HuggingFace MoE model with the standard `gate`/`experts` pattern is
automatically detected and quantized:

import modelopt.torch.quantization as mtq

# Any HuggingFace MoE model (Mixtral, Qwen3Moe, DeepSeek, etc.)
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-30B-A3B")

mtq.quantize(model, mtq.INT8_DEFAULT_CFG, forward_loop)

# During export, an .moe.html report with per-expert token counts is
saved automatically

## Testing
unittest, also test exporting qwen MOE

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added expert token count visualization for Mixture of Experts models,
exported as HTML reports during model export.
* Enhanced sparse MoE quantization with improved calibration-aware
routing and automatic model block detection.

* **Tests**
* Added comprehensive test suite for sparse MoE quantization validation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-02-19 19:46:58 +00:00
realAsma c4b662fbc8 [Bug fix] Fake quantized model save after HF accelerate hooks are added (#906)
## What does this PR do?

**Type of change:** Bug fix

**Overview:** Fix `AttributeError: Can't get local object
'add_hook_to_module.<locals>.new_forward'` when saving a quantized model
a second time after restoring it with `device_map="auto"`.

When a model is loaded with `device_map="auto"`, accelerate's
`add_hook_to_module` patches every submodule (including
`TensorQuantizer` instances) and injects three instance attributes:
`_hf_hook`, `_old_forward`, and `forward` (a `functools.partial`
wrapping a local function). These are not picklable and were leaking
into the modelopt state dict collected by `get_modelopt_state()`,
causing `torch.save` to fail.

This PR adds the three accelerate-injected attributes to
`TensorQuantizer._skip_properties_for_save_restore` so they are excluded
from the serialized state, matching the existing pattern used for
HuggingFace and DeepSpeed attributes.

## Usage

```python
mto.enable_huggingface_checkpointing()

# Quantize and save
model = AutoModelForCausalLM.from_pretrained(name, device_map="auto")
model = mtq.quantize(model, mtq.FP8_DEFAULT_CFG, forward_loop=forward_loop)
model.save_pretrained(save_dir)

# Restore and save again (this previously failed)
model2 = AutoModelForCausalLM.from_pretrained(save_dir, device_map="auto")
model2.save_pretrained(save_dir_round2)  # now works
```

## Testing

- Added unit test
`test_tensor_quantizer_modelopt_state_with_accelerate_hook` in
`tests/unit/torch/quantization/plugins/test_accelerate.py` that verifies
accelerate hook attributes are excluded from modelopt state and the
state dict remains picklable.

## Before your PR is "*Ready for review*"

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes — only adds entries to a
skip set; existing saved checkpoints are unaffected.
- **Did you write any new necessary tests?**: Yes
- **Did you add or update any necessary documentation?**: No (internal
fix, no API change)
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No

## Additional Information

The root cause is in accelerate's `add_hook_to_module`, which defines
`new_forward` as a local function and binds it via `functools.partial`
onto `module.forward`. Since local functions cannot be pickled, any
`TensorQuantizer` that has been hooked by accelerate becomes
unserializable unless these attributes are excluded.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced compatibility with accelerate library by excluding
framework-specific hooks and attributes from model state serialization,
preventing issues during save/restore operations.

* **Tests**
* Added test to validate that accelerate-related attributes are properly
excluded from model state and that the state remains picklable.

* **Public API**
  * TensorQuantizer is now publicly exported.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: realAsma <akuriparambi@nvidia.com>
2026-02-19 05:39:14 -08:00
h-guo18 eb99488da1 Fix: restore requires_grad in transformers5 reloading (#907)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Patch transformers 5.x parameter loading to preserve original
`requires_grad` settings.

In transformers v5.x, loading a checkpoint forcibly sets parameters'
requires_grad,
which unintentionally unfreeze frozen parameters (e.g. Base model in
eagle training).

This leads to optimizer initialization error since the restored
optimizer expected more parameter than the checkpoint.

This monkey-patch restores the original`requires_grad` after loading
parameters.

Reference:

https://github.com/huggingface/transformers/blob/v5.0.0.rc1-release/src/transformers/core_model_loading.py#L640

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed model parameter loading in speculative decoding to properly
preserve gradient requirements for each parameter when using HuggingFace
Transformers 5.x, ensuring correct behavior during checkpoint resumption
and model initialization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-19 01:25:47 +00:00
jingyu-ml 3dd52bf106 Diffusion export bug fixed for model_index.json (#901)
## What does this PR do?

**Type of change:** Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 

Updated diffusers export to preserve the original model_index.json
instead of always rebuilding a minimal one.
The export now uses a simple fallback order: copy original
model_index.json from source path if available, otherwise call
`pipe.save_config(export_dir)`, and only then generate a minimal
model_index.json as last resort.
Non-diffusers export behavior is unchanged.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved Diffusers pipeline export with enhanced model configuration
handling. The export process now better preserves original pipeline
configurations and uses fallback strategies to ensure complete
configuration files are generated.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-02-18 23:43:48 +00:00
h-guo18 b8a4586702 Refactor: Eagle data loading (#668)
## What does this PR do?

**Type of change:** Refactor <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** 
Jira ticket: https://jirasw.nvidia.com/browse/OMNIML-2955

Main changes :

- Consolidate Eagle data loading with @ChenhanYu 's implementation of
`transformers_dataset.py`

- Refactor: baked the following logics from `example/main.py` to
`modelopt/torch` for cleaner example entrance:
  - default config selecting and merging with custom config
  - tokenizer post-processor (chat template and pad_tok_id)
  - d2t loading 
- Implementation refactor: In HF workflow, reuse base modfel's input
hidden states as input_embedding, instead of calculating from input_ids.
This has two main benefits:
    - Easier VLM support, which has various embedding processing logics.
    - Training effieicy.
    
- Deprecating eagle1 from the example. It is still available by setting
custom config.
  
 - Other minor fixes and readme updates.

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

Tested that training curves after changes (both online&offline) is
identical with original branch:
<img width="1073" height="634" alt="image"
src="https://github.com/user-attachments/assets/abfd7bea-c82c-48a7-8181-68c5a9e4da8d"
/>


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added draft vocabulary cache support for EAGLE model training,
enabling runtime vocabulary customization via `--draft_vocab_cache`
parameter
* Introduced new data loading utilities with sharding, streaming, and
tokenization support for large-scale training
  * Added optional `--log_steps` configuration to training launcher

* **Documentation**
* Updated EAGLE configuration guides with draft vocabulary cache setup
instructions and examples

* **Refactor**
* Restructured data pipeline for offline training with improved dataset
handling and batching
* Updated command-line arguments across training scripts (`--input-data`
replaces `--input-file`)

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-02-18 14:04:06 -08:00
kaix-nv 9e38041d34 [OMNIML-2850] [3/n] Adds sparse attention calibration (#538)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature

**Overview:** ?
- This PR adds the sparse attention calibration algorithm
- Chunked prefill to support long ctx_len
- Separated calibration for prefill and decode

## Usage
<!-- You can potentially add a usage example below. -->

```python
import modelopt.torch.sparsity.attention_sparsity as mtsa

# Apply sparse attention with calibration
model = mtsa.sparsify(model, config=SKIP_SOFTMAX_CALIB)

# Print summary - now shows actual thresholds
mtsa.print_sparse_attention_summary(model)
# Output:
# Method: flash_skip_softmax, Threshold: Dynamic (λ=437.395926)

# Or llm_eval integration
# HuggingFace sparse attention example
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
    --pyt_ckpt_path Qwen/Qwen3-4B \
    --sparse_attn skip_softmax_calib 
```

# The calibration method
## Calibration Algorithm
- Implemented the Inverse Power model: scale_factor = k / (1 -
sparsity)^p
- Fit model parameters (k, p) per phase using scipy.optimize.curve_fit
- At inference: threshold = k / (1 - target_sparsity)^p / seqlen

## Why Choosing the Inverse Power model?
The inverse power model better fits the relationship between sparsity
ratio and threshold_scale_factor.
<img width="2388" height="1082" alt="sparsity_model_analysis"
src="https://github.com/user-attachments/assets/4dfb45d4-8c16-4f15-a878-c8e08a9b6128"
/>

## Runtime Flexibility
- Target sparsity can be changed at inference time without recalibration
- Users can adjust module._sparse_method_instance.target_sparse_ratio
dynamically
- Threshold automatically adapts to sequence length

## Testing
<!-- Mention how have you tested your change if applicable. -->
The calibration results for `Qwen/Qwen3-30B-A3B-Thinking-2507` are shown
below and are mostly consistent with the ground-truth numbers collected
from the kernel side.
```
Prefill Calibration Results:
  Model: scale_factor = k / (1 - sparsity)^p
  Fitted k: 1003.3990
  Fitted p: 1.2589
  R-squared: 0.827549

Scale factors for different target sparsities:
  Target     Scale Factor
  ---------- ---------------
  50%        2401.35
  70%        4568.26
  80%        7610.98
  90%        18214.70
  95%        43591.65
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-02-18 02:18:12 +00:00
Chenjie Luo 3801923e9d Support MiniMax M2.1 (FP8 checkpoint) (#817)
## What does this PR do?

**Type of change:** ? new feature

**Overview:** ?

Support loading the MiniMax M2.1 (FP8) checkpoint for PTQ.

## Usage
scripts/huggingface_example.sh --model <minimax checkpoint> --quant
nvfp4 --trust_remote_code


## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
  * Added MiniMax M2.1 model quantization support with nvfp4 format.
* Extended FP8 quantization capabilities with configurable dtype
parameter for enhanced precision control.

* **Improvements**
  * Enhanced detection of quantized linear module variants.
  * Improved weight unpacking for FP8-based linear modules.

* **Documentation**
  * Updated supported models table to include MiniMax M2.1.

<sub>✏️ Tip: You can customize this high-level summary in your review
settings.</sub>

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>
2026-02-17 18:47:22 +00:00
Jenny Chen 590f9fc662 Mamba MOE Quant Configs + Fix Export Bug (#882)
## What does this PR do?

**Type of change:** ? Bug fix

**Overview:** ?
- Fix a bug in MCore export `exclude_modules` where the layers had an
extra period at the end
- Add custom quant configs for mamba moes

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added four new Mamba MOE quantization configurations: aggressive and
conservative variants for both FP8 and NVFP4 quantization schemes,
providing enhanced flexibility in quantization options for different use
cases.

* **Bug Fixes**
* Improved quantization export module exclusion pattern handling to
properly normalize trailing dots from exclude patterns during export.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-02-17 16:45:41 +00:00
Gal Hubara-Agam 9763505981 [fix][5889686] AutoCast: Fix logger (#890)
## What does this PR do?

**Type of change:** Bug fix
**Overview:**
Previously relied on quantization logger, which caused logs to be
suppressed when onnx.autocast was used directly
Instead:
- Inherit format and level if called from onnx.quantization
- Configure independent format and level if called from onnx.autocast

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes <!--- If No, explain why.
-->
- **Did you write any new necessary tests?**: No
- **Did you add or update any necessary documentation?**: No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
No <!--- Only for new features, API changes, critical bug fixes or bw
breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved logging configuration to ensure consistent behavior across
modules with enhanced file and console output management.
* Fixed log file handling to automatically create required directories.
  * Enhanced logger propagation logic for more reliable output routing.

* **Chores**
* Refined logging initialization for better automatic configuration at
startup.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Gal Hubara Agam <96368689+galagam@users.noreply.github.com>
2026-02-17 14:16:00 +02:00