Commit Graph
741 Commits
Author SHA1 Message Date
Chenjie LuoandKeval Morabia 168cd828c1 Add qwen3 moe experts only test (#1274)
## Summary
- Add unit test for Qwen3 MoE HF export with `NVFP4_EXPERTS_ONLY_CFG`
quantization config
- Verifies that `hf_quant_config.json` correctly reports `quant_algo:
NVFP4` and that non-expert modules (`self_attn`, `lm_head`) appear in
`exclude_modules` while routed expert layers (`mlp.experts.*`) do not
- Reference:
https://huggingface.co/nvidia/Qwen3.5-397B-A17B-NVFP4/blob/main/hf_quant_config.json

Type of change: New tests

### Known issue
On `transformers>=5.0`, fused MoE experts (`_QuantFusedExperts`) are not
recognized by `get_quant_config`, causing `quant_algo=None` in the
exported config. This test currently **fails** on transformers 5.x and
is intended to be fixed by a follow-up change.

## Testing
- **transformers 4.57.6**: PASSED
- **transformers 5.5.4**: FAILED (`quant_algo` is `None` due to fused
expert export gap)

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Tests**
* Added GPU test coverage for exporting Qwen3 Mixture-of-Experts models
with NVFP4 quantization.
* Verifies the exported checkpoint records the NVFP4 quantization
algorithm and that module exclusion patterns correctly exclude attention
and LM head components while not excluding routed expert paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-05-01 17:56:04 +00:00
Chenjie Luo 3ad4f4f093 [Fix] Re-expand target_input on OOM in get_max_batch_size (#1374)
## Summary
- `get_max_batch_size` halved `target_data_batch` on
`torch.cuda.OutOfMemoryError` but never rebuilt `target_input`, so each
retry re-fed the same too-large tensor — the retry loop was effectively
a no-op.
- Refactor the expand logic into an `_expand_to(batch)` helper, rebuild
`target_input` after halving, and call `torch.cuda.empty_cache()`
between attempts.

## Test plan
- [x] New unit test `test_get_max_batch_size_oom_retry_shrinks_input`
mocks `torch.cuda.*` and asserts the second retry receives the halved
tensor (shapes seen: `[1, 10, 5]`, regulated result `4`).
- [x] `pytest tests/unit/torch/utils/test_dataset_utils.py` — 14/14 pass
(skipping the network-only minipile test).
- [x] `pre-commit` (ruff, mypy, bandit, license headers) clean on
commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Enhanced GPU memory management during batch size detection. When
out-of-memory errors occur during the initial probing phase, the system
now properly adapts input tensors to smaller batch sizes and clears GPU
cache before retry attempts, resulting in more reliable recovery and
stable batch sizing across diverse hardware environments.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-30 18:25:28 +00:00
Keval MorabiaandClaude Sonnet 4.6 bb08094ff1 Add Nemotron-Nano-9B-v2 → Pruned 7B e2e tutorial: Prune + Distill + Eval + Quantize + vLLM deployment (#1325)
## Summary

End-to-end optimization walkthrough for Nemotron-Nano-9B-v2 showing how
ModelOpt techniques stack:

- **Pruning** — Minitron structured pruning 9B → 7B
- **Distillation** — Megatron-Bridge knowledge distillation up to 80B
tokens; near-parity with official 9B on MMLU Pro, GPQA, LCB, AIME, Math
500, IFEval, SciCode
- **Evaluation** - using nemo-evaluator
- **Quantization** — FP8 PTQ via \`hf_ptq.py\`; checkpoint deployable on
vLLM/TRT-LLM/SGLang with no extra flags (quantization auto-detected from
\`config.json\`)
- **vLLM Throughput** — BF16 vs FP8 benchmark on single H100

<img width="2085" height="1740" alt="image"
src="https://github.com/user-attachments/assets/8620a019-5c09-4a6b-a5d2-ca164aaa5d87"
/>

<img width="2085" height="810" alt="image"
src="https://github.com/user-attachments/assets/742c8035-f1fb-4394-b11b-0c6c3ac4e843"
/>


### Files changed

- `examples/pruning/minitron/README.md` — index page for Minitron
end-to-end tutorials
- `examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/README.md` —
full repro doc with 6 sections: data prep, pruning, distillation,
evaluation, FP8 quantization, vLLM benchmarking
-
`examples/pruning/minitron/NVIDIA-Nemotron-Nano-9B-v2/nemo_evaluator.yaml`
— NeMo Evaluator config used for all benchmark numbers
- `examples/pruning/puzzletron/README.md` — index page for Puzzletron
distillation results
- `examples/pruning/puzzletron/Llama-3.1-8B-Instruct.md` — Puzzletron
distillation results (renamed from puzzletron.md)
- `examples/pruning/README.md` — updated Results section with direct
links to new locations
- `examples/megatron_bridge/README.md` — updated results link to point
to `examples/pruning/`
- `examples/puzzletron/README.md` — updated distillation results link
- `examples/dataset/MEGATRON_DATA_PREP.md` — tokenization commands for
all datasets used in the data blend

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Documentation

* **New end-to-end tutorial** for model optimization covering Minitron
pruning, knowledge distillation, FP8 quantization, and vLLM deployment
with reproducibility steps and benchmark results
* **Dataset preparation guide** with ready-to-run tokenization templates
for Nemotron HuggingFace datasets
* **Evaluation configuration** and results documentation including
ablation studies across multiple benchmarks
* **Updated navigation** across pruning, distillation, and dataset
examples to streamline user workflows

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-30 14:47:16 +05:30
dthienan-nv 378366b8c9 [6034518] Remove return statement preventing remote auto tuning (#1361)
### What does this PR do?

Remove return statement from the code checking remote auto tuning config
arguments since that results in skipping adding the actual remote tuning
config to the trtexec cmd.

**Root cause**: The necessary flags do not get added to
`self._base_cmd.extend(trtexec_args)` when remote autotuning is enabled.

**Before fix**:
```
['trtexec', '--avgRuns=100', '--iterations=100', '--warmUp=50', '--stronglyTyped', \
    '--saveEngine=engine.trt', '--timingCacheFile=trtexec_timing.cache', \
    '--onnx=baseline.onnx']
```
**After fix**:
```
['trtexec', '--avgRuns=100', '--iterations=100', '--warmUp=50', '--stronglyTyped', \
    '--saveEngine=engine.trt', '--timingCacheFile=trtexec_timing.cache', \
    '--remoteAutoTuningConfig=$CONFIG', '--safe', '--skipInference', \
    '--onnx=baseline.onnx']
```
Notice that the remote autotuning and related flags are now included in
the `trtexec` command.

**Related PR**: https://github.com/NVIDIA/Model-Optimizer/pull/1259

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes
* Fixed an issue where remote autotuning configuration arguments were
not being properly included in benchmark commands, ensuring all remote
autotuning settings are now correctly applied during execution.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: dmoodie <dmoodie@nvidia.com>
2026-04-30 04:32:18 +00:00
vishalpandya1990 a492fa9a14 Ensure removal of temp files on error in ONNX INT4 quantization (#1359)
### What does this PR do?

Type of change: Minor bug fix


- Put quantization steps inside try-finally to ensure removal of temp
files on error in ONNX INT4 quantization.
- To avoid redundancy between awq_lite() and awq_clip() methods, created
a utility _remove_augmented_onnx() for exception-handling based removal
of augmented onnx file and its data file.

### Testing

- Locally performed ONNX INT4 awq-lite and awq-clip quantization with
Llama 1B model.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Improved reliability of the quantization pipeline by ensuring
temporary conversion artifacts are always removed, making cleanup more
robust.
* Consolidated handling of external-data companions and added safer
deletion behavior that logs failures instead of raising errors.
* Ensured consistent session teardown and forced memory collection to
reduce resource leakage and intermittent errors during model conversion.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: vipandya <vipandya@nvidia.com>
2026-04-30 09:22:36 +05:30
h-guo18 c07ac215ef [Fix]: Relax Dflash Rregression Test Threshold fo 2GPUs (#1373)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

The offline dflash regression test can be runned on 1 or 2 gpus. For 2
gpus, the total steps is half of 1 gpu. This PR relax the failing
threshold for 2 gpu tests.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-29 20:11:13 +00:00
ynankani 9bb917d57c [BUG6108338] Update windows documentation for onnxruntime quantization with Cuda13.x (#1368)
### What does this PR do?

Type of change: ? documentation

<!-- Details about the change. -->
Update windows documentation for onnxruntime quantization with Cuda13.x




### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A <!---
Mandatory -->
- Did you write any new necessary tests?: N/A <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated Windows installation guide with CUDA 13.x-specific setup
instructions for GPU-accelerated dependencies, including CuPy and ONNX
Runtime configuration with nightly builds.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ynankani <ynankani@nvidia.com>
2026-04-29 16:15:40 +05:30
Wei-Ming Chen 077e29a48c [NVBug 6108145] Fix PTQ calibration and export for fused-experts MoE (Qwen3.5-MoE VLM) (#1340)
### What does this PR do?

Type of change: Bug fix

Fixes a 4-bug cascade that caused silent PTQ failure on Qwen3.5-MoE VLMs
(Qwen3.6-35B-A3B): calibration
appeared to succeed but produced token-salad at inference. Root cause:
HF's @use_experts_implementation
dispatches expert forward to torch._grouped_mm / torch.bmm, bypassing
the F.linear hook that captures
activations — so gate_up_proj_input_quantizer /
down_proj_input_quantizer never calibrated and no input_scale
  tensors were emitted.

  Changes:
- examples/llm_ptq/hf_ptq.py — force config._experts_implementation =
"eager" (recursing into text_config /
vision_config / …) so per-expert F.linear calls are visible to the
calibration hook.
- modelopt/torch/quantization/conversion.py — normalize plural
ModuleList quantizer names (weight_quantizers.N
→ weight_quantizer) before fnmatch, so wildcards like
*mlp.experts*weight_quantizer match fused-expert
  quantizers.
- modelopt/torch/export/unified_export_hf.py — hoist the
_QuantFusedExperts export branch above the
get_quantization_format() gate so _export_fused_experts() runs even when
the top-level format query returns
  QUANTIZATION_NONE (happens for experts-only recipes).
- modelopt_recipes/general/ptq/nvfp4_experts_only-fp8_kv.yaml —
layerwise: false (VLM nested layer structure
  breaks the layerwise walker).

<!-- Details about the change. -->

### Usage

```python
  python examples/llm_ptq/hf_ptq.py \
      --pyt_ckpt_path Qwen/Qwen3.6-35B-A3B \
      --qformat nvfp4 \
      --kv_cache_qformat fp8 \
      --calib_size 512 \
      --export_path Qwen3.6-35B-A3B-NVFP4
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

  Testing

End-to-end PTQ → vLLM deploy → NEL eval on Qwen3.6-35B-A3B (256 experts
× 40 layers, 35B params):

Hook-call diagnostic: 0 → 6720 per-expert F.linear calls during
calibration after the fix; 0 → 30720
  input_scale tensors emitted in the exported checkpoint.

FP8 fused-MoE path still produces gibberish — separate follow-up (vLLM
per-expert weight_scale handling).
* vLLM full-FP8: the FlashInfer TRTLLM Fp8MoE loader doesn't stack the
256 per-expert scalar weight_scale tensors
into a [num_experts] per-expert vector — it ends up applying one
expert's scale across all 256, so every
routed expert dequants with the wrong amplitude → coherent token stream
collapses into multilingual gibberish.

* SGLang full-FP8: qwen3_5.py::_make_packed_weight_loader rejects with
AssertionError: Unexpected scalar for
tuple shard load: loaded_shard_id=(0,1,2), split_sizes=[1,1,1] — its
packed-loader has no path for "N
independent per-tensor source scalars combining into one fused-shard
parameter," so the fused QKV (or
in_proj_qkvz) load is structurally refused and the model never finishes
loading.
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Better fused-expert export flow, a plugin to force eager expert
execution during calibration/export, and a representative quantizer
discovery utility.

* **Bug Fixes**
* Reliable matching/discovery of per-expert indexed quantizers enabling
correct calibration and mixed-precision export; fixes for calibration in
nested decoder layouts.

* **Documentation**
  * Clarified PTQ config guidance on layerwise calibration.

* **Tests**
* Added fused-experts calibration, export, and name-normalization tests.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-04-29 00:01:01 -07:00
Wei-Ming Chen e5ce0ae83a [NVBug 6102977] Add _disable_use_cache context manager to fix PTQ AttributeError on custom configs (#1324)
### What does this PR do?

Type of change: Bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

- Summary: Running hf_ptq.py on stepfun-ai/Step-3.5-Flash (and any model
whose custom HF config doesn't assign use_cache) crashed in
get_max_batch_size() with AttributeError: 'Step3p5Config' object has no
attribute 'use_cache' before calibration could start.
- Extract the existing "disable KV cache during calibration" logic into
a _disable_use_cache(model) context manager, apply it to both
get_max_batch_size and _forward_loop. The CM sets config.use_cache =
False unconditionally (not only when the attribute exists) and restores
the prior value on exit if one was set.
- Behavior unchanged for normal configs; the NemotronH hybrid-cache
correctness guarantee from #1251 is preserved.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

Step-3.5-Flash PTQ now passes get_max_batch_size


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Improved memory handling during model evaluation and calibration by
consistently disabling KV cache for both single-batch probes and full
dataloader runs, simplifying and stabilizing inference flow and ensuring
cache state is managed reliably.

* **Tests**
* Added unit tests verifying cache-state handling across models with and
without cache settings, including correct restoration behavior even when
errors occur.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
2026-04-28 22:22:50 +00:00
Jenny Chen 8eec6d4459 Support EP mcore import for TE Spec and Fix mamba moe config (#1342)
### What does this PR do?

Type of change: Bug fix

- Enable EP (expert parallelism) import for HF to MCore when using TE
Spec
- Fix bug in mamba moe config which doesn't skip attention layers
properly in MCore (Mcore uses different naming for attention layers than
HF)
- Add getter for Quant Config (used in MLM modelopt examples to get
quant cfg fields)

### Usage

```python
# In Megatron-LM/examples/post_training/modelopt
 MLM_EXTRA_ARGS="--export-default-te-spec --trust-remote-code --moe-router-dtype fp32" EP=4 HF_MODEL_CKPT=</path/to/hf> MLM_MODEL_SAVE=<save/path> ./convert.sh nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Corrected expert-slice assignment so each expert-parallel rank loads
the proper expert slice.
* Improved detection of pipeline-parallel layer indices in submodule
names.

* **Improvements**
* Relaxed constraints between local and global expert counts for
grouped-local-expert imports.
* Added typed helpers for managing quantization configuration entries
and expanded quantizer disable patterns.
  * Exporter now accepts an additional hybrid model type when available.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
2026-04-28 06:39:41 -07:00
Grzegorz K. Karch 6d330784a6 Add required keys to attention pruning config (#1360)
### What does this PR do?

Type of change: ? Bug fix

The config
`examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/pruning/attn_pruning.yaml`
didn't have required keys to use attention pruning in the example
`examples/puzzletron/main.py`

### Usage


### Testing

In
`examples/puzzletron/configs/llama-3_1-8B_pruneffn_memory/Llama-3_1-8B.yaml`
change `ffn_pruning` to `attn_pruning`

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Updated pruning configuration for improved KV-head pruning support,
including enhanced importance hook settings and attention output
handling for memory optimization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
2026-04-28 13:04:51 +00:00
Asha Anoosheh 6e08b13fe6 Fix regex capture for Megatron KD PP layer renaming (#1355)
### What does this PR do?

Type of change: Bug fix

Previously the regex we had looked for a dot after the integer layer
number, but it might not exist sometimes.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved detection and handling of pipeline-parallel layer indices in
model names to correctly support layer identifiers positioned at the end
of submodule names, enhancing compatibility with various naming
conventions in distillation workflows.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Asha Anoosheh <aanoosheh@nvidia.com>
2026-04-27 20:42:45 +00:00
Chenjie Luo 47a33db9b6 [NVBUG: 6103846] Fix nvfp4_awq export for uncalibrated MoE experts (#1354)
## Summary

- NVBug: [6103846](https://nvbugspro.nvidia.com/bug/6103846) —
`Qwen3-30B-A3B nvfp4_awq` quantization fails at export with
`AssertionError: Modules have different quantization formats`.
- Root cause: in `model_calib.awq_lite`, MoE experts that end up
disabled (NaN in act/weight scales, or no search-pass tokens) get
`max_calibrate`-d but no `pre_quant_scale`. `get_quantization_format`
then returns `nvfp4` for those experts while siblings stay `nvfp4_awq`.
`unified_export_hf.requantize_resmooth_fused_llm_layers` groups all 128
experts of each linear name (gate_proj/down_proj/up_proj) and calls
`preprocess_linear_fusion(..., resmooth_only=True)`, which asserts
uniform format → fires for any single mismatched expert.
- Fix: unify the disabled-expert paths in the awq_lite postprocess loop
so any expert with `is_enabled == False` (no cache hits, NaN scales, or
no search-pass tokens) receives `max_calibrate` + a neutral all-ones
`pre_quant_scale`, matching the existing behavior for `num_cache_steps
== 0`. Emit a warning so users notice that calibration coverage is
incomplete and accuracy may degrade.

## Test plan

- [x] `pytest tests/unit/torch/quantization/test_calib.py -k 'awq'` → 5
passed
- [x] End-to-end on `Qwen/Qwen3-30B-A3B` with `NVFP4_AWQ_LITE_CFG` and a
small calib set that leaves many experts uncalibrated:
- All 6144 gate_proj/up_proj/down_proj expert linears report `nvfp4_awq`
(no mismatch)
  - `export_hf_checkpoint` succeeds with no `AssertionError`
- The new "Forcing pre_quant_scale=1 ... may degrade accuracy" warning
fires for each affected expert
- [x] Re-run via `examples/llm_ptq/hf_ptq.py` with the bug-report CLI
(cnn_dailymail, batch_size=8, calib_size=64 — scaled down from 512 to
fit budget) on B200:
- 36 "the second time did not forward data through
..experts.X.{gate,up,down}_proj" warnings — i.e. the exact
bug-triggering condition from the original NVBug log naturally
reproduces
- 2058 "Forcing pre_quant_scale=1" warnings — fix path activates for
uncalibrated/disabled experts
  - 0 `AssertionError`s — export completes
- `Quantized model exported to: /tmp/test_plan_qwen3-30b-a3b-nvfp4_awq`
and post-PTQ generation works

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-27 12:40:14 -07:00
Keval Morabiaandcoderabbitai[bot] 4f378ad275 Add release-cherry-pick Claude Code skill (#1352)
Adds `.claude/skills/release-cherry-pick/SKILL.md` — a Claude Code skill
for cherry-picking labeled PRs to a release branch.

Invoke with `/release-cherry-pick <version>`.

See this PR created with the skill:
https://github.com/NVIDIA/Model-Optimizer/pull/1350

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added automated release cherry-pick workflow to streamline selecting
and applying multiple PRs into release branches.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-04-28 00:42:50 +05:30
Keval Morabia 1306841989 Enable SonarQube Static Application Security Testing (SAST) (#1349)
Enable SonarQube as a Nvidia recommended and more comprehensive code
scanning tools compared to Bandit we currently use in pre-commit hook
(still left for now)

Tested pipeline in internal gitlab and it works and results are uploaded
in internal SonarQube website

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Added CI jobs to run SonarQube analysis and generate a vulnerability
report, with scheduled and branch-triggered runs.
* Configured scans to preserve full git history, use caching, and
auto-cancel interruptible runs.
* Added an ignore rule to exclude generated analysis artifacts from
version control.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-28 00:34:27 +05:30
fe3042b5bb fix incomplete mapping of safetensors in generated puzzletron checkpoint (#1330)
### What does this PR do?

Type of change: ? Bug fix

Fixes
`https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/puzzletron/main.py`
where multi-GPU run caused only part of the file
`model.safetensors.index.json` to be written to disk.

### Usage

does not apply

### Testing

Follow [instructions, step
3](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/puzzletron#compress-the-model)
- run with `--nproc_per_node 2`

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a public checkpoint-saving entry that consolidates distributed
sharded model shards into a single filesystem checkpoint; retains direct
saving for single-process runs.

* **Refactor**
* Validation/evaluation tooling now uses the consolidated
checkpoint-saving flow when persisting realized model checkpoints during
runs.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Signed-off-by: Grzegorz K. Karch <grzegorz-k-karch@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-27 18:01:27 +00:00
Ajinkya Rasane 5b41ba4c34 chore: Move FP8 MHA quantization entry from 0.45 to 0.44 in CHANGELOG (#1351)
## Summary

PR #1289 (FP8 MHA quantization for ViT) was merged to `main` after
`0.44.0rc1` was tagged, so the rc1 wheel ships without the
`nn.LayerNorm` registration that the example's `_FP8_MHA_OVERRIDE` now
references — surfaced as nvbug 6114983 (`ValueError: parent_class
'nn.LayerNorm' not found in QuantModuleRegistry` when running
`torch_quant_to_onnx.py --quantize_mode=fp8`). PR #1289 is labeled
`cherry-pick-0.44.0` and will be cherry-picked to `release/0.44.0` for
the next rc, so the feature ships in 0.44 — this PR moves the
corresponding release-notes bullet from the `0.45 (Future)` section to
`0.44 (2026-05-xx)` to match.

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-04-27 14:39:43 +00:00
github-actions[bot]andgithub-actions[bot] 06860061c6 [chore]: weekly bump of uv.lock on main (2026-04-27) (#1347)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

<details>
<summary>uv lock --upgrade output</summary>

```
Using CPython 3.12.13 interpreter at: /opt/hostedtoolcache/Python/3.12.13/x64/bin/python3
Resolved 180 packages in 7.12s
Updated certifi v2026.2.25 -> v2026.4.22
Updated click v8.3.2 -> v8.3.3
Updated filelock v3.28.0 -> v3.29.0
Updated huggingface-hub v1.11.0 -> v1.12.0
Updated idna v3.11 -> v3.13
Updated onnx-ir v0.2.0 -> v0.2.1
Updated onnxscript v0.6.2 -> v0.7.0
Updated packaging v26.1 -> v26.2
Updated pathspec v1.0.4 -> v1.1.1
Updated pyarrow v23.0.1 -> v24.0.0
Updated pybind11 v3.0.3 -> v3.0.4
Updated pydantic v2.13.2 -> v2.13.3
Updated pydantic-core v2.46.2 -> v2.46.3
Updated pydantic-settings v2.13.1 -> v2.14.0
Updated transformers v5.5.4 -> v5.6.2
Updated typer v0.24.1 -> v0.25.0
Updated tzdata v2026.1 -> v2026.2
Updated uvicorn v0.44.0 -> v0.46.0
Updated wheel v0.46.3 -> v0.47.0
Updated xxhash v3.6.0 -> v3.7.0
```

</details>

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-04-27 10:46:40 +00:00
Zhiyu 4b5549b663 [1/N] Polish evaluation skills and common skills based on an E2E workflow testing, vendor two Claude skills from NeMo Evaluator (#1239)
### What does this PR do?

Type of change: Documentation / Skills

Polishes evaluation and common skills based on end-to-end experience
quantizing + deploying + evaluating LLMs. Vendors the two upstream
Claude skills from NeMo Evaluator, splits shared credential setup into
its own doc, and applies reviewer feedback.

**Status:** ✅ Approved by @mxinO; all CI passing.

### Changes

**New files**

- `.claude/skills/launching-evals/` — vendored verbatim from
[NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator) @
commit `8fa16b2` (latest). Covers run / check / debug / analyze flows
for NEL evaluations.
- `.claude/skills/accessing-mlflow/SKILL.md` — vendored verbatim from
the same upstream. Queries MLflow runs via the `mlflow-mcp` MCP server.
- `.claude/scripts/sync-upstream-skills.sh` — re-vendors the two skills
above at a pinned SHA. Idempotent; re-applies our provenance frontmatter
on each run.
- `.claude/skills/common/credentials.md` — shared HF / NGC / Docker
credential setup, referenced from slurm-setup.md. Generic (not
NVIDIA-internal) — public NEL SLURM-executor users rely on the same
NGC/HF setup. Includes a "check what's already set first" detection
section so the agent skips already-configured credentials.

**Updated files**

- `.claude/skills/common/slurm-setup.md` — NGC credential block
collapsed to a one-paragraph pointer at `credentials.md`.
- `.claude/skills/common/remote-execution.md` — reframed "Checkpoint and
storage availability" as "Staging checkpoints from your workstation".
Drops the misleading login-vs-compute framing and the dlcluster-specific
row.
- `.claude/skills/common/workspace-management.md` — drops stale pointer
to the deleted e2e doc.
- `.claude/skills/evaluation/SKILL.md` — workspace-integration intro
trimmed; NEL CI section removed (content moved to
Model-Optimizer-Internal). Monitoring block replaced via #1252 merge
with joint trim text that routes to `monitor`, `launching-evals`, and
`accessing-mlflow`.
- `.claude/skills/ptq/SKILL.md` — "Next steps" block refined.
- `.markdownlint-cli2.yaml` — excludes vendored upstream skills from
markdownlint so they stay byte-identical to upstream.

**Deleted files**

- `.claude/skills/common/end-to-end-workflow.md` — per @kaix-nv and
@mxinO: redundant with skill descriptions that already handle
cross-skill routing.
- `.claude/skills/evaluation/references/nel-ci-guide.md` — per
@shengliangxu: NVIDIA-internal. Moved to
`Model-Optimizer-Internal:agent/evaluation_guide.md` (MR !57).

### Review status

| Reviewer | Concern | How addressed |
|---|---|---|
| @shengliangxu | `nel-ci-guide.md` contains NVIDIA-internal content |
Moved to Model-Optimizer-Internal MR !57 (renamed `evaluation_guide.md`
since it covers both NEL SLURM executor and NEL CI). Refreshed against
current upstream NEL / NEL-CI (current cluster list, Sybil/non-Sybil
distinction, prerequisites checklist, `nel-ci-cli` preferred trigger). |
| @kaix-nv, @mxinO | e2e workflow doc unnecessary | Deleted. Skill
descriptions already route between ptq / deployment / evaluation. |
| @kaix-nv (#1252) | Overlap on `evaluation/SKILL.md` Monitoring section
| Coordinated via comment on #1252. @kaix-nv incorporated the joint trim
text referencing `monitor` + `launching-evals` + `accessing-mlflow`.
Landed via #1252 merge. |
| @mxinO | `remote-execution.md` compute-node framing is misleading |
Reframed as workstation→cluster staging. |
| @mxinO, CodeRabbit, Copilot | slurm-setup NGC creds aren't
SLURM-specific; `$oauthtoken` literal clarification; overwrite safety |
New `credentials.md` with NGC / HF / Docker setup. Append-if-missing
pattern (`grep -q … \|\| echo … >>`). Explicitly calls out `$oauthtoken`
as literal, kept unexpanded via single quotes. |
| @mxinO | `credentials.md`: check what's already set first; document
`hf auth login` | Added detection section at the top covering
`HF_TOKEN`, `~/.cache/huggingface/token`, Docker config, enroot creds.
HF section now documents `hf auth login` as the recommended interactive
path; env-var path kept as option 2 for scripts/CI. |
| @kevalmorabia97 | Internal `gitlab-master.nvidia.com` container URL in
`launching-evals/SKILL.md` | Auto-fixed by bumping pinned SHA from
`01899f8` to `8fa16b2` — upstream PR
[NVIDIA-NeMo/Evaluator#920](https://github.com/NVIDIA-NeMo/Evaluator/pull/920)
already replaced it with the public
`nvcr.io/nvidia/eval-factory/simple-evals:26.03`. SHA bump also picked
up upstream's new `nemo-evaluator-launcher resume` command and a tighter
"MANDATORY monitor after every `nel run`" directive. |
| @kevalmorabia97 | Internal `PPP` terminology +
`/lustre/fsw/portfolios/coreai/...` path | Still in upstream — vendoring
verbatim means we can't sanitize locally without breaking the "verbatim"
property. Filed upstream as
[NVIDIA-NeMo/Evaluator#938](https://github.com/NVIDIA-NeMo/Evaluator/issues/938)
to genericize. Next sync-script SHA bump will pick up the fix
automatically. |

### Related

- **Depends on:** #1236 (`deployment/references/unsupported-models.md`,
merged)
- **Coordinated with:** #1252 (monitor skill, merged — joint trim text
incorporated)
- **Internal counterpart:** [Model-Optimizer-Internal MR
!57](https://gitlab-master.nvidia.com/omniml/Model-Optimizer-Internal/-/merge_requests/57)
— `agent/evaluation_guide.md`
- **Upstream coordination:** vendored skills synced from
[NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator) @
`8fa16b2`. Follow-up issue:
[NVIDIA-NeMo/Evaluator#938](https://github.com/NVIDIA-NeMo/Evaluator/issues/938).

### Motivation

Learnings from running end-to-end PTQ → Deploy → Eval on
Devstral-Small-2-24B (FP8 VLM → NVFP4 MLP-only) on dlcluster B100, plus
prior NEL CI experience on oci-hsg.

### Testing

Validated end-to-end: PTQ (6 min) → vLLM deployment (3 debug iterations)
→ NEL evaluation (MMLU 77.4%, GSM8K 80%, GPQA 40% on
`limit_samples=10`).

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (documentation / skills only)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅
- `.claude/skills/launching-evals/` and
`.claude/skills/accessing-mlflow/` are vendored verbatim from
[NVIDIA-NeMo/Evaluator](https://github.com/NVIDIA-NeMo/Evaluator)
(Apache-2.0). Provenance SHA pinned in each `SKILL.md` frontmatter and
in `.claude/scripts/sync-upstream-skills.sh`.
- Did you write any new necessary tests?: N/A (skill documentation)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
2026-04-27 04:23:14 +00:00
h-guo18 1ec931c2c7 [2/3][Feat]: Offline DFlash training (#1343)
### What does this PR do?

Type of change: new feature

Part 2 of a 3-PR series splitting #1271:
- **[1/3] #1296**: File reorg + deprecate `ParallelDraft`
- **[2/3] this PR**: Offline DFlash training (depends on #1296)
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- Add `dflash_offline` flag to `DFlashConfig` for training from
pre-computed hidden states; deletes base model layers to save memory.
- Add Pydantic validators on `DFlashConfig`:
- `_derive_dflash_offline` — auto-derive `dflash_offline` from
`data_args.offline_data_path` in validation context. Not
user-configurable: any user-supplied value is overridden by the derived
value.
- `_resolve_mask_token_id` — auto-detect `dflash_mask_token_id` from
`tokenizer.mask_token_id`.
  - `_check_mask_token_id` — fail fast if unset after resolution.
- `HFDFlashModel.modify()`: select `num_orig_hidden_layers` when
offline; pick `_base_model_lm_head` device when no base layers present;
drop base-model `layers` module.
- `HFDFlashModel.forward()`: add offline branch — consumes precomputed
`base_model_outputs` via `DFlashBaseModelOutput.from_offline_dict`, and
when `dflash_self_logit_distillation` is enabled with
`base_model_logits` absent, recomputes logits from
`base_model_hidden_states` via `_base_model_lm_head`. Raises a clear
error from the non-training / `pseudo_speculative_generate` paths when
`dflash_offline=True`, since base-model layers have been deleted.
- `DFlashBaseModelOutput` dataclass in `modeling_dflash.py` (with
`from_offline_dict` classmethod) to unify online/offline output shapes.
`aux_hidden_states` is required in `from_offline_dict` so missing keys
fail fast at the entry point rather than deeper in the forward.
- `examples/speculative_decoding/main.py`: replace inline
`mask_token_id` auto-detect with
`DFlashConfig.model_validate(dflash_cfg, context={"tokenizer":
tokenizer, "data_args": data_args})`.

### Silent bug fix — `add_generation_template` → `add_generation_prompt`

The pre-refactor `compute_hidden_states_hf.py` passed
`add_generation_template=False` to `tokenizer.apply_chat_template`. This
kwarg does not exist on HF `apply_chat_template` and was being silently
ignored, so the intended "don't append a generation prompt" behavior was
never actually applied. The new `tokenize_with_loss_mask` helper in
`examples/speculative_decoding/collect_hidden_states/common.py` uses the
correct `add_generation_prompt=False`. **This is a real behavior
change** for anyone re-dumping hidden states: trailing generation
prompts that were previously appended to the tokenized sequences will no
longer be included.


### Testing
- New tests:
- `tests/unit/torch/speculative/plugins/test_hf_dflash_offline.py` — CPU
unit tests for convert path (online keeps base layers, offline deletes
them; `num_orig_hidden_layers` drives `target_layer_ids` in offline
mode) and `DFlashConfig._derive_dflash_offline` validator.
- `TestDFlashOfflineForwardGPU` in
`tests/gpu/torch/speculative/plugins/test_hf_dflash.py` — GPU forward
smoke with precomputed `base_model_outputs`, plus the
`dflash_self_logit_distillation` logit-recompute path.

- training test:
<img width="454" height="317" alt="image"
src="https://github.com/user-attachments/assets/79b92790-4d15-4313-bb9b-f35665b012e6"
/> <img width="456" height="310" alt="image"
src="https://github.com/user-attachments/assets/4558559f-9c35-49ed-b36e-82fbc99eab23"
/>


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — additive `dflash_offline`
flag defaulting to `False`; validators fall through when context not
provided.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — see Testing section above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### TODO (follow-up)

- [x] Update
`examples/speculative_decoding/collect_hidden_states/compute_hidden_states_*.py`
to support DFlash offline data. Current scripts are Eagle-specific —
they hardcode the `[2, N/2, N-3]` aux-layer selection and emit
`{input_ids, hidden_states, aux_hidden_states}`. DFlash offline needs:
- Aux layer indices driven by
`build_target_layer_ids(num_orig_hidden_layers, num_draft_layers)` (or a
configurable list), not the Eagle triplet.
- `base_model_hidden_states` key (last-layer hidden) so
`DFlashBaseModelOutput.from_offline_dict` + the
`dflash_self_logit_distillation` recompute path can consume it.
- Optional `base_model_logits` dump so offline training can skip the
self-distillation logit recomputation when logits are available.

### Additional Information

Base branch is #1296 (file reorg). Retarget to `main` once #1296 merges.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Offline DFlash speculative-decoding training from precomputed
base-model hidden states
* Answer-only-loss training with persisted loss masks and optional
chat-template support
* Flexible auxiliary-layer selection via CLI and an exposed default
aux-layer helper
* Auto-derived offline flag in config and automatic memory optimization
during offline conversion

* **Documentation**
* Updated guides for offline pipeline, aux-layer selection, and
loss-masking options

* **Tests**
* New unit, GPU, and regression tests covering offline conversion,
training, and config derivation
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-25 18:38:35 -07:00
h-guo18 7c80d85751 [1/3][Refactor]: File reorg; deprecate ParallelDraft (#1296)
### What does this PR do?

Type of change: refactoring

Part 1 of a 3-PR series splitting #1271:
- **[1/3] this PR**: File reorg + deprecate `ParallelDraft`
- **[2/3] #1295**: Offline DFlash training
- **[3/3] #1297**: Extract `HFSpecDecMixin`

Changes:
- **File reorg**: `transformers.py` → `hf_eagle.py`; extract
`HFMedusaModel` → `hf_medusa.py`; extract `EagleModule` /
`EagleBaseModelOutput` → `modeling_eagle.py`; extract `DFlashModule` /
`DFlashAttention` / `DFlashDecoderLayer` / `build_target_layer_ids` /
`apply_rotary_pos_emb` → `modeling_dflash.py`.
- **Deprecate `ParallelDraft`**: remove `parallel_draft_step`,
`parallel_draft_heads_num_layers`, and the `ParallelDraft` module from
HF Eagle; remove the `EagleMedusaExporter` branch from
`HFEagleModel.get_exporter()` (the `EagleMedusaExporter` class itself
still lives in `hf_spec_export.py` for Megatron parity).
- **Rename**: `_draft_model_config` → `eagle_config` in export plugin.
- Update imports in `examples/speculative_decoding/` and
`modelopt/torch/speculative/utils.py` to follow the module rename.

### Testing

Validated with existing Eagle and DFlash training scripts (re-run after
`9ae5302729 revert behavior change`).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ — renames
`modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`; removes
`parallel_draft_step` / `parallel_draft_heads_num_layers` from Eagle
config; renames `_draft_model_config` → `eagle_config` in export plugin.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — pure refactor; existing
tests updated for the rename. `test_hf_spec_rope_export.py` assertions
were also corrected to reflect the actual production path (the old
assertions were masked by `MagicMock` not invoking the
`_draft_model_config` `@property`).
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌

### Additional Information

Breaking changes:
- `modelopt.torch.speculative.plugins.transformers` → `.hf_eagle`
- `parallel_draft_step` / `parallel_draft_heads_num_layers` removed from
Eagle config
- `_draft_model_config` → `eagle_config` in export plugin

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactoring**
* Reorganized speculative-decoding plugins into focused modules,
converting the legacy "transformers" entry into a deprecated shim that
re-exports the new plugin surface.
* Consolidated DFlash implementation into a shared modeling component
and introduced a dedicated EAGLE decoder module.

* **New Features**
* Added a Medusa speculative-decoding plugin with configurable heads and
combined-loss training behavior.

* **Chores**
  * Updated pre-commit license-hook exclusion and feature-flag wiring.

* **Tests**
  * Updated export tests to expect rope-scaling fallback semantics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
2026-04-24 14:46:10 -07:00
Liana Mikaelyan 946639aa19 Fix PTQ for VLMs with image calibration (#1318)
### What does this PR do?

This PR fixes PTQ with image claibration for VLMs. 

### Usage

```python
python3 examples/llm_ptq/hf_ptq.py --pyt_ckpt_path Qwen/Qwen3-VL-8B-Instruct --qformat fp8 --export_path Qwen3-VL-8B-Instruct-fp8 --trust_remote_code --kv_cache_qformat none --calib_with_images --calib_size 512
```

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Image-text calibration now extends support to additional model
architectures when image calibration is enabled.
* Improved tokenizer truncation handling in multimodal dataset
processing to prevent configuration conflicts when image inputs are
present.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Liana Mikaelyan <lmikaelyan@nvidia.com>
2026-04-24 13:29:59 -07:00
Javier De Jesus 5887410a6d fix: EAGLE mix_hidden_states in-place op crash (#1088) (#1104)
### Type of change

- [x] Bug fix (non-breaking change which fixes an issue)

### Description

Fixes #1088 — `RuntimeError: one of the variables needed for gradient
computation has been modified by an inplace operation:
IndexPutBackward0` when training with `eagle_mix_hidden_states=True`.

**Root cause:** In `HFEagleModel._eagle_training_forward`, the indexed
assignment at line 991–994 modifies `eagle_input_hiddens` in-place while
it is still part of the autograd computation graph.

**Fix:** Clone the tensor before the in-place assignment. This is the
same pattern already used in the Megatron backend at
`megatron_eagle.py:1201-1202`:

```python
# Clone to avoid inplace modification of view created in no_grad mode
eagle_module_input_hidden_states = eagle_module_input_hidden_states.clone()
```

The HF backend was missing this clone.

### Usage

```python
config["eagle_mix_hidden_states"] = True
config["eagle_ttt_steps"] = 2
mtsp.convert(model, mode=[("eagle", config)])
model.train()
outputs = model(input_ids=input_ids, labels=labels)
outputs.loss.backward()  # no longer crashes
```

### Testing

Added `test_eagle_mix_hidden_states_backward` parametrized over
`eagle_ttt_steps` [1, 2] that:
- Converts a tiny LLaMA to EAGLE with `eagle_mix_hidden_states=True`
- Runs forward + backward pass
- Asserts loss is not None and gradients flow to `eagle_module`

```
pytest tests/unit/torch/speculative/plugins/test_hf_speculative.py::test_eagle_mix_hidden_states_backward -v
```

### Checklist

- [x] I have read the [contributor guidelines](CONTRIBUTING.md) and
signed my commits
- [x] I have followed the [security best practices](SECURITY.md)
- [x] This change is backward compatible
- [x] I have followed third-party code and dependency guidelines
- [x] I have added tests that prove my fix is effective

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed gradient computation issue in speculative decoding during model
training to ensure proper autograd behavior.

* **Tests**
* Added regression test to validate gradient computation in speculative
decoding scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: javierdejesusda <javier.dejesusj9@gmail.com>
2026-04-23 20:44:23 +00:00
Keval Morabia 0a1ca5d48a Fix unified_export_megatron for transformers 5.6 (#1335)
### What does this PR do?

Type of change: Bug fix

Broaden the exception handler around `AutoTokenizer.from_pretrained` in
`GPTModelExporter.save_pretrained` to also catch `ValueError` and
`ImportError`.

In `transformers` 4.x, attempting to load a tokenizer from a directory
that contains only `config.json` (no `tokenizer.json` /
`tokenizer.model` / `tokenizer_config.json`) raised `OSError`, which was
already handled. In `transformers` 5.x the resolution path now reaches
`PreTrainedTokenizerFast.__init__` and raises a `ValueError` ("Couldn't
instantiate the backend tokenizer from one of: ...") when none of the
three backend sources are available. This caused
`export_mcore_gpt_to_hf` to hard-fail for checkpoint directories that
don't carry tokenizer files — including
`tests/gpu_megatron/torch/export/test_unified_export_megatron.py`, which
writes only a minimal `config.json`.

The broadened `except (OSError, TypeError, ValueError, ImportError)`
mirrors the pattern already used just below for
`AutoProcessor.from_pretrained` and keeps tokenizer export best-effort,
as originally intended.

### Usage

No API change. Existing call sites continue to work:

```python
from modelopt.torch.export import export_mcore_gpt_to_hf

export_mcore_gpt_to_hf(
    model,
    pretrained_model_name_or_path,  # may or may not contain tokenizer files
    dtype=torch.bfloat16,
    export_dir=export_dir,
)
```

### Testing

- Reproduced the failure on `transformers==5.6` with:
  ```
pytest
tests/gpu_megatron/torch/export/test_unified_export_megatron.py::test_unified_export_megatron[llama-LlamaForCausalLM-medusa-None-None]
  ```
which failed with `ValueError: Couldn't instantiate the backend
tokenizer...` raised from `unified_export_megatron.py:299`.
- After the fix, the same parametrization passes, and the other `llama`
/ `nemotron` / `eagle` / `medusa` parametrizations in the same test file
remain green.
- No behavioral change on `transformers` 4.x: the `OSError` path is
still caught.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- Existing test
`test_unified_export_megatron` already covers this path; the fix makes
it pass on transformers 5.x. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

Triggered by the upgrade to `transformers` 5.6 (used in
`nvcr.io/nvidia/nemo:26.04`, which is the container for the
`gpu_megatron` nox session). The error message from `transformers` —
*"You need to have sentencepiece or tiktoken installed..."* — is a
misleading generic fallback; `sentencepiece` and `tiktoken` are already
pulled in via the `[hf]` extras, and the real cause is the missing
tokenizer files in the export source directory.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Improved error handling in the export process to gracefully manage
additional exception types.

* **Tests**
* Enhanced test validation for Megatron export by using actual tokenizer
artifacts, ensuring model vocabulary size matches test tokenizer
configuration.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-23 19:23:04 +00:00
Chenjie Luo fda0899e40 feat(recipes): add KV cache cast variants (fp8_cast / nvfp4_cast) (#1334)
## Summary

- Adds three built-in PTQ recipes that express the KV-cache *cast*
variants directly in YAML, using the existing `use_constant_amax: true`
quantizer field. These are recipe equivalents of
`--kv_cache_qformat=fp8_cast` / `nvfp4_cast`:
  - `general/ptq/fp8_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-fp8_cast_kv`
  - `general/ptq/nvfp4_default-nvfp4_cast_kv`
- Makes `--recipe` authoritative in `examples/llm_ptq/hf_ptq.py`: the
post-hoc `_set_kv_cache_constant_amax` override now only runs when
`--recipe is None`, so a recipe YAML fully determines KV-cache config
instead of being silently overridden by the default
`--kv_cache_qformat=fp8_cast`. Updated help text on both flags.
- Extends the recipe loader smoke test to cover the three new recipes.

## Motivation

Before this change, the cast variants lived only in argparse
(`_KV_CAST_FORMATS = {"fp8_cast", "nvfp4_cast"}`) and were layered on
top of any recipe-loaded config. That meant `--recipe
nvfp4_default-fp8_kv` would silently become a cast recipe due to the
`--kv_cache_qformat` default. Now the recipe is self-contained: its YAML
either sets `use_constant_amax: true` on the `*[kv]_bmm_quantizer` entry
(cast) or doesn't (data-driven calibration).

## Test plan

- [x] `pytest tests/unit/recipe/test_loader.py` — all 24 tests pass,
including the three new parametrized recipes.
- [x] Verified each new recipe round-trips through `load_recipe()` with
`use_constant_amax: True` surviving Pydantic validation on the KV entry.
- [x] End-to-end run on `/models/Qwen/Qwen3-8B` (RTX 6000 Ada, 4
samples, seq_len=128) for all three new recipes:
- After `mtq.quantize(model, recipe.quantize.model_dump(),
forward_loop=...)`, all 72 `k_bmm_quantizer` / `v_bmm_quantizer` modules
have `_use_constant_amax=True` and `_get_amax()` returns `448.0` (FP8
E4M3 max).
- Weight quantizers still calibrate from data normally (sample amax
values: q_proj=0.5508, k_proj=0.6250, v_proj=0.1689, o_proj=0.7266).
- [x] Verified the `--recipe` authoritative behavior change:
- Non-cast recipe + default `--kv_cache_qformat=fp8_cast` → KV entry
does NOT get `use_constant_amax` (no silent override).
- Cast recipe + contradictory `--kv_cache_qformat=fp8` → KV entry keeps
`use_constant_amax=True` (recipe wins).

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Fixed CLI to respect KV cache quantization settings from recipe YAML
instead of overriding them.

* **New Features**
* Added three new post-training quantization recipe configurations for
FP8 and NVFP4 with optimized KV cache handling.

* **Documentation**
* Enhanced CLI help text for recipe and KV cache quantization options
with configuration examples.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 18:45:16 +00:00
Chenjie Luo 01788bb007 Deprecate Mllama support in llm_ptq/vlm_ptq examples (#1332)
## Summary

- Removes Mllama (Llama 3.2 Vision) model-type branches from the
`llm_ptq` example (`hf_ptq.py`, `example_utils.py`) and drops the
now-unused `MllamaImageProcessor` wrapper from `modelopt/torch/utils/`.
- Drops the legacy `MllamaImageProcessor` path in
`modelopt/torch/utils/vlm_dataset_utils.py`; the generic HF
ProcessorMixin path handles the remaining cases.
- Adds a CHANGELOG entry under 0.44 Backward Breaking Changes.

## Test plan

- [x] CI lint / unit tests pass
- [x] Smoke-run ``examples/llm_ptq/scripts/huggingface_example.sh
--model <llm> --quant fp8`` (text-only path, non-mllama)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Removed Mllama (Llama 3.2 Vision) support from quantization examples.
This includes removal of dedicated image processor implementation,
specialized model handling, and related calibration logic.
* Updated VLM image-text calibration guidance to use
`--calib_with_images` flag with other supported VLMs instead of
Mllama-specific processing paths.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-04-23 23:24:42 +05:30
Michael Feil 8663678f12 fix: bug hf_ptq.py max_length setting ignored for LLMs (#1311)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage
- fixes a bug in example script. We were trying why our models were not
that strong at long context. Seems like a recent refactor did not
implement max seq length. so 512 is used by default.

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **Improvements**
* Calibration data loading now enforces a maximum sequence/sample length
during dataset preparation, ensuring calibration inputs adhere to
configured length limits. This yields more predictable calibration
behavior, reduces peak memory usage during calibration runs, and
improves consistency of quantization preprocessing.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Michael Feil <63565275+michaelfeil@users.noreply.github.com>
2026-04-23 09:02:31 -07:00
Ajinkya Rasane e4e3508f51 [OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT (#1289)
## Summary
Enables TensorRT attention-v2 fusion for vision transformers when
exported to ONNX with FP8 Q/DQ. The core library changes are
architecture-agnostic (drop-in for any FP8 ONNX export); coverage is
exercised by the existing `examples/torch_onnx/torch_quant_to_onnx.py`
pipeline.

- **`modelopt/onnx/export/fp8_exporter.py`** — new post-processing
passes: move attention-scaling `Mul` and K `Transpose` to the Q-side so
DQ feeds MatMul directly, pre-transpose constant weights, and insert FP8
Q/DQ on Softmax outputs (fixed `1/448` scale, data-independent) for
MHA-v2 fusion. Rewrites only fire when every downstream consumer is a
MatMul so non-attention branches are never perturbed.
- **`modelopt/onnx/utils.py`** — `fold_dq_fp32_to_fp16_casts` /
`fold_q_fp16_to_fp32_casts` remove the Cast nodes
`convert_float_to_float16` inserts around Q/DQ and rewrite scale
initializers to FP16 so TRT fuses DQ into the downstream GEMM. Guarded
behind opset >= 19 (FP16 Q/DQ scale requirement). Warns on FP16
overflow/underflow.
- **`modelopt/torch/_deploy/utils/torch_onnx.py`** — calls the fold
helpers for FP8-quantized models after `convert_float_to_float16`.
- **`modelopt/torch/quantization/export_onnx.py`** — keeps FP8 Q/DQ
scale in the native input dtype so no Cast is emitted between graph and
Q/DQ. Removes the now-unused `trt_high_precision_dtype` parameter from
`_fp8_quantize`/`_fp8_dequantize`.
- **`modelopt/torch/quantization/nn/modules/quant_layernorm.py`** (new)
— registers `nn.LayerNorm` in `QuantModuleRegistry` so LayerNorm output
quantizers are honored.
- **`modelopt/torch/quantization/plugins/huggingface.py`** — skips
`*Attention` wrappers whose children are also `*Attention` per-instance
(not per-class) to avoid double-patching `eager_attention_forward` (e.g.
`ViTAttention` vs `ViTSelfAttention`).
- **`examples/torch_onnx/torch_quant_to_onnx.py`** — adds a
`_FP8_MHA_OVERRIDE` config block to FP8 mode that enables LayerNorm
output quantizer + disables its input quantizer for TRT attention
fusion.
- **Unit tests** (12 CPU tests, ~1.2s total) — fp8_exporter rewrites +
fanout safety, fold-cast helpers + opset guard, LayerNorm quant-wrapper
identity, per-instance nested-attention detection.

## Benchmarks
ViT-base-patch16-224, RTX 6000 Ada, strongly-typed FP8 via `trtexec`.
Accuracy on 2 000 ImageNet-1k validation samples (streaming).

**Batch = 1 (latency-bound)**
| Model | Top-1 | Top-5 | TRT latency | Speedup |
|---|---|---|---|---|
| FP16 baseline | 80.96% | 95.80% | 0.722 ms | 1.00x |
| Torch FP8 MHA | 80.66% | 95.75% | 0.657 ms | **1.10x** |
| ONNX PTQ FP8 | — | — | 0.589 ms | **1.23x** |

**Batch = 64 (throughput-bound, realistic inference)**
| Model | TRT latency | Speedup | Images/s |
|---|---|---|---|
| FP16 baseline | 23.40 ms | 1.00x | 1152 |
| Torch FP8 MHA | 15.89 ms | **1.47x** | 1152 |
| ONNX PTQ FP8 | 15.89 ms | **1.47x** | 1216 |

Top-1 accuracy stays within 0.30 pp of FP16; at batch=64 the Torch FP8
MHA path matches ONNX PTQ wall-time — attention is the bottleneck there
and both paths achieve full FP8 attention fusion (36/36 attention
MatMuls with QDQ in ViT-base).

## Test plan
- [x] CPU unit tests (new): \`python -m pytest
tests/unit/onnx/quantization/test_fp8_mha_exporter.py
tests/unit/onnx/test_fold_casts.py
tests/unit/torch/quantization/test_quant_layernorm.py
tests/unit/torch/quantization/plugins/test_nested_attention_skip.py\`
- [x] Existing ONNX / quantization unit suites unaffected: \`python -m
pytest tests/unit/onnx tests/unit/torch/quantization\`
- [x] End-to-end ViT FP8 export: \`python
examples/torch_onnx/torch_quant_to_onnx.py --timm_model_name
vit_base_patch16_224 --quantize_mode fp8 --onnx_save_path
vit_base_fp8.onnx\` — expect log lines \`Folded 48 weight Transpose
nodes\`, \`Inserted FP8 weight DequantizeLinear for 1 Conv nodes\`, and
\`Attention QDQ rewrites: ... inserted QDQ on 12 Softmax outputs\`
- [x] trtexec FP8 strongly-typed build: \`trtexec
--onnx=vit_base_fp8.onnx --fp8 --stronglyTyped\`
- [x] Accuracy within ~0.3 pp of FP16 baseline on ImageNet-1k subset

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
2026-04-23 11:38:31 -04:00
Keval Morabia 7265ca6793 Fix lm_eval version checking (#1321)
### What does this PR do?

Type of change: bug fix <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

`lm_eval` does not have `__version__` attribute

### Additional Information
<!-- E.g. related issue. -->
NVBug 6102101


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced the package version detection system to improve overall
reliability and stability of the application while reducing unnecessary
external dependencies. All functionality, including version gating and
system warnings, continues to operate exactly as expected with no impact
on the user experience.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-23 16:57:54 +05:30
Grzegorz K. KarchandKeval Morabia 2564da7324 Update vLLM deployment docs for heterogeneous models (#1317)
### What does this PR do?

Type of change: ? documentation.

This PR updates vLLM deployment instructions, taking into account
heterogenous models created with AnyModel.

### Usage

Does not apply.

### Testing

Run the updated instructions in the documentation.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Replaced a benchmarking-focused section with a deployment guide for
running compressed models on vLLM.
* Added step-by-step setup for using an AnyModel-enabled vLLM fork,
including checkout and install guidance and required model config edits
(with optional architecture metadata).
* Simplified runtime to a single vllm serve command, removing manual
model rearrangement steps.
* Restored inference benchmarking as a subsection, retaining vllm bench
latency/throughput examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Grzegorz Karch <gkarch@nvidia.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-23 06:22:37 +00:00
jingyu-ml c7966119eb Reorg the sparse/quant/common kernel dir (#1303)
### What does this PR do?

Type of change: re-org code <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ We changed the import path
<!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ❌ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:✅
<!--- Only for new features, API changes, critical bug fixes or backward
incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Calibration support for skip-softmax multi-threshold measurement in
sparse attention.
  * N:M sparse softmax masking and helpers for sparsity-aware attention.

* **Chores**
* Reorganized and consolidated kernel/backends for quantization and
sparsity to a unified kernels layout, updating tests and examples to
match.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
2026-04-22 23:34:56 +00:00
Keval MorabiaandClaude Sonnet 4.6 e56682e34a docs: update installation pages with legal-approved license notices (#1322)
## Summary

- Replaces the old pip license notice ("Please review the license terms
of ModelOpt and any dependencies before use") with the Legal-approved
wording: "Model Optimizer will download and install additional
third-party open source software projects. Review the license terms of
these open source projects before use."
- Adds a generic container license review notice ("Before pulling and
using the container images, please review their respective license
terms.") to the Linux installation doc (Docker tab) and README.
- Adds a `.. note::` with the pip notice to the Windows installation
page (covers both standalone and Olive child pages).
- Expands the README container section to explicitly list all four
recommended NVIDIA container images (`pytorch`, `nemo`, `tensorrt-llm`,
`tensorrt`).

## Test plan

- [x] Verify rendered docs look correct (`nox -s docs`)
- [x] Confirm legal notices appear in Linux, Windows, and README install
sections

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated installation guides with explicit references to supported
NVIDIA container images (PyTorch, NeMo, TensorRT-LLM and variants),
clarified pre-installed Model Optimizer in some images, and added notes
to review each container’s license terms; clarified conditional
environment setup wording and local install license guidance.
* **Chores**
  * Updated project license header year.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-22 22:49:23 +05:30
kinjalpatel27 0678136335 Fix vLLM fakequant MoE megatron export bug (#1305)
### What does this PR do?

Type of change: Bug fix                     
Fixes two bugs in the vLLM + Megatron-Core MoE export path and cleans up
the related weight-collection helper:
    
1. **`_QuantFusedMoEBase` (vllm.py)**: The weight-quantizer path in
`_invoke_fused_moe_quantized_function` was temporarily mutating
`self.w13_weight` / `self.w2_weight` to the quantized tensor, then
restoring them via `finally`. This exposed a stale quantized tensor on
`self` between the mutation and the kernel call. Fixed by computing the
quantized weight directly into a local `B` without touching `self.*`
attributes.
2. **`GPTModelExporter` / `VllmFqGPTModelExporter`
(unified_export_megatron.py / vllm_fakequant_megatron.py)**:
`expert_bias` (present in grouped MoE layers) was silently dropped
during export because the bias collection ran after the early-return on
missing `weight`. Extracted a `_get_weight_bias` helper that collects
weight, bias, and expert_bias together, so bias/expert_bias are captured
even when weight is absent or zero-element.

### Usage

```python
# No API change; export pipelines pick this up automatically.
# export_mcore_gpt_to_hf_vllm_fq / export_mcore_gpt_to_hf now correctly                   
# export expert_bias for grouped-MoE checkpoints.    
```

### Testing
Step 1 — Quantize (run from Megatron-LM
examples/post_training/modelopt):
```
HF_MODEL_CKPT=<path/to/hf/weights> MLM_MODEL_SAVE=<quant-ckpt-name> \
bash quantize.sh <hf-model-id> NVFP4_DEFAULT_CFG
```
Step 2 — Export for vLLM fakequant:
```
MLM_EXTRA_ARGS=--export-vllm-fq \ 
HF_MODEL_CKPT=<path/to/hf/weights> \ 
MLM_MODEL_CKPT=<quant-ckpt-name> \ 
EXPORT_DIR=<export-dir> \ 
bash export.sh <hf-model-id> 
```
Step 3 — Serve (run from examples/vllm_serve):
```
 QUANT_CFG=NVFP4_DEFAULT_CFG \ 
 QUANT_FILE_PATH=<export-dir>/quantizer_state.pth \ 
 python3 vllm_serve_fakequant.py <export-dir> \ 
 -tp 1 --served-model-name <model-name> \ 
 --host 0.0.0.0 --port 8000 \ 
--trust-remote-code --enforce-eager \ 
 --disable-custom-all-reduce \ 
--gpu-memory-utilization 0.8          
```
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ❌
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Refactor**
* Centralized weight/bias/expert-bias extraction and export to a single
helper for consistent handling.
* Standardized quantized-weight flow to temporarily swap and restore
parameter tensors during computation.

* **Bug Fixes**
* Prevented missing or incorrect weight/bias exports by unifying
extraction logic.
* Broadened checkpoint key matching to preserve more quantizer state
during reloads.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-22 09:24:25 -07:00
nv-samcheng c417e6f4d9 Exclude small-k and small-n Matmul nodes from Int8 quantization (#1256)
### What does this PR do?
Exclude small-dimension MatMul nodes from INT8 quantization. MatMuls
with N or K < 16 cannot efficiently use INT8, causing performance
regressions.



### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization exclusions so MatMul/Gemm ops with derived K<16
or N<16 are skipped, honoring Gemm transB, using inferred and
runtime-determined shapes, and avoiding duplicate outputs.

* **Tests**
* Expanded unit tests to cover constant, inferred, and runtime-derived
shapes, Gemm transB behavior, small-dimension edge cases, and output
deduplication.

* **Documentation**
* Added changelog entry documenting the new small-dimension exclusion
thresholds and transB handling.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: samcheng <samcheng@nvidia.com>
2026-04-22 15:47:43 +05:30
Keval Morabia 785d3a2df6 [CI] Bump test containers to latest (#1299)
- Use latest containers for testing in CICD

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Chores**
* Bumped TensorRT-LLM Docker images to 1.3.0rc12 in example and GPU test
workflows.
  * Updated PyTorch container image from 26.01 to 26.03 for GPU tests.
* Captured uv lock upgrade output to a temp file, inlined it into PR
bodies, and adjusted workflow heredoc/templating and step behavior.

* **Documentation**
* Clarified an inline comment and simplified a warning message for an
ONNX quantization extension.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-21 23:44:12 -04:00
Keval Morabia c51c1762b3 fix: prevent gh-pages repo bloat from doc preview artifacts (#1309)
### What does this PR do?

Type of change: Bug fix

Fixes gh-pages branch bloat that grew from ~26 MB to ~441 MB in four
weeks (nvbug 6099503). Three compounding causes were identified and
addressed:

1. **Sphinx `.doctrees/` cache published to gh-pages** — `sphinx-build`
was writing its build cache inside `build/html/` which was then uploaded
verbatim. Accounts for ~3.3 GB uncompressed across history.
2. **`JamesIves/github-pages-deploy-action` appending a commit on every
push** — main-site files accumulated forever with `single-commit: false`
(default).
3. **PR preview deploying on every `synchronize` event for all PRs** —
`rossjrw/pr-preview-action` re-deployed the full site for every push to
any PR regardless of whether docs changed (e.g. PR #1128 triggered 64
preview deploys × ~11 MB each).

Changes:
- Pass `-d /tmp/doctrees` to `sphinx-build` so `.doctrees/` is never
written into `build/html/`
- Add `paths: [docs/**, modelopt/**]` filter to `pull_request` trigger
so the docs workflow only runs on PRs that touch docs or source code
- Set `single-commit: true` on the deploy action so main-site pushes
squash into one commit
- Deduplicate docs build: `deploy-preview` now downloads the artifact
from `build-docs` instead of running a second `sphinx-build`
- Set `retention-days: 1` on the artifact since it is only needed for
the duration of the workflow run

The one-time cleanup (force-push squashed orphan to gh-pages) was
already applied separately — repo is now ~59 MB for a full clone vs ~441
MB before.

### Usage

N/A — CI/workflow change only.

### Testing

- Workflow logic reviewed manually.
- The one-time cleanup was verified: `git rev-list --objects
--disk-usage origin/gh-pages` now reports ~28 MB; full clone is ~59 MB.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A

### Additional Information

nvbug 6099503

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Optimized documentation build and deployment workflow in CI/CD
pipeline.
* Improved pull request documentation preview handling with faster build
timeouts and refined artifact management.
* Enhanced GitHub Pages deployment configuration for better consistency.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-21 20:19:08 +00:00
sychen52 5ffb8487d9 add gptq fused kernel (#1291)
### What does this PR do?

Add gptq fused kernel to improve speed.

### Usage

check unittest

### Testing
added a unittest

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Fused GPTQ backend for faster blockwise weight updates, toggleable via
a new "fused" option.
  * Shared NVFP4 quantization primitives exposed for reuse.

* **Refactor**
* Consolidated FP4 scale/quantization logic into reusable utilities and
centralized Hessian inversion handling.

* **Tests**
* Expanded GPU tests comparing fused vs unfused GPTQ, added
Triton-availability gating and a local benchmark entrypoint.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-04-21 10:50:04 -07:00
yeyu-nvidiaandClaude Opus 4.6 2fef374ded fix: auto-compute dp_replicate_size from world_size (#1302)
## Summary
- When `dp_shard_size < world_size` (e.g., `dp_shard_size=4` on 8 GPUs
across 2 nodes), `ParallelismConfig` raises `total_size (4) does not
match num_processes (8)` because `dp_replicate_size` defaults to 1
- Auto-compute `dp_replicate_size = world_size // (dp_shard_size *
cp_size)` so intra-node FSDP2 sharding + inter-node data-parallel
replication works without manual config
- This enables `dp_shard_size` to be set to per-node GPU count (better
NVLink utilization) while automatically creating replicas across nodes

## Test plan
- [ ] Verify single-node training (dp_shard_size == world_size,
dp_replicate_size == 1) unchanged
- [ ] Verify multi-node with dp_shard_size < world_size creates correct
replica groups
- [ ] Verify existing EAGLE3/DFlash configs still work

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Enhanced parallelism configuration initialization in the speculative
decoding example to better handle distributed training scenarios.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 20:39:36 +00:00
Chenhan D. YuandClaude Opus 4.6 355c6b7883 fix: PTQ 1GPU, export PP divisibility, hidden states conversations key (#1293)
## Summary
- **megatron_lm_ptq.yaml**: Qwen3-8B PTQ to single GPU for L40 clusters
(TP=1, all tasks)
- **quantize.sh**: Auto-find largest PP dividing model's
`num_hidden_layers` for export step. Qwen3-8B has 36 layers which isn't
divisible by 8, causing `AssertionError` on 8-GPU nodes
- **compute_hidden_states_trtllm.py**: Use `messages` with
`conversations` fallback, matching the HF version. Fixes `KeyError:
'conversations'` when data uses OpenAI `messages` format

## Test plan
- [x] Qwen3-8B PTQ runs on single L40 GPU
- [x] Export PP auto-selects valid divisor (36 layers → PP=6 on 8 GPUs,
PP=4 on 4 GPUs, PP=1 on 1 GPU)
- [x] EAGLE3 offline pipeline reads data with `messages` field

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Dataset input handling now supports multiple field formats for
enhanced compatibility.

* **Bug Fixes**
* Optimized GPU resource allocation during model quantization with
improved pipeline parallelism computation.
* Updated quantization configuration for more efficient resource
utilization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenhan Yu <chenhany@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
0.45.0dev
2026-04-20 12:43:58 -07:00
yeyu-nvidiaandClaude Opus 4.6 289a239ca5 fix: use data_dir for directory paths in ShardedDataset (#1301)
## Summary
- `datasets`' `resolve_pattern` only matches entries with
`type=="file"`, so passing a bare directory path as `data_files` to
`load_dataset` results in `FileNotFoundError` even when the directory
exists on disk
- Detect directory paths in `ShardedDataset._load_dataset()` and pass
them via `data_dir` instead of `data_files`

## Reproduction
```python
from datasets import load_dataset
# This fails with FileNotFoundError:
load_dataset("json", data_files="/path/to/data_directory")
# This works:
load_dataset("json", data_dir="/path/to/data_directory")
```

## Test plan
- [ ] Verify existing EAGLE3/DFlash training pipelines that pass
directory paths work
- [ ] Verify file path and glob patterns still work (falls through to
`data_files`)
- [ ] Verify `data_files=None` (no data_files arg) still works

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Bug Fixes

* Fixed an issue with dataset loading that prevented proper handling of
directory-based data sources. Directories are now correctly detected and
processed during dataset initialization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-20 19:10:02 +00:00
Frida HouandKeval Morabia 97d153118e [minor] Add custom calibration backend registry (#1281)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a public backend-specific calibrator registration API to support
FP8 scale-sweep calibration, allowing backends to supply custom
calibrators used during FP8 tuning.

* **Tests**
* Added unit tests confirming registry insertion/overwrite, that
registered calibrators are invoked when FP8 scale-sweep is enabled, are
not invoked when disabled, and that calibration falls back to defaults
when no backend is registered.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Fridah-nv <201670829+Fridah-nv@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-20 09:24:17 -07:00
kinjalpatel27 010b220dc0 vLLM fakequant export update for AWQ checkpoint (#1242)
### What does this PR do?

Type of change: Bug

Enables end-to-end AWQ checkpoint export and reload in the vLLM
fake-quant serving path (`MODELOPT_STATE_PATH`). Previously, the
`input_quantizer` was using incorrect `pre_quant_scale` especially with
grouped quantizers like `qkv_proj`, using simply the first
`input_quantizer.pre_quant_scale`. This MR adds
`_resmooth_experts_for_export` that non-mutatively averages
`pre_quant_scale` across MoE experts and unifies input `_amax`, required
because vLLM uses a single input quantizer per expert group. Adds
`merge_amax_tensors_for_group` (element-wise max for same-shape, `cat`
for GQA, scalar-max fallback) replacing the scalar-collapsing
`torch.stack().max()` that dropped per-channel `_amax` structure.

### Usage

```python
# Export AWQ checkpoint from HF model
  from modelopt.torch.export.plugins.vllm_fakequant_hf import export_hf_vllm_fq_checkpoint
  export_hf_vllm_fq_checkpoint(model, export_dir="./awq_vllm_checkpoint")      
```

### Testing
**Step 1 — Export the quantized checkpoint:**
  ```bash                    
  python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <MODEL_PATH> \
    --recipe <AWQ_RECIPE> \
--calib_size 512 \
--export_path <EXPORT_DIR> \
    --vllm_fakequant_export
```
  This produces `<EXPORT_DIR>/vllm_fq_modelopt_state.pth` with the averaged per-expert                                                                                                                           
  pre_quant_scale and unified _amax now included.                                                                                                                                                              
   

 Step 2 — Serve via vLLM fakequant worker:                                                                                                                                                                    
```bash
  MODELOPT_STATE_PATH=<EXPORT_DIR>/vllm_fq_modelopt_state.pth \
python examples/vllm_serve/vllm_serve_fakequant.py \
      <EXPORT_DIR> --tensor-parallel-size <TP>   
```

Tested for quantization configurations:
```
FP8_DEFAULT_CFG
FP8_DEFAULT_CFG (input_q disabled)
INT8_SMOOTHQUANT_CFG
INT8_WEIGHT_ONLY_CFG
NVFP4_DEFAULT_CFG
NVFP4_AWQ_LITE_CFG
INT4_AWQ_CFG
NVFP4_AWQ_CFG
NVFP4_DEFAULT_CFG (input_q disabled)
```

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A 
- Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A 

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai -->
## Summary by CodeRabbit

* **New Features**
  * Added Nemotron-style MoE export support and group-aware AWQ resmoothing with optional requantization during export.
  * Improved handling for shared-input / expert groups and tensor-parallel sharding of pre-quantization scales.

* **Bug Fixes**
  * Removed AWQ reload limitation from known issues; improved checkpoint validation and safer save/load behavior.
  * Better detection and handling of enabled weight-quantizers and clearer warnings for mismatched checkpoint keys.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-04-19 22:12:45 -07:00
jingyu-ml 26ae8da517 [2/3] Implicit Gemm NVFP4 (#1227)
### What does this PR do?

Type of change: new feature <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

- Add Conv3D implicit GEMM kernel with BF16 WMMA tensor cores and fused
NVFP4 activation quantization for video diffusion VAE layers
- Integrate into _QuantConv3d via QuantModuleRegistry — automatically
dispatched when NVFP4 quantization is applied to nn.Conv3d
- Move kernel from `experimental/conv/ to modelopt/torch/kernels/conv/`;
move tests to `tests/gpu/torch/quantization/kernels/`

### Testing
<!-- Mention how have you tested your change if applicable. -->

- Added test cases to measure the difference between cuDNN and our CUDA
implicit GEMM kernel
- Added an NVFP4 fake quantization test using CUDA code

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ <!--- If ❌, explain why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ <!---
Mandatory -->
- Did you write any new necessary tests?: ✅ <!--- Mandatory for new
features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ <!--- Only for new features, API changes, critical bug fixes or
backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Per-backbone quantization/export in a single run with per-backbone
checkpoints and backbone-aware quant filters
* Configurable NVFP4 block-size via CLI/config; improved NVFP4 Conv3D
inference path and Wan 2.2 quantization support
* **Bug Fixes**
* Video-model calibration now respects extra params and forces video
decoding during calibration
* **Documentation**
* Added comprehensive Conv3D implicit‑GEMM kernel documentation; removed
experimental Conv3D prototype docs/benchmark
* **Tests**
* New Wan 2.2 quantization/export tests and expanded Conv3D/FP4 kernel
test coverage
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jingyu Xin <jingyux@nvidia.com>
0.44.0rc0
2026-04-19 12:20:14 +05:30
kaix-nv c20f9c411d Add a standalone monitor skill for persistent job tracking (#1252)
### What does this PR do?

Type of change: ? <!-- Use one of the following: Bug fix, new feature,
new example, new tests, documentation. -->
Add a standalone monitor skill for persistent job tracking across
sessions, and integrate it with PTQ, evaluation, and deployment skills.

Problem: Each skill had ad-hoc inline monitoring (squeue polling, nel
status checks) that didn't survive session restarts and couldn't track
multiple jobs. Users had to manually ask "check status" every time.

Solution: A centralized monitor skill with:
- Job registry (.claude/active_jobs.json): single source of truth for
all active jobs
- Durable recurring cron: polls every 15 min, survives session restarts,
self-cleans when all jobs complete
- User-initiated mode: works in new conversations by reading the
registry
- Aggregated reporting: "2 of 4 completed" instead of per-job noise

### Usage
After any skill submits a job, the monitor skill automatically:

1. Registers the job in .claude/active_jobs.json
2. Sets up a durable cron to poll status every 15 minutes

User can also trigger manually:
User: "check my eval status" → reads registry, reports current state
User: "is the PTQ done?"            → finds job, checks status
User: "what jobs are running?"      → lists all registered jobs

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added monitor skill for tracking SLURM jobs, NEL evaluations, and
launcher experiments with persistent job registry.

* **Documentation**
* Updated deployment, evaluation, and PTQ documentation to use the new
monitor skill.
  * Simplified diagnostic and troubleshooting instructions.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-04-19 04:44:17 +00:00
github-actions[bot]andgithub-actions[bot] e9a49890f1 [chore]: weekly bump of uv.lock on main (2026-04-18) (#1292)
## Summary
Automated weekly update of uv.lock file for nSpect Scanning:
- `uv.lock` — upgraded all transitive dependencies to latest compatible
versions

Signed-off-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-04-18 17:22:23 +00:00
Keval MorabiaandClaude Sonnet 4.6 3d0f0db49e [CI] Replace tox with nox, use nemo:26.04 for megatron tests, and simplify CI workflows (#1286)
### What does this PR do?

Type of change: New feature / infrastructure improvement

Follow-up to #1285 for correct CI test environment for megatron based
tests

Replaces `tox` + `tox-current-env` with `nox` for all test, lint, docs,
and wheel build sessions. The primary motivation was that
`tox-current-env` is incompatible with uv venvs in NGC containers (e.g.
NeMo's `/opt/venv`) — it picks the system Python via
`sys._base_executable` instead of the container's venv Python which has
megatron packages pre-installed.

Key changes:
- **`noxfile.py`** replaces `tox.ini` with GPU, CPU unit,
partial-install, pre-commit, docs, and wheel sessions
- **GPU sessions** use `venv_backend="none"` (run directly in container
env) and `python -m pip/pytest` to avoid PATH mismatches
- **uv** is set as the default venv backend (if available) for CPU
sessions (faster installs)

Also includes CI workflow simplifications:
- **`_pr_gate.yml`** new reusable workflow centralizing file-change
detection + linux-check wait logic (was duplicated across 3 workflow
files)
- **Collapsed pr/non-pr job pairs** into single jobs with conditional
`runs-on` in `gpu_tests.yml`, `example_tests.yml`,
`regression_tests.yml`
- **Collapsed `multi-py` / `multi-torch` / `multi-transformers`** into a
single `multi-version` matrix job in `unit_tests.yml`
- **PR path filtering** for unit test secondary jobs (multi-version,
launcher, partial-install) — skipped if no relevant files changed
- **Fixed schedule/workflow_dispatch skipping** — jobs with `needs:
[pr-gate]` were incorrectly skipped when all pr-gate internal jobs were
skipped; fixed by making the gate job always run
- **multi-version, launcher, partial-install** now also run on
`schedule` / `workflow_dispatch`

### Usage

```bash
python -m pip install nox uv                                                    # install nox and uv (once)
nox -l                                                                          # list all sessions
nox -s gpu_megatron                                                             # run a GPU session (inside container)
nox -s "unit-3.12(torch_211, tf_latest)"                                        # run a specific unit test combination
nox -s "unit-3.12(torch_211, tf_latest)" -R                                     # force-recreate venv (e.g. after dep changes)
COVERAGE_PROCESS_START=pyproject.toml nox -s "unit-3.12(torch_211, tf_latest)"  # with coverage
```

### Testing
- Ran `nox -l` to verify all session names
- Ran `gpu_megatron` session locally inside NeMo container — confirmed
it uses `/opt/venv/bin/python` correctly
- Manually triggered nightly-runs:
- Unit:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608013657
- GPU:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608018763
- Examples:
https://github.com/NVIDIA/Model-Optimizer/actions/runs/24608017322

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A — CI infrastructure only
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ (added `nox`
and `uv` to `dev-test`, both Apache-2.0)
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no user-facing changes

### Additional Information
Supersedes the tox-current-env workaround in the parent branch.

---------

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-18 16:56:02 +00:00
Keval Morabia 2b315eda47 Replace mip package with pulp (#663)
Replace mip package with more popular pulp package for puzzle mip
solving. Both use the CBC solver under the hood

## Testing

- Results very close for Qwen3-8B and Nemotron-Nano-12B-v2

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Chores**
* Simplified GPU test environment setup by removing unnecessary system
dependency installation
* Updated internal optimization solver dependencies in the puzzletron
module

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 21:19:36 +05:30
Farid AdilazuardaandKeval Morabia 2004779a67 Update README.md for DMS (fix cd experimental/DMS to cd Model-Optimizer/experimental/DMS) (#879)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

**Overview:** ?

## Usage
<!-- You can potentially add a usage example below. -->

```python
# Add a code snippet demonstrating how to use this
```

## Testing
<!-- Mention how have you tested your change if applicable. -->

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated DMS installation instructions to reflect the repository
structure and correct directory navigation during setup.
* Clarified the setup steps so users follow the accurate directory
change before running installation commands.
* Small wording improvements to reduce confusion during the installation
process.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Farid Adilazuarda <42537562+faridlazuarda@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:55:42 +00:00
760c980727 Add ResNet50 support for torch_onnx quantization workflow (#1263)
## Summary
- Add end-to-end ResNet50 support in the torch_onnx quantization → ONNX
export → TRT engine pipeline
- Fix multiple Conv2d-related export issues that blocked Conv2d-heavy
models from working with FP8/INT8/MXFP8/NVFP4/auto quantization modes
- Fix `configure_linear_module_onnx_quantizers` to handle all modules
with block quantization (not just `nn.Linear`), fixing NVFP4/MXFP8
export for models with quantized non-Linear modules
- Add `--trt_build` flag to `torch_quant_to_onnx.py` and simplify test
infrastructure

### Files Changed
- `modelopt/torch/_deploy/utils/torch_onnx.py` — Disable FP8 Conv2d
weight quantizers and autocast during ONNX export
- `modelopt/torch/quantization/export_onnx.py` — Fix
`configure_linear_module_onnx_quantizers` for all module types with
block quantization
- `examples/torch_onnx/torch_quant_to_onnx.py` — Add `--trt_build` flag,
calibration for FP8 override quantizers, Conv2d→FP8 override for auto
mode, filter_func updates
- `examples/torch_onnx/README.md` — Add ResNet50 to supported models
table
- `tests/examples/torch_onnx/test_torch_quant_to_onnx.py` — Add ResNet50
test entry, simplify using `--trt_build`
- `tests/_test_utils/torch/vision_models.py` — Add ResNet50 to timm
model registry

### Quantization modes passing
- ✅ FP8, INT8, MXFP8, NVFP4, Auto (all 5 modes pass export + TRT build)
- INT4_AWQ excluded (pre-existing limitation for all models)

## Test plan
- [x] All 5 resnet50 test modes pass: `pytest
tests/examples/torch_onnx/test_torch_quant_to_onnx.py -k resnet50` (5/5
passed)
- [x] Full regression: 18 passed, 2 failed (pre-existing swinv2_tiny
fp8/int8 failures)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added ResNet50 to supported ONNX export vision models with FP8, INT8,
MXFP8, and NVFP4 support.
  * Optional TensorRT engine build after export via a new CLI flag.

* **Improvements**
* Enhanced quantization calibration and export flows for FP8/INT8
models, including broader block-quantization support across module types
and safer export handling.
  * Tests updated to include ResNet50 in the model matrix.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Signed-off-by: ajrasane <arasane@nvidia.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-04-18 14:30:45 +00:00
bkartal-dev 92622a9aa6 Add nvfp4_mse and nvfp4_local_hessian options to the ptq script. (#1113)
### What does this PR do?

Type of change: Bugfix

<!-- Details about the change. -->

Add newly added quant configs to the example PTQ script.

### Testing

I have locally run auto_quantize with these two quant_configs, and
obtained successfully exported HF artifacts.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added support for three new quantization formats: nvfp4_mse,
nvfp4_local_hessian, and nvfp4_experts_only, expanding available export
options when using auto-quantize.

* **Bug Fixes / UX**
* Updated the invalid-quantization error message to include the newly
accepted format identifiers.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Bilal Kartal <bkartal@nvidia.com>
Signed-off-by: bkartal-dev <bkartal@nvidia.com>
2026-04-18 07:28:49 +00:00