mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: New feature
Add `NVFP4_EXPERTS_ONLY_CFG` quantization config that targets only MoE
expert layers (`*mlp.experts*` and `*block_sparse_moe*`) with NVFP4
(W4A4) quantization, leaving all other layers (including non-expert MLP)
unquantized. This is useful for MoE models where selectively quantizing
only expert layers provides a good accuracy-performance tradeoff.
Changes:
- Refactored `_nvfp4_experts_only_quant_cfg` as a reusable building
block in `config.py`, with `_nvfp4_mlp_only_quant_cfg` now composing on
top of it
- Added `NVFP4_EXPERTS_ONLY_CFG` to the Python config choices
- Added corresponding `nvfp4_experts_only-fp8_kv.yml` YAML recipe to the
new recipe system (`modelopt_recipes/general/ptq/`)
- Updated `hf_ptq.py`, `multinode_ptq.py`, example scripts, and README
to include the new config
### Usage
```python
import modelopt.torch.quantization as mtq
model = mtq.quantize(model, mtq.NVFP4_EXPERTS_ONLY_CFG, forward_loop)
```
Or via the YAML recipe system:
```python
from modelopt.recipe import load_recipe
recipe = load_recipe("general/ptq/nvfp4_experts_only-fp8_kv")
```
### Testing
- Verified the YAML recipe matches the Python config definition
- Existing unit tests cover the quantization config infrastructure
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <\!-- Config is exercised
by existing quantization tests -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ <\!-- Minor config addition -->
### Additional Information
The `experts_only` config is a subset of `mlp_only`: it quantizes
`*mlp.experts*` and `*block_sparse_moe*` patterns but not the broader
`*mlp*` pattern. The Python config was refactored so
`_nvfp4_mlp_only_quant_cfg` composes on top of
`_nvfp4_experts_only_quant_cfg`.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an "experts-only" NVFP4 quantization option that selectively
quantizes MoE expert layers (preserving dense MLP/attention) for
improved PTQ accuracy.
* Added a corresponding PTQ recipe enabling expert-only W4A4
quantization with FP8 KV cache support.
* **Documentation**
* Updated README, examples, scripts, and changelog to document and
surface the new experts-only quantization choice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Signed-off-by: Chenjie Luo <108829653+cjluo-nv@users.noreply.github.com>