Files
Model-Optimizer/modelopt
Chenjie Luo 42482b1b0f Add nvfp4_omlp_only config and simplify the config.py (#973)
### What does this PR do?

Type of change: ? new feature

1) Add vfp4_omlp_only config == nvfp4_mlp_only + o_proj quant
2) Add block sparse MOE to mlp only config
3) Simplfiy config.py
4) Update readme in llm_ptq mention these two configs for better
accuracy.

### Usage

huggingface_script.sh ... --quant nvfp4_omlp_only


### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, using
`torch.load(..., weights_only=True)`, avoiding `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other source, did you follow IP policy in
[CONTRIBUTING.md](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md#-copying-code-from-other-sources)?:
✅ / ❌ / N/A <!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added nvfp4_omlp_only quantization format for NVFP4, enabling
selective quantization of MLP and output projection layers while
preserving attention QKV projection accuracy.

* **Changed**
* pass_through_bwd now defaults to True; set to False if using STE with
zeroed outlier gradients for better QAT accuracy.

* **Documentation**
* Updated post-training quantization guidance with NVFP4-specific
configuration recommendations and usage examples.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
2026-03-06 01:03:25 +00:00
..