Files
Model-Optimizer/examples
mxinO ca7eb64ad0 DSV4 PTQ example with dequant on the fly (#1341)
### What does this PR do?

Type of change: new example <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->

Add deepseek v4 official modeling ptq example

### Usage

See readme, and it requires the vllm PR:
https://github.com/vllm-project/vllm/pull/42209

### Testing
Tested with ptq and export of dsv4 flash and served with vllm.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added DeepSeek‑V4 routed‑expert post‑training quantization and an
NVFP4 checkpoint conversion utility.

* **Documentation**
* Expanded DeepSeek quantization guide with directory layout, updated
V3/V3.2 workflows, and detailed V4 routed‑expert calibration,
single/multi‑node examples, and export guidance.

* **Chores**
  * Made example quantization scripts location‑independent.
* Updated pre‑commit license hook to skip DeepSeek example quantization
files.

<!-- review_stack_entry_start -->

[![Review Change
Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/NVIDIA/Model-Optimizer/pull/1341?utm_source=github_walkthrough&utm_medium=github&utm_campaign=change_stack)

<!-- review_stack_entry_end -->
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Meng Xin <mxin@nvidia.com>
2026-06-04 05:02:30 +00:00
..