Files
Model-Optimizer/tools
Jenny ChenandClaude Opus 4.8 e2c4d083d4 [OMNIML-4922] Four over Six PTQ & Updating Nemotron Ultra Example (#1684)
### What does this PR do?

[Four Over Six](https://arxiv.org/pdf/2512.02010) PTQ implementation for
weight-only quantization. Four Over Six was used to produce the Nemotron
3 Ultra NVFP4 checkpoint. Also updates the Ultra PTQ example in the
launcher to use this new 4/6 config
`huggingface/nvidia/Nemotron-3-Ultra-550B-A55B/ptq/ultra-nvfp4-46-max`

### Usage

```bash
uv run launch.py --yaml examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml --yes
```

### Testing

- [x] Unit tests pass
- [x] Run launcher example

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Only for new features, API changes, critical bug fixes
or backward incompatible changes. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added NVFP4 Four‑Over‑Six (4/6) adaptive per‑block weight scaling and
a configurable FP8 normalization option for FP4/FP8 quantization.

* **Documentation**
* Added PTQ recipes/presets and updated config docs to document FP8 max
variants and the Four‑Over‑Six option.

* **Tests**
* Added unit and GPU tests validating 4/6 selection, normalization
threading, scaling behavior, and reconstruction error checks.

* **Chores**
* Updated a PTQ pipeline example to use the NVFP4‑46‑max recipe and
bumped a container image; minor project config tweak.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 15:00:25 +00:00
..