Files
Model-Optimizer/examples
sychen52 d0c01a4e96 Add p quantization to our triton fa kernel (#1757)
### What does this PR do?

Type of change: new feature

- New P_QDQ feature in the Triton FA kernel: fake-quant softmax P before
P·V, modes fp8/nvfp4; API attention(..., p_qdq,
  p_qdq_scale); denominator unquantized, backward is STE.
- New quantization/attention/p_qdq.py +
quantization/common/fp8_quant.py; reuses nvfp4_quant FP4 rounding.
- _QuantAttention: softmax_quantizer→p_bmm_quantizer, dispatches
FP8/NVFP4 to the kernel (no kitchen); adds
TensorQuantizer.is_fp8/is_nvfp4_dynamic; envelope guards for unsupported
cases.
- Recipe wildcard + vLLM reload updated for the rename; nvfp4_tensor.py
comment typo fixed; ruff ignores + tests added.

### Usage

### Testing
added unit tests

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ 
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Summary

* **New Features**
* Added softmax probability quant-dequant (P_QDQ) support via `p_qdq`
and `p_qdq_scale`.
* When enabled, quantized attention can route through the Triton P_QDQ
path.
* Added stricter Triton attention “envelope” validation with
`validate_triton_attention_envelope`.

* **Bug Fixes**
* Improved handling of quantizer keys to correctly skip softmax-P
(`p_bmm_quantizer`) entries.

* **Tests**
* Added GPU forward/backward coverage for FP8 (E4M3) and NVFP4 (E2M1),
including reference comparisons and invalid-parameter cases.

* **Documentation**
* Updated attention quantization configs to use `p_bmm` quantizers
instead of `softmax_quantizer`.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-06-24 12:06:03 -07:00
..
2026-06-08 22:11:37 +00:00
2026-06-23 21:42:37 +05:30
2026-06-23 21:42:37 +05:30
2026-06-23 21:42:37 +05:30
2026-06-23 21:42:37 +05:30
2026-06-23 21:42:37 +05:30
2026-06-23 21:42:37 +05:30