mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: Bug fix The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute >= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block scale below 1e-5 with 1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static Triton kernel, the CUDA extension fallback and NVFP4 export have no such floor; the floor was only guarding division by zero. The Triton kernel had its own copy of the scale/round code instead of the shared `nvfp4_scalar_quant`. - Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block scale). - Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule. - Both: a zero, inf or NaN global amax uses a unit block scale, like the CUDA extension (the conv kernels returned NaN for inf/NaN before). Blocks with scale >= 1e-5 are unchanged. - New tests: power-of-two scaling of input and global amax (2^-10, 2^-20) scales the output by the same factor (Triton, standalone conv FP4, fused conv3d); invalid global amax gives unit-scale rounding. The conv test's Python reference drops the floor. - B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization` NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ). - Negative control on `main`: the small-input tests fail on both kernels; conv also fails for inf/NaN global amax. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ (only blocks with scale < 1e-5 or an inf/NaN global amax change, from zeros/NaN to correct values) - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * FP4 quantization now preserves proportional output when inputs and valid global scales are reduced together, including at small scales. * Zero, infinite, or NaN global scales use a safe fallback, preserving inputs already representable in FP4. * Small positive block scales are no longer discarded by an absolute scale threshold; subnormal scale handling is covered across quantization paths. * **Tests** * Added coverage for scale consistency across input types and block sizes, invalid global scales, and subnormal FP8 block scales. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shiyang Chen <shiychen@nvidia.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>