Files
Model-Optimizer/tests
sychen52andClaude Opus 5.5 834c90d7a1 Fix NVFP4 fake quant zeroing blocks with small scales (#2549)
### What does this PR do?

  Type of change: Bug fix

The dynamic NVFP4 Triton kernel (`fp4_fake_quant_block`, used on compute
>= 8.9) and the Conv3D implicit-GEMM CUDA kernels replaced any FP8 block
scale below 1e-5 with
1.0, so every block whose max |x| was below ~6e-5 was zeroed. The static
Triton kernel, the CUDA extension fallback and NVFP4 export have no such
floor; the floor was
only guarding division by zero. The Triton kernel had its own copy of
the scale/round code instead of the shared `nvfp4_scalar_quant`.

- Triton: use the shared `nvfp4_scalar_quant` (zero only on a zero block
scale).
- Conv3D CUDA (fused kernel and standalone `fp4_fake_quant`): same rule.
- Both: a zero, inf or NaN global amax uses a unit block scale, like the
CUDA extension (the conv kernels returned NaN for inf/NaN before).

  Blocks with scale >= 1e-5 are unchanged.

- New tests: power-of-two scaling of input and global amax (2^-10,
2^-20) scales the output by the same factor (Triton, standalone conv
FP4, fused conv3d); invalid
global amax gives unit-scale rounding. The conv test's Python reference
drops the floor.
- B200: 505 passed / 31 skipped (`tests/gpu/torch/quantization`
NVFP4/FP4 files) and 180 passed (conv implicit GEMM + attention P-QDQ).
- Negative control on `main`: the small-input tests fail on both
kernels; conv also fails for inf/NaN global amax.

  ### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (only blocks with scale < 1e-5
or an inf/NaN global amax change, from zeros/NaN to correct values)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
  - Did you write any new necessary tests?: ✅
  - Did you update Changelog?: N/A
  - Did you get Claude approval on this PR?: ❌




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* FP4 quantization now preserves proportional output when inputs and
valid global scales are reduced together, including at small scales.
* Zero, infinite, or NaN global scales use a safe fallback, preserving
inputs already representable in FP4.
* Small positive block scales are no longer discarded by an absolute
scale threshold; subnormal scale handling is covered across quantization
paths.
* **Tests**
* Added coverage for scale consistency across input types and block
sizes, invalid global scales, and subnormal FP8 block scales.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-29 10:31:31 -07:00
..