Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)

### What does this PR do?

Type of change: documentation

`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.

This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.

It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.

**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.

### Usage

No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:

```python
from modelopt.recipe import load_recipe

cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```

### Testing

Docs-only change; no code paths touched, so no new or updated tests.

- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
  including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
  and the existing `CHANGELOG.rst` entry, so the documented defaults
  (`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
  NVFP4-input-quantizer-only scope match the code.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->

### Additional Information

The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Chenjie Luo
2026-09-16 10:56:23 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent 6a4b3f147e
commit 655f94c207
+27 -1
View File
@@ -171,6 +171,31 @@ How the quantization scales are searched. The default (no suffix) is `max`.
(amax) calibrated as in the default recipes. Costs more calibration time but
recovers accuracy NVFP4 W4A4 can lose under plain max. Reach for it when a
`max` recipe regresses.
- **`nvfp4_act_headroom`** (`nvfp4_act_headroom-kv_fp8_cast`) — leaves headroom
on the NVFP4 **activation** global scale. Plain `max` anchors that scale to the
largest per-block amax seen during calibration, so any larger activation at
inference saturates. This variant instead sets `amax = max(rho * anchor, upper)`
from the per-block amaxes at `anchor_percentile` (default 1) and
`upper_percentile` (default 99.99), with headroom factor `rho` (default 16384):
calibrated blocks sit low in the FP8 block-scale range, leaving the rest for
unseen outliers. Costs one calibration pass and exports a standard NVFP4
checkpoint; coverage matches `nvfp4_default-kv_fp8_cast`, and the nested
`weight_scale_algorithm` still accepts `mse` / `local_hessian`. Affects NVFP4
**input** quantizers only — a no-op for FP8 and weight-only recipes.
Reach for it when a W4A4 recipe regresses and the **activations**, not the
weights, are to blame: an A16 ablation of the same scope clears it while `mse`
does not; the symptom is behavioral (verbose or runaway generations, hitting
the generation cap) rather than a flat score drop; inference contexts run
longer than the calibration set; or a few rare blocks dominate the activation
error. MoE experts-only scopes are the common case — one observer covers every
expert in a layer, and `down_proj` inputs clip first. On a GLM-5.3-Flash
experts-only W4A4 study (SciCode, temperature 1.0) it cut the median
generation-length regression from +38% to +19% and the mean from +19% to +4%
with no capped generations: the best strict-W4A4 result there, but still short
of the p50/p75 gate. A strong first lever, not a guaranteed fix — and sweep
`rho`, since headroom above the calibrated range costs resolution inside it.
- **`input_scale1`** (`nvfp4_experts_only_input_scale1-kv_fp8_cast`) — pins the
expert **activation** per-tensor amax to a constant `2688.0`
(= E2M1_MAX × E4M3_MAX = 6 × 448) via `constant_amax`, so the exported NVFP4
@@ -220,7 +245,8 @@ These can also be **stacked** when a single method isn't enough — e.g. `mse` +
target requires, checking accuracy as you go.
3. **Recover accuracy via calibration before backing off the scope.** If a
wider-scope recipe regresses, switch its `max` to the `mse` variant before
retreating to a narrower scope.
retreating to a narrower scope. If `mse` doesn't clear it but an A16 ablation
does, the **activations** are the problem: try `nvfp4_act_headroom` next.
4. **Pick KV by deployment.** `kv_fp8_cast` is the safe default (usually as
accurate as calibrated `kv_fp8`); use `kv_nvfp4_cast` for maximum KV
compression.