Files
Model-Optimizer/.github
kaix-nv d23030f91d [1/n] Adds skip-softmax calibration through the vLLM serving path (#1992)
### What does this PR do?

Type of change: new feature

Calibrates skip-softmax thresholds through the vLLM V1 execution path
for FlashAttention and FlashInfer. Calibration measures the paged
KV-cache path used at serving time, aggregates raw skipped/total tile
counts across tensor-parallel head shards, fits separate prefill and
decode curves, and exports the existing `sparse_attention_config`
checkpoint schema.

This uses raw counts rather than averaging per-rank sparsity ratios
because TP ranks can contribute different tile populations; summing
numerators and denominators before division preserves the global
tile-weighted result. The vLLM adapter lives in
`plugins/sparse_attn_calibration.py` rather than
`SparseAttentionStatsManager`: the latter records module-local ratios
for the HF calibration flow and has no aligned cross-process merge
contract, while this path must merge per-sample raw counts from vLLM
workers. Fitting and export still reuse `DynamicThresholdCalibrator` and
the canonical conversion helpers so the model and checkpoint schema do
not fork.

Skip decisions depend on tile geometry. The common Triton launch
boundary fixes the KV tile at 128 tokens and the prefill query tile at
128 tokens, including for direct kernel callers. Single-query decode can
use a 16x128 compute tile without changing its skip decision.
Measurement bypasses autotuning; serving still tunes warp and
pipeline-stage counts while keeping the decision geometry fixed.

### Usage

```bash
python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \
  --prompts_file prompts.txt \
  --target_sparse_ratio 0.7 \
  --fit_logspace \
  --tensor_parallel_size 4 \
  --decode_tokens 32 \
  --update_checkpoint_config
```

Calibration supports tensor parallelism and requires pipeline-parallel
and data-parallel sizes of 1. It always writes
`sparse_attention_config.json`; `--update_checkpoint_config` also merges
the result into `<CKPT>/config.json`.

### Testing

Latest revision `1e969cb380` (rebased onto main `02b58eb146`,
2026-09-17):

- Calibration/count-fitting unit tests: **33 passed**
(`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`).
- Paged and contiguous calibration GPU suite: **33 passed**
(`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including
NHD/HND equivalence, partial query tiles, decode counts, and
malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX
A6000; local GPU 0 was unavailable.
- Calibration CLI tests: **21 passed**
(`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`).
- `pre-commit run --files <four changed files>`: passed, including Ruff,
mypy, and Bandit.
- The new regression tests reproduced the skipped-counter truncation and
missing cache-boundary checks before the fix. Calibration arithmetic and
the 20-point threshold grid are unchanged.

Historical validation from earlier revisions (not rerun end-to-end for
this update):

- `PYTHONPATH="$PWD" pytest -q
tests/examples/vllm_serve/test_calibrate_sparse_attn.py
tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py`
— 37 passed.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py`
— 65 passed, including kv-first, blocks-first, and packed FlashAttention
cache layouts.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py
tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py`
— 33 passed.
- `PYTHONPATH="$PWD" pytest -q
tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py
tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py
tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py`
— 31 passed, 1 skipped because the GPU lacks enough shared memory for
the fp32 tile.
- `pre-commit run --files <changed files>` — passed.
- Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4,
48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill
`(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus
the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy
fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by
only -0.201%, while `a` retains the known grid-weighting shift.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ❌ Active skip-softmax fixes the
calibrated decision geometry (serving still tunes warp/stage counts),
and sparse-only vLLM installs fail fast for unsupported DCP,
DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead
of installing silently.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no
copied code or new dependency.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pipeline parallelism is rejected during calibration because the current
count-merging contract aligns records across tensor-parallel head
shards, not across pipeline stages with disjoint attention layers. The
unrelated HF padded-query behavior change was removed from this PR so it
can be reviewed independently with its own compatibility test.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added vLLM skip-softmax calibration for paged attention, including
prefill/decode support and checkpoint configuration generation.
- Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3
conversion, and NVFP4 activation headroom calibration recipes.
- Added calibration statistics aggregation, phase-specific fitting,
threshold validation, and preservation of existing sparse-attention
settings.

- **Bug Fixes**
  - Improved NVFP4 CPU/ONNX scale validation and clamping.
- Added clearer handling for unsupported quantization, cache, CUDA
graph, and engine configurations.
  - Standardized serving and calibration tile behavior.

- **Documentation**
- Expanded vLLM serving guidance, calibration instructions,
compatibility requirements, and sparse-attention limitations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-09-18 12:31:10 -07:00
..
2026-09-17 20:08:12 +00:00