mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do? Type of change: new feature Calibrates skip-softmax thresholds through the vLLM V1 execution path for FlashAttention and FlashInfer. Calibration measures the paged KV-cache path used at serving time, aggregates raw skipped/total tile counts across tensor-parallel head shards, fits separate prefill and decode curves, and exports the existing `sparse_attention_config` checkpoint schema. This uses raw counts rather than averaging per-rank sparsity ratios because TP ranks can contribute different tile populations; summing numerators and denominators before division preserves the global tile-weighted result. The vLLM adapter lives in `plugins/sparse_attn_calibration.py` rather than `SparseAttentionStatsManager`: the latter records module-local ratios for the HF calibration flow and has no aligned cross-process merge contract, while this path must merge per-sample raw counts from vLLM workers. Fitting and export still reuse `DynamicThresholdCalibrator` and the canonical conversion helpers so the model and checkpoint schema do not fork. Skip decisions depend on tile geometry. The common Triton launch boundary fixes the KV tile at 128 tokens and the prefill query tile at 128 tokens, including for direct kernel callers. Single-query decode can use a 16x128 compute tile without changing its skip decision. Measurement bypasses autotuning; serving still tunes warp and pipeline-stage counts while keeping the decision geometry fixed. ### Usage ```bash python examples/vllm_serve/calibrate_sparse_attn.py <CKPT> \ --prompts_file prompts.txt \ --target_sparse_ratio 0.7 \ --fit_logspace \ --tensor_parallel_size 4 \ --decode_tokens 32 \ --update_checkpoint_config ``` Calibration supports tensor parallelism and requires pipeline-parallel and data-parallel sizes of 1. It always writes `sparse_attention_config.json`; `--update_checkpoint_config` also merges the result into `<CKPT>/config.json`. ### Testing Latest revision `1e969cb380` (rebased onto main `02b58eb146`, 2026-09-17): - Calibration/count-fitting unit tests: **33 passed** (`test_sparse_attn_calibration.py` and `test_calibrator_fitting.py`). - Paged and contiguous calibration GPU suite: **33 passed** (`test_paged_calibrate.py` and `test_triton_fa_calibrate.py`), including NHD/HND equivalence, partial query tiles, decode counts, and malformed-cache rejection. Run with `CUDA_VISIBLE_DEVICES=1` on an RTX A6000; local GPU 0 was unavailable. - Calibration CLI tests: **21 passed** (`tests/examples/vllm_serve/test_calibrate_sparse_attn.py`). - `pre-commit run --files <four changed files>`: passed, including Ruff, mypy, and Bandit. - The new regression tests reproduced the skipped-counter truncation and missing cache-boundary checks before the fix. Calibration arithmetic and the 20-point threshold grid are unchanged. Historical validation from earlier revisions (not rerun end-to-end for this update): - `PYTHONPATH="$PWD" pytest -q tests/examples/vllm_serve/test_calibrate_sparse_attn.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_calibration.py` — 37 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_calibration.py tests/gpu_vllm/torch/sparsity/attention_sparsity/test_sparse_attn_worker.py` — 65 passed, including kv-first, blocks-first, and packed FlashAttention cache layouts. - `PYTHONPATH="$PWD" pytest -q tests/gpu_vllm/torch/sparsity/attention_sparsity/test_vllm_runtime.py tests/unit/torch/sparsity/attention_sparsity/test_sparse_attn_config.py` — 33 passed. - `PYTHONPATH="$PWD" pytest -q tests/gpu/torch/kernels/sparsity/attention/test_paged_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_calibrate.py tests/gpu/torch/kernels/sparsity/attention/test_triton_fa_skip_softmax.py` — 31 passed, 1 skipped because the GPU lacks enough shared memory for the fp32 tile. - `pre-commit run --files <changed files>` — passed. - Historical end-to-end Nemotron 3 Ultra (GCP job `558552`), TP4, FA4, 48 RULER prompts, and 20 threshold trials: completed `0:0` with prefill `(a, b) = (9.9104, 10.8881)`, respectively +0.147% and -0.066% versus the matching 20-point reference `(9.8958, 10.8953)`. The supplied legacy fit `(14.47, 10.91)` used a different threshold grid; its `b` differs by only -0.201%, while `a` retains the known grid-weighting shift. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ❌ Active skip-softmax fixes the calibrated decision geometry (serving still tunes warp/stage counts), and sparse-only vLLM installs fail fast for unsupported DCP, DBO/ubatching, speculative decoding, and FULL mixed-batch graphs instead of installing silently. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no copied code or new dependency. - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ ### Additional Information Pipeline parallelism is rejected during calibration because the current count-merging contract aligns records across tensor-parallel head shards, not across pipeline stages with disjoint attention layers. The unrelated HF padded-query behavior change was removed from this PR so it can be reviewed independently with its own compatibility test. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added vLLM skip-softmax calibration for paged attention, including prefill/decode support and checkpoint configuration generation. - Added Muse Glimmer AutoQuantize, Alpamayo QAD, streaming Kimi-K3 conversion, and NVFP4 activation headroom calibration recipes. - Added calibration statistics aggregation, phase-specific fitting, threshold validation, and preservation of existing sparse-attention settings. - **Bug Fixes** - Improved NVFP4 CPU/ONNX scale validation and clamping. - Added clearer handling for unsupported quantization, cache, CUDA graph, and engine configurations. - Standardized serving and calibration tile behavior. - **Documentation** - Expanded vLLM serving guidance, calibration instructions, compatibility requirements, and sparse-attention limitations. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kai Xu <kaix@nvidia.com>