Files
kaix-nv 9e38041d34 [OMNIML-2850] [3/n] Adds sparse attention calibration (#538)
## What does this PR do?

**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature

**Overview:** ?
- This PR adds the sparse attention calibration algorithm
- Chunked prefill to support long ctx_len
- Separated calibration for prefill and decode

## Usage
<!-- You can potentially add a usage example below. -->

```python
import modelopt.torch.sparsity.attention_sparsity as mtsa

# Apply sparse attention with calibration
model = mtsa.sparsify(model, config=SKIP_SOFTMAX_CALIB)

# Print summary - now shows actual thresholds
mtsa.print_sparse_attention_summary(model)
# Output:
# Method: flash_skip_softmax, Threshold: Dynamic (λ=437.395926)

# Or llm_eval integration
# HuggingFace sparse attention example
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
    --pyt_ckpt_path Qwen/Qwen3-4B \
    --sparse_attn skip_softmax_calib 
```

# The calibration method
## Calibration Algorithm
- Implemented the Inverse Power model: scale_factor = k / (1 -
sparsity)^p
- Fit model parameters (k, p) per phase using scipy.optimize.curve_fit
- At inference: threshold = k / (1 - target_sparsity)^p / seqlen

## Why Choosing the Inverse Power model?
The inverse power model better fits the relationship between sparsity
ratio and threshold_scale_factor.
<img width="2388" height="1082" alt="sparsity_model_analysis"
src="https://github.com/user-attachments/assets/4dfb45d4-8c16-4f15-a878-c8e08a9b6128"
/>

## Runtime Flexibility
- Target sparsity can be changed at inference time without recalibration
- Users can adjust module._sparse_method_instance.target_sparse_ratio
dynamically
- Threshold automatically adapts to sequence length

## Testing
<!-- Mention how have you tested your change if applicable. -->
The calibration results for `Qwen/Qwen3-30B-A3B-Thinking-2507` are shown
below and are mostly consistent with the ground-truth numbers collected
from the kernel side.
```
Prefill Calibration Results:
  Model: scale_factor = k / (1 - sparsity)^p
  Fitted k: 1003.3990
  Fitted p: 1.2589
  R-squared: 0.827549

Scale factors for different target sparsities:
  Target     Scale Factor
  ---------- ---------------
  50%        2401.35
  70%        4568.26
  80%        7610.98
  90%        18214.70
  95%        43591.65
```

## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->

- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->

## Additional Information
<!-- E.g. related issue. -->

---------

Signed-off-by: Kai Xu <kaix@nvidia.com>
2026-02-18 02:18:12 +00:00

79 lines
2.7 KiB
Python

# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Utilities for sparse attention integration with llm_eval."""
import modelopt.torch.sparsity.attention_sparsity as mtsa
def _extract_model(model_obj):
"""Extract actual model from wrapper (HFLM or EvalModel)."""
if hasattr(model_obj, "gpt2"):
return model_obj.gpt2
elif hasattr(model_obj, "model"):
return model_obj.model
else:
return model_obj
def sparsify_model(
model,
sparse_cfg: str,
backend=None,
):
"""Apply sparse attention to model with optional RULER calibration.
Args:
model: Model wrapper (HFLM or EvalModel) or raw model
sparse_cfg: Sparse attention config name or dict
backend: Backend to use (optional, overrides config backend)
Returns:
The model with sparse attention applied
Note:
Calibration is automatically triggered if the config contains a 'calibration' field.
The calibration will auto-generate RULER dataset from the model's tokenizer.
"""
# Extract actual model
net = _extract_model(model)
# Resolve config
if isinstance(sparse_cfg, str):
# Get config from mtsa module (e.g., SKIP_SOFTMAX_CALIB, SKIP_SOFTMAX_DEFAULT)
mtsa_cfg = getattr(mtsa, sparse_cfg, None)
if mtsa_cfg is None:
raise ValueError(f"Unknown sparse_cfg: {sparse_cfg}.")
else:
mtsa_cfg = sparse_cfg
# Override backend if specified
if backend:
if isinstance(mtsa_cfg, dict) and "sparse_cfg" in mtsa_cfg:
modified_sparse_cfg = {}
for pattern, cfg in mtsa_cfg["sparse_cfg"].items():
modified_cfg = cfg.copy() if isinstance(cfg, dict) else cfg
if isinstance(modified_cfg, dict):
modified_cfg["backend"] = backend
modified_sparse_cfg[pattern] = modified_cfg
mtsa_cfg = {"sparse_cfg": modified_sparse_cfg}
# Apply sparsification
print(f"\nApplying sparse attention with config: {sparse_cfg}")
mtsa.sparsify(net, mtsa_cfg)
print("Sparse attention applied successfully!")
return model