mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## What does this PR do?
**Type of change:** ? <!-- Use one of the following: Bug fix, new
feature, new example, new tests, documentation. -->
new feature
**Overview:** ?
- This PR adds the sparse attention calibration algorithm
- Chunked prefill to support long ctx_len
- Separated calibration for prefill and decode
## Usage
<!-- You can potentially add a usage example below. -->
```python
import modelopt.torch.sparsity.attention_sparsity as mtsa
# Apply sparse attention with calibration
model = mtsa.sparsify(model, config=SKIP_SOFTMAX_CALIB)
# Print summary - now shows actual thresholds
mtsa.print_sparse_attention_summary(model)
# Output:
# Method: flash_skip_softmax, Threshold: Dynamic (λ=437.395926)
# Or llm_eval integration
# HuggingFace sparse attention example
python examples/llm_sparsity/attention_sparsity/hf_sa.py \
--pyt_ckpt_path Qwen/Qwen3-4B \
--sparse_attn skip_softmax_calib
```
# The calibration method
## Calibration Algorithm
- Implemented the Inverse Power model: scale_factor = k / (1 -
sparsity)^p
- Fit model parameters (k, p) per phase using scipy.optimize.curve_fit
- At inference: threshold = k / (1 - target_sparsity)^p / seqlen
## Why Choosing the Inverse Power model?
The inverse power model better fits the relationship between sparsity
ratio and threshold_scale_factor.
<img width="2388" height="1082" alt="sparsity_model_analysis"
src="https://github.com/user-attachments/assets/4dfb45d4-8c16-4f15-a878-c8e08a9b6128"
/>
## Runtime Flexibility
- Target sparsity can be changed at inference time without recalibration
- Users can adjust module._sparse_method_instance.target_sparse_ratio
dynamically
- Threshold automatically adapts to sequence length
## Testing
<!-- Mention how have you tested your change if applicable. -->
The calibration results for `Qwen/Qwen3-30B-A3B-Thinking-2507` are shown
below and are mostly consistent with the ground-truth numbers collected
from the kernel side.
```
Prefill Calibration Results:
Model: scale_factor = k / (1 - sparsity)^p
Fitted k: 1003.3990
Fitted p: 1.2589
R-squared: 0.827549
Scale factors for different target sparsities:
Target Scale Factor
---------- ---------------
50% 2401.35
70% 4568.26
80% 7610.98
90% 18214.70
95% 43591.65
```
## Before your PR is "*Ready for review*"
<!-- If you haven't finished some of the above items you can still open
`Draft` PR. -->
- **Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CONTRIBUTING.md)**
and your commits are signed.
- **Is this change backward compatible?**: Yes/No <!--- If No, explain
why. -->
- **Did you write any new necessary tests?**: Yes/No
- **Did you add or update any necessary documentation?**: Yes/No
- **Did you update
[Changelog](https://github.com/NVIDIA/TensorRT-Model-Optimizer/blob/main/CHANGELOG.rst)?**:
Yes/No <!--- Only for new features, API changes, critical bug fixes or
bw breaking changes. -->
## Additional Information
<!-- E.g. related issue. -->
---------
Signed-off-by: Kai Xu <kaix@nvidia.com>
79 lines
2.7 KiB
Python
79 lines
2.7 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
"""Utilities for sparse attention integration with llm_eval."""
|
|
|
|
import modelopt.torch.sparsity.attention_sparsity as mtsa
|
|
|
|
|
|
def _extract_model(model_obj):
|
|
"""Extract actual model from wrapper (HFLM or EvalModel)."""
|
|
if hasattr(model_obj, "gpt2"):
|
|
return model_obj.gpt2
|
|
elif hasattr(model_obj, "model"):
|
|
return model_obj.model
|
|
else:
|
|
return model_obj
|
|
|
|
|
|
def sparsify_model(
|
|
model,
|
|
sparse_cfg: str,
|
|
backend=None,
|
|
):
|
|
"""Apply sparse attention to model with optional RULER calibration.
|
|
|
|
Args:
|
|
model: Model wrapper (HFLM or EvalModel) or raw model
|
|
sparse_cfg: Sparse attention config name or dict
|
|
backend: Backend to use (optional, overrides config backend)
|
|
|
|
Returns:
|
|
The model with sparse attention applied
|
|
|
|
Note:
|
|
Calibration is automatically triggered if the config contains a 'calibration' field.
|
|
The calibration will auto-generate RULER dataset from the model's tokenizer.
|
|
"""
|
|
# Extract actual model
|
|
net = _extract_model(model)
|
|
|
|
# Resolve config
|
|
if isinstance(sparse_cfg, str):
|
|
# Get config from mtsa module (e.g., SKIP_SOFTMAX_CALIB, SKIP_SOFTMAX_DEFAULT)
|
|
mtsa_cfg = getattr(mtsa, sparse_cfg, None)
|
|
if mtsa_cfg is None:
|
|
raise ValueError(f"Unknown sparse_cfg: {sparse_cfg}.")
|
|
else:
|
|
mtsa_cfg = sparse_cfg
|
|
|
|
# Override backend if specified
|
|
if backend:
|
|
if isinstance(mtsa_cfg, dict) and "sparse_cfg" in mtsa_cfg:
|
|
modified_sparse_cfg = {}
|
|
for pattern, cfg in mtsa_cfg["sparse_cfg"].items():
|
|
modified_cfg = cfg.copy() if isinstance(cfg, dict) else cfg
|
|
if isinstance(modified_cfg, dict):
|
|
modified_cfg["backend"] = backend
|
|
modified_sparse_cfg[pattern] = modified_cfg
|
|
mtsa_cfg = {"sparse_cfg": modified_sparse_cfg}
|
|
|
|
# Apply sparsification
|
|
print(f"\nApplying sparse attention with config: {sparse_cfg}")
|
|
mtsa.sparsify(net, mtsa_cfg)
|
|
print("Sparse attention applied successfully!")
|
|
|
|
return model
|