Files
Model-Optimizer/tests
kinjalpatel27 02b58eb146 Fix vLLM fakequant calibration for hybrid attention models (#2414)
### What does this PR do?

Type of change: Bug fix

Fix fakequant calibration for hybrid attention/Mamba models, including
NVIDIA Nemotron-3-Nano, on vLLM 0.26 and 0.28.

The manual calibration scheduler path previously submitted requests with
empty KV-cache block tables. Hybrid models require scheduler-compatible
cache state during prefill; on current vLLM releases the empty tables
caused the Mamba state to use the reserved null block and calibration
activations became NaN. Request cleanup also no longer matched the vLLM
0.28 execution lifecycle, which could leave request-scoped state in the
persistent batch.

This PR:

- Allocates non-null scratch blocks for every KV-cache group using the
vLLM warmup reservation policy.
- Supports both the vLLM 0.28 reservation helper and the equivalent vLLM
0.26 calculation.
- Passes newly allocated blocks through `new_block_ids_to_zero` when
that scheduler field is available.
- Validates that the calibration batch fits in the configured cache and
reports how to reduce calibration demand if it does not.
- Cleans up calibration requests through a zero-token scheduler step on
current vLLM, with a direct cleanup fallback for older runners.
- Updates the example Dockerfile to default to vLLM 0.28.0 while
retaining vLLM 0.26.0 through `VLLM_VERSION`.
- Documents the validated Nemotron-3-Nano NVFP4 KV-cache workflow and
clarifies that reducing `--max-num-batched-tokens` is not required.

### Usage

Build the default vLLM 0.28.0 image:

```bash
docker build -f examples/vllm_serve/Dockerfile \
  -t vllm-modelopt:v0.28.0 .
```

Build with vLLM 0.26.0:

```bash
docker build --build-arg VLLM_VERSION=0.26.0 \
  -f examples/vllm_serve/Dockerfile \
  -t vllm-modelopt:v0.26.0 .
```

Calibrate and serve Nemotron-3-Nano with NVFP4 KV-cache fakequant:

```bash
KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \
  python examples/vllm_serve/vllm_serve_fakequant.py \
  <nemotron3_nano_model_path> \
  --trust-remote-code --enforce-eager -tp 8 \
  --max-model-len 8192 --host 0.0.0.0 --port 8000
```

### Testing

Validated on omniml-a0 with `NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`,
tensor parallel size 8, `NVFP4_KV_CFG`, `QUANT_CALIB_SIZE=512`, and
`--max-model-len 8192`. No `--max-num-batched-tokens` override was used.

- vLLM 0.28.0:
  - All 512 calibration samples completed.
  - No NaNs or cache-cleanup warnings were observed.
  - The server started and `/health` passed.
- An OpenAI-compatible completion request returned coherent generated
text.
- vLLM 0.26.0:
- Repeated the same 512-sample TP8 calibration with the official
`vllm/vllm-openai:v0.26.0` image.
  - No NaNs were observed.
- The server started, passed `/health`, and returned coherent generated
text.
- Docker:
  - Built and verified the updated vLLM 0.28.0 image.
- Focused tests:
  - `tests/examples/vllm_serve/test_vllm_mlflow_utils.py`: 32 passed.
- Cleanup failure, missing legacy API, and legacy fallback tests: 5
passed on both vLLM 0.26.0 and 0.28.0.
- Repository hooks:
- Targeted pre-commit hooks for every changed Python, Markdown, and
Docker file: passed.
  - `git diff --check`: passed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — added focused coverage for
fail-closed cleanup, exception chaining, and the legacy cleanup
fallback; the full regression was also validated end to end.
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

The change is quantization-format agnostic. It corrects the calibration
scheduler and cache lifecycle rather than special-casing `NVFP4_KV_CFG`
or using an NVFP4 cast path.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added support for configuring the vLLM version through `VLLM_VERSION`,
with vLLM 0.28.0 as the default.
- Added calibration and serving guidance for hybrid attention/Mamba
models, including Nemotron 3 Nano with NVFP4 KV-cache fake quantization.

- **Bug Fixes**
  - Improved calibration block handling across supported vLLM versions.
- Improved calibration cleanup to preserve original errors and provide
reliable fallback behavior when standard cleanup is unavailable.

- **Documentation**
- Documented tested versions, direct installation commands, ModelOpt
setup, serving options, and guidance to avoid NaNs during batched
serving.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
2026-09-17 14:12:34 -07:00
..