Files
Shiyang Chen 439ad4bcf5 Add vLLM NVFP4 MLA KV-cache fake quant; fix FP8 and CUDA-graph serving
- Add the kv_nvfp4_mla recipe unit (NVFP4 on MLA kv_c, global scale 1).
- Fix update_kv_cfg_for_mla on vLLM >= 0.16 so KV_QUANT_CFG reaches MLA.
- Pass through layers vLLM already quantizes (e.g. FP8 checkpoints).
- Accept quantized KV-cache dtypes such as --kv-cache-dtype fp8.
- Make KV fake quant CUDA-graph safe: move quantizers to the GPU and
  build the use_constant_amax amax on the device.
- Default VLLM_DISABLE_COMPILE_CACHE=1: the torch.compile cache could
  reuse a graph without the fake quant.

Signed-off-by: Shiyang Chen <shiychen@nvidia.com>
2026-10-01 11:15:17 -07:00
..