mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
- Add the kv_nvfp4_mla recipe unit (NVFP4 on MLA kv_c, global scale 1). - Fix update_kv_cfg_for_mla on vLLM >= 0.16 so KV_QUANT_CFG reaches MLA. - Pass through layers vLLM already quantizes (e.g. FP8 checkpoints). - Accept quantized KV-cache dtypes such as --kv-cache-dtype fp8. - Make KV fake quant CUDA-graph safe: move quantizers to the GPU and build the use_constant_amax amax on the device. - Default VLLM_DISABLE_COMPILE_CACHE=1: the torch.compile cache could reuse a graph without the fake quant. Signed-off-by: Shiyang Chen <shiychen@nvidia.com>