Disable QKV NVFP4 quantization for Qwen3 MOE (#735)

## What does this PR do?

**Type of change:** ? Recipe improvement

**Overview:** ?

Disable QKV NVFP4 quantization for Qwen3 MOE models following the Qwen3
Next recipe for accuracy recovery

## Testing
Model accuracy benchmarking

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
This commit is contained in:
Chenjie Luo
2026-01-02 11:06:25 -08:00
committed by GitHub
parent b655321d87
commit d541324e84
2 changed files with 2 additions and 1 deletions
+1 -1
View File
@@ -180,7 +180,7 @@ def build_quant_cfg(
quant_cfg["quant_cfg"]["*image*"] = {"enable": False}
quant_cfg["quant_cfg"]["*vision*"] = {"enable": False}
if model_type == "qwen3next" and qformat == "nvfp4":
if model_type in ["qwen3moe", "qwen3next"] and qformat == "nvfp4":
# Disable the attention projection layers to retain accuracy
quant_cfg["quant_cfg"]["model*.*attn*in_proj*"] = {"enable": False}
quant_cfg["quant_cfg"]["model*.*attn*q_proj*"] = {"enable": False}
+1
View File
@@ -29,6 +29,7 @@ MODEL_NAME_TO_TYPE = {
"MPT": "mpt",
"Bloom": "bloom",
"ChatGLM": "chatglm",
"Qwen3Moe": "qwen3moe",
"Qwen3Next": "qwen3next",
"QWen": "qwen",
"RecurrentGemma": "recurrentgemma",