mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add W4A16 NVFP4-MSE Qwen3.5 dense/MoE PTQ recipes (#1620)
### What does this PR do? Type of change: new feature (PTQ recipe) Adds an MSE-calibrated counterpart of the existing `w4a16_nvfp4-fp8_attn-kv_fp8_cast` PTQ recipe for the Qwen3.5 family (dense `qwen3_5` and MoE `qwen3_5_moe`). New files: - `modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml` — shared `quant_cfg` snippet - `modelopt_recipes/huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` — dense recipe - `modelopt_recipes/huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` — MoE recipe The only difference from the `max` variant: NVFP4 MLP / `lm_head` weight scales come from an MSE FP8-scale sweep (`method: mse`, `fp8_scale_sweep: true`, `nvfp4_static`) instead of max calibration. FP8 attention / linear-attention projections and the FP8 KV cast are unchanged. The dense and MoE families share a single `quant_cfg` snippet under `qwen3_5/ptq`, matching the existing recipe's layout. ### Usage ```bash # Dense python hf_ptq.py --recipe huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ... # MoE python hf_ptq.py --recipe huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast ... ``` ### Testing - Both recipes load and resolve their `$import`s via `modelopt.recipe.loader.load_recipe`. - The `check-modelopt-recipes` pre-commit validator passes on all three files. - Verified against the source that weight-only MSE (`method: mse` + `fp8_scale_sweep: true`) is supported for W4A16: `mse_calibrate` refines only weight quantizers (`iter_weights_for_calibration`) and needs no input/activation quantizers, and the `nvfp4_static` numeric satisfies `is_nvfp4_static` so the FP8 scale sweep engages on the MLP/lm_head weights. No code changes were required. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A (config-only; covered by existing recipe-loader validation) - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: ❌ (not yet) 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit ## New Features * Added PTQ recipe configurations for Qwen3.5 and Qwen3.5-MoE model families * Supports W4A16 quantization with NVFP4 static weights and MSE-based calibration * Enables FP8 precision for self-attention and KV-cache optimization for improved model performance and reduced memory footprint <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chenjie Luo <chenjiel@nvidia.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
7f23d0f691
commit
106781659e
+100
@@ -0,0 +1,100 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Shared `quant_cfg` snippet for the Qwen3.5 family's
|
||||
# `w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast` recipe. Imported by both
|
||||
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (dense `qwen3_5`)
|
||||
# and `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml` (MoE
|
||||
# `qwen3_5_moe`); the two families share the hybrid linear-attention +
|
||||
# softmax-attention architecture, so the wildcard rules apply identically.
|
||||
# MoE-only patterns inside `default_disabled_quantizers`
|
||||
# (`*block_sparse_moe.gate*`, `*mlp.shared_expert_gate.*`, `*router*`) are
|
||||
# no-ops on dense.
|
||||
#
|
||||
# MSE variant of `w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`: the NVFP4
|
||||
# MLP/lm_head weight quantizers use static scales (`nvfp4_static`) populated by
|
||||
# the MSE FP8-scale sweep (`method: mse`, `fp8_scale_sweep: true` in the parent
|
||||
# recipes) instead of dynamic per-call scaling.
|
||||
|
||||
# modelopt-schema: modelopt.torch.quantization.config.QuantizerCfgListConfig
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
fp8: configs/numerics/fp8
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
nvfp4_static: configs/numerics/nvfp4_static
|
||||
---
|
||||
- $import: base_disable_all
|
||||
|
||||
# W4A16 NVFP4 on MLP projection targets, with static weight scales from the
|
||||
# MSE FP8-scale sweep. The gate/up/down projection patterns cover dense MLPs,
|
||||
# shared experts, and fused MoE expert quantizers
|
||||
# (e.g. gate_up_proj_weight_quantizers.N).
|
||||
- quantizer_name: '*mlp*gate_proj*weight_quantizer*'
|
||||
cfg: {$import: nvfp4_static}
|
||||
- quantizer_name: '*mlp*up_proj*weight_quantizer*'
|
||||
cfg: {$import: nvfp4_static}
|
||||
- quantizer_name: '*mlp*down_proj*weight_quantizer*'
|
||||
cfg: {$import: nvfp4_static}
|
||||
|
||||
# FP8 self-attention projections.
|
||||
- quantizer_name: '*self_attn*weight_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*self_attn*input_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
|
||||
# FP8 large linear-attention projections. in_proj_a and in_proj_b are
|
||||
# re-disabled explicitly below; conv1d stays disabled via base_disable_all
|
||||
# (no rule re-enables it).
|
||||
- quantizer_name: '*linear_attn.in_proj_qkv*weight_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*linear_attn.in_proj_qkv*input_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*linear_attn.in_proj_z*weight_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*linear_attn.in_proj_z*input_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*linear_attn.out_proj*weight_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
- quantizer_name: '*linear_attn.out_proj*input_quantizer'
|
||||
cfg: {$import: fp8}
|
||||
|
||||
# FP8 KV cache with constant amax.
|
||||
- $import: kv_fp8_cast
|
||||
|
||||
# Standard exclusions (BatchNorm, LeakyReLU, gates, routers, conv1d, output
|
||||
# heads, etc.). Includes `*lm_head*` disable, which is re-enabled below.
|
||||
- $import: default_disabled_quantizers
|
||||
|
||||
# Qwen-specific exclusions: linear-attention sub-modules that are not in the
|
||||
# reference recipe, and any visual / MTP siblings on multimodal releases.
|
||||
- quantizer_name: '*linear_attn.in_proj_a*'
|
||||
enable: false
|
||||
- quantizer_name: '*linear_attn.in_proj_b*'
|
||||
enable: false
|
||||
- quantizer_name: '*visual*'
|
||||
enable: false
|
||||
- quantizer_name: '*vision_tower*'
|
||||
enable: false
|
||||
# Name-match for "mtp"; complementary runtime path in hf_ptq.py catches
|
||||
# MTP layers identified by index instead.
|
||||
- quantizer_name: '*mtp*'
|
||||
enable: false
|
||||
|
||||
# Re-enable NVFP4 on lm_head weights (static MSE scales). Must come after
|
||||
# default_disabled_quantizers, which disables `*lm_head*`.
|
||||
- quantizer_name: '*lm_head*weight_quantizer'
|
||||
cfg: {$import: nvfp4_static}
|
||||
@@ -0,0 +1,44 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
# recipe for HuggingFace `qwen3_5` (dense) models. Covers Qwen3.5 and Qwen3.6
|
||||
# dense releases, which share the `qwen3_5` model_type and hybrid
|
||||
# linear-attention + softmax-attention architecture. MSE variant of
|
||||
# `w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`: NVFP4 weight scales come from an MSE
|
||||
# FP8-scale sweep instead of max calibration. Shares its `quant_cfg` with the
|
||||
# MoE counterpart at
|
||||
# `huggingface/qwen3_5_moe/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`;
|
||||
# the snippet lives under
|
||||
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
imports:
|
||||
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
recipe for HuggingFace `qwen3_5` (dense) models: NVFP4 with static scales
|
||||
from an MSE FP8-scale sweep for MLP projection weights and lm_head; FP8 for
|
||||
self-attention and the large linear-attention projections; FP8 KV cache with
|
||||
constant amax.
|
||||
quantize:
|
||||
algorithm:
|
||||
method: mse
|
||||
fp8_scale_sweep: true
|
||||
layerwise: false
|
||||
quant_cfg:
|
||||
- $import: shared_quant_cfg
|
||||
+44
@@ -0,0 +1,44 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
# recipe for HuggingFace `qwen3_5_moe` models. Covers Qwen3.5-MoE and
|
||||
# Qwen3.6-MoE releases, which share the `qwen3_5_moe` model_type and hybrid
|
||||
# linear-attention + softmax-attention MoE architecture. MSE variant of
|
||||
# `w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`: NVFP4 weight scales come from an MSE
|
||||
# FP8-scale sweep instead of max calibration. Shares its `quant_cfg` with the
|
||||
# dense counterpart at
|
||||
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.yaml`; the
|
||||
# snippet lives under
|
||||
# `huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
imports:
|
||||
shared_quant_cfg: huggingface/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
recipe for HuggingFace `qwen3_5_moe` models (Qwen3.5-MoE and Qwen3.6-MoE
|
||||
releases): NVFP4 with static scales from an MSE FP8-scale sweep for MoE /
|
||||
shared-expert MLP projection weights and lm_head; FP8 for self-attention and
|
||||
the large linear-attention projections; FP8 KV cache with constant amax.
|
||||
quantize:
|
||||
algorithm:
|
||||
method: mse
|
||||
fp8_scale_sweep: true
|
||||
layerwise: false
|
||||
quant_cfg:
|
||||
- $import: shared_quant_cfg
|
||||
Reference in New Issue
Block a user