mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
feat(recipes): add kv_fp8_cast variants for partial-NVFP4 and weight-only PTQ recipes (#1652)
### What does this PR do?
Type of change: new feature (recipes)
Several `general/ptq` recipe families shipped a data-driven FP8 KV-cache
(`-kv_fp8`) variant but lacked the constant-amax `kv_fp8_cast` companion
that `fp8_default` and `nvfp4_default` already have. This PR adds the
missing cast variants so every KV-quantizing (and the weight-only)
family offers the calibration-free FP8 KV-cache option:
- `general/ptq/nvfp4_experts_only-kv_fp8_cast`
- `general/ptq/nvfp4_mlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_omlp_only-kv_fp8_cast`
- `general/ptq/nvfp4_weight_only-kv_fp8_cast`
Each new recipe composes the exact same model-quant config as its
existing sibling and swaps the `kv_fp8` unit for the shared
`kv_fp8_cast` unit (constant-amax FP8 KV cache; no KV calibration
forward pass). The docs guide table/tree and the changelog are updated
to match.
### Usage
```bash
python examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path <model> \
--recipe general/ptq/nvfp4_mlp_only-kv_fp8_cast
```
### Testing
Extended the built-in PTQ smoke test
`tests/unit/recipe/test_loader.py::test_load_recipe_all_builtins` with
the four new recipe paths; all four load into a valid
`ModelOptPTQRecipe` with a populated `quantize` section.
```
$ python -m pytest tests/unit/recipe/test_loader.py tests/unit/recipe/test_presets.py -q
180 passed
```
`pre-commit` (including the `validate modelopt recipes` hook) passes on
all changed files.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ (additive — only new recipe
files)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (extended the builtin recipe
smoke test)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ (not yet)
### Additional Information
The two weight-only families were discussed for scope;
`nvfp4_weight_only` is included (it already names a KV mode, `kv_fp16`),
while `int4_blockwise_weight_only` is intentionally left untouched since
it carries no `-kv_` composition.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added four new NVFP4 PTQ (Post-Training Quantization) recipe variants:
experts-only, MLP-only, OMLP-only, and weight-only configurations.
* All new recipes include FP8 KV-cache cast mode support for improved
inference performance.
* **Documentation**
* Updated built-in recipes guide with new NVFP4 recipe options and
repository layout.
* **Tests**
* Expanded recipe loader test coverage for new recipe configurations.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.8
parent
aec72ffa68
commit
56c4af2333
@@ -35,6 +35,7 @@ Changelog
|
||||
- Add ``--cast_mxfp4_to_nvfp4`` flag to ``examples/llm_ptq/hf_ptq.py`` for closed-form, bit-exact MXFP4 → NVFP4 weight conversion. Supports the GPT-OSS family (``openai/gpt-oss-20b``, ``openai/gpt-oss-120b``). See `examples/llm_ptq/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq#mxfp4--nvfp4-cast-for-gpt-oss>`__ for usage.
|
||||
- DeepSeek PTQ (``examples/deepseek/ptq.py``) now defaults to native top-k calibration with post-hoc per-layer peer-max sync of expert ``input_quantizer.amax``; the all-experts path is preserved behind ``--calib_all_experts``.
|
||||
- Add NVFP4 W4A16 weight-only quantization (``w4a16_nvfp4``): FP4 weights with group_size=16, BF16 activations, no calibration forward pass required. Use ``mtq.W4A16_NVFP4_CFG`` or ``--qformat w4a16_nvfp4`` in ``hf_ptq.py``. vLLM deployment support is in progress.
|
||||
- Add FP8 KV-cache cast variants for the partial-NVFP4 and weight-only general PTQ recipes: ``general/ptq/nvfp4_mlp_only-kv_fp8_cast``, ``general/ptq/nvfp4_experts_only-kv_fp8_cast``, ``general/ptq/nvfp4_omlp_only-kv_fp8_cast``, and ``general/ptq/nvfp4_weight_only-kv_fp8_cast``. These compose the same model-quant configs as their ``-kv_fp8`` siblings with the ``kv_fp8_cast`` unit (constant-amax FP8 KV cache, no KV calibration forward pass).
|
||||
- Add Megatron Core export/import mapping for Qwen3-VL (``Qwen3VLForConditionalGeneration``) vision-language models. The mapping handles the ``model.language_model.`` weight prefix used by Qwen3-VL.
|
||||
- Add active-MoE cost accounting for ``mtq.auto_quantize`` effective-bits search. Set ``constraints={"effective_bits": ..., "cost_model": "active_moe", "cost": {"active_moe_expert_ratio": ...}}`` to weight routed MoE expert costs by active experts per token while keeping shared experts fully counted. The ``hf_ptq.py`` AutoQuant path exposes this via ``--auto_quantize_cost_model active_moe`` and ``--auto_quantize_active_moe_expert_ratio``.
|
||||
- Add ``DATASET_COMBOS`` to ``modelopt.torch.utils.dataset_utils`` — single ``--dataset`` tokens that fan out to multiple registered datasets; per-entry ``num_samples`` is split evenly across the members. Initial combos: ``cnn_nemotron_v2_mix`` (``cnn_dailymail`` + ``nemotron-post-training-dataset-v2``, used by ``hf_ptq.py`` when no ``--dataset`` is provided) and ``nemotron-post-training-v3`` (the seven ``nvidia/Nemotron-*`` SFT datasets added in #1498, mirroring the `nemotron-post-training-v3 collection <https://huggingface.co/collections/nvidia/nemotron-post-training-v3>`_). Combo names are listed by ``get_supported_datasets()`` and surfaced in ``--dataset`` help. ``get_dataset_dataloader`` rejects inputs that mix a combo with one of its member datasets (e.g. ``cnn_dailymail,cnn_nemotron_v2_mix``) to avoid double-sampling, and ``get_dataset_samples`` rejects combo names so callers route through the dataloader. ``hf_ptq.py`` default ``--calib_size`` is bumped from ``512`` to ``1024`` so the total calibration sample count under the new default combo matches the previous two-dataset fallback.
|
||||
|
||||
@@ -499,14 +499,22 @@ General PTQ recipes are model-agnostic and apply to any supported architecture:
|
||||
- NVFP4 W4A4, FP8 KV cache with data-driven calibration
|
||||
* - ``general/ptq/nvfp4_default-kv_nvfp4_cast``
|
||||
- NVFP4 W4A4, NVFP4 KV cache with constant amax, max calibration
|
||||
* - ``general/ptq/nvfp4_mlp_only-kv_fp8_cast``
|
||||
- NVFP4 for MLP layers only, FP8 KV cache with constant amax
|
||||
* - ``general/ptq/nvfp4_mlp_only-kv_fp8``
|
||||
- NVFP4 for MLP layers only, FP8 KV cache
|
||||
* - ``general/ptq/nvfp4_experts_only-kv_fp8_cast``
|
||||
- NVFP4 for MoE expert layers only, FP8 KV cache with constant amax
|
||||
* - ``general/ptq/nvfp4_experts_only-kv_fp8``
|
||||
- NVFP4 for MoE expert layers only, FP8 KV cache
|
||||
* - ``general/ptq/nvfp4_experts_only-kv_fp8_layerwise``
|
||||
- NVFP4 for MoE expert layers only, FP8 KV cache, layerwise calibration
|
||||
* - ``general/ptq/nvfp4_omlp_only-kv_fp8_cast``
|
||||
- NVFP4 for output projection + MLP layers, FP8 KV cache with constant amax
|
||||
* - ``general/ptq/nvfp4_omlp_only-kv_fp8``
|
||||
- NVFP4 for output projection + MLP layers, FP8 KV cache
|
||||
* - ``general/ptq/nvfp4_weight_only-kv_fp8_cast``
|
||||
- NVFP4 W4A16 weight-only, FP8 KV cache with constant amax
|
||||
|
||||
Model-specific recipes
|
||||
----------------------
|
||||
@@ -668,10 +676,14 @@ The ``modelopt_recipes/`` package is organized as follows:
|
||||
| +-- nvfp4_default-kv_fp8_cast.yaml
|
||||
| +-- nvfp4_default-kv_fp8.yaml
|
||||
| +-- nvfp4_default-kv_nvfp4_cast.yaml
|
||||
| +-- nvfp4_mlp_only-kv_fp8_cast.yaml
|
||||
| +-- nvfp4_mlp_only-kv_fp8.yaml
|
||||
| +-- nvfp4_experts_only-kv_fp8_cast.yaml
|
||||
| +-- nvfp4_experts_only-kv_fp8.yaml
|
||||
| +-- nvfp4_experts_only-kv_fp8_layerwise.yaml
|
||||
| +-- nvfp4_omlp_only-kv_fp8_cast.yaml
|
||||
| +-- nvfp4_omlp_only-kv_fp8.yaml
|
||||
| +-- nvfp4_weight_only-kv_fp8_cast.yaml
|
||||
+-- huggingface/ # Model-specific recipes
|
||||
| +-- <model_type>/ # see modelopt_recipes/huggingface/README.md
|
||||
| +-- <task>/
|
||||
|
||||
@@ -0,0 +1,51 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Composed PTQ recipe for expert-only dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to expert-layer weight and input quantizers, plus FP8 KV-cache cast
|
||||
mode using constant amax; uses max calibration.
|
||||
quantize:
|
||||
algorithm:
|
||||
method: max
|
||||
# Max calibration is fast and does not typically need checkpointing.
|
||||
# layerwise=false required for VLMs where the decoder layers are nested under
|
||||
# `model.language_model.layers` (layerwise_calibrate can't find them otherwise).
|
||||
layerwise: false
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- quantizer_name: '*mlp.experts*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*mlp.experts*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- $import: kv_fp8_cast
|
||||
- $import: default_disabled_quantizers
|
||||
@@ -0,0 +1,52 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Composed PTQ recipe for MLP/MoE-only dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to MLP/MoE weight and input quantizers, plus FP8 KV-cache cast mode
|
||||
using constant amax; uses max calibration.
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- quantizer_name: '*mlp*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*mlp*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*.experts.*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*.experts.*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- $import: kv_fp8_cast
|
||||
- $import: default_disabled_quantizers
|
||||
@@ -0,0 +1,52 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Composed PTQ recipe for output-projection and MLP/MoE dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 to output-projection and MLP/MoE weight and input quantizers, plus
|
||||
FP8 KV-cache cast mode using constant amax; uses max calibration.
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- quantizer_name: '*mlp*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*mlp*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*block_sparse_moe*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*o_proj*weight_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- quantizer_name: '*o_proj*input_quantizer'
|
||||
cfg:
|
||||
$import: nvfp4
|
||||
- $import: kv_fp8_cast
|
||||
- $import: default_disabled_quantizers
|
||||
@@ -0,0 +1,35 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Composed PTQ recipe for NVFP4 W4A16 weight-only quantization with FP8 KV-cache cast mode.
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
w4a16_nvfp4: configs/ptq/units/w4_nvfp4
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
NVFP4 W4A16 weight-only, BF16 activations, plus FP8 KV-cache cast mode using constant amax; uses
|
||||
max calibration. No calibration forward pass required.
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- $import: base_disable_all
|
||||
- $import: w4a16_nvfp4
|
||||
- $import: kv_fp8_cast
|
||||
- $import: default_disabled_quantizers
|
||||
@@ -161,9 +161,13 @@ _BUILTIN_PTQ_RECIPES = [
|
||||
"general/ptq/nvfp4_default-kv_nvfp4_cast",
|
||||
"general/ptq/nvfp4_default-kv_none-gptq",
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8",
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"general/ptq/nvfp4_experts_only-kv_fp8_layerwise",
|
||||
"general/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
"general/ptq/nvfp4_mlp_only-kv_fp8_cast",
|
||||
"general/ptq/nvfp4_omlp_only-kv_fp8",
|
||||
"general/ptq/nvfp4_omlp_only-kv_fp8_cast",
|
||||
"general/ptq/nvfp4_weight_only-kv_fp8_cast",
|
||||
]
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user