mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Add ModelOpt recipe for DeepSeek-V4-Pro-0813 NVFP4 and --recipe to its PTQ script (#2287)
### What does this PR do?
Type of change: new feature
The quantization config for `nvidia/DeepSeek-V4-Pro-0813-NVFP4` existed
only as Python inside `_build_nvfp4_experts_cfg()`, so the released
checkpoint had **no entry in `modelopt_recipes/`** and could not be
looked up by name the way every other published model can.
`modelopt_recipes/README.md` states the goal directly — a recipe is
*"the single, version-controlled source of truth for how a model is
optimized … expressed as data instead of code"* — and this model was the
exception.
This adds:
-
`modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml`,
composed from the existing `configs/ptq/units/base_disable_all` and
`configs/numerics/nvfp4` units.
- An optional `--recipe` flag on `examples/deepseek/deepseek_v4/ptq.py`.
It follows **`examples/kimi/kimi_k3`**, the closest precedent: a very
large MoE whose source already ships MXFP4 routed experts, converted via
`--cast_mxfp4_to_nvfp4` rather than through `examples/hf_ptq`, and
already wired to `--recipe` with a published YAML.
### Usage
```sh
torchrun --nproc-per-node 8 deepseek_v4/ptq.py \
--model_path <mp8_checkpoint> \
--config <DeepSeek-V4-Pro-0813>/inference/config.json \
--calib_size 512 \
--calib_seq 4096 \
--output_path <amax_dump> \
--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only
```
Omitting `--recipe` keeps the previous behaviour exactly.
### Testing
- `load_recipe` resolves the YAML and yields `num_bits (2, 1)` with
`block_sizes {-1: 16, type: dynamic, scale_bits: (4, 3)}` — identical to
the hardcoded config.
- **Equivalence checked behaviourally**, not by eyeballing dicts: both
configs were resolved against representative quantizer names using
last-match-wins, and agree on all of them.
| quantizer | hardcoded | recipe |
| --- | --- | --- |
| `...ffn.experts.17.w1_weight_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.experts.17.w2_input_quantizer` | enabled, NVFP4 | enabled,
NVFP4 |
| `...ffn.shared_experts.w1_weight_quantizer` | disabled | disabled |
| `...attn.wq_weight_quantizer` | disabled | disabled |
| `mtp.0.ffn.experts.2.w1_weight_quantizer` | disabled | disabled |
| `lm_head_weight_quantizer` | disabled | disabled |
- `mtq.quantize` documents `algorithm` as a string **or** a dict keyed
on `method`, so the recipe's `{'method': 'max'}` needs no translation.
- `pre-commit` clean, including `validate modelopt recipes`.
No GPU run: this changes config plumbing only, and the default path is
byte-identical to before.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `--recipe` is optional and
defaults to `None`; without it `_build_nvfp4_experts_cfg()` is used
exactly as before.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies; `modelopt.recipe` is already a first-party import.
- Did you write any new necessary tests?: N/A — no new logic;
equivalence to the existing config is the property that matters and is
documented above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — recipe library addition; recent recipe/example PRs add no entry.
- Did you get Claude approval on this PR?: ❌ — not yet run.
### Additional Information
The recipe covers the **quant config only**. `--calib_seq` — the setting
that mattered most for this checkpoint, since the 512 default does not
cover long-context activation ranges — is a dataloader argument rather
than part of the `mtq` config, so it stays on the CLI. Worth knowing if
the recipe is ever treated as a complete reproduction of the released
checkpoint: it is not, on its own.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added post-training quantization support for DeepSeek-V4-Pro-0813
routed experts using NVFP4.
* Added an optional recipe path for selecting equivalent quantization
settings.
* Preserved source formats for shared experts, attention, embeddings,
output layers, and MTP components.
* **Bug Fixes**
* Improved validation for missing or malformed quantization
configurations.
* Added safeguards against unsupported formats, scopes, algorithms, and
enabled MTP quantizers.
* **Documentation**
* Documented checkpoint conversion behavior, calibration requirements,
and supported quantization workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
ab0eba9ac1
commit
8810eb5e31
@@ -128,6 +128,10 @@ python ${DS_V4}/inference/convert.py \
|
||||
|
||||
### Calibrate routed experts
|
||||
|
||||
The quantization config defaults to the built-in routed-expert NVFP4 setup. Pass
|
||||
`--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only`
|
||||
to load the same config from [modelopt_recipes](../../modelopt_recipes/ptq.md) instead.
|
||||
|
||||
Single node:
|
||||
|
||||
```bash
|
||||
|
||||
@@ -78,6 +78,7 @@ from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import copy
|
||||
import fnmatch
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
@@ -93,6 +94,7 @@ from safetensors.torch import load_model
|
||||
from transformers import AutoTokenizer
|
||||
|
||||
import modelopt.torch.quantization as mtq
|
||||
from modelopt.recipe import load_recipe
|
||||
from modelopt.torch.quantization.nn import TensorQuantizer
|
||||
from modelopt.torch.quantization.qtensor.mxfp4_tensor import MXFP4QTensor
|
||||
from modelopt.torch.utils.dataset_utils import get_dataset_dataloader
|
||||
@@ -273,6 +275,96 @@ def load_deepseek_v4(
|
||||
return model
|
||||
|
||||
|
||||
_PUBLISHED_RECIPE = "huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only"
|
||||
|
||||
|
||||
_MTP_PROBE = "mtp.0.ffn.experts.0.w1_weight_quantizer"
|
||||
|
||||
|
||||
def _effective_enable(quant_cfg, name: str) -> bool:
|
||||
"""Whether ``name`` ends up enabled; mtq applies rules in order, so later ones win."""
|
||||
enabled = False
|
||||
for entry in quant_cfg:
|
||||
if isinstance(entry, dict) and "quantizer_name" in entry:
|
||||
if fnmatch.fnmatch(name, entry["quantizer_name"]):
|
||||
enabled = entry.get("enable", True) is True
|
||||
return enabled
|
||||
|
||||
|
||||
def _quant_cfg_from_recipe(recipe_path: str) -> dict:
|
||||
"""Load a ModelOpt recipe and return an ``mtq.quantize`` config.
|
||||
|
||||
``mtq.quantize`` accepts ``algorithm`` as a string or as a dict keyed on
|
||||
``method``, which is the shape a recipe yields, so no translation is needed.
|
||||
|
||||
The recipe is validated against what the rest of this pipeline can actually
|
||||
export. ``save_amax_and_quant_config`` persists only routed-expert quantizer
|
||||
state and writes ``quantized_layers_manifest.json`` hardcoded to
|
||||
``NVFP4_W4A4``/``num_bits [2, 1]``/``block_size 16``/``scale_bits [4, 3]``,
|
||||
and ``quantize_to_nvfp4.py`` re-derives expert paths itself and emits E4M3
|
||||
block scales unconditionally. A recipe that enables anything else, or asks
|
||||
for a different numeric format, would silently drop quantizers and describe
|
||||
the result with a manifest that does not match the amax it ships beside.
|
||||
Mirrors the guard in ``examples/kimi/kimi_k3/quantize_to_nvfp4.py``.
|
||||
"""
|
||||
recipe = load_recipe(recipe_path)
|
||||
quantize = getattr(recipe, "quantize", None)
|
||||
if quantize is None:
|
||||
raise ValueError(f"recipe {recipe_path!r} has no 'quantize' section; expected a PTQ recipe")
|
||||
cfg = quantize.model_dump()
|
||||
|
||||
algo = cfg.get("algorithm")
|
||||
method = algo.get("method") if isinstance(algo, dict) else algo
|
||||
if method != "max":
|
||||
raise ValueError(f"recipe must use the 'max' calibration algorithm, got {method!r}")
|
||||
|
||||
enabled_experts = False
|
||||
for entry in cfg.get("quant_cfg", []):
|
||||
if not isinstance(entry, dict) or entry.get("enable") is not True:
|
||||
continue
|
||||
name = entry.get("quantizer_name", "")
|
||||
if "ffn.experts." not in name:
|
||||
raise ValueError(
|
||||
f"recipe enables {name!r}, but this pipeline exports only routed-expert "
|
||||
"quantizers (*ffn.experts.*); it would be dropped from the amax dump "
|
||||
"and the manifest without any error"
|
||||
)
|
||||
qcfg = entry.get("cfg") or {}
|
||||
if not isinstance(qcfg, dict):
|
||||
raise ValueError(
|
||||
f"recipe entry {name!r} has a non-dict 'cfg' ({type(qcfg).__name__}); "
|
||||
"this pipeline supports only a single NVFP4 config per quantizer"
|
||||
)
|
||||
block_sizes = qcfg.get("block_sizes") or {}
|
||||
if (
|
||||
tuple(qcfg.get("num_bits") or ()) != (2, 1)
|
||||
or block_sizes.get(-1) != 16
|
||||
or block_sizes.get("type") != "dynamic"
|
||||
or tuple(block_sizes.get("scale_bits") or ()) != (4, 3)
|
||||
):
|
||||
raise ValueError(
|
||||
f"recipe entry {name!r} must be block-16 NVFP4 (num_bits (2, 1), "
|
||||
f"block_sizes[-1] 16, type 'dynamic', scale_bits (4, 3)) to match "
|
||||
f"the exported manifest, got num_bits={qcfg.get('num_bits')!r} "
|
||||
f"block_sizes={block_sizes!r}"
|
||||
)
|
||||
enabled_experts = True
|
||||
|
||||
if not enabled_experts:
|
||||
raise ValueError("recipe enables no routed-expert quantizers (*ffn.experts.*)")
|
||||
|
||||
# The scope check above is a substring test, so a pattern such as
|
||||
# *mtp.*ffn.experts.* satisfies it. MTP experts match save_amax's
|
||||
# \.experts\.\d+\.w[123]_ regex too, so resolve the rules the way mtq applies
|
||||
# them (in order, last match wins) and reject a recipe that leaves them on.
|
||||
if _effective_enable(cfg.get("quant_cfg", []), _MTP_PROBE):
|
||||
raise ValueError(
|
||||
f"recipe enables MTP quantizers (matched {_MTP_PROBE!r}); this pipeline "
|
||||
"leaves the MTP/DSpark block in its source format"
|
||||
)
|
||||
return cfg
|
||||
|
||||
|
||||
def _build_nvfp4_experts_cfg() -> dict:
|
||||
"""Quant config: NVFP4 weight + NVFP4 input, routed experts only.
|
||||
|
||||
@@ -322,6 +414,7 @@ def ptq(
|
||||
calib_size: int,
|
||||
calib_datasets: list[str],
|
||||
calib_seq: int = 512,
|
||||
recipe: str | None = None,
|
||||
):
|
||||
world_size = int(os.getenv("WORLD_SIZE", "1"))
|
||||
rank = int(os.getenv("RANK", "0"))
|
||||
@@ -357,7 +450,8 @@ def ptq(
|
||||
dist.barrier()
|
||||
_trace("pre-calib barrier done")
|
||||
|
||||
mtq_cfg = _build_nvfp4_experts_cfg()
|
||||
mtq_cfg = _quant_cfg_from_recipe(recipe) if recipe else _build_nvfp4_experts_cfg()
|
||||
_trace(f"quant config from {'recipe' if recipe else 'built-in default'}")
|
||||
_trace("calling mtq.quantize")
|
||||
model = mtq.quantize(model, mtq_cfg, calibrate_loop)
|
||||
_trace("mtq.quantize returned")
|
||||
@@ -375,9 +469,9 @@ def save_amax_and_quant_config(model, output_path: str):
|
||||
*not* rely on ``modelopt.torch.export.quant_utils.get_quant_config``
|
||||
because its introspection doesn't see weights stored on nested
|
||||
submodules — ``QuantExpert``'s ``w{1,2,3}`` are submodules of the
|
||||
container, not direct parameters. The downstream export script
|
||||
(``quantize_to_nvfp4.py``) uses this manifest as ground truth for which
|
||||
tensor paths to replace with NVFP4 packed weight + scales.
|
||||
container, not direct parameters. The manifest is a descriptive record of
|
||||
what was calibrated: ``quantize_to_nvfp4.py`` does not read it, it
|
||||
re-derives the expert tensor paths from the source checkpoint itself.
|
||||
"""
|
||||
|
||||
def _trace(msg):
|
||||
@@ -462,6 +556,15 @@ def main():
|
||||
help="where to dump amax files and quantized_layers_manifest.json",
|
||||
)
|
||||
p.add_argument("--batch_size", type=int, default=4)
|
||||
p.add_argument(
|
||||
"--recipe",
|
||||
default=None,
|
||||
help=(
|
||||
"ModelOpt recipe describing the quant config, loaded by path relative to "
|
||||
f"modelopt_recipes/ (published: {_PUBLISHED_RECIPE}). Defaults to the "
|
||||
"built-in routed-experts-only NVFP4 config, which this recipe mirrors."
|
||||
),
|
||||
)
|
||||
p.add_argument("--calib_size", type=int, default=64)
|
||||
p.add_argument(
|
||||
"--calib_seq",
|
||||
@@ -509,7 +612,13 @@ def main():
|
||||
args.model_path, trust_remote_code=args.trust_remote_code
|
||||
)
|
||||
model = ptq(
|
||||
model, tokenizer, args.batch_size, args.calib_size, args.calib_datasets, args.calib_seq
|
||||
model,
|
||||
tokenizer,
|
||||
args.batch_size,
|
||||
args.calib_size,
|
||||
args.calib_datasets,
|
||||
args.calib_seq,
|
||||
args.recipe,
|
||||
)
|
||||
save_amax_and_quant_config(model, args.output_path)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user