mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
323 lines
13 KiB
Python
323 lines
13 KiB
Python
# SPDX-FileCopyrightText: Copyright (c) 2024 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
|
# SPDX-License-Identifier: Apache-2.0
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
"""Pre-commit hook: validate ModelOpt recipes.
|
|
|
|
Pre-commit passes changed file paths as arguments. This script resolves each
|
|
file to its parent recipe (single-file or directory format), deduplicates, and
|
|
validates each recipe exactly once.
|
|
|
|
Checks performed:
|
|
|
|
1. ``quant_cfg`` must use the list-of-dicts format with explicit
|
|
``quantizer_name`` keys (legacy dict format is rejected). PTQ recipes only.
|
|
2. PTQ recipes must use ``quantize`` as the top-level key
|
|
(not ``ptq_cfg`` or other variants).
|
|
3. Each recipe (PTQ, EAGLE, DFlash, Medusa) is loaded via ``load_recipe()``
|
|
to catch structural and Pydantic-validation errors (skipped if modelopt is
|
|
not installed).
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import re
|
|
import sys
|
|
from pathlib import Path
|
|
|
|
import yaml
|
|
|
|
_YAML_PARSE_ERROR = object()
|
|
|
|
# Recipe types reached through the LEGACY metadata.recipe_type path only. A recipe that
|
|
# declares a ``# modelopt-schema:`` comment, or that delegates with ``$import``, is
|
|
# validated whether or not its kind appears here -- _is_recipe_file returns True on those
|
|
# branches before this set is consulted. So a new kind that declares a schema (the
|
|
# recommended form) needs no change here; only a new kind still using the deprecated
|
|
# metadata.recipe_type would. Mirrors RecipeType in modelopt.recipe.config; kept as a
|
|
# literal set so the hook can run without importing modelopt (which is also why
|
|
# _try_load_recipe gates on ImportError).
|
|
_SUPPORTED_RECIPE_TYPES = frozenset(
|
|
{"ptq", "speculative_eagle", "speculative_dflash", "speculative_medusa"}
|
|
)
|
|
|
|
# A recipe usually declares its kind with a ``# modelopt-schema:`` comment naming its
|
|
# schema class rather than with ``metadata.recipe_type`` (see modelopt/recipe/loader.py).
|
|
# Matched here by name so the hook keeps working without importing modelopt.
|
|
_SCHEMA_COMMENT_RE = re.compile(
|
|
r"^\s*#\s*modelopt-schema:\s*modelopt\.recipe\.config\.ModelOpt\w+Recipe\s*$",
|
|
re.MULTILINE,
|
|
)
|
|
|
|
# Any ``# modelopt-schema:`` declaration, recipe or not. The reusable snippets under
|
|
# ``modelopt_recipes/configs/`` declare non-recipe schemas -- QuantizerAttributeConfig,
|
|
# LayerPatternList and friends -- and a snippet is allowed a top-level ``$import`` of its
|
|
# own, which would otherwise make it indistinguishable here from a delegating alias.
|
|
_ANY_SCHEMA_COMMENT_RE = re.compile(
|
|
r"^\s*#\s*modelopt-schema:\s*\S+\s*$",
|
|
re.MULTILINE,
|
|
)
|
|
|
|
|
|
def _declares_recipe_schema(path: Path) -> bool:
|
|
"""Whether *path* names one of the recipe schema classes in a ``# modelopt-schema:`` comment.
|
|
|
|
Searched over the whole file, deliberately laxer than the loader's
|
|
``_parse_modelopt_schema``, which stops at the first non-comment line. A file carrying
|
|
the comment *below* its YAML body is therefore a recipe to this hook and not to the
|
|
loader -- which is the outcome we want: the hook hands it to ``load_recipe``, which
|
|
rejects it with "does not say what kind of recipe it is" rather than the file being
|
|
skipped silently. ``test_shipped_modelopt_schema_comments_are_in_the_preamble`` keeps
|
|
the shipped tree free of that shape.
|
|
"""
|
|
try:
|
|
return bool(_SCHEMA_COMMENT_RE.search(path.read_text(encoding="utf-8")))
|
|
except OSError:
|
|
return False
|
|
|
|
|
|
def _declares_non_recipe_schema(path: Path) -> bool:
|
|
"""Whether *path* declares a ``# modelopt-schema:`` that is not a recipe schema.
|
|
|
|
That is the signature of a reusable snippet (a quantizer attribute, a layer-pattern
|
|
list), which ``load_recipe`` cannot load and should never be handed.
|
|
"""
|
|
try:
|
|
text = path.read_text(encoding="utf-8")
|
|
except OSError:
|
|
return False
|
|
return bool(_ANY_SCHEMA_COMMENT_RE.search(text)) and not bool(_SCHEMA_COMMENT_RE.search(text))
|
|
|
|
|
|
def _check_quant_cfg(quant_cfg, label: str) -> list[str]:
|
|
"""Validate quant_cfg format. *label* is used in error messages."""
|
|
errors: list[str] = []
|
|
if isinstance(quant_cfg, dict):
|
|
errors.append(
|
|
f"{label}: quant_cfg uses the legacy dict format. "
|
|
"Use the list-of-dicts format with explicit 'quantizer_name' keys instead. "
|
|
"See https://nvidia.github.io/Model-Optimizer/guides/_quant_cfg.html for the format specification."
|
|
)
|
|
elif isinstance(quant_cfg, list):
|
|
for i, entry in enumerate(quant_cfg):
|
|
if not isinstance(entry, dict):
|
|
errors.append(
|
|
f"{label}: quant_cfg[{i}] must be a dict with "
|
|
f"'quantizer_name' or '$import', got {type(entry).__name__}. "
|
|
"See https://nvidia.github.io/Model-Optimizer/guides/_quant_cfg.html"
|
|
)
|
|
continue
|
|
# {$import: name} entries are resolved at load time
|
|
if "$import" in entry:
|
|
ref = entry["$import"]
|
|
if not isinstance(ref, (str, list)) or (
|
|
isinstance(ref, list) and not all(isinstance(r, str) for r in ref)
|
|
):
|
|
errors.append(
|
|
f"{label}: quant_cfg[{i}] '$import' must be a string or list of strings, "
|
|
f"got {type(ref).__name__}: {ref!r}"
|
|
)
|
|
continue
|
|
if "quantizer_name" not in entry:
|
|
errors.append(
|
|
f"{label}: quant_cfg[{i}] is missing 'quantizer_name'. "
|
|
"Each entry must have an explicit 'quantizer_name' or '$import' key. "
|
|
"See https://nvidia.github.io/Model-Optimizer/guides/_quant_cfg.html"
|
|
)
|
|
return errors
|
|
|
|
|
|
def _load_yaml(path: Path):
|
|
"""Load the first YAML document, returning _YAML_PARSE_ERROR on parse failure."""
|
|
try:
|
|
docs = list(yaml.safe_load_all(path.read_text(encoding="utf-8")))
|
|
except Exception:
|
|
return _YAML_PARSE_ERROR
|
|
return docs[0] if docs else None
|
|
|
|
|
|
def _check_single_file_recipe(path: Path) -> list[str]:
|
|
"""Check a single-file recipe (metadata + quantize in one file)."""
|
|
errors: list[str] = []
|
|
label = str(path)
|
|
data = _load_yaml(path)
|
|
if data is _YAML_PARSE_ERROR:
|
|
return [f"{label}: failed to parse YAML"]
|
|
if not isinstance(data, dict):
|
|
return [] # not a recipe file
|
|
|
|
metadata = data.get("metadata")
|
|
if not isinstance(metadata, dict) and not _declares_recipe_schema(path):
|
|
return [] # not a recipe file
|
|
|
|
if "ptq_cfg" in data:
|
|
errors.append(
|
|
f"{label}: uses 'ptq_cfg' as the top-level key. "
|
|
"PTQ recipes must use 'quantize' instead."
|
|
)
|
|
if "quantize" in data:
|
|
quant_section = data["quantize"]
|
|
elif "ptq_cfg" in data:
|
|
quant_section = data["ptq_cfg"]
|
|
else:
|
|
return errors
|
|
|
|
if isinstance(quant_section, dict):
|
|
quant_cfg = quant_section.get("quant_cfg")
|
|
if quant_cfg is not None:
|
|
errors.extend(_check_quant_cfg(quant_cfg, label))
|
|
|
|
return errors
|
|
|
|
|
|
def _check_dir_recipe(dir_path: Path) -> list[str]:
|
|
"""Check a directory-format recipe (metadata.yml + quantize.yml)."""
|
|
errors: list[str] = []
|
|
|
|
for name in ("quantize.yml", "quantize.yaml"):
|
|
quantize_file = dir_path / name
|
|
if quantize_file.is_file():
|
|
data = _load_yaml(quantize_file)
|
|
if data is _YAML_PARSE_ERROR:
|
|
errors.append(f"{quantize_file}: failed to parse YAML")
|
|
elif isinstance(data, dict):
|
|
quant_cfg = data.get("quant_cfg")
|
|
if quant_cfg is not None:
|
|
errors.extend(_check_quant_cfg(quant_cfg, str(quantize_file)))
|
|
break
|
|
|
|
return errors
|
|
|
|
|
|
def _try_load_recipe(path: str) -> list[str]:
|
|
"""Try loading a recipe via modelopt; return errors or []."""
|
|
try:
|
|
from modelopt.recipe.loader import load_recipe
|
|
except ImportError:
|
|
return [] # modelopt not installed, skip
|
|
|
|
try:
|
|
load_recipe(path)
|
|
except Exception as exc:
|
|
return [f"{path}: recipe failed to load: {exc}"]
|
|
return []
|
|
|
|
|
|
def _is_dir_recipe(dir_path: Path) -> bool:
|
|
"""Return True if *dir_path* is a directory-format recipe."""
|
|
return any((dir_path / n).is_file() for n in ("metadata.yml", "metadata.yaml"))
|
|
|
|
|
|
def _is_recipe_file(path: Path) -> bool:
|
|
"""Return True if *path* looks like a recipe file that should be validated.
|
|
|
|
Three ways in, checked in this order: a ``# modelopt-schema:`` comment, a top-level
|
|
``$import`` (a delegating alias, whose kind comes from what it imports), and finally
|
|
the deprecated ``metadata.recipe_type`` gated on ``_SUPPORTED_RECIPE_TYPES``. Only
|
|
that last branch consults the set, so a recipe of any kind that declares a schema is
|
|
validated here -- including kinds deliberately absent from the set, such as
|
|
``auto_quantize``. ``load_recipe`` handles those, so this is intended.
|
|
|
|
Malformed or unparseable files return True so that ``load_recipe()`` can
|
|
report the actual error.
|
|
"""
|
|
data = _load_yaml(path)
|
|
if data is _YAML_PARSE_ERROR:
|
|
return True # let load_recipe report the parse error
|
|
if not isinstance(data, dict):
|
|
return False # not a recipe file at all
|
|
if _declares_recipe_schema(path):
|
|
return True
|
|
if _declares_non_recipe_schema(path):
|
|
# A snippet, not a recipe -- and snippets may carry a top-level ``$import`` of
|
|
# their own (see ``test_import_cross_file_same_name_no_conflict``). Without this
|
|
# the next branch would claim it and ``load_recipe`` would reject it with "does
|
|
# not say what kind of recipe it is", which is a confusing way to learn that a
|
|
# fragment was never meant to be loaded as a recipe.
|
|
return False
|
|
if "$import" in data:
|
|
# A delegating alias declares neither a schema comment nor a recipe_type: its
|
|
# kind comes from the recipe it imports. Validate it so a typo in ``imports:``
|
|
# or a ``$import`` naming an undeclared import fails here rather than at use.
|
|
return True
|
|
metadata = data.get("metadata")
|
|
if not isinstance(metadata, dict) or "recipe_type" not in metadata:
|
|
return False # not a recipe file at all
|
|
return metadata["recipe_type"] in _SUPPORTED_RECIPE_TYPES
|
|
|
|
|
|
def _is_metadata_file(path: Path) -> bool:
|
|
"""Return True if *path* looks like a directory recipe metadata file.
|
|
|
|
Directory-format recipes are PTQ-only (speculative-decoding recipes are
|
|
always single YAML files), so the check is limited to ``recipe_type: ptq``.
|
|
"""
|
|
data = _load_yaml(path)
|
|
if data is _YAML_PARSE_ERROR:
|
|
return True # let load_recipe report the parse error
|
|
if not isinstance(data, dict):
|
|
return False
|
|
return data.get("recipe_type") == "ptq"
|
|
|
|
|
|
def _resolve_recipes(changed_files: list[str]) -> dict[Path, str]:
|
|
"""Resolve changed files to recipes. Returns {recipe_path: kind} mapping.
|
|
|
|
Non-recipe YAML files are silently skipped.
|
|
kind is "file" for single-file recipes or "dir" for directory-format recipes.
|
|
"""
|
|
recipes: dict[Path, str] = {}
|
|
for f in changed_files:
|
|
path = Path(f)
|
|
|
|
# Check if this file is inside a directory-format recipe.
|
|
if _is_dir_recipe(path.parent):
|
|
# Directory recipes have a metadata.yml with top-level metadata fields.
|
|
for name in ("metadata.yml", "metadata.yaml"):
|
|
candidate = path.parent / name
|
|
if candidate.is_file() and _is_metadata_file(candidate):
|
|
recipes.setdefault(path.parent, "dir")
|
|
break
|
|
elif path.is_file() and path.suffix in (".yml", ".yaml"):
|
|
if _is_recipe_file(path):
|
|
recipes.setdefault(path, "file")
|
|
|
|
return recipes
|
|
|
|
|
|
def main() -> int:
|
|
"""Validate changed recipes passed as CLI args, exit 1 on errors."""
|
|
recipes = _resolve_recipes(sys.argv[1:])
|
|
errors: list[str] = []
|
|
|
|
for recipe_path, kind in recipes.items():
|
|
if kind == "dir":
|
|
recipe_errors = _check_dir_recipe(recipe_path)
|
|
else:
|
|
recipe_errors = _check_single_file_recipe(recipe_path)
|
|
|
|
errors.extend(recipe_errors)
|
|
if not recipe_errors:
|
|
errors.extend(_try_load_recipe(str(recipe_path)))
|
|
|
|
if errors:
|
|
for e in errors:
|
|
print(f"ERROR: {e}", file=sys.stderr)
|
|
return 1
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|