Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)

### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Shengliang Xu
2026-09-21 17:12:31 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent 051d6adb20
commit d0142c9dca
96 changed files with 1360 additions and 129 deletions
+83
View File
@@ -0,0 +1,83 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""CI check: every models/<org>/<model_id>/ptq/ entry names a real Hugging Face Hub repo.
Not a pre-commit hook -- this makes one network call per entry, and pre-commit hooks
(tools/precommit/) have to stay usable offline. Wired into CI instead; see
.github/workflows/check_model_hub_orgs.yml, which only runs it when
modelopt_recipes/models/** changes.
A recipe's own text can't catch a wrong org -- it's typically self-consistent with
the (wrong) directory it lives in, since both were written by the same mistake. Hub
existence is the one fact that isn't derivable from the recipe itself.
Scoped to the ptq/ task, matching tests/unit/recipe/test_recipe_docs.py's
test_every_model_specific_ptq_dir_is_mentioned: a models/<org>/<model_id>/ entry with
only e.g. an auto_quantize/ recipe (Muse-Glimmer-30B) is exempt from the ptq.md
membership check for the same reason it should be exempt here -- models/README.md
allows a checkpoint mirror for a *planned*, not-yet-released checkpoint, and ptq/ is
where the backfilled, already-published aliases this check targets actually live.
"""
from __future__ import annotations
import sys
from pathlib import Path
from huggingface_hub import HfApi
from huggingface_hub.utils import RepositoryNotFoundError
MODELS_ROOT = Path(__file__).resolve().parents[2] / "modelopt_recipes" / "models"
def _checkpoint_ids() -> list[str]:
"""<org>/<model_id> for every models/<org>/<model_id>/ptq/ entry, deduplicated."""
ids = {
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
for p in MODELS_ROOT.glob("*/*/ptq/*.yaml")
}
return sorted(ids)
def _error_if_missing(api: HfApi, repo_id: str) -> str | None:
try:
api.model_info(repo_id)
except RepositoryNotFoundError:
return (
f"{repo_id}: no such Hugging Face Hub repo. modelopt_recipes/models/<org>/"
"<model_id>/ must be keyed by the SOURCE checkpoint's hub path (what you pass "
"to from_pretrained(...)), not a quantized derivative published elsewhere -- "
"see modelopt_recipes/models/README.md."
)
return None
def main() -> int:
"""Verify every checkpoint id against the Hub, exit 1 if any is missing."""
checkpoint_ids = _checkpoint_ids()
api = HfApi()
errors = [e for e in (_error_if_missing(api, r) for r in checkpoint_ids) if e is not None]
if errors:
for e in errors:
print(f"ERROR: {e}", file=sys.stderr)
return 1
print(f"Verified {len(checkpoint_ids)} models/<org>/<model_id>/ptq/ hub paths.")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+77 -6
View File
@@ -32,6 +32,7 @@ Checks performed:
from __future__ import annotations
import re
import sys
from pathlib import Path
@@ -39,13 +40,65 @@ import yaml
_YAML_PARSE_ERROR = object()
# Recipe types this hook validates via load_recipe(). Mirrors RecipeType in
# modelopt.recipe.config; kept as a literal set so the hook can run without
# importing modelopt (which is also why _try_load_recipe gates on ImportError).
# Recipe types reached through the LEGACY metadata.recipe_type path only. A recipe that
# declares a ``# modelopt-schema:`` comment, or that delegates with ``$import``, is
# validated whether or not its kind appears here -- _is_recipe_file returns True on those
# branches before this set is consulted. So a new kind that declares a schema (the
# recommended form) needs no change here; only a new kind still using the deprecated
# metadata.recipe_type would. Mirrors RecipeType in modelopt.recipe.config; kept as a
# literal set so the hook can run without importing modelopt (which is also why
# _try_load_recipe gates on ImportError).
_SUPPORTED_RECIPE_TYPES = frozenset(
{"ptq", "speculative_eagle", "speculative_dflash", "speculative_medusa"}
)
# A recipe usually declares its kind with a ``# modelopt-schema:`` comment naming its
# schema class rather than with ``metadata.recipe_type`` (see modelopt/recipe/loader.py).
# Matched here by name so the hook keeps working without importing modelopt.
_SCHEMA_COMMENT_RE = re.compile(
r"^\s*#\s*modelopt-schema:\s*modelopt\.recipe\.config\.ModelOpt\w+Recipe\s*$",
re.MULTILINE,
)
# Any ``# modelopt-schema:`` declaration, recipe or not. The reusable snippets under
# ``modelopt_recipes/configs/`` declare non-recipe schemas -- QuantizerAttributeConfig,
# LayerPatternList and friends -- and a snippet is allowed a top-level ``$import`` of its
# own, which would otherwise make it indistinguishable here from a delegating alias.
_ANY_SCHEMA_COMMENT_RE = re.compile(
r"^\s*#\s*modelopt-schema:\s*\S+\s*$",
re.MULTILINE,
)
def _declares_recipe_schema(path: Path) -> bool:
"""Whether *path* names one of the recipe schema classes in a ``# modelopt-schema:`` comment.
Searched over the whole file, deliberately laxer than the loader's
``_parse_modelopt_schema``, which stops at the first non-comment line. A file carrying
the comment *below* its YAML body is therefore a recipe to this hook and not to the
loader -- which is the outcome we want: the hook hands it to ``load_recipe``, which
rejects it with "does not say what kind of recipe it is" rather than the file being
skipped silently. ``test_shipped_modelopt_schema_comments_are_in_the_preamble`` keeps
the shipped tree free of that shape.
"""
try:
return bool(_SCHEMA_COMMENT_RE.search(path.read_text(encoding="utf-8")))
except OSError:
return False
def _declares_non_recipe_schema(path: Path) -> bool:
"""Whether *path* declares a ``# modelopt-schema:`` that is not a recipe schema.
That is the signature of a reusable snippet (a quantizer attribute, a layer-pattern
list), which ``load_recipe`` cannot load and should never be handed.
"""
try:
text = path.read_text(encoding="utf-8")
except OSError:
return False
return bool(_ANY_SCHEMA_COMMENT_RE.search(text)) and not bool(_SCHEMA_COMMENT_RE.search(text))
def _check_quant_cfg(quant_cfg, label: str) -> list[str]:
"""Validate quant_cfg format. *label* is used in error messages."""
@@ -105,7 +158,7 @@ def _check_single_file_recipe(path: Path) -> list[str]:
return [] # not a recipe file
metadata = data.get("metadata")
if not isinstance(metadata, dict) or "recipe_type" not in metadata:
if not isinstance(metadata, dict) and not _declares_recipe_schema(path):
return [] # not a recipe file
if "ptq_cfg" in data:
@@ -169,8 +222,12 @@ def _is_dir_recipe(dir_path: Path) -> bool:
def _is_recipe_file(path: Path) -> bool:
"""Return True if *path* looks like a recipe file that should be validated.
Covers PTQ + speculative-decoding (EAGLE/DFlash/Medusa) recipes; extend
``_SUPPORTED_RECIPE_TYPES`` for new types (e.g. QAT).
Three ways in, checked in this order: a ``# modelopt-schema:`` comment, a top-level
``$import`` (a delegating alias, whose kind comes from what it imports), and finally
the deprecated ``metadata.recipe_type`` gated on ``_SUPPORTED_RECIPE_TYPES``. Only
that last branch consults the set, so a recipe of any kind that declares a schema is
validated here -- including kinds deliberately absent from the set, such as
``auto_quantize``. ``load_recipe`` handles those, so this is intended.
Malformed or unparseable files return True so that ``load_recipe()`` can
report the actual error.
@@ -180,6 +237,20 @@ def _is_recipe_file(path: Path) -> bool:
return True # let load_recipe report the parse error
if not isinstance(data, dict):
return False # not a recipe file at all
if _declares_recipe_schema(path):
return True
if _declares_non_recipe_schema(path):
# A snippet, not a recipe -- and snippets may carry a top-level ``$import`` of
# their own (see ``test_import_cross_file_same_name_no_conflict``). Without this
# the next branch would claim it and ``load_recipe`` would reject it with "does
# not say what kind of recipe it is", which is a confusing way to learn that a
# fragment was never meant to be loaded as a recipe.
return False
if "$import" in data:
# A delegating alias declares neither a schema comment nor a recipe_type: its
# kind comes from the recipe it imports. Validate it so a typo in ``imports:``
# or a ``$import`` naming an undeclared import fails here rather than at use.
return True
metadata = data.get("metadata")
if not isinstance(metadata, dict) or "recipe_type" not in metadata:
return False # not a recipe file at all