Files
Model-Optimizer/tests/unit/recipe/test_recipe_docs.py
T
Shengliang XuandClaude Opus 5 d0142c9dca Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?

**Type of change:** New feature (recipe loading) + one bug fix

Two things, the second built on the first:

1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.

#### Declaring what kind of recipe a file is

`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.

It is now optional, and the loader takes the first of these that
answers:

1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.

Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.

Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.

A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)

#### Checkpoint aliases

Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:

-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.

(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)

Editing the base recipe changes every alias that points at it; nothing
is duplicated.

#### One fix

- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.

### Usage

A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:

```bash
python examples/hf_ptq/hf_ptq.py \
    --pyt_ckpt_path <checkpoint> \
    --recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
    --export_path <output>
```

A recipe that reuses another whole — the shape the aliases use:

```yaml
imports:
  base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast

$import: base
metadata:
  description: What this checkpoint uses the base recipe for.
```

### Testing

- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.

Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
  * Added checkpoint-specific recipes and MLflow experiment references.

* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
  * Nemotron-3 recipes keep MTP blocks in BF16.

* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 17:12:31 -07:00

258 lines
12 KiB
Python

# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Consistency checks between the shipped recipe YAML files and modelopt_recipes/ptq.md.
These tests force recipe additions, removals, and renames to be reflected in the
PTQ recipe guide so the doc never drifts from the files on disk.
"""
import re
from importlib.resources import files
from pathlib import Path
import pytest
RECIPES_DIR = Path(str(files("modelopt_recipes")))
GENERAL_PTQ_DIR = RECIPES_DIR / "general" / "ptq"
PTQ_MD = RECIPES_DIR / "ptq.md"
def _ptq_md_text() -> str:
return PTQ_MD.read_text(encoding="utf-8")
def _general_ptq_stems() -> list[str]:
return sorted(p.stem for p in GENERAL_PTQ_DIR.glob("*.yaml"))
def test_every_general_ptq_recipe_is_documented():
"""Every general/ptq/*.yaml recipe must be mentioned (backticked) in ptq.md."""
doc = _ptq_md_text()
missing = [stem for stem in _general_ptq_stems() if f"`{stem}`" not in doc]
assert not missing, (
f"Recipes under modelopt_recipes/general/ptq/ are missing from "
f"modelopt_recipes/ptq.md: {missing}. When adding a recipe, add a row to "
"the 'shipped recipes' table in ptq.md (and describe any new scheme, "
"KV mode, or calibration variant in the matching section)."
)
def test_documented_general_ptq_recipes_exist_on_disk():
"""Every recipe row in the ptq.md shipped-recipes table must exist on disk.
Catches renames/removals that leave stale rows behind. Rows are identified
by a first cell that is a single backticked token, which only occurs in the
shipped-recipes table.
"""
doc = _ptq_md_text()
documented = re.findall(r"^\| `([^`]+)` \|", doc, flags=re.MULTILINE)
assert documented, "ptq.md shipped-recipes table not found — was it reformatted?"
stale = [name for name in documented if not (GENERAL_PTQ_DIR / f"{name}.yaml").is_file()]
assert not stale, (
f"modelopt_recipes/ptq.md documents general/ptq recipes that do not exist "
f"on disk: {stale}. Update the 'shipped recipes' table after renaming or "
"removing a recipe."
)
def test_general_ptq_recipe_count_in_ptq_md():
"""The 'All N general/ptq/ recipes' summary line must match the file count."""
doc = _ptq_md_text()
match = re.search(r"All (\d+) <code>general/ptq/</code> recipes", doc)
assert match, (
"Could not find the 'All N <code>general/ptq/</code> recipes' summary "
"line in modelopt_recipes/ptq.md — keep that phrasing so this check can "
"verify the recipe count."
)
documented_count = int(match.group(1))
actual_count = len(_general_ptq_stems())
assert documented_count == actual_count, (
f"modelopt_recipes/ptq.md says 'All {documented_count} general/ptq/ "
f"recipes' but modelopt_recipes/general/ptq/ contains {actual_count} "
"recipes. Update the count and the table in ptq.md."
)
def test_documented_recipe_paths_resolve():
"""Every ``general/ptq/<name>`` a doc names must exist on disk.
The existing checks run doc-from-disk: they catch a recipe that no doc mentions.
This is the other direction -- a doc naming a recipe that was never landed, or was
moved to another branch after the doc row was written. A recipe path is the
user-facing ``--recipe`` interface, so a phantom row sends users to
``Recipe path '...' is not a valid YAML file or directory``.
"""
docs = {
"modelopt_recipes/ptq.md": PTQ_MD,
"docs/source/guides/10_recipes.rst": Path(__file__).resolve().parents[3]
/ "docs"
/ "source"
/ "guides"
/ "10_recipes.rst",
}
missing = []
for label, path in docs.items():
if not path.is_file():
continue
# encoding= is required, not decorative: these docs contain non-ASCII (em dashes
# among others) and a bare read_text() decodes with the locale codepage, which is
# cp1252 on the Windows runners -- UnicodeDecodeError on the first such byte.
for name in sorted(
set(re.findall(r"general/ptq/([A-Za-z0-9._-]+)", path.read_text(encoding="utf-8")))
):
stem = name.removesuffix(".yaml").removesuffix(".yml")
if (GENERAL_PTQ_DIR / f"{stem}.yaml").is_file():
continue
# Prose also names a recipe *family* -- e.g. ``general/ptq/nvfp4_mlp_only``
# standing for its -kv_* variants -- which is not a phantom path.
if any(GENERAL_PTQ_DIR.glob(f"{stem}-*.yaml")):
continue
missing.append(f"{label} -> general/ptq/{name}")
assert not missing, (
"Docs name general/ptq recipes that do not exist on disk:\n "
+ "\n ".join(missing)
+ "\nAdd the recipe, or remove the row if it belongs to a different change."
)
def test_every_model_specific_ptq_dir_is_mentioned():
"""Every model-specific PTQ recipe must be identifiable in ptq.md.
``model_type/<model_type>/ptq/`` recipes are checked by their ``model_type``
(e.g. ``gemma4``); ``models/<org>/<model_id>/ptq/`` recipes are checked by their
full ``<org>/<model_id>`` hub path (e.g. ``nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16``), so
the org — the whole point of the top-level tier — is verified too and an org
re-key (e.g. ``step3p5`` → ``stepfun-ai``) can't silently drift from the doc.
"""
doc = _ptq_md_text()
# model_type recipes: model_type/<model_type>/ptq/<recipe>.yaml -> <model_type>
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "model_type").glob("**/ptq/*.yaml")}
# checkpoint recipes: models/<org>/<model_id>/ptq/<recipe>.yaml -> <org>/<model_id>
model_ids = {
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
for p in (RECIPES_DIR / "models").glob("**/ptq/*.yaml")
}
identifiers = sorted(hf_ids | model_ids)
assert identifiers, "No model-specific PTQ recipes found under model_type/ or models/"
missing = [name for name in identifiers if name not in doc]
assert not missing, (
f"Model-specific PTQ recipe folders are missing from "
f"modelopt_recipes/ptq.md: {missing}. Add them to the model-specific "
"recipes section (kinds table and/or the matching subsection)."
)
def test_checkpoint_recipes_live_in_the_top_level_models_tier():
"""Lock in the model_type-vs-checkpoint split and the backward-compat symlinks.
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``model_type/``
(formerly ``huggingface/``) holds only per-``model_type`` recipes. Two
backward-compatibility **symlinks** are kept so old ``--recipe`` paths still
resolve: the top-level ``huggingface`` -> ``model_type`` rename alias, and the
nested ``model_type/models`` -> ``../models`` alias for the old
``huggingface/models/<org>/<model_id>/...`` checkpoint paths. Both must stay
symlinks and never become real directories that hold recipes. A checkpoint recipe
nested under a ``model_type`` (e.g. ``model_type/<model_type>/<checkpoint>/<task>/``)
still fails loudly here instead of silently shipping both tiers — e.g. on a bad
merge that re-adds the old layout.
"""
model_type = RECIPES_DIR / "model_type"
models = RECIPES_DIR / "models"
hf_alias = RECIPES_DIR / "huggingface"
mt_models = model_type / "models"
# Top-level huggingface -> model_type rename alias.
assert hf_alias.is_symlink(), (
"modelopt_recipes/huggingface must be a backward-compat symlink to model_type/ "
"(the rename alias), not a real directory."
)
assert hf_alias.resolve() == model_type.resolve(), (
f"huggingface must resolve to the model_type/ tier; resolves to "
f"{hf_alias.resolve()} instead of {model_type.resolve()}."
)
# Nested model_type/models -> ../models alias for the old huggingface/models/... paths.
assert mt_models.is_symlink(), (
"model_type/models must be a symlink to the top-level modelopt_recipes/models/ "
"tier (a backward-compat alias for the old --recipe huggingface/models/... paths), "
"not a real directory."
)
assert mt_models.resolve() == models.resolve(), (
f"model_type/models must resolve to the top-level models/ tier; resolves to "
f"{mt_models.resolve()} instead of {models.resolve()}."
)
# Every recipe under model_type/ must be <model_type>/<task>/<file> (3 parts);
# anything deeper is a checkpoint nested under a model_type and belongs in models/.
# Skip the model_type/models symlink so the models/ recipes it aliases (4 parts)
# aren't miscounted as nested here.
nested = sorted(
str(p.relative_to(RECIPES_DIR))
for ext in ("*.yaml", "*.yml")
for p in model_type.glob(f"**/{ext}")
if mt_models not in p.parents and len(p.relative_to(model_type).parts) != 3
)
assert not nested, (
f"Recipes under model_type/ must be <model_type>/<task>/<file>; found nested "
f"paths (a checkpoint recipe belongs under models/<org>/<model_id>/): {nested}"
)
# Every recipe under models/ must be <org>/<model_id>/<task>/<file> (4 parts) so the
# path is exactly the model-hub path; a different depth breaks that convention.
misplaced = sorted(
str(p.relative_to(RECIPES_DIR))
for ext in ("*.yaml", "*.yml")
for p in models.glob(f"**/{ext}")
if len(p.relative_to(models).parts) != 4
)
assert not misplaced, (
f"Recipes under models/ must be <org>/<model_id>/<task>/<file>; found: {misplaced}"
)
def test_launcher_yaml_recipe_paths_resolve():
"""Every modelopt_recipes recipe path a launcher example selects must resolve on disk.
Guards against a recipe rename — e.g. keying ``models/nvidia/<id>`` by the canonical Hub id,
which carries the ``NVIDIA-`` prefix — drifting from the launcher YAML that loads it. The
depth/doc tests can't catch a launcher pointing at a recipe path that no longer exists.
"""
repo_root = Path(__file__).resolve().parents[3]
launcher_dir = repo_root / "tools" / "launcher" / "examples"
if not launcher_dir.is_dir():
pytest.skip("tools/launcher/examples not available in this checkout")
def _resolves(rel: str) -> bool:
return any((RECIPES_DIR / f"{rel}{suffix}").exists() for suffix in ("", ".yaml", ".yml"))
# ``--recipe <p>`` / ``QUANT_CFG: <p>`` are modelopt_recipes-relative — only tier-prefixed
# values are recipe paths; bare names like ``auto`` or ``FP8_DEFAULT_CFG`` are not. The
# ``modelopt_recipes/<p>.yaml`` form (e.g. ``--config``) embeds the path directly.
tier = r"(?:general|model_type|models|configs)/[A-Za-z0-9._/-]+"
rel_re = re.compile(rf"(?:--recipe\s+|QUANT_CFG:\s*)({tier})")
abs_re = re.compile(rf"modelopt_recipes/({tier}\.ya?ml)")
missing = []
for yaml_path in sorted(launcher_dir.rglob("*.yaml")):
text = yaml_path.read_text(encoding="utf-8")
candidates = set(rel_re.findall(text)) | {
re.sub(r"\.ya?ml$", "", m) for m in abs_re.findall(text)
}
missing.extend(
f"{yaml_path.relative_to(repo_root)} -> {rel}"
for rel in sorted(candidates)
if not _resolves(rel)
)
assert not missing, "Launcher YAMLs reference recipe paths that do not resolve:\n" + "\n".join(
missing
)