From de3eda8a11131122f4c352340311983dc2a6eef4 Mon Sep 17 00:00:00 2001 From: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com> Date: Tue, 1 Sep 2026 10:22:27 -0700 Subject: [PATCH] Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ### What does this PR do? **Type of change:** Refactor (recipe-library layout) + documentation — backward-breaking for saved `--recipe` paths. Separate the two kinds of built-in Hugging Face recipes that were previously mixed under `modelopt_recipes/huggingface/`: - **`huggingface//`** — architecture recipes keyed by the transformers `model_type`; one recipe covers every checkpoint of that architecture. **Unchanged.** - **`models///`** — a *new top-level tier* for recipes that mirror one specific published checkpoint, keyed by its **model-hub path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk path equals the hub path. Concretely, the model-instance recipes move out of `huggingface/` to the top level: - `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` → `models/mistralai/…`, `models/nvidia/…` - `huggingface/step3p5/Step3.5-Flash/…` → `models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo id [`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash) — org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`) **Why:** `modelopt_recipes/README.md` already documented a top-level `models/` tier, but the files lived under `huggingface/models/` and instance-specific recipes were awkwardly nested under the per-`model_type` tree. This aligns the filesystem with the documented layout and makes the instance tier hub-addressable — given a checkpoint id you can find (or place) its recipe with no lookup table. `load_recipe` resolves paths directly under `modelopt_recipes/`, so a top-level `models/` sibling of `general/` and `huggingface/` works identically. The move is metadata-only — all recipe YAML content is byte-identical (`R100` renames). Everything else is updating references (nvidia launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`, plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the `10_recipes.rst` guide, which no longer describe instances under `huggingface/`. ### Usage Recipe paths for the moved checkpoint recipes lose the `huggingface/` prefix (and Step 3.5 Flash is keyed by its hub id): ```python from modelopt.recipe import load_recipe # before load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only") # after load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only") ``` The same rename applies to `--recipe …` CLI values and launcher `QUANT_CFG:` entries. Architecture recipes under `huggingface//` are unaffected. ### Testing - **Recipe resolution (torch-free):** parsed every recipe under `models/` and confirmed all `$import` targets resolve against the recipe root — 0 dangling across the tier. - **Docs consistency:** re-ran the `tests/unit/recipe/test_recipe_docs.py` logic; it now globs both `huggingface/` and `models/`, and every model dir (incl. `Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq` recipe is still mentioned in `ptq.md`. - **Reference sweep:** repo-wide grep confirms no remaining references to the old paths outside the intentional historical CHANGELOG entries (released 0.44 / 0.45). - **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit` hooks pass on the changed files. - Note: the full `pytest` suite was not run in my environment (no `torch`), so `test_recipe_docs.py` / `test_loader.py` should be exercised in CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--recipe` / `load_recipe` paths for the checkpoint-mirror tier change (drop the `huggingface/` prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`). Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the only *released* old paths affected shipped in 0.45. A clean break was chosen over a symlink or loader-alias shim. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: ✅ — updated `test_recipe_docs.py` to also glob the top-level `models/` tier so instance recipes stay covered by the doc-consistency check. - Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking Changes** entry. - Did you get Claude approval on this PR?: ❌ ### Additional Information Design note: an earlier iteration nested everything under `huggingface/model_type/` + `huggingface/models/`; the final layout keeps `huggingface/` flat (per-`model_type`) and lifts instances to a top-level `models/` tier, matching what `modelopt_recipes/README.md` already documented. The `Step3p5*` architecture class names (from the model's `trust_remote_code` modeling code) are unrelated to the recipe path and are left unchanged. ## Summary by CodeRabbit * **New Features** * Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5, and NVIDIA Nemotron models. * Added a Nemotron speculative-decoding warm-start recipe. * **Documentation** * Clarified recipe selection and directory organization. * Documented checkpoint naming conventions and updated usage examples. * **Bug Fixes** * Updated launcher configurations and examples to reference the new recipe locations and corrected model names. * **Tests** * Improved automatic recipe discovery and validation of documented recipe paths. --------- Signed-off-by: Shengliang Xu --- CHANGELOG.rst | 3 +- MANIFEST.in | 2 + docs/source/guides/10_recipes.rst | 19 ++- examples/deepseek/README.md | 2 +- examples/deepseek/deepseek_v4/ptq.py | 2 +- examples/kimi/README.md | 4 +- examples/kimi/kimi_k3/quantize_to_nvfp4.py | 4 +- modelopt/recipe/loader.py | 9 ++ modelopt_recipes/README.md | 24 ++-- modelopt_recipes/huggingface/README.md | 50 +++----- modelopt_recipes/huggingface/models | 1 + modelopt_recipes/models/README.md | 96 ++++++++++++++ .../ptq/nvfp4_experts_only.yaml | 0 .../ptq/nvfp4-max-calib.yaml | 0 .../ptq/nvfp4_experts-fp8_pb_attention.yaml | 0 .../ptq/nvfp4_w4a16.yaml | 0 .../ptq/nvfp4-max-calib.yaml | 0 .../ptq/nvfp4-mse.yaml | 0 .../ptq/nvfp4-4o6.yaml | 0 .../ptq/w4a16_nvfp4_4o6.yaml | 0 .../dspark_warmstart.yaml | 0 .../Step-3.5-Flash}/ptq/nvfp4-mlp-only.yaml | 0 modelopt_recipes/ptq.md | 25 ++-- pyproject.toml | 13 ++ tests/unit/recipe/test_kimi_k3_recipe.py | 2 +- tests/unit/recipe/test_loader.py | 93 ++++++++------ tests/unit/recipe/test_recipe_docs.py | 117 ++++++++++++++++-- .../megatron_lm_ptq.yaml | 4 +- .../megatron_lm_ptq.yaml | 4 +- .../hf_streaming_dspark_warmstart.yaml | 4 +- .../mbridge_qad.yaml | 2 +- .../mbridge_quantize.yaml | 2 +- .../megatron_lm_qad.yaml | 2 +- 33 files changed, 357 insertions(+), 127 deletions(-) create mode 100644 MANIFEST.in create mode 120000 modelopt_recipes/huggingface/models create mode 100644 modelopt_recipes/models/README.md rename modelopt_recipes/{huggingface => }/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml (100%) rename modelopt_recipes/{huggingface => }/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml (100%) rename modelopt_recipes/{huggingface => }/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16 => models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16}/ptq/nvfp4_w4a16.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16 => models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16}/ptq/nvfp4-max-calib.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16 => models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16}/ptq/nvfp4-mse.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16 => models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16}/ptq/nvfp4-4o6.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16 => models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16}/ptq/w4a16_nvfp4_4o6.yaml (100%) rename modelopt_recipes/{huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16 => models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16}/speculative_decoding/dspark_warmstart.yaml (100%) rename modelopt_recipes/{huggingface/step3p5/Step3.5-Flash => models/stepfun-ai/Step-3.5-Flash}/ptq/nvfp4-mlp-only.yaml (100%) diff --git a/CHANGELOG.rst b/CHANGELOG.rst index 77a8fffe1..59f01f678 100755 --- a/CHANGELOG.rst +++ b/CHANGELOG.rst @@ -30,7 +30,8 @@ Changelog **Backward Breaking Changes** -- Move the Mistral Medium 3.5 checkpoint-mirror recipe from ``huggingface/models/nvidia/Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib`` to ``huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib``, keying it by the canonical Hugging Face base model. Update any saved ``--recipe`` paths to the new location. +- Move the checkpoint-mirror recipe tier from ``huggingface/models///`` to the top-level ``models///``, keyed by each recipe's canonical Hugging Face Hub id — so the Step 3.5 Flash recipe moves to ``models/stepfun-ai/Step-3.5-Flash/ptq/`` and the NVIDIA Nemotron recipes gain the ``NVIDIA-`` prefix (e.g. ``models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse``). Update any saved ``--recipe`` paths for these checkpoint recipes accordingly; the per-``model_type`` recipes under ``huggingface/`` are unchanged. +- Move the Mistral Medium 3.5 checkpoint-mirror recipe from ``huggingface/models/nvidia/Mistral-Medium-3.5-128B-NVFP4/ptq/nvfp4-max-calib`` to ``models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib``, keying it by the canonical Hugging Face base model. Update any saved ``--recipe`` paths to the new location. - Remove the ``--auto_quantize_bits``, ``--auto_quantize_method``, ``--auto_quantize_score_size``, ``--auto_quantize_cost_model`` and ``--auto_quantize_active_moe_expert_ratio`` flags from ``examples/hf_ptq`` (deprecated in 0.46). Use an AutoQuantize ``--recipe`` from ``modelopt_recipes/general/auto_quantize/`` instead. Those recipes now also splice in the shared base ``cost_excluded_layers`` unit, which the removed CLI applied unconditionally, so a VL model keeps its vision tower and MTP layers out of the effective-bits denominator. On a VL model this changes the per-layer cost weights, so an existing ``--auto_quantize_checkpoint`` from an earlier release is rejected with "Use a different checkpoint path"; delete or repoint it to re-run the search. - Remove the ``examples/llm_ptq`` symlink and the ``examples/vlm_ptq`` forwarder (both deprecated in 0.46). Use ``examples/hf_ptq``, passing ``--vlm`` for vision-language models. - Remove the backward-compat ``--qformat`` / ``--quant_cfg`` short names ``int8_sq``, ``int8_wo``, ``w4a8_awq``, ``nvfp4_awq``, ``nvfp4_mse``, ``nvfp4_local_hessian``, ``fp8_pb_wo`` and ``fp8_pc_pt`` (deprecated in 0.45). Use the preset basename under ``modelopt_recipes/configs/ptq/presets/model/`` instead: ``int8_smoothquant``, ``int8_weight_only``, ``w4a8_awq_beta``, ``nvfp4_awq_lite``, ``nvfp4_w4a4_weight_mse_fp8_sweep``, ``nvfp4_w4a4_weight_local_hessian``, ``fp8_2d_blockwise_weight_only`` and ``fp8_per_channel_per_token``. The ``modelopt.recipe.presets.QFORMAT_ALIASES`` table and the ``aliases`` argument of ``load_quant_cfg_choices()`` are removed along with them. diff --git a/MANIFEST.in b/MANIFEST.in new file mode 100644 index 000000000..658f16f2f --- /dev/null +++ b/MANIFEST.in @@ -0,0 +1,2 @@ +exclude modelopt_recipes/huggingface/models +prune modelopt_recipes/huggingface/models diff --git a/docs/source/guides/10_recipes.rst b/docs/source/guides/10_recipes.rst index 0bd437821..4ee7a4138 100644 --- a/docs/source/guides/10_recipes.rst +++ b/docs/source/guides/10_recipes.rst @@ -519,11 +519,14 @@ General PTQ recipes are model-agnostic and apply to any supported architecture: Model-specific recipes ---------------------- -Model-specific recipes are tuned for a particular Hugging Face ``model_type`` -(or a specific released model) and live under -``huggingface//[/]/``. See +Model-specific recipes come in two tiers: architecture recipes keyed by a +Hugging Face ``model_type`` under ``huggingface///``, and +checkpoint mirrors keyed by a model-hub path under +``models////``. See `modelopt_recipes/huggingface/README.md `_ -for the layout convention and recipe-lookup order. +and +`modelopt_recipes/models/README.md `_ +for the layout conventions and recipe-lookup order. .. list-table:: :header-rows: 1 @@ -531,7 +534,7 @@ for the layout convention and recipe-lookup order. * - Recipe path - Description - * - ``huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only`` + * - ``models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only`` - NVFP4 MLP-only for Step 3.5 Flash MoE model * - ``huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts`` - MXFP8 language-model base with MSE-calibrated NVFP4 routed experts for MiniMax-M3 @@ -686,10 +689,14 @@ The ``modelopt_recipes/`` package is organized as follows: | +-- nvfp4_omlp_only-kv_fp8_cast.yaml | +-- nvfp4_omlp_only-kv_fp8.yaml | +-- nvfp4_weight_only-kv_fp8_cast.yaml - +-- huggingface/ # Model-specific recipes + +-- huggingface/ # Architecture-specific recipes (by model_type) | +-- / # see modelopt_recipes/huggingface/README.md | +-- / | +-- .yaml + +-- models/ # Checkpoint-specific recipes (by model-hub path) + | +-- // # see modelopt_recipes/models/README.md + | +-- / + | +-- .yaml +-- configs/ # Reusable config snippets (imported via $import) +-- numerics/ # Numeric format definitions | +-- fp8.yaml diff --git a/examples/deepseek/README.md b/examples/deepseek/README.md index e71361e30..27a4786b6 100644 --- a/examples/deepseek/README.md +++ b/examples/deepseek/README.md @@ -129,7 +129,7 @@ python ${DS_V4}/inference/convert.py \ ### Calibrate routed experts The quantization config defaults to the built-in routed-expert NVFP4 setup. Pass -`--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only` +`--recipe models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only` to load the same config from [modelopt_recipes](../../modelopt_recipes/ptq.md) instead. Single node: diff --git a/examples/deepseek/deepseek_v4/ptq.py b/examples/deepseek/deepseek_v4/ptq.py index 57c19a2d2..89debc5d8 100644 --- a/examples/deepseek/deepseek_v4/ptq.py +++ b/examples/deepseek/deepseek_v4/ptq.py @@ -275,7 +275,7 @@ def load_deepseek_v4( return model -_PUBLISHED_RECIPE = "huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only" +_PUBLISHED_RECIPE = "models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only" _MTP_PROBE = "mtp.0.ffn.experts.0.w1_weight_quantizer" diff --git a/examples/kimi/README.md b/examples/kimi/README.md index 70364db2c..3cca677ab 100644 --- a/examples/kimi/README.md +++ b/examples/kimi/README.md @@ -16,7 +16,7 @@ one safetensors shard at a time: vision tower, and `lm_head` remain BF16. The exact quantization map is recorded in -`modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml`. +`modelopt_recipes/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml`. Because the source checkpoint uses packed MXFP4 tensors rather than ordinary Hugging Face `Linear.weight` tensors, run the streaming converter instead of passing this recipe to `hf_ptq.py`: @@ -25,7 +25,7 @@ passing this recipe to `hf_ptq.py`: python examples/kimi/kimi_k3/quantize_to_nvfp4.py \ --source_ckpt /models/moonshotai/Kimi-K3 \ --output_ckpt /models/Kimi-K3-NVFP4 \ - --recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \ + --recipe models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \ --jobs 8 ``` diff --git a/examples/kimi/kimi_k3/quantize_to_nvfp4.py b/examples/kimi/kimi_k3/quantize_to_nvfp4.py index bfe0be30c..5eb612c02 100644 --- a/examples/kimi/kimi_k3/quantize_to_nvfp4.py +++ b/examples/kimi/kimi_k3/quantize_to_nvfp4.py @@ -72,7 +72,7 @@ Usage (CPU partition, no GPU needed; ``--jobs`` shards convert in parallel): python quantize_to_nvfp4.py \\ --source_ckpt /path/to/Kimi-K3 \\ --output_ckpt /path/to/Kimi-K3-NVFP4 \\ - --recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \\ + --recipe models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \\ --jobs 8 """ @@ -161,7 +161,7 @@ _FP8_MAX = 448.0 _FP8_PB_BLOCK = 128 _NVFP4_BLOCK = 16 # NVFP4 block size (elements) -_PUBLISHED_RECIPE = "huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention" +_PUBLISHED_RECIPE = "models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention" def _conversion_settings_from_recipe(recipe_path: str) -> dict[str, Any]: diff --git a/modelopt/recipe/loader.py b/modelopt/recipe/loader.py index 6af6d0a8a..7366bd0d0 100644 --- a/modelopt/recipe/loader.py +++ b/modelopt/recipe/loader.py @@ -58,6 +58,15 @@ def _resolve_recipe_path(recipe_path: str | Path | Traversable) -> Path | Traver isinstance(recipe_path, Path) and recipe_path.is_absolute() ): rp_str = str(recipe_path) + # Backward-compat alias: checkpoint-mirror recipes moved from the old + # ``huggingface/models///`` layout to the top-level ``models/`` + # tier. A source checkout also keeps a ``huggingface/models`` -> ``../models`` + # symlink, but symlinks don't survive into built wheels, so rewrite the old + # prefix here too — that keeps saved ``--recipe huggingface/models/...`` paths + # working for pip-installed users, not just source checkouts. + _bc_prefix = "huggingface/models/" + if rp_str.replace("\\", "/").startswith(_bc_prefix): + rp_str = "models/" + rp_str.replace("\\", "/")[len(_bc_prefix) :] suffixes = [""] if rp_str.endswith((".yml", ".yaml")) else ["", ".yml", ".yaml"] for suffix in suffixes: candidate = BUILTIN_RECIPES_LIB.joinpath(rp_str + suffix) diff --git a/modelopt_recipes/README.md b/modelopt_recipes/README.md index b366e4cc6..5d1c1a310 100644 --- a/modelopt_recipes/README.md +++ b/modelopt_recipes/README.md @@ -42,13 +42,14 @@ huggingface/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast`. | Directory | What lives here | |-----------|-----------------| | `general/` | **Model-agnostic** recipes — a good starting point for any model. PTQ combos, speculative-decoding training, and distillation. | -| `huggingface//` | **Model-specific** recipes keyed by a HF `model_type`, optionally nested by released checkpoint. Use these first if your model has an entry. | -| `models//` | **Instance-specific** recipes that mirror a particular published checkpoint's quantization config. | +| `huggingface//` | **Architecture-specific** recipes keyed by a HF `model_type`; one recipe covers every checkpoint of that architecture. | +| `models///` | **Checkpoint-specific** recipes that mirror a particular published checkpoint, keyed by its model-hub path (e.g. `nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16`). | | `configs/` | Shared building blocks (`numerics/`, `ptq/units/`, `ptq/presets/`) that recipes compose from via `$import`. Not run directly. | -**Choosing where to look:** check `huggingface//` (then any nested -`/`) for your model first; if there's no entry, fall back to -`general/`. The presence of a model folder signals a recommended, tuned recipe. +**Choosing where to look:** check `models///` for your exact +checkpoint first, then `huggingface//` for its architecture; if +neither has an entry, fall back to `general/`. The presence of a model folder +signals a recommended, tuned recipe. --- @@ -65,7 +66,7 @@ Other general recipe families are documented inside their own folders: --- -## `huggingface/` — model-specific recipes +## `huggingface/` — architecture-specific recipes Each lives under its HF `model_type`. The point of a model folder is to capture **what differs from the generic preset** — usually an algorithm tweak or a @@ -77,9 +78,12 @@ how the model-specific recipes compare to the general ones and why they deviate. ## `models/` — checkpoint-specific recipes -These mirror a single **published checkpoint's** quantization config exactly — -a per-component mixed-precision scheme tuned to match a specific release. Browse -[`models/`](models/) for the available checkpoints. +These mirror a single **published checkpoint's** quantization config exactly — a +per-component mixed-precision scheme tuned to match a specific release. Each is +keyed by the checkpoint's **model-hub path** `/` (as on the +Hugging Face Hub, ModelScope, etc.). Browse [`models/`](models/) for the +available checkpoints; see [`models/README.md`](models/README.md) for the naming +convention. --- @@ -90,6 +94,6 @@ a per-component mixed-precision scheme tuned to match a specific release. Browse - **Tuned for a HF architecture** → `huggingface///`, with a `README.md` documenting the delta from the generic preset. Verify the exact `model_type` against the checkpoint's `config.json` before placing it. -- **Mirrors a specific released checkpoint** → `models//`. +- **Mirrors a specific released checkpoint** → `models///` (its model-hub path). - Share reused bodies via a `# modelopt-schema:`-tagged snippet and `$import` it; keep recipe wrappers thin. diff --git a/modelopt_recipes/huggingface/README.md b/modelopt_recipes/huggingface/README.md index c0361f50d..054493431 100644 --- a/modelopt_recipes/huggingface/README.md +++ b/modelopt_recipes/huggingface/README.md @@ -1,21 +1,23 @@ -# Model-specific recipes for Hugging Face models +# Architecture-specific recipes for Hugging Face models This folder holds model-optimization recipes (e.g. PTQ recipes) whose -behavior is tied to a **specific Hugging Face model architecture or model instance**. +behavior is tied to a **specific Hugging Face `model_type` (architecture)** — one +recipe covers every checkpoint of that architecture. Recipes tuned to a single +*published checkpoint* live in the sibling [`../models/`](../models/) tier +instead. ## Choosing a recipe -Built-in recipes live in two places: `modelopt_recipes/huggingface//` -for model-specific recipes and `modelopt_recipes/general/` for model-agnostic -ones. When deciding which to use: +Built-in recipes live in three tiers — pick the most specific that applies: -1. **Look in `huggingface//` first** for the target model's - Hugging Face `model_type`, and inside it for a nested - `/` folder if the recipe is tuned for one released - checkpoint rather than every checkpoint of that `model_type`. The - presence of a folder here signals that there is a recommended recipe - for that `model_type` or model instance. -2. **Fall back to `general/`** if no `/` folder applies. The +1. **[`../models///`](../models/)** first, if there is an entry + for your **exact** published checkpoint (keyed by its model-hub path). It + mirrors a validated, per-checkpoint scheme. +2. **`huggingface//`** for the target model's Hugging Face + `model_type` — an architecture-level recipe that applies to every checkpoint + of that `model_type`. The presence of a folder here signals a recommended + recipe for that architecture. +3. **Fall back to `general/`** if no `/` folder applies. The general recipes are a good starting point for any model — and the recommended starting point for a model architecture that does not yet have a model-specific entry. @@ -34,9 +36,6 @@ modelopt_recipes/huggingface/ .yaml [..yaml] # optional snippet helpers (see below) [README.md] # optional; describes what's model-specific - models/// - / - .yaml # exact published-checkpoint mirror ``` `` is the model-optimization workflow the recipe targets (e.g. @@ -46,11 +45,6 @@ Selecting a recipe at runtime uses the path relative to `modelopt_recipes/`, e.g. `--recipe huggingface///`. -Recipes that reproduce one exact published checkpoint may instead live under -`huggingface/models///`. This layout records the canonical -source checkpoint directly and avoids implying that the recipe applies to -every checkpoint sharing the same `model_type`. - ### Verifying a model's `model_type` The authoritative source for a model's `model_type` is the released @@ -75,17 +69,13 @@ include the field name the snippet represents as a secondary suffix is its natural canonical home; other importers reference it by the same relative path under `modelopt_recipes/`. -### Per-family nested layout for specific model variants +### Checkpoint-specific recipes -If a recipe is tuned for one specific released model rather than every -checkpoint under a `model_type`, nest the model name as an extra level: - -```text -/ - / - / - .yaml -``` +If a recipe is tuned for one specific released checkpoint rather than every +checkpoint of a `model_type`, it does not belong here — it lives in the +top-level [`../models/`](../models/) tier, keyed by the checkpoint's model-hub +path `/`. See [`../models/README.md`](../models/README.md) for +that convention. ### Per-folder READMEs diff --git a/modelopt_recipes/huggingface/models b/modelopt_recipes/huggingface/models new file mode 120000 index 000000000..1e266b1b5 --- /dev/null +++ b/modelopt_recipes/huggingface/models @@ -0,0 +1 @@ +../models \ No newline at end of file diff --git a/modelopt_recipes/models/README.md b/modelopt_recipes/models/README.md new file mode 100644 index 000000000..73055b7f9 --- /dev/null +++ b/modelopt_recipes/models/README.md @@ -0,0 +1,96 @@ +# Recipes for specific model-hub checkpoints + +This folder holds model-optimization recipes (e.g. PTQ recipes) tuned for a +**specific published model instance** — one checkpoint released on a model hub +such as the [Hugging Face Hub](https://huggingface.co/), +[ModelScope](https://modelscope.cn/), or similar. Unlike +[`../huggingface/`](../huggingface/), which keys recipes by a transformers +`model_type` (an architecture shared by many checkpoints), a recipe here mirrors +**one checkpoint's** quantization scheme verbatim. + +## Folder structure + +Each instance is keyed by its **model-hub path** — the same `/` +you pass to `from_pretrained(...)` or find in the hub URL. The on-disk path +mirrors the hub path exactly: + +```text +modelopt_recipes/models/ + / # hub namespace / organization, e.g. nvidia, mistralai + / # hub model id, e.g. Nemotron-3-Nano-4B-BF16 + / # optimization workflow, e.g. ptq + .yaml + [..yaml] # optional $import snippet helpers (see below) + [README.md] # optional; describes what's checkpoint-specific +``` + +For example, the recipe for the hub checkpoint `nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16` +(`https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16`) lives at +`models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/`. Because the folder path *is* the +hub path, you can go straight from a checkpoint id to its recipe — and back — +with no lookup table. + +`` is the optimization workflow the recipe targets (e.g. `ptq` for +post-training quantization). + +### Naming the `/` folders + +Use the checkpoint's exact hub `/`, including casing. When the +same weights are published on more than one hub (e.g. the Hugging Face Hub and +ModelScope) under the same `/`, a single folder serves them all. +When a recipe was tuned against one **canonical / base** checkpoint but also +applies to its mirrors, key it by that base model's id. + +## Choosing a recipe + +Prefer the most specific entry that applies to your model: + +1. **`models///`** — if there is an entry for your **exact** + checkpoint. It reproduces a validated, often per-component mixed-precision + scheme for that release; use it to match a published quantized checkpoint. +2. **[`huggingface//`](../huggingface/)** — an architecture-level + recipe that applies to every checkpoint of that `model_type`. +3. **[`general/`](../general/)** — model-agnostic recipes; a good starting point + for any model without a more specific entry. + +## Selecting a recipe at runtime + +Use the path relative to `modelopt_recipes/`: + +```text +--recipe models//// +``` + +or from Python: + +```python +from modelopt.recipe import load_recipe + +recipe = load_recipe("models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") +``` + +## What belongs here + +A recipe earns a place here only when it mirrors **one specific released (or +planned) checkpoint** — a hand-mapped, usually per-layer or per-component +precision scheme tuned to match that exact release. If the tuning generalizes to +every checkpoint of an architecture, it belongs under +[`../huggingface//`](../huggingface/) instead; if it is +model-agnostic, it belongs under [`../general/`](../general/). See +[`../ptq.md`](../ptq.md) for what each checkpoint mirror does and how it compares +to its general baseline. + +## Sharing content across recipes + +When several recipes reuse the same body, extract it into a sibling **snippet** +file with a `# modelopt-schema:` header and `$import` it, keeping each recipe +wrapper thin. Name snippets so they are obviously not runnable recipes (e.g. +`..yaml`), and reference them by their path relative to +`modelopt_recipes/`. + +## Per-folder READMEs + +Each `/` folder may contain a short `README.md` describing exactly what is +checkpoint-specific — which layers deviate, the calibration used, and the +reference checkpoint it mirrors — so reviewers and users don't have to diff the +YAML against the generic presets to see the intent. diff --git a/modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml b/modelopt_recipes/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml rename to modelopt_recipes/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only.yaml diff --git a/modelopt_recipes/huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml b/modelopt_recipes/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml rename to modelopt_recipes/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib.yaml diff --git a/modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml b/modelopt_recipes/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml rename to modelopt_recipes/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-max-calib.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-max-calib.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-max-calib.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-max-calib.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6.yaml diff --git a/modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml b/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml similarity index 100% rename from modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml rename to modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml diff --git a/modelopt_recipes/huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only.yaml b/modelopt_recipes/models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only.yaml similarity index 100% rename from modelopt_recipes/huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only.yaml rename to modelopt_recipes/models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only.yaml diff --git a/modelopt_recipes/ptq.md b/modelopt_recipes/ptq.md index 4ce5e0faf..589718444 100644 --- a/modelopt_recipes/ptq.md +++ b/modelopt_recipes/ptq.md @@ -3,10 +3,9 @@ This doc walks through the **PTQ quantization schemes** in two parts: the model-agnostic recipes under [`general/ptq/`](general/ptq/) (the recommended starting point for any model), and then the -[model-specific recipes](#model-specific-recipes-huggingface) under -`huggingface/` — per-`model_type` folders plus the -`huggingface/models///` tier — comparing each to its general -baseline and explaining why it deviates. +[model-specific recipes](#model-specific-recipes) — per-`model_type` folders +under `huggingface/` plus the checkpoint-mirror `models///` +tier — comparing each to its general baseline and explaining why it deviates. --- @@ -232,14 +231,14 @@ These can also be **stacked** when a single method isn't enough — e.g. `mse` + --- -## Model-specific recipes (`huggingface/`) +## Model-specific recipes The general recipes above are **model-agnostic**: they select layers by wildcard (`*mlp*`, `*self_attn*`, `*[kv]_bmm_quantizer`) and lean on the shared `default_disabled_quantizers` exclusions, so the same file works on any architecture whose module names follow the usual conventions. A recipe only earns a place under `huggingface//` or -`huggingface/models///` when a model has to **deviate** from +`models///` when a model has to **deviate** from that baseline. The deviations come in four kinds: | Kind | What changes vs. the general recipe | Examples | @@ -247,7 +246,7 @@ that baseline. The deviations come in four kinds: | **Architecture-aware `quant_cfg`** | Per-sub-module format choices a single wildcard scheme can't express | `minimax_m3_vl`, `qwen3_5`, `qwen3_5_moe`, `vit`, `nemotron_llama` | | **Algorithm override** | Same numerics & scope, but the *calibration algorithm* is tweaked because the default breaks or regresses | `gemma`, `gemma4`, `mpt` | | **Extra exclusions** | Adds disabled-quantizer patterns so non-language branches stay full precision | `nemotron_vl`, `diffusion_gemma` | -| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/Nemotron-3-*`, `models/mistralai/Mistral-Medium-3.5-128B` | +| **Checkpoint mirror** | A mixed-precision map reproducing one published checkpoint exactly | `models/nvidia/NVIDIA-Nemotron-3-*`, `models/mistralai/Mistral-Medium-3.5-128B` | The numerics and standard exclusions are still inherited from `configs/` wherever possible — the model folder captures *only* the delta. Each `/` @@ -302,7 +301,7 @@ general recipes never enable output quantizers, and the pattern must stay scoped to GEMM outputs — a `DynamicQuantize` on non-GEMM outputs (embedding lookup, pooling) fails to compile in TensorRT. -A lighter case: **`step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only`** is close to +A lighter case: **`models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only`** is close to `general/ptq/nvfp4_mlp_only` (NVFP4 on MoE/MLP weights+inputs, FP8 KV) but pinned to one released checkpoint and carrying instance-specific disables (`share_expert`, `moe.gate`, the conv1d branches). @@ -349,7 +348,7 @@ everything else matches the general recipe. ### Checkpoint mirrors — `models//` -The `huggingface/models/` tier reproduces a **single published (or planned) +The `models/` tier reproduces a **single published (or planned) checkpoint's** quant config verbatim: - **`models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention`** mirrors @@ -373,7 +372,7 @@ checkpoint's** quant config verbatim: `nvidia/Mistral-Medium-3.5-128B-NVFP4`: decoder MLP layers 4–86 use NVFP4 W4A4, edge MLP layers 0–3 and 87 use FP8 W8A8, and all attention projections and the KV cache use FP8. It uses max calibration. -- **`Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse`** mirrors +- **`models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse`** mirrors `nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4` exactly — a hybrid **Mamba-MoE** with a hand-mapped, **per-component** precision scheme: - MoE routed experts → NVFP4 W4A4, `group_size 16`, **static** weight scales @@ -384,17 +383,17 @@ checkpoint's** quant config verbatim: `nvfp4-mse.yaml` uses MSE calibration with an FP8-scale sweep (matches the release); `nvfp4-max-calib.yaml` is the identical layer map under plain `max` calibration, kept for comparison. -- **`Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6`** follows the same Super-style +- **`models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6`** follows the same Super-style component map (routed experts NVFP4 W4A4 block-16; shared experts + Mamba `in/out_proj` + KV cache FP8; everything else BF16), but the routed-expert weights use **Four-over-Six (4/6)** NVFP4: an MSE search picks each weight's amax multiplier from `[1.0, 1.5]` (M=6 vs. M=4). Activations stay dynamic NVFP4 (not MSE-calibrated). -- **`Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6`** applies +- **`models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6`** applies Four-over-Six NVFP4 W4A16 to routed experts, shared experts, and the language model head; Mamba `in/out_proj` weights and inputs plus the KV cache use FP8, while attention remains BF16. -- **`Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16`** mirrors the GGUF **Q4_K_M** bit +- **`models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16`** mirrors the GGUF **Q4_K_M** bit allocation of the Nemotron-H hybrid, mapped onto NVFP4/FP8 **per layer**: Q4_K/Q5_0 linears → NVFP4 W4A4 (attention q/k/v/o kept uniform so export can fuse them), the Q6_K MLP `down_proj` layers → FP8 W8A8, embeddings → NVFP4 diff --git a/pyproject.toml b/pyproject.toml index 599171302..4e6aea509 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -146,6 +146,19 @@ include = ["modelopt*"] modelopt = ["**/*.h", "**/*.cpp", "**/*.cu"] modelopt_recipes = ["**/*.yml", "**/*.yaml"] +[tool.setuptools.exclude-package-data] +# huggingface/models is a backward-compat symlink to ../models. The recursive +# package-data glob above follows it, so drop the aliased copies here to avoid +# shipping every checkpoint recipe twice; setuptools' exclude glob is non-recursive, +# hence the explicit org/model/task depths. MANIFEST.in prunes the symlink dir entry +# itself (build_py can't copy a symlink-to-dir). Old huggingface/models/... --recipe +# paths keep working via the loader alias in modelopt/recipe/loader.py. +modelopt_recipes = [ + "huggingface/models", + "huggingface/models/*/*/*/*.yaml", "huggingface/models/*/*/*/*.yml", + "huggingface/models/*/*/*/*/*.yaml", "huggingface/models/*/*/*/*/*.yml", +] + [tool.uv] managed = true # override-dependencies = ["torch; sys_platform == 'never'"] diff --git a/tests/unit/recipe/test_kimi_k3_recipe.py b/tests/unit/recipe/test_kimi_k3_recipe.py index 104102198..87f4bee7e 100644 --- a/tests/unit/recipe/test_kimi_k3_recipe.py +++ b/tests/unit/recipe/test_kimi_k3_recipe.py @@ -17,7 +17,7 @@ from modelopt.recipe import load_recipe -RECIPE = "huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention" +RECIPE = "models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention" def test_kimi_k3_recipe_matches_published_quantization_map(): diff --git a/tests/unit/recipe/test_loader.py b/tests/unit/recipe/test_loader.py index 3dfb9906a..5945187dd 100644 --- a/tests/unit/recipe/test_loader.py +++ b/tests/unit/recipe/test_loader.py @@ -157,27 +157,58 @@ def test_load_recipe_builtin_description(): assert len(recipe.description) > 0 -_BUILTIN_PTQ_RECIPES = [ - "general/ptq/fp8_default-kv_fp8", - "general/ptq/fp8_default-kv_fp8_cast", - "general/ptq/int4_blockwise_weight_only", - "general/ptq/nvfp4_act_headroom-kv_fp8_cast", - "general/ptq/nvfp4_default-kv_fp8", - "general/ptq/nvfp4_default-kv_fp8_cast", - "general/ptq/nvfp4_default-kv_nvfp4_cast", - "general/ptq/nvfp4_default-kv_none-gptq", - "general/ptq/nvfp4_experts_only-kv_fp8", - "general/ptq/nvfp4_experts_only-kv_fp8_cast", - "general/ptq/nvfp4_experts_only-kv_fp8_layerwise", - "huggingface/models/mistralai/Mistral-Medium-3.5-128B/ptq/nvfp4-max-calib", - "general/ptq/nvfp4_mlp_only-kv_fp8", - "general/ptq/nvfp4_mlp_only-novit-kv_fp8", - "general/ptq/nvfp4_mlp_only-kv_fp8_cast", - "general/ptq/nvfp4_omlp_only-kv_fp8", - "general/ptq/nvfp4_omlp_only-kv_fp8_cast", - "general/ptq/nvfp4_weight_only-kv_fp16", - "general/ptq/nvfp4_weight_only-kv_fp8_cast", -] +def test_load_recipe_huggingface_models_backward_compat_alias(): + """Old ``huggingface/models///...`` recipe paths resolve to the + top-level ``models/`` tier. + + The restructure keeps a ``huggingface/models`` -> ``../models`` source symlink, but + symlinks don't survive into built wheels, so the loader rewrites the prefix directly. + This guards that saved ``--recipe huggingface/models/...`` paths keep working for + pip-installed users, not just source checkouts. + """ + from modelopt.recipe.loader import _resolve_recipe_path + + root = Path(str(files("modelopt_recipes"))) + sample = next(root.glob("models/*/*/ptq/*.yaml")) + new_path = str(sample.relative_to(root).with_suffix("")) # models///ptq/ + old_path = "huggingface/" + new_path # huggingface/models///ptq/ + + assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path)) + recipe = load_recipe(old_path) + assert recipe.recipe_type == RecipeType.PTQ + assert isinstance(recipe, ModelOptPTQRecipe) + + +def _all_shipped_ptq_recipe_paths(): + """Every shipped PTQ recipe, discovered from disk rather than a hardcoded list.""" + root = files("modelopt_recipes") + paths = [] + for path in sorted(Path(str(root)).rglob("*.yaml")): + rel = path.relative_to(str(root)) + # Units/presets under configs/ are fragments, not standalone recipes. + if rel.parts[0] == "configs": + continue + raw = _load_raw_config(path) + # List-shaped fragments (layer-pattern units) are not recipes. + if not isinstance(raw, dict): + continue + if (raw.get("metadata") or {}).get("recipe_type") == "ptq": + paths.append(str(rel.with_suffix(""))) + return paths + + +# Discovered from disk (not hardcoded) so the smoke tests cover every shipped PTQ +# recipe — general/, huggingface//, and models/// — and +# never drift as recipes are added, moved, or removed. +_BUILTIN_PTQ_RECIPES = _all_shipped_ptq_recipe_paths() + + +def test_ptq_recipes_are_discovered(): + """Discovery must find recipes; otherwise the parametrized smoke tests below get an + empty parameter set and silently *skip* (pytest default) instead of running.""" + assert _BUILTIN_PTQ_RECIPES, ( + "No shipped PTQ recipes discovered under modelopt_recipes/ — recipe discovery is broken." + ) @pytest.mark.parametrize("recipe_path", _BUILTIN_PTQ_RECIPES) @@ -1925,25 +1956,7 @@ def test_load_recipe_autoquantize_builtin_general(recipe_path): assert recipe.auto_quantize.cost_excluded_layers == ["*visual*", "*mtp*", "*vision_tower*"] -def _all_shipped_ptq_recipe_paths(): - """Every shipped PTQ recipe, discovered from disk rather than a hardcoded list.""" - root = files("modelopt_recipes") - paths = [] - for path in sorted(Path(str(root)).rglob("*.yaml")): - rel = path.relative_to(str(root)) - # Units/presets under configs/ are fragments, not standalone recipes. - if rel.parts[0] == "configs": - continue - raw = _load_raw_config(path) - # List-shaped fragments (layer-pattern units) are not recipes. - if not isinstance(raw, dict): - continue - if (raw.get("metadata") or {}).get("recipe_type") == "ptq": - paths.append(str(rel.with_suffix(""))) - return paths - - -@pytest.mark.parametrize("recipe_path", _all_shipped_ptq_recipe_paths()) +@pytest.mark.parametrize("recipe_path", _BUILTIN_PTQ_RECIPES) def test_shipped_ptq_recipe_algorithm_config_constructs(recipe_path): """Every shipped PTQ recipe's ``algorithm`` must build its calibration config class. diff --git a/tests/unit/recipe/test_recipe_docs.py b/tests/unit/recipe/test_recipe_docs.py index adeb10dce..d584d1ecf 100644 --- a/tests/unit/recipe/test_recipe_docs.py +++ b/tests/unit/recipe/test_recipe_docs.py @@ -23,6 +23,8 @@ import re from importlib.resources import files from pathlib import Path +import pytest + RECIPES_DIR = Path(str(files("modelopt_recipes"))) GENERAL_PTQ_DIR = RECIPES_DIR / "general" / "ptq" PTQ_MD = RECIPES_DIR / "ptq.md" @@ -85,22 +87,115 @@ def test_general_ptq_recipe_count_in_ptq_md(): def test_every_model_specific_ptq_dir_is_mentioned(): - """Every model dir under huggingface/ with PTQ recipes must appear in ptq.md. + """Every model-specific PTQ recipe must be identifiable in ptq.md. - The identifier checked is the directory containing the ptq/ folder — the - HF model_type (e.g. ``gemma4``), a nested checkpoint dir (e.g. - ``Step3.5-Flash``), or a models// leaf (e.g. - ``Nemotron-3-Nano-4B``). + ``huggingface//ptq/`` recipes are checked by their ``model_type`` + (e.g. ``gemma4``); ``models///ptq/`` recipes are checked by their + full ``/`` hub path (e.g. ``nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16``), so + the org — the whole point of the top-level tier — is verified too and an org + re-key (e.g. ``step3p5`` → ``stepfun-ai``) can't silently drift from the doc. """ doc = _ptq_md_text() - hf_dir = RECIPES_DIR / "huggingface" - model_dirs = sorted( - {yaml_path.parent.parent.name for yaml_path in hf_dir.glob("**/ptq/*.yaml")} - ) - assert model_dirs, "No model-specific PTQ recipes found under huggingface/" - missing = [name for name in model_dirs if name not in doc] + # model_type recipes: huggingface//ptq/.yaml -> + hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "huggingface").glob("**/ptq/*.yaml")} + # checkpoint recipes: models///ptq/.yaml -> / + model_ids = { + f"{p.parent.parent.parent.name}/{p.parent.parent.name}" + for p in (RECIPES_DIR / "models").glob("**/ptq/*.yaml") + } + identifiers = sorted(hf_ids | model_ids) + assert identifiers, "No model-specific PTQ recipes found under huggingface/ or models/" + missing = [name for name in identifiers if name not in doc] assert not missing, ( f"Model-specific PTQ recipe folders are missing from " f"modelopt_recipes/ptq.md: {missing}. Add them to the model-specific " "recipes section (kinds table and/or the matching subsection)." ) + + +def test_checkpoint_recipes_live_in_the_top_level_models_tier(): + """Lock in the model_type-vs-checkpoint split. + + Checkpoint-mirror recipes belong at ``models///``; ``huggingface/`` + holds only per-``model_type`` recipes. ``huggingface/models`` is kept as a + backward-compatibility **symlink** to the top-level ``models/`` tier, so the old + ``--recipe huggingface/models///...`` paths still resolve; it must + stay a symlink that points at ``../models`` and never become a real directory that + holds recipes. A checkpoint recipe nested under a ``model_type`` (e.g. + ``huggingface////``) still fails loudly here instead + of silently shipping both tiers — e.g. on a bad merge that re-adds the old layout. + """ + hf = RECIPES_DIR / "huggingface" + models = RECIPES_DIR / "models" + hf_models = hf / "models" + assert hf_models.is_symlink(), ( + "huggingface/models must be a symlink to the top-level modelopt_recipes/models/ " + "tier (a backward-compat alias for the old --recipe paths), not a real directory." + ) + assert hf_models.resolve() == models.resolve(), ( + f"huggingface/models must resolve to the top-level models/ tier; resolves to " + f"{hf_models.resolve()} instead of {models.resolve()}." + ) + # Every recipe under huggingface/ must be // (3 parts); + # anything deeper is a checkpoint nested under a model_type and belongs in models/. + # Skip the huggingface/models symlink so the models/ recipes it aliases (4 parts) + # aren't miscounted as nested here. + nested = sorted( + str(p.relative_to(RECIPES_DIR)) + for ext in ("*.yaml", "*.yml") + for p in hf.glob(f"**/{ext}") + if hf_models not in p.parents and len(p.relative_to(hf).parts) != 3 + ) + assert not nested, ( + f"Recipes under huggingface/ must be //; found nested " + f"paths (a checkpoint recipe belongs under models///): {nested}" + ) + # Every recipe under models/ must be /// (4 parts) so the + # path is exactly the model-hub path; a different depth breaks that convention. + misplaced = sorted( + str(p.relative_to(RECIPES_DIR)) + for ext in ("*.yaml", "*.yml") + for p in models.glob(f"**/{ext}") + if len(p.relative_to(models).parts) != 4 + ) + assert not misplaced, ( + f"Recipes under models/ must be ///; found: {misplaced}" + ) + + +def test_launcher_yaml_recipe_paths_resolve(): + """Every modelopt_recipes recipe path a launcher example selects must resolve on disk. + + Guards against a recipe rename — e.g. keying ``models/nvidia/`` by the canonical Hub id, + which carries the ``NVIDIA-`` prefix — drifting from the launcher YAML that loads it. The + depth/doc tests can't catch a launcher pointing at a recipe path that no longer exists. + """ + repo_root = Path(__file__).resolve().parents[3] + launcher_dir = repo_root / "tools" / "launcher" / "examples" + if not launcher_dir.is_dir(): + pytest.skip("tools/launcher/examples not available in this checkout") + + def _resolves(rel: str) -> bool: + return any((RECIPES_DIR / f"{rel}{suffix}").exists() for suffix in ("", ".yaml", ".yml")) + + # ``--recipe

`` / ``QUANT_CFG:

`` are modelopt_recipes-relative — only tier-prefixed + # values are recipe paths; bare names like ``auto`` or ``FP8_DEFAULT_CFG`` are not. The + # ``modelopt_recipes/

.yaml`` form (e.g. ``--config``) embeds the path directly. + tier = r"(?:general|huggingface|models|configs)/[A-Za-z0-9._/-]+" + rel_re = re.compile(rf"(?:--recipe\s+|QUANT_CFG:\s*)({tier})") + abs_re = re.compile(rf"modelopt_recipes/({tier}\.ya?ml)") + + missing = [] + for yaml_path in sorted(launcher_dir.rglob("*.yaml")): + text = yaml_path.read_text(encoding="utf-8") + candidates = set(rel_re.findall(text)) | { + re.sub(r"\.ya?ml$", "", m) for m in abs_re.findall(text) + } + missing.extend( + f"{yaml_path.relative_to(repo_root)} -> {rel}" + for rel in sorted(candidates) + if not _resolves(rel) + ) + assert not missing, "Launcher YAMLs reference recipe paths that do not resolve:\n" + "\n".join( + missing + ) diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml index 6ec49dd7a..77ecb14c4 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/megatron_lm_ptq.yaml @@ -29,7 +29,7 @@ pipeline: - --calib-size 32 environment: - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 - - QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse + - QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 # MMLU + Export run as separate tasks; quantize.sh does quantize only. - RUN_MMLU: "false" @@ -52,7 +52,7 @@ pipeline: script: common/megatron_lm/export/export.sh environment: - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 - - QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse + - QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 - TP: "1" - PP: "4" diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml index dfe363da6..a9c493caf 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/megatron_lm_ptq.yaml @@ -30,7 +30,7 @@ pipeline: - --calib-size 32 environment: - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 - - QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6 + - QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6 - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 # MMLU + Export run as separate tasks; quantize.sh does quantize only. - RUN_MMLU: "false" @@ -53,7 +53,7 @@ pipeline: script: common/megatron_lm/export/export.sh environment: - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 - - QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6 + - QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6 - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 - TP: "1" - PP: "12" diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_streaming_dspark_warmstart.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_streaming_dspark_warmstart.yaml index 9b2334b2d..674201006 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_streaming_dspark_warmstart.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/hf_streaming_dspark_warmstart.yaml @@ -6,7 +6,7 @@ # # Unlike the other streaming examples, the drafter architecture is NOT overridden here: it # all lives in the recipe this points at, -# modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ +# modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ # speculative_decoding/dspark_warmstart.yaml # because every one of those fields is transcribed from the released checkpoint's own # config.json and must match it exactly. Keep drafter shape/behaviour changes in the recipe @@ -78,7 +78,7 @@ pipeline: args: # Drafter architecture, warm-start source, block size, mask token, SWA window, # causal attention and attention sink all come from this recipe — see header. - - --config modules/Model-Optimizer/modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml + - --config modules/Model-Optimizer/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml - model.model_name_or_path=<> - dflash.dflash_init_checkpoint=<> - data.data_path=/scratchspace/data/train.jsonl diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml index 833c854be..9d4cd54fd 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml @@ -65,7 +65,7 @@ pipeline: --tp_size 1 --pp_size 1 --ep_size 1 - --recipe huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 + --recipe models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 --calib_batch_size 1 --calib_num_samples 1000 --seq_length 32768 diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml index 7f9556bb3..799adfc8a 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_quantize.yaml @@ -22,7 +22,7 @@ pipeline: --hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 --trust_remote_code --tp_size 1 - --recipe huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 + --recipe models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 --calib_batch_size 8 --calib_num_samples 256 --seq_length 512 diff --git a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml index 8519647b1..4f543b0b2 100644 --- a/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml +++ b/tools/launcher/examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml @@ -53,7 +53,7 @@ pipeline: - --export-default-te-spec environment: - MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 - - QUANT_CFG: huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 + - QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6 - MLM_MODEL_CKPT: /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-MCore - MLM_MODEL_SAVE: /cicd/megatron-lm/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-W4A16 - HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16