Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)

### What does this PR do?

**Type of change:** Refactor (recipe-library layout) + documentation —
backward-breaking for saved `--recipe` paths.

Separate the two kinds of built-in Hugging Face recipes that were
previously mixed under `modelopt_recipes/huggingface/`:

- **`huggingface/<model_type>/`** — architecture recipes keyed by the
transformers `model_type`; one recipe covers every checkpoint of that
architecture. **Unchanged.**
- **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes
that mirror one specific published checkpoint, keyed by its **model-hub
path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk
path equals the hub path.

Concretely, the model-instance recipes move out of `huggingface/` to the
top level:

- `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` →
`models/mistralai/…`, `models/nvidia/…`
- `huggingface/step3p5/Step3.5-Flash/…` →
`models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo
id
[`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash)
— org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`)

**Why:** `modelopt_recipes/README.md` already documented a top-level
`models/` tier, but the files lived under `huggingface/models/` and
instance-specific recipes were awkwardly nested under the
per-`model_type` tree. This aligns the filesystem with the documented
layout and makes the instance tier hub-addressable — given a checkpoint
id you can find (or place) its recipe with no lookup table.
`load_recipe` resolves paths directly under `modelopt_recipes/`, so a
top-level `models/` sibling of `general/` and `huggingface/` works
identically.

The move is metadata-only — all recipe YAML content is byte-identical
(`R100` renames). Everything else is updating references (nvidia
launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`,
plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the
`10_recipes.rst` guide, which no longer describe instances under
`huggingface/`.

### Usage

Recipe paths for the moved checkpoint recipes lose the `huggingface/`
prefix (and Step 3.5 Flash is keyed by its hub id):

```python
from modelopt.recipe import load_recipe

# before
load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only")

# after
load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16")
load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only")
```

The same rename applies to `--recipe …` CLI values and launcher
`QUANT_CFG:` entries. Architecture recipes under
`huggingface/<model_type>/` are unaffected.

### Testing

- **Recipe resolution (torch-free):** parsed every recipe under
`models/` and confirmed all `$import` targets resolve against the recipe
root — 0 dangling across the tier.
- **Docs consistency:** re-ran the
`tests/unit/recipe/test_recipe_docs.py` logic; it now globs both
`huggingface/` and `models/`, and every model dir (incl.
`Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq`
recipe is still mentioned in `ptq.md`.
- **Reference sweep:** repo-wide grep confirms no remaining references
to the old paths outside the intentional historical CHANGELOG entries
(released 0.44 / 0.45).
- **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit`
hooks pass on the changed files.
- Note: the full `pytest` suite was not run in my environment (no
`torch`), so `test_recipe_docs.py` / `test_loader.py` should be
exercised in CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — `--recipe` / `load_recipe`
paths for the checkpoint-mirror tier change (drop the `huggingface/`
prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`).
Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the
only *released* old paths affected shipped in 0.45. A clean break was
chosen over a symlink or loader-alias shim.
- If you copied code from any other sources or added a new PIP
dependency …: N/A
- Did you write any new necessary tests?: ✅ — updated
`test_recipe_docs.py` to also glob the top-level `models/` tier so
instance recipes stay covered by the doc-consistency check.
- Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking
Changes** entry.
- Did you get Claude approval on this PR?: ❌ <!-- run /claude review -->

### Additional Information

Design note: an earlier iteration nested everything under
`huggingface/model_type/` + `huggingface/models/`; the final layout
keeps `huggingface/` flat (per-`model_type`) and lifts instances to a
top-level `models/` tier, matching what `modelopt_recipes/README.md`
already documented. The `Step3p5*` architecture class names (from the
model's `trust_remote_code` modeling code) are unrelated to the recipe
path and are left unchanged.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5,
and NVIDIA Nemotron models.
  * Added a Nemotron speculative-decoding warm-start recipe.

* **Documentation**
  * Clarified recipe selection and directory organization.
  * Documented checkpoint naming conventions and updated usage examples.

* **Bug Fixes**
* Updated launcher configurations and examples to reference the new
recipe locations and corrected model names.

* **Tests**
* Improved automatic recipe discovery and validation of documented
recipe paths.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-01 10:22:27 -07:00
committed by GitHub
parent 8810eb5e31
commit de3eda8a11
33 changed files with 357 additions and 127 deletions
@@ -29,7 +29,7 @@ pipeline:
- --calib-size 32
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse
- QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
# MMLU + Export run as separate tasks; quantize.sh does quantize only.
- RUN_MMLU: "false"
@@ -52,7 +52,7 @@ pipeline:
script: common/megatron_lm/export/export.sh
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse
- QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
- TP: "1"
- PP: "4"
@@ -30,7 +30,7 @@ pipeline:
- --calib-size 32
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6
- QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
# MMLU + Export run as separate tasks; quantize.sh does quantize only.
- RUN_MMLU: "false"
@@ -53,7 +53,7 @@ pipeline:
script: common/megatron_lm/export/export.sh
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- QUANT_CFG: huggingface/models/nvidia/Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6
- QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16/ptq/nvfp4-4o6
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
- TP: "1"
- PP: "12"
@@ -6,7 +6,7 @@
#
# Unlike the other streaming examples, the drafter architecture is NOT overridden here: it
# all lives in the recipe this points at,
# modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/
# modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/
# speculative_decoding/dspark_warmstart.yaml
# because every one of those fields is transcribed from the released checkpoint's own
# config.json and must match it exactly. Keep drafter shape/behaviour changes in the recipe
@@ -78,7 +78,7 @@ pipeline:
args:
# Drafter architecture, warm-start source, block size, mask token, SWA window,
# causal attention and attention sink all come from this recipe — see header.
- --config modules/Model-Optimizer/modelopt_recipes/huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml
- --config modules/Model-Optimizer/modelopt_recipes/models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/speculative_decoding/dspark_warmstart.yaml
- model.model_name_or_path=<<global_vars.hf_model>>
- dflash.dflash_init_checkpoint=<<global_vars.draft_model>>
- data.data_path=/scratchspace/data/train.jsonl
@@ -65,7 +65,7 @@ pipeline:
--tp_size 1
--pp_size 1
--ep_size 1
--recipe huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
--recipe models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
--calib_batch_size 1
--calib_num_samples 1000
--seq_length 32768
@@ -22,7 +22,7 @@ pipeline:
--hf_model_name_or_path nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
--trust_remote_code
--tp_size 1
--recipe huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
--recipe models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
--calib_batch_size 8
--calib_num_samples 256
--seq_length 512
@@ -53,7 +53,7 @@ pipeline:
- --export-default-te-spec
environment:
- MLM_MODEL_CFG: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- QUANT_CFG: huggingface/models/nvidia/Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
- QUANT_CFG: models/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/ptq/w4a16_nvfp4_4o6
- MLM_MODEL_CKPT: /cicd/megatron-lm-bf16/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-MCore
- MLM_MODEL_SAVE: /cicd/megatron-lm/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16-W4A16
- HF_MODEL_CKPT: /hf-local/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16