mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Restructure recipes: split per-model_type recipes from model-hub checkpoint recipes (#2219)
### What does this PR do? **Type of change:** Refactor (recipe-library layout) + documentation — backward-breaking for saved `--recipe` paths. Separate the two kinds of built-in Hugging Face recipes that were previously mixed under `modelopt_recipes/huggingface/`: - **`huggingface/<model_type>/`** — architecture recipes keyed by the transformers `model_type`; one recipe covers every checkpoint of that architecture. **Unchanged.** - **`models/<org>/<model_id>/`** — a *new top-level tier* for recipes that mirror one specific published checkpoint, keyed by its **model-hub path** (as on the Hugging Face Hub, ModelScope, etc.) so the on-disk path equals the hub path. Concretely, the model-instance recipes move out of `huggingface/` to the top level: - `huggingface/models/mistralai/…`, `huggingface/models/nvidia/…` → `models/mistralai/…`, `models/nvidia/…` - `huggingface/step3p5/Step3.5-Flash/…` → `models/stepfun-ai/Step-3.5-Flash/…` (re-keyed to the canonical HF repo id [`stepfun-ai/Step-3.5-Flash`](https://huggingface.co/stepfun-ai/Step-3.5-Flash) — org `step3p5`→`stepfun-ai`, id `Step3.5-Flash`→`Step-3.5-Flash`) **Why:** `modelopt_recipes/README.md` already documented a top-level `models/` tier, but the files lived under `huggingface/models/` and instance-specific recipes were awkwardly nested under the per-`model_type` tree. This aligns the filesystem with the documented layout and makes the instance tier hub-addressable — given a checkpoint id you can find (or place) its recipe with no lookup table. `load_recipe` resolves paths directly under `modelopt_recipes/`, so a top-level `models/` sibling of `general/` and `huggingface/` works identically. The move is metadata-only — all recipe YAML content is byte-identical (`R100` renames). Everything else is updating references (nvidia launcher YAMLs, `test_loader.py`) and docs: a new `models/README.md`, plus `huggingface/README.md`, root `README.md`, `ptq.md`, and the `10_recipes.rst` guide, which no longer describe instances under `huggingface/`. ### Usage Recipe paths for the moved checkpoint recipes lose the `huggingface/` prefix (and Step 3.5 Flash is keyed by its hub id): ```python from modelopt.recipe import load_recipe # before load_recipe("huggingface/models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("huggingface/step3p5/Step3.5-Flash/ptq/nvfp4-mlp-only") # after load_recipe("models/nvidia/Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a16") load_recipe("models/stepfun-ai/Step-3.5-Flash/ptq/nvfp4-mlp-only") ``` The same rename applies to `--recipe …` CLI values and launcher `QUANT_CFG:` entries. Architecture recipes under `huggingface/<model_type>/` are unaffected. ### Testing - **Recipe resolution (torch-free):** parsed every recipe under `models/` and confirmed all `$import` targets resolve against the recipe root — 0 dangling across the tier. - **Docs consistency:** re-ran the `tests/unit/recipe/test_recipe_docs.py` logic; it now globs both `huggingface/` and `models/`, and every model dir (incl. `Step-3.5-Flash`, `Nemotron-3-Nano-4B-BF16`, …) plus every `general/ptq` recipe is still mentioned in `ptq.md`. - **Reference sweep:** repo-wide grep confirms no remaining references to the old paths outside the intentional historical CHANGELOG entries (released 0.44 / 0.45). - **pre-commit:** `markdownlint-cli2`, license-insert, and `bandit` hooks pass on the changed files. - Note: the full `pytest` suite was not run in my environment (no `torch`), so `test_recipe_docs.py` / `test_loader.py` should be exercised in CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — `--recipe` / `load_recipe` paths for the checkpoint-mirror tier change (drop the `huggingface/` prefix; `step3p5/Step3.5-Flash` → `stepfun-ai/Step-3.5-Flash`). Documented as a Backward Breaking Change in `CHANGELOG.rst` (0.47); the only *released* old paths affected shipped in 0.45. A clean break was chosen over a symlink or loader-alias shim. - If you copied code from any other sources or added a new PIP dependency …: N/A - Did you write any new necessary tests?: ✅ — updated `test_recipe_docs.py` to also glob the top-level `models/` tier so instance recipes stay covered by the doc-consistency check. - Did you update Changelog?: ✅ — added a 0.47 **Backward Breaking Changes** entry. - Did you get Claude approval on this PR?: ❌ <!-- run /claude review --> ### Additional Information Design note: an earlier iteration nested everything under `huggingface/model_type/` + `huggingface/models/`; the final layout keeps `huggingface/` flat (per-`model_type`) and lifts instances to a top-level `models/` tier, matching what `modelopt_recipes/README.md` already documented. The `Step3p5*` architecture class names (from the model's `trust_remote_code` modeling code) are unrelated to the recipe path and are left unchanged. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added checkpoint-specific PTQ recipes for Kimi-K3, Mistral Medium 3.5, and NVIDIA Nemotron models. * Added a Nemotron speculative-decoding warm-start recipe. * **Documentation** * Clarified recipe selection and directory organization. * Documented checkpoint naming conventions and updated usage examples. * **Bug Fixes** * Updated launcher configurations and examples to reference the new recipe locations and corrected model names. * **Tests** * Improved automatic recipe discovery and validation of documented recipe paths. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
@@ -129,7 +129,7 @@ python ${DS_V4}/inference/convert.py \
|
||||
### Calibrate routed experts
|
||||
|
||||
The quantization config defaults to the built-in routed-expert NVFP4 setup. Pass
|
||||
`--recipe huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only`
|
||||
`--recipe models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only`
|
||||
to load the same config from [modelopt_recipes](../../modelopt_recipes/ptq.md) instead.
|
||||
|
||||
Single node:
|
||||
|
||||
@@ -275,7 +275,7 @@ def load_deepseek_v4(
|
||||
return model
|
||||
|
||||
|
||||
_PUBLISHED_RECIPE = "huggingface/models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only"
|
||||
_PUBLISHED_RECIPE = "models/deepseek-ai/DeepSeek-V4-Pro-0813/ptq/nvfp4_experts_only"
|
||||
|
||||
|
||||
_MTP_PROBE = "mtp.0.ffn.experts.0.w1_weight_quantizer"
|
||||
|
||||
@@ -16,7 +16,7 @@ one safetensors shard at a time:
|
||||
vision tower, and `lm_head` remain BF16.
|
||||
|
||||
The exact quantization map is recorded in
|
||||
`modelopt_recipes/huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml`.
|
||||
`modelopt_recipes/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention.yaml`.
|
||||
Because the source checkpoint uses packed MXFP4 tensors rather than ordinary
|
||||
Hugging Face `Linear.weight` tensors, run the streaming converter instead of
|
||||
passing this recipe to `hf_ptq.py`:
|
||||
@@ -25,7 +25,7 @@ passing this recipe to `hf_ptq.py`:
|
||||
python examples/kimi/kimi_k3/quantize_to_nvfp4.py \
|
||||
--source_ckpt /models/moonshotai/Kimi-K3 \
|
||||
--output_ckpt /models/Kimi-K3-NVFP4 \
|
||||
--recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \
|
||||
--recipe models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \
|
||||
--jobs 8
|
||||
```
|
||||
|
||||
|
||||
@@ -72,7 +72,7 @@ Usage (CPU partition, no GPU needed; ``--jobs`` shards convert in parallel):
|
||||
python quantize_to_nvfp4.py \\
|
||||
--source_ckpt /path/to/Kimi-K3 \\
|
||||
--output_ckpt /path/to/Kimi-K3-NVFP4 \\
|
||||
--recipe huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \\
|
||||
--recipe models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention \\
|
||||
--jobs 8
|
||||
"""
|
||||
|
||||
@@ -161,7 +161,7 @@ _FP8_MAX = 448.0
|
||||
_FP8_PB_BLOCK = 128
|
||||
_NVFP4_BLOCK = 16 # NVFP4 block size (elements)
|
||||
|
||||
_PUBLISHED_RECIPE = "huggingface/models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention"
|
||||
_PUBLISHED_RECIPE = "models/moonshotai/Kimi-K3/ptq/nvfp4_experts-fp8_pb_attention"
|
||||
|
||||
|
||||
def _conversion_settings_from_recipe(recipe_path: str) -> dict[str, Any]:
|
||||
|
||||
Reference in New Issue
Block a user