Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)

### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
Shengliang Xu
2026-09-15 12:16:12 -07:00
committed by GitHub
parent 30f89908f0
commit c7ed23a103
75 changed files with 426 additions and 178 deletions
+3 -3
View File
@@ -292,7 +292,7 @@ def test_autoquant_recipe_cost_excluded_layers_map_into_cost(monkeypatch):
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
)
aq = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
).auto_quantize
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(aq, args)
@@ -313,12 +313,12 @@ def test_autoquant_recipe_maps_module_search_spaces(monkeypatch):
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
)
recipe = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
)
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(
recipe.auto_quantize, args, fixed_quantize_config=recipe.quantize
)
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
assert inputs["quantization_formats"] == []
assert inputs["fixed_quantization_config"] == model_ptq.quantize.model_dump()
@@ -33,7 +33,7 @@ def hf_ptq(monkeypatch):
("recipe", "extracts_language_model"),
[
(None, True),
("huggingface/qwen3_vl/ptq/fp8_vision-kv_none", False),
("model_type/qwen3_vl/ptq/fp8_vision-kv_none", False),
],
)
def test_image_calibration_model_target_follows_recipe(
@@ -160,10 +160,10 @@ def test_image_calibration_uses_full_vlm_forward(hf_ptq, monkeypatch):
@pytest.mark.parametrize(
"recipe",
[
"huggingface/qwen3_vl/ptq/fp8_vision-kv_none",
"huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
"huggingface/qwen3_5/ptq/fp8_vision-kv_none",
"huggingface/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
"model_type/qwen3_vl/ptq/fp8_vision-kv_none",
"model_type/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
"model_type/qwen3_5/ptq/fp8_vision-kv_none",
"model_type/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
],
)
def test_vision_recipe_requires_image_calibration(hf_ptq, recipe):
@@ -45,12 +45,12 @@ _TINY_CONFIG = {
[
(
"embedding",
"huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj",
"model_type/nemotron_llama/ptq/nvfp4_output_quant_proj",
"TRT_FP4DynamicQuantize",
),
(
"reranking",
"huggingface/nemotron_llama/ptq/fp8_output_quant_proj",
"model_type/nemotron_llama/ptq/fp8_output_quant_proj",
"QuantizeLinear",
),
],
@@ -19,7 +19,7 @@ from _test_utils.torch.transformers_models import create_tiny_vit_dir
# Recipe variants the example ships.
_RECIPES = [
"huggingface/vit/ptq/fp8",
"model_type/vit/ptq/fp8",
]
@@ -60,7 +60,7 @@ def test_qwen_vision_recipe_calibrates_and_exports(
):
model = _get_tiny_qwen_vlm(model_type).to("cuda").eval()
model.config.architectures = [architecture]
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
vision_config = model.config.vision_config
pixel_width = (
vision_config.in_channels * vision_config.temporal_patch_size * vision_config.patch_size**2
+132 -13
View File
@@ -37,8 +37,13 @@ from modelopt.recipe.config import (
ModelOptPTQRecipe,
RecipeType,
)
from modelopt.recipe.loader import _apply_dotlist, load_config, load_recipe
from modelopt.torch.opt.config_loader import _load_raw_config, _schema_type
from modelopt.recipe.loader import _apply_dotlist, _resolve_recipe_path, load_config, load_recipe
from modelopt.torch.opt.config_loader import (
_alias_builtin_recipe_prefix,
_load_raw_config,
_resolve_config_path,
_schema_type,
)
from modelopt.torch.quantization.config import QuantizerAttributeConfig, normalize_quant_cfg_list
from modelopt.torch.quantization.mode import CalibrateModeRegistry, get_modelike_from_algo_cfg
@@ -160,28 +165,142 @@ def test_load_recipe_builtin_description():
assert len(recipe.description) > 0
def _first_builtin_ptq_recipe(root: Path, glob_pattern: str) -> Path:
"""Deterministically pick the first built-in *PTQ* recipe matching *glob_pattern*.
``glob`` order is filesystem-dependent (NTFS returns entries sorted, ext4 does not), and
recipe directories also hold non-recipe ``$import`` fragments (e.g. ``*.quant_cfg.yaml``,
``disabled_quantizers.yaml``) that are not loadable on their own. Sorting makes the pick
stable across platforms; skipping anything that does not load as a PTQ recipe keeps those
fragments from being mistaken for one.
"""
for path in sorted(root.glob(glob_pattern)):
rel = str(path.relative_to(root).with_suffix(""))
try:
if load_recipe(rel).recipe_type == RecipeType.PTQ:
return path
except Exception:
continue
raise AssertionError(f"no built-in PTQ recipe matched {glob_pattern!r} under {root}")
def test_load_recipe_huggingface_arch_backward_compat_alias():
"""Old ``huggingface/<model_type>/...`` recipe paths resolve to the renamed
``model_type/`` tier.
``huggingface/`` was renamed to ``model_type/``. A source checkout keeps a
``huggingface`` -> ``model_type`` symlink, but symlinks don't survive into built
wheels, so the loader rewrites the prefix directly. This guards that saved
``--recipe huggingface/<model_type>/...`` paths keep working for pip-installed
users, not just source checkouts.
"""
root = Path(str(files("modelopt_recipes")))
sample = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
new_path = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
old_path = "huggingface/" + new_path[len("model_type/") :] # huggingface/<arch>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(old_path)
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_huggingface_models_backward_compat_alias():
"""Old ``huggingface/models/<org>/<model_id>/...`` recipe paths resolve to the
top-level ``models/`` tier.
The restructure keeps a ``huggingface/models`` -> ``../models`` source symlink, but
symlinks don't survive into built wheels, so the loader rewrites the prefix directly.
This guards that saved ``--recipe huggingface/models/...`` paths keep working for
pip-installed users, not just source checkouts.
The ``huggingface/models/`` prefix is more specific than the ``huggingface/`` ->
``model_type/`` rename and must win: checkpoint mirrors moved all the way out to
the top-level ``models/`` tier. A source checkout keeps the
``huggingface`` -> ``model_type`` -> ``models`` symlink chain, but symlinks don't
survive into built wheels, so the loader rewrites the prefix directly. This guards
that saved ``--recipe huggingface/models/...`` paths keep working for pip-installed
users, not just source checkouts.
"""
from modelopt.recipe.loader import _resolve_recipe_path
root = Path(str(files("modelopt_recipes")))
sample = next(root.glob("models/*/*/ptq/*.yaml"))
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
new_path = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
old_path = "huggingface/" + new_path # huggingface/models/<org>/<model>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(old_path)
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
recipe = load_recipe(old_path)
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_model_type_models_alias_resolves_like_wheel():
"""``model_type/models/<org>/<model_id>/...`` resolves to the top-level ``models/`` tier.
``model_type/models`` is a source-only ``../models`` symlink that packaging prunes, so
without the loader alias the path would resolve in a checkout but 404 from a built wheel.
The alias rewrites the prefix to ``models/`` so both behave identically.
"""
root = Path(str(files("modelopt_recipes")))
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
canonical = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
aliased = "model_type/" + canonical # model_type/models/<org>/<model>/ptq/<file>
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
recipe = load_recipe(aliased)
assert str(_resolve_recipe_path(aliased)) == str(_resolve_recipe_path(canonical))
assert recipe.recipe_type == RecipeType.PTQ
assert isinstance(recipe, ModelOptPTQRecipe)
def test_load_recipe_local_tree_overrides_builtin_even_on_name_collision(tmp_path, monkeypatch):
"""A local recipe tree overrides a built-in of the same name — even when the name
collides with a shipped ``model_type``.
``_resolve_recipe_path`` probes the filesystem before the built-in library (matching
``config_loader._resolve_config_path``), so a user who keeps their own recipe tree on disk
is never silently shadowed by the deprecated-tier alias. This uses a *shipped* recipe's
exact relative path, spelled with the old ``huggingface/`` prefix that aliases to it, to
prove the local file wins over the built-in — the collision case a non-shipped name misses.
"""
root = Path(str(files("modelopt_recipes")))
shipped = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
# The path a user would keep locally: the shipped recipe's own relative path, but under the
# deprecated ``huggingface/`` tier that the alias rewrites to ``model_type/``.
old_rel = Path("huggingface") / shipped.relative_to(root).relative_to("model_type")
local = tmp_path / old_rel
local.parent.mkdir(parents=True)
local.write_text(
"metadata:\n recipe_type: ptq\nquantize:\n quant_cfg: {}\n algorithm: max\n"
)
monkeypatch.chdir(tmp_path)
resolved = _resolve_recipe_path(str(old_rel.with_suffix("")))
assert Path(resolved).resolve() == local.resolve()
def test_import_resolution_honors_huggingface_alias():
"""``$import`` resolution rewrites deprecated tier prefixes just like ``load_recipe``.
``$import`` paths go through ``config_loader._resolve_config_path`` (not the recipe-path
alias), so a custom recipe that imports a shipped snippet by its old ``huggingface/...``
path must still resolve from a wheel where the ``huggingface`` symlink is gone.
"""
# Prefix-rewrite mapping: architecture rename plus both checkpoint-mirror aliases.
assert _alias_builtin_recipe_prefix("huggingface/qwen3_vl/ptq/x") == "model_type/qwen3_vl/ptq/x"
assert (
_alias_builtin_recipe_prefix("huggingface/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
)
assert (
_alias_builtin_recipe_prefix("model_type/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
)
assert _alias_builtin_recipe_prefix("general/ptq/x") == "general/ptq/x" # untouched
root = Path(str(files("modelopt_recipes")))
sample = next(root.glob("model_type/*/ptq/*.yaml"))
canonical = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
old = "huggingface/" + canonical[len("model_type/") :] # huggingface/<arch>/ptq/<file>
assert str(_resolve_config_path(old)) == str(_resolve_config_path(canonical))
def _all_shipped_ptq_recipe_paths():
"""Every shipped PTQ recipe, discovered from disk rather than a hardcoded list."""
root = files("modelopt_recipes")
@@ -201,7 +320,7 @@ def _all_shipped_ptq_recipe_paths():
# Discovered from disk (not hardcoded) so the smoke tests cover every shipped PTQ
# recipe — general/, huggingface/<model_type>/, and models/<org>/<model_id>/ — and
# recipe — general/, model_type/<model_type>/, and models/<org>/<model_id>/ — and
# never drift as recipes are added, moved, or removed.
_BUILTIN_PTQ_RECIPES = _all_shipped_ptq_recipe_paths()
@@ -1905,10 +2024,10 @@ def test_load_recipe_autoquantize_builtin_active_moe():
def test_load_recipe_autoquantize_module_search_spaces():
"""Qwen recipe separates its fixed PTQ baseline from explicit search spaces."""
recipe = load_recipe(
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
)
aq = recipe.auto_quantize
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
assert recipe.quantize is not None
assert recipe.quantize == model_ptq.quantize
assert aq.candidate_formats == []
+1 -1
View File
@@ -61,7 +61,7 @@ class _MiniMaxModel(nn.Module):
def test_mxfp8_nvfp4_experts_recipe_quantizer_precedence():
model = _MiniMaxModel()
register_fused_experts_on_the_fly(model)
recipe = load_recipe("huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
recipe = load_recipe("model_type/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
config = recipe.quantize.model_dump()
assert config["algorithm"]["layerwise"]["enable"] is True
config["algorithm"] = None
+1 -1
View File
@@ -45,7 +45,7 @@ def test_qwen_vision_recipes_select_expected_quantizers(
):
model = _get_tiny_qwen_vlm(model_type)
assert model.config.model_type == model_type
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
mtq.quantize(model, quant_cfg, forward_loop=None)
modules = dict(model.named_modules())
+41 -27
View File
@@ -89,22 +89,22 @@ def test_general_ptq_recipe_count_in_ptq_md():
def test_every_model_specific_ptq_dir_is_mentioned():
"""Every model-specific PTQ recipe must be identifiable in ptq.md.
``huggingface/<model_type>/ptq/`` recipes are checked by their ``model_type``
``model_type/<model_type>/ptq/`` recipes are checked by their ``model_type``
(e.g. ``gemma4``); ``models/<org>/<model_id>/ptq/`` recipes are checked by their
full ``<org>/<model_id>`` hub path (e.g. ``nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16``), so
the org — the whole point of the top-level tier — is verified too and an org
re-key (e.g. ``step3p5`` → ``stepfun-ai``) can't silently drift from the doc.
"""
doc = _ptq_md_text()
# model_type recipes: huggingface/<model_type>/ptq/<recipe>.yaml -> <model_type>
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "huggingface").glob("**/ptq/*.yaml")}
# model_type recipes: model_type/<model_type>/ptq/<recipe>.yaml -> <model_type>
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "model_type").glob("**/ptq/*.yaml")}
# checkpoint recipes: models/<org>/<model_id>/ptq/<recipe>.yaml -> <org>/<model_id>
model_ids = {
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
for p in (RECIPES_DIR / "models").glob("**/ptq/*.yaml")
}
identifiers = sorted(hf_ids | model_ids)
assert identifiers, "No model-specific PTQ recipes found under huggingface/ or models/"
assert identifiers, "No model-specific PTQ recipes found under model_type/ or models/"
missing = [name for name in identifiers if name not in doc]
assert not missing, (
f"Model-specific PTQ recipe folders are missing from "
@@ -114,40 +114,54 @@ def test_every_model_specific_ptq_dir_is_mentioned():
def test_checkpoint_recipes_live_in_the_top_level_models_tier():
"""Lock in the model_type-vs-checkpoint split.
"""Lock in the model_type-vs-checkpoint split and the backward-compat symlinks.
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``huggingface/``
holds only per-``model_type`` recipes. ``huggingface/models`` is kept as a
backward-compatibility **symlink** to the top-level ``models/`` tier, so the old
``--recipe huggingface/models/<org>/<model_id>/...`` paths still resolve; it must
stay a symlink that points at ``../models`` and never become a real directory that
holds recipes. A checkpoint recipe nested under a ``model_type`` (e.g.
``huggingface/<model_type>/<checkpoint>/<task>/``) still fails loudly here instead
of silently shipping both tiers — e.g. on a bad merge that re-adds the old layout.
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``model_type/``
(formerly ``huggingface/``) holds only per-``model_type`` recipes. Two
backward-compatibility **symlinks** are kept so old ``--recipe`` paths still
resolve: the top-level ``huggingface`` -> ``model_type`` rename alias, and the
nested ``model_type/models`` -> ``../models`` alias for the old
``huggingface/models/<org>/<model_id>/...`` checkpoint paths. Both must stay
symlinks and never become real directories that hold recipes. A checkpoint recipe
nested under a ``model_type`` (e.g. ``model_type/<model_type>/<checkpoint>/<task>/``)
still fails loudly here instead of silently shipping both tiers — e.g. on a bad
merge that re-adds the old layout.
"""
hf = RECIPES_DIR / "huggingface"
model_type = RECIPES_DIR / "model_type"
models = RECIPES_DIR / "models"
hf_models = hf / "models"
assert hf_models.is_symlink(), (
"huggingface/models must be a symlink to the top-level modelopt_recipes/models/ "
"tier (a backward-compat alias for the old --recipe paths), not a real directory."
hf_alias = RECIPES_DIR / "huggingface"
mt_models = model_type / "models"
# Top-level huggingface -> model_type rename alias.
assert hf_alias.is_symlink(), (
"modelopt_recipes/huggingface must be a backward-compat symlink to model_type/ "
"(the rename alias), not a real directory."
)
assert hf_models.resolve() == models.resolve(), (
f"huggingface/models must resolve to the top-level models/ tier; resolves to "
f"{hf_models.resolve()} instead of {models.resolve()}."
assert hf_alias.resolve() == model_type.resolve(), (
f"huggingface must resolve to the model_type/ tier; resolves to "
f"{hf_alias.resolve()} instead of {model_type.resolve()}."
)
# Every recipe under huggingface/ must be <model_type>/<task>/<file> (3 parts);
# Nested model_type/models -> ../models alias for the old huggingface/models/... paths.
assert mt_models.is_symlink(), (
"model_type/models must be a symlink to the top-level modelopt_recipes/models/ "
"tier (a backward-compat alias for the old --recipe huggingface/models/... paths), "
"not a real directory."
)
assert mt_models.resolve() == models.resolve(), (
f"model_type/models must resolve to the top-level models/ tier; resolves to "
f"{mt_models.resolve()} instead of {models.resolve()}."
)
# Every recipe under model_type/ must be <model_type>/<task>/<file> (3 parts);
# anything deeper is a checkpoint nested under a model_type and belongs in models/.
# Skip the huggingface/models symlink so the models/ recipes it aliases (4 parts)
# Skip the model_type/models symlink so the models/ recipes it aliases (4 parts)
# aren't miscounted as nested here.
nested = sorted(
str(p.relative_to(RECIPES_DIR))
for ext in ("*.yaml", "*.yml")
for p in hf.glob(f"**/{ext}")
if hf_models not in p.parents and len(p.relative_to(hf).parts) != 3
for p in model_type.glob(f"**/{ext}")
if mt_models not in p.parents and len(p.relative_to(model_type).parts) != 3
)
assert not nested, (
f"Recipes under huggingface/ must be <model_type>/<task>/<file>; found nested "
f"Recipes under model_type/ must be <model_type>/<task>/<file>; found nested "
f"paths (a checkpoint recipe belongs under models/<org>/<model_id>/): {nested}"
)
# Every recipe under models/ must be <org>/<model_id>/<task>/<file> (4 parts) so the
@@ -181,7 +195,7 @@ def test_launcher_yaml_recipe_paths_resolve():
# ``--recipe <p>`` / ``QUANT_CFG: <p>`` are modelopt_recipes-relative — only tier-prefixed
# values are recipe paths; bare names like ``auto`` or ``FP8_DEFAULT_CFG`` are not. The
# ``modelopt_recipes/<p>.yaml`` form (e.g. ``--config``) embeds the path directly.
tier = r"(?:general|huggingface|models|configs)/[A-Za-z0-9._/-]+"
tier = r"(?:general|model_type|models|configs)/[A-Za-z0-9._/-]+"
rel_re = re.compile(rf"(?:--recipe\s+|QUANT_CFG:\s*)({tier})")
abs_re = re.compile(rf"modelopt_recipes/({tier}\.ya?ml)")
+6 -6
View File
@@ -121,8 +121,8 @@ def _enabled(module, quantizer="weight_quantizer"):
@pytest.mark.parametrize(
"recipe_name",
[
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
],
)
def test_routed_experts_are_quantized(recipe_name):
@@ -144,8 +144,8 @@ def test_routed_experts_are_quantized(recipe_name):
@pytest.mark.parametrize(
"recipe_name",
[
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
],
)
def test_router_shared_expert_and_head_stay_bf16(recipe_name):
@@ -160,7 +160,7 @@ def test_router_shared_expert_and_head_stay_bf16(recipe_name):
def test_experts_only_leaves_dense_mlp_bf16():
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
dense_mlp = model.model.language_model.layers[1].mlp
for proj in ("gate_proj", "up_proj", "down_proj"):
@@ -168,7 +168,7 @@ def test_experts_only_leaves_dense_mlp_bf16():
def test_mlp_only_also_quantizes_dense_mlp():
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
dense_mlp = model.model.language_model.layers[1].mlp
for proj in ("gate_proj", "up_proj", "down_proj"):
@@ -40,6 +40,26 @@ from modelopt.torch.export.quant_aware_conversion import (
BLOCK = 16
def _set_scope_attr(transform, name, value):
"""Set an optional scoped-match attribute that only some transformers versions expose.
transformers>=5.9 dropped ``base_model_prefix`` from ``WeightTransform``'s ``__slots__``
(scoped matching now keys off ``scope_prefix`` alone); older supported versions still
carry it. Production ``_scope_prefixes`` reads it via ``getattr(..., None)``, so skipping
the assignment where the slot is absent is equivalent — and lets these tests run across
the whole supported transformers range instead of ``AttributeError``-ing on the setattr.
The suppression is scoped to that one known version-dependent slot: a setattr failure for
any other name (a typo or a future rename) still raises instead of silently no-op-ing.
"""
try:
setattr(transform, name, value)
except AttributeError:
if name != "base_model_prefix":
raise
# Tiny Mixtral shaped to match the synthetic expert tensors built by ``_nvfp4_linear`` below.
_MIXTRAL_KWARGS = {
"hidden_size": 32,
@@ -341,7 +361,7 @@ def test_scoped_submodel_prefix_change_does_not_capture_siblings():
prefix_change = PrefixChange(prefix_to_remove="vision_model")
prefix_change.scope_prefix = "model.vision_tower"
prefix_change.base_model_prefix = "model"
_set_scope_attr(prefix_change, "base_model_prefix", "model")
model._weight_conversions = [prefix_change]
state_dict = {
@@ -381,7 +401,7 @@ def test_scoped_rule_maps_config_module_names_consistently():
prefix_change = PrefixChange(prefix_to_remove="vision_model")
prefix_change.scope_prefix = "model.vision_tower"
prefix_change.base_model_prefix = "model"
_set_scope_attr(prefix_change, "base_model_prefix", "model")
model._weight_conversions = [prefix_change]
mapper = build_reverse_name_mapper(model)
@@ -420,7 +440,7 @@ def test_root_scoped_rule_still_faces_shadowing_guard():
)
# Root scope: reaches every key, exactly like an unscoped rule.
renaming.scope_prefix = ""
renaming.base_model_prefix = ""
_set_scope_attr(renaming, "base_model_prefix", "")
model._weight_conversions = [renaming]
state_dict = {
@@ -454,7 +474,7 @@ def test_scoped_weight_converter_is_refused():
operations=[Chunk(dim=0)],
)
conv.scope_prefix = "model.language_model"
conv.base_model_prefix = "model"
_set_scope_attr(conv, "base_model_prefix", "model")
model = types.SimpleNamespace(_weight_conversions=[conv])
sd = _nvfp4_linear("model.language_model.layers.0.mlp.gate_up_proj", 8, 16)