mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
This commit is contained in:
@@ -292,7 +292,7 @@ def test_autoquant_recipe_cost_excluded_layers_map_into_cost(monkeypatch):
|
||||
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
|
||||
)
|
||||
aq = load_recipe(
|
||||
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
|
||||
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_at_6p0bits-active_moe"
|
||||
).auto_quantize
|
||||
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(aq, args)
|
||||
|
||||
@@ -313,12 +313,12 @@ def test_autoquant_recipe_maps_module_search_spaces(monkeypatch):
|
||||
monkeypatch, "--pyt_ckpt_path", "dummy", "--kv_cache_qformat", "none"
|
||||
)
|
||||
recipe = load_recipe(
|
||||
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
|
||||
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
|
||||
)
|
||||
inputs = hf_ptq._mtq_inputs_from_auto_quantize_config(
|
||||
recipe.auto_quantize, args, fixed_quantize_config=recipe.quantize
|
||||
)
|
||||
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
|
||||
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
|
||||
|
||||
assert inputs["quantization_formats"] == []
|
||||
assert inputs["fixed_quantization_config"] == model_ptq.quantize.model_dump()
|
||||
|
||||
@@ -33,7 +33,7 @@ def hf_ptq(monkeypatch):
|
||||
("recipe", "extracts_language_model"),
|
||||
[
|
||||
(None, True),
|
||||
("huggingface/qwen3_vl/ptq/fp8_vision-kv_none", False),
|
||||
("model_type/qwen3_vl/ptq/fp8_vision-kv_none", False),
|
||||
],
|
||||
)
|
||||
def test_image_calibration_model_target_follows_recipe(
|
||||
@@ -160,10 +160,10 @@ def test_image_calibration_uses_full_vlm_forward(hf_ptq, monkeypatch):
|
||||
@pytest.mark.parametrize(
|
||||
"recipe",
|
||||
[
|
||||
"huggingface/qwen3_vl/ptq/fp8_vision-kv_none",
|
||||
"huggingface/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
|
||||
"huggingface/qwen3_5/ptq/fp8_vision-kv_none",
|
||||
"huggingface/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
|
||||
"model_type/qwen3_vl/ptq/fp8_vision-kv_none",
|
||||
"model_type/qwen3_vl/ptq/fp8_vision_lm-kv_fp8_cast",
|
||||
"model_type/qwen3_5/ptq/fp8_vision-kv_none",
|
||||
"model_type/qwen3_5/ptq/fp8_vision_lm-kv_fp8_cast",
|
||||
],
|
||||
)
|
||||
def test_vision_recipe_requires_image_calibration(hf_ptq, recipe):
|
||||
|
||||
@@ -45,12 +45,12 @@ _TINY_CONFIG = {
|
||||
[
|
||||
(
|
||||
"embedding",
|
||||
"huggingface/nemotron_llama/ptq/nvfp4_output_quant_proj",
|
||||
"model_type/nemotron_llama/ptq/nvfp4_output_quant_proj",
|
||||
"TRT_FP4DynamicQuantize",
|
||||
),
|
||||
(
|
||||
"reranking",
|
||||
"huggingface/nemotron_llama/ptq/fp8_output_quant_proj",
|
||||
"model_type/nemotron_llama/ptq/fp8_output_quant_proj",
|
||||
"QuantizeLinear",
|
||||
),
|
||||
],
|
||||
|
||||
@@ -19,7 +19,7 @@ from _test_utils.torch.transformers_models import create_tiny_vit_dir
|
||||
|
||||
# Recipe variants the example ships.
|
||||
_RECIPES = [
|
||||
"huggingface/vit/ptq/fp8",
|
||||
"model_type/vit/ptq/fp8",
|
||||
]
|
||||
|
||||
|
||||
|
||||
@@ -60,7 +60,7 @@ def test_qwen_vision_recipe_calibrates_and_exports(
|
||||
):
|
||||
model = _get_tiny_qwen_vlm(model_type).to("cuda").eval()
|
||||
model.config.architectures = [architecture]
|
||||
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
|
||||
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
|
||||
vision_config = model.config.vision_config
|
||||
pixel_width = (
|
||||
vision_config.in_channels * vision_config.temporal_patch_size * vision_config.patch_size**2
|
||||
|
||||
@@ -37,8 +37,13 @@ from modelopt.recipe.config import (
|
||||
ModelOptPTQRecipe,
|
||||
RecipeType,
|
||||
)
|
||||
from modelopt.recipe.loader import _apply_dotlist, load_config, load_recipe
|
||||
from modelopt.torch.opt.config_loader import _load_raw_config, _schema_type
|
||||
from modelopt.recipe.loader import _apply_dotlist, _resolve_recipe_path, load_config, load_recipe
|
||||
from modelopt.torch.opt.config_loader import (
|
||||
_alias_builtin_recipe_prefix,
|
||||
_load_raw_config,
|
||||
_resolve_config_path,
|
||||
_schema_type,
|
||||
)
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig, normalize_quant_cfg_list
|
||||
from modelopt.torch.quantization.mode import CalibrateModeRegistry, get_modelike_from_algo_cfg
|
||||
|
||||
@@ -160,28 +165,142 @@ def test_load_recipe_builtin_description():
|
||||
assert len(recipe.description) > 0
|
||||
|
||||
|
||||
def _first_builtin_ptq_recipe(root: Path, glob_pattern: str) -> Path:
|
||||
"""Deterministically pick the first built-in *PTQ* recipe matching *glob_pattern*.
|
||||
|
||||
``glob`` order is filesystem-dependent (NTFS returns entries sorted, ext4 does not), and
|
||||
recipe directories also hold non-recipe ``$import`` fragments (e.g. ``*.quant_cfg.yaml``,
|
||||
``disabled_quantizers.yaml``) that are not loadable on their own. Sorting makes the pick
|
||||
stable across platforms; skipping anything that does not load as a PTQ recipe keeps those
|
||||
fragments from being mistaken for one.
|
||||
"""
|
||||
for path in sorted(root.glob(glob_pattern)):
|
||||
rel = str(path.relative_to(root).with_suffix(""))
|
||||
try:
|
||||
if load_recipe(rel).recipe_type == RecipeType.PTQ:
|
||||
return path
|
||||
except Exception:
|
||||
continue
|
||||
raise AssertionError(f"no built-in PTQ recipe matched {glob_pattern!r} under {root}")
|
||||
|
||||
|
||||
def test_load_recipe_huggingface_arch_backward_compat_alias():
|
||||
"""Old ``huggingface/<model_type>/...`` recipe paths resolve to the renamed
|
||||
``model_type/`` tier.
|
||||
|
||||
``huggingface/`` was renamed to ``model_type/``. A source checkout keeps a
|
||||
``huggingface`` -> ``model_type`` symlink, but symlinks don't survive into built
|
||||
wheels, so the loader rewrites the prefix directly. This guards that saved
|
||||
``--recipe huggingface/<model_type>/...`` paths keep working for pip-installed
|
||||
users, not just source checkouts.
|
||||
"""
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
sample = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
|
||||
new_path = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
|
||||
old_path = "huggingface/" + new_path[len("model_type/") :] # huggingface/<arch>/ptq/<file>
|
||||
|
||||
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
|
||||
recipe = load_recipe(old_path)
|
||||
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
|
||||
assert recipe.recipe_type == RecipeType.PTQ
|
||||
assert isinstance(recipe, ModelOptPTQRecipe)
|
||||
|
||||
|
||||
def test_load_recipe_huggingface_models_backward_compat_alias():
|
||||
"""Old ``huggingface/models/<org>/<model_id>/...`` recipe paths resolve to the
|
||||
top-level ``models/`` tier.
|
||||
|
||||
The restructure keeps a ``huggingface/models`` -> ``../models`` source symlink, but
|
||||
symlinks don't survive into built wheels, so the loader rewrites the prefix directly.
|
||||
This guards that saved ``--recipe huggingface/models/...`` paths keep working for
|
||||
pip-installed users, not just source checkouts.
|
||||
The ``huggingface/models/`` prefix is more specific than the ``huggingface/`` ->
|
||||
``model_type/`` rename and must win: checkpoint mirrors moved all the way out to
|
||||
the top-level ``models/`` tier. A source checkout keeps the
|
||||
``huggingface`` -> ``model_type`` -> ``models`` symlink chain, but symlinks don't
|
||||
survive into built wheels, so the loader rewrites the prefix directly. This guards
|
||||
that saved ``--recipe huggingface/models/...`` paths keep working for pip-installed
|
||||
users, not just source checkouts.
|
||||
"""
|
||||
from modelopt.recipe.loader import _resolve_recipe_path
|
||||
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
sample = next(root.glob("models/*/*/ptq/*.yaml"))
|
||||
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
|
||||
new_path = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
|
||||
old_path = "huggingface/" + new_path # huggingface/models/<org>/<model>/ptq/<file>
|
||||
|
||||
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
|
||||
recipe = load_recipe(old_path)
|
||||
assert str(_resolve_recipe_path(old_path)) == str(_resolve_recipe_path(new_path))
|
||||
recipe = load_recipe(old_path)
|
||||
assert recipe.recipe_type == RecipeType.PTQ
|
||||
assert isinstance(recipe, ModelOptPTQRecipe)
|
||||
|
||||
|
||||
def test_load_recipe_model_type_models_alias_resolves_like_wheel():
|
||||
"""``model_type/models/<org>/<model_id>/...`` resolves to the top-level ``models/`` tier.
|
||||
|
||||
``model_type/models`` is a source-only ``../models`` symlink that packaging prunes, so
|
||||
without the loader alias the path would resolve in a checkout but 404 from a built wheel.
|
||||
The alias rewrites the prefix to ``models/`` so both behave identically.
|
||||
"""
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
sample = _first_builtin_ptq_recipe(root, "models/*/*/ptq/*.yaml")
|
||||
canonical = str(sample.relative_to(root).with_suffix("")) # models/<org>/<model>/ptq/<file>
|
||||
aliased = "model_type/" + canonical # model_type/models/<org>/<model>/ptq/<file>
|
||||
|
||||
with pytest.warns(FutureWarning, match="deprecated recipe-tier prefix"):
|
||||
recipe = load_recipe(aliased)
|
||||
assert str(_resolve_recipe_path(aliased)) == str(_resolve_recipe_path(canonical))
|
||||
assert recipe.recipe_type == RecipeType.PTQ
|
||||
assert isinstance(recipe, ModelOptPTQRecipe)
|
||||
|
||||
|
||||
def test_load_recipe_local_tree_overrides_builtin_even_on_name_collision(tmp_path, monkeypatch):
|
||||
"""A local recipe tree overrides a built-in of the same name — even when the name
|
||||
collides with a shipped ``model_type``.
|
||||
|
||||
``_resolve_recipe_path`` probes the filesystem before the built-in library (matching
|
||||
``config_loader._resolve_config_path``), so a user who keeps their own recipe tree on disk
|
||||
is never silently shadowed by the deprecated-tier alias. This uses a *shipped* recipe's
|
||||
exact relative path, spelled with the old ``huggingface/`` prefix that aliases to it, to
|
||||
prove the local file wins over the built-in — the collision case a non-shipped name misses.
|
||||
"""
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
shipped = _first_builtin_ptq_recipe(root, "model_type/*/ptq/*.yaml")
|
||||
# The path a user would keep locally: the shipped recipe's own relative path, but under the
|
||||
# deprecated ``huggingface/`` tier that the alias rewrites to ``model_type/``.
|
||||
old_rel = Path("huggingface") / shipped.relative_to(root).relative_to("model_type")
|
||||
|
||||
local = tmp_path / old_rel
|
||||
local.parent.mkdir(parents=True)
|
||||
local.write_text(
|
||||
"metadata:\n recipe_type: ptq\nquantize:\n quant_cfg: {}\n algorithm: max\n"
|
||||
)
|
||||
monkeypatch.chdir(tmp_path)
|
||||
|
||||
resolved = _resolve_recipe_path(str(old_rel.with_suffix("")))
|
||||
assert Path(resolved).resolve() == local.resolve()
|
||||
|
||||
|
||||
def test_import_resolution_honors_huggingface_alias():
|
||||
"""``$import`` resolution rewrites deprecated tier prefixes just like ``load_recipe``.
|
||||
|
||||
``$import`` paths go through ``config_loader._resolve_config_path`` (not the recipe-path
|
||||
alias), so a custom recipe that imports a shipped snippet by its old ``huggingface/...``
|
||||
path must still resolve from a wheel where the ``huggingface`` symlink is gone.
|
||||
"""
|
||||
# Prefix-rewrite mapping: architecture rename plus both checkpoint-mirror aliases.
|
||||
assert _alias_builtin_recipe_prefix("huggingface/qwen3_vl/ptq/x") == "model_type/qwen3_vl/ptq/x"
|
||||
assert (
|
||||
_alias_builtin_recipe_prefix("huggingface/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
|
||||
)
|
||||
assert (
|
||||
_alias_builtin_recipe_prefix("model_type/models/nvidia/m/ptq/x") == "models/nvidia/m/ptq/x"
|
||||
)
|
||||
assert _alias_builtin_recipe_prefix("general/ptq/x") == "general/ptq/x" # untouched
|
||||
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
sample = next(root.glob("model_type/*/ptq/*.yaml"))
|
||||
canonical = str(sample.relative_to(root).with_suffix("")) # model_type/<arch>/ptq/<file>
|
||||
old = "huggingface/" + canonical[len("model_type/") :] # huggingface/<arch>/ptq/<file>
|
||||
|
||||
assert str(_resolve_config_path(old)) == str(_resolve_config_path(canonical))
|
||||
|
||||
|
||||
def _all_shipped_ptq_recipe_paths():
|
||||
"""Every shipped PTQ recipe, discovered from disk rather than a hardcoded list."""
|
||||
root = files("modelopt_recipes")
|
||||
@@ -201,7 +320,7 @@ def _all_shipped_ptq_recipe_paths():
|
||||
|
||||
|
||||
# Discovered from disk (not hardcoded) so the smoke tests cover every shipped PTQ
|
||||
# recipe — general/, huggingface/<model_type>/, and models/<org>/<model_id>/ — and
|
||||
# recipe — general/, model_type/<model_type>/, and models/<org>/<model_id>/ — and
|
||||
# never drift as recipes are added, moved, or removed.
|
||||
_BUILTIN_PTQ_RECIPES = _all_shipped_ptq_recipe_paths()
|
||||
|
||||
@@ -1905,10 +2024,10 @@ def test_load_recipe_autoquantize_builtin_active_moe():
|
||||
def test_load_recipe_autoquantize_module_search_spaces():
|
||||
"""Qwen recipe separates its fixed PTQ baseline from explicit search spaces."""
|
||||
recipe = load_recipe(
|
||||
"huggingface/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
|
||||
"model_type/qwen3_6_moe/auto_quantize/w4a16_nvfp4_fp8_module_spaces_at_6p0bits-active_moe"
|
||||
)
|
||||
aq = recipe.auto_quantize
|
||||
model_ptq = load_recipe("huggingface/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
|
||||
model_ptq = load_recipe("model_type/qwen3_5_moe/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast")
|
||||
assert recipe.quantize is not None
|
||||
assert recipe.quantize == model_ptq.quantize
|
||||
assert aq.candidate_formats == []
|
||||
|
||||
@@ -61,7 +61,7 @@ class _MiniMaxModel(nn.Module):
|
||||
def test_mxfp8_nvfp4_experts_recipe_quantizer_precedence():
|
||||
model = _MiniMaxModel()
|
||||
register_fused_experts_on_the_fly(model)
|
||||
recipe = load_recipe("huggingface/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
|
||||
recipe = load_recipe("model_type/minimax_m3_vl/ptq/mxfp8_nvfp4_experts")
|
||||
config = recipe.quantize.model_dump()
|
||||
assert config["algorithm"]["layerwise"]["enable"] is True
|
||||
config["algorithm"] = None
|
||||
|
||||
@@ -45,7 +45,7 @@ def test_qwen_vision_recipes_select_expected_quantizers(
|
||||
):
|
||||
model = _get_tiny_qwen_vlm(model_type)
|
||||
assert model.config.model_type == model_type
|
||||
quant_cfg = load_recipe(f"huggingface/{model_type}/ptq/{recipe}").quantize.model_dump()
|
||||
quant_cfg = load_recipe(f"model_type/{model_type}/ptq/{recipe}").quantize.model_dump()
|
||||
|
||||
mtq.quantize(model, quant_cfg, forward_loop=None)
|
||||
modules = dict(model.named_modules())
|
||||
|
||||
@@ -89,22 +89,22 @@ def test_general_ptq_recipe_count_in_ptq_md():
|
||||
def test_every_model_specific_ptq_dir_is_mentioned():
|
||||
"""Every model-specific PTQ recipe must be identifiable in ptq.md.
|
||||
|
||||
``huggingface/<model_type>/ptq/`` recipes are checked by their ``model_type``
|
||||
``model_type/<model_type>/ptq/`` recipes are checked by their ``model_type``
|
||||
(e.g. ``gemma4``); ``models/<org>/<model_id>/ptq/`` recipes are checked by their
|
||||
full ``<org>/<model_id>`` hub path (e.g. ``nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16``), so
|
||||
the org — the whole point of the top-level tier — is verified too and an org
|
||||
re-key (e.g. ``step3p5`` → ``stepfun-ai``) can't silently drift from the doc.
|
||||
"""
|
||||
doc = _ptq_md_text()
|
||||
# model_type recipes: huggingface/<model_type>/ptq/<recipe>.yaml -> <model_type>
|
||||
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "huggingface").glob("**/ptq/*.yaml")}
|
||||
# model_type recipes: model_type/<model_type>/ptq/<recipe>.yaml -> <model_type>
|
||||
hf_ids = {p.parent.parent.name for p in (RECIPES_DIR / "model_type").glob("**/ptq/*.yaml")}
|
||||
# checkpoint recipes: models/<org>/<model_id>/ptq/<recipe>.yaml -> <org>/<model_id>
|
||||
model_ids = {
|
||||
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
|
||||
for p in (RECIPES_DIR / "models").glob("**/ptq/*.yaml")
|
||||
}
|
||||
identifiers = sorted(hf_ids | model_ids)
|
||||
assert identifiers, "No model-specific PTQ recipes found under huggingface/ or models/"
|
||||
assert identifiers, "No model-specific PTQ recipes found under model_type/ or models/"
|
||||
missing = [name for name in identifiers if name not in doc]
|
||||
assert not missing, (
|
||||
f"Model-specific PTQ recipe folders are missing from "
|
||||
@@ -114,40 +114,54 @@ def test_every_model_specific_ptq_dir_is_mentioned():
|
||||
|
||||
|
||||
def test_checkpoint_recipes_live_in_the_top_level_models_tier():
|
||||
"""Lock in the model_type-vs-checkpoint split.
|
||||
"""Lock in the model_type-vs-checkpoint split and the backward-compat symlinks.
|
||||
|
||||
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``huggingface/``
|
||||
holds only per-``model_type`` recipes. ``huggingface/models`` is kept as a
|
||||
backward-compatibility **symlink** to the top-level ``models/`` tier, so the old
|
||||
``--recipe huggingface/models/<org>/<model_id>/...`` paths still resolve; it must
|
||||
stay a symlink that points at ``../models`` and never become a real directory that
|
||||
holds recipes. A checkpoint recipe nested under a ``model_type`` (e.g.
|
||||
``huggingface/<model_type>/<checkpoint>/<task>/``) still fails loudly here instead
|
||||
of silently shipping both tiers — e.g. on a bad merge that re-adds the old layout.
|
||||
Checkpoint-mirror recipes belong at ``models/<org>/<model_id>/``; ``model_type/``
|
||||
(formerly ``huggingface/``) holds only per-``model_type`` recipes. Two
|
||||
backward-compatibility **symlinks** are kept so old ``--recipe`` paths still
|
||||
resolve: the top-level ``huggingface`` -> ``model_type`` rename alias, and the
|
||||
nested ``model_type/models`` -> ``../models`` alias for the old
|
||||
``huggingface/models/<org>/<model_id>/...`` checkpoint paths. Both must stay
|
||||
symlinks and never become real directories that hold recipes. A checkpoint recipe
|
||||
nested under a ``model_type`` (e.g. ``model_type/<model_type>/<checkpoint>/<task>/``)
|
||||
still fails loudly here instead of silently shipping both tiers — e.g. on a bad
|
||||
merge that re-adds the old layout.
|
||||
"""
|
||||
hf = RECIPES_DIR / "huggingface"
|
||||
model_type = RECIPES_DIR / "model_type"
|
||||
models = RECIPES_DIR / "models"
|
||||
hf_models = hf / "models"
|
||||
assert hf_models.is_symlink(), (
|
||||
"huggingface/models must be a symlink to the top-level modelopt_recipes/models/ "
|
||||
"tier (a backward-compat alias for the old --recipe paths), not a real directory."
|
||||
hf_alias = RECIPES_DIR / "huggingface"
|
||||
mt_models = model_type / "models"
|
||||
# Top-level huggingface -> model_type rename alias.
|
||||
assert hf_alias.is_symlink(), (
|
||||
"modelopt_recipes/huggingface must be a backward-compat symlink to model_type/ "
|
||||
"(the rename alias), not a real directory."
|
||||
)
|
||||
assert hf_models.resolve() == models.resolve(), (
|
||||
f"huggingface/models must resolve to the top-level models/ tier; resolves to "
|
||||
f"{hf_models.resolve()} instead of {models.resolve()}."
|
||||
assert hf_alias.resolve() == model_type.resolve(), (
|
||||
f"huggingface must resolve to the model_type/ tier; resolves to "
|
||||
f"{hf_alias.resolve()} instead of {model_type.resolve()}."
|
||||
)
|
||||
# Every recipe under huggingface/ must be <model_type>/<task>/<file> (3 parts);
|
||||
# Nested model_type/models -> ../models alias for the old huggingface/models/... paths.
|
||||
assert mt_models.is_symlink(), (
|
||||
"model_type/models must be a symlink to the top-level modelopt_recipes/models/ "
|
||||
"tier (a backward-compat alias for the old --recipe huggingface/models/... paths), "
|
||||
"not a real directory."
|
||||
)
|
||||
assert mt_models.resolve() == models.resolve(), (
|
||||
f"model_type/models must resolve to the top-level models/ tier; resolves to "
|
||||
f"{mt_models.resolve()} instead of {models.resolve()}."
|
||||
)
|
||||
# Every recipe under model_type/ must be <model_type>/<task>/<file> (3 parts);
|
||||
# anything deeper is a checkpoint nested under a model_type and belongs in models/.
|
||||
# Skip the huggingface/models symlink so the models/ recipes it aliases (4 parts)
|
||||
# Skip the model_type/models symlink so the models/ recipes it aliases (4 parts)
|
||||
# aren't miscounted as nested here.
|
||||
nested = sorted(
|
||||
str(p.relative_to(RECIPES_DIR))
|
||||
for ext in ("*.yaml", "*.yml")
|
||||
for p in hf.glob(f"**/{ext}")
|
||||
if hf_models not in p.parents and len(p.relative_to(hf).parts) != 3
|
||||
for p in model_type.glob(f"**/{ext}")
|
||||
if mt_models not in p.parents and len(p.relative_to(model_type).parts) != 3
|
||||
)
|
||||
assert not nested, (
|
||||
f"Recipes under huggingface/ must be <model_type>/<task>/<file>; found nested "
|
||||
f"Recipes under model_type/ must be <model_type>/<task>/<file>; found nested "
|
||||
f"paths (a checkpoint recipe belongs under models/<org>/<model_id>/): {nested}"
|
||||
)
|
||||
# Every recipe under models/ must be <org>/<model_id>/<task>/<file> (4 parts) so the
|
||||
@@ -181,7 +195,7 @@ def test_launcher_yaml_recipe_paths_resolve():
|
||||
# ``--recipe <p>`` / ``QUANT_CFG: <p>`` are modelopt_recipes-relative — only tier-prefixed
|
||||
# values are recipe paths; bare names like ``auto`` or ``FP8_DEFAULT_CFG`` are not. The
|
||||
# ``modelopt_recipes/<p>.yaml`` form (e.g. ``--config``) embeds the path directly.
|
||||
tier = r"(?:general|huggingface|models|configs)/[A-Za-z0-9._/-]+"
|
||||
tier = r"(?:general|model_type|models|configs)/[A-Za-z0-9._/-]+"
|
||||
rel_re = re.compile(rf"(?:--recipe\s+|QUANT_CFG:\s*)({tier})")
|
||||
abs_re = re.compile(rf"modelopt_recipes/({tier}\.ya?ml)")
|
||||
|
||||
|
||||
@@ -121,8 +121,8 @@ def _enabled(module, quantizer="weight_quantizer"):
|
||||
@pytest.mark.parametrize(
|
||||
"recipe_name",
|
||||
[
|
||||
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
],
|
||||
)
|
||||
def test_routed_experts_are_quantized(recipe_name):
|
||||
@@ -144,8 +144,8 @@ def test_routed_experts_are_quantized(recipe_name):
|
||||
@pytest.mark.parametrize(
|
||||
"recipe_name",
|
||||
[
|
||||
"huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
"model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast",
|
||||
"model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8",
|
||||
],
|
||||
)
|
||||
def test_router_shared_expert_and_head_stay_bf16(recipe_name):
|
||||
@@ -160,7 +160,7 @@ def test_router_shared_expert_and_head_stay_bf16(recipe_name):
|
||||
|
||||
|
||||
def test_experts_only_leaves_dense_mlp_bf16():
|
||||
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
|
||||
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast")
|
||||
dense_mlp = model.model.language_model.layers[1].mlp
|
||||
|
||||
for proj in ("gate_proj", "up_proj", "down_proj"):
|
||||
@@ -168,7 +168,7 @@ def test_experts_only_leaves_dense_mlp_bf16():
|
||||
|
||||
|
||||
def test_mlp_only_also_quantizes_dense_mlp():
|
||||
model = _quantize_with_recipe("huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
|
||||
model = _quantize_with_recipe("model_type/step3p7/ptq/nvfp4_mlp_only-kv_fp8")
|
||||
dense_mlp = model.model.language_model.layers[1].mlp
|
||||
|
||||
for proj in ("gate_proj", "up_proj", "down_proj"):
|
||||
|
||||
@@ -40,6 +40,26 @@ from modelopt.torch.export.quant_aware_conversion import (
|
||||
|
||||
BLOCK = 16
|
||||
|
||||
|
||||
def _set_scope_attr(transform, name, value):
|
||||
"""Set an optional scoped-match attribute that only some transformers versions expose.
|
||||
|
||||
transformers>=5.9 dropped ``base_model_prefix`` from ``WeightTransform``'s ``__slots__``
|
||||
(scoped matching now keys off ``scope_prefix`` alone); older supported versions still
|
||||
carry it. Production ``_scope_prefixes`` reads it via ``getattr(..., None)``, so skipping
|
||||
the assignment where the slot is absent is equivalent — and lets these tests run across
|
||||
the whole supported transformers range instead of ``AttributeError``-ing on the setattr.
|
||||
|
||||
The suppression is scoped to that one known version-dependent slot: a setattr failure for
|
||||
any other name (a typo or a future rename) still raises instead of silently no-op-ing.
|
||||
"""
|
||||
try:
|
||||
setattr(transform, name, value)
|
||||
except AttributeError:
|
||||
if name != "base_model_prefix":
|
||||
raise
|
||||
|
||||
|
||||
# Tiny Mixtral shaped to match the synthetic expert tensors built by ``_nvfp4_linear`` below.
|
||||
_MIXTRAL_KWARGS = {
|
||||
"hidden_size": 32,
|
||||
@@ -341,7 +361,7 @@ def test_scoped_submodel_prefix_change_does_not_capture_siblings():
|
||||
|
||||
prefix_change = PrefixChange(prefix_to_remove="vision_model")
|
||||
prefix_change.scope_prefix = "model.vision_tower"
|
||||
prefix_change.base_model_prefix = "model"
|
||||
_set_scope_attr(prefix_change, "base_model_prefix", "model")
|
||||
model._weight_conversions = [prefix_change]
|
||||
|
||||
state_dict = {
|
||||
@@ -381,7 +401,7 @@ def test_scoped_rule_maps_config_module_names_consistently():
|
||||
|
||||
prefix_change = PrefixChange(prefix_to_remove="vision_model")
|
||||
prefix_change.scope_prefix = "model.vision_tower"
|
||||
prefix_change.base_model_prefix = "model"
|
||||
_set_scope_attr(prefix_change, "base_model_prefix", "model")
|
||||
model._weight_conversions = [prefix_change]
|
||||
|
||||
mapper = build_reverse_name_mapper(model)
|
||||
@@ -420,7 +440,7 @@ def test_root_scoped_rule_still_faces_shadowing_guard():
|
||||
)
|
||||
# Root scope: reaches every key, exactly like an unscoped rule.
|
||||
renaming.scope_prefix = ""
|
||||
renaming.base_model_prefix = ""
|
||||
_set_scope_attr(renaming, "base_model_prefix", "")
|
||||
model._weight_conversions = [renaming]
|
||||
|
||||
state_dict = {
|
||||
@@ -454,7 +474,7 @@ def test_scoped_weight_converter_is_refused():
|
||||
operations=[Chunk(dim=0)],
|
||||
)
|
||||
conv.scope_prefix = "model.language_model"
|
||||
conv.base_model_prefix = "model"
|
||||
_set_scope_attr(conv, "base_model_prefix", "model")
|
||||
model = types.SimpleNamespace(_weight_conversions=[conv])
|
||||
|
||||
sd = _nvfp4_linear("model.language_model.layers.0.mlp.gate_up_proj", 8, 16)
|
||||
|
||||
Reference in New Issue
Block a user