mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
Reuse a whole recipe via $import, deprecate recipe_type, and start the published-checkpoint backfill with two aliases (#2376)
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
051d6adb20
commit
d0142c9dca
@@ -0,0 +1,28 @@
|
||||
name: Recipe Hub Org Check
|
||||
|
||||
on:
|
||||
pull_request:
|
||||
branches: [main, release/*, feature/*]
|
||||
paths:
|
||||
- "modelopt_recipes/models/**"
|
||||
- "tools/ci/check_model_hub_orgs.py"
|
||||
- ".github/workflows/check_model_hub_orgs.yml"
|
||||
workflow_dispatch:
|
||||
# On-demand
|
||||
|
||||
concurrency:
|
||||
# Cancel previous runs if new commit is pushed to the same PR
|
||||
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.sha }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
check-model-hub-orgs:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
steps:
|
||||
- uses: actions/checkout@v6
|
||||
- uses: ./.github/actions/ubuntu-setup
|
||||
- name: Verify models/<org>/<model_id> entries against the Hugging Face Hub
|
||||
run: |
|
||||
pip install huggingface_hub
|
||||
python tools/ci/check_model_hub_orgs.py
|
||||
@@ -13,6 +13,8 @@ Changelog
|
||||
|
||||
*Quantization*
|
||||
|
||||
- A recipe can now **delegate its whole body to another recipe** with a top-level ``$import``; any top-level key given alongside it overrides the imported one. ``metadata.recipe_type`` became optional along with it: a recipe states its kind with a ``# modelopt-schema:`` comment, with ``metadata.recipe_type``, or by delegating to a recipe that does, and only a recipe that another file imports has to carry the schema comment. Whatever a recipe does state must be true: a schema comment and a ``recipe_type`` must agree, and so must a recipe and the recipe it delegates to. ``modelopt_recipes/models/`` uses this for checkpoint entries that a portable recipe already reproduces: the entry aliases that recipe instead of copying it.
|
||||
- Backfill the recipes behind NVIDIA's already-published checkpoints under ``modelopt_recipes/models/``, so a released checkpoint's quantization scheme is reachable from its own model-hub path rather than only from the general tier. For example, ``moonshotai/Kimi-K2.6`` (published as ``nvidia/Kimi-K2.6-NVFP4``) and ``Qwen/Qwen3.5-397B-A17B`` (published as ``nvidia/Qwen3.5-397B-A17B-NVFP4-V2``) each alias a portable recipe wholesale -- the general expert-only NVFP4 recipe and the ``qwen3_5_moe`` architecture recipe respectively -- rather than copying its body; other checkpoints follow in separate changes.
|
||||
- Add ``layerwise.export_dir``: layerwise calibration writes each decoder layer to its own quantized checkpoint shard as it finishes, so no separate ``export_hf_checkpoint()`` pass is needed and, with ``layerwise.checkpoint_dir``, an interrupted run resumes without redoing finished layers. Calibration writes the layer shards; ``finalize()`` on the exporter left on the model adds the tail shard, the index and the config artifacts, and the checkpoint does not load until it runs. ``examples/hf_ptq`` does this for you. Supports FP8 and NVFP4 on single-process models, resident or offloaded, including multimodal models and models with MTP layers; other formats and placements raise ``NotImplementedError`` before calibration starts.
|
||||
- Add support for quantizing and calibrating enabled operators outside the transformer layers, such as ``lm_head``, when using layerwise calibration.
|
||||
- Add an end-to-end BEVFormer ONNX PTQ example with temporal calibration data generation, INT8 and FP8 quantization, TensorRT engine building, and nuScenes accuracy evaluation. See `examples/onnx_ptq/bevformer/README.md <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/onnx_ptq/bevformer>`_ for details.
|
||||
@@ -68,6 +70,7 @@ Changelog
|
||||
- The single-format quantization CLI flags are deprecated in favour of ``--recipe`` and will be removed in a future release; passing one now emits a ``FutureWarning``. ``examples/hf_ptq``: ``--qformat`` and ``--kv_cache_qformat``. ``examples/megatron_bridge/quantize.py``: ``--quant_cfg``, ``--kv_cache_quant`` and ``--weight_only``. ``examples/torch_onnx/torch_quant_to_onnx.py``: ``--qformat``. A recipe carries the quantization config, the calibration algorithm and the KV-cache setting in one file, so they cannot drift apart the way separate flags can -- and ``--recipe`` already took precedence over all six, silently on ``hf_ptq`` and with a warning on ``megatron_bridge`` -- with one gap the recipe closes rather than inherits: a weight AutoQuantize recipe that omits ``kv_cache`` still falls back to ``--kv_cache_qformat``, so set ``kv_cache`` in the recipe when migrating. Use a recipe from ``modelopt_recipes/general/ptq/``, an architecture-specific one under ``modelopt_recipes/model_type/<model_type>/``, or a checkpoint-specific one under ``modelopt_recipes/models/``. The warning fires only when a flag is passed explicitly: ``--qformat`` defaults to ``fp8`` and ``--kv_cache_qformat`` to ``fp8_cast``, so warning on the defaults would fire on every run, including runs that correctly use ``--recipe``. ``examples/speculative_decoding/scripts/quantize_drafter.py`` keeps ``--qformat`` undeprecated: it has no ``--recipe`` alternative yet.
|
||||
- The TensorRT-LLM checkpoint export format is deprecated and will be removed in 0.49.0: ``export_tensorrt_llm_checkpoint`` and ``torch_to_tensorrt_llm_checkpoint`` now emit a ``DeprecationWarning`` on use. Use ``export_hf_checkpoint``, which exports a unified Hugging Face checkpoint deployable on TensorRT-LLM, vLLM and SGLang. Its implementation moved to ``modelopt.torch.export.trtllm``, so import those two functions from there and the ``ModelConfig`` dataclasses from ``modelopt.torch.export.trtllm.model_config``; both functions remain importable from ``modelopt.torch.export`` for this release only.
|
||||
- The ``evaluation`` agent skill no longer supports GDPVal: its task recipe, example config and Apptainer SIF helper are removed, and the shared NeMo Gym machinery it carried now lives at ``references/gym.md`` with the launcher wrapper renamed ``scripts/nel-gym.sh``. GDPVal is an AA-suite member, so an "AA" request now generates the ``aa/`` tasks only -- report per-task scores rather than an aggregate compared against a published AA Index.
|
||||
- Deprecate ``metadata.recipe_type`` in recipe YAML, for both single-file recipes and a directory recipe's ``metadata.yml``. A recipe now says what kind it is with a ``# modelopt-schema:`` comment naming its schema class, or by delegating to a recipe that does (single-file recipes only -- a directory recipe has no body to delegate through); ``recipe_type`` is still read and still honoured, so a recipe outside this repo keeps working unchanged, but new recipes should leave it out and every recipe shipped here has been converted. Where both are present they must agree, and so must a recipe and the recipe it delegates to -- a disagreement is an error rather than a silent preference.
|
||||
|
||||
**Bug Fixes**
|
||||
|
||||
|
||||
@@ -380,7 +380,9 @@ Reusable snippets are stored under ``modelopt_recipes/configs/``:
|
||||
Metadata section
|
||||
================
|
||||
|
||||
Every recipe must contain a ``metadata`` mapping with at least a ``recipe_type`` field:
|
||||
Every recipe contains a ``metadata`` mapping. ``recipe_type`` is optional and
|
||||
**deprecated** -- a recipe may instead declare its kind with a ``# modelopt-schema:``
|
||||
comment, or inherit it from the recipe it delegates to; see `Declaring a recipe's kind`_:
|
||||
|
||||
.. list-table::
|
||||
:header-rows: 1
|
||||
@@ -390,7 +392,7 @@ Every recipe must contain a ``metadata`` mapping with at least a ``recipe_type``
|
||||
- Required
|
||||
- Description
|
||||
* - ``recipe_type``
|
||||
- Yes
|
||||
- No
|
||||
- The optimization category. Determines which configuration sections are
|
||||
expected (e.g., ``"ptq"`` expects a ``quantize`` section). See
|
||||
:class:`~modelopt.recipe.config.RecipeType` for supported values.
|
||||
@@ -403,7 +405,7 @@ Type-specific configuration sections
|
||||
=====================================
|
||||
|
||||
Each recipe type defines its own configuration section. The section name and
|
||||
schema depend on the ``recipe_type`` value in the metadata.
|
||||
schema depend on the recipe's kind, however it is declared.
|
||||
|
||||
PTQ (``recipe_type: ptq``)
|
||||
--------------------------
|
||||
@@ -516,18 +518,28 @@ General PTQ recipes are model-agnostic and apply to any supported architecture:
|
||||
* - ``general/ptq/nvfp4_weight_only-kv_fp8_cast``
|
||||
- NVFP4 W4A16 weight-only, FP8 KV cache with constant amax
|
||||
|
||||
See `modelopt_recipes/ptq.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/ptq.md>`_
|
||||
for the full list and for guidance on choosing between them.
|
||||
|
||||
Model-specific recipes
|
||||
----------------------
|
||||
|
||||
Model-specific recipes come in two tiers: architecture recipes keyed by a
|
||||
Hugging Face ``model_type`` under ``model_type/<model_type>/<task>/``, and
|
||||
checkpoint mirrors keyed by a model-hub path under
|
||||
checkpoint entries keyed by a model-hub path under
|
||||
``models/<org>/<model_id>/<task>/``. See
|
||||
`modelopt_recipes/model_type/README.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/model_type/README.md>`_
|
||||
and
|
||||
`modelopt_recipes/models/README.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/models/README.md>`_
|
||||
for the layout conventions and recipe-lookup order.
|
||||
|
||||
A checkpoint entry comes in two forms. A **mirror** carries its own body, because the
|
||||
release uses a per-layer scheme no portable recipe expresses. An **alias** has no body of
|
||||
its own: a general or architecture recipe already produces that scheme, so the entry
|
||||
imports that recipe wholesale and exists to make it findable from the checkpoint's hub
|
||||
path. See `modelopt_recipes/ptq.md <https://github.com/NVIDIA/Model-Optimizer/blob/main/modelopt_recipes/ptq.md>`_
|
||||
for what each checkpoint entry does.
|
||||
|
||||
.. note::
|
||||
|
||||
``model_type/`` was previously named ``huggingface/``. Old
|
||||
@@ -547,6 +559,44 @@ for the layout conventions and recipe-lookup order.
|
||||
- MXFP8 language-model base with MSE-calibrated NVFP4 routed experts for MiniMax-M3
|
||||
|
||||
|
||||
Delegating to another recipe
|
||||
----------------------------
|
||||
|
||||
A recipe can hand its whole body to another recipe with a top-level ``$import`` and keep
|
||||
only its own ``metadata``. Keys given alongside the ``$import`` override the imported
|
||||
ones, so the body -- ``quantize``, its algorithm and every ``quant_cfg`` entry -- is
|
||||
inherited unchanged:
|
||||
|
||||
.. code-block:: yaml
|
||||
|
||||
imports:
|
||||
base: general/ptq/nvfp4_default-kv_fp8_cast
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: What this checkpoint uses the base recipe for.
|
||||
|
||||
.. _Declaring a recipe's kind:
|
||||
|
||||
Note the missing ``recipe_type``. A recipe states its kind in whichever of these it
|
||||
likes, and the loader takes the first that answers: a ``# modelopt-schema:`` comment
|
||||
naming its schema class, ``metadata.recipe_type``, or -- as here -- the recipe it
|
||||
delegates to. ``metadata.recipe_type`` is **deprecated**: it is still read and still
|
||||
honoured, so no existing recipe needs changing, but new recipes should declare their
|
||||
schema instead. A directory-format recipe's ``metadata.yml`` can declare a schema
|
||||
comment too, the same way; delegation is the one source that does not apply to it,
|
||||
since a directory recipe has no top-level ``$import`` to inherit a kind from.
|
||||
|
||||
Stating more than one is allowed, but they must agree, and that extends across the
|
||||
import: a recipe and the recipe it delegates to must be the same kind, since the import
|
||||
takes over the whole body. A PTQ recipe importing an EAGLE one is rejected as a kind
|
||||
mismatch rather than left to fail on whatever the spliced sections do to the schema.
|
||||
|
||||
What *is* required: a recipe that another file imports must carry the schema comment,
|
||||
since ``$import`` resolution needs it to validate the imported payload. A recipe nothing
|
||||
imports needs no comment at all. This is how the checkpoint aliases under ``models/``
|
||||
name the recipe behind a release without duplicating its body.
|
||||
|
||||
Loading recipes
|
||||
===============
|
||||
|
||||
@@ -555,7 +605,7 @@ Python API
|
||||
|
||||
Use :func:`~modelopt.recipe.load_recipe` to load a recipe. The path is resolved
|
||||
against the built-in library first, then the filesystem. The returned object's
|
||||
type depends on the ``recipe_type`` in the metadata:
|
||||
type depends on the recipe's kind, however it is declared:
|
||||
|
||||
.. code-block:: python
|
||||
|
||||
|
||||
+55
-11
@@ -19,7 +19,7 @@ from __future__ import annotations
|
||||
|
||||
import warnings
|
||||
from enum import Enum
|
||||
from typing import Literal
|
||||
from typing import ClassVar, Literal
|
||||
|
||||
from pydantic import Field, field_validator, model_validator
|
||||
|
||||
@@ -67,9 +67,16 @@ _DEFAULT_RECIPE_DESCRIPTION = "Model optimization recipe."
|
||||
class RecipeMetadataConfig(ModeloptBaseConfig):
|
||||
"""YAML shape of the recipe metadata section."""
|
||||
|
||||
recipe_type: RecipeType = Field(
|
||||
recipe_type: RecipeType | None = ModeloptField(
|
||||
default=None,
|
||||
title="Recipe type",
|
||||
description="The type of the recipe (e.g. PTQ).",
|
||||
description="The type of the recipe (e.g. PTQ). **Deprecated** in recipe YAML: "
|
||||
"the ``# modelopt-schema:`` comment naming the recipe's schema class already says "
|
||||
"which kind it is -- and it is the same declaration that makes the file "
|
||||
"``$import``-able -- so the class fills this in. Still read and still honoured, so "
|
||||
"no existing recipe needs changing, but new recipes should leave it out -- including "
|
||||
"in a directory-format recipe's ``metadata.yml``, which supports the same comment. "
|
||||
"When both are present they must agree.",
|
||||
)
|
||||
description: str = ModeloptField(
|
||||
default=_DEFAULT_RECIPE_DESCRIPTION,
|
||||
@@ -78,10 +85,10 @@ class RecipeMetadataConfig(ModeloptBaseConfig):
|
||||
)
|
||||
|
||||
|
||||
def _metadata_field(recipe_type: RecipeType):
|
||||
"""Build the metadata Pydantic field with the recipe_type baked into the default."""
|
||||
def _metadata_field():
|
||||
"""Build a metadata Pydantic field that defaults to the owning class's recipe type."""
|
||||
return ModeloptField(
|
||||
default={"recipe_type": recipe_type, "description": _DEFAULT_RECIPE_DESCRIPTION},
|
||||
default={"description": _DEFAULT_RECIPE_DESCRIPTION},
|
||||
title="Metadata",
|
||||
description="Recipe metadata containing the recipe type and description.",
|
||||
validate_default=True,
|
||||
@@ -94,16 +101,43 @@ class ModelOptRecipeBase(ModeloptBaseConfig):
|
||||
If a layer name matches ``"*output_layer*"``, the attributes will be replaced with ``{"enable": False}``.
|
||||
"""
|
||||
|
||||
#: The kind of recipe this class *is*. Set on every concrete subclass; it is the
|
||||
#: single source of truth for ``metadata.recipe_type``, which the validator below
|
||||
#: fills in so a recipe file never has to repeat what its schema already states.
|
||||
RECIPE_TYPE: ClassVar[RecipeType | None] = None
|
||||
|
||||
metadata: RecipeMetadataConfig = Field(
|
||||
title="Metadata",
|
||||
description="Recipe metadata containing the recipe type and description. "
|
||||
"Required: a recipe without a ``metadata`` section is rejected so that a "
|
||||
"missing section can't silently fall back to a default recipe type.",
|
||||
"recipe always says what it is for.",
|
||||
)
|
||||
|
||||
@model_validator(mode="after")
|
||||
def _resolve_recipe_type(self):
|
||||
"""Fill ``metadata.recipe_type`` from the schema class, or reject a mismatch.
|
||||
|
||||
The schema class already determines the kind, so a recipe file that declares its
|
||||
schema needs no ``recipe_type``. One that states it anyway must state the truth --
|
||||
a silent disagreement between the two would make the file mean different things
|
||||
to the loader and to a reader.
|
||||
"""
|
||||
if self.RECIPE_TYPE is None:
|
||||
return self
|
||||
if self.metadata.recipe_type is None:
|
||||
self.metadata.recipe_type = self.RECIPE_TYPE
|
||||
elif self.metadata.recipe_type != self.RECIPE_TYPE:
|
||||
raise ValueError(
|
||||
f"metadata.recipe_type is {self.metadata.recipe_type.value!r} but this recipe "
|
||||
f"is a {type(self).__name__}, which is {self.RECIPE_TYPE.value!r}. Drop the "
|
||||
"recipe_type (the schema declares it) or correct it."
|
||||
)
|
||||
return self
|
||||
|
||||
@property
|
||||
def recipe_type(self) -> RecipeType:
|
||||
"""Return the recipe type from metadata."""
|
||||
assert self.metadata.recipe_type is not None, "recipe_type was not resolved"
|
||||
return self.metadata.recipe_type
|
||||
|
||||
@property
|
||||
@@ -115,6 +149,8 @@ class ModelOptRecipeBase(ModeloptBaseConfig):
|
||||
class ModelOptPTQRecipe(ModelOptRecipeBase):
|
||||
"""Our config class for PTQ recipes."""
|
||||
|
||||
RECIPE_TYPE: ClassVar[RecipeType] = RecipeType.PTQ
|
||||
|
||||
quantize: QuantizeConfig = Field(
|
||||
title="PTQ config",
|
||||
description="PTQ config containing quant_cfg and algorithm. Required: a PTQ "
|
||||
@@ -308,7 +344,9 @@ class AutoQuantizeConfig(ModeloptBaseConfig):
|
||||
class ModelOptAutoQuantizeRecipe(ModelOptRecipeBase):
|
||||
"""Our config class for AutoQuantize recipes."""
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field(RecipeType.AUTO_QUANTIZE)
|
||||
RECIPE_TYPE: ClassVar[RecipeType] = RecipeType.AUTO_QUANTIZE
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field()
|
||||
|
||||
quantize: QuantizeConfig | None = ModeloptField(
|
||||
default=None,
|
||||
@@ -380,7 +418,9 @@ class ModelOptSpeculativeRecipeBase(ModelOptRecipeBase):
|
||||
class ModelOptEagleRecipe(ModelOptSpeculativeRecipeBase):
|
||||
"""Our config class for EAGLE speculative decoding recipes."""
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field(RecipeType.SPECULATIVE_EAGLE)
|
||||
RECIPE_TYPE: ClassVar[RecipeType] = RecipeType.SPECULATIVE_EAGLE
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field()
|
||||
|
||||
eagle: EagleConfig = ModeloptField(
|
||||
default=EagleConfig(),
|
||||
@@ -409,7 +449,9 @@ class ModelOptEagleRecipe(ModelOptSpeculativeRecipeBase):
|
||||
class ModelOptDFlashRecipe(ModelOptSpeculativeRecipeBase):
|
||||
"""Our config class for DFlash speculative decoding recipes."""
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field(RecipeType.SPECULATIVE_DFLASH)
|
||||
RECIPE_TYPE: ClassVar[RecipeType] = RecipeType.SPECULATIVE_DFLASH
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field()
|
||||
|
||||
dflash: DFlashConfig = ModeloptField(
|
||||
default=DFlashConfig(),
|
||||
@@ -431,7 +473,9 @@ class ModelOptDFlashRecipe(ModelOptSpeculativeRecipeBase):
|
||||
class ModelOptMedusaRecipe(ModelOptSpeculativeRecipeBase):
|
||||
"""Our config class for Medusa speculative decoding recipes."""
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field(RecipeType.SPECULATIVE_MEDUSA)
|
||||
RECIPE_TYPE: ClassVar[RecipeType] = RecipeType.SPECULATIVE_MEDUSA
|
||||
|
||||
metadata: RecipeMetadataConfig = _metadata_field()
|
||||
|
||||
medusa: MedusaConfig = ModeloptField(
|
||||
default=MedusaConfig(),
|
||||
|
||||
+221
-16
@@ -26,7 +26,11 @@ from pathlib import Path
|
||||
from omegaconf import OmegaConf
|
||||
|
||||
from modelopt.torch.opt.config_loader import BUILTIN_CONFIG_ROOT as BUILTIN_RECIPES_LIB
|
||||
from modelopt.torch.opt.config_loader import _alias_builtin_recipe_prefix, load_config
|
||||
from modelopt.torch.opt.config_loader import (
|
||||
_alias_builtin_recipe_prefix,
|
||||
load_config,
|
||||
peek_declared_schema,
|
||||
)
|
||||
from modelopt.torch.quantization.config import QuantizeConfig
|
||||
|
||||
from .config import (
|
||||
@@ -114,6 +118,42 @@ def load_recipe(
|
||||
``eagle`` (EAGLE speculative decoding), ``dflash`` (DFlash speculative
|
||||
decoding) or ``medusa`` (Medusa speculative decoding) sections. The suffix
|
||||
may be omitted and will be probed automatically.
|
||||
|
||||
.. _recipe-alias:
|
||||
|
||||
A recipe can instead **alias an existing recipe**. A top-level ``$import``
|
||||
brings the imported recipe in whole, so the file needs nothing else::
|
||||
|
||||
imports:
|
||||
base: general/ptq/nvfp4_default-kv_fp8_cast
|
||||
|
||||
$import: base
|
||||
|
||||
That resolves to exactly the imported recipe -- its body, its algorithm and
|
||||
its metadata alike -- under a second name.
|
||||
|
||||
Keys given alongside the ``$import`` override the imported ones, so an alias
|
||||
can say what it is for while inheriting everything else::
|
||||
|
||||
imports:
|
||||
base: general/ptq/nvfp4_default-kv_fp8_cast
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: What this checkpoint uses the base recipe for.
|
||||
|
||||
An override replaces a top-level key outright rather than merging into it, so
|
||||
a partial ``metadata`` supplies the whole section. Nothing needs restating in
|
||||
either form: the recipe's kind comes from the imported recipe (see
|
||||
:func:`_peek_recipe_type`), and the two must agree if the alias states one.
|
||||
|
||||
The one requirement is on the other side -- the imported recipe must carry a
|
||||
``# modelopt-schema:`` comment naming its schema class, as every
|
||||
``$import``-able file must.
|
||||
|
||||
This is how the checkpoint entries under ``models/<org>/<model_id>/`` record
|
||||
that a released checkpoint is reproduced by an existing recipe without
|
||||
duplicating it.
|
||||
* A directory containing ``metadata.yml`` and ``quantize.yml`` —
|
||||
**PTQ recipes only**. Speculative-decoding recipes are always single YAML files.
|
||||
|
||||
@@ -162,19 +202,111 @@ def _apply_dotlist(data: dict, overrides: list[str]) -> dict:
|
||||
return OmegaConf.to_container(merged, resolve=False)
|
||||
|
||||
|
||||
def _peek_recipe_type(recipe_file: Path | Traversable) -> RecipeType | None:
|
||||
"""Extract ``metadata.recipe_type`` from a recipe YAML without resolving $imports.
|
||||
#: Recipe schema classes by their fully-qualified path, for resolving a
|
||||
#: ``# modelopt-schema:`` comment to the recipe kind it names.
|
||||
_RECIPE_SCHEMA_PATHS: dict[str, RecipeType] = {
|
||||
f"{cls.__module__}.{cls.__qualname__}": rtype for rtype, cls in RECIPE_TYPE_TO_CLASS.items()
|
||||
}
|
||||
|
||||
|
||||
_UNPARSED = object() # "caller supplied no pre-parsed body", distinct from a body of ``None``
|
||||
|
||||
|
||||
def _peek_recipe_type(
|
||||
recipe_file: Path | Traversable,
|
||||
_seen: frozenset[str] | None = None,
|
||||
_cycle: list[tuple[str, str]] | None = None,
|
||||
_raw: object = _UNPARSED,
|
||||
) -> RecipeType | None:
|
||||
"""Determine a recipe's kind without resolving its ``$import`` references.
|
||||
|
||||
Needed so :func:`load_config` can be called with the correct ``schema_type`` for
|
||||
typed-list ``$import`` resolution before the full recipe is constructed.
|
||||
typed-list ``$import`` resolution before the full recipe is constructed -- which is
|
||||
why this cannot simply wait for the imports to resolve.
|
||||
|
||||
Neither way of saying it is mandatory; a recipe just has to say it *somehow*, and
|
||||
the three sources are checked in this order:
|
||||
|
||||
1. a ``# modelopt-schema:`` comment naming the recipe's schema class. Only files
|
||||
that are **imported** by another one need this -- it is what
|
||||
``$import`` resolution requires of any snippet -- so a leaf recipe never has to
|
||||
carry it;
|
||||
2. ``metadata.recipe_type`` in the YAML body. **Deprecated** -- still read and still
|
||||
honoured, so no existing recipe has to change, but a new one should declare its
|
||||
schema instead;
|
||||
3. the recipe this one **delegates to** via a top-level ``$import``. A checkpoint
|
||||
alias states neither of the above: its kind is whatever its base is, and the base
|
||||
must declare a schema to be importable at all, so the walk terminates.
|
||||
|
||||
When more than one source is present they must agree --
|
||||
:class:`~modelopt.recipe.config.ModelOptRecipeBase` rejects a recipe whose
|
||||
``metadata.recipe_type`` contradicts its schema class.
|
||||
"""
|
||||
import yaml
|
||||
|
||||
key = str(recipe_file)
|
||||
_seen = (_seen or frozenset()) | {key}
|
||||
|
||||
try:
|
||||
declared = peek_declared_schema(recipe_file)
|
||||
except ValueError: # multiple modelopt-schema comments; load_config reports it
|
||||
declared = None
|
||||
if declared in _RECIPE_SCHEMA_PATHS:
|
||||
return _RECIPE_SCHEMA_PATHS[declared]
|
||||
|
||||
if _raw is not _UNPARSED:
|
||||
# Supplied by _load_recipe_from_file, which has already parsed this file; only the
|
||||
# top-level call can reuse it, the recursion below still reads each import itself.
|
||||
raw = _raw
|
||||
else:
|
||||
try:
|
||||
raw = yaml.safe_load(recipe_file.read_text(encoding="utf-8"))
|
||||
except yaml.YAMLError:
|
||||
return None
|
||||
if not isinstance(raw, dict):
|
||||
return None
|
||||
|
||||
try:
|
||||
raw = yaml.safe_load(recipe_file.read_text())
|
||||
return RecipeType(raw["metadata"]["recipe_type"])
|
||||
except (TypeError, KeyError, ValueError):
|
||||
return None
|
||||
pass
|
||||
|
||||
for base in _delegated_recipe_paths(raw):
|
||||
if str(base) in _seen:
|
||||
# Skipping is what stops the recursion, but a caller that ends up with no kind
|
||||
# at all needs to know a cycle is *why* -- otherwise it reports "nothing
|
||||
# declared anywhere" and sends the author to fix something that is not wrong.
|
||||
# Recorded rather than raised: another `$import` may still resolve the kind,
|
||||
# and a cycle that does not prevent resolution is not worth mentioning.
|
||||
if _cycle is not None:
|
||||
_cycle.append((key, str(base)))
|
||||
continue
|
||||
rtype = _peek_recipe_type(base, _seen, _cycle)
|
||||
if rtype is not None:
|
||||
return rtype
|
||||
return None
|
||||
|
||||
|
||||
def _delegated_recipe_paths(raw: dict) -> list[Path | Traversable]:
|
||||
"""Resolve the recipe files a raw recipe body delegates to via top-level ``$import``.
|
||||
|
||||
Returns an empty list for an ordinary recipe. Names that do not appear in the
|
||||
``imports`` section, or that do not resolve to a file, are skipped here and reported
|
||||
by ``$import`` resolution itself with a better message.
|
||||
"""
|
||||
ref = raw.get("$import")
|
||||
imports = raw.get("imports") or {}
|
||||
if ref is None or not isinstance(imports, dict):
|
||||
return []
|
||||
paths = []
|
||||
for name in ref if isinstance(ref, list) else [ref]:
|
||||
target = imports.get(name)
|
||||
if not target:
|
||||
continue
|
||||
resolved = _resolve_recipe_path(target)
|
||||
if resolved.is_file():
|
||||
paths.append(resolved)
|
||||
return paths
|
||||
|
||||
|
||||
def _load_recipe_from_file(
|
||||
@@ -183,26 +315,76 @@ def _load_recipe_from_file(
|
||||
) -> ModelOptRecipeBase:
|
||||
"""Load a recipe from a YAML file, optionally applying dotlist overrides.
|
||||
|
||||
The file must contain a ``metadata`` section with at least ``recipe_type``,
|
||||
plus the algorithm-specific section (``quantize`` / ``eagle`` / ``dflash`` / ``medusa``).
|
||||
The file has to say what kind of recipe it is -- see :func:`_peek_recipe_type` for
|
||||
the three ways to do that -- and to supply the matching body section (``quantize`` /
|
||||
``eagle`` / ``dflash`` / ``medusa``), either directly or through a top-level
|
||||
``$import`` of a recipe that has one.
|
||||
"""
|
||||
rtype = _peek_recipe_type(recipe_file)
|
||||
# Parsed once and threaded into the peek below, which would otherwise read and parse
|
||||
# the same file a second time. A YAMLError is held rather than raised here so the
|
||||
# kind-resolution errors below still come first, as they did when the peek owned the
|
||||
# only parse -- a malformed file that nonetheless declares its schema in a comment
|
||||
# keeps resolving, and one that does not still reports the parse error.
|
||||
import yaml
|
||||
|
||||
yaml_error: Exception | None = None
|
||||
try:
|
||||
raw_top: object = yaml.safe_load(recipe_file.read_text(encoding="utf-8"))
|
||||
except yaml.YAMLError as exc:
|
||||
raw_top, yaml_error = None, exc
|
||||
|
||||
cycle: list[tuple[str, str]] = []
|
||||
rtype = _peek_recipe_type(recipe_file, _cycle=cycle, _raw=raw_top)
|
||||
if rtype is None:
|
||||
raise ValueError(f"Recipe file {recipe_file} must contain a 'metadata.recipe_type' field.")
|
||||
if cycle:
|
||||
# The delegation remedy below is the one thing the author cannot do here --
|
||||
# they already delegated, and that is the problem. Name the cycle instead.
|
||||
edges = ", ".join(f"{src} -> {dst}" for src, dst in cycle)
|
||||
raise ValueError(
|
||||
f"Recipe file {recipe_file} delegates to a recipe that delegates back to "
|
||||
f"it, so its kind cannot be resolved (cycle: {edges}). Break the cycle, or "
|
||||
"declare the kind directly with a '# modelopt-schema: "
|
||||
"modelopt.recipe.config.ModelOpt<Kind>Recipe' comment."
|
||||
)
|
||||
raise ValueError(
|
||||
f"Recipe file {recipe_file} does not say what kind of recipe it is. Set "
|
||||
"'metadata.recipe_type', or declare a '# modelopt-schema: "
|
||||
"modelopt.recipe.config.ModelOpt<Kind>Recipe' comment, or delegate to a recipe "
|
||||
"that does with a top-level '$import'."
|
||||
)
|
||||
schema_class = RECIPE_TYPE_TO_CLASS.get(rtype)
|
||||
if schema_class is None:
|
||||
raise ValueError(f"Unsupported recipe type: {rtype!r}")
|
||||
|
||||
if yaml_error is not None:
|
||||
raise yaml_error
|
||||
raw = raw_top if isinstance(raw_top, dict) else {}
|
||||
|
||||
# A recipe that delegates inherits the imported recipe's body wholesale, so the two
|
||||
# have to be the same kind. Checked here, and against the *declared* kind on both
|
||||
# sides, so a mismatch reads as a mismatch -- otherwise it surfaces as whatever
|
||||
# pydantic makes of, say, an ``eagle`` section spliced into a PTQ schema.
|
||||
for base in _delegated_recipe_paths(raw):
|
||||
base_type = _peek_recipe_type(base)
|
||||
if base_type is not None and base_type != rtype:
|
||||
raise ValueError(
|
||||
f"Recipe file {recipe_file} is a {rtype.value!r} recipe but imports "
|
||||
f"{base}, which is a {base_type.value!r} recipe. A top-level '$import' "
|
||||
"takes over the whole body, so both must be the same kind."
|
||||
)
|
||||
|
||||
# Pre-flight check on the *raw* YAML so the user sees a clear loader-level error
|
||||
# rather than a generic pydantic missing-field error. Speculative recipes' body
|
||||
# sections have field-level defaults, so this check is what keeps their loader
|
||||
# semantics consistent with PTQ.
|
||||
required_section = _REQUIRED_SECTION_PER_RECIPE_TYPE.get(rtype)
|
||||
if required_section is not None:
|
||||
import yaml
|
||||
|
||||
raw = yaml.safe_load(recipe_file.read_text()) or {}
|
||||
if not isinstance(raw, dict) or required_section not in raw:
|
||||
# A recipe may delegate its whole body to another recipe with a top-level
|
||||
# ``$import`` and override only ``metadata`` -- see :ref:`recipe-alias`. The
|
||||
# body section then arrives during import resolution, so it cannot be required
|
||||
# in the raw YAML; pydantic still rejects the result if the import does not
|
||||
# supply one.
|
||||
if "$import" not in raw and required_section not in raw:
|
||||
# Strip only the ``speculative_`` prefix so multi-word non-speculative types
|
||||
# (e.g. ``auto_quantize``) keep their full name: AUTO_QUANTIZE, not QUANTIZE.
|
||||
kind = rtype.value.removeprefix("speculative_").upper()
|
||||
@@ -251,12 +433,35 @@ def _load_recipe_from_dir(recipe_dir: Path | Traversable) -> ModelOptRecipeBase:
|
||||
Each file is loaded independently. The file name provides the recipe
|
||||
section key: ``metadata.yml`` becomes metadata, and ``quantize.yml`` becomes
|
||||
quantize.
|
||||
|
||||
``metadata.yml``'s kind is resolved the same way a single-file recipe's is (see
|
||||
:func:`_peek_recipe_type`), minus the ``$import``-delegation source: a directory
|
||||
recipe's metadata has nowhere to delegate from. A ``# modelopt-schema:`` comment is
|
||||
checked first, then the (deprecated) ``recipe_type`` field; if both are present,
|
||||
:class:`~modelopt.recipe.config.ModelOptRecipeBase` rejects a disagreement.
|
||||
"""
|
||||
metadata_file = _find_recipe_section_file(recipe_dir, "metadata")
|
||||
metadata = load_config(metadata_file, schema_type=RecipeMetadataConfig)
|
||||
|
||||
if metadata.recipe_type == RecipeType.PTQ:
|
||||
try:
|
||||
declared = peek_declared_schema(metadata_file)
|
||||
except ValueError: # multiple modelopt-schema comments; load_config reports it
|
||||
declared = None
|
||||
# dict.get(declared, ...) is what ruff (SIM401) wants here, but its key type is str,
|
||||
# not str | None like `declared` -- mypy rejects that call, so the if/else stays.
|
||||
if declared in _RECIPE_SCHEMA_PATHS: # noqa: SIM401
|
||||
rtype = _RECIPE_SCHEMA_PATHS[declared]
|
||||
else:
|
||||
rtype = metadata.recipe_type
|
||||
if rtype is None:
|
||||
raise ValueError(
|
||||
f"Recipe directory {recipe_dir}: {metadata_file} must set 'recipe_type' or "
|
||||
"declare a '# modelopt-schema: modelopt.recipe.config.ModelOpt<Kind>Recipe' "
|
||||
"comment. A directory recipe's metadata.yml is the only place its kind can "
|
||||
"come from."
|
||||
)
|
||||
if rtype == RecipeType.PTQ:
|
||||
quantize_file = _find_recipe_section_file(recipe_dir, "quantize")
|
||||
quantize_cfg = load_config(quantize_file, schema_type=QuantizeConfig)
|
||||
return ModelOptPTQRecipe(metadata=metadata, quantize=quantize_cfg)
|
||||
raise ValueError(f"Unsupported recipe type: {metadata.recipe_type!r}")
|
||||
raise ValueError(f"Unsupported recipe type: {rtype!r}")
|
||||
|
||||
@@ -219,6 +219,18 @@ def _parse_modelopt_schema(text: str, config_path: Path | Traversable) -> str |
|
||||
return schema
|
||||
|
||||
|
||||
def peek_declared_schema(config_file: str | Path | Traversable) -> str | None:
|
||||
"""Return the ``# modelopt-schema:`` path a config file declares, if any.
|
||||
|
||||
Reads only the comment preamble -- no YAML parsing and no ``$import`` resolution --
|
||||
so a caller can find out what kind of config a file is *before* it can be loaded.
|
||||
:func:`modelopt.recipe.load_recipe` uses this to pick a recipe's schema class, which
|
||||
is why a recipe can inherit its body from another one via a top-level ``$import``.
|
||||
"""
|
||||
config_path = _resolve_config_path(config_file)
|
||||
return _parse_modelopt_schema(config_path.read_text(encoding="utf-8"), config_path)
|
||||
|
||||
|
||||
def _load_raw_config_with_schema(config_file: str | Path | Traversable) -> _RawConfig:
|
||||
"""Load a config YAML without resolving ``$import`` references."""
|
||||
config_path = _resolve_config_path(config_file)
|
||||
|
||||
@@ -24,7 +24,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: Layer-wise FP8-cast/NVFP4-cast KV-cache AutoQuantize at 5.4 bits using forward KL.
|
||||
|
||||
auto_quantize:
|
||||
|
||||
@@ -23,7 +23,6 @@ imports:
|
||||
fp8: configs/ptq/presets/model/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: Mixed NVFP4 + FP8 per-layer search at 5.4 effective bits.
|
||||
|
||||
auto_quantize:
|
||||
|
||||
@@ -25,7 +25,6 @@ imports:
|
||||
fp8: configs/ptq/presets/model/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: Mixed NVFP4 + FP8 per-layer search at 5.4 effective bits, kl_div scoring (no backprop).
|
||||
|
||||
auto_quantize:
|
||||
|
||||
@@ -23,7 +23,6 @@ imports:
|
||||
fp8: configs/ptq/presets/model/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: Mixed NVFP4 (weight-MSE + FP8 sweep) + FP8 per-layer search at 6.0 effective bits.
|
||||
|
||||
auto_quantize:
|
||||
|
||||
@@ -24,7 +24,6 @@ imports:
|
||||
w4a16_nvfp4: configs/ptq/presets/model/w4a16_nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: >-
|
||||
Mixed FP8 + NVFP4-weight-only per-layer search at 6.0 effective bits with the
|
||||
active-MoE cost model (expert ratio 0.03125).
|
||||
|
||||
@@ -23,7 +23,6 @@ imports:
|
||||
fp8: configs/ptq/presets/model/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: Mixed W4A8 AWQ-beta + FP8 per-layer search at 6.0 effective bits.
|
||||
|
||||
auto_quantize:
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for W8A8 FP8 E4M3 model quantization with FP8 KV-cache quantization.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes W8A8 FP8 E4M3 model quantization with FP8 KV-cache quantization; uses max calibration.
|
||||
quantize:
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for W8A8 FP8 E4M3 model quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes W8A8 FP8 E4M3 model quantization with FP8 KV-cache cast mode using constant amax; uses
|
||||
max calibration.
|
||||
|
||||
@@ -13,8 +13,8 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: INT4 blockwise weight-only (W4A16, block size 128), max calibration.
|
||||
|
||||
imports:
|
||||
|
||||
@@ -16,11 +16,11 @@
|
||||
# PTQ recipe wrapping the shipped mxfp4_mlp_weight_only preset (the config behind
|
||||
# mtq.MXFP4_MLP_WEIGHT_ONLY_CFG), so the two cannot drift apart.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
preset: configs/ptq/presets/model/mxfp4_mlp_weight_only
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic MXFP4 to MLP/MoE weight quantizers only (weight-only, activations untouched);
|
||||
no calibration forward pass is required.
|
||||
|
||||
@@ -16,6 +16,7 @@
|
||||
# using nvfp4_act_headroom calibration: NVFP4 weights on plain max, NVFP4 activation global
|
||||
# scales anchored with headroom. Module coverage matches nvfp4_default-kv_fp8_cast.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes dynamic NVFP4 W4A4 model quantization with FP8 KV-cache cast mode using constant
|
||||
amax; uses nvfp4_act_headroom calibration. Weight quantizers use plain max, while the NVFP4
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for dynamic NVFP4 W4A4 model quantization with FP8 KV-cache quantization.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes dynamic NVFP4 W4A4 model quantization with FP8 KV-cache quantization for PTQ,
|
||||
QAT, and QAD; uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for dynamic NVFP4 W4A4 model quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes dynamic NVFP4 W4A4 model quantization with FP8 KV-cache cast mode using constant amax;
|
||||
uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for NVFP4 W4A4 model quantization with KV quantizers disabled and GPTQ calibration.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies NVFP4 W4A4 with static weight scales, dynamic inputs, KV quantizers disabled, and GPTQ
|
||||
layerwise calibration.
|
||||
|
||||
@@ -14,12 +14,12 @@
|
||||
# limitations under the License.
|
||||
|
||||
# PTQ recipe wrapping the shipped NVFP4 W4A4 local-Hessian preset.
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
|
||||
imports:
|
||||
preset: configs/ptq/presets/model/nvfp4_w4a4_weight_local_hessian
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies NVFP4 W4A4 with static weight scales, dynamic inputs, KV quantizers disabled, and
|
||||
local-Hessian weight calibration with an FP8-scale sweep.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for dynamic NVFP4 W4A4 model quantization with NVFP4 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_nvfp4_cast: configs/ptq/units/kv_nvfp4_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Composes dynamic NVFP4 W4A4 model quantization with NVFP4 KV-cache cast mode using constant
|
||||
amax; uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for expert-only dynamic NVFP4 quantization with FP8 KV-cache quantization.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to expert-layer weight and input quantizers, plus FP8 KV-cache
|
||||
quantization; uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for expert-only dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to expert-layer weight and input quantizers, plus FP8 KV-cache cast
|
||||
mode using constant amax; uses max calibration.
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -20,7 +21,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: NVFP4 static weight and dynamic activation for expert layers only (W4A4), FP8 KV cache, max layerwise calibration.
|
||||
quantize:
|
||||
algorithm:
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -20,7 +21,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >
|
||||
NVFP4 static weight and dynamic activation for expert layers only (W4A4), FP8 KV cache,
|
||||
max layerwise calibration, exporting each decoder layer to a quantized checkpoint shard
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -20,7 +21,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >
|
||||
NVFP4 static weight and dynamic activation for expert layers only (W4A4), FP8 KV cache,
|
||||
max layerwise calibration with calib_mutates_weights=False for disk-offloaded single-GPU
|
||||
|
||||
@@ -22,6 +22,7 @@
|
||||
# ``amax / (E2M1_MAX * E4M3_MAX) = 1.0``; the per-block E4M3 activation scales remain dynamic.
|
||||
# No activation statistics are collected for these quantizers.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -29,7 +30,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies NVFP4 only to expert-layer weight and input quantizers, pinning the expert activation
|
||||
input_scale to 1.0 (constant amax = 2688, no activation calibration), plus FP8 KV-cache cast
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for expert-only NVFP4 quantization with MSE weight calibration and FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies static NVFP4 weight scales from MSE FP8-scale sweep and dynamic NVFP4 inputs to expert
|
||||
layers only, plus FP8 KV-cache cast mode.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for MLP/MoE-only dynamic NVFP4 quantization with FP8 KV-cache quantization.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to MLP/MoE weight and input quantizers, plus FP8 KV-cache
|
||||
quantization; uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for MLP/MoE-only dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to MLP/MoE weight and input quantizers, plus FP8 KV-cache cast mode
|
||||
using constant amax; uses max calibration.
|
||||
|
||||
@@ -30,6 +30,7 @@
|
||||
# `nvidia/Kimi-K2.5-NVFP4` checkpoint (which excludes the vision tower and
|
||||
# multimodal projector).
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -38,7 +39,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 only to MLP/MoE weight and input quantizers, plus FP8 KV-cache
|
||||
quantization; uses max calibration. Vision tower (model.visual.* / vision_tower.*)
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for MLP/MoE-only NVFP4 quantization with MSE weight calibration and FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies static NVFP4 weight scales from MSE FP8-scale sweep and dynamic NVFP4 inputs to MLP/MoE
|
||||
layers, plus FP8 KV-cache cast mode.
|
||||
|
||||
@@ -16,11 +16,11 @@
|
||||
# PTQ recipe wrapping the shipped nvfp4_mlp_weight_only preset (the config behind
|
||||
# mtq.NVFP4_MLP_WEIGHT_ONLY_CFG), so the two cannot drift apart.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
preset: configs/ptq/presets/model/nvfp4_mlp_weight_only
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies NVFP4 (block size 32) to MLP/MoE weight quantizers only (weight-only, activations
|
||||
untouched); uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for output-projection and MLP/MoE dynamic NVFP4 quantization with FP8 KV-cache quantization.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 to output-projection and MLP/MoE weight and input quantizers, plus
|
||||
FP8 KV-cache quantization; uses max calibration.
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for output-projection and MLP/MoE dynamic NVFP4 quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 to output-projection and MLP/MoE weight and input quantizers, plus
|
||||
FP8 KV-cache cast mode using constant amax; uses max calibration.
|
||||
|
||||
@@ -13,13 +13,13 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
w4a16_nvfp4: configs/ptq/units/w4_nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: NVFP4 W4A16 weight-only, BF16 activations, max calibration. No calibration forward pass required.
|
||||
quantize:
|
||||
algorithm: max
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Composed PTQ recipe for NVFP4 W4A16 weight-only quantization with FP8 KV-cache cast mode.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -22,7 +23,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
NVFP4 W4A16 weight-only, BF16 activations, plus FP8 KV-cache cast mode using constant amax; uses
|
||||
max calibration. No calibration forward pass required.
|
||||
|
||||
@@ -13,8 +13,8 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Learns separate pre-quantization and post-quantization NVFP4 weight scales with MSE
|
||||
initialization and FP8 scale sweep; uses dynamic NVFP4 activations and FP8 KV cache.
|
||||
|
||||
@@ -13,8 +13,8 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Learns one shared pre-quantization and post-quantization NVFP4 weight scale with MSE
|
||||
initialization and FP8 scale sweep; uses dynamic NVFP4 activations and FP8 KV cache.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# DFlash speculative-decoding training recipe. Override fields via OmegaConf dotlist on the CLI.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: DFlash training recipe (model/data/training/dflash bundled).
|
||||
|
||||
# maps to ModelArguments (main.py)
|
||||
|
||||
@@ -5,8 +5,8 @@
|
||||
# Online training is the default path (data.mode=online). Override fields via an
|
||||
# OmegaConf dotlist on the CLI.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: Domino training recipe (DFlash backbone + causal correction head).
|
||||
|
||||
# maps to ModelArguments (main.py)
|
||||
|
||||
@@ -8,8 +8,8 @@
|
||||
# loss: ce_alpha*CE + l1_alpha*TVD + conf_alpha*confidence_BCE. Online training is
|
||||
# the default path (data.mode=online). Override fields via an OmegaConf dotlist.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: DSpark training recipe (DFlash backbone + Markov head + confidence head).
|
||||
|
||||
# maps to ModelArguments (main.py)
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# EAGLE3 speculative-decoding training recipe. Override fields via OmegaConf dotlist on the CLI.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptEagleRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_eagle
|
||||
description: EAGLE3 training recipe (model/data/training/eagle bundled).
|
||||
|
||||
# maps to ModelArguments (main.py)
|
||||
|
||||
@@ -17,8 +17,8 @@
|
||||
# the composition of the cross-entropy weight — see the two-line override below.
|
||||
# Override fields via an OmegaConf dotlist.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: LiLiCorr training recipe (DFlash backbone + candidate-lattice reranker).
|
||||
|
||||
# maps to ModelArguments (main.py)
|
||||
|
||||
@@ -29,8 +29,8 @@
|
||||
# does not.
|
||||
#
|
||||
# Override fields via an OmegaConf dotlist.
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: >-
|
||||
LiLiCorr + grouped sublayer convolutions (DFlash backbone + candidate-lattice
|
||||
reranker + DFlash2 sublayer convolutions).
|
||||
|
||||
@@ -21,6 +21,7 @@
|
||||
# ``*self_conditioning*``. Matches the ``--qformat nvfp4_experts_only``
|
||||
# behavior in ``hf_ptq.py``, which defaults ``--kv_cache_qformat=fp8_cast``.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
experts_nvfp4: configs/ptq/units/experts_nvfp4
|
||||
@@ -28,7 +29,6 @@ imports:
|
||||
disabled_quantizers: model_type/diffusion_gemma/ptq/disabled_quantizers
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
DiffusionGemma PTQ recipe (nvfp4_experts_only): dynamic W4A4 NVFP4 on MoE
|
||||
experts only with FP8 KV-cache cast (constant-amax); attention Q/K/V/O
|
||||
|
||||
@@ -17,6 +17,7 @@
|
||||
# the SmoothQuant alpha from the default 1.0 to 0.5 to recover accuracy on
|
||||
# Gemma 7B (default alpha causes a regression).
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -25,7 +26,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Gemma INT8 SmoothQuant recipe with FP8 KV-cache cast: alpha=0.5 (instead
|
||||
of the default 1.0) to avoid accuracy regression on Gemma 7B, plus FP8
|
||||
|
||||
@@ -17,6 +17,7 @@
|
||||
# optimal-scale search (awq_lite with alpha_step=1) to avoid overflow observed
|
||||
# in TRT-LLM kernels when using the default AWQ search on Gemma.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -25,7 +26,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Gemma W4A8 AWQ recipe with FP8 KV-cache cast: INT4 block weights + FP8
|
||||
inputs, awq_lite with alpha_step=1 (coarser search) to avoid TRT-LLM
|
||||
|
||||
@@ -28,6 +28,7 @@
|
||||
# branch is kept in BF16 by the shared `default_disabled_quantizers` unit imported
|
||||
# below, which globally disables `*vision_tower*` / `*visual*` / `*embed_vision*`.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -36,7 +37,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Gemma 4 (multimodal) W4A8 AWQ recipe with FP8 KV-cache cast: INT4 block
|
||||
weights + FP8 inputs, awq_lite with alpha_step=1 (coarser search) to avoid
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -21,7 +22,6 @@ imports:
|
||||
nvfp4_static: configs/numerics/nvfp4_static
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
MiniMax-M3 mixed MXFP8 and NVFP4 PTQ. Routed experts use MSE-calibrated
|
||||
NVFP4 weights and dynamic NVFP4 activations with input_scale fixed to 1.0;
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -20,7 +21,6 @@ imports:
|
||||
nvfp4_static: configs/numerics/nvfp4_static
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
MiniMax-M3 routed-experts-only NVFP4 PTQ with MSE-calibrated static
|
||||
weights and dynamic activations; shared experts and the KV cache stay
|
||||
|
||||
@@ -17,6 +17,7 @@
|
||||
# optimal-scale search (awq_lite with alpha_step=1) to avoid overflow observed
|
||||
# in TRT-LLM kernels when using the default AWQ search on MPT.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -25,7 +26,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
MPT W4A8 AWQ recipe with FP8 KV-cache cast: INT4 block weights + FP8
|
||||
inputs, awq_lite with alpha_step=1 (coarser search) to avoid TRT-LLM
|
||||
|
||||
@@ -20,6 +20,7 @@
|
||||
# top of the FP16 activations, so the engine uses MORE activation memory than
|
||||
# an unquantized FP16 engine. [5726458]
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
|
||||
@@ -27,7 +28,6 @@ imports:
|
||||
fp8: configs/numerics/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: 'Llama-Nemotron embedding/reranking PTQ recipe (fp8 + projection output quantizers): same numerics as the general fp8 preset, plus per-tensor
|
||||
FP8 output quantizers on the projection Linears so TensorRT engines carry FP8 activations between layers.'
|
||||
quantize:
|
||||
|
||||
@@ -20,6 +20,7 @@
|
||||
# Output quantizers must stay scoped to GEMM outputs: a DynamicQuantize on
|
||||
# non-GEMM outputs (embeddings, pooling) fails to compile in TensorRT.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
|
||||
@@ -27,7 +28,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: 'Llama-Nemotron embedding PTQ recipe (nvfp4 + projection output quantizers): same numerics as the general nvfp4 preset, plus dynamic NVFP4
|
||||
output quantizers on the projection Linears so TensorRT engines carry FP4 activations between layers.'
|
||||
quantize:
|
||||
|
||||
@@ -17,6 +17,7 @@
|
||||
# Equivalent to the general `nvfp4` preset with quantization disabled
|
||||
# on non-language branches.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
|
||||
@@ -24,7 +25,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: 'Nemotron VL PTQ recipe (nvfp4): same numerics as the general nvfp4 preset, applied to the decoder (text generation) component only (vision/encoder
|
||||
branches are skipped).'
|
||||
quantize:
|
||||
|
||||
@@ -15,12 +15,12 @@
|
||||
|
||||
# Dense Qwen3.5 Vision Encoder FP8 with a high-precision language model and KV cache.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
FP8 quantization of dense Qwen3.5 Vision Encoder Linear layers, including any merger Linears;
|
||||
the language model, KV cache, patch embedding, and vision-attention operands remain in high
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Dense Qwen3.5 Vision Encoder and language model FP8 with FP8 KV-cache cast.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W8A8 FP8 quantization of dense Qwen3.5 Vision Encoder and language-model Linear layers,
|
||||
including any merger Linears, with FP8 KV-cache cast; patch embedding and vision-attention
|
||||
|
||||
@@ -21,11 +21,11 @@
|
||||
# the snippet lives under
|
||||
# `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights) MLP / FP8 attention / FP8 KV-cast PTQ recipe for
|
||||
HuggingFace `qwen3_5` (dense) models: NVFP4 for MLP projection weights and
|
||||
|
||||
@@ -24,11 +24,11 @@
|
||||
# the snippet lives under
|
||||
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
recipe for HuggingFace `qwen3_5` (dense) models: NVFP4 with static scales
|
||||
|
||||
@@ -19,6 +19,7 @@
|
||||
# Note: fp8_scale_sweep refines ONLY static NVFP4 weights with MSE; all FP8 / dynamic / KV
|
||||
# quantizers stay max-calibrated (modelopt default) -- i.e. MSE applies only to the NVFP4 layers.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -28,7 +29,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
NVFP4 (MSE static weights, dynamic inputs) on LM routed experts; ModelOpt-default FP8
|
||||
(max-calibrated) on all other Linear layers; FP8 KV cache; MTP block BF16.
|
||||
|
||||
@@ -21,11 +21,11 @@
|
||||
# `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.yaml`; the snippet lives
|
||||
# under `model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights) MLP / FP8 attention / FP8 KV-cast PTQ recipe for
|
||||
HuggingFace `qwen3_5_moe` models (Qwen3.5-MoE and Qwen3.6-MoE releases):
|
||||
|
||||
+1
-1
@@ -24,11 +24,11 @@
|
||||
# snippet lives under
|
||||
# `model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg.yaml`.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
shared_quant_cfg: model_type/qwen3_5/ptq/w4a16_nvfp4_mse-fp8_attn-kv_fp8_cast.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W4A16 (NVFP4 weights, MSE calibration) MLP / FP8 attention / FP8 KV-cast PTQ
|
||||
recipe for HuggingFace `qwen3_5_moe` models (Qwen3.5-MoE and Qwen3.6-MoE
|
||||
|
||||
-1
@@ -25,7 +25,6 @@ imports:
|
||||
w4a16_nvfp4: configs/ptq/presets/model/w4a16_nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: >-
|
||||
Qwen3.6 MoE: FP8 + NVFP4-weight-only per-layer search at 6.0 effective bits, active-MoE
|
||||
cost model (expert ratio 0.03125), with architecture-specific disabled layers.
|
||||
|
||||
-1
@@ -14,7 +14,6 @@ imports:
|
||||
w4a16_nvfp4: configs/ptq/presets/model/w4a16_nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: >-
|
||||
Qwen3.6 MoE module-specific search based on the Qwen3.5-MoE W4A16/FP8 PTQ
|
||||
recipe: shared experts, attention, and lm_head searched over W4A16 NVFP4
|
||||
|
||||
+1
-1
@@ -50,6 +50,7 @@
|
||||
# in_proj_b BF16 exclusion cannot be expressed here; `_gated_delta_net_slicing` in
|
||||
# modelopt/torch/export/unified_export_megatron.py re-splits it on export and writes those two
|
||||
# out in BF16.
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
@@ -59,7 +60,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Megatron-Core W4A4 (NVFP4 weights and activations, block 16) on MoE routed experts, shared
|
||||
experts and dense MLP; W4A4 NVFP4 on the output layer; FP8 W8A8 on softmax-attention and
|
||||
|
||||
@@ -15,12 +15,12 @@
|
||||
|
||||
# Qwen3-VL Vision Encoder FP8 with a high-precision language model and KV cache.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
vision_fp8: model_type/qwen3_vl/ptq/vision_fp8.quant_cfg
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
FP8 quantization of Qwen3-VL Vision Encoder Linear layers, including primary and deepstack
|
||||
merger Linears; the language model, KV cache, patch embedding, and vision-attention operands
|
||||
|
||||
@@ -15,6 +15,7 @@
|
||||
|
||||
# Qwen3-VL Vision Encoder and language model FP8 with FP8 KV-cache cast.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
W8A8 FP8 quantization of Qwen3-VL Vision Encoder and language-model Linear layers, including
|
||||
primary and deepstack merger Linears, with FP8 KV-cache cast; patch embedding and
|
||||
|
||||
@@ -20,6 +20,7 @@
|
||||
# `moe.{gate,up,down}_proj`, the router is `moe.gate`, and each MoE layer also has a dense
|
||||
# `share_expert` that stays in BF16.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -27,7 +28,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 to Step-3.7 routed-expert weight and input quantizers, plus FP8
|
||||
KV-cache cast mode using constant amax; the router, the shared expert and the dense MLPs
|
||||
|
||||
@@ -19,6 +19,7 @@
|
||||
# general `*mlp*` / `*.experts.*` patterns reach only the dense layers' `mlp` submodules and
|
||||
# leave the routed experts — the bulk of the model — in BF16.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -26,7 +27,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Applies dynamic NVFP4 to Step-3.7 routed-expert and dense-MLP weight and input
|
||||
quantizers, plus calibrated FP8 KV-cache quantization; the router and the shared expert
|
||||
|
||||
@@ -13,8 +13,12 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
ViT image classifiers: FP8 W8A8 on every linear plus the attention BMM and
|
||||
softmax-P quantizers, so the whole attention block runs in FP8 and Torch-TRT can
|
||||
compile it end to end; output quantizers disabled. Uses max calibration.
|
||||
imports:
|
||||
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
|
||||
attention_qkv_fp8: configs/ptq/units/attention_qkv_fp8
|
||||
|
||||
+1
-1
@@ -15,6 +15,7 @@
|
||||
|
||||
# Reproduces the quantization scheme of the published nvidia/MiniMax-M2.7-NVFP4
|
||||
# checkpoint from the MiniMaxAI/MiniMax-M2.7 input model.
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
@@ -23,7 +24,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
MiniMax M2.7 expert-only NVFP4 weights and activations with max calibration;
|
||||
FP8 KV-cache cast mode; attention projections, router, and LM head stay BF16.
|
||||
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Alias recipe for nvidia/Qwen3.5-397B-A17B-NVFP4-V2 (source Qwen/Qwen3.5-397B-A17B). The
|
||||
# qwen3_5_moe architecture recipe already produces this release's scheme -- NVFP4 (MSE static
|
||||
# weights, dynamic inputs) on the LM routed experts, ModelOpt-default FP8 elsewhere, and an FP8
|
||||
# KV cache -- so this entry imports that recipe wholesale instead of copying its body.
|
||||
|
||||
imports:
|
||||
base: model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: >-
|
||||
Qwen/Qwen3.5-397B-A17B quantized with the qwen3_5_moe architecture recipe -- NVFP4 (MSE
|
||||
static weights, dynamic inputs) on the LM routed experts, ModelOpt-default FP8 on every other
|
||||
Linear layer, and an FP8 KV cache -- as published in nvidia/Qwen3.5-397B-A17B-NVFP4-V2.
|
||||
+1
-1
@@ -49,6 +49,7 @@
|
||||
# BF16 before quantizers are inserted, so the scales below are calibrated against BF16
|
||||
# weights rather than against the shipped FP8.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
@@ -57,7 +58,6 @@ imports:
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
qwen3_5_moe_text: NVFP4 (MSE static weights, dynamic inputs) on the routed experts + FP8 (W8A8)
|
||||
on all self-attention AND all linear-attention projections; everything else BF16; FP8 KV cache
|
||||
|
||||
-1
@@ -18,7 +18,6 @@ imports:
|
||||
base: models/Qwen/Qwen3.8-27B/ptq/nvfp4_w4a4_mlp_fp8_attn_max
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Qwen3.8-27B local-Hessian PTQ using the NVFP4 W4A4 and FP8 W8A8 assignment exported by
|
||||
the 5.5-bit NVFP4-max AutoQuantize sweep.
|
||||
|
||||
@@ -5,8 +5,14 @@ This folder holds model-optimization recipes (e.g. PTQ recipes) tuned for a
|
||||
such as the [Hugging Face Hub](https://huggingface.co/),
|
||||
[ModelScope](https://modelscope.cn/), or similar. Unlike
|
||||
[`../model_type/`](../model_type/), which keys recipes by a transformers
|
||||
`model_type` (an architecture shared by many checkpoints), a recipe here mirrors
|
||||
**one checkpoint's** quantization scheme verbatim.
|
||||
`model_type` (an architecture shared by many checkpoints), an entry here is keyed
|
||||
to **one checkpoint**.
|
||||
|
||||
An entry takes one of two forms: a **mirror**, whose body reproduces a per-layer
|
||||
scheme no portable recipe can express, or an **alias**, which records that a
|
||||
`general/` or `model_type/<model_type>/` recipe already produces that
|
||||
checkpoint's scheme and imports it wholesale. See
|
||||
[What belongs here](#what-belongs-here) for which to write.
|
||||
|
||||
## Folder structure
|
||||
|
||||
@@ -46,8 +52,9 @@ applies to its mirrors, key it by that base model's id.
|
||||
Prefer the most specific entry that applies to your model:
|
||||
|
||||
1. **`models/<org>/<model_id>/`** — if there is an entry for your **exact**
|
||||
checkpoint. It reproduces a validated, often per-component mixed-precision
|
||||
scheme for that release; use it to match a published quantized checkpoint.
|
||||
checkpoint. Use it to match a published quantized checkpoint: it either
|
||||
reproduces a validated, often per-component mixed-precision scheme for that
|
||||
release, or aliases the portable recipe that does.
|
||||
2. **[`model_type/<model_type>/`](../model_type/)** — an architecture-level
|
||||
recipe that applies to every checkpoint of that `model_type`.
|
||||
3. **[`general/`](../general/)** — model-agnostic recipes; a good starting point
|
||||
@@ -71,7 +78,11 @@ recipe = load_recipe("models/nvidia/NVIDIA-Nemotron-3-Nano-4B-BF16/ptq/nvfp4_w4a
|
||||
|
||||
## What belongs here
|
||||
|
||||
A recipe earns a place here only when it mirrors **one specific released (or
|
||||
Two kinds of entry, and the difference matters when you add one.
|
||||
|
||||
### 1. Checkpoint mirrors — the recipe lives here
|
||||
|
||||
A recipe earns a **body** here only when it mirrors **one specific released (or
|
||||
planned) checkpoint** — a hand-mapped, usually per-layer or per-component
|
||||
precision scheme tuned to match that exact release. If the tuning generalizes to
|
||||
every checkpoint of an architecture, it belongs under
|
||||
@@ -80,6 +91,65 @@ model-agnostic, it belongs under [`../general/`](../general/). See
|
||||
[`../ptq.md`](../ptq.md) for what each checkpoint mirror does and how it compares
|
||||
to its general baseline.
|
||||
|
||||
### 2. Aliases — the recipe lives elsewhere, the *record* lives here
|
||||
|
||||
Many released checkpoints use a scheme a portable recipe already produces, with no
|
||||
checkpoint-specific changes at all. Those can still get an entry here, so the recipe
|
||||
is findable at the checkpoint's own hub path — but the entry is a thin **alias**
|
||||
that delegates its whole body to that recipe:
|
||||
|
||||
```yaml
|
||||
imports:
|
||||
base: general/ptq/nvfp4_default-kv_fp8_cast
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: >-
|
||||
meta-llama/Llama-3.1-8B-Instruct quantized with the general NVFP4 scheme and an
|
||||
FP8 KV cache in cast mode, as published in nvidia/Llama-3.1-8B-Instruct-NVFP4.
|
||||
```
|
||||
|
||||
A top-level `$import` brings in the whole imported recipe; the keys given
|
||||
alongside it override the imported ones, so the alias supplies its own
|
||||
`metadata` and inherits `quantize` — algorithm and every `quant_cfg` entry —
|
||||
unchanged. Nothing is duplicated: editing the base recipe changes every alias
|
||||
that points at it.
|
||||
|
||||
Note what the alias does *not* say. It has no `recipe_type`, because the kind is
|
||||
whatever the imported recipe's kind is; a recipe states its kind in whichever of
|
||||
these it likes, and the loader takes the first that answers:
|
||||
|
||||
1. a `# modelopt-schema:` comment naming its schema class,
|
||||
2. `metadata.recipe_type` — **deprecated**; still read, but new recipes should
|
||||
leave it out,
|
||||
3. the recipe it delegates to via a top-level `$import`.
|
||||
|
||||
Whatever a recipe does state has to be true. Declaring both a schema comment and
|
||||
a `recipe_type` is fine as long as they agree, and the same holds across a
|
||||
delegation: a recipe and the recipe it imports must be the same kind, since the
|
||||
import takes over the whole body. Any disagreement is an error, not a preference.
|
||||
|
||||
The one thing that *is* required: **a recipe another file imports must carry the
|
||||
schema comment**, because that is what `$import` resolution needs to validate the
|
||||
imported payload. A recipe nothing imports needs nothing — so aliasing a recipe
|
||||
for the first time means adding the comment to it in the same change.
|
||||
|
||||
**Which one to write.** Start by assuming an alias, and reach for a body here only
|
||||
once you have established that no portable recipe expresses the release's scheme —
|
||||
compare the release's own `hf_quant_config.json` (and, where the distinction matters,
|
||||
its exported scale tensors) against what the candidate recipe's `quant_cfg` would
|
||||
produce. A body duplicated from a general recipe is a maintenance liability: it stops
|
||||
tracking edits to the recipe it was copied from.
|
||||
|
||||
### When a release gets no entry at all
|
||||
|
||||
A release is backfilled here only when its model card describes a **post-training**
|
||||
quantization recipe. A card documenting quantization-aware distillation after PTQ is
|
||||
deliberately left out: no PTQ recipe reproduces that checkpoint, so an entry claiming to
|
||||
would be wrong. If you came looking for a published NVIDIA checkpoint and did not find
|
||||
it, this is the usual reason — check whether its card mentions QAD before assuming the
|
||||
entry is merely missing.
|
||||
|
||||
## Sharing content across recipes
|
||||
|
||||
When several recipes reuse the same body, extract it into a sibling **snippet**
|
||||
|
||||
+1
-1
@@ -25,12 +25,12 @@
|
||||
# are model.layers.<i>.ffn.experts.<j>.w{1,2,3}. The shared expert is the same
|
||||
# Expert class under a different MoE attribute, so it is disabled explicitly.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
DeepSeek-V4-Pro-0813 recipe quantizing only the routed-expert linears to
|
||||
NVFP4 (weights and activations). Shared experts, attention, router gates,
|
||||
|
||||
-1
@@ -14,7 +14,6 @@ imports:
|
||||
nvfp4_four_over_six: configs/numerics/nvfp4_four_over_six
|
||||
|
||||
metadata:
|
||||
recipe_type: auto_quantize
|
||||
description: >-
|
||||
Muse Glimmer language-model W4A16 NVFP4 Four-Over-Six, FP8, and BF16
|
||||
fallback search at 5.5 weight-effective bits. Use --kv_cache_qformat none.
|
||||
|
||||
@@ -13,6 +13,7 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -21,7 +22,6 @@ imports:
|
||||
kv_fp8: configs/ptq/units/kv_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
NVFP4 W4A4 on interior MLP layers, FP8 W8A8 on edge MLP and attention
|
||||
layers, and an FP8 KV cache; uses max calibration.
|
||||
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Alias recipe for nvidia/Kimi-K2.6-NVFP4 (source moonshotai/Kimi-K2.6). The general
|
||||
# expert-only NVFP4 scheme -- MSE-swept static weight scales, dynamic NVFP4 inputs, FP8
|
||||
# KV cache in cast mode -- reproduces this release with no checkpoint-specific deviation,
|
||||
# so this entry imports that recipe wholesale instead of copying its body.
|
||||
|
||||
imports:
|
||||
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: >-
|
||||
moonshotai/Kimi-K2.6 quantized with the general expert-only NVFP4 scheme (MSE static
|
||||
weight scales, dynamic inputs) and an FP8 KV cache in cast mode, as published in
|
||||
nvidia/Kimi-K2.6-NVFP4.
|
||||
@@ -18,6 +18,7 @@
|
||||
# applies this layout as a streaming checkpoint conversion rather than loading
|
||||
# the complete model through examples/hf_ptq/hf_ptq.py.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
default_disabled_quantizers: configs/ptq/units/default_disabled_quantizers
|
||||
@@ -25,7 +26,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Kimi-K3 checkpoint-mirror recipe with calibration-free MXFP4-to-NVFP4
|
||||
routed experts, expert input_scale fixed to 1.0, and 128x128 block-FP8
|
||||
|
||||
@@ -26,6 +26,7 @@
|
||||
# lm_head → FP8 W8A16 (weight-only; overrides its Q8_0 GGUF type)
|
||||
# everything else (F32: conv1d, norms, A_log/D/dt_bias) → bf16
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
w4a4_nvfp4_nvfp4: configs/ptq/units/w4a4_nvfp4_nvfp4
|
||||
@@ -34,7 +35,6 @@ imports:
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
GGUF Q4_K_M-mirrored Nemotron-H recipe: NVFP4 W4A4 for Q4_K/Q5_0 linears
|
||||
(attn q/k/v/o kept uniform for export fusion), FP8 W8A8 for Q6_K MLP
|
||||
|
||||
+1
-1
@@ -29,8 +29,8 @@
|
||||
# - SSM cache: FP32 (can be set to FP16 in VLLM)
|
||||
#
|
||||
# Calibration: amax/max calibration comparison variant
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: Super NVFP4 mixed precision — sparse MoE experts NVFP4 (W4A4, group_size 16); shared experts, mamba in/out_proj FP8 per-tensor; FP8 KV cache;
|
||||
everything else(lm_head/MTP/Latent MOE) stay BF16. Amax calibration comparison variant.
|
||||
quantize:
|
||||
|
||||
+1
-1
@@ -30,8 +30,8 @@
|
||||
# Calibration: weight MSE with FP8-scale sweep over the 128 e4m3 scale values
|
||||
# (NVFP4 weights use static block scales selected by MSE; FP8 per-tensor scales
|
||||
# are also chosen via MSE search instead of plain amax).
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: Super NVFP4 mixed precision — sparse MoE experts NVFP4 (W4A4, group_size 16); shared experts, mamba in/out_proj FP8 per-tensor; FP8 KV cache;
|
||||
everything else(lm_head/MTP/latent MOE) stay BF16. Weight-MSE calibration with FP8 scale sweep.
|
||||
quantize:
|
||||
|
||||
+1
-1
@@ -11,13 +11,13 @@
|
||||
# Weights: 4/6 NVFP4, per-block M=6 vs M=4 selected by MSE. Activations: dynamic NVFP4 (NOT MSE).
|
||||
# Shared experts + mamba in/out_proj + KV cache: FP8. Everything else: BF16.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
nvfp4_four_over_six: configs/numerics/nvfp4_four_over_six
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
fp8: configs/numerics/fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: Ultra NVFP4 with Four-Over-Six (4/6) MSE-selected weights on routed experts (W4A4, block 16); dynamic NVFP4 activations (not MSE); shared
|
||||
experts + mamba in/out_proj + KV cache FP8; everything else BF16.
|
||||
quantize:
|
||||
|
||||
+1
-1
@@ -11,11 +11,11 @@
|
||||
# - mixer.shared_experts.{up,down}_proj -> mlp.shared_experts.linear_fc{1,2}
|
||||
# - mixer.in_proj / out_proj -> mixer.in_proj / out_proj (same name; W4A16)
|
||||
# - lm_head -> output_layer (W4A16)
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
nvfp4_four_over_six: configs/numerics/nvfp4_four_over_six
|
||||
fp8_default: configs/numerics/fp8
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >
|
||||
Lightning 3.5 W4A16 PTQ, NVFP4 four_over_six (4/6) weight scales.
|
||||
Routed MoE experts, shared experts, and lm_head; FP8 Mamba in_proj/out_proj.
|
||||
|
||||
+1
-1
@@ -20,8 +20,8 @@
|
||||
# transformers-5 path (`mamba`->`linear_attention`, `attention`->`full_attention`, plus a
|
||||
# matching `hybrid_override_pattern`); that conversion is not covered here either.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
||||
metadata:
|
||||
recipe_type: speculative_dflash
|
||||
description: Warm-start fine-tune of the released Nemotron-3.5 DSpark drafter (streaming).
|
||||
|
||||
model:
|
||||
|
||||
@@ -15,12 +15,12 @@
|
||||
|
||||
# Model-specific PTQ recipe for Step3.5-Flash NVFP4 MLP/MoE quantization with FP8 KV cache.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
fp8: configs/numerics/fp8
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
Step3.5-Flash PTQ recipe that enables dynamic NVFP4 on MoE/MLP weight and input quantizers,
|
||||
enables FP8 KV-cache quantizers, and leaves other quantizers disabled.
|
||||
|
||||
+1
-1
@@ -53,13 +53,13 @@
|
||||
# than they are: `*.experts.*` needs a literal `.experts.`, so `mlp.shared_experts.*` does not
|
||||
# match; and `*mlp.gate_proj*` needs `mlp.gate_proj`, which the router `mlp.gate` lacks.
|
||||
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base_disable_all: configs/ptq/units/base_disable_all
|
||||
nvfp4: configs/numerics/nvfp4
|
||||
kv_fp8_cast: configs/ptq/units/kv_fp8_cast
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: >-
|
||||
GLM-5.3-Flash: NVFP4 (W4A4) on the routed experts and on the dense MLP of layers 0-2,
|
||||
plus an FP8 KV cache in cast mode using constant amax; max calibration. Shared experts,
|
||||
|
||||
+32
-1
@@ -153,6 +153,12 @@ of the body scheme. Quantizing the KV cache reduces memory at long context.
|
||||
- **`kv_fp8`** — FP8 E4M3 KV cache with **calibrated** per-tensor amax. The KV
|
||||
scales are measured during the calibration pass. Hopper+.
|
||||
|
||||
- **`kv_fp16`** — the KV cache is **not** quantized; it stays at the model's
|
||||
activation dtype (BF16/FP16). Combine with any body scheme when the deployment
|
||||
stack does not consume an FP8 KV cache, or when long-context memory is not the
|
||||
binding constraint. Some published checkpoints ship this way, so a body scheme
|
||||
gains a `-kv_fp16` variant when one of them needs it.
|
||||
|
||||
> **`kv_fp8_cast` vs `kv_fp8`:** both produce an FP8 KV cache. `_cast` uses a
|
||||
> fixed scale and skips the KV calibration step (faster, no extra data
|
||||
> dependence); plain `kv_fp8` calibrates the scale from data. The cast version
|
||||
@@ -269,7 +275,7 @@ The general recipes above are **model-agnostic**: they select layers by wildcard
|
||||
(`*mlp*`, `*self_attn*`, `*[kv]_bmm_quantizer`) and lean on the shared
|
||||
`default_disabled_quantizers` exclusions, so the same file works on any
|
||||
architecture whose module names follow the usual conventions. A recipe only
|
||||
earns a place under `model_type/<model_type>/` or
|
||||
earns a **body** under `model_type/<model_type>/` or
|
||||
`models/<org>/<checkpoint>/` when a model has to **deviate** from
|
||||
that baseline. The deviations come in four kinds:
|
||||
|
||||
@@ -288,6 +294,15 @@ The numerics and standard exclusions are still inherited from `configs/`
|
||||
wherever possible — the model folder captures *only* the delta. Each `<task>/`
|
||||
folder may carry a `README.md` spelling out that delta.
|
||||
|
||||
> **Not every `models/<org>/<checkpoint>/` folder holds a deviation.** When a
|
||||
> general (or `model_type/<model_type>`) recipe already produces a released
|
||||
> checkpoint's scheme, the folder holds a thin **alias** that imports that recipe
|
||||
> wholesale and overrides only `metadata` — no duplicated `quant_cfg` — so the
|
||||
> recipe is reachable from the checkpoint's own hub path. [Checkpoint
|
||||
> mirrors](#checkpoint-mirrors--modelsorgcheckpoint) covers the folders that
|
||||
> genuinely deviate and aliases the ones that don't; see
|
||||
> [`models/README.md`](models/README.md) for the alias format and when to write one.
|
||||
|
||||
### Architecture-aware `quant_cfg` — `minimax_m3_vl`, `qwen3_vl`, `qwen3_5`, `qwen3_5_moe`, `qwen3_6_moe`, `vit`, `nemotron_llama`
|
||||
|
||||
**`minimax_m3_vl/ptq/mxfp8_nvfp4_experts`** applies MXFP8 to the language-model
|
||||
@@ -510,6 +525,22 @@ is the exception: its deviation is a model-specific *scope* — a wildcard schem
|
||||
plus a load-bearing vision-tower exclusion and the VLM-required
|
||||
`layerwise.enable=false` — rather than a per-component precision map.
|
||||
|
||||
### Checkpoint aliases — `models/<org>/<checkpoint>`
|
||||
|
||||
Some published checkpoints use a scheme a portable recipe already produces, so their
|
||||
entry is a thin **alias** that imports that recipe wholesale and overrides only
|
||||
`metadata` (see [`models/README.md`](models/README.md) for the format):
|
||||
|
||||
- **`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`** aliases the general
|
||||
`general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` — expert-only NVFP4 with MSE-swept static
|
||||
weight scales and dynamic inputs, plus an FP8 KV cache in cast mode — as published in
|
||||
`nvidia/Kimi-K2.6-NVFP4`.
|
||||
- **`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`** aliases the
|
||||
`qwen3_5_moe` architecture recipe
|
||||
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4 (MSE static weights)
|
||||
on the routed experts, ModelOpt-default FP8 elsewhere, and an FP8 KV cache — as published in
|
||||
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
|
||||
|
||||
---
|
||||
|
||||
For the full catalog and how to pick a starting recipe for a given model, see
|
||||
|
||||
@@ -21,7 +21,6 @@ imports:
|
||||
w8a8_fp8_fp8: configs/ptq/units/w8a8_fp8_fp8
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: FP8 ResNet PTQ with quantized residual connections.
|
||||
|
||||
quantize:
|
||||
|
||||
@@ -21,7 +21,6 @@ imports:
|
||||
int8_per_channel: configs/numerics/int8_per_channel
|
||||
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: INT8 ResNet PTQ with quantized residual connections.
|
||||
|
||||
quantize:
|
||||
|
||||
@@ -26,8 +26,10 @@ from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
import modelopt.recipe.loader
|
||||
import modelopt.torch.quantization.config as qcfg
|
||||
from modelopt.recipe.config import (
|
||||
RECIPE_TYPE_TO_CLASS,
|
||||
AutoQuantizeConfig,
|
||||
AutoQuantizeConstraints,
|
||||
AutoQuantizeCost,
|
||||
@@ -35,14 +37,23 @@ from modelopt.recipe.config import (
|
||||
ModelOptDFlashRecipe,
|
||||
ModelOptEagleRecipe,
|
||||
ModelOptPTQRecipe,
|
||||
RecipeMetadataConfig,
|
||||
RecipeType,
|
||||
)
|
||||
from modelopt.recipe.loader import _apply_dotlist, _resolve_recipe_path, load_config, load_recipe
|
||||
from modelopt.recipe.loader import (
|
||||
_apply_dotlist,
|
||||
_peek_recipe_type,
|
||||
_resolve_recipe_path,
|
||||
load_config,
|
||||
load_recipe,
|
||||
)
|
||||
from modelopt.torch.opt.config_loader import (
|
||||
_MODELOPT_SCHEMA_RE,
|
||||
_alias_builtin_recipe_prefix,
|
||||
_load_raw_config,
|
||||
_resolve_config_path,
|
||||
_schema_type,
|
||||
peek_declared_schema,
|
||||
)
|
||||
from modelopt.torch.quantization.config import QuantizerAttributeConfig, normalize_quant_cfg_list
|
||||
from modelopt.torch.quantization.mode import CalibrateModeRegistry, get_modelike_from_algo_cfg
|
||||
@@ -301,8 +312,17 @@ def test_import_resolution_honors_huggingface_alias():
|
||||
assert str(_resolve_config_path(old)) == str(_resolve_config_path(canonical))
|
||||
|
||||
|
||||
_PTQ_SCHEMA = "modelopt.recipe.config.ModelOptPTQRecipe"
|
||||
|
||||
|
||||
def _all_shipped_ptq_recipe_paths():
|
||||
"""Every shipped PTQ recipe, discovered from disk rather than a hardcoded list."""
|
||||
"""Every shipped PTQ recipe, discovered from disk rather than a hardcoded list.
|
||||
|
||||
A recipe says it is PTQ with ``metadata.recipe_type``, with a ``# modelopt-schema:``
|
||||
comment naming :class:`ModelOptPTQRecipe`, or -- for a checkpoint alias, which states
|
||||
neither -- by delegating to a recipe that does. All three are picked up, so every
|
||||
shipped recipe is swept by the tests below.
|
||||
"""
|
||||
root = files("modelopt_recipes")
|
||||
paths = []
|
||||
for path in sorted(Path(str(root)).rglob("*.yaml")):
|
||||
@@ -314,7 +334,10 @@ def _all_shipped_ptq_recipe_paths():
|
||||
# List-shaped fragments (layer-pattern units) are not recipes.
|
||||
if not isinstance(raw, dict):
|
||||
continue
|
||||
if (raw.get("metadata") or {}).get("recipe_type") == "ptq":
|
||||
# Ask the loader's own resolver rather than re-deriving the rules here: a
|
||||
# top-level ``$import`` does not imply PTQ (an alias of a speculative recipe
|
||||
# delegates the same way), and this cannot drift from what load_recipe does.
|
||||
if _peek_recipe_type(path) == RecipeType.PTQ:
|
||||
paths.append(str(rel.with_suffix("")))
|
||||
return paths
|
||||
|
||||
@@ -476,6 +499,471 @@ def test_load_recipe_unsupported_type_raises(tmp_path):
|
||||
load_recipe(bad)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# load_recipe — whole-recipe delegation (checkpoint aliases)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
_BASE_RECIPE_FOR_ALIAS = """\
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: the base recipe
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- quantizer_name: '*'
|
||||
enable: false
|
||||
- quantizer_name: '*weight_quantizer'
|
||||
cfg: {num_bits: 8, axis: 0}
|
||||
"""
|
||||
|
||||
|
||||
def _write_alias_pair(tmp_path, alias_body: str):
|
||||
"""Write a base recipe plus an alias that delegates to it; return the alias path."""
|
||||
(tmp_path / "base.yaml").write_text(_BASE_RECIPE_FOR_ALIAS)
|
||||
alias = tmp_path / "alias.yaml"
|
||||
alias.write_text(alias_body.format(base=tmp_path / "base.yaml"))
|
||||
return alias
|
||||
|
||||
|
||||
def test_load_recipe_delegates_whole_body_to_import(tmp_path):
|
||||
"""A top-level ``$import`` supplies the body; local keys override the imported ones.
|
||||
|
||||
This is what the checkpoint aliases under ``modelopt_recipes/models/`` rely on: they
|
||||
name the recipe behind a published checkpoint without copying its body.
|
||||
"""
|
||||
alias = _write_alias_pair(
|
||||
tmp_path,
|
||||
"""\
|
||||
imports:
|
||||
base: {base}
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: the alias
|
||||
""",
|
||||
)
|
||||
recipe = load_recipe(alias)
|
||||
assert recipe.description == "the alias"
|
||||
assert recipe.quantize.model_dump() == load_recipe(tmp_path / "base.yaml").quantize.model_dump()
|
||||
|
||||
|
||||
def test_load_recipe_delegating_alias_can_override_the_body(tmp_path):
|
||||
"""A delegating recipe may also replace an imported section outright."""
|
||||
alias = _write_alias_pair(
|
||||
tmp_path,
|
||||
"""\
|
||||
imports:
|
||||
base: {base}
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: overridden body
|
||||
quantize:
|
||||
algorithm: max
|
||||
quant_cfg:
|
||||
- quantizer_name: '*input_quantizer'
|
||||
enable: false
|
||||
""",
|
||||
)
|
||||
quant_cfg = load_recipe(alias).quantize.model_dump()["quant_cfg"]
|
||||
assert [entry["quantizer_name"] for entry in quant_cfg] == ["*input_quantizer"]
|
||||
|
||||
|
||||
def test_load_recipe_infers_kind_from_schema_comment(tmp_path):
|
||||
"""A recipe that declares its schema needs no ``metadata.recipe_type``."""
|
||||
recipe = tmp_path / "r.yaml"
|
||||
recipe.write_text(_BASE_RECIPE_FOR_ALIAS.replace(" recipe_type: ptq\n", ""))
|
||||
loaded = load_recipe(recipe)
|
||||
assert loaded.recipe_type == RecipeType.PTQ
|
||||
assert isinstance(loaded, ModelOptPTQRecipe)
|
||||
|
||||
|
||||
def test_load_recipe_infers_kind_from_the_recipe_it_delegates_to(tmp_path):
|
||||
"""A pure alias states neither ``recipe_type`` nor a schema; it inherits both.
|
||||
|
||||
This is the shape the checkpoint aliases under ``modelopt_recipes/models/`` use, so
|
||||
a released checkpoint's entry carries nothing but a description and the import.
|
||||
"""
|
||||
alias = _write_alias_pair(
|
||||
tmp_path,
|
||||
"""\
|
||||
imports:
|
||||
base: {base}
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
description: nothing but a description
|
||||
""",
|
||||
)
|
||||
loaded = load_recipe(alias)
|
||||
assert loaded.recipe_type == RecipeType.PTQ
|
||||
assert loaded.description == "nothing but a description"
|
||||
|
||||
|
||||
def test_load_recipe_rejects_recipe_type_contradicting_its_schema(tmp_path):
|
||||
"""Stating a kind that disagrees with the schema class is an error, not a preference."""
|
||||
recipe = tmp_path / "r.yaml"
|
||||
recipe.write_text(
|
||||
_BASE_RECIPE_FOR_ALIAS.replace(" recipe_type: ptq\n", " recipe_type: speculative_eagle\n")
|
||||
)
|
||||
with pytest.raises(ValueError, match="recipe_type"):
|
||||
load_recipe(recipe)
|
||||
|
||||
|
||||
def test_load_recipe_without_any_kind_declaration_raises(tmp_path):
|
||||
"""No schema comment, no recipe_type and no delegation: the loader cannot dispatch."""
|
||||
recipe = tmp_path / "r.yaml"
|
||||
recipe.write_text(
|
||||
_BASE_RECIPE_FOR_ALIAS.replace(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n", ""
|
||||
).replace(" recipe_type: ptq\n", "")
|
||||
)
|
||||
with pytest.raises(ValueError, match="does not say what kind of recipe it is"):
|
||||
load_recipe(recipe)
|
||||
|
||||
|
||||
_EAGLE_RECIPE_FOR_ALIAS = """\
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptEagleRecipe
|
||||
metadata:
|
||||
description: an eagle recipe
|
||||
eagle: {}
|
||||
"""
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("declaration", "label"),
|
||||
[
|
||||
("# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n", "schema comment"),
|
||||
("", "metadata.recipe_type"),
|
||||
],
|
||||
)
|
||||
def test_load_recipe_rejects_delegating_to_a_different_kind(tmp_path, declaration, label):
|
||||
"""Importing a recipe of another kind is an error, however the kinds were declared.
|
||||
|
||||
A top-level ``$import`` takes over the whole body, so a PTQ recipe importing an
|
||||
EAGLE one would splice an ``eagle`` section into a PTQ schema. Caught as a kind
|
||||
mismatch rather than left to surface as whatever pydantic makes of the result.
|
||||
"""
|
||||
(tmp_path / "base.yaml").write_text(_EAGLE_RECIPE_FOR_ALIAS)
|
||||
alias = tmp_path / "alias.yaml"
|
||||
metadata = "metadata:\n description: a ptq recipe\n"
|
||||
if not declaration: # declare the kind the other way instead
|
||||
metadata = "metadata:\n recipe_type: ptq\n description: a ptq recipe\n"
|
||||
alias.write_text(
|
||||
f"{declaration}imports:\n base: {tmp_path / 'base.yaml'}\n\n$import: base\n{metadata}"
|
||||
)
|
||||
with pytest.raises(ValueError, match=r"is a 'ptq' recipe but imports .*'speculative_eagle'"):
|
||||
load_recipe(alias)
|
||||
|
||||
|
||||
def test_load_recipe_allows_delegating_within_the_same_kind(tmp_path):
|
||||
"""The matching case still loads -- the check rejects mismatches, not delegation."""
|
||||
alias = _write_alias_pair(
|
||||
tmp_path,
|
||||
"""\
|
||||
# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe
|
||||
imports:
|
||||
base: {base}
|
||||
|
||||
$import: base
|
||||
metadata:
|
||||
recipe_type: ptq
|
||||
description: agrees on every axis
|
||||
""",
|
||||
)
|
||||
assert load_recipe(alias).recipe_type == RecipeType.PTQ
|
||||
|
||||
|
||||
def test_load_recipe_reuses_a_whole_recipe_with_no_metadata(tmp_path):
|
||||
"""A recipe can be nothing but an import: two lines, everything inherited.
|
||||
|
||||
With no inline keys to override with, ``metadata`` arrives from the base along with
|
||||
``quantize`` -- so this is genuine whole-recipe reuse, not just body reuse. Useful
|
||||
when a second path should resolve to an existing recipe verbatim; a checkpoint alias
|
||||
keeps its own description instead, so it can say which release it stands for.
|
||||
"""
|
||||
alias = _write_alias_pair(tmp_path, "imports:\n base: {base}\n\n$import: base\n")
|
||||
aliased, original = load_recipe(alias), load_recipe(tmp_path / "base.yaml")
|
||||
assert aliased.recipe_type == RecipeType.PTQ
|
||||
assert aliased.quantize.model_dump() == original.quantize.model_dump()
|
||||
assert aliased.metadata.model_dump() == original.metadata.model_dump()
|
||||
assert aliased.description == "the base recipe"
|
||||
|
||||
|
||||
def test_load_recipe_delegation_chain_inherits_the_kind(tmp_path):
|
||||
"""Kind resolution follows a chain of delegations, not just one hop."""
|
||||
(tmp_path / "base.yaml").write_text(_BASE_RECIPE_FOR_ALIAS)
|
||||
(tmp_path / "middle.yaml").write_text(
|
||||
f"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
f"imports:\n base: {tmp_path / 'base.yaml'}\n\n$import: base\n"
|
||||
)
|
||||
leaf = tmp_path / "leaf.yaml"
|
||||
leaf.write_text(
|
||||
f"imports:\n mid: {tmp_path / 'middle.yaml'}\n\n$import: mid\n"
|
||||
"metadata:\n description: two hops from the body\n"
|
||||
)
|
||||
loaded = load_recipe(leaf)
|
||||
assert loaded.recipe_type == RecipeType.PTQ
|
||||
assert loaded.description == "two hops from the body"
|
||||
assert loaded.quantize.model_dump() == load_recipe(tmp_path / "base.yaml").quantize.model_dump()
|
||||
|
||||
|
||||
def test_load_recipe_delegation_cycle_is_reported_not_hung(tmp_path):
|
||||
"""Two recipes that delegate to each other fail cleanly instead of recursing forever.
|
||||
|
||||
Neither states a kind, so resolution has to walk the import to find one and would
|
||||
loop without the cycle guard. A ``ValueError`` rather than a ``RecursionError`` is
|
||||
the assertion that the guard is doing its job.
|
||||
|
||||
The message has to name the cycle, not fall back to the generic "does not say what
|
||||
kind" text: that text's third remedy is "delegate to a recipe that does", which is
|
||||
exactly what the author already did, so it sends them in a circle.
|
||||
"""
|
||||
a, b = tmp_path / "a.yaml", tmp_path / "b.yaml"
|
||||
a.write_text(f"imports:\n other: {b}\n\n$import: other\n")
|
||||
b.write_text(f"imports:\n other: {a}\n\n$import: other\n")
|
||||
with pytest.raises(ValueError, match="delegates back to it") as excinfo:
|
||||
load_recipe(a)
|
||||
assert "cycle:" in str(excinfo.value)
|
||||
assert str(b) in str(excinfo.value)
|
||||
# The generic advice must not be what the author is left holding.
|
||||
assert "does not say what kind of recipe it is" not in str(excinfo.value)
|
||||
|
||||
|
||||
def test_load_recipe_delegates_via_a_list_of_imports(tmp_path):
|
||||
"""``$import`` accepts a list; the kind comes from the first entry that is a recipe."""
|
||||
(tmp_path / "base.yaml").write_text(_BASE_RECIPE_FOR_ALIAS)
|
||||
(tmp_path / "extra.yaml").write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"metadata:\n description: extra\n"
|
||||
"quantize:\n algorithm: max\n quant_cfg: []\n"
|
||||
)
|
||||
alias = tmp_path / "alias.yaml"
|
||||
alias.write_text(
|
||||
f"imports:\n base: {tmp_path / 'base.yaml'}\n extra: {tmp_path / 'extra.yaml'}\n\n"
|
||||
"$import: [base, extra]\nmetadata:\n description: merged\n"
|
||||
)
|
||||
loaded = load_recipe(alias)
|
||||
assert loaded.recipe_type == RecipeType.PTQ
|
||||
# Later imports win, matching the dict-merge semantics of a multi-name $import.
|
||||
assert loaded.quantize.quant_cfg == []
|
||||
|
||||
|
||||
def test_load_recipe_ignores_a_non_recipe_schema_comment_when_dispatching(tmp_path):
|
||||
"""A schema comment naming something that is not a recipe falls through to metadata.
|
||||
|
||||
Only the recipe schema classes identify a recipe kind; anything else means the file
|
||||
is a snippet as far as dispatch is concerned, so ``metadata.recipe_type`` still has
|
||||
to answer.
|
||||
"""
|
||||
recipe = tmp_path / "r.yaml"
|
||||
recipe.write_text(
|
||||
"# modelopt-schema: modelopt.torch.quantization.config.QuantizeConfig\n"
|
||||
"metadata:\n recipe_type: ptq\n description: d\n"
|
||||
"quantize:\n algorithm: max\n quant_cfg: []\n"
|
||||
)
|
||||
assert load_recipe(recipe).recipe_type == RecipeType.PTQ
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# peek_declared_schema
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_peek_declared_schema_reads_the_preamble(tmp_path):
|
||||
"""The declared schema path is returned without parsing or resolving the file."""
|
||||
f = tmp_path / "c.yaml"
|
||||
f.write_text(
|
||||
"# a comment\n"
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"metadata:\n description: d\n"
|
||||
)
|
||||
assert peek_declared_schema(f) == "modelopt.recipe.config.ModelOptPTQRecipe"
|
||||
|
||||
|
||||
def test_peek_declared_schema_returns_none_without_a_comment(tmp_path):
|
||||
f = tmp_path / "c.yaml"
|
||||
f.write_text("metadata:\n recipe_type: ptq\n")
|
||||
assert peek_declared_schema(f) is None
|
||||
|
||||
|
||||
def test_peek_declared_schema_ignores_a_comment_below_the_preamble(tmp_path):
|
||||
"""A comment after the first YAML line is not a declaration -- and must not look like one.
|
||||
|
||||
This is the shape ``test_shipped_modelopt_schema_comments_are_in_the_preamble``
|
||||
guards the shipped recipes against.
|
||||
"""
|
||||
f = tmp_path / "c.yaml"
|
||||
f.write_text(
|
||||
"metadata:\n recipe_type: ptq\n"
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
)
|
||||
assert peek_declared_schema(f) is None
|
||||
|
||||
|
||||
def test_peek_declared_schema_rejects_two_declarations(tmp_path):
|
||||
f = tmp_path / "c.yaml"
|
||||
f.write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptEagleRecipe\n"
|
||||
"metadata:\n description: d\n"
|
||||
)
|
||||
with pytest.raises(ValueError, match="multiple modelopt-schema"):
|
||||
peek_declared_schema(f)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# metadata.recipe_type is derived from the schema class
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
|
||||
def test_metadata_config_recipe_type_is_optional(tmp_path):
|
||||
"""``RecipeMetadataConfig`` on its own no longer requires a recipe_type."""
|
||||
f = tmp_path / "metadata.yaml"
|
||||
f.write_text("description: no kind stated\n")
|
||||
metadata = load_config(f, schema_type=RecipeMetadataConfig)
|
||||
assert metadata.recipe_type is None
|
||||
assert metadata.description == "no kind stated"
|
||||
|
||||
|
||||
#: The body section each recipe class requires, so the metadata behaviour can be checked
|
||||
#: on every kind rather than only the one with the simplest body.
|
||||
_MINIMAL_BODIES: dict[RecipeType, dict] = {
|
||||
RecipeType.PTQ: {"quantize": {"algorithm": "max", "quant_cfg": []}},
|
||||
RecipeType.AUTO_QUANTIZE: {
|
||||
"auto_quantize": {
|
||||
"constraints": {"effective_bits": 4.8},
|
||||
"candidate_formats": [
|
||||
{"quant_cfg": [{"quantizer_name": "*", "enable": False}]},
|
||||
{"quant_cfg": [{"quantizer_name": "*weight_quantizer", "enable": True}]},
|
||||
],
|
||||
}
|
||||
},
|
||||
RecipeType.SPECULATIVE_EAGLE: {}, # body sections have field defaults
|
||||
RecipeType.SPECULATIVE_DFLASH: {},
|
||||
RecipeType.SPECULATIVE_MEDUSA: {},
|
||||
}
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("recipe_type", "schema_class"),
|
||||
sorted(RECIPE_TYPE_TO_CLASS.items(), key=lambda kv: kv[0].value),
|
||||
)
|
||||
def test_recipe_class_fills_in_its_own_recipe_type(recipe_type, schema_class):
|
||||
"""Every recipe class knows its kind and fills ``metadata.recipe_type`` from it."""
|
||||
assert recipe_type == schema_class.RECIPE_TYPE
|
||||
recipe = schema_class.model_validate(
|
||||
{"metadata": {"description": "d"}, **_MINIMAL_BODIES[recipe_type]}
|
||||
)
|
||||
assert recipe.metadata.recipe_type == recipe_type
|
||||
assert recipe.recipe_type == recipe_type
|
||||
|
||||
|
||||
def test_recipe_class_rejects_a_contradicting_recipe_type():
|
||||
"""Stating the wrong kind is rejected at validation, not silently overwritten."""
|
||||
with pytest.raises(ValueError, match="recipe_type"):
|
||||
ModelOptPTQRecipe.model_validate(
|
||||
{
|
||||
"metadata": {"recipe_type": "speculative_eagle", "description": "d"},
|
||||
"quantize": {"algorithm": "max", "quant_cfg": []},
|
||||
}
|
||||
)
|
||||
|
||||
|
||||
def test_load_recipe_dir_without_recipe_type_or_schema_raises(tmp_path):
|
||||
"""A directory recipe with neither a schema comment nor recipe_type is rejected."""
|
||||
(tmp_path / "metadata.yml").write_text("description: no kind stated\n")
|
||||
(tmp_path / "quantize.yml").write_text("algorithm: max\nquant_cfg: []\n")
|
||||
with pytest.raises(ValueError, match="recipe_type"):
|
||||
load_recipe(tmp_path)
|
||||
|
||||
|
||||
def test_load_recipe_dir_from_schema_comment(tmp_path):
|
||||
"""metadata.yml can declare its kind with a schema comment instead of recipe_type."""
|
||||
(tmp_path / "metadata.yml").write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"description: Dir test via schema comment.\n"
|
||||
)
|
||||
(tmp_path / "quantize.yml").write_text("algorithm: max\nquant_cfg: []\n")
|
||||
recipe = load_recipe(tmp_path)
|
||||
assert recipe.recipe_type == RecipeType.PTQ
|
||||
assert recipe.metadata.recipe_type == RecipeType.PTQ
|
||||
|
||||
|
||||
def test_load_recipe_dir_schema_comment_and_recipe_type_must_agree(tmp_path):
|
||||
"""A directory recipe stating both must agree, same as a single-file recipe."""
|
||||
(tmp_path / "metadata.yml").write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"recipe_type: speculative_eagle\n"
|
||||
"description: Disagreement.\n"
|
||||
)
|
||||
(tmp_path / "quantize.yml").write_text("algorithm: max\nquant_cfg: []\n")
|
||||
with pytest.raises(ValueError, match="recipe_type"):
|
||||
load_recipe(tmp_path)
|
||||
|
||||
|
||||
def test_shipped_modelopt_schema_comments_are_in_the_preamble():
|
||||
"""A ``modelopt-schema`` comment below the first YAML line is silently ignored.
|
||||
|
||||
:func:`_parse_modelopt_schema` stops at the first non-comment line, so a comment
|
||||
placed after e.g. ``metadata:`` parses as absent -- the file looks annotated but is
|
||||
not importable and cannot be dispatched from. Catch that here rather than at the
|
||||
point some future recipe tries to ``$import`` it.
|
||||
"""
|
||||
root = Path(str(files("modelopt_recipes")))
|
||||
# The parser matches line by line, so its own pattern is not multiline; recompile it
|
||||
# here to scan whole files, keeping the accepted syntax identical to the parser's.
|
||||
_schema_comment_anywhere = re.compile(_MODELOPT_SCHEMA_RE.pattern, re.MULTILINE)
|
||||
ignored = [
|
||||
str(path.relative_to(root))
|
||||
for path in sorted(root.rglob("*.yaml"))
|
||||
# Use the parser's own pattern so this can't drift from what it accepts, and so
|
||||
# prose that merely mentions the comment is not mistaken for one.
|
||||
if not path.is_symlink()
|
||||
and _schema_comment_anywhere.search(path.read_text(encoding="utf-8"))
|
||||
and peek_declared_schema(path) is None
|
||||
]
|
||||
assert not ignored, (
|
||||
"These files carry a modelopt-schema comment that the parser cannot see; move it "
|
||||
f"above the first YAML line: {ignored}"
|
||||
)
|
||||
|
||||
|
||||
def test_load_recipe_delegating_alias_still_needs_a_body(tmp_path):
|
||||
"""Delegation relaxes the raw-YAML check, it does not remove the requirement."""
|
||||
(tmp_path / "empty.yaml").write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.RecipeMetadataConfig\n"
|
||||
"recipe_type: ptq\ndescription: not a full recipe\n"
|
||||
)
|
||||
alias = tmp_path / "alias.yaml"
|
||||
alias.write_text(
|
||||
f"imports:\n base: {tmp_path / 'empty.yaml'}\n\n"
|
||||
"$import: base\nmetadata:\n recipe_type: ptq\n description: alias\n"
|
||||
)
|
||||
with pytest.raises(ValueError, match="quantize"):
|
||||
load_recipe(alias)
|
||||
|
||||
|
||||
def test_load_recipe_import_of_recipe_without_schema_raises(tmp_path):
|
||||
"""An imported recipe must declare its ``modelopt-schema``, like any snippet."""
|
||||
(tmp_path / "base.yaml").write_text(
|
||||
_BASE_RECIPE_FOR_ALIAS.replace(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n", ""
|
||||
)
|
||||
)
|
||||
alias = tmp_path / "alias.yaml"
|
||||
alias.write_text(
|
||||
f"imports:\n base: {tmp_path / 'base.yaml'}\n\n"
|
||||
"$import: base\nmetadata:\n recipe_type: ptq\n description: alias\n"
|
||||
)
|
||||
with pytest.raises(ValueError, match="modelopt-schema"):
|
||||
load_recipe(alias)
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# load_recipe — directory format
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -2141,3 +2629,57 @@ def test_shipped_ptq_recipe_algorithm_config_constructs(recipe_path):
|
||||
algorithm = load_recipe(recipe_path).quantize.algorithm
|
||||
for mode_name, mode_cfg in get_modelike_from_algo_cfg(algorithm):
|
||||
CalibrateModeRegistry[mode_name].config_class(**mode_cfg)
|
||||
|
||||
|
||||
def test_recipe_loader_never_reads_a_file_with_the_locale_encoding():
|
||||
"""Every text read in the recipe loader must pin ``encoding=``.
|
||||
|
||||
Shipped recipes contain non-ASCII -- em dashes in several descriptions -- and a bare
|
||||
``read_text()`` decodes with the locale codepage. On a cp1252 machine that does not
|
||||
raise; it silently yields mojibake, and for two of the Nemotron-3-Super recipes the
|
||||
corruption lands in parsed *values* rather than in a comment the YAML parser drops. A
|
||||
behavioural test cannot catch this on our UTF-8 CI, because there the locale encoding
|
||||
IS utf-8 and a bare read behaves identically -- so this asserts the source-level
|
||||
invariant instead, which holds on every platform.
|
||||
"""
|
||||
import ast
|
||||
|
||||
source = Path(modelopt.recipe.loader.__file__)
|
||||
tree = ast.parse(source.read_text(encoding="utf-8"))
|
||||
unpinned = []
|
||||
for node in ast.walk(tree):
|
||||
if not isinstance(node, ast.Call):
|
||||
continue
|
||||
func = node.func
|
||||
name = func.attr if isinstance(func, ast.Attribute) else getattr(func, "id", None)
|
||||
if name not in {"read_text", "write_text", "open"}:
|
||||
continue
|
||||
# A binary open needs no encoding; anything textual does.
|
||||
if any(
|
||||
isinstance(a, ast.Constant) and isinstance(a.value, str) and "b" in a.value
|
||||
for a in node.args
|
||||
):
|
||||
continue
|
||||
if not any(kw.arg == "encoding" for kw in node.keywords):
|
||||
unpinned.append(f"{name}() at {source.name}:{node.lineno}")
|
||||
assert not unpinned, "these reads use the locale encoding instead of utf-8: " + ", ".join(
|
||||
unpinned
|
||||
)
|
||||
|
||||
|
||||
def test_load_recipe_round_trips_non_ascii_description(tmp_path):
|
||||
"""A recipe's non-ASCII text survives loading byte-for-byte.
|
||||
|
||||
The companion to the source-level check above: that one pins how we read, this one
|
||||
pins what comes out, so a future rewrite that keeps ``encoding=`` but mangles the text
|
||||
some other way still fails.
|
||||
"""
|
||||
recipe = tmp_path / "em_dash.yaml"
|
||||
recipe.write_text(
|
||||
"# modelopt-schema: modelopt.recipe.config.ModelOptPTQRecipe\n"
|
||||
"metadata:\n"
|
||||
" description: NVFP4 for Qwen3.5-VL — vision tower left in BF16\n"
|
||||
"quantize:\n algorithm: max\n quant_cfg: []\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
assert "—" in load_recipe(recipe).metadata.description
|
||||
|
||||
@@ -86,6 +86,48 @@ def test_general_ptq_recipe_count_in_ptq_md():
|
||||
)
|
||||
|
||||
|
||||
def test_documented_recipe_paths_resolve():
|
||||
"""Every ``general/ptq/<name>`` a doc names must exist on disk.
|
||||
|
||||
The existing checks run doc-from-disk: they catch a recipe that no doc mentions.
|
||||
This is the other direction -- a doc naming a recipe that was never landed, or was
|
||||
moved to another branch after the doc row was written. A recipe path is the
|
||||
user-facing ``--recipe`` interface, so a phantom row sends users to
|
||||
``Recipe path '...' is not a valid YAML file or directory``.
|
||||
"""
|
||||
docs = {
|
||||
"modelopt_recipes/ptq.md": PTQ_MD,
|
||||
"docs/source/guides/10_recipes.rst": Path(__file__).resolve().parents[3]
|
||||
/ "docs"
|
||||
/ "source"
|
||||
/ "guides"
|
||||
/ "10_recipes.rst",
|
||||
}
|
||||
missing = []
|
||||
for label, path in docs.items():
|
||||
if not path.is_file():
|
||||
continue
|
||||
# encoding= is required, not decorative: these docs contain non-ASCII (em dashes
|
||||
# among others) and a bare read_text() decodes with the locale codepage, which is
|
||||
# cp1252 on the Windows runners -- UnicodeDecodeError on the first such byte.
|
||||
for name in sorted(
|
||||
set(re.findall(r"general/ptq/([A-Za-z0-9._-]+)", path.read_text(encoding="utf-8")))
|
||||
):
|
||||
stem = name.removesuffix(".yaml").removesuffix(".yml")
|
||||
if (GENERAL_PTQ_DIR / f"{stem}.yaml").is_file():
|
||||
continue
|
||||
# Prose also names a recipe *family* -- e.g. ``general/ptq/nvfp4_mlp_only``
|
||||
# standing for its -kv_* variants -- which is not a phantom path.
|
||||
if any(GENERAL_PTQ_DIR.glob(f"{stem}-*.yaml")):
|
||||
continue
|
||||
missing.append(f"{label} -> general/ptq/{name}")
|
||||
assert not missing, (
|
||||
"Docs name general/ptq recipes that do not exist on disk:\n "
|
||||
+ "\n ".join(missing)
|
||||
+ "\nAdd the recipe, or remove the row if it belongs to a different change."
|
||||
)
|
||||
|
||||
|
||||
def test_every_model_specific_ptq_dir_is_mentioned():
|
||||
"""Every model-specific PTQ recipe must be identifiable in ptq.md.
|
||||
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""CI check: every models/<org>/<model_id>/ptq/ entry names a real Hugging Face Hub repo.
|
||||
|
||||
Not a pre-commit hook -- this makes one network call per entry, and pre-commit hooks
|
||||
(tools/precommit/) have to stay usable offline. Wired into CI instead; see
|
||||
.github/workflows/check_model_hub_orgs.yml, which only runs it when
|
||||
modelopt_recipes/models/** changes.
|
||||
|
||||
A recipe's own text can't catch a wrong org -- it's typically self-consistent with
|
||||
the (wrong) directory it lives in, since both were written by the same mistake. Hub
|
||||
existence is the one fact that isn't derivable from the recipe itself.
|
||||
|
||||
Scoped to the ptq/ task, matching tests/unit/recipe/test_recipe_docs.py's
|
||||
test_every_model_specific_ptq_dir_is_mentioned: a models/<org>/<model_id>/ entry with
|
||||
only e.g. an auto_quantize/ recipe (Muse-Glimmer-30B) is exempt from the ptq.md
|
||||
membership check for the same reason it should be exempt here -- models/README.md
|
||||
allows a checkpoint mirror for a *planned*, not-yet-released checkpoint, and ptq/ is
|
||||
where the backfilled, already-published aliases this check targets actually live.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
from huggingface_hub import HfApi
|
||||
from huggingface_hub.utils import RepositoryNotFoundError
|
||||
|
||||
MODELS_ROOT = Path(__file__).resolve().parents[2] / "modelopt_recipes" / "models"
|
||||
|
||||
|
||||
def _checkpoint_ids() -> list[str]:
|
||||
"""<org>/<model_id> for every models/<org>/<model_id>/ptq/ entry, deduplicated."""
|
||||
ids = {
|
||||
f"{p.parent.parent.parent.name}/{p.parent.parent.name}"
|
||||
for p in MODELS_ROOT.glob("*/*/ptq/*.yaml")
|
||||
}
|
||||
return sorted(ids)
|
||||
|
||||
|
||||
def _error_if_missing(api: HfApi, repo_id: str) -> str | None:
|
||||
try:
|
||||
api.model_info(repo_id)
|
||||
except RepositoryNotFoundError:
|
||||
return (
|
||||
f"{repo_id}: no such Hugging Face Hub repo. modelopt_recipes/models/<org>/"
|
||||
"<model_id>/ must be keyed by the SOURCE checkpoint's hub path (what you pass "
|
||||
"to from_pretrained(...)), not a quantized derivative published elsewhere -- "
|
||||
"see modelopt_recipes/models/README.md."
|
||||
)
|
||||
return None
|
||||
|
||||
|
||||
def main() -> int:
|
||||
"""Verify every checkpoint id against the Hub, exit 1 if any is missing."""
|
||||
checkpoint_ids = _checkpoint_ids()
|
||||
api = HfApi()
|
||||
errors = [e for e in (_error_if_missing(api, r) for r in checkpoint_ids) if e is not None]
|
||||
|
||||
if errors:
|
||||
for e in errors:
|
||||
print(f"ERROR: {e}", file=sys.stderr)
|
||||
return 1
|
||||
print(f"Verified {len(checkpoint_ids)} models/<org>/<model_id>/ptq/ hub paths.")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -32,6 +32,7 @@ Checks performed:
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import re
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
@@ -39,13 +40,65 @@ import yaml
|
||||
|
||||
_YAML_PARSE_ERROR = object()
|
||||
|
||||
# Recipe types this hook validates via load_recipe(). Mirrors RecipeType in
|
||||
# modelopt.recipe.config; kept as a literal set so the hook can run without
|
||||
# importing modelopt (which is also why _try_load_recipe gates on ImportError).
|
||||
# Recipe types reached through the LEGACY metadata.recipe_type path only. A recipe that
|
||||
# declares a ``# modelopt-schema:`` comment, or that delegates with ``$import``, is
|
||||
# validated whether or not its kind appears here -- _is_recipe_file returns True on those
|
||||
# branches before this set is consulted. So a new kind that declares a schema (the
|
||||
# recommended form) needs no change here; only a new kind still using the deprecated
|
||||
# metadata.recipe_type would. Mirrors RecipeType in modelopt.recipe.config; kept as a
|
||||
# literal set so the hook can run without importing modelopt (which is also why
|
||||
# _try_load_recipe gates on ImportError).
|
||||
_SUPPORTED_RECIPE_TYPES = frozenset(
|
||||
{"ptq", "speculative_eagle", "speculative_dflash", "speculative_medusa"}
|
||||
)
|
||||
|
||||
# A recipe usually declares its kind with a ``# modelopt-schema:`` comment naming its
|
||||
# schema class rather than with ``metadata.recipe_type`` (see modelopt/recipe/loader.py).
|
||||
# Matched here by name so the hook keeps working without importing modelopt.
|
||||
_SCHEMA_COMMENT_RE = re.compile(
|
||||
r"^\s*#\s*modelopt-schema:\s*modelopt\.recipe\.config\.ModelOpt\w+Recipe\s*$",
|
||||
re.MULTILINE,
|
||||
)
|
||||
|
||||
# Any ``# modelopt-schema:`` declaration, recipe or not. The reusable snippets under
|
||||
# ``modelopt_recipes/configs/`` declare non-recipe schemas -- QuantizerAttributeConfig,
|
||||
# LayerPatternList and friends -- and a snippet is allowed a top-level ``$import`` of its
|
||||
# own, which would otherwise make it indistinguishable here from a delegating alias.
|
||||
_ANY_SCHEMA_COMMENT_RE = re.compile(
|
||||
r"^\s*#\s*modelopt-schema:\s*\S+\s*$",
|
||||
re.MULTILINE,
|
||||
)
|
||||
|
||||
|
||||
def _declares_recipe_schema(path: Path) -> bool:
|
||||
"""Whether *path* names one of the recipe schema classes in a ``# modelopt-schema:`` comment.
|
||||
|
||||
Searched over the whole file, deliberately laxer than the loader's
|
||||
``_parse_modelopt_schema``, which stops at the first non-comment line. A file carrying
|
||||
the comment *below* its YAML body is therefore a recipe to this hook and not to the
|
||||
loader -- which is the outcome we want: the hook hands it to ``load_recipe``, which
|
||||
rejects it with "does not say what kind of recipe it is" rather than the file being
|
||||
skipped silently. ``test_shipped_modelopt_schema_comments_are_in_the_preamble`` keeps
|
||||
the shipped tree free of that shape.
|
||||
"""
|
||||
try:
|
||||
return bool(_SCHEMA_COMMENT_RE.search(path.read_text(encoding="utf-8")))
|
||||
except OSError:
|
||||
return False
|
||||
|
||||
|
||||
def _declares_non_recipe_schema(path: Path) -> bool:
|
||||
"""Whether *path* declares a ``# modelopt-schema:`` that is not a recipe schema.
|
||||
|
||||
That is the signature of a reusable snippet (a quantizer attribute, a layer-pattern
|
||||
list), which ``load_recipe`` cannot load and should never be handed.
|
||||
"""
|
||||
try:
|
||||
text = path.read_text(encoding="utf-8")
|
||||
except OSError:
|
||||
return False
|
||||
return bool(_ANY_SCHEMA_COMMENT_RE.search(text)) and not bool(_SCHEMA_COMMENT_RE.search(text))
|
||||
|
||||
|
||||
def _check_quant_cfg(quant_cfg, label: str) -> list[str]:
|
||||
"""Validate quant_cfg format. *label* is used in error messages."""
|
||||
@@ -105,7 +158,7 @@ def _check_single_file_recipe(path: Path) -> list[str]:
|
||||
return [] # not a recipe file
|
||||
|
||||
metadata = data.get("metadata")
|
||||
if not isinstance(metadata, dict) or "recipe_type" not in metadata:
|
||||
if not isinstance(metadata, dict) and not _declares_recipe_schema(path):
|
||||
return [] # not a recipe file
|
||||
|
||||
if "ptq_cfg" in data:
|
||||
@@ -169,8 +222,12 @@ def _is_dir_recipe(dir_path: Path) -> bool:
|
||||
def _is_recipe_file(path: Path) -> bool:
|
||||
"""Return True if *path* looks like a recipe file that should be validated.
|
||||
|
||||
Covers PTQ + speculative-decoding (EAGLE/DFlash/Medusa) recipes; extend
|
||||
``_SUPPORTED_RECIPE_TYPES`` for new types (e.g. QAT).
|
||||
Three ways in, checked in this order: a ``# modelopt-schema:`` comment, a top-level
|
||||
``$import`` (a delegating alias, whose kind comes from what it imports), and finally
|
||||
the deprecated ``metadata.recipe_type`` gated on ``_SUPPORTED_RECIPE_TYPES``. Only
|
||||
that last branch consults the set, so a recipe of any kind that declares a schema is
|
||||
validated here -- including kinds deliberately absent from the set, such as
|
||||
``auto_quantize``. ``load_recipe`` handles those, so this is intended.
|
||||
|
||||
Malformed or unparseable files return True so that ``load_recipe()`` can
|
||||
report the actual error.
|
||||
@@ -180,6 +237,20 @@ def _is_recipe_file(path: Path) -> bool:
|
||||
return True # let load_recipe report the parse error
|
||||
if not isinstance(data, dict):
|
||||
return False # not a recipe file at all
|
||||
if _declares_recipe_schema(path):
|
||||
return True
|
||||
if _declares_non_recipe_schema(path):
|
||||
# A snippet, not a recipe -- and snippets may carry a top-level ``$import`` of
|
||||
# their own (see ``test_import_cross_file_same_name_no_conflict``). Without this
|
||||
# the next branch would claim it and ``load_recipe`` would reject it with "does
|
||||
# not say what kind of recipe it is", which is a confusing way to learn that a
|
||||
# fragment was never meant to be loaded as a recipe.
|
||||
return False
|
||||
if "$import" in data:
|
||||
# A delegating alias declares neither a schema comment nor a recipe_type: its
|
||||
# kind comes from the recipe it imports. Validate it so a typo in ``imports:``
|
||||
# or a ``$import`` naming an undeclared import fails here rather than at use.
|
||||
return True
|
||||
metadata = data.get("metadata")
|
||||
if not isinstance(metadata, dict) or "recipe_type" not in metadata:
|
||||
return False # not a recipe file at all
|
||||
|
||||
Reference in New Issue
Block a user