mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
**Type of change:** New feature (recipe loading) + one bug fix
Two things, the second built on the first:
1. **A recipe can now reuse another recipe whole.** A top-level
`$import` brings in the imported recipe's entire body; keys given
alongside it override the imported ones. `metadata.recipe_type` becomes
optional and is deprecated along the way.
2. **The deprecated `recipe_type` is swept out of every shipped recipe,
and the checkpoint backfill starts with two published checkpoints
recorded as aliases** that reuse a portable recipe wholesale — the first
users of the alias mechanism — plus a fix to two existing Nemotron NVFP4
recipes.
#### Declaring what kind of recipe a file is
`load_recipe` read `metadata.recipe_type` out of the raw YAML *before*
resolving imports, because it needs the schema class to hand to
`load_config`. That made the field impossible to inherit, so a recipe
reusing another had to restate a line it could only have copied.
It is now optional, and the loader takes the first of these that
answers:
1. a `# modelopt-schema:` comment naming the recipe's schema class,
2. `metadata.recipe_type` — **deprecated**; still read and still
honoured, so a recipe outside this repo keeps working unchanged,
3. the recipe it delegates to via a top-level `$import`.
Whatever a recipe *does* state must be true, in both directions. A
schema comment contradicting a `recipe_type` is rejected, and so is a
recipe importing a different kind of recipe — that used to surface as
whatever pydantic made of, say, an `eagle` section spliced into a PTQ
schema. The concrete recipe classes carry a `RECIPE_TYPE` ClassVar as
the single source of truth.
Only a recipe that another file **imports** needs the schema comment —
that is what `$import` resolution requires to validate the payload. The
sweep here drops `metadata.recipe_type` from all 78 shipped recipes that
carried it and gives the imported ones a `# modelopt-schema:` comment
instead, so nothing in-tree depends on the deprecated field.
A directory recipe's `metadata.yml` resolves its kind the same way —
schema comment first, `recipe_type` as the fallback — it just has no
`$import` to delegate through, since a directory recipe has no body of
its own to hand off. (Follow-up commit, after this PR's initial review:
it originally still required `recipe_type` unconditionally, the one
place the deprecation didn't reach.)
#### Checkpoint aliases
Two checkpoints NVIDIA has published in quantized form use a scheme a
portable recipe already produces, with no checkpoint-specific deviation,
so each is recorded as a thin **alias** (top-level `$import`, overriding
only `metadata`) at its own model-hub path -- the *source* checkpoint's
path, not the published quantized one's:
-
**`models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast`**
delegates to `general/ptq/nvfp4_experts_only_mse-kv_fp8_cast` —
expert-only NVFP4 (MSE static weights, dynamic inputs) with an FP8 KV
cache in cast mode — published as `nvidia/Kimi-K2.6-NVFP4`.
-
**`models/Qwen/Qwen3.5-397B-A17B/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8`**
delegates to the `qwen3_5_moe` architecture recipe
`model_type/qwen3_5_moe/ptq/nvfp4_experts_mse-fp8_rest-kv_fp8` — NVFP4
(MSE static weights) on the routed experts, ModelOpt-default FP8
elsewhere, FP8 KV cache — published as
`nvidia/Qwen3.5-397B-A17B-NVFP4-V2`.
(Follow-up commit, after this PR's initial review: the Qwen entry
originally lived at `models/nvidia/Qwen3.5-397B-A17B/` -- nvidia is the
*published* checkpoint's org, not Qwen3.5-397B-A17B's own. Moved to
match the source model's actual hub path, same as the Kimi-K2.6 entry
above.)
Editing the base recipe changes every alias that points at it; nothing
is duplicated.
#### One fix
- **The Nemotron-3 Super and Ultra NVFP4 recipes** quantized the MTP
block on the **Megatron-Core** path, where it is a live `model.mtp`
submodule their broad `*mixer.*` patterns matched into, contrary to
their own descriptions. They now disable `mtp.*` explicitly. Hugging
Face runs were unaffected — `NemotronHPreTrainedModel` sets
`_keys_to_ignore_on_load_unexpected = [r"mtp.*"]` and builds no MTP
module.
### Usage
A checkpoint alias resolves through `--recipe` to the recipe it
delegates to:
```bash
python examples/hf_ptq/hf_ptq.py \
--pyt_ckpt_path <checkpoint> \
--recipe models/moonshotai/Kimi-K2.6/ptq/nvfp4_experts_only_mse-kv_fp8_cast \
--export_path <output>
```
A recipe that reuses another whole — the shape the aliases use:
```yaml
imports:
base: general/ptq/nvfp4_experts_only_mse-kv_fp8_cast
$import: base
metadata:
description: What this checkpoint uses the base recipe for.
```
### Testing
- **`tests/unit/recipe/test_loader.py`** — 28 new cases covering
whole-recipe reuse with no `metadata` at all; kind resolution from each
of the three sources, from a delegation chain and from a `$import` list;
a delegation cycle failing with `ValueError` rather than recursing;
`peek_declared_schema` including a comment placed below the first YAML
line; `recipe_type` being optional, filled per class, and rejected when
it contradicts; a directory recipe resolving its kind from a schema
comment the same way, rejecting a comment/`recipe_type` disagreement,
and still requiring one or the other; and delegating across kinds being
an error.
- **`tests/unit/recipe/test_recipe_docs.py`** — the
model-specific-recipe check now also covers the two new alias folders,
which must be listed in `ptq.md` like every other
`models/<org>/<model_id>` entry.
- **Recipe validation** (`tools/precommit/check_modelopt_recipes.py`)
and **`pre-commit`** pass on the changed files. The full
`tests/unit/recipe/` suite is left to CI — a broken `transformer_engine`
in the local dev venv keeps the `mtq.quantize`-based cases from running
there.
Not covered: **numerics**. Nothing here asserts accuracy, or that
running one of these recipes reproduces a released checkpoint's weights.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — `metadata.recipe_type` is
still read and honoured for recipes outside this repo, the schema
comments are inert for direct loads, and the loader change only relaxes
a check.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ — no new
dependencies.
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — two feature entries, one deprecation, and one bug fix under 0.48.0.
- Did you get Claude approval on this PR?: ❌ — not yet run.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Recipes can delegate configurations, support checkpoint aliases, and
apply local metadata overrides.
* Recipe types can be inferred from schema declarations or delegated
recipes, with stronger consistency validation.
* Added unquantized KV-cache options, layerwise export, broader operator
calibration, and new PTQ examples.
* Added checkpoint-specific recipes and MLflow experiment references.
* **Bug Fixes**
* Improved ONNX calibration, FSDP2 export, and fused-MoE quantization
handling.
* Nemotron-3 recipes keep MTP blocks in BF16.
* **Documentation**
* Expanded guidance for aliases, delegation, schema declarations, and
recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
151 lines
6.4 KiB
YAML
151 lines
6.4 KiB
YAML
# LiLiCorr speculative-decoding training recipe (arXiv:2608.20530).
|
|
#
|
|
# LiLiCorr reuses the DFlash mode/pipeline and adds a reranker over the candidate
|
|
# lattice the parallel backbone already produces, selected via
|
|
# dflash_architecture_config.projector_type=lilicorr. The backbone keeps its
|
|
# top-k candidates per slot; a small transformer scores transitions between them
|
|
# and serving commits a path greedily, left to right. Trained jointly with the
|
|
# backbone, so the drafter learns to propose candidates that correlate into longer
|
|
# accepted sequences.
|
|
#
|
|
# The objective is the DFlash block loss plus three weighted terms:
|
|
# loss = dflash_loss + w_ce*CE + w_margin*hinge + w_pen*penalty
|
|
# No outer multiplier, so `loss == origin_loss + lilicorr_loss` holds exactly.
|
|
# Online training is required (the penalty reads the target model's logits).
|
|
#
|
|
# This file is the published `base` variant. The `margin` variant differs only in
|
|
# the composition of the cross-entropy weight — see the two-line override below.
|
|
# Override fields via an OmegaConf dotlist.
|
|
|
|
# modelopt-schema: modelopt.recipe.config.ModelOptDFlashRecipe
|
|
metadata:
|
|
description: LiLiCorr training recipe (DFlash backbone + candidate-lattice reranker).
|
|
|
|
# maps to ModelArguments (main.py)
|
|
model:
|
|
model_name_or_path:
|
|
trust_remote_code: false
|
|
use_fake_base_for_offline: false
|
|
|
|
# maps to DataArguments (main.py)
|
|
data:
|
|
# Online only: the distractor penalty weights each candidate by the target
|
|
# model's own logit gap, so there has to be a target model in the process.
|
|
mode: online
|
|
data_path:
|
|
# Jinja chat template with {% generation %} tags for answer_only_loss.
|
|
chat_template:
|
|
|
|
# maps to TrainingArguments (main.py)
|
|
training:
|
|
# --- commonly modified ---
|
|
output_dir:
|
|
num_train_epochs: 6
|
|
per_device_train_batch_size: 1
|
|
gradient_accumulation_steps: 1
|
|
learning_rate: 6.0e-4
|
|
warmup_ratio: 0.04
|
|
training_seq_len: 3072
|
|
logging_steps: 50
|
|
save_steps: 1000
|
|
seed: 42
|
|
cp_size: 1
|
|
dp_shard_size: 1
|
|
disable_tqdm: true
|
|
# Keep off: eval runs the DFlash backbone only (the reranker is not applied in
|
|
# pseudo_speculative_generate), so AR here would report the backbone alone and
|
|
# understate the trained model. Compare via export + a serving benchmark.
|
|
estimate_ar: false
|
|
ar_validate_steps: 0
|
|
answer_only_loss: true
|
|
|
|
# --- rarely modified ---
|
|
do_eval: false
|
|
# Cosine, unlike the sibling recipes' linear: the published variants were trained
|
|
# under a linear warmup into a cosine decay (specforge's CosineAnnealingWarmupLR),
|
|
# and the schedule is part of the recipe those numbers came from.
|
|
lr_scheduler_type: cosine
|
|
save_strategy: steps
|
|
weight_decay: 0.0
|
|
max_grad_norm: 1.0
|
|
dataloader_drop_last: true
|
|
bf16: true
|
|
tf32: true
|
|
remove_unused_columns: false
|
|
ddp_timeout: 1800
|
|
report_to: tensorboard
|
|
|
|
# maps to DFlashConfig (modelopt/torch/speculative/config.py).
|
|
dflash:
|
|
dflash_block_size: 16
|
|
dflash_num_anchors: 512
|
|
dflash_use_torch_compile: false
|
|
# The reranker's terms are added to the plain weighted cross-entropy. Turning KD
|
|
# on would replace that base term and change the objective the published
|
|
# checkpoints were trained under.
|
|
dflash_self_logit_distillation: false
|
|
# Static exponential position decay, gamma=7 for block_size=16: early in-block
|
|
# slots gate acceptance, so they carry more weight. Not 'dpace' — the published
|
|
# variants were trained on the static decay.
|
|
dflash_loss_objective: decay
|
|
dflash_loss_decay_factor: 7.0
|
|
# Qwen3 has no native mask token; 151669 is an unused id used by the reference.
|
|
dflash_mask_token_id: 151669
|
|
|
|
# Objective composition. Absolute weights, validated all-or-nothing.
|
|
# `base` (this file): w_ce 0.25, w_margin 0.0
|
|
# `margin` : w_ce 0.125, w_margin 0.125
|
|
# Both keep w_pen 0.25, so the head's total weight is 0.50 either way and the
|
|
# variants differ only in how the cross-entropy block is split.
|
|
# fp32 master weights for the draft. Compute stays bf16 under autocast; this keeps the
|
|
# master copy -- and therefore AdamW's moments -- in fp32, which a bf16 second moment
|
|
# cannot represent the updates of. The published results for this recipe were trained
|
|
# with this on; turning it off changes the optimizer's arithmetic, not just its memory.
|
|
dflash_fp32_master_weights: true
|
|
|
|
dflash_lilicorr_w_ce: 0.25
|
|
dflash_lilicorr_w_margin: 0.0
|
|
dflash_lilicorr_w_pen: 0.25
|
|
# Hinge width, in units of the log-potential. Unused while w_margin is 0.
|
|
dflash_lilicorr_margin: 2.0
|
|
|
|
dflash_architecture_config:
|
|
num_hidden_layers: 5
|
|
# Draft attention/MLP dims — set explicitly (the draft is an independent
|
|
# Qwen3 model and does NOT inherit these from the base). GQA: 8 KV heads.
|
|
num_attention_heads: 32
|
|
num_key_value_heads: 8
|
|
head_dim: 128
|
|
intermediate_size: 12288
|
|
projector_type: lilicorr
|
|
|
|
# Reranker geometry. Every field is required, never defaulted:
|
|
# candidate_topk sets the lattice width and the shape of rank_embedding, and
|
|
# logit_scale/vector_eps change the score without changing any tensor shape —
|
|
# so a guessed value builds a head that loads cleanly and scores a different
|
|
# function. K is otherwise free; the method works at any k.
|
|
# dflash_init_checkpoint restores weights only and reads geometry from here, so
|
|
# warm-starting reproduces a head only if every field matches the one the
|
|
# checkpoint was trained with. lilicorr_logit_scale and lilicorr_vector_eps are
|
|
# the two a mistake would not be caught on, having no effect on any shape.
|
|
lilicorr_candidate_topk: 8
|
|
lilicorr_hidden_size: 1024
|
|
lilicorr_factor_dim: 1024
|
|
lilicorr_num_layers: 2
|
|
lilicorr_num_heads: 8
|
|
lilicorr_mlp_ratio: 2.0
|
|
# The factors are cosines, so this temperature sets their usable range.
|
|
lilicorr_logit_scale: 8.0
|
|
lilicorr_vector_eps: 1.0e-4
|
|
|
|
# `data.chat_template` is supplied per run; use the shared
|
|
# tools/launcher/examples/Qwen/Qwen3-8B/chat_template_train.jinja, as the other
|
|
# speculative recipes do, so variants stay comparable to each other.
|
|
#
|
|
# It differs slightly from the mask the published checkpoints were trained under: the
|
|
# reference leaves the empty `<think>` preamble out of the assistant span and supervises
|
|
# `<|im_end|>`, where the shared template does the opposite. The token ids are identical
|
|
# either way, so this is 6 tokens of supervision per record (4 preamble, 2 end-of-turn)
|
|
# and nothing else. Noted because it is invisible in the data, not because it is
|
|
# expected to matter at this scale.
|