Files
Model-Optimizer/plugins/modelopt/skills/quant-recipe-search/references/recipe_iteration.md
T
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00

12 KiB

Recipe Iteration Reference

Problem

Quantization recipe search is a constrained optimization loop:

  • Improve the user's chosen objective: compute/throughput, memory/latency, or a custom score.
  • Keep every benchmark within the accepted accuracy-loss threshold. Default: less than 1 percentage point versus the matching BF16/FP16 baseline.
  • Keep accuracy, verbosity/token usage, runtime behavior, and active cost separate until the final decision.
  • Treat generated checkpoints as candidates. Evaluation and comparison decide whether a candidate is useful.

Ask for missing objective, primary quantization family, and benchmark set before planning candidates.

Search Space

Write the search space before launching PTQ. A recipe is a combination of these axes, not just a numeric format.

Axis Choices To Try Notes
Numeric format FP8/W8A8, NVFP4/W4A4, W4A16 NVFP4, INT4/AWQ, mixed formats Keep FP8/W8A8 as a near-lossless baseline unless it is the user's target.
Calibration/search algorithm Max, MSE, GPTQ, AWQ, AutoQuant scoring, calibration data variants Algorithm choice is independent from numeric format.
Selection method Manual/heuristic, sensitivity-guided manual, AutoQuant, hybrid Record how each candidate was selected.
Module family Attention, MLP, MoE experts, routers/gates, embeddings, lm_head, adapters, vision encoders Change one major family at a time for ablations.
Layer position First/last transformer-layer counts or explicit ordinal ranges kept in BF16 Treat as a controlled ablation, not a default. Start with first 3-4 and last 1-2 layers when evidence supports it.
Runtime constraints Fused attention groups, fused MoE expert projections, backend-supported formats Do not mix incompatible quantization inside a fused runtime group.
Calibration budget Dataset mix, sample count, sequence length, batch size Vary deliberately and record the budget.

Objective Axis

Use the user's objective to decide which candidates are worth testing.

  • Compute / throughput: typical data-center target. Favor activation quantization such as NVFP4 or FP8 when the runtime has fast kernels.
  • Memory / latency: typical edge or memory-pressure target. Favor W4A16 or weight-only recipes when they preserve accuracy. Prefer active bytes per forward/decode path over total checkpoint size for routed or sparse models.
  • Custom: use the user-provided score, for example checkpoint size, latency at a fixed concurrency, or product-specific memory budget.

If the user needs multiple objectives, maintain separate tables or define an explicit weighted score.

Numeric Format Axis

Common starting families:

  • FP8/W8A8: near-lossless baseline or explicit primary target.
  • NVFP4/W4A4: low-bit candidate family when activation quantization is part of the target.
  • W4A16 NVFP4: weight-only NVFP4 family for accuracy-preserving memory/latency searches.
  • INT4/AWQ: weight-only low-bit family for low-batch memory/latency targets.
  • Mixed formats: examples include NVFP4+FP8, W4A16 NVFP4 with FP8 attention, or model-specific recipe fragments.

KV-cache dtype, parser settings, token caps, and backend flags are runtime/eval controls unless the user explicitly makes them part of the recipe objective.

Calibration Algorithm Axis

Try calibration/search algorithms as independent recipe variants:

  • Max calibration: fast baseline for many FP8/NVFP4 formats.
  • MSE calibration: try when max calibration loses accuracy for low-bit weights or sensitive layers.
  • GPTQ: try for weight-quantized candidates where correction cost is acceptable.
  • AWQ: try for INT4 or other weight-only candidates.
  • AutoQuant scoring: use KL-divergence or gradient-based scoring when available to rank layers/modules and produce sensitivity reports.

Selection Method Axis

Choose which modules get which format by one of these methods:

  • Manual/heuristic: use prior experience, module-family cost, and controlled ablations.
  • Sensitivity-guided manual: generate or recover an AutoQuant sensitivity report, then protect sensitive families or quantize low-sensitivity high-cost families.
  • AutoQuant: search per-layer/per-module selections under constraints such as target bits, active-cost objective, or allowed formats. When AutoQuant is available, include at least one AutoQuant-generated candidate in the portfolio so its trade-off can be compared against manual recipes.
  • Hybrid: start from AutoQuant, then override known runtime constraints or known-sensitive fused groups manually.

Layer Position Axis

Use positional exclusions when sensitivity results or model behavior suggest that boundary transformer layers are unusually sensitive. Do not assume every model benefits from them.

  • Resolve ordinal positions from the model's transformer block sequence rather than assuming a model-specific module path.
  • Start with controlled candidates that keep the first 3 or 4 layers, the last 1 or 2 layers, or both ranges in BF16. Include zero exclusion as the control.
  • Change one boundary or count at a time when identifying the useful range.
  • For manual candidates, exclude the resolved block paths in the recipe. For AutoQuant or hybrid candidates, pass the positional exclusions into the selected implementation when supported; otherwise apply them as a manual override after selection and record that constraint.
  • Keep fused runtime groups internally compatible. Expand an exclusion to the complete fused group when the backend cannot mix precisions within it.
  • Compare each excluded-layer candidate with the same BF16 baseline, benchmarks, and acceptance threshold used for the rest of the search.

Design Workflow

  1. Recover existing evidence:

    • Result tables, checkpoints, recipe logs, AutoQuant states, sensitivity reports, and active jobs.
    • Use monitor, launching-evals, and compare-results for execution state and metric provenance.
  2. Define the target:

    • Objective, primary quantization family, benchmark set, accuracy-loss threshold, cost metric, and calibration budget.
    • Include scale storage and other quantization metadata in cost estimates.
  3. Pick baselines:

    • BF16/FP16 baseline.
    • FP8/W8A8 near-lossless baseline, unless FP8 is the final target.
    • Existing production recipe if one exists.
  4. Pick first candidates:

    • Start from modelopt_recipes when ModelOpt is available.
    • Prefer model-specific recipes, then general PTQ presets, then recipe fragments.
    • Add one AutoQuant candidate in the requested primary family when AutoQuant is available. Treat it as the expected best-search path, but validate it.
    • Add at least one manual or sensitivity-guided candidate for comparison and as a fallback if AutoQuant misses the benchmark frontier or produces a runtime-incompatible recipe.
    • Add positional BF16 exclusion candidates only when sensitivity or observed behavior warrants the ablation.
  5. Generate and validate:

    • Delegate checkpoint generation and validation to ptq.
    • Check checkpoint coverage, quantization metadata, and expected module coverage.
    • Pipe-clean serving only after checkpoint validation passes.
  6. Scale evaluation:

    • Run cheap screen benchmarks first.
    • Expand only candidates that pass screen evals and runtime gates.

Runtime Fusion Rules

Search by module family, but respect modules fused by the target runtime.

  • vLLM Qwen linear attention can fuse linear_attn.in_proj_qkv and linear_attn.in_proj_z into linear_attn.in_proj_qkvz; do not mix formats or algorithms across those shards unless the runtime supports it.
  • Fused MoE kernels can couple expert projections such as gate/up (w1/w3, or equivalent names); treat each fused expert group as one recipe unit unless deployment confirms mixed formats are supported.
  • Positional exclusions must preserve these grouping rules. If one boundary layer intersects a fused group, keep the whole group at a compatible precision or choose another boundary.
  • If a checkpoint is valid but deployment fails due to missing support, classify it as checkpoint-quality, recipe/runtime compatibility, or deployment implementation. For deployment implementation, try small patches or flags via deployment / debug before rejecting the recipe.

Iteration Loop

Use this loop after each candidate:

  1. Update the portfolio table with recipe axes, active cost, checkpoint path, eval logs, accuracy, verbosity, and decision.
  2. Compare against BF16/FP16 and FP8/W8A8 baselines.
  3. If accuracy drops:
    • Protect sensitive module families.
    • Try MSE, GPTQ, or AWQ variants.
    • Use AutoQuant sensitivity to choose manual overrides.
    • Test first-layer, last-layer, or combined BF16 exclusion candidates when sensitivity or model behavior points to boundary layers.
  4. If performance or active cost is insufficient:
    • Quantize the next high-cost active family.
    • Try a more aggressive format.
    • Revisit the active-cost objective or AutoQuant constraints.
  5. If verbosity changes:
    • Inspect output samples and generation stats.
    • Verify parser, token cap, sampling, backend, and KV-cache settings did not change.
  6. If results are close or noisy:
    • Rerun before labeling a benchmark regression.
  7. If AutoQuant gives repeated recipes:
    • Check achieved bits and recipe hashes.
    • Adjust objective, allowed formats, or constraints before larger sweeps.
  8. If AutoQuant underperforms manual recipes:
    • Compare the AutoQuant sensitivity report against manual ablation results.
    • Check whether AutoQuant protected high-active-cost modules, excluded the wrong families, optimized checkpoint size instead of active cost, or hit runtime-fusion constraints.
    • Keep the manual recipe in the table and use AutoQuant sensitivity to design the next hybrid/manual candidate.

Promote a recipe only when validated comparison shows it satisfies the user's objective and benchmark threshold.

Delegating To Existing Skills

Do not reimplement workflows that existing skills own:

Need Use
Generate/check a quantized checkpoint ptq
Serve a checkpoint or test backend flags deployment
Create or submit NEL configs evaluation
Resume/debug/analyze live eval runs launching-evals
Track active Slurm/NEL jobs monitor
Fetch MLflow artifacts accessing-mlflow
Compute baseline-vs-candidate deltas compare-results

Before launching PTQ in a ModelOpt repo, use the ptq skill; its current recipe paths and validation gates are authoritative.

ModelOpt Starting Points

When ModelOpt is available, start from modelopt_recipes:

  1. Check model-specific recipes first, for example modelopt_recipes/model_type/<model_family>/ptq/.
  2. Check general PTQ recipes and presets.
  3. Use recipe fragments to build controlled manual variants.
  4. Summarize include/exclude coverage before calibration. If a pattern misses the intended layer family, fix the recipe before launching.

Useful starting candidates:

  • Compute/throughput: FP8/W8A8, NVFP4/W4A4, mixed NVFP4+FP8 with activation quantization.
  • Memory/latency: W4A16 NVFP4, weight-only NVFP4, or W4A16 mixed with FP8 for sensitive modules.
  • MoE: experts-only or MLP-only recipes, then expand based on sensitivity and active-routing cost.

Candidate Record

For every candidate, record:

  • Objective and acceptance threshold.
  • Numeric formats and module-family coverage.
  • First/last BF16 exclusion counts or ordinal ranges.
  • Calibration/search algorithm and calibration data budget.
  • Selection method: manual, sensitivity-guided, AutoQuant, or hybrid.
  • Whether the candidate came from AutoQuant, manual ablation, or a hybrid override, so AutoQuant and manual trade-offs can be compared directly.
  • Runtime fusion assumptions.
  • Active bytes/token estimate including scales.
  • Checkpoint path and eval/log paths.
  • Accuracy and verbosity metrics.
  • Decision and next action.