Files
Model-Optimizer/plugins/modelopt/skills/quant-recipe-search/references/recipe_iteration.md
T
Shengliang Xu c7ed23a103 Rename modelopt_recipes/huggingface to model_type with backward-compat alias (#2328)
### What does this PR do?

Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.

Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.

- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
  source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
  `$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
  notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
  failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.

### Usage

```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none

# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
    --recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```

```python
from modelopt.recipe import load_recipe

load_recipe("model_type/vit/ptq/fp8")    # canonical
load_recipe("huggingface/vit/ptq/fp8")   # deprecated alias, resolves to the same recipe
```

### Testing

- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
  structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
  on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
  `model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
  the loader alias — including a recipe that pulls internal `$import`s.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.

### Additional Information

The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.

- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.

- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
  - Local recipe files now take precedence over built-in recipes.
  - Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
2026-09-15 12:16:12 -07:00

260 lines
12 KiB
Markdown

# Recipe Iteration Reference
## Problem
Quantization recipe search is a constrained optimization loop:
- Improve the user's chosen objective: compute/throughput, memory/latency, or a
custom score.
- Keep every benchmark within the accepted accuracy-loss threshold. Default:
less than 1 percentage point versus the matching BF16/FP16 baseline.
- Keep accuracy, verbosity/token usage, runtime behavior, and active cost
separate until the final decision.
- Treat generated checkpoints as candidates. Evaluation and comparison decide
whether a candidate is useful.
Ask for missing objective, primary quantization family, and benchmark set before
planning candidates.
## Search Space
Write the search space before launching PTQ. A recipe is a combination of these
axes, not just a numeric format.
| Axis | Choices To Try | Notes |
| --- | --- | --- |
| Numeric format | FP8/W8A8, NVFP4/W4A4, W4A16 NVFP4, INT4/AWQ, mixed formats | Keep FP8/W8A8 as a near-lossless baseline unless it is the user's target. |
| Calibration/search algorithm | Max, MSE, GPTQ, AWQ, AutoQuant scoring, calibration data variants | Algorithm choice is independent from numeric format. |
| Selection method | Manual/heuristic, sensitivity-guided manual, AutoQuant, hybrid | Record how each candidate was selected. |
| Module family | Attention, MLP, MoE experts, routers/gates, embeddings, `lm_head`, adapters, vision encoders | Change one major family at a time for ablations. |
| Layer position | First/last transformer-layer counts or explicit ordinal ranges kept in BF16 | Treat as a controlled ablation, not a default. Start with first 3-4 and last 1-2 layers when evidence supports it. |
| Runtime constraints | Fused attention groups, fused MoE expert projections, backend-supported formats | Do not mix incompatible quantization inside a fused runtime group. |
| Calibration budget | Dataset mix, sample count, sequence length, batch size | Vary deliberately and record the budget. |
### Objective Axis
Use the user's objective to decide which candidates are worth testing.
- **Compute / throughput:** typical data-center target. Favor activation
quantization such as NVFP4 or FP8 when the runtime has fast kernels.
- **Memory / latency:** typical edge or memory-pressure target. Favor W4A16 or
weight-only recipes when they preserve accuracy. Prefer active bytes per
forward/decode path over total checkpoint size for routed or sparse models.
- **Custom:** use the user-provided score, for example checkpoint size, latency
at a fixed concurrency, or product-specific memory budget.
If the user needs multiple objectives, maintain separate tables or define an
explicit weighted score.
### Numeric Format Axis
Common starting families:
- FP8/W8A8: near-lossless baseline or explicit primary target.
- NVFP4/W4A4: low-bit candidate family when activation quantization is part of
the target.
- W4A16 NVFP4: weight-only NVFP4 family for accuracy-preserving memory/latency
searches.
- INT4/AWQ: weight-only low-bit family for low-batch memory/latency targets.
- Mixed formats: examples include NVFP4+FP8, W4A16 NVFP4 with FP8 attention, or
model-specific recipe fragments.
KV-cache dtype, parser settings, token caps, and backend flags are runtime/eval
controls unless the user explicitly makes them part of the recipe objective.
### Calibration Algorithm Axis
Try calibration/search algorithms as independent recipe variants:
- Max calibration: fast baseline for many FP8/NVFP4 formats.
- MSE calibration: try when max calibration loses accuracy for low-bit weights
or sensitive layers.
- GPTQ: try for weight-quantized candidates where correction cost is acceptable.
- AWQ: try for INT4 or other weight-only candidates.
- AutoQuant scoring: use KL-divergence or gradient-based scoring when available
to rank layers/modules and produce sensitivity reports.
### Selection Method Axis
Choose which modules get which format by one of these methods:
- Manual/heuristic: use prior experience, module-family cost, and controlled
ablations.
- Sensitivity-guided manual: generate or recover an AutoQuant sensitivity report,
then protect sensitive families or quantize low-sensitivity high-cost families.
- AutoQuant: search per-layer/per-module selections under constraints such as
target bits, active-cost objective, or allowed formats. When AutoQuant is
available, include at least one AutoQuant-generated candidate in the portfolio
so its trade-off can be compared against manual recipes.
- Hybrid: start from AutoQuant, then override known runtime constraints or
known-sensitive fused groups manually.
### Layer Position Axis
Use positional exclusions when sensitivity results or model behavior suggest
that boundary transformer layers are unusually sensitive. Do not assume every
model benefits from them.
- Resolve ordinal positions from the model's transformer block sequence rather
than assuming a model-specific module path.
- Start with controlled candidates that keep the first 3 or 4 layers, the last
1 or 2 layers, or both ranges in BF16. Include zero exclusion as the control.
- Change one boundary or count at a time when identifying the useful range.
- For manual candidates, exclude the resolved block paths in the recipe. For
AutoQuant or hybrid candidates, pass the positional exclusions into the
selected implementation when supported; otherwise apply them as a manual
override after selection and record that constraint.
- Keep fused runtime groups internally compatible. Expand an exclusion to the
complete fused group when the backend cannot mix precisions within it.
- Compare each excluded-layer candidate with the same BF16 baseline, benchmarks,
and acceptance threshold used for the rest of the search.
## Design Workflow
1. Recover existing evidence:
- Result tables, checkpoints, recipe logs, AutoQuant states, sensitivity
reports, and active jobs.
- Use `monitor`, `launching-evals`, and `compare-results` for execution
state and metric provenance.
2. Define the target:
- Objective, primary quantization family, benchmark set, accuracy-loss
threshold, cost metric, and calibration budget.
- Include scale storage and other quantization metadata in cost estimates.
3. Pick baselines:
- BF16/FP16 baseline.
- FP8/W8A8 near-lossless baseline, unless FP8 is the final target.
- Existing production recipe if one exists.
4. Pick first candidates:
- Start from `modelopt_recipes` when ModelOpt is available.
- Prefer model-specific recipes, then general PTQ presets, then recipe
fragments.
- Add one AutoQuant candidate in the requested primary family when AutoQuant
is available. Treat it as the expected best-search path, but validate it.
- Add at least one manual or sensitivity-guided candidate for comparison and
as a fallback if AutoQuant misses the benchmark frontier or produces a
runtime-incompatible recipe.
- Add positional BF16 exclusion candidates only when sensitivity or observed
behavior warrants the ablation.
5. Generate and validate:
- Delegate checkpoint generation and validation to `ptq`.
- Check checkpoint coverage, quantization metadata, and expected module
coverage.
- Pipe-clean serving only after checkpoint validation passes.
6. Scale evaluation:
- Run cheap screen benchmarks first.
- Expand only candidates that pass screen evals and runtime gates.
## Runtime Fusion Rules
Search by module family, but respect modules fused by the target runtime.
- vLLM Qwen linear attention can fuse `linear_attn.in_proj_qkv` and
`linear_attn.in_proj_z` into `linear_attn.in_proj_qkvz`; do not mix formats or
algorithms across those shards unless the runtime supports it.
- Fused MoE kernels can couple expert projections such as gate/up (`w1`/`w3`, or
equivalent names); treat each fused expert group as one recipe unit unless
deployment confirms mixed formats are supported.
- Positional exclusions must preserve these grouping rules. If one boundary
layer intersects a fused group, keep the whole group at a compatible
precision or choose another boundary.
- If a checkpoint is valid but deployment fails due to missing support, classify
it as checkpoint-quality, recipe/runtime compatibility, or deployment
implementation. For deployment implementation, try small patches or flags via
`deployment` / `debug` before rejecting the recipe.
## Iteration Loop
Use this loop after each candidate:
1. Update the portfolio table with recipe axes, active cost, checkpoint path,
eval logs, accuracy, verbosity, and decision.
2. Compare against BF16/FP16 and FP8/W8A8 baselines.
3. If accuracy drops:
- Protect sensitive module families.
- Try MSE, GPTQ, or AWQ variants.
- Use AutoQuant sensitivity to choose manual overrides.
- Test first-layer, last-layer, or combined BF16 exclusion candidates when
sensitivity or model behavior points to boundary layers.
4. If performance or active cost is insufficient:
- Quantize the next high-cost active family.
- Try a more aggressive format.
- Revisit the active-cost objective or AutoQuant constraints.
5. If verbosity changes:
- Inspect output samples and generation stats.
- Verify parser, token cap, sampling, backend, and KV-cache settings did not
change.
6. If results are close or noisy:
- Rerun before labeling a benchmark regression.
7. If AutoQuant gives repeated recipes:
- Check achieved bits and recipe hashes.
- Adjust objective, allowed formats, or constraints before larger sweeps.
8. If AutoQuant underperforms manual recipes:
- Compare the AutoQuant sensitivity report against manual ablation results.
- Check whether AutoQuant protected high-active-cost modules, excluded the
wrong families, optimized checkpoint size instead of active cost, or hit
runtime-fusion constraints.
- Keep the manual recipe in the table and use AutoQuant sensitivity to design
the next hybrid/manual candidate.
Promote a recipe only when validated comparison shows it satisfies the user's
objective and benchmark threshold.
## Delegating To Existing Skills
Do not reimplement workflows that existing skills own:
| Need | Use |
| --- | --- |
| Generate/check a quantized checkpoint | `ptq` |
| Serve a checkpoint or test backend flags | `deployment` |
| Create or submit NEL configs | `evaluation` |
| Resume/debug/analyze live eval runs | `launching-evals` |
| Track active Slurm/NEL jobs | `monitor` |
| Fetch MLflow artifacts | `accessing-mlflow` |
| Compute baseline-vs-candidate deltas | `compare-results` |
Before launching PTQ in a ModelOpt repo, use the `ptq` skill; its current recipe
paths and validation gates are authoritative.
## ModelOpt Starting Points
When ModelOpt is available, start from `modelopt_recipes`:
1. Check model-specific recipes first, for example
`modelopt_recipes/model_type/<model_family>/ptq/`.
2. Check general PTQ recipes and presets.
3. Use recipe fragments to build controlled manual variants.
4. Summarize include/exclude coverage before calibration. If a pattern misses the
intended layer family, fix the recipe before launching.
Useful starting candidates:
- Compute/throughput: FP8/W8A8, NVFP4/W4A4, mixed NVFP4+FP8 with activation
quantization.
- Memory/latency: W4A16 NVFP4, weight-only NVFP4, or W4A16 mixed with FP8 for
sensitive modules.
- MoE: experts-only or MLP-only recipes, then expand based on sensitivity and
active-routing cost.
## Candidate Record
For every candidate, record:
- Objective and acceptance threshold.
- Numeric formats and module-family coverage.
- First/last BF16 exclusion counts or ordinal ranges.
- Calibration/search algorithm and calibration data budget.
- Selection method: manual, sensitivity-guided, AutoQuant, or hybrid.
- Whether the candidate came from AutoQuant, manual ablation, or a hybrid
override, so AutoQuant and manual trade-offs can be compared directly.
- Runtime fusion assumptions.
- Active bytes/token estimate including scales.
- Checkpoint path and eval/log paths.
- Accuracy and verbosity metrics.
- Decision and next action.