mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
### What does this PR do?
Type of change: Refactor + deprecation (recipe-library restructure,
backward compatible), plus an unrelated transformers-compat test fix.
Rename the architecture-specific recipe tier
`modelopt_recipes/huggingface/` to
`modelopt_recipes/model_type/`, making explicit that it holds recipes
**shared across
every checkpoint of a Hugging Face `model_type`** — as opposed to the
checkpoint-mirror
`models/<org>/<model_id>/` tier. The old `huggingface/` path keeps
working as a
deprecated backward-compat alias (a source-tree symlink plus a loader
alias), so no
saved `--recipe` path breaks.
- **Loader alias** (`modelopt/recipe/loader.py`): generalized so saved
`--recipe huggingface/<model_type>/...` paths rewrite to
`model_type/...`, alongside
the existing `huggingface/models/... -> models/...` rewrite (checked
first as the more
specific prefix). This keeps old paths resolving for pip-installed
wheels, where the
source-tree symlinks don't survive.
- **Internal `$import`s**: rewritten from `huggingface/... ->
model_type/...` inside the
shipped recipes so they resolve without the symlink — mandatory for
wheels, since
`$import` resolution goes through `config_loader` (no alias there).
- **Packaging** (`pyproject.toml`, `MANIFEST.in`): extended the
symlink-exclusion globs
so the recursive `**/*.yaml` package-data glob doesn't double-ship
recipes through the
`huggingface -> model_type` and `model_type/models -> ../models`
symlinks.
- **Docs / examples / skills / tests**: migrated all internal references
to the canonical
`model_type/`; `huggingface/` remains only in the deprecated-alias tests
and explanatory
notes.
- **Unrelated fix (2nd commit):**
`tests/unit/torch/export/test_quant_aware_conversion.py`
failed on transformers>=5.9, which dropped `base_model_prefix` from
`WeightTransform.__slots__` (the scoped-rule tests assigned it on the
now-slotted
object). Production `_scope_prefixes` already reads it via `getattr(...,
None)` and
degrades correctly, so there is no runtime change — the tests now set it
through a
helper that suppresses `AttributeError` across the supported
transformers range.
### Usage
```bash
# New canonical path
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe model_type/qwen3_vl/ptq/fp8_vision-kv_none
# Old path still works (deprecated backward-compat alias)
python examples/hf_ptq/hf_ptq.py --model <ckpt> \
--recipe huggingface/qwen3_vl/ptq/fp8_vision-kv_none
```
```python
from modelopt.recipe import load_recipe
load_recipe("model_type/vit/ptq/fp8") # canonical
load_recipe("huggingface/vit/ptq/fp8") # deprecated alias, resolves to the same recipe
```
### Testing
- `tests/unit/recipe/` — **336 passed**, including the new
`test_load_recipe_huggingface_arch_backward_compat_alias` and the
updated
structural/doc tests (`test_recipe_docs.py`).
- `tests/unit/torch/export/test_quant_aware_conversion.py` — **16
passed** (was 4 failed
on transformers 5.9.0).
- Built an sdist **and** a wheel and inspected both manifests: each
recipe ships exactly
once (29 `model_type/`, 13 `models/`, 2 `timm/`, 162 total) with
**zero** `huggingface/` or
`model_type/models/` duplicates and no build error on the symlinks.
- Simulated a wheel install (symlink-free extracted tree) and confirmed
`huggingface/<arch>/...`, `model_type/...`, and `huggingface/models/...`
all resolve via
the loader alias — including a recipe that pulls internal `$import`s.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅ — old `huggingface/...` recipe
paths keep resolving via the symlink + loader alias.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — backward-compat alias test
added; structural/doc tests updated to the new layout.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — Deprecations entry under 0.48.0. (The transformers-compat test fix
is not changelog-worthy.)
- Did you get Claude approval on this PR?: ❌ — not yet.
### Additional Information
The `model_type/models -> ../models` symlink is kept purely as a
backward-compat alias for
old `huggingface/models/<org>/<model_id>/...` paths; `model_type/` is
otherwise
architecture-only. If we ever want it strictly architecture-only, that
symlink can be
dropped later without breaking anything, since the loader rewrites
`huggingface/models/...`
straight to the top-level `models/` tier.
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **New Features**
- Added post-training quantization recipes for Gemma, Gemma 4,
MiniMax-M3, Nemotron, Qwen, Step-3.7, ViT, and other architectures.
- Added vision, multimodal, mixed-precision, and experts-only
quantization options.
- **Documentation**
- Standardized architecture-specific recipes under `model_type/` and
updated examples and guidance.
- **Compatibility**
- Legacy `huggingface/` recipe paths remain supported with deprecation
warnings.
- Local recipe files now take precedence over built-in recipes.
- Deprecated quantization-format flags warn when explicitly provided.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Shengliang Xu <shengliangx@nvidia.com>
260 lines
12 KiB
Markdown
260 lines
12 KiB
Markdown
# Recipe Iteration Reference
|
|
|
|
## Problem
|
|
|
|
Quantization recipe search is a constrained optimization loop:
|
|
|
|
- Improve the user's chosen objective: compute/throughput, memory/latency, or a
|
|
custom score.
|
|
- Keep every benchmark within the accepted accuracy-loss threshold. Default:
|
|
less than 1 percentage point versus the matching BF16/FP16 baseline.
|
|
- Keep accuracy, verbosity/token usage, runtime behavior, and active cost
|
|
separate until the final decision.
|
|
- Treat generated checkpoints as candidates. Evaluation and comparison decide
|
|
whether a candidate is useful.
|
|
|
|
Ask for missing objective, primary quantization family, and benchmark set before
|
|
planning candidates.
|
|
|
|
## Search Space
|
|
|
|
Write the search space before launching PTQ. A recipe is a combination of these
|
|
axes, not just a numeric format.
|
|
|
|
| Axis | Choices To Try | Notes |
|
|
| --- | --- | --- |
|
|
| Numeric format | FP8/W8A8, NVFP4/W4A4, W4A16 NVFP4, INT4/AWQ, mixed formats | Keep FP8/W8A8 as a near-lossless baseline unless it is the user's target. |
|
|
| Calibration/search algorithm | Max, MSE, GPTQ, AWQ, AutoQuant scoring, calibration data variants | Algorithm choice is independent from numeric format. |
|
|
| Selection method | Manual/heuristic, sensitivity-guided manual, AutoQuant, hybrid | Record how each candidate was selected. |
|
|
| Module family | Attention, MLP, MoE experts, routers/gates, embeddings, `lm_head`, adapters, vision encoders | Change one major family at a time for ablations. |
|
|
| Layer position | First/last transformer-layer counts or explicit ordinal ranges kept in BF16 | Treat as a controlled ablation, not a default. Start with first 3-4 and last 1-2 layers when evidence supports it. |
|
|
| Runtime constraints | Fused attention groups, fused MoE expert projections, backend-supported formats | Do not mix incompatible quantization inside a fused runtime group. |
|
|
| Calibration budget | Dataset mix, sample count, sequence length, batch size | Vary deliberately and record the budget. |
|
|
|
|
### Objective Axis
|
|
|
|
Use the user's objective to decide which candidates are worth testing.
|
|
|
|
- **Compute / throughput:** typical data-center target. Favor activation
|
|
quantization such as NVFP4 or FP8 when the runtime has fast kernels.
|
|
- **Memory / latency:** typical edge or memory-pressure target. Favor W4A16 or
|
|
weight-only recipes when they preserve accuracy. Prefer active bytes per
|
|
forward/decode path over total checkpoint size for routed or sparse models.
|
|
- **Custom:** use the user-provided score, for example checkpoint size, latency
|
|
at a fixed concurrency, or product-specific memory budget.
|
|
|
|
If the user needs multiple objectives, maintain separate tables or define an
|
|
explicit weighted score.
|
|
|
|
### Numeric Format Axis
|
|
|
|
Common starting families:
|
|
|
|
- FP8/W8A8: near-lossless baseline or explicit primary target.
|
|
- NVFP4/W4A4: low-bit candidate family when activation quantization is part of
|
|
the target.
|
|
- W4A16 NVFP4: weight-only NVFP4 family for accuracy-preserving memory/latency
|
|
searches.
|
|
- INT4/AWQ: weight-only low-bit family for low-batch memory/latency targets.
|
|
- Mixed formats: examples include NVFP4+FP8, W4A16 NVFP4 with FP8 attention, or
|
|
model-specific recipe fragments.
|
|
|
|
KV-cache dtype, parser settings, token caps, and backend flags are runtime/eval
|
|
controls unless the user explicitly makes them part of the recipe objective.
|
|
|
|
### Calibration Algorithm Axis
|
|
|
|
Try calibration/search algorithms as independent recipe variants:
|
|
|
|
- Max calibration: fast baseline for many FP8/NVFP4 formats.
|
|
- MSE calibration: try when max calibration loses accuracy for low-bit weights
|
|
or sensitive layers.
|
|
- GPTQ: try for weight-quantized candidates where correction cost is acceptable.
|
|
- AWQ: try for INT4 or other weight-only candidates.
|
|
- AutoQuant scoring: use KL-divergence or gradient-based scoring when available
|
|
to rank layers/modules and produce sensitivity reports.
|
|
|
|
### Selection Method Axis
|
|
|
|
Choose which modules get which format by one of these methods:
|
|
|
|
- Manual/heuristic: use prior experience, module-family cost, and controlled
|
|
ablations.
|
|
- Sensitivity-guided manual: generate or recover an AutoQuant sensitivity report,
|
|
then protect sensitive families or quantize low-sensitivity high-cost families.
|
|
- AutoQuant: search per-layer/per-module selections under constraints such as
|
|
target bits, active-cost objective, or allowed formats. When AutoQuant is
|
|
available, include at least one AutoQuant-generated candidate in the portfolio
|
|
so its trade-off can be compared against manual recipes.
|
|
- Hybrid: start from AutoQuant, then override known runtime constraints or
|
|
known-sensitive fused groups manually.
|
|
|
|
### Layer Position Axis
|
|
|
|
Use positional exclusions when sensitivity results or model behavior suggest
|
|
that boundary transformer layers are unusually sensitive. Do not assume every
|
|
model benefits from them.
|
|
|
|
- Resolve ordinal positions from the model's transformer block sequence rather
|
|
than assuming a model-specific module path.
|
|
- Start with controlled candidates that keep the first 3 or 4 layers, the last
|
|
1 or 2 layers, or both ranges in BF16. Include zero exclusion as the control.
|
|
- Change one boundary or count at a time when identifying the useful range.
|
|
- For manual candidates, exclude the resolved block paths in the recipe. For
|
|
AutoQuant or hybrid candidates, pass the positional exclusions into the
|
|
selected implementation when supported; otherwise apply them as a manual
|
|
override after selection and record that constraint.
|
|
- Keep fused runtime groups internally compatible. Expand an exclusion to the
|
|
complete fused group when the backend cannot mix precisions within it.
|
|
- Compare each excluded-layer candidate with the same BF16 baseline, benchmarks,
|
|
and acceptance threshold used for the rest of the search.
|
|
|
|
## Design Workflow
|
|
|
|
1. Recover existing evidence:
|
|
- Result tables, checkpoints, recipe logs, AutoQuant states, sensitivity
|
|
reports, and active jobs.
|
|
- Use `monitor`, `launching-evals`, and `compare-results` for execution
|
|
state and metric provenance.
|
|
|
|
2. Define the target:
|
|
- Objective, primary quantization family, benchmark set, accuracy-loss
|
|
threshold, cost metric, and calibration budget.
|
|
- Include scale storage and other quantization metadata in cost estimates.
|
|
|
|
3. Pick baselines:
|
|
- BF16/FP16 baseline.
|
|
- FP8/W8A8 near-lossless baseline, unless FP8 is the final target.
|
|
- Existing production recipe if one exists.
|
|
|
|
4. Pick first candidates:
|
|
- Start from `modelopt_recipes` when ModelOpt is available.
|
|
- Prefer model-specific recipes, then general PTQ presets, then recipe
|
|
fragments.
|
|
- Add one AutoQuant candidate in the requested primary family when AutoQuant
|
|
is available. Treat it as the expected best-search path, but validate it.
|
|
- Add at least one manual or sensitivity-guided candidate for comparison and
|
|
as a fallback if AutoQuant misses the benchmark frontier or produces a
|
|
runtime-incompatible recipe.
|
|
- Add positional BF16 exclusion candidates only when sensitivity or observed
|
|
behavior warrants the ablation.
|
|
|
|
5. Generate and validate:
|
|
- Delegate checkpoint generation and validation to `ptq`.
|
|
- Check checkpoint coverage, quantization metadata, and expected module
|
|
coverage.
|
|
- Pipe-clean serving only after checkpoint validation passes.
|
|
|
|
6. Scale evaluation:
|
|
- Run cheap screen benchmarks first.
|
|
- Expand only candidates that pass screen evals and runtime gates.
|
|
|
|
## Runtime Fusion Rules
|
|
|
|
Search by module family, but respect modules fused by the target runtime.
|
|
|
|
- vLLM Qwen linear attention can fuse `linear_attn.in_proj_qkv` and
|
|
`linear_attn.in_proj_z` into `linear_attn.in_proj_qkvz`; do not mix formats or
|
|
algorithms across those shards unless the runtime supports it.
|
|
- Fused MoE kernels can couple expert projections such as gate/up (`w1`/`w3`, or
|
|
equivalent names); treat each fused expert group as one recipe unit unless
|
|
deployment confirms mixed formats are supported.
|
|
- Positional exclusions must preserve these grouping rules. If one boundary
|
|
layer intersects a fused group, keep the whole group at a compatible
|
|
precision or choose another boundary.
|
|
- If a checkpoint is valid but deployment fails due to missing support, classify
|
|
it as checkpoint-quality, recipe/runtime compatibility, or deployment
|
|
implementation. For deployment implementation, try small patches or flags via
|
|
`deployment` / `debug` before rejecting the recipe.
|
|
|
|
## Iteration Loop
|
|
|
|
Use this loop after each candidate:
|
|
|
|
1. Update the portfolio table with recipe axes, active cost, checkpoint path,
|
|
eval logs, accuracy, verbosity, and decision.
|
|
2. Compare against BF16/FP16 and FP8/W8A8 baselines.
|
|
3. If accuracy drops:
|
|
- Protect sensitive module families.
|
|
- Try MSE, GPTQ, or AWQ variants.
|
|
- Use AutoQuant sensitivity to choose manual overrides.
|
|
- Test first-layer, last-layer, or combined BF16 exclusion candidates when
|
|
sensitivity or model behavior points to boundary layers.
|
|
4. If performance or active cost is insufficient:
|
|
- Quantize the next high-cost active family.
|
|
- Try a more aggressive format.
|
|
- Revisit the active-cost objective or AutoQuant constraints.
|
|
5. If verbosity changes:
|
|
- Inspect output samples and generation stats.
|
|
- Verify parser, token cap, sampling, backend, and KV-cache settings did not
|
|
change.
|
|
6. If results are close or noisy:
|
|
- Rerun before labeling a benchmark regression.
|
|
7. If AutoQuant gives repeated recipes:
|
|
- Check achieved bits and recipe hashes.
|
|
- Adjust objective, allowed formats, or constraints before larger sweeps.
|
|
8. If AutoQuant underperforms manual recipes:
|
|
- Compare the AutoQuant sensitivity report against manual ablation results.
|
|
- Check whether AutoQuant protected high-active-cost modules, excluded the
|
|
wrong families, optimized checkpoint size instead of active cost, or hit
|
|
runtime-fusion constraints.
|
|
- Keep the manual recipe in the table and use AutoQuant sensitivity to design
|
|
the next hybrid/manual candidate.
|
|
|
|
Promote a recipe only when validated comparison shows it satisfies the user's
|
|
objective and benchmark threshold.
|
|
|
|
## Delegating To Existing Skills
|
|
|
|
Do not reimplement workflows that existing skills own:
|
|
|
|
| Need | Use |
|
|
| --- | --- |
|
|
| Generate/check a quantized checkpoint | `ptq` |
|
|
| Serve a checkpoint or test backend flags | `deployment` |
|
|
| Create or submit NEL configs | `evaluation` |
|
|
| Resume/debug/analyze live eval runs | `launching-evals` |
|
|
| Track active Slurm/NEL jobs | `monitor` |
|
|
| Fetch MLflow artifacts | `accessing-mlflow` |
|
|
| Compute baseline-vs-candidate deltas | `compare-results` |
|
|
|
|
Before launching PTQ in a ModelOpt repo, use the `ptq` skill; its current recipe
|
|
paths and validation gates are authoritative.
|
|
|
|
## ModelOpt Starting Points
|
|
|
|
When ModelOpt is available, start from `modelopt_recipes`:
|
|
|
|
1. Check model-specific recipes first, for example
|
|
`modelopt_recipes/model_type/<model_family>/ptq/`.
|
|
2. Check general PTQ recipes and presets.
|
|
3. Use recipe fragments to build controlled manual variants.
|
|
4. Summarize include/exclude coverage before calibration. If a pattern misses the
|
|
intended layer family, fix the recipe before launching.
|
|
|
|
Useful starting candidates:
|
|
|
|
- Compute/throughput: FP8/W8A8, NVFP4/W4A4, mixed NVFP4+FP8 with activation
|
|
quantization.
|
|
- Memory/latency: W4A16 NVFP4, weight-only NVFP4, or W4A16 mixed with FP8 for
|
|
sensitive modules.
|
|
- MoE: experts-only or MLP-only recipes, then expand based on sensitivity and
|
|
active-routing cost.
|
|
|
|
## Candidate Record
|
|
|
|
For every candidate, record:
|
|
|
|
- Objective and acceptance threshold.
|
|
- Numeric formats and module-family coverage.
|
|
- First/last BF16 exclusion counts or ordinal ranges.
|
|
- Calibration/search algorithm and calibration data budget.
|
|
- Selection method: manual, sensitivity-guided, AutoQuant, or hybrid.
|
|
- Whether the candidate came from AutoQuant, manual ablation, or a hybrid
|
|
override, so AutoQuant and manual trade-offs can be compared directly.
|
|
- Runtime fusion assumptions.
|
|
- Active bytes/token estimate including scales.
|
|
- Checkpoint path and eval/log paths.
|
|
- Accuracy and verbosity metrics.
|
|
- Decision and next action.
|