### What does this PR do?
Type of change: documentation / repo housekeeping
Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.
```bash
repo-root/
├── .agents/ ← canonical source of truth
│ ├── README.md
│ ├── clusters.yaml.example
│ ├── scripts/
│ │ └── sync-upstream-skills.sh
│ └── skills/
│ ├── accessing-mlflow/ compare-results/ debug/
│ ├── deployment/ eagle3-new-model/ eagle3-review-logs/
│ ├── eagle3-triage/ eagle3-validate/ evaluation/
│ ├── launching-evals/ monitor/ ptq/
│ ├── quant-recipe-search/ release-cherry-pick/ common/
│
├── .claude/ ← back-compat (relative symlinks)
│ ├── clusters.yaml.example → ../.agents/clusters.yaml.example
│ ├── scripts → ../.agents/scripts
│ └── skills → ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
├── .codex/skills → ../.agents/skills
└── .cursor/skills → ../.agents/skills
```
### Why symlinks (and not "just point each agent's config at
`.agents/`")
Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.
### Changes
- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.
### Testing
- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.
### Additional Information
- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.
---------
Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
10 KiB
Recipe Iteration Reference
Problem
Quantization recipe search is a constrained optimization loop:
- Improve the user's chosen objective: compute/throughput, memory/latency, or a custom score.
- Keep every benchmark within the accepted accuracy-loss threshold. Default: less than 1 percentage point versus the matching BF16/FP16 baseline.
- Keep accuracy, verbosity/token usage, runtime behavior, and active cost separate until the final decision.
- Treat generated checkpoints as candidates. Evaluation and comparison decide whether a candidate is useful.
Ask for missing objective, primary quantization family, and benchmark set before planning candidates.
Search Space
Write the search space before launching PTQ. A recipe is a combination of these axes, not just a numeric format.
| Axis | Choices To Try | Notes |
|---|---|---|
| Numeric format | FP8/W8A8, NVFP4/W4A4, W4A16 NVFP4, INT4/AWQ, mixed formats | Keep FP8/W8A8 as a near-lossless baseline unless it is the user's target. |
| Calibration/search algorithm | Max, MSE, GPTQ, AWQ, AutoQuant scoring, calibration data variants | Algorithm choice is independent from numeric format. |
| Selection method | Manual/heuristic, sensitivity-guided manual, AutoQuant, hybrid | Record how each candidate was selected. |
| Module family | Attention, MLP, MoE experts, routers/gates, embeddings, lm_head, adapters, vision encoders |
Change one major family at a time for ablations. |
| Runtime constraints | Fused attention groups, fused MoE expert projections, backend-supported formats | Do not mix incompatible quantization inside a fused runtime group. |
| Calibration budget | Dataset mix, sample count, sequence length, batch size | Vary deliberately and record the budget. |
Objective Axis
Use the user's objective to decide which candidates are worth testing.
- Compute / throughput: typical data-center target. Favor activation quantization such as NVFP4 or FP8 when the runtime has fast kernels.
- Memory / latency: typical edge or memory-pressure target. Favor W4A16 or weight-only recipes when they preserve accuracy. Prefer active bytes per forward/decode path over total checkpoint size for routed or sparse models.
- Custom: use the user-provided score, for example checkpoint size, latency at a fixed concurrency, or product-specific memory budget.
If the user needs multiple objectives, maintain separate tables or define an explicit weighted score.
Numeric Format Axis
Common starting families:
- FP8/W8A8: near-lossless baseline or explicit primary target.
- NVFP4/W4A4: low-bit candidate family when activation quantization is part of the target.
- W4A16 NVFP4: weight-only NVFP4 family for accuracy-preserving memory/latency searches.
- INT4/AWQ: weight-only low-bit family for low-batch memory/latency targets.
- Mixed formats: examples include NVFP4+FP8, W4A16 NVFP4 with FP8 attention, or model-specific recipe fragments.
KV-cache dtype, parser settings, token caps, and backend flags are runtime/eval controls unless the user explicitly makes them part of the recipe objective.
Calibration Algorithm Axis
Try calibration/search algorithms as independent recipe variants:
- Max calibration: fast baseline for many FP8/NVFP4 formats.
- MSE calibration: try when max calibration loses accuracy for low-bit weights or sensitive layers.
- GPTQ: try for weight-quantized candidates where correction cost is acceptable.
- AWQ: try for INT4 or other weight-only candidates.
- AutoQuant scoring: use KL-divergence or gradient-based scoring when available to rank layers/modules and produce sensitivity reports.
Selection Method Axis
Choose which modules get which format by one of these methods:
- Manual/heuristic: use prior experience, module-family cost, and controlled ablations.
- Sensitivity-guided manual: generate or recover an AutoQuant sensitivity report, then protect sensitive families or quantize low-sensitivity high-cost families.
- AutoQuant: search per-layer/per-module selections under constraints such as target bits, active-cost objective, or allowed formats. When AutoQuant is available, include at least one AutoQuant-generated candidate in the portfolio so its trade-off can be compared against manual recipes.
- Hybrid: start from AutoQuant, then override known runtime constraints or known-sensitive fused groups manually.
Design Workflow
-
Recover existing evidence:
- Result tables, checkpoints, recipe logs, AutoQuant states, sensitivity reports, and active jobs.
- Use
monitor,launching-evals, andcompare-resultsfor execution state and metric provenance.
-
Define the target:
- Objective, primary quantization family, benchmark set, accuracy-loss threshold, cost metric, and calibration budget.
- Include scale storage and other quantization metadata in cost estimates.
-
Pick baselines:
- BF16/FP16 baseline.
- FP8/W8A8 near-lossless baseline, unless FP8 is the final target.
- Existing production recipe if one exists.
-
Pick first candidates:
- Start from
modelopt_recipeswhen ModelOpt is available. - Prefer model-specific recipes, then general PTQ presets, then recipe fragments.
- Add one AutoQuant candidate in the requested primary family when AutoQuant is available. Treat it as the expected best-search path, but validate it.
- Add at least one manual or sensitivity-guided candidate for comparison and as a fallback if AutoQuant misses the benchmark frontier or produces a runtime-incompatible recipe.
- Start from
-
Generate and validate:
- Delegate checkpoint generation and validation to
ptq. - Check checkpoint coverage, quantization metadata, and expected module coverage.
- Pipe-clean serving only after checkpoint validation passes.
- Delegate checkpoint generation and validation to
-
Scale evaluation:
- Run cheap screen benchmarks first.
- Expand only candidates that pass screen evals and runtime gates.
Runtime Fusion Rules
Search by module family, but respect modules fused by the target runtime.
- vLLM Qwen linear attention can fuse
linear_attn.in_proj_qkvandlinear_attn.in_proj_zintolinear_attn.in_proj_qkvz; do not mix formats or algorithms across those shards unless the runtime supports it. - Fused MoE kernels can couple expert projections such as gate/up (
w1/w3, or equivalent names); treat each fused expert group as one recipe unit unless deployment confirms mixed formats are supported. - If a checkpoint is valid but deployment fails due to missing support, classify
it as checkpoint-quality, recipe/runtime compatibility, or deployment
implementation. For deployment implementation, try small patches or flags via
deployment/debugbefore rejecting the recipe.
Iteration Loop
Use this loop after each candidate:
- Update the portfolio table with recipe axes, active cost, checkpoint path, eval logs, accuracy, verbosity, and decision.
- Compare against BF16/FP16 and FP8/W8A8 baselines.
- If accuracy drops:
- Protect sensitive module families.
- Try MSE, GPTQ, or AWQ variants.
- Use AutoQuant sensitivity to choose manual overrides.
- If performance or active cost is insufficient:
- Quantize the next high-cost active family.
- Try a more aggressive format.
- Revisit the active-cost objective or AutoQuant constraints.
- If verbosity changes:
- Inspect output samples and generation stats.
- Verify parser, token cap, sampling, backend, and KV-cache settings did not change.
- If results are close or noisy:
- Rerun before labeling a benchmark regression.
- If AutoQuant gives repeated recipes:
- Check achieved bits and recipe hashes.
- Adjust objective, allowed formats, or constraints before larger sweeps.
- If AutoQuant underperforms manual recipes:
- Compare the AutoQuant sensitivity report against manual ablation results.
- Check whether AutoQuant protected high-active-cost modules, excluded the wrong families, optimized checkpoint size instead of active cost, or hit runtime-fusion constraints.
- Keep the manual recipe in the table and use AutoQuant sensitivity to design the next hybrid/manual candidate.
Promote a recipe only when validated comparison shows it satisfies the user's objective and benchmark threshold.
Delegating To Existing Skills
Do not reimplement workflows that existing skills own:
| Need | Use |
|---|---|
| Generate/check a quantized checkpoint | ptq |
| Serve a checkpoint or test backend flags | deployment |
| Create or submit NEL configs | evaluation |
| Resume/debug/analyze live eval runs | launching-evals |
| Track active Slurm/NEL jobs | monitor |
| Fetch MLflow artifacts | accessing-mlflow |
| Compute baseline-vs-candidate deltas | compare-results |
Before launching PTQ in a ModelOpt repo, read the current PTQ skill from
.agents/skills/ptq/SKILL.md; recipe paths and validation gates can change.
ModelOpt Starting Points
When ModelOpt is available, start from modelopt_recipes:
- Check model-specific recipes first, for example
modelopt_recipes/huggingface/<model_family>/ptq/. - Check general PTQ recipes and presets.
- Use recipe fragments to build controlled manual variants.
- Summarize include/exclude coverage before calibration. If a pattern misses the intended layer family, fix the recipe before launching.
Useful starting candidates:
- Compute/throughput: FP8/W8A8, NVFP4/W4A4, mixed NVFP4+FP8 with activation quantization.
- Memory/latency: W4A16 NVFP4, weight-only NVFP4, or W4A16 mixed with FP8 for sensitive modules.
- MoE: experts-only or MLP-only recipes, then expand based on sensitivity and active-routing cost.
Candidate Record
For every candidate, record:
- Objective and acceptance threshold.
- Numeric formats and module-family coverage.
- Calibration/search algorithm and calibration data budget.
- Selection method: manual, sensitivity-guided, AutoQuant, or hybrid.
- Whether the candidate came from AutoQuant, manual ablation, or a hybrid override, so AutoQuant and manual trade-offs can be compared directly.
- Runtime fusion assumptions.
- Active bytes/token estimate including scales.
- Checkpoint path and eval/log paths.
- Accuracy and verbosity metrics.
- Decision and next action.