### What does this PR do?
Type of change: documentation / repo housekeeping
Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.
```bash
repo-root/
├── .agents/ ← canonical source of truth
│ ├── README.md
│ ├── clusters.yaml.example
│ ├── scripts/
│ │ └── sync-upstream-skills.sh
│ └── skills/
│ ├── accessing-mlflow/ compare-results/ debug/
│ ├── deployment/ eagle3-new-model/ eagle3-review-logs/
│ ├── eagle3-triage/ eagle3-validate/ evaluation/
│ ├── launching-evals/ monitor/ ptq/
│ ├── quant-recipe-search/ release-cherry-pick/ common/
│
├── .claude/ ← back-compat (relative symlinks)
│ ├── clusters.yaml.example → ../.agents/clusters.yaml.example
│ ├── scripts → ../.agents/scripts
│ └── skills → ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
├── .codex/skills → ../.agents/skills
└── .cursor/skills → ../.agents/skills
```
### Why symlinks (and not "just point each agent's config at
`.agents/`")
Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.
### Changes
- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.
### Testing
- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.
### Additional Information
- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.
---------
Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2.4 KiB
name, description, user_invocable
| name | description | user_invocable |
|---|---|---|
| eagle3-new-model | Add a new model to the EAGLE3 offline pipeline. Generates an hf_offline_eagle3.yaml launcher config for a new model checkpoint, choosing the right hidden state dump backend (TRT-LLM / HF / vLLM) and GPU configuration. Use when user wants to run EAGLE3 on a model that does not yet have a YAML in tools/launcher/examples/ or asks how to configure the pipeline for a new checkpoint. | true |
EAGLE3 New Model Configuration
Create tools/launcher/examples/<Org>/<Model>/hf_offline_eagle3.yaml by copying the
closest existing example and adapting it. Pick a reference with the same shape as the
target (dense vs MoE, similar size) from tools/launcher/examples/ — e.g. the Qwen3-8B
config for a dense model.
The pipeline is a 4-task config (task_0 data synthesis → task_1 hidden-state dump →
task_2 train → task_3 benchmark). The task structure, args, containers, and GPU/node
sizing are all visible in the existing examples — infer them from a reference rather than
hand-rolling. This file documents only the two things that are not obvious from the
examples: which dump backend to pick, and the model-specific gotchas.
Choosing the task_1 hidden-state dump backend
| Backend | Script | When to use |
|---|---|---|
| vLLM | common/eagle3/dump_offline_data_vllm.sh |
Default. Broad coverage via vLLM's native hidden-state extractor. |
| HF | common/eagle3/dump_offline_data_hf.sh |
VLMs / multimodal, custom-code models, sliding-window attention (TRT-LLM can't serve these). |
| TRT-LLM | common/eagle3/dump_offline_data.sh |
Pure-text models with TRT-LLM support; pass --tp <TP> and --moe-ep <EP>. |
Rule of thumb: HF if the model is a VLM or uses sliding-window attention; vLLM otherwise. TRT-LLM only when you specifically want its kernels for a supported plain-text model.
Model-specific adjustments
These are the non-obvious knobs that vary per model:
| Situation | What to change |
|---|---|
Requires --trust-remote-code |
Add to task_0 vLLM args (before the -- separator) and to task_3 benchmark args |
| MoE with large expert hidden dim | Increase intermediate_size in eagle_config.json to match moe_intermediate_size |
| Custom tokenizer (e.g. tiktoken) | Set TIKTOKEN_RS_CACHE_DIR env var in task_0 and task_1 |
After adapting the config, preview it with --dryrun before submitting.