Files
Model-Optimizer/.agents/skills/eagle3-new-model/SKILL.md
T
bdb793a42d [SKILL.md Chore] Make .agents/ the canonical agent-skills location (#1362)
### What does this PR do?

Type of change: documentation / repo housekeeping

Centralize agent-shared assets under **`.agents/`** as the single,
agent-agnostic source of truth, so the same `SKILL.md` (plus shared
scripts and cluster config) works across every coding agent without
maintaining N copies that drift out of sync. Claude Code discovers
skills only under `.claude/skills/`, so `.claude/` holds **relative
in-repo symlinks** into `.agents/` for back-compat.

```bash
repo-root/
├── .agents/                      ← canonical source of truth
│   ├── README.md
│   ├── clusters.yaml.example
│   ├── scripts/
│   │   └── sync-upstream-skills.sh
│   └── skills/
│       ├── accessing-mlflow/      compare-results/     debug/
│       ├── deployment/            eagle3-new-model/    eagle3-review-logs/
│       ├── eagle3-triage/         eagle3-validate/     evaluation/
│       ├── launching-evals/       monitor/             ptq/
│       ├── quant-recipe-search/   release-cherry-pick/ common/
│
├── .claude/                      ← back-compat (relative symlinks)
│   ├── clusters.yaml.example  →  ../.agents/clusters.yaml.example
│   ├── scripts                →  ../.agents/scripts
│   └── skills                 →  ../.agents/skills
│
└── (future agents — add a symlink/config, no copies)
    ├── .codex/skills           →  ../.agents/skills
    └── .cursor/skills          →  ../.agents/skills
```

### Why symlinks (and not "just point each agent's config at
`.agents/`")

Claude Code **only** auto-discovers project skills under
`.claude/skills/` — there is no setting/env var to redirect discovery to
an arbitrary path, and the plugin route would require committing a
`.claude/settings.json` + marketplace manifest, add a first-open
workspace-trust gate (breaks headless/CI runs), and namespace every
skill (`/ptq` → `/<plugin>:ptq`). A single relative in-repo symlink is
the smallest change that keeps `.agents/` canonical while satisfying
Claude Code's discovery requirement. This repo already commits relative
symlinks (`CLAUDE.md`, `tools/launcher/modules/Model-Optimizer`). See
the discussion thread for the full comparison.

### Changes

- Move `.claude/{skills,scripts,clusters.yaml.example}` → `.agents/`
(git renames preserve history).
- Add `.agents/README.md` documenting the convention and per-agent
wiring.
- Re-add `.claude/skills`, `.claude/scripts`,
`.claude/clusters.yaml.example` as relative symlinks into `.agents/`.
- Update internal path references and lint/sync config from `.claude/`
to `.agents/` (upstream provenance paths and `.claude/clusters.yaml`
back-compat lookups left intact).
- **Merged latest `main`** and folded in skills added there since branch
time — `compare-results`, `eagle3-new-model`, `eagle3-triage`,
`eagle3-review-logs`, `eagle3-validate`, `quant-recipe-search`, and new
`evaluation` recipes/tasks/references — into `.agents/skills/`.
- `main` converted `CLAUDE.md` into a symlink to a new agent-agnostic
`AGENTS.md`; this PR keeps that and moves the "skills live in
`.agents/`" guidance into `AGENTS.md`.

### Testing

- `.claude/skills` symlink resolves to all 15 skills; `ls
.claude/skills` and `ls .agents/skills` match.
- `pre-commit run check-symlinks --all-files` and `markdownlint-cli2
--all-files` pass.
- `bash -n` clean on `sync-upstream-skills.sh` and `remote_exec.sh`.
- `.claude/skills`, `.claude/scripts`, `.claude/clusters.yaml.example`,
and `CLAUDE.md` are all recorded as git symlinks (mode `120000`).

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅ `.claude/skills/`,
`.claude/scripts/`, and `.claude/clusters.yaml.example` continue to
resolve to the same content via symlinks; Claude Code auto-discovery is
unaffected; `remote_exec.sh` still accepts `.claude/clusters.yaml`.
- Did you write any new necessary tests?: N/A — directory move with
symlinks; verified as listed above.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — repo housekeeping only, no API/feature/bugfix change.

### Additional Information

- A few vendored skill files still carry internal details (Slurm account
names, lustre paths, internal `:5005` GitLab registry advice in
`launching-evals/`) worth scrubbing in a follow-up.

---------

Signed-off-by: Seonghee Lee <seongheel@nvidia.com>
Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Shengliang Xu <106840466+shengliangxu@users.noreply.github.com>
Co-authored-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-05 16:12:19 -07:00

2.4 KiB

name, description, user_invocable
name description user_invocable
eagle3-new-model Add a new model to the EAGLE3 offline pipeline. Generates an hf_offline_eagle3.yaml launcher config for a new model checkpoint, choosing the right hidden state dump backend (TRT-LLM / HF / vLLM) and GPU configuration. Use when user wants to run EAGLE3 on a model that does not yet have a YAML in tools/launcher/examples/ or asks how to configure the pipeline for a new checkpoint. true

EAGLE3 New Model Configuration

Create tools/launcher/examples/<Org>/<Model>/hf_offline_eagle3.yaml by copying the closest existing example and adapting it. Pick a reference with the same shape as the target (dense vs MoE, similar size) from tools/launcher/examples/ — e.g. the Qwen3-8B config for a dense model.

The pipeline is a 4-task config (task_0 data synthesis → task_1 hidden-state dump → task_2 train → task_3 benchmark). The task structure, args, containers, and GPU/node sizing are all visible in the existing examples — infer them from a reference rather than hand-rolling. This file documents only the two things that are not obvious from the examples: which dump backend to pick, and the model-specific gotchas.

Choosing the task_1 hidden-state dump backend

Backend Script When to use
vLLM common/eagle3/dump_offline_data_vllm.sh Default. Broad coverage via vLLM's native hidden-state extractor.
HF common/eagle3/dump_offline_data_hf.sh VLMs / multimodal, custom-code models, sliding-window attention (TRT-LLM can't serve these).
TRT-LLM common/eagle3/dump_offline_data.sh Pure-text models with TRT-LLM support; pass --tp <TP> and --moe-ep <EP>.

Rule of thumb: HF if the model is a VLM or uses sliding-window attention; vLLM otherwise. TRT-LLM only when you specifically want its kernels for a supported plain-text model.

Model-specific adjustments

These are the non-obvious knobs that vary per model:

Situation What to change
Requires --trust-remote-code Add to task_0 vLLM args (before the -- separator) and to task_3 benchmark args
MoE with large expert hidden dim Increase intermediate_size in eagle_config.json to match moe_intermediate_size
Custom tokenizer (e.g. tiktoken) Set TIKTOKEN_RS_CACHE_DIR env var in task_0 and task_1

After adapting the config, preview it with --dryrun before submitting.