Files
Model-Optimizer/plugins/modelopt
yeyu-nvidiaandClaude Opus 5 216f28a6e0 Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)
### What does this PR do?

Type of change: documentation

Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.

**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:

```
skills/eagle3-new-model/    skills/eagle3-review-logs/
skills/eagle3-triage/       skills/eagle3-validate/
```

Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.

**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:

```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md                      # router: stage table x algorithm table
└── references/
    ├── stages/                   # the procedure — algorithm-independent
    │   ├── configure.md          # <- eagle3-new-model
    │   ├── review-logs.md        # <- eagle3-review-logs
    │   ├── triage.md             # <- eagle3-triage
    │   └── validate.md           # <- eagle3-validate
    └── algorithms/               # the data sheet — per-algorithm
        ├── README.md             # contract: 6 required sections
        ├── eagle3.md
        ├── dflash.md
        ├── dspark.md             # DFlash variant — delta only
        └── domino.md             # DFlash variant — delta only
```

Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.

DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.

Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:

- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.

**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.

### Usage

```
/speculative-decoding
```

Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:

```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun   # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes      # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```

### Testing

- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`

### Additional Information

Follows #2025, which moved the skill tree into the installable plugin.

Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):

1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.

Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.

* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-17 11:09:14 -07:00
..
2026-09-10 17:30:37 +00:00