Consolidate speculative-decoding agent skills into one stage/algorithm tree (#2201)

### What does this PR do?

Type of change: documentation

Reorganizes the EAGLE3 agent skills into a single speculative-decoding
skill, then adds algorithm sheets for DFlash, DSpark, and Domino.

**The problem.** The four `eagle3-*` skills each baked the algorithm
into a *stage* of the same draft-model pipeline:

```
skills/eagle3-new-model/    skills/eagle3-review-logs/
skills/eagle3-triage/       skills/eagle3-validate/
```

Adding DFlash would have meant four more near-duplicate skills, since
the stages are shared and only the algorithm differs.

**The change.** One skill dir shaped like `ptq/` (SKILL.md +
references/), split along the two real axes:

```
plugins/modelopt/skills/speculative-decoding/
├── SKILL.md                      # router: stage table x algorithm table
└── references/
    ├── stages/                   # the procedure — algorithm-independent
    │   ├── configure.md          # <- eagle3-new-model
    │   ├── review-logs.md        # <- eagle3-review-logs
    │   ├── triage.md             # <- eagle3-triage
    │   └── validate.md           # <- eagle3-validate
    └── algorithms/               # the data sheet — per-algorithm
        ├── README.md             # contract: 6 required sections
        ├── eagle3.md
        ├── dflash.md
        ├── dspark.md             # DFlash variant — delta only
        └── domino.md             # DFlash variant — delta only
```

Stage docs cite algorithm-sheet sections by heading (*Pipeline tasks*,
*Success markers*, *Quality gate*, *Known failures*, ...), so a new
algorithm means one new file plus a table row — no stage edits. Every
recipe in `modelopt_recipes/general/speculative_decoding/` now has a
sheet.

DSpark and Domino are documented as **DFlash variants**, not separate
pipelines: same `recipe_type: speculative_dflash`, same training script,
same `dflash.*` config namespace, selected by
`dflash_architecture_config.projector_type`. Their sheets carry only the
delta.

Writing the sheets surfaced three things the old EAGLE3-only skills got
wrong or missed:

- **Task counts are not fixed.** The old skills hardcoded "4-step
pipeline, task_0 through task_3". DFlash offline is 2 tasks, DFlash
online is 3, Domino is 2. The stage docs no longer assume a count.
- **`--aux-layers` couples the dump to the draft.** For DFlash the
dump's layer count must equal the draft's `num_hidden_layers`; a
mismatch doesn't error, it silently captures the wrong layers. Recorded
under *Known failures*.
- **In-training AR is meaningless for DSpark and Domino.** Both recipes
pin `estimate_ar: false` / `ar_validate_steps: 0` because eval runs the
DFlash backbone with the new head bypassed. Each sheet says so under
*Quality gate* so nobody reads a backbone-only number as a result.

**Behavior change:** the four `/eagle3-*` slash commands are replaced by
one `/speculative-decoding`. This isn't optional —
`tools/precommit/sync_claude_skills.sh` iterates `.agents/skills/*/` one
level deep and plugin discovery is `skills/<name>/SKILL.md`, so a
directory is either one skill or a container of skills, not both.
`tools/launcher/docs/claude_code.md` is updated accordingly.

### Usage

```
/speculative-decoding
```

Or by description — the skill triggers on EAGLE3 / DFlash / DSpark /
draft model / acceptance rate. For a new model, follow the stages in
order:

```bash
# 1. Configure: copy the closest examples/<Org>/<Model>/hf_<mode>_<algo>.yaml and adapt
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun   # preview
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes      # submit
# 2. review-logs -> 3. triage (if anything failed) -> 4. validate
```

### Testing

- `claude plugin validate . --strict` and `claude plugin validate
plugins/modelopt --strict` — both pass
- `pre-commit run --files ...` over all changed files — passes,
including `markdownlint-cli2` and the `sync-claude-skills` symlink hook
(it agrees with the new `.claude/skills/speculative-decoding` symlink)
- Verified the new skill is discovered and its description loads
- Every relative link across the skill tree resolves; every repo path
cited in the sheets exists; no dangling `eagle3-*` reference remains
anywhere in the repo
- Each factual claim in the sheets was checked against its source — the
launcher example YAMLs, the four recipes, `dflash_online_training.sh`,
`vllm_smoke_test.sh`, `check_regression.py`, and
`plugins/hf_{dflash,dspark,domino}.py` — rather than written from memory

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — the four `/eagle3-*` slash
commands become `/speculative-decoding`. Agent tooling only; no library
or API surface is touched. The three YAML comment fixes are
comment-only, no behavior change.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation; covered
by plugin validation and pre-commit
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — agent tooling only, matching how #2025 handled it
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; 2
findings, both fixed in `9b3c568`

### Additional Information

Follows #2025, which moved the skill tree into the installable plugin.

Two stale in-repo comments were found while sourcing the sheets, and are
**fixed in this PR** (`c2696b1`, comment-only):

1. `modelopt_recipes/general/speculative_decoding/dflash.yaml` pointed
`chat_template` at a `chat_templates/` directory under
`modelopt_recipes` that does not exist — templates live per-model beside
each launcher example.
2. Both offline DFlash example YAMLs annotated `--aux-layers dflash`
with "Must match the draft model's num_hidden_layers". `--aux-layers` is
a preset keyword accepting only `eagle`, `dflash`, or an explicit id
list, so it carries no count. The constraint is real but belongs to the
draft depth the preset resolves to: `--num-draft-layers` on the vLLM
dump, and no override at all on the HF/TRT-LLM dumps, which hardcode 5
via `resolve_aux_layers`. This comment had already misled this PR's own
first draft, which is why it's fixed rather than just documented.

Because of (1), this PR now touches `modelopt_recipes/`, which adds
**@NVIDIA/modelopt-recipes-codeowners** to the required reviewers.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added comprehensive speculative-decoding guidance for configuration,
training, validation, troubleshooting, and supported algorithms.
* Added workflow references for DFlash, Domino, DSpark, and EAGLE3,
including quality checks and failure diagnosis.

* **Documentation**
* Generalized experiment-log review and pipeline triage across
algorithms.
* Clarified DFlash draft-depth configuration, resource sizing, task
recovery, and validation.
* Replaced the EAGLE3-specific workflow entry with the broader
speculative-decoding workflow.
* Removed standalone EAGLE3 skill documentation as guidance is now
consolidated under speculative decoding.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
yeyu-nvidia
2026-09-17 11:09:14 -07:00
committed by GitHub
co-authored by Claude Opus 5
parent 2ff2e1bc80
commit 216f28a6e0
21 changed files with 1043 additions and 374 deletions
-1
View File
@@ -1 +0,0 @@
../../.agents/skills/eagle3-new-model
-1
View File
@@ -1 +0,0 @@
../../.agents/skills/eagle3-review-logs
-1
View File
@@ -1 +0,0 @@
../../.agents/skills/eagle3-triage
-1
View File
@@ -1 +0,0 @@
../../.agents/skills/eagle3-validate
+1
View File
@@ -0,0 +1 @@
../../.agents/skills/speculative-decoding
@@ -1,46 +0,0 @@
---
name: eagle3-new-model
description: >
Add a new model to the EAGLE3 offline pipeline. Generates an hf_offline_eagle3.yaml
launcher config for a new model checkpoint, choosing the right hidden state dump
backend (TRT-LLM / HF / vLLM) and GPU configuration.
Use when user wants to run EAGLE3 on a model that does not yet have a YAML in
tools/launcher/examples/ or asks how to configure the pipeline for a new checkpoint.
user_invocable: true
---
# EAGLE3 New Model Configuration
Create `tools/launcher/examples/<Org>/<Model>/hf_offline_eagle3.yaml` by **copying the
closest existing example and adapting it**. Pick a reference with the same shape as the
target (dense vs MoE, similar size) from `tools/launcher/examples/` — e.g. the Qwen3-8B
config for a dense model.
The pipeline is a 4-task config (`task_0` data synthesis → `task_1` hidden-state dump →
`task_2` train → `task_3` benchmark). The task structure, args, containers, and GPU/node
sizing are all visible in the existing examples — infer them from a reference rather than
hand-rolling. This file documents only the two things that are **not** obvious from the
examples: which dump backend to pick, and the model-specific gotchas.
## Choosing the `task_1` hidden-state dump backend
| Backend | Script | When to use |
|---------|--------|-------------|
| vLLM | `common/eagle3/dump_offline_data_vllm.sh` | **Default.** Broad coverage via vLLM's native hidden-state extractor. |
| HF | `common/eagle3/dump_offline_data_hf.sh` | VLMs / multimodal, custom-code models, sliding-window attention (TRT-LLM can't serve these). |
| TRT-LLM | `common/eagle3/dump_offline_data.sh` | Pure-text models with TRT-LLM support; pass `--tp <TP>` and `--moe-ep <EP>`. |
Rule of thumb: **HF** if the model is a VLM or uses sliding-window attention; **vLLM**
otherwise. TRT-LLM only when you specifically want its kernels for a supported plain-text model.
## Model-specific adjustments
These are the non-obvious knobs that vary per model:
| Situation | What to change |
|---|---|
| Requires `--trust-remote-code` | Add to `task_0` vLLM args (before the `--` separator) and to `task_3` benchmark args |
| MoE with large expert hidden dim | Increase `intermediate_size` in `eagle_config.json` to match `moe_intermediate_size` |
| Custom tokenizer (e.g. tiktoken) | Set `TIKTOKEN_RS_CACHE_DIR` env var in `task_0` and `task_1` |
After adapting the config, preview it with `--dryrun` before submitting.
@@ -1,168 +0,0 @@
---
name: eagle3-triage
description: >
Triage a failed EAGLE3 pipeline run. Identifies which step failed (data synthesis,
hidden state dump, training, or benchmark), diagnoses root cause from logs, and
suggests fixes. Use when user reports an EAGLE3 pipeline failure or asks why a
specific step failed. Also helps debug new model support issues.
user_invocable: true
---
# EAGLE3 Pipeline Triage
Diagnose failures in the 4-step EAGLE3 offline pipeline. This skill walks through
each step, identifies the failure point, and provides actionable fixes.
## Pipeline Overview
| Step | Script | Purpose | Common failure area |
|------|--------|---------|---------------------|
| task_0 | `common/vllm/query.sh` | Data synthesis via vLLM server | Server startup, model loading, OOM |
| task_1 | `common/eagle3/dump_offline_data_vllm.sh` (or `_hf.sh` / `.sh`) | Dump hidden states | Backend selection, OOM, unsupported arch |
| task_2 | `common/eagle3/train_eagle.sh` | Train EAGLE3 draft head | Dependencies, training crash, export |
| task_3 | `common/specdec_bench/quick_check.sh` | Benchmark acceptance rate | Engine startup, draft model loading |
## Step 0 — Locate the experiment
Ask the user for one of:
- Experiment directory (e.g., the `--job-dir` passed to `launch.py` or `slurm.py`)
- The model name / YAML they ran
Find recent experiments under the job directory:
```bash
ls -td experiments/cicd/cicd_* | head -10
# or wherever --job-dir was pointed
```
Each experiment directory contains one subdirectory per task (task_0 through task_3),
each with a log file whose name varies by launch mode (Slurm: `sbatch_*.out`, local
Docker: `*.log`).
## Step 1 — Fetch logs for the failed task
Match the log files generally and read the tail of each — errors appear at the end:
```bash
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -200 "$f"; echo
done
```
Look for the first task with a non-zero exit code or error message.
## Step 2 — Diagnose by step
### task_0 failures (Data Synthesis)
**How it works:** Launches a vLLM OpenAI-compatible server, polls `/health` until ready,
then runs `query.py` to generate synthetic prompt/response pairs.
Output goes to `/scratchspace/data/`.
| Error pattern | Root cause | Fix |
|---|---|---|
| Server never becomes healthy (hangs at health check) | Model too large for allocated GPUs, or vLLM startup crash | Check BF16 weight size vs total allocated GPU memory; increase TP and/or nodes. |
| `CUDA out of memory` during model load | Insufficient GPU memory | Reduce `--max-model-len` or increase `--tensor-parallel-size` |
| `trust_remote_code` error | Model requires custom code but flag not set | Add `--trust-remote-code` before the `--` separator in task_0 args |
| Vocab / tokenizer error | Missing tokenizer cache (e.g., GPT-OSS-20B needs `TIKTOKEN_RS_CACHE_DIR`) | Set `TIKTOKEN_RS_CACHE_DIR` to a pre-populated cache path in the environment |
| Architecture not supported | vLLM version doesn't support this model | Try a newer vLLM container (`vllm/vllm-openai:latest`) |
| `CANCELLED ... DUE TO TIME LIMIT` | Job wall-clock limit too short | Increase Slurm `--time`. Note: `afterany` deps let task_1 still start. |
| Empty `/scratchspace/data/` | query.py ran but produced no output | Check `--data` path exists and contains prompts. Check query.py logs. |
### task_1 failures (Hidden State Dump)
**How it works:** Loads the target model and runs a forward pass on each conversation,
saving hidden states as `.pt` files in `/scratchspace/offline_hidden_states/`.
Three backends are available:
| Backend | Script | When to use |
|---------|--------|-------------|
| vLLM | `dump_offline_data_vllm.sh` | Broad model coverage; uses vLLM's native hidden-state extractor |
| HF | `dump_offline_data_hf.sh` | VLMs, custom-code models, SWA attention; uses `device_map="auto"` |
| TRT-LLM | `dump_offline_data.sh` | Pure-text models with TRT-LLM support; needs `--tp`/`--moe-ep` args |
| Error pattern | Root cause | Fix |
|---|---|---|
| `No such file or directory: dump_offline_data_vllm.sh` | Wrong script path in YAML | Use the correct path under `common/eagle3/` |
| `FileNotFoundError: /scratchspace/data` | task_0 failed or produced no output | Re-run task_0 first, or point `--input-data` to existing data |
| `CUDA out of memory` | Model too large | Switch to `_hf.sh` (device_map="auto") or increase TP |
| `RuntimeError` / unsupported arch | Model not supported by TRT-LLM backend | Switch to `dump_offline_data_hf.sh` or `dump_offline_data_vllm.sh` |
| `NCCL timeout` / `NCCL error` | Multi-node communication failure | Retry. Reduce EP. |
| No `.pt` files in output dir | Script ran but extraction produced nothing | Check `--max-seq-len` and input data format |
| `pyxis: child terminated with signal 15` | SIGTERM — likely OOM | Increase TP or switch backends |
### task_2 failures (Training)
**How it works:** Installs requirements, runs `launch_train.sh` (Accelerate + FSDP) with the
config from `modelopt_recipes/general/speculative_decoding/eagle3.yaml`, then exports via
`export_hf_checkpoint.py`. Output: `/scratchspace/eagle3/` and `/scratchspace/export/`.
| Error pattern | Root cause | Fix |
|---|---|---|
| `FileNotFoundError: /scratchspace/offline_hidden_states` | task_1 failed or produced no output | Re-run task_1 first |
| `CUDA out of memory` during training | Batch size too large | Reduce `training.train_bs` or `training.training_seq_len` |
| `KeyError` / `AttributeError` in model loading | Model architecture not recognized by EAGLE3 | Model may need code changes in modelopt for this architecture |
| Loss is NaN or diverges | LR too high or data quality issue | Reduce `training.lr`. Check hidden state data. |
| `export_hf_checkpoint.py` fails | Training produced incomplete checkpoint | Check `/scratchspace/eagle3/` for `model.safetensors` |
### task_3 failures (Benchmark)
**How it works:** Launches vLLM with the target + draft model, runs acceptance rate and
throughput benchmarks. Output: JSON files.
| Error pattern | Root cause | Fix |
|---|---|---|
| `FileNotFoundError: /scratchspace/export` | task_2 failed or export step failed | Re-run task_2. Check export output. |
| `trust_remote_code` error at benchmark | Model requires it but `quick_check.sh` doesn't forward the flag | Pass `--trust-remote-code` in task_3 args |
| Server fails with draft model | Draft model config incompatible with engine | Check `eagle_config.json` and engine version |
| AR below threshold / exit code 1 | Draft model quality too low | More epochs, data, or hyperparameter tuning |
| `CUDA out of memory` | Target + draft exceeds GPU memory | Increase TP |
| vLLM EAGLE3 not supported | vLLM version too old | Use a newer vLLM container |
## Step 3 — Check for new-model-specific issues
If the user is adding support for a new model, also check:
1. **Is the model a VLM?** → Use `dump_offline_data_hf.sh` (text-only path, no vision encoder invoked)
2. **Does the model use sliding window attention (SWA)?** → TRT-LLM backend won't work; use HF or vLLM
3. **Does the model need `trust_remote_code`?** → Add to task_0 args AND task_3 args
4. **Is the model MoE?** → Check `eagle_config.json` `intermediate_size` matches model's `moe_intermediate_size`
5. **Is the model architecture recognized by EAGLE3 training?** → may need code changes in `modelopt/torch/speculative/`
6. **Custom tokenizer?** → May need additional environment vars (e.g., `TIKTOKEN_RS_CACHE_DIR`)
## Step 4 — Suggest fix and next steps
After diagnosis, provide:
1. **Root cause** — one-line summary
2. **Fix** — specific config change, code edit, or command to run
3. **How to re-run** — skip earlier successful steps by pointing to existing scratchspace artifacts
To skip task_0 and task_1 and re-run from task_2:
```bash
uv run launch.py --yaml examples/<Org>/<Model>/hf_offline_eagle3.yaml \
pipeline.task_0.skip=true \
pipeline.task_1.skip=true \
--yes
```
To run only task_1 standalone (using existing task_0 data):
```bash
uv run launch.py --yaml examples/<Org>/<Model>/hf_offline_eagle3.yaml \
pipeline.task_0.skip=true \
pipeline.task_2.skip=true \
pipeline.task_3.skip=true \
--yes
```
If the fix requires code changes in ModelOpt (e.g., supporting a new model architecture),
note that a separate PR in the modelopt repo is needed.
## Step 5 — Record the failure pattern
If you encounter a failure pattern not seen before, capture it in the team's internal
triage tracker — the symptom, root cause, and fix — so the next engineer debugging the
same issue benefits.
@@ -1,119 +0,0 @@
---
name: eagle3-validate
description: >
Validate that an EAGLE3 pipeline run completed successfully end-to-end.
Checks all 4 steps produced expected artifacts, verifies acceptance rate
meets threshold (>= 2.1), and produces a summary report.
Use when user wants to verify a pipeline run or check benchmark results.
user_invocable: true
---
# EAGLE3 Pipeline Validation
Verify that an EAGLE3 pipeline run completed successfully and meets quality criteria.
## Step 0 — Identify the experiment
Find the most recent experiment directory (or ask the user for the path):
```bash
ls -td experiments/cicd/cicd_* | head -5
```
Each experiment directory has one subdirectory per task (numbered 0–3), each containing a
log file whose name varies by launch mode (Slurm: `sbatch_*.out`, local Docker: `*.log`).
## Step 1 — Check task outcomes
Match the log files generally and read the tail of each:
```bash
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -50 "$f"; echo
done
```
All 4 tasks must complete without error. Look for:
- `exit code: 0` or no error — success
- `DUE TO TIME LIMIT` — timeout
- `FAILED` / `signal` / exception traceback — failure
If any task failed, suggest running `/eagle3-triage` instead.
## Step 2 — Verify artifacts exist
Check each step produced the expected output (artifacts live on the cluster at `/scratchspace/`).
Confirm via log messages:
| Step | Expected log evidence | Artifact |
|------|-----------------------|----------|
| task_0 | "Saved N samples" or progress bar completing | `/scratchspace/data/*.jsonl` |
| task_1 | "Successfully processed N conversations" | `/scratchspace/offline_hidden_states/*.pt` |
| task_2 | Training loss decreasing, "export complete" | `/scratchspace/eagle3/model.safetensors`, `/scratchspace/export/` |
| task_3 | `Average Acceptance Length ... ratio: X.XX` | JSON result files |
## Step 3 — Check acceptance rate
In the task_3 log, find:
```text
Average Acceptance Length {'accept': X, 'count': Y, 'ratio': Z.ZZ}
```
The `ratio` field is the acceptance rate (AR).
| Criterion | Threshold | Status |
|-----------|-----------|--------|
| AR (MT-Bench) | >= 2.1 | PASS / FAIL |
If the log shows `AR ... < lower bound`, the run already triggered a threshold failure (exit code 1).
## Step 4 — Check training quality
In the task_2 log look for:
- **Final training loss** — should be decreasing, not NaN
- **AR validation during training** (if `training.ar_validate_steps` was set)
- **Number of training steps** — confirms full training duration
## Step 5 — Produce validation report
```markdown
## EAGLE3 Pipeline Validation Report
**Experiment:** <exp_dir>
**Model:** <model_name>
**Date:** <date>
**Pipeline config:** <yaml_path>
### Step Status
| Step | Task | Status | Notes |
|------|------|--------|-------|
| 0 | Data synthesis | PASS/FAIL/TIMEOUT | N samples generated |
| 1 | Hidden state dump | PASS/FAIL | N .pt files |
| 2 | Training + export | PASS/FAIL | Final loss: X.XX |
| 3 | Benchmark | PASS/FAIL | AR: X.XX |
### Acceptance Rate
- MT-Bench AR: X.XX (threshold: >= 2.1) — PASS/FAIL
### Training Summary
- Final loss: X.XX
- Training steps: N
- AR during training: X.XX (if validated)
### Overall: PASS / FAIL
<one-line summary>
```
## Step 6 — Suggest next steps
**If PASS:**
- Record the verified result (and checkpoint path) in the team's internal triage tracker
- This model is now a candidate to add as a launcher example in a dedicated PR
**If FAIL:**
- Identify which step or metric failed
- Suggest running `/eagle3-triage` for diagnosis
- For a low AR, diagnose the specific cause from the run (training loss curve, data
volume/quality, draft-head capacity, hyperparameters) and suggest fixes targeted to that
scenario — low AR can have many causes, so avoid a generic checklist.
@@ -0,0 +1,92 @@
---
name: speculative-decoding
description: >-
Train, debug, and validate a speculative decoding draft model (EAGLE3, DFlash,
DSpark, Domino) through the ModelOpt launcher pipeline. Use when the user wants
to add a new model to a draft-training pipeline, asks why a pipeline run failed,
wants experiment logs reviewed, or wants to check whether a run's acceptance rate
meets threshold. Triggers on "EAGLE3", "DFlash", "DSpark", "Domino", "draft model",
"acceptance rate", "speculative decoding pipeline". Do NOT use for quantizing a
model (use ptq) or serving a checkpoint (use deployment).
user_invocable: true
---
# Speculative Decoding Draft-Model Training
Everything needed to take a target model from "no draft head" to "validated
acceptance rate" lives in this directory. Work through the stages below in order
for a new model; jump straight to a stage when you already know which one you need.
## Two axes: stage and algorithm
The pipeline is the same shape for every draft-model algorithm — synthesize data,
dump base-model hidden states, train the draft, benchmark acceptance rate. What
changes between algorithms is which training script and recipe run, which knobs
matter, and which failures are typical.
So this skill is split along those two axes, and **you almost always read one file
from each**:
| Axis | Directory | What it holds |
| --- | --- | --- |
| Stage | `references/stages/` | The procedure — algorithm-independent |
| Algorithm | `references/algorithms/` | The data sheet — scripts, recipe, knobs, thresholds, known failures |
Read the stage file for *what to do*, and the algorithm sheet for *the values to
plug in*. When a stage file says "see the algorithm sheet", it means the section of
`references/algorithms/<algorithm>.md` with the matching heading.
## Stages
| Stage | Reference | Use when |
| --- | --- | --- |
| 1. Configure | `references/stages/configure.md` | Adding a model that has no pipeline YAML yet |
| 2. Review logs | `references/stages/review-logs.md` | A run finished (or died) and you want a pass/fail summary |
| 3. Triage | `references/stages/triage.md` | A task failed and you need root cause plus a fix |
| 4. Validate | `references/stages/validate.md` | All tasks passed and you need to confirm the acceptance rate gate |
Review-logs and triage overlap by design: review-logs is the fast sweep across all
tasks, triage is the deep dive into one failing task. Start with review-logs unless
the user already knows which task broke.
## Algorithms
All recipes live in `modelopt_recipes/general/speculative_decoding/<algorithm>.yaml`.
| Algorithm | Sheet | Family |
| --- | --- | --- |
| EAGLE3 | `references/algorithms/eagle3.md` | Autoregressive draft head |
| DFlash | `references/algorithms/dflash.md` | Block diffusion |
| DSpark | `references/algorithms/dspark.md` | DFlash backbone + Markov head + optional confidence head |
| Domino | `references/algorithms/domino.md` | DFlash backbone + GRU causal correction head |
DSpark and Domino are **DFlash variants**, not separate pipelines: same
`recipe_type: speculative_dflash`, same training script, same `dflash.*` config
namespace, selected by `dflash_architecture_config.projector_type`. Read
`references/algorithms/dflash.md` first, then the variant's sheet for the delta.
If the user's algorithm has no sheet yet, the stage procedures still apply — derive
the missing values from an existing launcher example for that algorithm
(`tools/launcher/examples/*/*/hf_*_<algorithm>.yaml`) and its recipe, then write the
sheet as you go. `references/algorithms/README.md` defines what a sheet must contain.
## End-to-end: a new model
1. Confirm the algorithm and find the closest existing launcher example.
2. **Configure** — write the pipeline YAML (`references/stages/configure.md`).
3. Preview with `--dryrun`, then submit:
```bash
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --yes
```
4. Register the job and set up monitoring per the **monitor skill**.
5. **Review logs** when it finishes (`references/stages/review-logs.md`).
6. **Triage** anything that failed (`references/stages/triage.md`), fix, re-run only
the failed tasks onward via `pipeline.task_N.skip=true`.
7. **Validate** once all tasks pass (`references/stages/validate.md`).
Model-support gaps that need code changes land in `modelopt/torch/speculative/` and
require a separate ModelOpt PR — the pipeline YAML alone cannot fix an unrecognized
architecture.
@@ -0,0 +1,43 @@
# Algorithm sheets
One file per draft-model algorithm. A sheet holds only what differs between
algorithms; the procedure lives in `../stages/`. `eagle3.md` and `dflash.md` are the
worked examples — read one before writing a new sheet.
When an algorithm is a **variant** of another (DSpark and Domino are both DFlash
backbones with a different head), its sheet documents only the delta and points at the
parent sheet. Don't restate the parent's pipeline, dump flags, or shared failures.
`dspark.md` and `domino.md` are the worked examples for that shape.
Keep sheets short. If something is true for every algorithm, it belongs in the
stage file instead.
## Required sections
Stage files reference these by heading, so use the headings verbatim.
| Heading | Contents |
| --- | --- |
| `## Pipeline tasks` | Table of task → script → purpose → output path. Task *count* varies per config — EAGLE3 offline is 4 tasks, DFlash offline is 2 — so describe the tasks this algorithm's examples actually use. |
| `## Recipe and training knobs` | The `modelopt_recipes/general/speculative_decoding/<algo>.yaml` path, plus the per-model overrides that usually need tuning. |
| `## Per-model adjustments` | The non-obvious knobs that vary by target model (attention type, MoE dims, tokenizer, `trust_remote_code`). |
| `## Success markers` | Per task, the log line that proves it worked, and the artifact it should leave behind. Consumed by review-logs and validate. |
| `## Quality gate` | The metric, where it appears in the log, and the pass threshold. |
| `## Known failures` | Error pattern → root cause → fix, for failures specific to this algorithm. Generic failures (OOM, NCCL, time limit) live in `../stages/triage.md`. |
## Adding a sheet
Source the facts from the repo rather than from memory:
- Launcher examples: `tools/launcher/examples/*/*/hf_*_<algo>.yaml` — task layout,
scripts, container images, GPU sizing.
- Scripts: `tools/launcher/common/` — `eagle3/` and `specdec/` hold the training and
hidden-state-dump entry points.
- Recipe: `modelopt_recipes/general/speculative_decoding/<algo>.yaml` — defaults.
- Implementation: `modelopt/torch/speculative/plugins/` — `hf_<algo>.py` and
`modeling_<algo>.py`.
Then add a row to the algorithm table in `../../SKILL.md`.
Every algorithm with a recipe in
`modelopt_recipes/general/speculative_decoding/` currently has a sheet.
@@ -0,0 +1,168 @@
# DFlash
Block-diffusion draft: predicts a whole block of `block_size` tokens in one forward
pass instead of autoregressively. Design details, results, and open items are in
`examples/speculative_decoding/doc/dflash.md`; the paper is arXiv:2602.06036.
Examples: `tools/launcher/examples/*/*/hf_online_dflash.yaml`,
`hf_offline_dflash.yaml`, `hf_streaming_dflash_multi_node.yaml`,
`specdec_bench_dflash_vllm.yaml`.
## Pipeline tasks
DFlash has three variants, and none uses EAGLE3 offline's 4-task shape: **online is 3
tasks, offline is 2, streaming is 3.** Read the task count off the config you're
using rather than assuming one.
**Online** (`hf_online_dflash.yaml`) — base model forwards during training:
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/specdec/dflash_online_training.sh` | Train the draft, then export | `<output_dir>/checkpoint-*`, `<output_dir>/exported-checkpoint-*` |
| task_1 | `common/specdec/vllm_smoke_test.sh` | Serve target + draft, verify responses | Smoke-test log |
| task_2 | `common/specdec/ar_eval_mtbench.sh` | MT-Bench per-category AR evaluation (1 GPU) | AR per category |
**Offline** (`hf_offline_dflash.yaml`) — for base models too large to forward
alongside training:
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/eagle3/dump_offline_data_vllm.sh` or `dump_offline_data_hf.sh` | Dump base hidden states | Hidden-state dump directory |
| task_1 | `common/specdec/dflash_online_training.sh` | Train on the dump, then export | `<output_dir>/exported-checkpoint-*` |
### Choosing the dump backend
The dump script is shared with EAGLE3, so the backend choice is the same three-way
pick described in `eagle3.md` (*Choosing the dump backend*). Both committed
offline examples are in play: MiniMax-M2.7 uses `dump_offline_data_vllm.sh`,
Qwen3-0.6B uses `dump_offline_data_hf.sh`.
For DFlash the choice is **not** cosmetic — it constrains the draft depth you can
capture. See below.
### Dump flags
- `--aux-layers dflash` selects DFlash's layer-selection **preset**. It is a keyword,
not a count: `--aux-layers` accepts only `eagle`, `dflash`, or an explicit
comma-separated id list (`collect_hidden_states/common.py`).
- **Draft depth is a separate flag, and only the vLLM backend exposes it.** The
captured ids come from `build_target_layer_ids(num_target_layers, num_draft_layers)`,
and `num_draft_layers` must equal the recipe's
`dflash.dflash_architecture_config.num_hidden_layers` or the dump silently captures
the wrong layers:
- **vLLM** — pass `--num-draft-layers <N>` (default 5).
- **HF / TRT-LLM** — no override exists; `resolve_aux_layers` hardcodes
`_DFLASH_DEFAULT_NUM_DRAFT_LAYERS = 5`. For a draft that is not 5 layers, you must
pass an explicit comma-separated id list to `--aux-layers`, or use the vLLM backend.
- `--answer-only-loss` and `--chat-template` — must agree with the training task's
`training.answer_only_loss` and `data.chat_template`.
> Both offline example YAMLs used to attach this constraint to `--aux-layers`, which
> carries no count. Their comments were corrected alongside this sheet — if you find
> the old wording anywhere else, the knob is `--num-draft-layers` (vLLM) or an
> explicit id list (HF / TRT-LLM).
Offline training additionally needs
`model.use_fake_base_for_offline=true` (loads only `lm_head` + `embed_tokens` rather
than the full base) and `data.offline_data_path` pointing at the dump. `data.mode` is
*derived* from which data-source field is set (`_check_mode_requirements` overwrites any
value passed in), so setting it is a no-op kept only for backward compatibility.
**Streaming** (`hf_streaming_dflash_multi_node.yaml`) — same NIXL RDMA transport as
streaming EAGLE3, splitting nodes into serve replicas plus DDP trainers. See
`common/eagle3/train_eagle_streaming.sh` for dispatch and sharding.
**Benchmark** (`specdec_bench_dflash_vllm.yaml`) — `common/specdec_bench/run.sh` with
`--speculative_algorithm DFLASH` and `--block_size`.
## Recipe and training knobs
`modelopt_recipes/general/speculative_decoding/dflash.yaml`, passed to
`dflash_online_training.sh` via `--config` with OmegaConf dotted overrides. Full table
in `examples/speculative_decoding/README.md#dflash-block-diffusion-for-speculative-decoding`.
| Override | Default | Note |
| --- | --- | --- |
| `dflash.dflash_block_size` | 8 | Tokens predicted per block. `training.training_seq_len` **must** be divisible by it. |
| `dflash.dflash_num_anchors` | 512 | Random anchor positions sampled per sequence |
| `dflash.dflash_loss_decay_factor` | 4.0 | Exponential decay gamma; 0 disables |
| `dflash.dflash_self_logit_distillation` | true | Logit distillation from the target |
| `dflash.dflash_architecture_config.num_hidden_layers` | 5 | Draft decoder layers — must equal the dump's draft depth (`--num-draft-layers` on vLLM; hardcoded 5 on HF / TRT-LLM) |
| `dflash.dflash_mask_token_id` | auto | See *Per-model adjustments* |
| `dflash.dflash_swa_window_size` | unset | Sliding-window attention for the draft; must be >= `dflash_block_size` |
| `dflash.dflash_export_rope_scaling` | `{}` | YaRN config injected at export so a short-window draft can serve long context |
| `training.learning_rate` | 6.0e-4 | |
| `training.training_seq_len` | 4096 | |
| `data.chat_template` | — | Required when `answer_only_loss=true` |
Export is automatic: after training, rank 0 exports every `checkpoint-<step>` to
`exported-checkpoint-<step>`, plus `exported-checkpoint-final` when
`modelopt_state.pth` sits directly in `output_dir`.
## Per-model adjustments
| Situation | What to change |
| --- | --- |
| Any model | Pin `dflash.dflash_mask_token_id` to a token that **already exists in the target's embedding** — the draft reuses the target's `embed_tokens`. Unset falls back to `tokenizer.mask_token_id`, which many tokenizers lack. MiniMax-M2.7 uses a reserved row (200054); Qwen3-8B uses 151669. |
| `answer_only_loss=true` (recipe default) | The chat template must contain `{% generation %}` / `{% endgeneration %}` tags. Most stock templates don't — supply one via `data.chat_template=<path>.jinja`. Each model keeps its own next to its example YAML (`examples/<Org>/<Model>/chat_template_train.jinja`); copy the closest one. |
| `trust_remote_code` MoE with an older transformers pin | Set `OVERRIDE_TRANSFORMERS` in the task environment (MiniMax-M2.7 needs 4.57.1). Set `ACCELERATE_CONFIG` when the model needs FSDP2 via accelerate config rather than transformers-native `ParallelismConfig`. |
| Very large MoE base | Use the offline variant with `model.use_fake_base_for_offline=true`; plain DDP suffices, so no FSDP2 patches. Set `MIXED_PRECISION: "no"` with `training.bf16=false` if the model requires it. |
| Draft trained at short context, served long | Set `dflash.dflash_export_rope_scaling` (YaRN); factor = target context / `training_seq_len`. |
| Multi-node | Set `NUM_NODES` in the environment; `HEAD_NODE_IP` is auto-detected from Slurm. |
## Success markers
| Task | Log evidence | Artifact |
| --- | --- | --- |
| Hidden-state dump (offline) | vLLM extraction completes over the input data | Dump directory populated |
| Training | `Training time: N seconds`, then `=== Exporting: <ckpt> → <export_dir> ===` and `=== Regression Check (...) ===` | `<output_dir>/checkpoint-*/trainer_state.json`, `<output_dir>/exported-checkpoint-*` |
| Smoke test | `Auto-detected draft model: ...`, `Server ready after Ns` | Smoke-test log with responses |
| AR eval | Per-category MT-Bench AR output | AR results |
| Benchmark | `Average_AL` in the saved results | JSON under `--save_dir` |
## Quality gate
DFlash gates in three places rather than on one acceptance-rate number.
**1. Training regression** — `common/check_regression.py` reads the latest
`trainer_state.json` and compares against env thresholds set in the YAML:
| Env var | Meaning |
| --- | --- |
| `MAX_FINAL_LOSS` | Final loss must be below this |
| `MIN_FINAL_ACC` | Final accuracy must be above this (any log key containing `acc`) |
Qwen3-8B online reference uses `MAX_FINAL_LOSS=5.0`, `MIN_FINAL_ACC=0.15`. Its
convergence baseline (8×B200, bs=1, seq_len=4096, 5-layer draft, block_size=16, 100K
samples, 1 epoch ≈ 12,500 steps) is in the YAML header — compare against it when
judging whether a run under-trained.
Note: `check_regression.py` is invoked with `|| true`, and it only warns when no
`trainer_state.json` exists. A green Slurm exit is not proof the gate ran — confirm
the `=== Regression Check ===` block is present in the log.
**2. Smoke test** — `MIN_ACCEPTANCE_LENGTH` env var (Qwen3-8B online uses 1.4) with
`NUM_SPEC_TOKENS` speculative tokens.
**3. Benchmark** — `Average_AL` (average acceptance length) from
`common/specdec_bench/run.sh`. Acceptance length is concurrency-independent, so it is
the primary metric even when the run trades timing fidelity for wall clock.
## Known failures
Generic infrastructure failures are in `../stages/triage.md`. These are
DFlash-specific:
| Error pattern | Root cause | Fix |
| --- | --- | --- |
| `seq_len (N) must be divisible by block_size (B)` | `training.training_seq_len` not a multiple of `dflash_block_size` | Adjust either value, or pad |
| `DFlash offline model cannot run eval/inference forward` | Offline conversion deletes base-model layers to save memory | Don't run eval on the offline model; reload the full base first |
| `DFlash offline model cannot run AR validation / pseudo_speculative_generate` | Same cause, hit via AR validation | Keep `training.estimate_ar=false` and `training.ar_validate_steps=0` in offline runs |
| `dflash_swa_window_size (N) must be >= dflash_block_size (B)` | Config validation | Raise the window or lower the block size |
| `The base model did not return hidden states required for DFlash training` | Base model's top-level forward ignores `output_hidden_states=True` | Usually a multimodal wrapper — needs a model-side fix |
| `ERROR: DRAFT_CKPT_DIR=... contains no exported-checkpoint-* directory` | Upstream training produced no draft | Fix training; do not chase the smoke test |
| vLLM rejects the speculative config / no DFlash method | DFlash landed in vLLM v0.22.0 (`vllm/v1/spec_decode/dflash.py`) | Use `vllm/vllm-openai:v0.22.1` or newer |
| Draft quality plateaus despite clean training | The dump captured the wrong layer *ids* while capturing the right **count**. A depth mismatch fails loudly — `DFlashModule` sizes its fusion layer from `len(config.target_layer_ids)`, so a count mismatch is a shape error. Equal-count-but-wrong-ids is the silent case, and `--aux-layers dflash` defaults to a 5-layer selection on **every** backend | Re-dump with `--num-draft-layers <N>` (vLLM), or an explicit `--aux-layers` id list (HF / TRT-LLM) matching what the draft recomputes from its own depth |
| Loss stalls high with `answer_only_loss=true` | Chat template lacks `{% generation %}` tags, so no positions contribute loss | Supply a template with generation tags |
| `dflash_dpace_alpha must be in (0, 1]` | Invalid D-PACE alpha | Correct the value |
| Qwen3-VL mRoPE / `mm_token_type_ids` errors | Qwen3-VL DFlash needs Transformers 5.3.0 or >=5.4.0 and the AutoProcessor's `mm_token_type_ids` | Match the version; don't drop processor outputs |
@@ -0,0 +1,115 @@
# Domino
**A DFlash variant, not a separate pipeline.** Domino is the DFlash draft backbone
plus a lightweight causal correction head — a GRU over the block's previously decoded
tokens producing a logit correction on the block suffix — selected with
`dflash_architecture_config.projector_type=domino`. It trains with a base/final dual
loss whose `lambda_base` weight decays from 1 to 0 over training (curriculum).
Read `dflash.md` first — the pipeline, dump flags, export behaviour, and generic
failure modes are all shared. This sheet covers only the delta.
Recipe: `modelopt_recipes/general/speculative_decoding/domino.yaml` (its
`metadata.recipe_type` is `speculative_dflash`, and every knob lives in the `dflash.*`
namespace).
Example: `tools/launcher/examples/Qwen/Qwen3-8B/hf_online_domino.yaml`. Reference:
SpecForge PR #571 (z-lab); drafter format `huggingface.co/Huang2020/Qwen3-8B-Domino-b16`.
## Pipeline tasks
The committed example is **online**, 2 tasks:
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/eagle3/make_dataset.sh` | Build training conversations (Daring-Anteater multi-turn SFT, 50K, `--full-conversations`) | `/scratchspace/data/train.jsonl` |
| task_1 | `common/specdec/dflash_online_training.sh` | Train the draft, then export | `<output_dir>/exported-checkpoint-*` |
`--full-conversations` matters: it keeps real assistant completions so
`answer_only_loss` has assistant spans to mask.
**The inference side is intentionally not wired up yet.** The Domino correction head
is not applied in `pseudo_speculative_generate` or in the serving stack, so the
example ships no vLLM smoke test and no MT-Bench AR eval. Do not treat their absence
as a broken config.
When that path lands, copy the smoke-test and AR-eval steps from
`hf_online_dflash.yaml` (its `task_1` and `task_2`) and append them to the Domino
config as **`task_2` and `task_3`** — Domino's `task_0`/`task_1` are already the
dataset build and training, so keep the source file's numbering and the destination's
distinct.
## Recipe and training knobs
Everything in `dflash.md` applies. Domino adds:
| Override | Recipe default | Note |
| --- | --- | --- |
| `dflash.dflash_architecture_config.projector_type` | `domino` | Selects the variant |
| `dflash.dflash_architecture_config.emb_dim` | 256 | GRU head embedding dim. **Required** |
| `dflash.dflash_architecture_config.gru_hidden_dim` | 1024 | GRU hidden dim. **Required** |
| `dflash.dflash_architecture_config.pure_draft_prefix_len` | 1 | Positions at block start kept as base logits only (no causal correction). Must be in `[0, block_size-1]` |
| `dflash.dflash_architecture_config.shift_label` | true | Next-token alignment — **only `true` is supported** |
| `dflash.dflash_lambda_base_start` | 1.0 | Curriculum start weight on the base loss |
| `dflash.dflash_lambda_base_decay_ratio` | 1.0 | Fraction of training over which `lambda_base` decays to 0 |
`dflash_self_logit_distillation` is **false** — Domino trains its own base/final CE
losses rather than distilling target logits. Recipe defaults also differ from
DFlash's: `block_size` 16, `num_anchors` 256, `num_train_epochs` 6,
`training_seq_len` 3072, `warmup_ratio` 0.04, `dflash_loss_decay_factor` 7.0.
`max_grad_norm: 1.0` is stated explicitly in the recipe but is *not* a delta — it is the
`transformers.TrainingArguments` default that `dflash.yaml` inherits by not setting it.
`ddp_find_unused_parameters: true` is **required**, not incidental: while
`lambda_base == 1` the head params are absent from the backward graph and DDP would
otherwise fail.
## Per-model adjustments
Everything in `dflash.md`'s table applies. Additionally:
| Situation | What to change |
| --- | --- |
| Any non-Qwen3 base | `domino.yaml` hardcodes `dflash_mask_token_id: 151669`, a Qwen3-specific unused id. `dflash.md`'s "unset falls back to `tokenizer.mask_token_id`" does **not** apply here — the pin is inherited, so a different base silently trains against a token that means something else. Override it. |
| Situation | What to change |
| --- | --- |
| Any model | **The Domino draft does not inherit the base model's GQA/FFN dims** — a fresh `Qwen3Config` already carries defaults, so `modify()`'s inherit-if-missing guard is a no-op. Set `num_attention_heads`, `num_key_value_heads`, `head_dim`, and `intermediate_size` explicitly. The Qwen3-8B reference drafter uses `32 / 8 / 128 / 12288`. |
| Any run | **Set `training.max_steps`.** The `lambda_base` curriculum is scheduled against `state.max_steps`; if it's unset the decay window collapses to one step and the curriculum is disabled (`lambda_base` 0 from the start). This warns rather than errors. The Qwen3-8B example sets `max_steps=2000`. |
## Success markers
Same as `dflash.md`: `Training time: N seconds`, then the `=== Exporting: ... ===`
and `=== Regression Check (...) ===` blocks, with `exported-checkpoint-*` on disk.
Because there is no smoke test or AR eval step, training completion plus a clean
export is the whole in-pipeline signal.
## Quality gate
**Do not trust in-training AR for Domino.** The recipe pins `estimate_ar: false` and
`ar_validate_steps: 0` deliberately: eval delegates to the DFlash backbone with the
correction head not applied, so reported acceptance rates are backbone-only. The code
logs this once as a warning — treat that warning as expected, not as a defect.
The training-regression gate from `dflash.md` applies and the Qwen3-8B example sets
it: `MAX_FINAL_LOSS=5.0`, `MIN_FINAL_ACC=0.15`, checked by `check_regression.py`
against `trainer_state.json`. Since no inference metric is produced, this is currently
the only automatic gate — and per `dflash.md` it is invoked with `|| true`, so confirm
the `=== Regression Check ===` block actually appears in the log.
## Known failures
Generic infrastructure failures are in `../stages/triage.md`; shared block-diffusion
failures (`seq_len` divisibility, offline eval, mask token, chat template) are in
`dflash.md`. Domino-specific:
| Error pattern | Root cause | Fix |
| --- | --- | --- |
| `Domino (projector_type='domino') requires ['emb_dim', 'gru_hidden_dim'] in dflash_architecture_config` | GRU head dims missing | Set both in `dflash_architecture_config` |
| `Domino currently supports shift_label=True (next-token alignment) only` | `shift_label=false` | Leave it at `true` |
| `pure_draft_prefix_len must be in [0, N] (block_size=B), got X` | Prefix length >= block size | Lower it below `block_size` |
| `DominoLambdaCallback: state.max_steps unset (<=0); lambda_base curriculum disabled` (warning) | `training.max_steps` not set | Set `training.max_steps`, else the curriculum never runs |
| `Domino eval uses the DFlash backbone only ...` (warning) | Correction head not applied at eval | Expected — do not chase it; evaluate after export once the inference path lands |
| DDP error about unused parameters | `ddp_find_unused_parameters` turned off | Keep it `true` |
| Draft trains but quality is poor | Draft dims left at `Qwen3Config` defaults instead of matching the base | Set the GQA/FFN dims explicitly |
@@ -0,0 +1,161 @@
# DSpark
**A DFlash variant, not a separate pipeline.** DSpark is the DFlash draft backbone
plus a lightweight sequential (Markov) head and an optional confidence head, selected
with `dflash_architecture_config.projector_type=dspark`. The Markov head adds a
prefix-dependent transition bias to the base logits, inducing a causal block
distribution (semi-autoregressive generation).
Read `dflash.md` first — the pipeline, dump flags, export behaviour, and generic
failure modes are all shared. This sheet covers only the delta.
Recipe: `modelopt_recipes/general/speculative_decoding/dspark.yaml` (its
`metadata.recipe_type` is `speculative_dflash`, and every knob lives in the `dflash.*`
namespace).
Examples: `tools/launcher/examples/*/*/hf_streaming_dspark_multi_node.yaml`,
`hf_streaming_dspark_warmstart.yaml`, `hf_dspark_ptq_*.yaml`.
## Pipeline tasks
Three committed shapes, all 2 tasks. No committed online or offline DSpark example;
the recipe defaults to `data.mode=online`, so an online run would use
`common/specdec/dflash_online_training.sh` exactly as DFlash does.
**Streaming, multi-node, from scratch** (Kimi-K2.6, MiniMax-M3):
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/eagle3/make_dataset.sh` | Build training conversations | `/scratchspace/data/train.jsonl` |
| task_1 | `common/eagle3/train_eagle_streaming.sh` | Streaming train (serve replicas + DDP trainers over NIXL RDMA), then export | `/scratchspace/export` |
`data.mode=streaming` with `model.use_fake_base_for_offline=true`.
**Streaming warm-start** (Nemotron-3.5-Lightning) — same two scripts, but continues a
**released** drafter from its published weights instead of training from scratch. The
key difference: the drafter architecture is **not** overridden in the YAML. It lives in
a per-model recipe under `modelopt_recipes/models/<org>/<model>/speculative_decoding/`,
because every field there is transcribed from the released checkpoint's own
`config.json` and must match it exactly. Keep drafter shape changes in that recipe and
cluster/serve knobs in the YAML.
**Drafter PTQ** (Qwen3-8B nvfp4; Nemotron-3.5-Lightning fp8 and nvfp4) — quantizes an
already-exported drafter. Calibration-free: every scale comes from the weights.
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/specdec/quantize_drafter.sh` | Quantize the drafter (CPU-only, ~1 min for an 8B-class draft) | `/scratchspace/export_quantized` |
| task_1 | `common/specdec_bench/quick_check.sh` | Acceptance length on MT-Bench, so the cost is visible | JSON results |
`DRAFTER_CKPT` accepts an exported checkpoint, an HF repo id, or a training
`output_dir` — for a directory it auto-selects the newest `exported-checkpoint-*`.
`w4a16_nvfp4` keeps activations in bf16; `nvfp4` does weight+activation with a fixed
`input_scale` of 1.0. To measure what quantization cost, re-run task_1 against the
unquantized drafter and compare acceptance length.
Streaming-specific environment (set in `task_1`, see
`common/eagle3/train_eagle_streaming.sh` for dispatch and rendezvous):
| Env var | Meaning |
| --- | --- |
| `EAGLE_CAPTURE_IDS` | Which base layers the serve side captures — the draft's target layer ids **+1**, plus the true final hidden layer |
| `SERVE_NODES` / `SERVE_TP` | How the node pool splits into serve replicas vs DDP trainers |
| `STREAMING_NUM_WORKERS` | Trainer-side streaming workers |
| `SERVE_MAX_MODEL_LEN`, `SERVE_MAX_NUM_SEQS`, `SERVE_GPU_MEM_UTIL`, `SERVE_READY_TIMEOUT` | Serve-replica limits |
| `EXPORT_EXTRA_ARGS` | Extra args at export (e.g. `--trust_remote_code`) |
No inference path is wired into these examples — neither ships a vLLM smoke test or
AR eval step.
## Recipe and training knobs
The general recipe is `modelopt_recipes/general/speculative_decoding/dspark.yaml`. A
warm-start run instead points `--config` at a **per-model** recipe under
`modelopt_recipes/models/<org>/<model>/speculative_decoding/`, which pins the drafter
shape to the released checkpoint.
Everything in `dflash.md` applies. DSpark adds:
| Override | Recipe default | Note |
| --- | --- | --- |
| `dflash.dflash_architecture_config.projector_type` | `dspark` | Selects the variant |
| `dflash.dflash_architecture_config.markov_rank` | 256 | Markov head low-rank dimension. **Required** and must be > 0 |
| `dflash.dflash_architecture_config.markov_head_type` | `vanilla` | `vanilla` (memoryless), `gated` (hidden-gated), or `rnn` (recurrent, closest to Domino's GRU) |
| `dflash.dflash_architecture_config.use_confidence_head` | true | Builds the per-position acceptance predictor |
| `dflash.dflash_ce_loss_alpha` | 0.1 | Cross-entropy term |
| `dflash.dflash_l1_loss_alpha` | 0.9 | TVD term — the DeepSpec defaults are L1/TVD-dominant |
| `dflash.dflash_confidence_head_alpha` | 1.0 | Confidence BCE term; requires `use_confidence_head=true` when > 0 |
Total loss is `ce_alpha*CE + l1_alpha*TVD + conf_alpha*confidence_BCE`.
`dflash_self_logit_distillation` is **false** for DSpark — it computes the target
distribution internally for the TVD and confidence terms, so the DFlash KD path is
unused. Recipe defaults also differ from DFlash's: `block_size` 16, `num_anchors` 256,
`num_train_epochs` 6, `training_seq_len` 3072, `warmup_ratio` 0.04.
## Per-model adjustments
Everything in `dflash.md`'s table applies. Additionally:
| Situation | What to change |
| --- | --- |
| Any non-Qwen3 base | `dspark.yaml` hardcodes `dflash_mask_token_id: 151669`, a Qwen3-specific unused id. `dflash.md`'s "unset falls back to `tokenizer.mask_token_id`" does **not** apply here — the pin is inherited, so a different base silently trains against a token that means something else. Override it. |
| Situation | What to change |
| --- | --- |
| Any model | **The DSpark draft does not inherit the base model's GQA/FFN dims.** Set `num_attention_heads`, `num_key_value_heads`, `head_dim`, and `intermediate_size` in `dflash_architecture_config` explicitly, or you get a silently wrong-shaped draft. Kimi-K2.6 uses `num_hidden_layers=6, num_key_value_heads=8, intermediate_size=18432`; MiniMax-M3 uses `intermediate_size=12288`. |
| Streaming | `EAGLE_CAPTURE_IDS` must be the draft's target layer ids +1 plus the final hidden layer. Kimi-K2.6: `[2,13,25,36,48,59,61]` for a 6-layer draft. Getting the final layer wrong caps acceptance length rather than erroring. |
| Sparse-attention base (e.g. MiniMax-M3 MSA) | Set `SERVE_BLOCK_SIZE` to the base's `sparse_block_size` (M3: 128) |
| Serve container lacks tensorboard | `training.report_to=none`, else trainer init crashes |
| Tokenizer can't emit assistant masks (e.g. Kimi slow tokenizer) | `training.answer_only_loss=true` still works — masks are recovered from token ids |
## Success markers
Same as `dflash.md`. Because the streaming examples have no smoke test or AR eval,
the only in-pipeline evidence is training progress plus the export landing in
`/scratchspace/export`.
## Quality gate
**Do not trust in-training AR for DSpark.** The recipe pins `estimate_ar: false` and
`ar_validate_steps: 0` deliberately: eval runs the DFlash backbone only, with the
Markov head not applied, so any reported AR reflects the backbone alone rather than
the trained model.
`pseudo_speculative_generate` *is* overridden for DSpark (unlike Domino), so a
non-offline model can generate correctly — but the offline/streaming path deletes base
layers and refuses. Evaluate by exporting and running the offline acceptance-length
harness separately.
Otherwise the training-regression gate from `dflash.md` (`MAX_FINAL_LOSS`,
`MIN_FINAL_ACC` via `check_regression.py`) applies; none of the streaming examples set
those thresholds.
The **PTQ** examples are the exception — they do produce a real inference metric, via
`quick_check.sh` on MT-Bench with `--speculative_algorithm DSPARK`. There is no fixed
pass threshold; the gate is relative. Run the same benchmark against the unquantized
drafter and compare acceptance length, so the quantization cost is explicit rather than
assumed. Published reference points on Qwen3-8B: the weight+activation `nvfp4` format
costs roughly 3.9% acceptance length, which is why the examples default to
`w4a16_nvfp4` (weights only, activations in bf16). Adding `fc` to the quantized set
costs about 0.7% acceptance length (3.0186 vs 3.0392) for ~3% more size saved.
## Known failures
Generic infrastructure failures are in `../stages/triage.md`; shared block-diffusion
failures (`seq_len` divisibility, offline eval, mask token, chat template) are in
`dflash.md`. DSpark-specific:
| Error pattern | Root cause | Fix |
| --- | --- | --- |
| `DSpark (projector_type='dspark') requires 'markov_rank' (> 0) in dflash_architecture_config` | Markov head dimension missing | Set `dflash_architecture_config.markov_rank` |
| `DSpark requires markov_rank > 0, got N` | Non-positive value | Set a positive rank |
| `Unsupported markov_head_type: '...'. Expected 'vanilla', 'gated' or 'rnn'` | Typo or unsupported head | Use one of the three |
| `dflash_confidence_head_alpha > 0 but the confidence head was not built` | Loss term enabled without the head | Set `dflash_architecture_config.use_confidence_head=true`, or set the alpha to 0 |
| `DSpark offline model cannot run AR validation / pseudo_speculative_generate` | Offline/streaming conversion deleted base layers | Keep `estimate_ar=false` and `ar_validate_steps=0`; evaluate after export |
| Draft trains but acceptance length is poor | Draft dims left at defaults instead of matching the base | Set the GQA/FFN dims explicitly (see *Per-model adjustments*) |
| Acceptance length capped despite clean training (streaming) | `EAGLE_CAPTURE_IDS` final layer wrong, or the vLLM aux-capture fix (vllm#46788) missing | Correct the ids; use a container with the fix |
| Trainer init crash on a serve container | tensorboard absent | `training.report_to=none` |
| Drafter PTQ produces a broken draft, or export fails on a packed tensor | q/k/v were quantized. DFlash-family drafters build their fused context-KV projection by reading `qkv_proj.weight` raw, which cannot be a packed tensor | The exclusions are **mandatory**, not tuning: `--exclude '*q_proj*' '*k_proj*' '*v_proj*' '*qkv_proj*'`. `o_proj` and the MLP still quantize; `*fc*` is optional (~3% more size saved for ~0.7% acceptance length on Qwen3-8B) |
| Benchmark reports a poor or nonsensical acceptance length | `--draft_length` was passed. DSPARK and DFLASH read `--block_size` instead, and it must match the drafter's block size | Pass `--block_size <N>` matching the drafter, plus `--speculative_algorithm DSPARK` |
| Warm-start run diverges or the drafter shape mismatches the released checkpoint | Drafter architecture overridden in the launcher YAML | Leave shape fields in the per-model recipe — they are transcribed from the released `config.json` and must match exactly |
@@ -0,0 +1,109 @@
# EAGLE3
Draft head trained on hidden states dumped from the target model. Examples:
`tools/launcher/examples/*/*/hf_offline_eagle3.yaml` (and the `hf_online_*`,
`hf_streaming_*` variants).
## Pipeline tasks
The offline configuration is 4 tasks; each passes artifacts to the next through a
shared `/scratchspace`.
| Task | Script | Purpose | Output |
| --- | --- | --- | --- |
| task_0 | `common/vllm/query.sh` or `common/tensorrt_llm/query.sh` | Data synthesis — serve the target model, generate prompt/response pairs | `/scratchspace/data/*.jsonl` |
| task_1 | `common/eagle3/dump_offline_data_vllm.sh` (or `_hf.sh` / `dump_offline_data.sh`) | Forward the target model, save hidden states | `/scratchspace/offline_hidden_states/*.pt` |
| task_2 | `common/eagle3/train_eagle.sh` | Train the draft head, then export | `/scratchspace/eagle3/model.safetensors`, `/scratchspace/export/` |
| task_3 | `common/specdec_bench/quick_check.sh` | Benchmark acceptance rate and throughput | JSON result files |
### Choosing the dump backend
| Backend | Script | When to use |
| --- | --- | --- |
| vLLM | `common/eagle3/dump_offline_data_vllm.sh` | **Default.** Broad coverage via vLLM's native hidden-state extractor. |
| HF | `common/eagle3/dump_offline_data_hf.sh` | VLMs / multimodal, custom-code models, sliding-window attention (TRT-LLM can't serve these). Uses `device_map="auto"`. |
| TRT-LLM | `common/eagle3/dump_offline_data.sh` | Pure-text models with TRT-LLM support; pass `--tp <TP>` and `--moe-ep <EP>`. |
Rule of thumb: **HF** if the model is a VLM or uses sliding-window attention; **vLLM**
otherwise. TRT-LLM only when you specifically want its kernels for a supported
plain-text model.
## Recipe and training knobs
`modelopt_recipes/general/speculative_decoding/eagle3.yaml`, passed to
`train_eagle.sh` via `--config` with dotted overrides:
| Override | Note |
| --- | --- |
| `model.model_name_or_path` | Target checkpoint |
| `data.offline_data_path` | task_1 output directory |
| `training.output_dir` | Draft checkpoint destination |
| `training.training_seq_len` | Lower it first when training OOMs |
| `training.per_device_train_batch_size` | Lower it next when training OOMs |
| `training.learning_rate` | Lower it when loss is NaN or diverging |
| `training.ar_validate_steps` | Set to run AR validation during training |
`task_3` selects the algorithm at benchmark time with
`--speculative_algorithm EAGLE3`.
## Per-model adjustments
| Situation | What to change |
| --- | --- |
| Requires `--trust-remote-code` | Add to `task_0` server args (before the `--` separator) **and** to `task_3` benchmark args |
| MoE with large expert hidden dim | Set `intermediate_size` under `eagle.eagle_architecture_config` in the recipe to match the model's `moe_intermediate_size`. There is no `eagle_config.json` — the draft architecture lives in the recipe |
| Custom tokenizer (e.g. tiktoken) | Set `TIKTOKEN_RS_CACHE_DIR` to a pre-populated cache path in `task_0` and `task_1` |
| VLM | Use `dump_offline_data_hf.sh` — the text-only path, no vision encoder invoked |
| Sliding-window attention | TRT-LLM backend won't work; use HF or vLLM |
| Architecture unrecognized by training | Needs code changes in `modelopt/torch/speculative/` — a separate ModelOpt PR |
## Success markers
| Task | Log evidence | Artifact |
| --- | --- | --- |
| task_0 | "Saved N samples", or a progress bar completing | `/scratchspace/data/*.jsonl` |
| task_1 | "Successfully processed N conversations" | `/scratchspace/offline_hidden_states/*.pt` |
| task_2 | Training loss decreasing, "export complete" | `/scratchspace/eagle3/model.safetensors`, `/scratchspace/export/` |
| task_3 | `Average Acceptance Length ... ratio: X.XX` | JSON result files |
## Quality gate
The `task_3` log prints:
```text
Average Acceptance Length {'accept': X, 'count': Y, 'ratio': Z.ZZ}
```
The `ratio` field is the acceptance rate (AR).
| Criterion | Threshold |
| --- | --- |
| AR (MT-Bench) | >= 2.1 |
**This gate is not self-enforcing.** `quick_check.sh` is a 27-line pass-through to
`specdec_bench/run.py`; neither reads a threshold nor exits non-zero on a low AR, and
no EAGLE3 launcher example sets one. `>= 2.1` is a human review threshold. Extract
`ratio` from the task_3 log and compare it yourself — a COMPLETED task_3 is not
evidence the AR passed. (Same caution as the DFlash regression gate below, which runs
under `|| true`.)
## Known failures
Generic infrastructure failures are in `../stages/triage.md`. These are
EAGLE3-specific:
| Error pattern | Root cause | Fix |
| --- | --- | --- |
| `No such file or directory: dump_offline_data_vllm.sh` | Wrong script path in YAML | Use the correct path under `common/eagle3/` |
| `FileNotFoundError: /scratchspace/data` | task_0 failed or produced no output | Re-run task_0, or point `--input-data` at existing data |
| `FileNotFoundError: /scratchspace/offline_hidden_states` | task_1 failed or produced no output | Re-run task_1 |
| `FileNotFoundError: /scratchspace/export` | task_2 or its export step failed | Re-run task_2; check export output |
| `RuntimeError` / unsupported arch during dump | Model not supported by the TRT-LLM backend | Switch to `dump_offline_data_hf.sh` or `dump_offline_data_vllm.sh` |
| No `.pt` files in the dump output dir | Extraction produced nothing | Check `--max-seq-len` and the input data format |
| `KeyError` / `AttributeError` loading the model in task_2 | Architecture not recognized by EAGLE3 training | Needs code changes in `modelopt/torch/speculative/` |
| Loss is NaN or diverging | LR too high, or data quality issue | Reduce `training.lr`; check the hidden-state data |
| `export_hf_checkpoint.py` fails | Training produced an incomplete checkpoint | Check `/scratchspace/eagle3/` for `model.safetensors` |
| Empty `/scratchspace/data/` after task_0 | `query.py` ran but wrote nothing | Check `--data` path exists and contains prompts; check `query.py` logs |
| Server fails to load the draft model in task_3 | Draft config incompatible with the engine | Check the exported draft's `config.json` against the engine version |
| vLLM reports EAGLE3 not supported | vLLM version too old | Use a newer vLLM container |
| AR below threshold / exit code 1 | Draft quality too low | More epochs or data, or hyperparameter tuning |
@@ -0,0 +1,53 @@
# Stage 1 — Configure a new model
Create `tools/launcher/examples/<Org>/<Model>/<config>.yaml` by **copying the closest
existing example and adapting it**. Pick a reference with the same algorithm and the
same shape as the target (dense vs MoE, similar size) from `tools/launcher/examples/`
— e.g. the Qwen3-8B config for a dense model.
The task structure, args, containers, and GPU/node sizing are all visible in the
existing examples — infer them from a reference rather than hand-rolling. This file
covers only what the examples don't make obvious.
## Step 1 — Pick the algorithm and the variant
Example filenames encode both: `hf_<mode>_<algorithm>.yaml`, where mode is `offline`
(dump hidden states first, then train on them), `online` (forward the base model at
training time), or `streaming`.
```bash
ls tools/launcher/examples/*/*/hf_*_<algorithm>.yaml
```
Offline is the default choice when the target model is too large to forward
alongside training. Task count follows from the variant, not from the algorithm — do
not assume a fixed number of tasks; copy the reference's layout.
## Step 2 — Fill in the algorithm-specific values
From the algorithm sheet (`../algorithms/<algorithm>.md`):
- **Pipeline tasks** — which script each task runs, and the artifact paths they pass
between each other.
- **Recipe and training knobs** — the recipe path for the training task, and which
overrides this model needs.
- **Per-model adjustments** — the non-obvious knobs that vary by target model.
For offline variants, the hidden-state dump task usually offers more than one
backend (vLLM / HF / TRT-LLM). The sheet's *Pipeline tasks* section says how to pick.
## Step 3 — Size the job
Copy node/GPU counts from the reference example, then sanity-check against the
target: the base model's BF16 weights must fit in the allocated GPU memory for the
serving and dump tasks, so scale `tensor-parallel-size`, `gpus_per_node`, or `nodes`
if the target is larger than the reference.
## Step 4 — Preview
```bash
cd tools/launcher
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml --dryrun
```
Check the resolved scripts, paths, and containers before submitting for real.
@@ -1,21 +1,13 @@
---
name: eagle3-review-logs
description: >
Review EAGLE3 pipeline experiment logs from the launcher's experiments/ directory.
Summarizes pass/fail status for all 4 tasks, diagnoses failures with root causes
and fixes, and flags warnings. Use when the user asks to review job logs,
check experiment results, or diagnose why a specific task failed.
user_invocable: true
---
# Stage 2 — Review experiment logs
# Review EAGLE3 Experiment Logs
Analyze output logs from a pipeline run launched via `launch.py` or `slurm.py`, and
produce a pass/fail summary across all tasks. For a deep dive into one failing task,
go to `triage.md` instead.
Analyze output logs from an EAGLE3 pipeline run launched via `launch.py` or `slurm.py`.
## Step 0 — Find the experiment
## Step 0 — Find experiment logs
Locate the experiment directory. The default is `experiments/` relative to the launcher root,
or wherever `--job-dir` was pointed.
The default job directory is `experiments/` relative to the launcher root, or
wherever `--job-dir` was pointed.
```bash
ls -td experiments/cicd/cicd_* | head -10
@@ -25,9 +17,9 @@ If no experiments exist, ask the user for the directory.
## Step 1 — Read all task logs
Each experiment has one subdirectory per task (0–3). Log filenames vary by launch mode
(Slurm writes `sbatch_*.out`, local Docker writes `*.log`), so match log files generally and
read the tail of each in a single Bash call — errors surface at the end:
Each experiment has one subdirectory per task. Log filenames vary by launch mode
(Slurm writes `sbatch_*.out`, local Docker writes `*.log`), so match log files
generally and read the tail of each in a single Bash call — errors surface at the end:
```bash
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
@@ -39,11 +31,12 @@ done
For each task log, check:
- **Exit / cancellation**: `DUE TO TIME LIMIT`, `FAILED`, signal (e.g., `signal 15`)
- **Python exceptions / tracebacks**: last exception is usually the root cause
- **Exit / cancellation**: `DUE TO TIME LIMIT`, `FAILED`, signal (e.g. `signal 15`)
- **Python exceptions / tracebacks**: the last exception is usually the root cause
- **CUDA errors**: OOM, NCCL timeout
- **Slurm state**: COMPLETED, FAILED, TIMEOUT, OUT_OF_MEMORY
- **Success indicators**: "Saved N samples", "Successfully processed N conversations", training loss line, AR output
- **Success indicators**: see *Success markers* in `../algorithms/<algorithm>.md` —
each task has a specific log line that proves it worked
## Step 3 — Produce report
@@ -52,14 +45,15 @@ Output a structured markdown report:
### Summary
- Overall status: PASSED / FAILED / MIXED / PARTIAL
- Task breakdown: e.g., task_0 TIMEOUT, task_1 FAIL, task_2 skipped, task_3 skipped
- Task breakdown: e.g. task_0 TIMEOUT, task_1 FAIL, task_2 skipped, task_3 skipped
### Task Results
For each task (0–3):
For each task:
**Task N — \<name\>: PASS / FAIL / TIMEOUT**
- Key output: (e.g., "3277/3295 samples generated" or "Script not found")
- Key output: (e.g. "3277/3295 samples generated" or "Script not found")
- Error (if failed): quoted error message, max 10 lines
- Root cause: one-line diagnosis
- Suggested fix: actionable step
@@ -70,26 +64,24 @@ Non-fatal issues worth noting (near-OOM, tokenizer warnings, slow throughput).
## Step 4 — Suggest next steps
Based on results:
- If a task failed due to a known issue, suggest the fix and how to re-run from that task:
- If a task failed due to a known issue, suggest the fix and how to re-run from that
task:
```bash
uv run launch.py --yaml examples/<Org>/<Model>/hf_offline_eagle3.yaml \
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml \
pipeline.task_0.skip=true \
--yes
```
- If the failure pattern looks new, suggest capturing it in the team's internal triage
tracker, and use `/eagle3-triage` for a deeper diagnosis.
- If all tasks passed, suggest running `/eagle3-validate` to confirm AR meets threshold.
- If the failure pattern looks new, suggest capturing it in the team's internal
triage tracker, and use `triage.md` for a deeper diagnosis.
- If all tasks passed, move to `validate.md` to confirm the quality gate.
## Known benign patterns (do NOT mark as failures)
| Pattern | Explanation |
|---|---|
| --- | --- |
| vLLM server exit code 143 | SIGTERM — server was killed after queries completed. Expected. |
| `CANCELLED AT ... DUE TO TASK FAILURE` after `exit code: 0` | Slurm cleanup of worker nodes after main task succeeded. |
| `CANCELLED AT ... DUE TO TASK FAILURE` after `exit code: 0` | Slurm cleanup of worker nodes after the main task succeeded. |
| `destroy_process_group() was not called` | Benign PyTorch shutdown warning. |
| `tokenizer class ... not equal to the registered tokenizer class` | Harmless tokenizer mismatch warning. |
@@ -0,0 +1,129 @@
# Stage 3 — Triage a failed run
Diagnose a failure in the draft-training pipeline: identify the failing task, find
the root cause, and give a fix plus a re-run command.
## Step 0 — Locate the experiment
Ask the user for one of:
- The experiment directory (e.g. the `--job-dir` passed to `launch.py` / `slurm.py`)
- The model name / YAML they ran
Find recent experiments under the job directory:
```bash
ls -td experiments/cicd/cicd_* | head -10
# or wherever --job-dir was pointed
```
Each experiment directory contains one subdirectory per task, each with a log file
whose name varies by launch mode (Slurm: `sbatch_*.out`, local Docker: `*.log`).
## Step 1 — Fetch logs for the failed task
Match the log files generally and read the tail of each — errors appear at the end:
```bash
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -200 "$f"; echo
done
```
Find the first task with a non-zero exit code or an error message. Later tasks
usually fail only because an upstream artifact is missing, so fix the first one.
## Step 2 — Diagnose
Work through two tables. Start here — these failures are independent of the
algorithm and account for most runs:
| Error pattern | Root cause | Fix |
| --- | --- | --- |
| Server never becomes healthy (hangs at the health check) | Model too large for the allocated GPUs, or a server startup crash | Compare BF16 weight size against total allocated GPU memory; increase TP and/or nodes |
| `CUDA out of memory` **while loading weights** (before the KV cache is allocated) | The weights themselves don't fit | Increase `--tensor-parallel-size`, add nodes, or switch backend. `--max-model-len` will **not** help — it doesn't change weight memory. |
| `CUDA out of memory` **after weights load** — KV-cache allocation, or during a forward pass | Activation / KV-cache pressure | Reduce `--max-model-len`, batch size, or concurrency; raising TP also helps by splitting the cache |
| `CUDA out of memory` during the hidden-state dump | Model too large for the chosen backend | Switch to a `device_map="auto"` backend, or increase TP |
| `CUDA out of memory` during training | Batch or sequence length too large | Reduce the recipe's training batch size or sequence length (see the algorithm sheet's *Recipe and training knobs*) |
| `CUDA out of memory` at benchmark | Target plus draft exceeds GPU memory | Increase TP |
| `pyxis: child terminated with signal 15` | SIGTERM — usually OOM | Increase TP or switch backends |
| `NCCL timeout` / `NCCL error` | Multi-node communication failure | Retry; reduce EP |
| `CANCELLED ... DUE TO TIME LIMIT` | Slurm wall-clock limit too short | Increase `--time`. Note that `afterany` dependencies let the next task start anyway. |
| `trust_remote_code` error | Model needs custom code but the flag isn't set for **that** task | Set it on **every** task that loads the model — see the spellings below |
| Vocab / tokenizer error | Missing tokenizer cache (e.g. a tiktoken cache) | Point the relevant cache env var at a pre-populated path |
| Architecture not supported by the serving engine | Engine version too old for this model | Try a newer container image |
### `trust_remote_code` spellings
The flag is spelled differently per task type, so setting it once is not enough — a
custom-code model needs it everywhere it is loaded:
| Task type | How to set it |
| --- | --- |
| Serving / benchmark | CLI flag before the `--` separator: `--trust-remote-code` (vLLM) or `--trust_remote_code` (trtllm-serve) |
| Hidden-state dump | Nothing to set — `dump_offline_data_hf.sh` and `dump_offline_data_vllm.sh` pass `--trust_remote_code` unconditionally. The TRT-LLM `dump_offline_data.sh` passes no such flag and offers no env var, so a custom-code model needs the HF or vLLM backend |
| Training | `model.trust_remote_code=true` as an OmegaConf override |
| Streaming serve replicas | `SERVE_EXTRA_ARGS: "--trust-remote-code"` |
| Export | `EXPORT_EXTRA_ARGS: "--trust_remote_code"` |
Then check *Known failures* in `../algorithms/<algorithm>.md` for failures specific
to this algorithm — wrong script paths, missing scratchspace artifacts, export
failures, draft-config incompatibilities.
## Step 3 — Check for new-model issues
If the user is adding support for a new model, re-read *Per-model adjustments* in
`../algorithms/<algorithm>.md` and confirm each applicable knob was set — attention
type, MoE dimensions, custom tokenizer, and `trust_remote_code` are the usual
offenders.
If the architecture isn't recognized by the training code at all, that needs changes
in `modelopt/torch/speculative/` and a separate ModelOpt PR — no YAML change fixes it.
## Step 4 — Suggest fix and next steps
Provide:
1. **Root cause** — one-line summary
2. **Fix** — the specific config change, code edit, or command
3. **How to re-run** — skip earlier successful tasks by pointing at the existing
scratchspace artifacts
Re-runs work by adding `pipeline.task_N.skip=true` for each task you want to skip.
**Read the task list out of the config first** — task count varies (EAGLE3 offline is
4, DFlash offline is 2, Domino is 2), so there is no fixed set of skip flags:
```bash
grep -n '^ task_[0-9]*:' examples/<Org>/<Model>/<config>.yaml
```
To resume from a failed task, skip every task before it. For a 4-task EAGLE3 offline
config whose `task_2` failed:
```bash
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml \
pipeline.task_0.skip=true \
pipeline.task_1.skip=true \
--yes
```
To run one task standalone, skip every other task in that config. For the same 4-task
config, running only `task_1`:
```bash
uv run launch.py --yaml examples/<Org>/<Model>/<config>.yaml \
pipeline.task_0.skip=true \
pipeline.task_2.skip=true \
pipeline.task_3.skip=true \
--yes
```
Both are EAGLE3-offline examples — translate the flags to the config at hand rather
than copying them verbatim.
## Step 5 — Record the failure pattern
If you hit a failure pattern not seen before, capture it in the team's internal
triage tracker — symptom, root cause, and fix — so the next engineer benefits. If
it's algorithm-specific, add a row to *Known failures* in the algorithm sheet; if it
applies to every algorithm, add it to Step 2 above.
@@ -0,0 +1,135 @@
# Stage 4 — Validate a completed run
Verify that a pipeline run completed successfully end-to-end and meets its quality
gate.
## Step 0 — Identify the experiment
Find the most recent experiment directory (or ask the user for the path):
```bash
ls -td experiments/cicd/cicd_* | head -5
```
Each experiment directory has one subdirectory per task, each containing a log file
whose name varies by launch mode (Slurm: `sbatch_*.out`, local Docker: `*.log`).
## Step 1 — Check task outcomes
Match the log files generally and read the tail of each:
```bash
find experiments/<exp_id>/ -type f \( -name '*.out' -o -name '*.log' \) | sort | while read -r f; do
echo "=== $f ==="; tail -50 "$f"; echo
done
```
Every task must complete without error. Look for:
- `exit code: 0` or no error — success
- `DUE TO TIME LIMIT` — timeout
- `FAILED` / `signal` / exception traceback — failure
If any task failed, go to `triage.md` instead.
## Step 2 — Verify artifacts exist
Check each task produced its expected output. The per-task log evidence and artifact
paths are in *Success markers* in `../algorithms/<algorithm>.md`.
A success line in a log is not proof the artifact survived — the next task reads it
from a shared `/scratchspace`, where it may be missing, empty, or unreadable. **When
you can reach the cluster, check the filesystem directly** and treat a missing or
zero-byte artifact as a validation failure:
```bash
test -s <artifact_path> && echo "ok: $(du -sh <artifact_path>)" || echo "MISSING/EMPTY"
ls -la <artifact_dir>/ | head
```
Fall back to log evidence only when the cluster isn't reachable, and say so in the
report rather than implying the artifacts were verified.
## Step 3 — Check the quality gate
Read *Quality gate* in `../algorithms/<algorithm>.md` first — **not every algorithm
produces an in-pipeline metric**, so what you check depends on the sheet:
- **Sheet defines a benchmark metric** (e.g. EAGLE3's MT-Bench AR, DFlash's
`Average_AL`) — extract it from the benchmark task's log and compare against the
threshold **yourself**. No benchmark step in this repo enforces one:
`common/specdec_bench/quick_check.sh` and `run.sh` pass straight through to
`specdec_bench/run.py`, which reads no threshold and never exits non-zero on a low
metric, and no launcher example sets one. A green benchmark task therefore says the
run finished, not that the metric passed. Reporting PASS without having read the
number off the log is a false PASS on the run's headline metric.
- **Sheet defines no inference metric** (currently Domino and DSpark — their eval path
runs the DFlash backbone with the new head bypassed, and Domino ships no benchmark
task at all) — do **not** go looking for a benchmark log. Report the training
regression gate instead, and state plainly that acceptance quality requires a
separate evaluation of the exported checkpoint. Never report a backbone-only
acceptance rate as the model's result.
Where the gate is the training regression check (`check_regression.py` against
`trainer_state.json`), confirm the `=== Regression Check ===` block is actually
present in the log — it is invoked with `|| true` and only warns when no
`trainer_state.json` exists, so a green exit does not prove it ran.
## Step 4 — Check training quality
In the training task's log look for:
- **Final training loss** — should be decreasing, not NaN
- **Metric validation during training** — if the recipe enabled periodic validation
- **Number of training steps** — confirms full training duration
## Step 5 — Produce validation report
```markdown
## Speculative Decoding Pipeline Validation Report
**Experiment:** <exp_dir>
**Model:** <model_name>
**Algorithm:** <algorithm>
**Date:** <date>
**Pipeline config:** <yaml_path>
### Task Status
| Task | Name | Status | Notes |
|------|------|--------|-------|
| 0 | Data synthesis | PASS/FAIL/TIMEOUT | N samples generated |
| 1 | Hidden state dump | PASS/FAIL | N .pt files |
| 2 | Training + export | PASS/FAIL | Final loss: X.XX |
| 3 | Benchmark | PASS/FAIL | AR: X.XX |
### Quality Gate
- <metric>: X.XX (threshold: <threshold>) — PASS/FAIL
### Training Summary
- Final loss: X.XX
- Training steps: N
- Metric during training: X.XX (if validated)
### Overall: PASS / FAIL
<one-line summary>
```
Adjust the task rows to the tasks this config actually ran — task count varies by
algorithm and variant.
## Step 6 — Suggest next steps
**If PASS:**
- Record the verified result (and checkpoint path) in the team's internal triage
tracker
- This model is now a candidate to add as a launcher example in a dedicated PR
**If FAIL:**
- Identify which task or metric failed
- Go to `triage.md` for diagnosis
- For a low acceptance rate, diagnose the specific cause from the run (training loss
curve, data volume/quality, draft capacity, hyperparameters) and suggest fixes
targeted to that scenario — a low rate can have many causes, so avoid a generic
checklist.
+1 -1
View File
@@ -80,7 +80,7 @@ Available skills:
|---|---|---|
| `/review-logs` | After job completion/failure | Analyze logs, diagnose failures, JUnit XML |
| `/wait-for-jobs` | After detached submission | Poll experiment status |
| `/eagle3-new-model` | Adding a new EAGLE3 model | Generate pipeline YAML |
| `/speculative-decoding` | Draft-model pipeline work (EAGLE3, DFlash, …) | Configure a new model, review logs, triage, validate |
## CI Integration
@@ -27,7 +27,9 @@ pipeline:
args:
- --input-data /hf-local/modelopt/MiniMax-M2.7-synthetic-data-clean-v2
- --output-dir /scratchspace/dflash_minimax_m2.7_hidden_states
# Must match the draft model's num_hidden_layers (recipe default: 5).
# Layer-selection preset (keyword, not a count). The draft depth it resolves
# to is --num-draft-layers (default 5), which must match the draft model's
# dflash_architecture_config.num_hidden_layers set on task_1 below.
- --aux-layers dflash
- --answer-only-loss
- --chat-template examples/MiniMax/MiniMax-M2.7-DFlash/chat_template_train.jinja
@@ -24,7 +24,13 @@ pipeline:
args:
- --input-data modules/Model-Optimizer/examples/dataset/synthetic_conversations_1k.jsonl
- --output-dir /scratchspace/dflash_qwen3_0.6b_hidden_states
# Must match the draft model's num_hidden_layers (recipe default: 5).
# Layer-selection preset (keyword, not a count). It resolves to a 5-layer
# draft, which must match the draft model's num_hidden_layers (recipe
# default: 5). The HF dump has no --num-draft-layers override, so a draft of a
# different depth needs an explicit id list here -- and the list is not
# free-form: the draft recomputes its own ids from its depth via
# build_target_layer_ids, so the list must equal that result or the dump
# captures the wrong layers while still having the right count.
- --aux-layers dflash
- --answer-only-loss
- --chat-template examples/Qwen/Qwen3-0.6B/chat_template_train.jinja