mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
skills(specdec): add DFlash2 and LiLiCorr algorithm sheets
Three recipes shipped after the skill tree landed in #2201 and had no sheet: dflash2.yaml, lilicorr.yaml and lilicorr_conv.yaml. The stage procedures already applied to them; only the per-algorithm values were missing. dflash2.md and lilicorr.md follow the variant shape README.md defines -- delta only, pointing at dflash.md for the shared pipeline. One sheet covers both LiLiCorr recipes, since lilicorr_conv.yaml is the same algorithm with DFlash2's sublayer convolutions rather than a separate one; the delta has its own section. Facts are sourced from the recipes, the launcher examples and the plugin validation paths rather than from the papers, so the Known failures tables carry the error strings a run actually prints. Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Ye Yu <yeyu@nvidia.com>
This commit is contained in:
@@ -59,12 +59,17 @@ All recipes live in `modelopt_recipes/general/speculative_decoding/<algorithm>.y
|
||||
| DFlash | `references/algorithms/dflash.md` | Block diffusion |
|
||||
| DSpark | `references/algorithms/dspark.md` | DFlash backbone + Markov head + optional confidence head |
|
||||
| Domino | `references/algorithms/domino.md` | DFlash backbone + GRU causal correction head |
|
||||
| DFlash2 | `references/algorithms/dflash2.md` | DFlash backbone + sublayer convolutions + candidate selector |
|
||||
| LiLiCorr | `references/algorithms/lilicorr.md` | DFlash backbone + candidate-lattice reranker (covers `lilicorr_conv.yaml`) |
|
||||
|
||||
DSpark and Domino are **DFlash variants**, not separate pipelines: same
|
||||
Everything below EAGLE3 is a **DFlash variant**, not a separate pipeline: same
|
||||
`recipe_type: speculative_dflash`, same training script, same `dflash.*` config
|
||||
namespace, selected by `dflash_architecture_config.projector_type`. Read
|
||||
`references/algorithms/dflash.md` first, then the variant's sheet for the delta.
|
||||
|
||||
LiLiCorr ships two recipes — `lilicorr.yaml` and `lilicorr_conv.yaml`, the latter
|
||||
adding DFlash2's sublayer convolutions — but one sheet covers both.
|
||||
|
||||
If the user's algorithm has no sheet yet, the stage procedures still apply — derive
|
||||
the missing values from an existing launcher example for that algorithm
|
||||
(`tools/launcher/examples/*/*/hf_*_<algorithm>.yaml`) and its recipe, then write the
|
||||
|
||||
@@ -40,4 +40,6 @@ Source the facts from the repo rather than from memory:
|
||||
Then add a row to the algorithm table in `../../SKILL.md`.
|
||||
|
||||
Every algorithm with a recipe in
|
||||
`modelopt_recipes/general/speculative_decoding/` currently has a sheet.
|
||||
`modelopt_recipes/general/speculative_decoding/` currently has a sheet. Where two
|
||||
recipes are the same algorithm under different settings — `lilicorr.yaml` and
|
||||
`lilicorr_conv.yaml` — one sheet covers both, with the delta in its own section.
|
||||
|
||||
@@ -0,0 +1,126 @@
|
||||
# DFlash2
|
||||
|
||||
**A DFlash variant, not a separate pipeline.** DFlash2 is the DFlash draft backbone
|
||||
plus two additions, selected with `dflash_architecture_config.projector_type=dflash2`:
|
||||
|
||||
- a **grouped dynamic depthwise convolution** wrapped around every attention and MLP
|
||||
sublayer, giving each block position a view of its predecessors inside the block.
|
||||
Taps are clipped at block boundaries and the wrapper is identity-initialized, so a
|
||||
fresh DFlash2 draft computes exactly what its DFlash backbone would;
|
||||
- a **low-rank candidate selector** that scores transitions between adjacent block
|
||||
positions' top-k candidates, so serving walks one coherent path instead of taking
|
||||
the per-slot argmax. Trained by an extra cross-entropy term weighted by
|
||||
`dflash_selector_loss_alpha`.
|
||||
|
||||
Read `dflash.md` first — the pipeline, dump flags, export behaviour, and generic
|
||||
failure modes are all shared. This sheet covers only the delta.
|
||||
|
||||
Recipe: `modelopt_recipes/general/speculative_decoding/dflash2.yaml` (its
|
||||
`metadata.recipe_type` is `speculative_dflash`, and every knob lives in the `dflash.*`
|
||||
namespace).
|
||||
|
||||
Examples: `tools/launcher/examples/Qwen/Qwen3-8B/hf_online_dflash2.yaml`,
|
||||
`hf_streaming_dflash2.yaml`.
|
||||
|
||||
Reference: <https://inco.ai/blog/dflash2>. Serving support is vLLM PR #52816.
|
||||
|
||||
## Pipeline tasks
|
||||
|
||||
Two committed shapes, both 2 tasks — identical in layout to DFlash's, only the
|
||||
`--config` recipe differs.
|
||||
|
||||
**Online** (`hf_online_dflash2.yaml`):
|
||||
|
||||
| Task | Script | Purpose | Output |
|
||||
| --- | --- | --- | --- |
|
||||
| task_0 | `common/eagle3/make_dataset.sh` | Build training conversations (Daring-Anteater multi-turn SFT, 50K, `--full-conversations`) | `/scratchspace/data/train.jsonl` |
|
||||
| task_1 | `common/specdec/dflash_online_training.sh` | Online train against a live base model, then export | `/scratchspace/export` |
|
||||
|
||||
`data.mode=online` is the recipe default. As committed the example is a short
|
||||
convergence check (`max_steps=2000`, 1 node x 8 GPUs). The header documents the full
|
||||
published A/B run: 3 epochs over the 1.96M-conversation Spec-Decoding-Dataset-v1,
|
||||
8 nodes x 8 H100, global batch 64, `save_steps=4000`.
|
||||
|
||||
**Streaming** (`hf_streaming_dflash2.yaml`) — same two scripts with
|
||||
`data.mode=streaming`; see `dspark.md` for the streaming environment variables, which
|
||||
are shared across the whole DFlash family.
|
||||
|
||||
No inference path is wired into either example. The selector is not applied in
|
||||
`pseudo_speculative_generate`, so acceptance measured from this repo reports the
|
||||
backbone (plus convolutions) alone.
|
||||
|
||||
## Recipe and training knobs
|
||||
|
||||
Everything in `dflash.md` applies. DFlash2 adds:
|
||||
|
||||
| Override | Recipe default | Note |
|
||||
| --- | --- | --- |
|
||||
| `dflash.dflash_architecture_config.projector_type` | `dflash2` | Selects the variant |
|
||||
| `dflash.dflash_architecture_config.conv_kernel_size` | 2 | Convolution taps; 2 = each position also sees its predecessor. Must be >= 1 and must not exceed `dflash_block_size` |
|
||||
| `dflash.dflash_architecture_config.conv_group_size` | 16 | Must be >= 1 and must divide the draft's `hidden_size` |
|
||||
| `dflash.dflash_architecture_config.selector_rank` | 256 | Rank of the transition codebooks. Must be >= 1 |
|
||||
| `dflash.dflash_architecture_config.selector_top_k` | 16 | How many backbone candidates per position the selector re-ranks. Must be in `[1, vocab_size]` |
|
||||
| `dflash.dflash_selector_loss_alpha` | 1.0 | Selector cross-entropy weight. **0 disables the selector** and trains backbone + convolutions only |
|
||||
| `dflash.dflash_lk_loss_type` | `lambda` | Anneals the block objective from cross-entropy toward acceptance as acceptance rises. Requires `dflash_self_logit_distillation: false` |
|
||||
| `dflash.dflash_lk_ce_scale` / `dflash_lk_ce_decay` | 1.0 / 1.0 | Shape of that anneal |
|
||||
|
||||
`dflash_self_logit_distillation` is **false**, and not merely as a default: the
|
||||
`lk_loss_type` anneal needs the draft's probability of the gold token, which the KD
|
||||
path never forms.
|
||||
|
||||
Recipe defaults that differ from DFlash's: `block_size` 16, `num_anchors` 256,
|
||||
`num_train_epochs` 6, `training_seq_len` 3072, `warmup_ratio` 0.04,
|
||||
`learning_rate` 6.0e-4, `loss_decay_factor` 7.0.
|
||||
|
||||
`training.ddp_find_unused_parameters` is `true` by design — the selector parameters
|
||||
are unused when `dflash_selector_loss_alpha == 0`, which would otherwise trip DDP.
|
||||
|
||||
## Per-model adjustments
|
||||
|
||||
Everything in `dflash.md`'s table applies, plus the DFlash-family items:
|
||||
|
||||
| Situation | What to change |
|
||||
| --- | --- |
|
||||
| Any model | **The draft does not inherit the base model's GQA/FFN dims.** Set `num_attention_heads`, `num_key_value_heads`, `head_dim` and `intermediate_size` in `dflash_architecture_config` explicitly, or you get a silently wrong-shaped draft |
|
||||
| Any non-Qwen3 base | `dflash2.yaml` hardcodes `dflash_mask_token_id: 151669`, a Qwen3-specific unused id. The pin is inherited, so a different base silently trains against a token that means something else. Override it |
|
||||
| Changing `dflash_block_size` | `conv_kernel_size` must not exceed it, and the convolution needs a sequence length divisible by it |
|
||||
| Changing the draft's `hidden_size` | `conv_group_size` must still divide it |
|
||||
|
||||
## Success markers
|
||||
|
||||
Same as `dflash.md`. Because neither example ships a smoke test or AR eval, the
|
||||
in-pipeline evidence is training progress plus the export landing in
|
||||
`/scratchspace/export`.
|
||||
|
||||
## Quality gate
|
||||
|
||||
**Do not trust in-training AR for DFlash2.** The recipe pins `estimate_ar: false`
|
||||
and `ar_validate_steps: 0` deliberately: eval takes a plain per-position argmax, so
|
||||
the candidate selector is not applied. The convolutions *are* — they run inside the
|
||||
backbone layers — so a reported AR measures backbone + convolutions and understates
|
||||
the trained model.
|
||||
|
||||
Evaluate by exporting and benchmarking on the serving stack (vLLM PR #52816), or via
|
||||
the offline acceptance-length harness.
|
||||
|
||||
The training-regression gate from `dflash.md` (`MAX_FINAL_LOSS`, `MIN_FINAL_ACC` via
|
||||
`check_regression.py`) applies; the online example sets `MAX_FINAL_LOSS=5.0` and
|
||||
`MIN_FINAL_ACC=0.15` for its 2000-step convergence check.
|
||||
|
||||
## Known failures
|
||||
|
||||
Generic infrastructure failures are in `../stages/triage.md`; shared block-diffusion
|
||||
failures (`seq_len` divisibility, offline eval, mask token, chat template) are in
|
||||
`dflash.md`. DFlash2-specific:
|
||||
|
||||
| Error pattern | Root cause | Fix |
|
||||
| --- | --- | --- |
|
||||
| `DFlash2 (projector_type='dflash2') requires <keys> in dflash_architecture_config` | Convolution or selector geometry missing | Supply all four: `conv_kernel_size`, `conv_group_size`, `selector_rank`, `selector_top_k` |
|
||||
| `DFlash2 (projector_type='dflash2') requires an integer '<name>' in dflash_architecture_config` | A geometry key is present but not an int | Fix the type — a YAML string will not be coerced |
|
||||
| `DFlash2 conv_kernel_size must be >= 1` / `must not exceed dflash_block_size (N)` | Tap count out of range | Set `1 <= conv_kernel_size <= dflash_block_size` |
|
||||
| `DFlash2 conv_group_size (N) must be >= 1 and divide hidden_size (H)` | Group size does not divide the draft hidden size | Pick a divisor of the draft's `hidden_size` |
|
||||
| `DFlash2 convolution needs a sequence length divisible by block_size (N), got M` | `training_seq_len` not a multiple of `dflash_block_size` | Same divisibility rule as DFlash — round `training_seq_len` to a multiple |
|
||||
| `DFlash2 selector_rank must be >= 1` / `selector_top_k must be in [1, vocab_size=V]` | Selector geometry out of range | Correct the value |
|
||||
| `dflash_lk_loss_type=... needs the draft's probability of the gold token, which the KD path never forms` | `dflash_self_logit_distillation` turned on alongside the `lk` anneal | Keep `dflash_self_logit_distillation: false`, or set `dflash_lk_loss_type: null` |
|
||||
| DDP hangs or complains about unused parameters | `dflash_selector_loss_alpha` set to 0, leaving the selector untrained | Keep `training.ddp_find_unused_parameters: true`, as the recipe does |
|
||||
| Benchmark reports a poor acceptance length | `--draft_length` was passed. The DFlash family reads `--block_size`, and it must match the drafter's | Pass `--block_size <N>` matching the drafter |
|
||||
@@ -0,0 +1,206 @@
|
||||
# LiLiCorr
|
||||
|
||||
**A DFlash variant, not a separate pipeline.** LiLiCorr is the DFlash draft backbone
|
||||
plus a **reranker over the candidate lattice the backbone already produces**, selected
|
||||
with `dflash_architecture_config.projector_type=lilicorr`. The backbone keeps its
|
||||
top-k tokens per slot; a small two-layer transformer scores transitions between
|
||||
adjacent slots' candidates, and serving commits a path greedily, left to right. It is
|
||||
trained jointly with the backbone, so the drafter learns to propose candidates that
|
||||
correlate into longer accepted sequences.
|
||||
|
||||
Read `dflash.md` first — the pipeline, dump flags, export behaviour, and generic
|
||||
failure modes are all shared. This sheet covers only the delta.
|
||||
|
||||
Recipes — **two files, one algorithm**:
|
||||
|
||||
| Recipe | What it is |
|
||||
| --- | --- |
|
||||
| `modelopt_recipes/general/speculative_decoding/lilicorr.yaml` | The published `base` variant. Trains today with no extra dependency |
|
||||
| `modelopt_recipes/general/speculative_decoding/lilicorr_conv.yaml` | The same, plus DFlash2's grouped sublayer convolutions. See *The conv recipe* below |
|
||||
|
||||
Recipes do not compose, so the second is a standalone file rather than an overlay —
|
||||
**keep shared fields in the two in sync when editing either.**
|
||||
|
||||
Example: `tools/launcher/examples/Qwen/Qwen3-8B/hf_online_lilicorr.yaml`.
|
||||
|
||||
Paper: <https://arxiv.org/abs/2608.20530>. Serving support is
|
||||
[sgl-project/sglang#37462](https://github.com/sgl-project/sglang/pull/37462) — the
|
||||
reranker's serving path is in SGLang, not vLLM.
|
||||
|
||||
## Pipeline tasks
|
||||
|
||||
One committed shape, 2 tasks — identical in layout to DFlash's online example, only
|
||||
the `--config` recipe differs.
|
||||
|
||||
| Task | Script | Purpose | Output |
|
||||
| --- | --- | --- | --- |
|
||||
| task_0 | `common/eagle3/make_dataset.sh` | Build training conversations (Daring-Anteater multi-turn SFT, 50K, `--full-conversations`) | `/scratchspace/data/train.jsonl` |
|
||||
| task_1 | `common/specdec/dflash_online_training.sh` | Online train against a live base model, then export | `/scratchspace/lilicorr_bs16` |
|
||||
|
||||
**Online only, and not by convention.** `data.mode: online` is a hard requirement:
|
||||
the distractor penalty weights every competing candidate by the target model's own
|
||||
logit gap, so a target model has to be in the process. An offline or streaming run
|
||||
fails at loss construction rather than training a weaker model.
|
||||
|
||||
As committed the example is a short convergence check (`max_steps=2000`,
|
||||
1 node x 8 GPUs). The published Qwen3-8B numbers come from 6 epochs at
|
||||
8 nodes x 8 H100, global batch 64.
|
||||
|
||||
No inference path is wired in. The reranker is not applied in
|
||||
`pseudo_speculative_generate`, so acceptance measured from this repo reports the
|
||||
backbone alone.
|
||||
|
||||
## Recipe and training knobs
|
||||
|
||||
Everything in `dflash.md` applies. LiLiCorr adds a three-term head objective on top
|
||||
of the DFlash block loss:
|
||||
|
||||
```text
|
||||
loss = dflash_loss + w_ce*CE + w_margin*hinge + w_pen*penalty
|
||||
```
|
||||
|
||||
No outer multiplier, so `loss == origin_loss + lilicorr_loss` holds exactly.
|
||||
|
||||
| Override | Recipe default | Note |
|
||||
| --- | --- | --- |
|
||||
| `dflash.dflash_architecture_config.projector_type` | `lilicorr` | Selects the variant |
|
||||
| `dflash.dflash_lilicorr_w_ce` | 0.25 | Cross-entropy term |
|
||||
| `dflash.dflash_lilicorr_w_margin` | 0.0 | Hinge term. **0 in the `base` variant** |
|
||||
| `dflash.dflash_lilicorr_w_pen` | 0.25 | Distractor penalty. Needs the target model's logits — hence online-only |
|
||||
| `dflash.dflash_lilicorr_margin` | 2.0 | Hinge width in log-potential units. Unused while `w_margin` is 0; must be > 0 when it isn't |
|
||||
| `dflash.dflash_fp32_master_weights` | true | The published numbers were trained with this on. Turning it off changes the optimizer's arithmetic, not just its memory |
|
||||
|
||||
**Two published variants**, differing only in how the cross-entropy block is split:
|
||||
|
||||
| Variant | `w_ce` | `w_margin` | `w_pen` |
|
||||
| --- | --- | --- | --- |
|
||||
| `base` (the shipped recipe) | 0.25 | 0.0 | 0.25 |
|
||||
| `margin` | 0.125 | 0.125 | 0.25 |
|
||||
|
||||
The head's total weight is 0.50 either way. The weights are **absolute and validated
|
||||
all-or-nothing** — at least one of the three must be above 0.
|
||||
|
||||
Reranker geometry, under `dflash_architecture_config`. **Every field is required and
|
||||
never defaulted**, which matters more here than elsewhere: `lilicorr_candidate_topk`
|
||||
sets the lattice width and the shape of `rank_embedding`, while `lilicorr_logit_scale`
|
||||
and `lilicorr_vector_eps` change the score **without changing any tensor shape**. A
|
||||
guessed value for those two builds a head that loads cleanly and scores a different
|
||||
function.
|
||||
|
||||
| Field | Default |
|
||||
| --- | --- |
|
||||
| `lilicorr_candidate_topk` | 8 |
|
||||
| `lilicorr_hidden_size` | 1024 |
|
||||
| `lilicorr_factor_dim` | 1024 |
|
||||
| `lilicorr_num_layers` | 2 |
|
||||
| `lilicorr_num_heads` | 8 |
|
||||
| `lilicorr_mlp_ratio` | 2.0 |
|
||||
| `lilicorr_logit_scale` | 8.0 |
|
||||
| `lilicorr_vector_eps` | 1.0e-4 |
|
||||
|
||||
`dflash_init_checkpoint` restores **weights only** and reads geometry from the config,
|
||||
so a warm start reproduces a head only if every field above matches the one the
|
||||
checkpoint was trained with.
|
||||
|
||||
Other recipe defaults that differ from DFlash's: `block_size` 16, `num_anchors` 512,
|
||||
`loss_objective` `decay` with `decay_factor` 7.0 (not `dpace` — the published variants
|
||||
were trained on the static decay), `lr_scheduler_type` **cosine** (not linear — the
|
||||
published variants used a linear warmup into cosine decay, and the schedule is part of
|
||||
the recipe those numbers came from), `num_train_epochs` 6, `training_seq_len` 3072.
|
||||
|
||||
`dflash_self_logit_distillation` is **false**: the reranker's terms are added to the
|
||||
plain weighted cross-entropy, and turning KD on would replace that base term and
|
||||
change the objective the published checkpoints were trained under.
|
||||
|
||||
## The conv recipe
|
||||
|
||||
`lilicorr_conv.yaml` is `lilicorr.yaml` with DFlash2's grouped dynamic depthwise
|
||||
convolution wrapped around every draft sublayer. The convolution is **shared code** —
|
||||
`DFlashGroupedConv` imported from the DFlash2 plugin, installed on the no-op sublayer
|
||||
seam `DFlashDecoderLayer` already exposes — so the two variants cannot drift apart
|
||||
arithmetically.
|
||||
|
||||
It adds three keys to `dflash_architecture_config`:
|
||||
|
||||
| Field | Default | Note |
|
||||
| --- | --- | --- |
|
||||
| `conv_kernel_size` | 2 | Tap count; must not exceed `dflash_block_size` |
|
||||
| `conv_group_size` | 16 | Must divide the draft's `hidden_size` |
|
||||
| `conv_projection_init_std` | 0.0 | **Zero means exact identity at init** — see below |
|
||||
|
||||
The two conv geometry keys are **all-or-nothing**: one alone is rejected rather than
|
||||
silently building a draft without convolutions.
|
||||
|
||||
`conv_projection_init_std: 0.0` zeroes `kernel_projection`, and `base_kernel` is
|
||||
identity at tap 0, so the wrapper is an *exact* identity at step 0. A run from this
|
||||
recipe begins as the plain reranker, and any difference is attributable to the
|
||||
convolutions rather than to a perturbed start. Raise the key only if you want a
|
||||
perturbed start deliberately. It is a separate key from `initializer_range`, which
|
||||
also seeds the reranker.
|
||||
|
||||
**Memory is the binding constraint.** The convolutions add ~42M trainable parameters
|
||||
(20 tensors for a 5-layer draft), and it is the activations they hold that bind. At an
|
||||
8B target, combined with `dflash_fp32_master_weights`, this is the memory worst case
|
||||
and may need `training.gradient_checkpointing: true` to fit on 80 GiB; at a 4B target
|
||||
it fits without. Checkpointing is mathematically neutral — same objective, same data
|
||||
order, same resulting model — but it trades step time for memory, so a run using it is
|
||||
not step-time-comparable with one that does not.
|
||||
|
||||
**Historical note:** this recipe used to be unusable on `main` because the DFlash2
|
||||
plugin lived on a branch. `modelopt/torch/speculative/plugins/modeling_dflash2.py` is
|
||||
now on `main`, so it trains today. If you hit the `ImportError` below, the install is
|
||||
older than that merge.
|
||||
|
||||
## Per-model adjustments
|
||||
|
||||
Everything in `dflash.md`'s table applies, plus the DFlash-family items:
|
||||
|
||||
| Situation | What to change |
|
||||
| --- | --- |
|
||||
| Any model | **The draft does not inherit the base model's GQA/FFN dims.** Set `num_attention_heads`, `num_key_value_heads`, `head_dim` and `intermediate_size` in `dflash_architecture_config` explicitly |
|
||||
| Any non-Qwen3 base | Both recipes hardcode `dflash_mask_token_id: 151669`, a Qwen3-specific unused id. Override it |
|
||||
| Reproducing published numbers | Use `tools/launcher/examples/Qwen/Qwen3-8B/chat_template_train.jinja`, as the other speculative recipes do. It differs from the reference mask by 6 tokens of supervision per record (the empty `<think>` preamble and the end-of-turn token); token ids are identical either way |
|
||||
| Warm start from a published head | Every reranker geometry field must match the checkpoint's, including `lilicorr_logit_scale` and `lilicorr_vector_eps`, which affect no shape and so fail silently |
|
||||
|
||||
## Success markers
|
||||
|
||||
Same as `dflash.md`. The example ships no smoke test or AR eval, so the in-pipeline
|
||||
evidence is training progress plus the export landing in `training.output_dir`.
|
||||
|
||||
## Quality gate
|
||||
|
||||
**Do not trust in-training AR for LiLiCorr.** The recipes pin `estimate_ar: false`
|
||||
and `ar_validate_steps: 0` deliberately: eval runs the DFlash backbone only, with the
|
||||
reranker not applied in `pseudo_speculative_generate`, so a reported AR describes the
|
||||
backbone alone and understates the trained model.
|
||||
|
||||
Evaluate by exporting the drafter and benchmarking it on SGLang
|
||||
(sgl-project/sglang#37462). Published Qwen3-8B acceptance lengths for context, all
|
||||
trained on one matched contract and served through SGLang on a single H100 at
|
||||
concurrency 1: LiLiCorr+conv 7.715 on gsm8k and 4.014 on mtbench, plain LiLiCorr
|
||||
7.557 / 3.939, against DFlash's head-free control at 6.341 / 3.478. Treat these as
|
||||
reference points for a reproduction, not as a pass threshold — see
|
||||
`../stages/validate.md` on why no threshold is enforced anywhere in this repo.
|
||||
|
||||
The training-regression gate from `dflash.md` (`MAX_FINAL_LOSS`, `MIN_FINAL_ACC` via
|
||||
`check_regression.py`) applies; the online example sets `MAX_FINAL_LOSS=5.0` and
|
||||
`MIN_FINAL_ACC=0.15` for its 2000-step convergence check.
|
||||
|
||||
## Known failures
|
||||
|
||||
Generic infrastructure failures are in `../stages/triage.md`; shared block-diffusion
|
||||
failures (`seq_len` divisibility, offline eval, mask token, chat template) are in
|
||||
`dflash.md`. LiLiCorr-specific:
|
||||
|
||||
| Error pattern | Root cause | Fix |
|
||||
| --- | --- | --- |
|
||||
| `All three LiLiCorr objective weights are 0, so the reranker would never ...` | `w_ce`, `w_margin` and `w_pen` all 0 | Set at least one above 0 |
|
||||
| `dflash_lilicorr_w_pen > 0 requires the target model's logits, which are ...` | Offline or streaming mode with the penalty enabled | Use `data.mode: online`, or set `w_pen: 0` |
|
||||
| `dflash_lilicorr_w_margin > 0 requires a positive dflash_lilicorr_margin` | Hinge enabled with a zero/negative width | Set `dflash_lilicorr_margin` above 0 |
|
||||
| `LiLiCorr block_size must be >= 2, got N` | Block too small for a lattice | Raise `dflash_block_size` |
|
||||
| `LiLiCorr candidate_topk must be >= 1` / `lilicorr_candidate_topk=K exceeds the vocabulary size V` | Lattice width out of range | Set `1 <= lilicorr_candidate_topk <= vocab_size` |
|
||||
| `... DFlashGroupedConv, which is not available in this installation. Remove conv_kernel_size and conv_group_size from dflash_architecture_config to ...` | `lilicorr_conv.yaml` on an install predating the DFlash2 merge | Update the install, or drop the two conv keys to train plain LiLiCorr |
|
||||
| `... requires conv_kernel_size and conv_group_size ... The grouped convolution needs the tap count ...` | Only one of the two conv keys supplied | They are all-or-nothing — set both or neither |
|
||||
| Warm-started head scores differently despite loading cleanly | `lilicorr_logit_scale` or `lilicorr_vector_eps` differ from the trained head; neither affects any tensor shape, so nothing errors | Transcribe every reranker field from the checkpoint's config |
|
||||
| Conv run OOMs at an 8B target | ~42M extra params plus their activations, on top of fp32 master weights | `training.gradient_checkpointing: true` — neutral to the result, costs step time |
|
||||
| Benchmark reports a poor acceptance length | `--draft_length` was passed, or the drafter was benchmarked on vLLM | The DFlash family reads `--block_size`; the reranker's serving path is SGLang |
|
||||
Reference in New Issue
Block a user