Files
Model-Optimizer/tools
c248dd5434 [Feat]: Domino support (#1710)
### What does this PR do?

Type of change: New feature

Adds **Domino** speculative decoding: the parallel DFlash draft backbone
plus a lightweight **GRU causal correction head**. The backbone produces
*base* logits for a full draft block in one forward; a GRU over the
block's teacher-forced tokens produces a causal state that is fused with
the backbone hidden state and projected to a vocab-sized logit
correction on the block suffix — injecting the intra-block causal
dependency the parallel backbone lacks. Trained with a dual loss
`(1-λ)*final + λ*base`, where `λ_base` decays linearly 1→0 (curriculum:
learn the parallel backbone first, then the correction).

Reuses the DFlash mode/config/recipe; selected via
`dflash_architecture_config.projector_type=domino` and routed to its own
registry so `HFDominoModel` does not shadow `HFDFlashModel`. Exports in
the z-lab/SpecForge drafter format (`prefix_gru.*` / `embed_proj.*`).

> Note: the inference side (vLLM / AR evaluation) is intentionally
**not** wired up yet — the correction head is not applied in serving. To
be added once the inference path lands.

### Usage

```bash
# Online training (recipe: projector_type=domino)
uv run launch.py --yaml examples/Qwen/Qwen3-8B/hf_online_domino.yaml --yes
```

### Testing

CPU unit tests in
`tests/unit/torch/speculative/plugins/test_hf_domino.py` cover
conversion routing, the training forward (dual loss + grads), the λ
schedule, and the export format. Online Qwen3-8B training validated
end-to-end (loss curve below).

<img width="1803" height="809" alt="image"
src="https://github.com/user-attachments/assets/7c9d2001-bd80-4dec-919b-443e61089cca"
/>

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ (opt-in via
`projector_type=domino`; DFlash path unchanged)
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A (no new
dependency)
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

### Additional Information

Reference: SpecForge PR #571 (z-lab); drafter format
`huggingface.co/Huang2020/Qwen3-8B-Domino-b16`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

## Release Notes

* **New Features**
* Added Domino speculative-decoding training with a decaying base/final
dual-loss curriculum and Domino-specific lambda scheduling
(training-only; inference wiring not yet included).
  * Added Domino draft-head export support for training checkpoints.

* **Documentation & Configuration**
* Added a Domino speculative-decoding training recipe and an HF Online
Domino launcher configuration for Qwen3-8B.

* **Refactor**
* Updated speculative model conversion/export to route to Domino
variants based on the configured projector type.

* **Tests**
* Added unit tests for Domino conversion, training loss/metrics, lambda
decay behavior, and exporter output layout.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: h-guo18 <67671475+h-guo18@users.noreply.github.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
2026-06-27 07:34:50 +00:00
..
2026-06-27 07:34:50 +00:00
2026-06-27 01:00:25 +05:30