Files
Model-Optimizer/tools
yeyu-nvidiaandClaude Opus 4.6 8025a3dc54 specdec(recipe): add MiniMax-M2.7-DFlash streaming multi-node pipeline (#1835)
## Summary

- Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE)
streaming DFlash training
- 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each)
over NIXL RDMA hidden-state transport
- Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)`
+ final layer output
- MiniMax-specific: trust_remote_code, FSDP2 via accelerate config,
mask_token=200054, YaRN rope_scaling factor=48
- Topology matches Kimi-K2.5 large-MoE streaming recipe

Resolves OMNIML-5221

## Test plan

- [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`)
- [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`)
- [ ] Full streaming training run

Signed-off-by: Ye Yu <yeyu@nvidia.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added a multi-node launcher configuration for MiniMax-M2.7 streaming
training with speculative decoding.
* Added an end-to-end workflow for conversational data preparation,
distributed training, checkpoint export, and vLLM smoke testing.
* Supports configurable serving and training settings across multiple
nodes.

* **Enhancements**
* Applies Transformers version overrides only to trainer or single-node
environments.
* Preserves the serving environment’s Transformers version on dedicated
serving nodes.
* Uses the launcher’s built-in bfloat16 mixed-precision configuration
for training.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-09-16 12:54:57 -07:00
..