mirror of
https://github.com/NVIDIA/Model-Optimizer.git
synced 2026-10-02 03:14:52 +08:00
## Summary - Add `hf_streaming_dflash_multi_node.yaml` for MiniMax-M2.7 (229B MoE) streaming DFlash training - 2 serve replicas (TP=4, whole node) + 2 trainer nodes (4 GPU each) over NIXL RDMA hidden-state transport - Capture IDs `[2,17,32,47,62,64]` from `build_target_layer_ids(64, 5)` + final layer output - MiniMax-specific: trust_remote_code, FSDP2 via accelerate config, mask_token=200054, YaRN rope_scaling factor=48 - Topology matches Kimi-K2.5 large-MoE streaming recipe Resolves OMNIML-5221 ## Test plan - [ ] Dry-run validation (`uv run launch.py --yaml ... --dry-run`) - [ ] Server-only smoke on CW-DFW (task_1 with `training.max_steps=1`) - [ ] Full streaming training run Signed-off-by: Ye Yu <yeyu@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added a multi-node launcher configuration for MiniMax-M2.7 streaming training with speculative decoding. * Added an end-to-end workflow for conversational data preparation, distributed training, checkpoint export, and vLLM smoke testing. * Supports configurable serving and training settings across multiple nodes. * **Enhancements** * Applies Transformers version overrides only to trainer or single-node environments. * Preserves the serving environment’s Transformers version on dedicated serving nodes. * Uses the launcher’s built-in bfloat16 mixed-precision configuration for training. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>