mirror of
https://github.com/radixark/miles.git
synced 2026-10-02 07:14:53 +08:00
The scripts were calling module-level functions that could only ever launch onto ray. Each one now asks its own config for a backend and calls that, which is what lets the same script reach a cluster later. The remaining command helpers move onto the backend with them, and the scripts settle on one shape for their subcommands: prepare, train and full_train.
Random fully-async example
Minimal sibling of examples/infra_features/fully_async/. Exercises the entire async
rollout ↔ trainer loop without any real dataset, real reward model, or
meaningful generation — useful as an agent infrastructure stress test
for bigger agentic workloads.
Quick start
# default (Qwen3.5-35B-A3B), in_place pause + broadcast weight transfer
python run_random_async_3node.py
# retract pause + p2p weight transfer
python run_random_async_3node.py \
--pause-generation-mode retract \
--update-weight-transfer-mode p2p
# swap in a different model
python run_random_async_3node.py \
--model-name Qwen3.5-35B-A3B --megatron-model-type qwen3.5-35B-A3B
Notes
--disable-rollout-global-datasetis on, so no--prompt-datafile is required. The rollout function ignores the data buffer and constructsSampleobjects from scratch.- The rollout uses Qwen3.5-35B's vocab size (151643) for the random
input_ids; any model with vocab ≥ that works without changes. ignore_eos=Truein the sampling params means SGLang generates until it hitsmax_new_tokens(drawn fromMAX_TOKENS_RANGE).Sample.statusis set toCOMPLETED.