Files
miles/examples/infra_features/random_async
fzyzcjy 8019f79062 Have every launch script talk to a backend object
The scripts were calling module-level functions that could only ever launch
onto ray. Each one now asks its own config for a backend and calls that, which
is what lets the same script reach a cluster later. The remaining command
helpers move onto the backend with them, and the scripts settle on one shape
for their subcommands: prepare, train and full_train.
2026-09-02 16:04:52 +08:00
..

Random fully-async example

Minimal sibling of examples/infra_features/fully_async/. Exercises the entire async rollout ↔ trainer loop without any real dataset, real reward model, or meaningful generation — useful as an agent infrastructure stress test for bigger agentic workloads.

Quick start

# default (Qwen3.5-35B-A3B), in_place pause + broadcast weight transfer
python run_random_async_3node.py

# retract pause + p2p weight transfer
python run_random_async_3node.py \
    --pause-generation-mode retract \
    --update-weight-transfer-mode p2p

# swap in a different model
python run_random_async_3node.py \
    --model-name Qwen3.5-35B-A3B --megatron-model-type qwen3.5-35B-A3B

Notes

  • --disable-rollout-global-dataset is on, so no --prompt-data file is required. The rollout function ignores the data buffer and constructs Sample objects from scratch.
  • The rollout uses Qwen3.5-35B's vocab size (151643) for the random input_ids; any model with vocab ≥ that works without changes.
  • ignore_eos=True in the sampling params means SGLang generates until it hits max_new_tokens (drawn from MAX_TOKENS_RANGE). Sample.status is set to COMPLETED.