mirror of
https://github.com/radixark/miles.git
synced 2026-10-02 07:14:53 +08:00
`--cluster-backend ray` accepts both `--worker-comm-backend ray` and `rpc`, but every e2e test takes the default, which is `ray`. The rpc wiring under a ray cluster -- `_ServeActorManager`, the per-worker http server, and the `RayObjectStore(frees_objects=True)` branch it selects -- therefore has no e2e coverage at all: it is only ever exercised by hand before a multi-deployment run, so a regression in it stays invisible until someone deploys. Generalize the gsm8k short scenario over the comm backend and add a second registration that runs it with `--worker-comm-backend rpc`. The new file is a shell: its module body is one `register_cuda_ci` call, and it loads the base test by path inside `__main__`, because the base module name contains dots and cannot be imported by name. The scenario takes the caller's `__file__` so the two variants land in separate wandb projects, and emits nothing when the backend is `ray`, so the base variant's command line is unchanged and a red base stays attributable to the base. The rpc variant registers CUDA only. The mi350 lane is expensive and the comm backend is not hardware specific, so one lane is enough to catch a regression.