mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Fixes #1692 Adds a SWE-Perf benchmarking workload to the benchmarking suite, exercising actor suspend/resume against a real SWE-bench task image (`astropy-7336`) rather than a synthetic load. The workload runs a recorded agent trajectory in configurable cycles, suspending and resuming the actor between each, so the benchmark measures lifecycle cost under a realistic in-sandbox server. ### What's here | Area | Change | |---|---| | `internal/benchmarking/boomer/sweperf/` | New `SweperfUser` boomer user class (+ tests) | | `internal/benchmarking/boomer/dynconfig/` | `sweperf_template`, `sweperf_total_steps`, `sweperf_num_cycles` knobs | | `benchmarking/locust/common/sweperf_config.py` | `--sweperf-*` Locust flags | | `benchmarking/locust/tests/sweperf.py` | Stub user class so the master attributes boomer's stats rows | | `benchmarking/workloads/manifests/` | `swebench-astropy-7336` ActorTemplate | Defaults: 21 trace steps partitioned into 4 cycles. ### Verification - `hack/verify-all.sh` — pass - `go test -race ./internal/benchmarking/... ./cmd/benchmarking/...` — pass - End-to-end on a GKE dev cluster (gVisor), 1 user: ``` state: running | users: 1 | fail_ratio: 0.0 NAME REQ FAIL AVG_ms CreateActor 1 0 1 CreateAtespace 1 0 1 ResumeActor 20 0 495 SuspendActor 20 0 947 Workload_Cycle_1 5 0 1996 Workload_Cycle_2 5 0 964 Workload_Cycle_3 5 0 4877 Workload_Cycle_4 5 0 3022 ``` 164 requests total, 0 failures, 0 errors. The ActorTemplate image is currently pinned to a personal Artifact Registry repo; a shared public registry for these SWE-Perf task images is planned. - [x] Tests pass - [ ] Appropriate changes to documentation are included in the PR