mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
## What this is We keep saying Substrate's sweet spot is "agents that are idle most of the time" — but none of our benchmarks actually behave like one. Glutton hammers one resource at a time, and sweperf needs an external image with a replay trace. This adds a workload that acts like the thing we're building for: a coding agent working through a task. Each locust user is one session. The actor gets told to do the kind of things a coding agent does — clone a repo, install deps, build, hit a failing test, fix it, write tests, refactor, package — twenty steps, each one costing the sandbox the CPU, memory, disk, and network the real action would. Between steps the "LLM is thinking," so the driver suspends the actor, and the next step's first request wakes it back up through the router. That parked wake (`WakeFirstTouch` in the stats) is the number this whole benchmark exists to measure. The part I care most about: the entire workload is one table in `internal/benchmarking/boomer/agentsession/script.go`. Every step says in plain English what the agent is doing and what it costs. If you want to know what step 6 does to the sandbox, you read step 6. If you want a different workload, you edit the table — tests will catch you if you write a step that reads a file nothing wrote, or blow the actor's memory budget. To act the steps out, glutton grew two RPCs: `BurnCPU` (compute-bound work) and `Ingest` (bytes that actually cross the network before hitting disk, so a "git clone" is a real download, not a local write). ## How it went when we ran it Validated on a fresh 2-node GKE cluster, micro-VM first, then gVisor on the same hardware. Smoke runs were clean on both classes (gVisor: 1340 requests over 7 full laps, zero failures, wake p50 1.4s; micro-VM: wake p50 2.1s). At 20 concurrent sessions with realistic think times, gVisor held 2.6% failures with wake p50 1.3s. Pushing past the knee (~7 concurrently-active sessions on 8 vCPUs) was also useful: it reproduced the ateom-socket-vanishing failure from #1133 and left six actors permanently wedged in DELETING — a live repro of #1665. Two things the first live run taught us are already folded in: RAM refills are in-place so repeat laps don't transiently double the guest heap (512Mi micro-VM actors OOM'd without this — use 1Gi), and the driver replaces an actor after three failed steps in a row, because a CRASHED actor never comes back on its own. ## Future changes Right now the script is compiled in — changing what steps do means editing the table and rebuilding the image. That's deliberate for this PR (one reviewable, test-guarded source of truth), but the follow-up we've agreed on is to make the script runtime-configurable: named script variants selectable per run first, then accepting a full script as a file so operators can define workloads without touching Go. That lands as its own PR once this one is in. --------- Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>