Files
Aditya ShantanuandAditya Shantanu bc3cbc519d benchmarking: agent-session workload — a scripted coding agent that suspends while the LLM thinks (#1934)
## What this is

We keep saying Substrate's sweet spot is "agents that are idle most of
the time" — but none of our benchmarks actually behave like one. Glutton
hammers one resource at a time, and sweperf needs an external image with
a replay trace. This adds a workload that acts like the thing we're
building for: a coding agent working through a task.

Each locust user is one session. The actor gets told to do the kind of
things a coding agent does — clone a repo, install deps, build, hit a
failing test, fix it, write tests, refactor, package — twenty steps,
each one costing the sandbox the CPU, memory, disk, and network the real
action would. Between steps the "LLM is thinking," so the driver
suspends the actor, and the next step's first request wakes it back up
through the router. That parked wake (`WakeFirstTouch` in the stats) is
the number this whole benchmark exists to measure.

The part I care most about: the entire workload is one table in
`internal/benchmarking/boomer/agentsession/script.go`. Every step says
in plain English what the agent is doing and what it costs. If you want
to know what step 6 does to the sandbox, you read step 6. If you want a
different workload, you edit the table — tests will catch you if you
write a step that reads a file nothing wrote, or blow the actor's memory
budget.

To act the steps out, glutton grew two RPCs: `BurnCPU` (compute-bound
work) and `Ingest` (bytes that actually cross the network before hitting
disk, so a "git clone" is a real download, not a local write).

## How it went when we ran it

Validated on a fresh 2-node GKE cluster, micro-VM first, then gVisor on
the same hardware. Smoke runs were clean on both classes (gVisor: 1340
requests over 7 full laps, zero failures, wake p50 1.4s; micro-VM: wake
p50 2.1s). At 20 concurrent sessions with realistic think times, gVisor
held 2.6% failures with wake p50 1.3s. Pushing past the knee (~7
concurrently-active sessions on 8 vCPUs) was also useful: it reproduced
the ateom-socket-vanishing failure from #1133 and left six actors
permanently wedged in DELETING — a live repro of #1665.

Two things the first live run taught us are already folded in: RAM
refills are in-place so repeat laps don't transiently double the guest
heap (512Mi micro-VM actors OOM'd without this — use 1Gi), and the
driver replaces an actor after three failed steps in a row, because a
CRASHED actor never comes back on its own.

## Future changes

Right now the script is compiled in — changing what steps do means
editing the table and rebuilding the image. That's deliberate for this
PR (one reviewable, test-guarded source of truth), but the follow-up
we've agreed on is to make the script runtime-configurable: named script
variants selectable per run first, then accepting a full script as a
file so operators can define workloads without touching Go. That lands
as its own PR once this one is in.

---------

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-09-30 23:54:36 +00:00
..