Nishanth Kotla 3f789b309a benchmarking/locust: harvest server-side telemetry from Prometheus (#1725)
Part of #1590

## What this PR does

Locust measures the client side only. This adds a post-run harvest of
server-side ground truth from Prometheus, written to a new
`server_summary.json` and summarized as one row in `stats.jsonl`.

## Proposed Changes

### Steady-state scoping

`server_telemetry.py` derives the steady-state window from
`stats_history.csv`: it starts at the first sample at 90% of peak user
count and ends at the last one. Ramp-up is excluded from every metric,
and the snapshot block also stops at the last full-load sample, so
teardown suspends are left out. Under a step-ladder load shape the
window covers only the top step.

### Harvested metrics

**Cluster packing**, from `ate_workerpool_workers`: busy workers
(`partial` + `at_capacity`) over the pool size at each sample, where the
pool size is the sum of all worker states, so it follows scale-ups and
scale-downs mid-run. Only the live ateapi is read: after a redeploy the
collector keeps re-exporting exited ateapi processes' last values for a
few minutes, which would otherwise inflate the counts. Reported as min,
p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries
so transient spikes remain visible.

**Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall
percentages for the node and for the worker pods, each as the same
percentile set over the window.

**Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both
in-window and cumulative, restore and checkpoint latency mean and
p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint
throughput as `checkpoint_mb_s`. The atelet exports metrics on an
interval, so its data reaches Prometheus late: the harvest waits
`--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default
export and a 10s scrape) and reads both window edges half that late.

### Constraints and failure behavior

The module uses only the standard library, since the locust image is
distroless, and every request carries a timeout. An unreachable
Prometheus records nulls and does not fail the run. `--prometheus-url`
overrides the in-cluster default, and `--atelet-lag-s` sets the wait for
the atelet's last export.

Unmeasured fields are `null` and a measured zero is `0`, consistent with
the rest of the runner. Prometheus exposes no byte counter on the
restore path, so no restore throughput field is emitted rather than
deriving one indirectly.

### Output

`server_summary.json` holds the full nested artifact. `stats.jsonl`
receives a single `server_summary` row with 45 flat keys for graphing:
packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot
means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`.

`status.json` is unchanged.

## How this was tested

14 unit tests in `test_server_telemetry.py` covering steady-state
detection, percentile boundaries, the range-query window guard,
malformed Prometheus responses, the packing and checkpoint arithmetic,
the per-sample pool size, ignoring exited ateapi series, a missing
denominator returning null, snapshot fields returning null rather than
zero, and telemetry surviving a missing stats CSV.

Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45
keys matched an independent recomputation from raw Prometheus, and the
snapshot block was cross-checked against atelet logs. A run with
`--prometheus-url` pointed at an unreachable address completes normally
with the affected fields null.

Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8
workers) from Prometheus with the new code. Runs with a steady pool
match the previous output exactly, and runs that followed a redeploy now
read the real pool size on every sample.

## References

[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-10-01 16:59:41 +00:00
2026-09-30 16:57:00 +00:00
2026-09-02 19:29:02 -07:00
2026-09-28 22:26:40 +00:00
2026-09-28 22:26:40 +00:00
2026-09-28 22:26:40 +00:00
2026-09-30 16:57:00 +00:00

Agent Substrate

License

What is Agent Substrate?

Agent Substrate is a secure-by-default agent execution runtime engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with native zero-trust kernel and network isolation. It supports multiple sandbox technologies including microVMs and gVisor, enabling consistent lifecycle operations for all sandbox types.

At its core, Agent Substrate maps a larger set of “actors” (applications such as agents) onto a smaller set of ready “workers”, relying on the fact that agent-like applications tend to be idle most of the time to achieve heavy multiplexing. It provides functionality to manage an actor’s lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real time, and to route incoming traffic to them.

Agent Substrate is intended to be a low-opinion system. The workloads it manages don't have to be literal AI agents, but those are the best example of the kind of applications it is designed for. It is not an SDK for building agents, but rather a system for running them at scale.

Agent Substrate leverages Kubernetes for the infrastructure provisioning and worker lifecycle management (Kubernetes Pods). It builds on top of Kubernetes features like Pods and Pod autoscaling, while Agent Substrate provides agent-specific scheduling and control to achieve lower latency. Using Kubernetes as the underlying system enables consistent infrastructure management across all workloads types that are required for end to end agentic deployments and allows holistic infrastructure optimizations for RL scenarios that span agentic, inference and training cycles.

Demo

Agent Substrate Demo

Watch the Agent Substrate cluster multiplex ~250 stateful actors across just 8 physical pods.

This demo highlights the core developer experience and "Agentic Infrastructure" capabilities of Substrate:

  1. Actor Teleport: High-performance suspend and resume of actors onto any available worker in the pool with sub-second activation.
  2. State Persistence: Persistent working memory (volatile RAM) and filesystem state preserved perfectly across hibernation cycles via full-state snapshots.
  3. Agent Multiplexing: Demonstrates 30x+ oversubscription by "juggling" a large registry of stateful actors onto a small pool of shared physical pods.

To reproduce this demo in your own cluster, please refer to the detailed walkthrough in the Counter Demo.

For more videos and walkthroughs, visit our YouTube channel: agent-substrate.

Framework Agnostic & Compatibility

Agent Substrate is designed to be framework and agent harness agnostic. Because it manages standard OCI containers at the kernel level (via gVisor), it can host agents built on any stack.

  • Agent Development Kit (ADK): Support for ADK agents with session state preservation across invocations as actor state. Ideal for all types of agents and stateful tool or subagent calls.
  • LangChain: Ideal execution environment for LangChain agents and tool calls.
  • Claude Code, CodeX, and Antigravity: Support for high-density, stateful coding environments that preserve system state and filesystem state across sessions.
  • Model Context Protocol (MCP): Support for deploying secure, sandboxed MCP servers as Substrate Actors to provide durable tools for any model.

Ecosystem & Examples

  • Agent Executor: A distributed agent runtime that demonstrates building a secure, hyper-scalable agent harness on Agent Substrate (see the announcement blog and integration guide).
  • kagent: A CNCF Sandbox project and Kubernetes-native framework for building, deploying, and managing AI agents that uses Agent Substrate to run sandboxed, stateful agent workloads (see the announcement blog).

Status and compatibility

Agent Substrate is pre-1.0. We are not making any guarantees about backward compatibility at this stage, and APIs and behavior may still change significantly.

Supported Kubernetes Releases

Currently we aim to support the latest stable release of Kubernetes, and the previous minor release.

Community

For announcements, technical discussions, and community support, please join the ate-dev Google Group.

We host a weekly community meeting every Thursday from 10:00am - 11:00am PST.

We also have channels in the CNCF slack; request an invite here if you don't have access.

Developing

Please see CONTRIBUTING.md for guidelines on contributing to the project. We welcome contributions of all kinds, but the project is VERY young. Our immediate focus is on building out the core system and demos, so we may not be able to review or merge contributions that don't align with those goals in the near term.

Quickstart (Development)

To quickly set up the complete environment:

  1. Make sure you have Go, kubectl, and docker installed and configured on your dev machine. We will automatically manage other dependencies via Go, including kind.

  2. Run the following steps:

# create cluster and local registry (IPv4; IP_FAMILY=dual|ipv6 overrides)
hack/create-kind-cluster.sh

# install ate, PostgreSQL, rustfs
hack/install-ate-kind.sh --deploy-ate-system

# install counter demo
hack/install-ate-kind.sh --deploy-demo-counter

# install kubectl-ate
go install ./cmd/kubectl-ate

# create a counter actor in the demo's atespace (--template names the
# actor template, resolved in the actor's atespace)
kubectl ate create actor my-counter-1 -a ate-demo-counter --template counter

# port-forward the network router to bind to local port `8000`
kubectl port-forward -n ate-system svc/atenet-router 8000:80
  1. In a separate terminal, send an HTTP request to increment the counter:
curl -X POST \
   -H "ate-target-actor: ate-demo-counter/my-counter-1" \
   -i http://localhost:8000/

Worker capacity is versioned: the dataplane (the atelet DaemonSet and the worker pods) schedules only on nodes that carry the ate.dev/substrate-version label, and the install stamps it on every node that exists when it runs. A node added later hosts no workers until you label it with the installed version (kubectl label node <node> ate.dev/substrate-version=<build version>). kubectl get ds -n ate-system -l app=atelet -L ate.dev/substrate-version prints the installed version, off the atelet DaemonSet the install created.

GKE Quickstart (Development)

  1. Create and configure your environment file:

    cp hack/ate-dev-env.sh.example .ate-dev-env.sh
    
    # Edit .ate-dev-env.sh to match your project and preferences, then source it:
    source .ate-dev-env.sh
    
  2. Enable application-default credentials for gcloud:

    gcloud auth application-default login --project=${PROJECT_ID}
    
  3. Provision the required GCP resources (GKE cluster, GCS, and IAM bindings):

    go run ./tools/setup-gcp bootstrap
    

    On a fresh project this step also creates the atelet Workload Identity IAM grants that snapshots depend on — see what create iam actually grants to audit them or apply them manually. If you bring your own cluster instead, note the required Kubernetes beta APIs can only be enabled at cluster creation — see the Create Cluster warning.

  4. Deploy the Agent Substrate system to your cluster:

    ./hack/install-ate.sh --deploy-ate-system
    

    Nodes that GKE adds later (autoscaling, auto-repair, node upgrades) are born with the node pool's labels, so the pool needs ate.dev/substrate-version too; see Node version labels.

  5. You can then deploy the sample applications. See demos/counter/README.md or demos/sandbox/README.md for detailed walkthroughs.

    ./hack/install-ate.sh --deploy-demo-counter
    

Custom Setup and Deployment

You can run individual setup steps to create GCP resources as needed. See go run ./tools/setup-gcp --help for available options. For example:

go run ./tools/setup-gcp create cluster
go run ./tools/setup-gcp create bucket

To run the PostgreSQL store backend on Cloud SQL — with IAM database authentication and no passwords — see tools/setup-gcp/cloud-sql.md.

Similarly, you can deploy or cleanup specific Agent Substrate components using the installation script. See ./hack/install-ate.sh --help for all options.

# Re-deploy only ate-apiserver of the ATE system
./hack/install-ate.sh --deploy-ate-apiserver

# Delete everything (core system and all demos)
./hack/install-ate.sh --delete-all

Tearing down resources (GCP)

If you need to delete the resources created by the setup script, you can use the provided script hack/teardown.sh. This script will delete resources in the reverse order of creation and handles partial failures gracefully.

./hack/teardown.sh --all

Or run individual teardown steps as needed (see ./hack/teardown.sh for available options).

Tearing down local kind resources

If you need to delete the local kind cluster and its registry (if it was created by hack/create-kind-cluster.sh):

./hack/delete-kind-cluster.sh

Demos

We provide several sample applications demonstrating Agent Substrate's capabilities:

  1. Counter Demo: A stateful Go HTTP server demonstrating state preservation across suspends/resumes, and on-demand actor resumption and routing via the Substrate router.
  2. Sandbox Demo (Antigravity): A secure, sandboxed execution environment (running Alpine Linux) that allows arbitrary shell execution while preserving filesystem state across sessions.
  3. Claude Code Multiplex: Demonstrates oversubscribing physical hardware by multiplexing multiple Claude Code agents onto a limited pool of workers.
  4. Multi-Template: Two ActorTemplates running different binaries share one WorkerPool, even though the templates live in different atespaces.
  5. Request Parking: An oversubscribed pool where the router holds inbound requests until a worker frees up, instead of returning 503.
  6. Autoscaled WorkerPool: Scales a WorkerPool on its assigned-worker count with an HPA fed by prometheus-adapter.

Documentation & Guides

Tour

Commands

  • cmd/ateapi: The core control plane API server exposing gRPC endpoints to manage actor and worker lifecycles.
  • cmd/atelet: A node-level DaemonSet that supervises physical worker pods, coordinates snapshotting, and manages state transfers.
  • cmd/atecontroller: A Kubernetes controller that reconciles WorkerPool custom resources.
  • cmd/atenet: A combined networking controller providing Envoy routing and proxy sidecars.
  • cmd/ateom-gvisor: An interior-pod helper running inside sandboxed worker pods to execute runsc checkpoint and restore commands.
  • cmd/ateom-microvm: The micro-VM peer of ateom-gvisor, running actors as cloud-hypervisor VMs.
  • cmd/podcertcontroller: A "polyfill" that provides Pod Certificate signers that will eventually ship in upstream Kubernetes (with different names).
  • cmd/kubectl-ate: A CLI tool for managing Agent Substrate resources. See its README.
  • cmd/benchmarking: Synthetic workloads used by the load tests, including glutton, which consumes RAM, disk, and file descriptors on demand.
  • tools/setup-gcp: A provisioning utility to set up the necessary GCP infrastructure resources (GKE, GCS, IAM).
  • demos/: Sample applications demonstrating Agent Substrate capabilities.

Star History

Star history chart for agent-substrate/substrate
S
Description
GitHub Trending: agent-substrate/substrate
Readme Apache-2.0
216 MiB
Languages
Go 92.8%
Python 3.5%
Shell 2.3%
Go Template 0.7%
HTML 0.2%
Other 0.4%