Part of #1590 ## What this PR does Locust measures the client side only. This adds a post-run harvest of server-side ground truth from Prometheus, written to a new `server_summary.json` and summarized as one row in `stats.jsonl`. ## Proposed Changes ### Steady-state scoping `server_telemetry.py` derives the steady-state window from `stats_history.csv`: it starts at the first sample at 90% of peak user count and ends at the last one. Ramp-up is excluded from every metric, and the snapshot block also stops at the last full-load sample, so teardown suspends are left out. Under a step-ladder load shape the window covers only the top step. ### Harvested metrics **Cluster packing**, from `ate_workerpool_workers`: busy workers (`partial` + `at_capacity`) over the pool size at each sample, where the pool size is the sum of all worker states, so it follows scale-ups and scale-downs mid-run. Only the live ateapi is read: after a redeploy the collector keeps re-exporting exited ateapi processes' last values for a few minutes, which would otherwise inflate the counts. Reported as min, p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries so transient spikes remain visible. **Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall percentages for the node and for the worker pods, each as the same percentile set over the window. **Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both in-window and cumulative, restore and checkpoint latency mean and p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint throughput as `checkpoint_mb_s`. The atelet exports metrics on an interval, so its data reaches Prometheus late: the harvest waits `--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default export and a 10s scrape) and reads both window edges half that late. ### Constraints and failure behavior The module uses only the standard library, since the locust image is distroless, and every request carries a timeout. An unreachable Prometheus records nulls and does not fail the run. `--prometheus-url` overrides the in-cluster default, and `--atelet-lag-s` sets the wait for the atelet's last export. Unmeasured fields are `null` and a measured zero is `0`, consistent with the rest of the runner. Prometheus exposes no byte counter on the restore path, so no restore throughput field is emitted rather than deriving one indirectly. ### Output `server_summary.json` holds the full nested artifact. `stats.jsonl` receives a single `server_summary` row with 45 flat keys for graphing: packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`. `status.json` is unchanged. ## How this was tested 14 unit tests in `test_server_telemetry.py` covering steady-state detection, percentile boundaries, the range-query window guard, malformed Prometheus responses, the packing and checkpoint arithmetic, the per-sample pool size, ignoring exited ateapi series, a missing denominator returning null, snapshot fields returning null rather than zero, and telemetry surviving a missing stats CSV. Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45 keys matched an independent recomputation from raw Prometheus, and the snapshot block was cross-checked against atelet logs. A run with `--prometheus-url` pointed at an unreachable address completes normally with the affected fields null. Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8 workers) from Prometheus with the new code. Runs with a steady pool match the previous output exactly, and runs that followed a redeploy now read the real pool size on every sample. ## References [Agent Substrate: Actor Density Benchmark Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0) - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
Agent Substrate
What is Agent Substrate?
Agent Substrate is a secure-by-default agent execution runtime engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with native zero-trust kernel and network isolation. It supports multiple sandbox technologies including microVMs and gVisor, enabling consistent lifecycle operations for all sandbox types.
At its core, Agent Substrate maps a larger set of “actors” (applications such as agents) onto a smaller set of ready “workers”, relying on the fact that agent-like applications tend to be idle most of the time to achieve heavy multiplexing. It provides functionality to manage an actor’s lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real time, and to route incoming traffic to them.
Agent Substrate is intended to be a low-opinion system. The workloads it manages don't have to be literal AI agents, but those are the best example of the kind of applications it is designed for. It is not an SDK for building agents, but rather a system for running them at scale.
Agent Substrate leverages Kubernetes for the infrastructure provisioning and worker lifecycle management (Kubernetes Pods). It builds on top of Kubernetes features like Pods and Pod autoscaling, while Agent Substrate provides agent-specific scheduling and control to achieve lower latency. Using Kubernetes as the underlying system enables consistent infrastructure management across all workloads types that are required for end to end agentic deployments and allows holistic infrastructure optimizations for RL scenarios that span agentic, inference and training cycles.
Demo
Watch the Agent Substrate cluster multiplex ~250 stateful actors across just 8 physical pods.
This demo highlights the core developer experience and "Agentic Infrastructure" capabilities of Substrate:
- Actor Teleport: High-performance suspend and resume of actors onto any available worker in the pool with sub-second activation.
- State Persistence: Persistent working memory (volatile RAM) and filesystem state preserved perfectly across hibernation cycles via full-state snapshots.
- Agent Multiplexing: Demonstrates 30x+ oversubscription by "juggling" a large registry of stateful actors onto a small pool of shared physical pods.
To reproduce this demo in your own cluster, please refer to the detailed walkthrough in the Counter Demo.
For more videos and walkthroughs, visit our YouTube channel: agent-substrate.
Framework Agnostic & Compatibility
Agent Substrate is designed to be framework and agent harness agnostic. Because it manages standard OCI containers at the kernel level (via gVisor), it can host agents built on any stack.
- Agent Development Kit (ADK): Support for ADK agents with session state preservation across invocations as actor state. Ideal for all types of agents and stateful tool or subagent calls.
- LangChain: Ideal execution environment for LangChain agents and tool calls.
- Claude Code, CodeX, and Antigravity: Support for high-density, stateful coding environments that preserve system state and filesystem state across sessions.
- Model Context Protocol (MCP): Support for deploying secure, sandboxed MCP servers as Substrate Actors to provide durable tools for any model.
Ecosystem & Examples
- Agent Executor: A distributed agent runtime that demonstrates building a secure, hyper-scalable agent harness on Agent Substrate (see the announcement blog and integration guide).
- kagent: A CNCF Sandbox project and Kubernetes-native framework for building, deploying, and managing AI agents that uses Agent Substrate to run sandboxed, stateful agent workloads (see the announcement blog).
Status and compatibility
Agent Substrate is pre-1.0. We are not making any guarantees about backward compatibility at this stage, and APIs and behavior may still change significantly.
Supported Kubernetes Releases
Currently we aim to support the latest stable release of Kubernetes, and the previous minor release.
Community
For announcements, technical discussions, and community support, please join the ate-dev Google Group.
We host a weekly community meeting every Thursday from 10:00am - 11:00am PST.
- Video call link: https://meet.google.com/uhq-cxvn-dhy
- Or dial: (US) +1 253-289-6971 PIN: 787 664 574 59#
- More phone numbers: https://tel.meet/uhq-cxvn-dhy?pin=9044088223662
- Meeting notes for the weekly sync meeting
- Recordings and transcripts of all community meetings
We also have channels in the CNCF slack; request an invite here if you don't have access.
- #substrate-users to discuss using substrate.
- #substrate-dev to discuss developing substrate.
Developing
Please see CONTRIBUTING.md for guidelines on contributing to the project. We welcome contributions of all kinds, but the project is VERY young. Our immediate focus is on building out the core system and demos, so we may not be able to review or merge contributions that don't align with those goals in the near term.
Quickstart (Development)
To quickly set up the complete environment:
-
Make sure you have Go,
kubectl, anddockerinstalled and configured on your dev machine. We will automatically manage other dependencies via Go, includingkind. -
Run the following steps:
# create cluster and local registry (IPv4; IP_FAMILY=dual|ipv6 overrides)
hack/create-kind-cluster.sh
# install ate, PostgreSQL, rustfs
hack/install-ate-kind.sh --deploy-ate-system
# install counter demo
hack/install-ate-kind.sh --deploy-demo-counter
# install kubectl-ate
go install ./cmd/kubectl-ate
# create a counter actor in the demo's atespace (--template names the
# actor template, resolved in the actor's atespace)
kubectl ate create actor my-counter-1 -a ate-demo-counter --template counter
# port-forward the network router to bind to local port `8000`
kubectl port-forward -n ate-system svc/atenet-router 8000:80
- In a separate terminal, send an HTTP request to increment the counter:
curl -X POST \
-H "ate-target-actor: ate-demo-counter/my-counter-1" \
-i http://localhost:8000/
Worker capacity is versioned: the dataplane (the atelet DaemonSet and the
worker pods) schedules only on nodes that carry the
ate.dev/substrate-version label, and the install stamps it on every node
that exists when it runs. A node added later hosts no workers until you label
it with the installed version
(kubectl label node <node> ate.dev/substrate-version=<build version>).
kubectl get ds -n ate-system -l app=atelet -L ate.dev/substrate-version
prints the installed version, off the atelet DaemonSet the install created.
GKE Quickstart (Development)
-
Create and configure your environment file:
cp hack/ate-dev-env.sh.example .ate-dev-env.sh # Edit .ate-dev-env.sh to match your project and preferences, then source it: source .ate-dev-env.sh -
Enable application-default credentials for gcloud:
gcloud auth application-default login --project=${PROJECT_ID} -
Provision the required GCP resources (GKE cluster, GCS, and IAM bindings):
go run ./tools/setup-gcp bootstrapOn a fresh project this step also creates the atelet Workload Identity IAM grants that snapshots depend on — see what
create iamactually grants to audit them or apply them manually. If you bring your own cluster instead, note the required Kubernetes beta APIs can only be enabled at cluster creation — see the Create Cluster warning. -
Deploy the Agent Substrate system to your cluster:
./hack/install-ate.sh --deploy-ate-systemNodes that GKE adds later (autoscaling, auto-repair, node upgrades) are born with the node pool's labels, so the pool needs
ate.dev/substrate-versiontoo; see Node version labels. -
You can then deploy the sample applications. See demos/counter/README.md or demos/sandbox/README.md for detailed walkthroughs.
./hack/install-ate.sh --deploy-demo-counter
Custom Setup and Deployment
You can run individual setup steps to create GCP resources as needed. See go run ./tools/setup-gcp --help for available options. For example:
go run ./tools/setup-gcp create cluster
go run ./tools/setup-gcp create bucket
To run the PostgreSQL store backend on Cloud SQL — with IAM database authentication and no passwords — see tools/setup-gcp/cloud-sql.md.
Similarly, you can deploy or cleanup specific Agent Substrate components using the installation script. See ./hack/install-ate.sh --help for all options.
# Re-deploy only ate-apiserver of the ATE system
./hack/install-ate.sh --deploy-ate-apiserver
# Delete everything (core system and all demos)
./hack/install-ate.sh --delete-all
Tearing down resources (GCP)
If you need to delete the resources created by the setup script, you can use the provided script hack/teardown.sh. This script will delete resources in the reverse order of creation and handles partial failures gracefully.
./hack/teardown.sh --all
Or run individual teardown steps as needed (see ./hack/teardown.sh for available options).
Tearing down local kind resources
If you need to delete the local kind cluster and its registry (if it was created by hack/create-kind-cluster.sh):
./hack/delete-kind-cluster.sh
Demos
We provide several sample applications demonstrating Agent Substrate's capabilities:
- Counter Demo: A stateful Go HTTP server demonstrating state preservation across suspends/resumes, and on-demand actor resumption and routing via the Substrate router.
- Sandbox Demo (Antigravity): A secure, sandboxed execution environment (running Alpine Linux) that allows arbitrary shell execution while preserving filesystem state across sessions.
- Claude Code Multiplex: Demonstrates oversubscribing physical hardware by multiplexing multiple Claude Code agents onto a limited pool of workers.
- Multi-Template: Two
ActorTemplates running different binaries share oneWorkerPool, even though the templates live in different atespaces. - Request Parking: An oversubscribed pool where the router holds inbound requests until a worker frees up, instead of returning
503. - Autoscaled WorkerPool: Scales a
WorkerPoolon its assigned-worker count with an HPA fed by prometheus-adapter.
Documentation & Guides
- Architecture: How the control plane, node supervisor, and networking stack fit together.
- API Configuration Guide: Detailed reference for configuring WorkerPools, ActorTemplates, Secrets, and Volumes.
- Full CLI Documentation: Installation and usage for
kubectl-ate. - Glossary: Core terms (Actor, Atespace, ActorTemplate, WorkerPool, Worker, ate-api-server, atenet, atelet, ateom) and how they relate.
- Integration Repositories: Where integrations live, how their repositories are named, and how fixes flow back to core.
- Observability Guide: Guide to actor logging, metrics, and distributed tracing.
- Authentication Guide: Configure trusted JWT providers and human credentials.
- Egress Traffic: Which protocols an Actor may reach the outside world with and which are blocked.
- Enabling man-in-the-middle (MITM) interception for Actor Egress policy: Egress policies such as header injection depend on MITM interception of Actor traffic. This guide explains how an Actor should be configured to enable interception.
- Request Parking: How the router parks requests through transient worker-pool saturation.
- Rolling Upgrade Runbook: Upgrade a running substrate node by node without losing actor state.
- Threat Model: Trust boundaries, assumptions, and known risks.
- Roadmap: Current limitations and what is planned next.
- Benchmarking Guide: Locust-based load tests, monitoring stack, and the orchestrated benchmark harness.
Tour
Commands
cmd/ateapi: The core control plane API server exposing gRPC endpoints to manage actor and worker lifecycles.cmd/atelet: A node-level DaemonSet that supervises physical worker pods, coordinates snapshotting, and manages state transfers.cmd/atecontroller: A Kubernetes controller that reconciles WorkerPool custom resources.cmd/atenet: A combined networking controller providing Envoy routing and proxy sidecars.cmd/ateom-gvisor: An interior-pod helper running inside sandboxed worker pods to executerunsccheckpoint and restore commands.cmd/ateom-microvm: The micro-VM peer ofateom-gvisor, running actors as cloud-hypervisor VMs.cmd/podcertcontroller: A "polyfill" that provides Pod Certificate signers that will eventually ship in upstream Kubernetes (with different names).cmd/kubectl-ate: A CLI tool for managing Agent Substrate resources. See its README.cmd/benchmarking: Synthetic workloads used by the load tests, includingglutton, which consumes RAM, disk, and file descriptors on demand.tools/setup-gcp: A provisioning utility to set up the necessary GCP infrastructure resources (GKE, GCS, IAM).demos/: Sample applications demonstrating Agent Substrate capabilities.
