Jonathan Jamroga 6f45aeb2dd Make substrate's namespace, Service names and ServiceAccount names configurable (#350)
# Description

**Substrate assumes the canonical install layout, and every deviation
fails closed.** The namespace `ate-system`, the Services `api` /
`atenet-router`, and the ServiceAccounts `atelet` / `atenet-router` are
compiled-in constants. An install in a per-developer namespace, or under
a deployment that prefixes resource names, breaks — and no failure
points at the naming.

This series makes the namespace, the Service names, and the
ServiceAccount names configurable. Every option defaults to the
canonical value. A canonical install is byte-for-byte unaffected.

| Hardcoded assumption | Failure | What it looks like instead |
|---|---|---|
| atelet's namespace in ateapi's SPIFFE check | ateapi rejects every
atelet | an mTLS handshake failure, not a naming error |
| atelet's namespace in the worker's broker check | no actor obtains a
certificate | `credential broker is not atelet` |
| NetworkPolicy's ingress namespace | the CNI drops every request to the
pool | a silent network fault |
| ateapi Service name and namespace in the client | the client cannot
authenticate | `services "api" not found`, then `invalid bearer token` |
| Resource names in `hack/install-ate.sh` and the e2e harness |
authorities land in the wrong namespace | suites die in preflight |

**A SPIFFE ID breaks on two axes.** It names a namespace and a
ServiceAccount. Both were constants, and the failure surfaces as a
rejected peer, not a missing object.

## Design notes

- **ate-controller passes the worker-side identities.** ateom already
took `--atunnel-client-identity` as a flag but relied on its default. A
new `--atunnel-broker-identity` flag alone would be inert: correct only
where the constant was already correct. The controller knows the control
plane's namespace, so it supplies both.
- **ateapi takes the whole expected identity, not a namespace.** The
dialer and `ateletauth` used the namespace only to build one string. One
place decides how atelet's identity is spelled.
- **`ateletauth` deduplicates without taking the dependency it was
avoiding.** Its constants are duplicated rather than imported so the
package does not depend on `controlapi` for three strings. `ateletdial`
would make a third copy, so the strings move to
`internal/installdefaults` instead — a leaf package of constants that
imports only the standard library, so consuming it does not reintroduce
the coupling the duplication was there to prevent.
- **The client reads environment variables, not flags.** It runs outside
the cluster. It has no downward API and nothing to discover from.
- **Namespaces come from the downward API, not flags.** Every supported
topology co-locates these components. Only Service and ServiceAccount
names, which a deployment may legitimately rename, get flags.
- **`hack/install-ate.sh` and the e2e harness are in scope.** A
relocated install cannot be bootstrapped or exercised without them. The
script refuses `ATE_NAMESPACE` or `ATE_API_SERVICE_NAME` overrides on
the manifest-applying subcommands, which would half-install:
`manifests/ate-install/` names `ate-system` and `api` literally.

## Why this is one PR

The series is stacked, not parallel. Commit 1 creates
`internal/installdefaults` and seeds seven imports; commit 2 adds
fourteen more; every later commit builds on those constants. Split into
separate PRs, each blocks on the previous merging and none reviews
independently.

The guard test cannot land first. Applied to `main` it fails on eleven
hardcoded literals across `ateletauth`, `controlapi/informer.go`,
`networkpolicy_controller.go`, both `ateom` mains, `ateclient`,
`ateletdial` and three e2e files. It is green only because the preceding
commits removed them.

A partial series still fails closed. Each commit's audit turned up more
places assuming `ate-system`, so landing commits 1-3 without 4-5 leaves
a relocated install broken later and less legibly than before.

`CONTRIBUTING.md` covers this case: when the intermediate steps are not
useful on their own, keep the change as one PR split into commits at
logical break points. Review commit by commit, and preserve the commits
on merge. If you would still rather split, the one defensible cut is by
axis — relocation (commits 1-3) then renaming (commits 4-5) with the
guard test rebased on top.

## Rebase notes (2026-09-15)

The series is rebased onto `main` at `d0d85c39`, ninety-two commits on
from the base the PR first carried:

- `main` moved actor JWT and certificate minting out of
`cmd/ateapi/internal/actoridentity` into `controlapi`, deleting the
package and dropping its inline atelet authorization check in favor of
the unified authorizer. This series previously threaded the expected
atelet identity into that check; the check no longer exists, so that
adaptation is gone. `ateletauth` survives with one caller,
`workerservice.SetWorkerCapacity`, which still takes the configured
identity.
- The same refactor replaced `BrokerConfig.ExpectedActorUID` with an
`ActorAtespace` / `ActorName` / `ActorUID` triple. `AteletSPIFFEID` is
additive and still required by `ateletdial.TLSConfig`.
- `ateletauth` (ateapi's side) and `ateletdial` (the worker's side) each
re-declared the hardcoded SPIFFE ID. Both take the expected identity as
a parameter.
- `main` deleted the atenet DNS subsystem. The series no longer touches
it, and `installdefaults` carries no DNS Service name.
- The controller passes `--atunnel-broker-identity` only when it differs
from the canonical default. An ateom old enough to predate the flag
exits on it, and `docs/upgrade.md` keeps such a pool serving during a
rolling upgrade. A relocated install gets the flag and necessarily runs
an image that accepts it.
- `main` added `internal/e2e/collector_metrics.go` (agentgateway CI
support), which addresses the router through the package-level
`routerNamespace` and `routerService` that this series replaces. Its two
call sites become `SystemNamespace()` and
`ResourceName("atenet-router")`, matching `router_client.go` and
`statusz.go`. A reviewer diffing against the older base sees those call
sites move; nothing else in that file changes.
- `main` added `cmd/credential-provider/kubernetes-secrets` (`#1335`),
whose `injectorSPIFFEID` constant is the mTLS peer check on the only
caller permitted to read Secrets. The comparison is an exact string
match, so a renamed install rejects every fetch at TLS. It becomes
`--injector-identity`, defaulting to
`installdefaults.EgressSPIFFEID(SystemNamespace)`, alongside a new
`EgressServiceAccount` constant and helper. The flag name follows
`--atunnel-client-identity`; the manifest beside it still names
`ate-system`, so a renaming deployment configures both.
- `main` replaced the dialer's worker indexer with
`DialForAteletOnNode`, dropped `WorkerPodInformer` from `controlapi`,
and removed `kataConfig` from ateom-microvm's `NewService`. The series
adapts to each narrower signature and keeps only its own added
parameter.
- The guard test's allowlist entry for the two inert `ate-system`
literals follows the code from `actoridentity.go` to
`controlapi/actor.go`. The refactor re-introduced exactly the class of
constant this series removes, so the guard earns its place.

# Testing

- `go test -race ./...` passes. `make verify` passes boilerplate,
codegen, go-modules, gofmt, golangci-lint, kube-api-linter, licenses,
metrics and postgresql-migrations. `proto-fmt` needs `clang-format`,
which is absent locally; this series touches no `.proto` files.
- Every commit builds and vets individually, via `git rebase -x 'go
build ./... && go vet ./...'`.
- On a fresh kind cluster, all twelve e2e suites pass: `capabilities`,
`combinedvolumes`, `demo`, `egressauthz`, `egressmitm`, `example`,
`identity`, `metrics`, `networking`, `networkpolicy`, `parking`,
`sizing`. The `demo`, `metrics`, `networkpolicy`, `parking` and
`networking` suites need the `--deploy-demo-counter` and
`--deploy-demo-egress` fixtures, and the MITM path needs
`--deploy-demo-egress-mitm`; without them they fail on `actor template
not found`, which reads as a control-plane fault rather than a missing
fixture.
- With the sdsmint egress gateway (`--deploy-atenet
--experimental-use-sdsmint`, `E2E_EGRESS_MITM=1`),
`TestActorEgressMITMTrust` and `TestActorEgressHTTPSByHostnameMITM` both
pass, and the `networking` suite is green at 25 passed. The two modes
are mutually exclusive by design:
`TestActorEgressHTTPSByHostnamePassthrough` covers the plain gateway and
skips under MITM, and the two MITM tests skip without it. Both modes
were run, so every case executed in one of them.
- The new unit tests use a relocated namespace and renamed
ServiceAccounts. Reintroducing each hardcoding makes them fail.

One e2e caveat that predates this series and misleads: the `identity`
suite calls `ReplaceEgressTrustPool`, which takes over the shared
`egress-mitm-ca-pool` Secret, overwrites it and registers no cleanup.
Any MITM test running afterwards fails with `certificate signed by
unknown authority`, which reads as a broken interception path rather
than a mutated fixture. Reinstall the gateway before running
`egressmitm`.

# Additional Notes

**Credential contents stay canonical.** `controlapi/actor.go` still
mints credentials naming `api.ate-system.svc`: the JWT issuer and the
certificate's Issuer CN. Relying parties validate these, so changing
them is a compatibility decision, not a lookup fix. The code carries a
TODO to make the issuer a globally unique, OIDC-compliant name.
2026-09-22 22:28:40 +00:00
…
…
…
…
…
…
…
…

Agent Substrate

License

NOTE: This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.

What is Agent Substrate?

Agent Substrate is a secure-by-default agent execution runtime engineered to run millions of sandboxes with 10x higher density than standard container runtimes. Purpose-built for the era of autonomous agents, Substrate delivers sub-500ms resume operations at over 500 suspend/resume activations per second with native zero-trust kernel and network isolation. It supports multiple sandbox technologies including microVMs and gVisor, enabling consistent lifecycle operations for all sandbox types.

At its core, Agent Substrate maps a larger set of “actors” (applications such as agents) onto a smaller set of ready “workers”, relying on the fact that agent-like applications tend to be idle most of the time to achieve heavy multiplexing. It provides functionality to manage an actor’s lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real time, and to route incoming traffic to them.

Agent Substrate is intended to be a low-opinion system. The workloads it manages don't have to be literal AI agents, but those are the best example of the kind of applications it is designed for. It is not an SDK for building agents, but rather a system for running them at scale.

Agent Substrate leverages Kubernetes for the infrastructure provisioning and worker lifecycle management (Kubernetes Pods). It builds on top of Kubernetes features like Pods and Pod autoscaling, while Agent Substrate provides agent-specific scheduling and control to achieve lower latency. Using Kubernetes as the underlying system enables consistent infrastructure management across all workloads types that are required for end to end agentic deployments and allows holistic infrastructure optimizations for RL scenarios that span agentic, inference and training cycles.

Demo

Agent Substrate Demo

Watch the Agent Substrate cluster multiplex ~250 stateful actors across just 8 physical pods.

This demo highlights the core developer experience and "Agentic Infrastructure" capabilities of Substrate:

  1. Actor Teleport: High-performance suspend and resume of actors onto any available worker in the pool with sub-second activation.
  2. State Persistence: Persistent working memory (volatile RAM) and filesystem state preserved perfectly across hibernation cycles via full-state snapshots.
  3. Agent Multiplexing: Demonstrates 30x+ oversubscription by "juggling" a large registry of stateful actors onto a small pool of shared physical pods.

To reproduce this demo in your own cluster, please refer to the detailed walkthrough in the Counter Demo.

For more videos and walkthroughs, visit our YouTube channel: agent-substrate.

Framework Agnostic & Compatibility

Agent Substrate is designed to be framework and agent harness agnostic. Because it manages standard OCI containers at the kernel level (via gVisor), it can host agents built on any stack.

  • Agent Development Kit (ADK): Support for ADK agents with session state preservation across invocations as actor state. Ideal for all types of agents and stateful tool or subagent calls.
  • LangChain: Ideal execution environment for LangChain agents and tool calls.
  • Claude Code, CodeX, and Antigravity: Support for high-density, stateful coding environments that preserve system state and filesystem state across sessions.
  • Model Context Protocol (MCP): Support for deploying secure, sandboxed MCP servers as Substrate Actors to provide durable tools for any model.

Ecosystem & Examples

  • Agent Executor: A distributed agent runtime that demonstrates building a secure, hyper-scalable agent harness on Agent Substrate (see the announcement blog and integration guide).
  • kagent: A CNCF Sandbox project and Kubernetes-native framework for building, deploying, and managing AI agents that uses Agent Substrate to run sandboxed, stateful agent workloads (see the announcement blog).

Status and compatibility

Agent Substrate is currently in early development. It is not ready for production use, and the APIs are almost guaranteed to change. We are not making any guarantees about backward compatibility at this stage, and everything in this project may be changed.

Supported Kubernetes Releases

Currently we aim to support the latest stable release of Kubernetes, and the previous minor release.

Community

For announcements, technical discussions, and community support, please join the ate-dev Google Group.

We host a weekly community meeting every Thursday from 10:00am - 11:00am PST.

We also have channels in the CNCF slack; request an invite here if you don't have access.

Developing

Please see CONTRIBUTING.md for guidelines on contributing to the project. We welcome contributions of all kinds, but the project is VERY young. Our immediate focus is on building out the core system and demos, so we may not be able to review or merge contributions that don't align with those goals in the near term.

Quickstart (Development)

To quickly set up the complete environment:

  1. Make sure you have Go, kubectl, and docker installed and configured on your dev machine. We will automatically manage other dependencies via Go, including kind.

  2. Run the following steps:

# create cluster and local registry (IPv4; IP_FAMILY=dual|ipv6 overrides)
hack/create-kind-cluster.sh

# install ate, PostgreSQL, rustfs
hack/install-ate-kind.sh --deploy-ate-system

# install counter demo
hack/install-ate-kind.sh --deploy-demo-counter

# install kubectl-ate
go install ./cmd/kubectl-ate

# create a counter actor in the demo's atespace (--template names the
# actor template, resolved in the actor's atespace)
kubectl ate create actor my-counter-1 -a ate-demo-counter --template counter

# port-forward the network router to bind to local port `8000`
kubectl port-forward -n ate-system svc/atenet-router 8000:80
  1. In a separate terminal, send an HTTP request to increment the counter:
curl -X POST \
   -H "ate-target-actor: ate-demo-counter/my-counter-1" \
   -i http://localhost:8000/

Worker capacity is versioned: the dataplane (the atelet DaemonSet and the worker pods) schedules only on nodes that carry the ate.dev/substrate-version label, and the install stamps it on every node that exists when it runs. A node added later hosts no workers until you label it with the installed version (kubectl label node <node> ate.dev/substrate-version=<build version>). kubectl get ds -n ate-system -l app=atelet -L ate.dev/substrate-version prints the installed version, off the atelet DaemonSet the install created.

GKE Quickstart (Development)

  1. Create and configure your environment file:

    cp hack/ate-dev-env.sh.example .ate-dev-env.sh
    
    # Edit .ate-dev-env.sh to match your project and preferences, then source it:
    source .ate-dev-env.sh
    
  2. Enable application-default credentials for gcloud:

    gcloud auth application-default login --project=${PROJECT_ID}
    
  3. Provision the required GCP resources (GKE cluster, GCS, and IAM bindings):

    go run ./tools/setup-gcp bootstrap
    

    On a fresh project this step also creates the atelet Workload Identity IAM grants that snapshots depend on — see what create iam actually grants to audit them or apply them manually. If you bring your own cluster instead, note the required Kubernetes beta APIs can only be enabled at cluster creation — see the Create Cluster warning.

  4. Deploy the Agent Substrate system to your cluster:

    ./hack/install-ate.sh --deploy-ate-system
    

    Nodes that GKE adds later (autoscaling, auto-repair, node upgrades) are born with the node pool's labels, so the pool needs ate.dev/substrate-version too; see Node version labels.

  5. You can then deploy the sample applications. See demos/counter/README.md or demos/sandbox/README.md for detailed walkthroughs.

    ./hack/install-ate.sh --deploy-demo-counter
    

Custom Setup and Deployment

You can run individual setup steps to create GCP resources as needed. See go run ./tools/setup-gcp --help for available options. For example:

go run ./tools/setup-gcp create cluster
go run ./tools/setup-gcp create bucket

To run the PostgreSQL store backend on Cloud SQL — with IAM database authentication and no passwords — see tools/setup-gcp/cloud-sql.md.

Similarly, you can deploy or cleanup specific Agent Substrate components using the installation script. See ./hack/install-ate.sh --help for all options.

# Re-deploy only ate-apiserver of the ATE system
./hack/install-ate.sh --deploy-ate-apiserver

# Delete everything (core system and all demos)
./hack/install-ate.sh --delete-all

Tearing down resources (GCP)

If you need to delete the resources created by the setup script, you can use the provided script hack/teardown.sh. This script will delete resources in the reverse order of creation and handles partial failures gracefully.

./hack/teardown.sh --all

Or run individual teardown steps as needed (see ./hack/teardown.sh for available options).

Tearing down local kind resources

If you need to delete the local kind cluster and its registry (if it was created by hack/create-kind-cluster.sh):

./hack/delete-kind-cluster.sh

Demos

We provide several sample applications demonstrating Agent Substrate's capabilities:

  1. Counter Demo: A stateful Go HTTP server demonstrating state preservation across suspends/resumes, and on-demand actor resumption and routing via the Substrate router.
  2. Sandbox Demo (Antigravity): A secure, sandboxed execution environment (running Alpine Linux) that allows arbitrary shell execution while preserving filesystem state across sessions.
  3. Claude Code Multiplex: Demonstrates oversubscribing physical hardware by multiplexing multiple Claude Code agents onto a limited pool of workers.
  4. Multi-Template: Two ActorTemplates running different binaries share one WorkerPool, even though the templates live in different atespaces.
  5. Request Parking: An oversubscribed pool where the router holds inbound requests until a worker frees up, instead of returning 503.
  6. Autoscaled WorkerPool: Scales a WorkerPool on its assigned-worker count with an HPA fed by prometheus-adapter.

Documentation & Guides

Tour

Commands

  • cmd/ateapi: The core control plane API server exposing gRPC endpoints to manage actor and worker lifecycles.
  • cmd/atelet: A node-level DaemonSet that supervises physical worker pods, coordinates snapshotting, and manages state transfers.
  • cmd/atecontroller: A Kubernetes controller that reconciles WorkerPool custom resources.
  • cmd/atenet: A combined networking controller providing Envoy routing and proxy sidecars.
  • cmd/ateom-gvisor: An interior-pod helper running inside sandboxed worker pods to execute runsc checkpoint and restore commands.
  • cmd/ateom-microvm: The micro-VM peer of ateom-gvisor, running actors as cloud-hypervisor VMs.
  • cmd/podcertcontroller: A "polyfill" that provides Pod Certificate signers that will eventually ship in upstream Kubernetes (with different names).
  • cmd/kubectl-ate: A CLI tool for managing Agent Substrate resources. See its README.
  • cmd/benchmarking: Synthetic workloads used by the load tests, including glutton, which consumes RAM, disk, and file descriptors on demand.
  • tools/setup-gcp: A provisioning utility to set up the necessary GCP infrastructure resources (GKE, GCS, IAM).
  • demos/: Sample applications demonstrating Agent Substrate capabilities.
S
Description
GitHub Trending: agent-substrate/substrate
Readme Apache-2.0
216 MiB
Languages
Go 92.8%
Python 3.5%
Shell 2.3%
Go Template 0.7%
HTML 0.2%
Other 0.4%