Add a Star History section to the README that embeds star-history.com's
live chart, with a dark variant for GitHub's dark theme.
star-history.com renders the image on request, so it stays current;
browsers and GitHub's image proxy cache it for up to a day.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Removes the "not an officially supported Google product" note from the
top of the README, and replaces the "early development / not ready for
production" wording with a plainer pre-1.0 compatibility statement.
The Google-supported product is
[ai-on-gke/substrate-gke](https://github.com/ai-on-gke/substrate-gke).
The Vulnerability Rewards Program statement is unchanged and stays in
[.github/SECURITY.md](https://github.com/agent-substrate/substrate/blob/main/.github/SECURITY.md);
only its "not eligible" link to the removed README note is dropped.
Record which egress protocols are allowed, which are allowed only under
policy controls, and which are blocked, along with the data path each
one takes. Having this written down gives users a single place to check
what Substrate lets an actor reach, and gives us a checklist to
implement and test against before GA.
Point readers at the issue tracker so that requests for traffic we do
not yet support arrive with a use case attached.
Address #1339
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## What
Updates the project overview language in the three docs that share it,
and fixes a long-standing typo.
- **`README.md`** — replaces the overview paragraph with the
secure-by-default positioning: density relative to standard container
runtimes, resume latency and activation throughput, and native
kernel/network isolation.
- **`docs/architecture.md`** — adopts the same lead sentence, keeping
the existing control-plane detail; `computer infrastructure` → `compute
infrastructure`.
- **`docs/roadmap.md`** — `computer infrastructure` → `compute
infrastructure`.
## Notes
The performance figures in the README paragraph (density multiple,
sub-500ms resume, activation rate) have been discussed and aligned
separately.
Docs-only change; no code or behavior is affected.
* Fix the column names to match the resource fields.
* Rename template flag name (fix TODO from CRD -> Substrate API
refactor).
* Fix bug in `create actor` which conflated the atespace flag for both
actor and actor template.
## What
Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.
Fixes#368 (docs side).
> This change was prepared with AI assistance; I have reviewed and
tested it.
- [x] Docs and comment-only change; no functional code changed, no tests
affected
Update the framework compatibility section to better describe how Agent
Substrate preserves session and system state across invocations for ADK,
LangChain, MCP, and coding agents (Claude Code, CodeX, and Antigravity).
Also simplify the langauge in the demo highlights.
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.
To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.
### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:**
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
The runbook for moving a running substrate to a new build without losing
actor state. An user runs the roll by hand with `kubectl`, `kubectl
ate`, `ate-setup`, `jq`, and `grpcurl`, and every piece of upgrade state
lives in cluster objects, so the roll can stop and resume at any point.
The order follows the [upgrade
design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A):
CRDs, then `ate-controller`, then the dataplane node by node, then
`ate-api-server` and `atenet`.
The dataplane moves by version label. The new atelet DaemonSet sits next
to the old one, each serving WorkerPool is cloned with the new worker
image and the new version pin, and each node is drained (workers marked
`DRAINING` through the `DrainWorker` RPC), emptied by suspending its
actors, relabeled, and cleared of old worker pods. The old objects stay
untouched until a separate retire step, so rollback is one label flip
per node.
Tested end to end locally on a GKE cluster: install at one build, run
counter actors, upgrade to a second build following the document, and
confirm the actors resume on the new pool with their state intact.
Rollback and retire were exercised the same way.
Fix#1272, fix#1273.
An sdsmint egress gateway terminates every TLS connection an actor opens
and re-originates it, so what the actor validates is a per-SNI leaf the
gateway minted rather than the origin's certificate. That leaf chains to
the gateway CA and to no public root, so an actor left on its default
trust store fails every HTTPS request it makes.
Fixes#1005
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#368 . Deletes ActorTemplate CRD and any references to it.
Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.
Verifications done:
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).
This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.
### One derivation for the version label
`internal/versionlabel` turns the build version into its two forms in
one place:
| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |
A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.
### A DaemonSet per version
`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.
### The install becomes version-aware
- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.
### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.
Fix#1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
certificates.k8s.io/v1beta1 podcertificaterequests and
clustertrustbundles cannot be enabled in place on an existing cluster -
the update is accepted but the APIs never become served, and the
install hangs waiting for ClusterTrustBundles. Warn in the create
cluster docs, show the bring-your-own-cluster flag, and name the
symptom.
On a fresh GCP project nothing documented which IAM bindings atelet
needs: setup-gcp creates them silently, and anyone who cannot run the
tool with project-level IAM permissions - or needs to audit what it
did - had to read cmd/iam.go and cmd/bucket.go. Spell out the exact
members, roles, and resources, the Workload Identity prerequisites,
and the gcloud equivalents, and point the README quickstart at it.
## Summary
A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.
- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable
## Benchmarking
Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:
-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing
---------
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
## Summary
Adds `docs/integration-repos.md`: where end-to-end integrations live,
how their
repositories are named, and how the fixes they need flow back into core.
The convention in one line — trivial demos stay in the core repo, each
non-trivial integration gets one dedicated repo under the
`agent-substrate`
org, and core gaps get closed by making core configurable with defaults
unchanged rather than by patching it downstream.
## Why now
We are about to create the first real, end-to-end integrations rather
than
counter-style demos: a code-execution sandbox, and an always-on agent.
Both are
large enough to need their own images, dependencies, and release
cadence.
Whichever repository gets created first will set the precedent for every
one
after it. This writes the convention down so that precedent is chosen
deliberately instead of inherited by accident.
## What it covers
- **Where code lives** — the core-repo/dedicated-repo split, the rough
test for
which side something falls on (API keys, external services, third-party
accounts), and why this is a set of peer repos rather than a second org.
- **Naming** — capability-named for general capabilities
(`code-execution-sandbox`), integration-named for specific third-party
products, named for the product rather than the vendor behind it. Plus
what to
avoid: over-broad names, names that clone a vendor's API or brand, and
the
redundant `-integration` suffix.
- **Third-party names** — allowed descriptively, with a non-affiliation
note in
the repo README, and brand/policy edge cases cleared before the repo
exists.
- **Upstreaming** — the part with teeth for this repo. Integration repos
that
accumulate local patches against core bitrot, and the gap they work
around
stays invisible to everyone else. So: prefer making core behavior
configurable
with defaults unchanged. #487 and #465 are linked as illustrations of
that
pattern — this PR does not depend on either, and branches from `main`.
- **Two worked examples** that validate the convention rather than just
following it, including the third-party-name edge case.
## Review
This was announced at the community meeting and circulated as a shared
design
doc with a 7-day review window, which has now closed. It synthesizes the
`#integrations` thread discussion. Comment history:
<https://docs.google.com/document/d/1Tb6u0b1XSvWrNpoyD4jdsQaJ58aAgDtQOM18uxujs-8/edit>
This PR is the trimmed version: doc-review scaffolding — status block,
reviewer
list, self-link — is dropped, and only the durable convention is carried
over.
## Left open
Two questions are deliberately out of scope, called out in the doc
rather than
answered. Both are maintainer calls and neither blocks the first
repositories:
- Governance tiers — whether to distinguish "official" from "community"
integrations with different review bars, as Home Assistant and Obsidian
do.
- Who creates integration repositories and grants per-integration
maintainer
access.
## Also in this PR
- README gets an entry in the docs list, matching every other file in
`docs/`.
- `CONTRIBUTING.md` gets one sentence pointing there, since "where does
my
integration go?" is a question a contributor asks before opening a PR.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
IP_FAMILY selects ipv4, ipv6 or dual and becomes networking.ipFamily, leaving
kind's per-family subnet defaults alone. The script also recreates a pre-IPv6
"kind" Docker network, fails fast if the daemon has IPv6 off, sets proxy_ndp
alongside proxy_arp for gVisor pod-to-pod traffic, and repoints an ipv6
kubeconfig from [::1] at localhost so a client outside the Docker host can
still reach the apiserver.
Tested on kind with all three families: node InternalIPs, Service ClusterIPs
and pod IPs land in the requested families, pod-to-pod and CoreDNS work on
them, and pods still pull through the local registry.
Fixed language describing project goals and relationship to Kubernetes
in readme.md, architecture.md and roadmap.md
- changed the top language to focus on project goals instead of
Kubernetes relations
- changed the language describing Kubernetes relationship to describe
value of the Agent Substrate layer and Kubernetes layer
The README lists 3 of the 6 demos and 5 of the 9 docs, and its command
tour skips `cmd/ateom-microvm` and `cmd/benchmarking`. The glossary
never defines Atespace, even though it is half of an actor's identity
and has its own API, and it omits SandboxConfig.
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.
API surface:
service SessionIdentity -> ActorIdentity
MintJWTRequest.session_id -> actor_id
MintJWTResponse.session_jwt -> actor_jwt
MintCertRequest.session_id -> actor_id
MintCertResponse.session_certificates -> actor_certificates
Go packages:
cmd/ateapi/internal/sessionidentity -> actoridentity
cmd/ateapi/internal/sessionidjwt -> actoridjwt
Flags and cluster resources:
--session-id-jwt-pool -> --actor-id-jwt-pool
--session-id-ca-pool -> --actor-id-ca-pool
Secrets, volumes and mount paths renamed to match, in both
manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
(--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).
Two credential identity values change with the rename:
JWT issuer https://broker.agentic-substrate-session-id-broker.svc
-> https://broker.agentic-substrate-actor-id-broker.svc
SPIFFE ID spiffe://substrate-session.local/app/../session/..
-> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.
BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
This demo is broken now that we require authn to ate-apiserver. The
fix is not trivial as it requires the actor to be able to authenticate, plus
it doesn't really adds much.
- `docs/observability.md`: log label keys were wrong. The code emits
`ate.dev/actor_id` and `ate.dev/actor_template_name`; the doc had
`actor_id` and `actor_template`. The Cloud Logging queries would return
no results.
- `docs/dev/best-practices/tracing.md`: container ports entry used
Docker Compose syntax (`"443:443"`). Changed to `containerPort: 443`.
- `docs/dev/valkey-direct-access.md`: the `kubectl exec` command had a
line break inside `--cert`, making it fail on copy-paste.
- `code-of-conduct.md`: linked to Contributor Covenant v1.4. Updated to
v2.1.
Update README.md and scripts to adopt atespaces for existing demos.
It's intentional to let user choose atespace in demos since actors are
also created by them.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Addresses #131
Adds `docs/glossary.md`: a single place that defines the terms used
across
Agent Substrate and shows how they relate.
### What it covers
- Resources (ActorTemplate, WorkerPool), control-plane records (Actor,
Worker),
components (ate-api-server, atecontroller, atelet, ateom, atenet,
podcertcontroller, kubectl-ate), runtime/sandbox concepts (gVisor,
runsc,
pause container, checkpoint/restore), snapshots (golden, last, storage),
and
the Uniform DNS Mesh.
- Three UML diagrams: a sequence diagram (activation/resume flow), a
class
diagram (resource/ownership model, grouped into Kubernetes-object and
control-plane-record packages), and a state-machine diagram (actor
lifecycle).
- Linked from the README documentation section.
### Scope
Per the issue thread, this documents the names as they exist in the
codebase;
rename proposals are intentionally left to a separate discussion, so
this PR is
documentation-only (not a renaming proposal).
### Accuracy
Every term and diagram was checked against the code and existing docs.
Project
terminology and casing are reused where things are already named (for
example
"Uniform DNS Mesh", the atelet "Herder", and the "interior gVisor"
coordinator),
and component responsibilities, RPC names, relationships, and the actor
status
lifecycle map to their sources.
- [x] Tests pass (documentation-only change)
- [x] Appropriate changes to documentation are included in the PR
This commit was automatically generated by the tools/reorg tool.
```
go run ./tools/reorg
go generate ./...
bash benchmarking/locust/generate_protos.sh
```
Fixes#55
I was following the demo but ran into the issues described in #55. The
command is failing because the curl command runs before the
port-forwarding is ready.
We've already been managing other go binaries with tools modules.
Doing this limits the number of setup steps needed, and allows our
scripts to make assumptions about the kind version.
`kind` has minimal dependencies and should compile very quickly.
Label the registry as created by our script, so we know if we should
delete it.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Incorporate Agent Executor as a demonstrative example of a distributed
agent runtime and harness built on Agent Substrate.
ISSUE=None
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Co-authored-by: Maya Wang <mymaya@google.com>
This is the initial release of the Agent Substrate.
Agent substrate is a system built on top of Kubernetes which manages agent-like
workloads to achieve higher scale and efficiency than Kubernetes alone can
offer, with lower latency. It builds on top of Kubernetes features like
Pods and Pod autoscaling, but takes the Kubernetes control-plane out of the
critical path to achieve lower latency.
It can run on any Kubernetes cluster and does not inhibit “regular” use of
Kubernetes in any way. Kubernetes provides the infrastructure provisioning and
management for all types of workloads, while Agent Substrate provides
agent-specific scheduling and control.
At its core, Agent Substrate maps a larger set of “actors” (applications such
as agents) onto a smaller set of ready “workers” (Kubernetes Pods), relying on
the fact that agent-like applications tend to be idle most of the time to
achieve heavy multiplexing. It provides functionality to manage an actor’s
lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real
time, and to route incoming traffic to them.
Agent Substrate is intended to be a low-opinion system. The workloads it
manages don't have to be literal AI agents, but those are the best example of
the kind of applications it is designed for. It is not an SDK for building
agents, but rather a system for running them at scale.
Agent Substrate is currently in VERY early development. It is not ready for
production use, and the APIs are almost guaranteed to change. We are not
making any guarantees about backward compatibility at this stage, and
everything in this project may be changed.
Co-authored-by: Alex Bulankou <alexbu@google.com>
Co-authored-by: Benjamin Elder <bentheelder@google.com>
Co-authored-by: Bowei Du <bowei@google.com>
Co-authored-by: Dmitry Berkovich <dberkov@google.com>
Co-authored-by: Fabricio Voznika <fvoznika@google.com>
Co-authored-by: Francisco Cabrera <fclieutier@google.com>
Co-authored-by: Haven Xia <haoyuxia@google.com>
Co-authored-by: Julian Gutierrez Oschmann <juliangut@google.com>
Co-authored-by: Kevin Steuer <ksteuer@google.com>
Co-authored-by: Max Smythe <smythe@google.com>
Co-authored-by: Maya Wang <mymaya@google.com>
Co-authored-by: Michael Taufen <mtaufen@google.com>
Co-authored-by: Shruti Nair <shrutinair@google.com>
Co-authored-by: Taahir Ahmed <taahm@google.com>
Co-authored-by: Tim Hockin <thockin@google.com>
Co-authored-by: Zoe Zhao <zoezhao@google.com>