Adds a credential provider for egress credential injection backed by
Google Cloud Secret Manager.
**It lives in its own Go module under `plugins/gcp-secret-manager`,
temporally hosted here until it moves to a repository of its own.**
**What it does**
- Serves `credproviderpb.CredentialProvider` over mTLS and admits only
the egress gateway's identity (`--injector-identity`).
- Resolves global and regional secrets, optionally picking one key out
of a JSON payload:
`ate-secret://secretmanager.googleapis.com/projects/<project>[/locations/<location>]/secrets/<secret>/versions/<version>[/keys/<key>]`
- Enforces a default-deny atespace→project policy
(`--project-policy-file`), the counterpart of the Kubernetes provider's
namespace policy.
- Returns a retryable 503 only for transient Secret Manager failures,
and caps each read with `--fetch-timeout` (default 3s).
**Repository changes**
- New top-level `plugins/` directory for self-contained plugins,
documented in `docs/dev/code-layout.md` and `AGENTS.md`. The module
imports only substrate's public `pkg/` packages; a test enforces this.
- CI runs the module's tests, `make verify` and golangci-lint.
govulncheck scans the module too, with the action pinned by SHA.
- `docs/egress-credential-injection.md` describes each provider's
credential URI format. The plugin's README covers installing and using
it.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
do not merge, not a draft cause i do want ci running
implements:
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching
tls_passhtrough is a followup.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## e2e: add e2e for egress credential injection
Adds the `egresscredinject` e2e suite: an actor fetches
`https://httpbin.org/headers`
through the sdsmint MITM gateway, and the echoed response proves the
injected
`Authorization` header actually reached the upstream.
### What it asserts
- **Injected**: echoed headers contain `Authorization: Bearer <token>`.
- **Overwritten**: an actor-pre-seeded `Authorization` is replaced by
the injected one.
- **Cleartext skip**: plain-HTTP fetch passes through with no
`Authorization`.
- **Fail closed**: nonexistent secret → 403, unserved provider → 500,
unauthorized namespace → 403.
### Additional changes
- Probe `/fetch` returns the response body, accepts
`header=<name>:<value>`
params, and no longer follows redirects (a cross-scheme redirect would
hop
between the cleartext and TLS legs). Covered by unit tests; egressmitm
rerun green.
- New `e2e.EgressInjectHeader` policy helper.
- New `e2e.DeployCredentialProvider` + `fixtures/credinject`: deploys
the real
provider manifest with a test Secret and namespace-policy ConfigMap,
restarts
the provider so it picks the policy up, cleans up on test end.
- CI: new envoy-lane steps after the MITM lanes, gated on
`E2E_EGRESS_CREDINJECT=1`.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:
1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).
The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.
Partial work for #1871
- [x] Tests pass
```
Run make build-junittool
make build-junittool
bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
shell: /usr/bin/bash -e {0}
env:
E2E_ATENET_DATAPLANE: agentgateway
ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
TOTAL
```
- [x] Appropriate changes to documentation are included in the PR
Removes the "not an officially supported Google product" note from the
top of the README, and replaces the "early development / not ready for
production" wording with a plainer pre-1.0 compatibility statement.
The Google-supported product is
[ai-on-gke/substrate-gke](https://github.com/ai-on-gke/substrate-gke).
The Vulnerability Rewards Program statement is unchanged and stays in
[.github/SECURITY.md](https://github.com/agent-substrate/substrate/blob/main/.github/SECURITY.md);
only its "not eligible" link to the removed README note is dropped.
This change addresses six ways CI pipeline could report false success or
hang on broken infrastructure. Importantly, this change does two things
to reduce load on infrastructure:
1. It prevents jobs from running for [the default action limit of
360m](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax#jobsjob_idtimeout-minutes)
by pinning the timeout to `45m`.
2. It limits pull_request jobs from continuing execution once superseded
by new changes.
- **Silent container test skipping**: Tests now explicitly fail if `CI`
or `REQUIRE_DOCKER` is set (while still skipping on local machines
lacking Docker), with the check implemented in `dockerenv` to avoid an
import cycle between `storetest` and `atepg`.
- **Unbounded trust bundle wait**: Enforced a shared 120-second timeout
across both bundles (overridable via `ATE_INSTALL_TRUST_BUNDLE_TIMEOUT`)
and added diagnostic dumping of bundles, controller pods, and logs
before returning a non-zero exit code.
- **Missing sandbox preflight validation**: Added early preflight checks
for `/dev/kvm` and `SandboxConfig/microvm` that fail fast and print
actionable remediation instructions.
- **Skipped migration checks on main**: Configured the migration
immutability check to run on pushes to main to catch modified migrations
at the point of merge.
- **Missing job timeouts and concurrency limits**: Defined explicit
timeout-minutes (45m and 120m bounds) and added concurrency groups that
automatically cancel superseded pull request runs without canceling runs
on main.
Fixes#1747
- [x] Tests pass
- Silent container skipping, unbounded trust bundles, and missing
sandbox were all forced locally and confirmed to exist with changes here
resolving each.
- The latter half of the scenarios exist in CI only due to being GH
Action trigger issues.
Run agentgateway data plane tests as a part of substrate CI
(non-blocking to start so we can confirm it's not flaky). Also, change
the `--atenet-router` flag to `--atenet-dataplane` to make it clearer
that the flag controls ingress and egress.
I've run the e2es locally across gVisor and microVM plus the MITM
variants for both. The only skip we do for agentgateway is
`TestIngressProtocolDowngrade` because 1. the behavior its testing only
exists on the non-CONNECT atunnel ingress path and agentgateway only
sends CONNECT to atunnel and 2. I'm not sure that we want this to be a
part of the contract that substrate is bound by (e.g. do we really want
to commit to atunnel always parsing HTTP?).
My goal with getting both dataplanes into CI is to start taking steps to
codify the proxy (router + egress PEP) contract for substrate. The
telemetry they emit, atunnel expectations, etc. are all important
contracts to explicitly call out so that they don't become too coupled
to a single dataplane implementation.
> It's a good idea to open an issue first for discussion.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
Migrate from basic markdown template to structured GitHub Issue Form
(.github/ISSUE_TEMPLATE/bug_report.yml). Includes fields for
reproduction steps, sandbox runtime (gVisor/microVM), Substrate version,
Kubernetes environment, and diagnostics/logs.
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.
## Summary
This change replaces startup schema setup with embedded, versioned SQL
migrations.
`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.
Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.
Closes#901.
Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing
## Migration behavior
`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.
One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.
If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.
Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.
## Changes
- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.
## Migration file policy
Migration files use sequential versions and contain exactly one Goose
`Up` section.
CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.
Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.
After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.
Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.
## Compatibility
No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.
Users must recreate databases from earlier PostgreSQL development
builds.
Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.
A binary rollback does not roll back the database schema.
## Testing
Tests cover:
- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
Fixes#368 . Deletes ActorTemplate CRD and any references to it.
Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.
Verifications done:
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
Switches all e2e tests that uses counter demo to validate the new
substrate proto ActorTemplate.
Made some changes to make the e2e test pass:
* cmd/ateapi/internal/controlapi/template_reconciler.go - golden actors
still need the golden atespace, because the logic in suspend actor
relies on it to always take FULL snapshot regardless of the config.
(filed https://github.com/agent-substrate/substrate/issues/1299)
* internal/ateattr/ateattr.go - Added placeholder label for now, will
fix in a follow up PR to wire the metric reporting when using new
ActorTemplate substrate object.
The networking suite's direct-access and arbitrary-port tests build
their actors from the counter fixture and behave the same whichever
egress gateway is deployed, so re-running them after the sdsmint swap
adds CI time without adding coverage. Only egressFixture() reads
E2E_EGRESS_MITM, so restrict both MITM lanes to the tests that use it.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Allow the `apitool validate` linter to define exemptions, then add
exemptions for current validation errors, and then enable this as part
of presubmit (GH action).
## Summary
A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.
- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable
## Benchmarking
Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:
-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing
---------
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
Run the networking suite a second time against the sdsmint (MITM)
egress gateway, so the egress path is covered on both gateway variants
rather than only on passthrough.
A part of #823
From
https://github.com/agent-substrate/substrate/actions/runs/32786529664/job/97619491954?pr=1155,
this change adds about 70 seconds to the `e2e-test` workflow:
- Step `Deploy MITM egress demo`: 8s
- Step `Run E2E tests (networking, MITM egress)`: 33s
- Step `Run E2E tests (networking, MITM egress, micro-VM)`: 30s
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data
source for SystemInfo volumes (#802) and the end-to-end proof that the
projected anchors work against the MITM egress gateway. Live refresh for
running actors (PR 2) and auto-injection (PR 3) come separately.
What this adds
A SystemInfo volume data source that projects the trust anchors of a
named trust bundle to a PEM file:
volumes:
- name: trust
systemInfo:
dataSources:
- trustBundle:
name: egress-mitm.ate.dev
path: egress-ca.pem
Inspired by the Kubernetes clusterTrustBundle projected volume source,
but source-neutral: the template names a bundle; where it's fetched from
is a deployment concern, not part of the API.
Design points
- Resolution lives on the node. The wire carries only {name, path};
atelet resolves the name at write time through an informer-backed lister
on ClusterTrustBundles and writes the sanitized PEM with the temp+rename
discipline from #803 (find-paths safe). Contents refresh on every
Run/Restore. ateapi is not involved, per review discussion — the same
informer is what live refresh (PR 2) will hang off.
- Allowlist in atelet, not the CRD schema. Today only
egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the
ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946)
derives from the egress-mitm-ca-pool Secret. The signer-linked object
name stays a backend detail; the future backend registry (#932) widens
the allowlist without an API change.
- The watch is scoped to the one backing object via a metadata.name
field selector — this informer runs on every node, so an unfiltered
watch would fan every ClusterTrustBundle in the cluster out to every
atelet. RBAC can't express this (resourceNames doesn't apply to
list/watch), so the field selector is the enforcement point.
get/list/watch on clustertrustbundles moves to the atelet ClusterRole.
- No availability probe. The informer registers unconditionally; a
cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1
blocks atelet startup at cache sync, with the reflector errors naming
the missing API (hack/create-kind-cluster.sh enables the gate).
- Fail-closed. Unknown names, missing bundles, and unusable bundles fail
actor start naming the bundle — an actor that declared a trust bundle
must not start without one.
- Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks
only, deduplicated, headers stripped, and anchors deliberately shuffled
so consumers can't grow a dependence on order.
- Schema note: dataSources MaxItems tightened 32→8 while adding the
trustBundle member. Vacuous in practice (the old schema couldn't admit
more than one entry), but flagged since it's ratchet-shaped.
E2E — delivery and consumption
Delivery (identity suite, both sandbox classes): provisions the
egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing
the bundle directly isn't possible — the reconciler reverts hand-edits),
asserts the projected file byte-exact, then rotates the pool across a
suspend/resume to prove refresh-on-restore. Since the probe fixture is
shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists
for whatever suite deploys it.
Consumption (new egressmitm suite, both sandbox classes): deploys the
sdsmint (MITM) egress gateway and proves an actor completes a TLS
handshake with the gateway's per-SNI minted leaf using ONLY the
projected anchors — plus a system-roots negative control that must fail.
The pair is unambiguous in both directions: the positive can't pass
under passthrough (the bundle holds no public CAs), and the negative
can't fail under passthrough.
CI: two steps appended to the existing e2e job after the standard lanes
(the gateway swap is cluster-wide and breaks passthrough assumptions):
--deploy-atenet --experimental-use-sdsmint redeploys only the atenet
components, then the egressmitm suite runs once per sandbox class.
Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each
suite deploys its own copy and drives one actor at a time, so the third
worker per copy was idle memory multiplied across suites on the one-node
CI cluster — pressure that has been killing sandboxes mid-test (runsc:
signal: killed, a vanished ateom socket) on this PR and on main's
identity suite. This reduces the pressure; right-sizing e2e concurrency
or worker-pod QoS cluster-wide is follow-up material.
Not in this PR
- Live refresh for running actors (#932 PR 2) — until then, a running
actor's file is the bundle as of its last Run/Restore, and correctness
rests on overlap rotation by the bundle publisher.
- Auto-injection of the egress trust volume (#932 PR 3).
- Configurable backend registry (#932) — the allowlist is the seam it
will replace.
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.
Validated with focused Go tests, shellcheck, and Kustomize renders.
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
This is one of the reasons that e2e tests are getting queued, and
~doubles the cost to e2e test every PR and merge.
The difference between token and cert e2e branches are small. Keeping
the cert one since it's the default in ate-apiserver.
The step went in preemptively alongside the micro-VM e2e (#327), not in
response to a run running out of space. Deleting `/usr/share/dotnet`,
`/usr/local/lib/android`, `/opt/ghc` and the CodeQL toolcache took
24-88s per matrix leg across the last few main runs, which is pure
wall-clock on every PR.
Easy to put back if a run does hit ENOSPC.
demos/egress is a small Actor that fetches a URL it is given and echoes the
upstream status and body back, which makes the egress path observable from
outside the sandbox. hack/install-demo-egress.sh registers it as a
--deploy-demo-egress fixture and hack/verify-egress-demo.sh drives it and
checks the atenet-egress logs for the corresponding authorized CONNECT.
TestActorEgress in the networking suite covers the same path automatically:
it creates an Actor from the demo template, POSTs a fetch request through
atenet-router, and asserts 200. The suite's actor helper is parameterised by
template so the ingress test keeps using the counter fixture.
When an e2e test failed, the evidence was deleted before anyone could
read it. The suite deleted every namespace it created on the way out,
taking the worker pods with it, and the workflow's post-failure dump
only looked at three fixed namespaces — never the suites' randomly-named
ones. So a failure inside an actor (#619: a micro-VM resume where the
guest died at boot) left nothing behind but the RPC error the test
printed.
Keep the namespaces when the suite failed, and dump every worker pod in
every namespace, so the ateom logs — which carry the guest's console
tail — reach the failed run's output.
Kept namespaces are nobody's to reclaim, and each holds a WorkerPool's
worth of running pods, so they now carry an ate.dev/e2e label and
hack/cleanup-e2e.sh deletes them once the logs have served their
purpose. CI throws its cluster away, but a development cluster
accumulates them run after run.
Fixes #<issue_number_goes_here>
Built while debugging
https://github.com/agent-substrate/substrate/issues/619
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The scan's other blocking finding on this file: with no permissions block a
job gets the repository's default token scopes, which for these jobs is more
than they use.
The org's zizmor scan audits every workflow file a PR touches, and
unpinned-uses is one of its mandatory checks: a mutable tag like @v5 can be
repointed at any commit, so the pin has to be a hash. The tags these were on
are kept as comments.
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.
API surface:
service SessionIdentity -> ActorIdentity
MintJWTRequest.session_id -> actor_id
MintJWTResponse.session_jwt -> actor_jwt
MintCertRequest.session_id -> actor_id
MintCertResponse.session_certificates -> actor_certificates
Go packages:
cmd/ateapi/internal/sessionidentity -> actoridentity
cmd/ateapi/internal/sessionidjwt -> actoridjwt
Flags and cluster resources:
--session-id-jwt-pool -> --actor-id-jwt-pool
--session-id-ca-pool -> --actor-id-ca-pool
Secrets, volumes and mount paths renamed to match, in both
manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
(--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).
Two credential identity values change with the rename:
JWT issuer https://broker.agentic-substrate-session-id-broker.svc
-> https://broker.agentic-substrate-actor-id-broker.svc
SPIFFE ID spiffe://substrate-session.local/app/../session/..
-> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.
BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
Most open-source projects include a `SECURITY.md` so contributors and
users know how to report vulnerabilities without opening a public issue.
This repo currently has no such file, which means reporters have no
obvious path and may default to opening a public issue instead, which
exposes the vulnerability before a fix is ready.
This PR adds a minimal `SECURITY.md` that covers:
- A private reporting path (GitHub Advisory + ate-dev mailing list)
- Severity-based response time targets, honest about the team size
- Supported versions (none yet, main only)
- Scope: what is and is not covered
Part of [#170](https://github.com/agent-substrate/substrate/issues/170)
Establishes mutual TLS between all ate system components, and updates
the certificate plumbing it depends on.
Main changes:
1. The atenet router now verifies ate apiserver' serving certificate,
and presents its client cert to ate apiserver. Previously the connection
used `InsecureSkipVerify`.
2. AteApi server verifies atelet's serving cert.
Minor bug fixes:
1. Prevent `servicednssigner` from signing a cert with no DNS SANs. Also
updated valkey cluster's cert configuration, because it was relying on
the cert with empty DNS.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fix#501
Tests that need root (overlay mounts, mknod, `trusted.*` xattrs, ...)
call `roottest.Require(t, ...)` from
[internal/roottest](internal/roottest) as their first statement. They
skip in a plain `go test ./...`; CI reruns every package whose tests
import that package under `sudo`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
- We no longer need to compile virtiofsd when installing for amd64, we
still need to for arm64 (no upstream release)
- Drop-in compatible, only updating the scripts that manage obtaining
the binaries, and the versions / hashes in the config.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Content-addressed pool of unpacked image layers, shared by every actor on
the node; actor rootfs becomes an overlayfs mount (cached layers as read-
only lowers, bundle-local upper) instead of a full re-untar per run.
atelet (no capabilities) pulls and unpacks; the privileged ateoms finalize
whiteouts and mount. Tag refs resolve via one HEAD and become cacheable;
pull memory is O(stream buffers); the cache survives restarts.
Phase 1 of #463. Fixes#437, #166, #228.
Validated: kind + GKE counter demos (gvisor and microvm), suspend/resume
(oci_unpack ~3ms vs ~15-20s), 411 SWE-bench-scale images pulled and
unpacked with 0 failures, root-gated unit tests for the privileged paths.
Known gaps: no GC yet (Phase 2, see internal/imagecache/README.md);
upgrade ordering — deploy new ateoms before/with the new atelet.
## Description
Closesagent-substrate/substrate#222.
Adds a portable JWT authentication path for ateapi clients while keeping
the
existing mTLS mode as the default.
### What changed
- Added `internal/ateapiauth` with:
- `mtls` and `jwt` auth modes
- server-side JWT gRPC interceptors
- client-side dial options for projected ServiceAccount tokens
- Wired JWT auth into:
- `ate-api-server`
- `ate-controller`
- `atenet router`
- `kubectl-ate` port-forward client path
- Updated Kubernetes JWT verification to support custom HTTP clients for
OIDC/JWKS discovery.
- Added JWT install overlays:
- `manifests/ate-install/jwt`
- `manifests/ate-install/kind-jwt`
- Updated `hack/install-ate.sh` to opt into JWT mode with:
- `ATE_API_AUTH_MODE=jwt`
- `--auth-mode=jwt`
### Notes
- Default install behavior remains `mtls`. I think this deserves a
second look as JWT will support more clusters.
- The JWT overlay projects short-lived ServiceAccount tokens with
audience
`api.ate-system.svc` and mounts the service-DNS trust bundle for ateapi
TLS
verification.
- The current discovery mechanism for JWT mode involves reading the
deployment which I really don't like, but we don't have a "config file"
concept so that bit is a massive TODO.
### Validation
```bash
KIND_CLUSTER_NAME=substrate-jwt ./hack/create-kind-cluster.sh
KO_DOCKER_REPO=localhost:5001 \
ATE_INSTALL_KIND=true \
./hack/install-ate.sh --auth-mode=jwt --deploy-ate-system
Validation results:
kubectl config current-context
# kind-substrate-jwt
kubectl get pods -n ate-system
# all runtime pods Running; init jobs Completed
go run ./cmd/kubectl-ate --context kind-substrate-jwt get actors
# succeeded, empty actor list
```
Free x86-64 ubuntu-latest runners expose /dev/kvm, so the existing kind
e2e job can exercise the micro-VM (kata + cloud-hypervisor) runtime
alongside gVisor on one shared control plane:
- enable /dev/kvm (udev rule) before create-kind-cluster.sh, which then
mounts it and labels the node for the microvm sandbox class;
- deploy the counter-microvm demo (run-microvm-demo-kind.sh) next to the
gVisor counter and run the demo lifecycle suite against both;
- cache the assembled micro-VM assets (keyed on assemble.sh) so the
expensive virtiofsd-from-source build only runs when the pins change;
push-to-main + a weekly schedule keep the cache warm for PRs.
The demo lifecycle suite (create -> suspend -> resume -> in-RAM
continuity) is reused for the micro-VM by parameterizing the source
ActorTemplate via E2E_TEMPLATE_NAMESPACE/NAME (default the gVisor
counter) and the golden-ready wait via E2E_TEMPLATE_READY_TIMEOUT.
PR AI assisted.
v3 is EOL — GitHub Actions runners flag it as too old to run.
v5 is the current stable major (latest release v5.0.1 as of 2026-05).
Two callsites (run-tests and e2e-test jobs), no behavioural change.
Signed-off-by: Davanum Srinivas <davanum@gmail.com>
In github actions:
- We manage kind with a script, no need to separately install at a
possibly incorrect version
- Use go.mod to specify version, instead of repeating the version in
actions config
In the code:
- Upgrade go.mod to use latest go release (which should be reflected in
CI)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
save some time in CI
follow-up https://github.com/agent-substrate/substrate/pull/63
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The existing trigger is pull_request only, so when a PR merges no CI
runs against the resulting main-branch commit. A bad squash, a
post-rebase test failure, or a flake-masked regression isn't surfaced
until the next PR opens, at which point bisecting "is it my change or
main?" wastes contributor time.
Add a push trigger restricted to main so the same run-tests and e2e-test
jobs run against the merged commit. No new jobs, no new permissions, no
change to PR behavior.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Signed-off-by: Davanum Srinivas <davanum@gmail.com>
- Add a trivial /readyz handler (which will only be served once we
finish initializing)
- We use the metrics handler because we already expose an HTTP port
there and there should be no conflict
- Add a readiness probe to the install manifest
- Update install script to wait for readiness on core components
Improves reliability for the demo flow, including the test flake seen in
#2
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This is the initial release of the Agent Substrate.
Agent substrate is a system built on top of Kubernetes which manages agent-like
workloads to achieve higher scale and efficiency than Kubernetes alone can
offer, with lower latency. It builds on top of Kubernetes features like
Pods and Pod autoscaling, but takes the Kubernetes control-plane out of the
critical path to achieve lower latency.
It can run on any Kubernetes cluster and does not inhibit “regular” use of
Kubernetes in any way. Kubernetes provides the infrastructure provisioning and
management for all types of workloads, while Agent Substrate provides
agent-specific scheduling and control.
At its core, Agent Substrate maps a larger set of “actors” (applications such
as agents) onto a smaller set of ready “workers” (Kubernetes Pods), relying on
the fact that agent-like applications tend to be idle most of the time to
achieve heavy multiplexing. It provides functionality to manage an actor’s
lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real
time, and to route incoming traffic to them.
Agent Substrate is intended to be a low-opinion system. The workloads it
manages don't have to be literal AI agents, but those are the best example of
the kind of applications it is designed for. It is not an SDK for building
agents, but rather a system for running them at scale.
Agent Substrate is currently in VERY early development. It is not ready for
production use, and the APIs are almost guaranteed to change. We are not
making any guarantees about backward compatibility at this stage, and
everything in this project may be changed.
Co-authored-by: Alex Bulankou <alexbu@google.com>
Co-authored-by: Benjamin Elder <bentheelder@google.com>
Co-authored-by: Bowei Du <bowei@google.com>
Co-authored-by: Dmitry Berkovich <dberkov@google.com>
Co-authored-by: Fabricio Voznika <fvoznika@google.com>
Co-authored-by: Francisco Cabrera <fclieutier@google.com>
Co-authored-by: Haven Xia <haoyuxia@google.com>
Co-authored-by: Julian Gutierrez Oschmann <juliangut@google.com>
Co-authored-by: Kevin Steuer <ksteuer@google.com>
Co-authored-by: Max Smythe <smythe@google.com>
Co-authored-by: Maya Wang <mymaya@google.com>
Co-authored-by: Michael Taufen <mtaufen@google.com>
Co-authored-by: Shruti Nair <shrutinair@google.com>
Co-authored-by: Taahir Ahmed <taahm@google.com>
Co-authored-by: Tim Hockin <thockin@google.com>
Co-authored-by: Zoe Zhao <zoezhao@google.com>