Most open-source projects include a `SECURITY.md` so contributors and
users know how to report vulnerabilities without opening a public issue.
This repo currently has no such file, which means reporters have no
obvious path and may default to opening a public issue instead, which
exposes the vulnerability before a fix is ready.
This PR adds a minimal `SECURITY.md` that covers:
- A private reporting path (GitHub Advisory + ate-dev mailing list)
- Severity-based response time targets, honest about the team size
- Supported versions (none yet, main only)
- Scope: what is and is not covered
Part of [#170](https://github.com/agent-substrate/substrate/issues/170)
Establishes mutual TLS between all ate system components, and updates
the certificate plumbing it depends on.
Main changes:
1. The atenet router now verifies ate apiserver' serving certificate,
and presents its client cert to ate apiserver. Previously the connection
used `InsecureSkipVerify`.
2. AteApi server verifies atelet's serving cert.
Minor bug fixes:
1. Prevent `servicednssigner` from signing a cert with no DNS SANs. Also
updated valkey cluster's cert configuration, because it was relying on
the cert with empty DNS.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fix#501
Tests that need root (overlay mounts, mknod, `trusted.*` xattrs, ...)
call `roottest.Require(t, ...)` from
[internal/roottest](internal/roottest) as their first statement. They
skip in a plain `go test ./...`; CI reruns every package whose tests
import that package under `sudo`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
- We no longer need to compile virtiofsd when installing for amd64, we
still need to for arm64 (no upstream release)
- Drop-in compatible, only updating the scripts that manage obtaining
the binaries, and the versions / hashes in the config.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Content-addressed pool of unpacked image layers, shared by every actor on
the node; actor rootfs becomes an overlayfs mount (cached layers as read-
only lowers, bundle-local upper) instead of a full re-untar per run.
atelet (no capabilities) pulls and unpacks; the privileged ateoms finalize
whiteouts and mount. Tag refs resolve via one HEAD and become cacheable;
pull memory is O(stream buffers); the cache survives restarts.
Phase 1 of #463. Fixes#437, #166, #228.
Validated: kind + GKE counter demos (gvisor and microvm), suspend/resume
(oci_unpack ~3ms vs ~15-20s), 411 SWE-bench-scale images pulled and
unpacked with 0 failures, root-gated unit tests for the privileged paths.
Known gaps: no GC yet (Phase 2, see internal/imagecache/README.md);
upgrade ordering — deploy new ateoms before/with the new atelet.
## Description
Closesagent-substrate/substrate#222.
Adds a portable JWT authentication path for ateapi clients while keeping
the
existing mTLS mode as the default.
### What changed
- Added `internal/ateapiauth` with:
- `mtls` and `jwt` auth modes
- server-side JWT gRPC interceptors
- client-side dial options for projected ServiceAccount tokens
- Wired JWT auth into:
- `ate-api-server`
- `ate-controller`
- `atenet router`
- `kubectl-ate` port-forward client path
- Updated Kubernetes JWT verification to support custom HTTP clients for
OIDC/JWKS discovery.
- Added JWT install overlays:
- `manifests/ate-install/jwt`
- `manifests/ate-install/kind-jwt`
- Updated `hack/install-ate.sh` to opt into JWT mode with:
- `ATE_API_AUTH_MODE=jwt`
- `--auth-mode=jwt`
### Notes
- Default install behavior remains `mtls`. I think this deserves a
second look as JWT will support more clusters.
- The JWT overlay projects short-lived ServiceAccount tokens with
audience
`api.ate-system.svc` and mounts the service-DNS trust bundle for ateapi
TLS
verification.
- The current discovery mechanism for JWT mode involves reading the
deployment which I really don't like, but we don't have a "config file"
concept so that bit is a massive TODO.
### Validation
```bash
KIND_CLUSTER_NAME=substrate-jwt ./hack/create-kind-cluster.sh
KO_DOCKER_REPO=localhost:5001 \
ATE_INSTALL_KIND=true \
./hack/install-ate.sh --auth-mode=jwt --deploy-ate-system
Validation results:
kubectl config current-context
# kind-substrate-jwt
kubectl get pods -n ate-system
# all runtime pods Running; init jobs Completed
go run ./cmd/kubectl-ate --context kind-substrate-jwt get actors
# succeeded, empty actor list
```
Free x86-64 ubuntu-latest runners expose /dev/kvm, so the existing kind
e2e job can exercise the micro-VM (kata + cloud-hypervisor) runtime
alongside gVisor on one shared control plane:
- enable /dev/kvm (udev rule) before create-kind-cluster.sh, which then
mounts it and labels the node for the microvm sandbox class;
- deploy the counter-microvm demo (run-microvm-demo-kind.sh) next to the
gVisor counter and run the demo lifecycle suite against both;
- cache the assembled micro-VM assets (keyed on assemble.sh) so the
expensive virtiofsd-from-source build only runs when the pins change;
push-to-main + a weekly schedule keep the cache warm for PRs.
The demo lifecycle suite (create -> suspend -> resume -> in-RAM
continuity) is reused for the micro-VM by parameterizing the source
ActorTemplate via E2E_TEMPLATE_NAMESPACE/NAME (default the gVisor
counter) and the golden-ready wait via E2E_TEMPLATE_READY_TIMEOUT.
PR AI assisted.
v3 is EOL — GitHub Actions runners flag it as too old to run.
v5 is the current stable major (latest release v5.0.1 as of 2026-05).
Two callsites (run-tests and e2e-test jobs), no behavioural change.
Signed-off-by: Davanum Srinivas <davanum@gmail.com>
In github actions:
- We manage kind with a script, no need to separately install at a
possibly incorrect version
- Use go.mod to specify version, instead of repeating the version in
actions config
In the code:
- Upgrade go.mod to use latest go release (which should be reflected in
CI)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
save some time in CI
follow-up https://github.com/agent-substrate/substrate/pull/63
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The existing trigger is pull_request only, so when a PR merges no CI
runs against the resulting main-branch commit. A bad squash, a
post-rebase test failure, or a flake-masked regression isn't surfaced
until the next PR opens, at which point bisecting "is it my change or
main?" wastes contributor time.
Add a push trigger restricted to main so the same run-tests and e2e-test
jobs run against the merged commit. No new jobs, no new permissions, no
change to PR behavior.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Signed-off-by: Davanum Srinivas <davanum@gmail.com>
- Add a trivial /readyz handler (which will only be served once we
finish initializing)
- We use the metrics handler because we already expose an HTTP port
there and there should be no conflict
- Add a readiness probe to the install manifest
- Update install script to wait for readiness on core components
Improves reliability for the demo flow, including the test flake seen in
#2
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This is the initial release of the Agent Substrate.
Agent substrate is a system built on top of Kubernetes which manages agent-like
workloads to achieve higher scale and efficiency than Kubernetes alone can
offer, with lower latency. It builds on top of Kubernetes features like
Pods and Pod autoscaling, but takes the Kubernetes control-plane out of the
critical path to achieve lower latency.
It can run on any Kubernetes cluster and does not inhibit “regular” use of
Kubernetes in any way. Kubernetes provides the infrastructure provisioning and
management for all types of workloads, while Agent Substrate provides
agent-specific scheduling and control.
At its core, Agent Substrate maps a larger set of “actors” (applications such
as agents) onto a smaller set of ready “workers” (Kubernetes Pods), relying on
the fact that agent-like applications tend to be idle most of the time to
achieve heavy multiplexing. It provides functionality to manage an actor’s
lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real
time, and to route incoming traffic to them.
Agent Substrate is intended to be a low-opinion system. The workloads it
manages don't have to be literal AI agents, but those are the best example of
the kind of applications it is designed for. It is not an SDK for building
agents, but rather a system for running them at scale.
Agent Substrate is currently in VERY early development. It is not ready for
production use, and the APIs are almost guaranteed to change. We are not
making any guarantees about backward compatibility at this stage, and
everything in this project may be changed.
Co-authored-by: Alex Bulankou <alexbu@google.com>
Co-authored-by: Benjamin Elder <bentheelder@google.com>
Co-authored-by: Bowei Du <bowei@google.com>
Co-authored-by: Dmitry Berkovich <dberkov@google.com>
Co-authored-by: Fabricio Voznika <fvoznika@google.com>
Co-authored-by: Francisco Cabrera <fclieutier@google.com>
Co-authored-by: Haven Xia <haoyuxia@google.com>
Co-authored-by: Julian Gutierrez Oschmann <juliangut@google.com>
Co-authored-by: Kevin Steuer <ksteuer@google.com>
Co-authored-by: Max Smythe <smythe@google.com>
Co-authored-by: Maya Wang <mymaya@google.com>
Co-authored-by: Michael Taufen <mtaufen@google.com>
Co-authored-by: Shruti Nair <shrutinair@google.com>
Co-authored-by: Taahir Ahmed <taahm@google.com>
Co-authored-by: Tim Hockin <thockin@google.com>
Co-authored-by: Zoe Zhao <zoezhao@google.com>