### Which issue(s) this PR is related to:
Fixes#721
Required for System Upgrade flow (#473)
### What this PR does / why we need it:
This PR implements graceful termination and zero-outage rolling updates
for `atenet-router` , building on the readiness probe and graceful drain
patterns established in #719 (`atelet`).
Draining a two-container networking pod (`atenet-router` Go
control-plane + `envoy` C++ proxy dataplane) introduces complex
lifecycle interdependencies. This PR resolves the dual-container SIGTERM
race, respects Envoy's `failClosed` `ext_proc` filter dependency,
preserves parked requests riding out worker pool saturation, and
accelerates idle deployments via an event-driven file handshake.
#### 1. Event-Driven Dataplane Synchronization (`emptyDir` Marker File)
Envoy fast-exits by default on SIGTERM. To keep Envoy alive while the Go
control-plane orchestrates the drain, we configured an IPC handshake
between containers via a pod-shared `emptyDir` volume mounted at
`/var/run/atenet`:
* **Go Router Container:** Removes any stale marker at startup, then
writes `/var/run/atenet/drain-complete` when its shutdown sequence
finishes.
* **Envoy Container `preStop` Hook:** Runs `while [ ! -f
/var/run/atenet/drain-complete ]; do sleep 0.5; done`.
**Outcome:** Envoy exits as soon as — the drain is done, rather than
wasting time in a fixed worst-case sleep. If the router crashes, Kubelet
terminates the `preStop` hook at `terminationGracePeriodSeconds:
60`—slower cleanup, never a wedge.
#### 2. The Multi-Container Shutdown Sequence (`drain.go`)
Because Kubernetes issues SIGTERM to both containers at once, a
coordination state machine ensures Envoy never drops a connection and
`ext_proc` is never stopped prematurely:
| Phase / (best-case ex.) Timeline | atenet-router | envoy | K8s /
Service Status |
| :--- | :--- | :--- | :--- |
| **SIGTERM Sent**<br>($t=0\text{s}$) | Catches SIGTERM<br>• Flips
`/readyz` $\rightarrow$ `503`.<br>• Starts 13s `drain-delay`. | Enters
`lifecycle.preStop` hook:<br>`while [ ! -f .../drain-complete ]; do
sleep 0.5; done`<br>• Kubelet holds SIGTERM back. | Deployment will
recreate the pod.<br>EndpointSlice controller begins dropping Old Pod
IP. |
| **Propagation**<br>($t=0\text{s}$ – $t=13\text{s}$) | Keeps serving
normally.<br>• `ext_proc` continues unparking/routing requests. | Runs
normally inside `preStop` loop.<br>• Serves active TCP connections. |
Service endpoint removal completes.<br>No new connections arrive at Old
Pod. |
| **Envoy drain**<br>($t=13\text{s}$) | Issues Envoy admin API
calls:<br>• `/healthcheck/fail`<br>•
`/drain_listeners?graceful&skip_exit`<br>• Polls `/stats` for active
downstream connections. | Begins graceful listener drain:<br>• GOAWAY /
`Connection: close` on established connections.<br>• In-flight requests
keep running. | In-flight HTTP requests finish executing through Envoy.
|
| **ext_proc drain**<br>($t\approx18\text{s}$) | Active Envoy
connections hit 0 (the poll exits early).<br>• Calls
`extproc.GracefulStop()`.<br>• Parked request streams finish. | All
downstream connections closed.<br>• Still waiting in the `preStop` loop.
| All client HTTP responses delivered. |
| **Handshake & Exit**<br>($t=18.5\text{s}$)| `ext_proc` drain
finishes.<br>• Writes `drain-complete` marker.<br>• Hard-stops xDS
(`Stop()`) & exits. | `preStop` loop detects marker file!<br>• `preStop`
exits 0.<br>• Kubelet sends SIGTERM to Envoy.<br>• Envoy exits
immediately. | Old Pod deleted cleanly at ~18.5s (well under the 60s
budget). |
Envoy ref:
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining
#### Testing & Verification
1. Unit Tests (`drain_test.go`): Added comprehensive unit tests
covering:
- Graceful completion of in-flight `ext_proc` streams within deadline.
- Force-stopping straggler streams past `--drain-timeout`.
- Envoy admin API interaction and connection polling.
- Stale marker cleanup and marker file creation.
2. Local testing on kind cluster (Observations):
Test script: `hack/verify-atenet-drain.sh`
- Steady state: `/readyz=200`, `/healthz=200`; the instant the pod
turned Terminating: `/readyz=503` while `/healthz=200` — `NotReady` but
alive, for the whole drain.
- The parked request survived the shutdown: fired at a busy 1-worker
pool → parked → pod deleted while parked → worker freed → HTTP=200,
total=1.89s, body hello from: 169.254.17.2 | preserved memory count: 1 |
preserved file counter: 1 — park → resume → route, served by the
Terminating pod during its drain-delay window.
- Pod terminated 14s after deletion — the idle floor, confirmed across
three runs: >13s drain-delay (sequence ran), ≪60s grace (marker
handshake released Envoy's preStop; no `SIGKILL`).
- Log sequence captured verbatim, drain-delay honored to the millisecond
(17:34:41.706 → 17:34:54.707):
>Shutdown signal received; draining
>(+13.000s) Draining Envoy {window: 15s}
>Envoy drained
>Starting ext_proc drain
>ext_proc drain completed within deadline
>Drain-complete marker written {path: /var/run/atenet/drain-complete}
>Shutdown complete
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
`setup-envtest use` with no version re-resolves "latest" against the
remote release index on every run, so the envtest-backed packages needed
network even with fully cached binaries.
This PR pins the control plane to `1.36.x`, and fold the three
per-package `envtest` into `internal/testenv.Start`, which also honors
`-short`: `go test -short ./...` skips the envtest-backed packages and
is the guaranteed-offline path.
Have Claude to help me fix it
Fixes#789
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Currently we are unable to easily automate the benchmarking of micro
VMs.
This PR:
* Refactors the microVM setup scripts to allow discrete installation of
uVM sanbox configs separate from counter demo
* Pipes a MicroVM container test config through the benchmarking system
* Updates default atelet daemonset to tolerate all kinds of sandboxClass
taints
OTEL_EXPORTER_OTLP_ENDPOINT is written in 9 places: hardcoded inline in
4 base workload manifests, then re-patched in 5 spots across the kind
overlays. Changing the collector address means editing all 9 and knowing
which install path renders which.
Replace them with one checked-in ConfigMap per environment, consumed by
every component via envFrom:
manifests/ate-install/ate-otel-config.yaml (GKE)
manifests/ate-install/kind/ate-otel-config.yaml (kind, same name)
The GKE copy is listed in manifests/ate-install/base, so token-client,
agentgateway and agentgateway-token-client all inherit it; the kind
overlay lists its own copy of the same ConfigMap name. No overlay builds
on both, so the two never collide.
A kustomize-only fix does not work here. The GKE path applies the base
directory raw and hack/install-ate.sh's targeted redeploys apply single
files with no Kustomize, so the mechanism has to survive `kubectl apply
-f <one-file>` -- which rules out configMapGenerator (hash-suffixed
names) and replacements. envFrom on a stable name reaches every path.
deploy_ate_system applies the ConfigMap before the rendered bundle, the
same way it already applies the namespace. A container whose envFrom
target is missing does not start, and a raw directory apply is ordered
by filename, so ate-api-server.yaml and ate-controller.yaml would
otherwise be created first and sit in CreateContainerConfigError until
the ConfigMap caught up.
This also fixes a latent bug in the targeted redeploys.
deploy_ate_apiserver, deploy_atelet and deploy_atenet apply raw files
even in kind mode, silently reverting the endpoint to the GKE value.
They now apply the environment's ConfigMap via a new apply_otel_config
helper, which picks the file by ATE_INSTALL_KIND rather than applying
the base copy unconditionally -- applying the base one on kind would
break telemetry for every component at once.
atenet-router needs no special handling: since a98f85b1 its
--otlp-collector-address defaults to $OTEL_EXPORTER_OTLP_ENDPOINT, so
Envoy's own spans follow the ConfigMap along with everything else.
OTEL_TRACES_SAMPLER deliberately stays an inline kind patch on
ate-api-server rather than moving into the ConfigMap. The ConfigMap is
shared by every component via envFrom, so putting parentbased_always_on
there would pin the router and the rest of the control plane to 100% and
undo the per-component ratios from 15eecd02.
Note that a ConfigMap edit does not roll the consuming pods the way an
inline env change did, since the pod template is unchanged. Callers must
follow a change with `kubectl rollout restart`. This is documented in
both ConfigMaps, in docs/observability.md, and in the tracing best
practices, which previously told authors to hardcode the variable.
Fixes#745
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#697
## Summary
`deploy_ate_system` gated `manifests/ate-install/generated/` behind
`ensure_crds`, whose existence check skips the directory on any upgrade
— stranding CRD schemas and, since `role.yaml` has no other apply path,
all ClusterRoles at first-install state. A controller image needing a
new permission then deadlocks on informer start while the rollout
reports success.
`deploy_ate_system` now calls `deploy_crds` unconditionally (`kubectl
apply` is idempotent). The demo scripts and per-component deploy flags
keep `ensure_crds`, where "make sure they exist" is the intended
semantic.
## Test plan
- Reproduced on a 4-day-old kind cluster: upgrading the control plane
deadlocked ate-controller on `cannot list networkpolicies`; manually
applying `role.yaml` unblocked it.
- With this fix, the same `install-ate-kind.sh --deploy-ate-system` run
prints the `deploy_crds` step and re-applied the drifted manifests
(`actortemplates.ate.dev configured`, `workerpools.ate.dev configured`,
ClusterRoles applied); cluster reconciles normally.
- `bash -n hack/install-ate.sh`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
Agentgateway is a popular [AAIF](https://aaif.io/) project designed for
agentic systems. It is of particular interest for substrate users as it
was designed specifically to serve the role of an egress proxy for
agents and other AI-adjacent workloads, with strong support for
credential injection, token exchanges, TLS MITM, and other
authentication and authorization schemes.
Agentgateway supports the ext_proc protocol, like Envoy, so we use that
here. There has been some debate in the community around the long term
end state of ext_proc vs other options, but for now this maintains the
status quo.
While Agentgateway can be configured dynamically over XDS, it does not
take the same configuration as Envoy. However, none of the config is
dynamic anyways, so we use just a static configuration file (configmap)
for now to keep things simple.
While we don't have specific per-component/file OWNERS in the project at
the moment, I can informally commit to myself and Eitan maintaining this
integration. Given the goals around velocity in the project we can
commit to either rapidly fixing any issues that may arise or
removing/disabling the integration if this ends up slowing things down.
When an e2e test failed, the evidence was deleted before anyone could
read it. The suite deleted every namespace it created on the way out,
taking the worker pods with it, and the workflow's post-failure dump
only looked at three fixed namespaces — never the suites' randomly-named
ones. So a failure inside an actor (#619: a micro-VM resume where the
guest died at boot) left nothing behind but the RPC error the test
printed.
Keep the namespaces when the suite failed, and dump every worker pod in
every namespace, so the ateom logs — which carry the guest's console
tail — reach the failed run's output.
Kept namespaces are nobody's to reclaim, and each holds a WorkerPool's
worth of running pods, so they now carry an ate.dev/e2e label and
hack/cleanup-e2e.sh deletes them once the logs have served their
purpose. CI throws its cluster away, but a development cluster
accumulates them run after run.
Fixes #<issue_number_goes_here>
Built while debugging
https://github.com/agent-substrate/substrate/issues/619
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.
API surface:
service SessionIdentity -> ActorIdentity
MintJWTRequest.session_id -> actor_id
MintJWTResponse.session_jwt -> actor_jwt
MintCertRequest.session_id -> actor_id
MintCertResponse.session_certificates -> actor_certificates
Go packages:
cmd/ateapi/internal/sessionidentity -> actoridentity
cmd/ateapi/internal/sessionidjwt -> actoridjwt
Flags and cluster resources:
--session-id-jwt-pool -> --actor-id-jwt-pool
--session-id-ca-pool -> --actor-id-ca-pool
Secrets, volumes and mount paths renamed to match, in both
manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
(--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).
Two credential identity values change with the rename:
JWT issuer https://broker.agentic-substrate-session-id-broker.svc
-> https://broker.agentic-substrate-actor-id-broker.svc
SPIFFE ID spiffe://substrate-session.local/app/../session/..
-> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.
BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
This introduces a new demo (`demos/autoscaled-workerpool`) demonstrating
how to dynamically autoscale a WorkerPool using HPA +
`prometheus-adapter` on a local Kind cluster.
As mentioned in https://github.com/agent-substrate/substrate/issues/198,
this approach may be too slow for some use cases, but this is a good
first milestone.
Plus, there are some limitations with scaling workerpools that still
need to be addressed, for example:
- Scale-down can strand paused actors. When a worker pod is removed, a
local-snapshot paused actor keeps its node pin pointing at a node that
may no longer have (free) worker pods.
- Scale-up gives no locality guarantee: If there are multiple
local-snapshot paused actors that can't resume because no worker pods
are available on their node, upscaling will not unblock them: the added
capacity can land anywhere, so the pool can grow without ever placing a
worker where a stranded actor needs it.
- Scale-down picks victims blind to actor assignment (can kill busy
pods).
- [x] Tests pass: n/a
- [x] Appropriate changes to documentation are included in the PR
run-e2e.sh always fail on my mac with ` line 106: go_test_args[@]:
unbound variable,exit`, asked Claude code and fix related Bash 3.2
problem in scripts.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
An oversubscribed WorkerPool (2 workers, several actors) that exercises the router parking path: requests to a saturated pool park and retry instead of failing fast, and are served once capacity frees up. Includes load.sh and --deploy-demo-parking / --delete-demo-parking wiring in hack/install-ate.sh.
This demo is broken now that we require authn to ate-apiserver. The
fix is not trivial as it requires the actor to be able to authenticate, plus
it doesn't really adds much.
The Deployment in `atenet-dns.yaml` is named `dns`; every other resource
in that file is `atenet-dns`. The installer waited on the filename
rather than the actual Deployment, so this step failed with NotFound on
every otherwise-successful deploy.
Part of [#170](https://github.com/agent-substrate/substrate/issues/170)
Establishes mutual TLS between all ate system components, and updates
the certificate plumbing it depends on.
Main changes:
1. The atenet router now verifies ate apiserver' serving certificate,
and presents its client cert to ate apiserver. Previously the connection
used `InsecureSkipVerify`.
2. AteApi server verifies atelet's serving cert.
Minor bug fixes:
1. Prevent `servicednssigner` from signing a cert with no DNS SANs. Also
updated valkey cluster's cert configuration, because it was relying on
the cert with empty DNS.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Initial part of #232.
Extends the ActorTemplate API to add support for per-actor external
volumes and adds control plane and atelet hooks following the Actor
lifecycle.
Callouts to volume operations are abstracted with a Volume interface.
Right now, only a mock volume plugin for testing has been implemented,
but we will add CSI support next.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR - Not
going to update documentation until we add CSI support.
Fix#501
Tests that need root (overlay mounts, mknod, `trusted.*` xattrs, ...)
call `roottest.Require(t, ...)` from
[internal/roottest](internal/roottest) as their first statement. They
skip in a plain `go test ./...`; CI reruns every package whose tests
import that package under `sudo`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
- We no longer need to compile virtiofsd when installing for amd64, we
still need to for arm64 (no upstream release)
- Drop-in compatible, only updating the scripts that manage obtaining
the binaries, and the versions / hashes in the config.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
On macOS, bash default has Bash 3.2 compatibility. mapfile isn't
available and the mapfile command is breaking the install-ate.sh. Use a
while loop to implement the exact behavior.
Best practices is not to put a -deployment deployment, -role role, etc;
the kind already declares the kind we don't need it in the name as well.
Additionally we did not do it consistently anyways.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Update README.md and scripts to adopt atespaces for existing demos.
It's intentional to let user choose atespace in demos since actors are
also created by them.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Description
Closesagent-substrate/substrate#222.
Adds a portable JWT authentication path for ateapi clients while keeping
the
existing mTLS mode as the default.
### What changed
- Added `internal/ateapiauth` with:
- `mtls` and `jwt` auth modes
- server-side JWT gRPC interceptors
- client-side dial options for projected ServiceAccount tokens
- Wired JWT auth into:
- `ate-api-server`
- `ate-controller`
- `atenet router`
- `kubectl-ate` port-forward client path
- Updated Kubernetes JWT verification to support custom HTTP clients for
OIDC/JWKS discovery.
- Added JWT install overlays:
- `manifests/ate-install/jwt`
- `manifests/ate-install/kind-jwt`
- Updated `hack/install-ate.sh` to opt into JWT mode with:
- `ATE_API_AUTH_MODE=jwt`
- `--auth-mode=jwt`
### Notes
- Default install behavior remains `mtls`. I think this deserves a
second look as JWT will support more clusters.
- The JWT overlay projects short-lived ServiceAccount tokens with
audience
`api.ate-system.svc` and mounts the service-DNS trust bundle for ateapi
TLS
verification.
- The current discovery mechanism for JWT mode involves reading the
deployment which I really don't like, but we don't have a "config file"
concept so that bit is a massive TODO.
### Validation
```bash
KIND_CLUSTER_NAME=substrate-jwt ./hack/create-kind-cluster.sh
KO_DOCKER_REPO=localhost:5001 \
ATE_INSTALL_KIND=true \
./hack/install-ate.sh --auth-mode=jwt --deploy-ate-system
Validation results:
kubectl config current-context
# kind-substrate-jwt
kubectl get pods -n ate-system
# all runtime pods Running; init jobs Completed
go run ./cmd/kubectl-ate --context kind-substrate-jwt get actors
# succeeded, empty actor list
```
Follow-up to https://github.com/agent-substrate/substrate/pull/287 /
#123
- align snapshot behavior: use a tmpfs (for writes) on top of read-only
viritio-fs mount for the container image rootfs instead of ext4 images
- enable multiple container support
- cleanup stale comments
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Assemble + stage the micro-VM runtime assets, an ateom-base image (debian-slim +
e2fsprogs for mkfs.ext4), and run-microvm-demo.sh to build + deploy the
counter-microvm demo end to end (overriding the worker base via KO_CONFIG_PATH so
no committed file is edited). Document the micro-VM sandbox class.
go-licenses only covers module dependencies and treats our own module —
including any copied-in / forked third-party SOURCE under in-tree
third_party/ dirs — as first-party, so those upstream LICENSE/NOTICE/
COPYING files are never collected.
Mirror them into LICENSES/third_party/, keyed by their path under
third_party/, the same way kubernetes' hack/update-vendor-licenses.sh
does. `git ls-files` is used so gitignored paths are skipped; vendor/ and
the output dir are excluded since they hold module deps / generated
output rather than forked source.
This is a no-op today (there is no in-tree third_party source yet) and
prepares the licenses tooling for forthcoming vendored third-party code.