176 Commits
Author SHA1 Message Date
Lior Lieberman 944096fa21 Passthrough dial address resolved by the gateway (#2045)
The passthrough chain from #2019 dialed the ORIGINAL_DST filter state,
which the CONNECT leg filled with the address the actor connected to.
The SNI picked the chain, the actor's own resolution picked the
destination, so a passthrough rule for one name let an actor reach any
IP by claiming that name in the ClientHello.

We fix it by:

- The passthrough chain is now `sni_dynamic_forward_proxy` then
`tcp_proxy` to a new raw forward-proxy cluster,
`egress_forward_proxy_passthrough`, on the shared egress_dns_cache. The
gateway resolves the SNI itself and sends the bytes to it.

- The CONNECT leg answers the dialed port under
dev.ate.egress:dialed_port; the outer chain copies it into
envoy.upstream.dynamic_port, shared with the inner listener, which both
the SNI filter and the cluster read before their configured port.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-10-01 18:50:20 +00:00
sfunkenhauser 445e1b384b Delete worker pods in terminal states (#2051)
Fixes #1935

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-10-01 14:45:19 +00:00
Lior Lieberman 13e8efff25 envoy dyn module followup (#1995)
do not merge, not a draft cause i do want ci running

implements: 
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching


tls_passhtrough is a followup. 



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-30 16:57:00 +00:00
Lior Lieberman 1240749483 ateapi: rename --egress-gateway-address to --default-egress-gateway-address (#1858)
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).

~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA

EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.

Related #1591


> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

/cc @bowei @EItanya
2026-09-30 15:46:52 +00:00
Davanum Srinivas afb62e18fb atelet: add CAP_DAC_OVERRIDE to reset dirs a non-root container wrote (#1910)
Follow-up to #1906. Removing a directory entry needs write permission on
the directory that holds it, and with all capabilities dropped uid 0
gets no exemption. A non-root container can create a subdirectory in its
`0777` durable dir; that subdirectory belongs to the container's uid,
typically `0755`, so root cannot empty it. `resetActorDirs` fails after
every checkpoint of such an actor and the suspend never completes. Chmod
first, as the bundle dir does, is no way out: chmod needs ownership or
`CAP_FOWNER`.

Why atelet and not ateom, which already holds the capability: an ateom
that dies before cleanup leaves the files behind and moves the hang to
the next resume on that node. atelet owns the actor directories and
needs it on the crash path either way.

Manifest only, so no unit test. Exercised on kind with the micro-VM
class: uid 65532 writes a durable dir, then suspend and resume. Without
the capability the same run hangs in `SUSPENDING`.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
2026-09-30 02:20:11 +00:00
yanavlasov 0f4572495a Egress policy plumbing (#1978)
Plumbing of actor's EgressPolicy into Envoy dataplane. This is the first
commit that only implements SNI allowlist, without wildcards. Followup
PRs will implement there rest of the policy:

* Wildcard matching.
* SNI passthrough policy.
* Allowing plaintext based on presence of `http` rules.
* Port verification.
* Making sdsmint the default and removing non-sdsmin config.

----

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Yan Avlasov <yavlasov@google.com>
2026-09-29 17:54:27 +00:00
Davanum Srinivas c7dbe9d672 nodepath: move the node state root to /var/lib/ate (#1926)
Fixes #1911

`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.

No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 18:27:55 +00:00
Lior Lieberman 6621f2b4b4 egress policy enhancements for GA (#1751)
**BREAKING** API changes prior in preparation for GA. 

~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow


- `EgressRule` is now a union of protocols: `http`, `https`,
  `tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
  Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
  listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
  request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
  `tls_passthrough`. Same wildcard grammar as before, documented once on
  `HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
  once on `HTTPRule.ports`.

- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
  placeholder value. Requests without the header pass unchanged.

- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.

- `google/protobuf/empty.proto` import dropped; nothing uses it now.

Questions

- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
  block is intercepted. is that clear enough?

follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
  the rule conflict check) and regenerate `zz_generated.validation.go`



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Fixes: #823 (there are probably more issues that i need to find)
2026-09-28 17:30:07 +00:00
Eitan Yarmush ed6d2a1fc8 Update agentgateway for split actor and ateom identities (#1923)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.

This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.

Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.

Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.

- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-26 23:02:14 +00:00
Max Smythe 8b5d9ff0ef Tolerations for Workload Types (#1874)
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-09-25 16:45:02 +00:00
Max Smythe d6d2a0fa1d Add the option to install a large-cluster manifest and cordon control plane to dedicated machines (#1632)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.

It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 23:39:01 +00:00
Tim Bai 93e691d9b1 serverboot: give the ateoms a LoggerProvider over the relay (#1857)
Step 2 of #1748, plus the consumer rule from step 1b.

`LoggingOptions` gains `ExporterConn` and `RelayCapable`, mirroring
`TracingOptions`. Both ateoms initialize logging after metrics with the
relay connection and defer a nil-guarded shutdown. `OTEL_LOGS_EXPORTER`
still defaults to `none`, so no ateom emits a record; this is the
provider the actor events land on when emission moves.

A relay-capable log resource carries `ate.otlp.relay` like traces and
metrics do. `newLoggerProvider` is split from `InitLogging` so the test
asserts that on an emitted record's resource, the way
`TestMeterProviderRelayAttribute` does for metrics.

`docs/observability.md` states the consumer rule that replaced the relay
denylist dropped in #1800: a lifecycle record is authoritative only
under ateapi's resource, because the relay admits only ateom resources.

Follow-up for the emission step, not here: the controller does not
inject `OTEL_LOGS_EXPORTER` into worker pods, so no environment
exercises log-over-relay yet.
2026-09-24 20:40:19 +00:00
yanavlasov 30e6d33daa Initial commit of Envoy Dynamic Modules (#1535)
Collecting initial feedback on integrating Envoy dynamic modules for
Substrate dataplane.

Envoy Dynamic Modules allow fast iteration and experimentation on
Substrate dataplane. Dynamic modules are used to improve feature
velocity. As requirements and solution are better understood they will
be generalized and moved into Envoy's main repository.

Important points:

- Dynamic modules introduce Rust toolchain into Substrate dev
environment. This PR does not add GH actions for testing Dynamic Modules
during CI builds. This will be added in a subsequent PR.
- This change brings in Dockerfiles and docker buildx. It produces an
image with a versioned Envoy binary (1.39) and accompanying dynamic
modules.
- Docker image registry URL in deployment spec for Envoy datplane is at
this point hardcoded to `localhost:5001`. This will break for anyone
using a different registry. Fixing this will require using envsubst or
equivalent. I'm open to other suggestions.


- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Yan Avlasov <yavlasov@google.com>
2026-09-24 17:18:48 +00:00
Yufan Su 32df553276 Add statusz page in k8s secret credential provider (#1811)
Follow-up change on a comment in #1335 to add a statusz page in the
credential provider service.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-23 18:45:57 +00:00
Huy Pham 92a84388b4 microvm: bump kata assets to 4.1.0 (#1708)
Kata 4.1.0 bundles virtiofsd 1.14.0 (required by Substrate). This
simplifies the dev process on arm64 because it removes the need to build
virtiofsd from source.

Verified on an arm64 KVM host (Lima + kind).

Fixes #1695

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-22 16:07:12 +00:00
yufan-su 57234a866a egress: add an example credential provider which reads from k8s secret (#1335)
Add an example gRPC credential-provider plugin which reads from k8s
secrets for the egress
credential-injection path. 

 - It resolves `ate-secret://k8s.io/default/<namespace>/<secret>/<key>`
URIs to Kubernetes Secret values, so Substrate never stores secrets — it
only
brokers a read that the provider is authorized to perform. 
 - It is the only
component in the injection path with Kubernetes secerts access; the
egress gateway and
its injector never read Secrets directly.

### What's included

- **`cmd/credential-provider/kubernetes-secrets`** — the gRPC service:
- Parses and validates `ate-secret://` URIs, resolving the requested
Secret
    (with single-key fallback when the URI omits a key).
- Enforces an **atespace→namespace authorization policy**
(default-deny),
    derived from the caller's attested actor SPIFFE ID.
- Serves over **mutual TLS**, requiring the caller's client cert to
chain to
    the trust bundle *and* carry the egress injector's SPIFFE SAN.
- **Manifests** (`manifests/egress-credential-injection/`) — Deployment,
Service, ServiceAccount + RBAC, a sample namespace-policy ConfigMap, and
a
  sample Secret.
- Renames the default provider address/service from `credprovider` to
`k8s-credential-provider` across `ate-setup`, install scripts, and the
egress
  injection overlay.

### atespace → namespace authorization

Beyond the mTLS check that only the egress injector may call the
provider, each
request is authorized against a **default-deny atespace→namespace
policy**. The
provider derives the requesting actor's atespace from its attested
SPIFFE ID and
resolves a Secret only if that atespace is explicitly granted access to
the URI's
namespace.
2026-09-22 10:13:25 +00:00
Huy Pham 8062bbafb4 microvm: drop the kata-config asset (#1704)
Today, `ateom` fetches the `kata-config` asset (configuration-clh.toml)
to read only 3 values `default_memory`, `default_vcpus` and
`kernel_params`. The first 2 are redundant because (1) the values never
change and (2) ateom has the same defaults. Only `kernel_params` is
relevant for the `kata-agent`; its value also changed once in the Kata
project history.

`ateom` now owns all three values, which also allows us to tune them
specifically for Substrate.

Verified e2e with a GKE cluster.

Fixes #1693

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-22 00:18:59 +00:00
Krisztian F 5f7c690108 (feat): Actor lifecycle events over otlp (#1658)
## What this does

ateapi writes a record every time an actor changes state. Until now
those records only went to the pod's stdout, and nothing reads stdout.
This sends the same records to a collector as OTLP log events.

Actor name and uid cannot be metric labels (too many values), and traces
are sampled at 1%. So these records are the only way to answer "what
state is this actor in, and since when".

## Changes

- `serverboot.InitLogging` sets up a LoggerProvider, next to the
existing tracer and meter ones.
- New `internal/actorevent` package builds the log records.
- ateapi emits at the two places that already write the stdout records.
- Two event names: `ate.actor.state_changed` and `ate.actor.crashed`.
- Both names are registered in `docs/metrics/registry/events.yaml`, so
`make verify` checks them.
- kind gets a logs pipeline and a count connector. The e2e suite reads
the counts back.
- Docs updated. `otel-collector.md` said substrate has no
LoggerProvider, which is no longer true.

## Opt-in

`OTEL_LOGS_EXPORTER` defaults to `none`. Only the kind overlay sets it
to `otlp`. The base ConfigMap is untouched, so no deployed environment
changes when this merges.

## Notes on the design

- **No slog bridge.** Only two call sites emit these records, so
emitting twice costs two lines. A bridge would also send every ateapi
log over the wire, could not set the event name, and would loop, because
SDK export errors are logged through slog.
- **Batching processor, not the simple one.** These records sit on the
actor resume path. A processor that exports inside the emit call would
add a blocking gRPC call there, so a slow collector would become control
plane latency.
- **Two event names, not one per state.** `ate.actor.state` already says
which transition happened. A crash gets its own name because it carries
two extra attributes and a higher severity.
- **Both copies are kept on purpose.** No collector in this repo reads
pod stdout, so nothing is duplicated today. `kubectl logs` keeps
working. If a filelog agent is ever added, drop one of the two. The
escape hatch is written down in `docs/metrics/substrate.yaml`.

## Dependencies

Adds `otel/log`, `otel/sdk/log` and `otlploggrpc`, all pinned at
v0.20.0. That is the release that matches the pinned `otel v1.44.0`.
v0.21.0 would pull the core modules to v1.45.0, which this change does
not need. The logs API has a
v1.47.0 release candidate upstream, so it is on its way to stable.

## Testing

- Unit tests for the exporter resolver, the record builder, and both
ateapi emit sites.
- The record builder test checks the attribute set matches what the
event name declares, in both directions.
- The ateapi tests check the OTLP record carries the same attributes as
the stdout record.
- Ran end to end on kind. Records arrive with the right event name,
severity, attributes, and with trace context on the record's own fields
rather than as attributes.
- Checked the off state too. With `OTEL_LOGS_EXPORTER` removed, the
collector receives no log records and stdout is unchanged.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-18 18:11:12 +00:00
yufan-su 85ce8ed531 egress: add credential injector (#1360)
Adds **egress credential injection**: on the sdsmint egress gateway's
decrypted MITM leg, when an actor's `EgressPolicy` rule matches and
carries an `inject_static_headers` effect, the gateway resolves the
referenced credential from a **credential provider** and sets it as a
request header (e.g. `Authorization: Bearer <token>`) before
re-originating upstream.

This implements the `applyEffects` in the egress handler. Injector runs
as part of the existing egress ext_proc handler that already
fetches/caches/evaluates each actor's `EgressPolicy`, so the MITM leg
gains injection without a second ext_proc hop.

### How it works

- New `CredentialProvider.FetchSecret` plugin API
(`pkg/proto/credproviderpb`): the gateway calls it with the credential
URI and the actor's attested SPIFFE identity; the provider authenticates
the gateway (mTLS) and resolves the secret.
- The egress handler dials the provider over mTLS
(`egress.DialProvider`) and injects on matched, allowed requests
(`egress.applyEffects`).
- Behavior:
- **TLS MITM leg + provider configured** → inject the credential
(overwriting any actor-set header).
- **Cleartext leg, or no provider configured** → skip injection and pass
the request through (never put a secret on a cleartext wire; don't block
allowed egress).
- **Attempted but failed** (unfetchable secret, empty/malformed secret,
credential URI of an unserved provider class) → fail closed.

## Installation

One flag stands up and wires the whole stack:
```
hack/install-ate.sh --deploy-ate-system --experimental-egress-credential-injection
```
`--experimental-egress-credential-injection` deploys credprovider and
the injector, then re-wires the sdsmint egress gateway to route through
the injector. It implies `--experimental-use-sdsmint`

### New install flags

Two flags configure which credential provider the injector targets:

| Flag | Purpose | Default |
|---|---|---|
| `--credential-provider-name` | Provider class, as a `ate-secret://`
prefix; a policy URI of any other class is refused |
`ate-secret://kubernetes.io` |
| `--credential-provider-address` | Where the injector dials the
provider | `credprovider.ate-system.svc:50051` |
2026-09-16 19:13:33 -04:00
Maya Wang 94f2285f92 atenet: raise the default route timeout, and size shutdown separately (#1529)
Addresses the timeout half of #1525. Deliberately not a closing keyword:
see Scope.

## The problem

`atenet-router` shipped with the workload route capped at 10s, plus 5s
of parking in front of it. That is below the normal duration of the
turns this platform exists to serve, since an agent relaying a model
completion holds the request open for the whole generation. A turn past
the ceiling returns `504 upstream request timeout`, and the caller
cannot safely retry it: the actor did receive the turn and is still
working, so a retry runs it twice.

Observed as 10 failures out of 10 on a burst-wake test during
acceptance, each at 15.9s to 18.4s.

The same ceiling bounds gRPC server-streaming and bidi RPCs, which is
what the `TODO(liorlieberman)` above the constant is about. This does
not resolve that TODO. It asks for streaming to get its ceiling without
imposing one on every workload, and a raised default is the blunt
version it warns against, so the TODO stays in place pointing at #1291.

## Why it was not a one-line change

`defaultRouteTimeout` was doing two jobs. It was the route default, and
it was the anchor the drain sequence sized its Envoy-drain window and
`--drain-timeout` from.

The `--route-timeout` flag never fed the drain, and `drainTimeout()`
carries a comment saying so. But the two constants were one, so raising
it for the route would have taken the derived drain from ~20s to over 5m
against a 60s grace period, and the kubelet would SIGKILL mid-drain.
This splits them so the route number can move without the drain number
following.

## What this does

Splits the constant.

- `defaultRouteTimeout` is now 5m and governs the route only.
- `drainRouteBudget` is new, stays at 10s, and is what shutdown plans
for. All the drain arithmetic is unchanged, in `drainTimeout()` and in
`dataplaneWindow`.

The behaviour that does not change: a turn still running after the drain
budget does not survive a shutdown. Operators who need one to raise
`--drain-timeout` and the grace period together, as before, and the
manifest still says so.

One thing to flag, since it is the weak point. Envoy's stream idle
default is also 5m and we never set it, so raising the route timeout
alone would not have worked: a turn that sends nothing while the actor
thinks is idle by that measure and would be reset before the route
timeout was ever reached. `routeIdleTimeout()` takes the larger of the
two, which leaves everything below 5m exactly as it was.

At exactly 5m, which is now the default, the two deadlines coincide and
either may fire first. An idle-triggered end reaches the client as a
stream reset rather than a 504, which is the worse of the two. Making
the idle timer strictly later would settle it, but that changes behavior
for anyone already running `--route-timeout` above 5m, so I left it
alone. Happy to take the other call.

The commented-out `--route-timeout=5m` comes out of the router manifest,
since it is now the default, and the comment above it is rewritten to
describe lowering rather than raising.

## Scope

**This does not fix the other half of #1525 and that issue should stay
open.** When a turn times out, the actor keeps working, finishes, and
the response has nowhere to return to. It is then delivered to whoever
speaks to that actor next, as a fluent `HTTP 200` answering somebody
else's question. We saw that on 6 of 10 actors. A longer ceiling makes
it rarer, not impossible, and anyone whose turns exceed 5m still reaches
it.

## Testing

`go test ./cmd/... ./internal/...` passes.

- `TestRouterConfigDrainTimeoutIndependentOfRouteTimeout` is new. It
pins the property this change turns on rather than the arithmetic: the
derived drain must not track the route timeout, and the full sequence
must fit inside the 60s grace period in the router manifest. It fails on
the naive version of this change.
- `TestXdsServer_RouteTimeout/IdleTimeoutAtTheDefaultRouteTimeout`
covers the equality case above, with
`IdleTimeoutTracksLongerRouteTimeout` and
`IdleTimeoutKeepsEnvoyDefaultWhenRouteTimeoutIsShorter` either side of
it.
- `SetterOverrides` now sets 30s instead of 5m, which was silently
passing against the new default whether or not the setter worked, and
lowering is the direction an operator capping turn length actually goes.
- The drain derivation cases and the idle-timeout cases now use values
that stay meaningful with the new default.

`make verify` does not complete in my environment:
`hack/verify/codegen.sh` fails creating the Locust codegen virtualenv
because `python3-venv` is not installed locally.
`hack/verify/boilerplate.sh`, `gofmt` and `go vet` are clean.
2026-09-15 16:18:57 -07:00
Benjamin Elder 06eb33f16f atelet: set resource requests and priority (#1651)
Reserve 50m CPU and 128Mi memory without limits, and prioritize atelet
above workers that depend on it.

I ran into issues while developing #1266 with the atelet being evicted
over the workers ... even without multi-actor workers we shouldn't let
that happen. atelet should be effectively system-critical like kubelet.
So we set maximum priority.
2026-09-15 15:03:01 -07:00
Keith Mattix II abd45ad081 Add agentgateway to CI & rename dataplane flag (#1598)
Run agentgateway data plane tests as a part of substrate CI
(non-blocking to start so we can confirm it's not flaky). Also, change
the `--atenet-router` flag to `--atenet-dataplane` to make it clearer
that the flag controls ingress and egress.

I've run the e2es locally across gVisor and microVM plus the MITM
variants for both. The only skip we do for agentgateway is
`TestIngressProtocolDowngrade` because 1. the behavior its testing only
exists on the non-CONNECT atunnel ingress path and agentgateway only
sends CONNECT to atunnel and 2. I'm not sure that we want this to be a
part of the contract that substrate is bound by (e.g. do we really want
to commit to atunnel always parsing HTTP?).

My goal with getting both dataplanes into CI is to start taking steps to
codify the proxy (router + egress PEP) contract for substrate. The
telemetry they emit, atunnel expectations, etc. are all important
contracts to explicitly call out so that they don't become too coupled
to a single dataplane implementation.

> It's a good idea to open an issue first for discussion.

- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-15 08:47:41 -07:00
Lior Lieberman c1d5b29f00 flip extproc order 2026-09-14 15:51:44 -07:00
Lior Lieberman 7034e79b10 atenet-egress: classify each tunnel and dial only what the policy allowed
The handler decides address rules at the CONNECT and every readable
request on its Host and the dialed address, so both gateways have to
give it those legs and act on what it decided.

The CONNECT chain's ext_proc may answer with dynamic metadata, and the
gateway admits exactly one namespace of it, dev.ate.egress. A
set_filter_state filter placed after the ext_proc copies
dev.ate.egress:passthrough_destination into Envoy's
envoy.network.transport_socket.original_dst_address and shares it with
the upstream connection. That is the whole mechanism carrying the
decision from the process that made it to the socket that acts on it,
with no header anywhere in between.

Inside the tunnel a connection lands on one of three chains, chosen by
tls_inspector and http_inspector under the 1s timeout that already
governs sdsmint:

  * egress_cleartext, on both gateways, matches raw_buffer with an HTTP
    ALPN and runs the policy ext_proc on every request. Its answer,
    dev.ate.egress:dial, picks the route: name goes to the
    dynamic_forward_proxy cluster, which resolves the Host, and address
    to an ORIGINAL_DST cluster fed by the same filter state as the
    passthrough chains, so the bytes go to what the matching rule
    checked. There is no route without a dial, and the answer clears the
    route cache, because Envoy picks a route before the filter runs. The
    chain accepts HTTP/1.0, which http_inspector steers here and the
    codec used to refuse before any filter ran.
  * egress_tls_mitm, sdsmint only, is that same per-request check and
    routing on the requests the gateway decrypted. Both of its dials
    re-originate TLS with the SNI and certificate check taken from the
    Host, so a Host the dialed origin cannot prove fails the handshake.
    Its filter sits below the existing #ATE_MITM_EXTPROC_FILTER markers,
    which keep working for an external processor spliced in by the
    installer.
  * the passthrough chains take what is left: on the plain gateway TLS it
    will not terminate, and on both gateways anything the inspectors
    could not classify, a client that said nothing before the timeout
    included. They run no filter. Their tcp_proxy points at an
    ORIGINAL_DST cluster with no address of its own, so it dials whatever
    the CONNECT decision put in filter state and closes the connection
    when there is nothing there. A destination no address rule allowed is
    therefore unreachable for traffic the gateway cannot read, which is
    what makes deferring a decision to the request legs safe.

    There is one such chain per transport protocol rather than a single
    chain matching nothing, because Envoy buckets filter chains by
    transport protocol and never falls out of a bucket that exists. With
    egress_cleartext holding the raw_buffer bucket and filling it with
    HTTP application protocols alone, a chain matching nothing would be
    unreachable and every opaque or unclassified connection would be
    closed as no_filter_chain_match, address rule or not. A manifest test
    now pins the invariant that each transport protocol the inspectors
    set has a chain, and that each has one catching what they could not
    name.

The plain gateway gains the inner listener it never had; until now it
spliced the tunnel straight to the IP:port and could police nothing
inside it. Its TLS stays end to end encrypted, since there is no MITM CA
there, so what it can enforce for HTTPS is the address, which the
manifest says in as many words. Server-speaks-first protocols pay the 1s
sniff timeout on this gateway now, as they already did on sdsmint.

The outer hop's connection pool is keyed per actor through
envoy.network.upstream_server_name, set from the same verified
certificate as the identity. Envoy captures shared filter state when a
pool is created and a string object takes no part in the pool's key, so
without that key a second actor's tunnel could be handed a pool carrying
the first actor's identity. The clusters that target an internal
listener also cap themselves at one request per connection.

The manifest tests pin the chain names, the attributes each ext_proc
requests, the metadata namespace every leg may answer in, the filter
that turns the CONNECT leg's answer into the passthrough address and the
absence of any other writer of that key, the two routes each request leg
has and the clusters they select, and a distinct stat_prefix per
ext_proc, because a mismatch fails closed at runtime and nothing else
would notice.
2026-09-14 15:51:44 -07:00
Julian Gutierrez Oschmann b8e403eedd Add missing validation to actor template container image. (#1614)
The container image must include the image digest. This validation was
dropped during the migration of the ActorTemplate resource from CRD to
Substrate API.
2026-09-11 10:22:28 -07:00
Keith Mattix II f16fc04fa0 Move from Host header to explicit headers for actor and atespace
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Da Huang 1ec75f9578 manifests: wire OTLP exporter settings into the egress ext-proc (#1460)
The egress ext-proc container never received the shared OTLP exporter
settings: unlike every other control plane component, its Deployment did
not consume ate-otel-config via envFrom, so OTEL_EXPORTER_OTLP_ENDPOINT
was unset and the OTel SDK fell back to localhost:4317, where nothing
listens. Every egress metric and span was silently dropped, with the SDK
logging an export-failure warning every ~20s — on a live GKE install
this was the largest warn/error source in the cluster, and the
registry-listed ext_proc instruments simply did not exist in GMP for the
egress gateway.

Give the ext-proc the same envFrom, POD_UID, and
OTEL_RESOURCE_ATTRIBUTES wiring the ingress router has, in both the
plain and sdsmint egress manifests. --otlp-collector-address stays
explicitly empty: that flag only controls Envoy-side tracing, and a span
per CONNECT tunnel is dataplane-volume noise; the comment now says so
instead of leaving the empty value looking accidental. Also add
atenet-egress to the consumer list in the ate-otel-config header
comment.

Verified on a live install: after applying the env change the SDK
export-failure warnings stop and target_info for the egress deployment
appears in Managed Prometheus.

Fixes #1459

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-03 19:49:31 -04:00
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
shrutiyam-glitch 616fd83431 feat: Support Cloud SQL via Auth Proxy for PostgreSQL backend (#996)
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.

To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.

### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:** 
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.


- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
2026-09-03 15:45:55 -04:00
Maya Wang 4dbc4e6b77 ateapi: give the api-server a startup probe covering its store-connect budget (#1413)
Partially addresses #1394 — the "spend the retry budget it configures
for itself" half of that issue's expected behavior. Candidate fix 4 from
the issue, done first because it is manifest-only.

## The problem, in one line

`ate-api-server` configures itself a 60s store-connect budget and is
killed by its own liveness probe at ~30s, halfway through it.

## Why the kill happens

`:9090` is bound near the end of boot — after `connectStore` and, since
#1196, after the goose migrations — so nothing answers `/healthz` for
the whole of that window. The liveness probe was the only probe
governing it: `initialDelaySeconds: 10` then probes at 10s, 20s and 30s,
so the default `failureThreshold: 3` is met at ~30s. The events show
`connect: connection refused`, not a probe returning failure — the port
is closed, the process is fine.

The kill then makes it worse than a plain restart.
`connectStore(shutdownCtx)` roots the retry loop in the signal context,
so SIGTERM cancels it mid-budget: the loop reports `context canceled` at
attempt 15 of 30 rather than exhausting the budget, `serverboot.Fatal`
exits, and CrashLoopBackOff holds the pod down well past the point where
the store came back.

A *running* replica rides out the same outage with `restarts=0` and
serves again ~15s after the store returns. A booting replica was
strictly less resilient than a running one, which is backwards.

## The change

One `startupProbe`, 24 × 5s. Liveness is suspended until `/healthz`
answers, so boot gets ~115s before the kubelet intervenes, against a
~60s connect budget.

| | before | after |
|---|---|---|
| kubelet kills at | ~30s | ~115s |
| connect attempts used | 15 of 30 | 30 of 30 |
| failure reported as | `context canceled` | `connect to PostgreSQL
after 30 attempts: …` |

A dead process is still restarted, just on a budget sized for boot
rather than one that never was. And because a startup probe also
suspends the *readiness* probe, the pod now stays out of Service
endpoints while booting instead of being latched Ready by the zero-value
`Readiness` — a small move in the right direction for #1395 rather than
against it, which is why this rather than simply raising
`failureThreshold`.

## What this deliberately does not do

- **It does not fix #1394.** The process still `Fatal`s once its own 60s
budget is gone. The motivating scenario there is an ungraceful node
loss, where GKE's PD force-detach alone is ~6 minutes; this raises
tolerance from 30s to 60s, and only "retry for the life of the process,
report through readiness" survives six minutes. That fix needs the
reversible readiness predicate #1395 describes.
- **It does not bound the migrations.** #1196 put goose on the boot
critical path behind a `pg_advisory_xact_lock`, so with two replicas
restarting together one waits on the other — inside the window where
nothing is listening. On an empty schema that is milliseconds; on a real
database during an upgrade it is not bounded by anything, and no probe
budget can bound it. Ordering can.
- **It does not change the boot ordering.** A slow boot and a dead
process are still indistinguishable while the port is closed. Moving the
health surface ahead of `connectStore` is candidate fix 1 and the
natural follow-up.

## Testing

`hack/verify-all.sh` passes (`metrics.sh` skipped locally — needs weaver
or docker; no metrics change here).

Note for anyone rebasing: `make test` currently fails on
`TestExternalVolumeRenders` in `cmd/ate-setup/internal/demos/counter` on
a clean `cbae8250`, unrelated to this change.
2026-09-03 12:41:57 -07:00
Dmitry Berkovich a9c1bd3894 atelet: pre-download sandbox assets from SandboxConfigs (#1358)
## Summary

atelet now watches SandboxConfig objects and pre-downloads their sandbox
assets into the node's content-addressed cache in the background,
instead of only fetching them inside the first Run/Restore on the node.

Related to #811 — this does **not** fix it, it is one more optimization
step toward it. It removes the cold-cache download+extract from the
first Run/Restore in the common case (node boot, or a release bump while
the node is idle), but the window between a SandboxConfig change and
prewarm completion still exists, and the extraction itself remains on
the critical path when a resume lands inside that window. The
uncancellable/silent extraction called out in #811 is unchanged by this
PR.

## How it works

- A handler on the existing shared informer factory enqueues
SandboxConfig add/update events; registration happens after cache sync,
so a freshly booted node replays every existing config as a synthetic
Add and prewarms immediately.
- A single background worker drains the queue, so concurrent prewarms
never compete for node bandwidth. Downloads are jittered by up to 30s so
a config rollout doesn't hit the bucket from every node in the fleet at
once.
- gVisor configs prewarm on every node. Micro-VM configs prewarm only
where `/dev/kvm` exists (`microvmNodeCapable`, the same signal the
device plugin advertises as `ate.dev/kvm`) — known at atelet startup,
before any WorkerPool schedules to the node, and avoids pulling
multi-hundred-MiB guest images onto nodes that can never run them.
- Prewarm is purely best-effort: failures are logged and the on-demand
fetch in `ensureSandboxAssets` remains the correctness path. Racing the
two is safe because both install content-addressed files via atomic
rename.
- The atelet ClusterRole gains get/list/watch on `sandboxconfigs`.

Left as TODOs in the code: retry-with-backoff on prewarm failure, and GC
of cached assets no longer referenced by any SandboxConfig or on-node
actor record.

## Pause image prewarm (cherry-picked)

This PR also carries `atelet: prewarm the pause image alongside sandbox
assets` (cherry-picked from 2c7dd31f). The prewarmer downloaded the
runtime binaries but not the SandboxConfig's pause image, so every node
still pulled it from the registry inside its first Run/Restore. At
benchmark start that is a synchronized fleet-wide pull burst —
registry.k8s.io answered a 1000-node run with 429s, failing Restores —
and a failed pull writes no cache record, so each subsequent Restore
re-pinged the registry and fed the rate limiter. The prewarm worker now
pulls the pause image into the image cache concurrently with the asset
downloads (different backends, so neither fetch waits on or fails the
other), inheriting the existing per-config jitter. It remains
best-effort: the pull inside `prepareOCIBundles` is still the
correctness path.

## Testing

- Unit tests cover the arch projection (`recordFromSandboxConfig`), the
enqueue filtering (KVM gating, unknown classes, full-queue drop), and
`microvmNodeCapable` negative cases.
- An end-to-end test drives a SandboxConfig from a fake clientset
through the real informer into the prewarm worker and verifies the asset
lands in the static-files cache.
- The pause-image commit adds `TestPrewarmPauseImage`, pulling from a
local test registry to verify the image lands in the image cache and
that an asset-fetch failure does not stop the pull.
- `go vet` and the full `cmd/atelet` test package pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-02 19:49:07 -07:00
Dmitry Berkovich 37da0cddbb Update gVisor to the 2026-09-02 nightly and switch to zstd tarballs (#1403)
## Summary

Improves the cold-node activation performance tracked in #811: the
gVisor release extract on a cold node currently exceeds the router
resume budget, so the first resume after a release bump always fails.

- Point the default gvisor SandboxConfig at the 2026-09-02 nightly,
using the `gvisor.tar.zstd` tarballs the nightlies now publish alongside
`gvisor.tar.bz2`. The win is decode speed, not just download size:
stdlib bzip2 is a pure-Go, single-threaded decoder, and extracting the
~165 MB release costs a node roughly 20 seconds of CPU on the first
actor operation it hosts — the dominant term of a cold-node activation.
With zstd the same extraction takes ~2.9 s, measured on a GKE node
running this change. The tarball is also ~22% smaller to download (129
MB vs 166 MB for x86_64).
- Teach `extractTarArchive` to decompress `.tar.zst` / `.tar.zstd`
archives using `github.com/klauspost/compress/zstd` (already a
dependency via ategcs), with test coverage for both suffixes.
- Update the `gvisor.tar.bz2` mentions in docs, manifests comments, and
the SandboxConfig API comment (generated CRD refreshed via
`hack/update/codegen.sh`).

The bucket only publishes sha512 checksums; the pinned sha256 sums were
computed from downloads verified against the published sha512 sums.

## Test plan

- `go test ./cmd/atelet/` (includes new `.tar.zst` / `.tar.zstd`
extraction cases)
- `go build ./cmd/atelet/...`, `go vet`, `gofmt` clean
- Verified the real x86_64 tarball lists `runsc` plus the `gvisor-bin/`
helpers
- Deployed to a GKE cluster: atelet fetched and extracted the 2026-09-02
x86_64 zstd tarball in ~2.9 s (vs ~20 s for bz2)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-02 11:10:52 -07:00
Benjamin Elder 7535180e5e Drop device passthrough until the API is designed
A GPU worker pool injected its devices into every container of every actor on it,
which only held up while a worker ran one actor.

A resource name and a count cannot replace it: that shape cannot express which
container gets a device, sharing one between containers, or resuming onto hardware
compatible with the snapshot. A nvidia.com/gpu pool limit is now an ordinary
extended resource -- it places the pod on a GPU node and nothing else reads it.

internal/cdi stays, unused: whatever the API becomes, an ateom still has to merge
CDI edits into an actor's bundle itself.
2026-09-02 10:27:14 -07:00
Jeremy Alvis f100c27286 ateapi: add versioned PostgreSQL schema migrations (#1196)
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.

## Summary

This change replaces startup schema setup with embedded, versioned SQL
migrations.

`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.

Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.

Closes #901.

Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing

## Migration behavior

`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.

One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.

If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.

Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.

## Changes

- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.

## Migration file policy

Migration files use sequential versions and contain exactly one Goose
`Up` section.

CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.

Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.

After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.

Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.

## Compatibility

No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.

Users must recreate databases from earlier PostgreSQL development
builds.

Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.

A binary rollback does not roll back the database schema.

## Testing

Tests cover:

- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
2026-09-02 10:56:19 -04:00
Zoe Zhao 120a519606 Delete ActorTemplate CRD (#1376)
Fixes #368 . Deletes ActorTemplate CRD and any references to it.

Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 10:48:42 -04:00
Haven Xia 9df8ca6abe [Upgrade] Version the install: node labels and a per-version atelet DaemonSet (#1357)
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).

This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.

### One derivation for the version label

`internal/versionlabel` turns the build version into its two forms in
one place:

| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |

A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.

### A DaemonSet per version

`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.

### The install becomes version-aware

- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.

### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.

Fix #1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
2026-09-01 12:40:49 -07:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
haiyanmeng e748574364 atenet: unify the ate.* dataplane attribute namespace (#1198)
unify the ate.* dataplane attribute namespace
2026-09-01 10:38:08 -07:00
Krisztian F 39df566dfa (otel): atelet and ateom telemetry now says which node it came from (#1363)
We want to be able to answer questions like:
 - image cache hit ratio per node
 - cold starts got slower on one node

but nothing we emit says which node anything is on, this PR fixes this.

Verified on a local kind cluster, where both atelet and ateom-gvisor now
show `k8s_node_name` on `target_info`, including ateom's, which arrives
through the OTLP relay.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-01 11:02:53 -04:00
Haven Xia 35344a2c68 api: make WorkerPool workerImage required again (#1334)
We should not make workerImage optional as it provide explicit
information for what image workers use, removing it will need
`atecontroller` to inject workerImage and leads to unnecessary control
plane & dataplane wiring.



Fixes #861 

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-31 10:19:41 -07:00
botengyao a0c4ee4622 atenet-egress: disable the CONNECT route timeout on the default gateway (#1328)
Envoy applies a route timeout to a CONNECT tunnel's whole lifetime, so
the 15s default capped every actor's outbound connection. The sdsmint
variant already disabled it; add the same line here, and a test over
both manifests so they cannot drift apart again.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-31 10:32:32 -04:00
Yuan Gao bb83206fd0 atenet/egress: resolve upstream names on both address families (#1060)
Part of #246

The egress Envoy pinned `dns_lookup_family` to `V4_ONLY` in both the
dynamic
forward proxy filter and the dynamic forward proxy cluster. On an
IPv6-only
cluster every upstream connection failed: Envoy reported a DNS
resolution
failure and the actor got a 503. Seven sites across the two egress
manifests,
including the sdsmint variant, now use `ALL`, which returns both
families and
enables Happy Eyeballs.

This is a prerequisite for IPv6 egress, not the fix on its own. An
actor's
connection is redirected by nftables and atunnel recovers the original
destination with `getsockopt(SOL_IP, SO_ORIGINAL_DST)`, which returns
ENOENT
for a v6-redirected connection — so egress fails before Envoy is ever
asked to
resolve anything.

## Testing

`make verify` is clean. A new test walks the shipped manifests and
requires
`ALL` on every `dns_cache_config`, so a new egress variant cannot
reintroduce
the pin. It has already paid for itself: the seventh site arrived with
the
sdsmint MITM leg while this was in review, still pinned to `V4_ONLY`,
and the
test caught it on rebase.

Measured on an IPv6-only kind cluster with #911, #958, #979 and #753
applied.
Each value was deployed, Envoy restarted, and the live `/config_dump`
checked
before running the suites:

| `dns_lookup_family` | `TestActorEgress` / `TestActorEgressHTTPS` |
| --- | --- |
| `ALL` (shipped) | pass / pass, `code=200` to
`[2606:4700:10::ac42:93f3]:80` |
| `V4_ONLY` (before) | fail / fail, `code=503 flags=DF` |
| `AUTO` | pass / pass |

`AUTO` is not distinguishable from `ALL` on this path: the dynamic
forward
proxy is handed the IP literal that atunnel recovered, never a hostname,
so
`dns_lookup_family` only decides which literal families it accepts.
`ALL` is
chosen as the value that strands neither family.

The IPv4 lane is unaffected: `TestActorEgress` and
`TestActorEgressHTTPS` pass
with the change in place, across two runs of the standard e2e job.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-31 09:10:05 -04:00
haiyanmeng 9ff766279e atenet: run every Envoy on one digest-pinned v1.39 (#1301)
Three Envoy deployments, three different notions of what version they
    were on: the passthrough egress gateway floated at v1.34-latest, the
MITM one was digest-pinned to v1.37, and the ingress router floated at
v1.39-latest. A behavior difference between two of them was therefore
    ambiguous, since it could come from the proxy rather than from the
configuration under test, and the floating tags let the gaps widen on
    their own with every upstream release.
    
    Put all three on v1.39-latest at the same multi-arch index digest.
Pinning the digest rather than tracking the tag keeps arm64 resolvable
while making a version change an explicit edit that shows up in review,
    which is the property the two floating references were missing.
    
Drop read_only from the actor-identity filter state along the way, since
v1.39 warns on it at every listener. Envoy removed the mutable versus
read-only distinction entirely -- the checks caused subtle runtime bugs
    and bought little.
    
    The passthrough, MITM, and additional-ext_proc egress renderings all
    validate against v1.39 with no warnings or deprecations, as does the
router bootstrap. The router's listeners and clusters arrive over ADS
    and are not covered by that check.

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-28 18:03:09 -07:00
Chuang Wang d909d69053 atenet-egress: deliver TLS certs via filesystem SDS so rotation works
Envoy honors watched_directory only on SDS-delivered secrets; on the
egress gateway's inline file-based certs it was silently ignored, so
kubelet's projected-certificate rotation never reached Envoy and the
gateway eventually served an expired certificate.

Move the serving cert (both egress manifests) and the spliced extproc
cluster's client cert and trust bundle (Go and shell emitters) to
filesystem SDS, with the flag-derived SAN pin kept inline via a
combined validation context. Fix the csi-deployment doc example that
showed the same broken pattern, and add tests for the emitters.
2026-08-28 14:52:33 -07:00
Haven Xia bdf494999b api: rename ateomImage to workerImage and make it optional (#1210)
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of #861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 12:46:00 -07:00
nybidari f45a0561c0 Bump default gvisor sandbox config to 2026-08-28 nightly (#1287)
Contains the gVisor fixes for #1090 #992
2026-08-28 12:15:26 -07:00
Lior Lieberman 80b652d1c6 remove https-h2 guard 2026-08-28 11:26:28 -07:00
Lior Lieberman b039c2f706 atenet: offer h2 on the HTTPS ingress and mirror the protocol to actors, behind a flag 2026-08-28 11:26:28 -07:00
Keith Mattix II f5517c0e8c Flesh out agentgateway implementation (#1265)
Flesh out the agentgateway implementation as a router and egress PEP.
This also changes the atunnel CONNECT port to do 8443. Most of this is
agentgateway contained, but the one other system we touch is localca;
agentgateway's MITM support requires the TLS cert and key in the secret
instead of the JSON pool (we don't rely on sdsmint).

- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-08-28 14:19:29 -04:00