182 Commits
Author SHA1 Message Date
Lior Lieberman 13e8efff25 envoy dyn module followup (#1995)
do not merge, not a draft cause i do want ci running

implements: 
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching


tls_passhtrough is a followup. 



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-30 16:57:00 +00:00
Steven Shriver 7c0c70c7a2 CI: create JUnit artifacts and gate on all tests passing (#1873)
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:

1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).

The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.

Partial work for #1871

- [x] Tests pass
   ```
  Run make build-junittool
  make build-junittool
  bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
  shell: /usr/bin/bash -e {0}
  env:
    E2E_ATENET_DATAPLANE: agentgateway
    ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0

/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1

/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
  TOTAL   
  ```
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 22:01:26 +00:00
Davanum Srinivas c7dbe9d672 nodepath: move the node state root to /var/lib/ate (#1926)
Fixes #1911

`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.

No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 18:27:55 +00:00
Benjamin Elder 22ba860e0f add get-contributors.sh for releases (#1870)
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.

v0.1.0 used a form of this, checking it in before v0.2.0

Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
2026-09-25 01:38:04 +00:00
Max Smythe d6d2a0fa1d Add the option to install a large-cluster manifest and cordon control plane to dedicated machines (#1632)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.

It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 23:39:01 +00:00
yanavlasov 30e6d33daa Initial commit of Envoy Dynamic Modules (#1535)
Collecting initial feedback on integrating Envoy dynamic modules for
Substrate dataplane.

Envoy Dynamic Modules allow fast iteration and experimentation on
Substrate dataplane. Dynamic modules are used to improve feature
velocity. As requirements and solution are better understood they will
be generalized and moved into Envoy's main repository.

Important points:

- Dynamic modules introduce Rust toolchain into Substrate dev
environment. This PR does not add GH actions for testing Dynamic Modules
during CI builds. This will be added in a subsequent PR.
- This change brings in Dockerfiles and docker buildx. It produces an
image with a versioned Envoy binary (1.39) and accompanying dynamic
modules.
- Docker image registry URL in deployment spec for Envoy datplane is at
this point hardcoded to `localhost:5001`. This will break for anyone
using a different registry. Fixing this will require using envsubst or
equivalent. I'm open to other suggestions.


- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Yan Avlasov <yavlasov@google.com>
2026-09-24 17:18:48 +00:00
Julian Gutierrez Oschmann 21b0b030bc Generate the benchmark Python gRPC clients at image build time (#1825)
protoc's Python output puts the serialized descriptor on a single line
which makes any two concurrent ateapi.proto changes conflict.

This is unfortunate as both proto and go generated code usually merges
with no conflict.

Stop checking in the *_pb2*.py files. The locust and nighthawk-ingress
images now generate them in a build stage via
benchmarking/locust/codegen/generate.sh.

Fixes #1824
2026-09-23 19:45:51 +00:00
Steven Shriver 898f6e6495 CI: Emit JUnit from test runners, pin their exec bounds (#1687)
Today CI depends on process exit codes. This leads to a few issues with
reporting results. For example, a green build cannot be distinguished
from one where no tests ran at all; a `-run` filter that matches
nothing; an e2e suite short-circuiting without `--e2e`; or a
testcontainer skip: all currently exit 0.
    
This change adds gotestsum as a pinned tool module allowing opt-in for
JUnit reporting in both test runners by using `E2E_JUNIT_FILE` and
`ROOT_JUNIT_FILE` respectively. When left unset, the runners exec `go
test` directly and gotestsum is never built leaving local runs unchanged
for now.
    
Additionally, this change pins two execution bounds that `go test` would
otherwise infer. Both remain overridable via environment variables as
listed below.
    
- `-p`: system defaults didn't align with cluster capacity so pinning
the value avoids over allocation. (override with `E2E_PARALLELISM`).
    
- `-timeout`: Suite defaults conflicted with some E2E test timings
causing some tests that are expected to timeout at the same time to
abort rather than fail as expected. Setting an 30m default avoids the
conlficting timeouts producing clean failure signals (override with
`E2E_TIMEOUT`).

Fixes #1685
2026-09-23 17:27:38 +00:00
Bowei Du 6c70d60c19 Update to the latest version of the hack scripts (#1785)
* Fixes some minor inconsistencies
* Deletes the shell scripts (at least some of them)
* Creates a temporary shim in the old shell to call the new ate-setup
command.
2026-09-22 18:28:34 +00:00
Huy Pham 92a84388b4 microvm: bump kata assets to 4.1.0 (#1708)
Kata 4.1.0 bundles virtiofsd 1.14.0 (required by Substrate). This
simplifies the dev process on arm64 because it removes the need to build
virtiofsd from source.

Verified on an arm64 KVM host (Lima + kind).

Fixes #1695

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-22 16:07:12 +00:00
yufan-su 57234a866a egress: add an example credential provider which reads from k8s secret (#1335)
Add an example gRPC credential-provider plugin which reads from k8s
secrets for the egress
credential-injection path. 

 - It resolves `ate-secret://k8s.io/default/<namespace>/<secret>/<key>`
URIs to Kubernetes Secret values, so Substrate never stores secrets — it
only
brokers a read that the provider is authorized to perform. 
 - It is the only
component in the injection path with Kubernetes secerts access; the
egress gateway and
its injector never read Secrets directly.

### What's included

- **`cmd/credential-provider/kubernetes-secrets`** — the gRPC service:
- Parses and validates `ate-secret://` URIs, resolving the requested
Secret
    (with single-key fallback when the URI omits a key).
- Enforces an **atespace→namespace authorization policy**
(default-deny),
    derived from the caller's attested actor SPIFFE ID.
- Serves over **mutual TLS**, requiring the caller's client cert to
chain to
    the trust bundle *and* carry the egress injector's SPIFFE SAN.
- **Manifests** (`manifests/egress-credential-injection/`) — Deployment,
Service, ServiceAccount + RBAC, a sample namespace-policy ConfigMap, and
a
  sample Secret.
- Renames the default provider address/service from `credprovider` to
`k8s-credential-provider` across `ate-setup`, install scripts, and the
egress
  injection overlay.

### atespace → namespace authorization

Beyond the mTLS check that only the egress injector may call the
provider, each
request is authorized against a **default-deny atespace→namespace
policy**. The
provider derives the requesting actor's atespace from its attested
SPIFFE ID and
resolves a Secret only if that atespace is explicitly granted access to
the URI's
namespace.
2026-09-22 10:13:25 +00:00
Huy Pham 8062bbafb4 microvm: drop the kata-config asset (#1704)
Today, `ateom` fetches the `kata-config` asset (configuration-clh.toml)
to read only 3 values `default_memory`, `default_vcpus` and
`kernel_params`. The first 2 are redundant because (1) the values never
change and (2) ateom has the same defaults. Only `kernel_params` is
relevant for the `kata-agent`; its value also changed once in the Kata
project history.

`ateom` now owns all three values, which also allows us to tune them
specifically for Substrate.

Verified e2e with a GKE cluster.

Fixes #1693

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-22 00:18:59 +00:00
Huy Pham 9f57d78f4f Add --enable-nested-virtualization to setup-gcp (#1669)
Nested virtualization exposes `/dev/kvm`, which microVMs need. The new
`--enable-nested-virtualization` flag turns it on for the node pool
`setup-gcp` creates. It defaults to on, so pass
`--enable-nested-virtualization=false` for a cluster that does not need
it.

Also, rename the env var `GVISOR_NODE_MACHINE_TYPE` to
`NODE_MACHINE_TYPE` to be consistent with other env vars and to avoid
confusion for microVMs. The old env var still works with a warning.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-18 04:27:38 +00:00
yufan-su 85ce8ed531 egress: add credential injector (#1360)
Adds **egress credential injection**: on the sdsmint egress gateway's
decrypted MITM leg, when an actor's `EgressPolicy` rule matches and
carries an `inject_static_headers` effect, the gateway resolves the
referenced credential from a **credential provider** and sets it as a
request header (e.g. `Authorization: Bearer <token>`) before
re-originating upstream.

This implements the `applyEffects` in the egress handler. Injector runs
as part of the existing egress ext_proc handler that already
fetches/caches/evaluates each actor's `EgressPolicy`, so the MITM leg
gains injection without a second ext_proc hop.

### How it works

- New `CredentialProvider.FetchSecret` plugin API
(`pkg/proto/credproviderpb`): the gateway calls it with the credential
URI and the actor's attested SPIFFE identity; the provider authenticates
the gateway (mTLS) and resolves the secret.
- The egress handler dials the provider over mTLS
(`egress.DialProvider`) and injects on matched, allowed requests
(`egress.applyEffects`).
- Behavior:
- **TLS MITM leg + provider configured** → inject the credential
(overwriting any actor-set header).
- **Cleartext leg, or no provider configured** → skip injection and pass
the request through (never put a secret on a cleartext wire; don't block
allowed egress).
- **Attempted but failed** (unfetchable secret, empty/malformed secret,
credential URI of an unserved provider class) → fail closed.

## Installation

One flag stands up and wires the whole stack:
```
hack/install-ate.sh --deploy-ate-system --experimental-egress-credential-injection
```
`--experimental-egress-credential-injection` deploys credprovider and
the injector, then re-wires the sdsmint egress gateway to route through
the injector. It implies `--experimental-use-sdsmint`

### New install flags

Two flags configure which credential provider the injector targets:

| Flag | Purpose | Default |
|---|---|---|
| `--credential-provider-name` | Provider class, as a `ate-secret://`
prefix; a policy URI of any other class is refused |
`ate-secret://kubernetes.io` |
| `--credential-provider-address` | Where the injector dials the
provider | `credprovider.ate-system.svc:50051` |
2026-09-16 19:13:33 -04:00
Eitan Yarmush a58481a18e Publish ActorTemplate golden snapshots as tags (#1523)
Fixes #1507

Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.

`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.

This PR is based directly on `main` and does not depend on #1521.

Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-16 09:24:42 -04:00
Keith Mattix II abd45ad081 Add agentgateway to CI & rename dataplane flag (#1598)
Run agentgateway data plane tests as a part of substrate CI
(non-blocking to start so we can confirm it's not flaky). Also, change
the `--atenet-router` flag to `--atenet-dataplane` to make it clearer
that the flag controls ingress and egress.

I've run the e2es locally across gVisor and microVM plus the MITM
variants for both. The only skip we do for agentgateway is
`TestIngressProtocolDowngrade` because 1. the behavior its testing only
exists on the non-CONNECT atunnel ingress path and agentgateway only
sends CONNECT to atunnel and 2. I'm not sure that we want this to be a
part of the contract that substrate is bound by (e.g. do we really want
to commit to atunnel always parsing HTTP?).

My goal with getting both dataplanes into CI is to start taking steps to
codify the proxy (router + egress PEP) contract for substrate. The
telemetry they emit, atunnel expectations, etc. are all important
contracts to explicitly call out so that they don't become too coupled
to a single dataplane implementation.

> It's a good idea to open an issue first for discussion.

- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-15 08:47:41 -07:00
Lior Lieberman f980a57d83 atenet: enforce the actor's EgressPolicy in the egress ext_proc handler
The egress handler so far authenticated the actor behind a CONNECT and
let everything through. It now authorizes the traffic against the actor's
EgressPolicy, on the leg where the destination is known. Which leg that
is comes from the filter chain name the dataplane asserts, the same
signal the mux already dispatches on.

  * egress, the outer CONNECT, is the only leg that sees a certificate,
    and the first to see the address the actor's kernel dialed. It keeps
    authenticating that certificate and decides the address rules,
    ip_blocks and all, against that address, once per connection. When
    one of them allows it the handler hands the address back as dynamic
    metadata, which is what lets the gateway dial it for traffic it
    cannot read and what tells the legs inside what was dialed. When
    none does but the policy names hostnames, the tunnel is opened and
    the requests inside are left to the legs that will see a name. When
    neither is true nothing inside the tunnel could ever be allowed, so
    the CONNECT is refused. A dataplane that calls out for the CONNECT
    alone sends no chain name and has no such legs behind it, so it gets
    the whole decision here and is refused rather than deferred.

  * egress_tls_mitm and egress_cleartext are the legs the gateway can
    read. Every request is decided the way the API describes: the rules
    in order, over the request's Host and the address the actor dialed,
    first match wins. It runs per request because the Host can change
    between requests on one connection. The answer also says where the
    request goes, so the bytes reach what the matching rule checked: a
    hostname match is resolved and dialed by name, an address or all
    match goes to the dialed address. Sending an address match by name
    would let an actor put any Host on an allowed address; sending a
    hostname match to the dialed address would turn the rule into a
    header check, since the actor picks both. A Host header that
    disagrees with :authority is refused rather than policed on one name
    and dialed on the other.

Identity on the request legs is the actor's SPIFFE ID, which the outer
chain sets as filter state from the certificate Envoy verified and
shares across the internal-listener hop; nothing inside the tunnel can
write it, and a callout without it is refused. The certificate's URI SAN
must name the same actor as its ActorIdentity extension, because the
CONNECT authenticates on the extension while the request legs attribute
traffic to the SAN.

Policies are read through the per-actor cache, whose TTL is a new
--egress-policy-cache-ttl flag defaulting to 10s, so a request inside a
tunnel is not a control-plane round trip.

Credential injection is not implemented yet. A matched rule that declares
one denies with 501 rather than forwarding without the secret the policy
promised.
2026-09-14 15:51:44 -07:00
Julian Gutierrez Oschmann 32e08c22d8 Print single-resource output as a bare object, not a list. (#1627)
Affects get <name>, create, resume, pause, and suspend for actors, actor
templates, atespaces, tags, and workers: JSON/YAML no longer wraps a
single resource in a List, matching kubectl's convention. Listing (no
name, or multiple names) still returns a List.

Also plumb the writer through cmd.OutOrStdout() instead of hardcoding
os.Stdout to honor Cobra's writer.
2026-09-14 09:40:37 -07:00
Michelle Au bce9e7e229 Enhancements to setup-gcp (#1616)
Multiple enhancements to setup-gcp

- Add boot disk customizations. Since snapshots are written to boot
disk, being able to increase the size will mitigate IO throttling.
- Require only one of project-id or project-number to be specified
- Verify region and cluster-location are compatible
2026-09-11 14:51:49 -07:00
Keith Mattix II 6bd89588dc Move to lowercase for header references
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Keith Mattix II f16fc04fa0 Move from Host header to explicit headers for actor and atespace
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Julian Gutierrez Oschmann 4003ad42d9 Some fixes to kubectl-ate CLI (#1536)
* Fix the column names to match the resource fields.
* Rename template flag name (fix TODO from CRD -> Substrate API
refactor).
* Fix bug in `create actor` which conflated the atespace flag for both
actor and actor template.
2026-09-08 12:28:51 -07:00
NekoPunch fda35d0c04 fix(hack): let teardown delete crashed actors (#942)
Fixes #794

A crashed actor makes `--delete-all` abort mid-teardown, leaving the
ActorTemplate, atespace and `ate-system` behind; the next
`--deploy-demo-counter` keeps the existing template, so the cluster goes
on serving the stale golden snapshot.

`prepare_actor_for_delete` treats `STATUS_CRASHED` as unexpected and
returns 1, which errexit turns into a hard abort, even though
DeleteActor accepts CRASHED and is idempotent for an actor already
marked DELETING (`workflow_delete.go`). Teardown now treats crashed and
already-deleting actors as deletable as-is.

The demo flag dispatch had the mirror-image defect: handlers ran from an
`if` condition, which suppresses errexit for their whole call tree, so
this same teardown and failed deploys like the issue's immutable-spec
apply exited 0. Handlers now run as plain commands and report an
unclaimed flag through `ate_demo_flag_unhandled`; the Jupyter handler
follows the same contract.

Repro: deploy the counter demo, resume an actor, delete its worker pod,
wait for the syncer to mark it `STATUS_CRASHED`, then run
`--delete-all`.
2026-09-07 12:13:26 +02:00
Tim Hockin 2fcfa64a73 Drop k8s third_party deps now that HEAD is updated (#1480)
The changes we need are live on HEAD of the k8s published repos. It was
a bit of a journey, since Go fought me all the way.

---

### Drop our third_party fork of k8s deps in tools
    
    The changes we need are released now (sort of - on HEAD anyway).

---

### Bump k8s.io/streaming to v0.37.0 (not rc) in tools

---

### Bump codegen deps to HEAD in tools
    
    GOPROXY=direct go get \
        k8s.io/apimachinery@master \
        k8s.io/code-generator@master
    
This pins them to HEAD of master. Go is terrible here: The HEAD is not
    actually tagged, so Go just uses the next "reachable" tag which is
    v0.36.0-alpha.  The datestamp is correct, though.

---

### Drop our third_party fork of k8s deps in root
    
    The changes we need are released now (sort of - on HEAD anyway).

---

### Bump k8s deps to v0.37.0 (not rc) in root

---

### Bump apimachinery dep in root to HEAD in root
    
GOPROXY=direct go mod edit -replace
k8s.io/apimachinery=k8s.io/apimachinery@master
    GOPROXY=direct go mod tidy
    GOPROXY=direct go mod vendor
    GOPROXY=direct go mod tidy
    
    This approach (-replace) is needed because Go is horrible here.
    
    The master branch of k8s.io/apimachinery is not tagged, per se, but
there is an OLDER tag which is "reachable" from HEAD. So go helpfully
decides to use that (v0.36.0-alpha.2). If we just `go get ... @master`
it works for that dep (pinned to the right date) but then it looks at
    transitive deps.  Because the tag seems to be 0.36 (older), it
recalculates all the OTHER dependencies and downgrades a whole tangle of
    things to versions that match 0.36, but we are ACTUALLY on 0.37+.
    
    This was the only approach that I (and Gemini) could find.  Blech.

---

### Run updated codegens

---

### Use DV's new maxBytes capability for `[]byte`

    Removes 1 custom.
2026-09-04 17:02:12 -07:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Haven Xia ac175376e4 Refactor the Python proto codegen into update/codegen.sh (#1483)
Fixes #1481

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 10:50:20 -07:00
Max Smythe a2233676bc install-ate.sh: allow overriding the JWT issuer ate-api-server expects (#1466)
Allow explicit overriding of JWT_ISSUER
2026-09-04 10:16:11 -07:00
Benjamin Elder 2c429a9906 multi-actor worker API (#1283)
Part of #1266 

This is a draft of the core API + data store changes.

It's still a large PR, apologies.

The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.

Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.

They're loadbearing for performance once we actually have multi-actor
workers.
2026-09-03 21:19:21 -07:00
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
Aditya ShantanuandAditya Shantanu 999dd3e449 docs: state the supported GKE versions for setup-gcp clusters (#1437)
The Create Cluster warning explained the beta-API creation-time
constraint but not which GKE versions work. Clarifies the two supported
configurations: **GKE 1.36 with the beta APIs enabled at cluster
creation**, or **GKE 1.37+** where `certificates.k8s.io/v1beta1` is
served by default (no beta enablement needed). Versions below 1.36 are
unsupported.

- [x] Tests pass (docs only)
- [x] Appropriate changes to documentation are included in the PR

---

Also updates `hack/ate-dev-env.sh.example`: it pinned
`1.35.5-gke.1163012` — below the supported floor and no longer offered,
so copying it verbatim failed cluster creation. Now pins the newest
offered 1.36 patch (`1.36.3-gke.1767000`) with a comment on how to pick
a fresh one when the pin ages out.

**Validated**: `setup-gcp create cluster` at `1.36.3-gke.1767000`
(us-west1-c) succeeded, and both `certificates.k8s.io/v1beta1` APIs
(`clustertrustbundles`, `podcertificaterequests`) are served on the
resulting cluster — `kubectl get clustertrustbundles` returns the
kube-apiserver-serving bundle. Throwaway cluster deleted after
verification.

---------

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-09-03 14:16:59 -07:00
shrutiyam-glitch 616fd83431 feat: Support Cloud SQL via Auth Proxy for PostgreSQL backend (#996)
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.

To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.

### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:** 
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.


- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
2026-09-03 15:45:55 -04:00
Aditya Shantanu cbae8250a5 Address substrate-gke installer feedback: AR API, teardown coverage, doc fix (#1408)
Three fixes from user feedback on the substrate-gke installer experience, consolidated per reviewer preference (replaces #1405, #1406, #1407):

**1. setup-gcp: enable `artifactregistry.googleapis.com` in bootstrap.** Bootstrap grants `roles/artifactregistry.reader` to the GKE node service account and the atelet workload-identity principal, and ko pushes the control-plane images through gcr.io's Artifact Registry backing — but `enableRequiredAPIs` never enabled the API, so a fresh project failed at image push/pull instead of step 1.

**2. teardown.sh: cover everything bootstrap creates, without a dev-env file.**
- Revoke atelet's project-level bindings (`roles/storage.objectAdmin`, `roles/artifactregistry.reader`), which `grant_atelet_permissions` adds but nothing removed.
- Delete the three Substrate monitoring dashboards, matched by the display names in `tools/setup-gcp/dashboards/`.
- Accept configuration from the environment when `.ate-dev-env.sh` is absent (missing variables are named specifically), so installers and one-liners can drive the script.
- Drop the kubectl cluster-admin precheck: every step talks to GCP, not the cluster, and it blocked tearing down a cluster that was already gone.
- `--all` now runs in the true reverse of bootstrap's setup order.

**3. commands.md: the `hack/install-demo-*.sh` scripts still exist.** The closing paragraph claimed they are gone, but `hack/install-ate.sh` still sources them and generates its `--deploy-demo-NAME` / `--delete-demo-NAME` flags from them.
2026-09-02 11:39:58 -07:00
Elijah Rodriguez-Beltran 7150d7aa2d e2e: automatically install CSI-nfs and enable in external volume lifecycle tests 2026-09-02 18:36:09 +00:00
Julian Gutierrez Oschmann 27b7b5904f Remove union discriminator field from ActorTemplate.volumes. (#1398)
Partially addresses #1397. Remove union discriminator `type` field from
`ActorTemplate.volumes`.
2026-09-02 09:29:19 -07:00
Haven Xia 6b76c6e89b [Upgrade] pin counter demo worker pools to the installed substrate version (#1393)
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-02 08:32:30 -07:00
Jeremy Alvis f100c27286 ateapi: add versioned PostgreSQL schema migrations (#1196)
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.

## Summary

This change replaces startup schema setup with embedded, versioned SQL
migrations.

`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.

Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.

Closes #901.

Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing

## Migration behavior

`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.

One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.

If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.

Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.

## Changes

- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.

## Migration file policy

Migration files use sequential versions and contain exactly one Goose
`Up` section.

CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.

Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.

After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.

Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.

## Compatibility

No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.

Users must recreate databases from earlier PostgreSQL development
builds.

Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.

A binary rollback does not roll back the database schema.

## Testing

Tests cover:

- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
2026-09-02 10:56:19 -04:00
Zoe Zhao 120a519606 Delete ActorTemplate CRD (#1376)
Fixes #368 . Deletes ActorTemplate CRD and any references to it.

Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 10:48:42 -04:00
Benjamin Elder 1d58b95a7f update grpc 1.83.2 in all modules (#1385)
picks up security fixes
2026-09-01 21:37:24 -07:00
Chuang Wang 07ebc56797 csi-hostpath: fix cert rotation in the testing Envoy config (#1336)
Follow-up to the review comments on #1288 (merged before they could land
there).

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Changes:
-
`hack/third_party/csi-driver-host-path/deploy/csi-hostpath-testing.yaml`:
same fix as #1288 — the inline `watched_directory` was a no-op, so the
serving cert and trust bundle move to filesystem SDS (SPIFFE SAN pin
kept inline via `combined_validation_context`). SDS requires a node id
and this sidecar passes no `--service-node`, so the bootstrap gains a
`node:` stanza; the `docs/csi-deployment.md` example had the same gap
and gets one too.
- `overlay_test.go` now asserts `watched_directory` appears nowhere in
the emitted cluster block or the patched `envoy.yaml` (comments
excluded) and is present in each SDS resource file, per review.
- Corrected the test comment claiming the sdsmint e2e suite is skipped
in CI: the sdsmint variant is deployed in the MITM lanes; only the
`--experimental-additional-egress-extproc-service` injection has no e2e
coverage.

Verification: `go test ./cmd/ate-setup/...` passes; `envoy --mode
validate` (v1.34) passes on the reworked csi-hostpath bootstrap with
dummy PEMs mounted at the referenced paths.
2026-09-01 20:01:22 -04:00
Elijah Rodriguez 53800cd045 Update CSI NFS server image to use kubernetes build 2026-09-01 14:57:42 -07:00
Zoe Zhao f6852b7754 Update existing demos and benchmark tests to use substrate ActorTemplate resource (#1355)
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.

Verifications done: 
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
2026-09-01 14:50:18 -07:00
Haven Xia 9df8ca6abe [Upgrade] Version the install: node labels and a per-version atelet DaemonSet (#1357)
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).

This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.

### One derivation for the version label

`internal/versionlabel` turns the build version into its two forms in
one place:

| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |

A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.

### A DaemonSet per version

`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.

### The install becomes version-aware

- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.

### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.

Fix #1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
2026-09-01 12:40:49 -07:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
haiyanmeng e748574364 atenet: unify the ate.* dataplane attribute namespace (#1198)
unify the ate.* dataplane attribute namespace
2026-09-01 10:38:08 -07:00
Benjamin Elder 8613dd9f7b hack: repair unusable Python venvs instead of trusting the directory (#1316)
Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no
bin/activate) and an interpreter upgrade (dangling bin/python3) as a
good venv, so the script dies in `source venv/bin/activate` and the only
fix is knowing to delete the directory. Probe that the venv runs, and
rebuild with --clear to relink the interpreter.

Also install requirements unconditionally, which is cheap. The license
verifier skipped the install whenever the venv already existed, so it
passed without ever seeing a newly added dependency.

Fixes failure encountered by @ahmedtd. Not filing an issue because this
is pretty trivial.

AI-assisted.
2026-08-31 09:52:29 -07:00
Benjamin Elder 07485f079a hack: let the kind node pull images in parallel [CI flake fix] (#1318)
Installing the control plane pulls roughly 570MB of third-party images
(postgres, prometheus, the otel collector, rustfs, envoy, jaeger) onto a
single node.

kubelet serializes image pulls by default, so they queue behind one
another and whichever workload lands at the back of the queue can miss
its readiness deadline.

e2e has failed repeatedly with postgres still in PodInitializing after
60s, its image not yet pulled, while everything else was already
running.

Sample failure:
https://github.com/agent-substrate/substrate/actions/runs/33258860897/job/99117220741

Note: This is after PR merge, and it's clearly unrelated to the PR.

AI-assisted.
2026-08-31 09:49:11 -07:00
Benjamin Elder 51f4114dc7 hack: wait for a pool's Deployment to exist before its rollout [CI flake fix] (#1320)
ate-controller creates a WorkerPool's Deployment, so it does not exist
yet when the apply returns.

`kubectl rollout status` reads the object before it starts watching and
errors on a missing one instead of waiting, so whether these waits
worked at all came down to beating the controller by a round trip.

e2e lost that race by 178ms, failing the deploy with NotFound while the
pool's pods were already being created.

Fixes yet another flake I observed here:
https://github.com/agent-substrate/substrate/actions/runs/33275880827/job/99162382080?pr=1283
2026-08-31 09:48:57 -07:00
Haven Xia 698289452b Bump Go to 1.27 2026-08-30 15:43:52 -07:00
Mesut Oezdil f936206dd5 ci: add kube-api-linter and exclude current findings (#416)
Part of #207, following the plan discussed there with @juli4n and
@kannon92: plumbing plus the non invasive fixes first, with the existing
invasive findings excluded so new APIs do not regress (as @kannon92
suggested in
https://github.com/agent-substrate/substrate/issues/207#issuecomment-4921051580).
The linter itself was originally suggested by @BenTheElder in #188.

Plumbing:
- hack/tools/kube-api-linter module with the
golangci-lint-kube-api-linter tool, same pattern as the other tools.
- .golangci-kal.yaml config. The pre-existing findings from rules that
need Go API changes (nomaps, nonpointerstructs, nophase, optionalfields,
requiredfields, ssatags) are excluded rather than disabled, and each
exclusion rule names the fields it excuses, so the rules still apply to
new API files and to new fields on the existing types. Fixing them on
the existing types is the follow up in #207.
- hack/verify/kube-api-linter.sh runs it against ./pkg/api/... and is
picked up by hack/verify-all.sh in CI.

Fixes in pkg/api/v1alpha1 (doc and marker only, no Go API change):
- commentstart: field godocs now start with the serialized field name.
- conditions: Conditions moved to the first position in
ActorTemplateStatus and got the listType, listMapKey and patch markers.
- defaultorrequired: SandboxConfigSpec.SandboxClass had both a default
and required; now optional with omitempty, matching the same field on
WorkerPoolSpec and ActorTemplateSpec. The ValidatingAdmissionPolicy in
sandboxconfig-validation.yaml is unaffected since CRD defaulting runs
before admission, so spec.sandboxClass is always set by then.
- optionalorrequired: missing +optional markers added on
ActorTemplateStatus fields.
- defaults: configured preferredDefaultMarker to kubebuilder:default
since CRDs are generated with controller-gen.

Generated CRDs regenerated. Structural schema changes are only the
conditions listType/listMapKey and sandboxClass no longer in the
required list (it is defaulted, so behavior is the same); the rest is
description text.

Note: my earlier count of 50 issues in #207 was capped by the default
golangci-lint issue limit. With the cap removed the real total is 86;
the extra ones are requiredfields (20) and ssatags (4), both in the
excluded set above.

Test plan
- hack/verify/kube-api-linter.sh passes (0 issues).
- Exclusions verified both ways: dropping them brings the pre-existing
findings back, and a scratch field added to WorkerPoolSpec is still
reported (optionalfields), which the earlier per-file exclusion
swallowed.
- TestSandboxConfigValidation passes against the regenerated CRD.
- go build ./... and go test ./pkg/api/... ./cmd/atecontroller/... pass.
- gofmt, shellcheck and the regular golangci-lint on pkg/api are clean.
2026-08-29 14:14:55 -07:00
Tim Hockin ebc7f3212c Use git to find the repo root (#1157)
Most scripts already do this.
2026-08-28 21:03:11 -07:00