do not merge, not a draft cause i do want ci running
implements:
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching
tls_passhtrough is a followup.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:
1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).
The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.
Partial work for #1871
- [x] Tests pass
```
Run make build-junittool
make build-junittool
bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
shell: /usr/bin/bash -e {0}
env:
E2E_ATENET_DATAPLANE: agentgateway
ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
TOTAL
```
- [x] Appropriate changes to documentation are included in the PR
Fixes#1911
`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.
No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.
v0.1.0 used a form of this, checking it in before v0.2.0
Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.
It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Collecting initial feedback on integrating Envoy dynamic modules for
Substrate dataplane.
Envoy Dynamic Modules allow fast iteration and experimentation on
Substrate dataplane. Dynamic modules are used to improve feature
velocity. As requirements and solution are better understood they will
be generalized and moved into Envoy's main repository.
Important points:
- Dynamic modules introduce Rust toolchain into Substrate dev
environment. This PR does not add GH actions for testing Dynamic Modules
during CI builds. This will be added in a subsequent PR.
- This change brings in Dockerfiles and docker buildx. It produces an
image with a versioned Envoy binary (1.39) and accompanying dynamic
modules.
- Docker image registry URL in deployment spec for Envoy datplane is at
this point hardcoded to `localhost:5001`. This will break for anyone
using a different registry. Fixing this will require using envsubst or
equivalent. I'm open to other suggestions.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Yan Avlasov <yavlasov@google.com>
protoc's Python output puts the serialized descriptor on a single line
which makes any two concurrent ateapi.proto changes conflict.
This is unfortunate as both proto and go generated code usually merges
with no conflict.
Stop checking in the *_pb2*.py files. The locust and nighthawk-ingress
images now generate them in a build stage via
benchmarking/locust/codegen/generate.sh.
Fixes#1824
Today CI depends on process exit codes. This leads to a few issues with
reporting results. For example, a green build cannot be distinguished
from one where no tests ran at all; a `-run` filter that matches
nothing; an e2e suite short-circuiting without `--e2e`; or a
testcontainer skip: all currently exit 0.
This change adds gotestsum as a pinned tool module allowing opt-in for
JUnit reporting in both test runners by using `E2E_JUNIT_FILE` and
`ROOT_JUNIT_FILE` respectively. When left unset, the runners exec `go
test` directly and gotestsum is never built leaving local runs unchanged
for now.
Additionally, this change pins two execution bounds that `go test` would
otherwise infer. Both remain overridable via environment variables as
listed below.
- `-p`: system defaults didn't align with cluster capacity so pinning
the value avoids over allocation. (override with `E2E_PARALLELISM`).
- `-timeout`: Suite defaults conflicted with some E2E test timings
causing some tests that are expected to timeout at the same time to
abort rather than fail as expected. Setting an 30m default avoids the
conlficting timeouts producing clean failure signals (override with
`E2E_TIMEOUT`).
Fixes#1685
* Fixes some minor inconsistencies
* Deletes the shell scripts (at least some of them)
* Creates a temporary shim in the old shell to call the new ate-setup
command.
Kata 4.1.0 bundles virtiofsd 1.14.0 (required by Substrate). This
simplifies the dev process on arm64 because it removes the need to build
virtiofsd from source.
Verified on an arm64 KVM host (Lima + kind).
Fixes#1695
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Add an example gRPC credential-provider plugin which reads from k8s
secrets for the egress
credential-injection path.
- It resolves `ate-secret://k8s.io/default/<namespace>/<secret>/<key>`
URIs to Kubernetes Secret values, so Substrate never stores secrets — it
only
brokers a read that the provider is authorized to perform.
- It is the only
component in the injection path with Kubernetes secerts access; the
egress gateway and
its injector never read Secrets directly.
### What's included
- **`cmd/credential-provider/kubernetes-secrets`** — the gRPC service:
- Parses and validates `ate-secret://` URIs, resolving the requested
Secret
(with single-key fallback when the URI omits a key).
- Enforces an **atespace→namespace authorization policy**
(default-deny),
derived from the caller's attested actor SPIFFE ID.
- Serves over **mutual TLS**, requiring the caller's client cert to
chain to
the trust bundle *and* carry the egress injector's SPIFFE SAN.
- **Manifests** (`manifests/egress-credential-injection/`) — Deployment,
Service, ServiceAccount + RBAC, a sample namespace-policy ConfigMap, and
a
sample Secret.
- Renames the default provider address/service from `credprovider` to
`k8s-credential-provider` across `ate-setup`, install scripts, and the
egress
injection overlay.
### atespace → namespace authorization
Beyond the mTLS check that only the egress injector may call the
provider, each
request is authorized against a **default-deny atespace→namespace
policy**. The
provider derives the requesting actor's atespace from its attested
SPIFFE ID and
resolves a Secret only if that atespace is explicitly granted access to
the URI's
namespace.
Today, `ateom` fetches the `kata-config` asset (configuration-clh.toml)
to read only 3 values `default_memory`, `default_vcpus` and
`kernel_params`. The first 2 are redundant because (1) the values never
change and (2) ateom has the same defaults. Only `kernel_params` is
relevant for the `kata-agent`; its value also changed once in the Kata
project history.
`ateom` now owns all three values, which also allows us to tune them
specifically for Substrate.
Verified e2e with a GKE cluster.
Fixes#1693
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Nested virtualization exposes `/dev/kvm`, which microVMs need. The new
`--enable-nested-virtualization` flag turns it on for the node pool
`setup-gcp` creates. It defaults to on, so pass
`--enable-nested-virtualization=false` for a cluster that does not need
it.
Also, rename the env var `GVISOR_NODE_MACHINE_TYPE` to
`NODE_MACHINE_TYPE` to be consistent with other env vars and to avoid
confusion for microVMs. The old env var still works with a warning.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Adds **egress credential injection**: on the sdsmint egress gateway's
decrypted MITM leg, when an actor's `EgressPolicy` rule matches and
carries an `inject_static_headers` effect, the gateway resolves the
referenced credential from a **credential provider** and sets it as a
request header (e.g. `Authorization: Bearer <token>`) before
re-originating upstream.
This implements the `applyEffects` in the egress handler. Injector runs
as part of the existing egress ext_proc handler that already
fetches/caches/evaluates each actor's `EgressPolicy`, so the MITM leg
gains injection without a second ext_proc hop.
### How it works
- New `CredentialProvider.FetchSecret` plugin API
(`pkg/proto/credproviderpb`): the gateway calls it with the credential
URI and the actor's attested SPIFFE identity; the provider authenticates
the gateway (mTLS) and resolves the secret.
- The egress handler dials the provider over mTLS
(`egress.DialProvider`) and injects on matched, allowed requests
(`egress.applyEffects`).
- Behavior:
- **TLS MITM leg + provider configured** → inject the credential
(overwriting any actor-set header).
- **Cleartext leg, or no provider configured** → skip injection and pass
the request through (never put a secret on a cleartext wire; don't block
allowed egress).
- **Attempted but failed** (unfetchable secret, empty/malformed secret,
credential URI of an unserved provider class) → fail closed.
## Installation
One flag stands up and wires the whole stack:
```
hack/install-ate.sh --deploy-ate-system --experimental-egress-credential-injection
```
`--experimental-egress-credential-injection` deploys credprovider and
the injector, then re-wires the sdsmint egress gateway to route through
the injector. It implies `--experimental-use-sdsmint`
### New install flags
Two flags configure which credential provider the injector targets:
| Flag | Purpose | Default |
|---|---|---|
| `--credential-provider-name` | Provider class, as a `ate-secret://`
prefix; a policy URI of any other class is refused |
`ate-secret://kubernetes.io` |
| `--credential-provider-address` | Where the injector dials the
provider | `credprovider.ate-system.svc:50051` |
Fixes#1507
Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.
`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.
This PR is based directly on `main` and does not depend on #1521.
Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Run agentgateway data plane tests as a part of substrate CI
(non-blocking to start so we can confirm it's not flaky). Also, change
the `--atenet-router` flag to `--atenet-dataplane` to make it clearer
that the flag controls ingress and egress.
I've run the e2es locally across gVisor and microVM plus the MITM
variants for both. The only skip we do for agentgateway is
`TestIngressProtocolDowngrade` because 1. the behavior its testing only
exists on the non-CONNECT atunnel ingress path and agentgateway only
sends CONNECT to atunnel and 2. I'm not sure that we want this to be a
part of the contract that substrate is bound by (e.g. do we really want
to commit to atunnel always parsing HTTP?).
My goal with getting both dataplanes into CI is to start taking steps to
codify the proxy (router + egress PEP) contract for substrate. The
telemetry they emit, atunnel expectations, etc. are all important
contracts to explicitly call out so that they don't become too coupled
to a single dataplane implementation.
> It's a good idea to open an issue first for discussion.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
The egress handler so far authenticated the actor behind a CONNECT and
let everything through. It now authorizes the traffic against the actor's
EgressPolicy, on the leg where the destination is known. Which leg that
is comes from the filter chain name the dataplane asserts, the same
signal the mux already dispatches on.
* egress, the outer CONNECT, is the only leg that sees a certificate,
and the first to see the address the actor's kernel dialed. It keeps
authenticating that certificate and decides the address rules,
ip_blocks and all, against that address, once per connection. When
one of them allows it the handler hands the address back as dynamic
metadata, which is what lets the gateway dial it for traffic it
cannot read and what tells the legs inside what was dialed. When
none does but the policy names hostnames, the tunnel is opened and
the requests inside are left to the legs that will see a name. When
neither is true nothing inside the tunnel could ever be allowed, so
the CONNECT is refused. A dataplane that calls out for the CONNECT
alone sends no chain name and has no such legs behind it, so it gets
the whole decision here and is refused rather than deferred.
* egress_tls_mitm and egress_cleartext are the legs the gateway can
read. Every request is decided the way the API describes: the rules
in order, over the request's Host and the address the actor dialed,
first match wins. It runs per request because the Host can change
between requests on one connection. The answer also says where the
request goes, so the bytes reach what the matching rule checked: a
hostname match is resolved and dialed by name, an address or all
match goes to the dialed address. Sending an address match by name
would let an actor put any Host on an allowed address; sending a
hostname match to the dialed address would turn the rule into a
header check, since the actor picks both. A Host header that
disagrees with :authority is refused rather than policed on one name
and dialed on the other.
Identity on the request legs is the actor's SPIFFE ID, which the outer
chain sets as filter state from the certificate Envoy verified and
shares across the internal-listener hop; nothing inside the tunnel can
write it, and a callout without it is refused. The certificate's URI SAN
must name the same actor as its ActorIdentity extension, because the
CONNECT authenticates on the extension while the request legs attribute
traffic to the SAN.
Policies are read through the per-actor cache, whose TTL is a new
--egress-policy-cache-ttl flag defaulting to 10s, so a request inside a
tunnel is not a control-plane round trip.
Credential injection is not implemented yet. A matched rule that declares
one denies with 501 rather than forwarding without the secret the policy
promised.
Affects get <name>, create, resume, pause, and suspend for actors, actor
templates, atespaces, tags, and workers: JSON/YAML no longer wraps a
single resource in a List, matching kubectl's convention. Listing (no
name, or multiple names) still returns a List.
Also plumb the writer through cmd.OutOrStdout() instead of hardcoding
os.Stdout to honor Cobra's writer.
Multiple enhancements to setup-gcp
- Add boot disk customizations. Since snapshots are written to boot
disk, being able to increase the size will mitigate IO throttling.
- Require only one of project-id or project-number to be specified
- Verify region and cluster-location are compatible
* Fix the column names to match the resource fields.
* Rename template flag name (fix TODO from CRD -> Substrate API
refactor).
* Fix bug in `create actor` which conflated the atespace flag for both
actor and actor template.
Fixes#794
A crashed actor makes `--delete-all` abort mid-teardown, leaving the
ActorTemplate, atespace and `ate-system` behind; the next
`--deploy-demo-counter` keeps the existing template, so the cluster goes
on serving the stale golden snapshot.
`prepare_actor_for_delete` treats `STATUS_CRASHED` as unexpected and
returns 1, which errexit turns into a hard abort, even though
DeleteActor accepts CRASHED and is idempotent for an actor already
marked DELETING (`workflow_delete.go`). Teardown now treats crashed and
already-deleting actors as deletable as-is.
The demo flag dispatch had the mirror-image defect: handlers ran from an
`if` condition, which suppresses errexit for their whole call tree, so
this same teardown and failed deploys like the issue's immutable-spec
apply exited 0. Handlers now run as plain commands and report an
unclaimed flag through `ate_demo_flag_unhandled`; the Jupyter handler
follows the same contract.
Repro: deploy the counter demo, resume an actor, delete its worker pod,
wait for the syncer to mark it `STATUS_CRASHED`, then run
`--delete-all`.
The changes we need are live on HEAD of the k8s published repos. It was
a bit of a journey, since Go fought me all the way.
---
### Drop our third_party fork of k8s deps in tools
The changes we need are released now (sort of - on HEAD anyway).
---
### Bump k8s.io/streaming to v0.37.0 (not rc) in tools
---
### Bump codegen deps to HEAD in tools
GOPROXY=direct go get \
k8s.io/apimachinery@master \
k8s.io/code-generator@master
This pins them to HEAD of master. Go is terrible here: The HEAD is not
actually tagged, so Go just uses the next "reachable" tag which is
v0.36.0-alpha. The datestamp is correct, though.
---
### Drop our third_party fork of k8s deps in root
The changes we need are released now (sort of - on HEAD anyway).
---
### Bump k8s deps to v0.37.0 (not rc) in root
---
### Bump apimachinery dep in root to HEAD in root
GOPROXY=direct go mod edit -replace
k8s.io/apimachinery=k8s.io/apimachinery@master
GOPROXY=direct go mod tidy
GOPROXY=direct go mod vendor
GOPROXY=direct go mod tidy
This approach (-replace) is needed because Go is horrible here.
The master branch of k8s.io/apimachinery is not tagged, per se, but
there is an OLDER tag which is "reachable" from HEAD. So go helpfully
decides to use that (v0.36.0-alpha.2). If we just `go get ... @master`
it works for that dep (pinned to the right date) but then it looks at
transitive deps. Because the tag seems to be 0.36 (older), it
recalculates all the OTHER dependencies and downgrades a whole tangle of
things to versions that match 0.36, but we are ACTUALLY on 0.37+.
This was the only approach that I (and Gemini) could find. Blech.
---
### Run updated codegens
---
### Use DV's new maxBytes capability for `[]byte`
Removes 1 custom.
Fixes#664
This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489
It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).
Now, an external snapshot is owned by a single resource:
- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.
Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.
this PR:
- Drops table actor_snapshots
- Keeps table actor_snapshot_tags
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1481
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #1266
This is a draft of the core API + data store changes.
It's still a large PR, apologies.
The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.
Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.
They're loadbearing for performance once we actually have multi-actor
workers.
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.
The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).
For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The Create Cluster warning explained the beta-API creation-time
constraint but not which GKE versions work. Clarifies the two supported
configurations: **GKE 1.36 with the beta APIs enabled at cluster
creation**, or **GKE 1.37+** where `certificates.k8s.io/v1beta1` is
served by default (no beta enablement needed). Versions below 1.36 are
unsupported.
- [x] Tests pass (docs only)
- [x] Appropriate changes to documentation are included in the PR
---
Also updates `hack/ate-dev-env.sh.example`: it pinned
`1.35.5-gke.1163012` — below the supported floor and no longer offered,
so copying it verbatim failed cluster creation. Now pins the newest
offered 1.36 patch (`1.36.3-gke.1767000`) with a comment on how to pick
a fresh one when the pin ages out.
**Validated**: `setup-gcp create cluster` at `1.36.3-gke.1767000`
(us-west1-c) succeeded, and both `certificates.k8s.io/v1beta1` APIs
(`clustertrustbundles`, `podcertificaterequests`) are served on the
resulting cluster — `kubectl get clustertrustbundles` returns the
kube-apiserver-serving bundle. Throwaway cluster deleted after
verification.
---------
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.
To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.
### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:**
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
Three fixes from user feedback on the substrate-gke installer experience, consolidated per reviewer preference (replaces #1405, #1406, #1407):
**1. setup-gcp: enable `artifactregistry.googleapis.com` in bootstrap.** Bootstrap grants `roles/artifactregistry.reader` to the GKE node service account and the atelet workload-identity principal, and ko pushes the control-plane images through gcr.io's Artifact Registry backing — but `enableRequiredAPIs` never enabled the API, so a fresh project failed at image push/pull instead of step 1.
**2. teardown.sh: cover everything bootstrap creates, without a dev-env file.**
- Revoke atelet's project-level bindings (`roles/storage.objectAdmin`, `roles/artifactregistry.reader`), which `grant_atelet_permissions` adds but nothing removed.
- Delete the three Substrate monitoring dashboards, matched by the display names in `tools/setup-gcp/dashboards/`.
- Accept configuration from the environment when `.ate-dev-env.sh` is absent (missing variables are named specifically), so installers and one-liners can drive the script.
- Drop the kubectl cluster-admin precheck: every step talks to GCP, not the cluster, and it blocked tearing down a cluster that was already gone.
- `--all` now runs in the true reverse of bootstrap's setup order.
**3. commands.md: the `hack/install-demo-*.sh` scripts still exist.** The closing paragraph claimed they are gone, but `hack/install-ate.sh` still sources them and generates its `--deploy-demo-NAME` / `--delete-demo-NAME` flags from them.
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.
## Summary
This change replaces startup schema setup with embedded, versioned SQL
migrations.
`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.
Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.
Closes#901.
Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing
## Migration behavior
`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.
One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.
If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.
Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.
## Changes
- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.
## Migration file policy
Migration files use sequential versions and contain exactly one Goose
`Up` section.
CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.
Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.
After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.
Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.
## Compatibility
No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.
Users must recreate databases from earlier PostgreSQL development
builds.
Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.
A binary rollback does not roll back the database schema.
## Testing
Tests cover:
- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
Fixes#368 . Deletes ActorTemplate CRD and any references to it.
Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Follow-up to the review comments on #1288 (merged before they could land
there).
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Changes:
-
`hack/third_party/csi-driver-host-path/deploy/csi-hostpath-testing.yaml`:
same fix as #1288 — the inline `watched_directory` was a no-op, so the
serving cert and trust bundle move to filesystem SDS (SPIFFE SAN pin
kept inline via `combined_validation_context`). SDS requires a node id
and this sidecar passes no `--service-node`, so the bootstrap gains a
`node:` stanza; the `docs/csi-deployment.md` example had the same gap
and gets one too.
- `overlay_test.go` now asserts `watched_directory` appears nowhere in
the emitted cluster block or the patched `envoy.yaml` (comments
excluded) and is present in each SDS resource file, per review.
- Corrected the test comment claiming the sdsmint e2e suite is skipped
in CI: the sdsmint variant is deployed in the MITM lanes; only the
`--experimental-additional-egress-extproc-service` injection has no e2e
coverage.
Verification: `go test ./cmd/ate-setup/...` passes; `envoy --mode
validate` (v1.34) passes on the reworked csi-hostpath bootstrap with
dummy PEMs mounted at the referenced paths.
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.
Verifications done:
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).
This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.
### One derivation for the version label
`internal/versionlabel` turns the build version into its two forms in
one place:
| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |
A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.
### A DaemonSet per version
`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.
### The install becomes version-aware
- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.
### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.
Fix#1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
This PR is large, I grouped changes to the following commits:
- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.
Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no
bin/activate) and an interpreter upgrade (dangling bin/python3) as a
good venv, so the script dies in `source venv/bin/activate` and the only
fix is knowing to delete the directory. Probe that the venv runs, and
rebuild with --clear to relink the interpreter.
Also install requirements unconditionally, which is cheap. The license
verifier skipped the install whenever the venv already existed, so it
passed without ever seeing a newly added dependency.
Fixes failure encountered by @ahmedtd. Not filing an issue because this
is pretty trivial.
AI-assisted.
Installing the control plane pulls roughly 570MB of third-party images
(postgres, prometheus, the otel collector, rustfs, envoy, jaeger) onto a
single node.
kubelet serializes image pulls by default, so they queue behind one
another and whichever workload lands at the back of the queue can miss
its readiness deadline.
e2e has failed repeatedly with postgres still in PodInitializing after
60s, its image not yet pulled, while everything else was already
running.
Sample failure:
https://github.com/agent-substrate/substrate/actions/runs/33258860897/job/99117220741
Note: This is after PR merge, and it's clearly unrelated to the PR.
AI-assisted.
ate-controller creates a WorkerPool's Deployment, so it does not exist
yet when the apply returns.
`kubectl rollout status` reads the object before it starts watching and
errors on a missing one instead of waiting, so whether these waits
worked at all came down to beating the controller by a round trip.
e2e lost that race by 178ms, failing the deploy with NotFound while the
pool's pods were already being created.
Fixes yet another flake I observed here:
https://github.com/agent-substrate/substrate/actions/runs/33275880827/job/99162382080?pr=1283
Part of #207, following the plan discussed there with @juli4n and
@kannon92: plumbing plus the non invasive fixes first, with the existing
invasive findings excluded so new APIs do not regress (as @kannon92
suggested in
https://github.com/agent-substrate/substrate/issues/207#issuecomment-4921051580).
The linter itself was originally suggested by @BenTheElder in #188.
Plumbing:
- hack/tools/kube-api-linter module with the
golangci-lint-kube-api-linter tool, same pattern as the other tools.
- .golangci-kal.yaml config. The pre-existing findings from rules that
need Go API changes (nomaps, nonpointerstructs, nophase, optionalfields,
requiredfields, ssatags) are excluded rather than disabled, and each
exclusion rule names the fields it excuses, so the rules still apply to
new API files and to new fields on the existing types. Fixing them on
the existing types is the follow up in #207.
- hack/verify/kube-api-linter.sh runs it against ./pkg/api/... and is
picked up by hack/verify-all.sh in CI.
Fixes in pkg/api/v1alpha1 (doc and marker only, no Go API change):
- commentstart: field godocs now start with the serialized field name.
- conditions: Conditions moved to the first position in
ActorTemplateStatus and got the listType, listMapKey and patch markers.
- defaultorrequired: SandboxConfigSpec.SandboxClass had both a default
and required; now optional with omitempty, matching the same field on
WorkerPoolSpec and ActorTemplateSpec. The ValidatingAdmissionPolicy in
sandboxconfig-validation.yaml is unaffected since CRD defaulting runs
before admission, so spec.sandboxClass is always set by then.
- optionalorrequired: missing +optional markers added on
ActorTemplateStatus fields.
- defaults: configured preferredDefaultMarker to kubebuilder:default
since CRDs are generated with controller-gen.
Generated CRDs regenerated. Structural schema changes are only the
conditions listType/listMapKey and sandboxClass no longer in the
required list (it is defaulted, so behavior is the same); the rest is
description text.
Note: my earlier count of 50 issues in #207 was capped by the default
golangci-lint issue limit. With the cap removed the real total is 86;
the extra ones are requiredfields (20) and ssatags (4), both in the
excluded set above.
Test plan
- hack/verify/kube-api-linter.sh passes (0 issues).
- Exclusions verified both ways: dropping them brings the pre-existing
findings back, and a scratch field added to WorkerPoolSpec is still
reported (optionalfields), which the earlier per-file exclusion
swallowed.
- TestSandboxConfigValidation passes against the regenerated CRD.
- go build ./... and go test ./pkg/api/... ./cmd/atecontroller/... pass.
- gofmt, shellcheck and the regular golangci-lint on pkg/api are clean.