Commit Graph
55 Commits
Author SHA1 Message Date
Luiz Oliveira 34f3003326 Add DV for GoldenSnapshotStatus and unbounded strings (#1831)
Fixes  #1168

* Add DV for GoldenSnapshotStatus and truncate its error_message to fit
* Require ExternalSnapshot.actor_template_uid to be a UUID
* Bound ExternalVolumeTemplate.capacity to 32 characters
2026-09-24 16:59:12 +00:00
Luiz Oliveira fb4b31529a Rename SnapshotsConfig to SnapshotConfig everywhere in the codebase (#1791)
related to -> #1378

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 14:39:09 +00:00
Julian Gutierrez Oschmann d277088bc1 Rename ActorTemplate.containers.readyz to wakeupProbe. (#1794)
The name `readyz` is not very descriptive. We also want to avoid
`readinessProbe` to prevent confusion with the Kubernetes concept, which
represents continuous traffic gating. Renaming this field to
`wakeupProbe` clarifies its actual behavior, and the fact that it only
operates when the actor is woken up.

Part of #1378
2026-09-22 23:32:01 +00:00
Julian Gutierrez Oschmann 6ec93a45f0 Make apitool validate fail if the exemptions are not sorted. (#1798)
The exemptions file is generated / maintained by the `apitool` so we
want to be strict and fail if it's not sorted (the tool uses
deterministic ordering already).

This is to prevent cases where entries are manually edited and
subsequent calls to `apitool validate --update` generate spurious
changes (like what happened in commit 1e56e66b).

Also update the exemptions to make sure the entries are properly sorted.
2026-09-22 18:33:27 +00:00
Shruti Nair b29f97778a Initialize OpenFGA server on ATE API Server bootstrap. (#1670)
Working on https://github.com/agent-substrate/substrate/issues/1563

## Summary
Initializes an embedded OpenFGA server within `ateapi` backed by
PostgreSQL and updates the top-level authorization scope.

## Key Changes
- **Rename scope to `global`**: Renamed `type cluster` and
`parent_cluster` to `type global` and `parent_global`, adding
`can_set_policy` and `can_get_policy` to `global`.
- **Embedded OpenFGA server (`internal/authz/server.go`)**:
  - Compiles `model.fga` into OpenFGA's protobuf representation.
  - Runs OpenFGA PostgreSQL schema migrations (`goose_db_version`).
- Serializes startup via `pg_advisory_lock` to avoid multi-replica
races.
- Connects OpenFGA's PostgreSQL adapter with a new dedicated
`*pgxpool.Pool`.
- Idempotently creates or reuses the `"substrate"` store and
authorization model.
- **Service wiring (`cmd/ateapi`)**: Exposes `Pool()` on
`*atepg.Persistence` and initializes `authz.NewServer` on bootstrap.

## Verification
- `openfga model test`: 103/103 checks passing (`model_test.fga.yaml`).
- `go test -mod=mod ./internal/authz/...`: Passes against PostgreSQL
(migrations, tuple writes, permission checks, and restart idempotency).
2026-09-18 23:30:41 +00:00
Huy Pham 9f57d78f4f Add --enable-nested-virtualization to setup-gcp (#1669)
Nested virtualization exposes `/dev/kvm`, which microVMs need. The new
`--enable-nested-virtualization` flag turns it on for the node pool
`setup-gcp` creates. It defaults to on, so pass
`--enable-nested-virtualization=false` for a cluster that does not need
it.

Also, rename the env var `GVISOR_NODE_MACHINE_TYPE` to
`NODE_MACHINE_TYPE` to be consistent with other env vars and to avoid
confusion for microVMs. The old env var still works with a warning.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-18 04:27:38 +00:00
Haven Xia d0a606f8d6 docs: tighten the upgrade runbook checks and rollback list
Phrase the first warning as the mistake, like the other two. Add
rollback entries for the new DaemonSet and the cloned pools, so an
operator who stopped before flipping any node finds their step in the
list. Say what the progress commands do not show instead of claiming
that steps leave no trace. Wait for a controller-triggered roll to
settle before cloning pools. Note that a single-node cluster makes the
per-node step a full stop. The setup-gcp README's eviction warning
still said 60 seconds; it is 30 minutes.
2026-09-14 12:28:30 -07:00
Michelle Au bce9e7e229 Enhancements to setup-gcp (#1616)
Multiple enhancements to setup-gcp

- Add boot disk customizations. Since snapshots are written to boot
disk, being able to increase the size will mitigate IO throttling.
- Require only one of project-id or project-number to be specified
- Verify region and cluster-location are compatible
2026-09-11 14:51:49 -07:00
Maya Wang ab0714729c docs: warn that worker node pools must not auto-upgrade or use spot nodes (#1528)
Documentation half of #1526.

## What

Two additions, no code:

- `tools/setup-gcp/README.md`, Create Cluster: a prerequisite saying
node auto-upgrade must be off on any pool that runs workers, and that
those nodes must not be spot or preemptible. It sits next to the
beta-API warning, which is the block the quickstart already points
bring-your-own-cluster users at.
- `docs/upgrade.md`: the runbook assumptions now state the same thing,
and ground rule 2 says what actually happens to an actor when a serving
worker pool is edited, rather than only that you should not edit one.
Scaling a serving pool down is called out explicitly, since that is not
obviously "editing" it.

## Why

An actor that is awake when its worker pod goes away moves to
`ACTOR_STATE_CRASHED`. That state is terminal: `resume` and `suspend`
are both refused, there is no recover verb, and the `latestSnapshot` the
actor still holds cannot be used to start it. The actor has to be
deleted and recreated.

`buildCreateClusterRequest` sets no `Management` on the node pool, so
GKE's default applies and auto-upgrade is on. Nothing in the install
documentation mentions it. A cluster built exactly as documented can
therefore lose actors on Google's maintenance schedule, and the operator
has no way to know that in advance.

Found during acceptance testing on `release-0.1` at `c48b3a3c`, where a
`WorkerPool` scaled from 12 replicas to 9 destroyed two of ten actors
permanently. That scale-down was our mistake and ground rule 2 already
covered it, which is why this PR is scoped to documentation. What the
docs did not cover is that the consequence is permanent, or that a
trigger exists which fires without an operator doing anything.

## Scope

This does not fix the underlying behaviour and is not a substitute for
#1526. Auto-repair, preemption and OOM kills reach the same path and no
prerequisite can close them. Telling operators to disable a standard
Kubernetes safety feature is a reasonable answer for this release and
not a reasonable one for 1.0, so the precondition relaxation in #1526 is
still owed.

## Testing

Docs only. `make test` passes. `make verify` does not complete in my
environment: `hack/verify/codegen.sh` fails creating the Locust codegen
virtualenv because `python3-venv` is not installed locally, which is
unrelated to this change.
2026-09-08 14:54:02 -07:00
Max Thompson 72e7906ad1 ateapi: validate system-info projection paths at template creation (#1439)
`CreateActorTemplate` accepted `../escape` and `/escaped` in
`ActorMetadataItem.path` and `TrustBundleDataSource.path`, and accepted
several data sources projecting to the same path. atelet only rejects
the traversal at actor start, so the template is created but the golden
actor sits in `RESUMING`. atelet never rejects duplicate paths at all:
it writes the files in order and the last one wins.

This PR moves both checks to template validation.

- The path rule lives in a new `internal/volumepath` package that both
ateapi and atelet call, so the two cannot drift. It splits on `/` and
rejects any segment that is empty, `.`, or `..`, caps depth at 16
segments, and rejects NUL bytes. Any other byte is allowed, so Unicode
paths pass. `ActorMetadataItem.path` and `TrustBundleDataSource.path`
apply it through a `+k8s:customValidation` hook, and atelet's
`writeSystemInfoFile` applies it again before touching the host
filesystem.
- `SystemInfoVolumeSource.data_sources` gets a hook that tracks
projected paths in a set and reports a duplicate on repeats, in the same
shape as the Kubernetes projected-volume validation.

The proto doc comment on `data_sources` also says at most one
`actor_metadata` entry may appear. Nothing enforces that, and this PR
leaves it alone.

New cases in `TestValidateSystemInfoVolumeSource` and
`TestValidateTrustBundleDataSource` cover the rejected shapes, and
`internal/volumepath` has its own table test. `hack/verify/codegen.sh`
and `hack/verify/golangci-lint.sh` pass.

#1231 moves `writeSystemInfoFile` into `systeminfovolume.go` and onto
`os.Root`. Whichever PR lands second points that copy at the shared
rule.

Fixes #1436
Fixes #1438
2026-09-04 16:32:34 -07:00
Luiz Oliveira 1e56e66b25 Rename ActorSnapshotTag to Tag (#1491)
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.

https://github.com/agent-substrate/substrate/issues/664

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 18:00:36 -04:00
Luiz Oliveira 6378b707e3 Call the generated validation fns for ActorSnapshotTag methods (#1488)
Plus, add DV for the other methods that were missing and remove obsolete
custom validation functions

https://github.com/agent-substrate/substrate/issues/664
https://github.com/agent-substrate/substrate/issues/1168



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 17:06:05 -04:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Michelle Au bbacb1ba8c Disable Filestore CSI when creating a GKE cluster. (#1463)
A custom version is required for now.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 11:59:48 -07:00
Tim Hockin dbdc4d25e3 Comment fixes 2026-09-04 09:58:49 -07:00
Tim Hockin 3b01db4717 Small API fixups after PR 1303 merged (#1382) 2026-09-04 10:07:01 -04:00
Benjamin Elder 2c429a9906 multi-actor worker API (#1283)
Part of #1266 

This is a draft of the core API + data store changes.

It's still a large PR, apologies.

The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.

Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.

They're loadbearing for performance once we actually have multi-actor
workers.
2026-09-03 21:19:21 -07:00
Da Huang cdd86eba1c setup-gcp: fix atelet dashboard filters that match no deployment (#1454)
The atelet Deployment is always named "atelet-<suffix>" (see
cmd/ate-setup/internal/steps/version.go), never plain "atelet", so every
dashboard widget filtering on top_level_controller_name="atelet" renders
empty even when the underlying metrics exist. The Snapshot Size & QPS
dashboard was entirely blank on a fresh install because of this; the
gRPC and E2E-latency dashboards had the same dead filter on their atelet
panels.

Match the deployment name prefix with a regex instead. Verified against
a live install (GMP PromQL): the equality filter returns 0 series while
=~"atelet-.*" returns all 12 atelet.snapshot.size series.

Fixes #1453

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 18:39:20 -04:00
Anna PendletonandAnna Pendleton b8fccb7ed0 setup-gcp: validate iam flags before applying any binding (#1451)
`create iam` checked --bucket from inside the bucket-bindings block,
after the GKE node and atelet grants had already run. Invoking it
without a bucket therefore wrote two project-level IAM bindings and only
then failed with a usage error, leaving the project partly configured by
a command that had rejected its own arguments.

Move the checks into validateIamFlags and call it before the first
grant, so an argument error is reported without touching the project.

Reproduced with:

  setup-gcp create iam --project-id=P --project-number=N --bucket=

  before: two grants applied, then
          Error: --bucket is required for bucket bindings
after: Error: --bucket is required for bucket bindings, no grants
applied

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [X] Appropriate changes to documentation are included in the PR

Co-authored-by: Anna Pendleton <13793293+annapendleton@users.noreply.github.com>
2026-09-03 14:35:55 -07:00
Aditya ShantanuandAditya Shantanu 999dd3e449 docs: state the supported GKE versions for setup-gcp clusters (#1437)
The Create Cluster warning explained the beta-API creation-time
constraint but not which GKE versions work. Clarifies the two supported
configurations: **GKE 1.36 with the beta APIs enabled at cluster
creation**, or **GKE 1.37+** where `certificates.k8s.io/v1beta1` is
served by default (no beta enablement needed). Versions below 1.36 are
unsupported.

- [x] Tests pass (docs only)
- [x] Appropriate changes to documentation are included in the PR

---

Also updates `hack/ate-dev-env.sh.example`: it pinned
`1.35.5-gke.1163012` — below the supported floor and no longer offered,
so copying it verbatim failed cluster creation. Now pins the newest
offered 1.36 patch (`1.36.3-gke.1767000`) with a comment on how to pick
a fresh one when the pin ages out.

**Validated**: `setup-gcp create cluster` at `1.36.3-gke.1767000`
(us-west1-c) succeeded, and both `certificates.k8s.io/v1beta1` APIs
(`clustertrustbundles`, `podcertificaterequests`) are served on the
resulting cluster — `kubectl get clustertrustbundles` returns the
kube-apiserver-serving bundle. Throwaway cluster deleted after
verification.

---------

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-09-03 14:16:59 -07:00
shrutiyam-glitch 616fd83431 feat: Support Cloud SQL via Auth Proxy for PostgreSQL backend (#996)
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.

To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.

### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:** 
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.


- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
2026-09-03 15:45:55 -04:00
Aditya Shantanu cbae8250a5 Address substrate-gke installer feedback: AR API, teardown coverage, doc fix (#1408)
Three fixes from user feedback on the substrate-gke installer experience, consolidated per reviewer preference (replaces #1405, #1406, #1407):

**1. setup-gcp: enable `artifactregistry.googleapis.com` in bootstrap.** Bootstrap grants `roles/artifactregistry.reader` to the GKE node service account and the atelet workload-identity principal, and ko pushes the control-plane images through gcr.io's Artifact Registry backing — but `enableRequiredAPIs` never enabled the API, so a fresh project failed at image push/pull instead of step 1.

**2. teardown.sh: cover everything bootstrap creates, without a dev-env file.**
- Revoke atelet's project-level bindings (`roles/storage.objectAdmin`, `roles/artifactregistry.reader`), which `grant_atelet_permissions` adds but nothing removed.
- Delete the three Substrate monitoring dashboards, matched by the display names in `tools/setup-gcp/dashboards/`.
- Accept configuration from the environment when `.ate-dev-env.sh` is absent (missing variables are named specifically), so installers and one-liners can drive the script.
- Drop the kubectl cluster-admin precheck: every step talks to GCP, not the cluster, and it blocked tearing down a cluster that was already gone.
- `--all` now runs in the true reverse of bootstrap's setup order.

**3. commands.md: the `hack/install-demo-*.sh` scripts still exist.** The closing paragraph claimed they are gone, but `hack/install-ate.sh` still sources them and generates its `--deploy-demo-NAME` / `--delete-demo-NAME` flags from them.
2026-09-02 11:39:58 -07:00
Benjamin Elder 1d58b95a7f update grpc 1.83.2 in all modules (#1385)
picks up security fixes
2026-09-01 21:37:24 -07:00
shrutiyam-glitchandMichelle Au 880456d2a8 Extend declarative validation to Actor (#1244)
Follow up on issue #1168 and the base PR #1215 
In this PR: 
**Actor read/lifecycle verbs — completes Actor end-to-end**
* `Get`/`Delete`/`Suspend`/`Pause`/`Resume` `ActorRequest`: dropped the
`+k8s:opaqueType` fence on the actor ref, added `+k8s:required` +
`+k8s:subfield(atespace)=+k8s:required`. Combined with the recursion
into `Validate_ObjectRef` this reproduces exactly what
`resources.ValidateObjectRef` enforced (ref required, atespace required,
name required, both DNS-1123 labels).
* `ListActorsRequest`: tagged to match `ListAtespacesRequest` (atespace
optional + short-name format, `page_size` minimum 1, `page_token` max
length 256).
* All six `validate*Request()` functions are now one-line calls to the
generated validators. With this, every RPC verb on `Actor` and
`Atespace` is fully declarative; `resources.ValidateObjectRef` has no
callers left in `actor.go`.
  
**Note:**`ListActorsRequest.page_token` now capped at 256 chars (parity
with `ListAtespaces`).
  
**Testing**
* Existing verb tests updated to the generated error shapes
(`field.ErrorMatcher` with origin matching).
*   New coverage: `page_token` boundary cases for `ListActors`; 
* `controlapi`, `functionaltest`, `actoridentity`, `store`, and `atepg`
suites all pass; `go vet` and `gofmt` clean.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Michelle Au <msau42@users.noreply.github.com>
2026-09-01 17:28:19 -04:00
Haven Xia 9df8ca6abe [Upgrade] Version the install: node labels and a per-version atelet DaemonSet (#1357)
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).

This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.

### One derivation for the version label

`internal/versionlabel` turns the build version into its two forms in
one place:

| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |

A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.

### A DaemonSet per version

`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.

### The install becomes version-aware

- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.

### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.

Fix #1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
2026-09-01 12:40:49 -07:00
shrutiyam-glitch 159de80b59 Declarative validation for ActorTemplate (#1303)
Continues the declarative-validation (DV) migration from #1215, covering
the `ActorTemplate` resource, its full spec tree, the create path, and
read/delete verbs.

**Key Changes**
* **DV Tags & Hooks**: Added validation tags and custom hooks across the
`ActorTemplate` spec (`Metadata`, `SandboxConfig`, `SnapshotsConfig`,
`Container`, `Volume`, `Resources`).
* **Create Path**: Replaced hand-written validation with generated
validators in `CreateActorTemplate`. Custom rules (e.g., `on_commit ⊆
on_pause`) are now handled via custom hooks.
* **Read/Delete Verbs**: Converted `Get`, `List`, and `Delete` requests
to use DV.
* **State Updates**: Refactored `atepg` metadata setters to update
in-place.
* Regenerated apitool exemptions (documented 15 previously-exempt
fields).

**Deliberate Behavior Changes**
The gRPC path now strictly enforces CRD rules. Specific tightenings
include:
* Empty `worker_selector` is now rejected.
* `EnvVar.name` is strictly required.
* Negative values in optional enum fields are rejected (`minimum=1`).
* `page_token` is capped at 256 characters (matching other list
requests).

**Deliberately Not Done**
* **No update DV / status tags**: `ActorTemplates` are immutable to
clients. The reconciler updates the status directly against the store,
so tagging the status subtree or adding update validation isn't
necessary right now.
* **`config_name` matching `sandbox_class`**: Enforcing this requires
calling SandboxConfigLister (ServiceImpl), so it's left as a TODO.

**Testing**
* Added ~70 positive and negative test cases in the validator tables. 
* Store-backed tests (`TestCreateActorTemplate*`,
`TestUpdateActorTemplateMetadata`, and the `atepg` suite) verified
against real Postgres via testcontainers.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-01 15:21:07 -04:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
shrutiyam-glitch c575055646 extend declarative validation to Workers (#1250)
Follow up on issue #1168  and the base PR #1215
 
**Overview:** 
Continues the declarative validation (DV) migration from #1215 by
converting the `Worker` resource. This moves immutable-field enforcement
from the storage layer up to the service layer. With `WorkerPoolSyncer`
now using RPCs, nearly all write paths share the exact same validation.

**Key Changes:**
* **Schema (`ateapi.proto`):** Added full DV tags (required, format,
immutable) to `Worker` pod-coordinates and metadata.
`WorkerStatus.state` is now strictly bounded to its enum range.
* **Service Layer:** Replaced ~60 lines of hand-written validation with
1-line generated calls. `CreateWorker` and `UpdateWorker` now scrub
server-owned fields and enforce immutability *before* hitting the store.
* **Storage Layer:** Removed `store.CheckWorkerMutation`.
`atepg.UpdateWorker` is optimized to only clone metadata instead of the
whole worker. Immutability contract tests were moved to the service
layer.
*   **Behavior Tweaks (Tightenings):** 
    *   `page_token` is now capped at 256 chars.
    *   `DeleteOptions.uid` must be a valid UUID.
* Global-ref atespace violations now correctly return `Forbidden`
instead of `Invalid`.
* Immutability checks now natively cover the `WorkerPoolSyncer` path.
* **Testing:** Added comprehensive positive/negative cases for all newly
tagged fields, boundary cases, and moved the immutability test suite.
All suites pass cleanly.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-01 10:34:18 -04:00
Da Huang dbed137302 Fix dashboard label names (#1362)
Fixes #1361

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-01 10:05:48 -04:00
Haven Xia 698289452b Bump Go to 1.27 2026-08-30 15:43:52 -07:00
Zoe Zhao e1adb33147 api: drop ActorTemplateStatus.sandbox_assets for now (#1300)
Remove `ActorTemplateStatus.sandbox_assets` along with the
`SandboxAssets` messages it referenced. Nothing ever wrote the field:
sandbox assets are resolved from the WorkerPool and SandboxConfig
objects at resume time and travel to atelet via ateletpb, so freezing
them into the template status never materialized.

This requires more thought, one option is to add it as a field of
GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a
repeated field to allow multiple goldens.
2026-08-29 10:11:49 -07:00
Julian Gutierrez Oschmann cb7c8385ef Make 'documented' rule handle validation-gen tags. (#1286)
When checking documentation, remove `validation-gen` tags so we don't
have false negative results if a field *only* has tags and no
documentation.
2026-08-28 10:09:08 -07:00
Aditya Shantanu ab12d95445 docs: k8s beta APIs must be enabled at GKE cluster creation
certificates.k8s.io/v1beta1 podcertificaterequests and
clustertrustbundles cannot be enabled in place on an existing cluster -
the update is accepted but the APIs never become served, and the
install hangs waiting for ClusterTrustBundles. Warn in the create
cluster docs, show the bring-your-own-cluster flag, and name the
symptom.
2026-08-28 08:43:30 -07:00
Aditya Shantanu 536dc3f51b docs: document the atelet Workload Identity grants setup-gcp creates
On a fresh GCP project nothing documented which IAM bindings atelet
needs: setup-gcp creates them silently, and anyone who cannot run the
tool with project-level IAM permissions - or needs to audit what it
did - had to read cmd/iam.go and cmd/bucket.go. Spell out the exact
members, roles, and resources, the Workload Identity prerequisites,
and the gcloud equivalents, and point the README quickstart at it.
2026-08-28 08:43:30 -07:00
Benjamin Elder b2c9459521 apitool: drop the exemptions for fields that are now documented (#1279)
https://github.com/agent-substrate/substrate/commit/48090979dc56324a9985a52de4557243b287d964#diff-51a595730fdeea0b1efc0394d1b3e866426a9acb000b6f2110262867b9aaeae7
added exemptions unrelated to the PR

https://github.com/agent-substrate/substrate/pull/1276 removed the need
for exemptions, which were currently failing on main

This PR cleans up the now failing defunct exemptions.
2026-08-27 22:06:26 -04:00
Zoe Zhao 48090979dc Update Actor Workflows to use substrate proto ActorTemplate (#1264)
Create/Suspend/Resume Actor workflows now uses the
`Actor.actor_template` ObjectRef when specified, otherwise falls back to
the CRD in cmd/ateapi/internal/controlapi/template_convert.go.
2026-08-27 19:35:50 -04:00
Julian Gutierrez Oschmann f43d4096d7 Allow apitool to declare exemptions and enable linter during presubmit (#1274)
Allow the `apitool validate` linter to define exemptions, then add
exemptions for current validation errors, and then enable this as part
of presubmit (GH action).
2026-08-27 12:49:24 -07:00
Walter Fender dad1ea3f5f Switch default region/zone from us-central1(-c) to us-west1(-c). (#1249)
Switch default location from us-central1 to us-west1.
2026-08-27 10:11:48 -07:00
Julian Gutierrez Oschmann a80edf4ce3 Introduce a apitool CLI for the substrate API. (#1162)
The idea is to bundle automation related to the API here. For now, there
is only a single `validate` command that runs a set of lint rules to
enforce that the API adheres to the style guide. This should help us
prevent regressions / deviations when modifying it.

These are not enforced as part of the presubmit checks, although we
might want to enforce that soon (right after we close the existing gaps?
see #1163). There is no support for violation exemptions for now.

In the future, we could add another command to, for example, generate a
reference documentation for the whole API.
2026-08-26 10:20:35 -07:00
igooch aeeca4df7c validate-image-cache: drive eviction through Store.EvictUnused + GC loop review follow-ups (#837)
Phase 2 of #463, following #735 and #836: `validate-image-cache` drives
eviction through `Store.EvictUnused`, deleting its pre-engine prototype.
Also carries #836's post-merge review follow-ups.

## Tool
- **Prototype evictor deleted** (mtime-sorted tree removal, then
dropping every manifest record — no refcounts, no two-phase rename, no
restore). The tool now asks the engine to reclaim the shortfall below
`--min-free-gb`, so corpus runs exercise production semantics.
`--evict-idle` maps to `WithMinAge`.
- **`--evict-all`**: one-shot flush, no refs file or auth. Unlike `rm
-rf`, running actors' images survive via bundle-spec rooting
(`WithActorsDir`). It is not synchronized with a running atelet (the
engine's locks are per-process): the pool can't be corrupted, but a
layer reused mid-pass can be evicted from under an actor — a
not-yet-mounted start fails once and heals; an already-mounted actor can
take EIO. Documented, warned at runtime; the intended end state is an
atelet-owned flush (RPC or trigger), not a second process in the pool.
- **`--evict-idle` floors at 1m when an actors dir exists**: min-age is
the only protection that applies across processes, so on a live node it
can't be tuned away. Validation hosts (no actors dir) keep full freedom.
- Gate-aware errors in both eviction paths (`ErrIncompleteEnumeration` =
"did nothing, repair the named path" vs per-item errors = stats still
print, exit 1); a 30s cooldown after fruitless passes so workers don't
serialize full engine passes when nothing is evictable.

## Carried #836 follow-ups
- Negative `--image-cache-gc-period` rejected (was silently disabling
the loop); `noteOutcome` extracted with the backoff cadence unit-tested;
engine per-pass line demoted to DEBUG (no-target ticks are now silent);
single-line sentinel wraps; `filepath.Abs` + table test for the
outside-BasePath warning; `runPass`/`Run` covered behind a `gcStore`
seam; flag help and README made truthful (`period=0` still runs startup
recovery; watermark reads ~5 pts above `df`; the total-size cap still
evicts everything evictable under sustained foreign pressure — retention
floor named as the extension).

## Testing
- Unit: `-race` clean; new tests for backoff cadence, flag validation,
path check, `runPass` skip/panic paths, and `Run` (immediate first pass
pinned deterministically; ticking covered separately).
- Kind: full e2e green; gated-pass arc verified live (one ERROR per tick
naming the corrupt record; startup scan gates and recovers; planted
orphan reclaimed at restart); `--evict-all` on a live node evicted 4
unrooted images while all rooted images survived.
- GKE (6 nodes): zero image-cache log lines fleet-wide at default log
level across ticks.
- **Corpus sweep** (7.9 GB volume, always-low-water, `--evict-idle=10s`
on a host with no actors dir, parallel pulls): 64/64 images validated
with two eviction episodes mid-run (62 img / 103 layers / 5.2 GB, then
46 img / 65 layers / 2.7 GB) — continuous eviction under concurrent real
pulls, zero races.
2026-08-12 10:34:22 -07:00
Krisztian F 0dbe1523e4 feat(otel): add cold start metrics (#776)
This PR implements the last two metrics from #433.

Right now, we can see that a resume was slow but not where.
`ate.actor.lifecycle.operation.duration` covers the whole ateapi
operation, and `atenet.router.route.duration` covers the edge, but
everything between ateapi-atelet-actors is one block that contains
fetching the manifest, downloading the snapshot, unpack the OCI image,
call to ateom.

We have `rpc.server.call.duration` that gives us the atelet restore
total time, but template, kind, and scope labels are missing, so today
we cannot really pinpoint why/where we have a regressions in latency.

In this PR I am adding per-phase histograms, here's an example of what
we can know after these changes:

```console
ateom_restore   522 ms   ###############################
download        8.9 ms   #
manifest_fetch  3.2 ms
oci_unpack      2.6 ms
---------------------------------------------------------
total           535 ms
```

It also fixes a gap #683 opened where a `data_on_golden` resume was
labeled identically to a plain one on the lifecycle histogram.

Things folks might want to argue with:
- total as a phase value. Partly duplicates `rpc.server.call.duration`,
but that one has no domain labels and gRPC-specifc. We can drop it, but
it's an inferior operational UX, so I'd rather have it here.
- Phases overlap, they are not a partition of total, because download
runs concurrently with the asset fetch and unpack. I called this out in
the metric description. Do not sum across phases.
- New `ate.snapshot.scope` key rather than a new `ate.snapshot.kind`
value for `data_on_golden`. A new value would collapse local and
external into one bucket, which is the biggest latency difference there
is. This does add a label to the already shipped lifecycle histogram.

Verified on kind, and all e2e suites pass, and the emitted series cover
every kind (golden, latest, local) on both metrics with no unknown
values.
 
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-07 10:15:19 -04:00
dberkov ee0dfd790f validate-image-cache: make eviction idle window configurable
The fixed 30-minute idle guard made nothing evictable while a fast corpus
filled a small disk (32-VM hicard sweep: local SSDs filled in ~12 minutes,
then every remaining image failed ENOSPC). Expose it as --evict-idle.
2026-07-22 21:15:39 -07:00
dberkov 610e3916ad Add node-local OCI image layer cache (imagecache), replacing memorypullcache
Content-addressed pool of unpacked image layers, shared by every actor on
the node; actor rootfs becomes an overlayfs mount (cached layers as read-
only lowers, bundle-local upper) instead of a full re-untar per run.
atelet (no capabilities) pulls and unpacks; the privileged ateoms finalize
whiteouts and mount. Tag refs resolve via one HEAD and become cacheable;
pull memory is O(stream buffers); the cache survives restarts.

Phase 1 of #463. Fixes #437, #166, #228.

Validated: kind + GKE counter demos (gvisor and microvm), suspend/resume
(oci_unpack ~3ms vs ~15-20s), 411 SWE-bench-scale images pulled and
unpacked with 0 failures, root-gated unit tests for the privileged paths.
Known gaps: no GC yet (Phase 2, see internal/imagecache/README.md);
upgrade ordering — deploy new ateoms before/with the new atelet.
2026-07-22 21:15:39 -07:00
John Howard a814761b6a Avoid redundant object naming suffixes (#441)
Best practices is not to put a -deployment deployment, -role role, etc;
the kind already declares the kind we don't need it in the name as well.
Additionally we did not do it consistently anyways.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-15 16:31:06 -07:00
Zoe Zhao 680dbea475 Update default GKE cluster version and enable Managed OTel feature (#352)
1) Always create a cluster with GKE Managed OpenTelemetry feature and
2) Update the default cluster version to 1.35.5-gke.1163012 which is the
current [regular channel default
version](https://docs.cloud.google.com/kubernetes-engine/docs/release-notes#current_versions)
to fix https://github.com/agent-substrate/substrate/issues/341

Tested:
* Successfully recreated GKE cluster, redeployed ate-system and ran
counter demo.
2026-06-30 16:20:05 -07:00
Bowei Du feb7e2cdc2 Enable dataplane v2 for network policy support
This enables dataplane v2 by default for the provisioned clusters so we
have a network policy provider.

Make getEnv generic
2026-06-26 14:11:22 -07:00
Bowei Du 35150f2508 Update documentation references 2026-06-26 10:18:53 -07:00
Bowei Du c945363566 Enable setup-gcp to take command line flags
- Document commands in README.md
- Restructure commands to be more idiomatic to Cobra
2026-06-26 10:18:53 -07:00
Haven Xia 91d1a75260 Move monitoring/dashboards under tools/setup-gcp. (#209)
Fixes #204

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-06-09 21:25:44 -04:00
Haven Xia 67bb44d090 Add Substrate Snapshot size metric and dashboard (#160)
This PR introduces a new OpenTelemetry histogram
`"atelet.snapshot.size"` from `atelet` over the existing
OTLP path. It records the uncompressed size in bytes of each gVisor
snapshot image written during checkpoint -- `checkpoint.img` ,
`pages.img` and `pages_meta.img` -- tagged by image `kind` and
`ActorTemplate`. This surfaces how large actor snapshots are and which
`ActorTemplate`s are expensive to checkpoint/restore.

Add a well-defined dashboard for snapshot metrics that contains 3
charts.
- snapshot image size p99 by `ActorTemplate`
- snapshot image size p50/p95/p99 (all kinds)
- snapshot QPS by `ActorTemplate`

<img width="996" height="717" alt="image"
src="https://github.com/user-attachments/assets/99bbab46-37dc-476e-bfd9-cdffeb174b4b"
/>


> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-06-05 11:01:49 -07:00