213 Commits
Author SHA1 Message Date
Joe Betz c544897d07 atepg: Guardrail to disallow changes to atespace tables that limit future partitioning (#1467)
This adds a test that serves as a guardrail to prevent changes to the
atespace-scoped tables that would make it difficult or impossible to
partition it in the future.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

cc @bowei @BenTheElder @juli4n @shrutiyam-glitch @thockin
2026-09-09 15:59:24 -04:00
Tim Bai e0e4bd5d86 docs: document the per-actor usage events channel (#1559)
Documents the per-actor usage events channel that #1206 added, closing
the docs request from that review.

New `Per-Actor Usage Events` section in the logging guide:
- an example record and the consumer contract (filter on `msg` + `kind`;
identity rides the same label group as lifecycle events and container
logs, so the guide's existing query dimensions apply unchanged, and one
`labels."ate.actor.uid"` filter returns an actor's output, transitions,
and usage interleaved);
- per-field semantics, including the part consumers must not get wrong:
`memory_current_bytes`/`memory_working_set_bytes` are point-in-time,
while `memory_peak_bytes`/`cpu_usage_usec` accumulate within an epoch
whose boundary depends on `source` (cgroup restarts on restore,
guest-agent survives it) — so window CPU is the increase between
samples, never a sum;
- the sampling knob (`--actor-stats-poll-interval`) and the delivery
contract (best-effort behind a bounded queue, independent of
`--log-level`);
- the cardinality rule in bold: log-based metrics over these events must
never label by actor identity.

The metrics section now names the `ate.actor.stats.*` instruments in its
registry pointer and cross-links here for per-actor detail.

Every technical claim is checked against the code and the
`WorkloadStatsSample` proto contract (epoch scoping, the
`memory.peak`/Linux 5.19 caveat, trace-context absence, drop-warning
text, flag semantics).
2026-09-09 15:07:23 -04:00
Jeff Luo 9ad967a1f3 metrics: bound ate.sandbox.class with one rule for every emitter (#1484)
Fixes #1475

`Control.CreateWorker` does not validate `Worker.sandbox_class`, so a
client can register a worker with an empty one (and any record written
before this keeps it). `RegisterWorkerCount` tallied that raw value, so
`ate.workerpool.workers` emitted `ate_sandbox_class=""` — not a member
of the `ate.sandbox.class` registry vocabulary — alongside the pool's
seeded series at 0.

This change normalizes the worker's class with the existing
`ateattr.NormalizeSandboxClass`, so an empty class reports as `unknown`.

New `ateattr.SandboxClassAttribute`, used by every emitter. This closes
the same gap on `ate.actor.crashes` and the `ate.actor.lifecycle.*`.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-09 18:53:42 +02:00
eliranw 353c21f203 docs: correct micro-VM rootfs description to host-backed overlay (#1531)
The architecture guide still described micro-VM container rootfs writes
as living in guest RAM via a tmpfs overlay.
#846 moved that overlay onto the host, served over the same single
virtio-fs share as the durable-dir volumes.
This updates the one stale sentence to match. 
Docs only.

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
2026-09-08 16:56:59 -07:00
Eitan Yarmush ee8d8faf09 Use Tag UIDs in snapshot storage paths (#1521)
Fixes #1508

Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.

- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.

Lint and code-generation verification also passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-08 18:27:12 -04:00
Maya Wang ab0714729c docs: warn that worker node pools must not auto-upgrade or use spot nodes (#1528)
Documentation half of #1526.

## What

Two additions, no code:

- `tools/setup-gcp/README.md`, Create Cluster: a prerequisite saying
node auto-upgrade must be off on any pool that runs workers, and that
those nodes must not be spot or preemptible. It sits next to the
beta-API warning, which is the block the quickstart already points
bring-your-own-cluster users at.
- `docs/upgrade.md`: the runbook assumptions now state the same thing,
and ground rule 2 says what actually happens to an actor when a serving
worker pool is edited, rather than only that you should not edit one.
Scaling a serving pool down is called out explicitly, since that is not
obviously "editing" it.

## Why

An actor that is awake when its worker pod goes away moves to
`ACTOR_STATE_CRASHED`. That state is terminal: `resume` and `suspend`
are both refused, there is no recover verb, and the `latestSnapshot` the
actor still holds cannot be used to start it. The actor has to be
deleted and recreated.

`buildCreateClusterRequest` sets no `Management` on the node pool, so
GKE's default applies and auto-upgrade is on. Nothing in the install
documentation mentions it. A cluster built exactly as documented can
therefore lose actors on Google's maintenance schedule, and the operator
has no way to know that in advance.

Found during acceptance testing on `release-0.1` at `c48b3a3c`, where a
`WorkerPool` scaled from 12 replicas to 9 destroyed two of ten actors
permanently. That scale-down was our mistake and ground rule 2 already
covered it, which is why this PR is scoped to documentation. What the
docs did not cover is that the consequence is permanent, or that a
trigger exists which fires without an operator doing anything.

## Scope

This does not fix the underlying behaviour and is not a substitute for
#1526. Auto-repair, preemption and OOM kills reach the same path and no
prerequisite can close them. Telling operators to disable a standard
Kubernetes safety feature is a reasonable answer for this release and
not a reasonable one for 1.0, so the precondition relaxation in #1526 is
still owed.

## Testing

Docs only. `make test` passes. `make verify` does not complete in my
environment: `hack/verify/codegen.sh` fails creating the Locust codegen
virtualenv because `python3-venv` is not installed locally, which is
unrelated to this change.
2026-09-08 14:54:02 -07:00
Youssuf Elshall 9ea39a5639 docs: remove stale references to ActorTemplate as a Kubernetes CRD (#1404)
## What

Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.

Fixes #368 (docs side).

> This change was prepared with AI assistance; I have reviewed and
tested it.

- [x] Docs and comment-only change; no functional code changed, no tests
affected
2026-09-04 16:40:04 -07:00
Max Thompson 72e7906ad1 ateapi: validate system-info projection paths at template creation (#1439)
`CreateActorTemplate` accepted `../escape` and `/escaped` in
`ActorMetadataItem.path` and `TrustBundleDataSource.path`, and accepted
several data sources projecting to the same path. atelet only rejects
the traversal at actor start, so the template is created but the golden
actor sits in `RESUMING`. atelet never rejects duplicate paths at all:
it writes the files in order and the last one wins.

This PR moves both checks to template validation.

- The path rule lives in a new `internal/volumepath` package that both
ateapi and atelet call, so the two cannot drift. It splits on `/` and
rejects any segment that is empty, `.`, or `..`, caps depth at 16
segments, and rejects NUL bytes. Any other byte is allowed, so Unicode
paths pass. `ActorMetadataItem.path` and `TrustBundleDataSource.path`
apply it through a `+k8s:customValidation` hook, and atelet's
`writeSystemInfoFile` applies it again before touching the host
filesystem.
- `SystemInfoVolumeSource.data_sources` gets a hook that tracks
projected paths in a set and reports a duplicate on repeats, in the same
shape as the Kubernetes projected-volume validation.

The proto doc comment on `data_sources` also says at most one
`actor_metadata` entry may appear. Nothing enforces that, and this PR
leaves it alone.

New cases in `TestValidateSystemInfoVolumeSource` and
`TestValidateTrustBundleDataSource` cover the rejected shapes, and
`internal/volumepath` has its own table test. `hack/verify/codegen.sh`
and `hack/verify/golangci-lint.sh` pass.

#1231 moves `writeSystemInfoFile` into `systeminfovolume.go` and onto
`os.Root`. Whichever PR lands second points that copy at the shared
rule.

Fixes #1436
Fixes #1438
2026-09-04 16:32:34 -07:00
Luiz Oliveira 1e56e66b25 Rename ActorSnapshotTag to Tag (#1491)
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.

https://github.com/agent-substrate/substrate/issues/664

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 18:00:36 -04:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
Krisztian F 645be8957c (otel): split infra vs. workload faults (#1445)
`ate.failure.reason` only ever named infrastructure faults, and the
crash log didn't say which actor crashed.

This PR adds the following:

- **`ate.failure.domain`** (`infrastructure` / `workload` / `unknown`)
is available next to every `ate.failure.reason`. It's a strict function
of the reason so it costs no series, but it's *emitted* rather than
derived downstream, so a component ahead of ateapi can report a reason
this build rejects, which collapses to `UNKNOWN`, and a consumer
matching on reason names would file every one of those as infra.
`UNKNOWN` now reports `unknown`, not `infrastructure`.
- **`Actor crashed`** record from ateapi, with full identity including
the uid, via a helper both crash sites share. It sits under the same
guard as the counter so the two can't disagree on how many crashes
happened. The counter is barred from carrying actor identity, so this
record is the only way to attribute a crash to one agent.
- **`WORKLOAD_NOT_READY`** on the readyz deadline, as the first
workload-domain reason. Deadline branch only: a cancellation is ateom
draining, not the actor failing.

Tested locally:
```
desc = WORKLOAD_NOT_READY: readyz for "counter" never returned 200 within 5s
```

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:53:27 -04:00
shrutiyam-glitch 616fd83431 feat: Support Cloud SQL via Auth Proxy for PostgreSQL backend (#996)
### Description
This PR introduces native, secure support for using Cloud SQL as the
PostgreSQL store backend for `ate-api-server`.

To ensure the highest level of security and ease of use in GCP
environments, this integration leverages the Cloud SQL Auth Proxy
sidecar with automatic IAM database authentication. This means transport
security (TLS 1.3 tunnel) is handled automatically, and database
sessions are authenticated using Workload Identity via short-lived OAuth
tokens, completely eliminating the need for database passwords.

### Key Changes
* **Cloud SQL Auth Proxy Sidecar:** Added
`manifests/ate-install/cloudsql-proxy-patch.yaml` to patch the sidecar
into the `ate-api-server` deployment when a Cloud SQL instance is
configured.
* **Automated Provisioning:** Extended `tools/setup-gcp` with a new
`cloudsql` command. This handles the idempotent creation of the Cloud
SQL instance, Google Service Accounts (GSA), IAM bindings, and Workload
Identity bindings.
* **Installation Script Updates:** Updated `hack/install-ate.sh` to
parse new environment variables (e.g.,
`ATE_API_POSTGRES_CLOUDSQL_INSTANCE`, `ATE_API_POSTGRES_CLOUDSQL_GSA`)
and correctly synthesize the passwordless DSN and ConfigMaps for the
proxy.
* **Security & Documentation:** 
* Added extensive documentation in `tools/setup-gcp/cloud-sql.md`
covering provisioning, schema privileges, deployment, and database
scaling.
* Updated `docs/threat-model.md` to reflect the new Cloud SQL egress
flows and Auth Proxy tunnel mechanics.
* **Dependencies:** Vendored required Google API clients (`sqladmin/v1`,
`servicenetworking/v1`, `iam/v1`) for the GCP setup tool.


- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
2026-09-03 15:45:55 -04:00
Haven Xia 2cd494384f docs: manual rolling upgrade runbook (clone workerpool approach) (#1392)
The runbook for moving a running substrate to a new build without losing
actor state. An user runs the roll by hand with `kubectl`, `kubectl
ate`, `ate-setup`, `jq`, and `grpcurl`, and every piece of upgrade state
lives in cluster objects, so the roll can stop and resume at any point.

The order follows the [upgrade
design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A):
CRDs, then `ate-controller`, then the dataplane node by node, then
`ate-api-server` and `atenet`.

The dataplane moves by version label. The new atelet DaemonSet sits next
to the old one, each serving WorkerPool is cloned with the new worker
image and the new version pin, and each node is drained (workers marked
`DRAINING` through the `DrainWorker` RPC), emptied by suspending its
actors, relabeled, and cleared of old worker pods. The old objects stay
untouched until a separate retire step, so rollback is one label flip
per node.

Tested end to end locally on a GKE cluster: install at one build, run
counter actors, upgrade to a second build following the document, and
confirm the actors resume on the new pool with their state intact.
Rollback and retire were exercised the same way.

Fix #1272, fix #1273.
2026-09-02 17:28:03 -07:00
haiyanmeng 539a17bf50 Guide for enabling man-in-the-middle (MITM) interception for Actor Egress policy (#1226)
An sdsmint egress gateway terminates every TLS connection an actor opens
and re-originates it, so what the actor validates is a per-SNI leaf the
gateway minted rather than the origin's certificate. That leaf chains to
the gateway CA and to no public root, so an actor left on its default
trust store fails every HTTPS request it makes.

Fixes #1005 

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 16:04:49 -07:00
Zoe Zhao 6340d8712a test: drop the stale Volume.Type assertion from the counter demo test (#1411)
#1398 removed the Volume.Type union discriminator, but #1008's counter
demo test still asserts the rendered manifest contains `type:
ExternalVolumeTemplate`, so run-tests fails on every main push since
both landed. Also drops the same stale line from the csi-volumes.md
example.
2026-09-02 13:17:49 -07:00
Julian Gutierrez Oschmann 6170b5f4fc Add a section to API style guide about union types. 2026-09-02 11:16:54 -07:00
Dmitry Berkovich 37da0cddbb Update gVisor to the 2026-09-02 nightly and switch to zstd tarballs (#1403)
## Summary

Improves the cold-node activation performance tracked in #811: the
gVisor release extract on a cold node currently exceeds the router
resume budget, so the first resume after a release bump always fails.

- Point the default gvisor SandboxConfig at the 2026-09-02 nightly,
using the `gvisor.tar.zstd` tarballs the nightlies now publish alongside
`gvisor.tar.bz2`. The win is decode speed, not just download size:
stdlib bzip2 is a pure-Go, single-threaded decoder, and extracting the
~165 MB release costs a node roughly 20 seconds of CPU on the first
actor operation it hosts — the dominant term of a cold-node activation.
With zstd the same extraction takes ~2.9 s, measured on a GKE node
running this change. The tarball is also ~22% smaller to download (129
MB vs 166 MB for x86_64).
- Teach `extractTarArchive` to decompress `.tar.zst` / `.tar.zstd`
archives using `github.com/klauspost/compress/zstd` (already a
dependency via ategcs), with test coverage for both suffixes.
- Update the `gvisor.tar.bz2` mentions in docs, manifests comments, and
the SandboxConfig API comment (generated CRD refreshed via
`hack/update/codegen.sh`).

The bucket only publishes sha512 checksums; the pinned sha256 sums were
computed from downloads verified against the published sha512 sums.

## Test plan

- `go test ./cmd/atelet/` (includes new `.tar.zst` / `.tar.zstd`
extraction cases)
- `go build ./cmd/atelet/...`, `go vet`, `gofmt` clean
- Verified the real x86_64 tarball lists `runsc` plus the `gvisor-bin/`
helpers
- Deployed to a GKE cluster: atelet fetched and extracted the 2026-09-02
x86_64 zstd tarball in ~2.9 s (vs ~20 s for bz2)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-02 11:10:52 -07:00
Benjamin Elder 7535180e5e Drop device passthrough until the API is designed
A GPU worker pool injected its devices into every container of every actor on it,
which only held up while a worker ran one actor.

A resource name and a count cannot replace it: that shape cannot express which
container gets a device, sharing one between containers, or resuming onto hardware
compatible with the snapshot. A nvidia.com/gpu pool limit is now an ordinary
extended resource -- it places the pod on a GPU node and nothing else reads it.

internal/cdi stays, unused: whatever the API becomes, an ateom still has to merge
CDI edits into an actor's bundle itself.
2026-09-02 10:27:14 -07:00
Jeremy Alvis f100c27286 ateapi: add versioned PostgreSQL schema migrations (#1196)
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.

## Summary

This change replaces startup schema setup with embedded, versioned SQL
migrations.

`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.

Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.

Closes #901.

Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing

## Migration behavior

`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.

One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.

If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.

Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.

## Changes

- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.

## Migration file policy

Migration files use sequential versions and contain exactly one Goose
`Up` section.

CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.

Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.

After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.

Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.

## Compatibility

No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.

Users must recreate databases from earlier PostgreSQL development
builds.

Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.

A binary rollback does not roll back the database schema.

## Testing

Tests cover:

- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
2026-09-02 10:56:19 -04:00
Zoe Zhao 120a519606 Delete ActorTemplate CRD (#1376)
Fixes #368 . Deletes ActorTemplate CRD and any references to it.

Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 10:48:42 -04:00
Chuang Wang 07ebc56797 csi-hostpath: fix cert rotation in the testing Envoy config (#1336)
Follow-up to the review comments on #1288 (merged before they could land
there).

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Changes:
-
`hack/third_party/csi-driver-host-path/deploy/csi-hostpath-testing.yaml`:
same fix as #1288 — the inline `watched_directory` was a no-op, so the
serving cert and trust bundle move to filesystem SDS (SPIFFE SAN pin
kept inline via `combined_validation_context`). SDS requires a node id
and this sidecar passes no `--service-node`, so the bootstrap gains a
`node:` stanza; the `docs/csi-deployment.md` example had the same gap
and gets one too.
- `overlay_test.go` now asserts `watched_directory` appears nowhere in
the emitted cluster block or the patched `envoy.yaml` (comments
excluded) and is present in each SDS resource file, per review.
- Corrected the test comment claiming the sdsmint e2e suite is skipped
in CI: the sdsmint variant is deployed in the MITM lanes; only the
`--experimental-additional-egress-extproc-service` injection has no e2e
coverage.

Verification: `go test ./cmd/ate-setup/...` passes; `envoy --mode
validate` (v1.34) passes on the reworked csi-hostpath bootstrap with
dummy PEMs mounted at the referenced paths.
2026-09-01 20:01:22 -04:00
Elijah Rodriguez 53800cd045 Update CSI NFS server image to use kubernetes build 2026-09-01 14:57:42 -07:00
Zoe Zhao f6852b7754 Update existing demos and benchmark tests to use substrate ActorTemplate resource (#1355)
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.

Verifications done: 
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
2026-09-01 14:50:18 -07:00
Joe Betz 3efac91283 api: delete the DebugClear RPC and the Debug service (#1346)
Fixes #999.
2026-09-01 17:46:05 -04:00
Krisztian F 85d404a823 (atelet): make the restore timing log a joinable per-actor latency record (#1364)
Restore timing breakdown is the only unsampled per-actor latency record
we emit, and since traces are head-sampled at 1% on the data plane, and
actor identity is not available from metric labels on purpose.

This adds this to logs:

 Before:
```json
 {"msg":"Restore timing breakdown",
  "actor":{"type":"*ateapipb.Actor","atespace":"demo","name":"counter-1"},
  "download":310000000,"oci_unpack":50000000,"ateom_restore":60000000,"total":420000000}
```

 After:
```json
 {"msg":"Restore timing breakdown",
  "ate.atespace":"ate-demo-counter-substrate","ate.actor.name":"join-probe",
  "ate.actor.uid":"3a5b4e86-1a77-4da7-84c4-c0993d9c83ec",
  "ate.template.atespace":"ate-demo-counter-substrate","ate.template.name":"counter",
  "ate.snapshot.scope":"full","ate.snapshot.kind":"golden","ate.sandbox.class":"gvisor",
  "ate.actor.restore.duration.volume_mount":0.000004583,
  "ate.actor.restore.duration.manifest_fetch":0.006111636,
  "ate.actor.restore.duration.sandbox_assets":0.000098958,
  "ate.actor.restore.duration.download":0.025386544,
  "ate.actor.restore.duration.oci_unpack":0.000954919,
  "ate.actor.restore.duration.ateom_restore":0.147448298,
  "ate.actor.restore.duration.total":0.181189689,
  "trace_id":"c9d7ea9a5f90a5eb4399bfa0b966d5ac","span_id":"89280ba761137c8e"}
```

The duration keys are the `ate.actor.restore.duration` instrument's name
plus an `ate.snapshot.phase` value, so a log key and the histogram it
mirrors are literally the same identifier in source, which is also why
they're seconds.

Also adds `ateattr.ActorLogAttrs`, the component-log twin of
`ActorLogLabels`, held to the same key set by a test.

 Verified on kind:
- the ateom Actor restoring record for the same actor carries the
identical `ate.actor.uid` and `trace_id`, this was impossible before
- `ate_actor_restore_duration_seconds_sum{phase="total"} = 0.181189689`,
matching the log to the nanosecond
- induced a download failure by deleting a snapshot object: record sets
`ate.failure.reason:` `FAILED_GET_EXTERNAL_OBJECT` and correctly omitted
`ateom_restore`; the metric for the same restore shows the same six
phases


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-01 16:18:28 -04:00
Haven Xia 9df8ca6abe [Upgrade] Version the install: node labels and a per-version atelet DaemonSet (#1357)
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).

This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.

### One derivation for the version label

`internal/versionlabel` turns the build version into its two forms in
one place:

| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |

A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.

### A DaemonSet per version

`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.

### The install becomes version-aware

- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.

### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.

Fix #1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
2026-09-01 12:40:49 -07:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
Krisztian F 39df566dfa (otel): atelet and ateom telemetry now says which node it came from (#1363)
We want to be able to answer questions like:
 - image cache hit ratio per node
 - cold starts got slower on one node

but nothing we emit says which node anything is on, this PR fixes this.

Verified on a local kind cluster, where both atelet and ateom-gvisor now
show `k8s_node_name` on `target_info`, including ateom's, which arrives
through the OTLP relay.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-01 11:02:53 -04:00
Zoe Zhao d4fc52b924 Report the actor template atespace in OTel tags (#1343)
* Renames the`ate.template.namespace` attribute to
`ate.template.atespace` and sources it from the actor's `actor_template`
ObjectRef instead of the legacy CRD namespace/name pair.
* Updated functional tests now to use substrate ActorTemplates and bind
actors through the `actor_template` ref.
2026-08-31 16:47:33 -04:00
Haven Xia 35344a2c68 api: make WorkerPool workerImage required again (#1334)
We should not make workerImage optional as it provide explicit
information for what image workers use, removing it will need
`atecontroller` to inject workerImage and leads to unnecessary control
plane & dataplane wiring.



Fixes #861 

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-31 10:19:41 -07:00
Tim Hockin 955a79f290 Add a doc on using DV to validate protobuf (#1327)
API reviewers should memorize this :)

@eitanya @juli4n @BenTheElder @dberkov
2026-08-31 13:10:13 -04:00
Chuang Wang d909d69053 atenet-egress: deliver TLS certs via filesystem SDS so rotation works
Envoy honors watched_directory only on SDS-delivered secrets; on the
egress gateway's inline file-based certs it was silently ignored, so
kubelet's projected-certificate rotation never reached Envoy and the
gateway eventually served an expired certificate.

Move the serving cert (both egress manifests) and the spliced extproc
cluster's client cert and trust bundle (Go and shell emitters) to
filesystem SDS, with the flag-derived SAN pin kept inline via a
combined validation context. Fix the csi-deployment doc example that
showed the same broken pattern, and add tests for the emitters.
2026-08-28 14:52:33 -07:00
Haven Xia 7face2c4d2 Rename the rest ateomImage to workerImage (#1298)
Some new file added also need the rename.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 13:56:05 -07:00
Haven Xia bdf494999b api: rename ateomImage to workerImage and make it optional (#1210)
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of #861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 12:46:00 -07:00
nybidari f45a0561c0 Bump default gvisor sandbox config to 2026-08-28 nightly (#1287)
Contains the gVisor fixes for #1090 #992
2026-08-28 12:15:26 -07:00
Jing 8eae7ed6a3 Emit a runtime-neutral OCI spec. (#1188)
atelet's spec was written for runsc: a hardcoded `runsc` hostname, CRI
pause/sandbox annotations and `dev.gvisor.spec.mount.*` hints. The
micro-VM shim compensated by rewriting every spec at runtime
(ensureKataCompatibleSpec), swapping the whole mount set and rebuilding
it from one hand-written builder per volume kind — so a volume kind
atelet learned about reached the guest only once the shim was taught
about it too.

atelet now emits a spec that names no runtime, and each ateom shapes it
for the runtime it drives. Whatever atelet adds therefore reaches both,
and a bind the micro-VM shaper cannot place in the guest is an error
rather than a silent drop.

Fixes #709.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-27 16:38:25 -07:00
hajiler b4529fcdb7 docs: add guide for CSI based external volumes (#1072)
#232

Add documentation guide for CSI volumes in Substrate


- [X ] Tests pass
- [ X] Appropriate changes to documentation are included in the PR
2026-08-27 16:03:53 -07:00
Benjamin Elder cb72d60bdf ateom microvm: drop privileged: true, using atelet device plugin for KVM (#1254)
This is the companion to #496 

In order to grant access to `/dev/kvm` we have to either:
- use a device plugin
- use a DRA driver
- use an NRI plugin

An NRI plugin is highly privileged in it's own right for all pods on the
host and is difficult to ship portably at the moment.
DRA is promising, but not enough functionality is GA yet at our current
1.35+ target.
Device plugin fits reasonably well. We do wind up publishing an
~arbitrarily high limit, which has some cost in kubelet memory, but
otherwise is relatively clean.

This approach is what kata uses currently. Their device plugin is not
available unbundled, and we anyhow have a per-node daemonset.

atlet is taught to sniff if /dev/kvm appears on the host at all, so we
can also stop using the manually labeled nodes for microVM class and
instead schedule to the KVM + TUN devices on nodes that advertise them.

Later we can migrate to device plugin by using 

I implemented that already, but I don't think it's worth merging at the
moment. We would want consumable capacity to be on by default. We can
migrate later without changing the pod spec by using
`extendedResourceName`.

https://github.com/agent-substrate/substrate/compare/main...BenTheElder:substrate:ateom-microvm-dra

NOTE: I confirmed with upstream that device plugin is not going anywhere
despite being "v1beta1", it's GA in all but name. It won't receive new
features but we don't really need anyhow. We'll move to DRA down the
line.

---

By doing this, we can drop `privileged: true` from the uVM ateom pods.

We can also drop the `ate.dev/sandboxClass` node label hacks, reducing
friction to deploy.
2026-08-27 12:23:16 -07:00
Jeff Luo 75c9b145fb docs: point at the new names of the verify scripts (#1236)
The rename of `hack/verify/verify-metrics.sh` to
`hack/verify/metrics.sh` left three references to the old path behind. A
reader who copies one of them gets a "No such file or directory" error.

Point `AGENTS.md`, `docs/observability.md`, and the header comment of
`docs/metrics/registry/metrics.yaml` at the new name.

Docs only; no change to the script or to the registry data.

A follow-up of https://github.com/agent-substrate/substrate/pull/1221
2026-08-26 21:51:40 -07:00
eliranw afcd5cd718 Per-container CPU and memory limits for micro-VM actors (#859)
Per-container CPU and memory limits for micro-VM actors. First slice of
#752.

An `ActorTemplate` container can cap its own cpu and memory so it cannot
starve or kill its siblings in the same actor. A container that exceeds
its memory limit is OOM-killed on its own; the rest of the actor is
unaffected.

```yaml
sandboxClass: microvm
containers:
  - name: trainer
    resources:
      limits: {memory: 1500Mi}
  - name: sidecar
    resources:
      limits: {memory: 256Mi, cpu: "0.2"}
```

#752 asks for per-container device selection as well. This PR does the
cpu and memory half; devices are the remaining half, and when they land,
GPU assignment should follow the same field rather than staying
actor-wide as it is today.

## Why micro-VM only

The two sandbox classes support opposite halves of #752. Micro-VM actors
have a real guest kernel, so each container gets its own cgroup and the
limits bind. gVisor applies cgroup limits at the sandbox level: one
sentry backs every container in the actor, so a per-container cgroup is
created and then stays empty (google/gvisor#190). Measured on a running
actor, the workload container's cgroup reported `memory.current=0` while
all 20 sandbox processes sat in the pause leaf. A template that sets
`resources` with `sandboxClass: gvisor` is rejected at admission.

## How a limit travels

`ActorTemplate` → CEL validation → `ate-api-server` resolves each
`resource.Quantity` once → `ateletpb` → atelet writes OCI
`linux.resources` → `ateom-microvm` merges kata's defaults and checks
the guest envelope → `SpecToAgentPB` → kata agent → guest cgroup.

The limit is carried as a standard OCI field rather than a
substrate-private concept, so a runtime that gains per-container
enforcement picks it up without new plumbing.

## Composition with #679

#679 sizes the sandbox itself from `ActorTemplate.spec.resources`: guest
RAM and vCPUs for a micro-VM, the sentry for gVisor. This PR subdivides
that sandbox. The two are different fields and complementary layers.

They collided in one place. #679 applied the actor-level size to
**every** container's OCI spec, overwriting the per-container limits
atelet writes into the same field, so a container asking for 64Mi
silently received the whole guest. This PR removes that call on the
micro-VM path. A container now gets a cgroup limit only when it declares
one; an undeclared container is bounded by guest RAM, which is the real
ceiling. On micro-VM the stamp never reached the guest at all:
`SpecToAgentPB` carried neither `Memory` nor `CPU.Quota`, so it stopped
at the bundle's `config.json`. Running the pre-PR pipeline shows
`Quota=50000 Memory=2Gi` on disk arriving as `Quota=0 Memory=<nil>` on
the wire.

gVisor is untouched. There the actor-level size is applied to the
`pause` container, and since one sentry backs every container, that is
the only cgroup that binds.

Two consequences worth calling out for reviewers of #679:

- The `sizing.SandboxSize` threaded into `buildActorContainers` and
`ensureKataCompatibleSpec` had no remaining reader, so it and the
`guestSize` helper are removed. `resolveGuestMemMiB` still sizes the VM
and still rejects a declared limit too small to boot.
- `internal/sizing`'s package doc claimed both runtimes shared
`ApplyToOCISpec` for container cgroups. That is now gVisor-only, and the
comments say so. No code in that file changed.

`checkResourceEnvelope` runs after `resolveGuestMemMiB`, so it validates
the per-container sum against the post-reserve guest rather than the
SandboxConfig default. Its error now names
`spec.resources.limits.memory` (or `.cpu`) when the actor declared a
size, and `SandboxConfig` when it did not.

`internal/e2e/suites/sizing` is unaffected: it asserts on `num_cpu` and
`mem_total_bytes`, both of which come from VM sizing, and only logs the
cgroup files.

## Verified on hardware

Same template, run twice, one commit apart on a micro-VM actor:

| | before | after |
| :--- | :--- | :--- |
| `hog_ovl/memory.max` | `max` | `67108864` (the declared 64Mi) |
| `bystander_ovl/memory.max` | `max` | `max` |
| hog allocates 128MB | survived | OOM-killed |
| bystander | alive | alive |

The bug in between: `SpecToAgentPB` converted only `Devices` and
`CPU.Shares` out of `Linux.Resources`, so a memory limit reached the
bundle's `config.json` and was dropped on the way to the agent. Every
unit test passed and the on-disk spec was correct while the feature did
nothing.

## Notes for review

- `ContainerResources` deliberately does not reuse
`corev1.ResourceRequirements`, which also carries `requests` and
`claims`. There is no scheduler inside an actor to hint at, and for
memory a soft request cannot express "must have this much to come back
at all". #679 reached the same conclusion from the other direction and
now rejects `spec.resources.requests` and `.claims` at admission.
- `Limits` is a bounded named type. A `MaxProperties` marker on the
field lands at the wrong schema level for a named map, and without a
bound the CEL cost estimator rejects the whole schema. This CRD is close
to its schema-wide CEL cost ceiling; a rule that iterates containers and
parses quantities exceeds it by more than 100x and takes the existing
rules down with it.
- A cpu limit below `10m` is raised to `10m`: the kernel rejects a CFS
quota under 1ms.
- `SpecToAgentPB` drops a non-positive quota and a zero period rather
than forwarding them. A non-positive quota means unlimited in OCI but
was sent as a literal zero, which the guest applies as no CPU at all; a
non-nil zero period overwrote the CFS default set for a live quota. Both
now read a spec the way `cpuLimitMillis` does, so the envelope check and
the conversion agree on the same input.
- `mergeKataResources` fills the gaps kata's defaults cover rather than
allowlisting known fields, so a field it does not recognise survives the
merge. That holds only as far as the merge: `SpecToAgentPB` converts
`Devices`, `Memory` and `CPU` only, so `Pids`, `BlockIO`,
`HugepageLimits` and `Network` stop at the ttrpc boundary. Nothing sets
them today.

## Known gaps

- An OOM-killed container is not reported above the guest. The actor
stays `STATUS_RUNNING` and nothing records which container died. Raised
on #550. Note that #961 settled workload stats at the sandbox level and
deliberately does not attribute per container, so surfacing which
container died needs a separate signal: the guest cgroup's
`memory.events`, not the stats path.
- A container's writable rootfs is a guest tmpfs, so filesystem writes
are charged to its memory limit and a container writing more than its
limit to `/tmp` is OOM-killed. Documented in `api-guide.md`; separating
the two budgets needs an `ephemeral-storage` field.
- Over-subscription across containers is caught when the actor starts,
not at apply time. An admission-time sum check is not implementable
within the CEL cost budget.
- The atelet hop that attaches resources to the bundle spec has no test:
replacing `ctr.GetResources()` with `nil` leaves the whole suite green.
Closing it needs an imagecache fixture for `prepareOCIDirectory`.
- Restored actors inherit the golden's cgroups rather than applying
their own spec. Correct today only because `ActorTemplateSpec` is
immutable.
- The GPU section of `api-guide.md` still says `ActorTemplate` has no
per-container resource fields, which this PR makes false. Correcting it
is left to the devices half of #752, which is what will change the GPU
behaviour that paragraph describes.

---

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
2026-08-25 16:43:46 -07:00
Jeff Luo 27a9412c10 observability: add a metric registry that CI checks (#1097)
`docs/observability.md` was the only description of the metrics of
Substrate, and nothing compared that text with the code. It listed 13
instruments; the code sends 21. The request-parking and actor
resource-usage instruments were
missing entirely.

This change adds an [OpenTelemetry
Weaver](https://github.com/open-telemetry/weaver) semantic convention
registry.

  ### What is in it
  
  | Path | Content |
  |---|---|
| `docs/metrics/registry/manifest.yaml` | The Weaver manifest: name and
schema URL. |
| `docs/metrics/registry/metrics.yaml` | 21 instruments and 12 attribute
groups: every label, its permitted values, and its buckets. |
  | `docs/metrics/substrate.yaml` | The rules Weaver cannot express. |
  | `hack/verify/verify-metrics.sh` | The check. |

Fixes #1093

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-25 15:49:37 -04:00
3a3ec62ff3 fix: land the three root-caused flaky-test fixes (identity, parking, relay) (#1160)
Fixes #675, fixes #1100, fixes #1146; addresses the CI flake in #1106.

## Why this PR

Flaky tests are the single biggest drag on this repo's velocity right
now: the three flakes fixed here account for the majority of red CI runs
over the last 7 days (identity: 35 failures, parking: 32, relay: 9 —
from the flake dashboard's cross-PR analysis of ~630 runs). Every red
run costs a contributor a rebase-and-rerun cycle and costs reviewers
signal. **This PR consolidates the three root-caused, in-flight fixes
into one change to get CI green now and unblock the community — the goal
is velocity, not authorship.**

## Credit where it's due

All three fixes were root-caused and written by others; this PR adopts
them onto latest main with their tests, unchanged in substance. Each
commit carries a `Co-authored-by` trailer:

| Commit | Original PR | Author | Root cause |
|---|---|---|---|
| e2e: give each probe fixture its own worker pool | #1147 |
@orangeCatDeveloper | identity/egressmitm/imagevolume suites share one
`workload: probe` pool label; cross-suite selection under concurrent
suite processes dials workers that are not there |
| atenet: never cancel an in-flight resume at the park budget | #991 |
@omeryahud | the park budget doubled as the ResumeActor RPC deadline; a
mid-restore cancel strands a RESUMING actor on a live worker |
| atunnel: close the relay's both ends before returning | #1101 |
@orangeCatDeveloper | the relay closed both ends from a
`context.AfterFunc` goroutine the test never waits for |

@Stevenjin8's #1107 correctly diagnosed the ateom readiness race in
#1106; the control-plane readiness gap it targets remains real and open
— this PR only removes the e2e-fixture contention that makes it fire
constantly in CI.

If maintainers prefer to land the original PRs individually instead,
closing this one is completely fine — the point is that the fixes land
somewhere, soon.

## Evidence the flakes are actually fixed

**TestRelayIngressCancellationClosesBothSides (unit, `-race`):**
- Unpatched main, `-count=3000`: **83 failures (2.8%)** — matches the
2.9% observed across 308 CI runs this week
- This branch, `-count=10000`: **0 failures**

**TestRequestParking (park-budget cancellation):**
- The new `InFlightAttemptRunsToCompletion` and
`LateRetryableErrorIsBudgetExhaustion` unit tests (from #991) encode the
exact failure mode from #675 and pass under `go test -race -count=100
./cmd/atenet/internal/router/ingress/`
- The pre-fix behavior (budget cancelling the in-flight RPC) is
deterministically reproduced by the old test it replaces

**TestActorIdentity_AfterRestore_IsOwnID_NotGolden (probe pool
isolation):**
- Not reproducible outside CI (needs concurrent suite processes on a
contended kind node), so verified statically: `${FIXTURE_SUFFIX}` is
always `-<suite>` (internal/e2e/sandbox.go:189,201 — never empty),
`probe-sized` already uses its own label, and no other manifest or
selector references `workload: probe`. #1147's CI data shows all three
failure signatures (missing `ateom.sock`, `runsc restore` killed, router
502/503) trace to cross-suite pool sharing; per-suite labels make the
selector suite-local by construction
- The definitive check is this PR's own CI plus the flake dashboard's
7-day window after merge — I will report the post-merge rates on #1106

Also run: `go build ./...`, `go vet` and the full `-race` suites of both
touched packages — all green.

## What this PR deliberately does NOT fix

`TestActorEgressHTTPS` (#1050, 4.6% this week, below the 5% flake
threshold) has no root-caused fix yet — the 503 `upstream connect error`
path needs investigation in a live cluster. #1103 (@orangeCatDeveloper)
tightens the related `TestActorArbitraryPortAccess` assertion so those
503s stop passing silently; it should land after #1050's cause is fixed,
or it converts hidden flakiness into visible red.


## Update (post-CI investigation)

The first e2e runs failed on `TestRequestParking/ParkThenServed`
(micro-VM lane). Investigation showed this is the **pre-existing
dominant mode** of #675 — identical failures in main-era runs
32305728993 / 32397291519 / 32487958077 — not a regression: on micro-VM,
`SuspendActor` returns before the snapshot upload completes, so the
worker legitimately isn't free within the 5s park budget and the
router's 503 is correct behavior. #991 fixes the *other* (mid-restore
cancellation/stranding) mode. Commit d637690d makes the subtest retry
while the worker is still freeing; a stranded worker still fails every
attempt, so the regression stays pinned.

**Additional validation:**
- CI e2e-test now **passes both lanes** (run 32754691017)
- Local kind cluster built from this branch: parking suite **10/10
consecutive passes**; identity + egressmitm + imagevolume run
**concurrently** (the exact contention behind the identity flake) × 3
iterations — **9/9 suite passes**

---------

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Co-authored-by: NekoPunch <engineer.jyao@gmail.com>
Co-authored-by: Omer Yahud <oyahud@nvidia.com>
2026-08-25 11:08:18 -07:00
Julian Gutierrez Oschmann fbe580e4f6 Fix resource names for ActorSnapshotTag methods. (#1167)
Also clarified the style-guide about the naming convention.

Fixes #1163
2026-08-25 09:48:12 -04:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Luiz Oliveira 31a2e3850a Simplify Update methods to do a whole object replace (#1108)
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.

#1011 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-24 15:18:55 -04:00
Haven Xia 640654df47 kubectl-ate: add --container log filters for actor logs (#1067)
Part of  #294 as #311 has been inactive for over a month.

Example

```
kubectl ate logs actors test -a demo                                      # default: all containers + lifecycle


kubectl ate logs actors test -a demo -c counter                           # only specified containers
```

Scoped this down to --container only after discussing with @BenTheElder
— we are not sure the naming of the container related logs, the
supervisor output may end up as a describe/events-style like how k8s
does?

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-21 19:06:40 -07:00
Max Thompson 074fd1c12b Add trustBundle as a SystemInfo volume data source (#941)
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data
source for SystemInfo volumes (#802) and the end-to-end proof that the
projected anchors work against the MITM egress gateway. Live refresh for
running actors (PR 2) and auto-injection (PR 3) come separately.

What this adds

A SystemInfo volume data source that projects the trust anchors of a
named trust bundle to a PEM file:

volumes:
- name: trust
  systemInfo:
    dataSources:
    - trustBundle:
        name: egress-mitm.ate.dev
        path: egress-ca.pem

Inspired by the Kubernetes clusterTrustBundle projected volume source,
but source-neutral: the template names a bundle; where it's fetched from
is a deployment concern, not part of the API.

Design points

- Resolution lives on the node. The wire carries only {name, path};
atelet resolves the name at write time through an informer-backed lister
on ClusterTrustBundles and writes the sanitized PEM with the temp+rename
discipline from #803 (find-paths safe). Contents refresh on every
Run/Restore. ateapi is not involved, per review discussion — the same
informer is what live refresh (PR 2) will hang off.
- Allowlist in atelet, not the CRD schema. Today only
egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the
ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946)
derives from the egress-mitm-ca-pool Secret. The signer-linked object
name stays a backend detail; the future backend registry (#932) widens
the allowlist without an API change.
- The watch is scoped to the one backing object via a metadata.name
field selector — this informer runs on every node, so an unfiltered
watch would fan every ClusterTrustBundle in the cluster out to every
atelet. RBAC can't express this (resourceNames doesn't apply to
list/watch), so the field selector is the enforcement point.
get/list/watch on clustertrustbundles moves to the atelet ClusterRole.
- No availability probe. The informer registers unconditionally; a
cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1
blocks atelet startup at cache sync, with the reflector errors naming
the missing API (hack/create-kind-cluster.sh enables the gate).
- Fail-closed. Unknown names, missing bundles, and unusable bundles fail
actor start naming the bundle — an actor that declared a trust bundle
must not start without one.
- Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks
only, deduplicated, headers stripped, and anchors deliberately shuffled
so consumers can't grow a dependence on order.
- Schema note: dataSources MaxItems tightened 32→8 while adding the
trustBundle member. Vacuous in practice (the old schema couldn't admit
more than one entry), but flagged since it's ratchet-shaped.

E2E — delivery and consumption

Delivery (identity suite, both sandbox classes): provisions the
egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing
the bundle directly isn't possible — the reconciler reverts hand-edits),
asserts the projected file byte-exact, then rotates the pool across a
suspend/resume to prove refresh-on-restore. Since the probe fixture is
shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists
for whatever suite deploys it.

Consumption (new egressmitm suite, both sandbox classes): deploys the
sdsmint (MITM) egress gateway and proves an actor completes a TLS
handshake with the gateway's per-SNI minted leaf using ONLY the
projected anchors — plus a system-roots negative control that must fail.
The pair is unambiguous in both directions: the positive can't pass
under passthrough (the bundle holds no public CAs), and the negative
can't fail under passthrough.

CI: two steps appended to the existing e2e job after the standard lanes
(the gateway swap is cluster-wide and breaks passthrough assumptions):
--deploy-atenet --experimental-use-sdsmint redeploys only the atenet
components, then the egressmitm suite runs once per sandbox class.

Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each
suite deploys its own copy and drives one actor at a time, so the third
worker per copy was idle memory multiplied across suites on the one-node
CI cluster — pressure that has been killing sandboxes mid-test (runsc:
signal: killed, a vanished ateom socket) on this PR and on main's
identity suite. This reduces the pressure; right-sizing e2e concurrency
or worker-pod QoS cluster-wide is follow-up material.

Not in this PR

- Live refresh for running actors (#932 PR 2) — until then, a running
actor's file is the bundle as of its last Run/Restore, and correctness
rests on overlap rotation by the bundle publisher.
- Auto-injection of the egress trust volume (#932 PR 3).
- Configurable backend registry (#932) — the allowlist is the seam it
will replace.
2026-08-21 19:53:31 -04:00
Sneha-at b7080602c6 Allow deleting an actor from any state. (#788)
Fixes #643 
This feature introduces the ability to delete Substrate actors from any
lifecycle state (e.g., RUNNING, PAUSED, PENDING), rather than requiring
them to be hibernated into SUSPENDED (or CRASHED) beforehand. This is
crucial for cleaning up stuck or failed actors.
#### 1. Control Plane & gRPC API (pkg/proto/ateapipb, ateapi)
• DeleteActorRequest.any_state: Added a new boolean field any_state to
DeleteActorRequest.
• When false (default): enforces the existing behavior where only actors
in ACTOR_STATE_SUSPENDED or ACTOR_STATE_CRASHED (or already
ACTOR_STATE_DELETING) can be deleted.
• When true: allows deleting an actor in any state (e.g., RUNNING,
PAUSED).
• Orchestrated Cleanup Workflow (workflow_delete.go):
1. Transitions actor state to ACTOR_STATE_DELETING.
2. Calls atelet.Terminate to stop live workloads on the node.
3. Detaches external volumes from the worker node (with fallback logic
using actor status if the ActorTemplate was deleted).
4. Releases the assigned physical worker Pod back to the WorkerPool.
5. Deletes external storage volumes through the storage plugins.
6. Finalizes removal of the actor from the persistence store.
#### 2. Node Herder Agent (atelet)
• Terminate RPC (main.go): Added a new Terminate method to AteomHerder:
• Calls ateom.TerminateWorkload on the target worker pod.
• Unmounts external volumes mounted for the actor on the node.
• Cleans up and resets actor runtime host directories (OCI bundles,
checkpoints, pid files).
#### 3. In-Pod Sandboxes (ateom-gvisor & ateom-microvm)
• TerminateWorkload RPC:
• gVisor (main.go): Stops and deletes active runsc containers, unmounts
bundle rootfs overlays, and tears down the actor's interior network
namespace.
• Micro-VM (checkpoint.go): Shuts down the Cloud Hypervisor VMM,
unmounts bundle rootfs overlays, and cleans up the network namespace.
• Emits an "Actor terminated" lifecycle log and clears the active actor
attribution so the worker can host new workloads.
#### 4. CLI (kubectl-ate)
• --any-state Flag (delete_actor.go):
kubectl-ate delete actor <actor-name> -a <atespace> --any-state

──────
  ## Verification 
  [x] Changes tested in local with 
 ```
• go build ./...: Passed
go test ./...: Passed (all unit and functional tests pass)
    make lint: Passed (no lint errors) 
```
Detailed manual tests                                                                                                                                                                                                                                                                                                                                           
 ```                                                                                                                                                                                                                                                                                                                                                
  ### Scenario 1: Force Delete a Running Actor with --any-state                                                                                                                                                                                                                                                                                                      
                                                                                                                                                                                                                                                                                                                                                                     
  1. Create the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate create actor manual-test-1 --template=ate-demo-counter-microvm/counter-microvm -a demo                                                                                                                                                                                                                                                               
                                                                                                                                                                                                                                                                                                                                                                     
  2. Resume the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate resume actor manual-test-1 -a demo                                                                                                                                                                                                                                                                                                                   
                                                                                                                                                                                                                                                                                                                                                                     
  3. Force delete the running actor:                                                                                                                                                                                                                                                                                                                                 
    kubectl ate delete actor manual-test-1 -a demo --any-state                                                                                                                                                                                                                                                                                                       
    actor "manual-test-1" deleted                                                                                                                                                                                                                                                                                                                                    
```
──────
### Scenario 2: Standard Delete a Suspended Actor
```                                                                                                                                                                                                                                                                                                                                  
  1. Create the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate create actor manual-test-2 --template=ate-demo-counter-microvm/counter-microvm -a demo                                                                                                                                                                                                                                                               
                                                                                                                                                                                                                                                                                                                                                                     
  2. Resume the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate resume actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                   
                                                                                                                                                                                                                                                                                                                                                                     
  3. Suspend the actor:                                                                                                                                                                                                                                                                                                                                              
    kubectl ate suspend actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                  
                                                                                                                                                                                                                                                                                                                                                                     
  4. Standard delete the suspended actor:                                                                                                                                                                                                                                                                                                                            
    kubectl ate delete actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                   
    actor "manual-test-2" deleted                                                                                                                                                                                                                                                                                                                                    
```
──────
### Unit & E2E Tests
# Unit & Functional Tests
go test ./cmd/ateapi/internal/controlapi -run
"TestDeleteActorWorkflow|TestEnsureMarkedDeleting"
# E2E Tests
E2E_TEMPLATE_NAMESPACE=ate-demo-counter-microvm
E2E_TEMPLATE_NAME=counter-microvm ./hack/run-e2e-kind.sh
./internal/e2e/suites/demo -run TestActorLifecycle
./hack/run-e2e-kind.sh ./internal/e2e/suites/demo -run
TestForceDeleteActorWithExternalVolume
```
- [ x] Appropriate changes to documentation are included in the PR
2026-08-20 13:17:59 -07:00
Youssuf Elshall 86a736c90c workerpool: propagate template metadata to workers (#1058)
Part of #212

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR

## Summary

- add bounded, validated labels and annotations to
`WorkerPoolPodTemplate`
- propagate that metadata to the generated Deployment and worker pod
template
- reserve the controller-owned `ate.dev/worker-pool` label
- regenerate the WorkerPool CRD and deepcopy code
- add API validation, controller tests, and documentation

I chose intentionally to avoid exposing a complete `PodTemplateSpec` as
per the discussion in the #212.
2026-08-20 10:07:51 -04:00