# Substrate API Guide: WorkerPool & ActorTemplate This guide explains how to configure Substrate resources to deploy high-density, stateful agents. ## 1. WorkerPool: The Physical Capacity The `WorkerPool` defines the pool of physical "warm" compute capacity. It manages a fleet of standby pods (herders) that are ready to receive and execute actor states. ### Specification (`WorkerPoolSpec`) | Field | Type | Description | | :--- | :--- | :--- | | `replicas` | `int32` | **Required.** Number of physical standby pods to maintain in the cluster. | | `workerImage` | `string` | **Required.** The container image for the `ateom` herder process (e.g. `ko://github.com/agent-substrate/substrate/cmd/ateom-gvisor`). | | `sandboxClass` | `string` | Optional. The sandbox runtime family for the pool: `gvisor` (default) or `microvm`. Drives the worker pod shape (e.g. KVM device mounts, node placement). The sandbox binaries themselves come from the [`SandboxConfig`](#3-sandboxconfig-the-sandbox-itself) each `ActorTemplate` selects. | | `template` | `WorkerPoolPodTemplate` | **Optional.** Metadata, scheduling, and resource settings for worker workloads. | #### `WorkerPoolPodTemplate` (`spec.template`) | Field | Type | Workload mapping | | :--- | :--- | :--- | | `labels` | `map[string]string` | Generated Deployment and `spec.template.metadata.labels` (max 64) | | `annotations` | `map[string]string` | Generated Deployment and `spec.template.metadata.annotations` (max 64) | | `nodeSelector` | `map[string]string` | `spec.nodeSelector` | | `tolerations` | `[]Toleration` | `spec.tolerations` (max 16) | | `priorityClassName` | `string` | `spec.priorityClassName` | | `nodeAffinity` | `NodeAffinity` | `spec.affinity.nodeAffinity` | | `resources` | `ResourceRequirements` | `spec.containers[].resources` | Keys in `ate.dev/` and its subdomains (for example, `policy.ate.dev/`) are reserved for controllers and cannot be set in `template.labels` or `template.annotations`. Metadata keys and label values must follow Kubernetes syntax. `template.labels` and `template.annotations` only configure Kubernetes workload metadata; they do not affect actor scheduling. Actor selectors match `WorkerPool.metadata.labels`, not `WorkerPool.spec.template.labels`. #### Pin pools to the installed substrate version (`template.nodeSelector`) The installation labels every node with the substrate build version it deployed. Set the same label as a `nodeSelector` on every pool, so its workers only run on nodes of that version: ```yaml spec: template: nodeSelector: ate.dev/substrate-version: "" ``` Read the installed version off the atelet DaemonSet: ```bash kubectl get ds -n ate-system -l app=atelet -L ate.dev/substrate-version ``` The pin is critical to make a rolling upgrade possible. An upgrade moves nodes to the new version one at a time, deleting each node's old worker pods once it moves. A pinned pool cannot put those pods back on a moved node, so the old version drains away node by node. An unpinned pool breaks this constraints. #### Worker Capacity (`spec.template.resources`) Setting `resources.limits` (CPU and Memory) on a `WorkerPool` establishes each worker pod's **capacity** — the envelope its actor sandboxes share, taken from the `ateom` container's limits. The scheduler only places an actor on a worker whose remaining capacity is `>=` the actor's declared resource limits (see [Sandbox Right-Sizing](#sandbox-right-sizing-resources) on the `ActorTemplate`). - Worker capacity is a shared budget: each actor placed on a worker subtracts its declared limits from what is left. Size a pool's `limits` for the actors it should host together. - A worker also has an actor limit, set by the ateom's `--max-actors` flag (default 1000). Placement stops at whichever runs out first. - Capacity is advisory for placement only: a worker that declares no CPU/memory limit reports zero capacity for that dimension, which the scheduler treats as **unconstrained** (placement is never blocked by missing data). The actual sandbox size still comes from the `ActorTemplate`. ### Example ```yaml apiVersion: ate.dev/v1alpha1 kind: WorkerPool metadata: name: agent-pool namespace: ate-demo labels: workload: secret-agent spec: replicas: 10 workerImage: ko://github.com/agent-substrate/substrate/cmd/ateom-gvisor template: labels: project: agent-platform annotations: policy.example.com/exemption: sandbox-host # sandboxClass defaults to gvisor. The sandbox binaries come from the # SandboxConfig each ActorTemplate selects, not from the pool. ``` ### Devices (GPUs) — temporarily unsupported Substrate does not pass devices through to actors. A GPU worker pod's devices were injected into every actor container; that was removed while the device API is designed, since a resource name and a count cannot express which container gets a device, sharing one, or resuming onto compatible hardware. A `nvidia.com/gpu` pool limit is now an ordinary extended resource: it places the worker pod on a GPU node and reserves the device, and nothing else reads it. ### Status (`WorkerPoolStatus`) | Field | Type | Description | | :--- | :--- | :--- | | `replicas` | `int32` | Total number of worker pods, mirrored from the managed Deployment. | | `selector` | `string` | Label selector for the worker pods. | --- ## 2. ActorTemplate: The Workload Blueprint The `ActorTemplate` defines the code, environment, and state-management policies for a specific type of agent. It is used to generate the "Golden Snapshot" from which all actors of this type are derived. ### Specification (`ActorTemplate`) | Field | Type | Description | | :--- | :--- | :--- | | `containers` | `[]Container` | **Required.** The workload definition — see [Container Fields](#container-fields) below. Each container may also declare an optional `wakeupProbe` HTTP probe — see [Container Wakeup Probe](#container-wakeup-probe-wakeupprobe). | | `sandboxConfig` | `SandboxConfig` | **Required.** The sandbox runtime selection: `sandboxClass` (**required**, `SANDBOX_CLASS_GVISOR` or `SANDBOX_CLASS_MICROVM`) picks the runtime family this template's actors require — only `WorkerPool`s whose `sandboxClass` matches are eligible — and `configName` (**required**) names the cluster-scoped [`SandboxConfig`](#3-sandboxconfig-the-sandbox-itself) object supplying the sandbox binaries. It must reference an existing config of the matching class; `CreateActorTemplate` rejects the template otherwise. | | `workerSelector` | `*Selector` | Optional. Gates which `WorkerPool`s actors from this template may use, by matching against each pool's labels (`matchLabels`). If unset, all pools are eligible (subject to the actor's own `worker_selector`). | | `snapshotConfig` | `SnapshotConfig` | **Required.** The base object-storage location snapshots are written under, plus the pause/commit/resume scopes. See [Snapshot Storage Layout](#snapshot-storage-layout). | | `volumes` | `[]Volume` | Optional. Volumes the containers may mount, each a `durableDir`, an `externalVolumeTemplate` (see [CSI Volumes Guide](csi-volumes.md)), or a `systemInfo` volume (see [SystemInfo Volumes](#systeminfo-volumes)). Every declared volume must be mounted by at least one container. A `microvm` template may declare several `durableDir` volumes; a `gvisor` template is limited to one. | | `resources` | `*ResourceRequirements` | Optional. Declares each actor's compute size via `limits` — see [Sandbox Right-Sizing](#sandbox-right-sizing-resources). Immutable, like the rest of the template. | The sandbox itself — the binaries (e.g. the gVisor `runsc` binary) and the `pauseImage` holding the sandbox's namespaces — comes from the cluster-scoped [`SandboxConfig`](#3-sandboxconfig-the-sandbox-itself) object the template names via `sandboxConfig.configName`. An actor always resolves the config from its current template — repointing the actor at another template requires the same config. Because a snapshot is not restorable across sandbox runtimes, `sandboxClass` is a **hard scheduling gate**: an actor is only ever placed on a `WorkerPool` of the matching class. It is AND'd with `workerSelector` (and the actor's `worker_selector`), which can only narrow the eligible pools further. It has no default — `sandboxConfig` is required and its `sandboxClass` must be set — and, like the rest of the template, is immutable, so each template's class is fixed at creation. ### Sandbox Right-Sizing (`resources`) Unlike a Pod, an actor is sized by its **`limits`** (CPU and Memory): the size is a property of the template, baked into snapshots, so it lives on the immutable `ActorTemplate`. Declared limits do three things: 1. **Size the sandbox.** The limits are supplied to the sandbox over the actor RPCs (control plane → atelet → ateom): - **gVisor (`ateom-gvisor`)** — applied to the container OCI spec: `limits.cpu` sets the cgroup v2 CPU quota (`cpu.max`) and the Sentry vCPU count (`--cpu-num-from-quota`); `limits.memory` sets the cgroup v2 memory limit (`memory.max`) and bounds the virtual total memory the sandbox reports (so JVM/Go do not over-allocate from host RAM). - **Micro-VM (`ateom-microvm`)** — `limits.cpu` sets Cloud Hypervisor `BootVcpus` / `MaxVcpus` (rounded up to whole vCPUs); `limits.memory` sets guest RAM, reserving a small configurable margin (default 256 MiB, `--vmm-mem-reserve-mib`) for the VMM and virtiofsd so the pod cgroup does not OOM. 2. **Gate scheduling.** An actor is only placed on a `WorkerPool` whose [worker capacity](#worker-capacity-spectemplateresources) is `>=` these limits. 3. **Fall back to runtime defaults.** A zero or absent limit leaves that dimension at the runtime default: unlimited for gVisor, and 2 GiB / 1 vCPU for the micro-VM. `requests` are not consulted today (an actor occupies its whole worker). Because the size is baked into snapshots, a **micro-VM FULL-scope restore reuses the size in the snapshot**; changing an actor's limits takes effect on its next cold boot. Container environment variables support literal `value` entries only. Values are not interpolated (`$(VAR)` references are not expanded), and Kubernetes `envFrom`/`valueFrom` sources are not supported. ### Workload Connectivity A higher-order system reaches an actor through the **Substrate Router** by setting `ate-target-actor` to `/`. This value selects the Actor; `Host` and HTTP/2 `:authority` remain application metadata. Normal HTTP requirements still apply: clients must send a valid `Host` or `:authority`, usually derived automatically from the request URL, and reverse proxies should preserve it when the application depends on the original authority. Clients that construct HTTP requests should add the routing header directly. For example, curl uses `-H`, Go uses `request.Header.Set`, and Python clients use their request `headers` mapping. WebSocket clients add the same header to the opening HTTP upgrade request. gRPC clients send it as outgoing metadata using the lowercase name `ate-target-actor`. For an HTTP `CONNECT` tunnel to a non-default Actor port, put the routing header on the outer `CONNECT` request and keep the target port in its authority. With curl, use `--proxy-header` instead of `-H`: ```bash curl --proxytunnel --proxy http://localhost:8001 \ --proxy-header "ate-target-actor: my-atespace/my-actor" \ http://actor-upstream:9090/ ``` Browser navigation cannot add custom request headers. Browser-based and other fixed clients must therefore connect through a user-controlled reverse proxy or policy enforcement point that overwrites the routing header before forwarding to `atenet-router`. The [Jupyter demo](../demos/jupyter/README.md) shows this pattern with NGINX. The proxy must derive the value from trusted configuration or authenticated request context rather than forwarding values supplied by an untrusted caller. ### SystemInfo Volumes To deliver identity information, including credentials, to a running actor, you can use a SystemInfo volume. Define it in `volumes`, and mount it into each container that needs it. Available information sources: #### actorMetadata The actorMetadata data source projects the actor's identity fields to files, one per item, analogous to the [Kubernetes downwardAPI volume](https://kubernetes.io/docs/concepts/storage/downward-api/). Each item selects a `field` — `name` (unique within an atespace), `atespace` (together with the name, the actor's full identity), or `uid` (server-generated, distinguishes incarnations of the same name) — and the `path` the value is written to, raw with no trailing newline. `path` is a clean relative path from the root of the volume (no leading `/`, no `.` or `..` segments, at most 16 segments) and must not repeat another path projected into the same volume. ```yaml spec: volumes: - name: system-info systemInfo: dataSources: - actorMetadata: items: - field: name path: actor-name - field: atespace path: atespace - field: uid path: actor-uid containers: - name: main # ... volumeMounts: - name: system-info mountPath: /run/ate # the actor reads e.g. /run/ate/actor-name ``` The values are delivered as files on a read-only per-actor bind mount, not environment variables, precisely so they carry the correct values after a resume from a shared snapshot — an env var (or a file baked into the image) would be frozen at the snapshot-source actor's values, since it lives in the checkpointed process memory, and would therefore be identical for every actor restored from that snapshot. The metadata fields themselves are fixed for the actor's lifetime, so workloads may cache them; future data sources that rotate (identity tokens and certificates) must be re-read at time of use. #### trustBundle The trustBundle data source projects the trust anchors of a named trust bundle to a single PEM file — inspired by the [Kubernetes clusterTrustBundle projected volume source](https://kubernetes.io/docs/concepts/storage/projected-volumes/#clustertrustbundle), but source-neutral: the name selects a bundle substrate knows how to fetch, and where it is fetched from is a deployment concern, not part of the API. Supported names are allowlisted. Today the only supported bundle is `egress-mitm.ate.dev` — the egress gateway CA bundle — resolved from the [ClusterTrustBundle](https://kubernetes.io/docs/reference/access-authn-authz/certificate-signing-requests/#cluster-trust-bundles) (`certificates.k8s.io/v1beta1`) that atecontroller's reconciler derives from the `egress-mitm-ca-pool` Secret in the `ate-system` namespace. A configurable backend registry may widen the allowlist later. ```yaml spec: volumes: - name: trust systemInfo: dataSources: - trustBundle: name: egress-mitm.ate.dev path: ca.pem containers: - name: main # ... volumeMounts: - name: trust mountPath: /run/substrate/certs # the actor reads /run/substrate/certs/ca.pem ``` atelet resolves the bundle on the node when the actor starts, reading the backing object through a cluster-wide watch (the same informer that drives live refresh) and sanitizing it the way kubelet does for projections: only `CERTIFICATE` PEM blocks are kept, deduplicated, with block headers stripped and the anchors deliberately shuffled, so consumers must not depend on their order. The actor itself never talks to any bundle backend. Starting the actor fails, with an error naming the bundle, if the name is not on the allowlist, the bundle's backend is unavailable in this deployment, or the resolved bundle is missing, empty, or contains no certificates. Bundle contents are re-resolved on every Run/Restore and refreshed while the actor runs: when the backing bundle changes, atelet rewrites the projected file atomically at the same path. As with the Kubernetes clusterTrustBundle projection, the application must re-read the file to pick up a rotation; a runtime that loads trust anchors once at startup sees the change at its next start or resume. If a refresh fails because the backing object was deleted or is unusable, the file keeps its last good contents. Bundle publishers should rotate with overlap: add the new CA before minting leaves under it, and keep the old CA until its last leaf expires. ### Container Fields Each entry in `containers` describes one process to run in the actor's sandbox. | Field | Type | Description | | :--- | :--- | :--- | | `name` | `string` | **Required.** DNS-label-safe container name. | | `image` | `string` | **Required.** Container image name; must include a digest (`name@sha256:...`). | | `command` | `[]string` | Optional. Entrypoint array. If unset, the image's `ENTRYPOINT` is used. If set, it replaces **both** the image's `ENTRYPOINT` and `CMD`. | | `args` | `[]string` | Optional. Arguments to the entrypoint. If unset, the image's `CMD` is used (unless `command` is set, which discards the image's `CMD`). If set, it replaces the image's `CMD`. | | `env` | `[]EnvVar` | Optional. Literal `value` entries. | | `wakeupProbe` | `ContainerWakeupProbe` | Optional. HTTP wakeup probe — see [Container Wakeup Probe](#container-wakeup-probe-wakeupprobe). | | `volumeMounts` | `[]VolumeMount` | Optional. Mounts a `volumes` entry (e.g. `durableDir`) into this container. | | `securityContext` | `SecurityContext` | Optional. Security settings for the container process — see [Container Capabilities](#container-capabilities-securitycontextcapabilities). | | `resources` | `ContainerResources` | Optional. Compute limits for this container, enforced inside the actor's sandbox. Only `limits` is supported, and only `cpu` and `memory`. See [Per-container limits](#per-container-limits). | `command` and `args` resolve against the container image's `ENTRYPOINT`/`CMD` the same way [Kubernetes Pod `command`/`args`](https://kubernetes.io/docs/tasks/inject-data-application/define-command-argument-container/) resolve against `ENTRYPOINT`/`CMD`. If the resolved argv is empty — the image sets neither `ENTRYPOINT` nor `CMD`, and the container sets neither `command` nor `args` — `Run`/`Restore` fails. ### Container Capabilities (`securityContext.capabilities`) Each container runs with a default set of Linux capabilities — `AUDIT_WRITE`, `KILL` and `NET_BIND_SERVICE`. `securityContext.capabilities` adjusts that set, mirroring `securityContext.capabilities` on a Kubernetes Pod container. | Field | Type | Description | | :--- | :--- | :--- | | `securityContext.capabilities.add` | `[]string` | Optional. Capabilities to grant on top of the default set. `ALL` is **not** accepted here. | | `securityContext.capabilities.drop` | `[]string` | Optional. Capabilities to remove from the default set. `ALL` drops the whole set. | - **Naming.** Capabilities are named **without** the `CAP_` prefix, as in Kubernetes — `NET_BIND_SERVICE`, not `CAP_NET_BIND_SERVICE`. The prefixed spelling is rejected at admission rather than silently granting nothing. - **Order.** `drop` is applied first, then `add`. A capability named in both is therefore **granted**. - **Exact sets.** Because `drop: ["ALL"]` clears the default set, combining it with `add` expresses an exact capability set rather than a relative one: ```yaml securityContext: capabilities: drop: ["ALL"] add: ["NET_BIND_SERVICE"] ``` - **`ALL` in `add` is rejected.** Kubernetes accepts it in the API and relies on PodSecurity admission to deny it; Substrate has no equivalent policy layer yet, so it is refused at admission instead. Name the capabilities the container needs. - **Ambient capabilities are not supported** ([gvisor#3166](https://github.com/google/gvisor/issues/3166)). The sandbox — gVisor or micro-VM — remains the isolation boundary; capabilities constrain the workload *inside* it. ### Per-container limits A container may cap its own CPU and memory so it cannot starve or kill its siblings in the same actor: ```yaml sandboxClass: microvm containers: - name: trainer resources: limits: {memory: 1500Mi} - name: sidecar resources: limits: {memory: 256Mi, cpu: "0.2"} ``` A container that exceeds its memory limit is OOM-killed on its own; the actor's other containers are unaffected. A `cpu` limit below `10m` is raised to `10m`, because the kernel rejects a CFS quota under 1ms. Per-container limits are micro-VM only today. gVisor applies cgroup limits at the sandbox level: one sentry backs every container in the actor, so a per-container cgroup is created and then stays empty ([google/gvisor#190](https://github.com/google/gvisor/issues/190)). A template that sets `resources` with `sandboxClass: gvisor` is rejected. These limits subdivide the sandbox that [`resources`](#sandbox-right-sizing-resources) already sized; a container that declares none is bounded by the guest as a whole, not by a copy of the actor's total. A micro-VM guest is sized from `resources.limits.memory` minus the VMM reserve, or from the template's [`SandboxConfig`](#3-sandboxconfig-the-sandbox-itself) when the template declares no actor-level limit. The CPU ceiling is the guest's vCPU count, which falls back to the config's `default_vcpus` (1 unless the `SandboxConfig` raises it), so a template that declares no `resources.limits.cpu` caps each container, and their sum, at `1000m`. A limit above either ceiling can never bind, so the actor fails to start with an error naming both the limit and the ceiling. Each limit is validated on its own at apply, but the sum across the actor's containers is only checked when the actor first runs, against the real guest size. A template whose limits do not fit is accepted by the API server and fails on its first actor. ### Container Wakeup Probe (`wakeupProbe`) Each entry in `containers` may declare an optional **HTTP wakeup probe** so `ResumeActor` only returns once the workload is actually serving traffic. Unlike a Kubernetes `readinessProbe`, which reruns for the life of the pod to gate traffic to it, this is a one-shot gate on each wakeup: `ResumeActor` blocks until the endpoint returns 200, and the probe does not run again while the actor stays `RUNNING`. | Field | Type | Description | | :--- | :--- | :--- | | `wakeupProbe.httpGet.path` | `string` | Optional. URL path to GET. Defaults to `/`. Must begin with `/` and contain only RFC 3986 path characters (no query string `?` or fragment `#`). | | `wakeupProbe.httpGet.port` | `int32` | **Required.** TCP port on the container to probe (`1..65535`). | | `wakeupProbe.timeoutSeconds` | `int32` | Optional. How long to keep polling before the wakeup fails (`1..3600`). Defaults to `30`. | How it behaves: - **When it runs.** On every `ResumeActor` that actually wakes the actor. A `ResumeActor` on an actor that is already `RUNNING` is a no-op and does not probe. - **Block-until-ready semantics.** `ResumeActor` returns successfully only after every container with a `wakeupProbe` has returned HTTP 200. If any container has not done so within `timeoutSeconds` (30s by default), `ResumeActor` fails and the actor moves to `ACTOR_STATE_CRASHED`, as it does for any other error while waking. A crashed actor cannot be resumed directly: `ResumeActor` rejects it with `FAILED_PRECONDITION`. Call `RevertActor` first, which returns it to `ACTOR_STATE_SUSPENDED` with its last snapshot, then resume it again. - **Aggressive polling.** The poll loop is tuned for single-millisecond detection latency: a keep-alive HTTP client with a 1ms interval and 250ms per-request timeout. While the workload is still booting, kernel `RST`s return in microseconds, so the loop spends almost no time blocked; once the listener is up, the next attempt completes on veth-local latency. - **Golden snapshot warm-up shortcut.** When **every** container in a template declares `wakeupProbe`, the actor template controller skips its default ~20s "give the workload time to settle" delay before taking the golden snapshot — `ResumeActor` already blocked until the workload reported 200, so the workload is known to be initialized. Templates that omit `wakeupProbe` on any container keep the 20s warm-up as a safety net. - **Snapshot/restore interaction.** The TCP listener is part of the checkpointed RAM, so on resume `wakeupProbe` typically returns 200 on the first attempt, with no observable latency penalty. If `wakeupProbe` is omitted from a container, `ResumeActor` returns as soon as the sandbox has started that container, which can be before the workload is listening. ### Example A protojson-shaped `ateapipb.ActorTemplate`, created through the ate API with `kubectl ate create actor-template -f secret-agent.yaml` (the `ate-demo` atespace must exist): ```yaml metadata: atespace: ate-demo name: secret-agent containers: - name: agent image: gcr.io/my-project/my-agent@sha256:7f28ab0e... # Optional: gate ResumeActor on the agent's HTTP wakeup probe endpoint. # See "Container Wakeup Probe (wakeupProbe)" above. wakeupProbe: httpGet: path: /readyz port: 80 workerSelector: matchLabels: workload: secret-agent # sandboxClass (required) picks the runtime family (set SANDBOX_CLASS_MICROVM # to require micro-VM pools); configName (required) names the cluster-scoped # SandboxConfig supplying the sandbox binaries (see section 3). # gvisor-default is the SandboxConfig that manifests/ate-install ships. sandboxConfig: sandboxClass: SANDBOX_CLASS_GVISOR configName: gvisor-default snapshotConfig: storageLocation: gs://my-bucket/secret-agent ``` ### Snapshot Storage Layout `snapshotConfig.storageLocation` is a **base prefix**, not the address of any one snapshot. Every external snapshot has exactly one owner: the actor that took it, or the tag that copied it. And the owner is part of the path, so an object's name says who it belongs to: ``` /atespaces//actors//snapshots/ /atespaces//tags/ ``` The objects of a snapshot (its manifest, memory image, durable-data tar) are named below it. So for the template above, a snapshot of an actor in atespace `team-a` is stored at `gs://my-bucket/secret-agent/atespaces/team-a/actors/3f8b…/snapshots/f47ac10b-…`, and the template's golden snapshot — the golden tag lives in the reserved `ate-golden` atespace — under `gs://my-bucket/secret-agent/atespaces/ate-golden/tags/`. An actor takes a series of snapshots over its life, so it gets a prefix of its own and each snapshot sits below it. A tag holds exactly one, so the tag's prefix *is* its snapshot's. Both owners are keyed on their UID, so recreating an actor or tag under the same name never inherits its predecessor's objects. A pending tag records its base location in `status.storageLocation`; together with its atespace and UID, this identifies any partial copy to collect if creation fails. An owner is collected by deleting everything under its prefix, and it can delete nothing else. That is what makes a borrowed snapshot safe: an actor created from a tag points at a URI under `tags/`, which its own prefix does not cover. See [Snapshot lifetime](#snapshot-lifetime). An `Actor` reports its current snapshot in the server-managed `status.externalSnapshot` and a `Tag` in `status.snapshot`, each an `ExternalSnapshot` carrying `snapshotUri`, `contentScope`, and `actorTemplateUid`. The URI is recorded when the snapshot is written. `actorTemplateUid` records the `ActorTemplate` whose sandbox the guest state was captured from, which is not always the template the actor points at now: an actor may be repointed while `SUSPENDED`, and the snapshot on disk still came from the old one. A resume that finds the two disagree restores the durable data only and boots the guest fresh, because memory captured under one sandbox image cannot be resumed under another. An `ActorTemplate` references its golden tag with the `ObjectRef` in `status.goldenSnapshotStatus.goldenTag`. These status fields are server-owned and ignored on input. Parse a URI only against the scheme above. An `ActorTemplate` belongs to one atespace, but one `storageLocation` still holds snapshots for many atespaces: the golden actor lives in the reserved `ate-golden` atespace, and a `PUBLISHED` snapshot may be cloned from other atespaces. The `` level exists so that access can be granted per tenant: an object-storage policy can only condition on an **object-name prefix**, and cannot read the identity recorded inside a snapshot's manifest. Binding a per-atespace grant on GCS looks like: ```yaml # Read-only on team-a's snapshots for this template, and nothing else. - members: ["serviceAccount:node-runtime@my-project.iam.gserviceaccount.com"] role: roles/storage.objectViewer condition: title: team-a-snapshots expression: > resource.name.startsWith( "projects/_/buckets/my-bucket/objects/secret-agent/atespaces/team-a/") ``` One consequence worth planning for: **a published snapshot is read from the atespace that took it.** Cloning across atespaces via a `PUBLISHED` tag reads the source atespace's prefix, so the reader needs a grant covering it — the target atespace's grant is not enough. --- ## 3. SandboxConfig: The Sandbox Itself `SandboxConfig` is a **cluster-scoped** resource that decouples the sandbox — its binaries (the gVisor `runsc` binary, or a micro-VM kernel/firmware/config) and the `pauseImage` that holds the sandbox's namespaces — from the workload definition in the `ActorTemplate`. An actor's cold boot resolves the sandbox binaries from the config its `ActorTemplate` names via `sandboxConfig.configName`. This means a single, cluster-managed config pins the sandbox runtime version for many templates: snapshots stay restorable because the version is recorded in each snapshot's manifest, and operators upgrade the runtime in one place. ### Specification (`SandboxConfigSpec`) | Field | Type | Description | | :--- | :--- | :--- | | `sandboxClass` | `string` | **Required.** Runtime family this config applies to: `gvisor` (default) or `microvm`. An `ActorTemplate` only uses `SandboxConfig`s whose `sandboxClass` matches its own. | | `pauseImage` | `string` | **Required.** The image for the sandbox's root container (e.g. `registry.k8s.io/pause`, or `gcr.io/gke-release/pause` on GKE). Must include a digest (`...@sha256:...`) — it is recorded in each snapshot's manifest so a restore rebuilds the sandbox from the same image. | | `assets` | `map[arch]map[name]AssetFile` | Optional. Content-addressed files atelet fetches, keyed by architecture (`amd64`, `arm64`) then asset name. gVisor expects a `gvisor` asset (the release's `gvisor.tar.zstd`), which atelet auto-extracts. A micro-VM backend expects several. Each `AssetFile` is a `{ url, sha256 }` pair. | A cluster-wide gVisor `SandboxConfig` (`gvisor-default`) is installed with the platform, so gVisor templates can name it via `sandboxConfig.configName` without any extra setup. ### Example ```yaml apiVersion: ate.dev/v1alpha1 kind: SandboxConfig metadata: name: gvisor-default spec: sandboxClass: gvisor pauseImage: "registry.k8s.io/pause:3.10.2@sha256:f548e0e8e3dc1896ca956272154dde3314e8cc4fde0a57577ee9fa1c63f5baf4" assets: amd64: gvisor: url: "gs://gvisor/releases/nightly/2026-09-02/x86_64/gvisor.tar.zstd" sha256: "d547d81401461fd1c679c5c4fa0a6c2b8ef7dc3c22ce23c9e25dcc4c69cfd06f" arm64: gvisor: url: "gs://gvisor/releases/nightly/2026-09-02/aarch64/gvisor.tar.zstd" sha256: "a64916f9813ce7e4841a30480a599337f7dda07b421c6bf0123db2212aa7d1df" ``` ### Micro-VM SandboxConfig A `microvm` `SandboxConfig` supplies the [Kata Containers](https://katacontainers.io/) + [Cloud Hypervisor](https://www.cloudhypervisor.org/) toolchain instead of `runsc`. Each architecture must define the full asset set — `cloud-hypervisor`, `virtiofsd`, `kata-kernel`, and `kata-image` — which a `ValidatingAdmissionPolicy` enforces at apply time. Worker pods for a micro-VM pool require `/dev/kvm` and nested-virtualization-capable nodes. The controller requests those devices on the pod automatically, and atelet advertises them only where they exist, so placement follows the hardware rather than a node label. Clusters that reserve nested-virt nodes with an `ate.dev/sandboxClass=microvm` taint are still tolerated: advertising a device attracts these pods to capable nodes but repels nothing else from them. The same convention applies to every class: worker pods of a pool tolerate `ate.dev/sandboxClass=:NoSchedule`, so a cluster can reserve a node pool per sandbox class with that taint, and the atelet DaemonSet tolerates the key for any value. See [`hack/microvm-assets/`](../hack/microvm-assets/) for scripts that assemble and stage these assets, plus a worked counter demo (`demos/counter/counter-microvm.yaml.tmpl`) that suspends and resumes an in-RAM counter across worker pods. --- ## 4. Operational Workflow ### The Golden Snapshot When an `ActorTemplate` is created: 1. Substrate creates and resumes a temporary golden actor in `ate-golden`. 2. It waits for readiness (or the warm-up interval), then suspends the actor. 3. It creates a published tag named after the template UID, copying the snapshot into tag-owned storage. 4. It deletes the golden actor and records the tag reference in the template status. `CreateActor` uses an explicit `sourceTag` when supplied; otherwise it resolves the template's golden tag and records that snapshot on the new actor. If the golden tag is not ready yet, the actor starts without a snapshot and cold-boots even if the tag becomes ready before its first resume. The default does not populate the caller-owned `sourceTag` field. Deleting the template collects its golden tag and any unfinished golden actor. ### Resumption Lifecycle Once a template is `Ready`, creating an actor logically (via `kubectl ate create actor`) allows it to be resumed instantly on any free worker in the referenced `WorkerPool`. Substrate bypasses the standard container boot and restores the process directly from its last saved state. --- ## 5. Best Practices * **Startup Logic:** Place expensive initialization (loading large models, establishing baseline connections) in your application's entry point. These will be captured in the Golden Snapshot and won't need to be repeated on every resumption. * **Placement:** Ensure your `ActorTemplate`'s `sandboxClass` matches your `WorkerPool`'s, and use the template's `workerSelector` to target specific pools — pool selection is by label match, not by namespace or RBAC. * **Version Management:** When updating code, create a new `ActorTemplate` (e.g. `v2`). Substrate treats each template as an immutable state root. * **Eviction:** When its worker pod is evicted, an actor gets `SIGTERM` and 30 minutes to be suspended. After that it is killed and moves to `ACTOR_STATE_CRASHED`, and everything since its last snapshot is lost. So an actor that runs for more than 30 minutes without a suspend can lose data. A `CRASHED` actor can be recovered back to `ACTOR_STATE_SUSPENDED` at its last external snapshot using `RevertActor` (`kubectl ate revert`). --- ## 6. Control Plane gRPC API The Substrate Control Plane (`ate-api-server`) exposes a gRPC interface for managing actors and workers. This is the primary API used by the `kubectl-ate` CLI and higher-level frameworks. ### Service: `ateapi.Control` #### `CreateActor` Registers a new logical actor in the system. * **Request:** `CreateActorRequest` * `actor`: `Actor` — the actor to create. Its `metadata` carries the atespace and name (name must be a DNS-1123 label); the `actor_template` ref (atespace + name) selects the `ActorTemplate`. * `actor.source_tag`: (Optional) `ObjectRef` of a `Tag` to seed the actor from. The tag must be taken under the same `ActorTemplate`, and either in the actor's own atespace or `PUBLISHED`. Nothing is copied: the new actor's `status.externalSnapshot` points at the tag's snapshot, under the tag's prefix, until its own first suspend. * **Response:** the initialized `Actor`. #### `UpdateActor` Replaces the mutable fields of an existing actor with the ones in the request. * **Request:** `UpdateActorRequest` * `actor`: `Actor` — the complete replacement actor. `metadata.atespace` and `metadata.name` identify the resource; `metadata.uid` and `metadata.version` are **required** preconditions. `metadata` and `status` are server-owned and whatever the request carries in them is ignored. `source_tag` is immutable. * `actor_template` may change only while the actor is `SUSPENDED`. The new template must name the same `SandboxConfig` and declare the same volumes and volume mounts. Once the actor owns an external snapshot, the new template must also store snapshots under the same `snapshotConfig.storageLocation`. * **Response:** the updated `Actor`. * **Errors:** `INVALID_ARGUMENT` if `uid` or `version` is unset, or if the request changes an immutable field — including by leaving one unset; `FAILED_PRECONDITION` if a change to `actor_template` breaks one of the rules above; `ABORTED` if either guard no longer matches the stored resource. Because the guards are required and only a read supplies them, an update is always a read-modify-write. To Update an `Actor`, you must first `GetActor`/`CreateActor`, instead of building a new one — see [§7.2 of the API style guide](api-style-guide.md#72-using-version-and-uid-to-guard-writes) for why reconstructing the message can silently drop data. #### `ResumeActor` Activates a suspended actor by restoring it onto a physical worker. * **Request:** `ResumeActorRequest` * `actor`: `ObjectRef` of the actor to resume. * **Response:** `ResumeActorResponse` containing the updated `Actor` object (including the physical worker placement in `status.worker_assignment`). #### `SuspendActor` Hibernate a running actor, capturing its current RAM and disk state into a snapshot. * **Request:** `SuspendActorRequest` * `actor`: `ObjectRef` of the actor to suspend. * **Response:** `SuspendActorResponse` containing the `Actor` object in `ACTOR_STATE_SUSPENDED`, with its snapshot in `status.externalSnapshot`. * A successful suspend releases the actor's previous external snapshot: an actor keeps one, and only tags outlive it. To keep the snapshot a suspend just wrote, tag it with `CreateTag` while the actor is still suspended. #### Snapshot lifetime Every external snapshot has exactly one owner, and the control plane deletes it when that owner lets go: | Owner | Released when | | :--- | :--- | | The actor that took it (`status.externalSnapshot`) | The actor's next successful suspend replaces it, or the actor is deleted. | | The tag that copied it (`status.snapshot`) | The tag is deleted. | An actor created from a tag borrows the tag's copy instead of taking one of its own. The borrowed URI sits under the tag's prefix, which the actor's own prefix does not cover, so neither suspending nor deleting the actor can reach it; its first own suspend writes a snapshot under the actor's prefix, and it owns its snapshots from then on. Deletion always runs before the database reference is dropped, and a failure fails the whole RPC. Clients are expected to retry with the same arguments: destinations are deterministic and every phase tolerates a partly-completed predecessor, so a retry resumes rather than duplicating work. The cost of that ordering is that a crash between the two can leave an external snapshot no row names; the reverse order would instead lose the handle needed to ever delete it. > **Do not delete a tag while actors created from it exist.** A clone borrows the tag's snapshot rather than copying it, and only stops borrowing at its own first suspend (its `status.externalSnapshot.snapshotUri` still names the tag's prefix while it is). Deleting the tag leaves such a clone unable to resume. This is not prevented today. #### `RevertActor` Discards an actor's live or crashed execution and transitions it to `ACTOR_STATE_SUSPENDED` at its last completed external snapshot (`status.externalSnapshot`). * **Request:** `RevertActorRequest` * `actor`: `ObjectRef` of the actor to revert. Accepted from `ACTOR_STATE_RUNNING`, `ACTOR_STATE_PAUSED`, and `ACTOR_STATE_CRASHED` (plus `ACTOR_STATE_REVERTING` for idempotent retries). Calling `RevertActor` on an already `ACTOR_STATE_SUSPENDED` actor returns `FAILED_PRECONDITION`. * **Response:** `RevertActorResponse` containing the reverted `Actor` in `ACTOR_STATE_SUSPENDED`. * Reverting terminates any bound worker sandbox, clears node-local pause checkpoints (`localSnapshot`), and garbage-collects any partial external snapshot left by an interrupted suspend while preserving the last committed `externalSnapshot`. * External volumes are not reverted. Their contents are never part of a snapshot, so a reverted actor comes back with its memory and root filesystem rewound but its volumes exactly as the discarded execution left them. #### `DeleteActor` Removes an actor from the registry and cleans up associated resources. * **Request:** `DeleteActorRequest` * `actor`: `ObjectRef` of the actor to delete. * `options`: (Optional) `DeleteOptions`. `uid` and `version` are preconditions checked against the actor as the caller last read it: a mismatch returns `ABORTED`, an omitted guard is skipped. They are checked before the workflow starts, so a guarded delete that fails part-way leaves the actor `ACTOR_STATE_DELETING` at a higher version. Retry it with the `uid` guard alone, or re-read first. * `any_state`: (Optional) If `true`, allows deleting the actor from any state (e.g. `RUNNING`, `PAUSED`), terminating active workloads, detaching volumes, and releasing worker allocations. By default (`false`), only actors in `ACTOR_STATE_SUSPENDED` or `ACTOR_STATE_CRASHED` (or already `ACTOR_STATE_DELETING`) can be deleted. * **Response:** the deleted `Actor`, as it was immediately before removal. * Deleting an actor also deletes the external snapshot it owns, along with one an interrupted suspend left behind. Snapshots it only borrows from a tag are left alone, and its tags are unaffected — they hold their own copies. #### `GetActor` / `ListActors` Query the state of logical actors. * **GetActor:** Retrieves a single actor by ID. * **ListActors:** Lists all actors currently tracked in the database. #### `ListWorkers` Query the physical resource pool. * **Request:** `ListWorkersRequest` * **Response:** `ListWorkersResponse` containing a list of `Worker` objects (Pods) and their current assignment status. --- ## 7. Advanced: Actor Identity Credentials Workloads can exchange their ephemeral Kubernetes credentials for stable **Actor Identity** credentials that persist even as the process migrates between different physical workers. This is distinct from the `actorMetadata` data source described under [SystemInfo Volumes](#systeminfo-volumes), which only tells an actor its own identity fields (name, atespace, uid). ### Service: `ateapi.ActorIdentity` * **`MintJWT`:** Generates an OIDC-compatible JWT identifying the Substrate Actor. * **`MintCert`:** Signs a Certificate Signing Request (CSR) to provide an mTLS identity for the actor. Both RPCs identify the actor the same way the rest of the API does, by `atespace` and `actor_name`. #### Who may call `MintCert` and `MintJWT` `MintCert` and `MintJWT` are not callable by actors directly. They must be called over mTLS with a Pod Certificate, and the broker only signs a CSR when all of the following hold: 1. The client certificate identifies the **`atelet`** service account (`spiffe://cluster.local/ns/ate-system/sa/atelet`) and carries a Pod Identity extension, which pins the calling atelet to a node. 2. The requested actor is **currently running**, per the actor database. 3. The worker Pod hosting that actor is on the **same node** as the calling atelet, and is still assigned to that actor. An atelet is therefore confined to minting credentials for the actors it is actually hosting. Callers that fail any of these checks receive `PERMISSION_DENIED` with no detail, so the RPC cannot be used to discover whether an actor exists or where it is running. An actor that exists but is suspended, paused, or crashed yields `FAILED_PRECONDITION`. The minted leaf certificate carries the SPIFFE URI `spiffe://substrate-actor.local/atespace/${atespace}/actor/${actor_name}`. --- ## 8. Framework & Ecosystem Integration Agent Substrate is designed to be the foundational execution layer for any agentic framework. ### Agent Development Kit (ADK) Substrate provides native support for ADK-compatible identities. Workloads can use the `ActorIdentity` service to mint JWTs that align with ADK's security model, ensuring seamless integration with ADK-managed tools and memory. ### LangChain Substrate is an ideal runtime for stateful LangChain agents. By defining a LangChain agent as an `ActorTemplate`, you can preserve the agent's internal "thought process" and conversation history in memory across hibernations, while sandboxing its tool execution for security. ### Claude Code & CodeX For developer-focused agents, Substrate enables massive multiplexing of coding environments. Each developer can have a dedicated, persistent terminal session (Actor) that preserves filesystem deltas, while the cluster only runs physical pods for active users.