This is needed because DeleteActor cleans up all snapshots under the
actor's location.
If we were to allow repointing an actor's snapshot location, we risked
leaking historical snapshots upon actor deletion.
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
Fixes#1266
Builds on #1689#1283 in particular.
Outstanding:
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~
EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Follow up on the comment -
https://github.com/agent-substrate/substrate/pull/1675#discussion_r4031619230
Stamps `actor_template_uid` onto `ExternalSnapshot` so template
provenance travels with the snapshot itself, rather than being inferred
from the single `ActorStatus.current_actor_template_uid`.
#### Motivation
`current_actor_template_uid` records the template the **last sprint
booted with** (`finalizeRunning` stamps it on every resume). It does
*not* record the template the **snapshot** was captured under. Those two
diverge as soon as an actor runs a sprint that does not produce a new
durable snapshot, and `loadActorForResume` was using the former to
decide whether to force a `DATA`-only restore instead of `FULL`.
Walking an actor through a repoint and a crash:
| Step | Action | `current_…_uid` | `ExternalSnapshot` |
|---|---|---|---|
| **a** | Suspend on template **A** | A | captured on **A** |
| **b** | `UpdateActor` repoints the spec to **B** | A | captured on
**A** |
| **c** | Resume → `A != B`, restores `DATA`, sprint boots on **B** |
**B** | still on **A** |
| **d** | Actor crashes (or is reverted) — no new snapshot taken | B |
still on **A** |
| **e** | Resume → `current == target == B`, so **no repoint is
detected** | B | still on **A** |
At step **e** the old check reports "template not replaced" and restores
the external snapshot in `FULL` — replaying a memory image captured
under **A**'s sandbox on **B**. That is exactly the case the guard
exists to prevent; it was silently defeated by the sprint at step **c**
advancing `current_actor_template_uid` past the snapshot.
*(Note: Paused actors do not face this divergence. Because `UpdateActor`
is only allowed in the `SUSPENDED` state, a paused actor cannot have its
template updated. Therefore, local checkpoints do not need separate
template provenance tracking).*
#### Changes
- **Data Model Updates:** Added the `actor_template_uid` field to
`ExternalSnapshot`.
- **Snapshot Stamping:** Updated the capture logic to stamp external
snapshots with the active sandbox's `ActorTemplate` UID when suspending
an actor or creating an actor from a tag.
- **Resume Evaluation Logic:** Modified the external resume logic to
compare the target template UID against the specific snapshot's UID
(`ExternalSnapshot.actor_template_uid`) rather than the actor's overall
status. Local restores bypass this check entirely since templates cannot
change during a pause.
**Tests**
- `TestResumeActor_AteletWireRequest` (unit): snapshot fixtures now
carry `actor_template_uid`, since provenance is read from the snapshot
rather than from actor status.
- Functional expectations for create/update/resume/pause updated for the
new field on the external snapshot golden files.
- `TestUpdateTemplateLifecycle` (e2e): asserts
`external_snapshot.actor_template_uid` after suspend and re-suspend.
- **e2e (Update Template)**: Added coverage for repoint detection after
a revert (verifying the `resume` → `revert` → `resume` path described in
the table above).
- **e2e (Combined Volumes)**: Added coverage to explicitly verify which
volumes a revert rewinds.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.
Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.
Follow up for the PR - #1675
Issue - #1556
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The name `readyz` is not very descriptive. We also want to avoid
`readinessProbe` to prevent confusion with the Kubernetes concept, which
represents continuous traffic gating. Renaming this field to
`wakeupProbe` clarifies its actual behavior, and the fact that it only
operates when the actor is woken up.
Part of #1378
Today, `ateom` fetches the `kata-config` asset (configuration-clh.toml)
to read only 3 values `default_memory`, `default_vcpus` and
`kernel_params`. The first 2 are redundant because (1) the values never
change and (2) ateom has the same defaults. Only `kernel_params` is
relevant for the `kata-agent`; its value also changed once in the Kata
project history.
`ateom` now owns all three values, which also allows us to tune them
specifically for Substrate.
Verified e2e with a GKE cluster.
Fixes#1693
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1507
Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.
`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.
This PR is based directly on `main` and does not depend on #1521.
Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Fixes https://github.com/agent-substrate/substrate/issues/1566
Today the Resume workflow resolves its restore source in the following
order
1. First check if the actor has node-local snapshot,
2. then its own durable external snapshot,
3. then the template's golden snapshot.
The boot flag was consulted at exactly one point in that chain, where it
suppressed using the golden-snapshot, which made its behavior much
narrower than "boot from scratch" suggests:
- Actor has its own external snapshot and boot=true: flag ignored,
restores the actor's snapshot.
- Actor has a local snapshot and boot=true: flag ignored, restores the
local snapshot.
- Actor has no snapshot, template has no golden snapshot: cold boot from
the spec regardless of the boot flag.
- Actor has no snapshot, template has a golden snapshot: boot=false
restores the golden, boot=true cold boots from the spec. This is the
only case where the flag is used.
The glutton benchmark was the only caller that set boot=true, on each
actor's first resume, to report true cold-start latency as a separate
ResumeActorColdStart stats row. @maxsmythe let me know if this is
required.
The proto field number and name were reserved.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The container image must include the image digest. This validation was
dropped during the migration of the ActorTemplate resource from CRD to
Substrate API.
Part of #932 (PR 2 of 3). PR 1 (#941) added the `trustBundle` SystemInfo
data source, resolved on the node at Run/Restore. This PR keeps those
projections current while the actor runs.
## Live refresh
A `systemInfoVolumeRefresher` in atelet, modeled on kubelet's projected
volumes: `collectData` is the one place that builds a volume's complete
contents from its spec, `write` applies them, and every lifecycle point
uses the pair. Run/Restore registers the actor's volumes, which writes
them fail-closed before the sandbox boots; ClusterTrustBundle events
rewrite them while it runs; Checkpoint, Terminate, and failed starts
deregister. No API, proto, or RBAC changes: the wire still carries only
`{name, path}`.
- Informer events only enqueue bundle names; a single run loop writes.
Failed writes requeue with backoff through a rate-limited workqueue;
resolution failures keep last-good contents and wait for the bundle's
next event. Resync (24h) is only a guard against missed watch events.
- Change detection hashes the raw backing contents, not the projected
output (sanitization shuffles). Files are replaced by temp-and-rename at
stable paths and byte-identical files are left untouched: a needless
rewrite replaces an inode that suspended guests re-bind on resume.
- Nothing is persisted. Registrations are in memory; after an atelet
restart, a running actor's files keep their last-written contents until
its next Run/Restore rewrites them from current cluster state.
## virtiofsd: `--migration-on-error=guest-error`
CI confirmed the hazard (3 of 3 micro-VM runs): a rotation renames over
a file whose inode the guest still references (a dcache reference is
enough, no held fd), the actor suspends before re-reading, and
find-paths serializes an inode with no findable path. Under the default
`abort`, the destination virtiofsd rejects the device state at
vm.restore: suspend succeeds, every restore fails, and the actor is
permanently stuck. Full analysis in the PR comments.
With `guest-error` the restore succeeds and only the stale reference is
faulty: EIO on access until a fresh lookup at the stable path heals it.
That matches the documented reader contract (the platform rewrites the
file, applications re-read it) and also covers guest-created
`O_TMPFILE`s held across suspend, which already fail restores under
`abort`. The flag is backend-process configuration, so existing
snapshots (template goldens included) restore unchanged. Trade-off: a
share-reconstruction bug now surfaces as post-resume EIO and a readyz
failure instead of a loud restore failure; virtiofsd names the faulty
inodes in the worker pod log.
## E2E
The identity suite covers live refresh on both sandbox classes: rotate
the pool under two running actors and wait for both to observe the new
bundle live; rotate/suspend/resume without waiting for propagation, so
the resume must deliver rotated contents whether the rewrite landed
before or after the guest went down (the exact `abort` repro); assert a
sibling that never cycled also got the rotation live. The probe fixture
now keeps its namespace when a test fails so the worker pod logs
survive, and the cloud-hypervisor client includes the vm.restore
response body in errors.
## Not in this PR
Auto-injected egress trust volume (#932 PR 3) and the configurable
backend registry (tracked in #932).
Fixes#1508
Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.
- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.
Lint and code-generation verification also passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
## What
Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.
Fixes#368 (docs side).
> This change was prepared with AI assistance; I have reviewed and
tested it.
- [x] Docs and comment-only change; no functional code changed, no tests
affected
`CreateActorTemplate` accepted `../escape` and `/escaped` in
`ActorMetadataItem.path` and `TrustBundleDataSource.path`, and accepted
several data sources projecting to the same path. atelet only rejects
the traversal at actor start, so the template is created but the golden
actor sits in `RESUMING`. atelet never rejects duplicate paths at all:
it writes the files in order and the last one wins.
This PR moves both checks to template validation.
- The path rule lives in a new `internal/volumepath` package that both
ateapi and atelet call, so the two cannot drift. It splits on `/` and
rejects any segment that is empty, `.`, or `..`, caps depth at 16
segments, and rejects NUL bytes. Any other byte is allowed, so Unicode
paths pass. `ActorMetadataItem.path` and `TrustBundleDataSource.path`
apply it through a `+k8s:customValidation` hook, and atelet's
`writeSystemInfoFile` applies it again before touching the host
filesystem.
- `SystemInfoVolumeSource.data_sources` gets a hook that tracks
projected paths in a set and reports a duplicate on repeats, in the same
shape as the Kubernetes projected-volume validation.
The proto doc comment on `data_sources` also says at most one
`actor_metadata` entry may appear. Nothing enforces that, and this PR
leaves it alone.
New cases in `TestValidateSystemInfoVolumeSource` and
`TestValidateTrustBundleDataSource` cover the rejected shapes, and
`internal/volumepath` has its own table test. `hack/verify/codegen.sh`
and `hack/verify/golangci-lint.sh` pass.
#1231 moves `writeSystemInfoFile` into `systeminfovolume.go` and onto
`os.Root`. Whichever PR lands second points that copy at the shared
rule.
Fixes#1436Fixes#1438
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.
https://github.com/agent-substrate/substrate/issues/664
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#664
This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489
It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).
Now, an external snapshot is owned by a single resource:
- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.
Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.
this PR:
- Drops table actor_snapshots
- Keeps table actor_snapshot_tags
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.
The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).
For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Summary
Improves the cold-node activation performance tracked in #811: the
gVisor release extract on a cold node currently exceeds the router
resume budget, so the first resume after a release bump always fails.
- Point the default gvisor SandboxConfig at the 2026-09-02 nightly,
using the `gvisor.tar.zstd` tarballs the nightlies now publish alongside
`gvisor.tar.bz2`. The win is decode speed, not just download size:
stdlib bzip2 is a pure-Go, single-threaded decoder, and extracting the
~165 MB release costs a node roughly 20 seconds of CPU on the first
actor operation it hosts — the dominant term of a cold-node activation.
With zstd the same extraction takes ~2.9 s, measured on a GKE node
running this change. The tarball is also ~22% smaller to download (129
MB vs 166 MB for x86_64).
- Teach `extractTarArchive` to decompress `.tar.zst` / `.tar.zstd`
archives using `github.com/klauspost/compress/zstd` (already a
dependency via ategcs), with test coverage for both suffixes.
- Update the `gvisor.tar.bz2` mentions in docs, manifests comments, and
the SandboxConfig API comment (generated CRD refreshed via
`hack/update/codegen.sh`).
The bucket only publishes sha512 checksums; the pinned sha256 sums were
computed from downloads verified against the published sha512 sums.
## Test plan
- `go test ./cmd/atelet/` (includes new `.tar.zst` / `.tar.zstd`
extraction cases)
- `go build ./cmd/atelet/...`, `go vet`, `gofmt` clean
- Verified the real x86_64 tarball lists `runsc` plus the `gvisor-bin/`
helpers
- Deployed to a GKE cluster: atelet fetched and extracted the 2026-09-02
x86_64 zstd tarball in ~2.9 s (vs ~20 s for bz2)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
A GPU worker pool injected its devices into every container of every actor on it,
which only held up while a worker ran one actor.
A resource name and a count cannot replace it: that shape cannot express which
container gets a device, sharing one between containers, or resuming onto hardware
compatible with the snapshot. A nvidia.com/gpu pool limit is now an ordinary
extended resource -- it places the pod on a GPU node and nothing else reads it.
internal/cdi stays, unused: whatever the API becomes, an ateom still has to merge
CDI edits into an actor's bundle itself.
Fixes#368 . Deletes ActorTemplate CRD and any references to it.
Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
A rolling upgrade needs two substrate versions running in one cluster,
split by node: `atelet` and `ateom` speak a node-local protocol with no
cross-version guarantees, so everything on a node has to come from one
build. Today nothing records which build a node runs, and the `atelet`
DaemonSet is a single fixed-name object, so a second version can only
replace the first in place, instead of our porposed node by node rolling
([design](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)).
This PR is the basis for upgrade. Nodes and the `atelet` DaemonSet get
keyed by an `ate.dev/substrate-version` label whose value is the build
version stamped into the binaries.
### One derivation for the version label
`internal/versionlabel` turns the build version into its two forms in
one place:
| form | grammar | used for |
|---|---|---|
| label value | k8s label value | node labels, DaemonSet labels,
nodeSelectors |
| name suffix | DNS-1123 | the DaemonSet name `atelet-<suffix>` |
A version that is not a valid label value is rejected, because the label
has to match what the `ldflags` stamp put into the binaries through
MAKEFILE and `ko`. It's also exposed shell, so the install script can
use it.
### A DaemonSet per version
`atelet-<suffix>` as the name, the version label on metadata, selector,
and pod template, and a pod nodeSelector on the same label. Two versions
run side by side on disjoint old/new node sets.
### The install becomes version-aware
- `install-ate.sh` reads the version from the same `make ldflags` output
it stamps the binaries with, then fills the manifest placeholders.
- It labels the nodes that exist at install time, nodes come after it
carry no version label yet (so add notes in README). Re-running the
install is idempotent, this is the basis for upgrade runbook later.
- The `README` documents the invariant, how to read the installed
version off the DaemonSet, and that a node added later hosts no workers
until it carries the label. On GKE the node pool label is the birth
default for new nodes - added in `tools/setup-gcp/README.md`.
### Follow-ups
- `ate-setup` need same modification
- minimize the affect for daily developer (pin the `VERSION` as
`{USER}-dev` instead of from git
- the manual rolling-upgrade runbook.
Fix#1270
Ref [design
doc](https://docs.google.com/document/d/1JduAyGZFyqdNp4UhKiv0EN-is3iW5tf5BwGTWbt2ouI/edit?usp=sharing&resourcekey=0-MXG12QCkleIxhY7cOB7g6A)
This PR is large, I grouped changes to the following commits:
- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.
Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
We should not make workerImage optional as it provide explicit
information for what image workers use, removing it will need
`atecontroller` to inject workerImage and leads to unnecessary control
plane & dataplane wiring.
Fixes#861
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.
This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.
Part of #861
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This is the companion to #496
In order to grant access to `/dev/kvm` we have to either:
- use a device plugin
- use a DRA driver
- use an NRI plugin
An NRI plugin is highly privileged in it's own right for all pods on the
host and is difficult to ship portably at the moment.
DRA is promising, but not enough functionality is GA yet at our current
1.35+ target.
Device plugin fits reasonably well. We do wind up publishing an
~arbitrarily high limit, which has some cost in kubelet memory, but
otherwise is relatively clean.
This approach is what kata uses currently. Their device plugin is not
available unbundled, and we anyhow have a per-node daemonset.
atlet is taught to sniff if /dev/kvm appears on the host at all, so we
can also stop using the manually labeled nodes for microVM class and
instead schedule to the KVM + TUN devices on nodes that advertise them.
Later we can migrate to device plugin by using
I implemented that already, but I don't think it's worth merging at the
moment. We would want consumable capacity to be on by default. We can
migrate later without changing the pod spec by using
`extendedResourceName`.
https://github.com/agent-substrate/substrate/compare/main...BenTheElder:substrate:ateom-microvm-dra
NOTE: I confirmed with upstream that device plugin is not going anywhere
despite being "v1beta1", it's GA in all but name. It won't receive new
features but we don't really need anyhow. We'll move to DRA down the
line.
---
By doing this, we can drop `privileged: true` from the uVM ateom pods.
We can also drop the `ate.dev/sandboxClass` node label hacks, reducing
friction to deploy.
Per-container CPU and memory limits for micro-VM actors. First slice of
#752.
An `ActorTemplate` container can cap its own cpu and memory so it cannot
starve or kill its siblings in the same actor. A container that exceeds
its memory limit is OOM-killed on its own; the rest of the actor is
unaffected.
```yaml
sandboxClass: microvm
containers:
- name: trainer
resources:
limits: {memory: 1500Mi}
- name: sidecar
resources:
limits: {memory: 256Mi, cpu: "0.2"}
```
#752 asks for per-container device selection as well. This PR does the
cpu and memory half; devices are the remaining half, and when they land,
GPU assignment should follow the same field rather than staying
actor-wide as it is today.
## Why micro-VM only
The two sandbox classes support opposite halves of #752. Micro-VM actors
have a real guest kernel, so each container gets its own cgroup and the
limits bind. gVisor applies cgroup limits at the sandbox level: one
sentry backs every container in the actor, so a per-container cgroup is
created and then stays empty (google/gvisor#190). Measured on a running
actor, the workload container's cgroup reported `memory.current=0` while
all 20 sandbox processes sat in the pause leaf. A template that sets
`resources` with `sandboxClass: gvisor` is rejected at admission.
## How a limit travels
`ActorTemplate` → CEL validation → `ate-api-server` resolves each
`resource.Quantity` once → `ateletpb` → atelet writes OCI
`linux.resources` → `ateom-microvm` merges kata's defaults and checks
the guest envelope → `SpecToAgentPB` → kata agent → guest cgroup.
The limit is carried as a standard OCI field rather than a
substrate-private concept, so a runtime that gains per-container
enforcement picks it up without new plumbing.
## Composition with #679#679 sizes the sandbox itself from `ActorTemplate.spec.resources`: guest
RAM and vCPUs for a micro-VM, the sentry for gVisor. This PR subdivides
that sandbox. The two are different fields and complementary layers.
They collided in one place. #679 applied the actor-level size to
**every** container's OCI spec, overwriting the per-container limits
atelet writes into the same field, so a container asking for 64Mi
silently received the whole guest. This PR removes that call on the
micro-VM path. A container now gets a cgroup limit only when it declares
one; an undeclared container is bounded by guest RAM, which is the real
ceiling. On micro-VM the stamp never reached the guest at all:
`SpecToAgentPB` carried neither `Memory` nor `CPU.Quota`, so it stopped
at the bundle's `config.json`. Running the pre-PR pipeline shows
`Quota=50000 Memory=2Gi` on disk arriving as `Quota=0 Memory=<nil>` on
the wire.
gVisor is untouched. There the actor-level size is applied to the
`pause` container, and since one sentry backs every container, that is
the only cgroup that binds.
Two consequences worth calling out for reviewers of #679:
- The `sizing.SandboxSize` threaded into `buildActorContainers` and
`ensureKataCompatibleSpec` had no remaining reader, so it and the
`guestSize` helper are removed. `resolveGuestMemMiB` still sizes the VM
and still rejects a declared limit too small to boot.
- `internal/sizing`'s package doc claimed both runtimes shared
`ApplyToOCISpec` for container cgroups. That is now gVisor-only, and the
comments say so. No code in that file changed.
`checkResourceEnvelope` runs after `resolveGuestMemMiB`, so it validates
the per-container sum against the post-reserve guest rather than the
SandboxConfig default. Its error now names
`spec.resources.limits.memory` (or `.cpu`) when the actor declared a
size, and `SandboxConfig` when it did not.
`internal/e2e/suites/sizing` is unaffected: it asserts on `num_cpu` and
`mem_total_bytes`, both of which come from VM sizing, and only logs the
cgroup files.
## Verified on hardware
Same template, run twice, one commit apart on a micro-VM actor:
| | before | after |
| :--- | :--- | :--- |
| `hog_ovl/memory.max` | `max` | `67108864` (the declared 64Mi) |
| `bystander_ovl/memory.max` | `max` | `max` |
| hog allocates 128MB | survived | OOM-killed |
| bystander | alive | alive |
The bug in between: `SpecToAgentPB` converted only `Devices` and
`CPU.Shares` out of `Linux.Resources`, so a memory limit reached the
bundle's `config.json` and was dropped on the way to the agent. Every
unit test passed and the on-disk spec was correct while the feature did
nothing.
## Notes for review
- `ContainerResources` deliberately does not reuse
`corev1.ResourceRequirements`, which also carries `requests` and
`claims`. There is no scheduler inside an actor to hint at, and for
memory a soft request cannot express "must have this much to come back
at all". #679 reached the same conclusion from the other direction and
now rejects `spec.resources.requests` and `.claims` at admission.
- `Limits` is a bounded named type. A `MaxProperties` marker on the
field lands at the wrong schema level for a named map, and without a
bound the CEL cost estimator rejects the whole schema. This CRD is close
to its schema-wide CEL cost ceiling; a rule that iterates containers and
parses quantities exceeds it by more than 100x and takes the existing
rules down with it.
- A cpu limit below `10m` is raised to `10m`: the kernel rejects a CFS
quota under 1ms.
- `SpecToAgentPB` drops a non-positive quota and a zero period rather
than forwarding them. A non-positive quota means unlimited in OCI but
was sent as a literal zero, which the guest applies as no CPU at all; a
non-nil zero period overwrote the CFS default set for a live quota. Both
now read a spec the way `cpuLimitMillis` does, so the envelope check and
the conversion agree on the same input.
- `mergeKataResources` fills the gaps kata's defaults cover rather than
allowlisting known fields, so a field it does not recognise survives the
merge. That holds only as far as the merge: `SpecToAgentPB` converts
`Devices`, `Memory` and `CPU` only, so `Pids`, `BlockIO`,
`HugepageLimits` and `Network` stop at the ttrpc boundary. Nothing sets
them today.
## Known gaps
- An OOM-killed container is not reported above the guest. The actor
stays `STATUS_RUNNING` and nothing records which container died. Raised
on #550. Note that #961 settled workload stats at the sandbox level and
deliberately does not attribute per container, so surfacing which
container died needs a separate signal: the guest cgroup's
`memory.events`, not the stats path.
- A container's writable rootfs is a guest tmpfs, so filesystem writes
are charged to its memory limit and a container writing more than its
limit to `/tmp` is OOM-killed. Documented in `api-guide.md`; separating
the two budgets needs an `ephemeral-storage` field.
- Over-subscription across containers is caught when the actor starts,
not at apply time. An admission-time sum check is not implementable
within the CEL cost budget.
- The atelet hop that attaches resources to the bundle spec has no test:
replacing `ctr.GetResources()` with `nil` leaves the whole suite green.
Closing it needs an imagecache fixture for `prepareOCIDirectory`.
- Restored actors inherit the golden's cgroups rather than applying
their own spec. Correct today only because `ActorTemplateSpec` is
immutable.
- The GPU section of `api-guide.md` still says `ActorTemplate` has no
per-container resource fields, which this PR makes false. Correcting it
is left to the devices half of #752, which is what will change the GPU
behaviour that paragraph describes.
---
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.
#1011
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data
source for SystemInfo volumes (#802) and the end-to-end proof that the
projected anchors work against the MITM egress gateway. Live refresh for
running actors (PR 2) and auto-injection (PR 3) come separately.
What this adds
A SystemInfo volume data source that projects the trust anchors of a
named trust bundle to a PEM file:
volumes:
- name: trust
systemInfo:
dataSources:
- trustBundle:
name: egress-mitm.ate.dev
path: egress-ca.pem
Inspired by the Kubernetes clusterTrustBundle projected volume source,
but source-neutral: the template names a bundle; where it's fetched from
is a deployment concern, not part of the API.
Design points
- Resolution lives on the node. The wire carries only {name, path};
atelet resolves the name at write time through an informer-backed lister
on ClusterTrustBundles and writes the sanitized PEM with the temp+rename
discipline from #803 (find-paths safe). Contents refresh on every
Run/Restore. ateapi is not involved, per review discussion — the same
informer is what live refresh (PR 2) will hang off.
- Allowlist in atelet, not the CRD schema. Today only
egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the
ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946)
derives from the egress-mitm-ca-pool Secret. The signer-linked object
name stays a backend detail; the future backend registry (#932) widens
the allowlist without an API change.
- The watch is scoped to the one backing object via a metadata.name
field selector — this informer runs on every node, so an unfiltered
watch would fan every ClusterTrustBundle in the cluster out to every
atelet. RBAC can't express this (resourceNames doesn't apply to
list/watch), so the field selector is the enforcement point.
get/list/watch on clustertrustbundles moves to the atelet ClusterRole.
- No availability probe. The informer registers unconditionally; a
cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1
blocks atelet startup at cache sync, with the reflector errors naming
the missing API (hack/create-kind-cluster.sh enables the gate).
- Fail-closed. Unknown names, missing bundles, and unusable bundles fail
actor start naming the bundle — an actor that declared a trust bundle
must not start without one.
- Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks
only, deduplicated, headers stripped, and anchors deliberately shuffled
so consumers can't grow a dependence on order.
- Schema note: dataSources MaxItems tightened 32→8 while adding the
trustBundle member. Vacuous in practice (the old schema couldn't admit
more than one entry), but flagged since it's ratchet-shaped.
E2E — delivery and consumption
Delivery (identity suite, both sandbox classes): provisions the
egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing
the bundle directly isn't possible — the reconciler reverts hand-edits),
asserts the projected file byte-exact, then rotates the pool across a
suspend/resume to prove refresh-on-restore. Since the probe fixture is
shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists
for whatever suite deploys it.
Consumption (new egressmitm suite, both sandbox classes): deploys the
sdsmint (MITM) egress gateway and proves an actor completes a TLS
handshake with the gateway's per-SNI minted leaf using ONLY the
projected anchors — plus a system-roots negative control that must fail.
The pair is unambiguous in both directions: the positive can't pass
under passthrough (the bundle holds no public CAs), and the negative
can't fail under passthrough.
CI: two steps appended to the existing e2e job after the standard lanes
(the gateway swap is cluster-wide and breaks passthrough assumptions):
--deploy-atenet --experimental-use-sdsmint redeploys only the atenet
components, then the egressmitm suite runs once per sandbox class.
Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each
suite deploys its own copy and drives one actor at a time, so the third
worker per copy was idle memory multiplied across suites on the one-node
CI cluster — pressure that has been killing sandboxes mid-test (runsc:
signal: killed, a vanished ateom socket) on this PR and on main's
identity suite. This reduces the pressure; right-sizing e2e concurrency
or worker-pod QoS cluster-wide is follow-up material.
Not in this PR
- Live refresh for running actors (#932 PR 2) — until then, a running
actor's file is the bundle as of its last Run/Restore, and correctness
rests on overlap rotation by the bundle publisher.
- Auto-injection of the egress trust volume (#932 PR 3).
- Configurable backend registry (#932) — the allowlist is the seam it
will replace.
Fixes#643
This feature introduces the ability to delete Substrate actors from any
lifecycle state (e.g., RUNNING, PAUSED, PENDING), rather than requiring
them to be hibernated into SUSPENDED (or CRASHED) beforehand. This is
crucial for cleaning up stuck or failed actors.
#### 1. Control Plane & gRPC API (pkg/proto/ateapipb, ateapi)
• DeleteActorRequest.any_state: Added a new boolean field any_state to
DeleteActorRequest.
• When false (default): enforces the existing behavior where only actors
in ACTOR_STATE_SUSPENDED or ACTOR_STATE_CRASHED (or already
ACTOR_STATE_DELETING) can be deleted.
• When true: allows deleting an actor in any state (e.g., RUNNING,
PAUSED).
• Orchestrated Cleanup Workflow (workflow_delete.go):
1. Transitions actor state to ACTOR_STATE_DELETING.
2. Calls atelet.Terminate to stop live workloads on the node.
3. Detaches external volumes from the worker node (with fallback logic
using actor status if the ActorTemplate was deleted).
4. Releases the assigned physical worker Pod back to the WorkerPool.
5. Deletes external storage volumes through the storage plugins.
6. Finalizes removal of the actor from the persistence store.
#### 2. Node Herder Agent (atelet)
• Terminate RPC (main.go): Added a new Terminate method to AteomHerder:
• Calls ateom.TerminateWorkload on the target worker pod.
• Unmounts external volumes mounted for the actor on the node.
• Cleans up and resets actor runtime host directories (OCI bundles,
checkpoints, pid files).
#### 3. In-Pod Sandboxes (ateom-gvisor & ateom-microvm)
• TerminateWorkload RPC:
• gVisor (main.go): Stops and deletes active runsc containers, unmounts
bundle rootfs overlays, and tears down the actor's interior network
namespace.
• Micro-VM (checkpoint.go): Shuts down the Cloud Hypervisor VMM,
unmounts bundle rootfs overlays, and cleans up the network namespace.
• Emits an "Actor terminated" lifecycle log and clears the active actor
attribution so the worker can host new workloads.
#### 4. CLI (kubectl-ate)
• --any-state Flag (delete_actor.go):
kubectl-ate delete actor <actor-name> -a <atespace> --any-state
──────
## Verification
[x] Changes tested in local with
```
• go build ./...: Passed
go test ./...: Passed (all unit and functional tests pass)
make lint: Passed (no lint errors)
```
Detailed manual tests
```
### Scenario 1: Force Delete a Running Actor with --any-state
1. Create the actor:
kubectl ate create actor manual-test-1 --template=ate-demo-counter-microvm/counter-microvm -a demo
2. Resume the actor:
kubectl ate resume actor manual-test-1 -a demo
3. Force delete the running actor:
kubectl ate delete actor manual-test-1 -a demo --any-state
actor "manual-test-1" deleted
```
──────
### Scenario 2: Standard Delete a Suspended Actor
```
1. Create the actor:
kubectl ate create actor manual-test-2 --template=ate-demo-counter-microvm/counter-microvm -a demo
2. Resume the actor:
kubectl ate resume actor manual-test-2 -a demo
3. Suspend the actor:
kubectl ate suspend actor manual-test-2 -a demo
4. Standard delete the suspended actor:
kubectl ate delete actor manual-test-2 -a demo
actor "manual-test-2" deleted
```
──────
### Unit & E2E Tests
# Unit & Functional Tests
go test ./cmd/ateapi/internal/controlapi -run
"TestDeleteActorWorkflow|TestEnsureMarkedDeleting"
# E2E Tests
E2E_TEMPLATE_NAMESPACE=ate-demo-counter-microvm
E2E_TEMPLATE_NAME=counter-microvm ./hack/run-e2e-kind.sh
./internal/e2e/suites/demo -run TestActorLifecycle
./hack/run-e2e-kind.sh ./internal/e2e/suites/demo -run
TestForceDeleteActorWithExternalVolume
```
- [ x] Appropriate changes to documentation are included in the PR
Part of #212
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Summary
- add bounded, validated labels and annotations to
`WorkerPoolPodTemplate`
- propagate that metadata to the generated Deployment and worker pod
template
- reserve the controller-owned `ate.dev/worker-pool` label
- regenerate the WorkerPool CRD and deepcopy code
- add API validation, controller tests, and documentation
I chose intentionally to avoid exposing a complete `PodTemplateSpec` as
per the discussion in the #212.
Part of #802 (first PR: the actorIdentity data source; does not close
the issue).
## What changed and why
Adds a systemInfo volume source to ActorTemplate — a read-only volume
whose files are generated by atelet on every Run/Restore, analogous to
Kubernetes projected volumes. The initial data source, actorIdentity,
writes the actor's own name to a configurable relative path:
```
spec:
volumes:
- name: system-info
systemInfo:
dataSources:
# Part 1 (this PR): own-metadata projection, downwardAPI-style
- actorMetadata:
items:
- field: name # enum: name | atespace | uid
path: actor-name
- field: atespace
path: atespace
- field: uid
path: actor-uid
containers:
- name: main
image: app@sha256:...
volumeMounts:
- name: system-info
mountPath: /run/ate
```
Because the files are regenerated before the sandbox starts, they carry
the resumed actor's own values regardless of what checkpointed state it
boots from — the property the old hardcoded /run/ate identity mount
provided, now as an explicit, extensible API that future data sources
(identity JWTs, certificates — see #802) can slot into.
**Behavior change:** the automatic /run/ate/actor-id mount is removed;
actors must opt in by declaring the volume (the e2e identity probe in
this PR is the reference example).
### Reviewer notes:
- Over half the diff is vendored + generated code
(cmd/atelet/internal/third_party/atomicwriter/, atelet.pb.go,
zz_generated.deepcopy.go, the CRD manifest). The hand-written surface is
~700 lines.
- System-info volume roots live under a new ActorPath/system-info/ host
dir, deliberately separate from durable-dir/: the micro-VM durable
machinery snapshots everything under the durable-dir root, and generated
identity files must never be captured into snapshots.
- Supports microVM as well as gVisor.
- The e2e identity suite exercises the new API end-to-end with unchanged
probe binary and assertions. It runs in the kind-cluster CI job (not run
locally).
## Checklist
- [x] Issue is linked above
- [x] Tests pass locally (go test ./...)
- [x] Root-gated tests pass if applicable (N/A — no root-gated packages
touched)
- [x] Documentation updated if behavior changed (docs/api-guide.md:
SystemInfo Volumes section with example)
---------
Co-authored-by: Taahir Ahmed <taahm@google.com>
Actors previously ran in sandboxes sized to the whole node; there was no
way to declare how much CPU/memory a given actor should get. This adds
an explicit, immutable sizing knob on the ActorTemplate and plumbs it
into the sandbox's OCI spec for both the gVisor and micro-VM runtimes.
**API** — `ActorTemplate.spec.resources`
(`*corev1.ResourceRequirements`). The `limits` size the sandbox and are
baked into the immutable spec; the CRD and generated code are
regenerated accordingly.
**internal/sizing** (shared by both runtimes) — new `SandboxSize` value
(`FromLimits` / `VCPUs` / `ApplyToOCISpec`). `ApplyToOCISpec` writes CPU
quota+period and the memory limit onto the OCI spec, and is a no-op when
neither dimension is set, so 0 means "unconstrained". `VCPUs` rounds
milliCPU up to whole
vCPUs.
**Plumbing** — ateapi reads the template limits (`actorResourceLimits` →
`tmpl.Spec.Resources`) and supplies `CpuMilli`/`MemoryBytes` over the
actor RPCs (ateapi → atelet → ateom); ateom applies them — gVisor via
the cgroup leaf (`runsc --cpu-num-from-quota` provisions the sentry vCPU
count), micro-VM via the guest VmConfig. The two fields are carried on
the proto messages.
**Scheduling** — worker capacity is taken from the WorkerPool's
per-worker limits and advertised on the Worker; the scheduler only
places an actor on a worker whose capacity >= the actor's declared
limits. A missing worker or actor
dimension is treated as unconstrained, so placement is never blocked by
absent data.
**Docs & demos** — document the model in `api-guide.md`; the counter,
sandbox, and micro-VM demos declare actor limits, and their WorkerPool
comments now describe the real model (worker limits size the worker pod
+ advertise scheduling capacity; the sandbox itself is sized by
`ActorTemplate.spec.resources`).
- [x] Tests pass (unit tests for sizing + scheduling; e2e suite in
`internal/e2e/suites/sizing` resumes an actor and asserts, via the probe
fixture's `/resources` endpoint, that the running sandbox observes the
declared CPU/memory from the inside)
- [x] Appropriate changes to documentation are included in the PR
This reverts the ActorTemplate valueFrom.secretKeyRef support added in
#20 (issue #15).
We don't want actors to have any access to secrets, they will be
injected on the egress route instead. Removing this now so we don't need
to copy secrets into substrate resources.
The pause image holds the sandbox's namespaces and runs no workload
code. It is an implementation detail of the sandbox, not something actor
authors pick, so it belongs with the sandbox binaries that already moved
off the ActorTemplate onto the cluster-scoped SandboxConfig.
It now travels with those binaries end to end: resolved from the pool's
SandboxConfig, carried on ateletpb.SandboxAssets rather than
WorkloadSpec, and recorded in the per-actor sandbox record so Checkpoint
pins it into the snapshot manifest and Restore rebuilds the sandbox from
the image the snapshot was taken with (the golden's on a DATA_ON_GOLDEN
restore). A record without one is rejected outright rather than pulling
an empty image.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#773
Some improvements to the way we create, track and propagate snapshot
ids.
* Consolidate snapshot id and name into a single concept.
* Use consistent terminology across the whole stack (i.e. remove
confusion about prefix vs id vs name).
* Get rid of internal only storage to hide the snapshot URI.
## Summary
- `atecontroller` propagates a pool's `nvidia.com/gpu` request onto the
`ateom` container and mounts the host NVIDIA toolkit read-only (path
overridable via `ATE_NVIDIA_TOOLKIT_HOST_PATH`)
- `ateom-gvisor` generates a CDI spec with `nvidia-ctk` and injects the
device nodes, driver-library mounts, and env into each actor container's
OCI spec
- runs the CDI `createContainer` hooks except `update-ldcache` which
needs a privileged ateom, staging the SONAME symlinks it would create
from each library's ELF `DT_SONAME`
- enables `runsc --nvproxy` at sandbox creation
Requesting `nvidia.com/gpu` on the pool is the only configuration
needed; a pool that requests N GPUs makes all N usable.
Two details of the CDI spec are worth calling out, because getting
either wrong fails at runtime rather than at parse time. `nvidia-ctk`
leaves `major`/`minor` unset — CDI delegates that to the OCI runtime —
so each device node is resolved by stat-ing the host; without it the
actor gets `0,0` char devices and NVML reports it cannot communicate
with the driver. And it emits per-index, per-UUID, and `all` devices
that repeat the same nodes, so only `all` is applied. The spec is plain
JSON, so `encoding/json` suffices and no CDI library is vendored.
`update-ldcache` is the one hook that cannot run here: its `ldconfig`
unshares a mount namespace and mounts a private `/proc`, which
`mount_too_revealing()` rejects under the pod's masked `/proc`.
Permitting it would need `procMount: Unmasked`, which Kubernetes only
allows with `hostUsers: false`, and that user namespace breaks the
per-actor cgroup delegation from #496. Skipping it avoids the whole
chain, so a GPU worker keeps the same posture as any other unprivileged
gVisor worker. `create-symlinks` and `enable-cuda-compat` still run
unmodified.
`--nvproxy` must be set when the sandbox is created — the `pause`
container, which holds no GPU devices — so runsc's auto-detection never
fires on its own; without the flag the GPU subcontainer crashes the
sentry on start. GPU detection matches any device index rather than
assuming `/dev/nvidia0`, since a worker sharing a multi-GPU node can be
assigned `/dev/nvidia2` and `/dev/nvidia3`.
GPU pools must set `spec.ateomImage` to a glibc build
(`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`)
because the distroless default cannot exec `nvidia-ctk`; the default
base is unchanged for every other pool. `atelet` also has to run on the
GPU nodes to restore actors there, so its DaemonSet needs a toleration
for whatever taint they carry. Both are documented in the API guide
rather than defaulted.
## Testing
- `make test`
- `env -u NO_COLOR make verify`
- Real GPU, Tesla T4 / driver 580.65.06, through the full actor flow:
`nvidia-smi`, `vectorAdd`, `nbody` at 3.77 TFLOP/s, PyTorch matmul via
cuBLAS at 3.5 TFLOP/s (T4 peak FP32 is ~8.1, so no measurable sandbox
penalty)
- Actor whose entrypoint runs the CUDA sample directly exits 0,
confirming the injected env reaches the workload without help from the
test harness
- Two workers holding two GPUs each on one 4-GPU node see disjoint
device sets
Snapshot and restore work when the workload holds no CUDA context. A
live CUDA context cannot be checkpointed - gVisor fails with `can't save
with live nvproxy clients` and the failed checkpoint terminates the
sandbox so a GPU actor can only be suspended between CUDA workloads.
Documented as a known limitation in the API
guide; a follow-up issue will track lifting it via `cuda-checkpoint`.
Fixes#627
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eliran Wolff <eliranw@nvidia.com>