213 Commits
Author SHA1 Message Date
Michael Taufen bbf9d440c6 Security status report SKILL (#680) 2026-08-20 01:33:00 -07:00
Luiz Oliveira fc02127bfe ateapi - Require preconditions in Update methods (#1014)
https://github.com/agent-substrate/substrate/issues/954

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 21:01:29 -04:00
Max ThompsonandTaahir Ahmed 5edb1af1fe feat: add systemInfo volume source with actorIdentity data source (#803)
Part of #802 (first PR: the actorIdentity data source; does not close
the issue).

## What changed and why

Adds a systemInfo volume source to ActorTemplate — a read-only volume
whose files are generated by atelet on every Run/Restore, analogous to
Kubernetes projected volumes. The initial data source, actorIdentity,
writes the actor's own name to a configurable relative path:

```
spec:
  volumes:
  - name: system-info
    systemInfo:
      dataSources:
      # Part 1 (this PR): own-metadata projection, downwardAPI-style
      - actorMetadata:
          items:
          - field: name              # enum: name | atespace | uid
            path: actor-name
          - field: atespace
            path: atespace
          - field: uid
            path: actor-uid
  
  containers:
  - name: main
    image: app@sha256:...
    volumeMounts:
    - name: system-info
      mountPath: /run/ate  
```

Because the files are regenerated before the sandbox starts, they carry
the resumed actor's own values regardless of what checkpointed state it
boots from — the property the old hardcoded /run/ate identity mount
provided, now as an explicit, extensible API that future data sources
(identity JWTs, certificates — see #802) can slot into.

**Behavior change:** the automatic /run/ate/actor-id mount is removed;
actors must opt in by declaring the volume (the e2e identity probe in
this PR is the reference example).

### Reviewer notes:

- Over half the diff is vendored + generated code
(cmd/atelet/internal/third_party/atomicwriter/, atelet.pb.go,
zz_generated.deepcopy.go, the CRD manifest). The hand-written surface is
~700 lines.
- System-info volume roots live under a new ActorPath/system-info/ host
dir, deliberately separate from durable-dir/: the micro-VM durable
machinery snapshots everything under the durable-dir root, and generated
identity files must never be captured into snapshots.
- Supports microVM as well as gVisor.
- The e2e identity suite exercises the new API end-to-end with unchanged
probe binary and assertions. It runs in the kind-cluster CI job (not run
locally).

## Checklist

- [x] Issue is linked above
- [x] Tests pass locally (go test ./...)
- [x] Root-gated tests pass if applicable (N/A — no root-gated packages
touched)
- [x] Documentation updated if behavior changed (docs/api-guide.md:
SystemInfo Volumes section with example)

---------

Co-authored-by: Taahir Ahmed <taahm@google.com>
2026-08-19 15:09:33 -07:00
Noureldin 01c92e9da3 Allow actors to opt in/out of Linux capabilities (#795)
Fixes #744
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 13:47:54 -07:00
Maya Wang c618baa593 docs: convention for integration repository structure and naming (#712)
## Summary

Adds `docs/integration-repos.md`: where end-to-end integrations live,
how their
repositories are named, and how the fixes they need flow back into core.

The convention in one line — trivial demos stay in the core repo, each
non-trivial integration gets one dedicated repo under the
`agent-substrate`
org, and core gaps get closed by making core configurable with defaults
unchanged rather than by patching it downstream.

## Why now

We are about to create the first real, end-to-end integrations rather
than
counter-style demos: a code-execution sandbox, and an always-on agent.
Both are
large enough to need their own images, dependencies, and release
cadence.

Whichever repository gets created first will set the precedent for every
one
after it. This writes the convention down so that precedent is chosen
deliberately instead of inherited by accident.

## What it covers

- **Where code lives** — the core-repo/dedicated-repo split, the rough
test for
which side something falls on (API keys, external services, third-party
accounts), and why this is a set of peer repos rather than a second org.
- **Naming** — capability-named for general capabilities
  (`code-execution-sandbox`), integration-named for specific third-party
products, named for the product rather than the vendor behind it. Plus
what to
avoid: over-broad names, names that clone a vendor's API or brand, and
the
  redundant `-integration` suffix.
- **Third-party names** — allowed descriptively, with a non-affiliation
note in
the repo README, and brand/policy edge cases cleared before the repo
exists.
- **Upstreaming** — the part with teeth for this repo. Integration repos
that
accumulate local patches against core bitrot, and the gap they work
around
stays invisible to everyone else. So: prefer making core behavior
configurable
with defaults unchanged. #487 and #465 are linked as illustrations of
that
  pattern — this PR does not depend on either, and branches from `main`.
- **Two worked examples** that validate the convention rather than just
  following it, including the third-party-name edge case.

## Review

This was announced at the community meeting and circulated as a shared
design
doc with a 7-day review window, which has now closed. It synthesizes the
`#integrations` thread discussion. Comment history:


<https://docs.google.com/document/d/1Tb6u0b1XSvWrNpoyD4jdsQaJ58aAgDtQOM18uxujs-8/edit>

This PR is the trimmed version: doc-review scaffolding — status block,
reviewer
list, self-link — is dropped, and only the durable convention is carried
over.

## Left open

Two questions are deliberately out of scope, called out in the doc
rather than
answered. Both are maintainer calls and neither blocks the first
repositories:

- Governance tiers — whether to distinguish "official" from "community"
integrations with different review bars, as Home Assistant and Obsidian
do.
- Who creates integration repositories and grants per-integration
maintainer
  access.

## Also in this PR

- README gets an entry in the docs list, matching every other file in
`docs/`.
- `CONTRIBUTING.md` gets one sentence pointing there, since "where does
my
  integration go?" is a question a contributor asks before opening a PR.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 11:24:41 -07:00
Krisztian F c3c6876985 (chore): consolidate logging to ateattr (#1035)
Actor logs used `ate.dev/actor_*` while spans and metrics use `ate.*`
registry in internal/ateattr. This PR makes ateattr the single source of
truth for everything telemetry-related.

- Renamed the six actor log labels onto the registry: 
- `ate.atespace`, `ate.actor.name`, `ate.actor.uid`,
`ate.template.namespace`, `ate.template.name`,
`ate.actor.container.name`
- Logs join traces now. Records set `trace_id`, `span_id` and
`trace_flags`, so you can go from Actor restored to the resume that
caused it. Our own lines only, not an actor's stdout: one goroutine
forwards a whole container stream and can't know which request produced
a given line. Per line correlation comes with #853.
- Actors can't fake platform labels. They already couldn't overwrite
ours, but they could invent new ones like `ate.tenant` that look
platform issued downstream. Anything under `ate.` from an actor is now
dropped.
- Fixed the asymmetry that was actually left: actor supplied label
values weren't stringified, and one non string value makes Cloud Logging
discard the labels for that whole entry.
- Note: the actor_uid bullet in the issue is stale, #841 fixed it
earlier. Lifecycle records still set five labels rather than six, on
purpose as they're about the actor, so no container produced them.

Fixes #886

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: krisztianfekete <git@krisztianfekete.org>
2026-08-19 11:02:10 -04:00
Julian Gutierrez Oschmann 4c1bd9d36b Add status fields to Substrate resources (#1025)
Add `status` fields to all Substrate resources that need it. Move
server-owned fields under it.

Fixes #1006 .
2026-08-18 11:58:31 -07:00
ILLUM1N0X 2af5308b4d ateapi: return ResourceExhausted when no worker is free (#908)
ResumeActor returned `FailedPrecondition` both when the fleet had no
free worker and when the actor was not in a resumable state. Only the
former is fixable by suspending the actor to drop its node pinning, and
a client driving that recovery could not tell the two apart.

`assignWorkerAttempt()` now returns `ResourceExhausted` for
`scheduling.ErrNoCapacity`. `FailedPrecondition` was the de-facto
capacity signal, so three router consumers move with it:

- `retryable()`: `ResourceExhausted` is parked, so saturation behaves as
before. `FailedPrecondition` stays retryable on purpose. It no longer
carries saturation, but it does cover states a concurrent operation can
move the actor out of.
- `mapResumeError()`: `ResourceExhausted` now preserves the gRPC
description, keeping saturation at 503 as today. Its former 429 "rate
limited" branch was written speculatively as nothing in substrate
produced the code until now, and the fleet being full is not the caller
sending too many requests.
- `classifyOutcome()`: `no_capacity` moves to `ResourceExhausted`, and
`FailedPrecondition` gets its own label so the capacity bucket stops
collecting non-capacity errors.

`TestResumeActor_RelocatesAfterSuspendFromPaused` covers the recovery
flow this signal exists for: a PAUSED actor pinned to a full node cannot
resume, and after a suspend clears the pinning it schedules onto a
worker on another node.

Fixes #660

- `make test` clean (except the pre-existing macOS-only
`internal/atunnel` failures, untouched by this PR)
- `make fmt` clean
- `go vet ./cmd/...` clean
2026-08-17 15:00:15 -07:00
Jeff LuoandKrisztian F 9bfc67e815 benchmarking: add observability benchmarking for the OTel collector (#749)
Part of #563

Substrate had no way to measure how much telemetry it sends. This PR
adds one, and builds it from the parts the benchmark suite already has.

The managed collector sends its self-metrics to Cloud Monitoring only,
and gives no data for each source. Thus you cannot see which component
makes the volume. A second collector, the telemetry meter, counts the
spans and the datapoints by `service.name`.

## Contents

- **`benchmarking/telemetry/meter.yaml`** — the meter. A tee: it counts
the telemetry by service, then sends it on to the managed collector.
Thus one run gives the volume for each service and the cost of the load
together.
- **`benchmarking/telemetry/README.md`** — the method: install, measure,
read, remove.
- **`hack/install-ate.sh --otlp-endpoint`** — sends all control plane
telemetry to the meter. Each component reads the endpoint from the
`ate-otel-config` ConfigMap through `envFrom`, and `ate-controller`
copies the value to the ateom worker pods, thus one patch is sufficient
for all of them.
- **`benchmarking/workloads/deploy.sh --otlp-endpoint`** — the same for
actor containers, which substrate configures separately.
- **`benchmarking/monitoring.yaml`** — scrapes the meter and cAdvisor.
- **`benchmarking/observability.md`** — prerequisites and the scenario
ladder.
- **`docs/dev/best-practices/otel-collector.md`** — the volume model and
the size guidance, corrected for the `ParentBased` defaults from #711.

## Correction

The `glutton` ActorTemplate had no `env`. Substrate puts no OTLP
configuration in an actor container, thus the exporter used the SDK
default `localhost:4317`, no process listens on that address in the
sandbox, and substrate discarded each actor span with no error message.
The template now sets `OTEL_EXPORTER_OTLP_ENDPOINT` and
`OTEL_RESOURCE_ATTRIBUTES`.

`ActorTemplate.spec` is immutable, thus `deploy.sh --deploy` deletes the
templates before it applies them.

## How I verified this

On a GKE cluster with managed OTel. The sequence:

```bash
# 1. Start from a clean state.
./hack/install-ate.sh --delete-benchmarks
./hack/install-ate.sh --delete-ate-system

# 2. Install substrate from this branch.
./hack/install-ate.sh --deploy-ate-system

# 3. Install the meter and Prometheus.
kubectl apply -f benchmarking/telemetry/meter.yaml
kubectl apply -f benchmarking/monitoring.yaml

# 4. Send the control plane to the meter.
METER=http://telemetry-meter.benchmarking.svc.cluster.local:4317
./hack/install-ate.sh --deploy-ate-system --otlp-endpoint "${METER}"

# 5. Install the workloads and locust, with the actors on the meter also.
./benchmarking/deploy_locust.sh --deploy --worker-count 10 \
    --otlp-endpoint "${METER}"

# 6. Read the volume.
kubectl port-forward -n benchmarking svc/prometheus 9090:9090
```

## Result

The idle floor, from one measurement on 10 workers:

| Service | datapoints/min |
|---|---|
| `ate-controller` | 234.8 |
| `ateapi` | 54.3 |
| `atelet` | 44.9 |
| `atenet-router` | 10.4 |
| `ateom-gvisor` | 5.2 |
| `glutton-actor` | 4.2 |
| **Total** | **353.7** |

The `glutton-actor` row is the correction above. Before it, actor
telemetry could not arrive at all.

Positive control for the same window:
`otelcol_receiver_accepted_metric_points` 10.5/s, and each
`otelcol_receiver_refused_*` counter at zero. Thus the zero values in
the table are true zeros, and not data that did not arrive.

Two values differ from the earlier manual measurement. `ate-controller`
is now the largest source, and `ateom-gvisor` is no longer zero. The new
imagecache counters from #831 are a probable cause. This difference is
the reason to keep results out of the repository.

### Traces

Trace volume follows the root sample decision, not the rate of each
component. `atelet` and `ateom-gvisor` sent zero spans in each run,
because they inherit the decision that `ateapi` makes and start no trace
of their own. At idle each component is near zero: a trace needs an
operation, but a metric pushes on a timer.

This run gives no span volume for a usual load. The actors crashed on
suspend, and the failed retries made an artificial rate of 13.7k
spans/s. That rate is not a measurement of substrate. It does show that
the meter accepted the spans with no refusals.

## Follow-up

A separate PR adds the ladder to `automation/tests.yaml` and writes the
results as artifacts of each run. The open items are at the end of
`observability.md`.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Krisztian F <103492698+krisztianfekete@users.noreply.github.com>
2026-08-17 16:06:09 -04:00
Jeff Luo 6dd24af84a ateapi: pair ate.workerpool.namespace with ate.workerpool.name (#953)
ate.actor.crashes, ate.actor.lifecycle.operation.duration and
ate.scheduler.assignment.duration named a WorkerPool by name alone. A
WorkerPool is namespaced, so same-named pools in different namespaces
merged into one series, and the three could not join the instruments
that already carry both keys.

ateattr.WorkerPoolAttributes now builds the pair, and omits both keys
when no pool is assigned, so a crash before the actor reached a worker
no longer reports an empty-string pool. This changes the series identity
of the three instruments.

Fixes #951

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-17 09:18:53 -04:00
Keith Mattix II 89b83a5c89 Support CONNECT in atenet router (#715)
added CONNECT support for atenet ingress to support arbitrary actor ports
2026-08-14 16:12:27 -07:00
Chenyi Wang c178f5b74e ateom: export telemetry through an atelet unix-socket OTLP relay (#809)
ateom runs inside the worker pod that hosts the actor, and exported OTLP
straight to the collector over the pod's network. This adds a node-local
relay: ateom pushes OTLP/gRPC over a unix socket that atelet serves and
forwards to the collector, so a worker pod needs no network path of its
own to export spans and metrics.

Four things motivate it:

- Blast radius. The pod runs untrusted agent code, so allowing it egress
to the collector makes the collector reachable to anything that escapes
the sandbox. A unix socket cannot leave the node.
- Connection count. Worker pods are heavily oversubscribed; N ateoms per
node each held their own collector connection. They collapse into
atelet's single per-node one.
- Interference. ateom transparently redirects actor egress to its own
atunnel listener, and its own outbound traffic has to stay clear of the
rules it installs. A unix socket is not IP traffic.
- Shutdown loss. Teardown frees the actor's network and then the pod
goes away, which is when the spans describing teardown are still queued
in the batch processor. atelet outlives the worker pod.

The relay forwards the OTLP request verbatim rather than decoding and
re-exporting, so each ateom's own resource (service.name,
service.instance.id) survives instead of being absorbed into atelet's.

It is best-effort: an ateom that finds no socket at startup logs it and
exports directly to OTEL_EXPORTER_OTLP_ENDPOINT as before, so this is a
no-op for a cluster running an older atelet. atelet likewise declines to
serve a relay when no collector is configured, since it would accept
spans only to drop them. Both halves stay off with
--otlp-relay-socket="".

The socket lives in ateompath.BasePath, the hostPath already mounted at
the same path into atelet and into every ateom pod, so no new volume or
controller change is needed.

Also includes:
  - End-to-end tests covering the full serverboot-to-collector path.
  - Observability documentation updates for Jaeger tracing.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-14 14:33:14 -07:00
Chenyi Wang 71b006309f Right-size actor sandboxes to declared ActorTemplate resource limits (#679)
Actors previously ran in sandboxes sized to the whole node; there was no
way to declare how much CPU/memory a given actor should get. This adds
an explicit, immutable sizing knob on the ActorTemplate and plumbs it
into the sandbox's OCI spec for both the gVisor and micro-VM runtimes.

**API** — `ActorTemplate.spec.resources`
(`*corev1.ResourceRequirements`). The `limits` size the sandbox and are
baked into the immutable spec; the CRD and generated code are
regenerated accordingly.

**internal/sizing** (shared by both runtimes) — new `SandboxSize` value
(`FromLimits` / `VCPUs` / `ApplyToOCISpec`). `ApplyToOCISpec` writes CPU
quota+period and the memory limit onto the OCI spec, and is a no-op when
neither dimension is set, so 0 means "unconstrained". `VCPUs` rounds
milliCPU up to whole
vCPUs.

**Plumbing** — ateapi reads the template limits (`actorResourceLimits` →
`tmpl.Spec.Resources`) and supplies `CpuMilli`/`MemoryBytes` over the
actor RPCs (ateapi → atelet → ateom); ateom applies them — gVisor via
the cgroup leaf (`runsc --cpu-num-from-quota` provisions the sentry vCPU
count), micro-VM via the guest VmConfig. The two fields are carried on
the proto messages.

**Scheduling** — worker capacity is taken from the WorkerPool's
per-worker limits and advertised on the Worker; the scheduler only
places an actor on a worker whose capacity >= the actor's declared
limits. A missing worker or actor
dimension is treated as unconstrained, so placement is never blocked by
absent data.

**Docs & demos** — document the model in `api-guide.md`; the counter,
sandbox, and micro-VM demos declare actor limits, and their WorkerPool
comments now describe the real model (worker limits size the worker pod
+ advertise scheduling capacity; the sandbox itself is sized by
`ActorTemplate.spec.resources`).

- [x] Tests pass (unit tests for sizing + scheduling; e2e suite in
`internal/e2e/suites/sizing` resumes an actor and asserts, via the probe
fixture's `/resources` endpoint, that the running sandbox observes the
declared CPU/memory from the inside)
- [x] Appropriate changes to documentation are included in the PR
2026-08-14 14:32:26 -07:00
Eitan Yarmush 2b3a4715c6 Configurable JWT authentication to ateapi (#757)
added configurable JWT authentication to ateapi.
2026-08-13 16:44:00 -07:00
Jeff Luo 1da82f3e6a imagecache: add ate.imagecache.requests hit and miss telemetry (#831) (#834)
Add the `ate.imagecache.requests` counter.

  **`ate.imagecache.outcome`** (new key in `internal/ateattr`)

  | Value | Meaning |
  |---|---|
| `hit` | the node holds a complete image record: each layer directory
that the record names is present |
  | `miss` | the lookup must pull |
| `error` | the lookup failed; the only outcome that carries
`error.type` |
| `cancelled`, `timeout` | the caller gave up, so the cache is not at
fault |

A failed lookup is neither a hit nor a miss, so it gets its own outcome,
as
`no_free_worker` does on `ate.scheduler.outcome`. `cancelled` and
`timeout` are
outcomes for the same reason they are on `ate.router.outcome`. The hit
ratio is
therefore `hit / (hit + miss)`, with failures and abandoned lookups out
of the
  denominator.

  **`error.type`** — set only on the `error` outcome.

  | Value | Meaning |
  |---|---|
| `404`, `401`, `429`, ... | the registry rejected the request; its own
HTTP status, reported verbatim |
  | `_OTHER` | the failure carries no status of its own |

Fixes #831

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-12 09:50:30 -04:00
Lucky Abolorunke bce9366b3d docs: add guide for running the microVM runtime locally (KVM / Lima) (#743)
### Why

The README Quickstart covers getting a local Substrate cluster running
with
the default gVisor runtime, but there's no documentation for the
micro-VM
runtime, where the friction can't be scripted away: it requires
`/dev/kvm`
(bare metal, nested virtualization, or Lima on macOS), and on Apple
Silicon
the Lima configuration and asset assembly are non-obvious. New
contributors
  have to reverse-engineer the `hack/` scripts.

  ### What
  
Adds `docs/dev/microvm-local.md` — "Running the microVM runtime locally"
— a
focused guide covering only the microVM delta, based on notes from real
onboarding runs. General setup is deferred to the README Quickstart
(listed
  as the guide's prerequisite) rather than duplicated:

- **Option A: Linux host with KVM** — verifying `/dev/kvm` and CPU virt
support, the rootless-Docker caveat for the KVM probe, cluster creation,
    and the one-shot `run-microvm-demo-kind.sh` bring-up.
- **Option B: Apple Silicon macOS via Lima** — nested virtualization
with the
guest image pinned to Ubuntu 25.10 (until the kernel issue in the
default
image is fixed), `vzNAT` networking, and assembling the arm64 assets
inside
the Lima guest (`assemble.sh` requires a Linux host of the target arch).
- **Trying it out** — defers to the next steps the demo script prints
and
links the counter demo's micro-VM variant, instead of duplicating those
    commands.
- **Troubleshooting** — symptom → root cause → fix entries actually hit
    during onboarding (rootless-Docker KVM probe failures, the `aws` CLI
staging requirement until #804 lands, arm64 `virtiofsd` build deps, M1
    lacking FEAT_NV2).
2026-08-12 08:39:41 -04:00
Zoe Zhao 4b3423c01a Remove the secretKeyRef env sources feature (#835)
This reverts the ActorTemplate valueFrom.secretKeyRef support added in
#20 (issue #15).

We don't want actors to have any access to secrets, they will be
injected on the egress route instead. Removing this now so we don't need
to copy secrets into substrate resources.
2026-08-11 16:13:30 -07:00
AngelaandJeff Luo d4f2cf1c51 feat: Add control plane telemetry - workerpool desired and ready workers gauges (#826)
Part of #564 (Part 4, Fixes #564)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

## Description
This PR implements Part 4 of #564 by adding OpenTelemetry gauge
instrumentation for `ate.workerpool.desired_workers` and
`ate.workerpool.ready_workers` in `atecontroller`.

## Key Changes:
* Added `ReadyReplicas int32` to `WorkerPoolStatus` with
`+kubebuilder:printcolumn:name="Ready"` annotation in
`pkg/api/v1alpha1/workerpool_types.go` and regenerated the CRD manifest.
* Updated `WorkerPoolReconciler.syncStatus` to synchronize
`dep.Status.ReadyReplicas` into `wp.Status.ReadyReplicas`.
* Implemented `InitMetrics(meter)` registering
`ate.workerpool.desired_workers` and `ate.workerpool.ready_workers` as
OpenTelemetry Observable UpDownCounters (`{worker}`) with an
asynchronous observer callback that samples controller-runtime's local
informer cache (`r.Client.List`).
* Labeled metric datapoints using centralized attributes
`ateattr.WorkerPoolNamespaceKey` and `ateattr.WorkerPoolNameKey`.

## Testing
* `go test -buildvcs=false ./cmd/ateapi/internal/controlapi/...`
* `go test -buildvcs=false ./cmd/atenet/internal/router/...`
* `make test`
## E2E  Test
* `./hack/create-kind-cluster.sh`
* `./hack/install-ate-kind.sh --deploy-ate-system --deploy-demo-counter`
* `./hack/run-e2e.sh ./internal/e2e/suites/metrics/...`

---------

Co-authored-by: Jeff Luo <jeffluoo@google.com>
2026-08-11 16:02:50 -04:00
Benjamin Elder da8414bbd2 api: move the pause image from ActorTemplate to SandboxConfig (#848)
The pause image holds the sandbox's namespaces and runs no workload
code. It is an implementation detail of the sandbox, not something actor
authors pick, so it belongs with the sandbox binaries that already moved
off the ActorTemplate onto the cluster-scoped SandboxConfig.

It now travels with those binaries end to end: resolved from the pool's
SandboxConfig, carried on ateletpb.SandboxAssets rather than
WorkloadSpec, and recorded in the per-actor sandbox record so Checkpoint
pins it into the snapshot manifest and Restore rebuilds the sandbox from
the image the snapshot was taken with (the golden's on a DATA_ON_GOLDEN
restore). A record without one is rejected outright rather than pulling
an empty image.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-11 10:14:03 -04:00
Lior Lieberman 24cc538dc1 fixes 2026-08-10 23:28:07 -07:00
Dmitry Berkovich 3a2d0c1d86 Support suspending a PAUSED actor without waking it (#816)
Part of #791 — the last planned piece of the [implementation
plan](https://github.com/agent-substrate/substrate/issues/791#issuecomment-5226674097)
(after #810, #812, #813). #791 stays open until #817 (gVisor
Full-capture → Data-commit conversion, blocked on #790) is done; this PR
covers micro-VM fully and gVisor for scope-matched suspends. Also part
of the actor state machine (#119) and a prerequisite for system upgrade
flows (#473).

`SuspendActor` now accepts a PAUSED actor: instead of checkpointing a
running workload, ateapi dials the atelet on the node holding the pause
snapshot and has it upload the node-local files to object storage, then
finalizes as usual — durable `ActorSnapshot`, `SUSPENDED` status, node
pinning cleared.

Two commits, reviewable independently:

## Commit 1 — atelet: `UploadPausedCheckpoint` RPC (dead code until
commit 2)

- New `AteomHerder` RPC: a pure disk→object-storage copy driven by the
snapshot's self-describing manifest — no ateom involved (the sandbox is
gone).
- **Scope conversion** dispatches per sandbox class
(`narrowFullCaptureToData`): a micro-VM FULL capture narrows to a DATA
upload by carving out `durable-dir.tar` (constant hoisted to
`ateompath`, shared with ateom-microvm); gVisor returns `Unimplemented`
until split checkpoints land (#790); DATA can never widen to FULL; a
scope-less manifest (older atelet) is rejected rather than guessed at.
- **Idempotent retry**: local files gone + remote manifest present ⇒ a
previous invocation committed, succeed; gone on both sides ⇒
unrecoverable (`LOCAL_SNAPSHOT_GONE`, crashes the actor). Upload
failures stay plain retryable errors; the manifest uploads last as the
commit marker, never in parallel.
- The golden atespace is rejected at validation (fully on the
`field.ErrorList` framework): golden actors are never paused.

## Commit 2 — control plane: enable suspend from PAUSED

- `FromPaused` discriminator: PAUSED status, or SUSPENDING with no
worker assignment and a `LocalSnapshotInfo` — the field alone is stale
on resumed-from-pause RUNNING actors, so the nil-assignment conjunct is
load-bearing.
- `MarkSuspendingStep` accepts PAUSED and rejects a Data-captured pause
against a Full commit *before* the actor leaves PAUSED (an upload cannot
fabricate memory), using the `content_scope` recorded at pause (#812)
with an `onPause` fallback.
- `CallAteletSuspendStep` paused branch dials by node
(`DialForAteletOnNode`, #813): missing node record ⇒ crash (the snapshot
can never be found); unreachable atelet ⇒ retryable; atelet's
`LOCAL_SNAPSHOT_GONE` ⇒ crash via `maybeCrashActor`.
`FinalizeSuspendedStep` needed no changes thanks to the #813 hoist.
- Root-cause guard: `MarkPausingStep` rejects pausing golden-atespace
actors.
- Docs: pause states + the new `PAUSED → SUSPENDING` edge in the
architecture state diagram; glossary Suspend entry covers both origins.

## Tests

- **atelet unit**: 10 upload-helper subtests (conversion matrix,
idempotency probe, data-loss crash, retryable upload failure) via a
recording object-storage fake; validation table.
- **control-plane unit**: discriminator table, scope-rejection table
(incl. onPause fallback), paused preconditions (no node ⇒ CRASHED, no
atelet ⇒ `ErrNoAteletOnNode` + still SUSPENDING), golden-pause
rejection, prerequisite matrix updated.
- **functional (envtest)**: `TestSuspendActor_FromPaused` (upload called
with the pause snapshot name, no Checkpoint RPC, SUSPENDED, pinning
cleared, ActorSnapshot at the upload destination) +
retry-after-failed-upload (same destination on retry).
- **e2e (demo suite)**: the lifecycle driver gains a suspend-from-PAUSED
mode; three durable-dir cases — Full/Full, Data/Data, and the micro-VM
Full→Data extraction — assert memory/file counters survive the full
pause→suspend→resume journey and the node pinning is gone.

`go test -race ./...` clean (except the pre-existing macOS-only
`internal/atunnel` unix-socket-path failures, untouched by this PR),
`gofmt`/`go vet` clean, protos regenerated via `hack/protoc.sh`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-10 22:34:03 -07:00
Julian Gutierrez Oschmann ddeee6c951 Some improvements to the way we allocate and track snapshot ids (#784)
Fixes #773 

Some improvements to the way we create, track and propagate snapshot
ids.

* Consolidate snapshot id and name into a single concept.
* Use consistent terminology across the whole stack (i.e. remove
confusion about prefix vs id vs name).
* Get rid of internal only storage to hide the snapshot URI.
2026-08-08 12:25:45 -07:00
eliranw 934c7418d0 feat: CDI-based NVIDIA GPU passthrough into gVisor actor containers (#502)
## Summary

- `atecontroller` propagates a pool's `nvidia.com/gpu` request onto the
`ateom` container and mounts the host NVIDIA toolkit read-only (path
overridable via `ATE_NVIDIA_TOOLKIT_HOST_PATH`)
- `ateom-gvisor` generates a CDI spec with `nvidia-ctk` and injects the
device nodes, driver-library mounts, and env into each actor container's
OCI spec
- runs the CDI `createContainer` hooks except `update-ldcache` which
needs a privileged ateom, staging the SONAME symlinks it would create
from each library's ELF `DT_SONAME`
- enables `runsc --nvproxy` at sandbox creation

Requesting `nvidia.com/gpu` on the pool is the only configuration
needed; a pool that requests N GPUs makes all N usable.

Two details of the CDI spec are worth calling out, because getting
either wrong fails at runtime rather than at parse time. `nvidia-ctk`
leaves `major`/`minor` unset — CDI delegates that to the OCI runtime —
so each device node is resolved by stat-ing the host; without it the
actor gets `0,0` char devices and NVML reports it cannot communicate
with the driver. And it emits per-index, per-UUID, and `all` devices
that repeat the same nodes, so only `all` is applied. The spec is plain
JSON, so `encoding/json` suffices and no CDI library is vendored.

`update-ldcache` is the one hook that cannot run here: its `ldconfig`
unshares a mount namespace and mounts a private `/proc`, which
`mount_too_revealing()` rejects under the pod's masked `/proc`.
Permitting it would need `procMount: Unmasked`, which Kubernetes only
allows with `hostUsers: false`, and that user namespace breaks the
per-actor cgroup delegation from #496. Skipping it avoids the whole
chain, so a GPU worker keeps the same posture as any other unprivileged
gVisor worker. `create-symlinks` and `enable-cuda-compat` still run
unmodified.

`--nvproxy` must be set when the sandbox is created — the `pause`
container, which holds no GPU devices — so runsc's auto-detection never
fires on its own; without the flag the GPU subcontainer crashes the
sentry on start. GPU detection matches any device index rather than
assuming `/dev/nvidia0`, since a worker sharing a multi-GPU node can be
assigned `/dev/nvidia2` and `/dev/nvidia3`.

GPU pools must set `spec.ateomImage` to a glibc build
(`KO_DEFAULTBASEIMAGE=debian:stable-slim ko build ./cmd/ateom-gvisor`)
because the distroless default cannot exec `nvidia-ctk`; the default
base is unchanged for every other pool. `atelet` also has to run on the
GPU nodes to restore actors there, so its DaemonSet needs a toleration
for whatever taint they carry. Both are documented in the API guide
rather than defaulted.

## Testing

- `make test`
- `env -u NO_COLOR make verify`
- Real GPU, Tesla T4 / driver 580.65.06, through the full actor flow:
`nvidia-smi`, `vectorAdd`, `nbody` at 3.77 TFLOP/s, PyTorch matmul via
cuBLAS at 3.5 TFLOP/s (T4 peak FP32 is ~8.1, so no measurable sandbox
penalty)
- Actor whose entrypoint runs the CUDA sample directly exits 0,
confirming the injected env reaches the workload without help from the
test harness
- Two workers holding two GPUs each on one 4-GPU node see disjoint
device sets

Snapshot and restore work when the workload holds no CUDA context. A
live CUDA context cannot be checkpointed - gVisor fails with `can't save
with live nvproxy clients` and the failed checkpoint terminates the
sandbox so a GPU actor can only be suspended between CUDA workloads.
Documented as a known limitation in the API
guide; a follow-up issue will track lifting it via `cuda-checkpoint`.

Fixes #627 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
2026-08-07 15:08:44 -07:00
shrutiyam-glitch 5f64ba4c26 feat(atenet): graceful termination of atenet-router (#774)
### Which issue(s) this PR is related to:
Fixes #721
Required for System Upgrade flow (#473)

### What this PR does / why we need it:
This PR implements graceful termination and zero-outage rolling updates
for `atenet-router` , building on the readiness probe and graceful drain
patterns established in #719 (`atelet`).

Draining a two-container networking pod (`atenet-router` Go
control-plane + `envoy` C++ proxy dataplane) introduces complex
lifecycle interdependencies. This PR resolves the dual-container SIGTERM
race, respects Envoy's `failClosed` `ext_proc` filter dependency,
preserves parked requests riding out worker pool saturation, and
accelerates idle deployments via an event-driven file handshake.

#### 1. Event-Driven Dataplane Synchronization (`emptyDir` Marker File)
Envoy fast-exits by default on SIGTERM. To keep Envoy alive while the Go
control-plane orchestrates the drain, we configured an IPC handshake
between containers via a pod-shared `emptyDir` volume mounted at
`/var/run/atenet`:

* **Go Router Container:** Removes any stale marker at startup, then
writes `/var/run/atenet/drain-complete` when its shutdown sequence
finishes.
* **Envoy Container `preStop` Hook:** Runs `while [ ! -f
/var/run/atenet/drain-complete ]; do sleep 0.5; done`.

**Outcome:** Envoy exits as soon as — the drain is done, rather than
wasting time in a fixed worst-case sleep. If the router crashes, Kubelet
terminates the `preStop` hook at `terminationGracePeriodSeconds:
60`—slower cleanup, never a wedge.

#### 2. The Multi-Container Shutdown Sequence (`drain.go`)
Because Kubernetes issues SIGTERM to both containers at once, a
coordination state machine ensures Envoy never drops a connection and
`ext_proc` is never stopped prematurely:

| Phase / (best-case ex.) Timeline | atenet-router | envoy | K8s /
Service Status |
| :--- | :--- | :--- | :--- |
| **SIGTERM Sent**<br>($t=0\text{s}$) | Catches SIGTERM<br>• Flips
`/readyz` $\rightarrow$ `503`.<br>• Starts 13s `drain-delay`. | Enters
`lifecycle.preStop` hook:<br>`while [ ! -f .../drain-complete ]; do
sleep 0.5; done`<br>• Kubelet holds SIGTERM back. | Deployment will
recreate the pod.<br>EndpointSlice controller begins dropping Old Pod
IP. |
| **Propagation**<br>($t=0\text{s}$ – $t=13\text{s}$) | Keeps serving
normally.<br>• `ext_proc` continues unparking/routing requests. | Runs
normally inside `preStop` loop.<br>• Serves active TCP connections. |
Service endpoint removal completes.<br>No new connections arrive at Old
Pod. |
| **Envoy drain**<br>($t=13\text{s}$) | Issues Envoy admin API
calls:<br>• `/healthcheck/fail`<br>•
`/drain_listeners?graceful&skip_exit`<br>• Polls `/stats` for active
downstream connections. | Begins graceful listener drain:<br>• GOAWAY /
`Connection: close` on established connections.<br>• In-flight requests
keep running. | In-flight HTTP requests finish executing through Envoy.
|
| **ext_proc drain**<br>($t\approx18\text{s}$) | Active Envoy
connections hit 0 (the poll exits early).<br>• Calls
`extproc.GracefulStop()`.<br>• Parked request streams finish. | All
downstream connections closed.<br>• Still waiting in the `preStop` loop.
| All client HTTP responses delivered. |
| **Handshake & Exit**<br>($t=18.5\text{s}$)| `ext_proc` drain
finishes.<br>• Writes `drain-complete` marker.<br>• Hard-stops xDS
(`Stop()`) & exits. | `preStop` loop detects marker file!<br>• `preStop`
exits 0.<br>• Kubelet sends SIGTERM to Envoy.<br>• Envoy exits
immediately. | Old Pod deleted cleanly at ~18.5s (well under the 60s
budget). |

Envoy ref:
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining

#### Testing & Verification

1. Unit Tests (`drain_test.go`): Added comprehensive unit tests
covering:
- Graceful completion of in-flight `ext_proc` streams within deadline.
- Force-stopping straggler streams past `--drain-timeout`.
- Envoy admin API interaction and connection polling.
- Stale marker cleanup and marker file creation.

2. Local testing on kind cluster (Observations):
Test script: `hack/verify-atenet-drain.sh`
- Steady state: `/readyz=200`, `/healthz=200`; the instant the pod
turned Terminating: `/readyz=503` while `/healthz=200` — `NotReady` but
alive, for the whole drain.
- The parked request survived the shutdown: fired at a busy 1-worker
pool → parked → pod deleted while parked → worker freed → HTTP=200,
total=1.89s, body hello from: 169.254.17.2 | preserved memory count: 1 |
preserved file counter: 1 — park → resume → route, served by the
Terminating pod during its drain-delay window.
- Pod terminated 14s after deletion — the idle floor, confirmed across
three runs: >13s drain-delay (sequence ran), ≪60s grace (marker
handshake released Envoy's preStop; no `SIGKILL`).
- Log sequence captured verbatim, drain-delay honored to the millisecond
(17:34:41.706 → 17:34:54.707):
  >Shutdown signal received; draining
  >(+13.000s) Draining Envoy  {window: 15s}
  >Envoy drained
  >Starting ext_proc drain
  >ext_proc drain completed within deadline
  >Drain-complete marker written  {path: /var/run/atenet/drain-complete}
  >Shutdown complete


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-07 13:36:08 -07:00
Krisztian F 0dbe1523e4 feat(otel): add cold start metrics (#776)
This PR implements the last two metrics from #433.

Right now, we can see that a resume was slow but not where.
`ate.actor.lifecycle.operation.duration` covers the whole ateapi
operation, and `atenet.router.route.duration` covers the edge, but
everything between ateapi-atelet-actors is one block that contains
fetching the manifest, downloading the snapshot, unpack the OCI image,
call to ateom.

We have `rpc.server.call.duration` that gives us the atelet restore
total time, but template, kind, and scope labels are missing, so today
we cannot really pinpoint why/where we have a regressions in latency.

In this PR I am adding per-phase histograms, here's an example of what
we can know after these changes:

```console
ateom_restore   522 ms   ###############################
download        8.9 ms   #
manifest_fetch  3.2 ms
oci_unpack      2.6 ms
---------------------------------------------------------
total           535 ms
```

It also fixes a gap #683 opened where a `data_on_golden` resume was
labeled identically to a plain one on the lifecycle histogram.

Things folks might want to argue with:
- total as a phase value. Partly duplicates `rpc.server.call.duration`,
but that one has no domain labels and gRPC-specifc. We can drop it, but
it's an inferior operational UX, so I'd rather have it here.
- Phases overlap, they are not a partition of total, because download
runs concurrently with the asset fetch and unpack. I called this out in
the metric description. Do not sum across phases.
- New `ate.snapshot.scope` key rather than a new `ate.snapshot.kind`
value for `data_on_golden`. A new value would collapse local and
external into one bucket, which is the biggest latency difference there
is. This does add a label to the already shipped lifecycle histogram.

Verified on kind, and all e2e suites pass, and the emitted series cover
every kind (golden, latest, local) on both metrics with no unknown
values.
 
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-07 10:15:19 -04:00
Angela 058b104dec feat: Add control plane telemetry - scheduler eligible workers histogram (#682)
Part of #564 (Part 3)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

## Description
This PR implements Part 3 of #564 by adding telemetry histogram
instrumentation for `ate.scheduler.eligible_workers` in `ateapi`. It
measures unassigned free worker capacity remaining after all scheduling
constraint filters are applied, sampled at every scheduling decision.

## Key Changes:
* Updated `Scheduler.Schedule()` to record eligible candidate workers
per pool (`recordEligibleWorkers`).
Defined `SchedulingConstraintKey` (`ate.scheduling.constraint`) and
constraint classification values (`none`, `required_nodes`, `selector`)
in `internal/ateattr/ateattr.go`.
* Added unit tests in `scheduling_test.go` covering candidate counts,
namespaced attributes, zero-capacity fleet states, empty fleets, sandbox
class mismatches, draining workers, and constraint classifications.

## Testing
* `go test -buildvcs=false ./cmd/ateapi/internal/controlapi/...`
* `go test -buildvcs=false ./cmd/atenet/internal/router/...`
* `make test`
## E2E  Test
* `./hack/create-kind-cluster.sh`
* `./hack/install-ate-kind.sh --deploy-ate-system --deploy-demo-counter`
* `./hack/run-e2e.sh ./internal/e2e/suites/metrics/...`
2026-08-07 09:29:20 -04:00
nybidari 5429ee8df7 Update gVisor release which supports multiple durable-dirs (#787)
- Update the gVisor release which has support for multiple durable-dirs.
- Modify the tests to run with gVisor (which were disabled before).
2026-08-06 19:08:02 -07:00
Krisztian F c155efd1ac feat(otel): onboard atecontroller to the OTLP path (#754)
atecontroller had no OTel at all. Dev-mode zap logger, and its
controller-runtime metrics were only available on a :8080 that we don't
scrape.

After this PR, logs go through the shared slog handler (plus a
`--log-level` flag to match the other binaries), and
controller-runtime's Prometheus registry is bridged onto the OTLP reader
so the reconcile/workqueue metrics actually reach the collector.
Filtering these out is a pipeline responsibility. Also added otelgrpc to
the ateapi client, which was untraced.

This unblocks #564 the workperpool metrics, cc @Angelawork, @JeffLuoo:
there's a working `MeterProvider` to use for
`ate.workerpool.desired_workers`/`ready_workers`.

Couple of things to mention for review:

-`InitMetricsPushOnly`, not `InitMetrics`, even though we do serve
:8080. That port is controller-runtime's own private registry, not the
global one `serverboot.metricsMux` serves, so a pull reader there would
collect into something we never expose.
- Bridge is pinned to v0.68.0 to match otelgrpc. Wanted to go to
v0.70.0, but that requires otel/sdk/metric 1.45.0 and pulls the whole
SDK up with it (406 vendor files instead of 81). Happy to do that bump
separately.
- zap/zapr fall out of go.mod since atecontroller was the last importer.
- I included an OTel collector image bump from the early 2024 (!) one to
latest, which was breaking exposing native histograms

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-06 10:25:03 -04:00
Etienne Perot 2fe26c162b Update gVisor build to fresher version using new release tarball asset. (#684)
This new release requires multiple files and is released as a tarball.
gVisor asset handling code is updated to automatically handle this
release format and extract it as necessary.

This build also contains the start of a series of upcoming
startup/memory performance optimizations to make Substrate sandbox churn
efficient.

Benchmarks between the previous gVisor build and this one:

```
CI benchmark suite:
                                                     │      old       │                    new                    │
                                                          │     sec/op     │     sec/op       vs base                  │
  ResumeActor/glutton_baseline_1_user/p50                   324.7m ±    3%    289.6m ±    2%  -10.79% (p=0.000 n=11)
  ResumeActor/glutton_baseline_1_user/p95                   373.4m ±    9%    328.6m ±    2%  -12.00% (p=0.000 n=11)
  ResumeActor/glutton_baseline_5_users/p50                  319.4m ±    1%    288.6m ±    1%   -9.65% (p=0.000 n=9)
  ResumeActor/glutton_baseline_5_users/p95                  367.6m ±    5%    341.7m ±   11%   -7.05% (p=0.024 n=9)
  ResumeActor/glutton_baseline_10_users/p50                 323.6m ±    1%    293.8m ±    2%   -9.22% (p=0.000 n=9+10)
  ResumeActor/glutton_baseline_10_users/p95                 396.2m ±    4%    372.1m ±    5%   -6.07% (p=0.001 n=9+10)
  ResumeActor/glutton_oversubscribe_15_users/p50            330.8m ±    2%    304.7m ±    1%   -7.89% (p=0.000 n=8+10)
  ResumeActor/glutton_oversubscribe_15_users/p95            402.7m ±    5%    390.4m ±    3%        ~ (p=0.083 n=8+10)
  ResumeActorColdStart/glutton_baseline_1_user/p50          227.0m ±   10%    211.6m ±   19%   -6.79% (p=0.003 n=11)
  ResumeActorColdStart/glutton_baseline_5_users/p50         231.5m ±    5%    216.3m ±    7%   -6.54% (p=0.004 n=9)
  ResumeActorColdStart/glutton_baseline_5_users/p95         246.4m ±    3%    230.2m ±    6%   -6.60% (p=0.000 n=9)
  ResumeActorColdStart/glutton_baseline_10_users/p50        217.4m ±   10%    193.5m ±    7%  -10.99% (p=0.001 n=9+10)
  ResumeActorColdStart/glutton_baseline_10_users/p95        241.8m ±    7%    231.2m ±    9%   -4.39% (p=0.022 n=9+10)
  ResumeActorColdStart/glutton_oversubscribe_15_users/p50   205.9m ±    6%    192.5m ±   28%        ~ (p=0.122 n=8+10)
  ResumeActorColdStart/glutton_oversubscribe_15_users/p95   247.1m ±    4%    229.6m ±  102%   -7.09% (p=0.043 n=8+10)
  SuspendActor/glutton_baseline_1_user/p50                  313.3m ±    7%    294.1m ±    4%   -6.14% (p=0.010 n=11)
  SuspendActor/glutton_baseline_5_users/p50                 304.6m ±    3%    289.1m ±    2%   -5.10% (p=0.000 n=9)
  SuspendActor/glutton_baseline_10_users/p50                309.7m ±    3%    296.9m ±    3%   -4.13% (p=0.000 n=9+10)
  SuspendActor/glutton_oversubscribe_15_users/p50           320.6m ±    3%    305.3m ±    2%   -4.75% (p=0.001 n=8+10)

Manual benchmarks (GKE on a single c3-standard-88 node):

- p50 sandbox lifecycle time (from issuing resume to checkpointed):
  ┌─────────┬──────────┬───────────┬────────┐
  │ workers │   old    │    new    │ delta  │
  ├─────────┼──────────┼───────────┼────────┤
  │ 1       │ 795 ms   │ 565 ms    │ −28.9% │
  ├─────────┼──────────┼───────────┼────────┤
  │ 2       │ 801 ms   │ 590 ms    │ −26.4% │
  ├─────────┼──────────┼───────────┼────────┤
  │ 8       │ 902 ms   │ 759 ms    │ −15.9% │
  ├─────────┼──────────┼───────────┼────────┤
  │ 64      │ 3.25 s   │ 3.14 s    │ −3.4%  │
  ├─────────┼──────────┼───────────┼────────┤
  │ 88      │ 4.47 s   │ 4.30 s    │ −3.8%  │
  └─────────┴──────────┴───────────┴────────┘

- Sandbox starts per second:
  ┌─────────┬────────────┬────────────┬────────┐
  │ workers │    old     │    new     │ delta  │
  ├─────────┼────────────┼────────────┼────────┤
  │ 1       │ 1.256 ± 1% │ 1.755 ± 2% │ +39.7% │
  ├─────────┼────────────┼────────────┼────────┤
  │ 2       │ 2.482 ± 1% │ 3.393 ± 1% │ +36.7% │
  ├─────────┼────────────┼────────────┼────────┤
  │ 8       │ 8.81 ± 2%  │ 10.50 ± 2% │ +19.2% │
  ├─────────┼────────────┼────────────┼────────┤
  │ 64      │ 19.85 ± 3% │ 20.45 ± 1% │ +3.0%  │
  ├─────────┼────────────┼────────────┼────────┤
  │ 88      │ 19.70 ± 1% │ 20.58 ± 2% │ +4.5%  │
  └─────────┴────────────┴────────────┴────────┘
```
2026-08-05 16:10:05 -07:00
Angela 1f8f842a64 feat: Add control plane telemetry - actor crashes counter and operation labels (#642)
Fixes #564 (Part 2)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

## Description
This PR implements the second part of #564 by adding telemetry counter
instrumentation for `ate.actor.crashes` in `ateapi` and establishing
bounded label validation for operations and crash failure causes.

## Key Changes:
- Updated `maybeCrashActor` and `crashActor` to reference
`ateattr.OperationName*` constants and normalize `opName`.
- Passed explicit failure reasons (`ReasonCorruptedAssignment`,
`ReasonWorkerPodGone`, `ReasonWorkerReassigned`) across
`workflow_resume.go`, `workflow_suspend.go`, and `workflow_pause.go`.
- Instrumented `WorkerPoolSyncer.releaseActorOnDeadWorker` in
`syncer.go` to record `recordActorCrash` when background pod deletion
crashes an actor.
- Instrumented `FinalizePausedStep` in `workflow_pause.go` to record
`recordActorCrash` when `nodeName` is missing during pause finalization.
- Added `ate_actor_crashes` to `PlatformMetricPrefixes` in
`collector_metrics.go`.
- Added pre-creation resource cleanup in `metrics_test.go` to prevent
`AlreadyExists` errors on test reruns.
- Updated `metrics_test.go` to trigger an actor crash via `UpdateActor`
and `ResumeActor`, asserting `ate_actor_crashes` emission and label
presence in OTel collector scrape outputs.

## Testing
* `go test -buildvcs=false ./cmd/ateapi/internal/controlapi/...`
* `go test -buildvcs=false ./cmd/atenet/internal/router/...`
* `make test`
## E2E  Test
* `./hack/create-kind-cluster.sh`
* `./hack/install-ate-kind.sh --deploy-ate-system --deploy-demo-counter`
* `./hack/run-e2e.sh ./internal/e2e/suites/metrics/...`
2026-08-05 16:31:43 -04:00
Da Huang 7a9bb4b1f2 manifests: centralize the OTLP endpoint in an ate-otel-config ConfigMap (#746)
OTEL_EXPORTER_OTLP_ENDPOINT is written in 9 places: hardcoded inline in
4 base workload manifests, then re-patched in 5 spots across the kind
overlays. Changing the collector address means editing all 9 and knowing
which install path renders which.

Replace them with one checked-in ConfigMap per environment, consumed by
every component via envFrom:

  manifests/ate-install/ate-otel-config.yaml       (GKE)
  manifests/ate-install/kind/ate-otel-config.yaml  (kind, same name)

The GKE copy is listed in manifests/ate-install/base, so token-client,
agentgateway and agentgateway-token-client all inherit it; the kind
overlay lists its own copy of the same ConfigMap name. No overlay builds
on both, so the two never collide.

A kustomize-only fix does not work here. The GKE path applies the base
directory raw and hack/install-ate.sh's targeted redeploys apply single
files with no Kustomize, so the mechanism has to survive `kubectl apply
-f <one-file>` -- which rules out configMapGenerator (hash-suffixed
names) and replacements. envFrom on a stable name reaches every path.

deploy_ate_system applies the ConfigMap before the rendered bundle, the
same way it already applies the namespace. A container whose envFrom
target is missing does not start, and a raw directory apply is ordered
by filename, so ate-api-server.yaml and ate-controller.yaml would
otherwise be created first and sit in CreateContainerConfigError until
the ConfigMap caught up.

This also fixes a latent bug in the targeted redeploys.
deploy_ate_apiserver, deploy_atelet and deploy_atenet apply raw files
even in kind mode, silently reverting the endpoint to the GKE value.
They now apply the environment's ConfigMap via a new apply_otel_config
helper, which picks the file by ATE_INSTALL_KIND rather than applying
the base copy unconditionally -- applying the base one on kind would
break telemetry for every component at once.

atenet-router needs no special handling: since a98f85b1 its
--otlp-collector-address defaults to $OTEL_EXPORTER_OTLP_ENDPOINT, so
Envoy's own spans follow the ConfigMap along with everything else.

OTEL_TRACES_SAMPLER deliberately stays an inline kind patch on
ate-api-server rather than moving into the ConfigMap. The ConfigMap is
shared by every component via envFrom, so putting parentbased_always_on
there would pin the router and the rest of the control plane to 100% and
undo the per-component ratios from 15eecd02.

Note that a ConfigMap edit does not roll the consuming pods the way an
inline env change did, since the pod template is unchanged. Callers must
follow a change with `kubectl rollout restart`. This is documented in
both ConfigMaps, in docs/observability.md, and in the tracing best
practices, which previously told authors to hardcode the variable.

Fixes #745

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-05 14:03:41 -04:00
Julian Gutierrez Oschmann 4f2fd4912f Wrap worker pod metadata into a WorkerAssignment message. (#737)
This change formalizes the invariant that all these fields are either
all set or not. This simplifies some client code that was just checking
random worker metadata fields to see if the actor has an assignment or
not.

It also fixes a bug in `{Pause,Suspend}Actor` where we were not cleaning
up the worker pod UID field.
2026-08-04 17:08:58 -07:00
Krisztian F 6f2b2178c1 feat(otel): add actor lifecycle + scheduler duration metrics (#514)
This PR is the second slice of the platform-metrics split (#433).

It adds two duration histograms emitted by ateapi:
- `ate.actor.lifecycle.operation.duration`:
create/resume/suspend/pause/delete, labeled by operation, template,
pool, sandbox class, and (on resume) snapshot kind
- I'd like to use this for user-facing latency for e.g. showing when
suspended actors can serve requests again. The existing
`rpc.server.call.duration` metric covers this, but the meaningful
dimensions are missing, so it's not really actionable. Extending its
labels with extra, domain-specific labels is an OTel anti-pattern, hence
the new metric.
- `ate.scheduler.assignment.duration`: worker-assignment step, labeled
by outcome (assigned / no_free_worker / error) and pool
- I'd like to use this to alert on no free worker situation, and have
proper SLOs via various percentiles for assigning latencies. There's no
RPC around this, so it's not something existing RPC metrics cover.

Tested e2e on a local kind cluste where: both metrics reach the
otel-system collector with the expected labels (resume shows
`snapshot_kind=golden`, scheduler shows `outcome=assigned`, no
`error.type` on success).

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-04 13:57:55 -04:00
Krisztian F 15eecd02c9 Define a production-ready default tracing policy (#711)
Today everything outside ateapi hardcodes `ParentBased(NeverSample)`,
and setting `OTEL_TRACES_SAMPLER` does nothing because an explicit
sampler silences the SDK's env handling. This makes troubleshooting
production issues impossible via traces.

This PR gives every component a sane default and makes the standard OTel
env vars work:

- Control plane (ateapi, atelet, ateom) defaults to
`parentbased_traceidratio` 0.1, the router (data plane root) to 0.01,
per the discussion on #584.
- `OTEL_TRACES_SAMPLER` / `OTEL_TRACES_SAMPLER_ARG` override any of this
without a rebuild. Invalid values keep the component default and log a
warning instead of inheriting the SDK's fall-open-to-100% behavior.
- Envoy's `RandomSampling` is derived from the router's resolved policy,
so the two root decisions cannot drift.
- kubectl-ate without `--trace` no longer installs a tracer provider at
all: the old `NeverSample` provider injected `sampled=0`, which pinned
every parent based sampler downstream and would have defeated the server
side ratios. `--trace` still forces a full end to end trace.
- ate-controller propagates the two env vars to the ateom worker pods it
creates, same as the metric export vars.
- kind pins ateapi to `parentbased_always_on`, so the local Jaeger flow
keeps showing every API call.
- The agentgateway integration follows what we have above. its
`randomSampling` changes from `true` to `0.01` to match the data plane
default, though as static config it does not follow
`OTEL_TRACES_SAMPLER` overrides, so the two need adjusting together.


Verified on a kind cluster e2e manually. It resolved samplers logged at
startup, a `--trace` resume produced one trace across ateapi, atelet,
and ateom, an unsampled CLI call got picked up server side, Envoy
continued a sampled traceparent while sampling 0 of 30 parentless
requests at the 1% default, and an invalid env value fell back to the
component default.

Gating who may use `--trace` stays a separate follow-up (and a
discussion), and the more sophisticaed tracing policies belongs in a
collector, not in substrate.

Fixes #584
cc. @git286 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-04 11:31:04 -04:00
John Howard 9e74b1d729 atenet: add option to run agentgateway
Agentgateway is a popular [AAIF](https://aaif.io/) project designed for
agentic systems. It is of particular interest for substrate users as it
was designed specifically to serve the role of an egress proxy for
agents and other AI-adjacent workloads, with strong support for
credential injection, token exchanges, TLS MITM, and other
authentication and authorization schemes.

Agentgateway supports the ext_proc protocol, like Envoy, so we use that
here. There has been some debate in the community around the long term
end state of ext_proc vs other options, but for now this maintains the
status quo.

While Agentgateway can be configured dynamically over XDS, it does not
take the same configuration as Envoy. However, none of the config is
dynamic anyways, so we use just a static configuration file (configmap)
for now to keep things simple.

While we don't have specific per-component/file OWNERS in the project at
the moment, I can informally commit to myself and Eitan maintaining this
integration. Given the goals around velocity in the project we can
commit to either rapidly fixing any issues that may arise or
removing/disabling the integration if this ends up slowing things down.
2026-08-03 14:23:16 -07:00
Dmitry Berkovich 3ed6aa07e1 microvm: API support for resume Data snapshots on the golden snapshot (#451, phase 1) (#683)
Phase 1 of #451, micro-VM only. gVisor support will follow in a separate
PR.

## What

Adds a `snapshotsConfig.onResume` block to the ActorTemplate, selecting
per snapshot situation what supplies the guest state at resume: each
field names what is being resumed *from*, the value names the boot
source. `onPause`/`onCommit` remain pure capture scopes (`Full | Data`),
and the `ActorSnapshot` records plain `Full`/`Data` content; the
golden-combine is strictly a restore-time behavior.

```yaml
snapshotsConfig:
  onPause: Data
  onCommit: Data
  onResume:
    fromData: Golden   # ColdBoot (default) | Golden — Data snapshots resume as
  location: gs://…     # golden memory + fs delta with the actor's data layered on
```

Under `fromData: Golden`, a data-only restore combines the template's
**golden snapshot** (guest memory + full fs delta) with the actor's
captured durable data — the actor comes back with the golden's warm
state over its own data instead of cold-booting. `fromData` applies to
every resume of a Data snapshot — commit or pause checkpoint. Future
situations (e.g. resuming a Full snapshot across a template upgrade,
#477) get sibling fields under `onResume` rather than overloading this
one.

## How

- **CRD** (`pkg/api/v1alpha1`): new `OnResumeConfig`/`ResumeSource`
types; `fromData: Golden` is CEL-gated to `sandboxClass: microvm`.
- **Control plane** (`ateapi`): resume derives the behavior from the
template's `onResume` configuration — for a Data durable snapshot or a
Data pause checkpoint it resolves the golden snapshot's location and
sends the restore-only `DATA_ON_GOLDEN` wire scope, failing early (with
an actionable error) if the golden snapshot is missing or not Full.
Golden actors always commit `Full` regardless of `onCommit` — their
snapshot is the base the combine needs.
- **Wire APIs** (`ateletpb`/`ateompb`): restore-only
`SNAPSHOT_SCOPE_DATA_ON_GOLDEN`, rejected on checkpoints;
`golden_snapshot_uri_prefix` is a top-level `RestoreRequest` field
because the actor's snapshot may be a local pause checkpoint while the
golden snapshot is always external.
- **atelet**: stages a single combined folder — the actor's files (the
durable-dir tar) win name collisions, the golden snapshot supplies the
rest. External restores download both halves concurrently; local (pause)
restores copy the actor's files from the local checkpoint dir while the
golden's files download in parallel. The golden manifest's pinned
sandbox binaries run the restored guest and are recorded on-node for
later checkpoints.
- **ateom-microvm**: `DATA_ON_GOLDEN` restores through the existing Full
path — cloud-hypervisor relaunches from the golden's guest files while
the durable virtio-fs share serves the actor's re-materialized data.
gVisor rejects the scope (defense-in-depth behind the CRD gate).

## Testing

- Unit: envtest CEL cases, converter tables, atelet request validation
(checkpoint rejection, golden-URI rules incl. local+golden),
combined-download test against a fake object store, resume-workflow
golden-resolution cases for both the durable and pause paths.
- e2e (`suites/demo`): OnGolden lifecycle cases for the commit path, the
pause path, and a two-durable-volume variant (micro-VM lane only),
asserting counters across pause/suspend and the recorded Data content
scope.
- Manually verified on a live GKE cluster: Data commit and Data pause
both resume with the golden guest's memory over the actor's data, on
different pods.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-01 20:13:37 -07:00
Lior Lieberman 0909650285 ActorIdentity: protos, cert issuance, actorIdentity x509 extension (#670)
Also Fixes #575 

(1) mintCert + MintJWT proto changes
(2) validate that atelet is the one that calls actorIdentity Service. 
(3) verify that this atelet is only requesting a cert for an actor on
its node.
(4) verify that the actor is still running
(5) create another substatex509 extension for actor identities

JWTs needs more love, lots as TODO there

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-01 19:23:53 -07:00
Lior Lieberman 860250bf71 diagram 2026-07-31 12:51:48 -07:00
Lior Lieberman d9ac4e24e6 networking: update docs and add an ingress e2e suite
Bring the architecture and threat-model docs in line with the atunnel
ingress path: the router no longer rewrites :authority to a worker pod IP
and forwards over plaintext port 80, it opens an mTLS tunnel to atunnel on
worker port 443, which forwards to the Actor over its private veth.

Drop cmd/atenet/atenet-diagram.png. It predates the ext_proc/ORIGINAL_DST
design and is now wrong in the part that matters most. The README section it
illustrated gains a short accurate note about the upstream hop instead of a
dangling image reference.

Add an e2e suite covering the change end to end: TestActorDirectAccess
asserts that the worker pod's port 80 is no longer a reachable Actor ingress
path (the DNAT rule is gone) and that the same Actor still answers /readyz
through atenet-router over the atunnel mTLS hop. It uses the counter demo as
its fixture, so it only needs --deploy-demo-counter.

internal/e2e/testmain.go picks up ParseSkippedFlags so that `go test` flags
(-run, -v, ...) survive pflag parsing and reach the suite.
2026-07-31 12:51:48 -07:00
Jeff Luo b255df38c0 docs: add an OTel Collector setup best-practices guide (#663)
Substrate documents how to instrument a service but not how to stand up
the collector it exports to, so pointing it at your own collector meant
reading six manifests.

Covers both topologies: GKE Managed OpenTelemetry as the default on GKE,
with why its gateway Deployment suits Substrate's elastic actor
workload, and a DaemonSet manifest for operators who need per-node
isolation, full config control, or a non-GKE cluster.

Part of #563.

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-31 15:35:42 -04:00
Alex Zakonov 6f44915bdf docs: fixed language describing project goals and relationship to Kubernetes (#524)
Fixed language describing project goals and relationship to Kubernetes
in readme.md, architecture.md and roadmap.md
- changed the top language to focus on project goals instead of
Kubernetes relations
- changed the language describing Kubernetes relationship to describe
value of the Agent Substrate layer and Kubernetes layer
2026-07-31 12:35:04 -07:00
Angela e70e1f21f8 feat: extend edge routing telemetry with resume and outcome labels (#591)
Fixes #564 (Part 1)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

## Description
This PR implements the first part of #564 by extending metric
`atenet.router.route.duration` with singleflight activation state
(`ate.router.resume`) and outcome status labels (`ate.router.outcome`).
All telemetry attribute keys and label value sets are centralized in
`internal/ateattr` to ensure consistency across components.

## Key Changes:
* Added `resumed` boolean to `ResumeActorResponse` in `ateapi.proto`
indicating whether a cold activation workflow occurred
(`!state.WasRunning`).
* Updated `workflow_resume_test.go` to assert `resumed == true` on cold
activations and `resumed == false` for already-running actors.
* Extended `classifyOutcome(err)` in Envoy `ext_proc` to classify errors
into metric labels: `ok`, `cancelled`, `timeout`, `no_capacity`,
`lock_conflict`, `not_found`, `unavailable`, `rate_limited`, and
`resume_error`.
* Added explicit handling for `codes.Unavailable` (`"unavailable"`) and
`codes.ResourceExhausted` / `StatusCode_TooManyRequests`
(`"rate_limited"`).
* Consolidated metric label names on `atenet.router.route.duration`
using standard keys from `internal/ateattr`:
     * `ate.template.namespace` (`TemplateNamespaceKey`)
     * `ate.template.name` (`TemplateNameKey`)
     * `ate.router.outcome` (`RouterOutcomeKey`)
     * `ate.router.resume` (`RouterResumeKey`)
* Updated `manifests/ate-install/kind/kustomization.yaml` to export
`atenet-router` OTLP metrics to `opentelemetry-collector`.

## Testing
* `go test -buildvcs=false ./cmd/ateapi/internal/controlapi/...`
* `go test -buildvcs=false ./cmd/atenet/internal/router/...`
* `make test`
## E2E  Test
* `./hack/create-kind-cluster.sh`
* `./hack/install-ate-kind.sh --deploy-ate-system --deploy-demo-counter`
* `./hack/run-e2e.sh ./internal/e2e/suites/metrics/...`
2026-07-31 13:41:56 -04:00
Benjamin Elder 97772a02f3 Support multiple durable-dir volumes on the micro-VM runtime
An ActorTemplate could declare only one durable-dir volume, and mount it
into a container only once. That is a gVisor limit, not a general one:
atelet declares the mount to gVisor through a single hardcoded annotation
key ("dev.gvisor.spec.mount.durabledir"), so a second volume would silently
overwrite the first.

The micro-VM runtime has no such constraint — every volume is a
subdirectory of the one writable virtio-fs share, and the snapshot tar
already archives that directory whole, so N volumes cost a subdirectory
each and round-trip through checkpoint/restore untouched. Gate the two
"at most one" CEL rules on sandboxClass so micro-VM templates may declare
several while gVisor keeps its cap, and say so in the messages.

What was actually missing was the volume NAME: ateom received mount paths
alone, so it inferred the single name by listing atelet's directory. Carry
each mount's name on the wire (Container.durable_dir_volume_mounts,
replacing the paths-only field, which is reserved rather than retyped since
atelet and ateom are separate images that can skew across a rollout), and
delete the inference. A container's binds now come from its own mounts, and
the actor-wide "has a durable share" question collapses to a bool.

The counter demo grows an optional --second-file-counter-directory, and a
micro-VM-only e2e case runs the lifecycle matrix against an Actor with two
durable volumes, asserting both counters advance together. gVisor skips it:
the template would be rejected at admission.
2026-07-30 20:03:55 -07:00
Eitan Yarmush ef7b29da44 ateapi: add ActorSnapshot lifecycle APIs 2026-07-30 19:40:59 -07:00
Mesut Oezdil 3240b6295f docs: list the missing demos, guides, and core terms (#631)
The README lists 3 of the 6 demos and 5 of the 9 docs, and its command
tour skips `cmd/ateom-microvm` and `cmd/benchmarking`. The glossary
never defines Atespace, even though it is half of an actor's identity
and has its own API, and it omits SandboxConfig.
2026-07-30 12:30:09 -04:00
Lior Lieberman 92f1aa7276 Rename SessionIdentity to ActorIdentity
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.

API surface:
  service SessionIdentity        -> ActorIdentity
  MintJWTRequest.session_id      -> actor_id
  MintJWTResponse.session_jwt    -> actor_jwt
  MintCertRequest.session_id     -> actor_id
  MintCertResponse.session_certificates -> actor_certificates

Go packages:
  cmd/ateapi/internal/sessionidentity -> actoridentity
  cmd/ateapi/internal/sessionidjwt    -> actoridjwt

Flags and cluster resources:
  --session-id-jwt-pool -> --actor-id-jwt-pool
  --session-id-ca-pool  -> --actor-id-ca-pool
  Secrets, volumes and mount paths renamed to match, in both
  manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
  (--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).

Two credential identity values change with the rename:
  JWT issuer https://broker.agentic-substrate-session-id-broker.svc
          -> https://broker.agentic-substrate-actor-id-broker.svc
  SPIFFE ID spiffe://substrate-session.local/app/../session/..
          -> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.

BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
2026-07-30 07:34:29 -07:00
Omer Yahud 6b1ed81e17 atenet/router: derive the ext_proc circuit breaker from the lot by default
Review feedback (bowei): rather than a fully independent flag, --extproc-max-requests now defaults to 0 = derive twice --parked-request-max, floored at Envoy's own default of 1024 — the lot always fits and keeps an equal share of fast-path headroom at any size, including a small or disabled lot. An absolute-percentage derivation like 120% under-provisions the fast path at small lots, which is why the headroom equals the lot instead. Explicit values still override and keep the >= lot validation, so operators who need a specific breaker retain control.
2026-07-28 23:55:51 -07:00
Omer Yahud 727181d026 atenet/router: make the ext_proc circuit breaker an explicit, validated flag
Add --extproc-max-requests (default 2048) and set circuit_breakers.max_requests on the ext_proc cluster from it, replacing the implicit Envoy default and the hand-maintained doc coupling with a guarantee. Every request's header exchange occupies one slot briefly and every parked request holds one for its entire wait, so startup validation enforces extproc-max-requests >= parked-request-max — a breaker below the lot silently truncates it with Envoy-generated 503s that bypass parking.rejected. The default leaves the lot's worth of fast-path headroom (1024 lot / 2048 breaker), so a saturated lot cannot starve requests to already-running actors.
2026-07-28 23:55:51 -07:00
Omer Yahud e794bfed0a atenet/router: size the default lot to Envoy's ext_proc circuit breaker
The ext_proc cluster sets no explicit circuit_breakers, so Envoy's default max_requests=1024 applies — and every parked request holds one ext_proc stream, i.e. one active request against that cluster. A 2048 lot was therefore half unreachable: requests 1025+ would be rejected by Envoy itself, with 503s that never reach the lot and never count in parking.rejected. Set the default to 1024 to match, and document the coupling at the constant, at buildCluster, and in the design doc, including what raising the flag beyond 1024 requires (an explicit circuit_breakers.max_requests on the cluster). Also drop Unavailable from the docs' non-retryable examples — it became retryable-while-parked in the previous commit.
2026-07-28 23:55:51 -07:00
Omer Yahud 4341b248d7 atenet/router: document the park budget as per-flight
The budget clock starts with a flight's first caller; requests de-duplicated onto an in-flight resume share its remaining budget and outcome, so a late joiner can see budget_exhausted after waiting far less than a full budget itself. That trade is inherent to collapsing a hot actor's requests into one control-plane RPC — state it explicitly in the design doc, the flight comment, and the flag help instead of implying a per-request guarantee. Also notes that wait-duration samples record each request's own parked time, so sub-budget budget_exhausted samples are expected under sustained saturation.
2026-07-28 23:55:51 -07:00