#732
* `UpdateActorRequest` now carries the resource itself + `update_mask`
* `worker_selector` is now applied via the update mask (making it
possible to clear the worker selector)
* `update_mask` is validated against an allowlist of paths. Currently,
only `worker_selector` is allowed.
* Added `uid` and `version` as optional guards.
I'll update UpdateActorSnapshotTag in a follow-up PR
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
OTEL_EXPORTER_OTLP_ENDPOINT is written in 9 places: hardcoded inline in
4 base workload manifests, then re-patched in 5 spots across the kind
overlays. Changing the collector address means editing all 9 and knowing
which install path renders which.
Replace them with one checked-in ConfigMap per environment, consumed by
every component via envFrom:
manifests/ate-install/ate-otel-config.yaml (GKE)
manifests/ate-install/kind/ate-otel-config.yaml (kind, same name)
The GKE copy is listed in manifests/ate-install/base, so token-client,
agentgateway and agentgateway-token-client all inherit it; the kind
overlay lists its own copy of the same ConfigMap name. No overlay builds
on both, so the two never collide.
A kustomize-only fix does not work here. The GKE path applies the base
directory raw and hack/install-ate.sh's targeted redeploys apply single
files with no Kustomize, so the mechanism has to survive `kubectl apply
-f <one-file>` -- which rules out configMapGenerator (hash-suffixed
names) and replacements. envFrom on a stable name reaches every path.
deploy_ate_system applies the ConfigMap before the rendered bundle, the
same way it already applies the namespace. A container whose envFrom
target is missing does not start, and a raw directory apply is ordered
by filename, so ate-api-server.yaml and ate-controller.yaml would
otherwise be created first and sit in CreateContainerConfigError until
the ConfigMap caught up.
This also fixes a latent bug in the targeted redeploys.
deploy_ate_apiserver, deploy_atelet and deploy_atenet apply raw files
even in kind mode, silently reverting the endpoint to the GKE value.
They now apply the environment's ConfigMap via a new apply_otel_config
helper, which picks the file by ATE_INSTALL_KIND rather than applying
the base copy unconditionally -- applying the base one on kind would
break telemetry for every component at once.
atenet-router needs no special handling: since a98f85b1 its
--otlp-collector-address defaults to $OTEL_EXPORTER_OTLP_ENDPOINT, so
Envoy's own spans follow the ConfigMap along with everything else.
OTEL_TRACES_SAMPLER deliberately stays an inline kind patch on
ate-api-server rather than moving into the ConfigMap. The ConfigMap is
shared by every component via envFrom, so putting parentbased_always_on
there would pin the router and the rest of the control plane to 100% and
undo the per-component ratios from 15eecd02.
Note that a ConfigMap edit does not roll the consuming pods the way an
inline env change did, since the pod template is unchanged. Callers must
follow a change with `kubectl rollout restart`. This is documented in
both ConfigMaps, in docs/observability.md, and in the tracing best
practices, which previously told authors to hardcode the variable.
Fixes#745
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
There's still some debate on whether this is the correct API for
configuring termination grace period. It's possible we only want to
allow users to configure the grace period per actor, instead of per
worker. We want to be careful about exposing the wrong API surface too
early, as it will need to be supported for a long time (potentially
forever).
For now let's hardcode a large value, and we can revisit this decision
later.
Refactors `WorkerPoolSyncer` to process events through a rate-limiting
workqueue instead of processing them synchronously inline.
Previously, store errors during Informer event handling, such as
- optimistic locking version conflicts (`store.ErrVersionConflict`)
- create races (`store.ErrAlreadyExists`)
- `WorkerPool` not yet present in the lister cache
were logged and dropped, leaving store state out of sync until a
subsequent pod event occurred.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
TestActorDirectAccess/via_ingress was flaky because the test probed the
ingress immediately after ResumeActor returned, before the Envoy xDS
update from the control plane had propagated — causing intermittent 503
connection timeouts.
Retry the GET with a 30 s deadline (1 s intervals), the same pattern
used by waitForActorStatus and callActor elsewhere in the e2e suite.
Fixes#723.
This change formalizes the invariant that all these fields are either
all set or not. This simplifies some client code that was just checking
random worker metadata fields to see if the actor has an assignment or
not.
It also fixes a bug in `{Pause,Suspend}Actor` where we were not cleaning
up the worker pod UID field.
This PR is the second slice of the platform-metrics split (#433).
It adds two duration histograms emitted by ateapi:
- `ate.actor.lifecycle.operation.duration`:
create/resume/suspend/pause/delete, labeled by operation, template,
pool, sandbox class, and (on resume) snapshot kind
- I'd like to use this for user-facing latency for e.g. showing when
suspended actors can serve requests again. The existing
`rpc.server.call.duration` metric covers this, but the meaningful
dimensions are missing, so it's not really actionable. Extending its
labels with extra, domain-specific labels is an OTel anti-pattern, hence
the new metric.
- `ate.scheduler.assignment.duration`: worker-assignment step, labeled
by outcome (assigned / no_free_worker / error) and pool
- I'd like to use this to alert on no free worker situation, and have
proper SLOs via various percentiles for assigning latencies. There's no
RPC around this, so it's not something existing RPC metrics cover.
Tested e2e on a local kind cluste where: both metrics reach the
otel-system collector with the expected labels (resume shows
`snapshot_kind=golden`, scheduler shows `outcome=assigned`, no
`error.type` on success).
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Today everything outside ateapi hardcodes `ParentBased(NeverSample)`,
and setting `OTEL_TRACES_SAMPLER` does nothing because an explicit
sampler silences the SDK's env handling. This makes troubleshooting
production issues impossible via traces.
This PR gives every component a sane default and makes the standard OTel
env vars work:
- Control plane (ateapi, atelet, ateom) defaults to
`parentbased_traceidratio` 0.1, the router (data plane root) to 0.01,
per the discussion on #584.
- `OTEL_TRACES_SAMPLER` / `OTEL_TRACES_SAMPLER_ARG` override any of this
without a rebuild. Invalid values keep the component default and log a
warning instead of inheriting the SDK's fall-open-to-100% behavior.
- Envoy's `RandomSampling` is derived from the router's resolved policy,
so the two root decisions cannot drift.
- kubectl-ate without `--trace` no longer installs a tracer provider at
all: the old `NeverSample` provider injected `sampled=0`, which pinned
every parent based sampler downstream and would have defeated the server
side ratios. `--trace` still forces a full end to end trace.
- ate-controller propagates the two env vars to the ateom worker pods it
creates, same as the metric export vars.
- kind pins ateapi to `parentbased_always_on`, so the local Jaeger flow
keeps showing every API call.
- The agentgateway integration follows what we have above. its
`randomSampling` changes from `true` to `0.01` to match the data plane
default, though as static config it does not follow
`OTEL_TRACES_SAMPLER` overrides, so the two need adjusting together.
Verified on a kind cluster e2e manually. It resolved samplers logged at
startup, a `--trace` resume produced one trace across ateapi, atelet,
and ateom, an unsampled CLI call got picked up server side, Envoy
continued a sampled traceparent while sampling 0 of 30 parentless
requests at the 1% default, and an invalid env value fell back to the
component default.
Gating who may use `--trace` stays a separate follow-up (and a
discussion), and the more sophisticaed tracing policies belongs in a
collector, not in substrate.
Fixes#584
cc. @git286
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fourth slice of #463 Phase 2, and the semantic pivot of the GC design:
the pull
now writes the image record **before** unpacking, touches it as each
layer
completes, rewrites it after the loop, and re-verifies every returned
layer dir.
Why record-first: the record's diffID list (known from the config before
any
download) is what the eviction engine's refcounts will trust — writing
it first
makes every layer referenced before it can exist, so pulls need no
separate
protection mechanism, and the protection survives atelet restarts. The
per-layer touch makes liveness progress-based: a wedged pull ages out
through
ordinary LRU; the end-of-pull rewrite means success always implies a
record;
the re-verify turns the residual race into a clean RPC retry instead of
a
bundle spec naming a missing lowerdir.
Also new: the config's diffID list is cross-checked against the
manifest's
layer count up front and against each layer's actual diffID as it lands,
so a
record can never reference something different from what is on disk.
**Note for @dberkov**: this inverts Phase 1's record-last ordering,
which
appears incidental (amendment 4 in the Phase 2 comment on #463 — please
contradict if there was intent). A record now explicitly means "known
image,
possibly partially present", which is what every reader already handles:
`cachedImage` verifies each layer and re-pulls only gaps. The one
behavior
change observable without GC: an interrupted pull leaves a valid partial
record — resumable progress — instead of nothing.
**Testing**: a gated registry holds real pulls mid-flight by blocking
chosen
blob downloads — interrupted pull leaves a resumable record and a retry
completes it; a record deleted mid-pull is rewritten on success;
per-layer
completions freshen the record while later layers are still gated; a
yanked
layer dir fails the pull with a retryable error. `go test -race` green.
> **Rescoped again.** This PR previously proposed
`--golden-snapshot-warmup`, a
> tunable wall-clock delay before the golden checkpoint. Per discussion,
that
> direction is dropped: the answer for a workload that cannot report
readiness
> is a readiness endpoint — a small sidecar where the workload itself
cannot be
> changed — not a longer timer. What survives is the piece that
discussion
> agreed on, and which the previous revision already flagged as a
follow-up:
> making the readyz deadline itself configurable.
>
> The warmup work is not in this branch. It is kept locally in case a
workload
> genuinely cannot be given a readiness signal before GA, and would come
back as
> its own PR if so.
## Summary
`readyz.Wait` polls until the container returns 200 or a **hardcoded
30s**
elapses. A workload that legitimately takes longer to bind its HTTP
server
cannot be accommodated without raising the ceiling for every actor in
the
cluster, and losing that race fails the actor start.
How long a workload takes to become ready is a property of that
workload, so
this makes the deadline a per-template setting rather than a package
constant.
Adds optional `timeoutSeconds` to `ContainerReadyz`. **Unset keeps
today's 30s**,
so no existing template changes behavior.
## Changes
The value rides on the existing probe, so it follows the chain the probe
already
takes and no call site needs to know about it:
`ContainerReadyz.timeoutSeconds` → `toAteletReadyz` → `ateletpb.Readyz`
→
`toAteomReadyz` → `ateompb.Readyz` → `readyz.Wait`
- `pkg/api/v1alpha1/actortemplate_types.go` — `TimeoutSeconds *int32`,
`+optional`, `Minimum=1`, `Maximum=3600`.
- `internal/proto/ateletpb/atelet.proto`,
`internal/proto/ateompb/ateom.proto` —
`int32 timeout_seconds = 2` on both `Readyz` messages.
- `cmd/ateapi/internal/controlapi/workload_spec.go`,
`cmd/atelet/main.go` — pass
it through the two conversions.
- `internal/readyz/readyz.go` — `OverallTimeout` becomes
`DefaultOverallTimeout` (still 30s) and `Wait` resolves its deadline
through a
new `overallTimeout(probe)` helper.
- Regenerated: both `.pb.go`, `zz_generated.deepcopy.go`, and the
`actortemplates` CRD.
None of the four `readyz.WaitAll` call sites change.
**On the zero value.** Unlike a warmup delay — where zero is a real
request
meaning "checkpoint immediately" — a zero readiness deadline could never
be met,
so it is never something a template author means. A non-positive value
on the
wire is therefore read as "unset" and falls back to the default, and the
CRD
field is a pointer with `Minimum=1` so the API rejects `0` outright
rather than
silently substituting 30s behind the author's back.
**On bounding**, which was the open question left on the previous
revision:
bounded at `3600`. A template asking to wait longer than an hour for
readiness
is expressing a broken workload, not a slow one, and the bound keeps a
typo from
pinning a worker for a day.
## Verification
- `go build ./...`, `go vet ./...`, `gofmt`, `go test ./...` — all pass.
- `internal/readyz/readyz_test.go` — `overallTimeout` resolves unset and
negative to the default and honors an explicit value; `Wait` against a
port
nothing binds gives up at the probe's 1s deadline rather than the 30s
default.
- `workload_spec_test.go`, `cmd/atelet/main_test.go` — the timeout
crosses both
conversions, and a probe without one stays zero on the wire.
- `actortemplate_validation_test.go` — the bounds are enforced by a real
API
server. This suite runs under envtest against the generated CRD
directory, so
it exercises the regenerated `actortemplates` CRD rather than the Go
markers:
`300` is accepted, unset is accepted, and `0`, `-1` and `3601` are all
rejected by apiserver schema validation.
- **On a real cluster, via CI.** `internal/e2e/fixtures/probe` now
declares a
`readyz` probe with `timeoutSeconds: 60`, pointed at the `/healthz` the
probe
binary already serves on `:80`. The kind e2e that runs on every PR
therefore
exercises the value crossing ateapi → atelet → ateom on real binaries,
across
the auth matrix, on both the run and restore paths. This is also the
readyz
path's first e2e coverage — no fixture declared a probe before.
Wire compatibility degrades safely in both skew directions:
`timeout_seconds` is
a new field 2 on a `Readyz` message that has only ever had field 1, so
an old
ateom ignores it and an old ateapi leaves it zero, which reads as the
30s
default.
No GKE run. What that would add over the above is a workload whose
readiness
genuinely exceeds 30s, and that is the readiness-sidecar work rather
than this
PR.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Maya Wang <mymaya@google.com>
Fixes#697
## Summary
`deploy_ate_system` gated `manifests/ate-install/generated/` behind
`ensure_crds`, whose existence check skips the directory on any upgrade
— stranding CRD schemas and, since `role.yaml` has no other apply path,
all ClusterRoles at first-install state. A controller image needing a
new permission then deadlocks on informer start while the rollout
reports success.
`deploy_ate_system` now calls `deploy_crds` unconditionally (`kubectl
apply` is idempotent). The demo scripts and per-component deploy flags
keep `ensure_crds`, where "make sure they exist" is the intended
semantic.
## Test plan
- Reproduced on a 4-day-old kind cluster: upgrading the control plane
deadlocked ate-controller on `cannot list networkpolicies`; manually
applying `role.yaml` unblocked it.
- With this fix, the same `install-ate-kind.sh --deploy-ate-system` run
prints the `deploy_crds` step and re-applied the drifted manifests
(`actortemplates.ate.dev configured`, `workerpools.ate.dev configured`,
ClusterRoles applied); cluster reconciles normally.
- `bash -n hack/install-ate.sh`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
Agentgateway is a popular [AAIF](https://aaif.io/) project designed for
agentic systems. It is of particular interest for substrate users as it
was designed specifically to serve the role of an egress proxy for
agents and other AI-adjacent workloads, with strong support for
credential injection, token exchanges, TLS MITM, and other
authentication and authorization schemes.
Agentgateway supports the ext_proc protocol, like Envoy, so we use that
here. There has been some debate in the community around the long term
end state of ext_proc vs other options, but for now this maintains the
status quo.
While Agentgateway can be configured dynamically over XDS, it does not
take the same configuration as Envoy. However, none of the config is
dynamic anyways, so we use just a static configuration file (configmap)
for now to keep things simple.
While we don't have specific per-component/file OWNERS in the project at
the moment, I can informally commit to myself and Eitan maintaining this
integration. Given the goals around velocity in the project we can
commit to either rapidly fixing any issues that may arise or
removing/disabling the integration if this ends up slowing things down.
I think it need to be closed when looking at what `netns` is in
https://github.com/vishvananda/netns.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Phase 1 of #451, micro-VM only. gVisor support will follow in a separate
PR.
## What
Adds a `snapshotsConfig.onResume` block to the ActorTemplate, selecting
per snapshot situation what supplies the guest state at resume: each
field names what is being resumed *from*, the value names the boot
source. `onPause`/`onCommit` remain pure capture scopes (`Full | Data`),
and the `ActorSnapshot` records plain `Full`/`Data` content; the
golden-combine is strictly a restore-time behavior.
```yaml
snapshotsConfig:
onPause: Data
onCommit: Data
onResume:
fromData: Golden # ColdBoot (default) | Golden — Data snapshots resume as
location: gs://… # golden memory + fs delta with the actor's data layered on
```
Under `fromData: Golden`, a data-only restore combines the template's
**golden snapshot** (guest memory + full fs delta) with the actor's
captured durable data — the actor comes back with the golden's warm
state over its own data instead of cold-booting. `fromData` applies to
every resume of a Data snapshot — commit or pause checkpoint. Future
situations (e.g. resuming a Full snapshot across a template upgrade,
#477) get sibling fields under `onResume` rather than overloading this
one.
## How
- **CRD** (`pkg/api/v1alpha1`): new `OnResumeConfig`/`ResumeSource`
types; `fromData: Golden` is CEL-gated to `sandboxClass: microvm`.
- **Control plane** (`ateapi`): resume derives the behavior from the
template's `onResume` configuration — for a Data durable snapshot or a
Data pause checkpoint it resolves the golden snapshot's location and
sends the restore-only `DATA_ON_GOLDEN` wire scope, failing early (with
an actionable error) if the golden snapshot is missing or not Full.
Golden actors always commit `Full` regardless of `onCommit` — their
snapshot is the base the combine needs.
- **Wire APIs** (`ateletpb`/`ateompb`): restore-only
`SNAPSHOT_SCOPE_DATA_ON_GOLDEN`, rejected on checkpoints;
`golden_snapshot_uri_prefix` is a top-level `RestoreRequest` field
because the actor's snapshot may be a local pause checkpoint while the
golden snapshot is always external.
- **atelet**: stages a single combined folder — the actor's files (the
durable-dir tar) win name collisions, the golden snapshot supplies the
rest. External restores download both halves concurrently; local (pause)
restores copy the actor's files from the local checkpoint dir while the
golden's files download in parallel. The golden manifest's pinned
sandbox binaries run the restored guest and are recorded on-node for
later checkpoints.
- **ateom-microvm**: `DATA_ON_GOLDEN` restores through the existing Full
path — cloud-hypervisor relaunches from the golden's guest files while
the durable virtio-fs share serves the actor's re-materialized data.
gVisor rejects the scope (defense-in-depth behind the CRD gate).
## Testing
- Unit: envtest CEL cases, converter tables, atelet request validation
(checkpoint rejection, golden-URI rules incl. local+golden),
combined-download test against a fake object store, resume-workflow
golden-resolution cases for both the durable and pause paths.
- e2e (`suites/demo`): OnGolden lifecycle cases for the commit path, the
pause path, and a two-durable-volume variant (micro-VM lane only),
asserting counters across pause/suspend and the recorded Data content
scope.
- Manually verified on a live GKE cluster: Data commit and Data pause
both resume with the golden guest's memory over the actor's data, on
different pods.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Also Fixes#575
(1) mintCert + MintJWT proto changes
(2) validate that atelet is the one that calls actorIdentity Service.
(3) verify that this atelet is only requesting a cert for an actor on
its node.
(4) verify that the actor is still running
(5) create another substatex509 extension for actor identities
JWTs needs more love, lots as TODO there
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Adds a dynamic `--log-level` flag (debug, info, warn, error;
case-insensitive) to
the server binaries, backed by a `slog.LevelVar` in `serverboot`.
Defaults to
`info` — no behavior change unless the flag is set.
- **serverboot**: the LevelVar behind `InitLogger`'s handler;
`SetLogLevel(string)`
to parse and set it (empty string = unset, a documented no-op);
`LogLevel()`
returns a `slog.Leveler` for binaries that build their own handler.
- **atelet, ateapi**: new pflag, fatal on an invalid value.
- **ateom-gvisor, ateom-microvm**: flag accepted by the binary; note
their
container args are built by the WorkerPool controller with no
pass-through yet,
so wiring it is a follow-up — and that follow-up must roll out only
after this
change's ateom images are deployed (older images reject unknown flags
and would
crashloop).
- **atenet (router + dns)**: replaces two pre-existing hand-rolled
`--log-level`
parsers that silently fell back to `info` on a bad value; both now use
`serverboot.SetLogLevel` and `serverboot.InitLogger`, which also adds
`ate.dev/trace-id` correlation to atenet logs.
- **Kind manifests**: atelet runs at `--log-level=debug` on kind
(dev/CI), so e2e
suites can assert on per-item log lines; production installs keep
`info`.
Motivation: groundwork for #463 Phase 2 observability — per-item GC log
lines can
ship at Debug and be enabled per-node instead of landing at Info.
**Testing**: unit tests pin the untouched default at exactly `info`,
dynamic
raise/lower, case-insensitivity, invalid-value rejection, and the
empty-string
no-op; `go test -race` green across serverboot, atenet, atelet, ateapi.
Don't try to rollback mounts during Resume/Run. Instead should introduce
a way to cleanup actor resources when an actor is in a failed state.
Discussion in #643.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Second foundation slice of #463 Phase 2 (GC), independent of #650 and
inert on its
own: nothing reads the new field and nothing deletes anything.
- **`OverlaySpec.ImageDigest`** (optional, `omitempty`): the manifest
digest the
bundle's image ref resolved to, populated in `prepareOCIDirectory`. The
GC
root-set scan will use it to root the whole image record while the
bundle
exists. Compatible in both directions: ateom consumers never read the
field
(old binaries drop the unknown JSON key), and specs written by older
atelets
keep parsing with the field empty — so there is no deploy-ordering
constraint.
- **Atomic `WriteSpec`** (temp file + rename in the bundle dir): the
root-set scan
will read specs concurrently with bundle preparation, and a torn spec
would
under-report the layers an actor is using — an error in the dangerous
direction (toward deleting in-use layers). With the rename, a reader
sees the
complete old spec or the complete new one, never a fragment.
Landing early matters: every bundle written from now on carries the
digest,
shrinking the digestless-spec population the GC's compatibility fallback
has to
cover.
**Testing**: round-trip + old-spec compat and no-temp-files-left unit
tests;
`go test -race` green on both affected packages.
Third foundation slice of #463 Phase 2 (GC), independent of #650 and
#656. The
primitives are uncalled until the eviction engine PR; the only live
behavior is
the startup sweep, which only ever sees `.rm-*` dirs the engine will
create.
- **`retireLayer`**: evicts a layer with one atomic rename to a `.rm-*`
name
inside the layer singleflight; the slow `RemoveAll` is the caller's job,
outside all locks. Existence/mtime pre-flight runs *outside* the flight
(a
retire must never block behind an in-progress download) and both checks
are
re-run inside it before the rename. Returns gone/vetoed/retired so the
engine
can distinguish "nothing there" from "must keep".
- **Reuse interlock**: `ensureLayer` now refreshes the layer dir mtime
inside
the same flight, so a retirement and a reuse cannot interleave — either
the
retire wins (the pull re-unpacks) or the touch wins (the retire vetoes).
- **Startup sweep**: `.rm-*` dirs (a crash between rename and removal)
are
reclaimed alongside the existing `.tmp-*` dirs, via `RemoveAllWritable`
since
layer trees legitimately contain read-only content.
- **`isLayerDirName`** validates names read from disk or records.
Divergence from the prototype: `retireLayer` reports no freed size —
size
crediting belongs to the engine's call sites and arrives with it.
**Testing**: unit tests for the status contract, startup-sweep recovery
and
non-interference, and a retire-vs-pull race test; `go test -race` green.
First slice of #463 Phase 2 (GC): the two facts eviction will need,
persisted in the
filesystem so they survive atelet restarts. Inert on its own — no
behavior change for
any current caller, and nothing deletes anything.
- **Layer sizes**: unpack counts the uncompressed tar stream and writes
a per-layer
`size` file before the atomic rename (best-effort; a missing file is
healed by a
lazy backfill). Backfill covers layers unpacked by older atelets: one
walk, once
ever, skipping unreadable dirs (atelet has no `CAP_DAC_READ_SEARCH`) and
preserving
the dir mtime.
- **Last use**: cache hits touch the image record's mtime — the
persisted LRU
timestamp — under a new `hitMu` (shared side only; the eviction pass
will be its
exclusive holder).
- **`Store.CacheSize()`**: sums recorded sizes; the accounting behind
the future
`--image-cache-max-bytes`.
On-disk changes are purely additive, so no cache-layout version bump and
no
migration: an older atelet reads this cache dir fine (ignores `size`
files), and
this atelet reads a pre-upgrade dir fine (backfill). Rollback-safe.
**Testing**: unit tests for size recording, backfill (incl. the
unreadable-dir
contract and mtime preservation), hit-path touch, and `CacheSize`; `go
test -race`
green standalone; the demo e2e suite passes unchanged on Kind against
this build,
with `size` files and record-mtime updates verified on the node.
Bring the architecture and threat-model docs in line with the atunnel
ingress path: the router no longer rewrites :authority to a worker pod IP
and forwards over plaintext port 80, it opens an mTLS tunnel to atunnel on
worker port 443, which forwards to the Actor over its private veth.
Drop cmd/atenet/atenet-diagram.png. It predates the ext_proc/ORIGINAL_DST
design and is now wrong in the part that matters most. The README section it
illustrated gains a short accurate note about the upstream hop instead of a
dangling image reference.
Add an e2e suite covering the change end to end: TestActorDirectAccess
asserts that the worker pod's port 80 is no longer a reachable Actor ingress
path (the DNAT rule is gone) and that the same Actor still answers /readyz
through atenet-router over the atunnel mTLS hop. It uses the counter demo as
its fixture, so it only needs --deploy-demo-counter.
internal/e2e/testmain.go picks up ParseSkippedFlags so that `go test` flags
(-run, -v, ...) survive pflag parsing and reach the suite.
The router used to rewrite :authority to the actor's worker pod IP and let
the dynamic_forward_proxy cluster resolve it, reaching the actor over
plaintext pod-IP:80. The worker no longer DNATs pod-IP:80 to the actor, so
that path is gone.
Instead the ext_proc resolves the actor to its worker's atunnel ingress
address (IP:443) and puts it in x-ate-original-dst. An ORIGINAL_DST cluster
dials exactly that address, which leaves the request Host as the actor DNS
name -- atunnel needs it to authorize the active actor. The header is set
with OVERWRITE_IF_EXISTS_OR_ADD so a client-supplied value can never
influence the address Envoy dials.
The upstream hop is mTLS: the cluster presents the router's podidentity
credential bundle as its client cert and validates the atunnel server
against the podidentity trust bundle. Validation matches the SPIFFE URI SAN
prefix rather than the dialed pod IP, because the atunnel cert carries only
a spiffe:// URI SAN and Envoy's default SAN check against an ephemeral pod
IP would never match. atunnel in turn only accepts
spiffe://cluster.local/ns/ate-system/sa/atenet-router.
The dynamic_forward_proxy cluster, HTTP filter and DNS cache config are
removed along with the :authority rewrite.
Every worker pod now hosts an atunnel ingress server on :443 and an
atunnel egress listener, both long-lived, with per-activation
Activate/Deactivate bracketing every Run, Restore and Checkpoint.
The nftables rules ateom installs change accordingly:
* The pod-IP:80 -> actor-veth:80 DNAT is gone. Worker port 80 is no
longer an Actor ingress path; the only way in is the mTLS listener
on :443, which authorizes the caller and checks the Actor is the one
currently assigned to this worker.
* A prerouting REDIRECT sends Actor TCP egress to the local atunnel
egress listener, preserving SO_ORIGINAL_DST. It is only installed
when the activation carries an egress gateway address, which nothing
populates yet, so the masquerade path is unchanged and Actor egress
behaves exactly as before.
The ateom Run/Restore protos gain the two fields that arm that path:
egress_gateway_address, which decides whether the redirect is installed
at all, and actor_version, the Actor resource version ate-api observed
when assigning the worker, which atunnel asserts to the egress gateway
as a lower bound on trustworthy Actor metadata. Both are consumed here
and left unset. atelet and ate-api start populating them in the egress
gateway change, which is the point at which they mean anything, so this
change adds no new requirement to the atelet wire contract.
atecontroller gives worker pods the podidentity credential + trust
bundles (the atunnel server identity), the servicedns trust bundle, and
container port 443. podidentitysigner now issues certs with
ExtKeyUsageServerAuth as well as ClientAuth, without which the worker
cannot present its podidentity cert as a TLS server cert and the
gateway handshake fails.
atunnel is the in-worker component that carries Actor traffic in both
directions over authenticated channels.
* Server: an mTLS HTTPS listener that terminates connections from the
ingress gateway, authorizes the peer by its SPIFFE identity, checks
that the request targets the Actor currently activated on this
worker, and reverse-proxies to the Actor over the private veth.
* Egress: an activation-aware TCP proxy for transparently intercepted
Actor connections. It resolves the original destination via
SO_ORIGINAL_DST and forwards through an EgressDialer, asserting the
Actor identity to the far end. Dormant until an egress gateway is
configured; the client and dialer land here so the package is
reviewed as one unit.
Activate/Deactivate bracket an Actor activation so traffic is only
carried while an Actor is assigned to the worker, and Deactivate drains
in-flight streams before the Actor network is torn down.
Substrate documents how to instrument a service but not how to stand up
the collector it exports to, so pointing it at your own collector meant
reading six manifests.
Covers both topologies: GKE Managed OpenTelemetry as the default on GKE,
with why its gateway Deployment suits Substrate's elastic actor
workload, and a DaemonSet manifest for operators who need per-node
isolation, full config control, or a non-GKE cluster.
Part of #563.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixed language describing project goals and relationship to Kubernetes
in readme.md, architecture.md and roadmap.md
- changed the top language to focus on project goals instead of
Kubernetes relations
- changed the language describing Kubernetes relationship to describe
value of the Agent Substrate layer and Kubernetes layer
Fixes#564 (Part 1)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Description
This PR implements the first part of #564 by extending metric
`atenet.router.route.duration` with singleflight activation state
(`ate.router.resume`) and outcome status labels (`ate.router.outcome`).
All telemetry attribute keys and label value sets are centralized in
`internal/ateattr` to ensure consistency across components.
## Key Changes:
* Added `resumed` boolean to `ResumeActorResponse` in `ateapi.proto`
indicating whether a cold activation workflow occurred
(`!state.WasRunning`).
* Updated `workflow_resume_test.go` to assert `resumed == true` on cold
activations and `resumed == false` for already-running actors.
* Extended `classifyOutcome(err)` in Envoy `ext_proc` to classify errors
into metric labels: `ok`, `cancelled`, `timeout`, `no_capacity`,
`lock_conflict`, `not_found`, `unavailable`, `rate_limited`, and
`resume_error`.
* Added explicit handling for `codes.Unavailable` (`"unavailable"`) and
`codes.ResourceExhausted` / `StatusCode_TooManyRequests`
(`"rate_limited"`).
* Consolidated metric label names on `atenet.router.route.duration`
using standard keys from `internal/ateattr`:
* `ate.template.namespace` (`TemplateNamespaceKey`)
* `ate.template.name` (`TemplateNameKey`)
* `ate.router.outcome` (`RouterOutcomeKey`)
* `ate.router.resume` (`RouterResumeKey`)
* Updated `manifests/ate-install/kind/kustomization.yaml` to export
`atenet-router` OTLP metrics to `opentelemetry-collector`.
## Testing
* `go test -buildvcs=false ./cmd/ateapi/internal/controlapi/...`
* `go test -buildvcs=false ./cmd/atenet/internal/router/...`
* `make test`
## E2E Test
* `./hack/create-kind-cluster.sh`
* `./hack/install-ate-kind.sh --deploy-ate-system --deploy-demo-counter`
* `./hack/run-e2e.sh ./internal/e2e/suites/metrics/...`
`--otlp-collector-address` defaults to
`os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT")`, unify all stacks to use
environment variables.
#563
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
A resume can fail with a bare `dial unix /run/vc/vm/<uid>/clh.sock: connect:
no such file or directory`, one minute after it started (#619). The socket
was there — the caller waits for it before dialing — so it went away while
we polled, and cloud-hypervisor unlinks it in the vsock device's shutdown:
the guest died during boot. Under a contended host the guest's boot stretches
badly (measured: systemd reaching its default target 25s in, versus about a
second when idle), and a boot that stalls long enough is torn down guest-side.
So treat ENOENT as what it is. Instead of polling a socket that cannot come
back, and reporting the last dial error as if it were the problem, give up at
once with an error that names the cause, and retry the cold boot: a guest that
never reached its agent ran none of the actor's containers, and the failure
path tears the whole attempt down, so starting over is safe — and it is the
only recovery, since the dead VM is not coming back. Each retry is logged with
the guest's boot diagnostics (the console tail, plus each virtiofsd's log,
because cloud-hypervisor also stops the VM when a vhost-user backend dies and
that leaves the console silent).
When an e2e test failed, the evidence was deleted before anyone could
read it. The suite deleted every namespace it created on the way out,
taking the worker pods with it, and the workflow's post-failure dump
only looked at three fixed namespaces — never the suites' randomly-named
ones. So a failure inside an actor (#619: a micro-VM resume where the
guest died at boot) left nothing behind but the RPC error the test
printed.
Keep the namespaces when the suite failed, and dump every worker pod in
every namespace, so the ateom logs — which carry the guest's console
tail — reach the failed run's output.
Kept namespaces are nobody's to reclaim, and each holds a WorkerPool's
worth of running pods, so they now carry an ate.dev/e2e label and
hack/cleanup-e2e.sh deletes them once the logs have served their
purpose. CI throws its cluster away, but a development cluster
accumulates them run after run.
Fixes #<issue_number_goes_here>
Built while debugging
https://github.com/agent-substrate/substrate/issues/619
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The scan's other blocking finding on this file: with no permissions block a
job gets the repository's default token scopes, which for these jobs is more
than they use.
The org's zizmor scan audits every workflow file a PR touches, and
unpinned-uses is one of its mandatory checks: a mutable tag like @v5 can be
repointed at any commit, so the pin has to be a hash. The tags these were on
are kept as comments.
An ActorTemplate could declare only one durable-dir volume, and mount it
into a container only once. That is a gVisor limit, not a general one:
atelet declares the mount to gVisor through a single hardcoded annotation
key ("dev.gvisor.spec.mount.durabledir"), so a second volume would silently
overwrite the first.
The micro-VM runtime has no such constraint — every volume is a
subdirectory of the one writable virtio-fs share, and the snapshot tar
already archives that directory whole, so N volumes cost a subdirectory
each and round-trip through checkpoint/restore untouched. Gate the two
"at most one" CEL rules on sandboxClass so micro-VM templates may declare
several while gVisor keeps its cap, and say so in the messages.
What was actually missing was the volume NAME: ateom received mount paths
alone, so it inferred the single name by listing atelet's directory. Carry
each mount's name on the wire (Container.durable_dir_volume_mounts,
replacing the paths-only field, which is reserved rather than retyped since
atelet and ateom are separate images that can skew across a rollout), and
delete the inference. A container's binds now come from its own mounts, and
the actor-wide "has a durable share" question collapses to a bool.
The counter demo grows an optional --second-file-counter-directory, and a
micro-VM-only e2e case runs the lifecycle matrix against an Actor with two
durable volumes, asserting both counters advance together. gVisor skips it:
the template would be rejected at admission.
Otherwise if the deferred `Close()` failed the caller will still
received a successful response from `copyFile()`.
Fixes#611.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #232
1. During CreateActor, the volumes to be created get instantiated in the
Actor API in PENDING status. The Actor is still instantiated as
SUSPENDED.
However, the CreateVolume operation is moved to the beginning of
ResumeActor. This is because CreateActor is intentionally designed to
not be idempotent. The downside is that the user won't know dependent
resource creation failed until the first resume.
2. Changes volume id to use actor uid instead of atespace+actorname and
adds "substrate" prefix
3. Adds a new DELETING status that gets persisted before we start to
delete volumes.
This shifts the existing DB DELETE precondition checks to the new
DELETING status. And now you can only delete the actor in DB from
DELETING status.
4. Refactor DeleteActor to use a workflow
In the future, we will also add a workflow that will do mandatory node
cleanup of volumes before starting to delete the volumes.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
`ate.workerpool.workers` tallies by pool name, state, and sandbox class,
but a WorkerPool is namespaced, so two pools that share a name sum into
a single series. This is reachable today: the counter demo and the
autoscaled-workerpool demo both create a pool named `counter`, and with
both installed the demo HPA scales on the other pool's assigned workers.
Adds `ate.workerpool.namespace` to the tally key and attributes, and
pins it in the demo's adapter query and HPA selector.
`syncWorkerToStore` checks `pod.Status.PodIP != ""` before
`DeletionTimestamp`, so a Terminating pod reporting no IP never reaches
`markWorkerDraining` and stays `STATE_ACTIVE`, schedulable for as long
as the pod lingers. Draining resolves the stored record by name and
never reads the IP, so this swaps the two guards.
The startup reconcile runs through the same function, so a pod already
Terminating when ate-api-server restarts is now marked too.
`markWorkerDraining` already runs on every update event for a
Terminating pod, so this adds no new call pattern.
Fixes#600
The README lists 3 of the 6 demos and 5 of the 9 docs, and its command
tour skips `cmd/ateom-microvm` and `cmd/benchmarking`. The glossary
never defines Atespace, even though it is half of an actor's identity
and has its own API, and it omits SandboxConfig.