Fixes#1472.
The Telemetry API rejects an exponential histogram point with no
positive buckets. Give such points one zero-count bucket instead.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
We should not override the whole actor trust store with SSL_CERT_DIR,
removing that to defer using SSL_CERT_FILE, which is additive.
Additionally, this does not cover all languages and runtimes, adding
guidance for other env vars required and other runtimes that need union
of the trust stores in a single file.
In a followup, we are going to remove all "sdsmint" and "plain gateway"
language and just have one gateway thats capable of minting or not based
on the sni.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This is needed because DeleteActor cleans up all snapshots under the
actor's location.
If we were to allow repointing an actor's snapshot location, we risked
leaking historical snapshots upon actor deletion.
Adds a credential provider for egress credential injection backed by
Google Cloud Secret Manager.
**It lives in its own Go module under `plugins/gcp-secret-manager`,
temporally hosted here until it moves to a repository of its own.**
**What it does**
- Serves `credproviderpb.CredentialProvider` over mTLS and admits only
the egress gateway's identity (`--injector-identity`).
- Resolves global and regional secrets, optionally picking one key out
of a JSON payload:
`ate-secret://secretmanager.googleapis.com/projects/<project>[/locations/<location>]/secrets/<secret>/versions/<version>[/keys/<key>]`
- Enforces a default-deny atespace→project policy
(`--project-policy-file`), the counterpart of the Kubernetes provider's
namespace policy.
- Returns a retryable 503 only for transient Secret Manager failures,
and caps each read with `--fetch-timeout` (default 3s).
**Repository changes**
- New top-level `plugins/` directory for self-contained plugins,
documented in `docs/dev/code-layout.md` and `AGENTS.md`. The module
imports only substrate's public `pkg/` packages; a test enforces this.
- CI runs the module's tests, `make verify` and golangci-lint.
govulncheck scans the module too, with the action pinned by SHA.
- `docs/egress-credential-injection.md` describes each provider's
credential URI format. The plugin's README covers installing and using
it.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## Benchamark(atelet, ateom-microvm): log per-actor checkpoint and
restore phase breakdowns
### Why
`SuspendActor` and `ResumeActor` latency is only attributable down to
the `ate.actor.{checkpoint,restore}.duration` histogram phases, and the
biggest buckets — `ateom_checkpoint`, `ateom_restore` — are opaque. When
a large memory benchmark suspend takes 6 s we want to able to tell where
the time is been spent during the suspend.
### What
One joinable, developer-facing log record per operation per layer, with
the full actor identity (allowed in logs, barred from metric labels),
the snapshot scope, and one float-seconds field per phase that ran.
**atelet**
- `Checkpoint` now writes a `Checkpoint timing breakdown` record, the
same way `Restore` has written `Restore timing breakdown` since #1364.
Phases: `sandbox_assets`, `ateom_checkpoint`, `persist`, `total`. One
slice feeds both the histogram and the record, so they cannot disagree.
- Both records carry `error.type` when the operation failed (the gRPC
code; context errors map to `DeadlineExceeded` / `Canceled`). The record
is written on the way out of a failure too, so its completed phases are
kept, and the marker lets a reader exclude a timed-out restore from a
latency distribution. This restores what the record lost when the
`ateerrors` taxonomy was deleted (#1817), using the `error.type`
convention the ateapi instruments already follow.
**ateom-microvm**
- New `phaselog.go`. `CheckpointWorkload` and `RestoreWorkload` emit
records with the same two messages, under
`ateom.actor.checkpoint.duration.<phase>` and
`ateom.actor.restore.duration.<phase>`, decomposing atelet's
`ateom_checkpoint` / `ateom_restore` buckets:
- checkpoint: `pause`, `snapshot`, `durable_dir`, `rootfs_upper`,
`teardown`, `total` — the three captures run concurrently on the paused
guest, so the paused window costs their max, not their sum.
- restore: `prep`, `bundles`, `upper_join`, `lowers`, `tap`,
`vmm_launch`, `vm_restore`, `resume`, `wakeup_probe`, `total` —
sequential; they partition the total. A Data-scope cold boot records
`total` only.
- These timings already existed as ad-hoc `slog.Duration` fields on the
`Actor checkpointed` / `Actor restore phases` lines; those lines are
kept. The record adds stable keys, identity, and seconds (the
histograms' unit).
- The phase names are deliberately private to the binary rather than
added to `internal/ateattr`, so they cannot be mistaken for
`ate.snapshot.phase` metric values. They are micro-VM specific;
ateom-gvisor is unchanged.
**docs/observability.md** is updated: the Restore record is no longer
the only per-actor latency record, and the ateom records are described.
No new instruments, no registry changes, no behavior change.
### Example
```json
{"msg":"Checkpoint timing breakdown","ate.actor.uid":"8f2a…","ate.template.name":"glutton",
"ate.snapshot.scope":"full",
"ateom.actor.checkpoint.duration.pause":0.003,
"ateom.actor.checkpoint.duration.snapshot":0.846,
"ateom.actor.checkpoint.duration.rootfs_upper":0.022,
"ateom.actor.checkpoint.duration.teardown":0.232,
"ateom.actor.checkpoint.duration.total":1.081}
```
Joined with atelet's record for the same actor, a run of the glutton
workload (1 GiB resident, microVM) attributes a 6.0 s p50 suspend as 75%
`persist`, 19% `ateom_checkpoint` (of which the CH `snapshot` is 0.85 s
and `teardown` 0.23 s), and a 5.3 s p50 resume as 79% `download`, 18%
`ateom_restore` (of which `vm_restore` is 0.73 s). The consumer that
produces those tables from pod logs is a separate
`benchmarking/analysis` PR.
### Testing
- `go test ./cmd/atelet/...` and `./cmd/ateom-microvm/...` pass; new
unit tests cover the record shape (seconds, identity keys, zero phases
absent, no duplicate keys), the scope mapping, and `error.type` for
gRPC, context and plain errors.
- `GOOS=linux go vet` clean for both binaries; boilerplate and gofmt
clean.
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:
1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).
The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.
Partial work for #1871
- [x] Tests pass
```
Run make build-junittool
make build-junittool
bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
shell: /usr/bin/bash -e {0}
env:
E2E_ATENET_DATAPLANE: agentgateway
ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
TOTAL
```
- [x] Appropriate changes to documentation are included in the PR
Fixes#1911
`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.
No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Step 3 of #1748. Vocabulary only: nothing emits the event or sets the
epoch yet.
**Event.** `ate.actor.usage_sampled` in `actorevent`, sixteen required
keys via `ateattr`: identity, pool, sandbox class, source,
`ate.stats.kind` (periodic, initial, final), the four measurements,
observation time, epoch. Measurement names follow the
`ate.actor.stats.*` instruments; units are in the registry briefs. Body
is `Actor usage sampled`, the accepted stdout shape change.
**Scope.** The usage stream has its own instrumentation scope, the
lifecycle scope with `/usage` appended, so a consumer can sample or drop
it without touching lifecycle records. `NewEmitter` takes the scope; the
package-level `Log` keeps the lifecycle one, so ateapi is unchanged.
**Epoch.** `ate.actor.epoch` and `WorkloadStatsSample.epoch_unix_nano`:
the unix-nano time an activation began. CPU time is per activation for
every source; lifetime is the sum over epochs of each epoch's highest
value. Memory usage and working set are absolute; peak is as the source
reports it. An ateom that leaves the epoch at zero still sends the raw
guest counter. The measurement comments on fields 8 to 11 say all of
this.
**Registry.** Event group in `events.yaml`, seven attributes in
`metrics.yaml`, `ate.actor.epoch` in the no-actor-identity forbidden
list. The one-record-one-place rule and the collector guide now name
both scopes.
Left for step 4: the event's code anchor moves to the ateom emit site,
atelet's poller comments and the guide's stdout usage section get
rewritten, and the controller starts passing `OTEL_LOGS_EXPORTER` to
worker pods.
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.
It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#1266
Builds on #1689#1283 in particular.
Outstanding:
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~
EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Step 2 of #1748, plus the consumer rule from step 1b.
`LoggingOptions` gains `ExporterConn` and `RelayCapable`, mirroring
`TracingOptions`. Both ateoms initialize logging after metrics with the
relay connection and defer a nil-guarded shutdown. `OTEL_LOGS_EXPORTER`
still defaults to `none`, so no ateom emits a record; this is the
provider the actor events land on when emission moves.
A relay-capable log resource carries `ate.otlp.relay` like traces and
metrics do. `newLoggerProvider` is split from `InitLogging` so the test
asserts that on an emitted record's resource, the way
`TestMeterProviderRelayAttribute` does for metrics.
`docs/observability.md` states the consumer rule that replaced the relay
denylist dropped in #1800: a lifecycle record is authoritative only
under ateapi's resource, because the relay admits only ateom resources.
Follow-up for the emission step, not here: the controller does not
inject `OTEL_LOGS_EXPORTER` into worker pods, so no environment
exercises log-over-relay yet.
Part of #1664, implements #1799.
## Changes
- `ate.worker.state` now has `idle`, `partial`, `at_capacity` and
`unschedulable`. `assigned` is removed.
- ateapi picks the state by comparing allocated actor slots with
capacity, the same check the scheduler makes.
- Draining workers and workers with no reported capacity are
`unschedulable`. This wins over occupancy, so the sum over the states is
still the pool size.
- Every known pool reports all four states, set to 0 when empty.
- Updated the registry, `docs/observability.md` and the
autoscaled-workerpool demo (`assigned` -> `at_capacity`).
## Open questions
- Are we fine with `unschedulable` as a fourth state? cc. @JeffLuoo
Every new worker starts with capacity 0, so `at_capacity` would make an
HPA on it scale up on its own scale-up.
- This breaks queries on `ate_worker_state="assigned"`. Should be still
fine.
- The demo HPA is only right while each worker holds one actor. Pool
utilization needs slot counts, which needs a new instrument (#1664).
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Follow up on the comment -
https://github.com/agent-substrate/substrate/pull/1675#discussion_r4031619230
Stamps `actor_template_uid` onto `ExternalSnapshot` so template
provenance travels with the snapshot itself, rather than being inferred
from the single `ActorStatus.current_actor_template_uid`.
#### Motivation
`current_actor_template_uid` records the template the **last sprint
booted with** (`finalizeRunning` stamps it on every resume). It does
*not* record the template the **snapshot** was captured under. Those two
diverge as soon as an actor runs a sprint that does not produce a new
durable snapshot, and `loadActorForResume` was using the former to
decide whether to force a `DATA`-only restore instead of `FULL`.
Walking an actor through a repoint and a crash:
| Step | Action | `current_…_uid` | `ExternalSnapshot` |
|---|---|---|---|
| **a** | Suspend on template **A** | A | captured on **A** |
| **b** | `UpdateActor` repoints the spec to **B** | A | captured on
**A** |
| **c** | Resume → `A != B`, restores `DATA`, sprint boots on **B** |
**B** | still on **A** |
| **d** | Actor crashes (or is reverted) — no new snapshot taken | B |
still on **A** |
| **e** | Resume → `current == target == B`, so **no repoint is
detected** | B | still on **A** |
At step **e** the old check reports "template not replaced" and restores
the external snapshot in `FULL` — replaying a memory image captured
under **A**'s sandbox on **B**. That is exactly the case the guard
exists to prevent; it was silently defeated by the sprint at step **c**
advancing `current_actor_template_uid` past the snapshot.
*(Note: Paused actors do not face this divergence. Because `UpdateActor`
is only allowed in the `SUSPENDED` state, a paused actor cannot have its
template updated. Therefore, local checkpoints do not need separate
template provenance tracking).*
#### Changes
- **Data Model Updates:** Added the `actor_template_uid` field to
`ExternalSnapshot`.
- **Snapshot Stamping:** Updated the capture logic to stamp external
snapshots with the active sandbox's `ActorTemplate` UID when suspending
an actor or creating an actor from a tag.
- **Resume Evaluation Logic:** Modified the external resume logic to
compare the target template UID against the specific snapshot's UID
(`ExternalSnapshot.actor_template_uid`) rather than the actor's overall
status. Local restores bypass this check entirely since templates cannot
change during a pause.
**Tests**
- `TestResumeActor_AteletWireRequest` (unit): snapshot fixtures now
carry `actor_template_uid`, since provenance is read from the snapshot
rather than from actor status.
- Functional expectations for create/update/resume/pause updated for the
new field on the external snapshot golden files.
- `TestUpdateTemplateLifecycle` (e2e): asserts
`external_snapshot.actor_template_uid` after suspend and re-suspend.
- **e2e (Update Template)**: Added coverage for repoint detection after
a revert (verifying the `resume` → `revert` → `resume` path described in
the table above).
- **e2e (Combined Volumes)**: Added coverage to explicitly verify which
volumes a revert rewinds.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.
Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.
Follow up for the PR - #1675
Issue - #1556
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Step 1 of #1748. The relay forwards `LogsService` the way it forwards
traces and metrics: source gate, verbatim payload, atelet's own
`OTEL_EXPORTER_OTLP_LOGS_HEADERS` upstream.
No content check on the records. Identity spoofing through the relay is
handled in #1748 by the consumer rule (a lifecycle record counts only
under ateapi's resource, which the source gate already enforces) and
per-pod sockets (#741), not by a denylist here.
**Endpoint and compression resolution** (second commit, from review).
The relay read the generic, traces, and metrics OTLP variables and
ignored the logs ones, so a logs-specific collector or compression
setting was silently overridden. All three signals now go through one
resolver: each falls back to the generic, and every resolved value must
agree, since the relay has one upstream connection. The conflict error
names the variable each value came from; a signal that falls back names
the generic variable, not its own unset one.
One behavior change: `traces=A, metrics=A, generic=B` with logs unset
used to pass and now fails at atelet startup, because logs falls back to
`B` and that is a two-collector configuration.
Tests: an ateom log batch arrives with the ateom's resource; the source
gate, mixed-batch, header-isolation, and per-signal header tests cover
logs; resolver tables gained logs rows including the fallback case. Docs
say the relay carries logs, traces, and metrics.
This is part of #1266 , opening now for discussion.
Stacked on https://github.com/agent-substrate/substrate/pull/1682 which
was slightly orthogonal.
This is loosely based on the mini proposal by @howardjohn as discussed
in the community meeting, and feedback from @bowei @EItanya
@LiorLieberman @aojea.
https://docs.google.com/document/d/1TycfQ3iiEpbI3rveMIj0S2PpPuLecb8I5R--yTpt9Ig/edit?resourcekey=0-kJbtEZ-KGzuL5eCjHDvBhg&tab=t.0#heading=h.ga9bfaf55ptk
Roughly:
1. Actor sandboxes each get their own netns.
- gVisor grabs all interfaces in the netns, and tap currently requires
running something like slipr, so for now we do two netns + a veth when
gVisor.
3. In the netns we directly intercept TCP => atunnel for general
traffic.
4. In the netns we serve a trivial TCP+UDP DNS relay to the pod
resolution.
- In the future we can insert policy here.
6. We consistently inject a modified resolv.conf instead of
bind-mounting it (gVisor) across both runtimes.
7. Readiness probe dialing happens in the actor netns.
Every actor gets the same fixed guest IP as before, which is only
visible to the actor.
All inbound/outbound traffic comes from atunnel / the DNS relay.
The actor no longer has any direct use of the pod interface, so we can
begin to consider ateom using the network itself.
When we add the rest of multi-actor changes, this greatly simplifies
thing.
Full multi-actor requires further changes, but this diff is already
large (suggest reading commit by commit) and can stand-alone. I'll file
more stacked changes when we've got consensus on this one.
The name `readyz` is not very descriptive. We also want to avoid
`readinessProbe` to prevent confusion with the Kubernetes concept, which
represents continuous traffic gating. Renaming this field to
`wakeupProbe` clarifies its actual behavior, and the fact that it only
operates when the actor is woken up.
Part of #1378
Closes#1802.
## What this does
Deletes the `ate.scheduler.eligible_workers` histogram. Nothing replaces
it in this PR.
## Why
**It is expensive on the resume path.** `Schedule` built a second slice
of the matching workers that scheduling never read, walked it again, and
repeated `HasRoom` for each entry. `HasRoom` parses resource quantity
strings up to three times for each worker. The metric roughly doubled
the per-worker arithmetic of a placement. The filter loop is now one
pass over the fleet and builds one slice.
**A metric cannot answer the question it was built for.** #564 wanted to
tell a full fleet apart from an empty intersection of the constraints,
such as a cordoned node behind a node requirement. That needs the
selectors, the node requirement and the actor.
`docs/metrics/substrate.yaml` bars all three from a metric label and
sends them to logs and spans, so no shape of this instrument reaches the
answer. A log record at the rejection is the replacement the issue
names, and it is not in this PR.
## Also removed
`ate.scheduling.constraint` goes with the instrument. No other signal
used the attribute.
## Scope of the change
- `cmd/ateapi/internal/scheduling/metrics.go` deleted, along with the
`WithMeter` option and the second filter pass in `Schedule`.
- The instrument and the attribute removed from
`docs/metrics/registry/metrics.yaml`. The `pool-keys-paired` exception
in `docs/metrics/substrate.yaml` dropped, because the empty pool pair
was this instrument's alone.
- `docs/observability.md`: the table row and the two label notes
removed.
- The e2e collector check and the label assertions for the instrument
removed.
- `TestSchedule_EligibleWorkersMetric` removed. The behaviors it covered
(draining workers, a sandbox class mismatch, busy workers) are already
in the `TestSchedule` table.
## Testing
`go test ./cmd/... ./internal/...` passes. `make verify` passes except
`hack/verify/metrics.sh`, which needs Weaver or Docker. Neither is
available on this machine, so CI is what checks the registry.
Fixes#1474
`atenet.router.route.duration` carries two labels that the router filled
in wrongly on a failure path.
## `ate.router.resume` said "none" for a failed resume
`none` means "the resume found the actor already running" — the warm
route. The router also gave `none` to every failed resume, every
canceled caller, and every caller shed by a full parking lot. Those
failures landed in the warm-route series, so they did not show in the
cold-start rate and they polluted the warm-route latency.
A gRPC code cannot recover the fact. A canceled leader's flight outlives
its request and keeps restoring the actor, and a `DeadlineExceeded` can
land in the middle of a restore. So the PR does not classify the error.
It reserves `none`, `triggered` and `joined` for a resume that
completed, and adds a fourth value, `unknown`, for every resume that did
not:
| Case | Before | After |
|---|---|---|
| Actor already running | `none` | `none` |
| Leader completed a cold activation | `triggered` | `triggered` |
| Joiner waited on a completed cold activation | `joined` | `joined` |
| Resume failed (leader and joiners) | `none` | `unknown` |
| Caller canceled before the flight finished | `none` | `unknown` |
| Caller shed by a full parking lot | `none` | `unknown` |
| Egress, which never resumes an actor | `none` | `unknown` |
## `ate.template.*` was the empty string on a failure
The registry marks `ate.template.atespace` and `ate.template.name`
required on this metric, but the router sent an empty string when it
failed before it resolved a template.
`ateattr.NormalizeTemplateDimension` now sends the constant `"unknown"`
instead. The constant does not come from the request, thus it adds no
caller-controlled label value — the cardinality rule in
`docs/metrics/substrate.yaml` still holds, and the PR records the
exception there.
## How to read the change on a dashboard
A query that already splits by `ate.router.resume` gains an `unknown`
series and loses the failures that used to hide inside `none`. A `none`
rate reads lower after the change and a cold-start rate is unaffected.
Do not read the time in the `unknown` series as an activation time.
## Documentation
- `docs/metrics/registry/metrics.yaml`: the `unknown` member of
`ate.router.resume`, the `"unknown"` fallback on the two template
labels, and a note on the metric.
- `docs/metrics/substrate.yaml`: the cardinality-rule exception.
- `docs/observability.md`: the label list for the metric.
## Tests
- `resumer_test.go`: a failed flight gives `unknown` to the leader and
to all nine joiners; a canceled caller gives `unknown`; a shed caller
and a shed joiner give `unknown`.
- `metrics_test.go`: empty template dimensions record as `"unknown"`;
`Result.resume()` defaults to `unknown`.
- `ateattr_test.go`: `NormalizeTemplateDimension` and the four
`RouterResume*` values.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Kata 4.1.0 bundles virtiofsd 1.14.0 (required by Substrate). This
simplifies the dev process on arm64 because it removes the need to build
virtiofsd from source.
Verified on an arm64 KVM host (Lima + kind).
Fixes#1695
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Today, `ateom` fetches the `kata-config` asset (configuration-clh.toml)
to read only 3 values `default_memory`, `default_vcpus` and
`kernel_params`. The first 2 are redundant because (1) the values never
change and (2) ateom has the same defaults. Only `kernel_params` is
relevant for the `kata-agent`; its value also changed once in the Kata
project history.
`ateom` now owns all three values, which also allows us to tune them
specifically for Substrate.
Verified e2e with a GKE cluster.
Fixes#1693
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Record which egress protocols are allowed, which are allowed only under
policy controls, and which are blocked, along with the data path each
one takes. Having this written down gives users a single place to check
what Substrate lets an actor reach, and gives us a checklist to
implement and test against before GA.
Point readers at the issue tracker so that requests for traffic we do
not yet support arrive with a use case attached.
Address #1339
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
An actor lifecycle record went to stdout and to OTLP in two separate
calls. The attributes were shared, but the severity was not, so the slog
level sat at the call site and the OTel severity sat on the `Event`.
Nothing made a caller write both copies either, so a new record could
reach stdout only and no test would notice.
Since this PR `actorevent.Log` now writes both copies. The level comes
from `Event.Severity`, so it is stored once. The OTLP only `Emit` and
the exported body constants are gone, so the dual write is the only way
out of the package.
Also normalizes `ate.actor.operation.name` on the crash event, which the
state change event already did.
Fixes#1744
Testing
- Unit tests for the level mapping, the dual write, and the registry
check.
- End to end on a fresh kind cluster. Both event names arrive with the
right severity (9 and 17), the right attributes, and trace context on
the record fields. The stdout copies match record for record.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
As we're getting closer to GA, it's important to start documenting the
network contracts, especially where we want things to be pluggable. This
PR takes a stab at documenting the contract between atunnel and the
egress PEP. It doesn't have a policy section since that's evolving, but
this aims to be a good first step.
/cc @EItanya @thockin @bowei @LiorLieberman
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
#### Summary
This PR introduces the `RevertActor` RPC for actor lifecycle management.
It includes the API definition, the corresponding workflow execution
logic, observability metrics, and updates to the authorization model to
support reverting actors.
Fixes#1556
Docs updated in #1711
#### Commit-wise Changes
**1. Add RevertActor RPC (`1145d99`)**
* Introduces the new `RevertActor` RPC to the API definitions.
* Updates the corresponding protobuf bindings (affecting files like
`ateapi.pb.go` and `ateapi_pb2.py`).
**2. Add the revertActor workflow and observability metrics
(`dcb1fae`)**
* Implements the core `revertActor` workflow logic, designed to be
idempotent and re-enterable. It progresses through the following steps:
* **Mark Reverting:** Validates that the actor is in a revertable state
(`RUNNING`, `PAUSED`, or `CRASHED`) and transitions its state to
`REVERTING`.
* **Discard Worker:** Safely tears down the execution environment by
terminating the workload, detaching volumes, and releasing the assigned
worker.
* **Collect In-Progress Snapshot:** Cleans up external object storage by
deleting any objects a previous suspend operation was partway through
writing.
* **Finalize:** Commits the actor to `SUSPENDED` and strips all
node-local and in-progress state pointers (clearing `WorkerAssignment`,
`LocalSnapshotInfo`, etc.), returning the actor to its untouched
external snapshot.
* Instruments the workflow with lifecycle operation metrics (e.g.,
updating `ate.actor.lifecycle.operation.duration` to track `revert`
operations).
* *Note/TODO:* Currently, when reverting a paused actor, the workflow
drops the pointer to the node-local state but does *not* actually prune
the local checkpoint bytes from the node (this is tracked in #641).
**3. Add `can_revert` to the authorization model (`2fc53ae`)**
* Adds the `can_revert` permission to the auth model, mirroring the
shape of `can_suspend` (editor tier of the parent atespace, plus a
direct grant so a machine identity can revert the actor it drives
without holding an atespace role).
**4. Serve RevertActor and add the CLI verb (`f04fab4`)**
* Wires the `Control.RevertActor` service method to the workflow
(replacing the generated stub that previously answered `Unimplemented`).
* Adds the `"ate revert actor"` CLI command, making the feature usable
end-to-end.
* Implements `Terminate` for the fake atelet. This was necessary because
reverting an actor from the `RUNNING` state is the first path to reach
this call in functional tests (previously, delete tests skipped this
step as they ran against actors with no worker assignment).
**5. Add a manual verify script for RevertActor (`8022db3`)**
* Adds a script to manually exercise `RevertActor` against a real
control plane, since unit and functional tests only run against a fake
atelet.
* Tests reverting from `CRASHED`, `RUNNING`, and `PAUSED` states, and
verifies that attempting to revert a `SUSPENDED` actor is properly
rejected.
* Simulates a crash by deleting the worker pod the actor runs on to
verify the workflow can handle the absence of a worker to terminate.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## What this does
ateapi writes a record every time an actor changes state. Until now
those records only went to the pod's stdout, and nothing reads stdout.
This sends the same records to a collector as OTLP log events.
Actor name and uid cannot be metric labels (too many values), and traces
are sampled at 1%. So these records are the only way to answer "what
state is this actor in, and since when".
## Changes
- `serverboot.InitLogging` sets up a LoggerProvider, next to the
existing tracer and meter ones.
- New `internal/actorevent` package builds the log records.
- ateapi emits at the two places that already write the stdout records.
- Two event names: `ate.actor.state_changed` and `ate.actor.crashed`.
- Both names are registered in `docs/metrics/registry/events.yaml`, so
`make verify` checks them.
- kind gets a logs pipeline and a count connector. The e2e suite reads
the counts back.
- Docs updated. `otel-collector.md` said substrate has no
LoggerProvider, which is no longer true.
## Opt-in
`OTEL_LOGS_EXPORTER` defaults to `none`. Only the kind overlay sets it
to `otlp`. The base ConfigMap is untouched, so no deployed environment
changes when this merges.
## Notes on the design
- **No slog bridge.** Only two call sites emit these records, so
emitting twice costs two lines. A bridge would also send every ateapi
log over the wire, could not set the event name, and would loop, because
SDK export errors are logged through slog.
- **Batching processor, not the simple one.** These records sit on the
actor resume path. A processor that exports inside the emit call would
add a blocking gRPC call there, so a slow collector would become control
plane latency.
- **Two event names, not one per state.** `ate.actor.state` already says
which transition happened. A crash gets its own name because it carries
two extra attributes and a higher severity.
- **Both copies are kept on purpose.** No collector in this repo reads
pod stdout, so nothing is duplicated today. `kubectl logs` keeps
working. If a filelog agent is ever added, drop one of the two. The
escape hatch is written down in `docs/metrics/substrate.yaml`.
## Dependencies
Adds `otel/log`, `otel/sdk/log` and `otlploggrpc`, all pinned at
v0.20.0. That is the release that matches the pinned `otel v1.44.0`.
v0.21.0 would pull the core modules to v1.45.0, which this change does
not need. The logs API has a
v1.47.0 release candidate upstream, so it is on its way to stable.
## Testing
- Unit tests for the exporter resolver, the record builder, and both
ateapi emit sites.
- The record builder test checks the attribute set matches what the
event name declares, in both directions.
- The ateapi tests check the OTLP record carries the same attributes as
the stdout record.
- Ran end to end on kind. Records arrive with the right event name,
severity, attributes, and with trace context on the record's own fields
rather than as attributes.
- Checked the off state too. With `OTEL_LOGS_EXPORTER` removed, the
collector receives no log records and stdout is unchanged.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Nested virtualization exposes `/dev/kvm`, which microVMs need. The new
`--enable-nested-virtualization` flag turns it on for the node pool
`setup-gcp` creates. It defaults to on, so pass
`--enable-nested-virtualization=false` for a cluster that does not need
it.
Also, rename the env var `GVISOR_NODE_MACHINE_TYPE` to
`NODE_MACHINE_TYPE` to be consistent with other env vars and to avoid
confusion for microVMs. The old env var still works with a warning.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1507
Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.
`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.
This PR is based directly on `main` and does not depend on #1521.
Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
The lot admitted every request before its resume lookup, so a full lot
shed traffic to already-RUNNING actors — exactly the starvation
docs/request-parking.md rules out (agent-substrate/substrate#1081), and
TestHandleRequestHeaders_ParkingLotFull pinned that behavior as expected.
Replace the resumer's singleflight.Group with a per-actor flight
registry whose flights signal their park transition (the first
retryable error). A caller now waits slot-free while its flight
resolves, acquires a slot only once the flight parks, and is shed with
the existing 503 "router at capacity" only at that transition — after
the single attempt that revealed it would have to wait. Requests
resolved on the first attempt never touch the lot, so a saturated lot
cannot starve running-actor traffic, and parking.active /
parking.wait.duration now count only genuinely parked requests.
The port-forward, CA bundle and token only need the running cluster,
so they can be prepared before the roll starts. Doing it there lets the
checklist confirm the installed ate-api-server serves DrainWorker
before anything changes, instead of finding out in the per-node step.
The intro said at most one node's worth of capacity is out of service
during the roll, which is only true when there is more than one node.
Say so, and point at the single-node note in the per-node step.
Phrase the first warning as the mistake, like the other two. Add
rollback entries for the new DaemonSet and the cloned pools, so an
operator who stopped before flipping any node finds their step in the
list. Say what the progress commands do not show instead of claiming
that steps leave no trace. Wait for a controller-triggered roll to
settle before cloning pools. Note that a single-node cluster makes the
per-node step a full stop. The setup-gcp README's eviction warning
still said 60 seconds; it is 30 minutes.
# Record actor state changes from ateapi
ateapi now writes a log record every time an actor changes state.
## Before
Nothing tells you what state an actor is in. Metrics can't carry actor
identity, and the control plane samples traces at 10%, so most
transitions leave nothing behind at all. If you want to know whether an
agent is running, paused or suspended, there's nowhere to look.
## After
One filter:
```text
ate.actor.uid="8f2a1c4e6b0d47f1" AND ate.actor.state!=""
```
Last record wins. That's the state it's in now, and the timestamp is
when it got there.
The record:
```json
{
"time": "…",
"level": "INFO",
"msg": "Actor state changed",
"ate.atespace": "ate-demo-counter",
"ate.actor.name": "counter-1",
"ate.actor.uid": "8f2a…",
"ate.template.atespace": "ate-demo-counter",
"ate.template.name": "counter",
"ate.actor.operation.name": "suspend",
"ate.actor.state": "suspended"
}
```
## The two new keys
`ate.actor.state` is the `ateapipb.ActorState` values lowercased, so the
log wording and the state machine can't drift apart.
`ate.actor.operation.name` says what caused the change. The state on its
own doesn't tell you: an actor lands in `suspended` from a suspend and
`paused` from a pause, and only one of those gives the worker back.
`Actor crashed` picks up the same state key. A crash is a state change
too, and it's the one an actor can reach without any operation
finishing.
## Notes for review
**Why ateapi.** It owns the state machine. It also sees transitions that
never reach a worker, like deleting an actor that was already suspended,
or a resume that fails in the scheduler.
**It can't log a state the store never had.** The record goes out after
the store commit, and every state commit has a version precondition, so
a losing writer in a concurrent update writes nothing.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes https://github.com/agent-substrate/substrate/issues/1566
Today the Resume workflow resolves its restore source in the following
order
1. First check if the actor has node-local snapshot,
2. then its own durable external snapshot,
3. then the template's golden snapshot.
The boot flag was consulted at exactly one point in that chain, where it
suppressed using the golden-snapshot, which made its behavior much
narrower than "boot from scratch" suggests:
- Actor has its own external snapshot and boot=true: flag ignored,
restores the actor's snapshot.
- Actor has a local snapshot and boot=true: flag ignored, restores the
local snapshot.
- Actor has no snapshot, template has no golden snapshot: cold boot from
the spec regardless of the boot flag.
- Actor has no snapshot, template has a golden snapshot: boot=false
restores the golden, boot=true cold boots from the spec. This is the
only case where the flag is used.
The glutton benchmark was the only caller that set boot=true, on each
actor's first resume, to report true cold-start latency as a separate
ResumeActorColdStart stats row. @maxsmythe let me know if this is
required.
The proto field number and name were reserved.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## What
Updates the project overview language in the three docs that share it,
and fixes a long-standing typo.
- **`README.md`** — replaces the overview paragraph with the
secure-by-default positioning: density relative to standard container
runtimes, resume latency and activation throughput, and native
kernel/network isolation.
- **`docs/architecture.md`** — adopts the same lead sentence, keeping
the existing control-plane detail; `computer infrastructure` → `compute
infrastructure`.
- **`docs/roadmap.md`** — `computer infrastructure` → `compute
infrastructure`.
## Notes
The performance figures in the README paragraph (density multiple,
sub-500ms resume, activation rate) have been discussed and aligned
separately.
Docs-only change; no code or behavior is affected.
The container image must include the image digest. This validation was
dropped during the migration of the ActorTemplate resource from CRD to
Substrate API.
Part of #932 (PR 2 of 3). PR 1 (#941) added the `trustBundle` SystemInfo
data source, resolved on the node at Run/Restore. This PR keeps those
projections current while the actor runs.
## Live refresh
A `systemInfoVolumeRefresher` in atelet, modeled on kubelet's projected
volumes: `collectData` is the one place that builds a volume's complete
contents from its spec, `write` applies them, and every lifecycle point
uses the pair. Run/Restore registers the actor's volumes, which writes
them fail-closed before the sandbox boots; ClusterTrustBundle events
rewrite them while it runs; Checkpoint, Terminate, and failed starts
deregister. No API, proto, or RBAC changes: the wire still carries only
`{name, path}`.
- Informer events only enqueue bundle names; a single run loop writes.
Failed writes requeue with backoff through a rate-limited workqueue;
resolution failures keep last-good contents and wait for the bundle's
next event. Resync (24h) is only a guard against missed watch events.
- Change detection hashes the raw backing contents, not the projected
output (sanitization shuffles). Files are replaced by temp-and-rename at
stable paths and byte-identical files are left untouched: a needless
rewrite replaces an inode that suspended guests re-bind on resume.
- Nothing is persisted. Registrations are in memory; after an atelet
restart, a running actor's files keep their last-written contents until
its next Run/Restore rewrites them from current cluster state.
## virtiofsd: `--migration-on-error=guest-error`
CI confirmed the hazard (3 of 3 micro-VM runs): a rotation renames over
a file whose inode the guest still references (a dcache reference is
enough, no held fd), the actor suspends before re-reading, and
find-paths serializes an inode with no findable path. Under the default
`abort`, the destination virtiofsd rejects the device state at
vm.restore: suspend succeeds, every restore fails, and the actor is
permanently stuck. Full analysis in the PR comments.
With `guest-error` the restore succeeds and only the stale reference is
faulty: EIO on access until a fresh lookup at the stable path heals it.
That matches the documented reader contract (the platform rewrites the
file, applications re-read it) and also covers guest-created
`O_TMPFILE`s held across suspend, which already fail restores under
`abort`. The flag is backend-process configuration, so existing
snapshots (template goldens included) restore unchanged. Trade-off: a
share-reconstruction bug now surfaces as post-resume EIO and a readyz
failure instead of a loud restore failure; virtiofsd names the faulty
inodes in the worker pod log.
## E2E
The identity suite covers live refresh on both sandbox classes: rotate
the pool under two running actors and wait for both to observe the new
bundle live; rotate/suspend/resume without waiting for propagation, so
the resume must deliver rotated contents whether the rewrite landed
before or after the guest went down (the exact `abort` repro); assert a
sibling that never cycled also got the rotation live. The probe fixture
now keeps its namespace when a test fails so the worker pod logs
survive, and the cloud-hypervisor client includes the vm.restore
response body in errors.
## Not in this PR
Auto-injected egress trust volume (#932 PR 3) and the configurable
backend registry (tracked in #932).
Component authors had the metric registry and the cardinality rules but
no single place that says how to add an instrument: which kind to pick,
how to name it, which labels are allowed, how the Go is shaped, how to
test it with a ManualReader, and how to register it. Each of those
conventions existed only as a pattern in one of the metrics.go files.
The guide sits beside tracing.md and otel-collector.md. It uses the
instruments already in the tree as the examples, ends with a worked
example of a node-local snapshot cache (a counter by outcome, a fill
histogram, a size gauge and an eviction counter, and what is
deliberately not a label). The registry snippets in the guide pass
`weaver registry check` when appended to the real registry.
observability.md and AGENTS.md now point at the guide from the place
that already told contributors to update the registry.
Fixes#1552
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR