They are useful for granular upgrade
They will not be used for install as we have `deploy ate-system`, they
are used for more granular steps in upgrade runbook.
Part of #1686
# benchmarking: actor telemetry for the sweperf workload
## What this does
Adds the suspend/resume and task-execution timings from the telemetry
proposal to the sweperf boomer workload.
**New locust rows**
| row | what it measures |
|---|---|
| `ResumeToFirstExec` | resume RPC start until the sandbox accepts the
cycle's `/execute` |
| `TaskCEL` | in-container execution time for one sweperf task (4
cycles), as reported by `replay.py` |
| `CycleCEL` | in-container execution time for one cycle |
| `TaskWallClock` | client wall clock for one task, excluding think time
|
| `<rpc>_rtt` | client round trip for each control-plane RPC |
The four derived rows are recorded under the `actor` method, the `_rtt`
rows under `grpc`.
We keep both the client RTT and the server-side elapsed for every
control-plane RPC so the network and queueing overhead stays visible
separately. The existing `ResumeActor` / `SuspendActor` rows keep the
server-side elapsed from the response trailer.
The liveness check at session start leaves the actor running, so the
first cycle's resume is a no-op. Its successful `ResumeActor`,
`ResumeActor_rtt` and `ResumeToFirstExec` samples are skipped (failures
are still recorded), the same way ateapi skips no-op resumes.
## Poll interval
The `/status` poll interval is now configurable with
`--sweperf-poll-interval-ms` (env `LOCUST_SWEPERF_POLL_INTERVAL_MS`),
default 100 ms (was a fixed 25 ms). It flows through dynconfig like the
other `--sweperf-*` flags and is read per job, so a mid-run change
applies from the next cycle.
## Documentation
`benchmarking/README.md` had no sweperf section, so this adds one
listing every row the workload emits and the poll interval flag,
matching the existing DurDir section.
## Testing
`go test -race ./internal/benchmarking/boomer/...`, plus runs on a real
cluster (sympy, 4 users / 2 workers, gVisor): 5 min at the default 100
ms poll interval and 2 min at 250 ms.
- Both had 0 failures and all expected rows.
- The min `ResumeActor` is 334 ms, so the first-cycle no-op is no longer
recorded.
- `ResumeActor` / `SuspendActor` counts match (197 / 197), and each
`_rtt` is within ~1 ms of its server-side row.
- Client overhead over CEL per cycle is 65 ms at 100 ms vs 141 ms at 250
ms, so the flag takes effect.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Both `Control` and `WorkerService` APIs need to access them so moving it
under a shared internal package.
This also aligns ate-apiserver with how atelet uses DV as serves as a
prefactor for upcoming changes to the `ServiceImpl` type.
Fixes#1823
do not merge, not a draft cause i do want ci running
implements:
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching
tls_passhtrough is a followup.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
FinalizeSuspended deleted the external snapshot a suspend replaced
before committing SUSPENDED. By then the new checkpoint is uploaded and
the worker's assignment row is released, so a transient object-store
error (429/503) on that delete aborted the workflow with the actor stuck
in SUSPENDING: ResumeActor rejects that state with FailedPrecondition,
and a retried suspend re-sends Checkpoint to the already-released
worker, which most likely crashes the actor despite a good snapshot in
storage.
Commit SUSPENDED first, then release the replaced snapshot best-effort,
logging a failure instead of returning it. A failed release now leaves
the old snapshot's objects in storage until the actor is deleted, which
removes its whole prefix, rather than stranding the actor. A commit lost
to a concurrent update releases nothing, since the record still names
the old snapshot.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).
~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA
EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.
Related #1591
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
/cc @bowei @EItanya
## Benchamark(atelet, ateom-microvm): log per-actor checkpoint and
restore phase breakdowns
### Why
`SuspendActor` and `ResumeActor` latency is only attributable down to
the `ate.actor.{checkpoint,restore}.duration` histogram phases, and the
biggest buckets — `ateom_checkpoint`, `ateom_restore` — are opaque. When
a large memory benchmark suspend takes 6 s we want to able to tell where
the time is been spent during the suspend.
### What
One joinable, developer-facing log record per operation per layer, with
the full actor identity (allowed in logs, barred from metric labels),
the snapshot scope, and one float-seconds field per phase that ran.
**atelet**
- `Checkpoint` now writes a `Checkpoint timing breakdown` record, the
same way `Restore` has written `Restore timing breakdown` since #1364.
Phases: `sandbox_assets`, `ateom_checkpoint`, `persist`, `total`. One
slice feeds both the histogram and the record, so they cannot disagree.
- Both records carry `error.type` when the operation failed (the gRPC
code; context errors map to `DeadlineExceeded` / `Canceled`). The record
is written on the way out of a failure too, so its completed phases are
kept, and the marker lets a reader exclude a timed-out restore from a
latency distribution. This restores what the record lost when the
`ateerrors` taxonomy was deleted (#1817), using the `error.type`
convention the ateapi instruments already follow.
**ateom-microvm**
- New `phaselog.go`. `CheckpointWorkload` and `RestoreWorkload` emit
records with the same two messages, under
`ateom.actor.checkpoint.duration.<phase>` and
`ateom.actor.restore.duration.<phase>`, decomposing atelet's
`ateom_checkpoint` / `ateom_restore` buckets:
- checkpoint: `pause`, `snapshot`, `durable_dir`, `rootfs_upper`,
`teardown`, `total` — the three captures run concurrently on the paused
guest, so the paused window costs their max, not their sum.
- restore: `prep`, `bundles`, `upper_join`, `lowers`, `tap`,
`vmm_launch`, `vm_restore`, `resume`, `wakeup_probe`, `total` —
sequential; they partition the total. A Data-scope cold boot records
`total` only.
- These timings already existed as ad-hoc `slog.Duration` fields on the
`Actor checkpointed` / `Actor restore phases` lines; those lines are
kept. The record adds stable keys, identity, and seconds (the
histograms' unit).
- The phase names are deliberately private to the binary rather than
added to `internal/ateattr`, so they cannot be mistaken for
`ate.snapshot.phase` metric values. They are micro-VM specific;
ateom-gvisor is unchanged.
**docs/observability.md** is updated: the Restore record is no longer
the only per-actor latency record, and the ateom records are described.
No new instruments, no registry changes, no behavior change.
### Example
```json
{"msg":"Checkpoint timing breakdown","ate.actor.uid":"8f2a…","ate.template.name":"glutton",
"ate.snapshot.scope":"full",
"ateom.actor.checkpoint.duration.pause":0.003,
"ateom.actor.checkpoint.duration.snapshot":0.846,
"ateom.actor.checkpoint.duration.rootfs_upper":0.022,
"ateom.actor.checkpoint.duration.teardown":0.232,
"ateom.actor.checkpoint.duration.total":1.081}
```
Joined with atelet's record for the same actor, a run of the glutton
workload (1 GiB resident, microVM) attributes a 6.0 s p50 suspend as 75%
`persist`, 19% `ateom_checkpoint` (of which the CH `snapshot` is 0.85 s
and `teardown` 0.23 s), and a 5.3 s p50 resume as 79% `download`, 18%
`ateom_restore` (of which `vm_restore` is 0.73 s). The consumer that
produces those tables from pod logs is a separate
`benchmarking/analysis` PR.
### Testing
- `go test ./cmd/atelet/...` and `./cmd/ateom-microvm/...` pass; new
unit tests cover the record shape (seconds, identity keys, zero phases
absent, no duplicate keys), the scope mapping, and `error.type` for
gRPC, context and plain errors.
- `GOOS=linux go vet` clean for both binaries; boilerplate and gofmt
clean.
Part of agent-substrate/substrate#1550.
Adds `kubectl ate delete egress-policy <actor-name> -a <atespace>`
The command removes the actor's policy. It exits 1 when the actor does
not exist
or has no policy. Also add optional `--uid` and `--version` flags, when
specified,
the server refuses the delete when either no longer matches.
### Testing
Tested on a kind cluster with the egress demo
and an API server built from this branch, using a CLI built from this
branch:
- guarded delete with a wrong version, then a wrong uid, then the right
version
- stale version after an `update` bumps the policy
- plain delete on an actor without a policy, and on an actor that does
not exist
🤖 This PR was developed with AI assistance. I have reviewed and tested
all changes.
Follow-up to #1906. Removing a directory entry needs write permission on
the directory that holds it, and with all capabilities dropped uid 0
gets no exemption. A non-root container can create a subdirectory in its
`0777` durable dir; that subdirectory belongs to the container's uid,
typically `0755`, so root cannot empty it. `resetActorDirs` fails after
every checkpoint of such an actor and the suspend never completes. Chmod
first, as the bundle dir does, is no way out: chmod needs ownership or
`CAP_FOWNER`.
Why atelet and not ateom, which already holds the capability: an ateom
that dies before cleanup leaves the files behind and moves the hang to
the next resume on that node. atelet owns the actor directories and
needs it on the crash path either way.
Manifest only, so no unit test. Exercised on kind with the micro-VM
class: uid 65532 writes a durable dir, then suspend and resume. Without
the capability the same run hangs in `SUSPENDING`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
Fixes a part of #1709
This adds declarative validation to atelet's `AteomSupport` RPCs.
Malformed requests are now rejected at atelet instead of being forwarded
to the control plane. Validation runs after authentication.
| RPC | Rules |
|---|---|
| `MintActorCertificate` | actor atespace/name are required short names;
actor UID is a required UUID; the CSR is required and at most 16 KiB |
| `RequestActorSuspend` | actor atespace/name are required short names;
actor UID is a required UUID |
| `SetWorkerCapacity` | `capacity` is required and must follow the
`WorkerResources` rules ateapi declares, shared through
`internal/resources` |
The `AteomHerder` RPCs and `WorkloadSpec` get their validation in a
follow-up PR.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## What
- A pre-built install (`ate-setup deploy --image-repo REPO --image-tag
TAG`) now pins `REPO/envoy-dataplane:TAG` like the ko images, instead of
building it with docker buildx.
- `make build-envoy-dataplane` publishes that image.
- `make build-release-images` publishes every image a pre-built install
needs (ko images, demos, envoy-dataplane), all tagged `$(VERSION)`.
## Why
envoy-dataplane (#1535) is built from a Dockerfile, and ate-setup built
it even for pre-built installs. Those installs have no registry to push
to, so they failed on the default envoy router after the control plane
was applied. It also couldn't join `ALL_IMAGES`, which only lists ko
builds, so nothing published it for a release.
## Testing
- New unit tests for pinning and rendering a pre-built egress manifest.
The rendering test fails on `main` with the production error (`invalid
tag "/envoy-dataplane:build-…"`).
- `make build-release-images` pushed all 13 images under one tag, and
`images.Prebuilt` pinned every one against the real registry.
- Cherry-picks cleanly onto `release-0.2`.
Fixes#1962
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Anna Pendleton <13793293+annapendleton@users.noreply.github.com>
## What
On the microVM restore path, ateom reseeds the guest kernel CRNG with
fresh entropy through the kata-agent's ReseedRandomDev RPC, right after
the guest resumes.
This implements the guest RNG reseed on restore that #1449 calls out as
an unfiled gap (its Strategy 2, inject freshness at restore). Scoped to
the kernel CRNG on the microVM runtime, so it does not close#1449.
## Why
Two actors restored from the same snapshot start with the same frozen
CRNG state. The kernel reseeds from ambient entropy on its own, but not
immediately. Measured on a live restore, two clones kept returning
identical /proc/sys/kernel/random/uuid values for up to about 36ms after
resume before that happened. A workload reading randomness in its first
moments after resume can land in that window and get duplicate values
across clones.
Cloud Hypervisor has no VmGenID device to signal the guest, so ateom
reseeds it directly. On each restore it hands a fresh 32-byte nonce to
the kata-agent, which mixes it into the guest CRNG. The nonce differs
per restore, so clones diverge regardless of the kernel's own timing.
The kata-agent already exposes this RPC, so the change is host-side in
ateom.
## Behavior
Fires on every restore (golden cold-start and resume alike). Cold boot
does not need it. Best-effort: a failure is logged, not fatal, so a
transient agent error never fails an otherwise good restore.
## Scope and limits
- Covers the kernel CRNG only. A workload's own userspace PRNG seed
lives in the checkpointed app memory and stays a workload concern
(Strategy 5 in #1449).
- The reseed runs just after Resume, through the kata-agent, so it lands
a few milliseconds after the vCPUs resume. In the live test the reseeded
window shrank from about 36ms to about 6ms but did not reach zero. The
first read or two after resume can still be frozen before the reseed
lands. Fully closing it needs freezing the workload cgroup across the
reseed, or a VMM VmGenID that acts before the vCPUs resume. A TODO in
the code points at the VmGenID path once Cloud Hypervisor gains the
device. Left as a follow-up.
- microVM runtime only. gVisor reads getrandom and urandom live from the
host with no checkpointed RNG state, so restored gVisor sandboxes
already get fresh entropy.
## Testing
- Unit test on the nonce generator (length, and that two nonces differ).
- Verified live on a GKE microVM cluster with a workload that
continuously records /proc/sys/kernel/random/uuid, so the restore
instant is observable. Two clones from one golden snapshot were
compared. Stock ateom returned identical CRNG output for up to about
36ms (7 consecutive reads) after resume, varying run to run. With this
change the shared window shrank to about 6ms (1 read), then diverged.
---------
Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
Plumbing of actor's EgressPolicy into Envoy dataplane. This is the first
commit that only implements SNI allowlist, without wildcards. Followup
PRs will implement there rest of the policy:
* Wildcard matching.
* SNI passthrough policy.
* Allowing plaintext based on presence of `http` rules.
* Port verification.
* Making sdsmint the default and removing non-sdsmin config.
----
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Yan Avlasov <yavlasov@google.com>
This PR let ateom-microvm reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.
The paths ateom-microvm still builds itself are put into microvm
specific`cmd/ateom-microvm/actorpath.go`.
The `ateompath` package as nothing uses it anymore.
Fix#1604
`run-dev.sh` refused to start unless the Running worker count was at
least `--actors`, on the assumption that each worker hosts one actor. It
does not: a worker advertises capacity for many actors and the scheduler
packs them.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## e2e: add e2e for egress credential injection
Adds the `egresscredinject` e2e suite: an actor fetches
`https://httpbin.org/headers`
through the sdsmint MITM gateway, and the echoed response proves the
injected
`Authorization` header actually reached the upstream.
### What it asserts
- **Injected**: echoed headers contain `Authorization: Bearer <token>`.
- **Overwritten**: an actor-pre-seeded `Authorization` is replaced by
the injected one.
- **Cleartext skip**: plain-HTTP fetch passes through with no
`Authorization`.
- **Fail closed**: nonexistent secret → 403, unserved provider → 500,
unauthorized namespace → 403.
### Additional changes
- Probe `/fetch` returns the response body, accepts
`header=<name>:<value>`
params, and no longer follows redirects (a cross-scheme redirect would
hop
between the cleartext and TLS legs). Covered by unit tests; egressmitm
rerun green.
- New `e2e.EgressInjectHeader` policy helper.
- New `e2e.DeployCredentialProvider` + `fixtures/credinject`: deploys
the real
provider manifest with a test Secret and namespace-policy ConfigMap,
restarts
the provider so it picks the policy up, cleans up on test end.
- CI: new envoy-lane steps after the MITM lanes, gated on
`E2E_EGRESS_CREDINJECT=1`.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Related: #1694
The `swebench-astropy-7336` ActorTemplate can't be deployed as shipped:
1. `image: ""` must be hand-edited into a tracked file before every run.
2. The command points at `/opt/swebench/replay.py` and
`/opt/swebench/astropy_trace.json`, which the sweperf image generator
never
produces. It writes `/replay.py` and `/trace.json` at the image root
(also
the image's ENTRYPOINT), so the container exits immediately.
This PR:
- takes the image from a `SWEPERF_IMAGE` env var, substituted by
`workloads/deploy.sh` like the other template variables;
- fixes both paths;
- documents the minimum sweperf commit (`c30c0d6`), since older images
have no
HTTP server and fail on the first `GET /status`.
Usage:
```sh
SWEPERF_IMAGE=<registry>/sweperf-astropy@sha256:... \
WORKLOAD_TEMPLATES="swebench-astropy-7336" \
benchmarking/workloads/deploy.sh --deploy ...
```
Verified on GKE with both gVisor and microVM worker pools: 30-minute
sweperf
runs at 1 user, all 4 cycles, 0 failures.
- [x] Tests pass (`bash -n`; exercised end-to-end as above)
- [x] Appropriate changes to documentation are included in the PR
(template comments)
This PR let ateom-gvisor reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.
The paths `ateom-gvisor` still builds itself are put into gvisor
specific `cmd/ateom-gvisor/actorpath.go`.
Resetting `runsc-state` and `pidfiles` moves from atelet's
`resetActorDirs` to ateom-gvisor's `main.go` to conduct at beginning of
`RunWorkload` and `RestoreWorkload`, which empties them at the start of
Run and Restore. Atelet doesn't actually read them, so #1730 drops them
from `ActorDirs`.
Added a new `ActorDirs` validation to reject a request with
`InvalidArgument` when `ActorDirs` is missing or any directory in it is
empty, relative or unclean.
The gVisor paths are removed from `internal/ateompath`; the rest goes
with ateom-microvm in the next PR.
Part of #1604.
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:
1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).
The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.
Partial work for #1871
- [x] Tests pass
```
Run make build-junittool
make build-junittool
bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
shell: /usr/bin/bash -e {0}
env:
E2E_ATENET_DATAPLANE: agentgateway
ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
TOTAL
```
- [x] Appropriate changes to documentation are included in the PR
Fixes#1911
`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.
No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Closes the remaining items in #1640: the atelet side of multi-actor
telemetry.
Since #1836 an ateom returns one `GetActiveWorkloadStats` entry per
hosted actor. atelet's sweep already folded every entry, but nothing
pinned the behavior at occupancy above one, and reviewing that fold
surfaced three CPU-accounting defects: two in how baselines survive a
pending sweep, one in the decrease rule.
## What changed
**Tests pin the multi-actor fold.** One worker, several actors:
same-template entries sum under one key, a different template gets its
own, a pending sibling contributes nothing to aggregates or events, CPU
baselines are per actor (independent advance, independent epoch reset),
one probe serves the whole response. Each test was verified to fail
against the corresponding mutation.
**Three CPU-accounting defects found and fixed along the way:**
1. *Undercount.* A pending entry (boot, restore, or a micro-VM guest the
ateom did not reach in its 45s budget) never wrote the actor's baseline,
and `lastCPU` is rebuilt each sweep — so measured → pending → measured
lost the whole interval. On the guest-agent source the counter survives
a restore, so that was real CPU time dropped, and on a full micro-VM
worker it would happen routinely. Pending now carries the baseline
forward; only actors absent from every response are dropped.
2. *Overcount opened by fix 1.* Carrying the baseline unconditionally
meant that during a restore spanning one sweep (source still reports the
actor measured at C1, destination reports it pending), goroutine order
could let the stale C0 overwrite C1, and the next sweep charged the
interval twice. The pending branch now writes only if no worker has
measured the actor this sweep.
3. *Overcount on a guest-agent decrease.* A CPU decrease was charged as
the whole new value. That is right for a cgroup counter, which restarts
at zero on restore. It is not always right for a guest-agent counter: a
FULL or DATA_ON_GOLDEN restore resumes a guest at that guest's counter
(after a golden restore, main billed the golden guest's CPU to the
actor), a cold boot restarts it at zero, and a container exit lowers the
sum. A sample cannot tell these apart, so a guest-agent decrease now
charges nothing and re-baselines. A cgroup decrease still charges the
new value, with or without a pending sweep between. **Trade-off, new
versus main:** a guest-agent decrease loses the usage from the event to
the first sample. Main counted it correctly only after a cold boot (a
DATA restore, or a Run after a crash); in every other case (golden
restore, FULL restore from an older snapshot, container exit) main
overcounted it.
Remaining limit, documented in the proto: an epoch that starts *above*
the last reported value cannot be detected from consecutive samples.
Main has this for guest-agent restores between sweeps; carrying the
baseline extends it to restores that span a pending sweep. The overcount
is bounded by the golden guest's counter minus the actor's last value,
so it needs an actor that used less CPU than the golden guest had at
snapshot.
The test for (2) forces both fold orders deterministically: the sweep
closes a probe's connection right after folding its response, so the
fixture's closer releases the gated probe from `Close` — no sleeps, no
polling.
**Timeout rationale reconciled.** The poll floor (50s) is one actor's
worst-case micro-VM read: 25 containers at 2s each. It never prevented
overlap, because sweeps are sequential; it bounds guest-agent load, and
the comment now says so. The RPC timeout (55s) was documented as
covering that single-actor worst case. With several actors per worker,
the micro-VM ateom caps its own read at `statsSweepBudget` (45s, chosen
to stay under atelet's 55s) and reports unreached guests as pending, so
the timeout stays a constant; its comment now points at that budget.
## Not in scope, recorded
- Two workers both *measuring* one actor in one sweep double-charges.
Pre-existing and unchanged here. It is rare: it needs the same probe
skew as defect 2, with the destination read after its restore finished.
- `ateom.proto` still says the discovery list holds "at most one entry
until multi-actor workers land". Stale; fixed in a separate change.
## Verification
`gofmt`, `go vet`, and race-enabled tests pass. Mutation checks:
dropping the carry, the measured-wins guard, or either side of the
per-source decrease rule each fails its test.
**BREAKING** API changes prior in preparation for GA.
~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow
- `EgressRule` is now a union of protocols: `http`, `https`,
`tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
`tls_passthrough`. Same wildcard grammar as before, documented once on
`HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
once on `HTTPRule.ports`.
- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
placeholder value. Requests without the header pass unchanged.
- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.
- `google/protobuf/empty.proto` import dropped; nothing uses it now.
Questions
- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
block is intercepted. is that clear enough?
follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
the rule conflict check) and regenerate `zz_generated.validation.go`
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes: #823 (there are probably more issues that i need to find)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.
This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.
Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.
Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.
- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Refactors the atelet Restore API so the DATA_ON_GOLDEN base snapshot
travels as a typed restore source instead of a bare string, preparing
for the node-local snapshot cache (#690, #1551): `RestoreRequest` gets
its own external snapshot *source* message, and read-side attributes of
a restore source have a home — today the URI, next (in the M2 cache
work) a sharing property set by the control plane.
**This change is not backward compatible.** `golden_snapshot_uri` is
removed from `RestoreRequest` with no transition: an ateapi from before
this change sends a DATA_ON_GOLDEN restore an updated atelet rejects (no
`base_config`), and an updated ateapi sends one an old atelet rejects
(no `golden_snapshot_uri`). ateapi and atelet must be rolled together.
Fresh starts, FULL/DATA restores of an actor's own snapshot, and local
pause restores without a golden base are unaffected.
Three commits, each buildable and tested:
1. **`ateletpb: give Restore its own external snapshot source message`**
— `ExternalRestoreConfiguration` replaces
`ExternalCheckpointConfiguration` in `RestoreRequest`'s config oneof
(same wire shape; the checkpoint-side write-destination message is
untouched), and `base_config` is added for the DATA_ON_GOLDEN base. No
behavior change: nothing sets or reads the new field yet.
2. **`ateapi, atelet: carry the restore base snapshot in base_config`**
— the cutover: ateapi sets `base_config` on both golden-data resume
paths (external data snapshot, and local pause checkpoint combining with
the golden); atelet reads and validates only `base_config`. The
characteristic test of the ateapi→atelet request and the golden-data
functional test pin the new field.
3. **`ateletpb: remove golden_snapshot_uri from RestoreRequest`** — the
old field is deleted and `base_config` takes its field number, keeping
the numbering dense.
Tested: full ateapi + atelet suites including the controlapi functional
tests; gofmt, golangci-lint, boilerplate clean.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
While benchmarking pause actor I ran into multiple bottlenecks.
I'm currently importing
https://github.com/agent-substrate/substrate/pull/935 into my branch but
https://github.com/agent-substrate/substrate/pull/1876 is another
alternative to solve the first bottleneck.
The 2nd bottleneck (primary target of this PR) is due to fsyncs being
called in writeFileAtomic() for sandbox-assets.json and SystemInfoVolume
files. Because these are new files that require ext4 filesystem metadata
updates, and we have ext4 mounted in ordered mode, calling fsync on the
file + directory causes all other pending writes (such as other
checkpoints) to be flushed. So the file creation ends up being blocked
behind other actor's checkpoint syncs.
These files do not actually need to be flushed to disk. SystemInfo is
recreatable. sandbox-assets.json is written during Run/Resume(), and
read during Checkpoint(). But once the checkpoint is finished, the
manifest is saved with the snapshot. If the node crashes between
Run/Resume and Checkpoint, the actor is crashed anyway.
### Single-Node Pause/Resume Benchmark Summary
**Configuration:** 17 workers, 15 actors, `glutton` (`512 MiB` RAM
capacity, `32 MiB` churn per cycle), `120s` run, `1.0s` wait, `100%`
pause (`0%` suspend)
| Metric | Baseline (`upstream/main`) | + Commit 1 (`os.Link` local
checkpoint) | + Commit 2 (remove `fsync` on ephemeral volumes) |
| :--- | :---: | :---: | :---: |
| **Completed Cycles** (`120s`) | `106.5` (`112` pauses / `101` resumes)
| `160.0` (`165` pauses / `155` resumes) | **`212.5`** (`214` pauses /
`211` resumes) |
| **Cycle Throughput** | `53.3 cycles/min` (`0.89/s`) | `80.0
cycles/min` (`1.33/s`, **+50.2%**) | **`106.3 cycles/min`** (`1.77/s`,
**2.00x**) |
| **`ResumeActor` p50** | `6,000 ms` | `4,100 ms` | **`210 ms`**
(**28.6x faster**) |
| **`ResumeActor` avg** | `7,023 ms` | `4,867 ms` | **`246 ms`**
(**28.5x faster**) |
| **`ResumeActor` p90 / p99** | `15,000 ms` / `19,000 ms` | `9,400 ms` /
`20,000 ms` | **`350 ms` / `590 ms`** |
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Jefftree <jeffrey.ying86@live.com>
The purpose field in MintActorCertificate (and the certs it issues) is
no longer needed. The distinction between actors and ateoms-for-actors
is now built into the SPIFFE URI.
Stacked over #1809
Add a Star History section to the README that embeds star-history.com's
live chart, with a dark variant for GitHub's dark theme.
star-history.com renders the image on request, so it stays current;
browsers and GitHub's image proxy cache it for up to a day.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Today `atelet` and `ateom` derive the per-actor directories (oci
bundles, runsc state, pid files, checkpoint and restore state,
durable-dir, system-info and volume roots) from the actor UID through
the same package, `internal/ateompath`.
This PR:
1. Add an `ActorDirs` message to RunWorkloadRequest,
RestoreWorkloadRequest, CheckpointWorkloadRequest and
TerminateWorkloadRequest, and have atelet fill it with the directories
it prepared. **The ateoms do not read it yet.**
2. Split `internal/ateompath` into:
- `cmd/atelet/internal/ateletpath`: atelet's pathes, the per-actor
directories (and `ActorDirs`, built from them) plus the directories only
atelet uses.
- `internal/nodepath`: shared pathes, the base dir both mount, the ateom
socket, the OTLP
sockets, and the netns name.
- `internal/ateompath`: what the ateoms still derive from the actor UID,
each function marked with the `ActorDirs` field it duplicates. atelet no
longer imports it.
---------
This is the first part of #1604 to decouple atelet and ateom shared
pathes. Behavior **does not change**: both sides still compute the same
paths, atelet from `ateletpath` and the ateoms from `ateompath`.
Followup PRs will remove ateompath by making `ateom-gvisor` and
`ateom-microvm` read the `ActorDirs` message from RPC.
Step 3 of #1748. Vocabulary only: nothing emits the event or sets the
epoch yet.
**Event.** `ate.actor.usage_sampled` in `actorevent`, sixteen required
keys via `ateattr`: identity, pool, sandbox class, source,
`ate.stats.kind` (periodic, initial, final), the four measurements,
observation time, epoch. Measurement names follow the
`ate.actor.stats.*` instruments; units are in the registry briefs. Body
is `Actor usage sampled`, the accepted stdout shape change.
**Scope.** The usage stream has its own instrumentation scope, the
lifecycle scope with `/usage` appended, so a consumer can sample or drop
it without touching lifecycle records. `NewEmitter` takes the scope; the
package-level `Log` keeps the lifecycle one, so ateapi is unchanged.
**Epoch.** `ate.actor.epoch` and `WorkloadStatsSample.epoch_unix_nano`:
the unix-nano time an activation began. CPU time is per activation for
every source; lifetime is the sum over epochs of each epoch's highest
value. Memory usage and working set are absolute; peak is as the source
reports it. An ateom that leaves the epoch at zero still sends the raw
guest counter. The measurement comments on fields 8 to 11 say all of
this.
**Registry.** Event group in `events.yaml`, seven attributes in
`metrics.yaml`, `ate.actor.epoch` in the no-actor-identity forbidden
list. The one-record-one-place rule and the collector guide now name
both scopes.
Left for step 4: the event's code anchor moves to the ateom emit site,
atelet's poller comments and the guide's stdout usage section get
rewritten, and the controller starts passing `OTEL_LOGS_EXPORTER` to
worker pods.
Part of #1756.
- Adds `--actor-jwt-issuer`. Unset, it defaults to
`https://idp.<namespace>.svc`, the Service ate-idp-server will serve
discovery from. Before this, `iss` was hardcoded to
`https://api.ate-system.svc`, which points at ateapi's gRPC port.
- Adds `oidcdiscovery.ParseIssuer`, which requires a canonical https URL
with no query, fragment, or user info and strips trailing slashes.
ate-idp-server will use it too, so both emit the same bytes.
- Actor JWT headers now include `typ: JWT`.
- Drops the `actorIdentityJWTIssuer` plumbing into `RPCService`, unused
since #1315.
The default `iss` changes, but `MintActorJWT` has no production caller
yet.
Testing: new unit tests for the parser, the default, the header, and
flag resolution. `TestMintActorJWT_Success` now checks `typ`, `iss`, and
`sub`. It needs Docker, so it hasn't run locally.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (docs
land later in the #1756 series)
This adds dual v1beta1 and v1 PCR support, favoring v1. We have seen
users on v1.37 without v1beta1 but with v1 seeing incomaptibility.
Given our usage of PCR is pretty scoped its not too bad to have the dual
version support.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
Fixes#1835
`atelet` chose the sandbox binaries and pause image for a restore from
the snapshot's `manifest.json` rather than from the control plane. This
PR makes ateapi send them on every restore, the way it already does for
Run.
**Changes**
- **proto:** add `sandbox_assets` to `atelet.RestoreRequest`.
- **ateapi:** `ensureAteletRestored` resolves the ActorTemplate's
SandboxConfig once and sends it on the local restore, the external
restore and Run. A missing SandboxConfig now fails before any atelet
call.
- **atelet:** `sandbox_assets` is required on Restore. Restore selects
the binaries and pause image with `recordFromRequest`, the same path Run
uses. The manifest still supplies the snapshot files, sandbox class and
metrics labels.
Note - The field is required, so after atelet is upgraded, resumes on
that node fail until ate-api-server is upgraded too.
**Testing**
- ateapi: every Run and Restore sends the resolved assets; a missing
SandboxConfig fails on both restore paths.
- atelet: `TestRecordFromRequest`, a validation case for a missing
field, and `TestRestoreUsesRequestSandboxAssets` (checkpoint with pause
image v1, restore with v2, and the actor runs with v2).
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.
v0.1.0 used a form of this, checking it in before v0.2.0
Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
Fixes the kind install on arm64 hosts (Apple Silicon), broken in #1535
- `BuildDockerfileImage` hard-codes `--platform=linux/amd64`, so the
dataplane image's Rust builder stage fails with `exec /bin/sh: exec
format
error` on mac with arm64
The ko images already follow `KO_DEFAULTPLATFORMS`, which the kind
install
sets to the host architecture; the Dockerfile images did not.
**Changes**
- Pass `KO_DEFAULTPLATFORMS` through to `docker buildx build --platform`
for
both Dockerfile images (envoy dataplane, claude-code-multiplex demo),
defaulting to `linux/amd64` when unset.
- Pin the Envoy base image by its multi-platform index digest so
`--platform`
selects the matching manifest.
- [x] Tested (`hack/install-ate-kind.sh --deploy-ate-system` on mac
succeed)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.
It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#1266
Builds on #1689#1283 in particular.
Outstanding:
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~
EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR