The passthrough chain from #2019 dialed the ORIGINAL_DST filter state,
which the CONNECT leg filled with the address the actor connected to.
The SNI picked the chain, the actor's own resolution picked the
destination, so a passthrough rule for one name let an actor reach any
IP by claiming that name in the ClientHello.
We fix it by:
- The passthrough chain is now `sni_dynamic_forward_proxy` then
`tcp_proxy` to a new raw forward-proxy cluster,
`egress_forward_proxy_passthrough`, on the shared egress_dns_cache. The
gateway resolves the SNI itself and sends the bytes to it.
- The CONNECT leg answers the dialed port under
dev.ate.egress:dialed_port; the outer chain copies it into
envoy.upstream.dynamic_port, shared with the inner listener, which both
the SNI filter and the cluster read before their configured port.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Follow up on #2007
This switches the remaining two `WorkerService` RPCs
(`RequestActorSuspend` and `SetWorkerCapacity`) to declarative
validation via `cmd/ateapi/internal/apivalidation`, completing
declarative request validation across `ateapi`.
Removes the hand-written checks in `workerservice` (`validateActorRef`,
`validateReportedCapacity`) and the now-unused
`resources.ValidateGlobalObjectRef`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## What this is
Follow-up to #1934, as discussed there: the agent-session script moves
out of Go and into YAML, so a workload can be read, edited, and selected
without touching the driver.
- The default script is now
`internal/benchmarking/boomer/agentsession/scripts/coding-session.yaml`,
embedded in the boomer worker. It was generated from the old Go table
and round-trips byte-for-byte on every op.
- `--agentsession-script <name>` picks a built-in variant (any
`scripts/<name>.yaml`). Resolved on the worker's first iteration (the
knob arrives with the first spawn message, after Init); a bad name is
logged and retried each iteration, so fixing the knob and re-swarming
recovers.
- Decoding is strict and runs the same invariants the old tests
enforced: unknown op kinds or fields, arguments a kind does not take,
reads before writes, walks before fills, duplicate steps, non-positive
think times.
- Each script declares `min_actor_memory`. The worker reads the glutton
template's memory limit at start and refuses to run against a smaller
actor, with a message naming `--actor-memory`. That replaces the
README's advice with a guard.
- README gains a "Writing an agent-session script" section with the
format and the op table.
## Not in this PR
Loading a script from a file on the worker (ConfigMap mount) is the next
PR, on top of this loader.
## Testing
`go test -race` on the benchmarking packages, boilerplate, gofmt, and
golangci-lint pass. New tests: every embedded script loads; default
script shape (20 steps, 1Gi floor, declared bytes under budget);
encode/decode round trip; 14 rejection cases; the template memory guard
in refuse, exact, and no-limit modes.
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Part of #1590
## What this PR does
Locust measures the client side only. This adds a post-run harvest of
server-side ground truth from Prometheus, written to a new
`server_summary.json` and summarized as one row in `stats.jsonl`.
## Proposed Changes
### Steady-state scoping
`server_telemetry.py` derives the steady-state window from
`stats_history.csv`: it starts at the first sample at 90% of peak user
count and ends at the last one. Ramp-up is excluded from every metric,
and the snapshot block also stops at the last full-load sample, so
teardown suspends are left out. Under a step-ladder load shape the
window covers only the top step.
### Harvested metrics
**Cluster packing**, from `ate_workerpool_workers`: busy workers
(`partial` + `at_capacity`) over the pool size at each sample, where the
pool size is the sum of all worker states, so it follows scale-ups and
scale-downs mid-run. Only the live ateapi is read: after a redeploy the
collector keeps re-exporting exited ateapi processes' last values for a
few minutes, which would otherwise inflate the counts. Reported as min,
p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries
so transient spikes remain visible.
**Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall
percentages for the node and for the worker pods, each as the same
percentile set over the window.
**Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both
in-window and cumulative, restore and checkpoint latency mean and
p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint
throughput as `checkpoint_mb_s`. The atelet exports metrics on an
interval, so its data reaches Prometheus late: the harvest waits
`--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default
export and a 10s scrape) and reads both window edges half that late.
### Constraints and failure behavior
The module uses only the standard library, since the locust image is
distroless, and every request carries a timeout. An unreachable
Prometheus records nulls and does not fail the run. `--prometheus-url`
overrides the in-cluster default, and `--atelet-lag-s` sets the wait for
the atelet's last export.
Unmeasured fields are `null` and a measured zero is `0`, consistent with
the rest of the runner. Prometheus exposes no byte counter on the
restore path, so no restore throughput field is emitted rather than
deriving one indirectly.
### Output
`server_summary.json` holds the full nested artifact. `stats.jsonl`
receives a single `server_summary` row with 45 flat keys for graphing:
packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot
means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`.
`status.json` is unchanged.
## How this was tested
14 unit tests in `test_server_telemetry.py` covering steady-state
detection, percentile boundaries, the range-query window guard,
malformed Prometheus responses, the packing and checkpoint arithmetic,
the per-sample pool size, ignoring exited ateapi series, a missing
denominator returning null, snapshot fields returning null rather than
zero, and telemetry surviving a missing stats CSV.
Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45
keys matched an independent recomputation from raw Prometheus, and the
snapshot block was cross-checked against atelet logs. A run with
`--prometheus-url` pointed at an unreachable address completes normally
with the affected fields null.
Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8
workers) from Prometheus with the new code. Runs with a steady pool
match the previous output exactly, and runs that followed a redeploy now
read the real pool size on every sample.
## References
[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1472.
The Telemetry API rejects an exponential histogram point with no
positive buckets. Give such points one zero-count bucket instead.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eric Curtin <eric.curtin@docker.com>
## What
Adds an optional `secretSelector` field so a credential-provider
namespace grant can be narrowed by label to specific Secrets instead of
the whole namespace:
```yaml
policies:
- atespace: team-a
allowedNamespaces: [team-a]
secretSelector: # new
matchLabels:
ate.dev/credential: "true"
```
No selector set → unchanged behavior (full-namespace grant).
## Why
A namespace-only grant lets any actor read *any* Secret in that
namespace, including Substrate's own CA/JWT/DB Secrets when they share a
namespace with actors (e.g. kagent as a Substrate subchart). RBAC
`resourceNames` can't scope this per-atespace, since the provider is a
single ServiceAccount. The policy is the only place with that context.
## Behavior
- Once a namespace policy is enforced, a missing Secret and one whose
labels don't match both return `PermissionDenied` — so a caller can't
probe what exists in the namespace.
- Malformed policies (bad label key/value, unknown YAML fields) fail to
load rather than silently allowing nothing.
## Tests
New unit tests cover label narrowing, the no-probing property, and
malformed-policy rejection. Existing tests pass unchanged. `go build
./...` and lint/verify scripts pass.
---------
Signed-off-by: Jonathan Jamroga <jonathan.jamroga@solo.io>
logFlagValues writes every flag to the log when the server starts, and
that included --postgres-connection-string verbatim. The bundled
database and Cloud SQL IAM auth use passwordless strings, but an
external database DSN can carry a password, and it was written to the
log on every restart.
Log the parsed, non-secret parts instead: host, port, database, user,
whether a password is set and whether TLS is on. pgconn.ParseConfig
handles both the URI and the libpq key=value forms, so no string
rewriting is needed and nothing is echoed. A string that does not parse
is logged as "<unparseable>"; connectStore still reports the real error.
Fixes#1832
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#687
Both ateoms build atunnel through the new `internal/ateomtunnel`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
We should not override the whole actor trust store with SSL_CERT_DIR,
removing that to defer using SSL_CERT_FILE, which is additive.
Additionally, this does not cover all languages and runtimes, adding
guidance for other env vars required and other runtimes that need union
of the trust stores in a single file.
In a followup, we are going to remove all "sdsmint" and "plain gateway"
language and just have one gateway thats capable of minting or not based
on the sni.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This is needed because DeleteActor cleans up all snapshots under the
actor's location.
If we were to allow repointing an actor's snapshot location, we risked
leaking historical snapshots upon actor deletion.
Implement passthrough egress policy.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Yan Avlasov <yavlasov@google.com>
Part of #1827
Moves `cmd/atelet/internal/ategcs` to `pkg/objectstorage` and renames
the package to `objectstorage`. No code changes besides the package
name, its import path and uses, the tracer name, and comments that named
the old package.
**Why:** the [Platform-Provided Snapshot API
Proposal](https://docs.google.com/document/d/1ac-hY6zsmX4FmtGOt9CXbNFHeluGouyYluer7kX6bto)
moves snapshot transfer out of atelet and ate-api-server into a snapshot
plugin binary that runs as a sidecar. The reference plugin (GCS and S3)
reuses this package's transfer code as is, and so can snapshot plugins
built outside this repository, which is why it goes in `pkg/` rather
than `internal/`.
**Naming:** `objectstorage`, not `objectstore`, because
`internal/objectstore` already exists. That package only lists, copies
and deletes objects for the control plane and is deliberately separate
from this one (see its package comment, updated here). The snapshot
plugin binary imports both.
This PR stands alone and changes no behavior.
**Reviewing:** `git show -M --stat` shows 20 renames as
`{cmd/atelet/internal/ategcs => pkg/objectstorage}/...`. Each changes
only its `package` clause, except:
- `objects.go`: tracer name `ategcs` → `objectstorage`
- `gcscompose_test.go`, `s3bench_test.go`: the `go test
./pkg/objectstorage` path in a comment; `gcscompose_test.go` also
renames its emulator test bucket
The other 4 files:
- `cmd/atelet/{main,main_test,sandbox_assets}.go`: import path and
`ategcs.` → `objectstorage.`, plus one comment in `main_test.go`
- `internal/objectstore/objectstore.go`: package comment now points at
`pkg/objectstorage`
- [x] Tests pass (`make test`, `make verify`)
- [x] Appropriate changes to documentation are included in the PR (none
needed)
Part of #1660.
- Adds `CredentialHeader.actor_jwt` (`ActorJWTSource`), a union member
alongside `credential_uri`, so a `replace_headers` entry can carry a
Substrate-issued actor JWT instead of a provider secret.
- `ActorJWTSource` takes the token's `audiences` and a required
`expiration_seconds` in [300, 3600], matching `MintActorJWTRequest`
(#1902).
- Until the gateway can mint tokens, it denies an `https` rule with an
`actor_jwt` entry with 501. Like `credential_uri` entries, `actor_jwt`
entries are skipped on `http` rules.
Testing: unit tests for validation and the gateway's 501 on `https` and
skip on `http`. The functional tests need Docker, so they haven't run
locally.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (the
proto comments are the API docs; the user guide lands later in the #1660
series)
## What this is
We keep saying Substrate's sweet spot is "agents that are idle most of
the time" — but none of our benchmarks actually behave like one. Glutton
hammers one resource at a time, and sweperf needs an external image with
a replay trace. This adds a workload that acts like the thing we're
building for: a coding agent working through a task.
Each locust user is one session. The actor gets told to do the kind of
things a coding agent does — clone a repo, install deps, build, hit a
failing test, fix it, write tests, refactor, package — twenty steps,
each one costing the sandbox the CPU, memory, disk, and network the real
action would. Between steps the "LLM is thinking," so the driver
suspends the actor, and the next step's first request wakes it back up
through the router. That parked wake (`WakeFirstTouch` in the stats) is
the number this whole benchmark exists to measure.
The part I care most about: the entire workload is one table in
`internal/benchmarking/boomer/agentsession/script.go`. Every step says
in plain English what the agent is doing and what it costs. If you want
to know what step 6 does to the sandbox, you read step 6. If you want a
different workload, you edit the table — tests will catch you if you
write a step that reads a file nothing wrote, or blow the actor's memory
budget.
To act the steps out, glutton grew two RPCs: `BurnCPU` (compute-bound
work) and `Ingest` (bytes that actually cross the network before hitting
disk, so a "git clone" is a real download, not a local write).
## How it went when we ran it
Validated on a fresh 2-node GKE cluster, micro-VM first, then gVisor on
the same hardware. Smoke runs were clean on both classes (gVisor: 1340
requests over 7 full laps, zero failures, wake p50 1.4s; micro-VM: wake
p50 2.1s). At 20 concurrent sessions with realistic think times, gVisor
held 2.6% failures with wake p50 1.3s. Pushing past the knee (~7
concurrently-active sessions on 8 vCPUs) was also useful: it reproduced
the ateom-socket-vanishing failure from #1133 and left six actors
permanently wedged in DELETING — a live repro of #1665.
Two things the first live run taught us are already folded in: RAM
refills are in-place so repeat laps don't transiently double the guest
heap (512Mi micro-VM actors OOM'd without this — use 1Gi), and the
driver replaces an actor after three failed steps in a row, because a
CRASHED actor never comes back on its own.
## Future changes
Right now the script is compiled in — changing what steps do means
editing the table and rebuilding the image. That's deliberate for this
PR (one reviewable, test-guarded source of truth), but the follow-up
we've agreed on is to make the script runtime-configurable: named script
variants selectable per run first, then accepting a full script as a
file so operators can define workloads without touching Go. That lands
as its own PR once this one is in.
---------
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Part of #1401.
A listing fails on one undecodable row. The error now names it.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1882
When a worker pod's ateom container restarts, every sandbox it was
hosting is lost, but the control plane kept reporting those Actors as
running.
This change:
- **Tracks ateom restarts:** each Worker gets an `epoch`, set by the
syncer from the ateom container's restart count. It can only increase.
- **Stamps each binding:** every Actor assignment records the Worker's
epoch at bind time.
- **Crashes stale Actors:** a new reconciler in ateapi watches for
Workers whose epoch has risen past `status.observed_epoch`. It crashes
the Actors bound during earlier epochs, releases their assignments, then
advances `observed_epoch`.
[ x ] Tests pass
[ x ] Appropriate changes to documentation are included in the PR
Adds a credential provider for egress credential injection backed by
Google Cloud Secret Manager.
**It lives in its own Go module under `plugins/gcp-secret-manager`,
temporally hosted here until it moves to a repository of its own.**
**What it does**
- Serves `credproviderpb.CredentialProvider` over mTLS and admits only
the egress gateway's identity (`--injector-identity`).
- Resolves global and regional secrets, optionally picking one key out
of a JSON payload:
`ate-secret://secretmanager.googleapis.com/projects/<project>[/locations/<location>]/secrets/<secret>/versions/<version>[/keys/<key>]`
- Enforces a default-deny atespace→project policy
(`--project-policy-file`), the counterpart of the Kubernetes provider's
namespace policy.
- Returns a retryable 503 only for transient Secret Manager failures,
and caps each read with `--fetch-timeout` (default 3s).
**Repository changes**
- New top-level `plugins/` directory for self-contained plugins,
documented in `docs/dev/code-layout.md` and `AGENTS.md`. The module
imports only substrate's public `pkg/` packages; a test enforces this.
- CI runs the module's tests, `make verify` and golangci-lint.
govulncheck scans the module too, with the action pinned by SHA.
- `docs/egress-credential-injection.md` describes each provider's
credential URI format. The plugin's README covers installing and using
it.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Part of #1756.
- Adds a required `MintActorJWTRequest.expiration_seconds`, validated to
[300, 3600]. Token exchange (#1661) wants longer-lived subject tokens.
- Adds `MintActorJWTResponse.expires_at`, equal to the `exp` claim, so
callers like the egress gateway's cache don't have to decode the token.
- Changes the subject to `actor/<atespace>/<name>`, matching the path of
the actor's SPIFFE ID. The token isn't a JWT-SVID.
- Rewrites the `MintActorJWTResponse` comment to match the settled
contract.
Field 8 skips 2, 3, 4, and 6, which the old ActorIdentity request used.
Testing: validation unit tests for the bounds. The functional test
checks the lifetime, `expires_at`, and `sub`; it needs Docker, so it
hasn't run locally.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR (docs
land later in the #1756 series)
They are useful for granular upgrade
They will not be used for install as we have `deploy ate-system`, they
are used for more granular steps in upgrade runbook.
Part of #1686
# benchmarking: actor telemetry for the sweperf workload
## What this does
Adds the suspend/resume and task-execution timings from the telemetry
proposal to the sweperf boomer workload.
**New locust rows**
| row | what it measures |
|---|---|
| `ResumeToFirstExec` | resume RPC start until the sandbox accepts the
cycle's `/execute` |
| `TaskCEL` | in-container execution time for one sweperf task (4
cycles), as reported by `replay.py` |
| `CycleCEL` | in-container execution time for one cycle |
| `TaskWallClock` | client wall clock for one task, excluding think time
|
| `<rpc>_rtt` | client round trip for each control-plane RPC |
The four derived rows are recorded under the `actor` method, the `_rtt`
rows under `grpc`.
We keep both the client RTT and the server-side elapsed for every
control-plane RPC so the network and queueing overhead stays visible
separately. The existing `ResumeActor` / `SuspendActor` rows keep the
server-side elapsed from the response trailer.
The liveness check at session start leaves the actor running, so the
first cycle's resume is a no-op. Its successful `ResumeActor`,
`ResumeActor_rtt` and `ResumeToFirstExec` samples are skipped (failures
are still recorded), the same way ateapi skips no-op resumes.
## Poll interval
The `/status` poll interval is now configurable with
`--sweperf-poll-interval-ms` (env `LOCUST_SWEPERF_POLL_INTERVAL_MS`),
default 100 ms (was a fixed 25 ms). It flows through dynconfig like the
other `--sweperf-*` flags and is read per job, so a mid-run change
applies from the next cycle.
## Documentation
`benchmarking/README.md` had no sweperf section, so this adds one
listing every row the workload emits and the poll interval flag,
matching the existing DurDir section.
## Testing
`go test -race ./internal/benchmarking/boomer/...`, plus runs on a real
cluster (sympy, 4 users / 2 workers, gVisor): 5 min at the default 100
ms poll interval and 2 min at 250 ms.
- Both had 0 failures and all expected rows.
- The min `ResumeActor` is 334 ms, so the first-cycle no-op is no longer
recorded.
- `ResumeActor` / `SuspendActor` counts match (197 / 197), and each
`_rtt` is within ~1 ms of its server-side row.
- Client overhead over CEL per cycle is 65 ms at 100 ms vs 141 ms at 250
ms, so the flag takes effect.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Both `Control` and `WorkerService` APIs need to access them so moving it
under a shared internal package.
This also aligns ate-apiserver with how atelet uses DV as serves as a
prefactor for upcoming changes to the `ServiceImpl` type.
Fixes#1823
do not merge, not a draft cause i do want ci running
implements:
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching
tls_passhtrough is a followup.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
FinalizeSuspended deleted the external snapshot a suspend replaced
before committing SUSPENDED. By then the new checkpoint is uploaded and
the worker's assignment row is released, so a transient object-store
error (429/503) on that delete aborted the workflow with the actor stuck
in SUSPENDING: ResumeActor rejects that state with FailedPrecondition,
and a retried suspend re-sends Checkpoint to the already-released
worker, which most likely crashes the actor despite a good snapshot in
storage.
Commit SUSPENDED first, then release the replaced snapshot best-effort,
logging a failure instead of returning it. A failed release now leaves
the old snapshot's objects in storage until the actor is deleted, which
removes its whole prefix, rather than stranding the actor. A commit lost
to a concurrent update releases nothing, since the record still names
the old snapshot.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).
~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA
EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.
Related #1591
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
/cc @bowei @EItanya
## Benchamark(atelet, ateom-microvm): log per-actor checkpoint and
restore phase breakdowns
### Why
`SuspendActor` and `ResumeActor` latency is only attributable down to
the `ate.actor.{checkpoint,restore}.duration` histogram phases, and the
biggest buckets — `ateom_checkpoint`, `ateom_restore` — are opaque. When
a large memory benchmark suspend takes 6 s we want to able to tell where
the time is been spent during the suspend.
### What
One joinable, developer-facing log record per operation per layer, with
the full actor identity (allowed in logs, barred from metric labels),
the snapshot scope, and one float-seconds field per phase that ran.
**atelet**
- `Checkpoint` now writes a `Checkpoint timing breakdown` record, the
same way `Restore` has written `Restore timing breakdown` since #1364.
Phases: `sandbox_assets`, `ateom_checkpoint`, `persist`, `total`. One
slice feeds both the histogram and the record, so they cannot disagree.
- Both records carry `error.type` when the operation failed (the gRPC
code; context errors map to `DeadlineExceeded` / `Canceled`). The record
is written on the way out of a failure too, so its completed phases are
kept, and the marker lets a reader exclude a timed-out restore from a
latency distribution. This restores what the record lost when the
`ateerrors` taxonomy was deleted (#1817), using the `error.type`
convention the ateapi instruments already follow.
**ateom-microvm**
- New `phaselog.go`. `CheckpointWorkload` and `RestoreWorkload` emit
records with the same two messages, under
`ateom.actor.checkpoint.duration.<phase>` and
`ateom.actor.restore.duration.<phase>`, decomposing atelet's
`ateom_checkpoint` / `ateom_restore` buckets:
- checkpoint: `pause`, `snapshot`, `durable_dir`, `rootfs_upper`,
`teardown`, `total` — the three captures run concurrently on the paused
guest, so the paused window costs their max, not their sum.
- restore: `prep`, `bundles`, `upper_join`, `lowers`, `tap`,
`vmm_launch`, `vm_restore`, `resume`, `wakeup_probe`, `total` —
sequential; they partition the total. A Data-scope cold boot records
`total` only.
- These timings already existed as ad-hoc `slog.Duration` fields on the
`Actor checkpointed` / `Actor restore phases` lines; those lines are
kept. The record adds stable keys, identity, and seconds (the
histograms' unit).
- The phase names are deliberately private to the binary rather than
added to `internal/ateattr`, so they cannot be mistaken for
`ate.snapshot.phase` metric values. They are micro-VM specific;
ateom-gvisor is unchanged.
**docs/observability.md** is updated: the Restore record is no longer
the only per-actor latency record, and the ateom records are described.
No new instruments, no registry changes, no behavior change.
### Example
```json
{"msg":"Checkpoint timing breakdown","ate.actor.uid":"8f2a…","ate.template.name":"glutton",
"ate.snapshot.scope":"full",
"ateom.actor.checkpoint.duration.pause":0.003,
"ateom.actor.checkpoint.duration.snapshot":0.846,
"ateom.actor.checkpoint.duration.rootfs_upper":0.022,
"ateom.actor.checkpoint.duration.teardown":0.232,
"ateom.actor.checkpoint.duration.total":1.081}
```
Joined with atelet's record for the same actor, a run of the glutton
workload (1 GiB resident, microVM) attributes a 6.0 s p50 suspend as 75%
`persist`, 19% `ateom_checkpoint` (of which the CH `snapshot` is 0.85 s
and `teardown` 0.23 s), and a 5.3 s p50 resume as 79% `download`, 18%
`ateom_restore` (of which `vm_restore` is 0.73 s). The consumer that
produces those tables from pod logs is a separate
`benchmarking/analysis` PR.
### Testing
- `go test ./cmd/atelet/...` and `./cmd/ateom-microvm/...` pass; new
unit tests cover the record shape (seconds, identity keys, zero phases
absent, no duplicate keys), the scope mapping, and `error.type` for
gRPC, context and plain errors.
- `GOOS=linux go vet` clean for both binaries; boilerplate and gofmt
clean.
Part of agent-substrate/substrate#1550.
Adds `kubectl ate delete egress-policy <actor-name> -a <atespace>`
The command removes the actor's policy. It exits 1 when the actor does
not exist
or has no policy. Also add optional `--uid` and `--version` flags, when
specified,
the server refuses the delete when either no longer matches.
### Testing
Tested on a kind cluster with the egress demo
and an API server built from this branch, using a CLI built from this
branch:
- guarded delete with a wrong version, then a wrong uid, then the right
version
- stale version after an `update` bumps the policy
- plain delete on an actor without a policy, and on an actor that does
not exist
🤖 This PR was developed with AI assistance. I have reviewed and tested
all changes.
Follow-up to #1906. Removing a directory entry needs write permission on
the directory that holds it, and with all capabilities dropped uid 0
gets no exemption. A non-root container can create a subdirectory in its
`0777` durable dir; that subdirectory belongs to the container's uid,
typically `0755`, so root cannot empty it. `resetActorDirs` fails after
every checkpoint of such an actor and the suspend never completes. Chmod
first, as the bundle dir does, is no way out: chmod needs ownership or
`CAP_FOWNER`.
Why atelet and not ateom, which already holds the capability: an ateom
that dies before cleanup leaves the files behind and moves the hang to
the next resume on that node. atelet owns the actor directories and
needs it on the crash path either way.
Manifest only, so no unit test. Exercised on kind with the micro-VM
class: uid 65532 writes a durable dir, then suspend and resume. Without
the capability the same run hangs in `SUSPENDING`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
Fixes a part of #1709
This adds declarative validation to atelet's `AteomSupport` RPCs.
Malformed requests are now rejected at atelet instead of being forwarded
to the control plane. Validation runs after authentication.
| RPC | Rules |
|---|---|
| `MintActorCertificate` | actor atespace/name are required short names;
actor UID is a required UUID; the CSR is required and at most 16 KiB |
| `RequestActorSuspend` | actor atespace/name are required short names;
actor UID is a required UUID |
| `SetWorkerCapacity` | `capacity` is required and must follow the
`WorkerResources` rules ateapi declares, shared through
`internal/resources` |
The `AteomHerder` RPCs and `WorkloadSpec` get their validation in a
follow-up PR.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## What
- A pre-built install (`ate-setup deploy --image-repo REPO --image-tag
TAG`) now pins `REPO/envoy-dataplane:TAG` like the ko images, instead of
building it with docker buildx.
- `make build-envoy-dataplane` publishes that image.
- `make build-release-images` publishes every image a pre-built install
needs (ko images, demos, envoy-dataplane), all tagged `$(VERSION)`.
## Why
envoy-dataplane (#1535) is built from a Dockerfile, and ate-setup built
it even for pre-built installs. Those installs have no registry to push
to, so they failed on the default envoy router after the control plane
was applied. It also couldn't join `ALL_IMAGES`, which only lists ko
builds, so nothing published it for a release.
## Testing
- New unit tests for pinning and rendering a pre-built egress manifest.
The rendering test fails on `main` with the production error (`invalid
tag "/envoy-dataplane:build-…"`).
- `make build-release-images` pushed all 13 images under one tag, and
`images.Prebuilt` pinned every one against the real registry.
- Cherry-picks cleanly onto `release-0.2`.
Fixes#1962
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Anna Pendleton <13793293+annapendleton@users.noreply.github.com>
## What
On the microVM restore path, ateom reseeds the guest kernel CRNG with
fresh entropy through the kata-agent's ReseedRandomDev RPC, right after
the guest resumes.
This implements the guest RNG reseed on restore that #1449 calls out as
an unfiled gap (its Strategy 2, inject freshness at restore). Scoped to
the kernel CRNG on the microVM runtime, so it does not close#1449.
## Why
Two actors restored from the same snapshot start with the same frozen
CRNG state. The kernel reseeds from ambient entropy on its own, but not
immediately. Measured on a live restore, two clones kept returning
identical /proc/sys/kernel/random/uuid values for up to about 36ms after
resume before that happened. A workload reading randomness in its first
moments after resume can land in that window and get duplicate values
across clones.
Cloud Hypervisor has no VmGenID device to signal the guest, so ateom
reseeds it directly. On each restore it hands a fresh 32-byte nonce to
the kata-agent, which mixes it into the guest CRNG. The nonce differs
per restore, so clones diverge regardless of the kernel's own timing.
The kata-agent already exposes this RPC, so the change is host-side in
ateom.
## Behavior
Fires on every restore (golden cold-start and resume alike). Cold boot
does not need it. Best-effort: a failure is logged, not fatal, so a
transient agent error never fails an otherwise good restore.
## Scope and limits
- Covers the kernel CRNG only. A workload's own userspace PRNG seed
lives in the checkpointed app memory and stays a workload concern
(Strategy 5 in #1449).
- The reseed runs just after Resume, through the kata-agent, so it lands
a few milliseconds after the vCPUs resume. In the live test the reseeded
window shrank from about 36ms to about 6ms but did not reach zero. The
first read or two after resume can still be frozen before the reseed
lands. Fully closing it needs freezing the workload cgroup across the
reseed, or a VMM VmGenID that acts before the vCPUs resume. A TODO in
the code points at the VmGenID path once Cloud Hypervisor gains the
device. Left as a follow-up.
- microVM runtime only. gVisor reads getrandom and urandom live from the
host with no checkpointed RNG state, so restored gVisor sandboxes
already get fresh entropy.
## Testing
- Unit test on the nonce generator (length, and that two nonces differ).
- Verified live on a GKE microVM cluster with a workload that
continuously records /proc/sys/kernel/random/uuid, so the restore
instant is observable. Two clones from one golden snapshot were
compared. Stock ateom returned identical CRNG output for up to about
36ms (7 consecutive reads) after resume, varying run to run. With this
change the shared window shrank to about 6ms (1 read), then diverged.
---------
Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
Plumbing of actor's EgressPolicy into Envoy dataplane. This is the first
commit that only implements SNI allowlist, without wildcards. Followup
PRs will implement there rest of the policy:
* Wildcard matching.
* SNI passthrough policy.
* Allowing plaintext based on presence of `http` rules.
* Port verification.
* Making sdsmint the default and removing non-sdsmin config.
----
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Yan Avlasov <yavlasov@google.com>
This PR let ateom-microvm reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.
The paths ateom-microvm still builds itself are put into microvm
specific`cmd/ateom-microvm/actorpath.go`.
The `ateompath` package as nothing uses it anymore.
Fix#1604
`run-dev.sh` refused to start unless the Running worker count was at
least `--actors`, on the assumption that each worker hosts one actor. It
does not: a worker advertises capacity for many actors and the scheduler
packs them.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## e2e: add e2e for egress credential injection
Adds the `egresscredinject` e2e suite: an actor fetches
`https://httpbin.org/headers`
through the sdsmint MITM gateway, and the echoed response proves the
injected
`Authorization` header actually reached the upstream.
### What it asserts
- **Injected**: echoed headers contain `Authorization: Bearer <token>`.
- **Overwritten**: an actor-pre-seeded `Authorization` is replaced by
the injected one.
- **Cleartext skip**: plain-HTTP fetch passes through with no
`Authorization`.
- **Fail closed**: nonexistent secret → 403, unserved provider → 500,
unauthorized namespace → 403.
### Additional changes
- Probe `/fetch` returns the response body, accepts
`header=<name>:<value>`
params, and no longer follows redirects (a cross-scheme redirect would
hop
between the cleartext and TLS legs). Covered by unit tests; egressmitm
rerun green.
- New `e2e.EgressInjectHeader` policy helper.
- New `e2e.DeployCredentialProvider` + `fixtures/credinject`: deploys
the real
provider manifest with a test Secret and namespace-policy ConfigMap,
restarts
the provider so it picks the policy up, cleans up on test end.
- CI: new envoy-lane steps after the MITM lanes, gated on
`E2E_EGRESS_CREDINJECT=1`.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Related: #1694
The `swebench-astropy-7336` ActorTemplate can't be deployed as shipped:
1. `image: ""` must be hand-edited into a tracked file before every run.
2. The command points at `/opt/swebench/replay.py` and
`/opt/swebench/astropy_trace.json`, which the sweperf image generator
never
produces. It writes `/replay.py` and `/trace.json` at the image root
(also
the image's ENTRYPOINT), so the container exits immediately.
This PR:
- takes the image from a `SWEPERF_IMAGE` env var, substituted by
`workloads/deploy.sh` like the other template variables;
- fixes both paths;
- documents the minimum sweperf commit (`c30c0d6`), since older images
have no
HTTP server and fail on the first `GET /status`.
Usage:
```sh
SWEPERF_IMAGE=<registry>/sweperf-astropy@sha256:... \
WORKLOAD_TEMPLATES="swebench-astropy-7336" \
benchmarking/workloads/deploy.sh --deploy ...
```
Verified on GKE with both gVisor and microVM worker pools: 30-minute
sweperf
runs at 1 user, all 4 cycles, 0 failures.
- [x] Tests pass (`bash -n`; exercised end-to-end as above)
- [x] Appropriate changes to documentation are included in the PR
(template comments)
This PR let ateom-gvisor reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.
The paths `ateom-gvisor` still builds itself are put into gvisor
specific `cmd/ateom-gvisor/actorpath.go`.
Resetting `runsc-state` and `pidfiles` moves from atelet's
`resetActorDirs` to ateom-gvisor's `main.go` to conduct at beginning of
`RunWorkload` and `RestoreWorkload`, which empties them at the start of
Run and Restore. Atelet doesn't actually read them, so #1730 drops them
from `ActorDirs`.
Added a new `ActorDirs` validation to reject a request with
`InvalidArgument` when `ActorDirs` is missing or any directory in it is
empty, relative or unclean.
The gVisor paths are removed from `internal/ateompath`; the rest goes
with ateom-microvm in the next PR.
Part of #1604.
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:
1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).
The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.
Partial work for #1871
- [x] Tests pass
```
Run make build-junittool
make build-junittool
bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
shell: /usr/bin/bash -e {0}
env:
E2E_ATENET_DATAPLANE: agentgateway
ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1
/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
TOTAL
```
- [x] Appropriate changes to documentation are included in the PR
Fixes#1911
`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.
No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Closes the remaining items in #1640: the atelet side of multi-actor
telemetry.
Since #1836 an ateom returns one `GetActiveWorkloadStats` entry per
hosted actor. atelet's sweep already folded every entry, but nothing
pinned the behavior at occupancy above one, and reviewing that fold
surfaced three CPU-accounting defects: two in how baselines survive a
pending sweep, one in the decrease rule.
## What changed
**Tests pin the multi-actor fold.** One worker, several actors:
same-template entries sum under one key, a different template gets its
own, a pending sibling contributes nothing to aggregates or events, CPU
baselines are per actor (independent advance, independent epoch reset),
one probe serves the whole response. Each test was verified to fail
against the corresponding mutation.
**Three CPU-accounting defects found and fixed along the way:**
1. *Undercount.* A pending entry (boot, restore, or a micro-VM guest the
ateom did not reach in its 45s budget) never wrote the actor's baseline,
and `lastCPU` is rebuilt each sweep — so measured → pending → measured
lost the whole interval. On the guest-agent source the counter survives
a restore, so that was real CPU time dropped, and on a full micro-VM
worker it would happen routinely. Pending now carries the baseline
forward; only actors absent from every response are dropped.
2. *Overcount opened by fix 1.* Carrying the baseline unconditionally
meant that during a restore spanning one sweep (source still reports the
actor measured at C1, destination reports it pending), goroutine order
could let the stale C0 overwrite C1, and the next sweep charged the
interval twice. The pending branch now writes only if no worker has
measured the actor this sweep.
3. *Overcount on a guest-agent decrease.* A CPU decrease was charged as
the whole new value. That is right for a cgroup counter, which restarts
at zero on restore. It is not always right for a guest-agent counter: a
FULL or DATA_ON_GOLDEN restore resumes a guest at that guest's counter
(after a golden restore, main billed the golden guest's CPU to the
actor), a cold boot restarts it at zero, and a container exit lowers the
sum. A sample cannot tell these apart, so a guest-agent decrease now
charges nothing and re-baselines. A cgroup decrease still charges the
new value, with or without a pending sweep between. **Trade-off, new
versus main:** a guest-agent decrease loses the usage from the event to
the first sample. Main counted it correctly only after a cold boot (a
DATA restore, or a Run after a crash); in every other case (golden
restore, FULL restore from an older snapshot, container exit) main
overcounted it.
Remaining limit, documented in the proto: an epoch that starts *above*
the last reported value cannot be detected from consecutive samples.
Main has this for guest-agent restores between sweeps; carrying the
baseline extends it to restores that span a pending sweep. The overcount
is bounded by the golden guest's counter minus the actor's last value,
so it needs an actor that used less CPU than the golden guest had at
snapshot.
The test for (2) forces both fold orders deterministically: the sweep
closes a probe's connection right after folding its response, so the
fixture's closer releases the gated probe from `Close` — no sleeps, no
polling.
**Timeout rationale reconciled.** The poll floor (50s) is one actor's
worst-case micro-VM read: 25 containers at 2s each. It never prevented
overlap, because sweeps are sequential; it bounds guest-agent load, and
the comment now says so. The RPC timeout (55s) was documented as
covering that single-actor worst case. With several actors per worker,
the micro-VM ateom caps its own read at `statsSweepBudget` (45s, chosen
to stay under atelet's 55s) and reports unreached guests as pending, so
the timeout stays a constant; its comment now points at that budget.
## Not in scope, recorded
- Two workers both *measuring* one actor in one sweep double-charges.
Pre-existing and unchanged here. It is rare: it needs the same probe
skew as defect 2, with the destination read after its restore finished.
- `ateom.proto` still says the discovery list holds "at most one entry
until multi-actor workers land". Stale; fixed in a separate change.
## Verification
`gofmt`, `go vet`, and race-enabled tests pass. Mutation checks:
dropping the carry, the measured-wins guard, or either side of the
per-source decrease rule each fails its test.