Fixes#920 partially.
Replaces per-write pg_notify with a worker_changes outbox table written
in the same transaction, and LISTEN with a 100ms polling watcher.
#### Motivation
* **Scalability Bottleneck**: `pg_notify` serializes the commits of all
notifying transactions through a global lock held across the commit
(including `fsync`). This artificially caps worker writes at ~600/s on
Cloud SQL regardless of instance size, whereas our target is O(10K)
worker updates/s.
* **Payload Limits**: Bypasses the 8KB `NOTIFY` payload limit that
previously caused writes to fail.
* **Reliability**: A cursor-based polling watcher survives reconnects
and failovers without missing events, which was a known flaw with the
ephemeral `LISTEN` approach.
*(Known Postgres pathology prior art:
[Recall.ai](https://www.recall.ai/blog/postgres-listen-notify-does-not-scale),
[DBOS](https://www.dbos.dev/blog/postgres-listen-notify-scalability)).*
### Performance Improvement
WorkerUpdate @ 1,000 QPS , 1M workers (preloaded) — before vs after the
change feed:
| | p50 | p90 | p95 | p99 |
|---|---|---|---|---|
| Before (per-update pg_notify) | 40.3s |55.6s | 61.2s | 63.8s |
| After (change-feed table) | 7.08 ms | 8.04 ms | 8.52 ms | 27.5 ms |
#### Changes Made
* **Schema**: Added transactional outbox table `worker_outbox`.
* **Write Path**: Worker writes now append to the `worker_outbox` feed
inside the same transaction instead of calling `pg_notify()`.
* **Watch Path**: Replaced `LISTEN` in `WatchWorkers` with a polling
watcher that queries the feed every 50ms.
* **Cleanup**: Implemented a janitor process during polling to
periodically delete old feed rows.
* **Tests**: Updated atomicity tests to verify feed inserts instead of
`pg_notify` payloads.
For full architecture:
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?tab=t.txqcjwvmhp3v
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Add an e2e test to cover egress on sending grpc requests and bidi
streaming in it.
What's here
- `internal/proto/grpcechopb` — a small Echo service with one method per
streaming shape: unary Echo, server-streaming EchoStream, bidirectional
EchoBidi. Generated files checked in, per the convention in the
neighbouring proto packages.
- `internal/e2e/fixtures/grpcecho` — the origin: a cleartext-HTTP/2
server, no TLS anywhere, with the standard health service so the pod can
use a grpc readinessProbe and nothing between the actor and the origin
parses HTTP. Pod + Service template, deployed per test into the suite's
namespace.
- `demos/egress` — the actor gains POST /grpc, which dials the target
and runs whichever RPCs the request asks for. It dials per request and
closes with it: the actor is checkpointed and restored, and an HTTP/2
connection opened before a snapshot does not survive one. The gRPC
status comes back as a string rather than being flattened into the HTTP
status, since that status is the trailer assertion.
- `internal/e2e/suites/networking/grpcegress_test.go` —
TestActorEgressGRPC drives all three shapes through nftables REDIRECT →
atunnel → atenet-egress → origin, then asserts the gateway's access log
recorded the CONNECT for that actor's certificate. Without that last
check everything above would also pass on masqueraded traffic that never
reached the gateway.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The old sdsmint suite includes three tests: TestSdsmintMintsALeafPerSN,
TestGatewayRefusesANonActorWorkload and
TestGatewayRefusesAnUnknownActor.
The egressmitm suite covers the e2e functionality of sdsmint, making
TestSdsmintMintsALeafPerSNI unnecessary.
TestGatewayRefusesANonActorWorkload and TestGatewayRefusesAnUnknownActor
used to live in the sdsmint suite. But their functionality does not
depend on sdsmint. So move them into a suite egressauthz.
New manifest_test.go verifies that the MITM CA is only mounted
on the `sdsmint` container in the `atenet-egress` Deployment.
The new `egressauthz` test suite takes about 13s to finish:
| Phase | Δ | Cumulative |
| --- | --- | --- |
| Framework init → `Creating namespace` | 1.45s | 1.45s |
| Namespace created, unknown-actor credential minted | 0.31s | 1.76s |
| `ko` build | 3.05s | 4.81s |
| `ko` publish to GCR | 1.92s | 6.73s |
| `kubectl apply` → pod created | 3.42s | 10.15s |
| Pod scheduled, image pulled, ready | 2.36s | 12.51s |
| `TestGatewayRefusesANonActorWorkload` | **0.49s** | 13.00s |
| `TestGatewayRefusesAnUnknownActor` | **0.05s** | 13.05s |
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This needs to be a property of the storage layer. It is in PG but not in
redis. This commit moves it out of the store interface and down into
redis. It's still not CORRECT but it constrains it.
This requires a lot of tests to create the atespace before creating
other resources. These tests are wrong in the face of a correct storage
layer, anyway.
xref #989
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.
#1011
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
GCS doesn't offer something like the S3 upload manager package
(https://github.com/agent-substrate/substrate/pull/1068) but the same
thing holds ... uploading decent sized chunks in parallel can really
speed up snapshot upload, we just have to implement it ourselves.
This PR implements parallel uploading for GCS.
Along the way, I discovered that atelet's current approach to
compression bottlenecks parallel upload, so the second commit splits
that into chunks as well.
For some test data, we wind up with epsilon the same bytes than before
but with a 451 MB working set we go from 4.66s (before this PR) to 1.51s
(after both commits).
For a tiny counter demo with snapshot mode full (NOTE: the default
counter demo we use for e2e uses only data mode for uploads, and full
only for pause), we go from 0.79s to 0.62s (22% faster), purely from the
parallelized compression (second commit) as it stays under the upload
chunk size.
Also relevant: https://github.com/agent-substrate/substrate/pull/1130
Part of #731 and follow up to #640. This PR makes improvements to the
Postgres store by addressing review comments on #940 and optimizing
Postgres updates as suggested in #988.
- Replaces Postgres row-locking updates for actors, templates, and
snapshot tags with bounded optimistic concurrency control using
UID/version CAS checks.
- Strengthens relational integrity with missing FKs and indexes, and
replace snapshot-tag creation pre-check with FK
- Removes unnecessary UID uniqueness constraints
- Standardize pagination and validation across stores, map malformed
tokens and invalid sizes to appropriate gRPC errors
- Improves actor resume race handling when workers disappear or
concurrent assignments exhaust their retry budget
- Improves reliability with startup retries, expired-lease cleanup,
worker-watch reconnection after malformed msg
- Expand Postgres, Redis, Contract, and Control API tests to cover new
constraints, error semantics, concurrency
---------
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
The router's dataplane health check and the drain sequence dial the
Envoy admin interface on the IPv4 loopback, which now depends on the
admin socket keeping ipv4_compat set alongside its `::` bind. Losing
that fails silently: the drain reads the refused dial as "Envoy already
exited" and reports a drain it never performed.
Both callers now dial localhost, which resolves to either loopback, so
the drain no longer depends on how the socket is spelled. ipv4_compat
stays for the egress kubelet probe, where a regression turns any IPv4
run red on the spot.
The router and egress manifests bind the IPv4 wildcard on every Envoy
socket, and neither gateway's Service asks for a second IP family. On a
dual-stack cluster the router answers on its IPv4 ClusterIP and on
nothing at all for IPv6; on an IPv6-primary cluster the kubelet cannot
probe the pod on its only address, so atenet-egress crashloops while
Envoy itself starts fine.
Both gateways now bind `::` as well and ask for PreferDualStack. That
makes them accept IPv6, not reach it: the egress dns_lookup_family, the
DNS AAAA path, and atunnel's original-destination lookup stay IPv4. The
experimental sdsmint egress variant is untouched.
`t.TempDir()` embeds the test name, and the four broker certificate
source tests have names long enough to push the socket past the ~104
byte sun_path limit on darwin, so every one of them failed to listen.
The test calls `setupTest`, `namespaceForTest` and `assertGrpcError`,
which are defined in the `functionaltest` package, but the file landed
in `controlapi` declaring package `controlapi`, so the `controlapi` test
binary did not build.
This code passed on the PR branch, but it logically conflicted with a
previous refactor without merge conflict.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #294 as #311 has been inactive for over a month.
Example
```
kubectl ate logs actors test -a demo # default: all containers + lifecycle
kubectl ate logs actors test -a demo -c counter # only specified containers
```
Scoped this down to --container only after discussing with @BenTheElder
— we are not sure the naming of the container related logs, the
supervisor output may end up as a describe/events-style like how k8s
does?
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #477.
Changes
* Implement ActorTemplate CRUD in the control API
* Added a shared generic ResourceRef behind the template and actor ref
types
Relevant:
* Following style guidances in
https://github.com/agent-substrate/substrate/issues/891
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data
source for SystemInfo volumes (#802) and the end-to-end proof that the
projected anchors work against the MITM egress gateway. Live refresh for
running actors (PR 2) and auto-injection (PR 3) come separately.
What this adds
A SystemInfo volume data source that projects the trust anchors of a
named trust bundle to a PEM file:
volumes:
- name: trust
systemInfo:
dataSources:
- trustBundle:
name: egress-mitm.ate.dev
path: egress-ca.pem
Inspired by the Kubernetes clusterTrustBundle projected volume source,
but source-neutral: the template names a bundle; where it's fetched from
is a deployment concern, not part of the API.
Design points
- Resolution lives on the node. The wire carries only {name, path};
atelet resolves the name at write time through an informer-backed lister
on ClusterTrustBundles and writes the sanitized PEM with the temp+rename
discipline from #803 (find-paths safe). Contents refresh on every
Run/Restore. ateapi is not involved, per review discussion — the same
informer is what live refresh (PR 2) will hang off.
- Allowlist in atelet, not the CRD schema. Today only
egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the
ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946)
derives from the egress-mitm-ca-pool Secret. The signer-linked object
name stays a backend detail; the future backend registry (#932) widens
the allowlist without an API change.
- The watch is scoped to the one backing object via a metadata.name
field selector — this informer runs on every node, so an unfiltered
watch would fan every ClusterTrustBundle in the cluster out to every
atelet. RBAC can't express this (resourceNames doesn't apply to
list/watch), so the field selector is the enforcement point.
get/list/watch on clustertrustbundles moves to the atelet ClusterRole.
- No availability probe. The informer registers unconditionally; a
cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1
blocks atelet startup at cache sync, with the reflector errors naming
the missing API (hack/create-kind-cluster.sh enables the gate).
- Fail-closed. Unknown names, missing bundles, and unusable bundles fail
actor start naming the bundle — an actor that declared a trust bundle
must not start without one.
- Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks
only, deduplicated, headers stripped, and anchors deliberately shuffled
so consumers can't grow a dependence on order.
- Schema note: dataSources MaxItems tightened 32→8 while adding the
trustBundle member. Vacuous in practice (the old schema couldn't admit
more than one entry), but flagged since it's ratchet-shaped.
E2E — delivery and consumption
Delivery (identity suite, both sandbox classes): provisions the
egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing
the bundle directly isn't possible — the reconciler reverts hand-edits),
asserts the projected file byte-exact, then rotates the pool across a
suspend/resume to prove refresh-on-restore. Since the probe fixture is
shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists
for whatever suite deploys it.
Consumption (new egressmitm suite, both sandbox classes): deploys the
sdsmint (MITM) egress gateway and proves an actor completes a TLS
handshake with the gateway's per-SNI minted leaf using ONLY the
projected anchors — plus a system-roots negative control that must fail.
The pair is unambiguous in both directions: the positive can't pass
under passthrough (the bundle holds no public CAs), and the negative
can't fail under passthrough.
CI: two steps appended to the existing e2e job after the standard lanes
(the gateway swap is cluster-wide and breaks passthrough assumptions):
--deploy-atenet --experimental-use-sdsmint redeploys only the atenet
components, then the egressmitm suite runs once per sandbox class.
Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each
suite deploys its own copy and drives one actor at a time, so the third
worker per copy was idle memory multiplied across suites on the one-node
CI cluster — pressure that has been killing sandboxes mid-test (runsc:
signal: killed, a vanished ateom socket) on this PR and on main's
identity suite. This reduces the pressure; right-sizing e2e concurrency
or worker-pod QoS cluster-wide is follow-up material.
Not in this PR
- Live refresh for running actors (#932 PR 2) — until then, a running
actor's file is the bundle as of its last Run/Restore, and correctness
rests on overlap rotation by the bundle publisher.
- Auto-injection of the egress trust volume (#932 PR 3).
- Configurable backend registry (#932) — the allowlist is the seam it
will replace.
## Problem
`ResumeActor` returns `ResourceExhausted` ("no free workers available")
when the target pool has no free worker. The control plane treats this
as a transient condition callers should wait out — the router's parking
resumer (`cmd/atenet/internal/router/ingress/resumer.go`) retries
exactly this error within its parking budget. The e2e suites are the
only resume caller that treats it as fatal.
Since suite packages run concurrently (up to `GOMAXPROCS`) and some
share pools (the demo and metrics suites both drive the 3-worker counter
demo pool), any unlucky timing overlap fails a lifecycle test on the
spot. This is currently the most frequent flake signature in the e2e
lane: recent failed main-branch runs contain 50+ occurrences each (e.g.
runs
[32516335516](https://github.com/agent-substrate/substrate/actions/runs/32516335516),
[32429542405](https://github.com/agent-substrate/substrate/actions/runs/32429542405)),
and it also produced red runs on #941 for code it doesn't touch.
## Change
- New `e2e.ResumeActorAwaitCapacity(t, ctx, clients, req)`: calls
`ResumeActor` and retries only `ResourceExhausted`, every 2s within a
90s budget (several neighbor-suite actor lifetimes). Every retry is
`t.Logf`'d so a pass that had to wait stays visible in the output — the
helper hides the coin-flip, not a capacity regression. Any other error,
or saturation outlasting the budget, returns to the caller unchanged.
- Mechanical sweep of the 24 resume call sites whose intent is "get the
actor running" (demo, termination, metrics, capabilities, identity,
networking, sizing, imagevolume, networkpolicy).
## Deliberately not swept
- The **parking suite**: park-on-saturation is the behavior it exists to
test; its calls stay raw.
- sdsmint's `createLiveActor`: the helper is deliberately
`*testing.T`-free (error-returning), and the suite skips in presubmit
lanes; not worth replumbing for the retry.
Scope note: this only removes the `ResourceExhausted` flake class. The
lane's other pre-existing failure class (sandbox processes killed under
node memory pressure) is separate and tracked independently.
This reverts commit 57c3e1927c.
Drops --experimental-additional-egress-extproc-service and the
#ATE_MITM_EXTPROC_FILTER / #ATE_MITM_EXTPROC_CLUSTER markers, so
atenet-egress-with-sdsmint.yaml renders exactly as committed again
and hack/experimental-additional-egress-extproc.sh goes away.
## Summary
Wait for the benchmark worker pool deployment to roll out before
returning from `benchmarking/workloads/deploy.sh`, preventing downstream
suites from racing against unready workers.
## Key Changes
- **Rollout Wait:** Added `kubectl wait --for=create` followed by
`kubectl rollout status` on `deployment/benchmark-ateom` in
`benchmarking/workloads/deploy.sh`.
- **Configurable Timeout:** Added `--wait-timeout DURATION` flag
(default: `300s`) to `workloads/deploy.sh` and forwarded it through
`deploy_locust.sh`.
- **Validation:** Added duration regex validation (`^([0-9]+(h|m|s))+$`)
to catch invalid formats/missing units early.
## Testing
- Verified successful rollout wait on GKE: `deploy.sh --deploy
--worker-count 2 --wait-timeout 300s`
- Verified timeout failure handling: `deploy.sh --deploy --worker-count
5 --wait-timeout 1s`
- Verified flag validation rejects invalid formats (`180`, `0`,
`invalid`).
## What Changed and Why
Durable-dir volumes were previously served to the guest by a second
per-actor `virtiofsd`. That arrangement predates the writable
`kataShared` share: when the rootfs share was a read-only lower, a
writable durable share had to be its own device. Since #846, the
`kataShared` tree is writable and served with `--announce-submounts`,
making the second daemon redundant—it cost an extra process and `vhost`
socket per actor, an extra `fs` device in every snapshot config, and a
restore-time revival of all three.
This PR folds the durable-dir volumes into the single existing share:
* **Subtree Bind Mounting (`_durable`):**
`kata.BindIntoShare` bind-mounts the `atelet`-owned volumes directory
into the served tree as its `_durable` subtree. The leading underscore
keeps it out of the container-ID namespace (container names are RFC 1123
labels and cannot begin with an underscore). The guest sees it as a
submount of the `kataShared` mount, and containers bind their volumes
from `<shared>/_durable/<volume>` exactly as they previously did from
the second share.
* **Streamlined VM Configuration:**
Cold boot no longer spawns the durable `virtiofsd`, and the VM
configuration carries exactly one `virtio-fs` device.
* **Unchanged Ownership Semantics:**
`atelet` still owns the directory (creates it before boot, wipes it on
actor reset). The bind is `ateom`-owned mount state, detached by
`CleanupSandboxState` before any removal, ensuring `atelet`'s data is
never touched through it.
* **Unchanged Snapshot Content:**
Checkpoints tar the host directory directly under every scope, exactly
as before.
> [!NOTE]
> The bind uses the same mechanism the per-container merged rootfs
mounts already use through this `virtiofsd` (submounts inside the served
tree, re-opened by `find-paths` on restore), introducing no new
guest-side behavior. `BindIntoShare`'s doc comment establishes this
pattern for future per-actor shares: **mount a reserved subtree rather
than adding a device** (relevant to in-flight work such as #803; #923
already follows this subtree approach).
---
## Compatibility with Existing Snapshots
Restore is self-describing in both directions:
* **Two-Share Era Snapshots:**
A snapshot whose `config.json` carries an `ateDurable` fs device
originates from the two-share era. The resumed guest still expects that
device, so restore revives the second `virtiofsd` exactly as before
(`stageLegacyDurableShare`). Such a lineage remains two-share across its
own re-checkpoints since `cloud-hypervisor` re-emits the device.
* **Single-Share Snapshots:**
A snapshot without the device gets the volumes re-bound into the shared
tree before `virtiofsd` starts, allowing `find-paths` to re-open the
guest's open durable files at their `_durable/...` paths, which the
restored tar reproduces exactly.
Detection logic lives in `rewriteSnapshotSocketPaths`, which already
inspects the configuration's `fs` devices. The configuration acts as the
authority because the device is what `cloud-hypervisor` re-opens,
regardless of the actor's spec.
---
## How This Was Tested
* **Unit Tests:**
* Added assertions verifying that `rewriteSnapshotSocketPaths`
classifies single-device and two-device snapshot configs correctly
(routing restore to the bind vs. the legacy share).
* Added a new `kata-package` test pinning `_durable` subtree invariants:
the underscore namespace reservation, the guest path residing inside the
single `kataShared` mount, and host/guest agreement on the relative path
re-opened by `find-paths`.
* Verified that existing durable-volume unit tests (tar/untar
round-trip, container mount construction, spec isolation) continue to
pass unchanged.
* **End-to-End Verification:**
* Ran the counter demo on a GKE cluster with nested virtualization:
verified cold boot, suspend, and resume of a durable-volume actor on
this branch (volume contents survived across worker pods).
* Verified restore of a legacy snapshot taken on `main` before this
change to exercise the legacy two-share fallback path.
## Summary
- Adds `.agents/skills/detect-flaky-tests/SKILL.md` — a new agent skill
for detecting flaky Go tests
- Uses cross-PR analysis over the last 7 days of `pr-workflow.yaml` runs
(the strongest flakiness signal)
- Low false-positive rate: a test is flagged only when `fail_count >= 2
AND pass_count >= 2 AND 0.05 < flake_rate < 0.95`
- Fully self-contained: no external storage, no dashboard writes — those
are cron-job concerns
## What the skill does
1. Fetches the last 7 days of completed `pr-workflow.yaml` runs via the
GitHub API
2. Downloads and parses `go test -v` output from the `run-tests` job
logs per run
3. Aggregates pass/fail counts per test name across all runs
4. Flags flaky tests at the threshold above
5. For each newly-detected flaky test (no matching open issue): creates
a GitHub issue with full evidence (run counts, links to failing and
passing runs)
6. Opens a **draft** fix PR targeting common Go flakiness patterns:
timing sleeps → `require.Eventually`, shared global state → `t.Cleanup`,
port conflicts, goroutine leaks, `t.TempDir()` for file-system races
## What the skill does NOT do
The skill deliberately has no awareness of BigQuery, dashboards, or any
external storage. Those are layered on top by the cron job that invokes
the skill, keeping it clean and reusable by anyone who just wants the
issue/PR workflow.
## Test plan
- [ ] Run the skill manually: confirm it fetches run IDs and parses log
output
- [ ] Verify the flakiness threshold correctly excludes reliably-failing
tests
- [ ] Confirm no duplicate issues are created for already-open flake
reports
- [ ] Confirm draft PRs are not opened when an existing fix PR is
already open
---------
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Follow-up to #749. One of three; the other two are independent of this
one.
This closes the review thread that stayed open on that PR:
krisztianfekete asked
for `memory_limiter` and `GOMEMLIMIT` on the meter, and I kept it open
to track.
The automation's test-cluster creation command did not enable Managed
OpenTelemetry, but the ate-otel-config ConfigMap applied by
hack/install-ate.sh and the runner Job both target
opentelemetry-collector.gke-managed-otel.svc.cluster.local:4317. Without
the addon that name does not resolve and all benchmark telemetry is
dropped.
`--experimental-additional-egress-extproc-service=NS/SVC:PORT` makes the
sdsmint egress gateway consult an external authorization processor on
the decrypted leg, where the request is a hostname, method, and path
rather than the IP:port the CONNECT checkpoint sees. The actor's
identity crosses into that callout as the `dev.ate.actor` filter state,
so the processor gets both halves of the decision.
The wiring is generated rather than checked in:
atenet-egress-with-sdsmint.yaml carries inert `#ATE_MITM_EXTPROC_FILTER`
and `#ATE_MITM_EXTPROC_CLUSTER` markers that the installer substitutes
only when the flag is set, so an install without it renders exactly the
manifest as committed.
The flag names any Service, so the processor itself is out of scope
here.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
mitm_listener classified everything that was not TLS as HTTP, so SSH and
every other non-HTTP protocol reached a connection manager that parsed the
first bytes as a request line, found none, and dropped the connection with
nothing in the access log. http_inspector now splits raw_buffer again, and
a third chain tcp_proxies what neither inspector could claim.
That chain cannot police what it carries -- an opaque stream has no Host
and no SNI -- so it dials the IP:port ext_proc already authorized on the
CONNECT, relayed across the internal hop as original_dst_address filter
state through an internal_upstream socket. Nothing is weakened: these
connections were not blocked before, they were accepted and mangled.
Server-speaks-first protocols would otherwise deadlock both inspectors, so
listener_filters_timeout is 1s with continue_on_listener_filters_timeout,
which lands a silent client on the passthrough chain where it belongs.
A long-lived stream was capped twice on the MITM leg: a route timeout runs
until the response is complete, which for SSE or a WebSocket is never, and
stream_idle_timeout reaped a stream that was merely quiet. Both drop to 0s
there, and the CONNECT route to 0s as well. The outer HCM's 5 minute
stream_idle_timeout still bounds an idle tunnel, and any byte in either
direction resets it. Both HTTP chains also gain a websocket upgrade_config,
because a leg that parses HTTP has to be told to proxy an upgrade rather
than answer it.
The MITM leaf now offers h2 as well as http/1.1, which grpc-go requires.
Upstream stays HTTP/1.1 by default, since only it can carry an Upgrade to
an origin that has no RFC 8441 extended CONNECT; gRPC needs HTTP/2 trailers
instead, so content-type routes it to a second forward proxy that takes its
protocol from upstream ALPN. The cleartext leg mirrors whatever the actor
spoke. Access logs carry grpc_status, since a failed RPC is still HTTP 200.
Fixes#643
This feature introduces the ability to delete Substrate actors from any
lifecycle state (e.g., RUNNING, PAUSED, PENDING), rather than requiring
them to be hibernated into SUSPENDED (or CRASHED) beforehand. This is
crucial for cleaning up stuck or failed actors.
#### 1. Control Plane & gRPC API (pkg/proto/ateapipb, ateapi)
• DeleteActorRequest.any_state: Added a new boolean field any_state to
DeleteActorRequest.
• When false (default): enforces the existing behavior where only actors
in ACTOR_STATE_SUSPENDED or ACTOR_STATE_CRASHED (or already
ACTOR_STATE_DELETING) can be deleted.
• When true: allows deleting an actor in any state (e.g., RUNNING,
PAUSED).
• Orchestrated Cleanup Workflow (workflow_delete.go):
1. Transitions actor state to ACTOR_STATE_DELETING.
2. Calls atelet.Terminate to stop live workloads on the node.
3. Detaches external volumes from the worker node (with fallback logic
using actor status if the ActorTemplate was deleted).
4. Releases the assigned physical worker Pod back to the WorkerPool.
5. Deletes external storage volumes through the storage plugins.
6. Finalizes removal of the actor from the persistence store.
#### 2. Node Herder Agent (atelet)
• Terminate RPC (main.go): Added a new Terminate method to AteomHerder:
• Calls ateom.TerminateWorkload on the target worker pod.
• Unmounts external volumes mounted for the actor on the node.
• Cleans up and resets actor runtime host directories (OCI bundles,
checkpoints, pid files).
#### 3. In-Pod Sandboxes (ateom-gvisor & ateom-microvm)
• TerminateWorkload RPC:
• gVisor (main.go): Stops and deletes active runsc containers, unmounts
bundle rootfs overlays, and tears down the actor's interior network
namespace.
• Micro-VM (checkpoint.go): Shuts down the Cloud Hypervisor VMM,
unmounts bundle rootfs overlays, and cleans up the network namespace.
• Emits an "Actor terminated" lifecycle log and clears the active actor
attribution so the worker can host new workloads.
#### 4. CLI (kubectl-ate)
• --any-state Flag (delete_actor.go):
kubectl-ate delete actor <actor-name> -a <atespace> --any-state
──────
## Verification
[x] Changes tested in local with
```
• go build ./...: Passed
go test ./...: Passed (all unit and functional tests pass)
make lint: Passed (no lint errors)
```
Detailed manual tests
```
### Scenario 1: Force Delete a Running Actor with --any-state
1. Create the actor:
kubectl ate create actor manual-test-1 --template=ate-demo-counter-microvm/counter-microvm -a demo
2. Resume the actor:
kubectl ate resume actor manual-test-1 -a demo
3. Force delete the running actor:
kubectl ate delete actor manual-test-1 -a demo --any-state
actor "manual-test-1" deleted
```
──────
### Scenario 2: Standard Delete a Suspended Actor
```
1. Create the actor:
kubectl ate create actor manual-test-2 --template=ate-demo-counter-microvm/counter-microvm -a demo
2. Resume the actor:
kubectl ate resume actor manual-test-2 -a demo
3. Suspend the actor:
kubectl ate suspend actor manual-test-2 -a demo
4. Standard delete the suspended actor:
kubectl ate delete actor manual-test-2 -a demo
actor "manual-test-2" deleted
```
──────
### Unit & E2E Tests
# Unit & Functional Tests
go test ./cmd/ateapi/internal/controlapi -run
"TestDeleteActorWorkflow|TestEnsureMarkedDeleting"
# E2E Tests
E2E_TEMPLATE_NAMESPACE=ate-demo-counter-microvm
E2E_TEMPLATE_NAME=counter-microvm ./hack/run-e2e-kind.sh
./internal/e2e/suites/demo -run TestActorLifecycle
./hack/run-e2e-kind.sh ./internal/e2e/suites/demo -run
TestForceDeleteActorWithExternalVolume
```
- [ x] Appropriate changes to documentation are included in the PR
This allows for faster signing of certs for clusters with large numbers
of atelets.
It also allows configurable ready wait times to compensate for larger
reconciliation waits generally.
Measures the max RPS atenet-router's ingress side sustains at a given Envoy CPU limit,
under a tail-latency SLO, using Nighthawk's adaptive load controller in
open-loop mode against the real routing path (Host-header routing via
ext_proc to warmed glutton actors).
- benchmarking/nighthawk-ingress/: runner Job that creates and warms the actor
fleet, drives nighthawk_service + nighthawk_adaptive_load_client with
Host rotation across actors, and uploads JSON/JSONL results to GCS.
- Search converges on three thresholds — tail latency (measured
mean+2stdev must stay under tailLatencySloMs), success-rate, and
send-rate — and records which one bounded the run. The client is
oversized (fixed event loops, large pools) so the harness is never
the ceiling.
- orchestrator.py: new `type: nighthawk-ingress` tests.yaml entries; pins the
router (cpu requests=limits, envoy --concurrency) before each run.
Validated end to end on a dev GKE cluster: ~8.9k RPS at 2 Envoy CPUs
under a 25ms tail-latency SLO.
Follow-up to #840.
Collapse `rewriteSnapshotSocketPaths` to only handle the single unified
`kataShared` virtiofs device, rejecting retired multi-share tags
(`ateDurable`, `ateCSI`) alongside `ateUpper`.
- Retire `DurableVirtiofsdSocketPath` and `CsiVirtiofsdSocketPath` in
`cmd/ateom-microvm/internal/kata/restore.go`.
- Update `restore_test.go` to assert that retired multi-share tags fail
loudly.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#783
Adds an `image` source to `ActorTemplate`'s `VolumeSource`: a container
can mount the contents of an OCI image it does not run. This is how
tooling gets into images built by third parties without rebuilding them.
```yaml
spec:
containers:
- name: sandbox
image: docker.io/example/benchmark@sha256:...
command: ["/ate/agent"]
volumeMounts:
- name: agent
mountPath: /ate
volumes:
- name: agent
image:
reference: registry.example.com/agent@sha256:...
```
## How it works
atelet pulls the image through the existing layer cache and records the
volume's layers in the bundle's overlay spec, next to the rootfs layers.
ateom composes the volume inside the bundle — the cached layers with no
writable layer on top, so the mount is read-only — and the container
binds it at the declared path. The volume is composed per container:
containers of one actor may mount the same volume, and each gets its own
mount point inside its own bundle, all backed by the same shared layers.
On resume the volume is re-composed the same way.
References must be digest-pinned, the same rule as container images: a
snapshot is only valid against the exact bytes it was taken with.
On micro-VMs the volume rides the same read-only virtio-fs share as the
container rootfs: ateom stages each composed volume beside the rootfs on
the host, and the guest binds it into the container at the declared
path.
Fixes#957
`ate.actor.lifecycle.operation.duration` carried the pool pair on
suspend and pause only when they failed, so per-pool dashboards saw
those two operations exclusively as failures.
Both workflows record the histogram from a defer that reads the `actor`
variable, and the happy path reassigns it to the finalized record. The
finalize step commits the new state and the cleared `WorkerAssignment`
in one update, so the defer found no assignment and dropped both keys. A
failure returns the pre-finalize record, which still names the worker.
Both now snapshot `lifecycleOpAttrs(...)` just before the finalize step
— the same snapshot-before-clear crash.go does for the crash counter.
Paths that end earlier keep the current computation.
`delete` stays without a pool: it only runs from SUSPENDED or CRASHED,
which already released the worker, so there is none to name.
TestLifecycleOpPoolAttributesOnSuccess drives a real suspend and pause
through the gRPC service; both subtests fail without the fix.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #212
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Summary
- add bounded, validated labels and annotations to
`WorkerPoolPodTemplate`
- propagate that metadata to the generated Deployment and worker pod
template
- reserve the controller-owned `ate.dev/worker-pool` label
- regenerate the WorkerPool CRD and deepcopy code
- add API validation, controller tests, and documentation
I chose intentionally to avoid exposing a complete `PodTemplateSpec` as
per the discussion in the #212.
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.
Validated with focused Go tests, shellcheck, and Kustomize renders.
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Snapshot upload is the largest single item in a micro-VM bake or suspend
— ~800 ms of a ~1.65 s server-side bake — and for the snapshots we
actually upload it is round-trip bound rather than byte bound.
The GCS client's resumable upload sends 16 MiB chunks one after another,
each paying a round trip. An idle micro-VM golden snapshot is ~24 MiB
compressed, so it spans two chunks and pays twice. Setting `ChunkSize`
to 64 MiB puts a typical snapshot in one request.
16MiB is reasonable for clients on poor connections sending small files.
It's not reasonable for data centers and large files.
### Measurements
Uploads through this exact streaming path (non-seekable body into a
`storage.Writer`), from a pod on a GKE worker node (c3-standard-4,
us-central1-f) using atelet's service account and the snapshot bucket.
Three runs at 24 MiB:
| chunk size | 24 MiB upload |
|---|---|
| 16 MiB (client default) | 425 / 438 / 530 ms |
| 32 MiB | 331 / 361 / 405 ms |
| **64 MiB** | **258 / 313 / 314 ms** |
| 128 MiB | 350 / 391 / 487 ms |
About 35% off an idle actor's snapshot upload. The client buffers at
most `min(object, ChunkSize)`, so a small object still costs only its
own bytes.
### What this deliberately does not do
Nothing for large snapshots. At 300 MiB a single stream measured 77–107
MiB/s for *every* chunk size from 16 to 128 MiB, because the transfer
dominates:
| config | 300 MiB |
|---|---|
| 16 / 32 / 64 / 128 MiB chunks | 82–107 MiB/s |
| composite, 2 parts | 150–163 MiB/s |
| composite, 4 parts | **233–257 MiB/s** |
| composite, 8 parts | 224–231 MiB/s |
Beating the single-stream ceiling needs parallel parts plus a compose,
which is a format change (each part has to be independently produced and
decodable), so it is left as a follow-up.
For context on why this is not a compression problem: on the same node
`zstd -1` compresses real guest memory at 549 MiB/s on two cores (167 ms
of a ~700 ms upload), while the whole pipeline moves ~105 MiB/s. Raising
the compression level trades CPU for fewer bytes on a wire that is not
the constraint at these sizes.
### Testing
- `go build`, `go vet`, `hack/verify/golangci-lint.sh` and `go test
./cmd/atelet/...` pass.
- The chunk-size effect is measured against the real GCS backend from a
real worker node, as above. I did not roll a patched atelet on the
shared cluster, so the end-to-end effect on a live suspend (~700 → ~400
ms) is inferred from the benchmark rather than observed in situ.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed; the rationale and numbers live in the code comment)
Part of #802 (first PR: the actorIdentity data source; does not close
the issue).
## What changed and why
Adds a systemInfo volume source to ActorTemplate — a read-only volume
whose files are generated by atelet on every Run/Restore, analogous to
Kubernetes projected volumes. The initial data source, actorIdentity,
writes the actor's own name to a configurable relative path:
```
spec:
volumes:
- name: system-info
systemInfo:
dataSources:
# Part 1 (this PR): own-metadata projection, downwardAPI-style
- actorMetadata:
items:
- field: name # enum: name | atespace | uid
path: actor-name
- field: atespace
path: atespace
- field: uid
path: actor-uid
containers:
- name: main
image: app@sha256:...
volumeMounts:
- name: system-info
mountPath: /run/ate
```
Because the files are regenerated before the sandbox starts, they carry
the resumed actor's own values regardless of what checkpointed state it
boots from — the property the old hardcoded /run/ate identity mount
provided, now as an explicit, extensible API that future data sources
(identity JWTs, certificates — see #802) can slot into.
**Behavior change:** the automatic /run/ate/actor-id mount is removed;
actors must opt in by declaring the volume (the e2e identity probe in
this PR is the reference example).
### Reviewer notes:
- Over half the diff is vendored + generated code
(cmd/atelet/internal/third_party/atomicwriter/, atelet.pb.go,
zz_generated.deepcopy.go, the CRD manifest). The hand-written surface is
~700 lines.
- System-info volume roots live under a new ActorPath/system-info/ host
dir, deliberately separate from durable-dir/: the micro-VM durable
machinery snapshots everything under the durable-dir root, and generated
identity files must never be captured into snapshots.
- Supports microVM as well as gVisor.
- The e2e identity suite exercises the new API end-to-end with unchanged
probe binary and assertions. It runs in the kind-cluster CI job (not run
locally).
## Checklist
- [x] Issue is linked above
- [x] Tests pass locally (go test ./...)
- [x] Root-gated tests pass if applicable (N/A — no root-gated packages
touched)
- [x] Documentation updated if behavior changed (docs/api-guide.md:
SystemInfo Volumes section with example)
---------
Co-authored-by: Taahir Ahmed <taahm@google.com>
Second piece of #896 (Phase 1 of #550): atelet polls every local ateom's
`GetActiveWorkloadStats` and turns the samples into template-level
metrics — the TSDB half of #174's cardinality split. After this PR, "how
much CPU and memory is this template using" is a Cloud Monitoring query.
## The reader
A `statsPoller` in atelet, driven by `--actor-stats-poll-interval`
(default 1m; `0` disables; nonzero values are clamped to a 50s floor,
the worst-case duration of sampling one micro-VM ateom — 25 containers ×
2s per guest-agent call).
- **Stateless discovery**: each tick lists `ateoms/*` (entry names are
worker pod UIDs — the same sockets the lifecycle RPCs dial) and probes
each with the parameterless discovery read. No worker-to-actor mapping,
no control-plane dependency, nothing to recover after an atelet restart.
Attribution comes solely from the identity echoed in each sample, per
the RPC's contract.
- **One tolerance rule**: any dial or call failure means "not a target
this tick" — covering stale directories of deleted workers, ateoms that
have not started listening, and teardowns mid-poll. `NO_WORKLOAD` /
`NOT_MEASURABLE_YET` are skips by the RPC's own contract.
- **Bounded concurrency**: distinct ateoms are probed 8 at a time (one
probe per guest, so nothing the interval floor defends against is
reintroduced); a node of stuck-but-accepting sockets degrades to
ceil(n/8) timeouts instead of n sequential ones.
- **WorkerPool enrichment**: one field+label-selected pod LIST per tick
maps pod UID → owning pool (`ate.dev/worker-pool`, the label the pool
controller stamps); unresolved pods group without pool labels rather
than vanish. Chosen over an informer deliberately — negligible apiserver
cost at this cadence, no cache-sync ordering, and the resolver sits
behind a function seam if that trade ever changes. Needs one new
Downward API env (`NODE_NAME`); the pods RBAC already existed.
## The metrics
Labels on every series: `ate.template.namespace/name`,
`ate.sandbox.class`, `ate.stats.source`, `ate.workerpool.namespace/name`
— all bounded sets; actor and atespace identity never reach a metric
label (they belong to the events channel, the next PR).
- `ate.actor.stats.sampled_actors`, `…memory_current_bytes`,
`…memory_working_set_bytes` — **observable gauges** over the latest
tick's snapshot: each collection observes exactly the groups that
currently exist, so a template whose actors leave a node disappears from
the export. (Synchronous gauges would re-export their last value until
process exit — stale memory for actors long gone.)
- `ate.actor.stats.cpu_usage` — **Float64Counter in seconds** (cAdvisor
/ OTel `*.cpu.time` convention; the wire stays µs). The raw
`cpu_usage_usec` is cumulative per-epoch per actor, so the poller
accumulates per-sweep *increases* against per-actor baselines: first
sight establishes a baseline and charges nothing (atelet cannot tell a
new actor from its own restart — re-charging epochs the previous process
counted would spike `rate()`), a decrease is an epoch reset charged from
the new value, and baselines are pruned to the actors seen. Bounded
imprecision (≤1 interval per actor across restarts; the pre-checkpoint
tail) is documented on the instrument; per-actor precision arrives with
the lifecycle events.
## Validated live on ate-dev
Both source families, simultaneously, with one actor per class:
<img width="2256" height="1180" alt="image"
src="https://github.com/user-attachments/assets/86ed3390-99b9-47f4-bbd2-a39ff1fd8d45"
/>
The same counter application reads 5-6× larger through the cgroup source
(whole sandbox: sentry heap, netstack, gofers) than through the guest
agent (workload containers only) — the concrete case for the
`ate.stats.source` label and its group-don't-sum rule. Also exercised
live: pool labels resolved on every series, restart-without-spike on the
CPU counter across a DaemonSet rollout, and OTLP delivery to the
gke-managed-otel collector with zero export errors.
## Out of scope
- Per-actor events + lifecycle first/final samples — next PR per #896.
- `k8s.node.name` resource attribute (one manifest line, any time).
- e2e coverage — the metrics e2e suite is the natural home once the
events channel lands.
Part of #896, toward #550.
## Summary
Adds `docs/integration-repos.md`: where end-to-end integrations live,
how their
repositories are named, and how the fixes they need flow back into core.
The convention in one line — trivial demos stay in the core repo, each
non-trivial integration gets one dedicated repo under the
`agent-substrate`
org, and core gaps get closed by making core configurable with defaults
unchanged rather than by patching it downstream.
## Why now
We are about to create the first real, end-to-end integrations rather
than
counter-style demos: a code-execution sandbox, and an always-on agent.
Both are
large enough to need their own images, dependencies, and release
cadence.
Whichever repository gets created first will set the precedent for every
one
after it. This writes the convention down so that precedent is chosen
deliberately instead of inherited by accident.
## What it covers
- **Where code lives** — the core-repo/dedicated-repo split, the rough
test for
which side something falls on (API keys, external services, third-party
accounts), and why this is a set of peer repos rather than a second org.
- **Naming** — capability-named for general capabilities
(`code-execution-sandbox`), integration-named for specific third-party
products, named for the product rather than the vendor behind it. Plus
what to
avoid: over-broad names, names that clone a vendor's API or brand, and
the
redundant `-integration` suffix.
- **Third-party names** — allowed descriptively, with a non-affiliation
note in
the repo README, and brand/policy edge cases cleared before the repo
exists.
- **Upstreaming** — the part with teeth for this repo. Integration repos
that
accumulate local patches against core bitrot, and the gap they work
around
stays invisible to everyone else. So: prefer making core behavior
configurable
with defaults unchanged. #487 and #465 are linked as illustrations of
that
pattern — this PR does not depend on either, and branches from `main`.
- **Two worked examples** that validate the convention rather than just
following it, including the third-party-name edge case.
## Review
This was announced at the community meeting and circulated as a shared
design
doc with a 7-day review window, which has now closed. It synthesizes the
`#integrations` thread discussion. Comment history:
<https://docs.google.com/document/d/1Tb6u0b1XSvWrNpoyD4jdsQaJ58aAgDtQOM18uxujs-8/edit>
This PR is the trimmed version: doc-review scaffolding — status block,
reviewer
list, self-link — is dropped, and only the durable convention is carried
over.
## Left open
Two questions are deliberately out of scope, called out in the doc
rather than
answered. Both are maintainer calls and neither blocks the first
repositories:
- Governance tiers — whether to distinguish "official" from "community"
integrations with different review bars, as Home Assistant and Obsidian
do.
- Who creates integration repositories and grants per-integration
maintainer
access.
## Also in this PR
- README gets an entry in the docs list, matching every other file in
`docs/`.
- `CONTRIBUTING.md` gets one sentence pointing there, since "where does
my
integration go?" is a question a contributor asks before opening a PR.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Make workers a gobal resource, identified by a unique name. Worker APIs
will be implemented in a follow up change.
Part of #730
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
Finishes up the vision from #715 to have atunnel serve CONNECT on the
ingress path. This will give us the option to hit actors on other ports
besides 80. I haven't wired up atenet router yet because it's
nontrivial; we should do that in a second step so we can have a baseline
for performance
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
Actor logs used `ate.dev/actor_*` while spans and metrics use `ate.*`
registry in internal/ateattr. This PR makes ateattr the single source of
truth for everything telemetry-related.
- Renamed the six actor log labels onto the registry:
- `ate.atespace`, `ate.actor.name`, `ate.actor.uid`,
`ate.template.namespace`, `ate.template.name`,
`ate.actor.container.name`
- Logs join traces now. Records set `trace_id`, `span_id` and
`trace_flags`, so you can go from Actor restored to the resume that
caused it. Our own lines only, not an actor's stdout: one goroutine
forwards a whole container stream and can't know which request produced
a given line. Per line correlation comes with #853.
- Actors can't fake platform labels. They already couldn't overwrite
ours, but they could invent new ones like `ate.tenant` that look
platform issued downstream. Anything under `ate.` from an actor is now
dropped.
- Fixed the asymmetry that was actually left: actor supplied label
values weren't stringified, and one non string value makes Cloud Logging
discard the labels for that whole entry.
- Note: the actor_uid bullet in the issue is stale, #841 fixed it
earlier. Lifecycle records still set five labels rather than six, on
purpose as they're about the actor, so no container produced them.
Fixes#886
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: krisztianfekete <git@krisztianfekete.org>
The in-memory version counter restarted at 1 on every router boot; if
Envoy reconnected still holding an identical version string the
snapshot cache saw a match and skipped the push, stranding Envoy on
pre-restart config. Versions now carry a per-process epoch (unix
seconds plus a random suffix) so no incarnation repeats an earlier
one's strings, even across clock jumps. Fixes#617.