**BREAKING** API changes prior in preparation for GA.
~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow
- `EgressRule` is now a union of protocols: `http`, `https`,
`tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
`tls_passthrough`. Same wildcard grammar as before, documented once on
`HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
once on `HTTPRule.ports`.
- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
placeholder value. Requests without the header pass unchanged.
- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.
- `google/protobuf/empty.proto` import dropped; nothing uses it now.
Questions
- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
block is intercepted. is that clear enough?
follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
the rule conflict check) and regenerate `zz_generated.validation.go`
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes: #823 (there are probably more issues that i need to find)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.
This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.
Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.
Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.
- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Refactors the atelet Restore API so the DATA_ON_GOLDEN base snapshot
travels as a typed restore source instead of a bare string, preparing
for the node-local snapshot cache (#690, #1551): `RestoreRequest` gets
its own external snapshot *source* message, and read-side attributes of
a restore source have a home — today the URI, next (in the M2 cache
work) a sharing property set by the control plane.
**This change is not backward compatible.** `golden_snapshot_uri` is
removed from `RestoreRequest` with no transition: an ateapi from before
this change sends a DATA_ON_GOLDEN restore an updated atelet rejects (no
`base_config`), and an updated ateapi sends one an old atelet rejects
(no `golden_snapshot_uri`). ateapi and atelet must be rolled together.
Fresh starts, FULL/DATA restores of an actor's own snapshot, and local
pause restores without a golden base are unaffected.
Three commits, each buildable and tested:
1. **`ateletpb: give Restore its own external snapshot source message`**
— `ExternalRestoreConfiguration` replaces
`ExternalCheckpointConfiguration` in `RestoreRequest`'s config oneof
(same wire shape; the checkpoint-side write-destination message is
untouched), and `base_config` is added for the DATA_ON_GOLDEN base. No
behavior change: nothing sets or reads the new field yet.
2. **`ateapi, atelet: carry the restore base snapshot in base_config`**
— the cutover: ateapi sets `base_config` on both golden-data resume
paths (external data snapshot, and local pause checkpoint combining with
the golden); atelet reads and validates only `base_config`. The
characteristic test of the ateapi→atelet request and the golden-data
functional test pin the new field.
3. **`ateletpb: remove golden_snapshot_uri from RestoreRequest`** — the
old field is deleted and `base_config` takes its field number, keeping
the numbering dense.
Tested: full ateapi + atelet suites including the controlapi functional
tests; gofmt, golangci-lint, boilerplate clean.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
While benchmarking pause actor I ran into multiple bottlenecks.
I'm currently importing
https://github.com/agent-substrate/substrate/pull/935 into my branch but
https://github.com/agent-substrate/substrate/pull/1876 is another
alternative to solve the first bottleneck.
The 2nd bottleneck (primary target of this PR) is due to fsyncs being
called in writeFileAtomic() for sandbox-assets.json and SystemInfoVolume
files. Because these are new files that require ext4 filesystem metadata
updates, and we have ext4 mounted in ordered mode, calling fsync on the
file + directory causes all other pending writes (such as other
checkpoints) to be flushed. So the file creation ends up being blocked
behind other actor's checkpoint syncs.
These files do not actually need to be flushed to disk. SystemInfo is
recreatable. sandbox-assets.json is written during Run/Resume(), and
read during Checkpoint(). But once the checkpoint is finished, the
manifest is saved with the snapshot. If the node crashes between
Run/Resume and Checkpoint, the actor is crashed anyway.
### Single-Node Pause/Resume Benchmark Summary
**Configuration:** 17 workers, 15 actors, `glutton` (`512 MiB` RAM
capacity, `32 MiB` churn per cycle), `120s` run, `1.0s` wait, `100%`
pause (`0%` suspend)
| Metric | Baseline (`upstream/main`) | + Commit 1 (`os.Link` local
checkpoint) | + Commit 2 (remove `fsync` on ephemeral volumes) |
| :--- | :---: | :---: | :---: |
| **Completed Cycles** (`120s`) | `106.5` (`112` pauses / `101` resumes)
| `160.0` (`165` pauses / `155` resumes) | **`212.5`** (`214` pauses /
`211` resumes) |
| **Cycle Throughput** | `53.3 cycles/min` (`0.89/s`) | `80.0
cycles/min` (`1.33/s`, **+50.2%**) | **`106.3 cycles/min`** (`1.77/s`,
**2.00x**) |
| **`ResumeActor` p50** | `6,000 ms` | `4,100 ms` | **`210 ms`**
(**28.6x faster**) |
| **`ResumeActor` avg** | `7,023 ms` | `4,867 ms` | **`246 ms`**
(**28.5x faster**) |
| **`ResumeActor` p90 / p99** | `15,000 ms` / `19,000 ms` | `9,400 ms` /
`20,000 ms` | **`350 ms` / `590 ms`** |
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Co-authored-by: Jefftree <jeffrey.ying86@live.com>
The purpose field in MintActorCertificate (and the certs it issues) is
no longer needed. The distinction between actors and ateoms-for-actors
is now built into the SPIFFE URI.
Stacked over #1809
Add a Star History section to the README that embeds star-history.com's
live chart, with a dark variant for GitHub's dark theme.
star-history.com renders the image on request, so it stays current;
browsers and GitHub's image proxy cache it for up to a day.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Today `atelet` and `ateom` derive the per-actor directories (oci
bundles, runsc state, pid files, checkpoint and restore state,
durable-dir, system-info and volume roots) from the actor UID through
the same package, `internal/ateompath`.
This PR:
1. Add an `ActorDirs` message to RunWorkloadRequest,
RestoreWorkloadRequest, CheckpointWorkloadRequest and
TerminateWorkloadRequest, and have atelet fill it with the directories
it prepared. **The ateoms do not read it yet.**
2. Split `internal/ateompath` into:
- `cmd/atelet/internal/ateletpath`: atelet's pathes, the per-actor
directories (and `ActorDirs`, built from them) plus the directories only
atelet uses.
- `internal/nodepath`: shared pathes, the base dir both mount, the ateom
socket, the OTLP
sockets, and the netns name.
- `internal/ateompath`: what the ateoms still derive from the actor UID,
each function marked with the `ActorDirs` field it duplicates. atelet no
longer imports it.
---------
This is the first part of #1604 to decouple atelet and ateom shared
pathes. Behavior **does not change**: both sides still compute the same
paths, atelet from `ateletpath` and the ateoms from `ateompath`.
Followup PRs will remove ateompath by making `ateom-gvisor` and
`ateom-microvm` read the `ActorDirs` message from RPC.
Step 3 of #1748. Vocabulary only: nothing emits the event or sets the
epoch yet.
**Event.** `ate.actor.usage_sampled` in `actorevent`, sixteen required
keys via `ateattr`: identity, pool, sandbox class, source,
`ate.stats.kind` (periodic, initial, final), the four measurements,
observation time, epoch. Measurement names follow the
`ate.actor.stats.*` instruments; units are in the registry briefs. Body
is `Actor usage sampled`, the accepted stdout shape change.
**Scope.** The usage stream has its own instrumentation scope, the
lifecycle scope with `/usage` appended, so a consumer can sample or drop
it without touching lifecycle records. `NewEmitter` takes the scope; the
package-level `Log` keeps the lifecycle one, so ateapi is unchanged.
**Epoch.** `ate.actor.epoch` and `WorkloadStatsSample.epoch_unix_nano`:
the unix-nano time an activation began. CPU time is per activation for
every source; lifetime is the sum over epochs of each epoch's highest
value. Memory usage and working set are absolute; peak is as the source
reports it. An ateom that leaves the epoch at zero still sends the raw
guest counter. The measurement comments on fields 8 to 11 say all of
this.
**Registry.** Event group in `events.yaml`, seven attributes in
`metrics.yaml`, `ate.actor.epoch` in the no-actor-identity forbidden
list. The one-record-one-place rule and the collector guide now name
both scopes.
Left for step 4: the event's code anchor moves to the ateom emit site,
atelet's poller comments and the guide's stdout usage section get
rewritten, and the controller starts passing `OTEL_LOGS_EXPORTER` to
worker pods.
Part of #1756.
- Adds `--actor-jwt-issuer`. Unset, it defaults to
`https://idp.<namespace>.svc`, the Service ate-idp-server will serve
discovery from. Before this, `iss` was hardcoded to
`https://api.ate-system.svc`, which points at ateapi's gRPC port.
- Adds `oidcdiscovery.ParseIssuer`, which requires a canonical https URL
with no query, fragment, or user info and strips trailing slashes.
ate-idp-server will use it too, so both emit the same bytes.
- Actor JWT headers now include `typ: JWT`.
- Drops the `actorIdentityJWTIssuer` plumbing into `RPCService`, unused
since #1315.
The default `iss` changes, but `MintActorJWT` has no production caller
yet.
Testing: new unit tests for the parser, the default, the header, and
flag resolution. `TestMintActorJWT_Success` now checks `typ`, `iss`, and
`sub`. It needs Docker, so it hasn't run locally.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (docs
land later in the #1756 series)
This adds dual v1beta1 and v1 PCR support, favoring v1. We have seen
users on v1.37 without v1beta1 but with v1 seeing incomaptibility.
Given our usage of PCR is pretty scoped its not too bad to have the dual
version support.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
Fixes#1835
`atelet` chose the sandbox binaries and pause image for a restore from
the snapshot's `manifest.json` rather than from the control plane. This
PR makes ateapi send them on every restore, the way it already does for
Run.
**Changes**
- **proto:** add `sandbox_assets` to `atelet.RestoreRequest`.
- **ateapi:** `ensureAteletRestored` resolves the ActorTemplate's
SandboxConfig once and sends it on the local restore, the external
restore and Run. A missing SandboxConfig now fails before any atelet
call.
- **atelet:** `sandbox_assets` is required on Restore. Restore selects
the binaries and pause image with `recordFromRequest`, the same path Run
uses. The manifest still supplies the snapshot files, sandbox class and
metrics labels.
Note - The field is required, so after atelet is upgraded, resumes on
that node fail until ate-api-server is upgraded too.
**Testing**
- ateapi: every Run and Restore sends the resolved assets; a missing
SandboxConfig fails on both restore paths.
- atelet: `TestRecordFromRequest`, a validation case for a missing
field, and `TestRestoreUsesRequestSandboxAssets` (checkpoint with pause
image v1, restore with v2, and the actor runs with v2).
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.
v0.1.0 used a form of this, checking it in before v0.2.0
Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
Fixes the kind install on arm64 hosts (Apple Silicon), broken in #1535
- `BuildDockerfileImage` hard-codes `--platform=linux/amd64`, so the
dataplane image's Rust builder stage fails with `exec /bin/sh: exec
format
error` on mac with arm64
The ko images already follow `KO_DEFAULTPLATFORMS`, which the kind
install
sets to the host architecture; the Dockerfile images did not.
**Changes**
- Pass `KO_DEFAULTPLATFORMS` through to `docker buildx build --platform`
for
both Dockerfile images (envoy dataplane, claude-code-multiplex demo),
defaulting to `linux/amd64` when unset.
- Pin the Envoy base image by its multi-platform index digest so
`--platform`
selects the matching manifest.
- [x] Tested (`hack/install-ate-kind.sh --deploy-ate-system` on mac
succeed)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.
It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#1266
Builds on #1689#1283 in particular.
Outstanding:
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~
EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #1590
First of two PRs. This one covers Kubernetes-side hardware discovery;
the follow-up adds the Prometheus harvest.
## What this PR does
The locust runner records no information about the hardware it runs on,
so trial results cannot be normalized across machine types or node
counts. This adds cluster capacity discovery and derives actor density
frontiers from it, persisting both the raw readings and the derived
ratios to `stats.jsonl`.
## Proposed Changes
### Cluster hardware discovery
New module `benchmarking/locust/cluster_facts.py`, kept separate from
`runner.py`. It reads node count, machine type, allocatable cores and
allocatable RAM, and the worker pod count via the official Kubernetes
Python client, which resolves in-cluster and local kubeconfig auth
without a kubectl subprocess.
Node and pod listing is proportional to cluster size, so
`--no-cluster-facts` disables discovery. Discovery is best effort in
either case: an unreachable API server or missing RBAC leaves the facts
null and does not fail the run.
No value is defaulted or inferred. Unmeasured fields are `null`; a
measured zero is `0`. The two remain distinguishable to any consumer.
### Density frontiers
Added to the `trial_summary` row in `stats.jsonl`:
| Field | Meaning |
|---|---|
| `actors_per_node` | peak actors over node count |
| `actors_per_vcpu` | peak actors over allocatable cores |
| `actors_per_gb_ram` | peak actors over allocatable RAM |
| `actors_per_pod_p50` / `_p90` / `_p99` | distribution of actors per
worker pod |
| `aggregate_failure_ratio` | failures over requests, all RPCs |
| `<operation>_failure_ratio` | one per operation Locust reported, e.g.
`dur_dir_write_failure_ratio` |
Actors per pod is reported as a distribution across the run's time
samples (`p50` / `p90` / `p99`) rather than one average, because ramp-up
and custom load shapes have no single user count to call steady.
The raw readings (`machine_type`, `node_count`, `allocatable_cores`,
`allocatable_ram_gb`, `worker_pod_count`) are written alongside the
derived ratios so they can be re-derived without re-running the trial.
### RBAC
Nodes are cluster scoped and require a ClusterRole with `list` on
`nodes`. Pods require only a namespaced Role with `list` in
`benchmark-workloads`. `locust.yaml` and `runner-job.yaml.tmpl` each
define their own separately named pair.
Note that `locust.yaml` now binds a Role in `benchmark-workloads`, so
that namespace must exist before the manifest is applied.
`benchmarking/workloads/deploy.sh` creates it.
`status.json` is unchanged.
## How this was tested
11 unit tests in `test_cluster_facts.py` covering percentile boundaries,
capacity scoped to only the nodes running worker pods, zero worker pods
treated as a valid reading, unreadable facts returning null, and flag
behavior.
Verified against a live GKE cluster. Discovery returned
`c3d-standard-8`, 1 node, 7.91 allocatable cores, 27.73 GB, 5 worker
pods, all matching `kubectl`. The emitted frontiers were re-derived by
hand from `stats_history.csv` and matched. `--no-cluster-facts` returns
all nulls.
Both manifests pass `kubectl apply --dry-run=client`, with no
cluster-scoped name collisions between them.
## References
[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Step 2 of #1748, plus the consumer rule from step 1b.
`LoggingOptions` gains `ExporterConn` and `RelayCapable`, mirroring
`TracingOptions`. Both ateoms initialize logging after metrics with the
relay connection and defer a nil-guarded shutdown. `OTEL_LOGS_EXPORTER`
still defaults to `none`, so no ateom emits a record; this is the
provider the actor events land on when emission moves.
A relay-capable log resource carries `ate.otlp.relay` like traces and
metrics do. `newLoggerProvider` is split from `InitLogging` so the test
asserts that on an emitted record's resource, the way
`TestMeterProviderRelayAttribute` does for metrics.
`docs/observability.md` states the consumer rule that replaced the relay
denylist dropped in #1800: a lifecycle record is authoritative only
under ateapi's resource, because the relay admits only ateom resources.
Follow-up for the emission step, not here: the controller does not
inject `OTEL_LOGS_EXPORTER` into worker pods, so no environment
exercises log-over-relay yet.
Part of #1664, implements #1799.
## Changes
- `ate.worker.state` now has `idle`, `partial`, `at_capacity` and
`unschedulable`. `assigned` is removed.
- ateapi picks the state by comparing allocated actor slots with
capacity, the same check the scheduler makes.
- Draining workers and workers with no reported capacity are
`unschedulable`. This wins over occupancy, so the sum over the states is
still the pool size.
- Every known pool reports all four states, set to 0 when empty.
- Updated the registry, `docs/observability.md` and the
autoscaled-workerpool demo (`assigned` -> `at_capacity`).
## Open questions
- Are we fine with `unschedulable` as a fourth state? cc. @JeffLuoo
Every new worker starts with capacity 0, so `at_capacity` would make an
HPA on it scale up on its own scale-up.
- This breaks queries on `ate_worker_state="assigned"`. Should be still
fine.
- The demo HPA is only right while each worker holds one actor. Pool
utilization needs slot counts, which needs a new instrument (#1664).
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes https://github.com/agent-substrate/substrate/issues/1755
Working on https://github.com/agent-substrate/substrate/issues/1563
Move the embedded OpenFGA authorization engine into
cmd/ateapi/internal/authz and unify its storage layer with ateapi's
PostgreSQL connection pool and Goose migration history:
- Add 000002_openfga.sql to cmd/ateapi/internal/store/atepg/migrations
so OpenFGA tables are created and versioned alongside Substrate tables
in the same schema, guarded by TestOpenFGAMigrationVersionGuard.
- Expose Pool() on atepg.Store and construct OpenFGA via
postgres.NewWithDB(pool, nil, ...) so it shares the existing
pgxpool.Pool without opening a second pool.
- Wrap the OpenFGA Postgres datastore in transactionalDatastore so Read,
ReadPage, ReadAuthorizationModel, and Write participate in
caller-supplied pgx.Tx transactions injected via ContextWithTx.
Part of agent-substrate/substrate#1550.
Adds `kubectl ate update egress-policy <actor-name> -a <atespace> -f
<manifest>`, which replaces an actor's egress policy and prints the
result. To prevent an edit from overwriting a change made since the
read, the server rejects a manifest with `metadata.uid` and
`metadata.version` that mismatch stored ones. On NotFound the command
says whether the actor or the policy is missing; on Aborted (due to
uid/version mismatch) it points back at `get -o yaml`.
### Testing
`go test -race ./cmd/kubectl-ate/...`, `go vet`, `gofmt`. On kind: get,
edit, update round trip; stale version, missing uid and version,
atespace mismatch, and update with no policy all exit 1 with the policy
unchanged.
🤖 This PR was developed with AI assistance. I have reviewed and tested
all changes.
Teardown killed the pause container, and with it the sandbox, before
deleting the actor's other containers. The dead sentry stays a zombie
until the ateom reaps it, which it defers while any runsc command runs,
and runsc delete mistakes the zombie for a live sandbox and fails. Leave
the pause container to cleanupContainers, which deletes it last.
Fixes what is a rare flake on main, that can become more common in
https://github.com/agent-substrate/substrate/pull/1836
Collecting initial feedback on integrating Envoy dynamic modules for
Substrate dataplane.
Envoy Dynamic Modules allow fast iteration and experimentation on
Substrate dataplane. Dynamic modules are used to improve feature
velocity. As requirements and solution are better understood they will
be generalized and moved into Envoy's main repository.
Important points:
- Dynamic modules introduce Rust toolchain into Substrate dev
environment. This PR does not add GH actions for testing Dynamic Modules
during CI builds. This will be added in a subsequent PR.
- This change brings in Dockerfiles and docker buildx. It produces an
image with a versioned Envoy binary (1.39) and accompanying dynamic
modules.
- Docker image registry URL in deployment spec for Envoy datplane is at
this point hardcoded to `localhost:5001`. This will break for anyone
using a different registry. Fixing this will require using envsubst or
equivalent. I'm open to other suggestions.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Yan Avlasov <yavlasov@google.com>
Removes the "not an officially supported Google product" note from the
top of the README, and replaces the "early development / not ready for
production" wording with a plainer pre-1.0 compatibility statement.
The Google-supported product is
[ai-on-gke/substrate-gke](https://github.com/ai-on-gke/substrate-gke).
The Vulnerability Rewards Program statement is unchanged and stays in
[.github/SECURITY.md](https://github.com/agent-substrate/substrate/blob/main/.github/SECURITY.md);
only its "not eligible" link to the removed README note is dropped.
Fixes #1168
* Add DV for GoldenSnapshotStatus and truncate its error_message to fit
* Require ExternalSnapshot.actor_template_uid to be a UUID
* Bound ExternalVolumeTemplate.capacity to 32 characters
Fixes #<issue_number_goes_here>
NA
This PR enhances the tests to validate volume lifecycle during state
transition failures
• **TestDeleteActor_VolumeDeletionFailure_RetrySuccess**- Verifies that
when volume deletion fails during DeleteActor, the actor transitions to
ACTOR_STATE_DELETING and its volumes to STATUS_DELETING. A
subsequent retry successfully finalizes volume deletion and deletes the
actor from the store.
• **TestResumeActor_VolumeAttachFailureAndRetry** - Verifies that when
volume attachment fails during ResumeActor, the actor remains in
ACTOR_STATE_RESUMING with worker assignment and provisioned volumes
intact. Retrying ResumeActor re-attempts volume attachment and
successfully transitions the actor to ACTOR_STATE_RUNNING.
• **TestResumeActor_VolumeAttachFailure_DeleteActor** - Verifies that an
actor stuck in ACTOR_STATE_RESUMING after an attach failure rejects
standard DeleteActor requests, but calling DeleteActor with
AnyState=true successfully cleans up worker assignments and provisioned
volumes.
• **TestResumeActor_MultiVolumePartialAttachFailure_Retry** - Verifies
that when attaching multiple volumes during ResumeActor and one volume
fails, previously attached volumes remain attached and the actor
stays in ACTOR_STATE_RESUMING. A subsequent retry attaches only the
remaining unattached volumes and transitions the actor to
ACTOR_STATE_RUNNING.
• **TestSuspendActor_VolumeDetachFailure_RetrySuccess** - Verifies that
volume detach failures during SuspendActor leave the actor in
ACTOR_STATE_SUSPENDING with its worker assignment held. A subsequent
retry
succeeds in detaching the volume, releasing the worker, and
transitioning the actor to ACTOR_STATE_SUSPENDED.
• **TestSuspendActor_VolumeDetachFailure_DeleteActorAnyState** -
Verifies that an actor stuck in ACTOR_STATE_SUSPENDING due to a volume
detach failure rejects standard deletion, but cleanly detaches volumes
and deletes the actor when called with AnyState=true.
• **TestPauseActor_VolumeLifecycle_DetachAndResumeAttach** - Validates
the volume lifecycle across pause and unpause, confirming external
volumes are detached when transitioning from RUNNING to PAUSED and re-
attached when resumed back to RUNNING.
• **TestPauseActor_VolumeDetachFailure_RetrySuccess** - Verifies that
volume detachment failure during PauseActor leaves the actor in
ACTOR_STATE_PAUSING. A retry of PauseActor re-attempts volume
detachment,
succeeds, and advances the actor state to ACTOR_STATE_PAUSED.
• **TestResumeActor_PausedLocalSnapshotMissing_Crashes** - Verifies that
if local snapshot files on the worker are missing when resuming a PAUSED
actor, the control plane marks the actor ACTOR_STATE_CRASHED
and frees the assigned worker pod.
• **TestDetachActorVolumes** - Tests control plane volume detachment
logic across multiple scenarios, including container mount filtering,
partial detach errors, unassigned workers, and idempotent not-found
responses from plugins.
• **TestCreateActorVolumes (Storage class parameters case)** - Verifies
that parameters configured on a StorageClass (such as disk and
filesystem types) are correctly propagated into the volume context of
newly created external volumes.
• **TestEnsureExternalSnapshotsReleased_DeletePrefixFailure** - Verifies
that transient object store errors during actor deletion preserve
un-deleted snapshot objects and return an error, which cleanly
completes deletion on a retry.
• **TestMountExternalVolumes** - Tests worker-side volume mounting,
verifying host mount directory creation, handling pre-existing paths,
plugin errors, and ensuring that partial failures during multi-volume
mounts do not unmount already mounted volumes.
• **TestVolumeHostDirectoryCleanup** - Verifies worker-side host
directory management, ensuring volume unmounting preserves mount
directories while directory reset removes empty directories and prevents
data
loss on non-empty ones.
• **TestExternalVolume_NodeMigration** - Verifies cross-node actor
migration with external volumes by suspending an actor, evicting its
worker pod, and resuming it on a new worker node while ensuring
application state is preserved
- [ x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## Problem
When `orchestrator.py` runs consecutive benchmark tests in Prow,
`teardown_substrate()` runs `hack/install-ate.sh --delete-ate-system`
(`DeleteAteSystem`), which previously returned immediately after issuing
asynchronous deletion on the `ate-system` namespace without waiting for
termination to complete.
Because draining `atelet` DaemonSets across 15 nodes and `postgres`
takes ~60–65 seconds:
1. The immediately following test runs `deploy_substrate()`
(`EnsureAteSystemNamespace`) ~1 second later while
`namespace/ate-system` is still `Terminating`.
2. `ApplyPath("ate-system-namespace.yaml")` merely updates the
terminating namespace object, and once the old namespace finishes
deleting and transitions to `NotFound`, `WaitNamespaceActive` never
re-applies the manifest and times out after 60s (`waiting for namespace
ate-system to become Active: context deadline exceeded`).
3. As a result, consecutive tests (such as the even-indexed `*_pause` M1
benchmarks) failed during `deploy_substrate()` in ~90% of Prow runs
before Locust could start.
## Fix
- **`DeleteAteSystem` (`cmd/ate-setup/internal/steps/delete.go`)**: Wait
for `namespace/<ns>` deletion (`WaitDeleted`) to complete before
returning from teardown.
- **`EnsureAteSystemNamespace`
(`cmd/ate-setup/internal/steps/env.go`)**: If the target namespace
already exists in `Terminating` phase when deploy starts, wait for it to
finish deleting before applying `ate-system-namespace.yaml` /
`EnsureNamespace`.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
## Summary
Replace broken/stuck `gluttonActor` instances in place when `resume()`
or `hibernate()` fails during benchmark iterations.
## Problem
When `SuspendActor` fails (e.g., due to transient GCS
`ResourceExhausted` errors during cold-bucket ramp-up), `ate-api-server`
can leave the actor permanently stuck in `ACTOR_STATE_SUSPENDING`.
Previously, `gluttonUser` kept the broken `gluttonActor` in its
`u.actors` slice for the remainder of the benchmark run:
- `hibernate()` ignored `SuspendActor` / `PauseActor` errors.
- Subsequent `ResumeActor` calls against that slot repeatedly failed
with `got: ACTOR_STATE_SUSPENDING` for the rest of the test, reducing
effective concurrency and generating continuous failure noise.
## Solution
- Have `hibernate()`, `pause()`, and `suspend()` return `bool`
indicating whether the operation succeeded.
- Add `(*gluttonUser).replaceActor(ctx, broken)` to delete the broken
actor (`AnyState: true`) and replace its slot in `u.actors` with a
freshly created actor (`sb-<uuid>`).
- Invoke `user.replaceActor(ctx, actor)` in `iterate()` when either
`actor.resume(ctx)` or `actor.hibernate(ctx)` fails.
- Add unit test coverage in `lifecycle_test.go`
(`TestGluttonIterate_ReplacesActorOnResumeFailure`).
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
The sweep for a sandbox's leftover cloud-hypervisor and virtiofsd
matched on binary names, but atelet stages both under content-addressed
names, so it never killed anything and a failed boot left its VMM
holding guest RAM. Match on the staged location instead.
Split out from https://github.com/agent-substrate/substrate/pull/1836 as
a small pre-existing fix.
Follow up on the comment -
https://github.com/agent-substrate/substrate/pull/1675#discussion_r4031619230
Stamps `actor_template_uid` onto `ExternalSnapshot` so template
provenance travels with the snapshot itself, rather than being inferred
from the single `ActorStatus.current_actor_template_uid`.
#### Motivation
`current_actor_template_uid` records the template the **last sprint
booted with** (`finalizeRunning` stamps it on every resume). It does
*not* record the template the **snapshot** was captured under. Those two
diverge as soon as an actor runs a sprint that does not produce a new
durable snapshot, and `loadActorForResume` was using the former to
decide whether to force a `DATA`-only restore instead of `FULL`.
Walking an actor through a repoint and a crash:
| Step | Action | `current_…_uid` | `ExternalSnapshot` |
|---|---|---|---|
| **a** | Suspend on template **A** | A | captured on **A** |
| **b** | `UpdateActor` repoints the spec to **B** | A | captured on
**A** |
| **c** | Resume → `A != B`, restores `DATA`, sprint boots on **B** |
**B** | still on **A** |
| **d** | Actor crashes (or is reverted) — no new snapshot taken | B |
still on **A** |
| **e** | Resume → `current == target == B`, so **no repoint is
detected** | B | still on **A** |
At step **e** the old check reports "template not replaced" and restores
the external snapshot in `FULL` — replaying a memory image captured
under **A**'s sandbox on **B**. That is exactly the case the guard
exists to prevent; it was silently defeated by the sprint at step **c**
advancing `current_actor_template_uid` past the snapshot.
*(Note: Paused actors do not face this divergence. Because `UpdateActor`
is only allowed in the `SUSPENDED` state, a paused actor cannot have its
template updated. Therefore, local checkpoints do not need separate
template provenance tracking).*
#### Changes
- **Data Model Updates:** Added the `actor_template_uid` field to
`ExternalSnapshot`.
- **Snapshot Stamping:** Updated the capture logic to stamp external
snapshots with the active sandbox's `ActorTemplate` UID when suspending
an actor or creating an actor from a tag.
- **Resume Evaluation Logic:** Modified the external resume logic to
compare the target template UID against the specific snapshot's UID
(`ExternalSnapshot.actor_template_uid`) rather than the actor's overall
status. Local restores bypass this check entirely since templates cannot
change during a pause.
**Tests**
- `TestResumeActor_AteletWireRequest` (unit): snapshot fixtures now
carry `actor_template_uid`, since provenance is read from the snapshot
rather than from actor status.
- Functional expectations for create/update/resume/pause updated for the
new field on the external snapshot golden files.
- `TestUpdateTemplateLifecycle` (e2e): asserts
`external_snapshot.actor_template_uid` after suspend and re-suspend.
- **e2e (Update Template)**: Added coverage for repoint detection after
a revert (verifying the `resume` → `revert` → `resume` path described in
the table above).
- **e2e (Combined Volumes)**: Added coverage to explicitly verify which
volumes a revert rewinds.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
protoc's Python output puts the serialized descriptor on a single line
which makes any two concurrent ateapi.proto changes conflict.
This is unfortunate as both proto and go generated code usually merges
with no conflict.
Stop checking in the *_pb2*.py files. The locust and nighthawk-ingress
images now generate them in a build stage via
benchmarking/locust/codegen/generate.sh.
Fixes#1824
Follow-up change on a comment in #1335 to add a statusz page in the
credential provider service.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.
Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.
Follow up for the PR - #1675
Issue - #1556
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This change addresses six ways CI pipeline could report false success or
hang on broken infrastructure. Importantly, this change does two things
to reduce load on infrastructure:
1. It prevents jobs from running for [the default action limit of
360m](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax#jobsjob_idtimeout-minutes)
by pinning the timeout to `45m`.
2. It limits pull_request jobs from continuing execution once superseded
by new changes.
- **Silent container test skipping**: Tests now explicitly fail if `CI`
or `REQUIRE_DOCKER` is set (while still skipping on local machines
lacking Docker), with the check implemented in `dockerenv` to avoid an
import cycle between `storetest` and `atepg`.
- **Unbounded trust bundle wait**: Enforced a shared 120-second timeout
across both bundles (overridable via `ATE_INSTALL_TRUST_BUNDLE_TIMEOUT`)
and added diagnostic dumping of bundles, controller pods, and logs
before returning a non-zero exit code.
- **Missing sandbox preflight validation**: Added early preflight checks
for `/dev/kvm` and `SandboxConfig/microvm` that fail fast and print
actionable remediation instructions.
- **Skipped migration checks on main**: Configured the migration
immutability check to run on pushes to main to catch modified migrations
at the point of merge.
- **Missing job timeouts and concurrency limits**: Defined explicit
timeout-minutes (45m and 120m bounds) and added concurrency groups that
automatically cancel superseded pull request runs without canceling runs
on main.
Fixes#1747
- [x] Tests pass
- Silent container skipping, unbounded trust bundles, and missing
sandbox were all forced locally and confirmed to exist with changes here
resolving each.
- The latter half of the scenarios exist in CI only due to being GH
Action trigger issues.
Today CI depends on process exit codes. This leads to a few issues with
reporting results. For example, a green build cannot be distinguished
from one where no tests ran at all; a `-run` filter that matches
nothing; an e2e suite short-circuiting without `--e2e`; or a
testcontainer skip: all currently exit 0.
This change adds gotestsum as a pinned tool module allowing opt-in for
JUnit reporting in both test runners by using `E2E_JUNIT_FILE` and
`ROOT_JUNIT_FILE` respectively. When left unset, the runners exec `go
test` directly and gotestsum is never built leaving local runs unchanged
for now.
Additionally, this change pins two execution bounds that `go test` would
otherwise infer. Both remain overridable via environment variables as
listed below.
- `-p`: system defaults didn't align with cluster capacity so pinning
the value avoids over allocation. (override with `E2E_PARALLELISM`).
- `-timeout`: Suite defaults conflicted with some E2E test timings
causing some tests that are expected to timeout at the same time to
abort rather than fail as expected. Setting an 30m default avoids the
conlficting timeouts producing clean failure signals (override with
`E2E_TIMEOUT`).
Fixes#1685
We need to draw a strong distinction between credentials that will be
wielded directly by actor logic, and credentials used by system
components on behalf of an actor. This will help us avoid attacks where
an actor pretends to be a system component.
Ateom certificates are requested by ateom/atelet and used for the
atunnel connection to the egress gateway. They have a SPIFFE URI like
`spiffe://${trustdomain}/ateom-for-actor/${atespace}/${actor}`.
Actor certificates are requested by the egress gateway and used for
opening outbound requests on behalf of the actor. They have a SPIFFE URI
like `spiffe://${trustdomain}/actor/${atespace}/${actor}`. Note, this
use case is not actually implemented yet.
Part of #1743, follow up of #1803 which added the `debug_redact` labels
but nothing was reading them yet.
The unary interceptors log every request and response body. Until now
the only redaction was to clear any field named `env`, so for example
`MintActorJWTResponse.actor_jwt` (a bearer token) was going to the
ateapi log on every mint call.
Now the interceptor follows the label instead of the field name. Before
logging it clones the message and masks every field with `[debug_redact
= true]` in the copy: singular strings become `[REDACTED]`, other kinds
(bytes, repeated, map, message) are cleared, and it recurses into nested
messages, repeated messages and map values (the old walker skipped
maps). The message returned to the client is never touched. The
hardcoded `env` check is gone since both `EnvVar.value` and
`EnvEntry.value` carry the label now.
One visible change: env var names now show up in the log next to
`[REDACTED]`, before the whole env list was dropped. I think this is
better for debugging and the values are still gone.
This is an incremental step and only fixes the known leak in the
interceptor. It keeps today's structure, clone + walk on every request
and response at this one call site. Two things are left for follow up
PRs on purpose: moving the redaction into the shared slog handler
(`internal/contextlogging`) so any slog call with a proto is covered,
and the performance work (per message type cache so clean messages are
not cloned at all, scan-then-rebuild handler). I benchmarked that
already, numbers are in the doc linked from #1743. Without the cache
this PR costs about the same as today on clean messages and a bit more
on messages with secrets, about 100ns per masked value, which is fine
for now.
The postgres connection string in ateapi's startup flag log is still
open from the P0 list, that will be a separate small PR.
Step 1 of #1748. The relay forwards `LogsService` the way it forwards
traces and metrics: source gate, verbatim payload, atelet's own
`OTEL_EXPORTER_OTLP_LOGS_HEADERS` upstream.
No content check on the records. Identity spoofing through the relay is
handled in #1748 by the consumer rule (a lifecycle record counts only
under ateapi's resource, which the source gate already enforces) and
per-pod sockets (#741), not by a denylist here.
**Endpoint and compression resolution** (second commit, from review).
The relay read the generic, traces, and metrics OTLP variables and
ignored the logs ones, so a logs-specific collector or compression
setting was silently overridden. All three signals now go through one
resolver: each falls back to the generic, and every resolved value must
agree, since the relay has one upstream connection. The conflict error
names the variable each value came from; a signal that falls back names
the generic variable, not its own unset one.
One behavior change: `traces=A, metrics=A, generic=B` with logs unset
used to pass and now fails at atelet startup, because logs falls back to
`B` and that is a two-collector configuration.
Tests: an ateom log batch arrives with the ateom's resource; the source
gate, mixed-batch, header-isolation, and per-signal header tests cover
logs; resolver tables gained logs rows including the fallback case. Docs
say the relay carries logs, traces, and metrics.
Fixes https://github.com/agent-substrate/substrate/issues/292
This flips the condition to crash actors: any errors returned by atelet
will cause actors to crash by default.
If we later discover specific error classes that are safe to retry, we
can explicitly mark them as retriable.