1065 Commits
Author SHA1 Message Date
Bowei Du 22efea18a9 Clean up packages in ateomnet (#1897)
* Move the sandbox DNS code into its own package
* Move the network namespace primitives into their own package
2026-09-28 17:36:52 +00:00
Lior Lieberman 6621f2b4b4 egress policy enhancements for GA (#1751)
**BREAKING** API changes prior in preparation for GA. 

~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow


- `EgressRule` is now a union of protocols: `http`, `https`,
  `tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
  Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
  listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
  request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
  `tls_passthrough`. Same wildcard grammar as before, documented once on
  `HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
  once on `HTTPRule.ports`.

- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
  placeholder value. Requests without the header pass unchanged.

- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.

- `google/protobuf/empty.proto` import dropped; nothing uses it now.

Questions

- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
  block is intercepted. is that clear enough?

follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
  the rule conflict check) and regenerate `zz_generated.validation.go`



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Fixes: #823 (there are probably more issues that i need to find)
2026-09-28 17:30:07 +00:00
Eitan Yarmush ed6d2a1fc8 Update agentgateway for split actor and ateom identities (#1923)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.

This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.

Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.

Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.

- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-26 23:02:14 +00:00
Dmitry Berkovich 1d7ca8ced0 ateletpb: carry the Restore base snapshot as a typed source, dropping golden_snapshot_uri (#1602)
Refactors the atelet Restore API so the DATA_ON_GOLDEN base snapshot
travels as a typed restore source instead of a bare string, preparing
for the node-local snapshot cache (#690, #1551): `RestoreRequest` gets
its own external snapshot *source* message, and read-side attributes of
a restore source have a home — today the URI, next (in the M2 cache
work) a sharing property set by the control plane.

**This change is not backward compatible.** `golden_snapshot_uri` is
removed from `RestoreRequest` with no transition: an ateapi from before
this change sends a DATA_ON_GOLDEN restore an updated atelet rejects (no
`base_config`), and an updated ateapi sends one an old atelet rejects
(no `golden_snapshot_uri`). ateapi and atelet must be rolled together.
Fresh starts, FULL/DATA restores of an actor's own snapshot, and local
pause restores without a golden base are unaffected.

Three commits, each buildable and tested:

1. **`ateletpb: give Restore its own external snapshot source message`**
— `ExternalRestoreConfiguration` replaces
`ExternalCheckpointConfiguration` in `RestoreRequest`'s config oneof
(same wire shape; the checkpoint-side write-destination message is
untouched), and `base_config` is added for the DATA_ON_GOLDEN base. No
behavior change: nothing sets or reads the new field yet.
2. **`ateapi, atelet: carry the restore base snapshot in base_config`**
— the cutover: ateapi sets `base_config` on both golden-data resume
paths (external data snapshot, and local pause checkpoint combining with
the golden); atelet reads and validates only `base_config`. The
characteristic test of the ateapi→atelet request and the golden-data
functional test pin the new field.
3. **`ateletpb: remove golden_snapshot_uri from RestoreRequest`** — the
old field is deleted and `base_config` takes its field number, keeping
the numbering dense.

Tested: full ateapi + atelet suites including the controlapi functional
tests; gofmt, golangci-lint, boilerplate clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-26 14:35:20 +00:00
Taahir Ahmed c7b5469968 ate-api: Move authn package into server-specific internal package (#1909)
Prefactor for unified config file
2026-09-26 01:34:20 +00:00
Taahir Ahmed fe82dae3f3 serverboot.Fatal: Allow additional slog arguments (#1908)
Prefactor for unified config file.
2026-09-26 01:29:52 +00:00
Michelle AuandJefftree d3aec58cdd Remove unnecessary fsync during ResumeActor (#1883)
While benchmarking pause actor I ran into multiple bottlenecks.

I'm currently importing
https://github.com/agent-substrate/substrate/pull/935 into my branch but
https://github.com/agent-substrate/substrate/pull/1876 is another
alternative to solve the first bottleneck.

The 2nd bottleneck (primary target of this PR) is due to fsyncs being
called in writeFileAtomic() for sandbox-assets.json and SystemInfoVolume
files. Because these are new files that require ext4 filesystem metadata
updates, and we have ext4 mounted in ordered mode, calling fsync on the
file + directory causes all other pending writes (such as other
checkpoints) to be flushed. So the file creation ends up being blocked
behind other actor's checkpoint syncs.

These files do not actually need to be flushed to disk. SystemInfo is
recreatable. sandbox-assets.json is written during Run/Resume(), and
read during Checkpoint(). But once the checkpoint is finished, the
manifest is saved with the snapshot. If the node crashes between
Run/Resume and Checkpoint, the actor is crashed anyway.

### Single-Node Pause/Resume Benchmark Summary
**Configuration:** 17 workers, 15 actors, `glutton` (`512 MiB` RAM
capacity, `32 MiB` churn per cycle), `120s` run, `1.0s` wait, `100%`
pause (`0%` suspend)

| Metric | Baseline (`upstream/main`) | + Commit 1 (`os.Link` local
checkpoint) | + Commit 2 (remove `fsync` on ephemeral volumes) |
| :--- | :---: | :---: | :---: |
| **Completed Cycles** (`120s`) | `106.5` (`112` pauses / `101` resumes)
| `160.0` (`165` pauses / `155` resumes) | **`212.5`** (`214` pauses /
`211` resumes) |
| **Cycle Throughput** | `53.3 cycles/min` (`0.89/s`) | `80.0
cycles/min` (`1.33/s`, **+50.2%**) | **`106.3 cycles/min`** (`1.77/s`,
**2.00x**) |
| **`ResumeActor` p50** | `6,000 ms` | `4,100 ms` | **`210 ms`**
(**28.6x faster**) |
| **`ResumeActor` avg** | `7,023 ms` | `4,867 ms` | **`246 ms`**
(**28.5x faster**) |
| **`ResumeActor` p90 / p99** | `15,000 ms` / `19,000 ms` | `9,400 ms` /
`20,000 ms` | **`350 ms` / `590 ms`** |


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Jefftree <jeffrey.ying86@live.com>
2026-09-25 23:52:04 +00:00
Taahir Ahmed 7f7a40c1b9 ateom/actor ID split: Clean up purpose from MintActorCertificate (#1810)
The purpose field in MintActorCertificate (and the certs it issues) is
no longer needed. The distinction between actors and ateoms-for-actors
is now built into the SPIFFE URI.

Stacked over #1809
2026-09-25 23:45:05 +00:00
botengyao d2b29a5fa2 docs: add a star history chart to the README (#1894)
Add a Star History section to the README that embeds star-history.com's
live chart, with a dark variant for GitHub's dark theme.
star-history.com renders the image on request, so it stays current;
browsers and GitHub's image proxy cache it for up to a day.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-25 21:44:01 +00:00
Taahir Ahmed 8d6be5fb4e ateom/actor ID split: Ateom uses MintAteomActorCertificate (#1809)
Update atelet/ateom/atunnel to use the new MintAteomActorCertificate,
and correct anything that was coded to expect the old actor SPIFFE URI.
2026-09-25 21:08:23 +00:00
Haven Xia a0b680d72e Add ActorDirs message to carry per-actor directories and and split the path packages by owner (#1730)
Today `atelet` and `ateom` derive the per-actor directories (oci
bundles, runsc state, pid files, checkpoint and restore state,
durable-dir, system-info and volume roots) from the actor UID through
the same package, `internal/ateompath`.

This PR: 

1. Add an `ActorDirs` message to RunWorkloadRequest,
RestoreWorkloadRequest, CheckpointWorkloadRequest and
TerminateWorkloadRequest, and have atelet fill it with the directories
it prepared. **The ateoms do not read it yet.**

2. Split `internal/ateompath` into:
- `cmd/atelet/internal/ateletpath`: atelet's pathes, the per-actor
directories (and `ActorDirs`, built from them) plus the directories only
atelet uses.
- `internal/nodepath`: shared pathes, the base dir both mount, the ateom
socket, the OTLP
     sockets, and the netns name.
- `internal/ateompath`: what the ateoms still derive from the actor UID,
each function marked with the `ActorDirs` field it duplicates. atelet no
longer imports it.

---------

This is the first part of #1604 to decouple atelet and ateom shared
pathes. Behavior **does not change**: both sides still compute the same
paths, atelet from `ateletpath` and the ateoms from `ateompath`.

Followup PRs will remove ateompath by making `ateom-gvisor` and
`ateom-microvm` read the `ActorDirs` message from RPC.
2026-09-25 21:03:44 +00:00
Tim Bai 03ba5821ec actorevent: define the actor usage event and the activation epoch (#1881)
Step 3 of #1748. Vocabulary only: nothing emits the event or sets the
epoch yet.

**Event.** `ate.actor.usage_sampled` in `actorevent`, sixteen required
keys via `ateattr`: identity, pool, sandbox class, source,
`ate.stats.kind` (periodic, initial, final), the four measurements,
observation time, epoch. Measurement names follow the
`ate.actor.stats.*` instruments; units are in the registry briefs. Body
is `Actor usage sampled`, the accepted stdout shape change.

**Scope.** The usage stream has its own instrumentation scope, the
lifecycle scope with `/usage` appended, so a consumer can sample or drop
it without touching lifecycle records. `NewEmitter` takes the scope; the
package-level `Log` keeps the lifecycle one, so ateapi is unchanged.

**Epoch.** `ate.actor.epoch` and `WorkloadStatsSample.epoch_unix_nano`:
the unix-nano time an activation began. CPU time is per activation for
every source; lifetime is the sum over epochs of each epoch's highest
value. Memory usage and working set are absolute; peak is as the source
reports it. An ateom that leaves the epoch at zero still sends the raw
guest counter. The measurement comments on fields 8 to 11 say all of
this.

**Registry.** Event group in `events.yaml`, seven attributes in
`metrics.yaml`, `ate.actor.epoch` in the no-actor-identity forbidden
list. The one-record-one-place rule and the collector guide now name
both scopes.

Left for step 4: the event's code anchor moves to the ateom emit site,
atelet's poller comments and the guide's stdout usage section get
rewritten, and the controller starts passing `OTEL_LOGS_EXPORTER` to
worker pods.
2026-09-25 19:54:15 +00:00
Max Thompson d3a556aae3 ateapi: make the actor JWT issuer configurable and add typ to the header (#1834)
Part of #1756.

- Adds `--actor-jwt-issuer`. Unset, it defaults to
`https://idp.<namespace>.svc`, the Service ate-idp-server will serve
discovery from. Before this, `iss` was hardcoded to
`https://api.ate-system.svc`, which points at ateapi's gRPC port.
- Adds `oidcdiscovery.ParseIssuer`, which requires a canonical https URL
with no query, fragment, or user info and strips trailing slashes.
ate-idp-server will use it too, so both emit the same bytes.
- Actor JWT headers now include `typ: JWT`.
- Drops the `actorIdentityJWTIssuer` plumbing into `RPCService`, unused
since #1315.

The default `iss` changes, but `MintActorJWT` has no production caller
yet.

Testing: new unit tests for the parser, the default, the header, and
flag resolution. `TestMintActorJWT_Success` now checks `typ`, `iss`, and
`sub`. It needs Docker, so it hasn't run locally.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (docs
land later in the #1756 series)
2026-09-25 19:22:25 +00:00
John Howard 10a1bfb2f5 PodCertificateRequest v1 support (#1829)
This adds dual v1beta1 and v1 PCR support, favoring v1. We have seen
users on v1.37 without v1beta1 but with v1 seeing incomaptibility.

Given our usage of PCR is pretty scoped its not too bad to have the dual
version support.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
v0.2.0
2026-09-25 17:29:49 +00:00
Max Smythe 8b5d9ff0ef Tolerations for Workload Types (#1874)
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-09-25 16:45:02 +00:00
shrutiyam-glitch 54cba01a8e feat(ateapi): send resolved sandbox assets during restore (#1866)
Fixes #1835 

`atelet` chose the sandbox binaries and pause image for a restore from
the snapshot's `manifest.json` rather than from the control plane. This
PR makes ateapi send them on every restore, the way it already does for
Run.

**Changes**
- **proto:** add `sandbox_assets` to `atelet.RestoreRequest`.
- **ateapi:** `ensureAteletRestored` resolves the ActorTemplate's
SandboxConfig once and sends it on the local restore, the external
restore and Run. A missing SandboxConfig now fails before any atelet
call.
- **atelet:** `sandbox_assets` is required on Restore. Restore selects
the binaries and pause image with `recordFromRequest`, the same path Run
uses. The manifest still supplies the snapshot files, sandbox class and
metrics labels.

Note - The field is required, so after atelet is upgraded, resumes on
that node fail until ate-api-server is upgraded too.

**Testing**
- ateapi: every Run and Restore sends the resolved assets; a missing
SandboxConfig fails on both restore paths.
- atelet: `TestRecordFromRequest`, a validation case for a missing
field, and `TestRestoreUsesRequestSandboxAssets` (checkpoint with pause
image v1, restore with v2, and the actor runs with v2).

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-25 16:06:02 +00:00
Luiz Oliveira e87c55fc2e Add the crash reason and timestamp to the ActorStatus (#1867)
Added a new field to actor status that has the crash reason (e.g., the
atelet response error) and the timestamp of the crash for better UX.
2026-09-25 15:33:08 +00:00
Luiz Oliveira d72edfbb3c Rename LocalSnapshotInfo to LocalSnapshot (#1856)
For consistency with ExternalSnapshot

#1378
2026-09-25 14:20:42 +00:00
Benjamin Elder 22ba860e0f add get-contributors.sh for releases (#1870)
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.

v0.1.0 used a form of this, checking it in before v0.2.0

Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
2026-09-25 01:38:04 +00:00
Yufan Su 5ebccf3897 Fix the kind install on arm64 hosts (#1869)
Fixes the kind install on arm64 hosts (Apple Silicon), broken in #1535

- `BuildDockerfileImage` hard-codes `--platform=linux/amd64`, so the
dataplane image's Rust builder stage fails with `exec /bin/sh: exec
format
  error` on mac with arm64

The ko images already follow `KO_DEFAULTPLATFORMS`, which the kind
install
sets to the host architecture; the Dockerfile images did not.

**Changes**

- Pass `KO_DEFAULTPLATFORMS` through to `docker buildx build --platform`
for
  both Dockerfile images (envoy dataplane, claude-code-multiplex demo),
  defaulting to `linux/amd64` when unset.
- Pin the Envoy base image by its multi-platform index digest so
`--platform`
  selects the matching manifest.

- [x] Tested (`hack/install-ate-kind.sh --deploy-ate-system` on mac
succeed)
2026-09-25 01:36:37 +00:00
Max Smythe d6d2a0fa1d Add the option to install a large-cluster manifest and cordon control plane to dedicated machines (#1632)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.

It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 23:39:01 +00:00
Benjamin Elder 14c0c136bc multi actor worker support (#1836)
Fixes #1266 

Builds on #1689 #1283 in particular.

Outstanding: 
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~

EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-24 21:25:22 +00:00
sfunkenhauser fa24dcc514 Add plumbing to allow ateom to suspend actors (#1779)
Part of #483 

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-09-24 20:59:46 +00:00
Nishanth Kotla 3198991d23 benchmarking/locust: record cluster hardware facts and density frontiers (#1723)
Part of #1590

First of two PRs. This one covers Kubernetes-side hardware discovery;
the follow-up adds the Prometheus harvest.

## What this PR does

The locust runner records no information about the hardware it runs on,
so trial results cannot be normalized across machine types or node
counts. This adds cluster capacity discovery and derives actor density
frontiers from it, persisting both the raw readings and the derived
ratios to `stats.jsonl`.

## Proposed Changes

### Cluster hardware discovery

New module `benchmarking/locust/cluster_facts.py`, kept separate from
`runner.py`. It reads node count, machine type, allocatable cores and
allocatable RAM, and the worker pod count via the official Kubernetes
Python client, which resolves in-cluster and local kubeconfig auth
without a kubectl subprocess.

Node and pod listing is proportional to cluster size, so
`--no-cluster-facts` disables discovery. Discovery is best effort in
either case: an unreachable API server or missing RBAC leaves the facts
null and does not fail the run.

No value is defaulted or inferred. Unmeasured fields are `null`; a
measured zero is `0`. The two remain distinguishable to any consumer.

### Density frontiers

Added to the `trial_summary` row in `stats.jsonl`:

| Field | Meaning |
|---|---|
| `actors_per_node` | peak actors over node count |
| `actors_per_vcpu` | peak actors over allocatable cores |
| `actors_per_gb_ram` | peak actors over allocatable RAM |
| `actors_per_pod_p50` / `_p90` / `_p99` | distribution of actors per
worker pod |
| `aggregate_failure_ratio` | failures over requests, all RPCs |
| `<operation>_failure_ratio` | one per operation Locust reported, e.g.
`dur_dir_write_failure_ratio` |

Actors per pod is reported as a distribution across the run's time
samples (`p50` / `p90` / `p99`) rather than one average, because ramp-up
and custom load shapes have no single user count to call steady.

The raw readings (`machine_type`, `node_count`, `allocatable_cores`,
`allocatable_ram_gb`, `worker_pod_count`) are written alongside the
derived ratios so they can be re-derived without re-running the trial.

### RBAC

Nodes are cluster scoped and require a ClusterRole with `list` on
`nodes`. Pods require only a namespaced Role with `list` in
`benchmark-workloads`. `locust.yaml` and `runner-job.yaml.tmpl` each
define their own separately named pair.

Note that `locust.yaml` now binds a Role in `benchmark-workloads`, so
that namespace must exist before the manifest is applied.
`benchmarking/workloads/deploy.sh` creates it.

`status.json` is unchanged.

## How this was tested

11 unit tests in `test_cluster_facts.py` covering percentile boundaries,
capacity scoped to only the nodes running worker pods, zero worker pods
treated as a valid reading, unreadable facts returning null, and flag
behavior.

Verified against a live GKE cluster. Discovery returned
`c3d-standard-8`, 1 node, 7.91 allocatable cores, 27.73 GB, 5 worker
pods, all matching `kubectl`. The emitted frontiers were re-derived by
hand from `stats_history.csv` and matched. `--no-cluster-facts` returns
all nulls.

Both manifests pass `kubectl apply --dry-run=client`, with no
cluster-scoped name collisions between them.

## References

[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-24 20:55:29 +00:00
Tim Bai 93e691d9b1 serverboot: give the ateoms a LoggerProvider over the relay (#1857)
Step 2 of #1748, plus the consumer rule from step 1b.

`LoggingOptions` gains `ExporterConn` and `RelayCapable`, mirroring
`TracingOptions`. Both ateoms initialize logging after metrics with the
relay connection and defer a nil-guarded shutdown. `OTEL_LOGS_EXPORTER`
still defaults to `none`, so no ateom emits a record; this is the
provider the actor events land on when emission moves.

A relay-capable log resource carries `ate.otlp.relay` like traces and
metrics do. `newLoggerProvider` is split from `InitLogging` so the test
asserts that on an emitted record's resource, the way
`TestMeterProviderRelayAttribute` does for metrics.

`docs/observability.md` states the consumer rule that replaced the relay
denylist dropped in #1800: a lifecycle record is authoritative only
under ateapi's resource, because the relay admits only ateom resources.

Follow-up for the emission step, not here: the controller does not
inject `OTEL_LOGS_EXPORTER` into worker pods, so no environment
exercises log-over-relay yet.
2026-09-24 20:40:19 +00:00
Krisztian F 945e44a5ec (ateapi): report worker occupancy by actor slots in ate.worker.state (#1854)
Part of #1664, implements #1799.

## Changes

- `ate.worker.state` now has `idle`, `partial`, `at_capacity` and
`unschedulable`. `assigned` is removed.
- ateapi picks the state by comparing allocated actor slots with
capacity, the same check the scheduler makes.
- Draining workers and workers with no reported capacity are
`unschedulable`. This wins over occupancy, so the sum over the states is
still the pool size.
- Every known pool reports all four states, set to 0 when empty.
- Updated the registry, `docs/observability.md` and the
autoscaled-workerpool demo (`assigned` -> `at_capacity`).

## Open questions

- Are we fine with `unschedulable` as a fourth state? cc. @JeffLuoo
Every new worker starts with capacity 0, so `at_capacity` would make an
HPA on it scale up on its own scale-up.
- This breaks queries on `ate_worker_state="assigned"`. Should be still
fine.
- The demo HPA is only right while each worker holds one actor. Pool
utilization needs slot counts, which needs a new instrument (#1664).

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-24 20:17:26 +00:00
Shruti Nair 5fb129d87d authz: manage OpenFGA schema in Goose migrations and share pgxpool/tx (#1839)
Fixes https://github.com/agent-substrate/substrate/issues/1755
Working on https://github.com/agent-substrate/substrate/issues/1563

Move the embedded OpenFGA authorization engine into
cmd/ateapi/internal/authz and unify its storage layer with ateapi's
PostgreSQL connection pool and Goose migration history:

- Add 000002_openfga.sql to cmd/ateapi/internal/store/atepg/migrations
so OpenFGA tables are created and versioned alongside Substrate tables
in the same schema, guarded by TestOpenFGAMigrationVersionGuard.
- Expose Pool() on atepg.Store and construct OpenFGA via
postgres.NewWithDB(pool, nil, ...) so it shares the existing
pgxpool.Pool without opening a second pool.
- Wrap the OpenFGA Postgres datastore in transactionalDatastore so Read,
ReadPage, ReadAuthorizationModel, and Write participate in
caller-supplied pgx.Tx transactions injected via ContextWithTx.
2026-09-24 19:36:07 +00:00
Yuan Gao d7d70420db kubectl-ate: add update egress-policy (#1813)
Part of agent-substrate/substrate#1550.

Adds `kubectl ate update egress-policy <actor-name> -a <atespace> -f
<manifest>`, which replaces an actor's egress policy and prints the
result. To prevent an edit from overwriting a change made since the
read, the server rejects a manifest with `metadata.uid` and
`metadata.version` that mismatch stored ones. On NotFound the command
says whether the actor or the policy is missing; on Aborted (due to
uid/version mismatch) it points back at `get -o yaml`.

### Testing

`go test -race ./cmd/kubectl-ate/...`, `go vet`, `gofmt`. On kind: get,
edit, update round trip; stale version, missing uid and version,
atespace mismatch, and update with no policy all exit 1 with the policy
unchanged.

🤖 This PR was developed with AI assistance. I have reviewed and tested
all changes.
2026-09-24 17:51:08 +00:00
Benjamin Elder 82a777186a ateom-gvisor: delete an actor's containers before its sandbox (#1847)
Teardown killed the pause container, and with it the sandbox, before
deleting the actor's other containers. The dead sentry stays a zombie
until the ateom reaps it, which it defers while any runsc command runs,
and runsc delete mistakes the zombie for a live sandbox and fails. Leave
the pause container to cleanupContainers, which deletes it last.

Fixes what is a rare flake on main, that can become more common in
https://github.com/agent-substrate/substrate/pull/1836
2026-09-24 17:38:37 +00:00
yanavlasov 30e6d33daa Initial commit of Envoy Dynamic Modules (#1535)
Collecting initial feedback on integrating Envoy dynamic modules for
Substrate dataplane.

Envoy Dynamic Modules allow fast iteration and experimentation on
Substrate dataplane. Dynamic modules are used to improve feature
velocity. As requirements and solution are better understood they will
be generalized and moved into Envoy's main repository.

Important points:

- Dynamic modules introduce Rust toolchain into Substrate dev
environment. This PR does not add GH actions for testing Dynamic Modules
during CI builds. This will be added in a subsequent PR.
- This change brings in Dockerfiles and docker buildx. It produces an
image with a versioned Envoy binary (1.39) and accompanying dynamic
modules.
- Docker image registry URL in deployment spec for Envoy datplane is at
this point hardcoded to `localhost:5001`. This will break for anyone
using a different registry. Fixing this will require using envsubst or
equivalent. I'm open to other suggestions.


- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Yan Avlasov <yavlasov@google.com>
2026-09-24 17:18:48 +00:00
Kevin Steuer 97d565f3eb docs: drop the Google product disclaimer from the README (#1862)
Removes the "not an officially supported Google product" note from the
top of the README, and replaces the "early development / not ready for
production" wording with a plainer pre-1.0 compatibility statement.

The Google-supported product is
[ai-on-gke/substrate-gke](https://github.com/ai-on-gke/substrate-gke).

The Vulnerability Rewards Program statement is unchanged and stays in
[.github/SECURITY.md](https://github.com/agent-substrate/substrate/blob/main/.github/SECURITY.md);
only its "not eligible" link to the removed README note is dropped.
2026-09-24 17:15:53 +00:00
Luiz Oliveira 34f3003326 Add DV for GoldenSnapshotStatus and unbounded strings (#1831)
Fixes  #1168

* Add DV for GoldenSnapshotStatus and truncate its error_message to fit
* Require ExternalSnapshot.actor_template_uid to be a UUID
* Bound ExternalVolumeTemplate.capacity to 32 characters
2026-09-24 16:59:12 +00:00
Sneha-at 444d7631d1 enhance tests to validate volume lifecycle for actor state changes (#1581)
Fixes #<issue_number_goes_here>
NA

This PR enhances the tests to validate volume lifecycle during state
transition failures

• **TestDeleteActor_VolumeDeletionFailure_RetrySuccess**- Verifies that
when volume deletion fails during DeleteActor, the actor transitions to
ACTOR_STATE_DELETING and its volumes to STATUS_DELETING. A
subsequent retry successfully finalizes volume deletion and deletes the
actor from the store.
• **TestResumeActor_VolumeAttachFailureAndRetry** - Verifies that when
volume attachment fails during ResumeActor, the actor remains in
ACTOR_STATE_RESUMING with worker assignment and provisioned volumes
intact. Retrying ResumeActor re-attempts volume attachment and
successfully transitions the actor to ACTOR_STATE_RUNNING.
• **TestResumeActor_VolumeAttachFailure_DeleteActor** - Verifies that an
actor stuck in ACTOR_STATE_RESUMING after an attach failure rejects
standard DeleteActor requests, but calling DeleteActor with
AnyState=true successfully cleans up worker assignments and provisioned
volumes.
• **TestResumeActor_MultiVolumePartialAttachFailure_Retry** - Verifies
that when attaching multiple volumes during ResumeActor and one volume
fails, previously attached volumes remain attached and the actor
stays in ACTOR_STATE_RESUMING. A subsequent retry attaches only the
remaining unattached volumes and transitions the actor to
ACTOR_STATE_RUNNING.
• **TestSuspendActor_VolumeDetachFailure_RetrySuccess** - Verifies that
volume detach failures during SuspendActor leave the actor in
ACTOR_STATE_SUSPENDING with its worker assignment held. A subsequent
retry
succeeds in detaching the volume, releasing the worker, and
transitioning the actor to ACTOR_STATE_SUSPENDED.
• **TestSuspendActor_VolumeDetachFailure_DeleteActorAnyState** -
Verifies that an actor stuck in ACTOR_STATE_SUSPENDING due to a volume
detach failure rejects standard deletion, but cleanly detaches volumes
and deletes the actor when called with AnyState=true.
• **TestPauseActor_VolumeLifecycle_DetachAndResumeAttach** - Validates
the volume lifecycle across pause and unpause, confirming external
volumes are detached when transitioning from RUNNING to PAUSED and re-
attached when resumed back to RUNNING.
• **TestPauseActor_VolumeDetachFailure_RetrySuccess** - Verifies that
volume detachment failure during PauseActor leaves the actor in
ACTOR_STATE_PAUSING. A retry of PauseActor re-attempts volume
detachment,
succeeds, and advances the actor state to ACTOR_STATE_PAUSED.
• **TestResumeActor_PausedLocalSnapshotMissing_Crashes** - Verifies that
if local snapshot files on the worker are missing when resuming a PAUSED
actor, the control plane marks the actor ACTOR_STATE_CRASHED
and frees the assigned worker pod.
• **TestDetachActorVolumes** - Tests control plane volume detachment
logic across multiple scenarios, including container mount filtering,
partial detach errors, unassigned workers, and idempotent not-found
responses from plugins.
• **TestCreateActorVolumes (Storage class parameters case)** - Verifies
that parameters configured on a StorageClass (such as disk and
filesystem types) are correctly propagated into the volume context of
newly created external volumes.
• **TestEnsureExternalSnapshotsReleased_DeletePrefixFailure** - Verifies
that transient object store errors during actor deletion preserve
un-deleted snapshot objects and return an error, which cleanly
completes deletion on a retry.
• **TestMountExternalVolumes** - Tests worker-side volume mounting,
verifying host mount directory creation, handling pre-existing paths,
plugin errors, and ensuring that partial failures during multi-volume
mounts do not unmount already mounted volumes.
• **TestVolumeHostDirectoryCleanup** - Verifies worker-side host
directory management, ensuring volume unmounting preserves mount
directories while directory reset removes empty directories and prevents
data
loss on non-empty ones.
• **TestExternalVolume_NodeMigration** - Verifies cross-node actor
migration with external volumes by suspending an actor, evicting its
worker pod, and resuming it on a new worker node while ensuring
  application state is preserved

- [ x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 16:25:33 +00:00
Chenyi Wang f3b53847ad ate-setup: wait for ate-system namespace termination during teardown/… (#1851)
## Problem
When `orchestrator.py` runs consecutive benchmark tests in Prow,
`teardown_substrate()` runs `hack/install-ate.sh --delete-ate-system`
(`DeleteAteSystem`), which previously returned immediately after issuing
asynchronous deletion on the `ate-system` namespace without waiting for
termination to complete.

Because draining `atelet` DaemonSets across 15 nodes and `postgres`
takes ~60–65 seconds:
1. The immediately following test runs `deploy_substrate()`
(`EnsureAteSystemNamespace`) ~1 second later while
`namespace/ate-system` is still `Terminating`.
2. `ApplyPath("ate-system-namespace.yaml")` merely updates the
terminating namespace object, and once the old namespace finishes
deleting and transitions to `NotFound`, `WaitNamespaceActive` never
re-applies the manifest and times out after 60s (`waiting for namespace
ate-system to become Active: context deadline exceeded`).
3. As a result, consecutive tests (such as the even-indexed `*_pause` M1
benchmarks) failed during `deploy_substrate()` in ~90% of Prow runs
before Locust could start.

## Fix
- **`DeleteAteSystem` (`cmd/ate-setup/internal/steps/delete.go`)**: Wait
for `namespace/<ns>` deletion (`WaitDeleted`) to complete before
returning from teardown.
- **`EnsureAteSystemNamespace`
(`cmd/ate-setup/internal/steps/env.go`)**: If the target namespace
already exists in `Terminating` phase when deploy starts, wait for it to
finish deleting before applying `ate-system-namespace.yaml` /
`EnsureNamespace`.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 08:50:04 +00:00
Chenyi Wang 8834f1bf29 glutton: replace stuck actors on resume/hibernate failure (#1759)
## Summary
Replace broken/stuck `gluttonActor` instances in place when `resume()`
or `hibernate()` fails during benchmark iterations.

## Problem
When `SuspendActor` fails (e.g., due to transient GCS
`ResourceExhausted` errors during cold-bucket ramp-up), `ate-api-server`
can leave the actor permanently stuck in `ACTOR_STATE_SUSPENDING`.

Previously, `gluttonUser` kept the broken `gluttonActor` in its
`u.actors` slice for the remainder of the benchmark run:
- `hibernate()` ignored `SuspendActor` / `PauseActor` errors.
- Subsequent `ResumeActor` calls against that slot repeatedly failed
with `got: ACTOR_STATE_SUSPENDING` for the rest of the test, reducing
effective concurrency and generating continuous failure noise.

## Solution
- Have `hibernate()`, `pause()`, and `suspend()` return `bool`
indicating whether the operation succeeded.
- Add `(*gluttonUser).replaceActor(ctx, broken)` to delete the broken
actor (`AnyState: true`) and replace its slot in `u.actors` with a
freshly created actor (`sb-<uuid>`).
- Invoke `user.replaceActor(ctx, actor)` in `iterate()` when either
`actor.resume(ctx)` or `actor.hibernate(ctx)` fails.
- Add unit test coverage in `lifecycle_test.go`
(`TestGluttonIterate_ReplacesActorOnResumeFailure`).

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 08:42:46 +00:00
Jaana Dogan 31a5e0ba29 Fix the broken link in ecosystem projects (#1845) 2026-09-24 02:20:05 +00:00
Benjamin Elder ddd6f64e51 ateom-microvm: find leftover VMM processes by their staged path (#1843)
The sweep for a sandbox's leftover cloud-hypervisor and virtiofsd
matched on binary names, but atelet stages both under content-addressed
names, so it never killed anything and a failed boot left its VMM
holding guest RAM. Match on the staged location instead.

Split out from https://github.com/agent-substrate/substrate/pull/1836 as
a small pre-existing fix.
2026-09-24 02:11:02 +00:00
shrutiyam-glitch 5334f16476 feat(ateapi): track actor_template_uid per snapshot on ExternalSnapshot (#1713)
Follow up on the comment -
https://github.com/agent-substrate/substrate/pull/1675#discussion_r4031619230

Stamps `actor_template_uid` onto `ExternalSnapshot` so template
provenance travels with the snapshot itself, rather than being inferred
from the single `ActorStatus.current_actor_template_uid`.

#### Motivation

`current_actor_template_uid` records the template the **last sprint
booted with** (`finalizeRunning` stamps it on every resume). It does
*not* record the template the **snapshot** was captured under. Those two
diverge as soon as an actor runs a sprint that does not produce a new
durable snapshot, and `loadActorForResume` was using the former to
decide whether to force a `DATA`-only restore instead of `FULL`.

Walking an actor through a repoint and a crash:

| Step | Action | `current_…_uid` | `ExternalSnapshot` | 
|---|---|---|---|
| **a** | Suspend on template **A** | A | captured on **A** | 
| **b** | `UpdateActor` repoints the spec to **B** | A | captured on
**A** |
| **c** | Resume → `A != B`, restores `DATA`, sprint boots on **B** |
**B** | still on **A** |
| **d** | Actor crashes (or is reverted) — no new snapshot taken | B |
still on **A** |
| **e** | Resume → `current == target == B`, so **no repoint is
detected** | B | still on **A** |

At step **e** the old check reports "template not replaced" and restores
the external snapshot in `FULL` — replaying a memory image captured
under **A**'s sandbox on **B**. That is exactly the case the guard
exists to prevent; it was silently defeated by the sprint at step **c**
advancing `current_actor_template_uid` past the snapshot.

*(Note: Paused actors do not face this divergence. Because `UpdateActor`
is only allowed in the `SUSPENDED` state, a paused actor cannot have its
template updated. Therefore, local checkpoints do not need separate
template provenance tracking).*

#### Changes

- **Data Model Updates:** Added the `actor_template_uid` field to
`ExternalSnapshot`.
- **Snapshot Stamping:** Updated the capture logic to stamp external
snapshots with the active sandbox's `ActorTemplate` UID when suspending
an actor or creating an actor from a tag.
- **Resume Evaluation Logic:** Modified the external resume logic to
compare the target template UID against the specific snapshot's UID
(`ExternalSnapshot.actor_template_uid`) rather than the actor's overall
status. Local restores bypass this check entirely since templates cannot
change during a pause.

**Tests**
- `TestResumeActor_AteletWireRequest` (unit): snapshot fixtures now
carry `actor_template_uid`, since provenance is read from the snapshot
rather than from actor status.
- Functional expectations for create/update/resume/pause updated for the
new field on the external snapshot golden files.
- `TestUpdateTemplateLifecycle` (e2e): asserts
`external_snapshot.actor_template_uid` after suspend and re-suspend.
- **e2e (Update Template)**: Added coverage for repoint detection after
a revert (verifying the `resume` → `revert` → `resume` path described in
the table above).
- **e2e (Combined Volumes)**: Added coverage to explicitly verify which
volumes a revert rewinds.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 23:58:06 +00:00
Zoe Zhao 74bbfc529c Delete the ateerrors failure-reason taxonomy (#1817)
Stacked on #1220 — review only the last commit until that merges
2026-09-23 20:33:21 +00:00
Julian Gutierrez Oschmann 21b0b030bc Generate the benchmark Python gRPC clients at image build time (#1825)
protoc's Python output puts the serialized descriptor on a single line
which makes any two concurrent ateapi.proto changes conflict.

This is unfortunate as both proto and go generated code usually merges
with no conflict.

Stop checking in the *_pb2*.py files. The locust and nighthawk-ingress
images now generate them in a build stage via
benchmarking/locust/codegen/generate.sh.

Fixes #1824
2026-09-23 19:45:51 +00:00
Julian Gutierrez Oschmann 985c200247 Add DeletePreconditions to all Delete<Resource> methods (#1807)
Allow clients to do optimistic locking on deletes, as well as prevent
unintended deletes on name reuse.

#866
2026-09-23 18:55:59 +00:00
Yufan Su 32df553276 Add statusz page in k8s secret credential provider (#1811)
Follow-up change on a comment in #1335 to add a statusz page in the
credential provider service.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-23 18:45:57 +00:00
shrutiyam-glitch 47b67574ac docs: Document RevertActor and drop "terminal" from CRASHED (#1711)
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.

Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.

Follow up for the PR - #1675 
Issue - #1556 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 17:51:56 +00:00
Steven Shriver 52c6a03c69 CI: Fail on missing preconditions and exceeded boundaries (#1754)
This change addresses six ways CI pipeline could report false success or
hang on broken infrastructure. Importantly, this change does two things
to reduce load on infrastructure:

1. It prevents jobs from running for [the default action limit of
360m](https://docs.github.com/en/actions/reference/workflows-and-actions/workflow-syntax#jobsjob_idtimeout-minutes)
by pinning the timeout to `45m`.
2. It limits pull_request jobs from continuing execution once superseded
by new changes.

- **Silent container test skipping**: Tests now explicitly fail if `CI`
or `REQUIRE_DOCKER` is set (while still skipping on local machines
lacking Docker), with the check implemented in `dockerenv` to avoid an
import cycle between `storetest` and `atepg`.
- **Unbounded trust bundle wait**: Enforced a shared 120-second timeout
across both bundles (overridable via `ATE_INSTALL_TRUST_BUNDLE_TIMEOUT`)
and added diagnostic dumping of bundles, controller pods, and logs
before returning a non-zero exit code.
- **Missing sandbox preflight validation**: Added early preflight checks
for `/dev/kvm` and `SandboxConfig/microvm` that fail fast and print
actionable remediation instructions.
- **Skipped migration checks on main**: Configured the migration
immutability check to run on pushes to main to catch modified migrations
at the point of merge.
- **Missing job timeouts and concurrency limits**: Defined explicit
timeout-minutes (45m and 120m bounds) and added concurrency groups that
automatically cancel superseded pull request runs without canceling runs
on main.

Fixes #1747


- [x] Tests pass
- Silent container skipping, unbounded trust bundles, and missing
sandbox were all forced locally and confirmed to exist with changes here
resolving each.
- The latter half of the scenarios exist in CI only due to being GH
Action trigger issues.
2026-09-23 17:45:42 +00:00
Steven Shriver 898f6e6495 CI: Emit JUnit from test runners, pin their exec bounds (#1687)
Today CI depends on process exit codes. This leads to a few issues with
reporting results. For example, a green build cannot be distinguished
from one where no tests ran at all; a `-run` filter that matches
nothing; an e2e suite short-circuiting without `--e2e`; or a
testcontainer skip: all currently exit 0.
    
This change adds gotestsum as a pinned tool module allowing opt-in for
JUnit reporting in both test runners by using `E2E_JUNIT_FILE` and
`ROOT_JUNIT_FILE` respectively. When left unset, the runners exec `go
test` directly and gotestsum is never built leaving local runs unchanged
for now.
    
Additionally, this change pins two execution bounds that `go test` would
otherwise infer. Both remain overridable via environment variables as
listed below.
    
- `-p`: system defaults didn't align with cluster capacity so pinning
the value avoids over allocation. (override with `E2E_PARALLELISM`).
    
- `-timeout`: Suite defaults conflicted with some E2E test timings
causing some tests that are expected to timeout at the same time to
abort rather than fail as expected. Setting an 30m default avoids the
conlficting timeouts producing clean failure signals (override with
`E2E_TIMEOUT`).

Fixes #1685
2026-09-23 17:27:38 +00:00
Taahir Ahmed 8ea4abe1a3 ateom/actor ID split: Define MintAteomActorCertificate (#1626)
We need to draw a strong distinction between credentials that will be
wielded directly by actor logic, and credentials used by system
components on behalf of an actor. This will help us avoid attacks where
an actor pretends to be a system component.
    
Ateom certificates are requested by ateom/atelet and used for the
atunnel connection to the egress gateway. They have a SPIFFE URI like
`spiffe://${trustdomain}/ateom-for-actor/${atespace}/${actor}`.
    
Actor certificates are requested by the egress gateway and used for
opening outbound requests on behalf of the actor. They have a SPIFFE URI
like `spiffe://${trustdomain}/actor/${atespace}/${actor}`. Note, this
use case is not actually implemented yet.
2026-09-23 17:17:35 +00:00
Da Huang 232d0f2227 ateinterceptors: redact debug_redact fields from RPC logs (#1822)
Part of #1743, follow up of #1803 which added the `debug_redact` labels
but nothing was reading them yet.

The unary interceptors log every request and response body. Until now
the only redaction was to clear any field named `env`, so for example
`MintActorJWTResponse.actor_jwt` (a bearer token) was going to the
ateapi log on every mint call.

Now the interceptor follows the label instead of the field name. Before
logging it clones the message and masks every field with `[debug_redact
= true]` in the copy: singular strings become `[REDACTED]`, other kinds
(bytes, repeated, map, message) are cleared, and it recurses into nested
messages, repeated messages and map values (the old walker skipped
maps). The message returned to the client is never touched. The
hardcoded `env` check is gone since both `EnvVar.value` and
`EnvEntry.value` carry the label now.

One visible change: env var names now show up in the log next to
`[REDACTED]`, before the whole env list was dropped. I think this is
better for debugging and the values are still gone.

This is an incremental step and only fixes the known leak in the
interceptor. It keeps today's structure, clone + walk on every request
and response at this one call site. Two things are left for follow up
PRs on purpose: moving the redaction into the shared slog handler
(`internal/contextlogging`) so any slog call with a proto is covered,
and the performance work (per message type cache so clean messages are
not cloned at all, scan-then-rebuild handler). I benchmarked that
already, numbers are in the doc linked from #1743. Without the cache
this PR costs about the same as today on clean messages and a bit more
on messages with secrets, about 100ns per masked value, which is fine
for now.

The postgres connection string in ateapi's startup flag log is still
open from the P0 list, that will be a separate small PR.
2026-09-23 17:05:55 +00:00
Tim Bai 35dde2ed3f otlprelay: forward log records (#1800)
Step 1 of #1748. The relay forwards `LogsService` the way it forwards
traces and metrics: source gate, verbatim payload, atelet's own
`OTEL_EXPORTER_OTLP_LOGS_HEADERS` upstream.

No content check on the records. Identity spoofing through the relay is
handled in #1748 by the consumer rule (a lifecycle record counts only
under ateapi's resource, which the source gate already enforces) and
per-pod sockets (#741), not by a denylist here.

**Endpoint and compression resolution** (second commit, from review).
The relay read the generic, traces, and metrics OTLP variables and
ignored the logs ones, so a logs-specific collector or compression
setting was silently overridden. All three signals now go through one
resolver: each falls back to the generic, and every resolved value must
agree, since the relay has one upstream connection. The conflict error
names the variable each value came from; a signal that falls back names
the generic variable, not its own unset one.

One behavior change: `traces=A, metrics=A, generic=B` with logs unset
used to pass and now fails at atelet startup, because logs falls back to
`B` and that is a two-collector configuration.

Tests: an ateom log batch arrives with the ateom's resource; the source
gate, mixed-batch, header-isolation, and per-signal header tests cover
logs; resolver tables gained logs rows including the fallback case. Docs
say the relay carries logs, traces, and metrics.
2026-09-23 17:02:55 +00:00
Zoe Zhao aae7df7810 Crash actor when Ateom Checkpoint or Restore fails; Also when the error is not marked as retriable (#1220)
Fixes https://github.com/agent-substrate/substrate/issues/292 

This flips the condition to crash actors: any errors returned by atelet
will cause actors to crash by default.
If we later discover specific error classes that are safe to retry, we
can explicitly mark them as retriable.
2026-09-23 16:40:27 +00:00
Eitan Yarmush d2da3609f5 Add Keith Mattix to the project maintainers (#1806)
Add @keithmattix as a maintainer of Substrate 🎉

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-23 15:33:50 +00:00