Commit Graph
1042 Commits
Author SHA1 Message Date
Haven Xia 15e765edf8 ate-setup: add deploy subcommands for podcertificate-controller and SandboxConfig (#2010)
They are useful for granular upgrade 

They will not be used for install as we have `deploy ate-system`, they
are used for more granular steps in upgrade runbook.

Part of #1686
2026-09-30 19:48:59 +00:00
Nishanth Kotla 9e525467d1 benchmarking: actor telemetry for the sweperf workload (#1848)
# benchmarking: actor telemetry for the sweperf workload

## What this does

Adds the suspend/resume and task-execution timings from the telemetry
proposal to the sweperf boomer workload.

**New locust rows**

| row | what it measures |
|---|---|
| `ResumeToFirstExec` | resume RPC start until the sandbox accepts the
cycle's `/execute` |
| `TaskCEL` | in-container execution time for one sweperf task (4
cycles), as reported by `replay.py` |
| `CycleCEL` | in-container execution time for one cycle |
| `TaskWallClock` | client wall clock for one task, excluding think time
|
| `<rpc>_rtt` | client round trip for each control-plane RPC |

The four derived rows are recorded under the `actor` method, the `_rtt`
rows under `grpc`.

We keep both the client RTT and the server-side elapsed for every
control-plane RPC so the network and queueing overhead stays visible
separately. The existing `ResumeActor` / `SuspendActor` rows keep the
server-side elapsed from the response trailer.

The liveness check at session start leaves the actor running, so the
first cycle's resume is a no-op. Its successful `ResumeActor`,
`ResumeActor_rtt` and `ResumeToFirstExec` samples are skipped (failures
are still recorded), the same way ateapi skips no-op resumes.

## Poll interval

The `/status` poll interval is now configurable with
`--sweperf-poll-interval-ms` (env `LOCUST_SWEPERF_POLL_INTERVAL_MS`),
default 100 ms (was a fixed 25 ms). It flows through dynconfig like the
other `--sweperf-*` flags and is read per job, so a mid-run change
applies from the next cycle.

## Documentation

`benchmarking/README.md` had no sweperf section, so this adds one
listing every row the workload emits and the poll interval flag,
matching the existing DurDir section.

## Testing

`go test -race ./internal/benchmarking/boomer/...`, plus runs on a real
cluster (sympy, 4 users / 2 workers, gVisor): 5 min at the default 100
ms poll interval and 2 min at 250 ms.

- Both had 0 failures and all expected rows.
- The min `ResumeActor` is 334 ms, so the first-cycle no-op is no longer
recorded.
- `ResumeActor` / `SuspendActor` counts match (197 / 197), and each
`_rtt` is within ~1 ms of its server-side row.
- Client overhead over CEL per cycle is 65 ms at 100 ms vs 141 ms at 250
ms, so the flag takes effect.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-30 19:23:53 +00:00
Julian Gutierrez Oschmann 262af570fb Move DV validators under apivalidations. (#2007)
Both `Control` and `WorkerService` APIs need to access them so moving it
under a shared internal package.

This also aligns ate-apiserver with how atelet uses DV as serves as a
prefactor for upcoming changes to the `ServiceImpl` type.

Fixes #1823
2026-09-30 18:34:15 +00:00
Lior Lieberman 13e8efff25 envoy dyn module followup (#1995)
do not merge, not a draft cause i do want ci running

implements: 
* tie brekaing
* wildcard support
* added cargo test to ci
* some fixes to hostname patterns and port matching


tls_passhtrough is a followup. 



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-30 16:57:00 +00:00
haiyanmeng 12cf471c28 ateapi: release a suspend's replaced snapshot after committing SUSPENDED (#2006)
FinalizeSuspended deleted the external snapshot a suspend replaced
before committing SUSPENDED. By then the new checkpoint is uploaded and
the worker's assignment row is released, so a transient object-store
error (429/503) on that delete aborted the workflow with the actor stuck
in SUSPENDING: ResumeActor rejects that state with FailedPrecondition,
and a retried suspend re-sends Checkpoint to the already-released
worker, which most likely crashes the actor despite a good snapshot in
storage.

Commit SUSPENDED first, then release the replaced snapshot best-effort,
logging a failure instead of returning it. A failed release now leaves
the old snapshot's objects in storage until the actor is deleted, which
removes its whole prefix, rather than stranding the actor. A commit lost
to a concurrent update releases nothing, since the record still names
the old snapshot.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-30 16:16:38 +00:00
Lior Lieberman 1240749483 ateapi: rename --egress-gateway-address to --default-egress-gateway-address (#1858)
Prefix the cluster-wide egress PEP address flag with `default-`
~'experimental-to-be-removed-' to indicate it is a temporary
cluster-level knob that will be removed~ once per-actor or per-atespace
egress gateway configuration is supported (#1591).

~I kept --egress-gateway-address as a deprecated alias for backward
compatibility but I really dont think we should. ~
I removed it, we dont need backward comp residuals pre GA

EDIT: I chatted with tim and taahir a the feedback was that its not
really experimental and can not really be removed. This flag would have
to be set for every GA user, and its not a nice experimental add on. We
dont have time to get into per-actor or per-atespace API discussions
pre-GA therfore we gotta have to live with it and discourage later if we
have a better story for that.

Related #1591


> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

/cc @bowei @EItanya
2026-09-30 15:46:52 +00:00
Lucky Abolorunke 9c1f0ba042 benchmark(atelet, ateom-microvm): log per-actor checkpoint and restor… (#1941)
## Benchamark(atelet, ateom-microvm): log per-actor checkpoint and
restore phase breakdowns

### Why

`SuspendActor` and `ResumeActor` latency is only attributable down to
the `ate.actor.{checkpoint,restore}.duration` histogram phases, and the
biggest buckets — `ateom_checkpoint`, `ateom_restore` — are opaque. When
a large memory benchmark suspend takes 6 s we want to able to tell where
the time is been spent during the suspend.

### What

One joinable, developer-facing log record per operation per layer, with
the full actor identity (allowed in logs, barred from metric labels),
the snapshot scope, and one float-seconds field per phase that ran.

**atelet**
- `Checkpoint` now writes a `Checkpoint timing breakdown` record, the
same way `Restore` has written `Restore timing breakdown` since #1364.
Phases: `sandbox_assets`, `ateom_checkpoint`, `persist`, `total`. One
slice feeds both the histogram and the record, so they cannot disagree.
- Both records carry `error.type` when the operation failed (the gRPC
code; context errors map to `DeadlineExceeded` / `Canceled`). The record
is written on the way out of a failure too, so its completed phases are
kept, and the marker lets a reader exclude a timed-out restore from a
latency distribution. This restores what the record lost when the
`ateerrors` taxonomy was deleted (#1817), using the `error.type`
convention the ateapi instruments already follow.

**ateom-microvm**
- New `phaselog.go`. `CheckpointWorkload` and `RestoreWorkload` emit
records with the same two messages, under
`ateom.actor.checkpoint.duration.<phase>` and
`ateom.actor.restore.duration.<phase>`, decomposing atelet's
`ateom_checkpoint` / `ateom_restore` buckets:
- checkpoint: `pause`, `snapshot`, `durable_dir`, `rootfs_upper`,
`teardown`, `total` — the three captures run concurrently on the paused
guest, so the paused window costs their max, not their sum.
- restore: `prep`, `bundles`, `upper_join`, `lowers`, `tap`,
`vmm_launch`, `vm_restore`, `resume`, `wakeup_probe`, `total` —
sequential; they partition the total. A Data-scope cold boot records
`total` only.
- These timings already existed as ad-hoc `slog.Duration` fields on the
`Actor checkpointed` / `Actor restore phases` lines; those lines are
kept. The record adds stable keys, identity, and seconds (the
histograms' unit).
- The phase names are deliberately private to the binary rather than
added to `internal/ateattr`, so they cannot be mistaken for
`ate.snapshot.phase` metric values. They are micro-VM specific;
ateom-gvisor is unchanged.

**docs/observability.md** is updated: the Restore record is no longer
the only per-actor latency record, and the ateom records are described.

No new instruments, no registry changes, no behavior change.

### Example

```json
{"msg":"Checkpoint timing breakdown","ate.actor.uid":"8f2a…","ate.template.name":"glutton",
 "ate.snapshot.scope":"full",
 "ateom.actor.checkpoint.duration.pause":0.003,
 "ateom.actor.checkpoint.duration.snapshot":0.846,
 "ateom.actor.checkpoint.duration.rootfs_upper":0.022,
 "ateom.actor.checkpoint.duration.teardown":0.232,
 "ateom.actor.checkpoint.duration.total":1.081}
```

Joined with atelet's record for the same actor, a run of the glutton
workload (1 GiB resident, microVM) attributes a 6.0 s p50 suspend as 75%
`persist`, 19% `ateom_checkpoint` (of which the CH `snapshot` is 0.85 s
and `teardown` 0.23 s), and a 5.3 s p50 resume as 79% `download`, 18%
`ateom_restore` (of which `vm_restore` is 0.73 s). The consumer that
produces those tables from pod logs is a separate
`benchmarking/analysis` PR.

### Testing

- `go test ./cmd/atelet/...` and `./cmd/ateom-microvm/...` pass; new
unit tests cover the record shape (seconds, identity keys, zero phases
absent, no duplicate keys), the scope mapping, and `error.type` for
gRPC, context and plain errors.
- `GOOS=linux go vet` clean for both binaries; boilerplate and gofmt
clean.
2026-09-30 10:44:20 +00:00
Yuan Gao 1282773d3a kubectl-ate: add delete egress-policy (#1808)
Part of agent-substrate/substrate#1550.

Adds `kubectl ate delete egress-policy <actor-name> -a <atespace>`

The command removes the actor's policy. It exits 1 when the actor does
not exist
or has no policy. Also add optional `--uid` and `--version` flags, when
specified,
the server refuses the delete when either no longer matches.


### Testing

Tested on a kind cluster with the egress demo
and an API server built from this branch, using a CLI built from this
branch:

- guarded delete with a wrong version, then a wrong uid, then the right
version
- stale version after an `update` bumps the policy
- plain delete on an actor without a policy, and on an actor that does
not exist


🤖 This PR was developed with AI assistance. I have reviewed and tested
all changes.
2026-09-30 04:59:43 +00:00
Davanum Srinivas afb62e18fb atelet: add CAP_DAC_OVERRIDE to reset dirs a non-root container wrote (#1910)
Follow-up to #1906. Removing a directory entry needs write permission on
the directory that holds it, and with all capabilities dropped uid 0
gets no exemption. A non-root container can create a subdirectory in its
`0777` durable dir; that subdirectory belongs to the container's uid,
typically `0755`, so root cannot empty it. `resetActorDirs` fails after
every checkpoint of such an actor and the suspend never completes. Chmod
first, as the bundle dir does, is no way out: chmod needs ownership or
`CAP_FOWNER`.

Why atelet and not ateom, which already holds the capability: an ateom
that dies before cleanup leaves the files behind and moves the hang to
the next resume on that node. atelet owns the actor directories and
needs it on the crash path either way.

Manifest only, so no unit test. Exercised on kind with the micro-VM
class: uid 65532 writes a durable dir, then suspend and resume. Without
the capability the same run hangs in `SUSPENDING`.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (none
needed)
2026-09-30 02:20:11 +00:00
Shruti Nair 6c60ec5d5a ateapi: [Part 1] Runtime Authorization check for Atespace CRUD (#1702)
Working on https://github.com/agent-substrate/substrate/issues/1563

Integrates runtime OpenFGA authorization checks for `Atespace` CRUD RPCs
(`CreateAtespace`, `ListAtespaces`, `GetAtespace`, `DeleteAtespace`) and
atomic OpenFGA tuple cleanup on atespace deletion, gated behind
`--enable-authz` (`ATE_API_ENABLE_AUTHZ`).

### What's Included
- Runtime Policy Decision Point
(`cmd/ateapi/internal/authz/authorizer.go`):
- gRPC Interceptor & Declarative Registry
(`cmd/ateapi/internal/authz/interceptor.go`, `registry.go`):
- Atomic Policy Cleanup on `DeleteAtespace` (`policy_manager.go`,
`atepg.go`):
- Server Wiring (`cmd/ateapi/main.go`):

### Next (PR 3)
- Policy management RPCs
2026-09-29 23:40:43 +00:00
shrutiyam-glitch 820591528e atelet: introduce declarative validation for the AteomSupport RPCs (#1903)
Fixes a part of #1709 

This adds declarative validation to atelet's `AteomSupport` RPCs.
Malformed requests are now rejected at atelet instead of being forwarded
to the control plane. Validation runs after authentication.
| RPC | Rules |
|---|---|
| `MintActorCertificate` | actor atespace/name are required short names;
actor UID is a required UUID; the CSR is required and at most 16 KiB |
| `RequestActorSuspend` | actor atespace/name are required short names;
actor UID is a required UUID |
| `SetWorkerCapacity` | `capacity` is required and must follow the
`WorkerResources` rules ateapi declares, shared through
`internal/resources` |


The `AteomHerder` RPCs and `WorkloadSpec` get their validation in a
follow-up PR.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-29 22:56:18 +00:00
Tim Flannagan fbfd50cd19 Fix several broken links to CODE_OF_CONDUCT.md (#1987)
Update markdown docs that point to the previous filename that was
introduced with #1312.

Signed-off-by: timflannagan <timflannagan@gmail.com>
2026-09-29 22:02:23 +00:00
Anna PendletonandAnna Pendleton 97f8e9aa88 Pin envoy-dataplane from the release on a pre-built install (#1964)
## What

- A pre-built install (`ate-setup deploy --image-repo REPO --image-tag
TAG`) now pins `REPO/envoy-dataplane:TAG` like the ko images, instead of
building it with docker buildx.
- `make build-envoy-dataplane` publishes that image.
- `make build-release-images` publishes every image a pre-built install
needs (ko images, demos, envoy-dataplane), all tagged `$(VERSION)`.

## Why

envoy-dataplane (#1535) is built from a Dockerfile, and ate-setup built
it even for pre-built installs. Those installs have no registry to push
to, so they failed on the default envoy router after the control plane
was applied. It also couldn't join `ALL_IMAGES`, which only lists ko
builds, so nothing published it for a release.

## Testing

- New unit tests for pinning and rendering a pre-built egress manifest.
The rendering test fails on `main` with the production error (`invalid
tag "/envoy-dataplane:build-…"`).
- `make build-release-images` pushed all 13 images under one tag, and
`images.Prebuilt` pinned every one against the real registry.
- Cherry-picks cleanly onto `release-0.2`.

Fixes #1962

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Anna Pendleton <13793293+annapendleton@users.noreply.github.com>
2026-09-29 21:58:58 +00:00
eliranw 2d5ef32d47 ateom-microvm: reseed the guest CRNG on restore (#1524)
## What

On the microVM restore path, ateom reseeds the guest kernel CRNG with
fresh entropy through the kata-agent's ReseedRandomDev RPC, right after
the guest resumes.

This implements the guest RNG reseed on restore that #1449 calls out as
an unfiled gap (its Strategy 2, inject freshness at restore). Scoped to
the kernel CRNG on the microVM runtime, so it does not close #1449.

## Why

Two actors restored from the same snapshot start with the same frozen
CRNG state. The kernel reseeds from ambient entropy on its own, but not
immediately. Measured on a live restore, two clones kept returning
identical /proc/sys/kernel/random/uuid values for up to about 36ms after
resume before that happened. A workload reading randomness in its first
moments after resume can land in that window and get duplicate values
across clones.

Cloud Hypervisor has no VmGenID device to signal the guest, so ateom
reseeds it directly. On each restore it hands a fresh 32-byte nonce to
the kata-agent, which mixes it into the guest CRNG. The nonce differs
per restore, so clones diverge regardless of the kernel's own timing.
The kata-agent already exposes this RPC, so the change is host-side in
ateom.

## Behavior

Fires on every restore (golden cold-start and resume alike). Cold boot
does not need it. Best-effort: a failure is logged, not fatal, so a
transient agent error never fails an otherwise good restore.

## Scope and limits

- Covers the kernel CRNG only. A workload's own userspace PRNG seed
lives in the checkpointed app memory and stays a workload concern
(Strategy 5 in #1449).
- The reseed runs just after Resume, through the kata-agent, so it lands
a few milliseconds after the vCPUs resume. In the live test the reseeded
window shrank from about 36ms to about 6ms but did not reach zero. The
first read or two after resume can still be frozen before the reseed
lands. Fully closing it needs freezing the workload cgroup across the
reseed, or a VMM VmGenID that acts before the vCPUs resume. A TODO in
the code points at the VmGenID path once Cloud Hypervisor gains the
device. Left as a follow-up.
- microVM runtime only. gVisor reads getrandom and urandom live from the
host with no checkpointed RNG state, so restored gVisor sandboxes
already get fresh entropy.

## Testing

- Unit test on the nonce generator (length, and that two nonces differ).
- Verified live on a GKE microVM cluster with a workload that
continuously records /proc/sys/kernel/random/uuid, so the restore
instant is observable. Two clones from one golden snapshot were
compared. Stock ateom returned identical CRNG output for up to about
36ms (7 consecutive reads) after resume, varying run to run. With this
change the shared window shrank to about 6ms (1 read), then diverged.

---------

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
2026-09-29 20:38:08 +00:00
yanavlasov 0f4572495a Egress policy plumbing (#1978)
Plumbing of actor's EgressPolicy into Envoy dataplane. This is the first
commit that only implements SNI allowlist, without wildcards. Followup
PRs will implement there rest of the policy:

* Wildcard matching.
* SNI passthrough policy.
* Allowing plaintext based on presence of `http` rules.
* Port verification.
* Making sdsmint the default and removing non-sdsmin config.

----

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Yan Avlasov <yavlasov@google.com>
2026-09-29 17:54:27 +00:00
Haven Xia 6b53cad92c ateom-microvm: take the actor directories from the request and remove ateompath package (#1872)
This PR let ateom-microvm reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.

The paths ateom-microvm still builds itself are put into microvm
specific`cmd/ateom-microvm/actorpath.go`.

The `ateompath` package as nothing uses it anymore.

Fix #1604
2026-09-29 17:42:32 +00:00
haiyanmeng 4672e95a20 nighthawk-ingress: let run-dev.sh run more actors than workers (#1982)
`run-dev.sh` refused to start unless the Running worker count was at
least `--actors`, on the assumption that each worker hosts one actor. It
does not: a worker advertises capacity for many actors and the scheduler
packs them.

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-29 17:38:32 +00:00
Yufan Su 4f9c293e48 e2e: add test for egress credential injection cujs (#1795)
## e2e: add e2e for egress credential injection

Adds the `egresscredinject` e2e suite: an actor fetches
`https://httpbin.org/headers`
through the sdsmint MITM gateway, and the echoed response proves the
injected
`Authorization` header actually reached the upstream.

### What it asserts

- **Injected**: echoed headers contain `Authorization: Bearer <token>`.
- **Overwritten**: an actor-pre-seeded `Authorization` is replaced by
the injected one.
- **Cleartext skip**: plain-HTTP fetch passes through with no
`Authorization`.
- **Fail closed**: nonexistent secret → 403, unserved provider → 500,
  unauthorized namespace → 403.

### Additional changes

- Probe `/fetch` returns the response body, accepts
`header=<name>:<value>`
params, and no longer follows redirects (a cross-scheme redirect would
hop
between the cleartext and TLS legs). Covered by unit tests; egressmitm
rerun green.
- New `e2e.EgressInjectHeader` policy helper.
- New `e2e.DeployCredentialProvider` + `fixtures/credinject`: deploys
the real
provider manifest with a test Secret and namespace-policy ConfigMap,
restarts
  the provider so it picks the policy up, cleans up on test end.
- CI: new envoy-lane steps after the MITM lanes, gated on
`E2E_EGRESS_CREDINJECT=1`.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-29 03:47:06 +00:00
Haowei Cai (Roy) f7ec6acfc0 fix(benchmarking): make the SWE-Perf ActorTemplate deployable (#1889)
Related: #1694

The `swebench-astropy-7336` ActorTemplate can't be deployed as shipped:

1. `image: ""` must be hand-edited into a tracked file before every run.
2. The command points at `/opt/swebench/replay.py` and
`/opt/swebench/astropy_trace.json`, which the sweperf image generator
never
produces. It writes `/replay.py` and `/trace.json` at the image root
(also
   the image's ENTRYPOINT), so the container exits immediately.

This PR:

- takes the image from a `SWEPERF_IMAGE` env var, substituted by
  `workloads/deploy.sh` like the other template variables;
- fixes both paths;
- documents the minimum sweperf commit (`c30c0d6`), since older images
have no
  HTTP server and fail on the first `GET /status`.

Usage:

```sh
SWEPERF_IMAGE=<registry>/sweperf-astropy@sha256:... \
WORKLOAD_TEMPLATES="swebench-astropy-7336" \
  benchmarking/workloads/deploy.sh --deploy ...
```

Verified on GKE with both gVisor and microVM worker pools: 30-minute
sweperf
runs at 1 user, all 4 cycles, 0 failures.

- [x] Tests pass (`bash -n`; exercised end-to-end as above)
- [x] Appropriate changes to documentation are included in the PR
(template comments)
2026-09-28 23:38:47 +00:00
Haven Xia 7380e230dc ateom-gvisor: take the actor directories from the request (#1865)
This PR let ateom-gvisor reads the per-actor directories from the
`ActorDirs` (added in #1730) in the request instead of deriving them
from the actor UID.

The paths `ateom-gvisor` still builds itself are put into gvisor
specific `cmd/ateom-gvisor/actorpath.go`.

Resetting `runsc-state` and `pidfiles` moves from atelet's
`resetActorDirs` to ateom-gvisor's `main.go` to conduct at beginning of
`RunWorkload` and `RestoreWorkload`, which empties them at the start of
Run and Restore. Atelet doesn't actually read them, so #1730 drops them
from `ActorDirs`.

Added a new `ActorDirs` validation to reject a request with
`InvalidArgument` when `ActorDirs` is missing or any directory in it is
empty, relative or unclean.

The gVisor paths are removed from `internal/ateompath`; the rest goes
with ateom-microvm in the next PR.

Part of #1604.
2026-09-28 22:29:32 +00:00
Benjamin Elderanddependabot[bot] 51a71cb365 upgrade otel log to 0.21 (#1958)
fixes https://github.com/agent-substrate/substrate/pull/1952

security bump stuck by breaking changes.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-28 22:26:40 +00:00
Yufan Su dd57a119ae doc: egress credential injection with sample k8s secret credential pr… (#1805)
add docs on egress credential injection and example k8s secret
credential provider.


- [x] Appropriate changes to documentation are included in the PR
2026-09-28 22:09:22 +00:00
Steven Shriver 7c0c70c7a2 CI: create JUnit artifacts and gate on all tests passing (#1873)
This change wires the `gotestsum` seam into the actual CI and ensures
that all tests _run_ and _pass_ using a new "gate" that verifies the
JUnit artifacts they produce. Specifically, the new gate works as
follows:

1. _registers_ that a test _should_ run and produce an artifact
2. _verifies_ that the artifact was produced and _all_ tests passed
(none were skipped).

The registration phase happens when a test is run in CI and requires
that the test be provided a `E2E_JUNIT_FILE` env variable.

Partial work for #1871

- [x] Tests pass
   ```
  Run make build-junittool
  make build-junittool
  bin/junittool verify -manifest "${ARTIFACTS}/expected-junit.txt"
  shell: /usr/bin/bash -e {0}
  env:
    E2E_ATENET_DATAPLANE: agentgateway
    ARTIFACTS: /home/runner/work/substrate/substrate/_artifacts
go -C tools/junittool build -o
/home/runner/work/substrate/substrate/bin//junittool .
FILE TESTS FAILURES ERRORS SKIPPED
/home/runner/work/substrate/substrate/_artifacts/e2e-gvisor.xml 84 0 0 4
/home/runner/work/substrate/substrate/_artifacts/e2e-microvm.xml 84 0 0
5
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm.xml 1 0 0 0
/home/runner/work/substrate/substrate/_artifacts/e2e-mitm-microvm.xml 1
0 0 0

/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm.xml
11 0 0 1

/home/runner/work/substrate/substrate/_artifacts/e2e-networking-mitm-microvm.xml
11 0 0 1
  TOTAL   
  ```
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 22:01:26 +00:00
Benjamin ElderandDu Bin 5b7417e85d Snapshot path validation (#1950)
Fixes https://github.com/agent-substrate/substrate/pull/1426 

Rebases that PR and adjusts to changes in main since.
https://github.com/agent-substrate/substrate/pull/1426#issuecomment-5804276427

Addresses outstanding review comments in additional commits.

I want to get this in M3. 

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Du Bin <dubin555@gmail.com>
2026-09-28 20:34:38 +00:00
Lior Lieberman 2e57c900a3 move agw tests to post-merge and release-* branches (#1949)
agw e2e are non blocking. moving to only run on release branches and
post-merge for now to save runners.

Ref:
https://cloud-native.slack.com/archives/C0B6M3E2J3D/p1790620259554579?thread_ts=1790453226.572009&cid=C0B6M3E2J3D

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 20:27:58 +00:00
Davanum Srinivas c7dbe9d672 nodepath: move the node state root to /var/lib/ate (#1926)
Fixes #1911

`/var/lib/ateom-gvisor` was named when gVisor was the only sandbox
class; worker pods of both classes mount it. This renames
`nodepath.BasePath` to `/var/lib/ate` everywhere it is spelled out:
atelet's manifest, the kind CSI scripts, `ate-setup`'s CSI step, and the
docs. The controller's worker pod mounts follow the constant.

No compatibility path, per the comment above. A rolling upgrade rolls
the pools once at the controller step, and a worker that lands on a node
whose atelet still uses the old path reaches it only once that node
moves; `docs/upgrade.md` says so. The old directory can be deleted
afterwards.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-28 18:27:55 +00:00
Tim Bai 7999d6f1d1 atelet: pin the multi-actor stats fold and keep CPU baselines across pending sweeps (#1885)
Closes the remaining items in #1640: the atelet side of multi-actor
telemetry.

Since #1836 an ateom returns one `GetActiveWorkloadStats` entry per
hosted actor. atelet's sweep already folded every entry, but nothing
pinned the behavior at occupancy above one, and reviewing that fold
surfaced three CPU-accounting defects: two in how baselines survive a
pending sweep, one in the decrease rule.

## What changed

**Tests pin the multi-actor fold.** One worker, several actors:
same-template entries sum under one key, a different template gets its
own, a pending sibling contributes nothing to aggregates or events, CPU
baselines are per actor (independent advance, independent epoch reset),
one probe serves the whole response. Each test was verified to fail
against the corresponding mutation.

**Three CPU-accounting defects found and fixed along the way:**

1. *Undercount.* A pending entry (boot, restore, or a micro-VM guest the
ateom did not reach in its 45s budget) never wrote the actor's baseline,
and `lastCPU` is rebuilt each sweep — so measured → pending → measured
lost the whole interval. On the guest-agent source the counter survives
a restore, so that was real CPU time dropped, and on a full micro-VM
worker it would happen routinely. Pending now carries the baseline
forward; only actors absent from every response are dropped.
2. *Overcount opened by fix 1.* Carrying the baseline unconditionally
meant that during a restore spanning one sweep (source still reports the
actor measured at C1, destination reports it pending), goroutine order
could let the stale C0 overwrite C1, and the next sweep charged the
interval twice. The pending branch now writes only if no worker has
measured the actor this sweep.
3. *Overcount on a guest-agent decrease.* A CPU decrease was charged as
the whole new value. That is right for a cgroup counter, which restarts
at zero on restore. It is not always right for a guest-agent counter: a
FULL or DATA_ON_GOLDEN restore resumes a guest at that guest's counter
(after a golden restore, main billed the golden guest's CPU to the
actor), a cold boot restarts it at zero, and a container exit lowers the
sum. A sample cannot tell these apart, so a guest-agent decrease now
charges nothing and re-baselines. A cgroup decrease still charges the
new value, with or without a pending sweep between. **Trade-off, new
versus main:** a guest-agent decrease loses the usage from the event to
the first sample. Main counted it correctly only after a cold boot (a
DATA restore, or a Run after a crash); in every other case (golden
restore, FULL restore from an older snapshot, container exit) main
overcounted it.

Remaining limit, documented in the proto: an epoch that starts *above*
the last reported value cannot be detected from consecutive samples.
Main has this for guest-agent restores between sweeps; carrying the
baseline extends it to restores that span a pending sweep. The overcount
is bounded by the golden guest's counter minus the actor's last value,
so it needs an actor that used less CPU than the golden guest had at
snapshot.

The test for (2) forces both fold orders deterministically: the sweep
closes a probe's connection right after folding its response, so the
fixture's closer releases the gated probe from `Close` — no sleeps, no
polling.

**Timeout rationale reconciled.** The poll floor (50s) is one actor's
worst-case micro-VM read: 25 containers at 2s each. It never prevented
overlap, because sweeps are sequential; it bounds guest-agent load, and
the comment now says so. The RPC timeout (55s) was documented as
covering that single-actor worst case. With several actors per worker,
the micro-VM ateom caps its own read at `statsSweepBudget` (45s, chosen
to stay under atelet's 55s) and reports unreached guests as pending, so
the timeout stays a constant; its comment now points at that budget.

## Not in scope, recorded

- Two workers both *measuring* one actor in one sweep double-charges.
Pre-existing and unchanged here. It is rare: it needs the same probe
skew as defect 2, with the destination read after its restore finished.
- `ateom.proto` still says the discovery list holds "at most one entry
until multi-actor workers land". Stale; fixed in a separate change.

## Verification

`gofmt`, `go vet`, and race-enabled tests pass. Mutation checks:
dropping the carry, the measured-wins guard, or either side of the
per-source decrease rule each fails its test.
2026-09-28 18:03:00 +00:00
Bowei Du 22efea18a9 Clean up packages in ateomnet (#1897)
* Move the sandbox DNS code into its own package
* Move the network namespace primitives into their own package
2026-09-28 17:36:52 +00:00
Lior Lieberman 6621f2b4b4 egress policy enhancements for GA (#1751)
**BREAKING** API changes prior in preparation for GA. 

~This PR has proto-only change to the EgressPolicy API. No
implementation; the gateway,
store contract, and e2e helpers stop compiling until the follow-up
lands.~ Note: Impl is only to make ci green. Full implementation of the
api is a fast follow


- `EgressRule` is now a union of protocols: `http`, `https`,
  `tls_passthrough`. Allow-only, deny by default.
- The `hostnames`, `cidrs`, and `all` rule kinds are removed. **No field
numbers
or names are reserved: we are pre-GA and existing policies must be
recreated.**
- `rules` is an **unordered set**. When more than one rule matches,
precedence is
decided by two criteria, in order: a pattern without a wildcard wins
over one
with a wildcard, then a port other than `"*"` wins over `"*"`. This
applies
within a protocol and across `https` and `tls_passthrough` on the same
SNI.
Two rules with the same pattern on the same port are rejected on write.
  Only the winning rule's effects apply.
- `http`: cleartext HTTP, matched per request on the authority and port.
Carries `effects`.
- `https`: MITMed HTTPS that the gateway intercepts. Matched on SNI and
port at the
ClientHello, then per request on the authority. Carries `effects`. MITM
is off unless a name is
  listed here.
- `tls_passthrough`: TLS forwarded without decryption, matched once per
connection on SNI and port. No effects. Implicit TLS only; STARTTLS does
not match.
- Protocols are told apart by what the Actor sends first, a ClientHello
or an HTTP
  request. Anything else, or a server-first protocol, is closed.
- Name fields are `host_patterns` on `http` and `https`, `sni_patterns`
on
  `tls_passthrough`. Same wildcard grammar as before, documented once on
  `HTTPRule.host_patterns`.
- `ports` on all three protocols are strings: a port number, or `"*"`
alone for any
port. `http` defaults to `["80"]` and `https` to `["443"]` when empty;
`tls_passthrough`
requires at least one. The port is the destination the Actor connected
to.
A port in a request's authority is neither matched nor dialed. Format
documented
  once on `HTTPRule.ports`.

- `inject_static_headers` is renamed to `replace_headers`, since the
behavior is
replacement and the name leaves room for other header operations later.
The
`CredentialHeaderInjection` message keeps its name. Replacement is
conditional:
a header is replaced only when the Actor's request already carries it,
with any
  placeholder value. Requests without the header pass unchanged.

- Unsupported in v1 rules for arbitrary TCP, non-TCP protocols.

- `google/protobuf/empty.proto` import dropped; nothing uses it now.

Questions

- The handler is called `https`, not `mitmHttps`. Everything in an
`https`
  block is intercepted. is that clear enough?

follow-ups
- align egress implementation
- named CONNECT, or some way to preserve what the actor dialed.
- add tcp support
- hand-written validators for the new fields (`host_patterns`,
`sni_patterns`, `ports`,
  the rule conflict check) and regenerate `zz_generated.validation.go`



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Fixes: #823 (there are probably more issues that i need to find)
2026-09-28 17:30:07 +00:00
Eitan Yarmush ed6d2a1fc8 Update agentgateway for split actor and ateom identities (#1923)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.

This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.

Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.

Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.

- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-26 23:02:14 +00:00
Dmitry Berkovich 1d7ca8ced0 ateletpb: carry the Restore base snapshot as a typed source, dropping golden_snapshot_uri (#1602)
Refactors the atelet Restore API so the DATA_ON_GOLDEN base snapshot
travels as a typed restore source instead of a bare string, preparing
for the node-local snapshot cache (#690, #1551): `RestoreRequest` gets
its own external snapshot *source* message, and read-side attributes of
a restore source have a home — today the URI, next (in the M2 cache
work) a sharing property set by the control plane.

**This change is not backward compatible.** `golden_snapshot_uri` is
removed from `RestoreRequest` with no transition: an ateapi from before
this change sends a DATA_ON_GOLDEN restore an updated atelet rejects (no
`base_config`), and an updated ateapi sends one an old atelet rejects
(no `golden_snapshot_uri`). ateapi and atelet must be rolled together.
Fresh starts, FULL/DATA restores of an actor's own snapshot, and local
pause restores without a golden base are unaffected.

Three commits, each buildable and tested:

1. **`ateletpb: give Restore its own external snapshot source message`**
— `ExternalRestoreConfiguration` replaces
`ExternalCheckpointConfiguration` in `RestoreRequest`'s config oneof
(same wire shape; the checkpoint-side write-destination message is
untouched), and `base_config` is added for the DATA_ON_GOLDEN base. No
behavior change: nothing sets or reads the new field yet.
2. **`ateapi, atelet: carry the restore base snapshot in base_config`**
— the cutover: ateapi sets `base_config` on both golden-data resume
paths (external data snapshot, and local pause checkpoint combining with
the golden); atelet reads and validates only `base_config`. The
characteristic test of the ateapi→atelet request and the golden-data
functional test pin the new field.
3. **`ateletpb: remove golden_snapshot_uri from RestoreRequest`** — the
old field is deleted and `base_config` takes its field number, keeping
the numbering dense.

Tested: full ateapi + atelet suites including the controlapi functional
tests; gofmt, golangci-lint, boilerplate clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-26 14:35:20 +00:00
Taahir Ahmed c7b5469968 ate-api: Move authn package into server-specific internal package (#1909)
Prefactor for unified config file
2026-09-26 01:34:20 +00:00
Taahir Ahmed fe82dae3f3 serverboot.Fatal: Allow additional slog arguments (#1908)
Prefactor for unified config file.
2026-09-26 01:29:52 +00:00
Michelle AuandJefftree d3aec58cdd Remove unnecessary fsync during ResumeActor (#1883)
While benchmarking pause actor I ran into multiple bottlenecks.

I'm currently importing
https://github.com/agent-substrate/substrate/pull/935 into my branch but
https://github.com/agent-substrate/substrate/pull/1876 is another
alternative to solve the first bottleneck.

The 2nd bottleneck (primary target of this PR) is due to fsyncs being
called in writeFileAtomic() for sandbox-assets.json and SystemInfoVolume
files. Because these are new files that require ext4 filesystem metadata
updates, and we have ext4 mounted in ordered mode, calling fsync on the
file + directory causes all other pending writes (such as other
checkpoints) to be flushed. So the file creation ends up being blocked
behind other actor's checkpoint syncs.

These files do not actually need to be flushed to disk. SystemInfo is
recreatable. sandbox-assets.json is written during Run/Resume(), and
read during Checkpoint(). But once the checkpoint is finished, the
manifest is saved with the snapshot. If the node crashes between
Run/Resume and Checkpoint, the actor is crashed anyway.

### Single-Node Pause/Resume Benchmark Summary
**Configuration:** 17 workers, 15 actors, `glutton` (`512 MiB` RAM
capacity, `32 MiB` churn per cycle), `120s` run, `1.0s` wait, `100%`
pause (`0%` suspend)

| Metric | Baseline (`upstream/main`) | + Commit 1 (`os.Link` local
checkpoint) | + Commit 2 (remove `fsync` on ephemeral volumes) |
| :--- | :---: | :---: | :---: |
| **Completed Cycles** (`120s`) | `106.5` (`112` pauses / `101` resumes)
| `160.0` (`165` pauses / `155` resumes) | **`212.5`** (`214` pauses /
`211` resumes) |
| **Cycle Throughput** | `53.3 cycles/min` (`0.89/s`) | `80.0
cycles/min` (`1.33/s`, **+50.2%**) | **`106.3 cycles/min`** (`1.77/s`,
**2.00x**) |
| **`ResumeActor` p50** | `6,000 ms` | `4,100 ms` | **`210 ms`**
(**28.6x faster**) |
| **`ResumeActor` avg** | `7,023 ms` | `4,867 ms` | **`246 ms`**
(**28.5x faster**) |
| **`ResumeActor` p90 / p99** | `15,000 ms` / `19,000 ms` | `9,400 ms` /
`20,000 ms` | **`350 ms` / `590 ms`** |


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

---------

Co-authored-by: Jefftree <jeffrey.ying86@live.com>
2026-09-25 23:52:04 +00:00
Taahir Ahmed 7f7a40c1b9 ateom/actor ID split: Clean up purpose from MintActorCertificate (#1810)
The purpose field in MintActorCertificate (and the certs it issues) is
no longer needed. The distinction between actors and ateoms-for-actors
is now built into the SPIFFE URI.

Stacked over #1809
2026-09-25 23:45:05 +00:00
botengyao d2b29a5fa2 docs: add a star history chart to the README (#1894)
Add a Star History section to the README that embeds star-history.com's
live chart, with a dark variant for GitHub's dark theme.
star-history.com renders the image on request, so it stays current;
browsers and GitHub's image proxy cache it for up to a day.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-25 21:44:01 +00:00
Taahir Ahmed 8d6be5fb4e ateom/actor ID split: Ateom uses MintAteomActorCertificate (#1809)
Update atelet/ateom/atunnel to use the new MintAteomActorCertificate,
and correct anything that was coded to expect the old actor SPIFFE URI.
2026-09-25 21:08:23 +00:00
Haven Xia a0b680d72e Add ActorDirs message to carry per-actor directories and and split the path packages by owner (#1730)
Today `atelet` and `ateom` derive the per-actor directories (oci
bundles, runsc state, pid files, checkpoint and restore state,
durable-dir, system-info and volume roots) from the actor UID through
the same package, `internal/ateompath`.

This PR: 

1. Add an `ActorDirs` message to RunWorkloadRequest,
RestoreWorkloadRequest, CheckpointWorkloadRequest and
TerminateWorkloadRequest, and have atelet fill it with the directories
it prepared. **The ateoms do not read it yet.**

2. Split `internal/ateompath` into:
- `cmd/atelet/internal/ateletpath`: atelet's pathes, the per-actor
directories (and `ActorDirs`, built from them) plus the directories only
atelet uses.
- `internal/nodepath`: shared pathes, the base dir both mount, the ateom
socket, the OTLP
     sockets, and the netns name.
- `internal/ateompath`: what the ateoms still derive from the actor UID,
each function marked with the `ActorDirs` field it duplicates. atelet no
longer imports it.

---------

This is the first part of #1604 to decouple atelet and ateom shared
pathes. Behavior **does not change**: both sides still compute the same
paths, atelet from `ateletpath` and the ateoms from `ateompath`.

Followup PRs will remove ateompath by making `ateom-gvisor` and
`ateom-microvm` read the `ActorDirs` message from RPC.
2026-09-25 21:03:44 +00:00
Tim Bai 03ba5821ec actorevent: define the actor usage event and the activation epoch (#1881)
Step 3 of #1748. Vocabulary only: nothing emits the event or sets the
epoch yet.

**Event.** `ate.actor.usage_sampled` in `actorevent`, sixteen required
keys via `ateattr`: identity, pool, sandbox class, source,
`ate.stats.kind` (periodic, initial, final), the four measurements,
observation time, epoch. Measurement names follow the
`ate.actor.stats.*` instruments; units are in the registry briefs. Body
is `Actor usage sampled`, the accepted stdout shape change.

**Scope.** The usage stream has its own instrumentation scope, the
lifecycle scope with `/usage` appended, so a consumer can sample or drop
it without touching lifecycle records. `NewEmitter` takes the scope; the
package-level `Log` keeps the lifecycle one, so ateapi is unchanged.

**Epoch.** `ate.actor.epoch` and `WorkloadStatsSample.epoch_unix_nano`:
the unix-nano time an activation began. CPU time is per activation for
every source; lifetime is the sum over epochs of each epoch's highest
value. Memory usage and working set are absolute; peak is as the source
reports it. An ateom that leaves the epoch at zero still sends the raw
guest counter. The measurement comments on fields 8 to 11 say all of
this.

**Registry.** Event group in `events.yaml`, seven attributes in
`metrics.yaml`, `ate.actor.epoch` in the no-actor-identity forbidden
list. The one-record-one-place rule and the collector guide now name
both scopes.

Left for step 4: the event's code anchor moves to the ateom emit site,
atelet's poller comments and the guide's stdout usage section get
rewritten, and the controller starts passing `OTEL_LOGS_EXPORTER` to
worker pods.
2026-09-25 19:54:15 +00:00
Max Thompson d3a556aae3 ateapi: make the actor JWT issuer configurable and add typ to the header (#1834)
Part of #1756.

- Adds `--actor-jwt-issuer`. Unset, it defaults to
`https://idp.<namespace>.svc`, the Service ate-idp-server will serve
discovery from. Before this, `iss` was hardcoded to
`https://api.ate-system.svc`, which points at ateapi's gRPC port.
- Adds `oidcdiscovery.ParseIssuer`, which requires a canonical https URL
with no query, fragment, or user info and strips trailing slashes.
ate-idp-server will use it too, so both emit the same bytes.
- Actor JWT headers now include `typ: JWT`.
- Drops the `actorIdentityJWTIssuer` plumbing into `RPCService`, unused
since #1315.

The default `iss` changes, but `MintActorJWT` has no production caller
yet.

Testing: new unit tests for the parser, the default, the header, and
flag resolution. `TestMintActorJWT_Success` now checks `typ`, `iss`, and
`sub`. It needs Docker, so it hasn't run locally.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR (docs
land later in the #1756 series)
2026-09-25 19:22:25 +00:00
John Howard 10a1bfb2f5 PodCertificateRequest v1 support (#1829)
This adds dual v1beta1 and v1 PCR support, favoring v1. We have seen
users on v1.37 without v1beta1 but with v1 seeing incomaptibility.

Given our usage of PCR is pretty scoped its not too bad to have the dual
version support.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
v0.2.0
2026-09-25 17:29:49 +00:00
Max Smythe 8b5d9ff0ef Tolerations for Workload Types (#1874)
We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-09-25 16:45:02 +00:00
shrutiyam-glitch 54cba01a8e feat(ateapi): send resolved sandbox assets during restore (#1866)
Fixes #1835 

`atelet` chose the sandbox binaries and pause image for a restore from
the snapshot's `manifest.json` rather than from the control plane. This
PR makes ateapi send them on every restore, the way it already does for
Run.

**Changes**
- **proto:** add `sandbox_assets` to `atelet.RestoreRequest`.
- **ateapi:** `ensureAteletRestored` resolves the ActorTemplate's
SandboxConfig once and sends it on the local restore, the external
restore and Run. A missing SandboxConfig now fails before any atelet
call.
- **atelet:** `sandbox_assets` is required on Restore. Restore selects
the binaries and pause image with `recordFromRequest`, the same path Run
uses. The manifest still supplies the snapshot files, sandbox class and
metrics labels.

Note - The field is required, so after atelet is upgraded, resumes on
that node fail until ate-api-server is upgraded too.

**Testing**
- ateapi: every Run and Restore sends the resolved assets; a missing
SandboxConfig fails on both restore paths.
- atelet: `TestRecordFromRequest`, a validation case for a missing
field, and `TestRestoreUsesRequestSandboxAssets` (checkpoint with pause
image v1, restore with v2, and the actor runs with v2).

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-25 16:06:02 +00:00
Luiz Oliveira e87c55fc2e Add the crash reason and timestamp to the ActorStatus (#1867)
Added a new field to actor status that has the crash reason (e.g., the
atelet response error) and the timestamp of the crash for better UX.
2026-09-25 15:33:08 +00:00
Luiz Oliveira d72edfbb3c Rename LocalSnapshotInfo to LocalSnapshot (#1856)
For consistency with ExternalSnapshot

#1378
2026-09-25 14:20:42 +00:00
Benjamin Elder 22ba860e0f add get-contributors.sh for releases (#1870)
NOTE: This looks like a script in sigs.k8s.io/kind because it is. And
yet it is not third_party, because I am the author of all current lines
in both locations (and the original author of the script), I am
licensing this form to both projects.

v0.1.0 used a form of this, checking it in before v0.2.0

Why do we want this? Because github will nicely render contributors *if*
you mention them.
But I've found that the auto-generated release notes failed to correctly
list everyone, and it feels bad when people are left out. Of course this
only attributes commits and not other forms of contributing ... but
that's another matter.
2026-09-25 01:38:04 +00:00
Yufan Su 5ebccf3897 Fix the kind install on arm64 hosts (#1869)
Fixes the kind install on arm64 hosts (Apple Silicon), broken in #1535

- `BuildDockerfileImage` hard-codes `--platform=linux/amd64`, so the
dataplane image's Rust builder stage fails with `exec /bin/sh: exec
format
  error` on mac with arm64

The ko images already follow `KO_DEFAULTPLATFORMS`, which the kind
install
sets to the host architecture; the Dockerfile images did not.

**Changes**

- Pass `KO_DEFAULTPLATFORMS` through to `docker buildx build --platform`
for
  both Dockerfile images (envoy dataplane, claude-code-multiplex demo),
  defaulting to `linux/amd64` when unset.
- Pin the Envoy base image by its multi-platform index digest so
`--platform`
  selects the matching manifest.

- [x] Tested (`hack/install-ate-kind.sh --deploy-ate-system` on mac
succeed)
2026-09-25 01:36:37 +00:00
Max Smythe d6d2a0fa1d Add the option to install a large-cluster manifest and cordon control plane to dedicated machines (#1632)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.

It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 23:39:01 +00:00
Benjamin Elder 14c0c136bc multi actor worker support (#1836)
Fixes #1266 

Builds on #1689 #1283 in particular.

Outstanding: 
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~

EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-24 21:25:22 +00:00
sfunkenhauser fa24dcc514 Add plumbing to allow ateom to suspend actors (#1779)
Part of #483 

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-09-24 20:59:46 +00:00