29 Commits
Author SHA1 Message Date
Eitan Yarmush ed6d2a1fc8 Update agentgateway for split actor and ateom identities (#1923)
Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.

This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.

Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.

Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.

- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-26 23:02:14 +00:00
Eitan Yarmush d2da3609f5 Add Keith Mattix to the project maintainers (#1806)
Add @keithmattix as a maintainer of Substrate 🎉

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-23 15:33:50 +00:00
Eitan Yarmush a58481a18e Publish ActorTemplate golden snapshots as tags (#1523)
Fixes #1507

Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.

`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.

This PR is based directly on `main` and does not depend on #1521.

Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-16 09:24:42 -04:00
Eitan Yarmush 02eef7e82b Make actor template resync interval configurable (#1522)
Actor template discovery currently uses a fixed 20-second resync
interval. This PR adds `--template-resync-interval` so deployments can
tune that delay, preserving the `20s` default and rejecting nonpositive
values. The setting also controls the reconciler's existing fallback
retry delay.

This is intentionally a small change to start a discussion about how
template builds should be dispatched as the template catalog grows.

The current resync fetches and decodes every template, including
completed ones. At 100,000 templates, a 20-second interval implies
roughly 5,000 template rows read per second per replica, assuming scans
finish quickly. This is an estimate from the code, not a benchmark.

Possible follow-ups:

- **Immediate enqueue:** start work after creation, retaining a slower
recovery scan for crashes between persistence and enqueue.
- **Outbox/watch:** consume changes instead of scanning the catalog. The
existing worker outbox still polls every 50 ms and broadcasts events to
each subscriber; write overhead and recovery scans need consideration.
- **Durable pending-build queue:** atomically record work, claim due
jobs using short `FOR UPDATE SKIP LOCKED` transactions, and recover
expired leases. This still polls, but queries pending work rather than
the full catalog. `LISTEN/NOTIFY` is not the proposed default because it
serializes notifying commits.
- **Imperative build / long-running operation:** give callers explicit
build control or a progress/completion handle. Either still needs
reliable execution underneath.

The main question is whether template builds need a broadcast change
feed or a queue where replicas claim different jobs. Neither alternative
is implemented here.

[Full research: database costs, execution options, recovery
requirements, and
sources](https://gist.github.com/EItanya/0a1d6893ede34f2f0e9d9d1929ecb823).

Validation: all Go race tests and repository verifiers passed using
module mode with `NO_COLOR` unset. CLI checks confirmed the default and
rejection of zero/negative intervals; the existing reconciliation test
checks a custom interval.


Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-15 16:18:13 -07:00
Eitan Yarmush ee8d8faf09 Use Tag UIDs in snapshot storage paths (#1521)
Fixes #1508

Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.

- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.

Lint and code-generation verification also passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-08 18:27:12 -04:00
Eitan Yarmush 1ab1bcb56c docs: add kagent to ecosystem examples (#1415)
## Summary

Adds [kagent](https://github.com/kagent-dev/kagent) to our ecosystem
examples, celebrating
its adoption of Agent Substrate for sandboxed, stateful agent workloads.

Read the [announcement](https://kagent.dev/blog/the-future-of-kagent) 🎉

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-04 09:59:20 -07:00
Eitan Yarmush 4e5a5a2883 Add Actor egress policy functional tests
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Eitan Yarmush 39f2eebc1e Test Actor egress policy behavior
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Eitan Yarmush ac609340d4 Implement Actor egress policy management
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Eitan Yarmush 01e92cca59 Add Actor egress policy API
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Eitan Yarmush a06f464e15 Remove ateapi token client mode (#1045)
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.

Validated with focused Go tests, shellcheck, and Kustomize renders.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-19 17:22:26 -07:00
Eitan Yarmush d944e624be atenet: add agentgateway egress support (#909) 2026-08-18 19:16:12 -07:00
Eitan Yarmush 2b3a4715c6 Configurable JWT authentication to ateapi (#757)
added configurable JWT authentication to ateapi.
2026-08-13 16:44:00 -07:00
Eitan Yarmush c9777b49c4 atunnel: broker actor certificates through atelet (#708)
Closes #706

## Summary
- broker short-lived actor certificates from atelet over a same-node
mTLS Unix socket
- keep the actor private key in atunnel and renew the certificate before
expiry
- authenticate egress CONNECT using the actor certificate instead of
bearer tokens

## Testing
- `make verify`
- `go test -race ./internal/atunnel`
2026-08-07 09:10:34 -07:00
Eitan Yarmush 8b2dd1bbc3 Fix/actor snapshot tag update cas (#755)
Removes unnecessary locking in `UpdateActorSnapshotTag`
2026-08-05 17:01:14 -04:00
Eitan Yarmush ef7b29da44 ateapi: add ActorSnapshot lifecycle APIs 2026-07-30 19:40:59 -07:00
Eitan Yarmush 38ade556ee ateapi: store self-describing snapshots under stable paths 2026-07-30 11:50:59 -07:00
Eitan Yarmush 3400f7fb82 feat: add jwt authentication mode for ateapi (#248)
## Description

  Closes agent-substrate/substrate#222.

Adds a portable JWT authentication path for ateapi clients while keeping
the
  existing mTLS mode as the default.

  ### What changed

  - Added `internal/ateapiauth` with:
    - `mtls` and `jwt` auth modes
    - server-side JWT gRPC interceptors
    - client-side dial options for projected ServiceAccount tokens
  - Wired JWT auth into:
    - `ate-api-server`
    - `ate-controller`
    - `atenet router`
    - `kubectl-ate` port-forward client path
- Updated Kubernetes JWT verification to support custom HTTP clients for
  OIDC/JWKS discovery.
  - Added JWT install overlays:
    - `manifests/ate-install/jwt`
    - `manifests/ate-install/kind-jwt`
  - Updated `hack/install-ate.sh` to opt into JWT mode with:
    - `ATE_API_AUTH_MODE=jwt`
    - `--auth-mode=jwt`

  ### Notes

- Default install behavior remains `mtls`. I think this deserves a
second look as JWT will support more clusters.
- The JWT overlay projects short-lived ServiceAccount tokens with
audience
`api.ate-system.svc` and mounts the service-DNS trust bundle for ateapi
TLS
  verification.
- The current discovery mechanism for JWT mode involves reading the
deployment which I really don't like, but we don't have a "config file"
concept so that bit is a massive TODO.

  ### Validation

```bash
KIND_CLUSTER_NAME=substrate-jwt ./hack/create-kind-cluster.sh

KO_DOCKER_REPO=localhost:5001 \
ATE_INSTALL_KIND=true \
./hack/install-ate.sh --auth-mode=jwt --deploy-ate-system

Validation results:

kubectl config current-context
# kind-substrate-jwt

kubectl get pods -n ate-system
# all runtime pods Running; init jobs Completed

go run ./cmd/kubectl-ate --context kind-substrate-jwt get actors
# succeeded, empty actor list
```
2026-06-30 17:20:56 -07:00
Eitan Yarmush a3f44744d3 Use link-local actor veth addresses 2026-06-10 22:37:09 -07:00
Eitan Yarmush b923428f46 Fix ateom cleanup failure handling 2026-06-10 22:37:09 -07:00
Eitan Yarmush b9ab32a50d Address actor network cleanup review 2026-06-10 22:37:09 -07:00
Eitan Yarmush ef1c9edde8 Fix gvisor cleanup issue
What was happening:

  1. runsc checkpoint pause succeeded.
  2. After that, substrate tried to run runsc state and runsc delete for the app container.
  3. But after checkpointing the root sandbox container, the runsc control server was no longer usable,
     so those post-checkpoint commands failed with connection refused.

  4. That made CheckpointWorkload return an error even though the actual checkpoint had already
     succeeded.

  5. The ActorTemplate stayed stuck in WaitGoldenActor.

  The change removes the post-checkpoint state/delete calls from ateom-gvisor. After checkpoint
  succeeds, it only cleans up the actor network and returns success. This matches the existing contract
  in the code: atelet owns resetting the actor directories after checkpoint/upload, and it already
  calls resetActorDirs(...).

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-06-10 22:37:09 -07:00
Eitan Yarmush 61318aa158 Update licenses for nftables dependency 2026-06-10 22:37:09 -07:00
Eitan Yarmush 62943d0dd2 Trim actor networking comment 2026-06-10 22:37:09 -07:00
Eitan Yarmush 9d9ad7369b Address actor inbound networking TODOs 2026-06-10 22:37:09 -07:00
Eitan Yarmush 709c346723 Group nftables rules with chain definitions 2026-06-10 22:37:09 -07:00
Eitan Yarmush 504cfd21ab Replace actor eth0 move with veth networking 2026-06-10 22:37:09 -07:00
Eitan Yarmush 0076b20c1a Support actortemplate secret env (#20)
A couple of TODOs/Discussions which I think are important to have based
on these changes.
1. Do we need a caching client for this to avoid calling the apiserver
every time?
2. Where/how do we want to resolve secrets generally? I think there are
good arguments to be made for the control-plane OR the "node" component.
In my mind this also related to:
https://github.com/agent-substrate/substrate/issues/18

Fixes #15 

> It's a good idea to open an issue first for discussion.

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-06-08 16:04:50 -07:00
Eitan Yarmush 600320f2e5 Switch ateom gVisor state path (#190)
## Summary

Switch the shared atelet/ateom gVisor state directory from
`/run/ateom-gvisor` to `/var/lib/ateom-
  gvisor`.

  ## Why

This path is not just a transient socket location. It is the shared
host/container state root for
downloaded `runsc` binaries, per-ateom sockets, OCI bundles, runsc
state, pidfiles, checkpoint state,
  and restore state.

`/run` is normally volatile runtime storage and may be tmpfs-backed,
which makes it a poor fit for
large or service-owned local state. Moving this under `/var/lib` better
matches Linux filesystem
conventions and avoids treating gVisor runtime/checkpoint artifacts as
ephemeral runtime-only files.

Additionally I was running into many permissions issues with apps trying
to use various folders mounted under `/root` before this change.

  ## Changes

  - Update `ateompath.BasePath` to `/var/lib/ateom-gvisor`.
- Update the WorkerPool controller to use `ateompath.BasePath` instead
of hard-coding the mount path.
  - Update the atelet install manifest to mount the new hostPath.

  ## Impact

This keeps `atelet`, `ateom-gvisor`, and dynamically created WorkerPool
Deployments aligned on the
same shared state path. Without that alignment, the system can fail to
find sockets, bundles, runsc
  state, or checkpoint/restore files.
2026-06-05 18:34:21 -07:00