Update the router and egress image to
`cr.agentgateway.dev/agentgateway:v0.0.0-alpha.9e78d1da` from [this
nightly](https://github.com/agentgateway/agentgateway/actions/runs/36271360584),
pinned to its multi-architecture digest.
This includes
[agentgateway/agentgateway#3677](https://github.com/agentgateway/agentgateway/pull/3677),
which reads the `ateom-for-actor` SPIFFE URI from the certificate and
sends the actor SPIFFE identity to credential providers. The currently
pinned image still expects the removed `ActorIdentity` extension, so it
rejects egress after the ateom/actor identity split.
Validation: all five agentgateway Kustomize overlays render with the
expected registry and digest. Fetching the tag and digest from
`cr.agentgateway.dev` returns the same image index as GHCR. Before the
registry-only change, `hack/verify-all.sh` and [upstream PR
CI](https://github.com/agent-substrate/substrate/actions/runs/36275291752)
passed, including both E2E dataplanes and the race tests. CI for the
registry change is pending.
Local `make verify` is blocked by `TestCAPoolCache_HitAndFileChange`: it
rewrites identical bytes and expects the file timestamp to advance,
which is not reliable on this machine's temporary filesystem. The first
run also hit a temporary-directory cleanup failure in
`TestSandboxAssetPrewarmDownloads`; that test passed on retry. Both
tests are unchanged from upstream.
- [ ] Tests pass (waiting for CI on the registry change; previous
revision passed)
- [x] Appropriate changes to documentation are included in the PR (none
needed)
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Fixes#1507
Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.
`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.
This PR is based directly on `main` and does not depend on #1521.
Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Actor template discovery currently uses a fixed 20-second resync
interval. This PR adds `--template-resync-interval` so deployments can
tune that delay, preserving the `20s` default and rejecting nonpositive
values. The setting also controls the reconciler's existing fallback
retry delay.
This is intentionally a small change to start a discussion about how
template builds should be dispatched as the template catalog grows.
The current resync fetches and decodes every template, including
completed ones. At 100,000 templates, a 20-second interval implies
roughly 5,000 template rows read per second per replica, assuming scans
finish quickly. This is an estimate from the code, not a benchmark.
Possible follow-ups:
- **Immediate enqueue:** start work after creation, retaining a slower
recovery scan for crashes between persistence and enqueue.
- **Outbox/watch:** consume changes instead of scanning the catalog. The
existing worker outbox still polls every 50 ms and broadcasts events to
each subscriber; write overhead and recovery scans need consideration.
- **Durable pending-build queue:** atomically record work, claim due
jobs using short `FOR UPDATE SKIP LOCKED` transactions, and recover
expired leases. This still polls, but queries pending work rather than
the full catalog. `LISTEN/NOTIFY` is not the proposed default because it
serializes notifying commits.
- **Imperative build / long-running operation:** give callers explicit
build control or a progress/completion handle. Either still needs
reliable execution underneath.
The main question is whether template builds need a broadcast change
feed or a queue where replicas claim different jobs. Neither alternative
is implemented here.
[Full research: database costs, execution options, recovery
requirements, and
sources](https://gist.github.com/EItanya/0a1d6893ede34f2f0e9d9d1929ecb823).
Validation: all Go race tests and repository verifiers passed using
module mode with `NO_COLOR` unset. CLI checks confirmed the default and
rejection of zero/negative intervals; the existing reconciliation test
checks a custom interval.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Fixes#1508
Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.
- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.
Lint and code-generation verification also passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.
Validated with focused Go tests, shellcheck, and Kustomize renders.
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
Closes#706
## Summary
- broker short-lived actor certificates from atelet over a same-node
mTLS Unix socket
- keep the actor private key in atunnel and renew the certificate before
expiry
- authenticate egress CONNECT using the actor certificate instead of
bearer tokens
## Testing
- `make verify`
- `go test -race ./internal/atunnel`
## Description
Closesagent-substrate/substrate#222.
Adds a portable JWT authentication path for ateapi clients while keeping
the
existing mTLS mode as the default.
### What changed
- Added `internal/ateapiauth` with:
- `mtls` and `jwt` auth modes
- server-side JWT gRPC interceptors
- client-side dial options for projected ServiceAccount tokens
- Wired JWT auth into:
- `ate-api-server`
- `ate-controller`
- `atenet router`
- `kubectl-ate` port-forward client path
- Updated Kubernetes JWT verification to support custom HTTP clients for
OIDC/JWKS discovery.
- Added JWT install overlays:
- `manifests/ate-install/jwt`
- `manifests/ate-install/kind-jwt`
- Updated `hack/install-ate.sh` to opt into JWT mode with:
- `ATE_API_AUTH_MODE=jwt`
- `--auth-mode=jwt`
### Notes
- Default install behavior remains `mtls`. I think this deserves a
second look as JWT will support more clusters.
- The JWT overlay projects short-lived ServiceAccount tokens with
audience
`api.ate-system.svc` and mounts the service-DNS trust bundle for ateapi
TLS
verification.
- The current discovery mechanism for JWT mode involves reading the
deployment which I really don't like, but we don't have a "config file"
concept so that bit is a massive TODO.
### Validation
```bash
KIND_CLUSTER_NAME=substrate-jwt ./hack/create-kind-cluster.sh
KO_DOCKER_REPO=localhost:5001 \
ATE_INSTALL_KIND=true \
./hack/install-ate.sh --auth-mode=jwt --deploy-ate-system
Validation results:
kubectl config current-context
# kind-substrate-jwt
kubectl get pods -n ate-system
# all runtime pods Running; init jobs Completed
go run ./cmd/kubectl-ate --context kind-substrate-jwt get actors
# succeeded, empty actor list
```
What was happening:
1. runsc checkpoint pause succeeded.
2. After that, substrate tried to run runsc state and runsc delete for the app container.
3. But after checkpointing the root sandbox container, the runsc control server was no longer usable,
so those post-checkpoint commands failed with connection refused.
4. That made CheckpointWorkload return an error even though the actual checkpoint had already
succeeded.
5. The ActorTemplate stayed stuck in WaitGoldenActor.
The change removes the post-checkpoint state/delete calls from ateom-gvisor. After checkpoint
succeeds, it only cleans up the actor network and returns success. This matches the existing contract
in the code: atelet owns resetting the actor directories after checkpoint/upload, and it already
calls resetActorDirs(...).
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
A couple of TODOs/Discussions which I think are important to have based
on these changes.
1. Do we need a caching client for this to avoid calling the apiserver
every time?
2. Where/how do we want to resolve secrets generally? I think there are
good arguments to be made for the control-plane OR the "node" component.
In my mind this also related to:
https://github.com/agent-substrate/substrate/issues/18Fixes#15
> It's a good idea to open an issue first for discussion.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
## Summary
Switch the shared atelet/ateom gVisor state directory from
`/run/ateom-gvisor` to `/var/lib/ateom-
gvisor`.
## Why
This path is not just a transient socket location. It is the shared
host/container state root for
downloaded `runsc` binaries, per-ateom sockets, OCI bundles, runsc
state, pidfiles, checkpoint state,
and restore state.
`/run` is normally volatile runtime storage and may be tmpfs-backed,
which makes it a poor fit for
large or service-owned local state. Moving this under `/var/lib` better
matches Linux filesystem
conventions and avoids treating gVisor runtime/checkpoint artifacts as
ephemeral runtime-only files.
Additionally I was running into many permissions issues with apps trying
to use various folders mounted under `/root` before this change.
## Changes
- Update `ateompath.BasePath` to `/var/lib/ateom-gvisor`.
- Update the WorkerPool controller to use `ateompath.BasePath` instead
of hard-coding the mount path.
- Update the atelet install manifest to mount the new hostPath.
## Impact
This keeps `atelet`, `ateom-gvisor`, and dynamically created WorkerPool
Deployments aligned on the
same shared state path. Without that alignment, the system can fail to
find sockets, bundles, runsc
state, or checkpoint/restore files.