Commit Graph
13 Commits
Author SHA1 Message Date
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
Zoe Zhao d4fc52b924 Report the actor template atespace in OTel tags (#1343)
* Renames the`ate.template.namespace` attribute to
`ate.template.atespace` and sources it from the actor's `actor_template`
ObjectRef instead of the legacy CRD namespace/name pair.
* Updated functional tests now to use substrate ActorTemplates and bind
actors through the `actor_template` ref.
2026-08-31 16:47:33 -04:00
Haven Xia bdf494999b api: rename ateomImage to workerImage and make it optional (#1210)
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of #861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 12:46:00 -07:00
Haven Xia 7e339137cf ateapi: reject unknown fields on the request, not just the resource (#1216)
Fix #1213, we should also reject unknown fields at request level.

Refactor all functionality back into a gRPC interceptor to check every
incoming request, no need to handle every new request method.

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-26 16:56:05 -07:00
Aditya ShantanuandAditya Shantanu 66300d5c47 workersync: gate worker registration on ateom readiness (#1243)
Fixes #1106. Takes over #1107 (adopting @Stevenjin8's approach and
commits from that PR — credit to him for the design; trailer omitted
only for CLA) and incorporates the review feedback it accumulated.

Workers were registered as schedulable the moment their pod had an IP —
long before ateom's gRPC server was listening, which is the scheduling
race in #1106 (actors dispatched to a cold ateom fail at dial time).
Now:

- ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port
(8080) once its gRPC listener is up, flipping to 503 on SIGTERM so
draining workers stop receiving work.
- The worker Deployment probes that endpoint (replaces the earlier TCP
probe against the TLS port, which spammed handshake errors).
- `isWorkerEligible` requires `PodReady=True` in addition to an IP.
Eligibility gates **registration only** — a readiness flap never
deregisters a worker with a bound actor (pinned by
`TestSyncer_ReadinessFlapDoesNotDeregister`).
- A readiness server that cannot come up exits the process rather than
leaving a pod that looks alive but can never receive work.

Beyond #1107's head: dropped the pre-migration
`controlapi/syncer_test.go` copy (didn't compile after #1203 moved the
syncer), fixed the `workersync` fixtures for the new gate, and added
`TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test
above).

## Testing

- `go test ./...`: green, zero failures.
- `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on
`controllers`, `-race` on all of `controlapi` incl. functional tests:
green.
- Green under a hostile ambient env
(`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage).

## Scope notes

- The one post-#1160 failure with a *post-restore* signature (run
32891995583: restored actor's own container `/healthz` timing out, actor
stuck RESUMING) is a different layer — ateom was demonstrably serving —
and is **not** claimed as fixed here. If it recurs, it deserves its own
issue. (The other suspected post-merge failure, run 32911674788, turned
out to be that PR's own x509 regression, not this flake.)
- Mixed-version caveat: this controller pointed at pods running a
pre-readyz ateom image leaves those workers unready/unregistered. The
supported install flows build both from the same source, so this only
matters for hand-rolled skew.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-08-26 17:32:03 -04:00
sfunkenhauser 37735a698d Move WorkerPoolSyncer to ate-controller (#1203)
Fixes #729 

- [ x  ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-08-26 14:18:59 -04:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Tim Hockin 87361ea519 Rename controlapi.Service -> RpcService (#1148)
Having been reading this code for a while, and in light of changes we
want to make, I think something like this will help clarify.

I am not particularly attached to the name or style `RpcService` - it's
not usual Go convention but it is protobuf style. If anyone feels
strongly it could be `RPCService` or `ControlService` or something else.



@juli4n @laoj2
2026-08-24 17:54:51 -04:00
Jeff Luo bc59db783a ateapi: stamp the worker pool on successful suspend and pause (#1075)
Fixes #957

`ate.actor.lifecycle.operation.duration` carried the pool pair on
suspend and pause only when they failed, so per-pool dashboards saw
those two operations exclusively as failures.

Both workflows record the histogram from a defer that reads the `actor`
variable, and the happy path reassigns it to the finalized record. The
finalize step commits the new state and the cleared `WorkerAssignment`
in one update, so the defer found no assignment and dropped both keys. A
failure returns the pre-finalize record, which still names the worker.

Both now snapshot `lifecycleOpAttrs(...)` just before the finalize step
— the same snapshot-before-clear crash.go does for the crash counter.
Paths that end earlier keep the current computation.

`delete` stays without a pool: it only runs from SUSPENDED or CRASHED,
which already released the worker, so there is none to name.

TestLifecycleOpPoolAttributesOnSuccess drives a real suspend and pause
through the gRPC service; both subtests fail without the fix.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-20 11:55:18 -04:00
Luiz Oliveira fc02127bfe ateapi - Require preconditions in Update methods (#1014)
https://github.com/agent-substrate/substrate/issues/954

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 21:01:29 -04:00
sfunkenhauser 43fa6ccbc7 Introduce the worker API (#1000)
Make workers a gobal resource, identified by a unique name. Worker APIs
will be implemented in a follow up change.

Part of #730 

- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
2026-08-19 12:06:51 -04:00
Luiz Oliveira 71a58cbf37 split controlapi functional tests into their own package (#918)
functional_test.go was also split into one file per resource, to match
what we're doing for the RPC handlers too. See #891

I also needed to make some changes to move the tests to a separate
package:

* `NewAteletDialer` now takes `options`, and `WithDialCredentials` lets
a test build its own transport credentials. The fake atelet is reached
over insecure transport. Didn't change existing callers. This is the
only non-test change.
* Moved a few unit tests that are actually functional tests to the
respective file under functionaltest/

Fixes #891


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-18 19:06:07 -04:00