This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.
The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).
For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This PR is large, I grouped changes to the following commits:
- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.
Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
* Renames the`ate.template.namespace` attribute to
`ate.template.atespace` and sources it from the actor's `actor_template`
ObjectRef instead of the legacy CRD namespace/name pair.
* Updated functional tests now to use substrate ActorTemplates and bind
actors through the `actor_template` ref.
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.
This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.
Part of #861
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fix#1213, we should also reject unknown fields at request level.
Refactor all functionality back into a gRPC interceptor to check every
incoming request, no need to handle every new request method.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#1106. Takes over #1107 (adopting @Stevenjin8's approach and
commits from that PR — credit to him for the design; trailer omitted
only for CLA) and incorporates the review feedback it accumulated.
Workers were registered as schedulable the moment their pod had an IP —
long before ateom's gRPC server was listening, which is the scheduling
race in #1106 (actors dispatched to a cold ateom fail at dial time).
Now:
- ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port
(8080) once its gRPC listener is up, flipping to 503 on SIGTERM so
draining workers stop receiving work.
- The worker Deployment probes that endpoint (replaces the earlier TCP
probe against the TLS port, which spammed handshake errors).
- `isWorkerEligible` requires `PodReady=True` in addition to an IP.
Eligibility gates **registration only** — a readiness flap never
deregisters a worker with a bound actor (pinned by
`TestSyncer_ReadinessFlapDoesNotDeregister`).
- A readiness server that cannot come up exits the process rather than
leaving a pod that looks alive but can never receive work.
Beyond #1107's head: dropped the pre-migration
`controlapi/syncer_test.go` copy (didn't compile after #1203 moved the
syncer), fixed the `workersync` fixtures for the new gate, and added
`TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test
above).
## Testing
- `go test ./...`: green, zero failures.
- `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on
`controllers`, `-race` on all of `controlapi` incl. functional tests:
green.
- Green under a hostile ambient env
(`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage).
## Scope notes
- The one post-#1160 failure with a *post-restore* signature (run
32891995583: restored actor's own container `/healthz` timing out, actor
stuck RESUMING) is a different layer — ateom was demonstrably serving —
and is **not** claimed as fixed here. If it recurs, it deserves its own
issue. (The other suspected post-merge failure, run 32911674788, turned
out to be that PR's own x509 regression, not this flake.)
- Mixed-version caveat: this controller pointed at pods running a
pre-readyz ateom image leaves those workers unready/unregistered. The
supported install flows build both from the same source, so this only
matters for hand-rolled skew.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
## Summary
A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.
- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable
## Benchmarking
Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:
-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing
---------
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
Having been reading this code for a while, and in light of changes we
want to make, I think something like this will help clarify.
I am not particularly attached to the name or style `RpcService` - it's
not usual Go convention but it is protobuf style. If anyone feels
strongly it could be `RPCService` or `ControlService` or something else.
@juli4n @laoj2
Fixes#957
`ate.actor.lifecycle.operation.duration` carried the pool pair on
suspend and pause only when they failed, so per-pool dashboards saw
those two operations exclusively as failures.
Both workflows record the histogram from a defer that reads the `actor`
variable, and the happy path reassigns it to the finalized record. The
finalize step commits the new state and the cleared `WorkerAssignment`
in one update, so the defer found no assignment and dropped both keys. A
failure returns the pre-finalize record, which still names the worker.
Both now snapshot `lifecycleOpAttrs(...)` just before the finalize step
— the same snapshot-before-clear crash.go does for the crash counter.
Paths that end earlier keep the current computation.
`delete` stays without a pool: it only runs from SUSPENDED or CRASHED,
which already released the worker, so there is none to name.
TestLifecycleOpPoolAttributesOnSuccess drives a real suspend and pause
through the gRPC service; both subtests fail without the fix.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Make workers a gobal resource, identified by a unique name. Worker APIs
will be implemented in a follow up change.
Part of #730
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
functional_test.go was also split into one file per resource, to match
what we're doing for the RPC handlers too. See #891
I also needed to make some changes to move the tests to a separate
package:
* `NewAteletDialer` now takes `options`, and `WithDialCredentials` lets
a test build its own transport credentials. The fake atelet is reached
over insecure transport. Didn't change existing callers. This is the
only non-test change.
* Moved a few unit tests that are actually functional tests to the
respective file under functionaltest/
Fixes#891
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR