mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Fixes #1106. Takes over #1107 (adopting @Stevenjin8's approach and commits from that PR — credit to him for the design; trailer omitted only for CLA) and incorporates the review feedback it accumulated. Workers were registered as schedulable the moment their pod had an IP — long before ateom's gRPC server was listening, which is the scheduling race in #1106 (actors dispatched to a cold ateom fail at dial time). Now: - ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port (8080) once its gRPC listener is up, flipping to 503 on SIGTERM so draining workers stop receiving work. - The worker Deployment probes that endpoint (replaces the earlier TCP probe against the TLS port, which spammed handshake errors). - `isWorkerEligible` requires `PodReady=True` in addition to an IP. Eligibility gates **registration only** — a readiness flap never deregisters a worker with a bound actor (pinned by `TestSyncer_ReadinessFlapDoesNotDeregister`). - A readiness server that cannot come up exits the process rather than leaving a pod that looks alive but can never receive work. Beyond #1107's head: dropped the pre-migration `controlapi/syncer_test.go` copy (didn't compile after #1203 moved the syncer), fixed the `workersync` fixtures for the new gate, and added `TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test above). ## Testing - `go test ./...`: green, zero failures. - `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on `controllers`, `-race` on all of `controlapi` incl. functional tests: green. - Green under a hostile ambient env (`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage). ## Scope notes - The one post-#1160 failure with a *post-restore* signature (run 32891995583: restored actor's own container `/healthz` timing out, actor stuck RESUMING) is a different layer — ateom was demonstrably serving — and is **not** claimed as fixed here. If it recurs, it deserves its own issue. (The other suspected post-merge failure, run 32911674788, turned out to be that PR's own x509 regression, not this flake.) - Mixed-version caveat: this controller pointed at pods running a pre-readyz ateom image leaves those workers unready/unregistered. The supported install flows build both from the same source, so this only matters for hand-rolled skew. - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>