Files
substrate/internal/serverboot
Aditya ShantanuandAditya Shantanu 66300d5c47 workersync: gate worker registration on ateom readiness (#1243)
Fixes #1106. Takes over #1107 (adopting @Stevenjin8's approach and
commits from that PR — credit to him for the design; trailer omitted
only for CLA) and incorporates the review feedback it accumulated.

Workers were registered as schedulable the moment their pod had an IP —
long before ateom's gRPC server was listening, which is the scheduling
race in #1106 (actors dispatched to a cold ateom fail at dial time).
Now:

- ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port
(8080) once its gRPC listener is up, flipping to 503 on SIGTERM so
draining workers stop receiving work.
- The worker Deployment probes that endpoint (replaces the earlier TCP
probe against the TLS port, which spammed handshake errors).
- `isWorkerEligible` requires `PodReady=True` in addition to an IP.
Eligibility gates **registration only** — a readiness flap never
deregisters a worker with a bound actor (pinned by
`TestSyncer_ReadinessFlapDoesNotDeregister`).
- A readiness server that cannot come up exits the process rather than
leaving a pod that looks alive but can never receive work.

Beyond #1107's head: dropped the pre-migration
`controlapi/syncer_test.go` copy (didn't compile after #1203 moved the
syncer), fixed the `workersync` fixtures for the new gate, and added
`TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test
above).

## Testing

- `go test ./...`: green, zero failures.
- `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on
`controllers`, `-race` on all of `controlapi` incl. functional tests:
green.
- Green under a hostile ambient env
(`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage).

## Scope notes

- The one post-#1160 failure with a *post-restore* signature (run
32891995583: restored actor's own container `/healthz` timing out, actor
stuck RESUMING) is a different layer — ateom was demonstrably serving —
and is **not** claimed as fixed here. If it recurs, it deserves its own
issue. (The other suspected post-merge failure, run 32911674788, turned
out to be that PR's own x509 regression, not this flake.)
- Mixed-version caveat: this controller pointed at pods running a
pre-readyz ateom image leaves those workers unready/unregistered. The
supported install flows build both from the same source, so this only
matters for hand-rolled skew.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-08-26 17:32:03 -04:00
..