### Which issue(s) this PR is related to: Fixes #721 Required for System Upgrade flow (#473) ### What this PR does / why we need it: This PR implements graceful termination and zero-outage rolling updates for `atenet-router` , building on the readiness probe and graceful drain patterns established in #719 (`atelet`). Draining a two-container networking pod (`atenet-router` Go control-plane + `envoy` C++ proxy dataplane) introduces complex lifecycle interdependencies. This PR resolves the dual-container SIGTERM race, respects Envoy's `failClosed` `ext_proc` filter dependency, preserves parked requests riding out worker pool saturation, and accelerates idle deployments via an event-driven file handshake. #### 1. Event-Driven Dataplane Synchronization (`emptyDir` Marker File) Envoy fast-exits by default on SIGTERM. To keep Envoy alive while the Go control-plane orchestrates the drain, we configured an IPC handshake between containers via a pod-shared `emptyDir` volume mounted at `/var/run/atenet`: * **Go Router Container:** Removes any stale marker at startup, then writes `/var/run/atenet/drain-complete` when its shutdown sequence finishes. * **Envoy Container `preStop` Hook:** Runs `while [ ! -f /var/run/atenet/drain-complete ]; do sleep 0.5; done`. **Outcome:** Envoy exits as soon as — the drain is done, rather than wasting time in a fixed worst-case sleep. If the router crashes, Kubelet terminates the `preStop` hook at `terminationGracePeriodSeconds: 60`—slower cleanup, never a wedge. #### 2. The Multi-Container Shutdown Sequence (`drain.go`) Because Kubernetes issues SIGTERM to both containers at once, a coordination state machine ensures Envoy never drops a connection and `ext_proc` is never stopped prematurely: | Phase / (best-case ex.) Timeline | atenet-router | envoy | K8s / Service Status | | :--- | :--- | :--- | :--- | | **SIGTERM Sent**<br>($t=0\text{s}$) | Catches SIGTERM<br>• Flips `/readyz` $\rightarrow$ `503`.<br>• Starts 13s `drain-delay`. | Enters `lifecycle.preStop` hook:<br>`while [ ! -f .../drain-complete ]; do sleep 0.5; done`<br>• Kubelet holds SIGTERM back. | Deployment will recreate the pod.<br>EndpointSlice controller begins dropping Old Pod IP. | | **Propagation**<br>($t=0\text{s}$ – $t=13\text{s}$) | Keeps serving normally.<br>• `ext_proc` continues unparking/routing requests. | Runs normally inside `preStop` loop.<br>• Serves active TCP connections. | Service endpoint removal completes.<br>No new connections arrive at Old Pod. | | **Envoy drain**<br>($t=13\text{s}$) | Issues Envoy admin API calls:<br>• `/healthcheck/fail`<br>• `/drain_listeners?graceful&skip_exit`<br>• Polls `/stats` for active downstream connections. | Begins graceful listener drain:<br>• GOAWAY / `Connection: close` on established connections.<br>• In-flight requests keep running. | In-flight HTTP requests finish executing through Envoy. | | **ext_proc drain**<br>($t\approx18\text{s}$) | Active Envoy connections hit 0 (the poll exits early).<br>• Calls `extproc.GracefulStop()`.<br>• Parked request streams finish. | All downstream connections closed.<br>• Still waiting in the `preStop` loop. | All client HTTP responses delivered. | | **Handshake & Exit**<br>($t=18.5\text{s}$)| `ext_proc` drain finishes.<br>• Writes `drain-complete` marker.<br>• Hard-stops xDS (`Stop()`) & exits. | `preStop` loop detects marker file!<br>• `preStop` exits 0.<br>• Kubelet sends SIGTERM to Envoy.<br>• Envoy exits immediately. | Old Pod deleted cleanly at ~18.5s (well under the 60s budget). | Envoy ref: https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining #### Testing & Verification 1. Unit Tests (`drain_test.go`): Added comprehensive unit tests covering: - Graceful completion of in-flight `ext_proc` streams within deadline. - Force-stopping straggler streams past `--drain-timeout`. - Envoy admin API interaction and connection polling. - Stale marker cleanup and marker file creation. 2. Local testing on kind cluster (Observations): Test script: `hack/verify-atenet-drain.sh` - Steady state: `/readyz=200`, `/healthz=200`; the instant the pod turned Terminating: `/readyz=503` while `/healthz=200` — `NotReady` but alive, for the whole drain. - The parked request survived the shutdown: fired at a busy 1-worker pool → parked → pod deleted while parked → worker freed → HTTP=200, total=1.89s, body hello from: 169.254.17.2 | preserved memory count: 1 | preserved file counter: 1 — park → resume → route, served by the Terminating pod during its drain-delay window. - Pod terminated 14s after deletion — the idle floor, confirmed across three runs: >13s drain-delay (sequence ran), ≪60s grace (marker handshake released Envoy's preStop; no `SIGKILL`). - Log sequence captured verbatim, drain-delay honored to the millisecond (17:34:41.706 → 17:34:54.707): >Shutdown signal received; draining >(+13.000s) Draining Envoy {window: 15s} >Envoy drained >Starting ext_proc drain >ext_proc drain completed within deadline >Drain-complete marker written {path: /var/run/atenet/drain-complete} >Shutdown complete - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
8.6 KiB
Request Parking (atenet router)
Summary
Request parking lets the atenet router hold ("park") an inbound request
whose target actor cannot be served yet because of transient worker-pool
saturation, retrying the resume until the actor becomes routable or a bounded
wait elapses — instead of immediately returning 503 to the client.
Motivation
When a request arrives for a suspended actor, the router resumes it before routing:
Envoy --(ext_proc RequestHeaders)--> router.handleRequestHeaders
--> ActorResumer.ResumeActor --> ateapi ResumeActor (gRPC)
ateapi's AssignWorkerStep claims a free worker from the actor's WorkerPool.
In an oversubscribed system — the core premise of Substrate, where many actors
multiplex onto few workers — a burst of traffic can momentarily exhaust the
pool. AssignWorkerStep then returns FailedPrecondition: "no free workers available".
Previously the router mapped that straight to an HTTP 503 and failed the
request. But such saturation is usually momentary: another actor suspends within
milliseconds and frees its worker. Failing fast turns a sub-second blip into a
user-visible error.
Behavior
With parking enabled (the default), the router treats FailedPrecondition and
Unavailable from ResumeActor as retryable conditions (alongside the
existing Aborted concurrent-resume conflict) — a parked request rides out
transient pool saturation and control-plane blips (e.g. an ateapi rolling
restart) alike. The request is parked: the resumer keeps retrying with
exponential backoff until either
- the resume succeeds (the actor is
RUNNINGand has a worker IP) — the request is then routed normally; or - the park budget (
--parked-request-budget, default5s) elapses — the underlying capacity error is returned, surfacing as503 "actor <id> unavailable: no free workers available".
To bound resource use and provide backpressure, the router admits requests to a
parking lot of fixed capacity (--parked-request-max, default 1024). Each
in-flight resume occupies one slot. When the lot is full, further requests are
shed immediately with 503 "actor <id> unavailable: router at capacity" rather
than queueing without bound.
Every parked request holds one ext_proc stream — one active request against
Envoy's ext_proc cluster — for its entire wait, while ordinary requests hold
one only for a millisecond-scale header exchange. The cluster's circuit breaker
is therefore the hard ceiling on concurrent parked requests. By default the
router derives it as twice --parked-request-max (minimum 1024), so the
lot always fits and an equal share of fast-path headroom remains — a
saturated lot cannot starve requests to already-running actors, at any lot
size. --extproc-max-requests overrides the derivation; explicit values are
validated >= --parked-request-max at startup, because a breaker below the lot
would silently truncate it — Envoy would reject the overflow itself, with 503s
that never reach the lot and never count in parking.rejected.
Concurrent requests for the same actor are de-duplicated by the resumer's
singleflight group: they share a single in-flight ResumeActor call and all
park on its result, so a hot actor consumes N parking slots but only one
control-plane RPC.
The park budget is per-flight, not per-request. The budget clock starts
when a flight's first caller begins the resume; every later request for the
same actor joins that flight and shares its remaining budget and outcome. A
request that joins late may therefore see budget_exhausted after waiting far
less than a full budget itself — the accepted cost of collapsing a hot actor's
requests into one control-plane call. (parking.wait.duration records each
request's own parked time, so sub-budget budget_exhausted samples are
expected under sustained saturation.)
What is not parked
Only transient conditions — capacity (FailedPrecondition), concurrency
(Aborted), and control-plane unavailability (Unavailable) — are parked.
Errors that will not resolve by waiting are returned immediately (fail fast):
| Resume result | Behavior |
|---|---|
OK |
Route to worker |
Aborted (concurrent resume) |
Retry (always) |
FailedPrecondition (no free worker) |
Park & retry (when enabled) |
Unavailable (control-plane blip) |
Park & retry (when enabled) |
NotFound |
Fail fast → 404 |
DeadlineExceeded |
Fail fast → 504 |
PermissionDenied / Unauthenticated |
Fail fast → 403 / 401 |
When parking is disabled (--parked-request-max=0), the router fails fast:
FailedPrecondition and Unavailable are returned immediately, there is no
admission cap, and only Aborted (concurrent-resume) conflicts are retried,
within a 15s budget.
Parked requests survive router shutdown
A request parked when the router pod receives SIGTERM is not reset: the
shutdown sequence keeps the ext_proc server (and, via a preStop handshake, the
Envoy sidecar) alive until in-flight streams finish, and the ext_proc drain
deadline (--drain-timeout) defaults to a value derived from
--parked-request-budget and is validated at startup to be >= the budget —
so a parked request always gets its full budget and a normal verdict (routed
200 or capacity 503) even mid-termination. See the graceful-shutdown knobs
(--drain-delay, --drain-timeout) in manifests/ate-install/atenet-router.yaml.
Configuration
| Flag | Default | Meaning |
|---|---|---|
--parked-request-budget |
5s |
Park budget per resume flight; requests de-duplicated onto an in-flight resume share its remaining budget (see Behavior). |
--parked-request-max |
1024 |
Max concurrent parked/in-flight resume requests; excess shed (503). 0 disables parking. |
--parked-request-retry-interval |
100ms |
Delay before a parked request's first resume retry. |
--parked-request-retry-factor |
1.1 |
Multiplier applied to the retry delay after each attempt (>= 1). |
--parked-request-retry-jitter |
0.1 |
Random fraction in [0, 1) added per retry to de-synchronize parked requests. |
--extproc-max-requests |
0 (auto) |
Envoy circuit-breaker max_requests for the ext_proc cluster. 0 derives twice --parked-request-max (min 1024); explicit values must be >= --parked-request-max (enforced at startup). The excess is fast-path headroom (see Behavior). |
The retry backoff deliberately has no cap and no attempt limit: the budget alone bounds the wait.
Observability
Metrics (OpenTelemetry, meter atenet-router):
-
atenet.router.parking.active— up/down counter: requests currently parked. -
atenet.router.parking.wait.duration— histogram (seconds) of time spent parked. Recorded exactly once per admitted request, at the moment its resume attempt completes; never recorded for shed requests (those only incrementparking.rejected) nor when parking is disabled. Theoutcomelabel says how the park ended:outcomeWhen it is set servedThe resume succeeded and the request was routed to its worker. budget_exhaustedThe park budget elapsed while the resume was still blocked on a retryable condition (pool saturated, a concurrent operation holding the actor, or the control plane unavailable) — the signal that capacity, not a fault, is the bottleneck. canceledThe client disconnected while parked (request context canceled). timeoutThe request's own deadline expired while parked (distinct from the park budget). errorThe resume failed with a non-retryable error ( NotFound,PermissionDenied, ...). -
atenet.router.parking.rejected— counter: requests shed because the lot was full.
Status page (/statusz): a "Request Parking" card shows whether parking is
enabled, the current vs. maximum parked count, and the max wait.