Files
substrate/docs/request-parking.md
T
shrutiyam-glitch 5f64ba4c26 feat(atenet): graceful termination of atenet-router (#774)
### Which issue(s) this PR is related to:
Fixes #721
Required for System Upgrade flow (#473)

### What this PR does / why we need it:
This PR implements graceful termination and zero-outage rolling updates
for `atenet-router` , building on the readiness probe and graceful drain
patterns established in #719 (`atelet`).

Draining a two-container networking pod (`atenet-router` Go
control-plane + `envoy` C++ proxy dataplane) introduces complex
lifecycle interdependencies. This PR resolves the dual-container SIGTERM
race, respects Envoy's `failClosed` `ext_proc` filter dependency,
preserves parked requests riding out worker pool saturation, and
accelerates idle deployments via an event-driven file handshake.

#### 1. Event-Driven Dataplane Synchronization (`emptyDir` Marker File)
Envoy fast-exits by default on SIGTERM. To keep Envoy alive while the Go
control-plane orchestrates the drain, we configured an IPC handshake
between containers via a pod-shared `emptyDir` volume mounted at
`/var/run/atenet`:

* **Go Router Container:** Removes any stale marker at startup, then
writes `/var/run/atenet/drain-complete` when its shutdown sequence
finishes.
* **Envoy Container `preStop` Hook:** Runs `while [ ! -f
/var/run/atenet/drain-complete ]; do sleep 0.5; done`.

**Outcome:** Envoy exits as soon as — the drain is done, rather than
wasting time in a fixed worst-case sleep. If the router crashes, Kubelet
terminates the `preStop` hook at `terminationGracePeriodSeconds:
60`—slower cleanup, never a wedge.

#### 2. The Multi-Container Shutdown Sequence (`drain.go`)
Because Kubernetes issues SIGTERM to both containers at once, a
coordination state machine ensures Envoy never drops a connection and
`ext_proc` is never stopped prematurely:

| Phase / (best-case ex.) Timeline | atenet-router | envoy | K8s /
Service Status |
| :--- | :--- | :--- | :--- |
| **SIGTERM Sent**<br>($t=0\text{s}$) | Catches SIGTERM<br>• Flips
`/readyz` $\rightarrow$ `503`.<br>• Starts 13s `drain-delay`. | Enters
`lifecycle.preStop` hook:<br>`while [ ! -f .../drain-complete ]; do
sleep 0.5; done`<br>• Kubelet holds SIGTERM back. | Deployment will
recreate the pod.<br>EndpointSlice controller begins dropping Old Pod
IP. |
| **Propagation**<br>($t=0\text{s}$ – $t=13\text{s}$) | Keeps serving
normally.<br>• `ext_proc` continues unparking/routing requests. | Runs
normally inside `preStop` loop.<br>• Serves active TCP connections. |
Service endpoint removal completes.<br>No new connections arrive at Old
Pod. |
| **Envoy drain**<br>($t=13\text{s}$) | Issues Envoy admin API
calls:<br>• `/healthcheck/fail`<br>•
`/drain_listeners?graceful&skip_exit`<br>• Polls `/stats` for active
downstream connections. | Begins graceful listener drain:<br>• GOAWAY /
`Connection: close` on established connections.<br>• In-flight requests
keep running. | In-flight HTTP requests finish executing through Envoy.
|
| **ext_proc drain**<br>($t\approx18\text{s}$) | Active Envoy
connections hit 0 (the poll exits early).<br>• Calls
`extproc.GracefulStop()`.<br>• Parked request streams finish. | All
downstream connections closed.<br>• Still waiting in the `preStop` loop.
| All client HTTP responses delivered. |
| **Handshake & Exit**<br>($t=18.5\text{s}$)| `ext_proc` drain
finishes.<br>• Writes `drain-complete` marker.<br>• Hard-stops xDS
(`Stop()`) & exits. | `preStop` loop detects marker file!<br>• `preStop`
exits 0.<br>• Kubelet sends SIGTERM to Envoy.<br>• Envoy exits
immediately. | Old Pod deleted cleanly at ~18.5s (well under the 60s
budget). |

Envoy ref:
https://www.envoyproxy.io/docs/envoy/latest/intro/arch_overview/operations/draining

#### Testing & Verification

1. Unit Tests (`drain_test.go`): Added comprehensive unit tests
covering:
- Graceful completion of in-flight `ext_proc` streams within deadline.
- Force-stopping straggler streams past `--drain-timeout`.
- Envoy admin API interaction and connection polling.
- Stale marker cleanup and marker file creation.

2. Local testing on kind cluster (Observations):
Test script: `hack/verify-atenet-drain.sh`
- Steady state: `/readyz=200`, `/healthz=200`; the instant the pod
turned Terminating: `/readyz=503` while `/healthz=200` — `NotReady` but
alive, for the whole drain.
- The parked request survived the shutdown: fired at a busy 1-worker
pool → parked → pod deleted while parked → worker freed → HTTP=200,
total=1.89s, body hello from: 169.254.17.2 | preserved memory count: 1 |
preserved file counter: 1 — park → resume → route, served by the
Terminating pod during its drain-delay window.
- Pod terminated 14s after deletion — the idle floor, confirmed across
three runs: >13s drain-delay (sequence ran), ≪60s grace (marker
handshake released Envoy's preStop; no `SIGKILL`).
- Log sequence captured verbatim, drain-delay honored to the millisecond
(17:34:41.706 → 17:34:54.707):
  >Shutdown signal received; draining
  >(+13.000s) Draining Envoy  {window: 15s}
  >Envoy drained
  >Starting ext_proc drain
  >ext_proc drain completed within deadline
  >Drain-complete marker written  {path: /var/run/atenet/drain-complete}
  >Shutdown complete


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-07 13:36:08 -07:00

8.6 KiB

Request Parking (atenet router)

Summary

Request parking lets the atenet router hold ("park") an inbound request whose target actor cannot be served yet because of transient worker-pool saturation, retrying the resume until the actor becomes routable or a bounded wait elapses — instead of immediately returning 503 to the client.

Motivation

When a request arrives for a suspended actor, the router resumes it before routing:

Envoy --(ext_proc RequestHeaders)--> router.handleRequestHeaders
    --> ActorResumer.ResumeActor --> ateapi ResumeActor (gRPC)

ateapi's AssignWorkerStep claims a free worker from the actor's WorkerPool. In an oversubscribed system — the core premise of Substrate, where many actors multiplex onto few workers — a burst of traffic can momentarily exhaust the pool. AssignWorkerStep then returns FailedPrecondition: "no free workers available".

Previously the router mapped that straight to an HTTP 503 and failed the request. But such saturation is usually momentary: another actor suspends within milliseconds and frees its worker. Failing fast turns a sub-second blip into a user-visible error.

Behavior

With parking enabled (the default), the router treats FailedPrecondition and Unavailable from ResumeActor as retryable conditions (alongside the existing Aborted concurrent-resume conflict) — a parked request rides out transient pool saturation and control-plane blips (e.g. an ateapi rolling restart) alike. The request is parked: the resumer keeps retrying with exponential backoff until either

  • the resume succeeds (the actor is RUNNING and has a worker IP) — the request is then routed normally; or
  • the park budget (--parked-request-budget, default 5s) elapses — the underlying capacity error is returned, surfacing as 503 "actor <id> unavailable: no free workers available".

To bound resource use and provide backpressure, the router admits requests to a parking lot of fixed capacity (--parked-request-max, default 1024). Each in-flight resume occupies one slot. When the lot is full, further requests are shed immediately with 503 "actor <id> unavailable: router at capacity" rather than queueing without bound.

Every parked request holds one ext_proc stream — one active request against Envoy's ext_proc cluster — for its entire wait, while ordinary requests hold one only for a millisecond-scale header exchange. The cluster's circuit breaker is therefore the hard ceiling on concurrent parked requests. By default the router derives it as twice --parked-request-max (minimum 1024), so the lot always fits and an equal share of fast-path headroom remains — a saturated lot cannot starve requests to already-running actors, at any lot size. --extproc-max-requests overrides the derivation; explicit values are validated >= --parked-request-max at startup, because a breaker below the lot would silently truncate it — Envoy would reject the overflow itself, with 503s that never reach the lot and never count in parking.rejected.

Concurrent requests for the same actor are de-duplicated by the resumer's singleflight group: they share a single in-flight ResumeActor call and all park on its result, so a hot actor consumes N parking slots but only one control-plane RPC.

The park budget is per-flight, not per-request. The budget clock starts when a flight's first caller begins the resume; every later request for the same actor joins that flight and shares its remaining budget and outcome. A request that joins late may therefore see budget_exhausted after waiting far less than a full budget itself — the accepted cost of collapsing a hot actor's requests into one control-plane call. (parking.wait.duration records each request's own parked time, so sub-budget budget_exhausted samples are expected under sustained saturation.)

What is not parked

Only transient conditions — capacity (FailedPrecondition), concurrency (Aborted), and control-plane unavailability (Unavailable) — are parked. Errors that will not resolve by waiting are returned immediately (fail fast):

Resume result Behavior
OK Route to worker
Aborted (concurrent resume) Retry (always)
FailedPrecondition (no free worker) Park & retry (when enabled)
Unavailable (control-plane blip) Park & retry (when enabled)
NotFound Fail fast → 404
DeadlineExceeded Fail fast → 504
PermissionDenied / Unauthenticated Fail fast → 403 / 401

When parking is disabled (--parked-request-max=0), the router fails fast: FailedPrecondition and Unavailable are returned immediately, there is no admission cap, and only Aborted (concurrent-resume) conflicts are retried, within a 15s budget.

Parked requests survive router shutdown

A request parked when the router pod receives SIGTERM is not reset: the shutdown sequence keeps the ext_proc server (and, via a preStop handshake, the Envoy sidecar) alive until in-flight streams finish, and the ext_proc drain deadline (--drain-timeout) defaults to a value derived from --parked-request-budget and is validated at startup to be >= the budget — so a parked request always gets its full budget and a normal verdict (routed 200 or capacity 503) even mid-termination. See the graceful-shutdown knobs (--drain-delay, --drain-timeout) in manifests/ate-install/atenet-router.yaml.

Configuration

Flag Default Meaning
--parked-request-budget 5s Park budget per resume flight; requests de-duplicated onto an in-flight resume share its remaining budget (see Behavior).
--parked-request-max 1024 Max concurrent parked/in-flight resume requests; excess shed (503). 0 disables parking.
--parked-request-retry-interval 100ms Delay before a parked request's first resume retry.
--parked-request-retry-factor 1.1 Multiplier applied to the retry delay after each attempt (>= 1).
--parked-request-retry-jitter 0.1 Random fraction in [0, 1) added per retry to de-synchronize parked requests.
--extproc-max-requests 0 (auto) Envoy circuit-breaker max_requests for the ext_proc cluster. 0 derives twice --parked-request-max (min 1024); explicit values must be >= --parked-request-max (enforced at startup). The excess is fast-path headroom (see Behavior).

The retry backoff deliberately has no cap and no attempt limit: the budget alone bounds the wait.

Observability

Metrics (OpenTelemetry, meter atenet-router):

  • atenet.router.parking.active — up/down counter: requests currently parked.

  • atenet.router.parking.wait.duration — histogram (seconds) of time spent parked. Recorded exactly once per admitted request, at the moment its resume attempt completes; never recorded for shed requests (those only increment parking.rejected) nor when parking is disabled. The outcome label says how the park ended:

    outcome When it is set
    served The resume succeeded and the request was routed to its worker.
    budget_exhausted The park budget elapsed while the resume was still blocked on a retryable condition (pool saturated, a concurrent operation holding the actor, or the control plane unavailable) — the signal that capacity, not a fault, is the bottleneck.
    canceled The client disconnected while parked (request context canceled).
    timeout The request's own deadline expired while parked (distinct from the park budget).
    error The resume failed with a non-retryable error (NotFound, PermissionDenied, ...).
  • atenet.router.parking.rejected — counter: requests shed because the lot was full.

Status page (/statusz): a "Request Parking" card shows whether parking is enabled, the current vs. maximum parked count, and the max wait.