- Adopt the ParkedRequest* vocabulary for parking flags and config (bowei's suggestion): --parked-request-budget / --parked-request-max, matching fields and default consts. - Make the parked-retry backoff configurable: --parked-request-retry-interval/-factor/-jitter, validated at startup (factor >= 1, jitter in [0,1)); the backoff still has no cap and no attempt limit, so the budget alone bounds the wait. - Resolve the effective parking config once in Run() so the resumer's retry loop and the Envoy ext_proc timeout always agree, even when the budget flag is set non-positive. - Drop timeline-relative wording from docs and identifiers (failFastResumeBudget, fail-fast behavior). - Guard the parking-lot counter against going negative, loudly. - Document exactly when the wait-duration metric is recorded and what each outcome label means.
5.9 KiB
Request Parking (atenet router)
Summary
Request parking lets the atenet router hold ("park") an inbound request
whose target actor cannot be served yet because of transient worker-pool
saturation, retrying the resume until the actor becomes routable or a bounded
wait elapses — instead of immediately returning 503 to the client.
Motivation
When a request arrives for a suspended actor, the router resumes it before routing:
Envoy --(ext_proc RequestHeaders)--> router.handleRequestHeaders
--> ActorResumer.ResumeActor --> ateapi ResumeActor (gRPC)
ateapi's AssignWorkerStep claims a free worker from the actor's WorkerPool.
In an oversubscribed system — the core premise of Substrate, where many actors
multiplex onto few workers — a burst of traffic can momentarily exhaust the
pool. AssignWorkerStep then returns FailedPrecondition: "no free workers available".
Previously the router mapped that straight to an HTTP 503 and failed the
request. But such saturation is usually momentary: another actor suspends within
milliseconds and frees its worker. Failing fast turns a sub-second blip into a
user-visible error.
Behavior
With parking enabled (the default), the router treats FailedPrecondition from
ResumeActor as a retryable condition (alongside the existing Aborted
concurrent-resume conflict). The request is parked: the resumer keeps retrying
with exponential backoff until either
- the resume succeeds (the actor is
RUNNINGand has a worker IP) — the request is then routed normally; or - the park budget (
--parked-request-budget, default5s) elapses — the underlying capacity error is returned, surfacing as503 "actor <id> unavailable: no free workers available".
To bound resource use and provide backpressure, the router admits requests to a
parking lot of fixed capacity (--parked-request-max, default 2048). Each
in-flight resume occupies one slot. When the lot is full, further requests are
shed immediately with 503 "actor <id> unavailable: router at capacity" rather
than queueing without bound.
Concurrent requests for the same actor are de-duplicated by the resumer's
singleflight group: they share a single in-flight ResumeActor call and all
park on its result, so a hot actor consumes N parking slots but only one
control-plane RPC.
What is not parked
Only transient capacity (FailedPrecondition) and concurrency (Aborted)
conditions are parked. Errors that will not resolve by waiting are returned
immediately (fail fast):
| Resume result | Behavior |
|---|---|
OK |
Route to worker |
Aborted (concurrent resume) |
Retry (always) |
FailedPrecondition (no free worker) |
Park & retry (when enabled) |
NotFound |
Fail fast → 404 |
Unavailable |
Fail fast → 503 |
DeadlineExceeded |
Fail fast → 504 |
PermissionDenied / Unauthenticated |
Fail fast → 403 / 401 |
When parking is disabled (--parked-request-max=0), the router fails fast:
FailedPrecondition is returned immediately, there is no admission cap, and
only Aborted (concurrent-resume) conflicts are retried, within a 15s budget.
Configuration
| Flag | Default | Meaning |
|---|---|---|
--parked-request-budget |
5s |
Max time a single request may stay parked awaiting resume. |
--parked-request-max |
2048 |
Max concurrent parked/in-flight resume requests; excess shed (503). 0 disables parking. |
--parked-request-retry-interval |
100ms |
Delay before a parked request's first resume retry. |
--parked-request-retry-factor |
1.1 |
Multiplier applied to the retry delay after each attempt (>= 1). |
--parked-request-retry-jitter |
0.1 |
Random fraction in [0, 1) added per retry to de-synchronize parked requests. |
The retry backoff deliberately has no cap and no attempt limit: the budget alone bounds the wait.
Observability
Metrics (OpenTelemetry, meter atenet-router):
-
atenet.router.parking.active— up/down counter: requests currently parked. -
atenet.router.parking.wait.duration— histogram (seconds) of time spent parked. Recorded exactly once per admitted request, at the moment its resume attempt completes; never recorded for shed requests (those only incrementparking.rejected) nor when parking is disabled. Theoutcomelabel says how the park ended:outcomeWhen it is set servedThe resume succeeded and the request was routed to its worker. budget_exhaustedThe park budget elapsed while the resume was still blocked on a retryable condition (pool saturated, or a concurrent operation holding the actor) — the signal that capacity, not a fault, is the bottleneck. canceledThe client disconnected while parked (request context canceled). timeoutThe request's own deadline expired while parked (distinct from the park budget). errorThe resume failed with a non-retryable error ( NotFound,Unavailable, ...). -
atenet.router.parking.rejected— counter: requests shed because the lot was full.
Status page (/statusz): a "Request Parking" card shows whether parking is
enabled, the current vs. maximum parked count, and the max wait.