multi actor worker support (#1836)

Fixes #1266 

Builds on #1689 #1283 in particular.

Outstanding: 
- We need to rethink how HPA will work, I've punted that from this PR,
but it's probably # 1 on the list for follow-up tasks.
~~- There are some resource management / cleanup bugs that are
pre-existing. I'm trying to keep this PR size down but do plan to submit
fixes. Both gVisor and uVM workers need improvements to dealing with
hanging sandbox processes from previous actors. This is more concerning
with multi-actor but not new.~~

EDIT:
1. HPA is just a POC right now anyhow, we think this is fine and we'll
need something more sophisticated later
2. I fixed most of these.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This commit is contained in:
Benjamin Elder
2026-09-24 21:25:22 +00:00
committed by GitHub
parent fa24dcc514
commit 14c0c136bc
43 changed files with 3305 additions and 1124 deletions
+3 -2
View File
@@ -61,9 +61,10 @@ moves. A pinned pool cannot put those pods back on a moved node, so the old
version drains away node by node. An unpinned pool breaks this constraints.
#### Worker Capacity (`spec.template.resources`)
Setting `resources.limits` (CPU and Memory) on a `WorkerPool` establishes each worker pod's **capacity** — the envelope available to host an actor sandbox, taken from the `ateom` container's limits. The scheduler only places an actor on a worker whose capacity is `>=` the actor's declared resource limits (see [Sandbox Right-Sizing](#sandbox-right-sizing-resources) on the `ActorTemplate`).
Setting `resources.limits` (CPU and Memory) on a `WorkerPool` establishes each worker pod's **capacity** — the envelope its actor sandboxes share, taken from the `ateom` container's limits. The scheduler only places an actor on a worker whose remaining capacity is `>=` the actor's declared resource limits (see [Sandbox Right-Sizing](#sandbox-right-sizing-resources) on the `ActorTemplate`).
- Size a pool's `limits` to the largest actor it should host. An actor occupies its whole worker, so worker capacity is the per-actor ceiling, not a shared budget.
- Worker capacity is a shared budget: each actor placed on a worker subtracts its declared limits from what is left. Size a pool's `limits` for the actors it should host together.
- A worker also has an actor limit, set by the ateom's `--max-actors` flag (default 1000). Placement stops at whichever runs out first.
- Capacity is advisory for placement only: a worker that declares no CPU/memory limit reports zero capacity for that dimension, which the scheduler treats as **unconstrained** (placement is never blocked by missing data). The actual sandbox size still comes from the `ActorTemplate`.
### Example
+2 -2
View File
@@ -46,8 +46,8 @@ for etcd.
status and snapshot references.
- **Worker**: a record representing one worker pod in a `WorkerPool`. A Worker
hosts at most one Actor at a time; many Actors are multiplexed across a pool
over time.
hosts several Actors at once, each in its own sandbox, up to its actor limit
and its compute capacity; many more are multiplexed across a pool over time.
## Components