(ateapi): report worker occupancy by actor slots in ate.worker.state (#1854)

Part of #1664, implements #1799.

## Changes

- `ate.worker.state` now has `idle`, `partial`, `at_capacity` and
`unschedulable`. `assigned` is removed.
- ateapi picks the state by comparing allocated actor slots with
capacity, the same check the scheduler makes.
- Draining workers and workers with no reported capacity are
`unschedulable`. This wins over occupancy, so the sum over the states is
still the pool size.
- Every known pool reports all four states, set to 0 when empty.
- Updated the registry, `docs/observability.md` and the
autoscaled-workerpool demo (`assigned` -> `at_capacity`).

## Open questions

- Are we fine with `unschedulable` as a fourth state? cc. @JeffLuoo
Every new worker starts with capacity 0, so `at_capacity` would make an
HPA on it scale up on its own scale-up.
- This breaks queries on `ate_worker_state="assigned"`. Should be still
fine.
- The demo HPA is only right while each worker holds one actor. Pool
utilization needs slot counts, which needs a new instrument (#1664).

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This commit is contained in:
Krisztian F
2026-09-24 20:17:26 +00:00
committed by GitHub
parent 5fb129d87d
commit 945e44a5ec
10 changed files with 216 additions and 79 deletions
+25 -8
View File
@@ -203,18 +203,32 @@ groups:
- id: ate.worker.state
stability: development
brief: >
The assignment state of a worker. The key starts with worker and not
with the name of the pool. Thus you can add related keys later.
How full a worker is, by actor slots. The key starts with worker and
not with the name of the pool. Thus you can add related keys later.
note: >
The state compares the allocated actors with the actor capacity. It
does not look at CPU or memory: whether those have room depends on
the size of the next actor. unschedulable wins over the other states.
type:
members:
- id: idle
stability: development
value: idle
brief: No actor has this worker.
- id: assigned
brief: The worker can take actors and holds none.
- id: partial
stability: development
value: assigned
brief: An actor has this worker.
value: partial
brief: The worker holds actors and has free actor slots.
- id: at_capacity
stability: development
value: at_capacity
brief: The worker has no free actor slot.
- id: unschedulable
stability: development
value: unschedulable
brief: >
The scheduler places no actor on the worker. It is not active,
or it has not reported an actor capacity yet.
- id: registry.ate.sandbox
type: attribute_group
@@ -645,9 +659,12 @@ groups:
ateapi reads this value with a callback. You can add the counts together.
The sum across the states is the size of the pool. The sum across the
pools is the size of the fleet. Thus this instrument is an UpDownCounter
and not a gauge. ateapi sets both states to 0 for each known pool. Thus a
and not a gauge. ateapi sets every state to 0 for each known pool. Thus a
full pool or an empty pool reports 0. An absent series would stop an alert
on idle == 0.
The states count workers, not actor slots. A worker with 1 of 20 slots
used and one with 19 of 20 are both partial. Do not read utilization of
the pool from this instrument.
A worker with the class unknown is not capacity of the pool: the scheduler
puts no actor on it. Thus a query that asks how much capacity a pool has
must keep ate.sandbox.class, because a sum across the classes adds that
@@ -985,7 +1002,7 @@ groups:
atecontroller reads this value with a callback. This is a separate
instrument and not a state key on a shared instrument. This agrees with
the Kubernetes conventions (k8s.deployment.desired_pods). The sum of
desired and ready has no meaning. But the sum of idle and assigned on
desired and ready has no meaning. But the sum across the states of
ate.workerpool.workers has a meaning.
annotations:
substrate:
+1 -1
View File
@@ -242,7 +242,7 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
`ate.workerpool.namespace`, `ate.workerpool.name`) |
| `ate.workerpool.ready_workers` | atecontroller | up/down counter | number of worker pods currently ready for a WorkerPool, from `status.readyReplicas` (labels
`ate.workerpool.namespace`, `ate.workerpool.name`) |
| `ate.workerpool.workers` | ateapi | up/down counter | live worker count per pool, split by state (`idle`/`assigned`) and sandbox class to provide fleet capacity and saturation at a glance |
| `ate.workerpool.workers` | ateapi | up/down counter | live worker count per pool, split by state (`idle`/`partial`/`at_capacity`/`unschedulable`) and sandbox class to provide fleet capacity and saturation at a glance |
| `ate.actor.lifecycle.operation.duration` | ateapi | histogram | how long each actor operation (create/resume/suspend/pause/delete/revert) takes and whether it failed (`error.type` present = failure, absent = success); labeled by operation, template, pool (`ate.workerpool.namespace` + `ate.workerpool.name`), sandbox class, and snapshot kind and scope on resume; already-running resume no-ops are not recorded so the histogram tracks actual activations, not router traffic |
| `ate.scheduler.assignment.duration` | ateapi | histogram | time it takes for an actor to be assigned to a worker, per attempt (version-conflict retries record only the final attempt), with the outcome (`assigned` / `no_free_worker` / `error`), the assigned pool (`ate.workerpool.namespace` + `ate.workerpool.name`) and sandbox class to catch scheduling latency and capacity starvation problems |
| `ate.actor.restore.duration` | atelet | histogram | how long each phase of a restore takes on the worker node, which is where cold-start latency actually goes once ateapi hands off (labels `ate.snapshot.phase`, `ate.snapshot.kind`, `ate.snapshot.scope`, `ate.template.atespace`, `ate.template.name`, `ate.sandbox.class`) |