Delete ate.scheduler.eligible_workers (#1804)

Closes #1802.

## What this does

Deletes the `ate.scheduler.eligible_workers` histogram. Nothing replaces
it in this PR.

## Why

**It is expensive on the resume path.** `Schedule` built a second slice
of the matching workers that scheduling never read, walked it again, and
repeated `HasRoom` for each entry. `HasRoom` parses resource quantity
strings up to three times for each worker. The metric roughly doubled
the per-worker arithmetic of a placement. The filter loop is now one
pass over the fleet and builds one slice.

**A metric cannot answer the question it was built for.** #564 wanted to
tell a full fleet apart from an empty intersection of the constraints,
such as a cordoned node behind a node requirement. That needs the
selectors, the node requirement and the actor.
`docs/metrics/substrate.yaml` bars all three from a metric label and
sends them to logs and spans, so no shape of this instrument reaches the
answer. A log record at the rejection is the replacement the issue
names, and it is not in this PR.

## Also removed

`ate.scheduling.constraint` goes with the instrument. No other signal
used the attribute.

## Scope of the change

- `cmd/ateapi/internal/scheduling/metrics.go` deleted, along with the
`WithMeter` option and the second filter pass in `Schedule`.
- The instrument and the attribute removed from
`docs/metrics/registry/metrics.yaml`. The `pool-keys-paired` exception
in `docs/metrics/substrate.yaml` dropped, because the empty pool pair
was this instrument's alone.
- `docs/observability.md`: the table row and the two label notes
removed.
- The e2e collector check and the label assertions for the instrument
removed.
- `TestSchedule_EligibleWorkersMetric` removed. The behaviors it covered
(draining workers, a sandbox class mismatch, busy workers) are already
in the `TestSchedule` table.

## Testing

`go test ./cmd/... ./internal/...` passes. `make verify` passes except
`hack/verify/metrics.sh`, which needs Weaver or Docker. Neither is
available on this machine, so CI is what checks the registry.
This commit is contained in:
Jeff Luo
2026-09-22 21:53:54 +00:00
committed by GitHub
parent 25cdbe00bc
commit 9e503c42b8
10 changed files with 19 additions and 560 deletions
-48
View File
@@ -474,23 +474,6 @@ groups:
stability: development
value: error
brief: The attempt failed. Only this outcome has an error.type key.
- id: ate.scheduling.constraint
stability: development
brief: The type of constraint in the scheduling request.
type:
members:
- id: none
stability: development
value: none
brief: The request has no constraint.
- id: selector
stability: development
value: selector
brief: The request has a label selector for an actor or a template.
- id: required_nodes
stability: development
value: required_nodes
brief: The request names specific node VMs.
- id: registry.ate.router
type: attribute_group
@@ -878,37 +861,6 @@ groups:
requirement_level:
conditionally_required: The outcome is error.
- id: metric.ate.scheduler.eligible_workers
type: metric
metric_name: ate.scheduler.eligible_workers
instrument: histogram
unit: "{worker}"
stability: development
brief: The number of free workers that remain after all the constraint filters, at each scheduling decision.
note: >
This is an early sign of low capacity. It warns you before the first
rejection. The rate of 503 errors tells you only after the users have a
fault. If no pool agrees with the constraints, ateapi sends one series
with the value 0. Both pool keys are empty in that series. Thus a
dashboard shows the "no workers are available" state with the series of
the pools. Only this instrument has the empty pair of pool keys.
annotations:
substrate:
emitted_by: [ateapi]
golden_signals: [saturation]
code_anchor: cmd/ateapi/internal/scheduling/metrics.go
buckets: [0, 1, 2, 3, 5, 10, 20, 50, 100, 250]
cuj: Resumes fail with a 503 error, but ate.workerpool.workers shows idle workers. Why?
attributes:
- ref: ate.workerpool.namespace
requirement_level: required
- ref: ate.workerpool.name
requirement_level: required
- ref: ate.sandbox.class
requirement_level: required
- ref: ate.scheduling.constraint
requirement_level: required
# ---------------------------------------------------------------------------
# Metrics. atelet
# ---------------------------------------------------------------------------
+1 -2
View File
@@ -202,8 +202,7 @@ cardinality_rules:
brief: >
A component sets ate.workerpool.namespace and ate.workerpool.name
together, or it sets no pool key. One key alone does not identify a pool.
ate.scheduler.eligible_workers is the one exception. It sends both keys
empty to show that no worker is available.
No instrument is an exception.
enforced: false
could_be_enforced_by: >
A Weaver Rego policy can make the two keys occur together in this
+1 -7
View File
@@ -246,7 +246,6 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
| `rpc.server.call.duration` | ateapi & atelet (gRPC servers, via `otelgrpc`) | histogram | per-method gRPC latency, request rate, and errors (labels `rpc.method`, `rpc.response.status_code`) |
| `ate.actor.crashes` | ateapi | counter | Number of times actors transitioned to `ACTOR_STATE_CRASHED` with failure reasons (labels `ate.actor.operation.name`, `ate.failure.reason`, `ate.failure.domain`, `ate.template.atespace`, `ate.template.name`, `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`) |
| `atenet.router.route.duration` | atenet-router | histogram | Substrate E2E — Envoy receiving a request to Envoy forwarding it to the resolved worker, excluding actor compute and the response (labels `ate.template.atespace`, `ate.template.name`, `ate.router.outcome`, `ate.router.resume`) |
| `ate.scheduler.eligible_workers` | ateapi | histogram | number of eligible unassigned workers available during scheduling given the constraint filters (labels `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`, `ate.scheduling.constraint`) |
| `atelet.snapshot.size` | atelet | histogram | uncompressed size in bytes of each gVisor snapshot image written during checkpoint (labels `file.name`, `ate.template.atespace`, `ate.template.name`) |
| `ate.workerpool.desired_workers` | atecontroller | up/down counter | number of worker pods requested for a WorkerPool, from `spec.replicas` (labels
`ate.workerpool.namespace`, `ate.workerpool.name`) |
@@ -269,18 +268,13 @@ For `atenet.router.route.duration`:
* `ate.router.outcome` categorizes the route attempt result: `ok`, `cancelled`, `timeout`, `no_capacity`, `failed_precondition`, `lock_conflict`, `not_found`, `unavailable`, `rate_limited`, or `resume_error`.
* `ate.router.resume` indicates the singleflight execution state of actor resumption: `none` (the resume found the actor already running), `triggered` (this request completed a cold activation), `joined` (this request waited on another request's resume, which activated the actor), or `unknown` (the resume did not complete, so whether an activation ran is unknown). `ate.template.atespace` and `ate.template.name` hold `unknown` when the router has no template to name.
For `ate.scheduler.eligible_workers`:
* `ate.scheduling.constraint` categorizes the scheduling request constraint type: `none` (unconstrained), `selector` (actor or template label selectors specified), or `required_nodes` (pinned to specific node VMs).
For `ate.imagecache.requests`:
* `ate.imagecache.outcome` is `hit` when the node holds a complete image record — every layer directory the record names is present — and `miss` when the lookup must pull. A failed lookup is neither: `error` is a failed lookup whatever the cause, and `cancelled` or `timeout` is the caller giving up, as on `ate.router.outcome`. So the hit ratio is `hit / (hit + miss)`, with failures and abandoned lookups out of the denominator.
* `error.type` is present only on the `error` outcome, and carries the registry's own HTTP status for its rejection, from a fixed set: `401`, `403`, `404`, `429`, `500`, `502`, `503`, `504`. The set is an allow-list because the registry client reports whatever the remote returned. Each other status, and each failure that carries no status, reports `_OTHER`.
`ate.workerpool.namespace` and `ate.workerpool.name` identify a pool together, on every instrument that names one. A WorkerPool is a namespaced resource, so the name on its own merges same-named pools from different namespaces into one series. The pair means that capacity (`ate.workerpool.workers`, `ate.workerpool.desired_workers`, `ate.workerpool.ready_workers`) joins to demand (`ate.scheduler.assignment.duration`, `ate.actor.lifecycle.operation.duration`, `ate.actor.crashes`) by pool.
Two states read differently:
* **No keys** means the operation has no pool. The actor-centric instruments omit the pair, so a crash before the actor reached a worker, or the `no_free_worker` outcome, names no pool.
* **Both keys empty** means no pool matched. Only `ate.scheduler.eligible_workers` reports it, as one zero-valued series that keeps "nothing is schedulable" on the same chart as the per-pool series.
**No keys** means the operation has no pool. The actor-centric instruments omit the pair, so a crash before the actor reached a worker, or the `no_free_worker` outcome, names no pool. No instrument sends the pair with both keys empty.
The three snapshot labels are orthogonal and mean the same thing on every histogram that carries them:
* `ate.snapshot.kind`: which snapshot the operation reads or writes. `local` (node-local, written by a pause), `latest` (the actor's own durable snapshot), `golden` (the template's image), or `boot` (from scratch, so it never appears on the atelet histograms).