Delete the ateerrors failure-reason taxonomy (#1817)

Stacked on #1220 — review only the last commit until that merges
This commit is contained in:
Zoe Zhao
2026-09-23 20:33:21 +00:00
committed by GitHub
parent 21b0b030bc
commit 74bbfc529c
37 changed files with 210 additions and 1317 deletions
+9 -17
View File
@@ -74,8 +74,8 @@ Rules of thumb:
type is a promise to the pipeline about what the datapoints are, and
rollups and the actor relay act on that promise without checking.
* **Do not add a failure counter next to a success counter.** One instrument,
with the failure on `error.type` or `ate.failure.reason`; the key's absence
means success. See [Reporting failures](#reporting-failures).
with the failure on `error.type`; the key's absence means success. See
[Reporting failures](#reporting-failures).
* **One histogram, many phases**, rather than one histogram per phase, when
the phases share dimensions. `ate.actor.restore.duration` carries
`ate.snapshot.phase` as a label so the phases sit on one chart.
@@ -165,8 +165,7 @@ ones every new metric meets:
failed before a worker was picked has no pool; the pool keys are absent, not
`""`. `ateattr.WorkerPoolAttributes` returns nil for that case.
* **Paired keys travel together.** `ate.workerpool.namespace` with
`ate.workerpool.name`; `ate.failure.reason` with `ate.failure.domain`. Use
the `ateattr` helpers that return the pair.
`ate.workerpool.name`. Use the `ateattr` helpers that return the pair.
Every label key is a constant in `internal/ateattr/ateattr.go`, and every
bounded value set is a group of constants there beside it. A metric never
@@ -191,18 +190,11 @@ the cluster.
## Reporting failures
One instrument carries success and failure. Success is the absence of the
failure key. Which key depends on where the error is classified:
* **`error.type`** at an RPC boundary, holding the gRPC status code
(`status.Code(err).String()`), as `ate.actor.lifecycle.operation.duration`
does. Also for a bounded set of protocol statuses, as
`ate.imagecache.requests` does with an allow-list of HTTP codes and `_OTHER`
for everything else.
* **`ate.failure.reason` and `ate.failure.domain`** together, when the failure
is one of substrate's own reasons (`internal/ateerrors`). Use
`ateattr.FailureAttributes(reason)`; never set the two keys by hand. This is
the right choice inside atelet and ateom handlers, where the gRPC status is
assigned only after the handler returns and would read `Unknown`.
failure key, `error.type`: at an RPC boundary it holds the gRPC status code
(`status.Code(err).String()`), as `ate.actor.lifecycle.operation.duration`
does. It also fits a bounded set of protocol statuses, as
`ate.imagecache.requests` does with an allow-list of HTTP codes and `_OTHER`
for everything else.
Separate a caller that gave up from a failure: `context.Canceled` and
`context.DeadlineExceeded` are their own outcomes (`cancelled`, `timeout`) on
@@ -451,7 +443,7 @@ author can find the saturation instruments without reading each brief), and
`buckets` for a histogram; a reader of the registry should not need the code.
Attributes that already exist (`ate.template.name`, `ate.snapshot.kind`,
`error.type`, `ate.failure.reason`) are referenced with `ref:`, not redefined.
`error.type`) are referenced with `ref:`, not redefined.
Mark a label `conditionally_required` and say when, rather than `required`, if
the code omits it on some path.
+1 -5
View File
@@ -81,7 +81,7 @@ groups:
substrate:
emitted_by: [ateapi]
code_anchor: cmd/ateapi/internal/controlapi/crash.go
cuj: Which actor was lost, and why?
cuj: Which actor was lost?
attributes:
- ref: ate.atespace
requirement_level: required
@@ -97,7 +97,3 @@ groups:
requirement_level: required
- ref: ate.actor.state
requirement_level: required
- ref: ate.failure.reason
requirement_level: required
- ref: ate.failure.domain
requirement_level: required
+5 -130
View File
@@ -106,7 +106,7 @@ groups:
- id: crashed
stability: development
value: crashed
brief: The actor is lost. ate.failure.reason says why.
brief: The actor is lost.
- id: deleting
stability: development
value: deleting
@@ -345,112 +345,6 @@ groups:
value: total
brief: The full operation. Use this value as the denominator. Do not add the phases together.
- id: registry.ate.failure
type: attribute_group
brief: >
The failure reasons of Substrate. These instruments use this key and not
error.type. The handler returns a domain error. The interceptor sets the
gRPC status only after the handler returns. Thus error.type shows Unknown
for almost each real failure.
attributes:
- id: ate.failure.reason
stability: development
brief: The cause of the failure.
type:
members:
- id: terminal_file_system_error
stability: development
value: TERMINAL_FILE_SYSTEM_ERROR
brief: A permanent file system fault on the node. Usually the disks of the node are full.
- id: invalid_sandbox_asset
stability: development
value: INVALID_SANDBOX_ASSET
brief: A sandbox asset is absent or bad.
- id: invalid_checkpoint_result
stability: development
value: INVALID_CHECKPOINT_RESULT
brief: ateom returned a checkpoint that is not valid.
- id: failed_save_snapshot
stability: development
value: FAILED_SAVE_SNAPSHOT
brief: atelet could not write the snapshot.
- id: invalid_object_url
stability: development
value: INVALID_OBJECT_URL
brief: The URL of the snapshot object is bad.
- id: failed_get_external_object
stability: development
value: FAILED_GET_EXTERNAL_OBJECT
brief: atelet could not read the snapshot from object storage. Usually the storage backend has a fault.
- id: invalid_container_config
stability: development
value: INVALID_CONTAINER_CONFIG
brief: The container configuration of the ActorTemplate is not valid.
- id: local_snapshot_gone
stability: development
value: LOCAL_SNAPSHOT_GONE
brief: The local snapshot on the node no longer exists.
- id: workload_not_ready
stability: development
value: WORKLOAD_NOT_READY
brief: >
A container started but never passed its wakeup probe before the
deadline of the probe. A slow node confounds this value: a cold
image or a throttled CPU also runs the probe out of time. The
phases of ate.actor.restore.duration separate the two.
- id: corrupted_assignment
stability: development
value: CORRUPTED_ASSIGNMENT
brief: Control plane. The worker assignment of the actor is not correct.
- id: worker_reassigned
stability: development
value: WORKER_REASSIGNED
brief: Control plane. A different actor got the worker.
- id: worker_pod_gone
stability: development
value: WORKER_POD_GONE
brief: Control plane. The worker pod no longer exists.
- id: unknown
stability: development
value: UNKNOWN
brief: >
A failure with no reason. Substrate does not make a value from
the error message. Thus the list of values stays short. This
value asserts no fault domain: read ate.failure.domain, which
reports unknown here and not infrastructure.
- id: ate.failure.domain
stability: development
brief: >
The side of the platform boundary the failure came from. A strict
function of ate.failure.reason, thus it adds no series. It exists
because an operator triages on the boundary and drills into the
reason, and because a consumer must not have to match on the names of
the reasons: a component ahead of the control plane can report a
reason that the control plane does not know, which becomes UNKNOWN,
and a name match would then count it as an infrastructure fault.
type:
members:
- id: infrastructure
stability: development
value: infrastructure
brief: Substrate failed. The node, the store, the object storage, the control plane.
- id: workload
stability: development
value: workload
brief: >
The owner of the actor fixes this, not the operator of the
platform. It covers a bad ActorTemplate as well as a process
that does not start. For the reasons that the actor reports
about itself, such as an exit or a probe that never passes,
this value is a report and not a proof: the actor chooses its
own exit. The reasons that Substrate measures, such as a
template that resolves to no runnable process, are not.
- id: unknown
stability: development
value: unknown
brief: The reason gives no domain. A gap in the taxonomy stays visible here and does not inflate one side.
- id: registry.ate.scheduler
type: attribute_group
brief: The labels of a scheduler decision.
@@ -707,14 +601,14 @@ groups:
instrument: counter
unit: "{crash}"
stability: development
brief: The number of actors that went to the ACTOR_STATE_CRASHED state, with the reasons.
brief: The number of actors that went to the ACTOR_STATE_CRASHED state.
note: >
This counter counts the moves into the CRASHED state of the state machine.
It does not count the crashes of a process. The CRASHED state is a loss of
the actor's live execution (recoverable back to SUSPENDED via RevertActor).
Substrate can also lose the data that it did not write. Some moves into
this state are not a loss of data. They are careful responses to a control
plane problem. The ate.failure.reason key keeps these groups separate.
plane problem.
ate.sandbox.class is unknown when ateapi could not read the class of the
worker: the worker record is already gone, or its assignment is already
clear.
@@ -723,14 +617,10 @@ groups:
emitted_by: [ateapi]
golden_signals: [errors]
code_anchor: cmd/ateapi/internal/controlapi/metrics.go
cuj: Some actors go to the CRASHED state. How many, and why?
cuj: Some actors go to the CRASHED state. How many?
attributes:
- ref: ate.actor.operation.name
requirement_level: required
- ref: ate.failure.reason
requirement_level: required
- ref: ate.failure.domain
requirement_level: required
- ref: ate.template.atespace
requirement_level: required
- ref: ate.template.name
@@ -875,10 +765,7 @@ groups:
note: >
This instrument shows where the cold start time goes after ateapi hands
the work to atelet. The phases occur at the same time. They do not add up
to the total. Use the total phase as the denominator. If the operation
fails, atelet puts ate.failure.reason on the phase that failed and on the
total. It puts the key on no other phase. Thus you can still query the
phases that were correct.
to the total. Use the total phase as the denominator.
annotations:
substrate:
emitted_by: [atelet]
@@ -901,12 +788,6 @@ groups:
- ref: ate.sandbox.class
requirement_level:
conditionally_required: The class is known.
- ref: ate.failure.reason
requirement_level:
conditionally_required: The operation failed. The key occurs only on the phase that failed and on the total.
- ref: ate.failure.domain
requirement_level:
conditionally_required: ate.failure.reason occurs. The two keys always occur together.
- id: metric.ate.actor.checkpoint.duration
type: metric
@@ -941,12 +822,6 @@ groups:
- ref: ate.sandbox.class
requirement_level:
conditionally_required: The class is known.
- ref: ate.failure.reason
requirement_level:
conditionally_required: The operation failed. The key occurs only on the phase that failed and on the total.
- ref: ate.failure.domain
requirement_level:
conditionally_required: ate.failure.reason occurs. The two keys always occur together.
- id: metric.atelet.snapshot.size
type: metric
+3 -32
View File
@@ -162,42 +162,13 @@ cardinality_rules:
reason. Only `weaver registry live-check` can count the values that occur.
- id: error-type-not-parallel-counters
brief: >
A component reports a failure with error.type or ate.failure.reason on the
same instrument. If the key is absent, the operation was correct. There
are no separate counters for the failures.
A component reports a failure with error.type on the same instrument. If
the key is absent, the operation was correct. There are no separate
counters for the failures.
enforced: false
could_be_enforced_by: >
A Weaver Rego policy can refuse a metric name that ends with errors_total
and has a counterpart without the suffix.
- id: failure-keys-paired
brief: >
A component sets ate.failure.reason and ate.failure.domain together, or it
sets neither. ateattr.FailureAttributes and ateattr.FailureLogAttrs return
the pair, and a producer uses one of them rather than the keys.
The domain is a strict function of the reason, so it adds no series. It is
a key and not a rule in the consumer because the reasons of a component
ahead of the control plane can be unknown to the control plane, which
makes them UNKNOWN, and a consumer that matched on the names of the
reasons would count each of those as an infrastructure fault.
enforced: false
could_be_enforced_by: >
A Weaver Rego policy can make the two keys occur together in this
registry. Only live-check can see a component that sends one key.
- id: workload-domain-is-a-report
brief: >
Some of the reasons in the workload domain are what the actor reported
about itself: an exit code, or a probe that never passes. The actor
chooses those, thus it decides the value. It can raise the count of the
failures of its own template and pool, and it can exit 0 to hide a fault.
A dashboard says "reported" for those, and an alert on them has a rate
limit; one report for each activation is the bound. The reasons that
Substrate measures, such as a template that resolves to no runnable
process, carry no such caveat.
enforced: false
could_be_enforced_by: >
Nothing. This is a property of a multi-tenant platform and not of the
registry. It is here so that a reader of the panel knows what the number
is.
- id: pool-keys-paired
brief: >
A component sets ate.workerpool.namespace and ate.workerpool.name
+6 -17
View File
@@ -137,7 +137,7 @@ atelet's `Restore timing breakdown` is the first of these, and the only unsample
"trace_id":"4bf92f…","span_id":"00f067…","trace_flags":"01"}
```
The duration keys are the [`ate.actor.restore.duration`](#the-metric-registry) instrument's name with an `ate.snapshot.phase` value appended, and they hold **seconds**, matching that instrument's declared unit. The same rules apply as on the histogram: a phase that never ran is absent rather than zero, phases overlap and do not sum to the total, and `ate.failure.reason` is present only when the restore failed. There is no `ate.snapshot.phase` key on the record — on a datapoint it names the one step timed, and this record carries them all.
The duration keys are the [`ate.actor.restore.duration`](#the-metric-registry) instrument's name with an `ate.snapshot.phase` value appended, and they hold **seconds**, matching that instrument's declared unit. The same rules apply as on the histogram: a phase that never ran is absent rather than zero, and phases overlap and do not sum to the total. There is no `ate.snapshot.phase` key on the record — on a datapoint it names the one step timed, and this record carries them all.
This is the record to use for a per-actor wake-up distribution. The histogram cannot answer that question at all, because actor identity is barred from metric labels; traces can, but the data plane is head-sampled at 1%.
@@ -170,11 +170,10 @@ Creating an actor counts as a change. A new actor is born suspended, so it gets
"ate.atespace":"ate-demo-counter","ate.actor.name":"counter-1","ate.actor.uid":"8f2a…",
"ate.template.atespace":"ate-demo-counter","ate.template.name":"counter",
"ate.actor.operation.name":"resume","ate.actor.state":"crashed",
"ate.failure.reason":"WORKER_POD_GONE","ate.failure.domain":"infrastructure",
"trace_id":"4bf92f…","span_id":"00f067…","trace_flags":"01"}
```
The counter carries the same reason but no actor identity, so this record is the only way to attribute a crash to one agent. The decision-point line that precedes it (`Setting Actor to crashed due to error`) carries only `ate.atespace` and `ate.actor.name`: it is written before the Actor is loaded, so no uid exists yet.
The counter carries no actor identity, so this record is the only way to attribute a crash to one agent. The decision-point line that precedes it (`Setting Actor to crashed due to error`) carries only `ate.atespace` and `ate.actor.name`: it is written before the Actor is loaded, so no uid exists yet.
#### The same records over OTLP
@@ -185,9 +184,9 @@ Two `event.name` values, which is the OTLP LogRecord's own field rather than an
| `event.name` | Body | Severity | Attributes |
|---|---|---|---|
| `ate.actor.state_changed` | `Actor state changed` | 9 | the five identity keys, `ate.actor.operation.name`, `ate.actor.state` |
| `ate.actor.crashed` | `Actor crashed` | 17 | the above plus `ate.failure.reason`, `ate.failure.domain` |
| `ate.actor.crashed` | `Actor crashed` | 17 | the same keys |
A crash is its own name because an event name promises a fixed set of attributes, and it carries two more. There is no name per state: `ate.actor.state` already says which transition happened, so a consumer still selects on that one attribute and needs no map from a name to a state. Both names are in [`docs/metrics/registry/events.yaml`](metrics/registry/events.yaml), which `make verify` checks.
A crash is its own name because an event name promises a fixed set of attributes and a crash has a different severity and shape. There is no name per state: `ate.actor.state` already says which transition happened, so a consumer still selects on that one attribute and needs no map from a name to a state. Both names are in [`docs/metrics/registry/events.yaml`](metrics/registry/events.yaml), which `make verify` checks.
The attributes are the same flat `ate.*` keys as the stdout copy, so they arrive as real log attributes with no transform in front of them. Trace context is not among them: it goes on the record's own `TraceId` and `SpanId` fields, where the stdout copy's top-level `trace_id`/`span_id` would be mapped to anyway. The instrumentation scope is `github.com/agent-substrate/substrate/internal/actorevent`, which is how you select this stream, or exclude it.
@@ -225,14 +224,6 @@ labels."ate.actor.uid"="8f2a…" AND jsonPayload.msg="Actor usage sample"
**Do not put actor identity on a log-based metric.** Aggregating these events in the log store is what they are for, but a log-based metric built over them must label only by the bounded set (template, sandbox class, source, pool) — promoting `ate.actor.uid` or `ate.actor.name` into a metric label reintroduces exactly the per-actor cardinality this split keeps out of the TSDB.
### Which side failed: `ate.failure.domain`
`ate.failure.reason` names the cause; `ate.failure.domain` names the side of the platform boundary it came from, as `infrastructure`, `workload`, or `unknown`. It rides on every signal that carries a reason, and the two are always emitted together — producers call `ateattr.FailureAttributes` or `ateattr.FailureLogAttrs` rather than setting either key directly.
The domain is a strict function of the reason, so it costs no series and no consumer needs it to disambiguate a value. It exists because the alternative is every consumer keeping its own map from reason to domain, and that map breaks silently: a component running ahead of the control plane can report a reason this build's `ateerrors.AllReasons` rejects, `ExtractReason` turns it into `UNKNOWN`, and a name-matching consumer would file every one of those as an infrastructure fault. `UNKNOWN` therefore reports `unknown` and not `infrastructure` — a gap in the taxonomy stays visible instead of inflating one side.
One caveat belongs on any panel built from this. A `workload` domain says what the actor **reported**, not what substrate measured: the actor picks its own exit and can hold memory until the kernel kills it, so it can raise the failure count for its own template and pool, or exit cleanly to hide a fault. See `workload-domain-is-a-report` in [`docs/metrics/substrate.yaml`](metrics/substrate.yaml).
---
## 2. Metrics
@@ -244,7 +235,7 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
| Metric | Emitted by | Type | Measures |
|--------|------------|------|----------|
| `rpc.server.call.duration` | ateapi & atelet (gRPC servers, via `otelgrpc`) | histogram | per-method gRPC latency, request rate, and errors (labels `rpc.method`, `rpc.response.status_code`) |
| `ate.actor.crashes` | ateapi | counter | Number of times actors transitioned to `ACTOR_STATE_CRASHED` with failure reasons (labels `ate.actor.operation.name`, `ate.failure.reason`, `ate.failure.domain`, `ate.template.atespace`, `ate.template.name`, `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`) |
| `ate.actor.crashes` | ateapi | counter | Number of times actors transitioned to `ACTOR_STATE_CRASHED` (labels `ate.actor.operation.name`, `ate.template.atespace`, `ate.template.name`, `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`) |
| `atenet.router.route.duration` | atenet-router | histogram | Substrate E2E — Envoy receiving a request to Envoy forwarding it to the resolved worker, excluding actor compute and the response (labels `ate.template.atespace`, `ate.template.name`, `ate.router.outcome`, `ate.router.resume`) |
| `atelet.snapshot.size` | atelet | histogram | uncompressed size in bytes of each gVisor snapshot image written during checkpoint (labels `file.name`, `ate.template.atespace`, `ate.template.name`) |
| `ate.workerpool.desired_workers` | atecontroller | up/down counter | number of worker pods requested for a WorkerPool, from `spec.replicas` (labels
@@ -254,7 +245,7 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
| `ate.workerpool.workers` | ateapi | up/down counter | live worker count per pool, split by state (`idle`/`assigned`) and sandbox class to provide fleet capacity and saturation at a glance |
| `ate.actor.lifecycle.operation.duration` | ateapi | histogram | how long each actor operation (create/resume/suspend/pause/delete/revert) takes and whether it failed (`error.type` present = failure, absent = success); labeled by operation, template, pool (`ate.workerpool.namespace` + `ate.workerpool.name`), sandbox class, and snapshot kind and scope on resume; already-running resume no-ops are not recorded so the histogram tracks actual activations, not router traffic |
| `ate.scheduler.assignment.duration` | ateapi | histogram | time it takes for an actor to be assigned to a worker, per attempt (version-conflict retries record only the final attempt), with the outcome (`assigned` / `no_free_worker` / `error`), the assigned pool (`ate.workerpool.namespace` + `ate.workerpool.name`) and sandbox class to catch scheduling latency and capacity starvation problems |
| `ate.actor.restore.duration` | atelet | histogram | how long each phase of a restore takes on the worker node, which is where cold-start latency actually goes once ateapi hands off (labels `ate.snapshot.phase`, `ate.snapshot.kind`, `ate.snapshot.scope`, `ate.template.atespace`, `ate.template.name`, `ate.sandbox.class`, plus `ate.failure.reason` and `ate.failure.domain` on failure) |
| `ate.actor.restore.duration` | atelet | histogram | how long each phase of a restore takes on the worker node, which is where cold-start latency actually goes once ateapi hands off (labels `ate.snapshot.phase`, `ate.snapshot.kind`, `ate.snapshot.scope`, `ate.template.atespace`, `ate.template.name`, `ate.sandbox.class`) |
| `ate.actor.checkpoint.duration` | atelet | histogram | the same phase breakdown for writing a snapshot, so a slow suspend can be attributed to ateom or to the upload (same labels as the restore histogram) |
| `ate.imagecache.requests` | atelet | counter | image lookups in the node-local image cache, by outcome (`ate.imagecache.outcome`), with `error.type` on the `error` outcome. A miss pays for the pull and the unpack, so the hit ratio per node is a leading indicator of resume latency |
@@ -283,8 +274,6 @@ The three snapshot labels are orthogonal and mean the same thing on every histog
**Phases overlap and do not sum to `total`.** The download runs concurrently with the asset fetch and OCI unpack, so each is an independent observation; use `total` as the denominator. A phase that never started is absent rather than zero.
On a failure, `ate.failure.reason` marks the phase that died and the `total`, and nothing else, so `ate.actor.restore.duration{ate.snapshot.phase="download", ate.failure.reason!=""}` says how often the download is what breaks and why, while the phases that succeeded stay queryable as successes. The atelet histograms classify with substrate's own reason taxonomy (the same one `ate.actor.crashes` uses) rather than `error.type`, because these handlers return wrapped domain errors and the gRPC status is only assigned after the handler returns, so a status code would read `Unknown` for nearly every real failure. A failure that carries no reason reports `UNKNOWN`. [`ate.failure.domain`](#which-side-failed-atefailuredomain) rides alongside wherever the reason is present, so a restore that failed because the actor never passed its wakeup probe is separable from one the node broke.
The `ate.*` control-plane metric labels are either fixed value sets (operation, outcome, state, class, kind, scope, phase) or scoped to the deployment catalog (template and pool names are operator-created, never derived from request payloads), and the label set varies per operation: resume carries the most dimensions, delete only the operation and error type. `ate.sandbox.class` is derived from the template (each template has exactly one class), so it adds no extra series next to the template labels; it exists so dashboards can aggregate by class without enumerating template names. High-cardinality actor identity (name/uid/atespace) stays off metrics entirely and lives on logs and traces instead — for resource usage, on the [per-actor usage events](#per-actor-usage-events).
### The metric registry