mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Delete the ateerrors failure-reason taxonomy (#1817)
Stacked on #1220 — review only the last commit until that merges
This commit is contained in:
@@ -74,8 +74,8 @@ Rules of thumb:
|
||||
type is a promise to the pipeline about what the datapoints are, and
|
||||
rollups and the actor relay act on that promise without checking.
|
||||
* **Do not add a failure counter next to a success counter.** One instrument,
|
||||
with the failure on `error.type` or `ate.failure.reason`; the key's absence
|
||||
means success. See [Reporting failures](#reporting-failures).
|
||||
with the failure on `error.type`; the key's absence means success. See
|
||||
[Reporting failures](#reporting-failures).
|
||||
* **One histogram, many phases**, rather than one histogram per phase, when
|
||||
the phases share dimensions. `ate.actor.restore.duration` carries
|
||||
`ate.snapshot.phase` as a label so the phases sit on one chart.
|
||||
@@ -165,8 +165,7 @@ ones every new metric meets:
|
||||
failed before a worker was picked has no pool; the pool keys are absent, not
|
||||
`""`. `ateattr.WorkerPoolAttributes` returns nil for that case.
|
||||
* **Paired keys travel together.** `ate.workerpool.namespace` with
|
||||
`ate.workerpool.name`; `ate.failure.reason` with `ate.failure.domain`. Use
|
||||
the `ateattr` helpers that return the pair.
|
||||
`ate.workerpool.name`. Use the `ateattr` helpers that return the pair.
|
||||
|
||||
Every label key is a constant in `internal/ateattr/ateattr.go`, and every
|
||||
bounded value set is a group of constants there beside it. A metric never
|
||||
@@ -191,18 +190,11 @@ the cluster.
|
||||
## Reporting failures
|
||||
|
||||
One instrument carries success and failure. Success is the absence of the
|
||||
failure key. Which key depends on where the error is classified:
|
||||
|
||||
* **`error.type`** at an RPC boundary, holding the gRPC status code
|
||||
(`status.Code(err).String()`), as `ate.actor.lifecycle.operation.duration`
|
||||
does. Also for a bounded set of protocol statuses, as
|
||||
`ate.imagecache.requests` does with an allow-list of HTTP codes and `_OTHER`
|
||||
for everything else.
|
||||
* **`ate.failure.reason` and `ate.failure.domain`** together, when the failure
|
||||
is one of substrate's own reasons (`internal/ateerrors`). Use
|
||||
`ateattr.FailureAttributes(reason)`; never set the two keys by hand. This is
|
||||
the right choice inside atelet and ateom handlers, where the gRPC status is
|
||||
assigned only after the handler returns and would read `Unknown`.
|
||||
failure key, `error.type`: at an RPC boundary it holds the gRPC status code
|
||||
(`status.Code(err).String()`), as `ate.actor.lifecycle.operation.duration`
|
||||
does. It also fits a bounded set of protocol statuses, as
|
||||
`ate.imagecache.requests` does with an allow-list of HTTP codes and `_OTHER`
|
||||
for everything else.
|
||||
|
||||
Separate a caller that gave up from a failure: `context.Canceled` and
|
||||
`context.DeadlineExceeded` are their own outcomes (`cancelled`, `timeout`) on
|
||||
@@ -451,7 +443,7 @@ author can find the saturation instruments without reading each brief), and
|
||||
`buckets` for a histogram; a reader of the registry should not need the code.
|
||||
|
||||
Attributes that already exist (`ate.template.name`, `ate.snapshot.kind`,
|
||||
`error.type`, `ate.failure.reason`) are referenced with `ref:`, not redefined.
|
||||
`error.type`) are referenced with `ref:`, not redefined.
|
||||
Mark a label `conditionally_required` and say when, rather than `required`, if
|
||||
the code omits it on some path.
|
||||
|
||||
|
||||
@@ -81,7 +81,7 @@ groups:
|
||||
substrate:
|
||||
emitted_by: [ateapi]
|
||||
code_anchor: cmd/ateapi/internal/controlapi/crash.go
|
||||
cuj: Which actor was lost, and why?
|
||||
cuj: Which actor was lost?
|
||||
attributes:
|
||||
- ref: ate.atespace
|
||||
requirement_level: required
|
||||
@@ -97,7 +97,3 @@ groups:
|
||||
requirement_level: required
|
||||
- ref: ate.actor.state
|
||||
requirement_level: required
|
||||
- ref: ate.failure.reason
|
||||
requirement_level: required
|
||||
- ref: ate.failure.domain
|
||||
requirement_level: required
|
||||
|
||||
@@ -106,7 +106,7 @@ groups:
|
||||
- id: crashed
|
||||
stability: development
|
||||
value: crashed
|
||||
brief: The actor is lost. ate.failure.reason says why.
|
||||
brief: The actor is lost.
|
||||
- id: deleting
|
||||
stability: development
|
||||
value: deleting
|
||||
@@ -345,112 +345,6 @@ groups:
|
||||
value: total
|
||||
brief: The full operation. Use this value as the denominator. Do not add the phases together.
|
||||
|
||||
- id: registry.ate.failure
|
||||
type: attribute_group
|
||||
brief: >
|
||||
The failure reasons of Substrate. These instruments use this key and not
|
||||
error.type. The handler returns a domain error. The interceptor sets the
|
||||
gRPC status only after the handler returns. Thus error.type shows Unknown
|
||||
for almost each real failure.
|
||||
attributes:
|
||||
- id: ate.failure.reason
|
||||
stability: development
|
||||
brief: The cause of the failure.
|
||||
type:
|
||||
members:
|
||||
- id: terminal_file_system_error
|
||||
stability: development
|
||||
value: TERMINAL_FILE_SYSTEM_ERROR
|
||||
brief: A permanent file system fault on the node. Usually the disks of the node are full.
|
||||
- id: invalid_sandbox_asset
|
||||
stability: development
|
||||
value: INVALID_SANDBOX_ASSET
|
||||
brief: A sandbox asset is absent or bad.
|
||||
- id: invalid_checkpoint_result
|
||||
stability: development
|
||||
value: INVALID_CHECKPOINT_RESULT
|
||||
brief: ateom returned a checkpoint that is not valid.
|
||||
- id: failed_save_snapshot
|
||||
stability: development
|
||||
value: FAILED_SAVE_SNAPSHOT
|
||||
brief: atelet could not write the snapshot.
|
||||
- id: invalid_object_url
|
||||
stability: development
|
||||
value: INVALID_OBJECT_URL
|
||||
brief: The URL of the snapshot object is bad.
|
||||
- id: failed_get_external_object
|
||||
stability: development
|
||||
value: FAILED_GET_EXTERNAL_OBJECT
|
||||
brief: atelet could not read the snapshot from object storage. Usually the storage backend has a fault.
|
||||
- id: invalid_container_config
|
||||
stability: development
|
||||
value: INVALID_CONTAINER_CONFIG
|
||||
brief: The container configuration of the ActorTemplate is not valid.
|
||||
- id: local_snapshot_gone
|
||||
stability: development
|
||||
value: LOCAL_SNAPSHOT_GONE
|
||||
brief: The local snapshot on the node no longer exists.
|
||||
- id: workload_not_ready
|
||||
stability: development
|
||||
value: WORKLOAD_NOT_READY
|
||||
brief: >
|
||||
A container started but never passed its wakeup probe before the
|
||||
deadline of the probe. A slow node confounds this value: a cold
|
||||
image or a throttled CPU also runs the probe out of time. The
|
||||
phases of ate.actor.restore.duration separate the two.
|
||||
- id: corrupted_assignment
|
||||
stability: development
|
||||
value: CORRUPTED_ASSIGNMENT
|
||||
brief: Control plane. The worker assignment of the actor is not correct.
|
||||
- id: worker_reassigned
|
||||
stability: development
|
||||
value: WORKER_REASSIGNED
|
||||
brief: Control plane. A different actor got the worker.
|
||||
- id: worker_pod_gone
|
||||
stability: development
|
||||
value: WORKER_POD_GONE
|
||||
brief: Control plane. The worker pod no longer exists.
|
||||
- id: unknown
|
||||
stability: development
|
||||
value: UNKNOWN
|
||||
brief: >
|
||||
A failure with no reason. Substrate does not make a value from
|
||||
the error message. Thus the list of values stays short. This
|
||||
value asserts no fault domain: read ate.failure.domain, which
|
||||
reports unknown here and not infrastructure.
|
||||
|
||||
- id: ate.failure.domain
|
||||
stability: development
|
||||
brief: >
|
||||
The side of the platform boundary the failure came from. A strict
|
||||
function of ate.failure.reason, thus it adds no series. It exists
|
||||
because an operator triages on the boundary and drills into the
|
||||
reason, and because a consumer must not have to match on the names of
|
||||
the reasons: a component ahead of the control plane can report a
|
||||
reason that the control plane does not know, which becomes UNKNOWN,
|
||||
and a name match would then count it as an infrastructure fault.
|
||||
type:
|
||||
members:
|
||||
- id: infrastructure
|
||||
stability: development
|
||||
value: infrastructure
|
||||
brief: Substrate failed. The node, the store, the object storage, the control plane.
|
||||
- id: workload
|
||||
stability: development
|
||||
value: workload
|
||||
brief: >
|
||||
The owner of the actor fixes this, not the operator of the
|
||||
platform. It covers a bad ActorTemplate as well as a process
|
||||
that does not start. For the reasons that the actor reports
|
||||
about itself, such as an exit or a probe that never passes,
|
||||
this value is a report and not a proof: the actor chooses its
|
||||
own exit. The reasons that Substrate measures, such as a
|
||||
template that resolves to no runnable process, are not.
|
||||
- id: unknown
|
||||
stability: development
|
||||
value: unknown
|
||||
brief: The reason gives no domain. A gap in the taxonomy stays visible here and does not inflate one side.
|
||||
|
||||
- id: registry.ate.scheduler
|
||||
type: attribute_group
|
||||
brief: The labels of a scheduler decision.
|
||||
@@ -707,14 +601,14 @@ groups:
|
||||
instrument: counter
|
||||
unit: "{crash}"
|
||||
stability: development
|
||||
brief: The number of actors that went to the ACTOR_STATE_CRASHED state, with the reasons.
|
||||
brief: The number of actors that went to the ACTOR_STATE_CRASHED state.
|
||||
note: >
|
||||
This counter counts the moves into the CRASHED state of the state machine.
|
||||
It does not count the crashes of a process. The CRASHED state is a loss of
|
||||
the actor's live execution (recoverable back to SUSPENDED via RevertActor).
|
||||
Substrate can also lose the data that it did not write. Some moves into
|
||||
this state are not a loss of data. They are careful responses to a control
|
||||
plane problem. The ate.failure.reason key keeps these groups separate.
|
||||
plane problem.
|
||||
ate.sandbox.class is unknown when ateapi could not read the class of the
|
||||
worker: the worker record is already gone, or its assignment is already
|
||||
clear.
|
||||
@@ -723,14 +617,10 @@ groups:
|
||||
emitted_by: [ateapi]
|
||||
golden_signals: [errors]
|
||||
code_anchor: cmd/ateapi/internal/controlapi/metrics.go
|
||||
cuj: Some actors go to the CRASHED state. How many, and why?
|
||||
cuj: Some actors go to the CRASHED state. How many?
|
||||
attributes:
|
||||
- ref: ate.actor.operation.name
|
||||
requirement_level: required
|
||||
- ref: ate.failure.reason
|
||||
requirement_level: required
|
||||
- ref: ate.failure.domain
|
||||
requirement_level: required
|
||||
- ref: ate.template.atespace
|
||||
requirement_level: required
|
||||
- ref: ate.template.name
|
||||
@@ -875,10 +765,7 @@ groups:
|
||||
note: >
|
||||
This instrument shows where the cold start time goes after ateapi hands
|
||||
the work to atelet. The phases occur at the same time. They do not add up
|
||||
to the total. Use the total phase as the denominator. If the operation
|
||||
fails, atelet puts ate.failure.reason on the phase that failed and on the
|
||||
total. It puts the key on no other phase. Thus you can still query the
|
||||
phases that were correct.
|
||||
to the total. Use the total phase as the denominator.
|
||||
annotations:
|
||||
substrate:
|
||||
emitted_by: [atelet]
|
||||
@@ -901,12 +788,6 @@ groups:
|
||||
- ref: ate.sandbox.class
|
||||
requirement_level:
|
||||
conditionally_required: The class is known.
|
||||
- ref: ate.failure.reason
|
||||
requirement_level:
|
||||
conditionally_required: The operation failed. The key occurs only on the phase that failed and on the total.
|
||||
- ref: ate.failure.domain
|
||||
requirement_level:
|
||||
conditionally_required: ate.failure.reason occurs. The two keys always occur together.
|
||||
|
||||
- id: metric.ate.actor.checkpoint.duration
|
||||
type: metric
|
||||
@@ -941,12 +822,6 @@ groups:
|
||||
- ref: ate.sandbox.class
|
||||
requirement_level:
|
||||
conditionally_required: The class is known.
|
||||
- ref: ate.failure.reason
|
||||
requirement_level:
|
||||
conditionally_required: The operation failed. The key occurs only on the phase that failed and on the total.
|
||||
- ref: ate.failure.domain
|
||||
requirement_level:
|
||||
conditionally_required: ate.failure.reason occurs. The two keys always occur together.
|
||||
|
||||
- id: metric.atelet.snapshot.size
|
||||
type: metric
|
||||
|
||||
@@ -162,42 +162,13 @@ cardinality_rules:
|
||||
reason. Only `weaver registry live-check` can count the values that occur.
|
||||
- id: error-type-not-parallel-counters
|
||||
brief: >
|
||||
A component reports a failure with error.type or ate.failure.reason on the
|
||||
same instrument. If the key is absent, the operation was correct. There
|
||||
are no separate counters for the failures.
|
||||
A component reports a failure with error.type on the same instrument. If
|
||||
the key is absent, the operation was correct. There are no separate
|
||||
counters for the failures.
|
||||
enforced: false
|
||||
could_be_enforced_by: >
|
||||
A Weaver Rego policy can refuse a metric name that ends with errors_total
|
||||
and has a counterpart without the suffix.
|
||||
- id: failure-keys-paired
|
||||
brief: >
|
||||
A component sets ate.failure.reason and ate.failure.domain together, or it
|
||||
sets neither. ateattr.FailureAttributes and ateattr.FailureLogAttrs return
|
||||
the pair, and a producer uses one of them rather than the keys.
|
||||
The domain is a strict function of the reason, so it adds no series. It is
|
||||
a key and not a rule in the consumer because the reasons of a component
|
||||
ahead of the control plane can be unknown to the control plane, which
|
||||
makes them UNKNOWN, and a consumer that matched on the names of the
|
||||
reasons would count each of those as an infrastructure fault.
|
||||
enforced: false
|
||||
could_be_enforced_by: >
|
||||
A Weaver Rego policy can make the two keys occur together in this
|
||||
registry. Only live-check can see a component that sends one key.
|
||||
- id: workload-domain-is-a-report
|
||||
brief: >
|
||||
Some of the reasons in the workload domain are what the actor reported
|
||||
about itself: an exit code, or a probe that never passes. The actor
|
||||
chooses those, thus it decides the value. It can raise the count of the
|
||||
failures of its own template and pool, and it can exit 0 to hide a fault.
|
||||
A dashboard says "reported" for those, and an alert on them has a rate
|
||||
limit; one report for each activation is the bound. The reasons that
|
||||
Substrate measures, such as a template that resolves to no runnable
|
||||
process, carry no such caveat.
|
||||
enforced: false
|
||||
could_be_enforced_by: >
|
||||
Nothing. This is a property of a multi-tenant platform and not of the
|
||||
registry. It is here so that a reader of the panel knows what the number
|
||||
is.
|
||||
- id: pool-keys-paired
|
||||
brief: >
|
||||
A component sets ate.workerpool.namespace and ate.workerpool.name
|
||||
|
||||
+6
-17
@@ -137,7 +137,7 @@ atelet's `Restore timing breakdown` is the first of these, and the only unsample
|
||||
"trace_id":"4bf92f…","span_id":"00f067…","trace_flags":"01"}
|
||||
```
|
||||
|
||||
The duration keys are the [`ate.actor.restore.duration`](#the-metric-registry) instrument's name with an `ate.snapshot.phase` value appended, and they hold **seconds**, matching that instrument's declared unit. The same rules apply as on the histogram: a phase that never ran is absent rather than zero, phases overlap and do not sum to the total, and `ate.failure.reason` is present only when the restore failed. There is no `ate.snapshot.phase` key on the record — on a datapoint it names the one step timed, and this record carries them all.
|
||||
The duration keys are the [`ate.actor.restore.duration`](#the-metric-registry) instrument's name with an `ate.snapshot.phase` value appended, and they hold **seconds**, matching that instrument's declared unit. The same rules apply as on the histogram: a phase that never ran is absent rather than zero, and phases overlap and do not sum to the total. There is no `ate.snapshot.phase` key on the record — on a datapoint it names the one step timed, and this record carries them all.
|
||||
|
||||
This is the record to use for a per-actor wake-up distribution. The histogram cannot answer that question at all, because actor identity is barred from metric labels; traces can, but the data plane is head-sampled at 1%.
|
||||
|
||||
@@ -170,11 +170,10 @@ Creating an actor counts as a change. A new actor is born suspended, so it gets
|
||||
"ate.atespace":"ate-demo-counter","ate.actor.name":"counter-1","ate.actor.uid":"8f2a…",
|
||||
"ate.template.atespace":"ate-demo-counter","ate.template.name":"counter",
|
||||
"ate.actor.operation.name":"resume","ate.actor.state":"crashed",
|
||||
"ate.failure.reason":"WORKER_POD_GONE","ate.failure.domain":"infrastructure",
|
||||
"trace_id":"4bf92f…","span_id":"00f067…","trace_flags":"01"}
|
||||
```
|
||||
|
||||
The counter carries the same reason but no actor identity, so this record is the only way to attribute a crash to one agent. The decision-point line that precedes it (`Setting Actor to crashed due to error`) carries only `ate.atespace` and `ate.actor.name`: it is written before the Actor is loaded, so no uid exists yet.
|
||||
The counter carries no actor identity, so this record is the only way to attribute a crash to one agent. The decision-point line that precedes it (`Setting Actor to crashed due to error`) carries only `ate.atespace` and `ate.actor.name`: it is written before the Actor is loaded, so no uid exists yet.
|
||||
|
||||
#### The same records over OTLP
|
||||
|
||||
@@ -185,9 +184,9 @@ Two `event.name` values, which is the OTLP LogRecord's own field rather than an
|
||||
| `event.name` | Body | Severity | Attributes |
|
||||
|---|---|---|---|
|
||||
| `ate.actor.state_changed` | `Actor state changed` | 9 | the five identity keys, `ate.actor.operation.name`, `ate.actor.state` |
|
||||
| `ate.actor.crashed` | `Actor crashed` | 17 | the above plus `ate.failure.reason`, `ate.failure.domain` |
|
||||
| `ate.actor.crashed` | `Actor crashed` | 17 | the same keys |
|
||||
|
||||
A crash is its own name because an event name promises a fixed set of attributes, and it carries two more. There is no name per state: `ate.actor.state` already says which transition happened, so a consumer still selects on that one attribute and needs no map from a name to a state. Both names are in [`docs/metrics/registry/events.yaml`](metrics/registry/events.yaml), which `make verify` checks.
|
||||
A crash is its own name because an event name promises a fixed set of attributes and a crash has a different severity and shape. There is no name per state: `ate.actor.state` already says which transition happened, so a consumer still selects on that one attribute and needs no map from a name to a state. Both names are in [`docs/metrics/registry/events.yaml`](metrics/registry/events.yaml), which `make verify` checks.
|
||||
|
||||
The attributes are the same flat `ate.*` keys as the stdout copy, so they arrive as real log attributes with no transform in front of them. Trace context is not among them: it goes on the record's own `TraceId` and `SpanId` fields, where the stdout copy's top-level `trace_id`/`span_id` would be mapped to anyway. The instrumentation scope is `github.com/agent-substrate/substrate/internal/actorevent`, which is how you select this stream, or exclude it.
|
||||
|
||||
@@ -225,14 +224,6 @@ labels."ate.actor.uid"="8f2a…" AND jsonPayload.msg="Actor usage sample"
|
||||
|
||||
**Do not put actor identity on a log-based metric.** Aggregating these events in the log store is what they are for, but a log-based metric built over them must label only by the bounded set (template, sandbox class, source, pool) — promoting `ate.actor.uid` or `ate.actor.name` into a metric label reintroduces exactly the per-actor cardinality this split keeps out of the TSDB.
|
||||
|
||||
### Which side failed: `ate.failure.domain`
|
||||
|
||||
`ate.failure.reason` names the cause; `ate.failure.domain` names the side of the platform boundary it came from, as `infrastructure`, `workload`, or `unknown`. It rides on every signal that carries a reason, and the two are always emitted together — producers call `ateattr.FailureAttributes` or `ateattr.FailureLogAttrs` rather than setting either key directly.
|
||||
|
||||
The domain is a strict function of the reason, so it costs no series and no consumer needs it to disambiguate a value. It exists because the alternative is every consumer keeping its own map from reason to domain, and that map breaks silently: a component running ahead of the control plane can report a reason this build's `ateerrors.AllReasons` rejects, `ExtractReason` turns it into `UNKNOWN`, and a name-matching consumer would file every one of those as an infrastructure fault. `UNKNOWN` therefore reports `unknown` and not `infrastructure` — a gap in the taxonomy stays visible instead of inflating one side.
|
||||
|
||||
One caveat belongs on any panel built from this. A `workload` domain says what the actor **reported**, not what substrate measured: the actor picks its own exit and can hold memory until the kernel kills it, so it can raise the failure count for its own template and pool, or exit cleanly to hide a fault. See `workload-domain-is-a-report` in [`docs/metrics/substrate.yaml`](metrics/substrate.yaml).
|
||||
|
||||
---
|
||||
|
||||
## 2. Metrics
|
||||
@@ -244,7 +235,7 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
|
||||
| Metric | Emitted by | Type | Measures |
|
||||
|--------|------------|------|----------|
|
||||
| `rpc.server.call.duration` | ateapi & atelet (gRPC servers, via `otelgrpc`) | histogram | per-method gRPC latency, request rate, and errors (labels `rpc.method`, `rpc.response.status_code`) |
|
||||
| `ate.actor.crashes` | ateapi | counter | Number of times actors transitioned to `ACTOR_STATE_CRASHED` with failure reasons (labels `ate.actor.operation.name`, `ate.failure.reason`, `ate.failure.domain`, `ate.template.atespace`, `ate.template.name`, `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`) |
|
||||
| `ate.actor.crashes` | ateapi | counter | Number of times actors transitioned to `ACTOR_STATE_CRASHED` (labels `ate.actor.operation.name`, `ate.template.atespace`, `ate.template.name`, `ate.workerpool.namespace`, `ate.workerpool.name`, `ate.sandbox.class`) |
|
||||
| `atenet.router.route.duration` | atenet-router | histogram | Substrate E2E — Envoy receiving a request to Envoy forwarding it to the resolved worker, excluding actor compute and the response (labels `ate.template.atespace`, `ate.template.name`, `ate.router.outcome`, `ate.router.resume`) |
|
||||
| `atelet.snapshot.size` | atelet | histogram | uncompressed size in bytes of each gVisor snapshot image written during checkpoint (labels `file.name`, `ate.template.atespace`, `ate.template.name`) |
|
||||
| `ate.workerpool.desired_workers` | atecontroller | up/down counter | number of worker pods requested for a WorkerPool, from `spec.replicas` (labels
|
||||
@@ -254,7 +245,7 @@ Agent Substrate emits foundational OpenTelemetry system and server metrics to mo
|
||||
| `ate.workerpool.workers` | ateapi | up/down counter | live worker count per pool, split by state (`idle`/`assigned`) and sandbox class to provide fleet capacity and saturation at a glance |
|
||||
| `ate.actor.lifecycle.operation.duration` | ateapi | histogram | how long each actor operation (create/resume/suspend/pause/delete/revert) takes and whether it failed (`error.type` present = failure, absent = success); labeled by operation, template, pool (`ate.workerpool.namespace` + `ate.workerpool.name`), sandbox class, and snapshot kind and scope on resume; already-running resume no-ops are not recorded so the histogram tracks actual activations, not router traffic |
|
||||
| `ate.scheduler.assignment.duration` | ateapi | histogram | time it takes for an actor to be assigned to a worker, per attempt (version-conflict retries record only the final attempt), with the outcome (`assigned` / `no_free_worker` / `error`), the assigned pool (`ate.workerpool.namespace` + `ate.workerpool.name`) and sandbox class to catch scheduling latency and capacity starvation problems |
|
||||
| `ate.actor.restore.duration` | atelet | histogram | how long each phase of a restore takes on the worker node, which is where cold-start latency actually goes once ateapi hands off (labels `ate.snapshot.phase`, `ate.snapshot.kind`, `ate.snapshot.scope`, `ate.template.atespace`, `ate.template.name`, `ate.sandbox.class`, plus `ate.failure.reason` and `ate.failure.domain` on failure) |
|
||||
| `ate.actor.restore.duration` | atelet | histogram | how long each phase of a restore takes on the worker node, which is where cold-start latency actually goes once ateapi hands off (labels `ate.snapshot.phase`, `ate.snapshot.kind`, `ate.snapshot.scope`, `ate.template.atespace`, `ate.template.name`, `ate.sandbox.class`) |
|
||||
| `ate.actor.checkpoint.duration` | atelet | histogram | the same phase breakdown for writing a snapshot, so a slow suspend can be attributed to ateom or to the upload (same labels as the restore histogram) |
|
||||
| `ate.imagecache.requests` | atelet | counter | image lookups in the node-local image cache, by outcome (`ate.imagecache.outcome`), with `error.type` on the `error` outcome. A miss pays for the pull and the unpack, so the hit ratio per node is a leading indicator of resume latency |
|
||||
|
||||
@@ -283,8 +274,6 @@ The three snapshot labels are orthogonal and mean the same thing on every histog
|
||||
|
||||
**Phases overlap and do not sum to `total`.** The download runs concurrently with the asset fetch and OCI unpack, so each is an independent observation; use `total` as the denominator. A phase that never started is absent rather than zero.
|
||||
|
||||
On a failure, `ate.failure.reason` marks the phase that died and the `total`, and nothing else, so `ate.actor.restore.duration{ate.snapshot.phase="download", ate.failure.reason!=""}` says how often the download is what breaks and why, while the phases that succeeded stay queryable as successes. The atelet histograms classify with substrate's own reason taxonomy (the same one `ate.actor.crashes` uses) rather than `error.type`, because these handlers return wrapped domain errors and the gRPC status is only assigned after the handler returns, so a status code would read `Unknown` for nearly every real failure. A failure that carries no reason reports `UNKNOWN`. [`ate.failure.domain`](#which-side-failed-atefailuredomain) rides alongside wherever the reason is present, so a restore that failed because the actor never passed its wakeup probe is separable from one the node broke.
|
||||
|
||||
The `ate.*` control-plane metric labels are either fixed value sets (operation, outcome, state, class, kind, scope, phase) or scoped to the deployment catalog (template and pool names are operator-created, never derived from request payloads), and the label set varies per operation: resume carries the most dimensions, delete only the operation and error type. `ate.sandbox.class` is derived from the template (each template has exactly one class), so it adds no extra series next to the template labels; it exists so dashboards can aggregate by class without enumerating template names. High-cardinality actor identity (name/uid/atespace) stays off metrics entirely and lives on logs and traces instead — for resource usage, on the [per-actor usage events](#per-actor-usage-events).
|
||||
|
||||
### The metric registry
|
||||
|
||||
Reference in New Issue
Block a user