benchmark(atelet, ateom-microvm): log per-actor checkpoint and restor… (#1941)

## Benchamark(atelet, ateom-microvm): log per-actor checkpoint and
restore phase breakdowns

### Why

`SuspendActor` and `ResumeActor` latency is only attributable down to
the `ate.actor.{checkpoint,restore}.duration` histogram phases, and the
biggest buckets — `ateom_checkpoint`, `ateom_restore` — are opaque. When
a large memory benchmark suspend takes 6 s we want to able to tell where
the time is been spent during the suspend.

### What

One joinable, developer-facing log record per operation per layer, with
the full actor identity (allowed in logs, barred from metric labels),
the snapshot scope, and one float-seconds field per phase that ran.

**atelet**
- `Checkpoint` now writes a `Checkpoint timing breakdown` record, the
same way `Restore` has written `Restore timing breakdown` since #1364.
Phases: `sandbox_assets`, `ateom_checkpoint`, `persist`, `total`. One
slice feeds both the histogram and the record, so they cannot disagree.
- Both records carry `error.type` when the operation failed (the gRPC
code; context errors map to `DeadlineExceeded` / `Canceled`). The record
is written on the way out of a failure too, so its completed phases are
kept, and the marker lets a reader exclude a timed-out restore from a
latency distribution. This restores what the record lost when the
`ateerrors` taxonomy was deleted (#1817), using the `error.type`
convention the ateapi instruments already follow.

**ateom-microvm**
- New `phaselog.go`. `CheckpointWorkload` and `RestoreWorkload` emit
records with the same two messages, under
`ateom.actor.checkpoint.duration.<phase>` and
`ateom.actor.restore.duration.<phase>`, decomposing atelet's
`ateom_checkpoint` / `ateom_restore` buckets:
- checkpoint: `pause`, `snapshot`, `durable_dir`, `rootfs_upper`,
`teardown`, `total` — the three captures run concurrently on the paused
guest, so the paused window costs their max, not their sum.
- restore: `prep`, `bundles`, `upper_join`, `lowers`, `tap`,
`vmm_launch`, `vm_restore`, `resume`, `wakeup_probe`, `total` —
sequential; they partition the total. A Data-scope cold boot records
`total` only.
- These timings already existed as ad-hoc `slog.Duration` fields on the
`Actor checkpointed` / `Actor restore phases` lines; those lines are
kept. The record adds stable keys, identity, and seconds (the
histograms' unit).
- The phase names are deliberately private to the binary rather than
added to `internal/ateattr`, so they cannot be mistaken for
`ate.snapshot.phase` metric values. They are micro-VM specific;
ateom-gvisor is unchanged.

**docs/observability.md** is updated: the Restore record is no longer
the only per-actor latency record, and the ateom records are described.

No new instruments, no registry changes, no behavior change.

### Example

```json
{"msg":"Checkpoint timing breakdown","ate.actor.uid":"8f2a…","ate.template.name":"glutton",
 "ate.snapshot.scope":"full",
 "ateom.actor.checkpoint.duration.pause":0.003,
 "ateom.actor.checkpoint.duration.snapshot":0.846,
 "ateom.actor.checkpoint.duration.rootfs_upper":0.022,
 "ateom.actor.checkpoint.duration.teardown":0.232,
 "ateom.actor.checkpoint.duration.total":1.081}
```

Joined with atelet's record for the same actor, a run of the glutton
workload (1 GiB resident, microVM) attributes a 6.0 s p50 suspend as 75%
`persist`, 19% `ateom_checkpoint` (of which the CH `snapshot` is 0.85 s
and `teardown` 0.23 s), and a 5.3 s p50 resume as 79% `download`, 18%
`ateom_restore` (of which `vm_restore` is 0.73 s). The consumer that
produces those tables from pod logs is a separate
`benchmarking/analysis` PR.

### Testing

- `go test ./cmd/atelet/...` and `./cmd/ateom-microvm/...` pass; new
unit tests cover the record shape (seconds, identity keys, zero phases
absent, no duplicate keys), the scope mapping, and `error.type` for
gRPC, context and plain errors.
- `GOOS=linux go vet` clean for both binaries; boilerplate and gofmt
clean.
This commit is contained in:
Lucky Abolorunke
2026-09-30 10:44:20 +00:00
committed by GitHub
parent 1282773d3a
commit 9c1f0ba042
8 changed files with 474 additions and 33 deletions
+3 -1
View File
@@ -123,7 +123,7 @@ An actor's **own** lines carry trace context only if the actor emits these field
A component's own `slog` output can also be about a specific actor. Those records take the identity keys from [`internal/ateattr`](../internal/ateattr) too, flat at the top level rather than inside a label group: a component writes no envelope, so a collector lifts the keys straight onto the log record's attributes. `ateattr.ActorLogAttrs` and `ateattr.ActorLogLabels` return the same five keys for this reason, and a test holds them together. Filtering on `ate.actor.uid` therefore finds a component record and an actor's own output alike.
atelet's `Restore timing breakdown` is the first of these, and the only unsampled per-actor latency record Substrate produces. It is emitted once per restore, whether the restore succeeded or failed:
atelet's `Restore timing breakdown` and `Checkpoint timing breakdown` are the unsampled per-actor latency records Substrate produces. Each is emitted once per operation, whether it succeeded or failed; a failed one also carries `error.type` (the gRPC code, `DeadlineExceeded` and `Canceled` for context errors), so a reader can leave it out of a latency distribution:
```json
{"time":"…","level":"INFO","msg":"Restore timing breakdown",
@@ -141,6 +141,8 @@ The duration keys are the [`ate.actor.restore.duration`](#the-metric-registry) i
This is the record to use for a per-actor wake-up distribution. The histogram cannot answer that question at all, because actor identity is barred from metric labels; traces can, but the data plane is head-sampled at 1%.
ateom-microvm writes records with the same two messages, from inside the `ateom_restore` and `ateom_checkpoint` phases, under `ateom.actor.restore.duration.<phase>` and `ateom.actor.checkpoint.duration.<phase>` (for example `vm_restore`, `wakeup_probe`, `prep`, `pause`, `snapshot`, `teardown`). Those keys are not instruments: the phases are implementation details of one runtime, so they stay a developer-facing record. The same identity keys make the two layers' records joinable per actor. The checkpoint record is written on failure too, with `error.type` and the elapsed time of the step that failed; the restore record is written on success only. The restore phases are sequential and partition the total; a checkpoint's `snapshot`, `durable_dir` and `rootfs_upper` run concurrently on the paused guest, so the paused window costs their maximum, not their sum.
ateapi's `Actor state changed` is written once per committed actor state transition. ateapi owns the state machine, so this is where an actor's state and the time it reached it come from:
```json