Files
Nishanth Kotla 3f789b309a benchmarking/locust: harvest server-side telemetry from Prometheus (#1725)
Part of #1590

## What this PR does

Locust measures the client side only. This adds a post-run harvest of
server-side ground truth from Prometheus, written to a new
`server_summary.json` and summarized as one row in `stats.jsonl`.

## Proposed Changes

### Steady-state scoping

`server_telemetry.py` derives the steady-state window from
`stats_history.csv`: it starts at the first sample at 90% of peak user
count and ends at the last one. Ramp-up is excluded from every metric,
and the snapshot block also stops at the last full-load sample, so
teardown suspends are left out. Under a step-ladder load shape the
window covers only the top step.

### Harvested metrics

**Cluster packing**, from `ate_workerpool_workers`: busy workers
(`partial` + `at_capacity`) over the pool size at each sample, where the
pool size is the sum of all worker states, so it follows scale-ups and
scale-downs mid-run. Only the live ateapi is read: after a redeploy the
collector keeps re-exporting exited ateapi processes' last values for a
few minutes, which would otherwise inflate the counts. Reported as min,
p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries
so transient spikes remain visible.

**Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall
percentages for the node and for the worker pods, each as the same
percentile set over the window.

**Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both
in-window and cumulative, restore and checkpoint latency mean and
p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint
throughput as `checkpoint_mb_s`. The atelet exports metrics on an
interval, so its data reaches Prometheus late: the harvest waits
`--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default
export and a 10s scrape) and reads both window edges half that late.

### Constraints and failure behavior

The module uses only the standard library, since the locust image is
distroless, and every request carries a timeout. An unreachable
Prometheus records nulls and does not fail the run. `--prometheus-url`
overrides the in-cluster default, and `--atelet-lag-s` sets the wait for
the atelet's last export.

Unmeasured fields are `null` and a measured zero is `0`, consistent with
the rest of the runner. Prometheus exposes no byte counter on the
restore path, so no restore throughput field is emitted rather than
deriving one indirectly.

### Output

`server_summary.json` holds the full nested artifact. `stats.jsonl`
receives a single `server_summary` row with 45 flat keys for graphing:
packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot
means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`.

`status.json` is unchanged.

## How this was tested

14 unit tests in `test_server_telemetry.py` covering steady-state
detection, percentile boundaries, the range-query window guard,
malformed Prometheus responses, the packing and checkpoint arithmetic,
the per-sample pool size, ignoring exited ateapi series, a missing
denominator returning null, snapshot fields returning null rather than
zero, and telemetry surviving a missing stats CSV.

Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45
keys matched an independent recomputation from raw Prometheus, and the
snapshot block was cross-checked against atelet logs. A run with
`--prometheus-url` pointed at an unreachable address completes normally
with the affected fields null.

Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8
workers) from Prometheus with the new code. Runs with a steady pool
match the previous output exactly, and runs that followed a redeploy now
read the real pool size on every sample.

## References

[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-10-01 16:59:41 +00:00
..