mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Part of #1590 ## What this PR does Locust measures the client side only. This adds a post-run harvest of server-side ground truth from Prometheus, written to a new `server_summary.json` and summarized as one row in `stats.jsonl`. ## Proposed Changes ### Steady-state scoping `server_telemetry.py` derives the steady-state window from `stats_history.csv`: it starts at the first sample at 90% of peak user count and ends at the last one. Ramp-up is excluded from every metric, and the snapshot block also stops at the last full-load sample, so teardown suspends are left out. Under a step-ladder load shape the window covers only the top step. ### Harvested metrics **Cluster packing**, from `ate_workerpool_workers`: busy workers (`partial` + `at_capacity`) over the pool size at each sample, where the pool size is the sum of all worker states, so it follows scale-ups and scale-downs mid-run. Only the live ateapi is read: after a redeploy the collector keeps re-exporting exited ateapi processes' last values for a few minutes, which would otherwise inflate the counts. Reported as min, p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries so transient spikes remain visible. **Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall percentages for the node and for the worker pods, each as the same percentile set over the window. **Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both in-window and cumulative, restore and checkpoint latency mean and p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint throughput as `checkpoint_mb_s`. The atelet exports metrics on an interval, so its data reaches Prometheus late: the harvest waits `--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default export and a 10s scrape) and reads both window edges half that late. ### Constraints and failure behavior The module uses only the standard library, since the locust image is distroless, and every request carries a timeout. An unreachable Prometheus records nulls and does not fail the run. `--prometheus-url` overrides the in-cluster default, and `--atelet-lag-s` sets the wait for the atelet's last export. Unmeasured fields are `null` and a measured zero is `0`, consistent with the rest of the runner. Prometheus exposes no byte counter on the restore path, so no restore throughput field is emitted rather than deriving one indirectly. ### Output `server_summary.json` holds the full nested artifact. `stats.jsonl` receives a single `server_summary` row with 45 flat keys for graphing: packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`. `status.json` is unchanged. ## How this was tested 14 unit tests in `test_server_telemetry.py` covering steady-state detection, percentile boundaries, the range-query window guard, malformed Prometheus responses, the packing and checkpoint arithmetic, the per-sample pool size, ignoring exited ateapi series, a missing denominator returning null, snapshot fields returning null rather than zero, and telemetry surviving a missing stats CSV. Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45 keys matched an independent recomputation from raw Prometheus, and the snapshot block was cross-checked against atelet logs. A run with `--prometheus-url` pointed at an unreachable address completes normally with the affected fields null. Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8 workers) from Prometheus with the new code. Runs with a steady pool match the previous output exactly, and runs that followed a redeploy now read the real pool size on every sample. ## References [Agent Substrate: Actor Density Benchmark Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0) - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR