Part of #1590 First of two PRs. This one covers Kubernetes-side hardware discovery; the follow-up adds the Prometheus harvest. ## What this PR does The locust runner records no information about the hardware it runs on, so trial results cannot be normalized across machine types or node counts. This adds cluster capacity discovery and derives actor density frontiers from it, persisting both the raw readings and the derived ratios to `stats.jsonl`. ## Proposed Changes ### Cluster hardware discovery New module `benchmarking/locust/cluster_facts.py`, kept separate from `runner.py`. It reads node count, machine type, allocatable cores and allocatable RAM, and the worker pod count via the official Kubernetes Python client, which resolves in-cluster and local kubeconfig auth without a kubectl subprocess. Node and pod listing is proportional to cluster size, so `--no-cluster-facts` disables discovery. Discovery is best effort in either case: an unreachable API server or missing RBAC leaves the facts null and does not fail the run. No value is defaulted or inferred. Unmeasured fields are `null`; a measured zero is `0`. The two remain distinguishable to any consumer. ### Density frontiers Added to the `trial_summary` row in `stats.jsonl`: | Field | Meaning | |---|---| | `actors_per_node` | peak actors over node count | | `actors_per_vcpu` | peak actors over allocatable cores | | `actors_per_gb_ram` | peak actors over allocatable RAM | | `actors_per_pod_p50` / `_p90` / `_p99` | distribution of actors per worker pod | | `aggregate_failure_ratio` | failures over requests, all RPCs | | `<operation>_failure_ratio` | one per operation Locust reported, e.g. `dur_dir_write_failure_ratio` | Actors per pod is reported as a distribution across the run's time samples (`p50` / `p90` / `p99`) rather than one average, because ramp-up and custom load shapes have no single user count to call steady. The raw readings (`machine_type`, `node_count`, `allocatable_cores`, `allocatable_ram_gb`, `worker_pod_count`) are written alongside the derived ratios so they can be re-derived without re-running the trial. ### RBAC Nodes are cluster scoped and require a ClusterRole with `list` on `nodes`. Pods require only a namespaced Role with `list` in `benchmark-workloads`. `locust.yaml` and `runner-job.yaml.tmpl` each define their own separately named pair. Note that `locust.yaml` now binds a Role in `benchmark-workloads`, so that namespace must exist before the manifest is applied. `benchmarking/workloads/deploy.sh` creates it. `status.json` is unchanged. ## How this was tested 11 unit tests in `test_cluster_facts.py` covering percentile boundaries, capacity scoped to only the nodes running worker pods, zero worker pods treated as a valid reading, unreadable facts returning null, and flag behavior. Verified against a live GKE cluster. Discovery returned `c3d-standard-8`, 1 node, 7.91 allocatable cores, 27.73 GB, 5 worker pods, all matching `kubectl`. The emitted frontiers were re-derived by hand from `stats_history.csv` and matched. `--no-cluster-facts` returns all nulls. Both manifests pass `kubectl apply --dry-run=client`, with no cluster-scoped name collisions between them. ## References [Agent Substrate: Actor Density Benchmark Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0) - [x] Tests pass - [x] Appropriate changes to documentation are included in the PR
Substrate Benchmarking
This is the nascent suite for benchmarking Substrate's performance at scale.
The suite also measures the telemetry volume and the capacity of the OTel collector: how much trace data and metric data substrate and its actors send, and if the collector can accept it. To make a measurement, read telemetry/README.md. For the prerequisites and the scenario ladder, read observability.md.
Deploy benchmarks
Important
Source the environment configuration file (e.g.,
source .ate-dev-env.sh) first soPROJECT_ID,BUCKET_NAME, etc. are set.
Note that deploying the benchmarks does not run them. You must visit Locust's web UI to start a test.
A single wrapper deploys the scale workloads, builds and pushes the Locust image, then deploys the Locust workers:
./benchmarking/deploy_locust.sh --deploy
Useful flags:
--worker-count N— number ofWorkerPoolreplicas (default 1).--skip-build— reuse the existing:latestlocust image (skip thedocker build && docker pushstep).
To tear everything down (locust then workloads, in reverse order):
./benchmarking/deploy_locust.sh --delete
The same operations are also reachable from the top-level installer for convenience:
./hack/install-ate.sh --deploy-benchmarks
./hack/install-ate.sh --delete-benchmarks
The installer accepts --benchmark-worker-count N (default 1).
--skip-build is only available when invoking
benchmarking/deploy_locust.sh directly.
Running Tests
Locust Web UI
- Run
kubectl port-forward svc/locust -n benchmarking 8089:8089 - Visit
http://localhost:8089in your browser to configure and start the load test.
The different user classes you can select are different types of load behaviors you can throw at the system. Note that the "CounterUser" load type requires that the counter demo be installed.
You can also configure things like the number of users, how quickly those users are spawned, the frequency with which requests are made and whether or not tracing is enabled.
User classes implemented in boomer rather than Python are selected at deploy time — the stack runs one per deployment:
./benchmarking/locust/deploy.sh --deploy --user-class durdir
Headless (automation only)
runner.py runs a test without the web UI, writing CSVs, logs and traces to
--dest. The nightly automation submits it as a Job on the test cluster; it is
not a local entry point. See automation/README.md.
python3 runner.py -f tests/<user-class>.py -t 1m -u 1 --name <run-name> --dest /tmp/bench
One flag controls the optional post-run measurements described in Benchmark output files:
--cluster-facts/--no-cluster-facts: read node capacity and worker pod count from the Kubernetes API once the run ends, to derive density frontiers. On by default. Pass--no-cluster-factsto skip Kubernetes API discovery.
Test-specific flags are appended to the same command; see the sections below.
DurDir Benchmark
The DurDir benchmark evaluates actor suspend/resume performance, disk persistence overhead, and state restoration latency when a durable directory is attached to the actor.
DurDir Configuration Knobs
--durdir-file-size-bytes: Size in bytes of the data file (default8388608= 8 MiB).--resume-mode: Resume trigger mode:explicit(default): Client invokes theResumeActorRPC before sending traffic.implicit: Client sends traffic through the router without an explicit wake RPC, testing traffic-triggered resume.
--durdir-read-mode: Verification read mode:data(default): Server returns full payload bytes for client-side SHA-256 verification.digest: Server hashes the file and returns size and digest, reducing network transfer.
--durdir-template: ActorTemplate name:glutton-durdir-data(default): Attaches a durable data directory without memory snapshot restore.glutton-durdir-full: Attaches a durable data directory and performs a full memory snapshot restore.
DurDir Reported Metrics
DurDirWrite: Initial truncate-write creating the data file.DurDirServeInitial: First read immediately following file creation.SuspendActor: Actor suspend latency (snapshot creation + persistence upload).ResumeActor: Actor resume latency.DurDirServeAfterResume: First read after resume (measures page faults / lazy load overhead on restored volume).DurDirServeWarm: Subsequent read within the same active cycle (cached state baseline).DurDirOverwrite: In-place file overwrite with checksum verification.
Viewing Traces
You must have enabled otel tracing for your cluster to view traces.
You can find trace IDs by viewing the logs tab in the Locust UI
Benchmark output files
A run writes the following to --dest. Each run produces them fresh; none of
them are checked into the repository.
status.json:locust_exit_codeandstats_generated. Deliberately just those two keys, because it is what CI orchestration reads to decide whether a trial ran at all.stats.csv,stats_history.csv,failures.csv,exceptions.csv: Locust's own CSV output.logs.txt,traces.txt: the runner log, and the trace IDs seen during the run.stats.jsonl: one JSON object per line, one per metric. Every row carries the same five keys:timestamp,tag,test_name,metric, and a flatmeasurementsmap holding that metric's numbers.
Density frontiers
With cluster discovery enabled, stats.jsonl gains a trial_summary row
describing how densely actors are packed onto the hardware. Its measurements
map holds the raw facts and the derived numbers side by side.
machine_type,node_count,allocatable_cores,allocatable_ram_gb(GiB),worker_pod_count: the measured facts, before any arithmetic. Capacity covers the nodes the worker pods are running on rather than the whole cluster, so a separate infrastructure pool is not counted. They are recorded so the ratios below can be re-derived later, or recomputed against a different denominator.actors_per_node,actors_per_vcpu,actors_per_gb_ram: the most actors Locust reported running, over the matching capacity. The-uflag only stands in when no sample was read.actors_per_pod_p50,actors_per_pod_p90,actors_per_pod_p99: actors per worker pod across the run. Reported as a distribution rather than one average, and it spans ramp-up too, because a custom load shape has no single user count to call steady.aggregate_failure_ratio: failures over requests for the run.<operation>_failure_ratio: the same ratio for every operation Locust reported, so each test carries its own names through. The operation name is lowercased with underscores, soDurDirWritebecomesdur_dir_write_failure_ratio. A key is absent when the test has no such row, and null when the row ran no requests.
The six actors_per_* ratios rest on three assumptions. Read them before
comparing numbers across runs:
- Actors are derived, not counted. Locust only sees virtual users, so the
numerator is the peak user count times
--actors-per-user. No server-side gauge counts resident actors:ate.actor.stats.sampled_actorsdrops any actor without a live resource measurement, so suspended ones fall out. - The denominators are read once, after the run. A cluster that autoscaled mid-run is measured at its final size, so the ratio pairs a peak from one moment with a capacity from another.
- The peak assumes every actor is alive at once. A workload that creates and deletes actors as it goes never holds them all at the same time, so its real density is lower than reported.
The Kubernetes API is not required. If it is unreachable, or discovery was
skipped, the affected fields are written as null and the run still
succeeds. A null means the value was not measured. It never means zero.
Optional: Prometheus + Grafana
Locust provides graphs, statistics, etc. via the UI. However, you can install Prometheus/Grafana if you want richer details or the ability to perform deeper analysis. Skip this section if you're only using the Locust web UI.
kubectl apply -f benchmarking/monitoring.yaml
Once installed:
- Run
kubectl port-forward svc/grafana -n benchmarking 3000:3000 - Visit
http://localhost:3000in your browser.
Development
Generating gRPC Python clients
The clients are not checked in. The locust and nighthawk-ingress images
generate them at build time, and hack/verify/python-protos.sh compiles them
on every PR, so a proto change needs no extra step. For local use, such as
editor completion, run benchmarking/locust/codegen/generate.sh. It manages
its own virtual environment under locust/codegen/venv.
Unit tests
locust/unit_tests covers the runner's helpers and needs no cluster. From the
repository root:
python3 -m unittest discover -s benchmarking/locust/unit_tests