We should enable the use of special taints to isolate workloads:
avoiding noisy neighbor issues with load generator and isolating worker
nodes from all other nodes.
- [ x ] Tests pass
- [ x ] Appropriate changes to documentation are included in the PR
Part of #1590
First of two PRs. This one covers Kubernetes-side hardware discovery;
the follow-up adds the Prometheus harvest.
## What this PR does
The locust runner records no information about the hardware it runs on,
so trial results cannot be normalized across machine types or node
counts. This adds cluster capacity discovery and derives actor density
frontiers from it, persisting both the raw readings and the derived
ratios to `stats.jsonl`.
## Proposed Changes
### Cluster hardware discovery
New module `benchmarking/locust/cluster_facts.py`, kept separate from
`runner.py`. It reads node count, machine type, allocatable cores and
allocatable RAM, and the worker pod count via the official Kubernetes
Python client, which resolves in-cluster and local kubeconfig auth
without a kubectl subprocess.
Node and pod listing is proportional to cluster size, so
`--no-cluster-facts` disables discovery. Discovery is best effort in
either case: an unreachable API server or missing RBAC leaves the facts
null and does not fail the run.
No value is defaulted or inferred. Unmeasured fields are `null`; a
measured zero is `0`. The two remain distinguishable to any consumer.
### Density frontiers
Added to the `trial_summary` row in `stats.jsonl`:
| Field | Meaning |
|---|---|
| `actors_per_node` | peak actors over node count |
| `actors_per_vcpu` | peak actors over allocatable cores |
| `actors_per_gb_ram` | peak actors over allocatable RAM |
| `actors_per_pod_p50` / `_p90` / `_p99` | distribution of actors per
worker pod |
| `aggregate_failure_ratio` | failures over requests, all RPCs |
| `<operation>_failure_ratio` | one per operation Locust reported, e.g.
`dur_dir_write_failure_ratio` |
Actors per pod is reported as a distribution across the run's time
samples (`p50` / `p90` / `p99`) rather than one average, because ramp-up
and custom load shapes have no single user count to call steady.
The raw readings (`machine_type`, `node_count`, `allocatable_cores`,
`allocatable_ram_gb`, `worker_pod_count`) are written alongside the
derived ratios so they can be re-derived without re-running the trial.
### RBAC
Nodes are cluster scoped and require a ClusterRole with `list` on
`nodes`. Pods require only a namespaced Role with `list` in
`benchmark-workloads`. `locust.yaml` and `runner-job.yaml.tmpl` each
define their own separately named pair.
Note that `locust.yaml` now binds a Role in `benchmark-workloads`, so
that namespace must exist before the manifest is applied.
`benchmarking/workloads/deploy.sh` creates it.
`status.json` is unchanged.
## How this was tested
11 unit tests in `test_cluster_facts.py` covering percentile boundaries,
capacity scoped to only the nodes running worker pods, zero worker pods
treated as a valid reading, unreadable facts returning null, and flag
behavior.
Verified against a live GKE cluster. Discovery returned
`c3d-standard-8`, 1 node, 7.91 allocatable cores, 27.73 GB, 5 worker
pods, all matching `kubectl`. The emitted frontiers were re-derived by
hand from `stats_history.csv` and matched. `--no-cluster-facts` returns
all nulls.
Both manifests pass `kubectl apply --dry-run=client`, with no
cluster-scoped name collisions between them.
## References
[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Benchmarking: multi-actor glutton VUs, live window, resume retries,
client-side latency
Glutton actors per VU. Each VU creates --actors-per-user actors on
startup and cycles through them round-robin, so one goroutine can drive
many mostly-idle actors. runner.py forwards the flag to boomer-glutton.
A crashed actor stays crashed for the run (ateapi never rehabilitates
it), so it is marked on the first Aborted "crashed" error, skipped from
then on, and counted in a CrashCount stat.
Live window. --min/--max-live-time (default 0-0) set how long an actor
stays resumed between its first ping and the suspend. Up to
--max-pings-per-wake pings (default 1) run inside it, spaced 0.2-1.0s
apart. The wait window (--min/--max-wait-time) remains the gap between
one actor's suspend and the VU's next resume, and every return path
sleeps it, so a failing startUser, resume, or crashed actor does not
spin on boomer's zero-delay re-entry. With the defaults the cycle is
resume, ping, suspend, wait: the same shape as before.
Resume retries. ResumeActor retries ateapi's transient "concurrent
update conflict" Aborted up to five times with a 50ms backoff, inside
the timed call, so the conflict no longer shows up as a failure.
Client-side latency. Every gRPC row in the locust stats now reports
client wall clock, which covers retries, queueing, and the network. The
server's elapsed-time trailer stays on the trace span only.
Worker robustness. The router HTTP client keeps up to 10000 idle
connections per host so each VU reuses its connection across wakes. A
failed dynconfig fetch after the first successful one keeps the last
fetched values instead of exiting the worker.
Orchestrator. util.run logs the duration of each shell command.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Run agentgateway data plane tests as a part of substrate CI
(non-blocking to start so we can confirm it's not flaky). Also, change
the `--atenet-router` flag to `--atenet-dataplane` to make it clearer
that the flag controls ingress and egress.
I've run the e2es locally across gVisor and microVM plus the MITM
variants for both. The only skip we do for agentgateway is
`TestIngressProtocolDowngrade` because 1. the behavior its testing only
exists on the non-CONNECT atunnel ingress path and agentgateway only
sends CONNECT to atunnel and 2. I'm not sure that we want this to be a
part of the contract that substrate is bound by (e.g. do we really want
to commit to atunnel always parsing HTTP?).
My goal with getting both dataplanes into CI is to start taking steps to
codify the proxy (router + egress PEP) contract for substrate. The
telemetry they emit, atunnel expectations, etc. are all important
contracts to explicitly call out so that they don't become too coupled
to a single dataplane implementation.
> It's a good idea to open an issue first for discussion.
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
Add support for comparing PauseActor vs SuspendActor performance in the
benchmarking suite. Specifically:
- Add LifecycleMode ("suspend" vs "pause") to boomer dynconfig and
Python Locust flag registration (--lifecycle-mode).
- Update GluttonUser and DurdirUser to hibernate actors using either
PauseActor or SuspendActor based on lifecycle_mode, and ensure
DeleteActor passes AnyState=true during teardown so paused actors are
cleanly deleted without precondition errors.
- Add unit test coverage for pause vs suspend execution and teardown.
- Add pause benchmark configurations to automation/tests.yaml for
baseline, memory scaling, concurrency, and DurableDir scenarios.
- Add PauseActor and comparison latency/QPS panels to monitoring
dashboards in monitoring.yaml.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.
The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).
For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#### What pr does
Makes resume measurements require the actor's memory to actually work:
adds `ReadRAM` — a glutton request that walks the working set (reads one
byte per 4KiB page across the requested size) before responding, plus a
`--mem-read` knob so the benchmark cycle performs that walk right after
every resume.
Today's cycle proves an actor is *reachable* after resume, not that its
memory is *usable*: the ping answers without touching the working set. A
real application must read its memory to serve requests. This matters
for where restore optimization is headed — a lazy/on-demand restore
would look great on a benchmark that never reads memory (resume returns
fast, ping returns fast) while real first-requests would stall faulting
pages back in. With the walk in the cycle, "resume + first response"
includes the cost of making memory usable, however the restore path
schedules that work: eager restore pays it during resume, lazy restore
would pay it during the walk — either way the total is in the tracked
numbers.
Also upgrades churn with `WRITE_MODE_OVERWRITE_ROTATE`: overwrite at a
per-key cursor that advances past each write and wraps, so repeated
churn walks the whole array over time instead of re-dirtying the same
prefix every cycle.
#### How it works
- `ReadRAM(key, size)` walks the first `size` bytes (suffixed string,
e.g. `"1Gi"`; empty walks the whole array) of a `WriteRAM` allocation,
one byte per 4KiB page — the cheapest touch that forces every page
resident. The response returns bytes walked plus an XOR checksum of the
sampled bytes so the reads are observable and can't be elided.
- Cycle order: resume → fill (once) → **walk** → churn → ping → suspend.
The walk runs *before* churn deliberately: it must read the memory as
restored, not pages churn just rewrote; churn then re-dirties after, so
the next snapshot still carries fresh pages.
- The walk reports as its own `GluttonReadRAM` stats row — it never
pollutes ping or resume latencies. Today (eager restore) it reads warm
memory in milliseconds; a jump in this row is the signal that restore
work got deferred onto the request path.
- Config travels the established channel: `--mem-read` in the suite's
locust `flags:` → `/boomer-config` → the Go worker, passed verbatim to
the wire; glutton is the only parser. Empty = disabled; the tracked
large-memory suites set it to the full target (walk everything —
strongest signal, simplest story). Existing suites unchanged.
#### Testing
- `go test -race` across `cmd/benchmarking/glutton` and
`internal/benchmarking/boomer/...`: PASS. New tests cover the walk's
byte count and checksum, missing-key/bad-size errors, cycle call order
(fill → read → churn with the right sizes and modes), rotate-mode cursor
wrap, disabled-by-default, and walk-before-fill as a no-op.
- Cluster verification (microvm, 1Gi target, full walk):
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.
Verifications done:
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.
This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.
Part of #861
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#### What this pr do
Re-dirties part of the glutton working set on every benchmark cycle, so
repeated suspends snapshot an actor whose memory is changing — like a
live application's — instead of a set that was filled once and never
touched again.
Each iteration (after the one-time fill), the GluttonUser sends one
`WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first
`mem_churn` bytes of the working set in place. Overwrite mode already
existed on the glutton API; this PR adds no proto changes and no glutton
server changes — it is driver + config wiring only.
This is the follow-up from #1130's: the original self-driving memload
continuously re-dirtied its pages, and that property was dropped in the
move to API-driven fill. This restores it in the API-driven shape with a
dialable amount instead of the old all-or-nothing full pass.
#### Why it matters
A fill-once working set is static: every suspend after the first packs
up identical memory. If snapshotting ever gains incremental / dirty-page
optimizations, a static benchmark would measure almost nothing from
cycle two onward — and would score "upload nothing" as an infinite win.
With churn, every cycle carries a known amount of freshly-dirtied pages,
and the knob is sweepable (64Mi, 256Mi, …) so a future
differential-snapshot optimization can be demonstrated as "upload cost
scales with churn size, not total size."
## What this PR does:
Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.
The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).
## How it's configured
The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.
Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom
## Behavior notes
- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.
## Testing
- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
The automation's test-cluster creation command did not enable Managed
OpenTelemetry, but the ate-otel-config ConfigMap applied by
hack/install-ate.sh and the runner Job both target
opentelemetry-collector.gke-managed-otel.svc.cluster.local:4317. Without
the addon that name does not resolve and all benchmark telemetry is
dropped.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Measures the max RPS atenet-router's ingress side sustains at a given Envoy CPU limit,
under a tail-latency SLO, using Nighthawk's adaptive load controller in
open-loop mode against the real routing path (Host-header routing via
ext_proc to warmed glutton actors).
- benchmarking/nighthawk-ingress/: runner Job that creates and warms the actor
fleet, drives nighthawk_service + nighthawk_adaptive_load_client with
Host rotation across actors, and uploads JSON/JSONL results to GCS.
- Search converges on three thresholds — tail latency (measured
mean+2stdev must stay under tailLatencySloMs), success-rate, and
send-rate — and records which one bounded the run. The client is
oversized (fixed event loops, large pools) so the harness is never
the ceiling.
- orchestrator.py: new `type: nighthawk-ingress` tests.yaml entries; pins the
router (cpu requests=limits, envoy --concurrency) before each run.
Validated end to end on a dev GKE cluster: ~8.9k RPS at 2 Envoy CPUs
under a 25ms tail-latency SLO.
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.
Validated with focused Go tests, shellcheck, and Kustomize renders.
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
#### benchmarking: right-size actor memory, default 256Mi
Benchmark `ActorTemplates` declared no resources, so microvm actors fell
through to the 2 GiB kata default guest — and everything the guest
kernel caches rides along in the memory snapshot, inflating snapshot
size and suspend/resume latency, which the benchmarks then measure.
Set `spec.resources.limits.memory` on the `glutton` and `sleep`
templates, parameterized as `ACTOR_MEMORY` with a `256Mi` default (the
smallest size microvm admits: 128Mi VMM reserve + 128Mi guest floor),
and thread it through the deploy chain:
- `--actor-memory` on `workloads/deploy.sh` and `deploy_locust.sh`
- `--benchmark-actor-memory` on `install-ate.sh`
- Optional per-test `actorMemory` field in the automation's `tests.yaml`
for future `WriteRAM` stress suites.
# Benchmark DurDir
Fixes#673
Adds a `DurDirUser` load-generation workload that exercises the
DurableDir
suspend/resume loop end to end, verifies every served byte against a
SHA-256
digest, and emits separable latency percentiles for each step.
## The loop
per VU, first iteration: create -> resume -> WriteDisk ->
ReadDisk+verify
steady state: suspend -> resume -> ReadDisk+verify (cold)
-> ReadDisk+verify (warm)
-> WriteDisk (overwrite, TRUNCATE)
Under `onCommit: Data` the container cold-boots from the OCI image and
process
memory is discarded, so a matching digest after resume can only have
come from
the restored DurableDir. That is the durability assertion.
## Results
All six scenarios ran on GKE/gvisor at 1 VU for 1m each: `Data` vs
`Full`
snapshot scope, explicit vs implicit resume, and a 5/10/64 MiB size
sweep.
**Zero failures across every run**, so every served byte matched its
digest in
all six.
**Snapshot growth over repeated overwrites:** `SuspendActor` latency
stayed flat
across consecutive overwrite-and-suspend cycles on the same DurableDir
volume.
The loop overwrites with `WRITE_MODE_TRUNCATE`, so the file is exactly X
bytes
after every write and the captured directory contents are the same size
every
cycle. Nothing accumulates across suspends. That is the growth question
the
issue asks about.
Latency percentiles per step are in the run artifacts.
## Change surface
- **glutton:** `WriteDisk` returns size + sha256; new `ReadDisk` with a
`READ_MODE_DIGEST_ONLY` mode for measuring restore cost without paying
wire
transfer; disk RPCs exposed over HTTP mode.
- **manifests:** two new ActorTemplates, `glutton-durdir-{data,full}`,
with a
`durableDir` volume and `onPause: Full` / `onCommit: {Data,Full}`.
- **boomer:** shared actor-lifecycle plumbing extracted from the ping
task, then
a `DurDirUser` task on top of it; a general `resume_mode` knob.
- **harnesses:** `--workload` selects the task at deploy time;
`durdir.py` stub,
typed dynconfig flags, six nightly scenarios.
The two boomer commits above are incremental extractions made as the
second
workload landed. The final package layout for
`internal/benchmarking/boomer`
lands as a follow-up PR.
## Testing
- `go build ./...`, `go test -race ./cmd/benchmarking/...
./internal/benchmarking/...`
clean, no race warnings.
- `hack/verify-all.sh`: all nine checks pass. Python protos regenerated,
tree
clean.
- Both ActorTemplates reach `Ready`; golden snapshots confirmed in the
bucket.
- **Manual durability proof:** 200 MiB payload written, suspended,
resumed, and
re-read with matching sha256 across every read. Process memory discarded
and
container cold-booted in between.
- **Scale validation:** DurableDir persistence verified up to 1 GiB with
0
failures.
- **Regression gate:** `glutton_baseline_5_users` ran 919 requests with
0
failures and 7 ms ping latency post-rebase. No regression on the
existing
benchmark.
## Deliberately out of scope
- **Image size over time:** `ActorSnapshot` has no size field, so there
is
nothing for a client to read. The issue permits deferring this.
`SuspendActor`
latency is the available proxy; `atelet.snapshot.size` is the
server-side one.
- **boomer package restructure:** landing as a follow-up so this PR's
files stay
reviewable in place.
- **RAM-backed variant:** glutton already has `WriteRAM`; small
follow-up.
This PR adds the `--junit-output` flag to
`benchmarking/automation/orchestrator.py` to record test execution
durations and export structured JUnit XML reports.
Benchmarking needed updates to fix auth bitrot and to enable readint
endpointslices. Also streamlined the ability to manually benchmark
microvm and removed Python Locust workers by default to simplify
self-serve benchmarking.
Currently we are unable to easily automate the benchmarking of micro
VMs.
This PR:
* Refactors the microVM setup scripts to allow discrete installation of
uVM sanbox configs separate from counter demo
* Pipes a MicroVM container test config through the benchmarking system
* Updates default atelet daemonset to tolerate all kinds of sandboxClass
taints
This PR introduces pluggable path configurations for the benchmark
orchestrator and adds support for optional Kubernetes ServiceAccount
token-based authentication in the benchmark runner.