Affects get <name>, create, resume, pause, and suspend for actors, actor
templates, atespaces, tags, and workers: JSON/YAML no longer wraps a
single resource in a List, matching kubectl's convention. Listing (no
name, or multiple names) still returns a List.
Also plumb the writer through cmd.OutOrStdout() instead of hardcoding
os.Stdout to honor Cobra's writer.
Fixes https://github.com/agent-substrate/substrate/issues/1566
Today the Resume workflow resolves its restore source in the following
order
1. First check if the actor has node-local snapshot,
2. then its own durable external snapshot,
3. then the template's golden snapshot.
The boot flag was consulted at exactly one point in that chain, where it
suppressed using the golden-snapshot, which made its behavior much
narrower than "boot from scratch" suggests:
- Actor has its own external snapshot and boot=true: flag ignored,
restores the actor's snapshot.
- Actor has a local snapshot and boot=true: flag ignored, restores the
local snapshot.
- Actor has no snapshot, template has no golden snapshot: cold boot from
the spec regardless of the boot flag.
- Actor has no snapshot, template has a golden snapshot: boot=false
restores the golden, boot=true cold boots from the spec. This is the
only case where the flag is used.
The glutton benchmark was the only caller that set boot=true, on each
actor's first resume, to report true cold-start latency as a separate
ResumeActorColdStart stats row. @maxsmythe let me know if this is
required.
The proto field number and name were reserved.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
This was originally a separate service to make it easy to apply separate
authentication and authorization interceptors.
It now seems clear that our authn/z framework will be strong enough to
support atelet and external callers in one system (based on OpenFGA).
This change moves the MintJWT and MintCert RPCs into the control API,
and removes some inline authz checks that will be handled by our unified
authorizer framework.
Fixes#1508
Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.
- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.
Lint and code-generation verification also passed.
---------
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
## What
Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.
Fixes#368 (docs side).
> This change was prepared with AI assistance; I have reviewed and
tested it.
- [x] Docs and comment-only change; no functional code changed, no tests
affected
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.
https://github.com/agent-substrate/substrate/issues/664
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#664
This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489
It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).
Now, an external snapshot is owned by a single resource:
- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.
Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.
this PR:
- Drops table actor_snapshots
- Keeps table actor_snapshot_tags
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#1481
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #1266
This is a draft of the core API + data store changes.
It's still a large PR, apologies.
The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.
Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.
They're loadbearing for performance once we actually have multi-actor
workers.
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.
The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).
For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes https://github.com/agent-substrate/substrate/issues/477 Implements
the following actor template update flow:
```
UpdateActor(): set .template = templ-v2
ResumeActor()
resumes using templ-v2
wait for readyz
if fail: return failure
.template = templ-v2; .status.current_template = templ-v1
(the next resume will attempt templ-v2 again)
set .status = STATUS_RUNNING
set .status.current_template = templ-v2
return success
.template = templ-v2; .status.current_template = templ-v2
```
Fixes#1248
Renamed boomer-glutton to boomer-worker.
Extracted the boomer shared utils that future non-glutton workloads may
use to the boomerutil package.
Updated durdir and glutton to import and use the changes.
- [ X] Tests pass
- [ X] Appropriate changes to documentation are included in the PR
Fixes#1351
`Bump Go to 1.27` (69828945) raised the root `go.mod` and every
`hack/tools/*/go.mod` to `go 1.27.0`. `benchmarking/deploy_locust.sh
--deploy`
now fails at the `boomer-glutton` build:
go: go.mod requires go >= 1.27.0 (running go 1.26.7; GOTOOLCHAIN=local)
The `goboomer` stage builds `FROM golang:1.26-bookworm`, and the
official
`golang` images set `GOTOOLCHAIN=local` in their image config. That
disables
Go's toolchain resolution, so the base image tag becomes a second place
where
the project's Go version is declared. It drifted out of sync with
`go.mod` at
the 1.27 bump and would drift again at 1.28.
Simply bumping the tag would fix today's build and leave that second
declaration in place. This makes `go.mod` the only place the version is
declared instead.
## Changes
- **`benchmarking/locust/Dockerfile`**: set `GOTOOLCHAIN=auto` in the
`goboomer` stage, undoing the base image's `local`. Go now reads the
`go`
directive from `go.mod` and fetches a matching toolchain when the base
image
falls behind, so a future minor bump needs no change here.
## Validation
No test covers this file. Verified by build.
- [x] `docker build --no-cache --platform linux/amd64 -f
benchmarking/locust/Dockerfile .`
compiles `boomer-glutton`, logging `go: downloading go1.27.0` where it
previously failed.
- [x] Same build with `GOTOOLCHAIN` left at the image default still
fails with
the error above, confirming that variable is the cause rather than the
base image version.
#### What pr does
Makes resume measurements require the actor's memory to actually work:
adds `ReadRAM` — a glutton request that walks the working set (reads one
byte per 4KiB page across the requested size) before responding, plus a
`--mem-read` knob so the benchmark cycle performs that walk right after
every resume.
Today's cycle proves an actor is *reachable* after resume, not that its
memory is *usable*: the ping answers without touching the working set. A
real application must read its memory to serve requests. This matters
for where restore optimization is headed — a lazy/on-demand restore
would look great on a benchmark that never reads memory (resume returns
fast, ping returns fast) while real first-requests would stall faulting
pages back in. With the walk in the cycle, "resume + first response"
includes the cost of making memory usable, however the restore path
schedules that work: eager restore pays it during resume, lazy restore
would pay it during the walk — either way the total is in the tracked
numbers.
Also upgrades churn with `WRITE_MODE_OVERWRITE_ROTATE`: overwrite at a
per-key cursor that advances past each write and wraps, so repeated
churn walks the whole array over time instead of re-dirtying the same
prefix every cycle.
#### How it works
- `ReadRAM(key, size)` walks the first `size` bytes (suffixed string,
e.g. `"1Gi"`; empty walks the whole array) of a `WriteRAM` allocation,
one byte per 4KiB page — the cheapest touch that forces every page
resident. The response returns bytes walked plus an XOR checksum of the
sampled bytes so the reads are observable and can't be elided.
- Cycle order: resume → fill (once) → **walk** → churn → ping → suspend.
The walk runs *before* churn deliberately: it must read the memory as
restored, not pages churn just rewrote; churn then re-dirties after, so
the next snapshot still carries fresh pages.
- The walk reports as its own `GluttonReadRAM` stats row — it never
pollutes ping or resume latencies. Today (eager restore) it reads warm
memory in milliseconds; a jump in this row is the signal that restore
work got deferred onto the request path.
- Config travels the established channel: `--mem-read` in the suite's
locust `flags:` → `/boomer-config` → the Go worker, passed verbatim to
the wire; glutton is the only parser. Empty = disabled; the tracked
large-memory suites set it to the full target (walk everything —
strongest signal, simplest story). Existing suites unchanged.
#### Testing
- `go test -race` across `cmd/benchmarking/glutton` and
`internal/benchmarking/boomer/...`: PASS. New tests cover the walk's
byte count and checksum, missing-key/bad-size errors, cycle call order
(fill → read → churn with the right sizes and modes), rotate-mode cursor
wrap, disabled-by-default, and walk-before-fill as a no-op.
- Cluster verification (microvm, 1Gi target, full walk):
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.
Verifications done:
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
This PR is large, I grouped changes to the following commits:
- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.
Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no
bin/activate) and an interpreter upgrade (dangling bin/python3) as a
good venv, so the script dies in `source venv/bin/activate` and the only
fix is knowing to delete the directory. Probe that the venv runs, and
rebuild with --clear to relink the interpreter.
Also install requirements unconditionally, which is cheap. The license
verifier skipped the install whenever the venv already existed, so it
passed without ever seeing a newly added dependency.
Fixes failure encountered by @ahmedtd. Not filing an issue because this
is pretty trivial.
AI-assisted.
Remove `ActorTemplateStatus.sandbox_assets` along with the
`SandboxAssets` messages it referenced. Nothing ever wrote the field:
sandbox assets are resolved from the WorkerPool and SandboxConfig
objects at resume time and travel to atelet via ateletpb, so freezing
them into the template status never materialized.
This requires more thought, one option is to add it as a field of
GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a
repeated field to allow multiple goldens.
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.
This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.
Part of #861
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#### What this pr do
Re-dirties part of the glutton working set on every benchmark cycle, so
repeated suspends snapshot an actor whose memory is changing — like a
live application's — instead of a set that was filled once and never
touched again.
Each iteration (after the one-time fill), the GluttonUser sends one
`WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first
`mem_churn` bytes of the working set in place. Overwrite mode already
existed on the glutton API; this PR adds no proto changes and no glutton
server changes — it is driver + config wiring only.
This is the follow-up from #1130's: the original self-driving memload
continuously re-dirtied its pages, and that property was dropped in the
move to API-driven fill. This restores it in the API-driven shape with a
dialable amount instead of the old all-or-nothing full pass.
#### Why it matters
A fill-once working set is static: every suspend after the first packs
up identical memory. If snapshotting ever gains incremental / dirty-page
optimizations, a static benchmark would measure almost nothing from
cycle two onward — and would score "upload nothing" as an infinite win.
With churn, every cycle carries a known amount of freshly-dirtied pages,
and the knob is sweepable (64Mi, 256Mi, …) so a future
differential-snapshot optimization can be demonstrated as "upload cost
scales with churn size, not total size."
## What this PR does:
Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.
The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).
## How it's configured
The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.
Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom
## Behavior notes
- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.
## Testing
- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
The Capabilities, SecurityContext, ImageVolumeSource and Volume.image
fields landed without doc comments, so apitool's documented rule fails
on them and they are not in the exemption backlog. Document them rather
than growing the backlog.
Fixes failure on main
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Part of #477.
The controller has 2 components
1. A "producer" that periodically lists ActorTemplates and add work
items to a client-go workqueue.
2. A "consumer" that gets work items from the queue, and reconcile it.
AT will have the FAILED condition if we encounter non-retriable errors
during the transition
The volume measurement in current benchmarking repo worked on GKE only.
There the managed collector cannot hold a connector, thus
benchmarking/telemetry/meter.yaml puts a second collector in front of it
and counts the data as it goes through.
This PR adds the support to kind cluster to count the telemetry data.
Follow-up to #749.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Summary
A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.
- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable
## Benchmarking
Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:
-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing
---------
Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.
#1011
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
## Summary
Wait for the benchmark worker pool deployment to roll out before
returning from `benchmarking/workloads/deploy.sh`, preventing downstream
suites from racing against unready workers.
## Key Changes
- **Rollout Wait:** Added `kubectl wait --for=create` followed by
`kubectl rollout status` on `deployment/benchmark-ateom` in
`benchmarking/workloads/deploy.sh`.
- **Configurable Timeout:** Added `--wait-timeout DURATION` flag
(default: `300s`) to `workloads/deploy.sh` and forwarded it through
`deploy_locust.sh`.
- **Validation:** Added duration regex validation (`^([0-9]+(h|m|s))+$`)
to catch invalid formats/missing units early.
## Testing
- Verified successful rollout wait on GKE: `deploy.sh --deploy
--worker-count 2 --wait-timeout 300s`
- Verified timeout failure handling: `deploy.sh --deploy --worker-count
5 --wait-timeout 1s`
- Verified flag validation rejects invalid formats (`180`, `0`,
`invalid`).
Follow-up to #749. One of three; the other two are independent of this
one.
This closes the review thread that stayed open on that PR:
krisztianfekete asked
for `memory_limiter` and `GOMEMLIMIT` on the meter, and I kept it open
to track.
The automation's test-cluster creation command did not enable Managed
OpenTelemetry, but the ate-otel-config ConfigMap applied by
hack/install-ate.sh and the runner Job both target
opentelemetry-collector.gke-managed-otel.svc.cluster.local:4317. Without
the addon that name does not resolve and all benchmark telemetry is
dropped.
Fixes #<issue_number_goes_here>
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Measures the max RPS atenet-router's ingress side sustains at a given Envoy CPU limit,
under a tail-latency SLO, using Nighthawk's adaptive load controller in
open-loop mode against the real routing path (Host-header routing via
ext_proc to warmed glutton actors).
- benchmarking/nighthawk-ingress/: runner Job that creates and warms the actor
fleet, drives nighthawk_service + nighthawk_adaptive_load_client with
Host rotation across actors, and uploads JSON/JSONL results to GCS.
- Search converges on three thresholds — tail latency (measured
mean+2stdev must stay under tailLatencySloMs), success-rate, and
send-rate — and records which one bounded the run. The client is
oversized (fixed event loops, large pools) so the harness is never
the ceiling.
- orchestrator.py: new `type: nighthawk-ingress` tests.yaml entries; pins the
router (cpu requests=limits, envoy --concurrency) before each run.
Validated end to end on a dev GKE cluster: ~8.9k RPS at 2 Envoy CPUs
under a 25ms tail-latency SLO.
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.
Validated with focused Go tests, shellcheck, and Kustomize renders.
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
#### benchmarking: right-size actor memory, default 256Mi
Benchmark `ActorTemplates` declared no resources, so microvm actors fell
through to the 2 GiB kata default guest — and everything the guest
kernel caches rides along in the memory snapshot, inflating snapshot
size and suspend/resume latency, which the benchmarks then measure.
Set `spec.resources.limits.memory` on the `glutton` and `sleep`
templates, parameterized as `ACTOR_MEMORY` with a `256Mi` default (the
smallest size microvm admits: 128Mi VMM reserve + 128Mi guest floor),
and thread it through the deploy chain:
- `--actor-memory` on `workloads/deploy.sh` and `deploy_locust.sh`
- `--benchmark-actor-memory` on `install-ate.sh`
- Optional per-test `actorMemory` field in the automation's `tests.yaml`
for future `WriteRAM` stress suites.
# Benchmark DurDir
Fixes#673
Adds a `DurDirUser` load-generation workload that exercises the
DurableDir
suspend/resume loop end to end, verifies every served byte against a
SHA-256
digest, and emits separable latency percentiles for each step.
## The loop
per VU, first iteration: create -> resume -> WriteDisk ->
ReadDisk+verify
steady state: suspend -> resume -> ReadDisk+verify (cold)
-> ReadDisk+verify (warm)
-> WriteDisk (overwrite, TRUNCATE)
Under `onCommit: Data` the container cold-boots from the OCI image and
process
memory is discarded, so a matching digest after resume can only have
come from
the restored DurableDir. That is the durability assertion.
## Results
All six scenarios ran on GKE/gvisor at 1 VU for 1m each: `Data` vs
`Full`
snapshot scope, explicit vs implicit resume, and a 5/10/64 MiB size
sweep.
**Zero failures across every run**, so every served byte matched its
digest in
all six.
**Snapshot growth over repeated overwrites:** `SuspendActor` latency
stayed flat
across consecutive overwrite-and-suspend cycles on the same DurableDir
volume.
The loop overwrites with `WRITE_MODE_TRUNCATE`, so the file is exactly X
bytes
after every write and the captured directory contents are the same size
every
cycle. Nothing accumulates across suspends. That is the growth question
the
issue asks about.
Latency percentiles per step are in the run artifacts.
## Change surface
- **glutton:** `WriteDisk` returns size + sha256; new `ReadDisk` with a
`READ_MODE_DIGEST_ONLY` mode for measuring restore cost without paying
wire
transfer; disk RPCs exposed over HTTP mode.
- **manifests:** two new ActorTemplates, `glutton-durdir-{data,full}`,
with a
`durableDir` volume and `onPause: Full` / `onCommit: {Data,Full}`.
- **boomer:** shared actor-lifecycle plumbing extracted from the ping
task, then
a `DurDirUser` task on top of it; a general `resume_mode` knob.
- **harnesses:** `--workload` selects the task at deploy time;
`durdir.py` stub,
typed dynconfig flags, six nightly scenarios.
The two boomer commits above are incremental extractions made as the
second
workload landed. The final package layout for
`internal/benchmarking/boomer`
lands as a follow-up PR.
## Testing
- `go build ./...`, `go test -race ./cmd/benchmarking/...
./internal/benchmarking/...`
clean, no race warnings.
- `hack/verify-all.sh`: all nine checks pass. Python protos regenerated,
tree
clean.
- Both ActorTemplates reach `Ready`; golden snapshots confirmed in the
bucket.
- **Manual durability proof:** 200 MiB payload written, suspended,
resumed, and
re-read with matching sha256 across every read. Process memory discarded
and
container cold-booted in between.
- **Scale validation:** DurableDir persistence verified up to 1 GiB with
0
failures.
- **Regression gate:** `glutton_baseline_5_users` ran 919 requests with
0
failures and 7 ms ping latency post-rebase. No regression on the
existing
benchmark.
## Deliberately out of scope
- **Image size over time:** `ActorSnapshot` has no size field, so there
is
nothing for a client to read. The issue permits deferring this.
`SuspendActor`
latency is the available proxy; `atelet.snapshot.size` is the
server-side one.
- **boomer package restructure:** landing as a follow-up so this PR's
files stay
reviewable in place.
- **RAM-backed variant:** glutton already has `WriteRAM`; small
follow-up.
This PR adds the `--junit-output` flag to
`benchmarking/automation/orchestrator.py` to record test execution
durations and export structured JUnit XML reports.