Commit Graph
92 Commits
Author SHA1 Message Date
Julian Gutierrez Oschmann 32e08c22d8 Print single-resource output as a bare object, not a list. (#1627)
Affects get <name>, create, resume, pause, and suspend for actors, actor
templates, atespaces, tags, and workers: JSON/YAML no longer wraps a
single resource in a List, matching kubectl's convention. Listing (no
name, or multiple names) still returns a List.

Also plumb the writer through cmd.OutOrStdout() instead of hardcoding
os.Stdout to honor Cobra's writer.
2026-09-14 09:40:37 -07:00
Zoe Zhao a536fabe22 Remove the top level --boot flag from ResumeActorRequest (#1548)
Fixes https://github.com/agent-substrate/substrate/issues/1566

Today the Resume workflow resolves its restore source in the following
order
1. First check if the actor has node-local snapshot, 
2. then its own durable external snapshot, 
3. then the template's golden snapshot. 
 
The boot flag was consulted at exactly one point in that chain, where it
suppressed using the golden-snapshot, which made its behavior much
narrower than "boot from scratch" suggests:

- Actor has its own external snapshot and boot=true: flag ignored,
restores the actor's snapshot.
- Actor has a local snapshot and boot=true: flag ignored, restores the
local snapshot.
- Actor has no snapshot, template has no golden snapshot: cold boot from
the spec regardless of the boot flag.
- Actor has no snapshot, template has a golden snapshot: boot=false
restores the golden, boot=true cold boots from the spec. This is the
only case where the flag is used.


The glutton benchmark was the only caller that set boot=true, on each
actor's first resume, to report true cold-start latency as a separate
ResumeActorColdStart stats row. @maxsmythe let me know if this is
required.

The proto field number and name were reserved.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-11 17:55:09 -04:00
haiyanmeng 5fb2c0a3cd ateapi: rename IPBlockRule to CIDRRule (#1597) 2026-09-11 13:49:51 -07:00
Taahir Ahmed dc675e8edd Identity: Move actor JWT/cert minting into the main API (#1315)
This was originally a separate service to make it easy to apply separate
authentication and authorization interceptors.

It now seems clear that our authn/z framework will be strong enough to
support atelet and external callers in one system (based on OpenFGA).

This change moves the MintJWT and MintCert RPCs into the control API,
and removes some inline authz checks that will be handled by our unified
authorizer framework.
2026-09-11 10:23:36 -07:00
Keith Mattix II 0b3d2d078f Address PR feedback
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Keith Mattix II 6bd89588dc Move to lowercase for header references
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Keith Mattix II f16fc04fa0 Move from Host header to explicit headers for actor and atespace
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Eitan Yarmush ee8d8faf09 Use Tag UIDs in snapshot storage paths (#1521)
Fixes #1508

Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.

- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.

Lint and code-generation verification also passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-08 18:27:12 -04:00
Youssuf Elshall 9ea39a5639 docs: remove stale references to ActorTemplate as a Kubernetes CRD (#1404)
## What

Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.

Fixes #368 (docs side).

> This change was prepared with AI assistance; I have reviewed and
tested it.

- [x] Docs and comment-only change; no functional code changed, no tests
affected
2026-09-04 16:40:04 -07:00
Luiz Oliveira 1e56e66b25 Rename ActorSnapshotTag to Tag (#1491)
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.

https://github.com/agent-substrate/substrate/issues/664

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 18:00:36 -04:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Haven Xia ac175376e4 Refactor the Python proto codegen into update/codegen.sh (#1483)
Fixes #1481

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 10:50:20 -07:00
Tim Hockin a9c84dc2ad API: Remove one layer of nesting in Worker status
Worker.Status.allocation.{capacity,allocated} -> Worker.status.{capacity,allocated}
2026-09-04 09:58:49 -07:00
Benjamin Elder 2c429a9906 multi-actor worker API (#1283)
Part of #1266 

This is a draft of the core API + data store changes.

It's still a large PR, apologies.

The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.

Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.

They're loadbearing for performance once we actually have multi-actor
workers.
2026-09-03 21:19:21 -07:00
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
Max Smythe f3af651567 add JQ to automation dockerfile (#1452)
This adds JQ to the automation image because it is now a critical
dependency for the installation scripts.
2026-09-03 15:12:16 -07:00
Zoe Zhao cad4ce2b7a ateapi: UpdateActor allows updating ActorTemplate. (#1365)
Fixes https://github.com/agent-substrate/substrate/issues/477 Implements
the following actor template update flow:

```
UpdateActor(): set .template = templ-v2

ResumeActor()
    resumes using templ-v2
    wait for readyz
    if fail: return failure
        .template = templ-v2; .status.current_template = templ-v1
        (the next resume will attempt templ-v2 again)
    set .status = STATUS_RUNNING
    set .status.current_template = templ-v2
    return success
    .template = templ-v2; .status.current_template = templ-v2
```
2026-09-02 16:41:32 -07:00
pmandewalkar cedb0147ff benchmarking: Refactor boomer based Locust benchmarking (#1295)
Fixes #1248 

Renamed boomer-glutton to boomer-worker.
Extracted the boomer shared utils that future non-glutton workloads may
use to the boomerutil package.
Updated durdir and glutton to import and use the changes.
- [ X] Tests pass
- [ X] Appropriate changes to documentation are included in the PR
2026-09-02 13:12:45 -07:00
Julian Gutierrez Oschmann 27b7b5904f Remove union discriminator field from ActorTemplate.volumes. (#1398)
Partially addresses #1397. Remove union discriminator `type` field from
`ActorTemplate.volumes`.
2026-09-02 09:29:19 -07:00
Sairaj Pokale 528d905b65 benchmarking: follow go.mod's Go version in the benchmarking images (#1352)
Fixes #1351

`Bump Go to 1.27` (69828945) raised the root `go.mod` and every
`hack/tools/*/go.mod` to `go 1.27.0`. `benchmarking/deploy_locust.sh
--deploy`
now fails at the `boomer-glutton` build:

go: go.mod requires go >= 1.27.0 (running go 1.26.7; GOTOOLCHAIN=local)

The `goboomer` stage builds `FROM golang:1.26-bookworm`, and the
official
`golang` images set `GOTOOLCHAIN=local` in their image config. That
disables
Go's toolchain resolution, so the base image tag becomes a second place
where
the project's Go version is declared. It drifted out of sync with
`go.mod` at
the 1.27 bump and would drift again at 1.28.

Simply bumping the tag would fix today's build and leave that second
declaration in place. This makes `go.mod` the only place the version is
declared instead.

## Changes

- **`benchmarking/locust/Dockerfile`**: set `GOTOOLCHAIN=auto` in the
`goboomer` stage, undoing the base image's `local`. Go now reads the
`go`
directive from `go.mod` and fetches a matching toolchain when the base
image
  falls behind, so a future minor bump needs no change here.


## Validation

No test covers this file. Verified by build.

- [x] `docker build --no-cache --platform linux/amd64 -f
benchmarking/locust/Dockerfile .`
compiles `boomer-glutton`, logging `go: downloading go1.27.0` where it
      previously failed.
- [x] Same build with `GOTOOLCHAIN` left at the image default still
fails with
the error above, confirming that variable is the cause rather than the
      base image version.
2026-09-02 03:01:52 -07:00
Lucky Abolorunke 811fac3e7b benchmarking: walk RAM after resume (--mem-read) and rotate the churn window (#1310)
#### What pr does

Makes resume measurements require the actor's memory to actually work:
adds `ReadRAM` — a glutton request that walks the working set (reads one
byte per 4KiB page across the requested size) before responding, plus a
`--mem-read` knob so the benchmark cycle performs that walk right after
every resume.

Today's cycle proves an actor is *reachable* after resume, not that its
memory is *usable*: the ping answers without touching the working set. A
real application must read its memory to serve requests. This matters
for where restore optimization is headed — a lazy/on-demand restore
would look great on a benchmark that never reads memory (resume returns
fast, ping returns fast) while real first-requests would stall faulting
pages back in. With the walk in the cycle, "resume + first response"
includes the cost of making memory usable, however the restore path
schedules that work: eager restore pays it during resume, lazy restore
would pay it during the walk — either way the total is in the tracked
numbers.

Also upgrades churn with `WRITE_MODE_OVERWRITE_ROTATE`: overwrite at a
per-key cursor that advances past each write and wraps, so repeated
churn walks the whole array over time instead of re-dirtying the same
prefix every cycle.

#### How it works

- `ReadRAM(key, size)` walks the first `size` bytes (suffixed string,
e.g. `"1Gi"`; empty walks the whole array) of a `WriteRAM` allocation,
one byte per 4KiB page — the cheapest touch that forces every page
resident. The response returns bytes walked plus an XOR checksum of the
sampled bytes so the reads are observable and can't be elided.
- Cycle order: resume → fill (once) → **walk** → churn → ping → suspend.
The walk runs *before* churn deliberately: it must read the memory as
restored, not pages churn just rewrote; churn then re-dirties after, so
the next snapshot still carries fresh pages.
- The walk reports as its own `GluttonReadRAM` stats row — it never
pollutes ping or resume latencies. Today (eager restore) it reads warm
memory in milliseconds; a jump in this row is the signal that restore
work got deferred onto the request path.
- Config travels the established channel: `--mem-read` in the suite's
locust `flags:` → `/boomer-config` → the Go worker, passed verbatim to
the wire; glutton is the only parser. Empty = disabled; the tracked
large-memory suites set it to the full target (walk everything —
strongest signal, simplest story). Existing suites unchanged.

#### Testing

- `go test -race` across `cmd/benchmarking/glutton` and
`internal/benchmarking/boomer/...`: PASS. New tests cover the walk's
byte count and checksum, missing-key/bad-size errors, cycle call order
(fill → read → churn with the right sizes and modes), rotate-mode cursor
wrap, disabled-by-default, and walk-before-fill as a no-op.
- Cluster verification (microvm, 1Gi target, full walk):
2026-09-02 02:27:56 -07:00
Zoe Zhao f6852b7754 Update existing demos and benchmark tests to use substrate ActorTemplate resource (#1355)
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.

Verifications done: 
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
2026-09-01 14:50:18 -07:00
Joe Betz 3efac91283 api: delete the DebugClear RPC and the Debug service (#1346)
Fixes #999.
2026-09-01 17:46:05 -04:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
Benjamin Elder 8613dd9f7b hack: repair unusable Python venvs instead of trusting the directory (#1316)
Guarding on `[ ! -d "$VENV_DIR" ]` treats an interrupted create (no
bin/activate) and an interpreter upgrade (dangling bin/python3) as a
good venv, so the script dies in `source venv/bin/activate` and the only
fix is knowing to delete the directory. Probe that the venv runs, and
rebuild with --clear to relink the interpreter.

Also install requirements unconditionally, which is cheap. The license
verifier skipped the install whenever the venv already existed, so it
passed without ever seeing a newly added dependency.

Fixes failure encountered by @ahmedtd. Not filing an issue because this
is pretty trivial.

AI-assisted.
2026-08-31 09:52:29 -07:00
Zoe Zhao e1adb33147 api: drop ActorTemplateStatus.sandbox_assets for now (#1300)
Remove `ActorTemplateStatus.sandbox_assets` along with the
`SandboxAssets` messages it referenced. Nothing ever wrote the field:
sandbox assets are resolved from the WorkerPool and SandboxConfig
objects at resume time and travel to atelet via ateletpb, so freezing
them into the template status never materialized.

This requires more thought, one option is to add it as a field of
GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a
repeated field to allow multiple goldens.
2026-08-29 10:11:49 -07:00
Tim Hockin cfeaf23f9e Repeated fields should be named plural 2026-08-28 22:08:45 -07:00
Tim Hockin ebc7f3212c Use git to find the repo root (#1157)
Most scripts already do this.
2026-08-28 21:03:11 -07:00
Eitan Yarmush 01e92cca59 Add Actor egress policy API
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Haven Xia bdf494999b api: rename ateomImage to workerImage and make it optional (#1210)
Rename `WorkerPoolSpec.AteomImage`to `WorkerImage` and drop the
constraints so the field can be left unset.

This is the basis for let an empty workerImage lets the controller
inject a versioned default image chosen by the pool's sandbox class.

Part of #861

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-28 12:46:00 -07:00
Lucky Abolorunke 1c431b1348 benchmarking: re-dirty the glutton working set each cycle (--mem-churn) (#1280)
#### What this pr do

Re-dirties part of the glutton working set on every benchmark cycle, so
repeated suspends snapshot an actor whose memory is changing — like a
live application's — instead of a set that was filled once and never
touched again.

Each iteration (after the one-time fill), the GluttonUser sends one
`WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first
`mem_churn` bytes of the working set in place. Overwrite mode already
existed on the glutton API; this PR adds no proto changes and no glutton
server changes — it is driver + config wiring only.

This is the follow-up from #1130's: the original self-driving memload
continuously re-dirtied its pages, and that property was dropped in the
move to API-driven fill. This restores it in the API-driven shape with a
dialable amount instead of the old all-or-nothing full pass.

#### Why it matters

A fill-once working set is static: every suspend after the first packs
up identical memory. If snapshotting ever gains incremental / dirty-page
optimizations, a static benchmark would measure almost nothing from
cycle two onward — and would score "upload nothing" as an infinite win.
With churn, every cycle carries a known amount of freshly-dirtied pages,
and the knob is sweepable (64Mi, 256Mi, …) so a future
differential-snapshot optimization can be demonstrated as "upload cost
scales with churn size, not total size."
2026-08-28 11:28:51 -07:00
Benjamin Elder 69428103ee verify python codegen (#1278)
That's twice today that we missed regenerating the proto clients:

1. mid-PR on https://github.com/agent-substrate/substrate/pull/1276
(caught by code review)
2. https://github.com/agent-substrate/substrate/pull/1130 (missed and
still missing)

So this PR does:

1. regenerate with latest after #1130
2. ensure verify will catch this (so CI should fail if we haven't run
it), and that update will run it when developing

fixes https://github.com/agent-substrate/substrate/issues/738

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-27 22:25:26 -04:00
Lucky Abolorunke 9cfe3dd690 benchmarking(memory usage phase 1 baseline): large-memory glutton workloads (--mem-targetLarge mem bench (#1130)
## What this PR does:

Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.

The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).

## How it's configured

The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.

Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom

## Behavior notes

- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.

## Testing

- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
2026-08-27 16:40:26 -07:00
Benjamin Elder d7ee171ec6 ateapi: document the container security context and image volume fields (#1276)
The Capabilities, SecurityContext, ImageVolumeSource and Volume.image
fields landed without doc comments, so apitool's documented rule fails
on them and they are not in the exemption backlog. Document them rather
than growing the backlog.

Fixes failure on main

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-27 18:23:28 -04:00
Zoe Zhao d8c4f8a4b9 Add controller in ate apiserver to reconcile ActorTemplate state (#1096)
Part of #477.

The controller has 2 components
1. A "producer" that periodically lists ActorTemplates and add work
items to a client-go workqueue.
2. A "consumer" that gets work items from the queue, and reconcile it.

AT will have the FAILED condition if we encounter non-retriable errors
during the transition
2026-08-26 16:42:05 -07:00
Jeff Luo 50bc287014 manifests: count telemetry by service on kind (#1070)
The volume measurement in current benchmarking repo worked on GKE only.
There the managed collector cannot hold a connector, thus
benchmarking/telemetry/meter.yaml puts a second collector in front of it
and counts the data as it goes through.

This PR adds the support to kind cluster to count the telemetry data.

Follow-up to #749.

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-25 09:48:44 -04:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Luiz Oliveira 31a2e3850a Simplify Update methods to do a whole object replace (#1108)
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.

#1011 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-24 15:18:55 -04:00
Max Smythe b8f2e0ca28 Have worker deploy script wait for worker availability (#1142)
For large numbers of workers, we should be able to have custom waits
2026-08-24 09:50:22 -07:00
Jeff Luo 4bc650a83d benchmarking: run the observability ladder as one test (#1071)
Follow-up to #749, for comment of adding the ladder design in
benchmarking for observability.
2026-08-22 01:16:25 -07:00
Sairaj Pokale d3cb3d3282 benchmarking: wait for worker pool rollout in deploy scripts (#1128)
## Summary

Wait for the benchmark worker pool deployment to roll out before
returning from `benchmarking/workloads/deploy.sh`, preventing downstream
suites from racing against unready workers.

## Key Changes

- **Rollout Wait:** Added `kubectl wait --for=create` followed by
`kubectl rollout status` on `deployment/benchmark-ateom` in
`benchmarking/workloads/deploy.sh`.
- **Configurable Timeout:** Added `--wait-timeout DURATION` flag
(default: `300s`) to `workloads/deploy.sh` and forwarded it through
`deploy_locust.sh`.
- **Validation:** Added duration regex validation (`^([0-9]+(h|m|s))+$`)
to catch invalid formats/missing units early.

## Testing

- Verified successful rollout wait on GKE: `deploy.sh --deploy
--worker-count 2 --wait-timeout 300s`
- Verified timeout failure handling: `deploy.sh --deploy --worker-count
5 --wait-timeout 1s`
- Verified flag validation rejects invalid formats (`180`, `0`,
`invalid`).
2026-08-21 14:36:29 -07:00
Jeff Luo 8ccf0edef7 benchmarking: hold the memory of the telemetry meter (#1069)
Follow-up to #749. One of three; the other two are independent of this
one.

This closes the review thread that stayed open on that PR:
krisztianfekete asked
for `memory_limiter` and `GOMEMLIMIT` on the meter, and I kept it open
to track.
2026-08-21 12:44:02 -04:00
Jeff Luo 52287c4b41 benchmarking: enable Managed OpenTelemetry on the benchmark test cluster (#887)
The automation's test-cluster creation command did not enable Managed
OpenTelemetry, but the ate-otel-config ConfigMap applied by
hack/install-ate.sh and the runner Job both target
opentelemetry-collector.gke-managed-otel.svc.cluster.local:4317. Without
the addon that name does not resolve and all benchmark telemetry is
dropped.
2026-08-21 12:42:00 -04:00
Max Smythe 4287e1b407 Allow benchmarks to customize substrate deployment per-test (#1110)
Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-20 18:39:11 -05:00
Chuang Wang 3eaf6c2f36 benchmarking: add Nighthawk ingress capacity benchmark
Measures the max RPS atenet-router's ingress side sustains at a given Envoy CPU limit,
under a tail-latency SLO, using Nighthawk's adaptive load controller in
open-loop mode against the real routing path (Host-header routing via
ext_proc to warmed glutton actors).

- benchmarking/nighthawk-ingress/: runner Job that creates and warms the actor
  fleet, drives nighthawk_service + nighthawk_adaptive_load_client with
  Host rotation across actors, and uploads JSON/JSONL results to GCS.
- Search converges on three thresholds — tail latency (measured
  mean+2stdev must stay under tailLatencySloMs), success-rate, and
  send-rate — and records which one bounded the run. The client is
  oversized (fixed event loops, large pools) so the harness is never
  the ceiling.
- orchestrator.py: new `type: nighthawk-ingress` tests.yaml entries; pins the
  router (cpu requests=limits, envoy --concurrency) before each run.

Validated end to end on a dev GKE cluster: ~8.9k RPS at 2 Envoy CPUs
under a 25ms tail-latency SLO.
2026-08-20 11:20:00 -07:00
Eitan Yarmush a06f464e15 Remove ateapi token client mode (#1045)
Removes the non-functional token/JWT mode for in-cluster ateapi clients.
Clients now always use mTLS certificates; related flags, install and
benchmark plumbing, tests, and overlays are deleted.

Validated with focused Go tests, shellcheck, and Kustomize renders.

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-19 17:22:26 -07:00
Lucky Abolorunke d279ef55a2 benchmarking: right-size actor memory, default 256Mi (#1046)
#### benchmarking: right-size actor memory, default 256Mi

Benchmark `ActorTemplates` declared no resources, so microvm actors fell
through to the 2 GiB kata default guest — and everything the guest
kernel caches rides along in the memory snapshot, inflating snapshot
size and suspend/resume latency, which the benchmarks then measure.

Set `spec.resources.limits.memory` on the `glutton` and `sleep`
templates, parameterized as `ACTOR_MEMORY` with a `256Mi` default (the
smallest size microvm admits: 128Mi VMM reserve + 128Mi guest floor),
and thread it through the deploy chain:
- `--actor-memory` on `workloads/deploy.sh` and `deploy_locust.sh`
- `--benchmark-actor-memory` on `install-ate.sh`
- Optional per-test `actorMemory` field in the automation's `tests.yaml`
for future `WriteRAM` stress suites.
2026-08-18 20:41:27 -07:00
Sairaj Pokale ba45517fa0 benchmarks: durdir benchmark that evaluates the efficacy and performance of DurDir (#907)
# Benchmark DurDir

Fixes #673

Adds a `DurDirUser` load-generation workload that exercises the
DurableDir
suspend/resume loop end to end, verifies every served byte against a
SHA-256
digest, and emits separable latency percentiles for each step.

## The loop

per VU, first iteration: create -> resume -> WriteDisk ->
ReadDisk+verify
steady state: suspend -> resume -> ReadDisk+verify (cold)
                                      -> ReadDisk+verify (warm)
                                      -> WriteDisk (overwrite, TRUNCATE)

Under `onCommit: Data` the container cold-boots from the OCI image and
process
memory is discarded, so a matching digest after resume can only have
come from
the restored DurableDir. That is the durability assertion.

## Results

All six scenarios ran on GKE/gvisor at 1 VU for 1m each: `Data` vs
`Full`
snapshot scope, explicit vs implicit resume, and a 5/10/64 MiB size
sweep.
**Zero failures across every run**, so every served byte matched its
digest in
all six.

**Snapshot growth over repeated overwrites:** `SuspendActor` latency
stayed flat
across consecutive overwrite-and-suspend cycles on the same DurableDir
volume.
The loop overwrites with `WRITE_MODE_TRUNCATE`, so the file is exactly X
bytes
after every write and the captured directory contents are the same size
every
cycle. Nothing accumulates across suspends. That is the growth question
the
issue asks about.

Latency percentiles per step are in the run artifacts.

## Change surface

- **glutton:** `WriteDisk` returns size + sha256; new `ReadDisk` with a
`READ_MODE_DIGEST_ONLY` mode for measuring restore cost without paying
wire
  transfer; disk RPCs exposed over HTTP mode.
- **manifests:** two new ActorTemplates, `glutton-durdir-{data,full}`,
with a
  `durableDir` volume and `onPause: Full` / `onCommit: {Data,Full}`.
- **boomer:** shared actor-lifecycle plumbing extracted from the ping
task, then
  a `DurDirUser` task on top of it; a general `resume_mode` knob.
- **harnesses:** `--workload` selects the task at deploy time;
`durdir.py` stub,
  typed dynconfig flags, six nightly scenarios.

The two boomer commits above are incremental extractions made as the
second
workload landed. The final package layout for
`internal/benchmarking/boomer`
lands as a follow-up PR.

## Testing

- `go build ./...`, `go test -race ./cmd/benchmarking/...
./internal/benchmarking/...`
  clean, no race warnings.
- `hack/verify-all.sh`: all nine checks pass. Python protos regenerated,
tree
  clean.
- Both ActorTemplates reach `Ready`; golden snapshots confirmed in the
bucket.
- **Manual durability proof:** 200 MiB payload written, suspended,
resumed, and
re-read with matching sha256 across every read. Process memory discarded
and
  container cold-booted in between.
- **Scale validation:** DurableDir persistence verified up to 1 GiB with
0
  failures.
- **Regression gate:** `glutton_baseline_5_users` ran 919 requests with
0
failures and 7 ms ping latency post-rebase. No regression on the
existing
  benchmark.

## Deliberately out of scope

- **Image size over time:** `ActorSnapshot` has no size field, so there
is
nothing for a client to read. The issue permits deferring this.
`SuspendActor`
latency is the available proxy; `atelet.snapshot.size` is the
server-side one.
- **boomer package restructure:** landing as a follow-up so this PR's
files stay
  reviewable in place.
- **RAM-backed variant:** glutton already has `WriteRAM`; small
follow-up.
2026-08-18 20:40:08 -07:00
Julian Gutierrez Oschmann 4c1bd9d36b Add status fields to Substrate resources (#1025)
Add `status` fields to all Substrate resources that need it. Move
server-owned fields under it.

Fixes #1006 .
2026-08-18 11:58:31 -07:00
Michał Kuliński a778517752 benchmarking/orchestrator: add --junit-output flag and record test durations (#990)
This PR adds the `--junit-output` flag to
`benchmarking/automation/orchestrator.py` to record test execution
durations and export structured JUnit XML reports.
2026-08-17 14:22:54 -07:00