87 Commits
Author SHA1 Message Date
Aditya ShantanuandAditya Shantanu 4e079e7b75 benchmarking(agentsession): define the session script in YAML, selectable per run (#2046)
## What this is

Follow-up to #1934, as discussed there: the agent-session script moves
out of Go and into YAML, so a workload can be read, edited, and selected
without touching the driver.

- The default script is now
`internal/benchmarking/boomer/agentsession/scripts/coding-session.yaml`,
embedded in the boomer worker. It was generated from the old Go table
and round-trips byte-for-byte on every op.
- `--agentsession-script <name>` picks a built-in variant (any
`scripts/<name>.yaml`). Resolved on the worker's first iteration (the
knob arrives with the first spawn message, after Init); a bad name is
logged and retried each iteration, so fixing the knob and re-swarming
recovers.
- Decoding is strict and runs the same invariants the old tests
enforced: unknown op kinds or fields, arguments a kind does not take,
reads before writes, walks before fills, duplicate steps, non-positive
think times.
- Each script declares `min_actor_memory`. The worker reads the glutton
template's memory limit at start and refuses to run against a smaller
actor, with a message naming `--actor-memory`. That replaces the
README's advice with a guard.
- README gains a "Writing an agent-session script" section with the
format and the op table.

## Not in this PR

Loading a script from a file on the worker (ConfigMap mount) is the next
PR, on top of this loader.

## Testing

`go test -race` on the benchmarking packages, boilerplate, gofmt, and
golangci-lint pass. New tests: every embedded script loads; default
script shape (20 steps, 1Gi floor, declared bytes under budget);
encode/decode round trip; 14 rejection cases; the template memory guard
in refuse, exact, and no-limit modes.

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-10-01 17:09:42 +00:00
Nishanth Kotla 3f789b309a benchmarking/locust: harvest server-side telemetry from Prometheus (#1725)
Part of #1590

## What this PR does

Locust measures the client side only. This adds a post-run harvest of
server-side ground truth from Prometheus, written to a new
`server_summary.json` and summarized as one row in `stats.jsonl`.

## Proposed Changes

### Steady-state scoping

`server_telemetry.py` derives the steady-state window from
`stats_history.csv`: it starts at the first sample at 90% of peak user
count and ends at the last one. Ramp-up is excluded from every metric,
and the snapshot block also stops at the last full-load sample, so
teardown suspends are left out. Under a step-ladder load shape the
window covers only the top step.

### Harvested metrics

**Cluster packing**, from `ate_workerpool_workers`: busy workers
(`partial` + `at_capacity`) over the pool size at each sample, where the
pool size is the sum of all worker states, so it follows scale-ups and
scale-downs mid-run. Only the live ateapi is read: after a redeploy the
collector keeps re-exporting exited ateapi processes' last values for a
few minutes, which would otherwise inflate the counts. Reported as min,
p50, p90, p95, p99, max and mean (`avg`), plus the underlying timeseries
so transient spikes remain visible.

**Kernel pressure**, from cAdvisor PSI: CPU, memory and IO stall
percentages for the node and for the worker pods, each as the same
percentile set over the window.

**Snapshots**: size mean and p50/p90/p95/p99, checkpoint counts both
in-window and cumulative, restore and checkpoint latency mean and
p50/p90/p95/p99 from the AteomHerder RPC histograms, and checkpoint
throughput as `checkpoint_mb_s`. The atelet exports metrics on an
interval, so its data reaches Prometheus late: the harvest waits
`--atelet-lag-s` (default 70s, enough for the OTel SDK's 60s default
export and a 10s scrape) and reads both window edges half that late.

### Constraints and failure behavior

The module uses only the standard library, since the locust image is
distroless, and every request carries a timeout. An unreachable
Prometheus records nulls and does not fail the run. `--prometheus-url`
overrides the in-cluster default, and `--atelet-lag-s` sets the wait for
the atelet's last export.

Unmeasured fields are `null` and a measured zero is `0`, consistent with
the rest of the runner. Prometheus exposes no byte counter on the
restore path, so no restore throughput field is emitted rather than
deriving one indirectly.

### Output

`server_summary.json` holds the full nested artifact. `stats.jsonl`
receives a single `server_summary` row with 45 flat keys for graphing:
packing, node PSI and pod PSI at p50/p90/p95/p99, plus the snapshot
means, percentiles, `checkpoints_in_window` and `checkpoint_mb_s`.

`status.json` is unchanged.

## How this was tested

14 unit tests in `test_server_telemetry.py` covering steady-state
detection, percentile boundaries, the range-query window guard,
malformed Prometheus responses, the packing and checkpoint arithmetic,
the per-sample pool size, ignoring exited ateapi series, a missing
denominator returning null, snapshot fields returning null rather than
zero, and telemetry surviving a missing stats CSV.

Verified on 2 user / 2 worker and 4 user / 2 worker sympy runs. All 45
keys matched an independent recomputation from raw Prometheus, and the
snapshot block was cross-checked against atelet logs. A run with
`--prometheus-url` pointed at an unreachable address completes normally
with the affected fields null.

Re-harvested 5 past runs (Glutton and SWE-perf, 3 × 60 and 30 × 8
workers) from Prometheus with the new code. Runs with a steady pool
match the previous output exactly, and runs that followed a redeploy now
read the real pool size on every sample.

## References

[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-10-01 16:59:41 +00:00
Aditya ShantanuandAditya Shantanu bc3cbc519d benchmarking: agent-session workload — a scripted coding agent that suspends while the LLM thinks (#1934)
## What this is

We keep saying Substrate's sweet spot is "agents that are idle most of
the time" — but none of our benchmarks actually behave like one. Glutton
hammers one resource at a time, and sweperf needs an external image with
a replay trace. This adds a workload that acts like the thing we're
building for: a coding agent working through a task.

Each locust user is one session. The actor gets told to do the kind of
things a coding agent does — clone a repo, install deps, build, hit a
failing test, fix it, write tests, refactor, package — twenty steps,
each one costing the sandbox the CPU, memory, disk, and network the real
action would. Between steps the "LLM is thinking," so the driver
suspends the actor, and the next step's first request wakes it back up
through the router. That parked wake (`WakeFirstTouch` in the stats) is
the number this whole benchmark exists to measure.

The part I care most about: the entire workload is one table in
`internal/benchmarking/boomer/agentsession/script.go`. Every step says
in plain English what the agent is doing and what it costs. If you want
to know what step 6 does to the sandbox, you read step 6. If you want a
different workload, you edit the table — tests will catch you if you
write a step that reads a file nothing wrote, or blow the actor's memory
budget.

To act the steps out, glutton grew two RPCs: `BurnCPU` (compute-bound
work) and `Ingest` (bytes that actually cross the network before hitting
disk, so a "git clone" is a real download, not a local write).

## How it went when we ran it

Validated on a fresh 2-node GKE cluster, micro-VM first, then gVisor on
the same hardware. Smoke runs were clean on both classes (gVisor: 1340
requests over 7 full laps, zero failures, wake p50 1.4s; micro-VM: wake
p50 2.1s). At 20 concurrent sessions with realistic think times, gVisor
held 2.6% failures with wake p50 1.3s. Pushing past the knee (~7
concurrently-active sessions on 8 vCPUs) was also useful: it reproduced
the ateom-socket-vanishing failure from #1133 and left six actors
permanently wedged in DELETING — a live repro of #1665.

Two things the first live run taught us are already folded in: RAM
refills are in-place so repeat laps don't transiently double the guest
heap (512Mi micro-VM actors OOM'd without this — use 1Gi), and the
driver replaces an actor after three failed steps in a row, because a
CRASHED actor never comes back on its own.

## Future changes

Right now the script is compiled in — changing what steps do means
editing the table and rebuilding the image. That's deliberate for this
PR (one reviewable, test-guarded source of truth), but the follow-up
we've agreed on is to make the script runtime-configurable: named script
variants selectable per run first, then accepting a full script as a
file so operators can define workloads without touching Go. That lands
as its own PR once this one is in.

---------

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
2026-09-30 23:54:36 +00:00
Nishanth Kotla 9e525467d1 benchmarking: actor telemetry for the sweperf workload (#1848)
# benchmarking: actor telemetry for the sweperf workload

## What this does

Adds the suspend/resume and task-execution timings from the telemetry
proposal to the sweperf boomer workload.

**New locust rows**

| row | what it measures |
|---|---|
| `ResumeToFirstExec` | resume RPC start until the sandbox accepts the
cycle's `/execute` |
| `TaskCEL` | in-container execution time for one sweperf task (4
cycles), as reported by `replay.py` |
| `CycleCEL` | in-container execution time for one cycle |
| `TaskWallClock` | client wall clock for one task, excluding think time
|
| `<rpc>_rtt` | client round trip for each control-plane RPC |

The four derived rows are recorded under the `actor` method, the `_rtt`
rows under `grpc`.

We keep both the client RTT and the server-side elapsed for every
control-plane RPC so the network and queueing overhead stays visible
separately. The existing `ResumeActor` / `SuspendActor` rows keep the
server-side elapsed from the response trailer.

The liveness check at session start leaves the actor running, so the
first cycle's resume is a no-op. Its successful `ResumeActor`,
`ResumeActor_rtt` and `ResumeToFirstExec` samples are skipped (failures
are still recorded), the same way ateapi skips no-op resumes.

## Poll interval

The `/status` poll interval is now configurable with
`--sweperf-poll-interval-ms` (env `LOCUST_SWEPERF_POLL_INTERVAL_MS`),
default 100 ms (was a fixed 25 ms). It flows through dynconfig like the
other `--sweperf-*` flags and is read per job, so a mid-run change
applies from the next cycle.

## Documentation

`benchmarking/README.md` had no sweperf section, so this adds one
listing every row the workload emits and the poll interval flag,
matching the existing DurDir section.

## Testing

`go test -race ./internal/benchmarking/boomer/...`, plus runs on a real
cluster (sympy, 4 users / 2 workers, gVisor): 5 min at the default 100
ms poll interval and 2 min at 250 ms.

- Both had 0 failures and all expected rows.
- The min `ResumeActor` is 334 ms, so the first-cycle no-op is no longer
recorded.
- `ResumeActor` / `SuspendActor` counts match (197 / 197), and each
`_rtt` is within ~1 ms of its server-side row.
- Client overhead over CEL per cycle is 65 ms at 100 ms vs 141 ms at 250
ms, so the flag takes effect.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-30 19:23:53 +00:00
Nishanth Kotla 3198991d23 benchmarking/locust: record cluster hardware facts and density frontiers (#1723)
Part of #1590

First of two PRs. This one covers Kubernetes-side hardware discovery;
the follow-up adds the Prometheus harvest.

## What this PR does

The locust runner records no information about the hardware it runs on,
so trial results cannot be normalized across machine types or node
counts. This adds cluster capacity discovery and derives actor density
frontiers from it, persisting both the raw readings and the derived
ratios to `stats.jsonl`.

## Proposed Changes

### Cluster hardware discovery

New module `benchmarking/locust/cluster_facts.py`, kept separate from
`runner.py`. It reads node count, machine type, allocatable cores and
allocatable RAM, and the worker pod count via the official Kubernetes
Python client, which resolves in-cluster and local kubeconfig auth
without a kubectl subprocess.

Node and pod listing is proportional to cluster size, so
`--no-cluster-facts` disables discovery. Discovery is best effort in
either case: an unreachable API server or missing RBAC leaves the facts
null and does not fail the run.

No value is defaulted or inferred. Unmeasured fields are `null`; a
measured zero is `0`. The two remain distinguishable to any consumer.

### Density frontiers

Added to the `trial_summary` row in `stats.jsonl`:

| Field | Meaning |
|---|---|
| `actors_per_node` | peak actors over node count |
| `actors_per_vcpu` | peak actors over allocatable cores |
| `actors_per_gb_ram` | peak actors over allocatable RAM |
| `actors_per_pod_p50` / `_p90` / `_p99` | distribution of actors per
worker pod |
| `aggregate_failure_ratio` | failures over requests, all RPCs |
| `<operation>_failure_ratio` | one per operation Locust reported, e.g.
`dur_dir_write_failure_ratio` |

Actors per pod is reported as a distribution across the run's time
samples (`p50` / `p90` / `p99`) rather than one average, because ramp-up
and custom load shapes have no single user count to call steady.

The raw readings (`machine_type`, `node_count`, `allocatable_cores`,
`allocatable_ram_gb`, `worker_pod_count`) are written alongside the
derived ratios so they can be re-derived without re-running the trial.

### RBAC

Nodes are cluster scoped and require a ClusterRole with `list` on
`nodes`. Pods require only a namespaced Role with `list` in
`benchmark-workloads`. `locust.yaml` and `runner-job.yaml.tmpl` each
define their own separately named pair.

Note that `locust.yaml` now binds a Role in `benchmark-workloads`, so
that namespace must exist before the manifest is applied.
`benchmarking/workloads/deploy.sh` creates it.

`status.json` is unchanged.

## How this was tested

11 unit tests in `test_cluster_facts.py` covering percentile boundaries,
capacity scoped to only the nodes running worker pods, zero worker pods
treated as a valid reading, unreadable facts returning null, and flag
behavior.

Verified against a live GKE cluster. Discovery returned
`c3d-standard-8`, 1 node, 7.91 allocatable cores, 27.73 GB, 5 worker
pods, all matching `kubectl`. The emitted frontiers were re-derived by
hand from `stats_history.csv` and matched. `--no-cluster-facts` returns
all nulls.

Both manifests pass `kubectl apply --dry-run=client`, with no
cluster-scoped name collisions between them.

## References

[Agent Substrate: Actor Density Benchmark
Specs](https://docs.google.com/document/d/1sv5aBXvGOQ69iaqFxdPZ-d6tNbh5AMuwKODDyYjzYQw/edit?tab=t.0)

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-24 20:55:29 +00:00
Julian Gutierrez Oschmann 21b0b030bc Generate the benchmark Python gRPC clients at image build time (#1825)
protoc's Python output puts the serialized descriptor on a single line
which makes any two concurrent ateapi.proto changes conflict.

This is unfortunate as both proto and go generated code usually merges
with no conflict.

Stop checking in the *_pb2*.py files. The locust and nighthawk-ingress
images now generate them in a build stage via
benchmarking/locust/codegen/generate.sh.

Fixes #1824
2026-09-23 19:45:51 +00:00
Julian Gutierrez Oschmann 985c200247 Add DeletePreconditions to all Delete<Resource> methods (#1807)
Allow clients to do optimistic locking on deletes, as well as prevent
unintended deletes on name reuse.

#866
2026-09-23 18:55:59 +00:00
shrutiyam-glitch 47b67574ac docs: Document RevertActor and drop "terminal" from CRASHED (#1711)
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.

Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.

Follow up for the PR - #1675 
Issue - #1556 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 17:51:56 +00:00
Taahir Ahmed 8ea4abe1a3 ateom/actor ID split: Define MintAteomActorCertificate (#1626)
We need to draw a strong distinction between credentials that will be
wielded directly by actor logic, and credentials used by system
components on behalf of an actor. This will help us avoid attacks where
an actor pretends to be a system component.
    
Ateom certificates are requested by ateom/atelet and used for the
atunnel connection to the egress gateway. They have a SPIFFE URI like
`spiffe://${trustdomain}/ateom-for-actor/${atespace}/${actor}`.
    
Actor certificates are requested by the egress gateway and used for
opening outbound requests on behalf of the actor. They have a SPIFFE URI
like `spiffe://${trustdomain}/actor/${atespace}/${actor}`. Note, this
use case is not actually implemented yet.
2026-09-23 17:17:35 +00:00
Luiz Oliveira fb4b31529a Rename SnapshotsConfig to SnapshotConfig everywhere in the codebase (#1791)
related to -> #1378

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 14:39:09 +00:00
Luiz Oliveira ee9aa9c247 Add DV for TagStatus + drop unused TagStatus.source_actor_uid field (#1772)
TagStatus.source_actor_uid is unused, we can add this back if we ever
need it.

#1168 
 
We were also missing some of the DV for TagStatus fields:

* Added maxLength=2048 requirement for ExternalSnapshot.snapshot_uri
(matching in_progress_snapshot_uri)
* Added validation for actor_template_uid
* Added validation for storage_location, matching what we have in
SnapshotsConfig.storage_location
2026-09-23 13:40:41 +00:00
Julian Gutierrez Oschmann 4ab35e1d6e Delete unused messages. (#1814)
These are not used any longer now that `ActorSnapshot` resource is gone
and the `ActorSnapshotTag` was renamed to `Tag`
2026-09-23 02:21:14 +00:00
Julian Gutierrez Oschmann d277088bc1 Rename ActorTemplate.containers.readyz to wakeupProbe. (#1794)
The name `readyz` is not very descriptive. We also want to avoid
`readinessProbe` to prevent confusion with the Kubernetes concept, which
represents continuous traffic gating. Renaming this field to
`wakeupProbe` clarifies its actual behavior, and the fact that it only
operates when the actor is woken up.

Part of #1378
2026-09-22 23:32:01 +00:00
Da Huang 25cdbe00bc proto: mark secret-bearing fields with debug_redact (#1803)
Label the four fields that carry credentials or user-supplied secrets
with the standard protobuf debug_redact option:

  - ateapi        EnvVar.value                      actor env var values
  - ateapi        MintActorJWTResponse.actor_jwt    a bearer token
  - atelet        EnvEntry.value                    actor env var values
  - credprovider  FetchSecretResponse.opaque_bytes  a fetched credential

The option is descriptor metadata only. It does not change the wire
format, the generated Go or Python API, or any runtime behaviour on its
own; Go's protobuf library does not act on it. It exists so that log
redaction can find sensitive fields by asking the schema instead of
matching field names, and so that new sensitive fields are labelled in
the same line that defines them. The interceptor that logs RPC bodies
today still clears only fields named "env"; making it honour this label
is a separate change.

Generated files are refreshed with hack/update/codegen.sh; the only
generated changes are the embedded descriptor bytes.

First step to fix #1743

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-22 21:40:30 +00:00
pmandewalkar cdac9baef8 benchmarking/locust: Add sweperf trajectory-based benchmarking workload (#1694)
Fixes #1692 

Adds a SWE-Perf benchmarking workload to the benchmarking suite,
exercising
actor suspend/resume against a real SWE-bench task image
(`astropy-7336`)
rather than a synthetic load.

The workload runs a recorded agent trajectory in configurable cycles,
suspending and resuming the actor between each, so the benchmark
measures
lifecycle cost under a realistic in-sandbox server.

### What's here

| Area | Change |
|---|---|
| `internal/benchmarking/boomer/sweperf/` | New `SweperfUser` boomer
user class (+ tests) |
| `internal/benchmarking/boomer/dynconfig/` | `sweperf_template`,
`sweperf_total_steps`, `sweperf_num_cycles` knobs |
| `benchmarking/locust/common/sweperf_config.py` | `--sweperf-*` Locust
flags |
| `benchmarking/locust/tests/sweperf.py` | Stub user class so the master
attributes boomer's stats rows |
| `benchmarking/workloads/manifests/` | `swebench-astropy-7336`
ActorTemplate |

Defaults: 21 trace steps partitioned into 4 cycles.

### Verification

- `hack/verify-all.sh` — pass
- `go test -race ./internal/benchmarking/... ./cmd/benchmarking/...` —
pass
- End-to-end on a GKE dev cluster (gVisor), 1 user:

```
state: running | users: 1 | fail_ratio: 0.0

NAME                     REQ  FAIL   AVG_ms
CreateActor                1     0        1
CreateAtespace             1     0        1
ResumeActor               20     0      495
SuspendActor              20     0      947
Workload_Cycle_1           5     0     1996
Workload_Cycle_2           5     0      964
Workload_Cycle_3           5     0     4877
Workload_Cycle_4           5     0     3022
```

164 requests total, 0 failures, 0 errors.

The ActorTemplate image is currently pinned to a personal Artifact
Registry
repo; a shared public registry for these SWE-Perf task images is
planned.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-22 01:08:59 +00:00
shrutiyam-glitch 944abe3278 feat(api): introduce RevertActor lifecycle RPC (#1675)
#### Summary
This PR introduces the `RevertActor` RPC for actor lifecycle management.
It includes the API definition, the corresponding workflow execution
logic, observability metrics, and updates to the authorization model to
support reverting actors.

Fixes #1556
Docs updated in #1711 

#### Commit-wise Changes

**1. Add RevertActor RPC (`1145d99`)**
* Introduces the new `RevertActor` RPC to the API definitions.
* Updates the corresponding protobuf bindings (affecting files like
`ateapi.pb.go` and `ateapi_pb2.py`).

**2. Add the revertActor workflow and observability metrics
(`dcb1fae`)**
* Implements the core `revertActor` workflow logic, designed to be
idempotent and re-enterable. It progresses through the following steps:
* **Mark Reverting:** Validates that the actor is in a revertable state
(`RUNNING`, `PAUSED`, or `CRASHED`) and transitions its state to
`REVERTING`.
* **Discard Worker:** Safely tears down the execution environment by
terminating the workload, detaching volumes, and releasing the assigned
worker.
* **Collect In-Progress Snapshot:** Cleans up external object storage by
deleting any objects a previous suspend operation was partway through
writing.
* **Finalize:** Commits the actor to `SUSPENDED` and strips all
node-local and in-progress state pointers (clearing `WorkerAssignment`,
`LocalSnapshotInfo`, etc.), returning the actor to its untouched
external snapshot.
* Instruments the workflow with lifecycle operation metrics (e.g.,
updating `ate.actor.lifecycle.operation.duration` to track `revert`
operations).
* *Note/TODO:* Currently, when reverting a paused actor, the workflow
drops the pointer to the node-local state but does *not* actually prune
the local checkpoint bytes from the node (this is tracked in #641).

**3. Add `can_revert` to the authorization model (`2fc53ae`)**
* Adds the `can_revert` permission to the auth model, mirroring the
shape of `can_suspend` (editor tier of the parent atespace, plus a
direct grant so a machine identity can revert the actor it drives
without holding an atespace role).

**4. Serve RevertActor and add the CLI verb (`f04fab4`)**
* Wires the `Control.RevertActor` service method to the workflow
(replacing the generated stub that previously answered `Unimplemented`).
* Adds the `"ate revert actor"` CLI command, making the feature usable
end-to-end.
* Implements `Terminate` for the fake atelet. This was necessary because
reverting an actor from the `RUNNING` state is the first path to reach
this call in functional tests (previously, delete tests skipped this
step as they ran against actors with no worker assignment).

**5. Add a manual verify script for RevertActor (`8022db3`)**
* Adds a script to manually exercise `RevertActor` against a real
control plane, since unit and functional tests only run against a fake
atelet.
* Tests reverting from `CRASHED`, `RUNNING`, and `PAUSED` states, and
verifies that attempting to revert a `SUSPENDED` actor is properly
rejected.
* Simulates a crash by deleting the worker pod the actor runs on to
verify the workflow can handle the absence of a worker to terminate.


- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-18 20:18:17 +00:00
Julian Gutierrez Oschmann af2477e45d When connecting to Atelet, always resolve the atelet IP from the node name (#1742)
Dialing by resolving the worker pod is problematic, because in some
cases (e.g. during Worker Pod deletion), we need to dial to atelet (to
properly call `Terminate`), but the worker is gone.

A much simpler model is to always dial by resolving the IP of the atelet
on a given node.
2026-09-18 16:07:11 +00:00
Max Smythe b25c23e4ad Glutton Benchmarking Enhancements (#1631)
Benchmarking: multi-actor glutton VUs, live window, resume retries,
client-side latency
    
    Glutton actors per VU. Each VU creates --actors-per-user actors on
startup and cycles through them round-robin, so one goroutine can drive
many mostly-idle actors. runner.py forwards the flag to boomer-glutton.
A crashed actor stays crashed for the run (ateapi never rehabilitates
it), so it is marked on the first Aborted "crashed" error, skipped from
    then on, and counted in a CrashCount stat.
    
Live window. --min/--max-live-time (default 0-0) set how long an actor
    stays resumed between its first ping and the suspend. Up to
--max-pings-per-wake pings (default 1) run inside it, spaced 0.2-1.0s
apart. The wait window (--min/--max-wait-time) remains the gap between
    one actor's suspend and the VU's next resume, and every return path
    sleeps it, so a failing startUser, resume, or crashed actor does not
    spin on boomer's zero-delay re-entry. With the defaults the cycle is
    resume, ping, suspend, wait: the same shape as before.
    
    Resume retries. ResumeActor retries ateapi's transient "concurrent
update conflict" Aborted up to five times with a 50ms backoff, inside
    the timed call, so the conflict no longer shows up as a failure.
    
    Client-side latency. Every gRPC row in the locust stats now reports
client wall clock, which covers retries, queueing, and the network. The
    server's elapsed-time trailer stays on the trace span only.
    
    Worker robustness. The router HTTP client keeps up to 10000 idle
connections per host so each VU reuses its connection across wakes. A
    failed dynconfig fetch after the first successful one keeps the last
    fetched values instead of exiting the worker.
    
    Orchestrator. util.run logs the duration of each shell command.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-17 08:52:54 -07:00
Luiz Oliveira fd99858fb8 Store the in-progress snapshot as a URI instead of a name (#1609)
To build a snapshot URI, we need the external storage location (which is
stored in the actor template), actor atespace and UID (which are stored
in the actor resource). If an actorTemplate was deleted, we'd leak the
in-progress snapshot, because the storage information was gone.

Fixes #1608

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-16 15:28:58 -04:00
Eitan Yarmush a58481a18e Publish ActorTemplate golden snapshots as tags (#1523)
Fixes #1507

Golden snapshots currently remain owned by the temporary golden actor,
so another resume/suspend cycle or actor deletion can collect a snapshot
still referenced by its template. The controller now copies the warmed
snapshot into a published tag, deletes the golden actor, and records the
tag reference on the template. Interrupted tag creation and cleanup
remain retryable; template deletion cleans up both resources.

`CreateActor` resolves an explicit `sourceTag` or the template's golden
tag into the actor's initial snapshot. Actors created before the golden
tag is ready retain their cold-boot behavior. The golden tag uses the
template UID as its name in `ate-golden`. The proto replaces
`golden_snapshot` with `golden_tag` at field 1, without backward
compatibility.

This PR is based directly on `main` and does not depend on #1521.

Follow-up recommendation: move the create → resume → wait → suspend →
tag → delete sequence into a golden-template workflow using the existing
workflow conventions. The reconciler now coordinates multiple
recoverable steps; it could retain scheduling and retries while
delegating that sequence to the workflow. This refactor is outside this
PR.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Validation: full `env -u NO_COLOR make verify` passed after rebasing
onto `main`. After the final proto field-number change, bindings were
regenerated and the control API unit/functional tests plus proto-format
and Go-format checks passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-16 09:24:42 -04:00
Michelle Au 6126a2b093 Add pause lifecycle mode support to benchmarking harness (#1645)
Add support for comparing PauseActor vs SuspendActor performance in the
benchmarking suite. Specifically:
- Add LifecycleMode ("suspend" vs "pause") to boomer dynconfig and
Python Locust flag registration (--lifecycle-mode).
- Update GluttonUser and DurdirUser to hibernate actors using either
PauseActor or SuspendActor based on lifecycle_mode, and ensure
DeleteActor passes AnyState=true during teardown so paused actors are
cleanly deleted without precondition errors.
- Add unit test coverage for pause vs suspend execution and teardown.
- Add pause benchmark configurations to automation/tests.yaml for
baseline, memory scaling, concurrency, and DurableDir scenarios.
- Add PauseActor and comparison latency/QPS panels to monitoring
dashboards in monitoring.yaml.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-14 16:45:17 -07:00
Zoe Zhao a536fabe22 Remove the top level --boot flag from ResumeActorRequest (#1548)
Fixes https://github.com/agent-substrate/substrate/issues/1566

Today the Resume workflow resolves its restore source in the following
order
1. First check if the actor has node-local snapshot, 
2. then its own durable external snapshot, 
3. then the template's golden snapshot. 
 
The boot flag was consulted at exactly one point in that chain, where it
suppressed using the golden-snapshot, which made its behavior much
narrower than "boot from scratch" suggests:

- Actor has its own external snapshot and boot=true: flag ignored,
restores the actor's snapshot.
- Actor has a local snapshot and boot=true: flag ignored, restores the
local snapshot.
- Actor has no snapshot, template has no golden snapshot: cold boot from
the spec regardless of the boot flag.
- Actor has no snapshot, template has a golden snapshot: boot=false
restores the golden, boot=true cold boots from the spec. This is the
only case where the flag is used.


The glutton benchmark was the only caller that set boot=true, on each
actor's first resume, to report true cold-start latency as a separate
ResumeActorColdStart stats row. @maxsmythe let me know if this is
required.

The proto field number and name were reserved.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-11 17:55:09 -04:00
haiyanmeng 5fb2c0a3cd ateapi: rename IPBlockRule to CIDRRule (#1597) 2026-09-11 13:49:51 -07:00
Taahir Ahmed dc675e8edd Identity: Move actor JWT/cert minting into the main API (#1315)
This was originally a separate service to make it easy to apply separate
authentication and authorization interceptors.

It now seems clear that our authn/z framework will be strong enough to
support atelet and external callers in one system (based on OpenFGA).

This change moves the MintJWT and MintCert RPCs into the control API,
and removes some inline authz checks that will be handled by our unified
authorizer framework.
2026-09-11 10:23:36 -07:00
Keith Mattix II 6bd89588dc Move to lowercase for header references
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Keith Mattix II f16fc04fa0 Move from Host header to explicit headers for actor and atespace
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Eitan Yarmush ee8d8faf09 Use Tag UIDs in snapshot storage paths (#1521)
Fixes #1508

Store tag snapshots at `<base>/atespaces/<atespace>/tags/<tag-uid>`.
Replace `in_progress_snapshot_uri` with immutable `storage_location`, so
pending and completed tags share UID-based cleanup independent of the
source actor or template.

- [x] Tests pass: race-enabled control API tests and PostgreSQL tag
contract tests.
- [x] Documentation updated.

Lint and code-generation verification also passed.

---------

Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-09-08 18:27:12 -04:00
Luiz Oliveira 1e56e66b25 Rename ActorSnapshotTag to Tag (#1491)
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.

https://github.com/agent-substrate/substrate/issues/664

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 18:00:36 -04:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Haven Xia ac175376e4 Refactor the Python proto codegen into update/codegen.sh (#1483)
Fixes #1481

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 10:50:20 -07:00
Tim Hockin a9c84dc2ad API: Remove one layer of nesting in Worker status
Worker.Status.allocation.{capacity,allocated} -> Worker.status.{capacity,allocated}
2026-09-04 09:58:49 -07:00
Benjamin Elder 2c429a9906 multi-actor worker API (#1283)
Part of #1266 

This is a draft of the core API + data store changes.

It's still a large PR, apologies.

The "as rows" commit could be split out, but this takes it to ~all of
the breaking changes we can't hide behind updating internals.

Same for the claimlock, but in both cases it seems these are worth
understanding when considering the API shape.

They're loadbearing for performance once we actually have multi-actor
workers.
2026-09-03 21:19:21 -07:00
Zoe Zhao cad4ce2b7a ateapi: UpdateActor allows updating ActorTemplate. (#1365)
Fixes https://github.com/agent-substrate/substrate/issues/477 Implements
the following actor template update flow:

```
UpdateActor(): set .template = templ-v2

ResumeActor()
    resumes using templ-v2
    wait for readyz
    if fail: return failure
        .template = templ-v2; .status.current_template = templ-v1
        (the next resume will attempt templ-v2 again)
    set .status = STATUS_RUNNING
    set .status.current_template = templ-v2
    return success
    .template = templ-v2; .status.current_template = templ-v2
```
2026-09-02 16:41:32 -07:00
pmandewalkar cedb0147ff benchmarking: Refactor boomer based Locust benchmarking (#1295)
Fixes #1248 

Renamed boomer-glutton to boomer-worker.
Extracted the boomer shared utils that future non-glutton workloads may
use to the boomerutil package.
Updated durdir and glutton to import and use the changes.
- [ X] Tests pass
- [ X] Appropriate changes to documentation are included in the PR
2026-09-02 13:12:45 -07:00
Julian Gutierrez Oschmann 27b7b5904f Remove union discriminator field from ActorTemplate.volumes. (#1398)
Partially addresses #1397. Remove union discriminator `type` field from
`ActorTemplate.volumes`.
2026-09-02 09:29:19 -07:00
Sairaj Pokale 528d905b65 benchmarking: follow go.mod's Go version in the benchmarking images (#1352)
Fixes #1351

`Bump Go to 1.27` (69828945) raised the root `go.mod` and every
`hack/tools/*/go.mod` to `go 1.27.0`. `benchmarking/deploy_locust.sh
--deploy`
now fails at the `boomer-glutton` build:

go: go.mod requires go >= 1.27.0 (running go 1.26.7; GOTOOLCHAIN=local)

The `goboomer` stage builds `FROM golang:1.26-bookworm`, and the
official
`golang` images set `GOTOOLCHAIN=local` in their image config. That
disables
Go's toolchain resolution, so the base image tag becomes a second place
where
the project's Go version is declared. It drifted out of sync with
`go.mod` at
the 1.27 bump and would drift again at 1.28.

Simply bumping the tag would fix today's build and leave that second
declaration in place. This makes `go.mod` the only place the version is
declared instead.

## Changes

- **`benchmarking/locust/Dockerfile`**: set `GOTOOLCHAIN=auto` in the
`goboomer` stage, undoing the base image's `local`. Go now reads the
`go`
directive from `go.mod` and fetches a matching toolchain when the base
image
  falls behind, so a future minor bump needs no change here.


## Validation

No test covers this file. Verified by build.

- [x] `docker build --no-cache --platform linux/amd64 -f
benchmarking/locust/Dockerfile .`
compiles `boomer-glutton`, logging `go: downloading go1.27.0` where it
      previously failed.
- [x] Same build with `GOTOOLCHAIN` left at the image default still
fails with
the error above, confirming that variable is the cause rather than the
      base image version.
2026-09-02 03:01:52 -07:00
Lucky Abolorunke 811fac3e7b benchmarking: walk RAM after resume (--mem-read) and rotate the churn window (#1310)
#### What pr does

Makes resume measurements require the actor's memory to actually work:
adds `ReadRAM` — a glutton request that walks the working set (reads one
byte per 4KiB page across the requested size) before responding, plus a
`--mem-read` knob so the benchmark cycle performs that walk right after
every resume.

Today's cycle proves an actor is *reachable* after resume, not that its
memory is *usable*: the ping answers without touching the working set. A
real application must read its memory to serve requests. This matters
for where restore optimization is headed — a lazy/on-demand restore
would look great on a benchmark that never reads memory (resume returns
fast, ping returns fast) while real first-requests would stall faulting
pages back in. With the walk in the cycle, "resume + first response"
includes the cost of making memory usable, however the restore path
schedules that work: eager restore pays it during resume, lazy restore
would pay it during the walk — either way the total is in the tracked
numbers.

Also upgrades churn with `WRITE_MODE_OVERWRITE_ROTATE`: overwrite at a
per-key cursor that advances past each write and wraps, so repeated
churn walks the whole array over time instead of re-dirtying the same
prefix every cycle.

#### How it works

- `ReadRAM(key, size)` walks the first `size` bytes (suffixed string,
e.g. `"1Gi"`; empty walks the whole array) of a `WriteRAM` allocation,
one byte per 4KiB page — the cheapest touch that forces every page
resident. The response returns bytes walked plus an XOR checksum of the
sampled bytes so the reads are observable and can't be elided.
- Cycle order: resume → fill (once) → **walk** → churn → ping → suspend.
The walk runs *before* churn deliberately: it must read the memory as
restored, not pages churn just rewrote; churn then re-dirties after, so
the next snapshot still carries fresh pages.
- The walk reports as its own `GluttonReadRAM` stats row — it never
pollutes ping or resume latencies. Today (eager restore) it reads warm
memory in milliseconds; a jump in this row is the signal that restore
work got deferred onto the request path.
- Config travels the established channel: `--mem-read` in the suite's
locust `flags:` → `/boomer-config` → the Go worker, passed verbatim to
the wire; glutton is the only parser. Empty = disabled; the tracked
large-memory suites set it to the full target (walk everything —
strongest signal, simplest story). Existing suites unchanged.

#### Testing

- `go test -race` across `cmd/benchmarking/glutton` and
`internal/benchmarking/boomer/...`: PASS. New tests cover the walk's
byte count and checksum, missing-key/bad-size errors, cycle call order
(fill → read → churn with the right sizes and modes), rotate-mode cursor
wrap, disabled-by-default, and walk-before-fill as a no-op.
- Cluster verification (microvm, 1Gi target, full walk):
2026-09-02 02:27:56 -07:00
Zoe Zhao f6852b7754 Update existing demos and benchmark tests to use substrate ActorTemplate resource (#1355)
This PR is very large since it updates all existing demos and benchmark
workloads to use the new ActorTemplate substrate proto.
Please use the "Commits" tab to review individual commits.

Verifications done: 
* Used this script: gpaste/5143788763348992 to verify that the change
from CRD -> proto are equivalent.
* The e2e tests are using the new susbtrate resources.
* Picked the parking demo to run e2e manually: gpaste/6193361380311040
2026-09-01 14:50:18 -07:00
Joe Betz 3efac91283 api: delete the DebugClear RPC and the Debug service (#1346)
Fixes #999.
2026-09-01 17:46:05 -04:00
Zoe Zhao fe9013a4a2 Full cutover: Drop the k8s CRD ActorTemplate fields in the ate apiserver, and update e2e tests (#1353)
This PR is large, I grouped changes to the following commits:

- 1d5d8f06: moves consumers off the CRD path: demos, ate-setup scripts
and the e2e suites address templates by the actor_template ref.
- e502cf60: Ate API changes: removes actor_template_namespace,
actor_template_name from Actor, ActorAssignment and ActorSnapshot in the
public API, drops the CRD conversion fallback.
- 8f6c1adb: atecontroller: removes the ActorTemplate CRD controller.

Once this PR is submitted, existing demos that still uses CRD will stop
working. Created https://github.com/agent-substrate/substrate/pull/1355
to update existing demos.
2026-09-01 14:02:59 -04:00
Zoe Zhao e1adb33147 api: drop ActorTemplateStatus.sandbox_assets for now (#1300)
Remove `ActorTemplateStatus.sandbox_assets` along with the
`SandboxAssets` messages it referenced. Nothing ever wrote the field:
sandbox assets are resolved from the WorkerPool and SandboxConfig
objects at resume time and travel to atelet via ateletpb, so freezing
them into the template status never materialized.

This requires more thought, one option is to add it as a field of
GoldenSnapshotStatus in the future, and make GoldenSnapshotStatus a
repeated field to allow multiple goldens.
2026-08-29 10:11:49 -07:00
Tim Hockin cfeaf23f9e Repeated fields should be named plural 2026-08-28 22:08:45 -07:00
Eitan Yarmush 01e92cca59 Add Actor egress policy API
Signed-off-by: Eitan Yarmush <eitan.yarmush@solo.io>
2026-08-28 13:36:23 -07:00
Lucky Abolorunke 1c431b1348 benchmarking: re-dirty the glutton working set each cycle (--mem-churn) (#1280)
#### What this pr do

Re-dirties part of the glutton working set on every benchmark cycle, so
repeated suspends snapshot an actor whose memory is changing — like a
live application's — instead of a set that was filled once and never
touched again.

Each iteration (after the one-time fill), the GluttonUser sends one
`WriteRAM` request with `WRITE_MODE_OVERWRITE`, re-randomizing the first
`mem_churn` bytes of the working set in place. Overwrite mode already
existed on the glutton API; this PR adds no proto changes and no glutton
server changes — it is driver + config wiring only.

This is the follow-up from #1130's: the original self-driving memload
continuously re-dirtied its pages, and that property was dropped in the
move to API-driven fill. This restores it in the API-driven shape with a
dialable amount instead of the old all-or-nothing full pass.

#### Why it matters

A fill-once working set is static: every suspend after the first packs
up identical memory. If snapshotting ever gains incremental / dirty-page
optimizations, a static benchmark would measure almost nothing from
cycle two onward — and would score "upload nothing" as an infinite win.
With churn, every cycle carries a known amount of freshly-dirtied pages,
and the knob is sweepable (64Mi, 256Mi, …) so a future
differential-snapshot optimization can be demonstrated as "upload cost
scales with churn size, not total size."
2026-08-28 11:28:51 -07:00
Benjamin Elder 69428103ee verify python codegen (#1278)
That's twice today that we missed regenerating the proto clients:

1. mid-PR on https://github.com/agent-substrate/substrate/pull/1276
(caught by code review)
2. https://github.com/agent-substrate/substrate/pull/1130 (missed and
still missing)

So this PR does:

1. regenerate with latest after #1130
2. ensure verify will catch this (so CI should fail if we haven't run
it), and that update will run it when developing

fixes https://github.com/agent-substrate/substrate/issues/738

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-27 22:25:26 -04:00
Lucky Abolorunke 9cfe3dd690 benchmarking(memory usage phase 1 baseline): large-memory glutton workloads (--mem-targetLarge mem bench (#1130)
## What this PR does:

Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.

The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).

## How it's configured

The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.

Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom

## Behavior notes

- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.

## Testing

- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
2026-08-27 16:40:26 -07:00
Benjamin Elder d7ee171ec6 ateapi: document the container security context and image volume fields (#1276)
The Capabilities, SecurityContext, ImageVolumeSource and Volume.image
fields landed without doc comments, so apitool's documented rule fails
on them and they are not in the exemption backlog. Document them rather
than growing the backlog.

Fixes failure on main

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-27 18:23:28 -04:00
Zoe Zhao d8c4f8a4b9 Add controller in ate apiserver to reconcile ActorTemplate state (#1096)
Part of #477.

The controller has 2 components
1. A "producer" that periodically lists ActorTemplates and add work
items to a client-go workqueue.
2. A "consumer" that gets work items from the queue, and reconcile it.

AT will have the FAILED condition if we encounter non-retriable errors
during the transition
2026-08-26 16:42:05 -07:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Luiz Oliveira 31a2e3850a Simplify Update methods to do a whole object replace (#1108)
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.

#1011 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-24 15:18:55 -04:00