37 Commits
Author SHA1 Message Date
Luiz Oliveira e8cad55675 Reject repointing an actor that owns a snapshot to a different storage location (#2014)
This is needed because DeleteActor cleans up all snapshots under the
actor's location.

If we were to allow repointing an actor's snapshot location, we risked
leaking historical snapshots upon actor deletion.
2026-10-01 03:53:35 +00:00
Luiz Oliveira b4dc55d5fe Remove reserved fields from proto and update back-compat instruction in AGENTS.md (#1985)
https://github.com/agent-substrate/substrate/issues/1378
2026-10-01 01:57:12 +00:00
Luiz Oliveira e87c55fc2e Add the crash reason and timestamp to the ActorStatus (#1867)
Added a new field to actor status that has the crash reason (e.g., the
atelet response error) and the timestamp of the crash for better UX.
2026-09-25 15:33:08 +00:00
Luiz Oliveira d72edfbb3c Rename LocalSnapshotInfo to LocalSnapshot (#1856)
For consistency with ExternalSnapshot

#1378
2026-09-25 14:20:42 +00:00
Luiz Oliveira 34f3003326 Add DV for GoldenSnapshotStatus and unbounded strings (#1831)
Fixes  #1168

* Add DV for GoldenSnapshotStatus and truncate its error_message to fit
* Require ExternalSnapshot.actor_template_uid to be a UUID
* Bound ExternalVolumeTemplate.capacity to 32 characters
2026-09-24 16:59:12 +00:00
Luiz Oliveira fb4b31529a Rename SnapshotsConfig to SnapshotConfig everywhere in the codebase (#1791)
related to -> #1378

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 14:39:09 +00:00
Luiz Oliveira ee9aa9c247 Add DV for TagStatus + drop unused TagStatus.source_actor_uid field (#1772)
TagStatus.source_actor_uid is unused, we can add this back if we ever
need it.

#1168 
 
We were also missing some of the DV for TagStatus fields:

* Added maxLength=2048 requirement for ExternalSnapshot.snapshot_uri
(matching in_progress_snapshot_uri)
* Added validation for actor_template_uid
* Added validation for storage_location, matching what we have in
SnapshotsConfig.storage_location
2026-09-23 13:40:41 +00:00
Luiz Oliveira 43a5c56838 Remove reserved fields from protos (#1793)
https://github.com/agent-substrate/substrate/issues/1378

We don't have a backwards compatibility requirement yet, so let's clean
this up for now.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 13:36:42 +00:00
Luiz Oliveira 7d241bfe1d Split atepg.go into one file per resource/table (#1777)
Fixes #1776

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-21 21:04:00 +00:00
Luiz Oliveira 0ab30264bc Move DeleteTag orchestration to a workflow (#1720)
The tag deletion logic is getting too complex, so let's move it to a
workflow, following what we do for other resources.

#1510 #664

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-20 14:22:49 +00:00
Luiz Oliveira fd99858fb8 Store the in-progress snapshot as a URI instead of a name (#1609)
To build a snapshot URI, we need the external storage location (which is
stored in the actor template), actor atespace and UID (which are stored
in the actor resource). If an actorTemplate was deleted, we'd leak the
in-progress snapshot, because the storage information was gone.

Fixes #1608

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-16 15:28:58 -04:00
Luiz Oliveira a0d306ee6c Pass the resync interval in the golden tag functional test (#1673)
#1522 added a resyncInterval parameter to NewActorTemplateReconciler and
#1523 added this call site. They landed without seeing each other, so
`main` doesn't compile.

SEt the resyncInterval to 7s, to match what we have in other tests

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-16 10:36:35 -04:00
Luiz Oliveira cfaf72a17b Stop leaking in-progress snapshots after worker deletion (#1541)
Fixes #1539

If an actor is `CRASHED`, it can't be resumed/paused/suspended anymore,
it can only be deleted.

An actor deletion will clean up everything under the prefix of
`Actor.external_snapshot.snapshot_uri` and
`Actor.in_progress_snapshot_name` (the latter will be empty if the
underlying worker was deleted).

Before, we risked leaking a snapshot after a worker deletion because we
were clearing the `in_progress_snapshot_name` field. We don't need to
clear it up because the CRASHED state is terminal.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-09 09:13:34 -04:00
Luiz Oliveira 1e56e66b25 Rename ActorSnapshotTag to Tag (#1491)
Renames the proto message, its status message and scope enum, the five
RPCs and their request and response messages, the actor_snapshot_tag
request fields, and Actor.source_snapshot_tag to source_tag. The store
interface, its Postgres table, the object-storage prefix segment and the
kubectl-ate verbs follow, so nothing keeps the old spelling.

https://github.com/agent-substrate/substrate/issues/664

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 18:00:36 -04:00
Luiz Oliveira 6378b707e3 Call the generated validation fns for ActorSnapshotTag methods (#1488)
Plus, add DV for the other methods that were missing and remove obsolete
custom validation functions

https://github.com/agent-substrate/substrate/issues/664
https://github.com/agent-substrate/substrate/issues/1168



> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 17:06:05 -04:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Luiz Oliveira 801e065ee8 Delete local snapshots when an actor is terminated (#1246)
We were leaking snapshots after actor termination.

Note that this is a temporary fix: it only removes local snapshots in
one node. We should clean up the copies on any other
NodeVmsWithLocalSnapshots. This is fine *as of the day this was written*
because today NodeVmsWithLocalSnapshots has at most one item.

Related to https://github.com/agent-substrate/substrate/issues/668 and
#664


- [x] Tests pass
- [] Appropriate changes to documentation are included in the PR
2026-08-27 10:32:33 -07:00
Luiz Oliveira 790ee35229 Reject object updates containing unknown fields (#1174)
https://github.com/agent-substrate/substrate/issues/1011

An update containing unknown fields during a RMW indicates there's a
version skew between client and server.

So, the server enforces a policy to fail explicitly when it receives an
unknown proto field: so that it does not silently drop fields that the
user intended to set but the current server version cannot process

This also helps to avoid behaviour drift during rolling upgrades: During
rolling upgrades, different components (like the data plane and control
plane) can be on different versions. If an older server replica . If an
older server replica processes a request coming from a newer client, it
will reject it (it doesn't know how to validate/handle it) and it's the
client's responsibility to retry.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-25 08:21:25 -07:00
Luiz Oliveira 31a2e3850a Simplify Update methods to do a whole object replace (#1108)
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.

#1011 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-24 15:18:55 -04:00
Luiz Oliveira fc02127bfe ateapi - Require preconditions in Update methods (#1014)
https://github.com/agent-substrate/substrate/issues/954

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 21:01:29 -04:00
Luiz Oliveira 71a58cbf37 split controlapi functional tests into their own package (#918)
functional_test.go was also split into one file per resource, to match
what we're doing for the RPC handlers too. See #891

I also needed to make some changes to move the tests to a separate
package:

* `NewAteletDialer` now takes `options`, and `WithDialCredentials` lets
a test build its own transport credentials. The fake atelet is reached
over insecure transport. Didn't change existing callers. This is the
only non-test change.
* Moved a few unit tests that are actually functional tests to the
respective file under functionaltest/

Fixes #891


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-18 19:06:07 -04:00
Luiz Oliveira 73aa9824c2 Split store contract tests into a function per resource (#955)
RunContractTests is getting large and resource tests are not always
grouped together. So, let's start with splitting them into separate
functions. As they grow, we may consider moving them to their own files,
like we did for functional tests in
https://github.com/agent-substrate/substrate/pull/918

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-14 15:13:35 -04:00
Luiz Oliveira 9eeb1c2482 Fix actor snapshot test to use the new CreateActorSnapshotTag (#937)
https://github.com/agent-substrate/substrate/pull/914 got merged in
between the time I sent
https://github.com/agent-substrate/substrate/pull/898 for review and its
submission.

#898 wasn't rebased, so tests broke.
2026-08-13 17:36:26 -04:00
Luiz Oliveira 5fbfd774e3 Add UNSPECIFIED zero value for ActorSnapshotTagScope (#898)
Fixes #897


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 16:50:21 -04:00
Luiz Oliveira a174f2880e Consolidate the actor controlapi handlers into one file (#913)
#891

No declaration added, removed, or renamed, import sets
unchanged, and every non-header line is identical to what it replaced.



- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 13:53:07 -04:00
Luiz Oliveira 953e38e03e Consolidate the atespace controlapi handlers into one file (#915)
No declaration added, removed, or renamed, import sets unchanged, and
every non-header line identical to what it replaced.

https://github.com/agent-substrate/substrate/issues/891

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 13:03:26 -04:00
Luiz Oliveira eda290e788 Consolidate the worker controlapi handlers into one file (#916)
Since we only have ListWorkers so far, this is just a file rename. The
rest of the handlers should be added to this file, instead of new
handler-specific files, as per #891

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 12:52:37 -04:00
Luiz Oliveira 7d80b3a3cc Only register demo-autoscaled-workerpool when running on kind clusters (#892)
This was breaking 

`hack/install-ate.sh --delete-all` with
`--delete-demo-autoscaled-workerpool is not supported on GKE`

Also removed the custom usage message for
--deploy-demo-autoscaled-workerpool. This was duplicate with the default
message:

Before:

```
Demo: demo-autoscaled-workerpool

  --deploy-demo-autoscaled-workerpool                         Deploy demo-autoscaled-workerpool
  --delete-demo-autoscaled-workerpool                         Delete demo-autoscaled-workerpool
  --deploy-demo-autoscaled-workerpool            Deploy autoscaled-workerpool demo (HPA + prometheus-adapter + counter workload)
```

After:

```
Demo: demo-autoscaled-workerpool

  --deploy-demo-autoscaled-workerpool                         Deploy demo-autoscaled-workerpool
  --delete-demo-autoscaled-workerpool                         Delete demo-autoscaled-workerpool
```

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-12 13:18:06 -04:00
Luiz Oliveira c538b68ba6 Fix TOCTOU race in UpdateActorSnapshotTag (#854)
This is a very similar fix to #829

We're also removing the precondition as the storage layer function
arguments and passing it as a wrapper around the closure functions.

https://github.com/agent-substrate/substrate/issues/763

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-12 10:00:22 -04:00
Luiz Oliveira 9b24617616 Fix toctou race in UpdateActor (#829)
UpdateActor now takes an ActorRef, an ActorPrecondition, and a mutate
callback. The store reads the stored actor inside the WATCH transaction,
checks the precondition against that value, and hands it to mutate,
which edits it in place.

The storage layer retries up to 5 times when a concurrent write
invalidates it, re-running mutate against the newer state

I still need to migrate the other callers of store.UpdateActor to avoid
calling GetActor outside of the closure function.

#763

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-10 17:20:34 -04:00
Luiz Oliveira aa9b7b82cd Make UpdateActorSnapshot follow substrate API guidelines (#758)
Fixes #732  

* `UpdateActorSnapshot` now carries the resource itself + `update_mask`
* `scope` is now applied via the update mask
* Moved `update_mask` to a separate file, so it can be reused by other
update RPCs.
* Added `uid` and `version` as optional guards. 


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-05 23:10:17 -04:00
Luiz Oliveira 2feda2ae44 Enable the Go race detector for the test make target and CI workflow (#765)
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-05 16:23:06 -07:00
Luiz Oliveira b11c6e07ff Make UpdateActor method to follow API guidelines (#742)
#732  

* `UpdateActorRequest` now carries the resource itself + `update_mask`
* `worker_selector` is now applied via the update mask (making it
possible to clear the worker selector)
* `update_mask` is validated against an allowlist of paths. Currently,
only `worker_selector` is allowed.
* Added `uid` and `version` as optional guards. 

I'll update UpdateActorSnapshotTag in a follow-up PR


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-05 16:28:59 -04:00
Luiz Oliveira 9490650566 Add autoscaled-workerpool demo using HPA and prometheus-adapter on Kind (#576)
This introduces a new demo (`demos/autoscaled-workerpool`) demonstrating
how to dynamically autoscale a WorkerPool using HPA +
`prometheus-adapter` on a local Kind cluster.

As mentioned in https://github.com/agent-substrate/substrate/issues/198,
this approach may be too slow for some use cases, but this is a good
first milestone.

Plus, there are some limitations with scaling workerpools that still
need to be addressed, for example:

- Scale-down can strand paused actors. When a worker pod is removed, a
local-snapshot paused actor keeps its node pin pointing at a node that
may no longer have (free) worker pods.
- Scale-up gives no locality guarantee: If there are multiple
local-snapshot paused actors that can't resume because no worker pods
are available on their node, upscaling will not unblock them: the added
capacity can land anywhere, so the pool can grow without ever placing a
worker where a stranded actor needs it.
- Scale-down picks victims blind to actor assignment (can kill busy
pods).

- [x] Tests pass: n/a
- [x] Appropriate changes to documentation are included in the PR
2026-07-29 14:48:22 -04:00
Luiz Oliveira 0dd73e1679 Add pod selector to the WorkerPool scale subresource (#547)
WorkerPool already served /scale, so `kubectl scale` worked, but a HPA
pointed at one did not: the subresource declared no labelSelectorPath,
and HPA needs a selector to find the pods whose metrics it is averaging
or to compute ambiguous selectors.

https://github.com/agent-substrate/substrate/issues/198

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-27 12:02:33 -04:00
Luiz Oliveira 3cc3b539f2 Rename nameentity to identity (#447)
nameentity seems to be a typo when renaming `id` to `name`
2026-07-17 12:58:08 -04:00
Luiz Oliveira c1bd202dab fix(scripts): correct jq JSON path selector in delete_demo_actors
Update the actor listing query in `hack/install-ate.sh` to extract `metadata.atespace` and `metadata.name` instead of the legacy root fields `atespace` and `actorId`.

`metadata` was introduced in https://github.com/agent-substrate/substrate/commit/af1f00522c63d24efed30ecb691676ce6fde847a
2026-07-15 14:23:05 -07:00