Commit Graph
19 Commits
Author SHA1 Message Date
Luiz Oliveira 31a2e3850a Simplify Update methods to do a whole object replace (#1108)
* Removed field_mask from the API
* Added a new protoupdate package to handle replacing mutable fields.
This makes sure that unknown fields in the server are not dropped by an
update from a stale/old client.

#1011 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-24 15:18:55 -04:00
Luiz Oliveira fc02127bfe ateapi - Require preconditions in Update methods (#1014)
https://github.com/agent-substrate/substrate/issues/954

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-19 21:01:29 -04:00
Luiz Oliveira 71a58cbf37 split controlapi functional tests into their own package (#918)
functional_test.go was also split into one file per resource, to match
what we're doing for the RPC handlers too. See #891

I also needed to make some changes to move the tests to a separate
package:

* `NewAteletDialer` now takes `options`, and `WithDialCredentials` lets
a test build its own transport credentials. The fake atelet is reached
over insecure transport. Didn't change existing callers. This is the
only non-test change.
* Moved a few unit tests that are actually functional tests to the
respective file under functionaltest/

Fixes #891


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-18 19:06:07 -04:00
Luiz Oliveira 73aa9824c2 Split store contract tests into a function per resource (#955)
RunContractTests is getting large and resource tests are not always
grouped together. So, let's start with splitting them into separate
functions. As they grow, we may consider moving them to their own files,
like we did for functional tests in
https://github.com/agent-substrate/substrate/pull/918

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-14 15:13:35 -04:00
Luiz Oliveira 9eeb1c2482 Fix actor snapshot test to use the new CreateActorSnapshotTag (#937)
https://github.com/agent-substrate/substrate/pull/914 got merged in
between the time I sent
https://github.com/agent-substrate/substrate/pull/898 for review and its
submission.

#898 wasn't rebased, so tests broke.
2026-08-13 17:36:26 -04:00
Luiz Oliveira 5fbfd774e3 Add UNSPECIFIED zero value for ActorSnapshotTagScope (#898)
Fixes #897


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 16:50:21 -04:00
Luiz Oliveira a174f2880e Consolidate the actor controlapi handlers into one file (#913)
#891

No declaration added, removed, or renamed, import sets
unchanged, and every non-header line is identical to what it replaced.



- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 13:53:07 -04:00
Luiz Oliveira 953e38e03e Consolidate the atespace controlapi handlers into one file (#915)
No declaration added, removed, or renamed, import sets unchanged, and
every non-header line identical to what it replaced.

https://github.com/agent-substrate/substrate/issues/891

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 13:03:26 -04:00
Luiz Oliveira eda290e788 Consolidate the worker controlapi handlers into one file (#916)
Since we only have ListWorkers so far, this is just a file rename. The
rest of the handlers should be added to this file, instead of new
handler-specific files, as per #891

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-13 12:52:37 -04:00
Luiz Oliveira 7d80b3a3cc Only register demo-autoscaled-workerpool when running on kind clusters (#892)
This was breaking 

`hack/install-ate.sh --delete-all` with
`--delete-demo-autoscaled-workerpool is not supported on GKE`

Also removed the custom usage message for
--deploy-demo-autoscaled-workerpool. This was duplicate with the default
message:

Before:

```
Demo: demo-autoscaled-workerpool

  --deploy-demo-autoscaled-workerpool                         Deploy demo-autoscaled-workerpool
  --delete-demo-autoscaled-workerpool                         Delete demo-autoscaled-workerpool
  --deploy-demo-autoscaled-workerpool            Deploy autoscaled-workerpool demo (HPA + prometheus-adapter + counter workload)
```

After:

```
Demo: demo-autoscaled-workerpool

  --deploy-demo-autoscaled-workerpool                         Deploy demo-autoscaled-workerpool
  --delete-demo-autoscaled-workerpool                         Delete demo-autoscaled-workerpool
```

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-12 13:18:06 -04:00
Luiz Oliveira c538b68ba6 Fix TOCTOU race in UpdateActorSnapshotTag (#854)
This is a very similar fix to #829

We're also removing the precondition as the storage layer function
arguments and passing it as a wrapper around the closure functions.

https://github.com/agent-substrate/substrate/issues/763

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-12 10:00:22 -04:00
Luiz Oliveira 9b24617616 Fix toctou race in UpdateActor (#829)
UpdateActor now takes an ActorRef, an ActorPrecondition, and a mutate
callback. The store reads the stored actor inside the WATCH transaction,
checks the precondition against that value, and hands it to mutate,
which edits it in place.

The storage layer retries up to 5 times when a concurrent write
invalidates it, re-running mutate against the newer state

I still need to migrate the other callers of store.UpdateActor to avoid
calling GetActor outside of the closure function.

#763

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-10 17:20:34 -04:00
Luiz Oliveira aa9b7b82cd Make UpdateActorSnapshot follow substrate API guidelines (#758)
Fixes #732  

* `UpdateActorSnapshot` now carries the resource itself + `update_mask`
* `scope` is now applied via the update mask
* Moved `update_mask` to a separate file, so it can be reused by other
update RPCs.
* Added `uid` and `version` as optional guards. 


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-05 23:10:17 -04:00
Luiz Oliveira 2feda2ae44 Enable the Go race detector for the test make target and CI workflow (#765)
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-05 16:23:06 -07:00
Luiz Oliveira b11c6e07ff Make UpdateActor method to follow API guidelines (#742)
#732  

* `UpdateActorRequest` now carries the resource itself + `update_mask`
* `worker_selector` is now applied via the update mask (making it
possible to clear the worker selector)
* `update_mask` is validated against an allowlist of paths. Currently,
only `worker_selector` is allowed.
* Added `uid` and `version` as optional guards. 

I'll update UpdateActorSnapshotTag in a follow-up PR


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-05 16:28:59 -04:00
Luiz Oliveira 9490650566 Add autoscaled-workerpool demo using HPA and prometheus-adapter on Kind (#576)
This introduces a new demo (`demos/autoscaled-workerpool`) demonstrating
how to dynamically autoscale a WorkerPool using HPA +
`prometheus-adapter` on a local Kind cluster.

As mentioned in https://github.com/agent-substrate/substrate/issues/198,
this approach may be too slow for some use cases, but this is a good
first milestone.

Plus, there are some limitations with scaling workerpools that still
need to be addressed, for example:

- Scale-down can strand paused actors. When a worker pod is removed, a
local-snapshot paused actor keeps its node pin pointing at a node that
may no longer have (free) worker pods.
- Scale-up gives no locality guarantee: If there are multiple
local-snapshot paused actors that can't resume because no worker pods
are available on their node, upscaling will not unblock them: the added
capacity can land anywhere, so the pool can grow without ever placing a
worker where a stranded actor needs it.
- Scale-down picks victims blind to actor assignment (can kill busy
pods).

- [x] Tests pass: n/a
- [x] Appropriate changes to documentation are included in the PR
2026-07-29 14:48:22 -04:00
Luiz Oliveira 0dd73e1679 Add pod selector to the WorkerPool scale subresource (#547)
WorkerPool already served /scale, so `kubectl scale` worked, but a HPA
pointed at one did not: the subresource declared no labelSelectorPath,
and HPA needs a selector to find the pods whose metrics it is averaging
or to compute ambiguous selectors.

https://github.com/agent-substrate/substrate/issues/198

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-27 12:02:33 -04:00
Luiz Oliveira 3cc3b539f2 Rename nameentity to identity (#447)
nameentity seems to be a typo when renaming `id` to `name`
2026-07-17 12:58:08 -04:00
Luiz Oliveira c1bd202dab fix(scripts): correct jq JSON path selector in delete_demo_actors
Update the actor listing query in `hack/install-ate.sh` to extract `metadata.atespace` and `metadata.name` instead of the legacy root fields `atespace` and `actorId`.

`metadata` was introduced in https://github.com/agent-substrate/substrate/commit/af1f00522c63d24efed30ecb691676ce6fde847a
2026-07-15 14:23:05 -07:00