Commit Graph
20 Commits
Author SHA1 Message Date
Julian Gutierrez Oschmann 2dc1fc6266 Introduce an ActorRef type to bundle actor atespace and name.
The pattern of passing a tuple of atespace and actor name together
is across the whole codebase. This simplifies both callers and callees
by leting them pass a single field that bundles both.

The ActorRef is the actor specific, typed,
in-process version of the ObjectRef we have in our gRPC API.
2026-07-28 12:33:23 -07:00
Julian Gutierrez Oschmann 3e6da2eb29 Extract scheduling logic into its own package. (#546)
The scheduling logic is all mixed up with the resume/suspend/pause
workflows. Move it into a dedicated package, behind an interface and add
tests.
2026-07-27 11:28:52 -07:00
Michelle Au 9300387fef Add basic per-actor external volume flow (#405)
Initial part of #232.

Extends the ActorTemplate API to add support for per-actor external
volumes and adds control plane and atelet hooks following the Actor
lifecycle.

Callouts to volume operations are abstracted with a Volume interface.
Right now, only a mock volume plugin for testing has been implemented,
but we will add CSI support next.

- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR - Not
going to update documentation until we add CSI support.
2026-07-24 14:40:26 -07:00
Haven Xia c7aa4a2559 Same crashing for workflow_pause.go when pod is gone (#508)
Should found it when writing #474..
And also correct an error in `workflow_resume.go`.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-23 17:37:34 -07:00
Zoe Zhao 4ecf3fbf6c Add CheckPrerequisite to WorkflowStep to validate state machine edges (#374)
Fixes #369

Add a CheckPrerequisite method to the WorkflowStep interface, called by
RunWorkflow after IsComplete returns false and before Execute, so that
each workflow validates its actor state-machine edge up front while
retried (reentrant) workflows still fast-forward past completed steps.
2026-07-20 19:21:34 -07:00
Eric Bishop 381b3ad52c fix: make atelet / ateom use actor UID for its internal bookkeeping (#438) 2026-07-17 12:32:23 -07:00
Mesut Oezdil fc1fe4d91c fix: skip empty node name in local snapshot info after pause finalization (#329)
When the worker pod is gone from the DB during pause finalization,
nodeName stays empty.
The code still wrote NodeVmsWithLocalSnapshots: []string{""}, so on the
next ResumeActor the scheduler looked for a worker with node name "",
found none, and returned "no free workers available" permanently.
Fix: only set NodeVmsWithLocalSnapshots when nodeName is non-empty.
2026-07-17 10:24:04 -07:00
Haven Xia 669965a735 Rename actor_id to actor_name in atelet/ateom protos.
Aligns the internal ateletpb/ateompb services with the API style guide's
actor naming (field tags unchanged, so the wire format is compatible).

internal APIs keep dedicated identity fields (not the public ObjectRef).
2026-07-10 12:54:16 -04:00
Haven Xia de5d1d93f9 Make worker assignment checks atespace and actor name, also updated locks. (#411)
Current we only check & lock on actor name.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-07-09 09:29:10 -04:00
Julian Gutierrez Oschmann af1f00522c Add ResourceMetadata to both actor and atespace resources.
Replace actor_id and atespace name with `metadata` field which
holds common fields.

This commit only changes the ate-apiserver to support this field, but
intentionally skips upstream / downstream systems (such as kubectl-ate,
atenet, atelet, etc) to keep the scope manageable. This will come in subsequent PRs.
2026-07-08 17:22:03 -07:00
Dmitry Berkovich b66b5c2007 refactor: remove redundant SnapshotType enum and use protobuf oneof d… (#370)
refactor: remove redundant SnapshotType enum and use protobuf oneof d…
2026-07-08 10:05:58 -07:00
Zoe Zhao 069ba0840a Detect CRASHED Actors during Checkpoint and Restore (#353)
Related to #292 and #119

Changes:

1. ResumeActor:  
1. During `ResumeActor` we'll put the actor in CRASHED state if the
Resume failed due to any errors related to the snapshot itself, which
includes:
      1. Missing or corrupted snapshot when downloading from GCS  
      2. Runsc command fails.  
2. All other types of errors will not put the Actor in CRASHED state.
2. SuspendActor or PauseActor  
1. During `SuspendActor` or `PauseActor` we'll put the actor in CRASHED
state if we cannot produce a valid snapshot, which includes:
      1. cannot connect to Ateom after retries,  
      2. runsc command fails.  
3. cannot upload the snapshot to GCS (during suspend) ; or cannot save
snapshot locally (during pause).
   2. All other types of errors will not put the Actor in CRASHED state.
2026-07-08 09:26:03 -07:00
Julian Gutierrez Oschmann dce688a224 Propagate (atespace/actor_id) to atelet.
This fixes an issue where atelet/ateom will clash the state
of actors with same ID within different atespaces.

Once substrate resources get UIDs, we should revisit the way
atelet / ateom does the bookkeeping and check in which cases it
should use UID instead of (atespace,actor_id).
2026-07-06 09:50:50 -07:00
Tim Hockin 100532777c Put Worker.actor fields into an "Assignment" message (#363)
This is a cleanup I noticed as I was reviewing other PRs. I tried to
keep the commits clean, so you can see them one by one.
2026-07-01 12:35:58 -04:00
Haven Xia 05f7e6a856 Scope actor store and handlers by atespace 2026-06-29 14:57:54 -07:00
Dmitry Berkovich 37b5006d92 DurableDir support (#295)
Implement a layer allowing to take durableDir snapshot only, without
taking snapshot of entire memory. This is a part of the #119 feature.

The "durableDir" support is introduced via adding Volumes support to
ActorTemplate.

## Knowns issues / limitations


1. No MicroVM support for homedir yet. Will be implemented in following
PR.

2. Resume from SNAPSHOT_SCOPE_DATA temporary behavior
Current behavior of resume from durableStorage returns immediately after
gVisor process is started, it means container's bootstrap is not
completed by the time resume API returned to the caller. It causes the
issue that first curl request for suspended actor throws an HTTP
exception, since webserver listening on the port inside container is not
running yet.
PR #330 fixes the issue, for a meantime the workaround is to call `
kubectl ate resume actor` and after it call to the curl itself.

- [X ] Unit tests  pass
- [X ] E2E tests are updated and pass
2026-06-26 20:05:03 -07:00
Dmitry Berkovich 125180ea65 feat: implement container readiness probes with custom HTTP endpoint … (#330)
Extends ActorTemplate.Container spec to allow defining a readyz probe. 
* ateoms (both gVisor & microVM) are updated and wait for readyz signal
if it was enabled on container when creating a new container or resume
from a snapshot.
* actorTemplates controller has been updated not to wait 20 sec during
golden snapshot creation, if readyz is enabled on actor. It assumes,
atelet call returns when all containers passed readyz.

Tested manually both counters: 
* gVisor 
* microvm 

Fixes #316 

- [x] e2e tests pass 
- [X] Appropriate changes to documentation are included in the PR
2026-06-26 18:59:56 -07:00
Julian Gutierrez Oschmann 9f8475493b Decouple ActorTemplates from WorkerPools. (#270)
This change removes the hard-link from templates to worker pools,
allowing pools to be selected using k8s labels.

Both actors and actor templates are augmented with a label selector
field.

When an actor is resumed, we resolve the set of *eligible* worker pools
by evaluating the intersection of both actor and template selectors.
Then we find a free worker from one of them. The overall design is
described in [1].

[1]:
https://github.com/agent-substrate/substrate/issues/47#issuecomment-4712334620
2026-06-18 10:58:34 -07:00
Benjamin Elder 00a90a8eda Decouple Sandbox Binaries / Config from ActorTemplate (#261)
Fixes https://github.com/agent-substrate/substrate/issues/253

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Should make something like
https://github.com/agent-substrate/substrate/pull/239 less awkward, and
also https://github.com/agent-substrate/substrate/issues/255 (we don't
have to cram more fields into actortemplate for the shim binary and
update _all_ of the demo actortemplates to include them).
2026-06-17 09:46:46 -07:00
Dmitry Berkovich c1b51134e7 Actor Pause/Resume flow (#227)
Implements  `PAUSED` state for issue #119.

The snapshot files are kept locally on node VM in a separate folder. At
resume time, scheduler uses a node VM hint and picks up a worker from
the same node where files were stored at suspend time.

The local file management solution is temporary and will be replaced
once @msau42 introduces a new component that is supposed to manage files
on the node VM.

- [X] Tests pass
- [X] Manual tests with counter demo
```
>kubectl ate create actor my-counter-1 --template ate-demo-counter/counter
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 1
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 2
>kubectl ate pause actor my-counter-1 -o json
> curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 3
```

### Breaking change
This PR introduces a breaking change in the Actor proto. All existing
actor needs to be recreated, prior testing PAUSE functionality.
2026-06-16 13:15:36 -07:00