The pattern of passing a tuple of atespace and actor name together
is across the whole codebase. This simplifies both callers and callees
by leting them pass a single field that bundles both.
The ActorRef is the actor specific, typed,
in-process version of the ObjectRef we have in our gRPC API.
Initial part of #232.
Extends the ActorTemplate API to add support for per-actor external
volumes and adds control plane and atelet hooks following the Actor
lifecycle.
Callouts to volume operations are abstracted with a Volume interface.
Right now, only a mock volume plugin for testing has been implemented,
but we will add CSI support next.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR - Not
going to update documentation until we add CSI support.
Should found it when writing #474..
And also correct an error in `workflow_resume.go`.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fixes#369
Add a CheckPrerequisite method to the WorkflowStep interface, called by
RunWorkflow after IsComplete returns false and before Execute, so that
each workflow validates its actor state-machine edge up front while
retried (reentrant) workflows still fast-forward past completed steps.
When the worker pod is gone from the DB during pause finalization,
nodeName stays empty.
The code still wrote NodeVmsWithLocalSnapshots: []string{""}, so on the
next ResumeActor the scheduler looked for a worker with node name "",
found none, and returned "no free workers available" permanently.
Fix: only set NodeVmsWithLocalSnapshots when nodeName is non-empty.
Aligns the internal ateletpb/ateompb services with the API style guide's
actor naming (field tags unchanged, so the wire format is compatible).
internal APIs keep dedicated identity fields (not the public ObjectRef).
Replace actor_id and atespace name with `metadata` field which
holds common fields.
This commit only changes the ate-apiserver to support this field, but
intentionally skips upstream / downstream systems (such as kubectl-ate,
atenet, atelet, etc) to keep the scope manageable. This will come in subsequent PRs.
Related to #292 and #119
Changes:
1. ResumeActor:
1. During `ResumeActor` we'll put the actor in CRASHED state if the
Resume failed due to any errors related to the snapshot itself, which
includes:
1. Missing or corrupted snapshot when downloading from GCS
2. Runsc command fails.
2. All other types of errors will not put the Actor in CRASHED state.
2. SuspendActor or PauseActor
1. During `SuspendActor` or `PauseActor` we'll put the actor in CRASHED
state if we cannot produce a valid snapshot, which includes:
1. cannot connect to Ateom after retries,
2. runsc command fails.
3. cannot upload the snapshot to GCS (during suspend) ; or cannot save
snapshot locally (during pause).
2. All other types of errors will not put the Actor in CRASHED state.
This fixes an issue where atelet/ateom will clash the state
of actors with same ID within different atespaces.
Once substrate resources get UIDs, we should revisit the way
atelet / ateom does the bookkeeping and check in which cases it
should use UID instead of (atespace,actor_id).
Implement a layer allowing to take durableDir snapshot only, without
taking snapshot of entire memory. This is a part of the #119 feature.
The "durableDir" support is introduced via adding Volumes support to
ActorTemplate.
## Knowns issues / limitations
1. No MicroVM support for homedir yet. Will be implemented in following
PR.
2. Resume from SNAPSHOT_SCOPE_DATA temporary behavior
Current behavior of resume from durableStorage returns immediately after
gVisor process is started, it means container's bootstrap is not
completed by the time resume API returned to the caller. It causes the
issue that first curl request for suspended actor throws an HTTP
exception, since webserver listening on the port inside container is not
running yet.
PR #330 fixes the issue, for a meantime the workaround is to call `
kubectl ate resume actor` and after it call to the curl itself.
- [X ] Unit tests pass
- [X ] E2E tests are updated and pass
Extends ActorTemplate.Container spec to allow defining a readyz probe.
* ateoms (both gVisor & microVM) are updated and wait for readyz signal
if it was enabled on container when creating a new container or resume
from a snapshot.
* actorTemplates controller has been updated not to wait 20 sec during
golden snapshot creation, if readyz is enabled on actor. It assumes,
atelet call returns when all containers passed readyz.
Tested manually both counters:
* gVisor
* microvm
Fixes#316
- [x] e2e tests pass
- [X] Appropriate changes to documentation are included in the PR
This change removes the hard-link from templates to worker pools,
allowing pools to be selected using k8s labels.
Both actors and actor templates are augmented with a label selector
field.
When an actor is resumed, we resolve the set of *eligible* worker pools
by evaluating the intersection of both actor and template selectors.
Then we find a free worker from one of them. The overall design is
described in [1].
[1]:
https://github.com/agent-substrate/substrate/issues/47#issuecomment-4712334620
Implements `PAUSED` state for issue #119.
The snapshot files are kept locally on node VM in a separate folder. At
resume time, scheduler uses a node VM hint and picks up a worker from
the same node where files were stored at suspend time.
The local file management solution is temporary and will be replaced
once @msau42 introduces a new component that is supposed to manage files
on the node VM.
- [X] Tests pass
- [X] Manual tests with counter demo
```
>kubectl ate create actor my-counter-1 --template ate-demo-counter/counter
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 1
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 2
>kubectl ate pause actor my-counter-1 -o json
> curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 3
```
### Breaking change
This PR introduces a breaking change in the Actor proto. All existing
actor needs to be recreated, prior testing PAUSE functionality.