34 Commits
Author SHA1 Message Date
Max Smythe d6d2a0fa1d Add the option to install a large-cluster manifest and cordon control plane to dedicated machines (#1632)
This change allows users to install Substrate on larger clusters. It
adds a `--cluster-size` flag to enable more t-shirt-style sizing in the
future to accommodate clusters of different sizes.

It also adds a --cordon-control-plane flag that allows
taints/tolerances/antiaffinity/node labels to have each control plane
element run on its own dedicated machine.

Fixes #<issue_number_goes_here>

> It's a good idea to open an issue first for discussion.

- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-09-24 23:39:01 +00:00
shrutiyam-glitch 47b67574ac docs: Document RevertActor and drop "terminal" from CRASHED (#1711)
`RevertActor` returns a RUNNING, PAUSED, or CRASHED actor to SUSPENDED
at its last external snapshot, so CRASHED is no longer a dead end that
only `DeleteActor` can clear. Docs and code comments still described it
as terminal and told operators to delete and recreate the actor, losing
its state.

Update the api-guide, architecture, upgrade guide, and kubectl-ate
README to cover the new verb, and correct the comments that justified
keeping a partial external snapshot by naming actor deletion as the only
remaining collector -- revert collects it too.

Follow up for the PR - #1675 
Issue - #1556 

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-23 17:51:56 +00:00
Alex Zakonov 0c5a1cddb3 docs: sharpen Substrate overview and fix "computer infrastructure" typo (#1619)
## What

Updates the project overview language in the three docs that share it,
and fixes a long-standing typo.

- **`README.md`** — replaces the overview paragraph with the
secure-by-default positioning: density relative to standard container
runtimes, resume latency and activation throughput, and native
kernel/network isolation.
- **`docs/architecture.md`** — adopts the same lead sentence, keeping
the existing control-plane detail; `computer infrastructure` → `compute
infrastructure`.
- **`docs/roadmap.md`** — `computer infrastructure` → `compute
infrastructure`.

## Notes

The performance figures in the README paragraph (density multiple,
sub-500ms resume, activation rate) have been discussed and aligned
separately.

Docs-only change; no code or behavior is affected.
2026-09-11 14:08:16 -07:00
Keith Mattix II 6bd89588dc Move to lowercase for header references
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
Keith Mattix II f16fc04fa0 Move from Host header to explicit headers for actor and atespace
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
2026-09-09 13:23:36 -07:00
eliranw 353c21f203 docs: correct micro-VM rootfs description to host-backed overlay (#1531)
The architecture guide still described micro-VM container rootfs writes
as living in guest RAM via a tmpfs overlay.
#846 moved that overlay onto the host, served over the same single
virtio-fs share as the durable-dir volumes.
This updates the one stale sentence to match. 
Docs only.

Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
2026-09-08 16:56:59 -07:00
Youssuf Elshall 9ea39a5639 docs: remove stale references to ActorTemplate as a Kubernetes CRD (#1404)
## What

Now that the ActorTemplate CRD has been deleted and its resources moved
to the substrate gRPC API and the control-plane store (created/managed
with `kubectl-ate`, persisted in PostgreSQL), several documents still
describe ActorTemplate as a Kubernetes CRD, or describe namespace/RBAC
relationships that no longer exist. This sweeps the docs for those stale
references.

Fixes #368 (docs side).

> This change was prepared with AI assistance; I have reviewed and
tested it.

- [x] Docs and comment-only change; no functional code changed, no tests
affected
2026-09-04 16:40:04 -07:00
Luiz Oliveira 9b333c6fce Garbage Collect snapshots and remove the snapshot resource (#1417)
Fixes #664 

This PR implements the idea described in
https://github.com/agent-substrate/substrate/issues/664#issuecomment-5499311489

It does more than Garbage Collection of snapshots, because we also got
rid of the Snapshot resource (from the DB/API).

Now, an external snapshot is owned by a single resource:

- An Actor owns the snapshot it writes at suspend
- A tag owns a copy taken at tag creation,
- An actor cloned from a tag borrows the tag's snapshot until its own
first suspend.

Garbage Collection: whoever created/owns the snapshot is the only one
who ever deletes them:
i.e., if an actor is deleted and it owns a snapshot. The underlying
snapshot is deleted with the actor.

this PR:

- Drops table actor_snapshots
- Keeps table actor_snapshot_tags 
- Adds an object copy at tag creation, and an owned versus borrowed
distinction on the Actor
- Adds synchronous external snapshot deletion at actor suspend, at actor
delete, and at tag delete

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-04 16:22:02 -04:00
Zoe Zhao 6d3afdd63b Resolve SandboxConfig from the ActorTemplate instead of the WorkerPool (#1446)
This PR moves sandbox config selection from the WorkerPool to the
ActorTemplate.

The existing behavior is preserved while we are designing the upgrade:
sandbox config still cannot be updated once set (ActorTemplates are
create-only and `sandbox_config` is immutable).

For now the ActorTemplate still *requires* `sandbox_config.config_name`
— there is no resolution of the cluster default (`spec.default`). This
is temporary while we figure out the defaulting design.

- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-03 16:13:07 -07:00
Zoe Zhao 120a519606 Delete ActorTemplate CRD (#1376)
Fixes #368 . Deletes ActorTemplate CRD and any references to it.

Note about atenet router: It had a k8sclient controller that monitors
ActorTemplate, but the results were not used. Deleted as well.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 10:48:42 -04:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Sneha-at b7080602c6 Allow deleting an actor from any state. (#788)
Fixes #643 
This feature introduces the ability to delete Substrate actors from any
lifecycle state (e.g., RUNNING, PAUSED, PENDING), rather than requiring
them to be hibernated into SUSPENDED (or CRASHED) beforehand. This is
crucial for cleaning up stuck or failed actors.
#### 1. Control Plane & gRPC API (pkg/proto/ateapipb, ateapi)
• DeleteActorRequest.any_state: Added a new boolean field any_state to
DeleteActorRequest.
• When false (default): enforces the existing behavior where only actors
in ACTOR_STATE_SUSPENDED or ACTOR_STATE_CRASHED (or already
ACTOR_STATE_DELETING) can be deleted.
• When true: allows deleting an actor in any state (e.g., RUNNING,
PAUSED).
• Orchestrated Cleanup Workflow (workflow_delete.go):
1. Transitions actor state to ACTOR_STATE_DELETING.
2. Calls atelet.Terminate to stop live workloads on the node.
3. Detaches external volumes from the worker node (with fallback logic
using actor status if the ActorTemplate was deleted).
4. Releases the assigned physical worker Pod back to the WorkerPool.
5. Deletes external storage volumes through the storage plugins.
6. Finalizes removal of the actor from the persistence store.
#### 2. Node Herder Agent (atelet)
• Terminate RPC (main.go): Added a new Terminate method to AteomHerder:
• Calls ateom.TerminateWorkload on the target worker pod.
• Unmounts external volumes mounted for the actor on the node.
• Cleans up and resets actor runtime host directories (OCI bundles,
checkpoints, pid files).
#### 3. In-Pod Sandboxes (ateom-gvisor & ateom-microvm)
• TerminateWorkload RPC:
• gVisor (main.go): Stops and deletes active runsc containers, unmounts
bundle rootfs overlays, and tears down the actor's interior network
namespace.
• Micro-VM (checkpoint.go): Shuts down the Cloud Hypervisor VMM,
unmounts bundle rootfs overlays, and cleans up the network namespace.
• Emits an "Actor terminated" lifecycle log and clears the active actor
attribution so the worker can host new workloads.
#### 4. CLI (kubectl-ate)
• --any-state Flag (delete_actor.go):
kubectl-ate delete actor <actor-name> -a <atespace> --any-state

──────
  ## Verification 
  [x] Changes tested in local with 
 ```
• go build ./...: Passed
go test ./...: Passed (all unit and functional tests pass)
    make lint: Passed (no lint errors) 
```
Detailed manual tests                                                                                                                                                                                                                                                                                                                                           
 ```                                                                                                                                                                                                                                                                                                                                                
  ### Scenario 1: Force Delete a Running Actor with --any-state                                                                                                                                                                                                                                                                                                      
                                                                                                                                                                                                                                                                                                                                                                     
  1. Create the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate create actor manual-test-1 --template=ate-demo-counter-microvm/counter-microvm -a demo                                                                                                                                                                                                                                                               
                                                                                                                                                                                                                                                                                                                                                                     
  2. Resume the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate resume actor manual-test-1 -a demo                                                                                                                                                                                                                                                                                                                   
                                                                                                                                                                                                                                                                                                                                                                     
  3. Force delete the running actor:                                                                                                                                                                                                                                                                                                                                 
    kubectl ate delete actor manual-test-1 -a demo --any-state                                                                                                                                                                                                                                                                                                       
    actor "manual-test-1" deleted                                                                                                                                                                                                                                                                                                                                    
```
──────
### Scenario 2: Standard Delete a Suspended Actor
```                                                                                                                                                                                                                                                                                                                                  
  1. Create the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate create actor manual-test-2 --template=ate-demo-counter-microvm/counter-microvm -a demo                                                                                                                                                                                                                                                               
                                                                                                                                                                                                                                                                                                                                                                     
  2. Resume the actor:                                                                                                                                                                                                                                                                                                                                               
    kubectl ate resume actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                   
                                                                                                                                                                                                                                                                                                                                                                     
  3. Suspend the actor:                                                                                                                                                                                                                                                                                                                                              
    kubectl ate suspend actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                  
                                                                                                                                                                                                                                                                                                                                                                     
  4. Standard delete the suspended actor:                                                                                                                                                                                                                                                                                                                            
    kubectl ate delete actor manual-test-2 -a demo                                                                                                                                                                                                                                                                                                                   
    actor "manual-test-2" deleted                                                                                                                                                                                                                                                                                                                                    
```
──────
### Unit & E2E Tests
# Unit & Functional Tests
go test ./cmd/ateapi/internal/controlapi -run
"TestDeleteActorWorkflow|TestEnsureMarkedDeleting"
# E2E Tests
E2E_TEMPLATE_NAMESPACE=ate-demo-counter-microvm
E2E_TEMPLATE_NAME=counter-microvm ./hack/run-e2e-kind.sh
./internal/e2e/suites/demo -run TestActorLifecycle
./hack/run-e2e-kind.sh ./internal/e2e/suites/demo -run
TestForceDeleteActorWithExternalVolume
```
- [ x] Appropriate changes to documentation are included in the PR
2026-08-20 13:17:59 -07:00
Julian Gutierrez Oschmann 4c1bd9d36b Add status fields to Substrate resources (#1025)
Add `status` fields to all Substrate resources that need it. Move
server-owned fields under it.

Fixes #1006 .
2026-08-18 11:58:31 -07:00
Keith Mattix II 89b83a5c89 Support CONNECT in atenet router (#715)
added CONNECT support for atenet ingress to support arbitrary actor ports
2026-08-14 16:12:27 -07:00
Benjamin Elder da8414bbd2 api: move the pause image from ActorTemplate to SandboxConfig (#848)
The pause image holds the sandbox's namespaces and runs no workload
code. It is an implementation detail of the sandbox, not something actor
authors pick, so it belongs with the sandbox binaries that already moved
off the ActorTemplate onto the cluster-scoped SandboxConfig.

It now travels with those binaries end to end: resolved from the pool's
SandboxConfig, carried on ateletpb.SandboxAssets rather than
WorkloadSpec, and recorded in the per-actor sandbox record so Checkpoint
pins it into the snapshot manifest and Restore rebuilds the sandbox from
the image the snapshot was taken with (the golden's on a DATA_ON_GOLDEN
restore). A record without one is rejected outright rather than pulling
an empty image.


- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-11 10:14:03 -04:00
Dmitry Berkovich 3a2d0c1d86 Support suspending a PAUSED actor without waking it (#816)
Part of #791 — the last planned piece of the [implementation
plan](https://github.com/agent-substrate/substrate/issues/791#issuecomment-5226674097)
(after #810, #812, #813). #791 stays open until #817 (gVisor
Full-capture → Data-commit conversion, blocked on #790) is done; this PR
covers micro-VM fully and gVisor for scope-matched suspends. Also part
of the actor state machine (#119) and a prerequisite for system upgrade
flows (#473).

`SuspendActor` now accepts a PAUSED actor: instead of checkpointing a
running workload, ateapi dials the atelet on the node holding the pause
snapshot and has it upload the node-local files to object storage, then
finalizes as usual — durable `ActorSnapshot`, `SUSPENDED` status, node
pinning cleared.

Two commits, reviewable independently:

## Commit 1 — atelet: `UploadPausedCheckpoint` RPC (dead code until
commit 2)

- New `AteomHerder` RPC: a pure disk→object-storage copy driven by the
snapshot's self-describing manifest — no ateom involved (the sandbox is
gone).
- **Scope conversion** dispatches per sandbox class
(`narrowFullCaptureToData`): a micro-VM FULL capture narrows to a DATA
upload by carving out `durable-dir.tar` (constant hoisted to
`ateompath`, shared with ateom-microvm); gVisor returns `Unimplemented`
until split checkpoints land (#790); DATA can never widen to FULL; a
scope-less manifest (older atelet) is rejected rather than guessed at.
- **Idempotent retry**: local files gone + remote manifest present ⇒ a
previous invocation committed, succeed; gone on both sides ⇒
unrecoverable (`LOCAL_SNAPSHOT_GONE`, crashes the actor). Upload
failures stay plain retryable errors; the manifest uploads last as the
commit marker, never in parallel.
- The golden atespace is rejected at validation (fully on the
`field.ErrorList` framework): golden actors are never paused.

## Commit 2 — control plane: enable suspend from PAUSED

- `FromPaused` discriminator: PAUSED status, or SUSPENDING with no
worker assignment and a `LocalSnapshotInfo` — the field alone is stale
on resumed-from-pause RUNNING actors, so the nil-assignment conjunct is
load-bearing.
- `MarkSuspendingStep` accepts PAUSED and rejects a Data-captured pause
against a Full commit *before* the actor leaves PAUSED (an upload cannot
fabricate memory), using the `content_scope` recorded at pause (#812)
with an `onPause` fallback.
- `CallAteletSuspendStep` paused branch dials by node
(`DialForAteletOnNode`, #813): missing node record ⇒ crash (the snapshot
can never be found); unreachable atelet ⇒ retryable; atelet's
`LOCAL_SNAPSHOT_GONE` ⇒ crash via `maybeCrashActor`.
`FinalizeSuspendedStep` needed no changes thanks to the #813 hoist.
- Root-cause guard: `MarkPausingStep` rejects pausing golden-atespace
actors.
- Docs: pause states + the new `PAUSED → SUSPENDING` edge in the
architecture state diagram; glossary Suspend entry covers both origins.

## Tests

- **atelet unit**: 10 upload-helper subtests (conversion matrix,
idempotency probe, data-loss crash, retryable upload failure) via a
recording object-storage fake; validation table.
- **control-plane unit**: discriminator table, scope-rejection table
(incl. onPause fallback), paused preconditions (no node ⇒ CRASHED, no
atelet ⇒ `ErrNoAteletOnNode` + still SUSPENDING), golden-pause
rejection, prerequisite matrix updated.
- **functional (envtest)**: `TestSuspendActor_FromPaused` (upload called
with the pause snapshot name, no Checkpoint RPC, SUSPENDED, pinning
cleared, ActorSnapshot at the upload destination) +
retry-after-failed-upload (same destination on retry).
- **e2e (demo suite)**: the lifecycle driver gains a suspend-from-PAUSED
mode; three durable-dir cases — Full/Full, Data/Data, and the micro-VM
Full→Data extraction — assert memory/file counters survive the full
pause→suspend→resume journey and the node pinning is gone.

`go test -race ./...` clean (except the pre-existing macOS-only
`internal/atunnel` unix-socket-path failures, untouched by this PR),
`gofmt`/`go vet` clean, protos regenerated via `hack/protoc.sh`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-10 22:34:03 -07:00
Lior Lieberman 860250bf71 diagram 2026-07-31 12:51:48 -07:00
Lior Lieberman d9ac4e24e6 networking: update docs and add an ingress e2e suite
Bring the architecture and threat-model docs in line with the atunnel
ingress path: the router no longer rewrites :authority to a worker pod IP
and forwards over plaintext port 80, it opens an mTLS tunnel to atunnel on
worker port 443, which forwards to the Actor over its private veth.

Drop cmd/atenet/atenet-diagram.png. It predates the ext_proc/ORIGINAL_DST
design and is now wrong in the part that matters most. The README section it
illustrated gains a short accurate note about the upstream hop instead of a
dangling image reference.

Add an e2e suite covering the change end to end: TestActorDirectAccess
asserts that the worker pod's port 80 is no longer a reachable Actor ingress
path (the DNAT rule is gone) and that the same Actor still answers /readyz
through atenet-router over the atunnel mTLS hop. It uses the counter demo as
its fixture, so it only needs --deploy-demo-counter.

internal/e2e/testmain.go picks up ParseSkippedFlags so that `go test` flags
(-run, -v, ...) survive pflag parsing and reach the suite.
2026-07-31 12:51:48 -07:00
Alex Zakonov 6f44915bdf docs: fixed language describing project goals and relationship to Kubernetes (#524)
Fixed language describing project goals and relationship to Kubernetes
in readme.md, architecture.md and roadmap.md
- changed the top language to focus on project goals instead of
Kubernetes relations
- changed the language describing Kubernetes relationship to describe
value of the Agent Substrate layer and Kubernetes layer
2026-07-31 12:35:04 -07:00
Benjamin Elder 97772a02f3 Support multiple durable-dir volumes on the micro-VM runtime
An ActorTemplate could declare only one durable-dir volume, and mount it
into a container only once. That is a gVisor limit, not a general one:
atelet declares the mount to gVisor through a single hardcoded annotation
key ("dev.gvisor.spec.mount.durabledir"), so a second volume would silently
overwrite the first.

The micro-VM runtime has no such constraint — every volume is a
subdirectory of the one writable virtio-fs share, and the snapshot tar
already archives that directory whole, so N volumes cost a subdirectory
each and round-trip through checkpoint/restore untouched. Gate the two
"at most one" CEL rules on sandboxClass so micro-VM templates may declare
several while gVisor keeps its cap, and say so in the messages.

What was actually missing was the volume NAME: ateom received mount paths
alone, so it inferred the single name by listing atelet's directory. Carry
each mount's name on the wire (Container.durable_dir_volume_mounts,
replacing the paths-only field, which is reserved rather than retyped since
atelet and ateom are separate images that can skew across a rollout), and
delete the inference. A container's binds now come from its own mounts, and
the actor-wide "has a durable share" question collapses to a bool.

The counter demo grows an optional --second-file-counter-directory, and a
micro-VM-only e2e case runs the lifecycle matrix against an Actor with two
durable volumes, asserting both counters advance together. gVisor skips it:
the template would be rejected at admission.
2026-07-30 20:03:55 -07:00
Eitan Yarmush ef7b29da44 ateapi: add ActorSnapshot lifecycle APIs 2026-07-30 19:40:59 -07:00
Lior Lieberman 92f1aa7276 Rename SessionIdentity to ActorIdentity
Sessions are no longer a concept in Substrate; Actor is the glossary
term. This completes the "s/Session/Actor" TODO that sat at the top of
ateapi.proto, and removes the TODO.

API surface:
  service SessionIdentity        -> ActorIdentity
  MintJWTRequest.session_id      -> actor_id
  MintJWTResponse.session_jwt    -> actor_jwt
  MintCertRequest.session_id     -> actor_id
  MintCertResponse.session_certificates -> actor_certificates

Go packages:
  cmd/ateapi/internal/sessionidentity -> actoridentity
  cmd/ateapi/internal/sessionidjwt    -> actoridjwt

Flags and cluster resources:
  --session-id-jwt-pool -> --actor-id-jwt-pool
  --session-id-ca-pool  -> --actor-id-ca-pool
  Secrets, volumes and mount paths renamed to match, in both
  manifests/ate-install/ate-api-server.yaml and hack/install-ate.sh
  (--create-session-id-ca-pool-secret -> --create-actor-id-ca-pool-secret).

Two credential identity values change with the rename:
  JWT issuer https://broker.agentic-substrate-session-id-broker.svc
          -> https://broker.agentic-substrate-actor-id-broker.svc
  SPIFFE ID spiffe://substrate-session.local/app/../session/..
          -> spiffe://substrate-actor.local/app/../actor/..
Tokens and certificates issued before this change will not validate
against the new issuer or trust domain.

BREAKING: the gRPC wire path moves from /ateapi.SessionIdentity/* to
/ateapi.ActorIdentity/*, and the Secrets must be recreated under their
new names before the new ate-api-server rolls out.
2026-07-30 07:34:29 -07:00
Benjamin Elder aa1d14a7b3 demos,e2e,docs: exercise durable-dir volumes on the micro-VM counter
Give the micro-VM counter demo the same durableDir volume the gVisor one
has, and drop TestDurableDirLifecycle's micro-VM skip so the scope matrix
(Full/Full, Data/Full, Data/Data) runs against both runtimes.

TestActorLifecycle no longer special-cases micro-VM to expect a file counter
of -1. That value was the demo failing to write: it counts in
/home/counter/a.txt, which nothing created in the micro-VM rootfs. The volume
mounts there now, so both runtimes assert the same counters.

Also correct the glossary: an ActorTemplate may declare one DurableDir
volume, not several — CEL has enforced that since the feature landed — and
Full scope captures the volumes whether or not the runtime keeps them inside
rootfs (the micro-VM does not).
2026-07-24 18:59:35 -07:00
Mesut Oezdil 9b70da80ab fix: correct errors in docs (#322)
- `docs/observability.md`: log label keys were wrong. The code emits
`ate.dev/actor_id` and `ate.dev/actor_template_name`; the doc had
`actor_id` and `actor_template`. The Cloud Logging queries would return
no results.
- `docs/dev/best-practices/tracing.md`: container ports entry used
Docker Compose syntax (`"443:443"`). Changed to `containerPort: 443`.
- `docs/dev/valkey-direct-access.md`: the `kubectl exec` command had a
line break inside `--cert`, making it fail on copy-paste.
- `code-of-conduct.md`: linked to Contributor Covenant v1.4. Updated to
v2.1.
2026-07-23 10:11:00 -07:00
Julian Gutierrez Oschmann a2d55e99e0 Update Actor ID -> Actor Name references (#455)
Update all old references to actor id to refer to actor name.

Also fix benchmarking framework to use the latest version of the API.
2026-07-17 16:49:19 -07:00
mesutoezdil 09f1c4673c fix: update stale actor DNS format to include atespace label 2026-07-06 10:14:14 -07:00
Benjamin Elder 3d53de85eb microvm MVP cleanup: align rootfs behavior to gvisor, ... (#313)
Follow-up to https://github.com/agent-substrate/substrate/pull/287 /
#123

- align snapshot behavior: use a tmpfs (for writes) on top of read-only
viritio-fs mount for the container image rootfs instead of ext4 images
- enable multiple container support
- cleanup stale comments

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-06-26 20:02:11 -07:00
Benjamin Elder 2c50d282e2 hack,demos,docs: micro-VM asset tooling, base image, demo runner, docs
Assemble + stage the micro-VM runtime assets, an ateom-base image (debian-slim +
e2fsprogs for mkfs.ext4), and run-microvm-demo.sh to build + deploy the
counter-microvm demo end to end (overriding the worker base via KO_CONFIG_PATH so
no committed file is edited). Document the micro-VM sandbox class.
2026-06-25 13:44:25 -07:00
Frederick F. Kautz IV bd33e07606 docs: name resource-model namespaces by owning server
Rename the class-diagram namespaces from the abstract categories
(KubernetesObjects / ControlPlaneRecords) to the servers that actually
own the records: kube-apiserver (ActorTemplate/WorkerPool CRDs,
Deployment, WorkerPod) and ate-api-server (Actor/Worker, the proto
records persisted in Redis). Per review feedback on #245.

Uses the hyphenated ate-api-server to match the deployed Service and
Deployment, the client dial address, and the rest of this doc, rather
than the ate-apiserver spelling that only appears in an install flag.
2026-06-24 22:46:46 -07:00
Frederick F. Kautz IV c8cc795898 docs: add architecture diagrams (resource model, activation, lifecycle)
Add the sequence, class, and state diagrams to architecture.md, where
people look for diagrams when modifying the system. Each lands in the
section whose prose already describes it: the resource model under API
Resource Models, and the activation flow and lifecycle state machine
under Actor Lifecycle.

This is the companion to #200, which drops the same diagrams from the
glossary to keep it a plain list of terms (per review feedback).

Every diagram claim was verified against the current implementation. The
activation note no longer implies idle-triggered suspend, since
SuspendActor is an explicit call with no idle-detection mechanism.
2026-06-24 22:46:46 -07:00
han2ni3bal-pixel 962ff6b1ae Add WorkerPool scheduling fields (#247)
**Updated Original PR Description**
## Summary

This PR adds a small, controlled set of scheduling fields to WorkerPool
so WorkerPool-managed Pods can be placed onto appropriate Kubernetes
nodes.

Added fields:

- `nodeSelector`
- `tolerations`
- `priorityClassName`
- `nodeAffinity`

## Scope

This intentionally does not expose a full `PodTemplateSpec`, and does
not add support for full `affinity`, `podAffinity`, or
`podAntiAffinity`.

Resource requests/limits are intentionally left out of scope for this CL
while the WorkerPool resource model is still being discussed separately.

The goal is to keep the first version small while addressing the
WorkerPool-to-Kubernetes-Node scheduling gap discussed in
[#212](https://github.com/agent-substrate/substrate/issues/212).

`nodeAffinity` is included to support heterogeneous node pools where
`nodeSelector` is too limited, such as requiring or preferring one of
several equivalent accelerator/local-SSD/cache node pools, or expressing
soft preferences during node pool migration.

`priorityClassName` is included to help preserve warm WorkerPool
capacity under cluster pressure. Since WorkerPool pods are the slots
available for actor resume, higher-priority pools can protect
latency-sensitive or interactive actor workloads from displacement by
lower-priority batch workloads, and allow different WorkerPools to
represent different service classes, e.g. interactive vs. batch.

## Related issue

Part of #212.
Related to #47.
2026-06-17 10:33:00 -04:00
Dmitry Berkovich c1b51134e7 Actor Pause/Resume flow (#227)
Implements  `PAUSED` state for issue #119.

The snapshot files are kept locally on node VM in a separate folder. At
resume time, scheduler uses a node VM hint and picks up a worker from
the same node where files were stored at suspend time.

The local file management solution is temporary and will be replaced
once @msau42 introduces a new component that is supposed to manage files
on the node VM.

- [X] Tests pass
- [X] Manual tests with counter demo
```
>kubectl ate create actor my-counter-1 --template ate-demo-counter/counter
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 1
>curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 2
>kubectl ate pause actor my-counter-1 -o json
> curl -X POST -H "Host: my-counter-1.actors.resources.substrate.ate.dev" http://localhost:8000
hello from: 169.254.17.2 | preserved memory count: 3
```

### Breaking change
This PR introduces a breaking change in the Actor proto. All existing
actor needs to be recreated, prior testing PAUSE functionality.
2026-06-16 13:15:36 -07:00
Tim Hockin d9773e7d03 Remove trailing space from md files 2026-05-31 19:45:36 -07:00
+7 af3c65088e Initial commit of Agent Substrate
This is the initial release of the Agent Substrate.

Agent substrate is a system built on top of Kubernetes which manages agent-like
workloads to achieve higher scale and efficiency than Kubernetes alone can
offer, with lower latency.  It builds on top of Kubernetes features like
Pods and Pod autoscaling, but takes the Kubernetes control-plane out of the
critical path to achieve lower latency.

It can run on any Kubernetes cluster and does not inhibit “regular” use of
Kubernetes in any way. Kubernetes provides the infrastructure provisioning and
management for all types of workloads, while Agent Substrate provides
agent-specific scheduling and control.

At its core, Agent Substrate maps a larger set of “actors” (applications such
as agents) onto a smaller set of ready “workers” (Kubernetes Pods), relying on
the fact that agent-like applications tend to be idle most of the time to
achieve heavy multiplexing.  It provides functionality to manage an actor’s
lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real
time, and to route incoming traffic to them.

Agent Substrate is intended to be a low-opinion system.  The workloads it
manages don't have to be literal AI agents, but those are the best example of
the kind of applications it is designed for.  It is not an SDK for building
agents, but rather a system for running them at scale.

Agent Substrate is currently in VERY early development.  It is not ready for
production use, and the APIs are almost guaranteed to change.  We are not
making any guarantees about backward compatibility at this stage, and
everything in this project may be changed.

Co-authored-by: Alex Bulankou <alexbu@google.com>
Co-authored-by: Benjamin Elder <bentheelder@google.com>
Co-authored-by: Bowei Du <bowei@google.com>
Co-authored-by: Dmitry Berkovich <dberkov@google.com>
Co-authored-by: Fabricio Voznika <fvoznika@google.com>
Co-authored-by: Francisco Cabrera <fclieutier@google.com>
Co-authored-by: Haven Xia <haoyuxia@google.com>
Co-authored-by: Julian Gutierrez Oschmann <juliangut@google.com>
Co-authored-by: Kevin Steuer <ksteuer@google.com>
Co-authored-by: Max Smythe <smythe@google.com>
Co-authored-by: Maya Wang <mymaya@google.com>
Co-authored-by: Michael Taufen <mtaufen@google.com>
Co-authored-by: Shruti Nair <shrutinair@google.com>
Co-authored-by: Taahir Ahmed <taahm@google.com>
Co-authored-by: Tim Hockin <thockin@google.com>
Co-authored-by: Zoe Zhao <zoezhao@google.com>
2026-05-19 16:57:14 -07:00