mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Remove trailing space from md files
This commit is contained in:
@@ -17,7 +17,7 @@ The AGENTS.md file at the root of the project may include the following sections
|
||||
- Testing instructions
|
||||
- Security considerations
|
||||
|
||||
AGENTS.md files in subfolders may be even more concise and specific to those subfolders.
|
||||
AGENTS.md files in subfolders may be even more concise and specific to those subfolders.
|
||||
|
||||
## Maximize AGENTS.md Performance
|
||||
|
||||
|
||||
@@ -109,7 +109,7 @@ kubectl ate get workers
|
||||
| `ASSIGNED ACTOR` | If `STATUS=ASSIGNED`, the actor reference `<namespace>/<template>/<actor-id>`. |
|
||||
|
||||
### Actor Lifecycle
|
||||
Manage the execution state of your workloads.
|
||||
Manage the execution state of your workloads.
|
||||
*(Note: Actors are identified by a user-provided ID, which must be a valid DNS-1123 label)*
|
||||
|
||||
```bash
|
||||
|
||||
+2
-2
@@ -58,7 +58,7 @@ metadata:
|
||||
spec:
|
||||
runsc:
|
||||
amd64:
|
||||
# Note: These values are from the 2026-05-19 nightly.
|
||||
# Note: These values are from the 2026-05-19 nightly.
|
||||
# For the latest verified versions, see: demos/counter/counter.yaml.tmpl
|
||||
url: "gs://gvisor/releases/nightly/2026-05-19/x86_64/runsc"
|
||||
sha256Hash: "a397be1abc2420d26bce6c70e6e2ff96c73aaaab929756c56f5e2089ea842b63"
|
||||
@@ -159,7 +159,7 @@ Workloads can exchange their ephemeral Kubernetes credentials for stable **Sessi
|
||||
|
||||
## 7. Framework & Ecosystem Integration
|
||||
|
||||
Agent Substrate is designed to be the foundational execution layer for any agentic framework.
|
||||
Agent Substrate is designed to be the foundational execution layer for any agentic framework.
|
||||
|
||||
### Agent Development Kit (ADK)
|
||||
Substrate provides native support for ADK-compatible identities. Workloads can use the `SessionIdentity` service to mint JWTs that align with ADK's security model, ensuring seamless integration with ADK-managed tools and memory.
|
||||
|
||||
@@ -293,7 +293,7 @@ The node-level subsystem manages the physical execution of sandboxes and the mov
|
||||
|
||||
Handles session-aware routing and automatic re-animation.
|
||||
|
||||
* **Uniform DNS Mesh**: Substrate provides a location-transparent actor discovery scheme via a global DNS suffix (`<id>.actors.resources.substrate.ate.dev`).
|
||||
* **Uniform DNS Mesh**: Substrate provides a location-transparent actor discovery scheme via a global DNS suffix (`<id>.actors.resources.substrate.ate.dev`).
|
||||
|
||||
* **Routing**: The `atenet` router (powered by Envoy and an External Processing server) intercepts traffic destined for the mesh. It extracts the actor ID from the `Host` header, queries the Control Plane to determine the actor's current location, and triggers a `ResumeActor` workflow if the session is currently suspended.
|
||||
|
||||
@@ -358,14 +358,14 @@ collected. The garbage collection process is not implemented yet.
|
||||
Agent Substrate distinguishes between two types of state, which are currently
|
||||
captured together in a single versioned snapshot:
|
||||
|
||||
1. **Memory Snapshot**: The exact RAM state of the process.
|
||||
1. **Memory Snapshot**: The exact RAM state of the process.
|
||||
|
||||
2. **Working Volume (Disk)**: The files written to the container's writable
|
||||
layer (the "working memory").
|
||||
layer (the "working memory").
|
||||
|
||||
In the current implementation, both memory and disk states are tied to the
|
||||
specific version of the code (ActorTemplate). This ensures strict consistency
|
||||
during resumption.
|
||||
during resumption.
|
||||
|
||||
Snapshots are stored durably in **Google Cloud Storage (GCS)**. This model
|
||||
allows the physical compute resources in the `WorkerPool` to be fully reclaimed
|
||||
@@ -385,7 +385,7 @@ Agent Substrate is built on a **Defense-in-Depth** model:
|
||||
versions.
|
||||
|
||||
* **Request Authorization**: The system currently performs **Identity-Aware
|
||||
Routing** by utilizing a uniform DNS routing scheme
|
||||
Routing** by utilizing a uniform DNS routing scheme
|
||||
(`<actor id>.actors.resources.substrate.ate.dev`)
|
||||
at the gateway to extract and validate actor identifiers from incoming traffic. This
|
||||
ensures requests are only routed to recognized, registered actors.
|
||||
|
||||
+62
-62
@@ -6,12 +6,12 @@ Agent Substrate is a Kubernetes-based system for building multitenant applicatio
|
||||
|
||||
This project is trying to move very quickly to find the right set of capabilities for the ever-changing agentic workloads market. Below are our priorities. Any efforts which are not aligned with these priorities should probably be deferred.
|
||||
|
||||
1. Pinning down architectural decisions which influence the rest of these priorities.
|
||||
2. Performance and reliability of wakeups, including suspend/resume and snapshot management.
|
||||
3. A fast, scalable control-plane, including the storage layer, reliability, and security.
|
||||
4. Useful identity, such that policies can be written in terms of identity.
|
||||
5. Useful policy, including how users can specify ingress, egress, and peer networking.
|
||||
6. Runtime modularity, such that different kinds of sandboxes (gVisor, microVMs, etc) can be used.
|
||||
1. Pinning down architectural decisions which influence the rest of these priorities.
|
||||
2. Performance and reliability of wakeups, including suspend/resume and snapshot management.
|
||||
3. A fast, scalable control-plane, including the storage layer, reliability, and security.
|
||||
4. Useful identity, such that policies can be written in terms of identity.
|
||||
5. Useful policy, including how users can specify ingress, egress, and peer networking.
|
||||
6. Runtime modularity, such that different kinds of sandboxes (gVisor, microVMs, etc) can be used.
|
||||
|
||||
Below is a collection of finer-grained efforts which we believe align with the above.
|
||||
|
||||
@@ -19,114 +19,114 @@ Below is a collection of finer-grained efforts which we believe align with the a
|
||||
|
||||
### Actor Management (Compute)
|
||||
|
||||
* Actor versioning via ActorTemplate (or new ActorDeployment API), with support for A/B testing and rollout strategies.
|
||||
* Namespaces or a similar concept for grouping related actors together for both management and convenience of writing actor-to-actor authorization policy.
|
||||
* Actor Forking/Cloning: Ability to branch a new logical actor from an existing checkpoint (the 'State Root') to support complex agent reasoning paths.
|
||||
* Worker horizontal autoscaling: Ability to rapidly scale up nodes and warm Pods to meet actor demand.
|
||||
* Clarify the actor lifecycle, including what data is retained across which events (e.g. gvisor upgrade \-\> loss of memory snapshot)
|
||||
* Decide: What mode(s) of actor runs do we support? Can we prioritize and phase this? Assuming we always persist “working” data
|
||||
* Clean start from OCI for every activation
|
||||
* Resume from golden for every activation
|
||||
* Persist rootfs (when possible), clean binary start for every activation
|
||||
* Actor versioning via ActorTemplate (or new ActorDeployment API), with support for A/B testing and rollout strategies.
|
||||
* Namespaces or a similar concept for grouping related actors together for both management and convenience of writing actor-to-actor authorization policy.
|
||||
* Actor Forking/Cloning: Ability to branch a new logical actor from an existing checkpoint (the 'State Root') to support complex agent reasoning paths.
|
||||
* Worker horizontal autoscaling: Ability to rapidly scale up nodes and warm Pods to meet actor demand.
|
||||
* Clarify the actor lifecycle, including what data is retained across which events (e.g. gvisor upgrade \-\> loss of memory snapshot)
|
||||
* Decide: What mode(s) of actor runs do we support? Can we prioritize and phase this? Assuming we always persist “working” data
|
||||
* Clean start from OCI for every activation
|
||||
* Resume from golden for every activation
|
||||
* Persist rootfs (when possible), clean binary start for every activation
|
||||
* Persist rootfs and memory (when possible), full process resume
|
||||
|
||||
### Worker management
|
||||
|
||||
* Worker class (e.g. free-tier on cheap hardware, premium on good hardware)
|
||||
* Worker class (e.g. free-tier on cheap hardware, premium on good hardware)
|
||||
* Fungible worker pools (any worker in a class can run any actor in that class)
|
||||
|
||||
### Control-plane
|
||||
|
||||
* User authorization
|
||||
* User authorization
|
||||
* Good CLI
|
||||
|
||||
### Networking
|
||||
|
||||
* Actor security boundary implementation, default deny with explicit ACLs at scale with low latency. This overlaps with some of the security items (see below).
|
||||
* Policy definition: between framework (outside) and Actors, between Actors, Actor Egress.
|
||||
* Actor security boundary implementation, default deny with explicit ACLs at scale with low latency. This overlaps with some of the security items (see below).
|
||||
* Policy definition: between framework (outside) and Actors, between Actors, Actor Egress.
|
||||
* Standardized DNS Mesh: Moving to a production-grade routing format (\<id\>.actors.resources.substrate.ate.dev) for location-transparent actor-to-actor communication.
|
||||
|
||||
### Storage
|
||||
|
||||
* Decide: Is Redis/ValKey the right answer for API storage?
|
||||
* gVisor snapshot/resume optimizations
|
||||
* storage tiering (local zswap, local SSD, peer-to-peer, blob)
|
||||
* incremental snapshots
|
||||
* Support for S3 (via plugin)
|
||||
* Distinct lifecycle for “rootfs” and memory snapshots vs. “working” space. Needs API surface of where to mount.
|
||||
*
|
||||
* ConfigMaps as volumes
|
||||
* Decide: Is Redis/ValKey the right answer for API storage?
|
||||
* gVisor snapshot/resume optimizations
|
||||
* storage tiering (local zswap, local SSD, peer-to-peer, blob)
|
||||
* incremental snapshots
|
||||
* Support for S3 (via plugin)
|
||||
* Distinct lifecycle for “rootfs” and memory snapshots vs. “working” space. Needs API surface of where to mount.
|
||||
*
|
||||
* ConfigMaps as volumes
|
||||
* Data locality in scheduling (needs to expose per-node available snapshots via API)
|
||||
|
||||
### Security
|
||||
|
||||
* Goal of two security boundaries between mutually untrusted actors that share the same Kubernetes node.
|
||||
* Secure mTLS authentication and authorization between all system components.
|
||||
* Secure actor-to-actor authentication and authorization policy that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
|
||||
* Ability to specify network policy for individual actors, including L7 protocol-aware filtering rules, that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
|
||||
* Credential injection, including actor identity, via proxies to eliminate exposure of cryptographic keys and bearer tokens to actors.
|
||||
* Audit logging on API and lifecycle operations.
|
||||
* Sandbox integrations for threat detection telemetry.
|
||||
* Harden actor networking to further isolate from the surrounding node (e.g. with current networking
|
||||
* Goal of two security boundaries between mutually untrusted actors that share the same Kubernetes node.
|
||||
* Secure mTLS authentication and authorization between all system components.
|
||||
* Secure actor-to-actor authentication and authorization policy that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
|
||||
* Ability to specify network policy for individual actors, including L7 protocol-aware filtering rules, that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
|
||||
* Credential injection, including actor identity, via proxies to eliminate exposure of cryptographic keys and bearer tokens to actors.
|
||||
* Audit logging on API and lifecycle operations.
|
||||
* Sandbox integrations for threat detection telemetry.
|
||||
* Harden actor networking to further isolate from the surrounding node (e.g. with current networking
|
||||
* Support for additional sandboxing technologies beyond gVisor, including at least one flavor of microVM.
|
||||
|
||||
### Observability
|
||||
|
||||
* Session-Aware Telemetry Correlation: Automated OTLP export where all logs, metrics, and traces are natively tagged with the Substrate Actor ID and worker ID.
|
||||
* Session-Aware Telemetry Correlation: Automated OTLP export where all logs, metrics, and traces are natively tagged with the Substrate Actor ID and worker ID.
|
||||
* Prometheus metrics
|
||||
|
||||
### Performance and Reliability
|
||||
|
||||
* Representative workloads/traffic patterns
|
||||
* Provisioning load test and benchmarking compute/infrastructure
|
||||
* Storage and visualization for benchmark results
|
||||
* Integrate debugging into load tests
|
||||
* State Store Scale: Horizontal sharding support (via Redis Hash Tags) to enable management of 1M+ concurrent actors.
|
||||
* Representative workloads/traffic patterns
|
||||
* Provisioning load test and benchmarking compute/infrastructure
|
||||
* Storage and visualization for benchmark results
|
||||
* Integrate debugging into load tests
|
||||
* State Store Scale: Horizontal sharding support (via Redis Hash Tags) to enable management of 1M+ concurrent actors.
|
||||
* Disk-Only Resume Policy: Support for cost-optimized hibernation where only the filesystem state is preserved, skipping the RAM restore for stateless or "cold" start-capable agents.
|
||||
|
||||
### Testing
|
||||
|
||||
* Support for local development and testing on KinD clusters.
|
||||
* Define sufficient test matrices (what container system, clouds, other configuration knobs should be tested)
|
||||
* Burn-in testing to detect memory/other resource leaks
|
||||
* Support for local development and testing on KinD clusters.
|
||||
* Define sufficient test matrices (what container system, clouds, other configuration knobs should be tested)
|
||||
* Burn-in testing to detect memory/other resource leaks
|
||||
* Mock LLMs/other dependencies
|
||||
|
||||
### Integrations
|
||||
|
||||
* Tight integration with Agent Executor for deploying AX on Kubernetes.
|
||||
* Agent Development Kit (ADK) Native Support: Developing first-class bindings for ADK, allowing developers to build stateful agents that natively leverage Substrate’s lifecycle management and persistent working memory.
|
||||
* LangChain Remote Execution Provider: A dedicated provider for LangChain to run complex, long-running agent tools in durable, sandboxed environments.
|
||||
* Native MCP Server Hosting: Built-in support for deploying Model Context Protocol (MCP) servers as managed Substrate Actors, creating a secure tool ecosystem for any LLM.
|
||||
* Actor-to-Actor (A2A) Calling Model: Standardized protocol for actors to discover and call other actors within the mesh via the gateway.
|
||||
* Tight integration with Agent Executor for deploying AX on Kubernetes.
|
||||
* Agent Development Kit (ADK) Native Support: Developing first-class bindings for ADK, allowing developers to build stateful agents that natively leverage Substrate’s lifecycle management and persistent working memory.
|
||||
* LangChain Remote Execution Provider: A dedicated provider for LangChain to run complex, long-running agent tools in durable, sandboxed environments.
|
||||
* Native MCP Server Hosting: Built-in support for deploying Model Context Protocol (MCP) servers as managed Substrate Actors, creating a secure tool ecosystem for any LLM.
|
||||
* Actor-to-Actor (A2A) Calling Model: Standardized protocol for actors to discover and call other actors within the mesh via the gateway.
|
||||
* Native MCP Tool Hosting: Ability to define and deploy standard Model Context Protocol (MCP) servers as managed Substrate Actors, providing a plug-and-play ecosystem for agentic tools.
|
||||
|
||||
### Operability
|
||||
|
||||
* Control-plane upgrades
|
||||
* Worker upgrades
|
||||
* Gvisor upgrades
|
||||
* API schema changes
|
||||
* Control-plane upgrades
|
||||
* Worker upgrades
|
||||
* Gvisor upgrades
|
||||
* API schema changes
|
||||
* Plan for growth from small to large (esp. wrt CP sharding)
|
||||
|
||||
## Additional ideas we are thinking about
|
||||
|
||||
### Actor Management (Compute)
|
||||
|
||||
* Vertical worker autoscaling and IPPU (e.g. retain the memory/CPU config of the previous snapshot and update the worker pod when we assign an actor)
|
||||
* Actor-\>worker selectors, taints, etc.
|
||||
* Automated Garbage Collection: Background cleanup of idle actors based on configurable TTL
|
||||
* Vertical worker autoscaling and IPPU (e.g. retain the memory/CPU config of the previous snapshot and update the worker pod when we assign an actor)
|
||||
* Actor-\>worker selectors, taints, etc.
|
||||
* Automated Garbage Collection: Background cleanup of idle actors based on configurable TTL
|
||||
|
||||
### Storage
|
||||
|
||||
* Policies to drive retention of old snapshots and controller(s) to do automated cleanups
|
||||
* Peer-to-peer OCI and snapshot sharing
|
||||
* K8s “projected” volumes, other than ConfigMap
|
||||
* Storage model supported by external storage plugins (via CSI drivers?).
|
||||
* Policies to drive retention of old snapshots and controller(s) to do automated cleanups
|
||||
* Peer-to-peer OCI and snapshot sharing
|
||||
* K8s “projected” volumes, other than ConfigMap
|
||||
* Storage model supported by external storage plugins (via CSI drivers?).
|
||||
* Shared writeable storage across actors (e.g. an NFS volume).
|
||||
|
||||
### Security
|
||||
|
||||
* Ability for actors to delegate downscoped rights to children and peers.
|
||||
* Ability for actors to delegate downscoped rights to children and peers.
|
||||
* Image-pull credentials
|
||||
|
||||
### Performance and Reliability
|
||||
@@ -139,4 +139,4 @@ Below is a collection of finer-grained efforts which we believe align with the a
|
||||
|
||||
### Administrative
|
||||
|
||||
* Ability to run multiple substrates in a single cluster
|
||||
* Ability to run multiple substrates in a single cluster
|
||||
|
||||
@@ -9,7 +9,7 @@ We want to resolve requests for <actor id>.actors.resources.substrate.ate.dev to
|
||||
Cluster resources:
|
||||
|
||||
* Deployment `ate-system:dns`. Label: app=dns
|
||||
* Service `ate-system:dns`.
|
||||
* Service `ate-system:dns`.
|
||||
* ConfigMap `ate-system:dns`.
|
||||
|
||||
These are defined in manifests/ate-install/atenet-dns.yaml.
|
||||
|
||||
Reference in New Issue
Block a user