Remove trailing space from md files

This commit is contained in:
Tim Hockin
2026-05-31 19:45:36 -07:00
committed by Bowei Du
parent f353ac79ab
commit d9773e7d03
6 changed files with 72 additions and 72 deletions
+1 -1
View File
@@ -17,7 +17,7 @@ The AGENTS.md file at the root of the project may include the following sections
- Testing instructions
- Security considerations
AGENTS.md files in subfolders may be even more concise and specific to those subfolders.
AGENTS.md files in subfolders may be even more concise and specific to those subfolders.
## Maximize AGENTS.md Performance
+1 -1
View File
@@ -109,7 +109,7 @@ kubectl ate get workers
| `ASSIGNED ACTOR` | If `STATUS=ASSIGNED`, the actor reference `<namespace>/<template>/<actor-id>`. |
### Actor Lifecycle
Manage the execution state of your workloads.
Manage the execution state of your workloads.
*(Note: Actors are identified by a user-provided ID, which must be a valid DNS-1123 label)*
```bash
+2 -2
View File
@@ -58,7 +58,7 @@ metadata:
spec:
runsc:
amd64:
# Note: These values are from the 2026-05-19 nightly.
# Note: These values are from the 2026-05-19 nightly.
# For the latest verified versions, see: demos/counter/counter.yaml.tmpl
url: "gs://gvisor/releases/nightly/2026-05-19/x86_64/runsc"
sha256Hash: "a397be1abc2420d26bce6c70e6e2ff96c73aaaab929756c56f5e2089ea842b63"
@@ -159,7 +159,7 @@ Workloads can exchange their ephemeral Kubernetes credentials for stable **Sessi
## 7. Framework & Ecosystem Integration
Agent Substrate is designed to be the foundational execution layer for any agentic framework.
Agent Substrate is designed to be the foundational execution layer for any agentic framework.
### Agent Development Kit (ADK)
Substrate provides native support for ADK-compatible identities. Workloads can use the `SessionIdentity` service to mint JWTs that align with ADK's security model, ensuring seamless integration with ADK-managed tools and memory.
+5 -5
View File
@@ -293,7 +293,7 @@ The node-level subsystem manages the physical execution of sandboxes and the mov
Handles session-aware routing and automatic re-animation.
* **Uniform DNS Mesh**: Substrate provides a location-transparent actor discovery scheme via a global DNS suffix (`<id>.actors.resources.substrate.ate.dev`).
* **Uniform DNS Mesh**: Substrate provides a location-transparent actor discovery scheme via a global DNS suffix (`<id>.actors.resources.substrate.ate.dev`).
* **Routing**: The `atenet` router (powered by Envoy and an External Processing server) intercepts traffic destined for the mesh. It extracts the actor ID from the `Host` header, queries the Control Plane to determine the actor's current location, and triggers a `ResumeActor` workflow if the session is currently suspended.
@@ -358,14 +358,14 @@ collected. The garbage collection process is not implemented yet.
Agent Substrate distinguishes between two types of state, which are currently
captured together in a single versioned snapshot:
1. **Memory Snapshot**: The exact RAM state of the process.
1. **Memory Snapshot**: The exact RAM state of the process.
2. **Working Volume (Disk)**: The files written to the container's writable
layer (the "working memory").
layer (the "working memory").
In the current implementation, both memory and disk states are tied to the
specific version of the code (ActorTemplate). This ensures strict consistency
during resumption.
during resumption.
Snapshots are stored durably in **Google Cloud Storage (GCS)**. This model
allows the physical compute resources in the `WorkerPool` to be fully reclaimed
@@ -385,7 +385,7 @@ Agent Substrate is built on a **Defense-in-Depth** model:
versions.
* **Request Authorization**: The system currently performs **Identity-Aware
Routing** by utilizing a uniform DNS routing scheme
Routing** by utilizing a uniform DNS routing scheme
(`<actor id>.actors.resources.substrate.ate.dev`)
at the gateway to extract and validate actor identifiers from incoming traffic. This
ensures requests are only routed to recognized, registered actors.
+62 -62
View File
@@ -6,12 +6,12 @@ Agent Substrate is a Kubernetes-based system for building multitenant applicatio
This project is trying to move very quickly to find the right set of capabilities for the ever-changing agentic workloads market. Below are our priorities. Any efforts which are not aligned with these priorities should probably be deferred.
1. Pinning down architectural decisions which influence the rest of these priorities.
2. Performance and reliability of wakeups, including suspend/resume and snapshot management.
3. A fast, scalable control-plane, including the storage layer, reliability, and security.
4. Useful identity, such that policies can be written in terms of identity.
5. Useful policy, including how users can specify ingress, egress, and peer networking.
6. Runtime modularity, such that different kinds of sandboxes (gVisor, microVMs, etc) can be used.
1. Pinning down architectural decisions which influence the rest of these priorities.
2. Performance and reliability of wakeups, including suspend/resume and snapshot management.
3. A fast, scalable control-plane, including the storage layer, reliability, and security.
4. Useful identity, such that policies can be written in terms of identity.
5. Useful policy, including how users can specify ingress, egress, and peer networking.
6. Runtime modularity, such that different kinds of sandboxes (gVisor, microVMs, etc) can be used.
Below is a collection of finer-grained efforts which we believe align with the above.
@@ -19,114 +19,114 @@ Below is a collection of finer-grained efforts which we believe align with the a
### Actor Management (Compute)
* Actor versioning via ActorTemplate (or new ActorDeployment API), with support for A/B testing and rollout strategies.
* Namespaces or a similar concept for grouping related actors together for both management and convenience of writing actor-to-actor authorization policy.
* Actor Forking/Cloning: Ability to branch a new logical actor from an existing checkpoint (the 'State Root') to support complex agent reasoning paths.
* Worker horizontal autoscaling: Ability to rapidly scale up nodes and warm Pods to meet actor demand.
* Clarify the actor lifecycle, including what data is retained across which events (e.g. gvisor upgrade \-\> loss of memory snapshot)
* Decide: What mode(s) of actor runs do we support? Can we prioritize and phase this? Assuming we always persist “working” data
* Clean start from OCI for every activation
* Resume from golden for every activation
* Persist rootfs (when possible), clean binary start for every activation
* Actor versioning via ActorTemplate (or new ActorDeployment API), with support for A/B testing and rollout strategies.
* Namespaces or a similar concept for grouping related actors together for both management and convenience of writing actor-to-actor authorization policy.
* Actor Forking/Cloning: Ability to branch a new logical actor from an existing checkpoint (the 'State Root') to support complex agent reasoning paths.
* Worker horizontal autoscaling: Ability to rapidly scale up nodes and warm Pods to meet actor demand.
* Clarify the actor lifecycle, including what data is retained across which events (e.g. gvisor upgrade \-\> loss of memory snapshot)
* Decide: What mode(s) of actor runs do we support? Can we prioritize and phase this? Assuming we always persist “working” data
* Clean start from OCI for every activation
* Resume from golden for every activation
* Persist rootfs (when possible), clean binary start for every activation
* Persist rootfs and memory (when possible), full process resume
### Worker management
* Worker class (e.g. free-tier on cheap hardware, premium on good hardware)
* Worker class (e.g. free-tier on cheap hardware, premium on good hardware)
* Fungible worker pools (any worker in a class can run any actor in that class)
### Control-plane
* User authorization
* User authorization
* Good CLI
### Networking
* Actor security boundary implementation, default deny with explicit ACLs at scale with low latency. This overlaps with some of the security items (see below).
* Policy definition: between framework (outside) and Actors, between Actors, Actor Egress.
* Actor security boundary implementation, default deny with explicit ACLs at scale with low latency. This overlaps with some of the security items (see below).
* Policy definition: between framework (outside) and Actors, between Actors, Actor Egress.
* Standardized DNS Mesh: Moving to a production-grade routing format (\<id\>.actors.resources.substrate.ate.dev) for location-transparent actor-to-actor communication.
### Storage
* Decide: Is Redis/ValKey the right answer for API storage?
* gVisor snapshot/resume optimizations
* storage tiering (local zswap, local SSD, peer-to-peer, blob)
* incremental snapshots
* Support for S3 (via plugin)
* Distinct lifecycle for “rootfs” and memory snapshots vs. “working” space. Needs API surface of where to mount.
*
* ConfigMaps as volumes
* Decide: Is Redis/ValKey the right answer for API storage?
* gVisor snapshot/resume optimizations
* storage tiering (local zswap, local SSD, peer-to-peer, blob)
* incremental snapshots
* Support for S3 (via plugin)
* Distinct lifecycle for “rootfs” and memory snapshots vs. “working” space. Needs API surface of where to mount.
*
* ConfigMaps as volumes
* Data locality in scheduling (needs to expose per-node available snapshots via API)
### Security
* Goal of two security boundaries between mutually untrusted actors that share the same Kubernetes node.
* Secure mTLS authentication and authorization between all system components.
* Secure actor-to-actor authentication and authorization policy that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
* Ability to specify network policy for individual actors, including L7 protocol-aware filtering rules, that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
* Credential injection, including actor identity, via proxies to eliminate exposure of cryptographic keys and bearer tokens to actors.
* Audit logging on API and lifecycle operations.
* Sandbox integrations for threat detection telemetry.
* Harden actor networking to further isolate from the surrounding node (e.g. with current networking
* Goal of two security boundaries between mutually untrusted actors that share the same Kubernetes node.
* Secure mTLS authentication and authorization between all system components.
* Secure actor-to-actor authentication and authorization policy that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
* Ability to specify network policy for individual actors, including L7 protocol-aware filtering rules, that can be deployed in-band with actor lifecycle, maintaining low latency deployment.
* Credential injection, including actor identity, via proxies to eliminate exposure of cryptographic keys and bearer tokens to actors.
* Audit logging on API and lifecycle operations.
* Sandbox integrations for threat detection telemetry.
* Harden actor networking to further isolate from the surrounding node (e.g. with current networking
* Support for additional sandboxing technologies beyond gVisor, including at least one flavor of microVM.
### Observability
* Session-Aware Telemetry Correlation: Automated OTLP export where all logs, metrics, and traces are natively tagged with the Substrate Actor ID and worker ID.
* Session-Aware Telemetry Correlation: Automated OTLP export where all logs, metrics, and traces are natively tagged with the Substrate Actor ID and worker ID.
* Prometheus metrics
### Performance and Reliability
* Representative workloads/traffic patterns
* Provisioning load test and benchmarking compute/infrastructure
* Storage and visualization for benchmark results
* Integrate debugging into load tests
* State Store Scale: Horizontal sharding support (via Redis Hash Tags) to enable management of 1M+ concurrent actors.
* Representative workloads/traffic patterns
* Provisioning load test and benchmarking compute/infrastructure
* Storage and visualization for benchmark results
* Integrate debugging into load tests
* State Store Scale: Horizontal sharding support (via Redis Hash Tags) to enable management of 1M+ concurrent actors.
* Disk-Only Resume Policy: Support for cost-optimized hibernation where only the filesystem state is preserved, skipping the RAM restore for stateless or "cold" start-capable agents.
### Testing
* Support for local development and testing on KinD clusters.
* Define sufficient test matrices (what container system, clouds, other configuration knobs should be tested)
* Burn-in testing to detect memory/other resource leaks
* Support for local development and testing on KinD clusters.
* Define sufficient test matrices (what container system, clouds, other configuration knobs should be tested)
* Burn-in testing to detect memory/other resource leaks
* Mock LLMs/other dependencies
### Integrations
* Tight integration with Agent Executor for deploying AX on Kubernetes.
* Agent Development Kit (ADK) Native Support: Developing first-class bindings for ADK, allowing developers to build stateful agents that natively leverage Substrate’s lifecycle management and persistent working memory.
* LangChain Remote Execution Provider: A dedicated provider for LangChain to run complex, long-running agent tools in durable, sandboxed environments.
* Native MCP Server Hosting: Built-in support for deploying Model Context Protocol (MCP) servers as managed Substrate Actors, creating a secure tool ecosystem for any LLM.
* Actor-to-Actor (A2A) Calling Model: Standardized protocol for actors to discover and call other actors within the mesh via the gateway.
* Tight integration with Agent Executor for deploying AX on Kubernetes.
* Agent Development Kit (ADK) Native Support: Developing first-class bindings for ADK, allowing developers to build stateful agents that natively leverage Substrate’s lifecycle management and persistent working memory.
* LangChain Remote Execution Provider: A dedicated provider for LangChain to run complex, long-running agent tools in durable, sandboxed environments.
* Native MCP Server Hosting: Built-in support for deploying Model Context Protocol (MCP) servers as managed Substrate Actors, creating a secure tool ecosystem for any LLM.
* Actor-to-Actor (A2A) Calling Model: Standardized protocol for actors to discover and call other actors within the mesh via the gateway.
* Native MCP Tool Hosting: Ability to define and deploy standard Model Context Protocol (MCP) servers as managed Substrate Actors, providing a plug-and-play ecosystem for agentic tools.
### Operability
* Control-plane upgrades
* Worker upgrades
* Gvisor upgrades
* API schema changes
* Control-plane upgrades
* Worker upgrades
* Gvisor upgrades
* API schema changes
* Plan for growth from small to large (esp. wrt CP sharding)
## Additional ideas we are thinking about
### Actor Management (Compute)
* Vertical worker autoscaling and IPPU (e.g. retain the memory/CPU config of the previous snapshot and update the worker pod when we assign an actor)
* Actor-\>worker selectors, taints, etc.
* Automated Garbage Collection: Background cleanup of idle actors based on configurable TTL
* Vertical worker autoscaling and IPPU (e.g. retain the memory/CPU config of the previous snapshot and update the worker pod when we assign an actor)
* Actor-\>worker selectors, taints, etc.
* Automated Garbage Collection: Background cleanup of idle actors based on configurable TTL
### Storage
* Policies to drive retention of old snapshots and controller(s) to do automated cleanups
* Peer-to-peer OCI and snapshot sharing
* K8s “projected” volumes, other than ConfigMap
* Storage model supported by external storage plugins (via CSI drivers?).
* Policies to drive retention of old snapshots and controller(s) to do automated cleanups
* Peer-to-peer OCI and snapshot sharing
* K8s “projected” volumes, other than ConfigMap
* Storage model supported by external storage plugins (via CSI drivers?).
* Shared writeable storage across actors (e.g. an NFS volume).
### Security
* Ability for actors to delegate downscoped rights to children and peers.
* Ability for actors to delegate downscoped rights to children and peers.
* Image-pull credentials
### Performance and Reliability
@@ -139,4 +139,4 @@ Below is a collection of finer-grained efforts which we believe align with the a
### Administrative
* Ability to run multiple substrates in a single cluster
* Ability to run multiple substrates in a single cluster
+1 -1
View File
@@ -9,7 +9,7 @@ We want to resolve requests for <actor id>.actors.resources.substrate.ate.dev to
Cluster resources:
* Deployment `ate-system:dns`. Label: app=dns
* Service `ate-system:dns`.
* Service `ate-system:dns`.
* ConfigMap `ate-system:dns`.
These are defined in manifests/ate-install/atenet-dns.yaml.