30 Commits
Author SHA1 Message Date
Benjamin Elderanddependabot[bot] 51a71cb365 upgrade otel log to 0.21 (#1958)
fixes https://github.com/agent-substrate/substrate/pull/1952

security bump stuck by breaking changes.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-28 22:26:40 +00:00
Shruti Nair b29f97778a Initialize OpenFGA server on ATE API Server bootstrap. (#1670)
Working on https://github.com/agent-substrate/substrate/issues/1563

## Summary
Initializes an embedded OpenFGA server within `ateapi` backed by
PostgreSQL and updates the top-level authorization scope.

## Key Changes
- **Rename scope to `global`**: Renamed `type cluster` and
`parent_cluster` to `type global` and `parent_global`, adding
`can_set_policy` and `can_get_policy` to `global`.
- **Embedded OpenFGA server (`internal/authz/server.go`)**:
  - Compiles `model.fga` into OpenFGA's protobuf representation.
  - Runs OpenFGA PostgreSQL schema migrations (`goose_db_version`).
- Serializes startup via `pg_advisory_lock` to avoid multi-replica
races.
- Connects OpenFGA's PostgreSQL adapter with a new dedicated
`*pgxpool.Pool`.
- Idempotently creates or reuses the `"substrate"` store and
authorization model.
- **Service wiring (`cmd/ateapi`)**: Exposes `Pool()` on
`*atepg.Persistence` and initializes `authz.NewServer` on bootstrap.

## Verification
- `openfga model test`: 103/103 checks passing (`model_test.fga.yaml`).
- `go test -mod=mod ./internal/authz/...`: Passes against PostgreSQL
(migrations, tuple writes, permission checks, and restart idempotency).
2026-09-18 23:30:41 +00:00
Krisztian F 5f7c690108 (feat): Actor lifecycle events over otlp (#1658)
## What this does

ateapi writes a record every time an actor changes state. Until now
those records only went to the pod's stdout, and nothing reads stdout.
This sends the same records to a collector as OTLP log events.

Actor name and uid cannot be metric labels (too many values), and traces
are sampled at 1%. So these records are the only way to answer "what
state is this actor in, and since when".

## Changes

- `serverboot.InitLogging` sets up a LoggerProvider, next to the
existing tracer and meter ones.
- New `internal/actorevent` package builds the log records.
- ateapi emits at the two places that already write the stdout records.
- Two event names: `ate.actor.state_changed` and `ate.actor.crashed`.
- Both names are registered in `docs/metrics/registry/events.yaml`, so
`make verify` checks them.
- kind gets a logs pipeline and a count connector. The e2e suite reads
the counts back.
- Docs updated. `otel-collector.md` said substrate has no
LoggerProvider, which is no longer true.

## Opt-in

`OTEL_LOGS_EXPORTER` defaults to `none`. Only the kind overlay sets it
to `otlp`. The base ConfigMap is untouched, so no deployed environment
changes when this merges.

## Notes on the design

- **No slog bridge.** Only two call sites emit these records, so
emitting twice costs two lines. A bridge would also send every ateapi
log over the wire, could not set the event name, and would loop, because
SDK export errors are logged through slog.
- **Batching processor, not the simple one.** These records sit on the
actor resume path. A processor that exports inside the emit call would
add a blocking gRPC call there, so a slow collector would become control
plane latency.
- **Two event names, not one per state.** `ate.actor.state` already says
which transition happened. A crash gets its own name because it carries
two extra attributes and a higher severity.
- **Both copies are kept on purpose.** No collector in this repo reads
pod stdout, so nothing is duplicated today. `kubectl logs` keeps
working. If a filelog agent is ever added, drop one of the two. The
escape hatch is written down in `docs/metrics/substrate.yaml`.

## Dependencies

Adds `otel/log`, `otel/sdk/log` and `otlploggrpc`, all pinned at
v0.20.0. That is the release that matches the pinned `otel v1.44.0`.
v0.21.0 would pull the core modules to v1.45.0, which this change does
not need. The logs API has a
v1.47.0 release candidate upstream, so it is on its way to stable.

## Testing

- Unit tests for the exporter resolver, the record builder, and both
ateapi emit sites.
- The record builder test checks the attribute set matches what the
event name declares, in both directions.
- The ateapi tests check the OTLP record carries the same attributes as
the stdout record.
- Ran end to end on kind. Records arrive with the right event name,
severity, attributes, and with trace context on the record's own fields
rather than as attributes.
- Checked the off state too. With `OTEL_LOGS_EXPORTER` removed, the
collector receives no log records and stdout is unchanged.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-18 18:11:12 +00:00
Tim Hockin 2fcfa64a73 Drop k8s third_party deps now that HEAD is updated (#1480)
The changes we need are live on HEAD of the k8s published repos. It was
a bit of a journey, since Go fought me all the way.

---

### Drop our third_party fork of k8s deps in tools
    
    The changes we need are released now (sort of - on HEAD anyway).

---

### Bump k8s.io/streaming to v0.37.0 (not rc) in tools

---

### Bump codegen deps to HEAD in tools
    
    GOPROXY=direct go get \
        k8s.io/apimachinery@master \
        k8s.io/code-generator@master
    
This pins them to HEAD of master. Go is terrible here: The HEAD is not
    actually tagged, so Go just uses the next "reachable" tag which is
    v0.36.0-alpha.  The datestamp is correct, though.

---

### Drop our third_party fork of k8s deps in root
    
    The changes we need are released now (sort of - on HEAD anyway).

---

### Bump k8s deps to v0.37.0 (not rc) in root

---

### Bump apimachinery dep in root to HEAD in root
    
GOPROXY=direct go mod edit -replace
k8s.io/apimachinery=k8s.io/apimachinery@master
    GOPROXY=direct go mod tidy
    GOPROXY=direct go mod vendor
    GOPROXY=direct go mod tidy
    
    This approach (-replace) is needed because Go is horrible here.
    
    The master branch of k8s.io/apimachinery is not tagged, per se, but
there is an OLDER tag which is "reachable" from HEAD. So go helpfully
decides to use that (v0.36.0-alpha.2). If we just `go get ... @master`
it works for that dep (pinned to the right date) but then it looks at
    transitive deps.  Because the tag seems to be 0.36 (older), it
recalculates all the OTHER dependencies and downgrades a whole tangle of
    things to versions that match 0.36, but we are ACTUALLY on 0.37+.
    
    This was the only approach that I (and Gemini) could find.  Blech.

---

### Run updated codegens

---

### Use DV's new maxBytes capability for `[]byte`

    Removes 1 custom.
2026-09-04 17:02:12 -07:00
Haven Xia 07f08a3ab1 [vuln] Bump golang.org/x/crypto to v0.56.0 (#1412)
Vuln check failed at main
https://github.com/agent-substrate/substrate/actions/runs/33678429903

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-09-02 13:57:10 -07:00
Jeremy Alvis f100c27286 ateapi: add versioned PostgreSQL schema migrations (#1196)
> [!WARNING]
> Recreate PostgreSQL databases from earlier development builds.

## Summary

This change replaces startup schema setup with embedded, versioned SQL
migrations.

`ateapi` uses Goose to apply migrations before readiness. Goose stores
one ledger record for each applied migration.

Goose runs each migration and inserts its ledger record in one
PostgreSQL transaction.

Closes #901.

Based on this design:
https://docs.google.com/document/d/13ixDKRoAIFXeLxS-_1nikNcAy76ca8m0eVgobfib93E/edit?usp=sharing

## Migration behavior

`ateapi` gets a session advisory lock for the configured schema before
it applies pending migrations.

One replica applies migrations while other replicas wait. Goose reads
the ledger again after it gets the lock.

If a migration fails, PostgreSQL rolls back its SQL and ledger record.
Earlier successful migrations remain applied and recorded.

Kubernetes restarts the failed replica. The next startup resumes from
the first migration without a ledger record.

## Changes

- Add Goose and a per-migration ledger.
- Replace the initial up and down files with one transactional, up-only
migration.
- Keep migration 1 aligned with the current schema, including actor
egress policy storage.
- Remove existence guards and explicit transaction statements from the
baseline migration.
- Apply all pending migrations before `ateapi` becomes ready.
- Serialize each migration run with a PostgreSQL session advisory lock.
- Start without changes when the database schema is current or ahead.
- Reject application tables that do not have a migration ledger.
- Log the starting, current, and latest versions.
- Log the applied migration count and duration.
- Retry only initial database connection failures.
- Return schema and migration errors without a retry.
- Add `--postgres-schema` and `ATE_API_POSTGRES_SCHEMA`.
- Use `public` as the default PostgreSQL schema.
- Use the configured schema for the main and watch pools.
- Restrict outbox partition maintenance to the configured schema.
- Let the installer use an external PostgreSQL database.
- Add the migration design and recovery policy to the repository.

## Migration file policy

Migration files use sequential versions and contain exactly one Goose
`Up` section.

CI rejects down migrations, nontransactional migrations, environment
substitution, explicit transaction control, and `IF NOT EXISTS` guards.

Before the first stable v1 release, developers can change or squash
migrations. Developers must recreate databases after migration history
changes.

After that release, CI rejects changes or deletions against the latest
stable release tag that contains migrations.

Goose does not store migration checksums. The binary embeds each
migration file, and release-tag checks protect released migration
history.

## Compatibility

No release includes PostgreSQL support. The `v0.0.0` release predates
the PostgreSQL backend.

Users must recreate databases from earlier PostgreSQL development
builds.

Every committed migration prefix must work with the current and previous
`ateapi` releases. This rule supports rolling upgrades and temporary
binary rollback.

A binary rollback does not roll back the database schema.

## Testing

Tests cover:

- Fresh database migration.
- Concurrent startup.
- Advisory lock waits.
- Current and ahead database schemas.
- Rejection of application tables without a migration ledger.
- Atomic rollback of a failed migration.
- Retention of earlier successful migrations.
- Resume from the failed migration after restart.
- Configured schema isolation.
- Outbox partition isolation.
- Migration file policy checks.
- Stable release migration immutability.
2026-09-02 10:56:19 -04:00
Benjamin Elder 1d58b95a7f update grpc 1.83.2 in all modules (#1385)
picks up security fixes
2026-09-01 21:37:24 -07:00
Benjamin Elder f3faf7b02b Ateom child reaper (#1293)
Necessary for #1266, but functionally useful regardless.

This eliminates the MPL go-reap dependency in favor of our own
implementation tuned to our needs.

Our package has utilities for spawning child processes without holding a
RWMutex to avoid convoying lots of concurrent activations in the future.
2026-08-28 21:08:37 -07:00
Bowei Du bac7736b84 Go-based hack scripts (#1190)
Adds ate-setup, a Go-implemented version of the hack scripts
    
    This was done with heavily machine-translated code iterated
    with testing.
2026-08-27 15:06:13 -07:00
Benjamin Elder cb72d60bdf ateom microvm: drop privileged: true, using atelet device plugin for KVM (#1254)
This is the companion to #496 

In order to grant access to `/dev/kvm` we have to either:
- use a device plugin
- use a DRA driver
- use an NRI plugin

An NRI plugin is highly privileged in it's own right for all pods on the
host and is difficult to ship portably at the moment.
DRA is promising, but not enough functionality is GA yet at our current
1.35+ target.
Device plugin fits reasonably well. We do wind up publishing an
~arbitrarily high limit, which has some cost in kubelet memory, but
otherwise is relatively clean.

This approach is what kata uses currently. Their device plugin is not
available unbundled, and we anyhow have a per-node daemonset.

atlet is taught to sniff if /dev/kvm appears on the host at all, so we
can also stop using the manually labeled nodes for microVM class and
instead schedule to the KVM + TUN devices on nodes that advertise them.

Later we can migrate to device plugin by using 

I implemented that already, but I don't think it's worth merging at the
moment. We would want consumable capacity to be on by default. We can
migrate later without changing the pod spec by using
`extendedResourceName`.

https://github.com/agent-substrate/substrate/compare/main...BenTheElder:substrate:ateom-microvm-dra

NOTE: I confirmed with upstream that device plugin is not going anywhere
despite being "v1beta1", it's GA in all but name. It won't receive new
features but we don't really need anyhow. We'll move to DRA down the
line.

---

By doing this, we can drop `privileged: true` from the uVM ateom pods.

We can also drop the `ate.dev/sandboxClass` node label hacks, reducing
friction to deploy.
2026-08-27 12:23:16 -07:00
NekoPunch 7b02f51320 deps: bump testcontainers-go to v0.44.0 (#1237)
Part of #1230

When reusing an already-running Ryuk reaper, testcontainers-go v0.43.0
waits only for its Docker port mapping. Docker exposes that port before
Ryuk is listening, so a second package can connect too early and lose
the handshake with `read ack: EOF`. That is
[testcontainers-go#3743](https://github.com/testcontainers/testcontainers-go/issues/3743);
v0.44.0 also waits for the reaper's `Started` log line
([#3761](https://github.com/testcontainers/testcontainers-go/pull/3761)).

A package whose handshake fails is not counted as a Ryuk client but its
containers still carry the shared session label, so another package
exiting can delete a database that is still in use — the failure
reported in #1230.

## Scope

This helps local runs, where the reaper stays on. **It does not fix CI
on its own**, so it is deliberately separate from #1235, which disables
the reaper for `run-tests`.

Measured on a 4-core Linux VM with cold build caches, which staggers
package start times the way CI does:

```
v0.43.0, 3 rounds   handshake failures in 3/3 rounds; one round cascaded
v0.44.0, 3 rounds   no handshake failures; database tests still unavailable in 3/3 rounds
```

Under v0.44.0 the failures move to `wait for reaper <id>: context
deadline exceeded` — a late package finds a reaper that is already
shutting down, and one readiness probe (`defaultStartupTimeout`, 60s)
outlives the whole reaper retry budget (`MaxElapsedTime`, 20s), so the
retry loop never gets a second attempt. The remaining fail-open behind
all of this is
[#3827](https://github.com/testcontainers/testcontainers-go/issues/3827),
still open upstream.

## Diff size

Four lines of `go.mod`. The rest is `go mod vendor` output: v0.44.0
pulls newer `moby/client`, `gopsutil`, and `otelhttp`, and `otelhttp`
moves `otel/semconv` from v1.39.0 to v1.41.0. Insertions and deletions
nearly cancel because most of it is a directory swap and one generated
`httpsnoop` file being merged into another.

```
go.mod, go.sum      62 lines
vendor/             61 files, 17356 +/17490 -
```

`hack/verify/go-modules.sh`, `licenses.sh`, `boilerplate.sh`, and
`gofmt.sh` all pass; `go test -race ./cmd/ateapi/...` is green with no
silently skipped database tests.

---

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-26 21:51:25 -07:00
Haven Xia d32a74e21b vuln: bump github.com/moby/go-archive to v0.3.0 to fix GO-2026-6253 (#1222)
vuln check error in my merge
https://github.com/agent-substrate/substrate/actions/runs/32926359035


- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
2026-08-25 21:12:44 -07:00
Tim Hockin a8d83ee1b5 Use Kubernetes declarative validation (#1215)
This mostly eliminates the need to hand-write validation code.

This PR is a long series of commits which add DV for most of Actor and
all of Atespace.

Here is a map to the commits:
* The first few take a dep on a new Kubernetes tag, import the code into
third_party, and apply a single patch. Because we have different Go
modules for tools, I had to do it twice. When that patch lands, we can
revert these commits, but that won't be until the 1.38 cycle in a few
months.
* The next commits slowly add DV support, so a human can review each of
them in a reasonable amount of time. The emphasis is on great test
cases.
* I made a bad choice early on as to where to generate the code into, so
I moved it. Rebasing on that was exceedingly hard, so I left it as a
move.
* This required changing update/go-generate -> update/codegen -- we need
to get the ordering of tools right, which `go generate` does not
guarantee.
* Then I added a "middle" layer called "ServiceImpl" between the RPC and
storage layers. This allows things like workflow to call the same
business logic as the RPC layer, including validation. Lots of test
fixes.
* Then I finished the Create() and Update() paths for Actor. Those
represent the "right" (or closest to) way to implement resources, and
tests for validation.

I strongly encourage reviewers to read it commit-by-commit. Rebasing
this is VERY tedious, so the sooner it lands or dies completely, the
better. Then we can start converting the rest.

`./hack/run-tool.sh validation-gen --docs` will produce some docs on the
tool and the available tags.

@laoj2 @juli4n @EItanya @HavenXia
@lalitc375 @yongruilin @jpbetz  FYI
2026-08-25 20:24:59 -07:00
Jet Chiang 60073ecd57 Replace ateredis with atepg (#940)
## Summary

A follow up to #640 where we introduced PostgreSQL as an alternative
storage backend, selected conditionally in ateapi.

- Deleted ateredis, its tests, and its dependencies
- Removed Redis backend selection and configuration so ateapi always
connects to Postgres
- Replaced Valkey resources with Postgres in the standard and Kind
deployment paths and simplified install script
- Replaced miniredis fixtures with isolated Postgres testcontainers and
added centralized helpers for seeding resources
- Renamed Redis-specific debug flush command to backend-neutral
`debug-clear-store` in CLI
- Updated comments and docs where applicable

## Benchmarking

Extensive benchmarking have been performed to evaluate Redis vs
Postgres, and results can be found in these two documents:

-
https://docs.google.com/document/d/10K0wB6aTeFkJCL4HN3NbLJCdFGoLYdhIcqkFnqFHkKc/edit?usp=sharing
-
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-25 07:23:40 -04:00
Benjamin Elder 7ae9d4a44e atelet: stream snapshot uploads to S3 in parallel parts (#1068) 2026-08-20 18:21:24 -07:00
Keith Mattix II 89b83a5c89 Support CONNECT in atenet router (#715)
added CONNECT support for atenet ingress to support arbitrary actor ports
2026-08-14 16:12:27 -07:00
Jet Chiang 4b12ce67b0 ateapi: Support Postgres as alternative persistence backend (#640)
Adds Postgres as an alternative persistence backend for ateapi as
proposed in #731. Currently this is opt-in with flags in ateapi or use
the `--store-backend=postgres` option in the install script.

Slack discussion:
https://cloud-native.slack.com/archives/C0B6M3E2J3D/p1785183769252679

Performance benchmark results for Postgres vs Redis:
https://docs.google.com/document/d/12-ko_BFHcBo_nJkx9f4B7zMbiiWKC2saGhMhZG3aQ-s/edit?usp=sharing

Benchmarks are not included in this PR, but you can reproduce the
results by following the doc above

Limitations and questions:

- This draft does not yet handle migrations, this might come as a follow
up
- It currently uses `testcontainers` for testing Postgres, but as a
result it adds a lot of dependencies

---------

Signed-off-by: Jet Chiang <pokyuen.jetchiang-ext@solo.io>
2026-08-12 14:00:10 -04:00
Elijah Rodriguez-Beltran 231a5d2059 Add CSI boilerplate 2026-08-07 18:20:22 +00:00
Krisztian F c155efd1ac feat(otel): onboard atecontroller to the OTLP path (#754)
atecontroller had no OTel at all. Dev-mode zap logger, and its
controller-runtime metrics were only available on a :8080 that we don't
scrape.

After this PR, logs go through the shared slog handler (plus a
`--log-level` flag to match the other binaries), and
controller-runtime's Prometheus registry is bridged onto the OTLP reader
so the reconcile/workqueue metrics actually reach the collector.
Filtering these out is a pipeline responsibility. Also added otelgrpc to
the ateapi client, which was untraced.

This unblocks #564 the workperpool metrics, cc @Angelawork, @JeffLuoo:
there's a working `MeterProvider` to use for
`ate.workerpool.desired_workers`/`ready_workers`.

Couple of things to mention for review:

-`InitMetricsPushOnly`, not `InitMetrics`, even though we do serve
:8080. That port is controller-runtime's own private registry, not the
global one `serverboot.metricsMux` serves, so a pull reader there would
collect into something we never expose.
- Bridge is pinned to v0.68.0 to match otelgrpc. Wanted to go to
v0.70.0, but that requires otel/sdk/metric 1.45.0 and pulls the whole
SDK up with it (406 vendor files instead of 81). Happy to do that bump
separately.
- zap/zapr fall out of go.mod since atecontroller was the last importer.
- I included an OTel collector image bump from the early 2024 (!) one to
latest, which was breaking exposing native histograms

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-06 10:25:03 -04:00
Haven Xia 1afaaf5725 Update google.golang.org/grpc to address vulnerability: GO-2026-6061 (#551)
The weekly govulncheck run on main is failing:
[GO-2026-6061](https://pkg.go.dev/vuln/GO-2026-6061) (vulnerabilities in
the xDS RBAC engine and the HTTP/2 transport server) in
google.golang.org/grpc v1.81.0.

Bump grpc to the fixed v1.82.
2026-07-27 14:51:28 -04:00
Tim Bai 6afafed093 feat(kubectl-ate): implement kubectl ate top workers command (#515) (#516)
##  Feature Implementation
This PR implements the `kubectl ate top workers` command in
`kubectl-ate` to display real-time CPU and Memory resource utilization
for worker pods.

###  Changes Included:
- **Command Structure**: Added `top.go` and `top_workers.go` under
`cmd/kubectl-ate/internal/cmd/`.
- **Supported Flags**:
- `-n, --namespace <ns>`: Scope output to a specific Kubernetes
namespace.
- `-a, --atespace <space>`: Filter worker pods hosting actors in a
specific atespace.
  - `-l, --selector <labels>`: Filter by worker pool labels.
  - `-o, --output table|json|yaml`: Output format option.
- **Metrics Client**: Added `NewMetricsClientset` to connect to K8s
`metrics.k8s.io/v1beta1`.
- **Data Join & Fallback**: Joined Substrate `ListWorkers` gRPC RPC with
K8s `PodMetrics`, gracefully displaying `metrics unavailable` if
`metrics-server` is missing/unavailable.
- **Printer & Unit Tests**: Implemented table, JSON, and YAML printers
and comprehensive unit tests.

Closes #515
2026-07-24 11:57:18 -07:00
Jaana Dogan cecf4856e7 Update x/text (#489) to address vulnerability: GO-2026-5970
More info: https://pkg.go.dev/vuln/GO-2026-5970
2026-07-22 09:41:18 -07:00
Dmitry Berkovich 228fde4949 Bump go-containerregistry to v0.21.7 to fix OOM during image extract (#464)
## Summary

Bumps `github.com/google/go-containerregistry` v0.21.5 → v0.21.7. Phase
0 of #463.

v0.21.6 fixed the memory spike in `mutate.Extract` described in #120
(upstream fix: google/go-containerregistry#2190, merged as
google/go-containerregistry@38d6e4087c): extraction is now refactored
into a per-layer `extractLayer` function, so each layer's decompressor
and HTTP response body are closed as soon as that layer is processed,
instead of all being held open via deferred `Close` calls until the
entire (multi-GB) extraction completes. atelet's pull path goes through
`memorypullcache.Fetch` → `mutate.Extract`, so this directly caps its
extract-time memory.

Note this is only interim relief: #437's unbounded cache retention is
untouched and is addressed by the redesign in #463.

## Changes

- `go-containerregistry` v0.21.5 → v0.21.7, plus required transitive
bumps: `docker/cli`, `klauspost/compress`, `golang.org/x/sync`,
`golang.org/x/sys`
- `go mod tidy` + `go mod vendor`; the
`containerd/stargz-snapshotter/estargz` vendor tree drops out (upstream
removed the estargz integration; nothing in substrate used it), along
with `vbatts/tar-split` and `mitchellh/go-homedir`

## Testing

- `go build ./...` passes
- Full `go test ./...` passes
- Verified the layer-close fix is present in the vendored
`pkg/v1/mutate/mutate.go`

Fixes #120

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-20 10:25:47 -04:00
dependabot[bot] 8a63e4b585 build(deps): bump golang.org/x/crypto from 0.51.0 to 0.52.0
Bumps [golang.org/x/crypto](https://github.com/golang/crypto) from 0.51.0 to 0.52.0.
- [Commits](https://github.com/golang/crypto/compare/v0.51.0...v0.52.0)

---
updated-dependencies:
- dependency-name: golang.org/x/crypto
  dependency-version: 0.52.0
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-08 14:54:20 -07:00
Max Smythe 354a53a614 Switch to golang for glutton locust worker using boomer 2026-06-30 14:32:47 -07:00
Benjamin Elder 7abf9343bf vendor: add go-toml/v2 and ttrpc
go-toml/v2 parses the kata configuration.toml; ttrpc (and its log dependency)
backs the kata-agent client the micro-VM runtime drives. Includes their licenses.
2026-06-25 13:44:25 -07:00
Eitan Yarmush 504cfd21ab Replace actor eth0 move with veth networking 2026-06-10 22:37:09 -07:00
Benjamin Elder 210b4395c3 Bump golang.org/x/net to v0.55.0 to fix govulncheck failures (#143)
Fixes GO-2026-5026 and GO-2026-4918 in golang.org/x/net, reached via the
kubectl-ate log streaming path through the Kubernetes client. Also pulls
forward sibling golang.org/x/{crypto,sys,term,text}.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR

Follow-up #97
2026-06-02 14:13:22 -04:00
John Howard e1ff0ad5ff Handle Kubernetes 1.36 pod certificate requests (#8)
Kubernetes 1.36 stopped putting the subject key in spec.pkixPublicKey
for PodCertificateRequest and now sends a stub PKCS#10 request instead.
The old controller parsed the deprecated field unconditionally, so new
PCRs had pkixPublicKey=null and failed with an ASN.1 'sequence
truncated' error before any certificate could be issued.

Add a shared helper that extracts the public key from
spec.stubPKCS10Request first and falls back to spec.pkixPublicKey for
older clusters. Both pod identity and service DNS signers use that path,
so their behavior stays consistent across Kubernetes versions. Bump the
Kubernetes dependencies to expose the new field and update
controller-runtime to the matching generation.

Without this, the `ate-system` pod all fail to start due to missing
certificates.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-05-20 11:05:13 -07:00
+7 af3c65088e Initial commit of Agent Substrate
This is the initial release of the Agent Substrate.

Agent substrate is a system built on top of Kubernetes which manages agent-like
workloads to achieve higher scale and efficiency than Kubernetes alone can
offer, with lower latency.  It builds on top of Kubernetes features like
Pods and Pod autoscaling, but takes the Kubernetes control-plane out of the
critical path to achieve lower latency.

It can run on any Kubernetes cluster and does not inhibit “regular” use of
Kubernetes in any way. Kubernetes provides the infrastructure provisioning and
management for all types of workloads, while Agent Substrate provides
agent-specific scheduling and control.

At its core, Agent Substrate maps a larger set of “actors” (applications such
as agents) onto a smaller set of ready “workers” (Kubernetes Pods), relying on
the fact that agent-like applications tend to be idle most of the time to
achieve heavy multiplexing.  It provides functionality to manage an actor’s
lifecycle (e.g. create/destroy, suspend/resume), to assign actors to workers in real
time, and to route incoming traffic to them.

Agent Substrate is intended to be a low-opinion system.  The workloads it
manages don't have to be literal AI agents, but those are the best example of
the kind of applications it is designed for.  It is not an SDK for building
agents, but rather a system for running them at scale.

Agent Substrate is currently in VERY early development.  It is not ready for
production use, and the APIs are almost guaranteed to change.  We are not
making any guarantees about backward compatibility at this stage, and
everything in this project may be changed.

Co-authored-by: Alex Bulankou <alexbu@google.com>
Co-authored-by: Benjamin Elder <bentheelder@google.com>
Co-authored-by: Bowei Du <bowei@google.com>
Co-authored-by: Dmitry Berkovich <dberkov@google.com>
Co-authored-by: Fabricio Voznika <fvoznika@google.com>
Co-authored-by: Francisco Cabrera <fclieutier@google.com>
Co-authored-by: Haven Xia <haoyuxia@google.com>
Co-authored-by: Julian Gutierrez Oschmann <juliangut@google.com>
Co-authored-by: Kevin Steuer <ksteuer@google.com>
Co-authored-by: Max Smythe <smythe@google.com>
Co-authored-by: Maya Wang <mymaya@google.com>
Co-authored-by: Michael Taufen <mtaufen@google.com>
Co-authored-by: Shruti Nair <shrutinair@google.com>
Co-authored-by: Taahir Ahmed <taahm@google.com>
Co-authored-by: Tim Hockin <thockin@google.com>
Co-authored-by: Zoe Zhao <zoezhao@google.com>
2026-05-19 16:57:14 -07:00