Flesh out the agentgateway implementation as a router and egress PEP.
This also changes the atunnel CONNECT port to do 8443. Most of this is
agentgateway contained, but the one other system we touch is localca;
agentgateway's MITM support requires the TLS cert and key in the secret
instead of the JSON pool (we don't rely on sdsmint).
- [X] Tests pass
- [X] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Keith Mattix II <keithmattix2@gmail.com>
certificates.k8s.io/v1beta1 podcertificaterequests and
clustertrustbundles cannot be enabled in place on an existing cluster -
the update is accepted but the APIs never become served, and the
install hangs waiting for ClusterTrustBundles. Warn in the create
cluster docs, show the bring-your-own-cluster flag, and name the
symptom.
On a fresh GCP project nothing documented which IAM bindings atelet
needs: setup-gcp creates them silently, and anyone who cannot run the
tool with project-level IAM permissions - or needs to audit what it
did - had to read cmd/iam.go and cmd/bucket.go. Spell out the exact
members, roles, and resources, the Workload Identity prerequisites,
and the gcloud equivalents, and point the README quickstart at it.
## What this PR does:
Adds large-memory benchmark suites: glutton actors that hold a resident
1–2Gi working set, so the suspend/resume path is measured at realistic
application sizes on both runtimes — {1Gi, 2Gi} × {gvisor, microvm},
tracked.
The boomer GluttonUser fills each actor to a configured target via
chunked `WriteRAM` calls (64Mi per keyed allocation — the proto size
field is int32) after resume and **before the first suspend**, so every
snapshot from cycle one onward carries the full working set. `WriteRAM`
writes incompressible random bytes, so snapshots are genuinely
target-sized rather than zstd-compressing away (a zero-filled working
set compresses ~35:1 and would shrink the upload/download phases to a
few MB).
## How it's configured
The target flows through the established boomer dynconfig channel, like
the durdir knobs: `--mem-target-bytes` in the suite's locust `flags:` →
`/boomer-config` → the Go worker. Because it's per-user-class runtime
config, heterogeneous and changing workload shapes are expressible with
no redeploy, and the deploy stack needs no changes.
Changes:
- `cmd/benchmarking/glutton`: route `WriteRAM` on the HTTP-mode mux (it
was gRPC-only, unreachable in `--mode=http` deployments)
- `internal/benchmarking/boomer/glutton`: `ensureRAMFilled` on the user
cycle — runs once per actor (glutton holds the allocations across
suspend/resume), retried on failure, reported as its own
`GluttonFillRAM` stats row so it never pollutes ping/resume numbers
- `internal/benchmarking/boomer/dynconfig` + `common/boomer_config.py`:
new `mem_target_bytes` knob
- `tests.yaml`: the four tracked suites, with `actorMemory` sized above
the target for headroom
## Behavior notes
- The golden template snapshot and each actor's cold boot stay small:
the working set exists from the first fill onward. All steady-state
suspend/resume cycles measure at size; `ResumeActorColdStart` does not.
- Each actor's first suspend uploads the first at-size snapshot and
follows the fill — cycle one is an expected outlier on the dashboards.
- A mid-run change to `mem_target_bytes` applies to newly spawned users;
already-filled actors keep their size, so each actor's cycles stay
comparable.
- Existing suites are untouched: with no flag the target is 0 and the
fill is a no-op.
## Testing
- `go test -race` across `cmd/benchmarking/glutton`,
`internal/benchmarking/boomer/...`, and the glutton fake: PASS. New
tests cover chunking to an exact target, disabled-by-default, and
fail-then-retry.
- Cluster verification (microvm, 1Gi target): verified that the RAM fill
completes before the initial suspend, snapshot upload reflects the
expected ~1Gi payload, and steady-state suspend/resume cycles succeed
without error.
atelet's spec was written for runsc: a hardcoded `runsc` hostname, CRI
pause/sandbox annotations and `dev.gvisor.spec.mount.*` hints. The
micro-VM shim compensated by rewriting every spec at runtime
(ensureKataCompatibleSpec), swapping the whole mount set and rebuilding
it from one hand-written builder per volume kind — so a volume kind
atelet learned about reached the guest only once the shim was taught
about it too.
atelet now emits a spec that names no runtime, and each ateom shapes it
for the runtime it drives. Whatever atelet adds therefore reaches both,
and a bind the micro-VM shaper cannot place in the guest is an error
rather than a silent drop.
Fixes#709.
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Create/Suspend/Resume Actor workflows now uses the
`Actor.actor_template` ObjectRef when specified, otherwise falls back to
the CRD in cmd/ateapi/internal/controlapi/template_convert.go.
The counter tutorial's headline claim — state preserved across
suspend/resume — is false on the default path: with `onCommit: Data`,
suspend excludes process memory, so the in-memory counter resets while
only the durable file counter continues.
**Reproduced live** (GKE, template config identical to main): memory 5 →
**1**,2,3 after suspend/resume; suspend manifest `scope:"data"`,
`pages.img` 71 B (vs 2.2 MB golden Full). The micro-VM variant has the
same broken promise via a different path: `fromData: Golden` restarts
the count from the golden state — the e2e Golden case itself asserts
memory=1 after suspend.
**Fix** (per maintainer decision — config *and* docs, both variants
matching): `onCommit: Full` on both counter templates, drop the now-dead
`onResume` block on the micro-VM one, and make the README precise about
which scope preserves what.
**Verified live**: with `Full`, memory count runs 5 → **6**,7,8 across
suspend/resume; snapshot manifest `scope:"full"`, `pages.img` 2.4 MB.
No coverage lost: `TestDurableDirLifecycle` exercises every
`onCommit`/`onPause`/`fromData` combination (including Data and Golden)
with its own parameterized templates, independent of the demo defaults —
and its `Full/Full` case already asserts the memory count continues, in
both sandbox classes.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
The Capabilities, SecurityContext, ImageVolumeSource and Volume.image
fields landed without doc comments, so apitool's documented rule fails
on them and they are not in the exemption backlog. Document them rather
than growing the backlog.
Fixes failure on main
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Allow the `apitool validate` linter to define exemptions, then add
exemptions for current validation errors, and then enable this as part
of presubmit (GH action).
This is the companion to #496
In order to grant access to `/dev/kvm` we have to either:
- use a device plugin
- use a DRA driver
- use an NRI plugin
An NRI plugin is highly privileged in it's own right for all pods on the
host and is difficult to ship portably at the moment.
DRA is promising, but not enough functionality is GA yet at our current
1.35+ target.
Device plugin fits reasonably well. We do wind up publishing an
~arbitrarily high limit, which has some cost in kubelet memory, but
otherwise is relatively clean.
This approach is what kata uses currently. Their device plugin is not
available unbundled, and we anyhow have a per-node daemonset.
atlet is taught to sniff if /dev/kvm appears on the host at all, so we
can also stop using the manually labeled nodes for microVM class and
instead schedule to the KVM + TUN devices on nodes that advertise them.
Later we can migrate to device plugin by using
I implemented that already, but I don't think it's worth merging at the
moment. We would want consumable capacity to be on by default. We can
migrate later without changing the pod spec by using
`extendedResourceName`.
https://github.com/agent-substrate/substrate/compare/main...BenTheElder:substrate:ateom-microvm-dra
NOTE: I confirmed with upstream that device plugin is not going anywhere
despite being "v1beta1", it's GA in all but name. It won't receive new
features but we don't really need anyhow. We'll move to DRA down the
line.
---
By doing this, we can drop `privileged: true` from the uVM ateom pods.
We can also drop the `ate.dev/sandboxClass` node label hacks, reducing
friction to deploy.
Adds a temporary converter that allows each Actor workflow to use the
substrate resource rather than CRD, to prevent regression while we are
in the parallel state.
* Allows each workflow to use the ActorTemplate substrate proto under
the hood, in preparation for the full cut over, e2e tests start testing
the new code path.
* Changes any functions that consumes the ActorTemplate CRD to use the
substrate proto instead.
* The converter will be deleted once we fully cutover to the substrate
resource.
The egress e2e suites stood up in-cluster helper pods from two
separately-built ko images: `grpcecho` (a gRPC echo origin) and
`egressprobe` (an HTTP binary that both drives the gateway via
`/handshake` and doubles as a plain HTTP target via `/healthz`). This
collapses both into a single `internal/e2e/fixtures/testserver` binary
that dispatches on a subcommand:
- `testserver grpc` — cleartext-h2 gRPC echo origin (was grpcecho)
- `testserver http` — plain HTTP origin serving /healthz (the httpTarget
role)
- `testserver probe` — the gateway client-driver (was egressprobe)
Each pod runs one subcommand on one listener, so per-pod wire behavior
is unchanged.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
#1017
We were leaking snapshots after actor termination.
Note that this is a temporary fix: it only removes local snapshots in
one node. We should clean up the copies on any other
NodeVmsWithLocalSnapshots. This is fine *as of the day this was written*
because today NodeVmsWithLocalSnapshots has at most one item.
Related to https://github.com/agent-substrate/substrate/issues/668 and
#664
- [x] Tests pass
- [] Appropriate changes to documentation are included in the PR
AteletDialer caches one grpc.ClientConn per atelet pod UID in a
1024-entry LRU with no eviction function, so a conn pushed out to make
room was forgotten, never closed. grpc does not reclaim an un-Closed
conn: its goroutines and buffers stay alive for the life of the process.
Close evicted conns via the cache's eviction hook, the same fix the
atelet ateom dialer received.
Fixes#470
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [x] Appropriate changes to documentation are included in the PR
The rename of `hack/verify/verify-metrics.sh` to
`hack/verify/metrics.sh` left three references to the old path behind. A
reader who copies one of them gets a "No such file or directory" error.
Point `AGENTS.md`, `docs/observability.md`, and the header comment of
`docs/metrics/registry/metrics.yaml` at the new name.
Docs only; no change to the script or to the registry data.
A follow-up of https://github.com/agent-substrate/substrate/pull/1221
Part of #1230
When reusing an already-running Ryuk reaper, testcontainers-go v0.43.0
waits only for its Docker port mapping. Docker exposes that port before
Ryuk is listening, so a second package can connect too early and lose
the handshake with `read ack: EOF`. That is
[testcontainers-go#3743](https://github.com/testcontainers/testcontainers-go/issues/3743);
v0.44.0 also waits for the reaper's `Started` log line
([#3761](https://github.com/testcontainers/testcontainers-go/pull/3761)).
A package whose handshake fails is not counted as a Ryuk client but its
containers still carry the shared session label, so another package
exiting can delete a database that is still in use — the failure
reported in #1230.
## Scope
This helps local runs, where the reaper stays on. **It does not fix CI
on its own**, so it is deliberately separate from #1235, which disables
the reaper for `run-tests`.
Measured on a 4-core Linux VM with cold build caches, which staggers
package start times the way CI does:
```
v0.43.0, 3 rounds handshake failures in 3/3 rounds; one round cascaded
v0.44.0, 3 rounds no handshake failures; database tests still unavailable in 3/3 rounds
```
Under v0.44.0 the failures move to `wait for reaper <id>: context
deadline exceeded` — a late package finds a reaper that is already
shutting down, and one readiness probe (`defaultStartupTimeout`, 60s)
outlives the whole reaper retry budget (`MaxElapsedTime`, 20s), so the
retry loop never gets a second attempt. The remaining fail-open behind
all of this is
[#3827](https://github.com/testcontainers/testcontainers-go/issues/3827),
still open upstream.
## Diff size
Four lines of `go.mod`. The rest is `go mod vendor` output: v0.44.0
pulls newer `moby/client`, `gopsutil`, and `otelhttp`, and `otelhttp`
moves `otel/semconv` from v1.39.0 to v1.41.0. Insertions and deletions
nearly cancel because most of it is a directory swap and one generated
`httpsnoop` file being merged into another.
```
go.mod, go.sum 62 lines
vendor/ 61 files, 17356 +/17490 -
```
`hack/verify/go-modules.sh`, `licenses.sh`, `boilerplate.sh`, and
`gofmt.sh` all pass; `go test -race ./cmd/ateapi/...` is green with no
silently skipped database tests.
---
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Pick up the latest kind release for local dev clusters.
Includes fix relevant to
https://github.com/agent-substrate/substrate/issues/1133 (kindnetd
memory)
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Fix#1213, we should also reject unknown fields at request level.
Refactor all functionality back into a gRPC interceptor to check every
incoming request, no need to handle every new request method.
> It's a good idea to open an issue first for discussion.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Part of #477.
The controller has 2 components
1. A "producer" that periodically lists ActorTemplates and add work
items to a client-go workqueue.
2. A "consumer" that gets work items from the queue, and reconcile it.
AT will have the FAILED condition if we encounter non-retriable errors
during the transition
Part of #1131.
The substrate ActorTemplate's Volume union (#824, #1026) predates the
systemInfo volume source (#803, trustBundle in #941): it only had
durableDir and externalVolumeTemplate, so a template created through the
substrate API could not carry SystemInfo projections. This PR adds:
- `SystemInfoVolumeSource` with the `actorMetadata` and `trustBundle`
data sources, mirrored from the CRD types in the plain-field union shape
(#962) that `Volume` already uses ("SystemInfo" joins the `type`
discriminator).
- Declarative validation tags (#1215) mirroring the CRD's kubebuilder
rules: the SystemInfoDataSource one-of union, list bounds, string
lengths, and unique projected fields, with generated validators and
tests in controlapi.
- Store contract round-trip coverage: the shared fixture now carries a
systemInfo volume through both the redis and postgres backends.
Not in this PR (tracked in #1131): the CRD's cross-item CEL rules
(duplicate paths across data sources, at most one actorMetadata entry)
and the clean-relative-path checks, which have no declarative tag
equivalent and belong in the controlapi create/update handlers when
ActorTemplate serving lands (#477), and conversion of the substrate
resource in the actor-start path.
Fixes#1106. Takes over #1107 (adopting @Stevenjin8's approach and
commits from that PR — credit to him for the design; trailer omitted
only for CLA) and incorporates the review feedback it accumulated.
Workers were registered as schedulable the moment their pod had an IP —
long before ateom's gRPC server was listening, which is the scheduling
race in #1106 (actors dispatched to a cold ateom fail at dial time).
Now:
- ateom (gvisor + microvm) serves HTTP `/readyz` on a dedicated port
(8080) once its gRPC listener is up, flipping to 503 on SIGTERM so
draining workers stop receiving work.
- The worker Deployment probes that endpoint (replaces the earlier TCP
probe against the TLS port, which spammed handshake errors).
- `isWorkerEligible` requires `PodReady=True` in addition to an IP.
Eligibility gates **registration only** — a readiness flap never
deregisters a worker with a bound actor (pinned by
`TestSyncer_ReadinessFlapDoesNotDeregister`).
- A readiness server that cannot come up exits the process rather than
leaving a pod that looks alive but can never receive work.
Beyond #1107's head: dropped the pre-migration
`controlapi/syncer_test.go` copy (didn't compile after #1203 moved the
syncer), fixed the `workersync` fixtures for the new gate, and added
`TestSyncer_PodWithIPButNotReadyIsNotRegistered` (+ the flap test
above).
## Testing
- `go test ./...`: green, zero failures.
- `-race -count=10` on `workersync` + `serverboot`, `-race -count=5` on
`controllers`, `-race` on all of `controlapi` incl. functional tests:
green.
- Green under a hostile ambient env
(`KUBECTL_CONTEXT`/`KO_DOCKER_REPO`/`PROJECT_ID` set to garbage).
## Scope notes
- The one post-#1160 failure with a *post-restore* signature (run
32891995583: restored actor's own container `/healthz` timing out, actor
stuck RESUMING) is a different layer — ateom was demonstrably serving —
and is **not** claimed as fixed here. If it recurs, it deserves its own
issue. (The other suspected post-merge failure, run 32911674788, turned
out to be that PR's own x509 regression, not this flake.)
- Mixed-version caveat: this controller pointed at pods running a
pre-readyz ateom image leaves those workers unready/unregistered. The
supported install flows build both from the same source, so this only
matters for hand-rolled skew.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
This commit fleshes out localca's rotation support by changing the
interface so that signing operations should always go through a pool.
All existing in-tree callers are converted over to this pattern. This
ensures that signing can properly continue as a CA pool is rotated,
without requiring any process restarts.
The pool is periodically reloaded from disk in the course of signing
operations, with the loaded data being cached for up to a minute.
Some code for storing the CA state using PEM files was removed because
it seemed to be dead.
Follow-on changes will give localjwt a similar treatment, and add
administrative commands for running rotations on pools.
The image reference is pinned by digest, so the tag is discarded on
pull. Building it from ${WEAVER_VERSION} suggested that bumping that one
variable moved the image, when a new digest is needed as well.
Write the tag literally instead, next to the digest, as
`hack/third_party/kubernetes/verify-shellcheck.sh` does.
A follow-up of
https://github.com/agent-substrate/substrate/pull/1097#discussion_r3857057767
The idea is to bundle automation related to the API here. For now, there
is only a single `validate` command that runs a set of lint rules to
enforce that the API adheres to the style guide. This should help us
prevent regressions / deviations when modifying it.
These are not enforced as part of the presubmit checks, although we
might want to enforce that soon (right after we close the existing gaps?
see #1163). There is no support for violation exemptions for now.
In the future, we could add another command to, for example, generate a
reference documentation for the whole API.
## Summary
This change updates the boilerplate verification to accept either of the
approved copyright attribution forms:
- `Copyright YYYY Google LLC`
- `Copyright YYYY The Agent Substrate Authors`
## Why
The repo now allows either owner attribution, and the boilerplate
checker should enforce that both valid forms pass without rejecting the
other.
I retained the Google LLC so that we didn't have to re-copyright all the
existing files.
## Verification
- Ran the focused matcher assertions for both allowed strings and an
invalid holder.
- Ran `hack/verify/boilerplate.sh` successfully.
AteomDialer caches one grpc.ClientConn per worker pod UID in a 256-entry
LRU, but the cache had no eviction function: a conn pushed out to make
room was forgotten, never closed. grpc does not reclaim an un-Closed
conn -- measured against a dead socket, each one holds 4 goroutines, and
the channel idle timeout later parks only one of them, so ~3 goroutines
and their buffers leak per eviction for the life of the process.
Today evictions need >256 distinct worker pods dialed since atelet
started, so the leak is a slow drip reset by restarts. Close evicted
conns via the cache's eviction hook. The trade is that an RPC still in
flight on a conn that aged to the LRU tail now fails visibly instead of
completing on a leaked conn; that needs the same >256-pod churn, and the
failure is retryable.
Fixes#1009.
- [ ] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
This mostly eliminates the need to hand-write validation code.
This PR is a long series of commits which add DV for most of Actor and
all of Atespace.
Here is a map to the commits:
* The first few take a dep on a new Kubernetes tag, import the code into
third_party, and apply a single patch. Because we have different Go
modules for tools, I had to do it twice. When that patch lands, we can
revert these commits, but that won't be until the 1.38 cycle in a few
months.
* The next commits slowly add DV support, so a human can review each of
them in a reasonable amount of time. The emphasis is on great test
cases.
* I made a bad choice early on as to where to generate the code into, so
I moved it. Rebasing on that was exceedingly hard, so I left it as a
move.
* This required changing update/go-generate -> update/codegen -- we need
to get the ordering of tools right, which `go generate` does not
guarantee.
* Then I added a "middle" layer called "ServiceImpl" between the RPC and
storage layers. This allows things like workflow to call the same
business logic as the RPC layer, including validation. Lots of test
fixes.
* Then I finished the Create() and Update() paths for Actor. Those
represent the "right" (or closest to) way to implement resources, and
tests for validation.
I strongly encourage reviewers to read it commit-by-commit. Rebasing
this is VERY tedious, so the sooner it lands or dies completely, the
better. Then we can start converting the rest.
`./hack/run-tool.sh validation-gen --docs` will produce some docs on the
tool and the available tags.
@laoj2 @juli4n @EItanya @HavenXia
@lalitc375 @yongruilin @jpbetz FYI
Per-container CPU and memory limits for micro-VM actors. First slice of
#752.
An `ActorTemplate` container can cap its own cpu and memory so it cannot
starve or kill its siblings in the same actor. A container that exceeds
its memory limit is OOM-killed on its own; the rest of the actor is
unaffected.
```yaml
sandboxClass: microvm
containers:
- name: trainer
resources:
limits: {memory: 1500Mi}
- name: sidecar
resources:
limits: {memory: 256Mi, cpu: "0.2"}
```
#752 asks for per-container device selection as well. This PR does the
cpu and memory half; devices are the remaining half, and when they land,
GPU assignment should follow the same field rather than staying
actor-wide as it is today.
## Why micro-VM only
The two sandbox classes support opposite halves of #752. Micro-VM actors
have a real guest kernel, so each container gets its own cgroup and the
limits bind. gVisor applies cgroup limits at the sandbox level: one
sentry backs every container in the actor, so a per-container cgroup is
created and then stays empty (google/gvisor#190). Measured on a running
actor, the workload container's cgroup reported `memory.current=0` while
all 20 sandbox processes sat in the pause leaf. A template that sets
`resources` with `sandboxClass: gvisor` is rejected at admission.
## How a limit travels
`ActorTemplate` → CEL validation → `ate-api-server` resolves each
`resource.Quantity` once → `ateletpb` → atelet writes OCI
`linux.resources` → `ateom-microvm` merges kata's defaults and checks
the guest envelope → `SpecToAgentPB` → kata agent → guest cgroup.
The limit is carried as a standard OCI field rather than a
substrate-private concept, so a runtime that gains per-container
enforcement picks it up without new plumbing.
## Composition with #679#679 sizes the sandbox itself from `ActorTemplate.spec.resources`: guest
RAM and vCPUs for a micro-VM, the sentry for gVisor. This PR subdivides
that sandbox. The two are different fields and complementary layers.
They collided in one place. #679 applied the actor-level size to
**every** container's OCI spec, overwriting the per-container limits
atelet writes into the same field, so a container asking for 64Mi
silently received the whole guest. This PR removes that call on the
micro-VM path. A container now gets a cgroup limit only when it declares
one; an undeclared container is bounded by guest RAM, which is the real
ceiling. On micro-VM the stamp never reached the guest at all:
`SpecToAgentPB` carried neither `Memory` nor `CPU.Quota`, so it stopped
at the bundle's `config.json`. Running the pre-PR pipeline shows
`Quota=50000 Memory=2Gi` on disk arriving as `Quota=0 Memory=<nil>` on
the wire.
gVisor is untouched. There the actor-level size is applied to the
`pause` container, and since one sentry backs every container, that is
the only cgroup that binds.
Two consequences worth calling out for reviewers of #679:
- The `sizing.SandboxSize` threaded into `buildActorContainers` and
`ensureKataCompatibleSpec` had no remaining reader, so it and the
`guestSize` helper are removed. `resolveGuestMemMiB` still sizes the VM
and still rejects a declared limit too small to boot.
- `internal/sizing`'s package doc claimed both runtimes shared
`ApplyToOCISpec` for container cgroups. That is now gVisor-only, and the
comments say so. No code in that file changed.
`checkResourceEnvelope` runs after `resolveGuestMemMiB`, so it validates
the per-container sum against the post-reserve guest rather than the
SandboxConfig default. Its error now names
`spec.resources.limits.memory` (or `.cpu`) when the actor declared a
size, and `SandboxConfig` when it did not.
`internal/e2e/suites/sizing` is unaffected: it asserts on `num_cpu` and
`mem_total_bytes`, both of which come from VM sizing, and only logs the
cgroup files.
## Verified on hardware
Same template, run twice, one commit apart on a micro-VM actor:
| | before | after |
| :--- | :--- | :--- |
| `hog_ovl/memory.max` | `max` | `67108864` (the declared 64Mi) |
| `bystander_ovl/memory.max` | `max` | `max` |
| hog allocates 128MB | survived | OOM-killed |
| bystander | alive | alive |
The bug in between: `SpecToAgentPB` converted only `Devices` and
`CPU.Shares` out of `Linux.Resources`, so a memory limit reached the
bundle's `config.json` and was dropped on the way to the agent. Every
unit test passed and the on-disk spec was correct while the feature did
nothing.
## Notes for review
- `ContainerResources` deliberately does not reuse
`corev1.ResourceRequirements`, which also carries `requests` and
`claims`. There is no scheduler inside an actor to hint at, and for
memory a soft request cannot express "must have this much to come back
at all". #679 reached the same conclusion from the other direction and
now rejects `spec.resources.requests` and `.claims` at admission.
- `Limits` is a bounded named type. A `MaxProperties` marker on the
field lands at the wrong schema level for a named map, and without a
bound the CEL cost estimator rejects the whole schema. This CRD is close
to its schema-wide CEL cost ceiling; a rule that iterates containers and
parses quantities exceeds it by more than 100x and takes the existing
rules down with it.
- A cpu limit below `10m` is raised to `10m`: the kernel rejects a CFS
quota under 1ms.
- `SpecToAgentPB` drops a non-positive quota and a zero period rather
than forwarding them. A non-positive quota means unlimited in OCI but
was sent as a literal zero, which the guest applies as no CPU at all; a
non-nil zero period overwrote the CFS default set for a live quota. Both
now read a spec the way `cpuLimitMillis` does, so the envelope check and
the conversion agree on the same input.
- `mergeKataResources` fills the gaps kata's defaults cover rather than
allowlisting known fields, so a field it does not recognise survives the
merge. That holds only as far as the merge: `SpecToAgentPB` converts
`Devices`, `Memory` and `CPU` only, so `Pids`, `BlockIO`,
`HugepageLimits` and `Network` stop at the ttrpc boundary. Nothing sets
them today.
## Known gaps
- An OOM-killed container is not reported above the guest. The actor
stays `STATUS_RUNNING` and nothing records which container died. Raised
on #550. Note that #961 settled workload stats at the sandbox level and
deliberately does not attribute per container, so surfacing which
container died needs a separate signal: the guest cgroup's
`memory.events`, not the stats path.
- A container's writable rootfs is a guest tmpfs, so filesystem writes
are charged to its memory limit and a container writing more than its
limit to `/tmp` is OOM-killed. Documented in `api-guide.md`; separating
the two budgets needs an `ephemeral-storage` field.
- Over-subscription across containers is caught when the actor starts,
not at apply time. An admission-time sum check is not implementable
within the CEL cost budget.
- The atelet hop that attaches resources to the bundle spec has no test:
replacing `ctr.GetResources()` with `nil` leaves the whole suite green.
Closing it needs an imagecache fixture for `prepareOCIDirectory`.
- Restored actors inherit the golden's cgroups rather than applying
their own spec. Correct today only because `ActorTemplateSpec` is
immutable.
- The GPU section of `api-guide.md` still says `ActorTemplate` has no
per-container resource fields, which this PR makes false. Correcting it
is left to the devices half of #752, which is what will change the GPU
behaviour that paragraph describes.
---
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
---------
Signed-off-by: Eliran Wolff <eliranw@nvidia.com>
Envoy's ext_proc message_timeout defaults to 200ms, and the egress
CONNECT filter never set it. That default suits a sidecar answering out
of its own memory; this one authorizes each CONNECT with a GetActor
round trip to the ate API, and the first RPC after a router restart also
pays the lazy gRPC channel's DNS, TCP, and mTLS handshake. Past 200ms
with failure_mode_allow: false, Envoy abandons the request and answers
504 Gateway Timeout, discarding the verdict the handler was about to
give -- so a deny arrives as a 504 rather than the handler's 403, and a
real actor's first egress after a restart can eat the same 504.
Set message_timeout to 5s.
`docs/observability.md` was the only description of the metrics of
Substrate, and nothing compared that text with the code. It listed 13
instruments; the code sends 21. The request-parking and actor
resource-usage instruments were
missing entirely.
This change adds an [OpenTelemetry
Weaver](https://github.com/open-telemetry/weaver) semantic convention
registry.
### What is in it
| Path | Content |
|---|---|
| `docs/metrics/registry/manifest.yaml` | The Weaver manifest: name and
schema URL. |
| `docs/metrics/registry/metrics.yaml` | 21 instruments and 12 attribute
groups: every label, its permitted values, and its buckets. |
| `docs/metrics/substrate.yaml` | The rules Weaver cannot express. |
| `hack/verify/verify-metrics.sh` | The check. |
Fixes#1093
> It's a good idea to open an issue first for discussion.
- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
Towards #372
Recognize any HTTP 404 response (including bare 404s and `NoSuchBucket`)
in addition to typed `NoSuchKey` errors as object absence in the S3
client, mapping them to `ateerrors.ReasonFailedGetExternalObject`. This
ensures consistent absence classification across AWS S3 and
S3-compatible endpoints (e.g., MinIO, R2, Ceph). HTTP 403
(`AccessDenied`) remains an unclassified error.
AI has assisted with this PR and I have verified all the changes.
- [x] Tests pass
- [ ] Appropriate changes to documentation are included in the PR
Fixes#675, fixes#1100, fixes#1146; addresses the CI flake in #1106.
## Why this PR
Flaky tests are the single biggest drag on this repo's velocity right
now: the three flakes fixed here account for the majority of red CI runs
over the last 7 days (identity: 35 failures, parking: 32, relay: 9 —
from the flake dashboard's cross-PR analysis of ~630 runs). Every red
run costs a contributor a rebase-and-rerun cycle and costs reviewers
signal. **This PR consolidates the three root-caused, in-flight fixes
into one change to get CI green now and unblock the community — the goal
is velocity, not authorship.**
## Credit where it's due
All three fixes were root-caused and written by others; this PR adopts
them onto latest main with their tests, unchanged in substance. Each
commit carries a `Co-authored-by` trailer:
| Commit | Original PR | Author | Root cause |
|---|---|---|---|
| e2e: give each probe fixture its own worker pool | #1147 |
@orangeCatDeveloper | identity/egressmitm/imagevolume suites share one
`workload: probe` pool label; cross-suite selection under concurrent
suite processes dials workers that are not there |
| atenet: never cancel an in-flight resume at the park budget | #991 |
@omeryahud | the park budget doubled as the ResumeActor RPC deadline; a
mid-restore cancel strands a RESUMING actor on a live worker |
| atunnel: close the relay's both ends before returning | #1101 |
@orangeCatDeveloper | the relay closed both ends from a
`context.AfterFunc` goroutine the test never waits for |
@Stevenjin8's #1107 correctly diagnosed the ateom readiness race in
#1106; the control-plane readiness gap it targets remains real and open
— this PR only removes the e2e-fixture contention that makes it fire
constantly in CI.
If maintainers prefer to land the original PRs individually instead,
closing this one is completely fine — the point is that the fixes land
somewhere, soon.
## Evidence the flakes are actually fixed
**TestRelayIngressCancellationClosesBothSides (unit, `-race`):**
- Unpatched main, `-count=3000`: **83 failures (2.8%)** — matches the
2.9% observed across 308 CI runs this week
- This branch, `-count=10000`: **0 failures**
**TestRequestParking (park-budget cancellation):**
- The new `InFlightAttemptRunsToCompletion` and
`LateRetryableErrorIsBudgetExhaustion` unit tests (from #991) encode the
exact failure mode from #675 and pass under `go test -race -count=100
./cmd/atenet/internal/router/ingress/`
- The pre-fix behavior (budget cancelling the in-flight RPC) is
deterministically reproduced by the old test it replaces
**TestActorIdentity_AfterRestore_IsOwnID_NotGolden (probe pool
isolation):**
- Not reproducible outside CI (needs concurrent suite processes on a
contended kind node), so verified statically: `${FIXTURE_SUFFIX}` is
always `-<suite>` (internal/e2e/sandbox.go:189,201 — never empty),
`probe-sized` already uses its own label, and no other manifest or
selector references `workload: probe`. #1147's CI data shows all three
failure signatures (missing `ateom.sock`, `runsc restore` killed, router
502/503) trace to cross-suite pool sharing; per-suite labels make the
selector suite-local by construction
- The definitive check is this PR's own CI plus the flake dashboard's
7-day window after merge — I will report the post-merge rates on #1106
Also run: `go build ./...`, `go vet` and the full `-race` suites of both
touched packages — all green.
## What this PR deliberately does NOT fix
`TestActorEgressHTTPS` (#1050, 4.6% this week, below the 5% flake
threshold) has no root-caused fix yet — the 503 `upstream connect error`
path needs investigation in a live cluster. #1103 (@orangeCatDeveloper)
tightens the related `TestActorArbitraryPortAccess` assertion so those
503s stop passing silently; it should land after #1050's cause is fixed,
or it converts hidden flakiness into visible red.
## Update (post-CI investigation)
The first e2e runs failed on `TestRequestParking/ParkThenServed`
(micro-VM lane). Investigation showed this is the **pre-existing
dominant mode** of #675 — identical failures in main-era runs
32305728993 / 32397291519 / 32487958077 — not a regression: on micro-VM,
`SuspendActor` returns before the snapshot upload completes, so the
worker legitimately isn't free within the 5s park budget and the
router's 503 is correct behavior. #991 fixes the *other* (mid-restore
cancellation/stranding) mode. Commit d637690d makes the subtest retry
while the worker is still freeing; a stranded worker still fails every
attempt, so the regression stays pinned.
**Additional validation:**
- CI e2e-test now **passes both lanes** (run 32754691017)
- Local kind cluster built from this branch: parking suite **10/10
consecutive passes**; identity + egressmitm + imagevolume run
**concurrently** (the exact contention behind the identity flake) × 3
iterations — **9/9 suite passes**
---------
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Co-authored-by: NekoPunch <engineer.jyao@gmail.com>
Co-authored-by: Omer Yahud <oyahud@nvidia.com>