Files
Max Thompson 074fd1c12b Add trustBundle as a SystemInfo volume data source (#941)
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data
source for SystemInfo volumes (#802) and the end-to-end proof that the
projected anchors work against the MITM egress gateway. Live refresh for
running actors (PR 2) and auto-injection (PR 3) come separately.

What this adds

A SystemInfo volume data source that projects the trust anchors of a
named trust bundle to a PEM file:

volumes:
- name: trust
  systemInfo:
    dataSources:
    - trustBundle:
        name: egress-mitm.ate.dev
        path: egress-ca.pem

Inspired by the Kubernetes clusterTrustBundle projected volume source,
but source-neutral: the template names a bundle; where it's fetched from
is a deployment concern, not part of the API.

Design points

- Resolution lives on the node. The wire carries only {name, path};
atelet resolves the name at write time through an informer-backed lister
on ClusterTrustBundles and writes the sanitized PEM with the temp+rename
discipline from #803 (find-paths safe). Contents refresh on every
Run/Restore. ateapi is not involved, per review discussion — the same
informer is what live refresh (PR 2) will hang off.
- Allowlist in atelet, not the CRD schema. Today only
egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the
ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946)
derives from the egress-mitm-ca-pool Secret. The signer-linked object
name stays a backend detail; the future backend registry (#932) widens
the allowlist without an API change.
- The watch is scoped to the one backing object via a metadata.name
field selector — this informer runs on every node, so an unfiltered
watch would fan every ClusterTrustBundle in the cluster out to every
atelet. RBAC can't express this (resourceNames doesn't apply to
list/watch), so the field selector is the enforcement point.
get/list/watch on clustertrustbundles moves to the atelet ClusterRole.
- No availability probe. The informer registers unconditionally; a
cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1
blocks atelet startup at cache sync, with the reflector errors naming
the missing API (hack/create-kind-cluster.sh enables the gate).
- Fail-closed. Unknown names, missing bundles, and unusable bundles fail
actor start naming the bundle — an actor that declared a trust bundle
must not start without one.
- Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks
only, deduplicated, headers stripped, and anchors deliberately shuffled
so consumers can't grow a dependence on order.
- Schema note: dataSources MaxItems tightened 32→8 while adding the
trustBundle member. Vacuous in practice (the old schema couldn't admit
more than one entry), but flagged since it's ratchet-shaped.

E2E — delivery and consumption

Delivery (identity suite, both sandbox classes): provisions the
egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing
the bundle directly isn't possible — the reconciler reverts hand-edits),
asserts the projected file byte-exact, then rotates the pool across a
suspend/resume to prove refresh-on-restore. Since the probe fixture is
shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists
for whatever suite deploys it.

Consumption (new egressmitm suite, both sandbox classes): deploys the
sdsmint (MITM) egress gateway and proves an actor completes a TLS
handshake with the gateway's per-SNI minted leaf using ONLY the
projected anchors — plus a system-roots negative control that must fail.
The pair is unambiguous in both directions: the positive can't pass
under passthrough (the bundle holds no public CAs), and the negative
can't fail under passthrough.

CI: two steps appended to the existing e2e job after the standard lanes
(the gateway swap is cluster-wide and breaks passthrough assumptions):
--deploy-atenet --experimental-use-sdsmint redeploys only the atenet
components, then the egressmitm suite runs once per sandbox class.

Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each
suite deploys its own copy and drives one actor at a time, so the third
worker per copy was idle memory multiplied across suites on the one-node
CI cluster — pressure that has been killing sandboxes mid-test (runsc:
signal: killed, a vanished ateom socket) on this PR and on main's
identity suite. This reduces the pressure; right-sizing e2e concurrency
or worker-pod QoS cluster-wide is follow-up material.

Not in this PR

- Live refresh for running actors (#932 PR 2) — until then, a running
actor's file is the bundle as of its last Run/Restore, and correctness
rests on overlap rotation by the bundle publisher.
- Auto-injection of the egress trust volume (#932 PR 3).
- Configurable backend registry (#932) — the allowlist is the seam it
will replace.
2026-08-21 19:53:31 -04:00
..