mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data source for SystemInfo volumes (#802) and the end-to-end proof that the projected anchors work against the MITM egress gateway. Live refresh for running actors (PR 2) and auto-injection (PR 3) come separately. What this adds A SystemInfo volume data source that projects the trust anchors of a named trust bundle to a PEM file: volumes: - name: trust systemInfo: dataSources: - trustBundle: name: egress-mitm.ate.dev path: egress-ca.pem Inspired by the Kubernetes clusterTrustBundle projected volume source, but source-neutral: the template names a bundle; where it's fetched from is a deployment concern, not part of the API. Design points - Resolution lives on the node. The wire carries only {name, path}; atelet resolves the name at write time through an informer-backed lister on ClusterTrustBundles and writes the sanitized PEM with the temp+rename discipline from #803 (find-paths safe). Contents refresh on every Run/Restore. ateapi is not involved, per review discussion — the same informer is what live refresh (PR 2) will hang off. - Allowlist in atelet, not the CRD schema. Today only egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946) derives from the egress-mitm-ca-pool Secret. The signer-linked object name stays a backend detail; the future backend registry (#932) widens the allowlist without an API change. - The watch is scoped to the one backing object via a metadata.name field selector — this informer runs on every node, so an unfiltered watch would fan every ClusterTrustBundle in the cluster out to every atelet. RBAC can't express this (resourceNames doesn't apply to list/watch), so the field selector is the enforcement point. get/list/watch on clustertrustbundles moves to the atelet ClusterRole. - No availability probe. The informer registers unconditionally; a cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1 blocks atelet startup at cache sync, with the reflector errors naming the missing API (hack/create-kind-cluster.sh enables the gate). - Fail-closed. Unknown names, missing bundles, and unusable bundles fail actor start naming the bundle — an actor that declared a trust bundle must not start without one. - Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks only, deduplicated, headers stripped, and anchors deliberately shuffled so consumers can't grow a dependence on order. - Schema note: dataSources MaxItems tightened 32→8 while adding the trustBundle member. Vacuous in practice (the old schema couldn't admit more than one entry), but flagged since it's ratchet-shaped. E2E — delivery and consumption Delivery (identity suite, both sandbox classes): provisions the egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing the bundle directly isn't possible — the reconciler reverts hand-edits), asserts the projected file byte-exact, then rotates the pool across a suspend/resume to prove refresh-on-restore. Since the probe fixture is shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists for whatever suite deploys it. Consumption (new egressmitm suite, both sandbox classes): deploys the sdsmint (MITM) egress gateway and proves an actor completes a TLS handshake with the gateway's per-SNI minted leaf using ONLY the projected anchors — plus a system-roots negative control that must fail. The pair is unambiguous in both directions: the positive can't pass under passthrough (the bundle holds no public CAs), and the negative can't fail under passthrough. CI: two steps appended to the existing e2e job after the standard lanes (the gateway swap is cluster-wide and breaks passthrough assumptions): --deploy-atenet --experimental-use-sdsmint redeploys only the atenet components, then the egressmitm suite runs once per sandbox class. Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each suite deploys its own copy and drives one actor at a time, so the third worker per copy was idle memory multiplied across suites on the one-node CI cluster — pressure that has been killing sandboxes mid-test (runsc: signal: killed, a vanished ateom socket) on this PR and on main's identity suite. This reduces the pressure; right-sizing e2e concurrency or worker-pod QoS cluster-wide is follow-up material. Not in this PR - Live refresh for running actors (#932 PR 2) — until then, a running actor's file is the bundle as of its last Run/Restore, and correctness rests on overlap rotation by the bundle publisher. - Auto-injection of the egress trust volume (#932 PR 3). - Configurable backend registry (#932) — the allowlist is the seam it will replace.