mirror of
https://github.com/agent-substrate/substrate.git
synced 2026-10-02 03:24:42 +08:00
Part of #932 (PR 1 of 3). Adds the user-declarable trustBundle data source for SystemInfo volumes (#802) and the end-to-end proof that the projected anchors work against the MITM egress gateway. Live refresh for running actors (PR 2) and auto-injection (PR 3) come separately. What this adds A SystemInfo volume data source that projects the trust anchors of a named trust bundle to a PEM file: volumes: - name: trust systemInfo: dataSources: - trustBundle: name: egress-mitm.ate.dev path: egress-ca.pem Inspired by the Kubernetes clusterTrustBundle projected volume source, but source-neutral: the template names a bundle; where it's fetched from is a deployment concern, not part of the API. Design points - Resolution lives on the node. The wire carries only {name, path}; atelet resolves the name at write time through an informer-backed lister on ClusterTrustBundles and writes the sanitized PEM with the temp+rename discipline from #803 (find-paths safe). Contents refresh on every Run/Restore. ateapi is not involved, per review discussion — the same informer is what live refresh (PR 2) will hang off. - Allowlist in atelet, not the CRD schema. Today only egress-mitm.ate.dev (the egress gateway CA bundle, #823), mapped to the ClusterTrustBundle that atecontroller's EgressMITMTrustReconciler (#946) derives from the egress-mitm-ca-pool Secret. The signer-linked object name stays a backend detail; the future backend registry (#932) widens the allowlist without an API change. - The watch is scoped to the one backing object via a metadata.name field selector — this informer runs on every node, so an unfiltered watch would fan every ClusterTrustBundle in the cluster out to every atelet. RBAC can't express this (resourceNames doesn't apply to list/watch), so the field selector is the enforcement point. get/list/watch on clustertrustbundles moves to the atelet ClusterRole. - No availability probe. The informer registers unconditionally; a cluster that doesn't serve the feature-gated certificates.k8s.io/v1beta1 blocks atelet startup at cache sync, with the reflector errors naming the missing API (hack/create-kind-cluster.sh enables the gate). - Fail-closed. Unknown names, missing bundles, and unusable bundles fail actor start naming the bundle — an actor that declared a trust bundle must not start without one. - Kubelet-parity sanitization (internal/pemutil): CERTIFICATE blocks only, deduplicated, headers stripped, and anchors deliberately shuffled so consumers can't grow a dependence on order. - Schema note: dataSources MaxItems tightened 32→8 while adding the trustBundle member. Vacuous in practice (the old schema couldn't admit more than one entry), but flagged since it's ratchet-shaped. E2E — delivery and consumption Delivery (identity suite, both sandbox classes): provisions the egress-mitm-ca-pool Secret and drives the real #946 reconciler (writing the bundle directly isn't possible — the reconciler reverts hand-edits), asserts the projected file byte-exact, then rotates the pool across a suspend/resume to prove refresh-on-restore. Since the probe fixture is shared and fail-closed, e2e.DeployProbe itself ensures the bundle exists for whatever suite deploys it. Consumption (new egressmitm suite, both sandbox classes): deploys the sdsmint (MITM) egress gateway and proves an actor completes a TLS handshake with the gateway's per-SNI minted leaf using ONLY the projected anchors — plus a system-roots negative control that must fail. The pair is unambiguous in both directions: the positive can't pass under passthrough (the bundle holds no public CAs), and the negative can't fail under passthrough. CI: two steps appended to the existing e2e job after the standard lanes (the gateway swap is cluster-wide and breaks passthrough assumptions): --deploy-atenet --experimental-use-sdsmint redeploys only the atenet components, then the egressmitm suite runs once per sandbox class. Flake mitigation: the probe fixture pool drops from 3 workers to 2. Each suite deploys its own copy and drives one actor at a time, so the third worker per copy was idle memory multiplied across suites on the one-node CI cluster — pressure that has been killing sandboxes mid-test (runsc: signal: killed, a vanished ateom socket) on this PR and on main's identity suite. This reduces the pressure; right-sizing e2e concurrency or worker-pod QoS cluster-wide is follow-up material. Not in this PR - Live refresh for running actors (#932 PR 2) — until then, a running actor's file is the bundle as of its last Run/Restore, and correctness rests on overlap rotation by the bundle publisher. - Auto-injection of the egress trust volume (#932 PR 3). - Configurable backend registry (#932) — the allowlist is the seam it will replace.
154 lines
7.5 KiB
YAML
154 lines
7.5 KiB
YAML
# Copyright 2026 Google LLC
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
name: pr-workflow
|
|
on:
|
|
pull_request:
|
|
push:
|
|
branches: [main]
|
|
schedule:
|
|
# Weekly run on the default branch keeps the micro-VM asset cache warm (GitHub
|
|
# evicts caches idle for 7 days). The push-to-main run populates the cache that
|
|
# PRs inherit; this refreshes it so PRs keep hitting it (no asset re-download).
|
|
- cron: '37 4 * * 1'
|
|
# Nothing here writes to the repository: the jobs check the tree out, build it,
|
|
# and run it in a throwaway cluster.
|
|
permissions:
|
|
contents: read
|
|
jobs:
|
|
run-tests:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@fbc6f3992d24b796d5a048ff273f7fcc4a7b6c09 # v5.1.0
|
|
- name: Setup Go
|
|
uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0
|
|
with:
|
|
go-version-file: 'go.mod'
|
|
- run: go test -race -v ./...
|
|
# Root-gated tests (overlay mounts, whiteout mknod, trusted.* xattrs, ...)
|
|
# skip for the unprivileged runner user above; rerun the packages that
|
|
# contain them (any test importing internal/roottest) under sudo.
|
|
- name: root-gated tests
|
|
run: hack/run-root-tests.sh -race -v
|
|
- name: verify
|
|
run: hack/verify-all.sh
|
|
# One kind cluster exercises BOTH runtimes. Free x86-64 ubuntu-latest runners
|
|
# expose /dev/kvm (with a udev rule), so create-kind-cluster.sh mounts it and the
|
|
# micro-VM (kata + cloud-hypervisor) sandbox class works alongside gVisor.
|
|
e2e-test:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@fbc6f3992d24b796d5a048ff273f7fcc4a7b6c09 # v5.1.0
|
|
- name: Setup Go
|
|
uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0
|
|
with:
|
|
go-version-file: 'go.mod'
|
|
- name: Cache micro-VM assets
|
|
id: microvm-assets
|
|
uses: actions/cache@0057852bfaa89a56745cba8c7296529d2fc39830 # v4.3.0
|
|
with:
|
|
# Assembling the assets (download kata-static + cloud-hypervisor + the
|
|
# prebuilt virtiofsd — all downloads on amd64) is fully pinned by assemble.sh,
|
|
# so key the cache on its hash. The push-to-main run populates the
|
|
# default-branch cache PRs inherit; the weekly schedule refreshes it.
|
|
# run-microvm-demo-kind.sh skips assembling when these are present.
|
|
path: bin/microvm-assets/amd64
|
|
key: microvm-assets-amd64-${{ hashFiles('hack/microvm-assets/assemble.sh') }}
|
|
- name: Enable KVM
|
|
# Grant the runner access to /dev/kvm so create-kind-cluster.sh mounts it into
|
|
# the node and labels it for the micro-VM sandbox class.
|
|
run: |
|
|
echo 'KERNEL=="kvm", GROUP="kvm", MODE="0666", OPTIONS+="static_node=kvm"' \
|
|
| sudo tee /etc/udev/rules.d/99-kvm4all.rules
|
|
sudo udevadm control --reload-rules
|
|
sudo udevadm trigger --name-match=kvm
|
|
- name: Create cluster
|
|
run: hack/create-kind-cluster.sh
|
|
- name: Install Agent Substrate
|
|
run: hack/install-ate-kind.sh --deploy-ate-system --store-backend=postgres
|
|
- name: Deploy micro-VM counter demo
|
|
# Stages the (cached) assets into the cluster's rustfs and applies the
|
|
# counter-microvm demo onto the control plane installed above.
|
|
run: hack/run-microvm-demo-kind.sh
|
|
- name: Deploy gVisor counter demo
|
|
run: hack/install-ate-kind.sh --deploy-demo-counter
|
|
- name: Deploy egress demos
|
|
# TestActorEgress in the networking suite builds its Actor from the egress
|
|
# ActorTemplate for the class under test, so both fixtures have to exist
|
|
# before their respective lane runs.
|
|
run: |
|
|
hack/install-ate-kind.sh --deploy-demo-egress
|
|
hack/install-ate-kind.sh --deploy-demo-egress-microvm
|
|
- name: Wait for micro-VM golden snapshot
|
|
run: |
|
|
kubectl --context kind-kind wait --for=condition=Ready \
|
|
actortemplate/counter-microvm -n ate-demo-counter-microvm --timeout=600s
|
|
- name: Run E2E tests (gVisor)
|
|
run: hack/run-e2e-kind.sh -v -args --no-color
|
|
- name: Run E2E tests (micro-VM)
|
|
# The same suites again, with every fixture repointed at its micro-VM
|
|
# variant by the single E2E_SANDBOX_CLASS knob (see internal/e2e/sandbox.go).
|
|
# Sequential with the gVisor lane above, not concurrent: a suite releases
|
|
# its namespace — and with it its worker pods — as each test passes, so the
|
|
# two runs do not contend for the one kind node.
|
|
env:
|
|
E2E_SANDBOX_CLASS: microvm
|
|
run: hack/run-e2e-kind.sh -v -args --no-color
|
|
- name: Deploy MITM egress (sdsmint)
|
|
# Swap the passthrough egress gateway for the sdsmint variant, which
|
|
# mints per-SNI leaves from the egress-mitm-ca-pool (created here if
|
|
# missing). --deploy-atenet redeploys only the atenet components (the
|
|
# rest of the control plane is unchanged), keeping this step cheap.
|
|
# Cluster-wide, so it must come AFTER the standard lanes: once egress
|
|
# TLS is intercepted, their passthrough assumptions
|
|
# (TestActorEgressHTTPS's end-to-end TLS with the origin) no longer hold.
|
|
run: hack/install-ate-kind.sh --deploy-atenet --experimental-use-sdsmint
|
|
- name: Run E2E tests (egress MITM trust)
|
|
# The consumption half of the trust-bundle chain: an actor does TLS with
|
|
# the MITM gateway's minted leaf using ONLY the projected bundle, plus a
|
|
# system-roots negative control proving interception is real (see
|
|
# internal/e2e/suites/egressmitm).
|
|
env:
|
|
E2E_EGRESS_MITM: "1"
|
|
run: hack/run-e2e-kind.sh ./internal/e2e/suites/egressmitm -v -args --no-color
|
|
- name: Run E2E tests (egress MITM trust, micro-VM)
|
|
# The same proof with the probe on the micro-VM runtime. Trust DELIVERY
|
|
# differs per sandbox class (gVisor RO bind vs the micro-VM unified
|
|
# virtio-fs share), so the handshake is proven on both. Uses the
|
|
# micro-VM deps staged earlier in this job.
|
|
env:
|
|
E2E_EGRESS_MITM: "1"
|
|
E2E_SANDBOX_CLASS: microvm
|
|
run: hack/run-e2e-kind.sh ./internal/e2e/suites/egressmitm -v -args --no-color
|
|
- name: Dump diagnostics on failure
|
|
if: failure()
|
|
run: |
|
|
kubectl --context kind-kind get actortemplate,workerpool,pods -A -o wide || true
|
|
dump() {
|
|
echo "=== logs: $1/$2 ==="
|
|
kubectl --context kind-kind logs -n "$1" "$2" --all-containers --tail=300 2>/dev/null || true
|
|
}
|
|
for p in $(kubectl --context kind-kind get pods -n ate-system -o name 2>/dev/null); do
|
|
dump ate-system "$p"
|
|
done
|
|
# Every worker pod in any namespace: the demo pools plus the e2e suites'
|
|
# randomly-named per-test namespaces, which the suites keep on failure.
|
|
# The failing actor runs in one of these, so this is where its ateom logs
|
|
# (and, for a micro-VM worker, the guest console tail) live.
|
|
kubectl --context kind-kind get pods -A -l ate.dev/worker-pool \
|
|
-o 'custom-columns=:.metadata.namespace,:.metadata.name' --no-headers 2>/dev/null \
|
|
| while read -r ns name; do dump "$ns" "$name"; done
|