atelet: live refresh of projected trust bundles for running actors (#1231)

Part of #932 (PR 2 of 3). PR 1 (#941) added the `trustBundle` SystemInfo
data source, resolved on the node at Run/Restore. This PR keeps those
projections current while the actor runs.

## Live refresh

A `systemInfoVolumeRefresher` in atelet, modeled on kubelet's projected
volumes: `collectData` is the one place that builds a volume's complete
contents from its spec, `write` applies them, and every lifecycle point
uses the pair. Run/Restore registers the actor's volumes, which writes
them fail-closed before the sandbox boots; ClusterTrustBundle events
rewrite them while it runs; Checkpoint, Terminate, and failed starts
deregister. No API, proto, or RBAC changes: the wire still carries only
`{name, path}`.

- Informer events only enqueue bundle names; a single run loop writes.
Failed writes requeue with backoff through a rate-limited workqueue;
resolution failures keep last-good contents and wait for the bundle's
next event. Resync (24h) is only a guard against missed watch events.
- Change detection hashes the raw backing contents, not the projected
output (sanitization shuffles). Files are replaced by temp-and-rename at
stable paths and byte-identical files are left untouched: a needless
rewrite replaces an inode that suspended guests re-bind on resume.
- Nothing is persisted. Registrations are in memory; after an atelet
restart, a running actor's files keep their last-written contents until
its next Run/Restore rewrites them from current cluster state.

## virtiofsd: `--migration-on-error=guest-error`

CI confirmed the hazard (3 of 3 micro-VM runs): a rotation renames over
a file whose inode the guest still references (a dcache reference is
enough, no held fd), the actor suspends before re-reading, and
find-paths serializes an inode with no findable path. Under the default
`abort`, the destination virtiofsd rejects the device state at
vm.restore: suspend succeeds, every restore fails, and the actor is
permanently stuck. Full analysis in the PR comments.

With `guest-error` the restore succeeds and only the stale reference is
faulty: EIO on access until a fresh lookup at the stable path heals it.
That matches the documented reader contract (the platform rewrites the
file, applications re-read it) and also covers guest-created
`O_TMPFILE`s held across suspend, which already fail restores under
`abort`. The flag is backend-process configuration, so existing
snapshots (template goldens included) restore unchanged. Trade-off: a
share-reconstruction bug now surfaces as post-resume EIO and a readyz
failure instead of a loud restore failure; virtiofsd names the faulty
inodes in the worker pod log.

## E2E

The identity suite covers live refresh on both sandbox classes: rotate
the pool under two running actors and wait for both to observe the new
bundle live; rotate/suspend/resume without waiting for propagation, so
the resume must deliver rotated contents whether the rewrite landed
before or after the guest went down (the exact `abort` repro); assert a
sibling that never cycled also got the rotation live. The probe fixture
now keeps its namespace when a test fails so the worker pod logs
survive, and the cloud-hypervisor client includes the vm.restore
response body in errors.

## Not in this PR

Auto-injected egress trust volume (#932 PR 3) and the configurable
backend registry (tracked in #932).
This commit is contained in:
Max Thompson
2026-09-10 16:13:55 -07:00
committed by GitHub
parent c3a3f3bdd0
commit f3845f3ce6
14 changed files with 1247 additions and 357 deletions
+3 -1
View File
@@ -229,7 +229,9 @@ spec:
mountPath: /run/substrate/certs # the actor reads /run/substrate/certs/ca.pem
```
atelet resolves the bundle on the node when the actor starts, reading the backing object through a cluster-wide watch (the same informer dynamic refresh will later hang off) and sanitizing it the way kubelet does for projections: only `CERTIFICATE` PEM blocks are kept, deduplicated, with block headers stripped and the anchors deliberately shuffled — order carries no meaning, so consumers must not depend on it. The actor itself never talks to any bundle backend. Starting the actor fails, with an error naming the bundle, if the name is not on the allowlist, the bundle's backend is unavailable in this deployment, or the resolved bundle is missing, empty, or contains no certificates. Bundle contents are re-resolved on every Run/Restore.
atelet resolves the bundle on the node when the actor starts, reading the backing object through a cluster-wide watch (the same informer that drives live refresh) and sanitizing it the way kubelet does for projections: only `CERTIFICATE` PEM blocks are kept, deduplicated, with block headers stripped and the anchors deliberately shuffled, so consumers must not depend on their order. The actor itself never talks to any bundle backend. Starting the actor fails, with an error naming the bundle, if the name is not on the allowlist, the bundle's backend is unavailable in this deployment, or the resolved bundle is missing, empty, or contains no certificates.
Bundle contents are re-resolved on every Run/Restore and refreshed while the actor runs: when the backing bundle changes, atelet rewrites the projected file atomically at the same path. As with the Kubernetes clusterTrustBundle projection, the application must re-read the file to pick up a rotation; a runtime that loads trust anchors once at startup sees the change at its next start or resume. If a refresh fails because the backing object was deleted or is unusable, the file keeps its last good contents. Bundle publishers should rotate with overlap: add the new CA before minting leaves under it, and keep the old CA until its last leaf expires.
### Container Fields