3 Commits
Author SHA1 Message Date
Lucky Abolorunke ad830889ad microvm: exclude overlay workdirs from rootfs snapshots (#1030)
#### microvm: Exclude Overlay Workdirs from Rootfs Snapshots

With `index=off` pinned on the merged rootfs mount, a restored workdir
is inert: overlayfs wipes and rebuilds it at mount time, and staging
recreates the directory regardless.

Archiving `<cid>/work` was therefore dead weight on the suspend path—and
not always trivial weight, since a copy-up in flight at pause can leave
file-sized temporary data there.

---

#### Key Changes:

* **`tarutil.CreateFiltered`**: Added a skip predicate over `Create`
(skipping a directory prunes its entire subtree).
* **Workdir Pruning**: Drops exactly the second-level `work` entries
during snapshot archive creation.
* **Backward Compatibility**: Extraction logic remains unchanged;
existing snapshots containing `workdir` entries will still restore
without issue.

---

> **Follow-up to #846 review note:** With `index=off` pinned, a restored
workdir is inert and overlayfs rebuilds it at mount, making its archival
unnecessary overhead on the suspend path.
2026-08-17 18:17:47 -07:00
Benjamin Elder 2d6befa396 ateom-microvm: fix graceful termination, cleanup dead code (#1021)
Fixes #1016

1. Fix graceful termination, which was broken by the drop of the
"carrier container" hack when we switched to overlay for the rootfs
writes
2. Drop dead code attempting to enable compatibility before that change.
It's a lot of complexity and we're not there yet

> It's a good idea to open an issue first for discussion.

- [x] Tests pass
- [x] Appropriate changes to documentation are included in the PR
2026-08-17 19:03:49 -04:00
Lucky Abolorunke c1339e5f02 microvm: write container rootfs to host disk instead of guest memory (#846)
### Problem

Each container's rootfs upper previously lived on a tmpfs inside the
guest: every written byte consumed guest RAM 1:1, and writes failed with
`ENOSPC` at the tmpfs cap (20% of guest RAM nominally; **~264 MiB
effective** on the standard 2 GiB guest once `/run`'s other occupants
are counted). Write-heavy actors crash; the rest degrade worker density.

### Change

The rootfs overlay is now **assembled on the host** — the conventional
arrangement for VM-isolated containers: lower = the read-only OCI image
bundle, upper/work = per-actor host directories, merged by the host
kernel and served to the guest over the **single existing virtio-fs
share** (`cache=auto`). The guest runs each container directly on the
merged directory in one step; it never mounts an overlay, needs no
special mount options, and no additional virtiofsd exists. The mount
pins `metacopy=off,index=off` so every copy-up is a full data copy — the
upper never contains file-handle references to lower inodes, which would
go stale when restore rebuilds the lower from the image (review
finding).

**Snapshot/restore integration:**

- **FULL checkpoints** archive the upper as `rootfs-upper.tar`, captured
while the guest is paused (the share is write-through, so the tar is
coherent) and **concurrently** with the cloud-hypervisor memory snapshot
— the paused window costs the slowest artifact, not the sum.
- **Restores** re-materialize the upper in the background (overlapped
with bundle preparation), re-mount the merged trees at the frozen
find-paths locations, then start the share.
- **Restore is self-describing:** the tar's presence routes it.
Snapshots from the current memory implementation restore unchanged
(bare-image bind at `cache=always`; their upper rides inside the
restored guest memory), and their re-checkpoints stay correct.
- **DATA scope unchanged:** rootfs state discarded, as today.

**Enabling change — `tarutil` now round-trips overlay deletion
metadata:**
Whiteout device nodes (`mknodat`, parent-fd contained) and `user.*` /
`trusted.overlay.*` xattrs (PAX `SCHILY.xattr` records, restored through
the extraction root so a crafted archive cannot write outside it).
Without these, deleted files and directories silently reappeared after
resume.

### How it was tested

- **Unit** (privileged tests run as root, no skips): `tarutil`
round-trips including a literal `0:0` whiteout device and
`trusted.overlay.opaque`; extraction-escape rejections; overlay layout
and snapshot-config rewrite regression tests.
- **End-to-end** (live cluster, counter actor with a durable volume — so
every run also covers durable-dir coexistence):
  - In-guest writes visible in the host upper.
- In-guest `rm` producing a genuine kernel whiteout (`char 0:0, nlink
2`) and directory replacement producing `trusted.overlay.opaque`, both
surviving `FULL` suspend → restore verified *from inside the guest*.
- Multi-cycle suspend/resume with a byte-identical 64 MiB payload
(`md5`).
  - `DATA`-scope commits emit no rootfs tar.
- `DATA_ON_GOLDEN` stitches the golden's rootfs tar with the actor's
durable data.
  - Legacy (tmpfs-era) snapshots restore and re-checkpoint correctly.
- **Benchmark**: Compared head-to-head against the memory baseline on
GKE with real GCS snapshot storage ($N=25$ full-suspend cycles per side,
same node/bucket) — suspend and restore latency at parity, no
regression.
2026-08-17 11:15:24 -07:00