Files
Ruslan KonviserandClaude Opus 5 2018bb5787 fix(ci): stop the e2e suite re-installing 9 GB of node_modules in every job
Run 31889575135 ran for 12 h 34 m and produced ZERO test results. Every shard
was killed at its 120-minute timeout inside `Install Packages & Bootstrap
(cache fallback)` — shard 2 spent 1 h 58 m there and never reached step 5.

The cause is a cache that cannot fit. The job split (deps -> build -> e2e)
exists so `yarn bootstrap` is paid once, and it handed the tree over through
`actions/cache` as an EXPLODED ~9 GB, ~900k-file tree. The repository's Actions
cache budget is 10 GB, and `build.yml` runs on every pull_request with five
`actions/setup-node … cache: yarn` jobs, each storing a 4.66 GB yarn cache PER
REF — develop, stage and every open PR hold their own copy of identical
content. Two of those fill the budget by themselves. Measured, not assumed:

  - ACROSS runs the e2e cache never survived: deps paid a cold bootstrap every
    time (3 h 20 m on run 31848093657, 3 h 46 m on 31889575135).
  - WITHIN a run it was a coin flip: on 31848093657 all four shards hit it; on
    31889575135 the build job hit it and shard 2 missed an hour later.

So the tree now travels as ONE compressed archive. An artifact is the handoff —
it is scoped to the run and cannot be evicted mid-run by another workflow's
cache write — and the same file is also saved to `actions/cache` as a
best-effort cross-run fast path, where ~3 GB has a chance that 9 GB never had.
Downstream jobs download and extract it, keeping the full-bootstrap fallback
for the case where the artifact is unavailable. zstd with a gzip fallback,
since only the archive's contents are compressible and one big file avoids the
cache action's per-file cost in both directions. A size floor rejects a
truncated archive rather than publishing one that would send every downstream
job back to a multi-hour bootstrap.

Also adds a cache janitor. It keeps exactly one copy per cache key — the
default branch's if present, else the most recently used — and deletes the
duplicates that differ only by ref. Verified against the live cache list: it
selects the redundant `refs/pull/9992/merge` copy of a key already held on
`refs/heads/stage` and reclaims 4.66 GB, touching nothing else. Cache entries
are reconstructible by definition, so pruning duplicates is strictly better
than letting GitHub's LRU pick a victim at random, which is what happens today.
This leaves `build.yml` alone deliberately: it is the PR gate, and reclaiming
the budget does not require changing how it builds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 08:59:57 +02:00

120 lines
4.7 KiB
YAML

name: Cache Janitor
# The repository's Actions cache budget is 10 GB. `build.yml` runs on every pull_request with five
# `actions/setup-node@v4 … cache: yarn` jobs, and setup-node saves its ~4.66 GB yarn cache PER REF —
# so `refs/heads/develop`, `refs/heads/stage` and every open PR's `refs/pull/N/merge` each hold their
# own copy of the SAME content, keyed identically. Two of those fill the budget on their own.
#
# The damage is not theoretical. On 2026-08-15 the Playwright suite's dependency cache was gone
# within the hour it was written, so run 31889575135 paid a cold `yarn bootstrap` in the deps job
# (3 h 46 m) and its shards started doing the same again — the exact cost the deps/build/e2e job
# split exists to pay only once.
#
# Cache entries are, by definition, reconstructible: the worst case for deleting one is that the next
# run that wants it re-downloads. That makes pruning duplicates strictly better than letting GitHub's
# LRU choose a victim at random, which is what happens today.
#
# RULE: for each cache key, keep exactly ONE copy — the one on the default branch if it exists,
# otherwise the most recently used. Delete the rest. Branch-scoped restores are allowed to read the
# default branch's caches, so the kept copy still serves every PR.
on:
schedule:
# Every 6 hours. Frequent enough that duplicates never accumulate for long, rare enough that it
# never races a build that is mid-save.
- cron: '17 */6 * * *'
workflow_dispatch:
inputs:
dry_run:
description: 'List what would be deleted without deleting anything'
type: boolean
default: false
permissions:
contents: read
actions: write # required to delete cache entries
concurrency:
group: ${{ github.workflow }}
cancel-in-progress: false
jobs:
prune:
name: Prune duplicate caches
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- name: Prune
shell: bash
env:
GH_TOKEN: ${{ github.token }}
REPO: ${{ github.repository }}
DEFAULT_BRANCH_REF: refs/heads/${{ github.event.repository.default_branch }}
DRY_RUN: ${{ inputs.dry_run }}
run: |
set -euo pipefail
echo "Default branch ref: $DEFAULT_BRANCH_REF"
before=$(gh api "repos/$REPO/actions/cache/usage" \
--jq '(.active_caches_size_in_bytes/1073741824*100|floor)/100')
echo "Cache usage before: ${before} GB"
# `gh cache list` pages at 100; --limit 100 is the documented maximum per call.
gh cache list --repo "$REPO" --limit 100 \
--json id,key,ref,sizeInBytes,lastAccessedAt > caches.json
echo "Entries: $(jq length caches.json)"
# For each key: keep the default-branch copy, else the most recently used. Emit the rest.
jq -r --arg def "$DEFAULT_BRANCH_REF" '
group_by(.key)
| map(
(map(select(.ref == $def)) | first) as $keep
| (sort_by(.lastAccessedAt) | reverse | first) as $newest
| ($keep // $newest) as $winner
| map(select(.id != $winner.id))
)
| flatten
| .[] | "\(.id)\t\(.ref)\t\((.sizeInBytes/1073741824*100|floor)/100)GB\t\(.key)"
' caches.json > victims.tsv
count=$(wc -l < victims.tsv | tr -d ' ')
if [ "$count" -eq 0 ]; then
echo "No duplicate caches to prune."
exit 0
fi
echo "Duplicates to remove ($count):"
cat victims.tsv
if [ "${DRY_RUN:-false}" = "true" ]; then
echo "::notice::dry run — nothing deleted"
exit 0
fi
# A cache can vanish between the list and the delete (another janitor run, GitHub's own
# LRU, a branch deletion). That is the desired end state, so a failed delete must not fail
# the job.
while IFS=$'\t' read -r id ref size key; do
[ -z "${id:-}" ] && continue
if gh api --method DELETE "repos/$REPO/actions/caches/$id" >/dev/null 2>&1; then
echo "deleted $size $ref $key"
else
echo "::warning::could not delete cache $id ($ref $key) — already gone?"
fi
done < victims.tsv
after=$(gh api "repos/$REPO/actions/cache/usage" \
--jq '(.active_caches_size_in_bytes/1073741824*100|floor)/100')
echo "Cache usage after: ${after} GB"
{
echo "### Cache janitor"
echo ""
echo "| | GB |"
echo "|---|---|"
echo "| before | ${before} |"
echo "| after | ${after} |"
echo ""
echo "Removed **${count}** duplicate entries (kept one copy per key)."
} >> "$GITHUB_STEP_SUMMARY"