mirror of
https://github.com/ever-co/ever-gauzy.git
synced 2026-10-02 01:54:50 +08:00
Run 31889575135 ran for 12 h 34 m and produced ZERO test results. Every shard
was killed at its 120-minute timeout inside `Install Packages & Bootstrap
(cache fallback)` — shard 2 spent 1 h 58 m there and never reached step 5.
The cause is a cache that cannot fit. The job split (deps -> build -> e2e)
exists so `yarn bootstrap` is paid once, and it handed the tree over through
`actions/cache` as an EXPLODED ~9 GB, ~900k-file tree. The repository's Actions
cache budget is 10 GB, and `build.yml` runs on every pull_request with five
`actions/setup-node … cache: yarn` jobs, each storing a 4.66 GB yarn cache PER
REF — develop, stage and every open PR hold their own copy of identical
content. Two of those fill the budget by themselves. Measured, not assumed:
- ACROSS runs the e2e cache never survived: deps paid a cold bootstrap every
time (3 h 20 m on run 31848093657, 3 h 46 m on 31889575135).
- WITHIN a run it was a coin flip: on 31848093657 all four shards hit it; on
31889575135 the build job hit it and shard 2 missed an hour later.
So the tree now travels as ONE compressed archive. An artifact is the handoff —
it is scoped to the run and cannot be evicted mid-run by another workflow's
cache write — and the same file is also saved to `actions/cache` as a
best-effort cross-run fast path, where ~3 GB has a chance that 9 GB never had.
Downstream jobs download and extract it, keeping the full-bootstrap fallback
for the case where the artifact is unavailable. zstd with a gzip fallback,
since only the archive's contents are compressible and one big file avoids the
cache action's per-file cost in both directions. A size floor rejects a
truncated archive rather than publishing one that would send every downstream
job back to a multi-hour bootstrap.
Also adds a cache janitor. It keeps exactly one copy per cache key — the
default branch's if present, else the most recently used — and deletes the
duplicates that differ only by ref. Verified against the live cache list: it
selects the redundant `refs/pull/9992/merge` copy of a key already held on
`refs/heads/stage` and reclaims 4.66 GB, touching nothing else. Cache entries
are reconstructible by definition, so pruning duplicates is strictly better
than letting GitHub's LRU pick a victim at random, which is what happens today.
This leaves `build.yml` alone deliberately: it is the PR gate, and reclaiming
the budget does not require changing how it builds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
120 lines
4.7 KiB
YAML
120 lines
4.7 KiB
YAML
name: Cache Janitor
|
|
|
|
# The repository's Actions cache budget is 10 GB. `build.yml` runs on every pull_request with five
|
|
# `actions/setup-node@v4 … cache: yarn` jobs, and setup-node saves its ~4.66 GB yarn cache PER REF —
|
|
# so `refs/heads/develop`, `refs/heads/stage` and every open PR's `refs/pull/N/merge` each hold their
|
|
# own copy of the SAME content, keyed identically. Two of those fill the budget on their own.
|
|
#
|
|
# The damage is not theoretical. On 2026-08-15 the Playwright suite's dependency cache was gone
|
|
# within the hour it was written, so run 31889575135 paid a cold `yarn bootstrap` in the deps job
|
|
# (3 h 46 m) and its shards started doing the same again — the exact cost the deps/build/e2e job
|
|
# split exists to pay only once.
|
|
#
|
|
# Cache entries are, by definition, reconstructible: the worst case for deleting one is that the next
|
|
# run that wants it re-downloads. That makes pruning duplicates strictly better than letting GitHub's
|
|
# LRU choose a victim at random, which is what happens today.
|
|
#
|
|
# RULE: for each cache key, keep exactly ONE copy — the one on the default branch if it exists,
|
|
# otherwise the most recently used. Delete the rest. Branch-scoped restores are allowed to read the
|
|
# default branch's caches, so the kept copy still serves every PR.
|
|
|
|
on:
|
|
schedule:
|
|
# Every 6 hours. Frequent enough that duplicates never accumulate for long, rare enough that it
|
|
# never races a build that is mid-save.
|
|
- cron: '17 */6 * * *'
|
|
workflow_dispatch:
|
|
inputs:
|
|
dry_run:
|
|
description: 'List what would be deleted without deleting anything'
|
|
type: boolean
|
|
default: false
|
|
|
|
permissions:
|
|
contents: read
|
|
actions: write # required to delete cache entries
|
|
|
|
concurrency:
|
|
group: ${{ github.workflow }}
|
|
cancel-in-progress: false
|
|
|
|
jobs:
|
|
prune:
|
|
name: Prune duplicate caches
|
|
runs-on: ubuntu-latest
|
|
timeout-minutes: 15
|
|
|
|
steps:
|
|
- name: Prune
|
|
shell: bash
|
|
env:
|
|
GH_TOKEN: ${{ github.token }}
|
|
REPO: ${{ github.repository }}
|
|
DEFAULT_BRANCH_REF: refs/heads/${{ github.event.repository.default_branch }}
|
|
DRY_RUN: ${{ inputs.dry_run }}
|
|
run: |
|
|
set -euo pipefail
|
|
|
|
echo "Default branch ref: $DEFAULT_BRANCH_REF"
|
|
before=$(gh api "repos/$REPO/actions/cache/usage" \
|
|
--jq '(.active_caches_size_in_bytes/1073741824*100|floor)/100')
|
|
echo "Cache usage before: ${before} GB"
|
|
|
|
# `gh cache list` pages at 100; --limit 100 is the documented maximum per call.
|
|
gh cache list --repo "$REPO" --limit 100 \
|
|
--json id,key,ref,sizeInBytes,lastAccessedAt > caches.json
|
|
echo "Entries: $(jq length caches.json)"
|
|
|
|
# For each key: keep the default-branch copy, else the most recently used. Emit the rest.
|
|
jq -r --arg def "$DEFAULT_BRANCH_REF" '
|
|
group_by(.key)
|
|
| map(
|
|
(map(select(.ref == $def)) | first) as $keep
|
|
| (sort_by(.lastAccessedAt) | reverse | first) as $newest
|
|
| ($keep // $newest) as $winner
|
|
| map(select(.id != $winner.id))
|
|
)
|
|
| flatten
|
|
| .[] | "\(.id)\t\(.ref)\t\((.sizeInBytes/1073741824*100|floor)/100)GB\t\(.key)"
|
|
' caches.json > victims.tsv
|
|
|
|
count=$(wc -l < victims.tsv | tr -d ' ')
|
|
if [ "$count" -eq 0 ]; then
|
|
echo "No duplicate caches to prune."
|
|
exit 0
|
|
fi
|
|
|
|
echo "Duplicates to remove ($count):"
|
|
cat victims.tsv
|
|
|
|
if [ "${DRY_RUN:-false}" = "true" ]; then
|
|
echo "::notice::dry run — nothing deleted"
|
|
exit 0
|
|
fi
|
|
|
|
# A cache can vanish between the list and the delete (another janitor run, GitHub's own
|
|
# LRU, a branch deletion). That is the desired end state, so a failed delete must not fail
|
|
# the job.
|
|
while IFS=$'\t' read -r id ref size key; do
|
|
[ -z "${id:-}" ] && continue
|
|
if gh api --method DELETE "repos/$REPO/actions/caches/$id" >/dev/null 2>&1; then
|
|
echo "deleted $size $ref $key"
|
|
else
|
|
echo "::warning::could not delete cache $id ($ref $key) — already gone?"
|
|
fi
|
|
done < victims.tsv
|
|
|
|
after=$(gh api "repos/$REPO/actions/cache/usage" \
|
|
--jq '(.active_caches_size_in_bytes/1073741824*100|floor)/100')
|
|
echo "Cache usage after: ${after} GB"
|
|
{
|
|
echo "### Cache janitor"
|
|
echo ""
|
|
echo "| | GB |"
|
|
echo "|---|---|"
|
|
echo "| before | ${before} |"
|
|
echo "| after | ${after} |"
|
|
echo ""
|
|
echo "Removed **${count}** duplicate entries (kept one copy per key)."
|
|
} >> "$GITHUB_STEP_SUMMARY"
|