mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 02:07:25 +08:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every pull request runs the Trusted PR CI workflow before merge > - The test suites roughly tripled in six weeks, and shard balance did not keep up, so PR runs crept from ~4 to ~17 minutes > - Slow CI delays every merge and every contributor > - This pull request rebalances the shards from fresh measurements, splits the largest test files, reuses the Rust build cache in three more jobs, and takes the policy job off the critical path > - The benefit is a PR wall clock near 6 minutes with the same coverage ## Linked Issues or Issue Description **What existing behavior does this improve?** PR CI wall clock. A typical green run took 16-17 minutes. Two months ago it took about 4 minutes. **Subsystem affected** The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the shard-duration manifests, the vitest shard runner scripts, the `paperclip-runner` package scripts, and the dry-run branch of `release.sh`. **Current behavior** The shard-duration manifests were stale. The general-server manifest had durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29 specs. Stale median weights made shard steps range 417s-806s (server) and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release build. Every test lane waited ~60s for the policy job before it could start. **Proposed behavior** All lanes finish in a narrow ~200-290s band. The manifests carry fresh measured durations for every suite. The three largest test files are split so no single file caps a shard. The Rust cache restore runs in every job that builds the Runner binary. Test lanes start as soon as the gate resolves. **Reason and benefit** Merges stop waiting on CI. The projected wall clock is ~6 minutes for the same test coverage. ## What Changed - Rebuild `scripts/general-server-shard-durations.json` (646 suites) and `scripts/e2e-shard-durations.json` (all specs) from per-suite completion timestamps in runs 35036001734 and 35024948947. - Move the PR server lane to the release-verify shape: `general-server-without-chat` across twelve duration-balanced shards, plus the chat integration suite split by collected test location across three dedicated lanes. - Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and `-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions` and `-projects` specs. Each pair shares fixtures through a `.shared.ts` module. Playwright collects the same test sets (39 and 20 tests). - Raise e2e shards to eight and serialized shards to nine. - Run the runner package's `check:all` as four matrix lanes: `check:static`, `check:runner`, and two native vitest `--shard` halves. The union is exactly `check:all`. - Add the read-only Rust cache restore (toolchain pin, `save-if: false`) to the Canary Dry Run, Build, and Typecheck jobs. - Make release.sh preview publish payloads concurrently in batches of eight during `--dry-run`. The real publish path stays strictly serial. - Drop the policy-job lockfile artifact chain. Each lane installs with `--frozen-lockfile` and falls back to an inline `--resolution-only` regeneration. The policy job stays a required check through the `verify` and `e2e` aggregates. - Update the shard-count mirrors and workflow assertions in the partition and gate tests. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/e2e-shard.test.mjs` — 30 pass. - `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass. - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass. - `playwright test --list` collects 39 tests across the chat-adapters split and 20 across the agent-chat split, equal to the original files. - A local vitest collection of the chat suite partitions 995 tests into 498/497 line shards. - Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s x8, serialized ~216s x9. ## Risks - The split spec files reorder tests relative to the original files. Every describe seeds its own company, so the specs stay independent; a hidden cross-describe dependency would surface as a deterministic failure in one shard. - The inline lockfile fallback changes install behavior for manifest-changing and stacked PRs. The policy job still validates resolution as a required check. - `release.sh` changes are confined to the `--dry-run` preview branch. The publish loop is untouched. `bash -n` passes and the release dry-run tests pass. - One PR now schedules ~44 fleet runners. If the RunsOn fleet caps concurrency, queueing may absorb part of the gain; watch the first runs. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with tool use (shell, file edits) in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge
38 lines
2.6 KiB
JSON
38 lines
2.6 KiB
JSON
{
|
|
"$comment": "Per-spec wall-clock durations (ms) for the Playwright e2e lane, used by scripts/e2e-shard.mjs to balance specs across the PR shard matrix. Averaged from two real PR runs of .github/workflows/pr.yml (actions runs 35036001734 and 35024948947, 2026-09-15) by diffing consecutive test-completion timestamps in the 'Run e2e tests' logs, attributing each gap to the finishing test's spec — that captures fixture and navigation time the list-reporter durations alone miss. Smoke Lab tests report under tests/e2e/smoke-lab.shared.ts; their time is mapped back to the p1-p4/p5-p7 spec by scenario tag. Specs missing here get the median weight, so the manifest only needs occasional refreshes: regenerate the same way when e2e shard steps drift apart. chat-adapters-ui and agent-chat were split 2026-09-15 into providers/messaging specs; their weights are the per-describe sums from the same two runs. agent-chat halves are weighted by their per-test sums scaled to the measured spec total.",
|
|
"unit": "ms",
|
|
"durations": {
|
|
"tests/e2e/acp-stop-continuation.spec.ts": 44152,
|
|
"tests/e2e/action-test-agent-picker.spec.ts": 10050,
|
|
"tests/e2e/agent-chat-projects.spec.ts": 159100,
|
|
"tests/e2e/agent-chat-sessions.spec.ts": 151200,
|
|
"tests/e2e/agentmail.spec.ts": 8944,
|
|
"tests/e2e/app-not-connected.spec.ts": 15684,
|
|
"tests/e2e/application-delete-screenshot.spec.ts": 3585,
|
|
"tests/e2e/applications-crud.spec.ts": 10924,
|
|
"tests/e2e/apps-dark-mode-shots.spec.ts": 27798,
|
|
"tests/e2e/apps-prosumer-mcp-flow.spec.ts": 19557,
|
|
"tests/e2e/archived-company-url.spec.ts": 11149,
|
|
"tests/e2e/board-attachment-receipts.spec.ts": 102785,
|
|
"tests/e2e/chat-adapters-ui-messaging.spec.ts": 241600,
|
|
"tests/e2e/chat-adapters-ui-providers.spec.ts": 200200,
|
|
"tests/e2e/composer-stop.spec.ts": 57673,
|
|
"tests/e2e/conference-room-typing-intro.spec.ts": 15196,
|
|
"tests/e2e/connection-intents.spec.ts": 14451,
|
|
"tests/e2e/connection-reviews.spec.ts": 141760,
|
|
"tests/e2e/legacy-failure-continuation.spec.ts": 74050,
|
|
"tests/e2e/mcp-user-stories.spec.ts": 38899,
|
|
"tests/e2e/nux-phase4-screenshots.spec.ts": 18776,
|
|
"tests/e2e/onboarding.spec.ts": 21295,
|
|
"tests/e2e/paused-composer.spec.ts": 33211,
|
|
"tests/e2e/pipelines-tutorial-flow.spec.ts": 30545,
|
|
"tests/e2e/planning-mode-visual-verification.spec.ts": 26193,
|
|
"tests/e2e/project-repositories.spec.ts": 21634,
|
|
"tests/e2e/runner-e2e-dashboard.spec.ts": 2028,
|
|
"tests/e2e/sidebar-takeover.spec.ts": 19034,
|
|
"tests/e2e/signoff-policy.spec.ts": 12121,
|
|
"tests/e2e/smoke-lab-p1-p4.spec.ts": 95529,
|
|
"tests/e2e/smoke-lab-p5-p7.spec.ts": 64328
|
|
}
|
|
}
|