mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 10:16:20 +08:00
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Browser tests check chat state across navigation, reload, and agent runs. > - Runtime tests check that a service is ready before Paperclip publishes its address. > - CI for #13891 failed in these paths, then passed attempt 3 with the same code. > - The browser traces stopped during development asset startup. A separate sidebar assertion used an unstable focus path through the rich editor. > - This PR gives browser tests fresh built assets and a direct keyboard path to the star button. It adds service-worker reload coverage under CPU throttling. > - A later browser failure exposed a real reload race: comments can change between the first query and the first live subscription. The UI now refreshes active queries when that subscription opens. > - Runtime fixtures now have separate registry state and better failure evidence. The readiness deadline remains unchanged. > - A later CI run exposed a wall-clock backoff assertion and three authorization cases sharing one test lifecycle. The PR anchors the assertion to transport time and separates the cases. ## Linked Issues or Issue Description Refs #13891. Related runtime ownership and cleanup work: #11791, #11389, #11278. Evidence: [original CI run, attempt 3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3). Attempt 3 passed both affected shards without code changes. This PR is separate from the wake-payload change, which has since merged. | Failure | Diagnosis and classification | | --- | --- | | Sidebar blank page and retry text missing after reload in CI | Both traces show a blank document before React startup. Vite connects, but the failing page makes no application API requests. The service worker forwards the unbundled development module graph. The retry response is already stored and its process-adapter run succeeded before reload. This places the failure in browser bootstrap, not wake payload or reply persistence. The exact reason the development module graph stopped is not established by the retained trace. The harness now serves built assets, and the new test covers startup and controlled reload at 4x CPU throttling. | | Star opacity remains zero on macOS | Reproduced on unchanged master `f55759942b`. The old test clicked the rich editor, focused the star, then used Tab and Shift+Tab. An instrumented baseline run captured the sequence: Tab moved from the star to the next sidebar link, then the editor bundle called `focus()` on its contenteditable before Shift+Tab. That key reached the composer, where it is a work-mode shortcut. This confirms a pending editor selection update stole focus; it was not the browser skipping sidebar buttons. The assertion did not prove the star retained focus. The test now moves the pointer away, focuses the preceding sidebar link, presses Tab once, and asserts both actual focus and opacity. This is test synchronization and keyboard traversal, not a demonstrated CSS defect. | | Runtime readiness exceeds 10 seconds | The old error only says `fetch failed`. There is no child startup output in that failure, so it cannot distinguish slow process startup from a refused or stalled probe. It passed unchanged locally and took 3.416 seconds in attempt 3. Resource contention is plausible but unproved. This PR does not claim a proven historical runtime root cause: it isolates fixture registry/log files, checks ports before spawn and after stop, checks live backends at publication, and preserves transport errors, probe count, elapsed time, and fixture startup timestamps for the next occurrence. | | Later CI: recovery reply disappears after reload | Product synchronization bug, distinct from the blank-page bootstrap failure. Reproduced locally with a trace: the comments request started at `21:56:07.443`, the server saved the reply at `.520`, and the first live subscription started at `.541`. The reload fetched a successful run and its complete log, but missed the comment event. The provider refreshed after reconnects only. It now refreshes active queries on the first connection too. A deterministic regression test fails before the fix. No browser assertion was changed. | | Later CI: Slack backoff and authorization tests | The 30-second backoff assertion required more than 25 seconds to remain when it read the saved action. CI spent 10.311 seconds in the test, exceeding that five-second allowance. A local six-second read delay reproduces the failure; the new transport-anchored lower and upper bounds pass the same fault injection. The neighbouring authorization test ran three independent fixtures in one test and timed out at 15 seconds. Each case now has its own fixture cleanup and the normal per-test deadline, so earlier cases do not remain active during later global worker sweeps. No specific production slow call was established. | | Self-hosted runner loses communication or shuts down | The original lost-communication failure has no assertion. On final-head [attempt 1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1), Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet instances. All received a runner shutdown signal at `22:04:55 UTC`, within 35 milliseconds, then cancellation. Server shard 6/12 received the same shutdown signal one minute later. All four were Spot `m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build error preceded those stops. This is infrastructure interruption; the reason the fleet stopped the runners is not available in job logs. | ## What Changed - Build the browser fixture UI into the static server's preferred directory, `server/ui-dist`, and disable Vite middleware. Ignore these generated assets. Start the source CLI directly from the repository root, as required by the CLI invocation safety contract. - Keep all existing chat assertions. Add first takeover and three service-worker-controlled reloads under CPU throttling. Assert that the page uses built module assets and that opening it creates no chat task. - Use forward keyboard traversal from the Zeta link to its star. Check focus before checking the reveal style. - Give each runtime exposure test a temporary Paperclip home and restore environment state after process cleanup. - Check that reserved ports are free before spawn, serve the fixture response before exposure, and become free after stop. - Include the nested fetch error, probe count, and elapsed time in readiness failures. Add a unit test for this diagnostic contract. - Refresh active queries when the first live connection opens, covering events missed during initial page loading. Keep reconnect toast suppression unchanged. - Measure Slack retry timing from the transport attempt and recovery completion. This checks the full provider-requested delay without spending a small wall-clock allowance on unrelated processing. - Run each Slack authorization-revocation scenario as a separate test, with cleanup between cases. All assertions remain. - Document the browser fixture's build and serving mode. No configured assertion deadline, readiness deadline, retry count, or skip was added. ## Verification - Final-head [Linux CI, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2): **green**. All 52 check runs passed; two conditional checks were skipped. The legacy Snyk status also passed. No pending or failed checks remain. - `pnpm -r typecheck` passed. - `pnpm build` passed again after the final UI fix. - Focused runtime suites: 38 passed, 3 existing platform skips. - CLI invocation safety suite: 39 passed. - Slack timing negative control: inserting a six-second delay before reading the saved action fails the old assertion. All four revised authorization/backoff cases pass with that same delay. The diagnostic delay is not committed. - Live update suites: 92 passed, including the new regression. UI typecheck and token gates passed. - Recovery browser negative control: the unchanged tests reproduced the missing reply (1 failed, 9 passed). After the first-connection fix, all six recovery paths passed twice (12 passed). Browser assertions and deadlines are unchanged. - Server typecheck passed after the Slack test adjustment. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat --shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong to the remaining shards; collection verified exact coverage. - Final direct-source CLI launch: all five chat session tests passed. - `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed, including nine controlled reloads at 4x CPU throttling. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-server-without-chat --shard-index 10 --shard-count 12`: 55 files passed, 867 tests passed, 21 existing skips. An earlier run could not initialize PostgreSQL because this Mac exhausted its System V shared-memory slots. After reclaiming the orphaned segment from this task's stopped browser server, the full shard passed. - On final commit `1f1fafc08d`, [browser 4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947) passed all 16 tests; [browser 8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155) passed all 23 tests; [server 11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328) passed 887 tests with one existing skip. The readiness lifecycle case took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by runner shutdowns passed unchanged on their single rerun. - Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments. - Full local `pnpm test:run` was attempted and stopped after confirming failures outside this patch: a sibling `skills` directory shadows bundled Slack/AgentMail skills; macOS rejects rename of read-only skill-cache directories (`EACCES`, also reproduced in an isolated filesystem probe); and the host exhausts PostgreSQL System V shared-memory slots. Focused reruns confirmed these limits. The complete Linux CI run is the repository-wide verification; the full local run is not green. - Baseline evidence: the original sidebar test failed on unchanged master; a development-mode run at 4x CPU throttling passed six selected cases, so CPU pressure alone did not reproduce the CI bootstrap stall. ## Risks - Default browser tests now exercise the shipped static UI. They no longer implicitly cover Vite middleware or HMR; use the development server for those checks. - The new browser startup test uses Chromium CDP, matching the only configured browser project. - Opening a live subscription now causes one active-query refresh to close the initial event gap. This adds startup API reads but no recurring poll. - Runtime behavior and deadlines are unchanged except for error details. The historical readiness stall remains unconfirmed; a green rerun alone cannot establish its cause. - Local verification runs on macOS. The runtime lifecycle checks also passed on Linux CI. ## Model Used - OpenAI GPT-6 through Codex. The session identifies the model family as GPT-6; an exact served model ID and context window size are not exposed. Used reasoning, repository inspection, shell tools, code editing, and test execution. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted suites passed; full local host limits are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
100 lines
4.6 KiB
TypeScript
100 lines
4.6 KiB
TypeScript
import fs from "node:fs";
|
|
import os from "node:os";
|
|
import path from "node:path";
|
|
import { defineConfig } from "@playwright/test";
|
|
|
|
// Use a dedicated port so e2e tests always start their own server in local_trusted mode,
|
|
// even when the dev server is running on :3100 in authenticated mode.
|
|
const PORT = Number(process.env.PAPERCLIP_E2E_PORT ?? 3199);
|
|
const BASE_URL = `http://127.0.0.1:${PORT}`;
|
|
const PAPERCLIP_HOME = fs.mkdtempSync(path.join(os.tmpdir(), "paperclip-e2e-home-"));
|
|
const PAPERCLIP_INSTANCE_ID = "playwright-e2e";
|
|
const PAPERCLIP_CONFIG = path.join(PAPERCLIP_HOME, "instances", PAPERCLIP_INSTANCE_ID, "config.json");
|
|
const PAPERCLIP_AGENT_JWT_SECRET = process.env.PAPERCLIP_AGENT_JWT_SECRET ?? "playwright-e2e-agent-jwt-secret";
|
|
const PAPERCLIP_DECISION_SIGNING_SECRET =
|
|
process.env.PAPERCLIP_DECISION_SIGNING_SECRET ?? "playwright-e2e-decision-signing-secret";
|
|
const PAPERCLIP_TOOL_ACTION_SIGNING_SECRET =
|
|
process.env.PAPERCLIP_TOOL_ACTION_SIGNING_SECRET ?? "playwright-e2e-tool-action-signing-secret";
|
|
const PLAYWRIGHT_CHANNEL = process.env.PAPERCLIP_PLAYWRIGHT_CHANNEL;
|
|
|
|
process.env.PAPERCLIP_HOME = PAPERCLIP_HOME;
|
|
process.env.PAPERCLIP_CONFIG = PAPERCLIP_CONFIG;
|
|
// Worker processes reload this config; retain the main process's server path
|
|
// for specs that seed historical database state in the throwaway instance.
|
|
process.env.PAPERCLIP_E2E_SERVER_CONFIG ??= PAPERCLIP_CONFIG;
|
|
// Specs that mint agent JWTs in-process (via createLocalAgentJwt) must derive
|
|
// the same per-instance signing key as the webServer, or verification fails
|
|
// with a 401 instead of authenticating as the agent.
|
|
process.env.PAPERCLIP_INSTANCE_ID = PAPERCLIP_INSTANCE_ID;
|
|
process.env.PAPERCLIP_AGENT_JWT_SECRET = PAPERCLIP_AGENT_JWT_SECRET;
|
|
process.env.PAPERCLIP_DECISION_SIGNING_SECRET = PAPERCLIP_DECISION_SIGNING_SECRET;
|
|
process.env.PAPERCLIP_TOOL_ACTION_SIGNING_SECRET = PAPERCLIP_TOOL_ACTION_SIGNING_SECRET;
|
|
|
|
export default defineConfig({
|
|
testDir: ".",
|
|
testMatch: "**/*.spec.ts",
|
|
// These suites target dedicated multi-user configurations/ports and are
|
|
// intentionally not part of the default local_trusted e2e run.
|
|
testIgnore: ["in-feed-native/**", "multi-user.spec.ts", "multi-user-authenticated.spec.ts"],
|
|
timeout: 60_000,
|
|
retries: 0,
|
|
// All specs share one throwaway server, and several toggle instance-level
|
|
// state (the `enableConferenceRoomChat` experimental flag) that changes
|
|
// which UI variant renders. Run files serially so a flag flip in one spec
|
|
// can't change the wizard/thread under another spec mid-flight.
|
|
workers: 1,
|
|
use: {
|
|
baseURL: BASE_URL,
|
|
headless: true,
|
|
screenshot: "only-on-failure",
|
|
trace: "on-first-retry",
|
|
},
|
|
projects: [
|
|
{
|
|
name: "chromium",
|
|
use: {
|
|
browserName: "chromium",
|
|
...(PLAYWRIGHT_CHANNEL ? { channel: PLAYWRIGHT_CHANNEL } : {}),
|
|
},
|
|
},
|
|
],
|
|
// The webServer directive bootstraps a throwaway instance and then starts it.
|
|
// `onboard --yes --run` works in a non-interactive temp PAPERCLIP_HOME.
|
|
webServer: {
|
|
cwd: path.resolve(import.meta.dirname, "../.."),
|
|
// Exercise the shipped UI. Source-checkout onboarding otherwise enables
|
|
// Vite middleware: every reload traverses thousands of modules, including
|
|
// service-worker-intercepted requests, before React can even start.
|
|
// Build the server's first-choice static directory so a prior package build
|
|
// cannot shadow the UI under test with stale server/ui-dist assets.
|
|
command: "pnpm --filter @paperclipai/ui build --outDir ../server/ui-dist --emptyOutDir && node cli/node_modules/tsx/dist/cli.mjs cli/src/index.ts onboard --yes --run",
|
|
url: `${BASE_URL}/api/health`,
|
|
// Always boot a dedicated throwaway instance for e2e so browser tests
|
|
// never attach to the developer's active Paperclip home/server.
|
|
reuseExistingServer: false,
|
|
timeout: 120_000,
|
|
stdout: "pipe",
|
|
stderr: "pipe",
|
|
env: {
|
|
...process.env,
|
|
NODE_ENV: "test",
|
|
PAPERCLIP_UI_DEV_MIDDLEWARE: "false",
|
|
NODE_OPTIONS: `${process.env.NODE_OPTIONS ?? ""} --import=${path.resolve(import.meta.dirname, "fixtures/agent-chat-github.mjs")}`,
|
|
PORT: String(PORT),
|
|
PAPERCLIP_OPEN_ON_LISTEN: "false",
|
|
PAPERCLIP_API_URL: BASE_URL,
|
|
PAPERCLIP_HOME,
|
|
PAPERCLIP_INSTANCE_ID,
|
|
PAPERCLIP_CONFIG,
|
|
PAPERCLIP_AGENT_JWT_SECRET,
|
|
PAPERCLIP_DECISION_SIGNING_SECRET,
|
|
PAPERCLIP_TOOL_ACTION_SIGNING_SECRET,
|
|
PAPERCLIP_BIND: "loopback",
|
|
PAPERCLIP_DEPLOYMENT_MODE: "local_trusted",
|
|
PAPERCLIP_DEPLOYMENT_EXPOSURE: "private",
|
|
},
|
|
},
|
|
outputDir: "./test-results",
|
|
reporter: [["list"], ["html", { open: "never", outputFolder: "./playwright-report" }]],
|
|
});
|