mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 02:07:25 +08:00
1b86af995ff4b69b56cdf170f753e84c5af33042
3708
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1b86af995f |
build(deps-dev): bump @storybook/addon-docs from 10.5.0 to 10.5.8 (#11513)
Bumps [@storybook/addon-docs](https://github.com/storybookjs/storybook/tree/HEAD/code/addons/docs) from 10.5.0 to 10.5.8. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/storybookjs/storybook/releases">@storybook/addon-docs's releases</a>.</em></p> <blockquote> <h2>v10.5.8</h2> <h2>10.5.8</h2> <ul> <li>React: Fix RDT tsconfig selection for Vite project references - <a href="https://redirect.github.com/storybookjs/storybook/pull/35743">#35743</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>Tanstack React: Remove <code>@cloudflare/vite-plugin</code> from the inherited Vite config - <a href="https://redirect.github.com/storybookjs/storybook/pull/35706">#35706</a>, thanks <a href="https://github.com/FrancoKaddour"><code>@FrancoKaddour</code></a>!</li> <li>Tanstack: Wait for router to load before rendering - <a href="https://redirect.github.com/storybookjs/storybook/pull/35784">#35784</a>, thanks <a href="https://github.com/huang-julien"><code>@huang-julien</code></a>!</li> <li>Test: Fix Illegal invocation when reading prototype.focus - <a href="https://redirect.github.com/storybookjs/storybook/pull/35528">#35528</a>, thanks <a href="https://github.com/FrancoKaddour"><code>@FrancoKaddour</code></a>!</li> </ul> <h2>v10.5.7</h2> <h2>10.5.7</h2> <ul> <li>Angular: Serve ancestor node_modules for addon-vitest in browser mode - <a href="https://redirect.github.com/storybookjs/storybook/pull/35600">#35600</a>, thanks <a href="https://github.com/brandonroberts"><code>@brandonroberts</code></a>!</li> <li>Refactor: Update getVersionedPackages method to handle non-Storybook packages correctly - <a href="https://redirect.github.com/storybookjs/storybook/pull/35769">#35769</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> </ul> <h2>v10.5.6</h2> <h2>10.5.6</h2> <ul> <li>Dependencies: Pin `@testing-library/jest-dom` to `6.9.1` - <a href="https://redirect.github.com/storybookjs/storybook/pull/35614">#35614</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>ESLint Plugin: Add plugin meta and document oxlint usage - <a href="https://redirect.github.com/storybookjs/storybook/pull/35655">#35655</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> <li>Vue: Skip docgen for module ids carrying a query - <a href="https://redirect.github.com/storybookjs/storybook/pull/35598">#35598</a>, thanks <a href="https://github.com/seanogdev"><code>@seanogdev</code></a>!</li> </ul> <h2>v10.5.5</h2> <h2>10.5.5</h2> <ul> <li>CLI: Update AI setup instructions to msw-storybook-addon v3 - <a href="https://redirect.github.com/storybookjs/storybook/pull/35512">#35512</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> <li>Core: Upgrade `ws` to fix security advisories - <a href="https://redirect.github.com/storybookjs/storybook/pull/35584">#35584</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>ReactNative: Telemetry framework detection fix - <a href="https://redirect.github.com/storybookjs/storybook/pull/35560">#35560</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>SyntaxHighlighter: Fix PrismJS dark mode mismatch - <a href="https://redirect.github.com/storybookjs/storybook/pull/35541">#35541</a>, thanks <a href="https://github.com/hxy-asdw"><code>@hxy-asdw</code></a>!</li> <li>TanStack: Preserve explicit route ids on pathful clones - <a href="https://redirect.github.com/storybookjs/storybook/pull/35499">#35499</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> <li>TanStack: Resolve mock redirects through Vite's resolver - <a href="https://redirect.github.com/storybookjs/storybook/pull/35501">#35501</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> <li>TanStack: Respect routeOverrides component overrides in stories - <a href="https://redirect.github.com/storybookjs/storybook/pull/35497">#35497</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> </ul> <h2>v10.5.4</h2> <h2>10.5.4</h2> <ul> <li>ReactNative: Telemetry framework detection fix - <a href="https://redirect.github.com/storybookjs/storybook/pull/35560">#35560</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>SyntaxHighlighter: Fix PrismJS dark mode mismatch - <a href="https://redirect.github.com/storybookjs/storybook/pull/35541">#35541</a>, thanks <a href="https://github.com/hxy-asdw"><code>@hxy-asdw</code></a>!</li> </ul> <h2>v10.5.3</h2> <h2>10.5.3</h2> <ul> <li>Dependencies: Upgrade TypeScript to 6.0.3 - <a href="https://redirect.github.com/storybookjs/storybook/pull/34971">#34971</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> </ul> <h2>v10.5.2</h2> <h2>10.5.2</h2> <ul> <li>Angular-Vite: Drop <code>@angular/platform-browser-dynamic</code> peer dependency - <a href="https://redirect.github.com/storybookjs/storybook/pull/35457">#35457</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> <li>Angular-Vite: Widen TypeScript peer dependency range to support TypeScript 6 - <a href="https://redirect.github.com/storybookjs/storybook/pull/35455">#35455</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> <li>Core: Include chromatic packages in ecosystem identifier - <a href="https://redirect.github.com/storybookjs/storybook/pull/35170">#35170</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> <li>TanStack: Fix createServerFn validator mock - <a href="https://redirect.github.com/storybookjs/storybook/pull/35185">#35185</a>, thanks <a href="https://github.com/sjh9714"><code>@sjh9714</code></a>!</li> </ul> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/storybookjs/storybook/blob/next/CHANGELOG.md">@storybook/addon-docs's changelog</a>.</em></p> <blockquote> <h2>10.5.8</h2> <ul> <li>React: Fix RDT tsconfig selection for Vite project references - <a href="https://redirect.github.com/storybookjs/storybook/pull/35743">#35743</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>Tanstack React: Remove <code>@cloudflare/vite-plugin</code> from the inherited Vite config - <a href="https://redirect.github.com/storybookjs/storybook/pull/35706">#35706</a>, thanks <a href="https://github.com/FrancoKaddour"><code>@FrancoKaddour</code></a>!</li> <li>Tanstack: Wait for router to load before rendering - <a href="https://redirect.github.com/storybookjs/storybook/pull/35784">#35784</a>, thanks <a href="https://github.com/huang-julien"><code>@huang-julien</code></a>!</li> <li>Test: Fix Illegal invocation when reading prototype.focus - <a href="https://redirect.github.com/storybookjs/storybook/pull/35528">#35528</a>, thanks <a href="https://github.com/FrancoKaddour"><code>@FrancoKaddour</code></a>!</li> </ul> <h2>10.5.7</h2> <ul> <li>Angular: Serve ancestor node_modules for addon-vitest in browser mode - <a href="https://redirect.github.com/storybookjs/storybook/pull/35600">#35600</a>, thanks <a href="https://github.com/brandonroberts"><code>@brandonroberts</code></a>!</li> <li>Refactor: Update getVersionedPackages method to handle non-Storybook packages correctly - <a href="https://redirect.github.com/storybookjs/storybook/pull/35769">#35769</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> </ul> <h2>10.5.6</h2> <ul> <li>Dependencies: Pin <code>@testing-library/jest-dom</code> to <code>6.9.1</code> - <a href="https://redirect.github.com/storybookjs/storybook/pull/35614">#35614</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>ESLint Plugin: Add plugin meta and document oxlint usage - <a href="https://redirect.github.com/storybookjs/storybook/pull/35655">#35655</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> <li>Vue: Skip docgen for module ids carrying a query - <a href="https://redirect.github.com/storybookjs/storybook/pull/35598">#35598</a>, thanks <a href="https://github.com/seanogdev"><code>@seanogdev</code></a>!</li> </ul> <h2>10.5.5</h2> <ul> <li>CLI: Update AI setup instructions to msw-storybook-addon v3 - <a href="https://redirect.github.com/storybookjs/storybook/pull/35512">#35512</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> <li>Core: Upgrade <code>ws</code> to fix security advisories - <a href="https://redirect.github.com/storybookjs/storybook/pull/35584">#35584</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>ReactNative: Telemetry framework detection fix - <a href="https://redirect.github.com/storybookjs/storybook/pull/35560">#35560</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>SyntaxHighlighter: Fix PrismJS dark mode mismatch - <a href="https://redirect.github.com/storybookjs/storybook/pull/35541">#35541</a>, thanks <a href="https://github.com/hxy-asdw"><code>@hxy-asdw</code></a>!</li> <li>TanStack: Preserve explicit route ids on pathful clones - <a href="https://redirect.github.com/storybookjs/storybook/pull/35499">#35499</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> <li>TanStack: Resolve mock redirects through Vite's resolver - <a href="https://redirect.github.com/storybookjs/storybook/pull/35501">#35501</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> <li>TanStack: Respect routeOverrides component overrides in stories - <a href="https://redirect.github.com/storybookjs/storybook/pull/35497">#35497</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> </ul> <h2>10.5.4</h2> <ul> <li>ReactNative: Telemetry framework detection fix - <a href="https://redirect.github.com/storybookjs/storybook/pull/35560">#35560</a>, thanks <a href="https://github.com/ndelangen"><code>@ndelangen</code></a>!</li> <li>SyntaxHighlighter: Fix PrismJS dark mode mismatch - <a href="https://redirect.github.com/storybookjs/storybook/pull/35541">#35541</a>, thanks <a href="https://github.com/hxy-asdw"><code>@hxy-asdw</code></a>!</li> </ul> <h2>10.5.3</h2> <ul> <li>Dependencies: Upgrade TypeScript to 6.0.3 - <a href="https://redirect.github.com/storybookjs/storybook/pull/34971">#34971</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> </ul> <h2>10.5.2</h2> <ul> <li>TanStack: Fix createServerFn validator mock - <a href="https://redirect.github.com/storybookjs/storybook/pull/35185">#35185</a>, thanks <a href="https://github.com/sjh9714"><code>@sjh9714</code></a>!</li> <li>TanStack: Support pathless layout routes (id-only) in story routing - <a href="https://redirect.github.com/storybookjs/storybook/pull/35465">#35465</a>, thanks <a href="https://github.com/unpunnyfuns"><code>@unpunnyfuns</code></a>!</li> <li>Tanstack-react: Add missing Hydrate export - <a href="https://redirect.github.com/storybookjs/storybook/pull/35111">#35111</a>, thanks <a href="https://github.com/arun-357"><code>@arun-357</code></a>!</li> <li>Tanstack-react: Keep JSX-only component references during dead-code elimination - <a href="https://redirect.github.com/storybookjs/storybook/pull/35206">#35206</a>, thanks <a href="https://github.com/yatishgoel"><code>@yatishgoel</code></a>!</li> <li>Vitest: Fix coverage toggle crash on Vite 6 by clearing closed Vitest instance on restart - <a href="https://redirect.github.com/storybookjs/storybook/pull/35461">#35461</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> </ul> <h2>10.5.1</h2> <ul> <li>Angular-Vite: Drop <code>@angular/platform-browser-dynamic</code> peer dependency - <a href="https://redirect.github.com/storybookjs/storybook/pull/35457">#35457</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> <li>Angular-Vite: Widen TypeScript peer dependency range to support TypeScript 6 - <a href="https://redirect.github.com/storybookjs/storybook/pull/35455">#35455</a>, thanks <a href="https://github.com/valentinpalkovic"><code>@valentinpalkovic</code></a>!</li> <li>Core: Include chromatic packages in ecosystem identifier - <a href="https://redirect.github.com/storybookjs/storybook/pull/35170">#35170</a>, thanks <a href="https://github.com/yannbf"><code>@yannbf</code></a>!</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/storybookjs/storybook/commit/6ef7d1ae816ebd5fb8bf84b8dec7d4a92410d73c"><code>6ef7d1a</code></a> Bump version from "10.5.7" to "10.5.8" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/7c6fb3a5ecf4495d73de6d70f802251934e079bd"><code>7c6fb3a</code></a> Bump version from "10.5.6" to "10.5.7" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/3126f0a14a351e971679f61cd0bd5609d086ed57"><code>3126f0a</code></a> Bump version from "10.5.5" to "10.5.6" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/05a52b7a888c6b85c3f8aa6765ed9a0a69a79e4c"><code>05a52b7</code></a> Bump version from "10.5.4" to "10.5.5" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/3327dc44697304275e28ceaa2cd34d9bede4e333"><code>3327dc4</code></a> Bump version from "10.5.3" to "10.5.4" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/9ac273930a49ad33b6a331f1f36dc472f5c36054"><code>9ac2739</code></a> Bump version from "10.5.2" to "10.5.3" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/b4b00f27662caf7328330b5ab47e9b903218f43a"><code>b4b00f2</code></a> Merge pull request <a href="https://github.com/storybookjs/storybook/tree/HEAD/code/addons/docs/issues/34971">#34971</a> from storybookjs/valentin/upgrade-typescript-6</li> <li><a href="https://github.com/storybookjs/storybook/commit/518f711cb367d8df184be22f1fab9a218b2743df"><code>518f711</code></a> Bump version from "10.5.1" to "10.5.2" [skip ci]</li> <li><a href="https://github.com/storybookjs/storybook/commit/c253a0667d39899a0f7a09fa90f262f69ad4ae90"><code>c253a06</code></a> Bump version from "10.5.0" to "10.5.1" [skip ci]</li> <li>See full diff in <a href="https://github.com/storybookjs/storybook/commits/v10.5.8/code/addons/docs">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
93548e7f77 |
build(deps-dev): bump @playwright/test from 1.61.1 to 1.62.1 (#10734)
Bumps [@playwright/test](https://github.com/microsoft/playwright) from 1.61.1 to 1.62.1. <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/microsoft/playwright/releases">@playwright/test's releases</a>.</em></p> <blockquote> <h2>v1.62.1</h2> <h3>Bug Fixes</h3> <ul> <li><a href="https://redirect.github.com/microsoft/playwright/issues/41989">#41989</a> [Regression]: tsconfig "extends" bare specifier isn't resolved via node_modules walk-up like tsc (fatal since 1.62)</li> <li><a href="https://redirect.github.com/microsoft/playwright/issues/41998">#41998</a> [Regression]: directory-form tsconfig project references ("path": "../pkg") fail to resolve (fatal since 1.62)</li> <li><a href="https://redirect.github.com/microsoft/playwright/issues/41985">#41985</a> Accessibility snapshot drops button name when text is nested inside spans with aria-hidden SVG</li> <li><a href="https://redirect.github.com/microsoft/playwright/issues/42000">#42000</a> [Regression]: page.evaluate() arg of a branded primitive type (string & { brand }) no longer type-checks since 1.62</li> <li><a href="https://redirect.github.com/microsoft/playwright/issues/42013">#42013</a> [BUG]Image-type actionable elements are not presented in the snapshot.</li> </ul> <h2>v1.62.0</h2> <h2>🧱 New component testing model</h2> <p><a href="https://playwright.dev/docs/test-components">Component testing</a> moves to a <strong>stories and galleries</strong> model. A <strong>story</strong> wraps your component in one specific scenario — hard-coded props, mock data, providers — and a <strong>gallery</strong> page that you serve renders stories on demand. The new <a href="https://playwright.dev/docs/api/class-fixtures#fixtures-mount">fixtures.mount()</a> fixture navigates to the gallery, mounts a story by id, and returns a <a href="https://playwright.dev/docs/api/class-locator">Locator</a> scoped to the story's root element:</p> <pre lang="js"><code>test('click should expand', async ({ mount }) => { const component = await mount('components/Expandable/Stateful'); await component.getByRole('button').click(); await expect(component.getByTestId('expanded')).toHaveValue('true'); }); </code></pre> <p>Pass a story type as a template argument to type-check its props, and use <code>update(props)</code> / <code>unmount()</code> on the returned locator to re-render or tear down within a test.</p> <h2>🛑 Cancel operations with AbortSignal</h2> <p>Most operations and web-first assertions now accept a <code>signal</code> option that takes an <a href="https://developer.mozilla.org/en-US/docs/Web/API/AbortSignal"><code>AbortSignal</code></a>, letting you cancel long-running actions, navigations, waits, and assertions:</p> <pre lang="js"><code>const controller = new AbortController(); setTimeout(() => controller.abort(), 1000); <p>await page.getByRole('button', { name: 'Submit' }).click({ signal: controller.signal }); await expect(page.getByText('Done')).toBeVisible({ signal: controller.signal }); </code></pre></p> <p>Providing a signal does not disable the default timeout; pass <code>timeout: 0</code> to disable it.</p> <h2>🖼️ WebP screenshots</h2> <p><a href="https://playwright.dev/docs/api/class-pageassertions#page-assertions-to-have-screenshot-1">expect(page).toHaveScreenshot()</a> and <a href="https://playwright.dev/docs/api/class-locatorassertions#locator-assertions-to-have-screenshot-1">expect(locator).toHaveScreenshot()</a> can now store snapshots in the WebP format — just give the snapshot a <code>.webp</code> name:</p> <pre lang="js"><code>// Visual comparisons store the golden snapshot as lossless WebP. await expect(page).toHaveScreenshot('homepage.webp'); <p>// Standalone screenshots can trade quality for size with lossy WebP. await page.screenshot({ path: 'homepage.webp', quality: 50 }); </tr></table> </code></pre></p> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/microsoft/playwright/commit/26a9e470a7b3c7822084b09fb7f13902c5f37b51"><code>26a9e47</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/42043">#42043</a>): docs: release notes for v1.62 Python, Java, and .NET (<a href="https://redirect.github.com/microsoft/playwright/issues/4">#4</a>...</li> <li><a href="https://github.com/microsoft/playwright/commit/0a81d5d09b10eeefe228fe745c3f80c7368a239b"><code>0a81d5d</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/42040">#42040</a>): docs(release-notes): mention the isolated headless clipb...</li> <li><a href="https://github.com/microsoft/playwright/commit/83768264e64a821bcef9e634b8e5c33897f2b032"><code>8376826</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/42034">#42034</a>): fix(aria): keep icon-only clickable elements in ai snaps...</li> <li><a href="https://github.com/microsoft/playwright/commit/66c5cc92a60ce20ab3abe779339e1f90d2e2e888"><code>66c5cc9</code></a> chore: mark v1.62.1 (<a href="https://redirect.github.com/microsoft/playwright/issues/42020">#42020</a>)</li> <li><a href="https://github.com/microsoft/playwright/commit/9672bc3f2a7098cb6a9791ca97222187363a3037"><code>9672bc3</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/42009">#42009</a>): fix(types): support branded primitives in evaluate argum...</li> <li><a href="https://github.com/microsoft/playwright/commit/4325804427a214aa0c8c39bb1352f4ac4f712fd1"><code>4325804</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/41988">#41988</a>): fix(aria): preserve names from collapsed text contributors</li> <li><a href="https://github.com/microsoft/playwright/commit/9632f8ecbc2accba140ea342f1070ccfdd5f5d41"><code>9632f8e</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/42005">#42005</a>): fix(tsconfig): do not throw when "extends"/"references" ...</li> <li><a href="https://github.com/microsoft/playwright/commit/e3950d9c140d007bd52853b45813c6274b24e36f"><code>e3950d9</code></a> chore: mark v1.62.0 (<a href="https://redirect.github.com/microsoft/playwright/issues/41981">#41981</a>)</li> <li><a href="https://github.com/microsoft/playwright/commit/f07e0f720fbe6691cc3d3d66ff9f3e58139e804c"><code>f07e0f7</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/41940">#41940</a>): docs: release notes for v1.62 (<a href="https://redirect.github.com/microsoft/playwright/issues/41967">#41967</a>)</li> <li><a href="https://github.com/microsoft/playwright/commit/05a306c78f11767535fd986eebab5d4c4dad4614"><code>05a306c</code></a> cherry-pick(<a href="https://redirect.github.com/microsoft/playwright/issues/41964">#41964</a>): Revert "feat(routeFromHar): add interceptAPIRequests opt...</li> <li>Additional commits viewable in <a href="https://github.com/microsoft/playwright/compare/v1.61.1...v1.62.1">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
aad97d93fe |
fix(hermes): surface real reasoning text from reasoning.available events (#9237)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When an agent runs through the Hermes gateway adapter, its stdout is parsed line-by-line into transcript entries that the issue chat renders (the UI fetches the adapter's `./ui-parser` from `/api/adapters/:type/ui-parser.js` and runs `parseStdoutLine` client-side) > - Reasoning-capable models emit a `reasoning.available` gateway event carrying the model's reasoning text, and the chat renders `thinking` parts as expandable chain-of-thought > - The gateway parser mapped `reasoning.available` to a hardcoded `"Hermes reasoning available"` string and discarded the event payload, so the "thinking" part had no real content — the indicator looked static and expanding it revealed nothing (#9209) > - This pull request extracts the actual reasoning text from the event payload and uses it as the `thinking` part's text, keeping the old string only as a fallback for payloads that carry no text > - The benefit is that the "Hermes reasoning available" indicator now surfaces the model's real reasoning, which the existing expandable-thinking UI can display ## Linked Issues or Issue Description Fixes: #9209 ## What Changed - `packages/adapters/hermes/src/gateway/ui/parse-stdout.ts`: the `reasoning.available` handler now extracts the reasoning text from the event `data` via a small helper (`extractReasoningText`), checking the plausible field names (`reasoning`, `reasoning_text`, `thinking`, `text`, `summary`, `content`) and recursing one level into nested `data` / `payload` records, with ANSI stripped. The prior `"Hermes reasoning available"` string is kept only as a fallback when no text field is present. - `packages/adapters/hermes/gateway-ui-parser.cjs`: applied the identical logical change to the committed CommonJS mirror (exported as `./gateway/ui-parser`), keeping the two files in sync. - `packages/adapters/hermes/src/gateway/ui/parse-stdout.test.ts` (new): unit tests for the gateway parser (there were none) covering direct-field, `summary`, nested `data`/`payload` extraction, the no-text fallback, and regression guards for `message.delta` and plain stdout. ## Verification Ran from `packages/adapters/hermes`: - `node_modules/.bin/vitest run src/gateway/ui/parse-stdout.test.ts` → **8/8 passed**. - Negative control: stashed the source changes and re-ran the same test file against the current (pre-patch) parser → **4/8 failed** (exactly the reasoning-extraction assertions), then restored — confirming the tests are discriminating, not vacuous. - `npx tsc --noEmit -p .` → clean. Real-behavior proof (driving the actual shipped `gateway-ui-parser.cjs` `parseStdoutLine`) is in the block below. ## Risks - **Low risk.** Behavior is unchanged for events that carry no recognizable text field — the `"Hermes reasoning available"` fallback is preserved (verified). Only the `reasoning.available` branch changed; `message.delta`, `run.failed`/`run.error`, and the generic/system/stdout branches are untouched. - The exact field name in a real `reasoning.available` payload is defined by the external Hermes gateway and is not present anywhere in this repo, so the extraction is intentionally defensive across several plausible field names rather than pinned to one. If the real event nests the text differently than `data` / `payload`, it will fall back to the existing placeholder (i.e. no regression vs. today). Happy to tighten the field list against real gateway traffic if a maintainer can share a sample. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`) via Claude Code, with tool use and local test execution (ran vitest/tsc against the change). Planning, code review, and the real-behavior proof were done with Claude (Opus 4.8) in the same session. ## Real behavior proof **Behavior addressed:** A `reasoning.available` Hermes gateway event now produces a `thinking` transcript part containing the model's real reasoning text, instead of a static `"Hermes reasoning available"` placeholder with no content behind it (#9209). **Real environment tested:** Drove the actual shipped production artifact — `packages/adapters/hermes/gateway-ui-parser.cjs`, the exact module the UI loads via `/api/adapters/hermes-gateway/ui-parser.js` and runs to parse gateway stdout — on Node v24.16.0, macOS. The input is a raw stdout line in the exact format emitted by `packages/adapters/hermes/src/gateway/server/execute.ts` (`[hermes-gateway:event] run=… event=reasoning.available data=…`). Only the external gateway boundary (the raw line) is synthesized; the parser code path is the real one. **Exact steps or command run after this patch:** ``` # BEFORE = git show HEAD:…/gateway-ui-parser.cjs ; AFTER = patched artifact node proof.cjs # requires each parser build and calls parseStdoutLine(line, ts) # line = [hermes-gateway:event] run=run-abc123 event=reasoning.available \ # data={"text":"Checking whether the cache key includes the tenant id before I refactor the lookup."} ``` **Evidence after fix:** ``` ===== BEFORE (master / old code) ===== [ { "kind": "thinking", "ts": "…", "text": "Hermes reasoning available" } ] thinking part carries real reasoning text? -> NO (static placeholder, nothing for the UI to expand) ===== AFTER (this patch) ===== [ { "kind": "thinking", "ts": "…", "text": "Checking whether the cache key includes the tenant id before I refactor the lookup." } ] thinking part carries real reasoning text? -> YES ``` Additional cases through the same shipped artifact after the patch: ``` -- nested payload (data.payload.reasoning) -- {"kind":"thinking","ts":"…","text":"Weighing two migration orders."} -- bare signal, no text field (regression guard) -- {"kind":"thinking","ts":"…","text":"Hermes reasoning available"} # fallback preserved -- message.delta still works (regression guard) -- {"kind":"assistant","ts":"…","text":"Hello","delta":true} ``` **Observed result after fix:** The `reasoning.available` event yields a `thinking` part carrying the model's real reasoning text (top-level or nested), which the existing expandable-thinking rendering in the chat can display. Events with no text field still yield the original placeholder, and unrelated events are unaffected. **What was not tested:** I did not run against a live Hermes gateway — Paperclip's Hermes gateway binary and its credentials aren't available on this machine, and no captured real `reasoning.available` payload exists in the repo, so the exact wire field name is inferred (hence the defensive multi-field extraction + safe fallback). I also did not render the full React chat component in jsdom; the change is confined to the parser, and the chat's expandable `thinking` rendering already exists (`ui/src/components/IssueChatThread.tsx`). CI / unit tests here are supplemental to the runtime proof above. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (searched `9209 in:body` and keyword variants — none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (`fix/hermes-reasoning-available-payload`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (no user-facing docs describe this behavior; none needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (will confirm once CI runs on the PR) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (will address on review) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
120ae5428f |
feat(server): add a one-click relink action for detached custom-image templates (#11641)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip environments can use captured custom images for agent runs > - A configuration fingerprint change can detach a valid custom-image template > - Operators need a safe way to confirm that the image still matches the boot source > - This pull request adds a guarded relink action with drift classification and audit logging > - The benefit is a deliberate relink without a new sandbox boot or provider snapshot ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting environment, server, and UI behavior. **Problem or motivation** A custom-image template detaches when the environment configuration fingerprint changes. The runtime then uses the base image, even when the boot source did not change. The only prior remedy required a full re-capture. **Proposed solution** Add an operator-triggered relink action. Classify configuration drift from a server-owned boot-relevant snapshot. Relink knob-only drift without confirmation. Require explicit confirmation for boot-source or unclassified drift. Guard the route for instance administrators and record a safe activity event. **Alternatives considered** Keep requiring a full re-capture. This adds a sandbox boot and provider snapshot for cases where the image remains correct. **Roadmap alignment** The roadmap has no matching custom-image relink item. This change addresses an environment operation gap. **Additional context** The relink response exposes raw drift values only in the transient 409 response to the instance administrator. The service never persists or logs fingerprints or configuration values. Reserved identity-path segments fail closed. ## What Changed - Add `relinkActiveTemplate` with drift classification and conditional fingerprint update. - Persist a server-owned boot-relevant configuration snapshot during capture. - Add the guarded relink route with strict request validation and activity logging. - Add the relink action and confirmation flow to the environment page. - Add service, route, UI, and OpenAPI coverage. ## Verification - Run the focused service suite: `pnpm vitest run server/src/services/environment-custom-images-service.test.ts`. - Run the focused route suite: `pnpm vitest run server/src/routes/environment-custom-image-routes.test.ts`. - Run the focused UI suite: `pnpm vitest run ui/src/pages/CompanyEnvironments.test.tsx`. - Run server and UI TypeScript checks. - Confirm the OpenAPI snapshot matches the new route. - Confirm all required GitHub checks pass on commit `e46fdcfe94a719be854adf8849d30714e5b70b93`. - Confirm Greptile reports 5/5 with no unresolved review threads. ## Risks The relink action can keep an image after configuration drift. The service requires explicit confirmation for boot-source or unclassified drift. Reserved path segments produce a safe unresolved marker and never enter stored values. ## Model Used OpenAI GPT-5 Codex. The model used repository inspection, GitHub operations, and PR preparation with tool use and code execution. The runtime did not expose a context-window value or a separate reasoning-mode value. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.818.0-canary.12 |
||
|
|
fe803bedf1 |
feat(server): add a configurable cooldown to the terminal workspace reaper (#11642)
## Thinking Path > - Paperclip is an open source app that manages AI agents for work. > - The server manages execution workspaces and their worktrees. > - The terminal workspace reaper removes a workspace when its issue tree reaches a terminal state. > - Immediate removal prevents a person from reopening recently completed work. > - This pull request adds a configurable cooldown before the reaper archives the workspace. > - The cooldown keeps recent work available and keeps immediate cleanup available with value `0`. ## Linked Issues or Issue Description Refs: #7790 **Problem** The reaper archives an execution workspace and deletes its worktree as soon as the issue tree becomes terminal. A person cannot reopen recent work without extra effort. **Expected behavior** The reaper should keep a recently completed workspace during a configurable cooldown window. It should archive older work and support immediate cleanup when the value is `0`. **Proposed solution** Read the cooldown from `PAPERCLIP_WORKSPACE_REAPER_COOLDOWN_DAYS`. Use a seven-day default. Use the latest terminal timestamp in the source issue tree as the cooldown anchor. ## What Changed - Add `PAPERCLIP_WORKSPACE_REAPER_COOLDOWN_DAYS` with a seven-day default. - Treat `0` as no cooldown and use the default for negative or non-numeric values. - Use the latest `completedAt` or `cancelledAt` value in the source issue tree. - Use `updatedAt` when a terminal timestamp is null. - Skip candidates inside the cooldown and report them in `skippedCooldown`. - Recheck the cutoff during the guarded archive operation. - Document the environment variable and add focused tests. ## Verification - Run `npx vitest run server/src/__tests__/execution-workspaces-service.test.ts`. - Confirm that the test run passes 66 tests. - Confirm that the tests cover a recent tree, an old tree, value `0`, and a null terminal timestamp. - Confirm that the changed files pass `tsc --noEmit`. ## Risks The default changes terminal workspace cleanup from immediate removal to a seven-day delay. A value of `0` preserves immediate cleanup. The guarded archive check limits race risk during concurrent lifecycle changes. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap. ## Model Used OpenAI Codex, GPT-5, tool use and code execution. This model assisted with the implementation review and PR preparation. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
393da0f67c |
fix(adapter-utils): graft unrelated imported histories instead of failing the run (#11638)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - At run finalize, the host imports the sandbox git history and reconciles it with the local worktree in `integrateImportedGitHead` > - Transported workspaces are depth-1 shallow clones, so the boundary commit reads as parentless inside the sandbox > - A `git commit --amend` there rewrites the boundary commit into a root commit, and the re-imported history no longer connects to the host history > - `git merge-tree` has no common base to merge against, so the sync throws "Failed to merge concurrent remote git histories" and the run fails with its work stranded in the sandbox > - This pull request grafts the imported tree onto the current head as a single commit instead of failing > - The benefit is that a history rewrite inside the sandbox can no longer lose a run's work ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "unrelated histories", "Failed to merge concurrent", and "shallow". Depends on #11637 (merged; the graft commit reuses its identity constant). This PR is now rebased onto `master`. **What happened?** An agent run amended a commit inside its sandbox workspace to address review feedback. The sandbox clone is depth-1 shallow, so git treated the boundary commit as parentless and the amend produced a root commit. At finalize, the host-side sync failed with `Failed to merge concurrent remote git histories for <sha>` and the run was marked failed. A follow-up run had to repair the branch by hand: fetch the true parent from origin and rebuild the commit with `git commit-tree`. **Expected behavior** The sync must never strand completed work. When the imported history shares no ancestor with the local one, the imported tree should still land on the current head, with the imported message preserved and the graft recorded. **Steps to reproduce** 1. Start a run whose workspace transport uses the shallow clone path (`withShallowGitWorkspaceClone`, depth 1). 2. Inside the sandbox workspace, run `git commit --amend` on the boundary commit. The result is a parentless root commit. 3. Finish the run. The host-side `integrateImportedGitHead` finds no merge base, `merge-tree` fails, and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `createUnrelatedHistoryGraftCommit` helper. It reads the imported head's tree and message, and creates one commit on top of the current head with the deterministic sync identity and a trailer that records the graft and both shas. - `integrateImportedGitHead` (both the remote-git-sync version and the SSH copy in `ssh.ts`): when `merge-base` reports no common ancestor, graft instead of throwing. The ref update keeps the same compare-and-swap and concurrent-retry semantics as the merge path. - The graft is gated on `git merge-base` exiting with status 1 — the no-ancestor signal. Operational failures (timeout, missing object, repository error) keep the loud merge failure instead of rewriting the tip. - New regression tests: one builds the exact shallow-amend shape (a root commit rebuilt from the base tree) and asserts the graft lands on the current head with the imported tree, subject, and graft trailer; one integrates a well-formed sha the repository does not hold and asserts the integration still throws with the branch tip unchanged. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 19/19 pass (includes the new graft test and the merge-base failure-discrimination test). - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks - Behavioral shift: unrelated imported histories previously failed the integration; now they land as a squash-graft. In this degenerate case there is no base to merge against, so the imported tree is taken wholesale and concurrent local-only tree changes are superseded at the tip. The local commits keep their place in the graft's ancestry, and the trailer records both shas, so nothing is unrecoverable. The old behavior lost the imported work instead, which is the worse failure for an autonomous run. - The graft reuses the imported head's commit message, so branch history still reads naturally after a sandbox rewrite. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.818.0-canary.11 |
||
|
|
9ea8143c87 |
fix(adapter-utils): give sync-created merge commits a deterministic git identity (#11637)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Runs execute in transported workspaces; at finalize, the host syncs the sandbox git history back into the local worktree > - When both sides advanced, `integrateImportedGitHead` reconciles them with `git merge-tree` plus `git commit-tree` on the host > - Execution hosts are often containers with no git config and no resolvable hostname, so `commit-tree` fails with "Author identity unknown" > - That one local command failure marks the whole run as failed, even though the run's work succeeded > - This pull request gives sync-created merge commits an explicit, deterministic identity at the call site > - The benefit is that workspace finalize no longer depends on ambient host git configuration ## Linked Issues or Issue Description No existing issue found. I searched issues and PRs for "Author identity unknown", "unrelated histories", and "commit-tree identity". **What happened?** A run finished its work, but workspace finalize failed. The host-side sync ran `git commit-tree <tree> -p <localHead> -p <importedHead> -m "Paperclip remote git sync merge <sha>"`. Git exited with `Author identity unknown ... fatal: unable to auto-detect email address (got 'node@<container-id>.(none)')`. The adapter recorded the whole run as failed, and the host worktree kept the stale head. Any container deployment without a global gitconfig reproduces this; I observed it on a Paperclip Cloud stack. **Expected behavior** Commits that the sync machinery itself creates must not depend on ambient host git configuration. The merge commit is machine-authored, so it should carry a deterministic Paperclip identity. **Steps to reproduce** 1. Run the Paperclip server in a container with no `user.name`/`user.email` git config and a hostname git cannot turn into an email. 2. Let a run's sandbox branch diverge from the host worktree, so both sides advance. 3. Workspace finalize calls `integrateImportedGitHead`. The `git commit-tree` step fails with "Author identity unknown" and the run fails. ## What Changed - `git-workspace-sync.ts`: new exported `GIT_SYNC_COMMIT_IDENTITY_ARGS` (`-c user.name=Paperclip -c user.email=noreply@paperclip.ing`), applied to the `commit-tree` call in `integrateImportedGitHead`. - `ssh.ts`: the SSH-sync copy of `integrateImportedGitHead` applies the same identity args to its `commit-tree` call. - New regression test: builds divergent histories in a repo with no configured identity and asserts the sync merge commit is created with the deterministic identity, correct parents, and merged tree. ## Verification - `pnpm vitest run packages/adapter-utils/src/git-workspace-sync.test.ts` — 18/18 pass. - `pnpm --filter @paperclipai/adapter-utils typecheck` — clean. - Negative proof: with the source fix stashed, the new test fails on the identity assertion. - Full `pnpm vitest run packages/adapter-utils`: every file passes except `local-process-sandbox.test.ts`, which fails identically on an untouched `master` checkout on macOS (bubblewrap-dependent, pre-existing, unrelated). ## Risks Low risk. The change only adds `-c` identity flags to two machine-generated commit invocations. `GIT_AUTHOR_*` / `GIT_COMMITTER_*` environment variables still take precedence over `-c` when an operator sets them, so existing deployments that configure an identity keep their behavior. ## Model Used - Claude Fable 5 (`claude-fable-5`), extended thinking, via Claude Code CLI (tool use for code exploration, test runs, and verification). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no doc surface describes this internal sync path) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.818.0-canary.10 |
||
|
|
0e9b03832d |
fix(server): skip host provisionCommand for sandbox-driver environments (#11626)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The run orchestrator prepares an execution environment before an agent starts > - A sandbox driver creates a remote folder before the adapter uploads repository content > - The orchestrator ran the host `provisionCommand` in that empty folder > - The command failed with exit 127 before the adapter could run its `stage.sync` step > - This pull request skips host provisioning for sandbox drivers and keeps the existing local and SSH behavior > - The benefit is that sandbox runs reach the adapter sync step without an empty-folder setup failure ## Linked Issues or Issue Description **What happened?** A sandbox environment ran the host `provisionCommand` before the adapter uploaded repository content. The command ran in an empty remote folder and failed with exit 127. **Expected behavior** The orchestrator should skip host provisioning for a sandbox driver. The adapter should upload the provisioned tree during its `stage.sync` step. **Steps to reproduce** 1. Configure an environment with the `sandbox` driver and a host `provisionCommand`. 2. Start a run that uses this environment. 3. Observe that the command runs in the empty sandbox folder and the run fails with `setup_failed`. **Paperclip version or commit** Reproduced on the current `master` commit before this change. **Deployment mode** Built from source with a sandbox environment. **Installation method** Built from source with pnpm. **Agent adapter(s) involved** Not adapter-specific (core bug). The sandbox adapter syncs the tree after environment setup. **Database mode** Not database-related. **Additional context** Related context: [#11091](https://github.com/paperclipai/paperclip/pull/11091) changes provision behavior for reused workspaces. This pull request covers the separate sandbox ordering failure. ## What Changed - Skip the orchestrator provision step when `environment.driver` is `sandbox`. - Keep the existing skip for `local` and the provision step for `ssh`. - Log one info message when a sandbox skip drops a present command. - Keep the existing `plugin` path because it has no `stage.sync` step and runs against the host filesystem. - Add tests for sandbox, local, SSH, plugin, logging, and provision failures. ## Verification - Run `./node_modules/.bin/vitest run server/src/__tests__/environment-run-orchestrator.test.ts`. - Confirm that the test run passes all 10 tests. - Confirm that CI checks pass on this pull request. ## Risks - Low risk. The change affects only the provision gate for sandbox drivers. - SSH and local behavior stays unchanged. - The plugin driver stays on its current path. - The new log line makes a sandbox skip visible to operators. ## Model Used OpenAI, GPT-5, exact runtime model `gpt-5`, with tool use and code review support. The implementation author used this model to inspect code, edit source and tests, and run the targeted test suite. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (no exact duplicate found; related PR #11091 reviewed) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes (no documentation change applies to this internal gate correction) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.818.0-canary.9 |
||
|
|
6b8e42168e |
Add governed secret alias confirmation cards (#11486)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents need scoped secret bindings to use external services safely. > - Agents could not request an existing secret under a new config name without an internal secret identifier. > - Existing binding proposals were only visible in Settings and did not create an issue-thread approval path. > - A confirmation card could record acceptance without proving that the binding was created. > - This pull request extends the existing secret proposal system with safe source references and governed issue-thread confirmation cards. > - The benefit is a one-click flow that creates the binding or shows a clear failure without exposing secret material. ## Linked Issues or Issue Description Related prerequisite: #11482. **Subsystem affected** Cross-cutting: server REST APIs, shared interaction contracts, database proposal schema, and issue-thread UI. **Problem or motivation** An agent can need an existing bound secret under a second config name. The agent cannot safely discover the internal secret identifier. The existing proposal is also easy for the operator to miss because it only appears in Settings. A generic confirmation can record acceptance without executing the binding. **Proposed solution** Let an agent create a binding proposal from one of its existing config paths. Mint a server-owned, human-only confirmation card on the checked-out issue. Recheck the operator's target-agent permission under the proposal row lock. Execute the existing proposal transaction after card acceptance. Store an `executed` or `failed` result on the card. Render the complete lifecycle in the issue thread and attention resolver. **Alternatives considered** A new alias subsystem would duplicate proposal quotas, expiry, authorization, and binding synchronization. A text-only issue comment would not provide a governed action or an execution result. An agent-supplied card payload would permit metadata smuggling. This change uses the existing proposal transaction and a server-owned payload instead. **Roadmap alignment** This change extends the completed "Secrets Manager with per-agent access" roadmap item. It preserves scoped bindings and audited resolution. The required GitHub search found no other open duplicate issue or pull request. ## What Changed - Added safe source-config-path binding proposals and preserved user-secret ownership checks. - Added a proposal-to-interaction link and an idempotent database migration. - Minted human-only `request_confirmation` cards with server-owned `secretProposal` metadata. - Rejected agent-supplied governed metadata and agent addressees. - Rechecked `agent_config:update` authority under the proposal lock before execution. - Recorded `executed` or `failed` results and posted a failure comment when no binding was created. - Settled failed accepted proposals atomically and mirrored rejection, withdrawal, and expiry in both directions. - Emitted `secret.binding.created` for new agent binding writes. - Added a dedicated issue-thread card for pending, executed, failed, rejected, withdrawn, and expired states. - Showed only the source label, target agent, config path, skeptical justification, expiry, and safe failure code. - Replaced resolved attention-query entries immediately with the stitched server result. - Added focused server, database, UI, and state-transition tests. - Added Storybook fixtures for every review state and documented the API and agent behavior. ## Verification - `pnpm exec vitest run ui/src/components/IssueThreadInteractionCard.test.tsx ui/src/components/AttentionInteractionResolver.test.ts` — 58 passed. - `pnpm --filter @paperclipai/ui typecheck` - `pnpm check:token-gates` - `pnpm build-storybook` - `pnpm --filter @paperclipai/shared typecheck` - `pnpm --filter @paperclipai/db typecheck` - `pnpm --filter @paperclipai/server typecheck` - `pnpm --filter @paperclipai/db check:migrations` - `NODE_ENV=test pnpm exec vitest run server/src/__tests__/issue-thread-interaction-routes.test.ts server/src/__tests__/secret-proposals-routes.test.ts server/src/__tests__/secrets-routes.test.ts server/src/__tests__/agents-service-secret-bindings.test.ts` — 142 passed. - `NODE_ENV=test pnpm --filter @paperclipai/db exec vitest run src/company-secret-proposals-migration.test.ts --silent` — 1 passed. - `pnpm -r typecheck` - `pnpm test:run` — server 4,175 passed, UI 4,109 passed; the CLI AWS-doctor case passes 8/8 with runtime-injected static AWS credential variables unset. - `pnpm build` - `git diff --check origin/master...HEAD` ## Risks - Migration `0221` adds one nullable foreign key and one index. It uses idempotent guards. - The accept route performs a governed write after it records card acceptance. A failed write is visible and settles the proposal as rejected. - Concurrent proposal and card resolution must use proposal-before-interaction lock order. A race test covers direct approval against card rejection. - The new audit event increases activity rows for newly added agent bindings. It does not include secret values or fingerprints. - The card includes only safe proposal metadata. It does not include secret value, fingerprint, version, or internal secret identifiers. - The UI uses the stitched resolution result. Focused tests cover immediate cache replacement and every terminal state. > This work extends an existing completed roadmap capability. The GitHub duplicate search returned no other open related work. ## Model Used - OpenAI Codex with model ID `gpt-5`. The runtime did not expose its context-window size. Reasoning, repository tools, code execution, database integration tests, UI rendering, and GitHub tools were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.818.0-canary.8 |
||
|
|
b446ff59bf |
refactor(acpx-engine): coordinator-owned ACP run lifecycle with a typed resource ledger (#11576)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent adapters run agent sessions through the ACPX engine > - The ACPX engine handled one run attempt as a long implicit procedure > - That shape made resource ownership, cleanup order, and failure behavior hard to verify > - This pull request gives the attempt a coordinator, a typed resource ledger, separate run sites, and explicit turn and settlement sequences > - The benefit is clear ownership, one cleanup path, safer session reuse, and testable failure behavior ## Linked Issues or Issue Description **What existing behavior does this improve?** The ACPX engine manages startup, turn execution, session reuse, and cleanup inside one large run procedure. **Current behavior** The run procedure owns several resources through implicit control flow. Cleanup and session reuse behavior depend on lane-specific branches and error paths. **Proposed behavior** The coordinator owns the run attempt. A typed ledger records six resources and their states. Host and sandbox run sites own lane-specific acquisition. Turn and settlement sequences expose typed outcomes. The engine emits allowlisted phase telemetry. **Reason and benefit** Explicit ownership makes cleanup and failure behavior easier to inspect. The fault matrix and characterization tests protect the external result while the refactor reduces hidden control flow. **Breaking changes** None to the public adapter contract. The host warm-save path now closes and relaunches the runtime because a transferred runtime could retain a run-scoped credential. A cold session-handshake failure now closes the created runtime. **Additional context** This pull request contains the ACPX engine lifecycle refactor, its tests, and the lifecycle document. ## What Changed - Add a run coordinator for startup, turn execution, settlement, and result reproduction. - Add a typed resource ledger with open, sealed, and consumed states. - Add host and sandbox run sites for lane-specific resource acquisition. - Replace separate runtime maps with a generic session reuse store. - Split session fingerprint identity from the outer session key. - Add typed turn and settlement sequences with one cleanup owner. - Add a closed allowlist for phase telemetry. - Add characterization tests and a 17-case fault matrix. - Add `doc/acp-run-lifecycle.md`. ## Verification - `npx vitest run packages/adapter-utils/src/acpx-engine/` passes 18 files and 286 tests at the submitted commit. - `pnpm --filter @paperclipai/adapter-utils typecheck` reports 0 errors at the submitted commit. - Run the full pull request checks after GitHub starts CI. - Run Greptile review after the pull request opens. ## Risks - The refactor changes internal control flow across the ACPX engine. - Host warm-save behavior now closes and relaunches the runtime. - Settlement changes the handling of a cold session-handshake failure from a leak to a close. - The characterization baselines and fault matrix reduce the risk of an external behavior change. > Paperclip is the open source app people use to manage AI agents for work > The adapter layer runs agent sessions through the ACPX engine > The engine needs explicit lifecycle ownership for reliable cleanup > This pull request adds coordinator-owned phases and a typed resource ledger > The result makes lifecycle behavior easier to test and review ## Model Used OpenAI GPT-5 Codex. Exact model ID: GPT-5. The model used tool execution, repository inspection, and code review support. The implementation author supplied the submitted code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.818.0-nightly.1 canary/v2026.818.0-canary.7 |
||
|
|
a1278e6ded |
fix(markdown): make rendered code blocks follow the active theme (#11591)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents report their work in task chat, and that chat renders markdown through `MarkdownBody` > - Agents post shell commands, API payloads and diffs, so a fenced code block is one of the most frequent things a person reads in the product > - The rendered code block was pinned to two literal colors, so it stayed dark in light mode and was the only dark surface on a light page > - The `prose` and the `prose-invert` variables held the same two values, so the theme could not change the block at all > - This pull request binds the block to the theme tokens that already exist and already carry a `.dark` override > - The benefit is that a code block now matches the page in both modes, and it follows any future change to the theme automatically ## Linked Issues or Issue Description No public issue exists. The problem is described below. **What happened?** A fenced code block in rendered markdown is always dark. It uses the background `#1e1e2e` and the text color `#cdd6f4` in light mode and in dark mode. In light mode the block is the only dark surface on the page. The rule is in `ui/src/index.css`. The variables `--tw-prose-pre-bg` and `--tw-prose-invert-pre-bg` are set to the same literal value, so `prose-invert` cannot change it. **Expected behavior** A code block uses a light surface with dark text in light mode. It uses a dark surface with light text in dark mode. It follows the theme like every other surface in the app. **Steps to reproduce** 1. Start the app and open a task that contains a fenced code block in its chat. 2. Set the theme to light. 3. Look at the code block. The block is dark. The page is light. 4. Set the theme to dark. The block does not change. **Paperclip version or commit** Reproduced on `master` at `4af55ba6b`. **Agent adapter(s) involved** Not adapter-specific (core bug). The defect is in the UI render path. **Deployment mode** Local development server. ## What Changed - Bind `.paperclip-markdown pre` to `--muted`, `--foreground`, `--border` and `--radius-lg` in `ui/src/index.css`. Remove the `#1e1e2e` and `#cdd6f4` literals. - Set the four `--tw-prose-*-pre-*` variables to the same tokens, so `prose` and `prose-invert` both follow the theme. - Apply the same four tokens to `.paperclip-mdxeditor-content pre`. - Change the fill of the copy button and the wrap button to `--background` in `ui/src/components/MarkdownBody.tsx`. The old fill was `color-mix(in oklab, var(--muted) 92%, var(--background) 8%)`. That value is almost equal to the block's new `--muted` surface, so the buttons would nearly disappear. - Update the stale comment above the rule. The comment said "Dark theme code blocks". - Add `ui/src/components/MarkdownCodeBlockStyles.test.ts`. It fails if a literal color returns to any themed code surface. - Export `codeBlockActionStyle` from `MarkdownBody.tsx` so the new test can read it. No new design token is added. Every token used here is already defined in `ui/src/index.css`, and each one already has a `.dark` override. ## Verification Automated: ```bash cd ui && npx vitest run --config vitest.config.ts src/components/MarkdownCodeBlockStyles.test.ts src/components/MarkdownBody.test.tsx src/components/MarkdownBody.wrap.test.tsx src/components/MarkdownAccentStyles.test.ts ``` 59 tests pass. `MarkdownCodeBlockStyles.test.ts` is new. It guards the defect directly. It reads `index.css` and asserts that each themed code surface rides a token and holds no hex, `rgb()` or `hsl()` literal. I confirmed the test fails on the original defect: restoring `#1e1e2e` and `#cdd6f4` fails 2 of the 5 tests. Manual. Open a task that has a fenced code block in its chat. Read the computed style of `.paperclip-markdown pre` in the browser. The measured values are: | Property | Light | Dark | | --- | --- | --- | | background | `oklch(0.97 0 0)` | `oklch(0.269 0 0)` | | color | `oklch(0.145 0 0)` | `oklch(0.985 0 0)` | | border | `oklch(0.922 0 0)` | `oklch(1 0 0 / 0.1)` | | radius | `8px` | `8px` | These values are `--muted`, `--foreground`, `--border` and `--radius-lg` for each mode. Toggle the theme and confirm the block changes with the page. ## Risks Low risk. The change is CSS and one inline style value. There is no migration and no schema change. Two behavioral notes for the reviewer: 1. The corner radius changes from `calc(var(--radius) - 3px)` (5px) to `var(--radius-lg)` (8px). This matches the approved design and removes an arbitrary offset. It is a small visual change on every code block. 2. The CodeMirror theme inside the MDXEditor still uses the Catppuccin literals (`ui/src/index.css`, the `.paperclip-mdxeditor .cm-editor` rules). That surface is an editor, not a rendered snippet, and a change there needs a full light CodeMirror theme. A code block therefore looks light when it is rendered and dark while a person edits it. This is intentional in this pull request. Tell me if you want it in scope. Syntax highlighting, line numbers and diff rows are not in this pull request. The repository has no highlighter dependency and no syntax or diff tokens. Those are separate changes. ## Model Used Claude Opus 5 (`claude-opus-5`), via Claude Code. Extended thinking was on. Tool use was on, with file edit, shell, browser automation and the Paper design tool. The design was produced first in Paper, then read back through the design tool for exact token values rather than from screenshots. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Notes on the unchecked boxes: - Branch name. The branch is `claude/paperclip-code-snippet-styling-202f7d`. It describes the change, but the `202f7d` suffix comes from the local worktree name. It carries no ticket id. I did not rename it, because the branch was already named when this work was requested. Tell me if you want it renamed before review. - CI and Greptile. Not yet run at the time of opening. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>canary/v2026.818.0-canary.6 |
||
|
|
1c366a9059 |
fix(server): reject invalid agent credentials instead of downgrading to the local user actor (#11589)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100): short sentences, one instruction per sentence, simple approved vocabulary, and the active voice. --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The server authenticates each agent request in `actorMiddleware` before it attributes chat comments > - When an agent bearer token failed verification, the middleware called `next()` with no error and the request continued without an agent actor > - The request then fell back to the local user actor, so the server stored agent replies as user comments > - The task chat UI renders user comments in blue bubbles, so agent messages appeared as blue user bubbles > - This pull request rejects invalid agent credentials with 401 instead of a silent downgrade > - The benefit is that agent messages keep agent attribution, and broken credentials fail loudly with a clear retry message ## Linked Issues or Issue Description **What happened?** A user cancelled an onboarding question card. The agent posted a follow-up reply. The reply appeared in a blue bubble, which the UI reserves for human messages. The agent run held an expired local agent JWT. The auth middleware could not verify the token, called `next()` without an actor, and the request fell back to the local user identity. The server stored the agent comment as a user comment. **Expected behavior** Agent messages always render as agent bubbles. A request with invalid agent credentials must fail with 401 so the adapter can refresh credentials and retry. It must not post content under a human identity. **Steps to reproduce** 1. Start a local Paperclip instance. 2. Give an agent run an expired or malformed agent JWT. 3. Let the agent post an issue comment through the API bridge. 4. Before this change: the comment is stored with the local user identity and renders as a blue bubble. After this change: the request fails with 401 and a message that tells the caller to obtain fresh credentials. ## What Changed - `server/src/middleware/auth.ts`: a bearer token that fails verification now produces a 401 `unauthorized` error instead of a silent fall-through to the anonymous/local-user actor. - The 401 message states the cause: expired token, unverifiable token, empty bearer token, missing agent record, agent record in another company, terminated agent, or agent pending approval. - The API-key path now also rejects an agent record whose company does not match the key. - `packages/adapter-utils/src/execution-target.ts`: the bridge proxy now writes a `comment id: <id>` marker to the run log for each posted issue comment, so misattributed comments can be traced to a run. - `ui/src/components/task-chat/task-chat-adapter.test.ts`: a regression test asserts that a recovered `local-board` comment with a derived agent author renders as an agent bubble, not a user bubble. - `server/src/__tests__/agent-auth-middleware.test.ts` and `packages/adapter-utils/src/execution-target-sandbox.test.ts`: new tests cover each rejection path and the log marker. ## Verification - Run `pnpm vitest run src/__tests__/agent-auth-middleware.test.ts` in `server/` — 14 tests pass. - Run `pnpm vitest run execution-target-sandbox` at the repo root — 44 tests pass. - Run `pnpm vitest run src/components/task-chat/task-chat-adapter.test.ts` in `ui/` — 4 tests pass. - Manual check: post an issue comment with an expired agent JWT; the API returns 401 with a retry message and no comment is stored. ## Risks - Behavioral shift: requests that previously continued as anonymous or local-user actors after a failed agent-token verification now receive 401. Any caller that relied on the silent downgrade must refresh its credentials. This is the intended fix, and the adapters already handle 401 with a credential refresh. - No schema or migration changes. Low risk otherwise. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, via Claude Code with extended thinking and tool use (agent harness with shell, file, and git tools). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [ ] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
664052f8ea |
feat(release): draft stable notes at beta publish, read them from master at promotion (#11567)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system promotes builds canary → nightly → beta → stable, and stable releases publish a GitHub Release from `releases/vYYYY.MDD.P.md` > - The stable lane requires that notes file to exist inside the promoted source commit, but the file is named for the promotion date, which is unknown when the source commit is created > - A promoted beta can therefore never pass the notes check: every happy-path stable is forced through the candidate-branch fix path, with a soak-gate justification, for a notes-only change > - This pull request drafts the notes automatically when the beta is published and lets the stable promotion read them from `master` > - The benefit is a walkable stable happy path: the soak gate stays exact, notes get a real review window during the soak, and the justification path returns to its real purpose (cherry-picked fixes) ## Linked Issues or Issue Description **What existing behavior does this improve?** The stable promotion path in the release channel system (`release.yml`, `scripts/release.sh`). **Current behavior** `release.sh stable` requires `releases/vYYYY.MDD.P.md` in the checked-out source tree, and `publish_stable` checks out the exact promoted SHA. The soak gate requires a `beta/v*` tag to point at that same SHA. No commit can satisfy both for a promoted beta, so a stable promotion must cut a candidate branch with a notes-only commit and bypass the soak gate with a written justification. Release notes are also written at promotion time, under time pressure, with no review window. **Proposed behavior** When a beta publishes, a `draft_stable_notes` job generates a grouped notes skeleton at `releases/beta/v<beta-version>.md` and pushes it to a machine-owned branch; a human opens the PR and edits it during the 3-day soak. The stable preflight resolves notes before the `npm-stable` approval gate: source-tree notes first (the candidate fix path, unchanged), then the merged beta-keyed file on `master`; it fails early with the missing path named when neither exists. After the stable ships, a canonicalization job pushes a branch that moves the file to `releases/vYYYY.MDD.P.md`. Related (not duplicates): #11006 and #11008 introduced the nightly and beta lanes this builds on; older changelog PRs (for example #10669) authored notes manually at promotion time, which is the flow this replaces. **Reason and benefit** The happy path becomes: promote the exact soaked SHA, no justification, notes reviewed during the soak instead of written at the gate. The `releases/vYYYY.MDD.P.md` invariant still holds durably via the canonicalization PR. ## What Changed - `scripts/release.sh`: new `--notes-file PATH` (stable only) overrides where the pre-publish notes check looks, so notes can live outside the source checkout without dirtying the worktree. - `scripts/create-github-release.sh`: same `--notes-file` override for the GitHub Release body. - `scripts/draft-stable-notes.sh` (new): deterministic skeleton generator — commit subjects from the newest stable tag (falling back to the previous beta, then full history) to the beta's source commit, grouped into Features / Fixes / Other. - `.github/workflows/release.yml`: - `draft_stable_notes` job after `publish_beta`: runs the generator and force-pushes `release-notes/v<beta-version>`; the job summary links the compare page. It recreates the beta tag locally if the tag push was rejected (the known workflows-permission case), so drafting is not blocked on manual tag recovery. - `preflight_stable`: computes the target stable version (`release.sh stable --print-version`) and resolves the notes source (`source_tree` → `master_beta` → fail early / warn on dry run); new outputs. - `publish_stable`: materializes `master`-side notes into `RUNNER_TEMP` and passes `--notes-file` to both scripts; outputs the published stable version. - `canonicalize_stable_notes` job: pushes the `git mv` branch after a stable that used `master`-side notes. - `doc/RELEASING.md`, `doc/RELEASE-CHECKLIST.md`: document the drafted-notes flow, the preflight resolution order, and the canonicalization step; the LLM changelog flow now targets the draft branch during the soak. - `.agents/skills/release-changelog/SKILL.md`, `.agents/skills/release-changelog-discord-message/SKILL.md`: the notes-authoring skills now describe this flow — range ends at the beta source commit (not `HEAD`), the file is beta-keyed on the `release-notes/v<beta-version>` branch (seeded with `scripts/draft-stable-notes.sh` for betas that predate the automation), and the canonicalization link caveat is called out for announcements. ## Verification - `node --test scripts/draft-stable-notes.test.mjs` — 6 tests, temp git-repo fixtures: grouping, stable-tag range, previous-beta and full-history fallbacks, default output path, malformed version, missing tag. - `node --test scripts/release-lib.test.mjs` — unchanged suite still green. - `bash -n` on both changed shell scripts; `release.yml` re-parsed as YAML. - `./scripts/release.sh stable --print-version` unchanged (prints the next stable version); `--notes-file` on a non-stable channel fails with a clear error. - Not exercised end-to-end: the new workflow jobs need a real beta publish to run. The first beta after merge is the live test; the draft job is additive and cannot affect the publish result (it runs after `publish_beta` completes). ## Risks - Low risk to publishing itself: `--notes-file` defaults preserve today's behavior everywhere; the draft and canonicalization jobs are additive and run after the publishes succeed. - The preflight now fails a real stable run when no notes are found. That is the intended fail-early behavior (it previously failed later, inside `publish_stable`, after the `npm-stable` approval). - `draft_stable_notes` force-pushes only the machine-owned `release-notes/v<beta-version>` branch; a beta re-cut regenerates it cleanly. - The stable version computed at preflight could differ from the published one if a run crosses UTC midnight between the two jobs; the materialized notes are passed by path, so the publish still succeeds, and the canonicalization job uses the actually-published version. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue templatecanary/v2026.818.0-canary.5 beta/v2026.818.0-beta.1 nightly/v2026.818.0-nightly.0 v2026.824.0 |
||
|
|
6691c57e54 |
docs(release): reconcile v2026.817.0 release notes to master (#11590)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Stable release notes live at `releases/vYYYY.MDD.P.md` on master as the durable record > - v2026.817.0 shipped from a candidate branch, so its notes file exists on that branch and in the GitHub Release, but not yet on master > - This pull request lands the exact shipped notes at the canonical path > - The benefit is a complete stable-notes record on master, and the FULL RELEASE NOTES link in announcements resolves ## Linked Issues or Issue Description **What existing behavior does this improve?** The `releases/` record on master. **Current behavior** `releases/` on master ends at `v2026.722.0.md`. The v2026.817.0 notes exist only on `candidate/release-2026.817.0` and in the published GitHub Release. **Proposed behavior** `releases/v2026.817.0.md` exists on master, byte-identical to what the release published. **Reason and benefit** Complete history at the canonical path; the standard notes link (`releases/v2026.817.0.md` on master) resolves for the announcement. ## What Changed - Add `releases/v2026.817.0.md`, taken verbatim from the candidate branch the release was published from (tag `v2026.817.0`, commit `213dabab4`). ## Verification - `git diff v2026.817.0 -- releases/v2026.817.0.md` against this branch is empty (byte-identical to the shipped file). - Docs-only; no code surface. ## Risks - None; docs-only. The candidate branch is deleted after this merges (the `v2026.817.0` tag keeps its commits reachable). ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue templatecanary/v2026.818.0-canary.4 |
||
|
|
49217aadf0 |
refactor: balance serialized server shards by recorded suite duration (#11528)
<!-- Write all pull request text in Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The PR verify workflow gates every pull request; its wall-clock time sets the feedback loop for all contributors > - In a recent successful PR run (actions run 32012408876), the slowest check was "Verify serialized server suites (1/5)" at 337s, while its four sibling shards finished in 212-238s > - The serialized lane assigns suites to shards round-robin over an alphabetical list, so the heavy heartbeat and issues suites cluster on one runner > - The general-server lane already solves this with a duration-aware LPT partition backed by a recorded manifest > - This pull request reuses that partitioner for the serialized lane with a fresh per-suite duration manifest > - The benefit is a balanced serialized matrix: the measured 968s suite total levels to about 194s per shard, which removes about 80-100s from the run's slowest check ## Linked Issues or Issue Description **What existing behavior does this improve?** The `Verify serialized server suites` shard matrix in `.github/workflows/pr.yml` distributes route/authz test suites across five runners. **Subsystem affected** CI / test infrastructure (`scripts/run-vitest-stable.mjs`). **Current behavior** `selectSerializedSuites` assigns suites round-robin (`index % shardCount`) over the alphabetically sorted file list. The heavy suites cluster on shard 1/5. In actions run 32012408876, shard 1/5 spent 291s in its test step while the other shards spent 170-201s, which made that job (337s total) the slowest check of the whole PR run. **Proposed behavior** Partition the serialized suites with the same duration-aware LPT algorithm the general-server lane already uses (`scripts/general-server-shard.mjs`), backed by a new per-suite duration manifest. All five shards then carry about 194s of measured test time. **Reason and benefit** The slowest check bounds PR feedback time. Balancing the serialized matrix removes about 80-100s from that bound without adding runners. **Breaking changes** None. The partition remains deterministic, complete, and non-overlapping; suites missing from the manifest get the median weight. ## What Changed - Added `scripts/serialized-shard-durations.json`: per-suite wall-clock durations (ms) for all 134 serialized suites, sampled from actions run 32012408876 by diffing consecutive per-suite label timestamps in the shard logs (captures vitest spawn overhead, not just reported test time) - `scripts/run-vitest-stable.mjs`: `selectSerializedSuites` now uses the existing LPT partitioner (`selectGeneralServerShard`) with the new manifest instead of round-robin - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: added a manifest-freshness test and a shard-balance test for the serialized lane, mirroring the general-server ones - `.github/workflows/pr.yml`: updated the serialized matrix comment with the new measurement and mechanism ## Verification - `node --test ./scripts/__tests__/run-vitest-stable-shard.test.mjs` passes (13 tests), including the existing test that the serialized shards form a complete, non-overlapping partition - Dry-run of all five shards shows estimated totals of 194/194/194/194/193s (round-robin was 276/175/160/172/187s): `node scripts/run-vitest-stable.mjs --mode serialized --shard-index N --shard-count 5 --dry-run` - The `Verify serialized server suites` jobs on this PR run the real partition end to end ## Risks - Low risk. Selection logic only; the vitest invocation per suite is unchanged - A stale manifest degrades gracefully: unknown suites get the median weight, and a dedicated test fails if fewer than half the current suites have recorded durations ## Model Used - Claude (Anthropic), model ID `claude-fable-5`, agentic coding session with tool use (Claude Code / Claude Agent SDK); no extended-thinking mode ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Related prior work: #10923 (split serialized tests into five shards), #10925 (general-server duration manifest), #11156 (workspaces-a native shards). Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8087661bb8 |
fix: bound workspace Git scans (#11572)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Workspaces let users and agents inspect files that belong to an issue > - Changed-file views use full-tree Git status scans > - Many issue views could start those scans at the same time and make the server unresponsive > - Route-level limits did not protect the process or coalesce work for one repository > - This pull request adds one bounded scheduler for every expensive workspace Git scan > - It also starts browser scans only when the file panel is open and visible > - The benefit is bounded child-process use and responsive health checks during request storms ## Linked Issues or Issue Description **What happened?** Many changed-file requests could start full `git status --porcelain=v1 -z --untracked-files=all` scans at the same time. One production incident produced about 270 direct Git child processes. The Node process stayed alive but stopped answering health requests in time. **Expected behavior** Paperclip must bound expensive Git work across all companies, actors, issues, repositories, and browser tabs. Duplicate requests for one worktree must share work. Excess requests must fail fast with a retryable response. Hidden or closed file panels must not start scans. **Steps to reproduce** 1. Open changed-file views for many issue and actor keys. 2. Send requests for two large workspace roots at the same time. 3. Observe that route-level limiter keys allow many full Git scans to run together. 4. Observe delayed health responses and accumulated Git children. **Paperclip version or commit** Reproduced on master before commit `43ab441f0f`. **Deployment mode** Self-hosted server with local workspace repositories. ## What Changed - Add a process-wide scheduler with configurable concurrency, queue capacity, timeout, and cache TTL. - Add fair admission, a bounded queue, canonical worktree keys, single-flight joins, and bounded result caching. - Add subprocess timeouts, TERM-to-KILL escalation, bounded output, waiter cancellation, and slot cleanup. - Route full-tree status work from file resources, workspace runtime, execution workspaces, and adapter overlay sync through the scheduler. - Return stable retryable `503` and `504` error codes for saturation and timeout. - Add structured logs with safe workspace hashes, durations, queue state, cache use, joins, and terminal outcomes. - Gate UI queries on panel and document visibility. Cancel queries on close, hide, unmount, and workspace change. - Disable focus and reconnect bursts. Keep one explicit refresh action and a retryable unavailable state. - Document the 10-second default freshness tradeoff and all configuration variables. - Add unit, route, UI, adapter, and deterministic 500-request load coverage. ## Verification - `pnpm -r typecheck` - `pnpm build` - `pnpm check:token-gates` - `pnpm --filter @paperclipai/server exec vitest run src/services/workspace-git-operation-scheduler.test.ts src/__tests__/file-resources-git-scan-load.test.ts --reporter=dot` — 16 tests passed. - `pnpm --filter @paperclipai/ui exec vitest run src/components/WorkspaceFileBrowser.test.tsx src/lib/page-visibility.test.ts --reporter=dot` — 38 tests passed. - `pnpm --filter @paperclipai/adapter-utils exec vitest run src/git-workspace-sync.test.ts --reporter=dot` — 16 tests passed. - Existing file-resource, workspace-runtime, and execution-workspace regression selections passed. - Two cleanup safety regressions prove failed scans preserve the worktree before archive and at the final deletion fence. - Before: the incident produced about 270 Git children and health requests timed out. - After: 500 concurrent requests across 500 issue keys, 73 actors, and two roots started two underlying scans. Peak scan concurrency was 2. All 500 requests succeeded. Health p99 was 4.94 ms. The harness found zero unreaped children. - The full local Vitest run passed 4,267 tests. Ten existing fixed-port HTTPS exposure tests could not run because this host already owns Tailnet listeners on ports 42000 and 52000. Clean GitHub CI is the final full-suite result. - Latest-head GitHub CI passed all required test, typecheck, build, canary, e2e, policy, and security gates. - Greptile completed at 5/5 with zero unresolved comments, recommendations, or follow-ups. ## Risks - Changed-file results can be up to 10 seconds old by default. Explicit refresh remains available. - A full queue returns a retryable `503` instead of waiting without a bound. - A scan that exceeds the default 8-second deadline returns a retryable `504` and terminates its process group. - Operators can tune all limits with documented environment variables. Safe defaults protect local and shared servers. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5 family. The runtime does not expose the exact deployment ID or context-window size. High reasoning, tool use, and code execution were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.818.0-canary.3 |
||
|
|
48f4ae16ac |
fix(codex-local): keep a promoted device-login credential when re-seeding the managed home (#11578)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The codex_local adapter supports a device login that runs in a trusted sandbox and promotes the credential into the per-company managed Codex home > - The managed-home re-seeding step treats every regular-file `auth.json` as apikey-mode residue and removes it so the shared-home symlink can be restored > - The promotion writes the company credential as a regular file, so the first environment Test or run after a successful login deletes it > - On a server with no shared Codex login — any containerized deployment — nothing replaces the file, and the UI reports that the sandbox has no ready authentication right after it reported a successful login > - This pull request makes the cleanup identity-anchored: a subscription credential whose identity the shared source does not hold survives re-seeding > - The benefit is that a device login stays usable after Test and runs, on hosts with and without a shared Codex login ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the bug template. Related public PRs: [#11237](https://github.com/paperclipai/paperclip/pull/11237) added the sandbox device login and the credential promotion, [#11097](https://github.com/paperclipai/paperclip/pull/11097) added its building blocks, and the `ensureSymlink` heal for stale copies came from the fix for #5028. [#9621](https://github.com/paperclipai/paperclip/pull/9621) touches the adjacent sandbox auth sync-back lane but not this defect. **Subsystem affected** packages/adapters/codex-local — managed `CODEX_HOME` seeding (`codex-home.ts`). **Current behavior** A successful device login promotes the subscription `auth.json` into the company Codex home as a regular file, and the UI reports the login as authenticated. The next `seedManagedCodexHome` call — the environment Test probe and every execute both run it — removes any regular-file `auth.json` when no API key is configured, because the cleanup assumes such a file is apikey-mode residue left by a previous run. It then symlinks `auth.json` from the shared source home. On a server whose shared home has no Codex login (a container image, for example), there is no source to symlink, so the home ends with no credential at all. The Test probe then reports "The sandbox has no ready authentication for this adapter" immediately after a successful login, and a fresh login repeats the same cycle. On a server whose shared home does hold a login, the symlink silently replaces the promoted account with the host account. **Expected behavior** The credential a device login promoted stays in the company home across Test probes and runs. The #5028 heal (a stale regular-file copy of the shared credential becomes a symlink to the live source) and the apikey-residue cleanup keep working. **Steps to reproduce** 1. Run the server in an environment whose shared Codex home (`$CODEX_HOME` or `~/.codex`) has no `auth.json`. 2. Complete a Codex device login for a company; the promotion writes the company home `auth.json` and the UI reports authenticated. 3. Click Test on a codex_local agent (or start a run). The probe reports no ready authentication, and the promoted `auth.json` is gone from the company home. **Proposed solution** Make the cleanup identity-anchored, the same rule the promotion and the cache vend already use. A regular-file `auth.json` survives re-seeding when it holds a usable subscription identity that the shared source does not also hold, and the shared symlink does not replace it. A same-identity regular file is still the #5028 stale copy and is still healed into the symlink, because the symlink serves the same account with live, rotating tokens. An apikey-mode or unreadable file is still removed. ## What Changed - `seedManagedCodexHome` reads the target `auth.json` before the cleanup and keeps it when `readSubscriptionAccountId` yields an identity the shared source `auth.json` does not hold. The kept file is excluded from the shared symlink pass, and the function logs a fixed line when it keeps the file. - The function doc comment states the kept-promoted-credential rule. - Four new `seedManagedCodexHome` test cases: a promoted credential with no shared auth, a promoted credential with a different shared identity, the same-identity #5028 heal, and apikey-mode residue removal. ## Verification ```sh cd packages/adapters/codex-local npx tsc --noEmit # clean npx vitest run # 28 files, 321 passed, 1 skipped ``` The four new cases fail on the previous code: the first two observed the promoted file deleted (and, with a shared login present, replaced by the shared symlink). ## Risks Low risk. The change narrows one deletion path. Deployments that never use the device login see no difference: without a promoted subscription file, the cleanup and the symlink behave exactly as before, and the #5028 heal is pinned by an existing test plus a new same-identity test. The one deliberate behavioral shift: after a device login, the promoted company credential now stays authoritative over the shared host login for that company — which is the promotion's documented contract ("the company credential slot"). ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
adfb357223 |
docs(skills): maintain the release-credits exclusion list (#11588)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Stable release notes end with a Contributors section that credits
community contributors
> - The release-changelog skill excluded founders from that list, but
had no rule for non-founder core contributors
> - A release-notes draft therefore credited two core contributors as
community contributors
> - This pull request makes the exclusion list canonical in the
changelog skill and adds the core contributors
> - The benefit is that release credits consistently mean "community",
with one list to maintain
## Linked Issues or Issue Description
**What existing behavior does this improve?**
The Contributors rules in `.agents/skills/release-changelog/SKILL.md`
and the Community-section rules in
`.agents/skills/release-changelog-discord-message/SKILL.md`.
**Current behavior**
The changelog skill says to exclude "Paperclip founders (e.g.
`cryppadotta`, `forgottendev`, `devinfoley`, `sockmonster`,
`scotttong`)". Core contributors who are not founders are not covered,
so drafts credit them in the community list. The Discord skill refers to
"founders" in two places.
**Proposed behavior**
The changelog skill holds the canonical exclusion list of specific folks
— now including `nguyenm7`, `nickyleach`, and `tonio-alucema` — and both
Discord-skill references defer to that one list.
**Reason and benefit**
Release credits consistently mean community contributions, and there is
exactly one place to update when the core team changes.
## What Changed
- `.agents/skills/release-changelog/SKILL.md`: the exclusion rule names
specific folks, is the canonical list, and adds `nguyenm7`,
`nickyleach`, and `tonio-alucema`.
- `.agents/skills/release-changelog-discord-message/SKILL.md`: both
references ("Community" template note and the final checklist) defer to
the changelog skill's canonical list.
## Verification
- Docs-only change; rendered and proofread. No workflow, script, or test
surface is touched.
## Risks
- None beyond docs accuracy. If the canonical-list wording lands after
#11567 (which touches other sections of the same files), Git merges them
cleanly — the edited regions do not overlap.
## Model Used
Claude Fable 5 (Claude Code)
## Pre-submission checklist
- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
canary/v2026.818.0-canary.2
|
||
|
|
962e98b1be |
feat(server): operator declaration for platform edge TLS termination on the Claude login guard (#11579)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports a setup-token subscription login, and its confidential routes pass a fail-closed transport guard > - The guard accepts direct socket TLS, a local_trusted loopback peer, or an allowlisted proxy peer that forwards https — and deliberately never reads the global `TRUST_PROXY` > - On a managed platform the edge terminates TLS, the app socket is always plain HTTP, and the edge-proxy peer addresses are not stable or documented, so none of the three cases can hold > - Every login on such a deployment shows the clear-text transport warning although the user's connection is HTTPS, and the agent-scoped confidential routes fail closed entirely > - This pull request adds a dedicated operator declaration that the platform edge terminates TLS, as a fourth guard case > - The benefit is a correct transport decision on managed platforms with the default posture unchanged everywhere else ## Linked Issues or Issue Description No public GitHub issue covers this. The problem is described in-PR following the enhancement template. Related public PRs: [#11347](https://github.com/paperclipai/paperclip/pull/11347) added the new-agent login flow and the non-blocking transport advisory, and [#11286](https://github.com/paperclipai/paperclip/pull/11286) added the setup-token login and the guard with its `CLAUDE_LOGIN_TRUSTED_PROXIES` allowlist. **Subsystem affected** server/ — the confidential transport guard for the Claude setup-token login (`services/setup-token-session.ts`, `routes/agents.ts`, `app.ts`). **Current behavior** The guard allows a confidential response on direct socket TLS, on a `local_trusted` loopback peer, or when the immediate peer is on the dedicated `CLAUDE_LOGIN_TRUSTED_PROXIES` allowlist and forwards `https`. Behind a managed platform's TLS-terminating edge (Railway, Render, Fly, and similar), the app socket is plain HTTP and the edge-proxy peer addresses are not operator-visible or stable, so the allowlist cannot express them — IPv6 entries match by exact string only. The result: the login panel shows "This connection is not encrypted" for a connection that is HTTPS to the user, and the agent-scoped confidential routes return the fixed no-secret error. **Proposed behavior** `CLAUDE_LOGIN_EDGE_TLS_TERMINATED=true` is an explicit, single-purpose operator declaration that every client request reaches the server through the platform's TLS-terminating edge. Under the declaration the guard treats a request as confidential unless the edge itself labels the client hop as plain `http` in `X-Forwarded-Proto`. The declaration is never derived from the global `TRUST_PROXY` setting, which the guard still never reads. Without the declaration, nothing changes. **Reason and benefit** The guard's spoofing concern does not apply to this deployment shape: a client cannot pick its transport, because the platform admits HTTPS only, and the header the guard consults is set by the platform edge, not the client. A blanket warning that is always wrong teaches users to ignore it. The declaration keeps the strict default for every deployment that does not opt in, and it keeps the allowlist as the precise tool for operators who do know their proxy addresses. ## What Changed - `ConfidentialTransportConfig` gains optional `edgeTlsTerminated` (default false), documented as the operator declaration for platform edge TLS termination. - `evaluateConfidentialTransport` adds the declaration as a guard case: allowed unless the forwarded protocol's first hop is explicitly `http` (reason `edge_labeled_plain_http` then; `operator_edge_tls_termination` when allowed). - `assessConfidentialStartup` reports `edge_tls_termination_declared`, so the startup log shows why forwarded requests pass. - `app.ts` parses `CLAUDE_LOGIN_EDGE_TLS_TERMINATED` (truthy: `1/true/yes/on`) and passes it to the agent routes; the routes build the guard config from it. - The SR-7 operator-requirement comment on the setup-token routes documents the new variable next to the allowlist. - Tests: five new guard unit cases and a route case asserting the prompt and code responses carry no `transportAdvisory` under the declaration. ## Verification ```sh cd server npx tsc --noEmit # clean npx vitest run src/services/setup-token-session.test.ts \ src/routes/setup-token-route.test.ts \ src/__tests__/openapi-routes.test.ts # 3 files, 89 passed ``` The new "keeps failing closed when the declaration is absent" case pins the unchanged default posture. ## Risks The declaration is an operator statement the server cannot verify; an operator who sets it on a deployment whose edge does not terminate TLS re-labels plain-HTTP requests as confidential. This is the same trust class as `CLAUDE_LOGIN_TRUSTED_PROXIES` (a wrong allowlist entry has the same effect) and is opt-in, off by default, and scoped to the login routes only. The guard still fails closed when the edge explicitly labels a request `http`. No schema change, no API shape change — `transportAdvisory` was already nullable. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking, with tool use and code execution — investigation, implementation, and tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.818.0-canary.1 |
||
|
|
2ec984502a |
fix(release): stop smoke_beta silently skipping on promote-mode betas (#11582)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system re-smokes every published beta as post-publish verification (`smoke_beta`) > - The candidate-branch beta lane (#11209) added `verify_beta_candidate` to `publish_beta`'s needs; that job is skipped on every normal promote-mode beta > - `smoke_beta`'s condition has no status-check function, so GitHub attaches an implicit `success()` that evaluates the needs chain transitively — a skipped ancestor makes it false > - This pull request makes the condition explicit so promote-mode betas smoke again, and pins the shape in the workflow wiring test > - The benefit is that the post-publish beta gate actually runs instead of silently skipping ## Linked Issues or Issue Description **What happened?** Beta `2026.818.0-beta.0` (run 32082007439) published successfully, but its post-publish `smoke_beta` job was skipped. No configuration or input asked for that: the run was a plain `channel: beta` dispatch with `dry_run` at its default `false`, and the same expression `!inputs.dry_run` evaluated true inside `publish_beta`'s own steps (the Docker dispatch step ran). **Expected behavior** Every non-dry-run beta publish is followed by the release smoke suite against the exact published version, as documented in `doc/RELEASING.md` and `doc/RELEASE-CHECKLIST.md`. **Steps to reproduce** Dispatch `release.yml` with `channel: beta` promoting a nightly (promote mode). `verify_beta_candidate` is skipped by design; `publish_beta` runs through its explicit `!cancelled()` condition; `smoke_beta` then skips because its implicit `success()` sees the skipped ancestor in the transitive needs chain (actions/runner#2205 semantics). The beta published on 2026-08-11 predated #11209, so this never surfaced before. **Paperclip version or commit** master at `43ab441f0` (workflow file, current head). Related (not duplicates): #11209 introduced the candidate lane whose skipped job triggers this; #11208 covers the adjacent tag-push failure playbooks. ## What Changed - `smoke_beta`'s condition becomes `!cancelled() && needs.publish_beta.result == 'success' && !inputs.dry_run` — an explicit status-check function suppresses the implicit `success()`, and the result check keeps the dependency on a successful publish. - A comment above the job records why the explicit form is load-bearing. - `scripts/__tests__/release-verify-workflow.test.mjs` pins the new shape so the implicit form cannot silently return. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs` — 8 pass, including the new assertion. - `release.yml` re-parsed as YAML. - The exact skip is visible on run 32082007439 (`smoke_beta: skipped` after `publish_beta: success`); the coverage gap for that beta was closed manually by dispatching `release-smoke.yml` with `paperclip_version: beta` (run 32084880767). - Not exercised end-to-end: the corrected condition needs the next real promote-mode beta to demonstrate; the expression change is minimal and the semantics are the documented actions/runner behavior. ## Risks - Low risk: condition-only change on one job plus a test. Dry runs still skip the smoke (`!inputs.dry_run` retained). Candidate-mode betas, where `verify_beta_candidate` actually runs, behave as before. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue templatecanary/v2026.818.0-canary.0 |
||
|
|
4af55ba6bd |
test(release-smoke): follow the onboarding wizard's chat-first rewrite (#11565)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release channel system publishes a nightly build only after the release smoke suite passes against the newest canary > - The scheduled nightly run has failed every night since August 12, so no nightly, and therefore no beta candidate, has shipped for six days > - The failures are stale test locators, not a product regression: the chat-first onboarding rewrite (#11101) changed wizard copy and the post-launch destination > - This pull request updates the smoke spec to match the current wizard > - The benefit is a green nightly lane and an unblocked beta promotion ## Linked Issues or Issue Description **What happened?** The scheduled `Release` nightly run fails in `smoke_nightly / smoke` every night since 2026-08-12. The failing spec is `tests/release-smoke/docker-auth-onboarding.spec.ts`. Four assertions no longer match the product after the chat-first onboarding rewrite (#11101): - The step-1 heading is now "Name your organization", not "Name your company". - The step-4 hire button is now "Connect", not "Give it a heartbeat". - The seeded first task is now titled "Paperclip onboarding". - A successful launch navigates to the seeded task's thread (`/issues/<ref>`), not `/dashboard`. **Expected behavior** The smoke suite passes against a canary that contains the current onboarding wizard, and the nightly lane publishes again. **Steps to reproduce** Run `.github/workflows/release-smoke.yml` against `paperclipai@canary` (any version at or after the rewrite), or dispatch `release.yml` with `channel: nightly`. Example red runs: 32014452506 (Aug 17), 31938284401 (Aug 16). **Paperclip version or commit** `2026.817.0-canary.12` Related (not duplicates): #11190 updated this same spec for the mission-first wizard; this PR is the follow-up for the chat-first rewrite that landed after it. ## What Changed - Update the step-1 wizard heading locator to "Name your organization". - Update the step-4 hire button locator to "Connect" and reword the step comment. - Update `FIRST_TASK_TITLE` to "Paperclip onboarding" (the wizard's current `DEFAULT_TASK_TITLE`). - Assert the post-launch URL is the seeded task's thread (`/issues/`), not `/dashboard`. ## Verification - Local run of the exact CI harness: `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.817.0-canary.12`, then `pnpm run test:release-smoke` against the container — 1 passed. - The suite's later API assertions (company, CEO agent, mission goal, seeded issue assignment, assignment-sourced heartbeat run) all pass unchanged against the current canary. ## Risks - Low risk: test-only change; no product code is touched. - The spec remains copy-coupled to the wizard. If wizard copy churn continues, a follow-up could add stable `data-testid` hooks to the wizard so the smoke spec stops breaking on wording changes. ## Model Used Claude Fable 5 (Claude Code) ## Pre-submission checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue templatecanary/v2026.817.0-canary.15 |
||
|
|
43ab441f0f |
Exempt plugin-managed issues from successful-run-handoff recovery (#9047)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Plugins can extend Paperclip with their own multi-step workflow/graph engines that own an issue's lifecycle across many agent handoffs > - When such a plugin-managed issue legitimately stays `in_progress` for a while (e.g. an anchor issue parked at a fan-out step, waiting on child issues it spawned), Paperclip's generic recovery mechanisms have no way to know that's intentional > - The first commit on this branch fixed one such mechanism (`decideSuccessfulRunHandoff`) to skip plugin-owned issues, and was deployed to a real instance to verify the fix > - Watching that same instance afterward, the identical symptom (repeated "give a disposition" nags, the agent repeating "completed", the plugin's own enforcement correctly reverting the status) recurred on the same class of issue — meaning a second, independent code path had the exact same gap > - Traced it to `reconcileStrandedAssignedIssues` in `service.ts`: it detects a stale successful-run-handoff corrective run via `isExhaustedSuccessfulRunHandoff` and, once "exhausted" (default max attempts is 1, so effectively immediate), escalates via `escalateStrandedAssignedIssue` — with no check on who owns the issue's lifecycle at all > - Rather than duplicate the `originKind` check inline a second time (which is exactly how it got missed the first time), extracted it into a shared, exported, unit-tested helper (`isPluginManagedIssueLifecycle`) that both recovery paths now call > - The benefit is the same as the first commit, but closing the second loop this pull request's earlier version left open: generic plugin-owned issues (any workflow/graph-engine plugin, not just one specific plugin) stop burning real agent-run cost in a loop that can never actually resolve, across both recovery mechanisms that can trigger it ## Linked Issues or Issue Description No existing public issue covers this — describing it directly, following the bug report fields: **What happened?** An issue owned by a workflow-engine-style plugin (`originKind` starting with `"plugin:"`) was correctly held at `in_progress` by the plugin while it waited on spawned child issues to finish. The assigned agent's heartbeat succeeded and posted a well-formed completion comment, but `issue.status` stayed `in_progress` (the plugin's own enforcement reverted it, correctly, since the underlying work wasn't done). - **Path 1 (fixed in the first commit):** `decideSuccessfulRunHandoff()` saw `status === "in_progress"` after a successful run and enqueued a "missing disposition" corrective wake. The agent responded again, the plugin reverted the status again, and the recovery re-triggered again. - **Path 2 (fixed in the second commit, found after deploying and verifying the first fix on a live instance):** separately, `reconcileStrandedAssignedIssues` periodically re-scans `in_progress` issues, sees the corrective run from Path 1 (or any prior successful-run-handoff wake) as "exhausted" evidence, and escalates the issue via `escalateStrandedAssignedIssue` regardless of plugin ownership — producing the same nag-revert-nag cycle through a completely different call path that the first commit's fix did not touch. **Expected behavior** Neither recovery mechanism should nag an agent for a disposition, or escalate for one, on an issue whose lifecycle is already owned and managed by a plugin — that plugin's own enforcement/recovery path is the correct owner of "what happens next," not these generic core mechanisms. **Steps to reproduce** 1. Install a plugin that creates/owns issues via the plugin host bridge (`ctx.issues.create`/`ctx.issues.update`) with an `originKind` of `"plugin:<pluginKey>"`. 2. Have the plugin's own graph/workflow logic hold an issue at `in_progress` while some multi-step process it owns is still pending (e.g. spawned child issues not yet complete). 3. Let an agent run a successful heartbeat on that issue that produces visible progress (a comment) but does not change `issue.status` away from `in_progress` in a way that sticks (the plugin's own logic reverts any change back to `in_progress` on the next event). 4. Observe `decideSuccessfulRunHandoff` enqueue a corrective handoff wake (Path 1), and/or `reconcileStrandedAssignedIssues` treat that wake's run as exhausted and escalate (Path 2). Either one repeats indefinitely on its own. **Paperclip version or commit** `eb2cb916be3271e3e7ab5f643ad3ca3eb7c34d01` (current `master` at time of the first commit; rebased onto `ad961227f` for the second) **Deployment mode** Self-hosted server ## What Changed - **First commit:** Added `originKind: issues.originKind` to the issue query in `heartbeat.ts` that feeds `decideSuccessfulRunHandoff`, and a skip condition there for `originKind` starting with `"plugin:"`. - **Second commit:** - Extracted the plugin-ownership check out of `decideSuccessfulRunHandoff` into a new exported helper, `isPluginManagedIssueLifecycle(issue)`, in `successful-run-handoff.ts`. - Added the same check to `reconcileStrandedAssignedIssues` in `service.ts`, immediately before it would otherwise escalate an issue based on `isExhaustedSuccessfulRunHandoff` evidence — skipping plugin-managed issues there too. - Added unit tests for the new helper directly (plugin-prefixed origin kinds → `true`; non-plugin/missing origin kinds → `false`), alongside the existing `decideSuccessfulRunHandoff` tests (updated to import and rely on the shared helper, behavior unchanged). ## Verification - `npx vitest run server/src/services/recovery/successful-run-handoff.test.ts` — 20/20 passed (18 pre-existing/from the first commit + 2 new for the extracted helper). - `npx vitest run server/src/__tests__/heartbeat-comment-wake-batching.test.ts server/src/__tests__/heartbeat-process-recovery.test.ts server/src/services/recovery/successful-run-handoff.test.ts` — full suite still passes with the refactor. - `pnpm --filter @paperclipai/server exec tsc --noEmit` — no new errors introduced by either commit (confirmed against a pre-existing baseline of unrelated `@paperclipai/plugin-sdk` module-resolution errors from the workspace, present identically with the changes stashed out). - Manually verified Path 1 against a real plugin-managed issue stuck in that loop: after deploying the first commit and restarting the server, the same agent posted the same completion comment again, and the corrective-handoff recovery did not re-trigger. - Path 2 was found live on the same instance after that first deploy (the loop recurred through the second, independent mechanism) — root-caused via direct inspection of `heartbeat_runs`/`agent_wakeup_requests`/issue comment history, then fixed in the second commit. Not yet re-verified live on the instance (pending redeploy of this updated branch). ## Risks - Low risk. Both changes are additive skip conditions — they only cause a recovery decision to return early for a specific, narrow case (`originKind` starting with `"plugin:"`) that previously fell through to escalation/enqueue. No existing skip conditions are changed or reordered in a way that affects non-plugin issues. - Behavioral shift: plugin-managed issues that are genuinely stuck (not just correctly mid-flight) will no longer get either of these corrective nags. This is intentional — the plugin owning the issue is expected to have its own recovery path — but it does mean these mechanisms are no longer a safety net for buggy plugins that leave issues stranded. Plugin authors should ensure their own enforcement handles stranded states. - The refactor (extracting `isPluginManagedIssueLifecycle`) is a pure code-motion change for the first commit's check — no behavior change there, only a new call site added in `service.ts`. - No migration required (query-shape and control-flow changes only, no schema change). ## Model Used Claude (Anthropic), model `claude-sonnet-5`, used within a coding-agent harness (Claude Code) with tool use (file edit, test execution, git operations, live production-instance debugging via SSH/SQL) and extended reasoning across two sessions: the first implemented and deployed the initial fix, the second discovered the second recovery path was still looping on a live instance, root-caused it, and implemented/tested this follow-up commit. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>beta/v2026.818.0-beta.0 nightly/v2026.817.0-nightly.0 canary/v2026.817.0-canary.14 |
||
|
|
8774909361 |
fix(heartbeat): reuse sessions across execution handoffs (#9917)
## Thinking Path > - Paperclip uses durable task sessions so local adapters can resume work across sequential heartbeat runs. > - `execution_review_requested` and `execution_changes_requested` are issue-local execution-policy handoffs, not new task assignments. > - The existing `agent_task_sessions` lookup, adapter session codec, workspace resolution, and effective config freshness checks already decide whether reuse is safe. > - Treating those two handoff wake reasons as unconditional fresh-session boundaries discards a valid saved task session before adapter resume can be attempted. > - This makes Dev → CodeReview → Dev loops repeatedly cold-start even when task, issue, agent, adapter, workspace, and config identity are unchanged. > - The fix is to let normal review/change-request handoffs reach the durable task-session path while preserving explicit fresh-session and unsafe-boundary resets. ## Linked Issues or Issue Description Fixes #8246. cc @cryppadotta — this is the narrow handoff-session policy change discussed there: normal `execution_review_requested` / `execution_changes_requested` wakes no longer force a fresh task session by wake reason alone, while assignment, approval, review-participant recovery, timer wakes, explicit `forceFreshSession`, and config/workspace/model/session freshness still keep their safety boundaries. ## What Changed - Removed normal `execution_review_requested` and `execution_changes_requested` from the unconditional task-session reset policy. - Kept fresh-session boundaries for: - `issue_assigned` - `execution_approval_requested` - `execution_review_participant_recovery` - `heartbeat_timer` - explicit `forceFreshSession` - existing config/model/workspace/session freshness reset paths - Updated heartbeat session-policy tests so execution handoffs are resume-eligible by wake reason alone. - Preserved PF-4 timer-wake behavior and its explicit reset reason. ## Verification - `npx pnpm@9.15.4 exec vitest run server/src/__tests__/heartbeat-workspace-session.test.ts server/src/__tests__/heartbeat-timer-wake-session-reset-pf4.test.ts server/src/__tests__/codex-local-execute.test.ts server/src/__tests__/issue-comment-reopen-routes.test.ts --reporter=verbose` — 220 tests passed. - `npx pnpm@9.15.4 --filter @paperclipai/server typecheck` — passed. - `git diff --check` — passed. - `coderabbit review --agent -t committed --base origin/master` — 0 findings. ## Risks - Moderate behavior change in session-boundary policy: normal review/change-request handoffs may now reuse a saved per-task session when the existing identity/freshness checks pass. - Safety boundaries remain in place for new assignments, approval gates, review-participant recovery, timer/discovery wakes, explicit fresh-session requests, and config/model/workspace/session drift. - If a saved session is stale or incompatible, existing freshness/resume fallback behavior still handles reset/fresh execution. > For core feature work, check ROADMAP.md first and discuss it in #dev before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See CONTRIBUTING.md. ## Model Used OpenAI GPT-5.5. Tool use and local verification were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket ID - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — N/A, server policy/test-only change - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: santastabber <184111696+santastabber@users.noreply.github.com> |
||
|
|
b90da4d115 |
fix: keep task sessions across issue comments (#10111)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Local session adapters persist task sessions so later wakes can resume the same conversation > - Session reuse correctly resets when effective execution configuration changes > - The workspace fingerprint currently includes the issue row's `updatedAt` timestamp > - Adding a comment advances that timestamp even though workspace configuration is unchanged > - The next same-issue wake therefore discards a valid task session and starts cold > - This pull request excludes that volatile timestamp while retaining actual workspace settings in the fingerprint > - The benefit is reliable same-task continuation without weakening configuration-freshness safety ## Linked Issues or Issue Description No public issue exists. The inline report below follows the bug report template. ### Pre-submission checklist - [x] I searched existing open and closed issues and found no duplicate. - [x] I reproduced the bug on the latest release and current `master`. - [x] I confirmed the error originates in Paperclip core fingerprinting, not an adapter, provider, or local configuration. ### What happened? On Paperclip 2026.720.0 and current `master`, a comment on an issue changes `issues.updated_at`. Heartbeat session fingerprinting includes that value under `workspaceConfig.issueConfigRevisionAt`, so the next wake for the same issue reports a workspace-config change and refuses the saved task session. ### Expected behavior Comment-only and other non-configuration issue updates should be delivered as wake deltas without invalidating the task session. Changes to the execution mode, issue workspace settings, project policy, environment, instructions, model, secrets, or other effective run configuration must still reset it. ### Steps to reproduce 1. Complete a local session-adapter run for an issue and retain its task session. 2. Add a comment to the issue without changing execution configuration. 3. Wake the same agent for the same issue. 4. Observe `changedCategories: ["workspaceConfig"]` and a fresh session. ### Paperclip version or commit Reproduced on Paperclip 2026.720.0 and current `master`. ### Deployment mode Self-hosted server. ### Installation method npm global install; also reproduced from the current source tree. ### Agent adapter(s) involved Codex exposed the symptom. The bug is in core fingerprint construction and is not adapter-specific. ### Database mode External Postgres. The bug is not database-specific. ### Access context Board comments trigger the timestamp change; the subsequent agent wake exposes the reset. ### Node.js version Node.js 22. ### Operating system Ubuntu 24.04. ### Relevant logs or output The next run records `changedCategories: ["workspaceConfig"]` and starts a fresh session after a comment-only mutation. ### Relevant config No unusual configuration is required. ### Additional context The regression test exercises the fingerprint directly on current `master`. ### Privacy checklist - [x] I reviewed the report for PII, credentials, private paths, company names, and instance-local identifiers. ## What Changed - Copy and sanitize the session workspace-fingerprint input before hashing. - Exclude only `issueConfigRevisionAt`, which reflects general issue mutation rather than workspace configuration. - Add regression coverage proving comment timestamps preserve the session while real workspace mode/settings changes still reset it. ## Verification - `pnpm exec vitest run server/src/__tests__/heartbeat-workspace-session.test.ts` - 120 tests passed. - `pnpm --filter @paperclipai/server typecheck` - passed. - `git diff --check` - passed. ## Risks Low risk. A general issue update no longer rotates the adapter session solely because its row timestamp changed. The fingerprint still includes issue workspace settings, issue adapter overrides, project workspace policy, environment, instructions, runtime skills, secrets, model profile, adapter configuration, and agent runtime configuration, so actual execution-config drift continues to reset. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected - check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, `gpt-5`, context-window size not exposed, reasoning and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Uliana Savostenko <ulia@MacBook-Air.local> |
||
|
|
7a4767017e |
fix(issues): base comment-wake decisions on post-insert issue state (#10068)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work; issues get commented on by both humans and agents, and the assignee is woken to act on new comments. > - #10050 added human-attributed issue comments for chat gateway plugins, with the host waking the issue's assignee the same way a board user's comment does. > - Greptile's review on #10050 flagged that the wakeup guard in `plugin-host-services.ts` decides whether to wake the assignee using the issue snapshot fetched *before* the comment was inserted. > - If another request closes, cancels, unassigns, or reassigns the issue in the window between that fetch and the wakeup call, the guard still acts on the stale snapshot — it can wake an agent for a now-terminal issue, or wake the old assignee instead of the new one. > - The PR discussion noted the HTTP add-comment route (`routes/issues.ts`) has the identical pattern outside its reopen/auto-approval branches, and deferred a fix to a follow-up covering both call sites — this PR is that follow-up. > - The fix re-fetches the issue immediately before the wake decision in both places, so the decision reflects the latest committed state instead of a pre-insert snapshot. ## Linked Issues or Issue Description Refs #10050 **Problem or motivation** Both the plugin-comment wakeup guard (`plugin-host-services.ts`) and the HTTP add-comment route's wakeup guard (`routes/issues.ts`, outside its reopen/auto-approval branches) decide whether to wake the issue's assignee using the issue state fetched before the comment was inserted. A concurrent close/unassign/reassign landing in that window is invisible to the guard, so it can enqueue a wakeup for a stale assignee or an issue that is no longer open. **Proposed solution** Re-fetch the issue immediately before the wake decision in both call sites, and base the assignee/status checks on that fresh read instead of the earlier snapshot. This shrinks the race window to essentially nothing (the fetch happens right before the fire-and-forget wakeup call), and any residual window is already covered by the heartbeat/checkout machinery re-validating issue status and assignee ownership when a woken run actually starts. **Alternatives considered** Wrap the whole comment-insert + wake-decision sequence in a single serializable transaction with row locking (rejected for this change — much larger blast radius across two already-complex handlers for a wakeup that is explicitly best-effort; the woken run's own re-validation already makes a stale wake degrade to a no-op rather than incorrect work). Leaving the plugin path fixed but not the HTTP route (rejected — that was the exact gap the original PR discussion flagged as needing a follow-up covering both call sites). **Roadmap alignment** Bug fix / hardening follow-up to #10050; no change to planned core roadmap items. ## What Changed - `server/src/services/plugin-host-services.ts`: `issues.createComment`'s assignee-wakeup guard now re-fetches the issue after the comment is inserted and bases the assignee/status checks on that fresh read, instead of the snapshot fetched before the insert. - `server/src/routes/issues.ts`: the `POST /issues/:id/comments` route's wakeup guard (outside the reopen/auto-approval branches, which already use post-mutation state) now does the same re-fetch before deciding whether — and whom — to wake. - Adds regression coverage for both: - `server/src/__tests__/plugin-orchestration-apis.test.ts`: a new embedded-Postgres test holds a row lock on the issue to deterministically force the race (comment-insert's internal update blocks until a concurrent transaction commits a cancellation), then asserts no wakeup is enqueued. - `server/src/__tests__/issue-comment-reopen-routes.test.ts`: two new mocked-service tests assert the route skips the wakeup when the fresh re-fetch shows the issue cancelled, and wakes the freshly reassigned agent (not the pre-insert snapshot's assignee) when the fresh re-fetch shows a different assignee. ## Verification - `pnpm --filter @paperclipai/server typecheck` — clean. - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/plugin-orchestration-apis.test.ts` — 13/13 (1 new). - `pnpm --filter @paperclipai/server exec vitest run src/__tests__/issue-comment-reopen-routes.test.ts` — 74/74 (2 new). - `pnpm --filter @paperclipai/plugin-sdk exec vitest run tests/host-client-factory.test.ts` — 14/14. - Broader sweep of 38 `routes/issues.ts`-adjacent test files (447 tests) — all passing, confirming the added re-fetch doesn't change behavior for any existing reopen/auto-approval/interrupt/scheduled-retry/dependency-wake scenario. ## Risks Low risk. Both changes are additive guards around an existing best-effort, fire-and-forget wakeup (failures already logged, not thrown) — no change to the comment-write path itself, response shape, or status codes. The HTTP route's fix only touches the plain (non-reopen, non-auto-approval) wake-decision path; the reopen and auto-approval branches already used post-mutation state for the reasons documented inline and are unchanged. Adds one extra `SELECT` per comment on each call site, negligible relative to the existing query volume in both handlers. ## Model Used Claude Sonnet 5 (`claude-sonnet-5`), extended thinking, tool use enabled, via Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I searched the GitHub PR list (open and recently closed) for similar PRs; found no duplicate — this is a direct follow-up to the review discussion on #10050 - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: anicca <annica@Michaels-Mac-Studio.local> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>canary/v2026.817.0-canary.13 |
||
|
|
7ef75f5636 |
fix(opencode-local): retry models preflight during transient contention (#9225)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Local CLI adapters are responsible for starting agent runtimes and validating that their configured models are usable before a run starts. > - The OpenCode local adapter checks `opencode models` during model discovery and preflight validation. > - On hosts with a shared Ollama daemon, that lightweight metadata call can transiently queue behind an active generation and time out or return a short failure. > - Treating that transient contention as a hard adapter failure prevents otherwise valid local OpenCode runs from starting. > - This pull request adds a small bounded retry/backoff around OpenCode model discovery while keeping the existing per-attempt timeout and surfacing a final failure when retries are exhausted. > - The benefit is fewer false adapter failures during local Ollama contention without changing shared Ollama configuration or hiding genuinely stuck model discovery. ## Linked Issues or Issue Description No public GitHub issue exists for this adapter reliability bug. Bug description: - What happened: `opencode models` can transiently time out or fail while a shared local Ollama daemon is busy serving another OpenCode generation, causing the adapter preflight to fail before the actual run starts. - Expected behavior: transient model-list contention should be retried briefly before declaring the adapter unavailable. - Steps to reproduce: run an OpenCode local adapter using an Ollama-backed model while another `opencode run` is actively generating against the same daemon, then trigger model discovery/preflight during that contention window. - Paperclip version/commit: observed on the current Paperclip master-line OpenCode local adapter before this change. - Deployment mode: local trusted / local CLI adapter execution with a shared local Ollama daemon. Related search: - Searched public GitHub issues for `opencode models preflight retry`; no matching issue found. - Searched public GitHub PRs for `opencode models preflight retry`; no matching PR found. The only search hit was unrelated OpenClaw gateway authentication work (#6121). ## What Changed - Added bounded retry/backoff to OpenCode model discovery: three total attempts with 2s and 4s waits between failures. - Preserved the existing 20s per-attempt `opencode models` timeout. - Retry covers timeout and non-zero process exits, while spawn-level failures still surface immediately. - Added unit coverage for transient fail -> timeout -> success behavior and exhausted retry behavior. - Updated existing OpenCode environment diagnostic tests with explicit timeouts for the intentional retry/backoff path. ## Verification - `pnpm --filter @paperclipai/adapter-opencode-local exec vitest run src/server/models.test.ts src/server/execute.test.ts` -> 2 files passed, 13 tests passed. - `pnpm --filter @paperclipai/adapter-opencode-local typecheck` -> passed. - `pnpm vitest run server/src/__tests__/opencode-local-adapter-environment.test.ts` -> 1 file passed, 3 tests passed. - Branch diff against current `upstream/master` is limited to `packages/adapters/opencode-local/src/server/models.ts`, `packages/adapters/opencode-local/src/server/models.test.ts`, and `server/src/__tests__/opencode-local-adapter-environment.test.ts`. ## Risks Low risk. This only changes OpenCode model discovery behavior and keeps the preflight bounded. A genuinely unavailable `opencode models` call still fails after three attempts, and command spawn failures are not masked. ## Model Used OpenAI Codex, GPT-5.5 coding agent, tool-enabled repository editing and shell verification in a local Paperclip workspace. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Test <test@paperclip.ing> |
||
|
|
d77eeb8914 |
fix(sandbox-bridge): allow the agent-hire skill's routes through the callback bridge (#8978)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Managed agents run inside a sandbox and reach the Paperclip server only through the sandbox callback bridge, which forwards a fixed route allowlist (`DEFAULT_SANDBOX_CALLBACK_BRIDGE_ROUTE_ALLOWLIST`) > - The `paperclip-create-agent` skill instructs an agent to call adapter/icon discovery endpoints, compare existing agent configurations, submit a hire request, and link the resulting approval to its source issue > - None of those routes were on the bridge allowlist, so a sandboxed agent following the skill correctly hit `Route not allowed` on every call — including the hire `POST` itself — making hiring impossible from inside a sandbox > - This pull request adds the six routes the skill uses to the bridge allowlist, while keeping direct agent creation (`POST /api/companies/:id/agents`) denied > - The benefit is that hiring works end-to-end for sandboxed agents through the approval-gated `agent-hires` path, without widening the bridge beyond what the skill needs ## Linked Issues or Issue Description No public issue exists; describing the bug in-PR (bug template fields): - **What happened:** A managed agent running in a sandbox followed the `paperclip-create-agent` skill and got `Route not allowed` from the callback bridge on every endpoint the skill documents — adapter discovery (`/llms/agent-configuration.txt`, `/llms/agent-configuration/:adapterType.txt`, `/llms/agent-icons.txt`), config comparison (`GET /api/companies/:id/agent-configurations`), the hire submission (`POST /api/companies/:id/agent-hires`), and approval linking (`POST /api/issues/:id/approvals`). - **Expected behavior:** An agent with hiring permission can complete the hire flow from inside a sandbox; the bridge forwards the skill's routes and the server enforces authorization (`canCreateAgents`). - **Impact:** Hiring by sandboxed agents was fully broken — the failure is in the transport allowlist, not permissions, so no configuration could work around it. Related: #8981 (companion fix making the `paperclip-create-agent` skill available to agents that can hire; supersedes #8823). The two changes serve the same end-to-end hire flow but are independently mergeable — this PR is purely the bridge transport allowlist. Supersedes #8853. ## What Changed - `packages/adapter-utils/src/sandbox-callback-bridge.ts`: add six routes used by the `paperclip-create-agent` skill to `DEFAULT_SANDBOX_CALLBACK_BRIDGE_ROUTE_ALLOWLIST` (three `GET /llms/...` discovery routes, `GET .../agent-configurations`, `POST .../agent-hires`, `POST /api/issues/:id/approvals`), with a comment documenting why direct agent creation stays denied - `packages/adapter-utils/src/sandbox-callback-bridge.test.ts`: assert the six routes are allowed, and add negative cases proving the regexes do not over-match (no `POST .../agents`, no non-`.txt` or arbitrary `/llms` files, no `agent-hires` sub-resources) ## Verification - `npx vitest run packages/adapter-utils/src/sandbox-callback-bridge.test.ts` — 13/13 tests pass locally - `npx tsc --noEmit -p packages/adapter-utils` — clean - Manual: run a managed agent in a sandbox, invoke the `paperclip-create-agent` skill, and confirm the discovery calls, hire `POST`, and approval linking all pass through the bridge; `POST /api/companies/:id/agents` still returns `Route not allowed` ## Risks - Low risk: additive allowlist entries only; anchored regexes with `[^/]+` segments prevent over-matching (covered by tests) - The bridge allowlist bounds surface area but does not replace server-side authorization — the hire `POST` remains approval-gated and permission-checked (`canCreateAgents`) on the server > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - Claude (Anthropic) — Fable 5 (`claude-fable-5`), extended thinking, agentic tool use via Claude Code; original diff authored with Claude Opus 4.8 (1M context) ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
93ff6a8771 |
fix(cursor-cloud): drop unreachable Paperclip API callback for remote… (#8546)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents run via per-adapter execute paths; the `cursor_cloud` adapter runs the agent in Cursor's cloud (remote), orchestrated server-side via the Cursor Agent SDK > - Local adapters receive a run-scoped Paperclip JWT (`supportsLocalAgentJwt=true`) injected as `PAPERCLIP_API_KEY` so the agent can call the Paperclip API; `cursor_cloud` is intentionally `supportsLocalAgentJwt=false` (no JWT minted for a remote worker) > - But `buildPaperclipEnv` always sets `PAPERCLIP_API_URL` (defaulting to the local runtime host), so the remote cloud worker is handed a callback URL it can neither reach nor authenticate against > - Any agent-initiated Paperclip API call from the cloud worker therefore fails with a 401 (or is unreachable), producing log noise and confusing failures > - This pull request drops the callback wiring when there is no usable key, so cloud-side Paperclip tools degrade to a clean no-op > - The benefit is no spurious 401s from remote cloud runs, with run results unaffected (delivered server-side via the Cursor Agent SDK) ## Linked Issues or Issue Description No existing public issue — describing the bug inline (per `.github/ISSUE_TEMPLATE/bug_report.yml`): **What happened** `cursor_cloud` runs emit 401s when the remote cloud agent attempts Paperclip API calls. Root cause: `buildPaperclipEnv` (`packages/adapter-utils/src/server-utils.ts`) always sets `PAPERCLIP_API_URL` (local runtime default), while `cursor_cloud` has `supportsLocalAgentJwt=false`, so no `PAPERCLIP_API_KEY` is minted — URL present, key absent → 401 / unreachable from `buildWakeEnv` in `packages/adapters/cursor-cloud/src/server/execute.ts`. **Expected behavior** A remote cloud worker that is not issued a run JWT should not attempt (and fail) Paperclip API callbacks. **Steps to reproduce** 1. Configure a `cursor_cloud` agent (runs in Cursor's cloud; `supportsLocalAgentJwt=false`). 2. Trigger a run that causes the cloud agent to make a Paperclip API call. 3. Observe a 401 (or unreachable) because `PAPERCLIP_API_URL` points at an unreachable local runtime and no key is present. **Paperclip version** Reproduced on current `master` (cutover base `e68188c43`). **Deployment mode** Self-hosted control plane, `cursor_cloud` adapter (remote execution in Cursor's cloud). **Related PRs (searched; none duplicate this fix):** - #8197 — `claude_local` opt-out of the sandbox *bridge* for direct-reachable remote SSH targets. Related family, but the opposite situation: that path keeps the callback because the remote is reachable **and** has a run token. `cursor_cloud` has neither, so here the callback is removed. - #8130, #4794, #8025 — `PAPERCLIP_API_URL`/loopback injection for **local** agents (distinct from the remote cloud worker case). - #401 — alternative agent-auth scheme (run-ID header when no bearer token); different approach, not overlapping with this targeted fix. ## What Changed - `packages/adapters/cursor-cloud/src/server/execute.ts`: in `buildWakeEnv`, when there is no usable `PAPERCLIP_API_KEY`, delete `PAPERCLIP_API_URL` and `PAPERCLIP_API_BRIDGE_MODE` so the remote worker performs no Paperclip API callbacks. Informational `PAPERCLIP_*` vars (run id, agent id, company id, task, wake reason) still flow. When a key *is* present (operator-provided), the URL is retained. - `packages/adapters/cursor-cloud/src/server/execute.test.ts`: new test asserting no callback vars are injected when no run JWT is present; positive assertion that the URL is retained when a key is present. ## Verification - `pnpm exec vitest run packages/adapters/cursor-cloud/src/server/execute.test.ts` → **5/5 pass**. - `pnpm --filter @paperclipai/adapter-cursor-cloud typecheck` → **green**. - Confirmed result delivery does not depend on this callback: `execute()` reads results server-side via `Agent.getRun()` and `run.wait()`. ## Risks - **Low risk.** Only affects the env handed to remote `cursor_cloud` workers. No schema/migration/behavioral change to result delivery (which is server-side). When an operator explicitly provides `PAPERCLIP_API_KEY`, the callback URL is retained, preserving intentional callback setups. ## Model Used - **Claude Opus 4.8** (Anthropic), extended/high reasoning mode, via the Cursor agent with tool use + code execution. Diagnosis grounded in the adapter/runtime code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above (none found) - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub references) - [x] My branch name describes the change (`fix/cursor-cloud-skip-unreachable-callback`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [ ] I have updated relevant documentation to reflect my changes (N/A — no documented behavior changes) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Sebastian Heyneman <sebastian@joinnova.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
2ae6fa51b1 |
fix: warn when a worktree-mode embedded-postgres data dir is in the OS temp dir (#8283)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work, and it can run as a self-hosted instance backed by an embedded PostgreSQL database. > - To let agents work in isolation, Paperclip supports worktree-local instances, gated by `PAPERCLIP_IN_WORKTREE` / `PAPERCLIP_HOME` (see `server/src/worktree-config.ts`, `cli/src/config/home.ts`). > - Those worktree env vars can leak into a *primary* instance's environment (inherited from an agent/worktree shell, or persisted into the instance env file). When they do, `paperclipai run` resolves the data root from `PAPERCLIP_HOME` and rewrites `config.json` to point the DB/backups/logs/storage at `$PAPERCLIP_HOME/instances/<id>/…`. > - If `$PAPERCLIP_HOME` is a throwaway dir under the OS temp dir, the primary instance boots a brand-new **empty** database. Every login then fails (`better-auth` logs `User not found`; the UI returns a generic `401`), so it looks like a *password* problem while the real data sits untouched in `~/.paperclip`. Nothing warns that the control-plane DB is ephemeral. > - This pull request makes that situation non-silent: the database preflight check (also surfaced by `doctor`) emits a `warn` when a worktree-mode instance's embedded-postgres data dir is inside the OS temp directory, with clear remediation. > - The benefit is that a confusing total lockout becomes an obvious, actionable warning the operator sees at every `run` and `doctor`. ## Linked Issues or Issue Description Refs #8282 Related PRs (not duplicates — complementary work on the same area): - #3030 — *stop leaking server worktree env into unrelated local adapter heartbeats* (tackles one **leak vector** of the same root cause; this PR adds **detection** of the resulting bad state). - #3899 — *fix(db): refuse side-started embedded migration instances* (adjacent embedded-postgres safety hardening). ## What Changed - `cli/src/checks/database-check.ts`: for `embedded-postgres` mode, emit `status: "warn"` when the resolved data dir is inside `os.tmpdir()` **and** `PAPERCLIP_IN_WORKTREE === "true"`. The message explains the ephemerality + likely env leak; the repair hint says to unset `PAPERCLIP_HOME` / `PAPERCLIP_IN_WORKTREE` (or pass `--data-dir`). Added a small `isInsideOsTmpDir()` helper. - Intentionally gated on worktree mode so deliberate ephemeral/CI instances that use a temp data dir without `PAPERCLIP_IN_WORKTREE` are **not** flagged. - `cli/src/__tests__/database-check.test.ts` (new): covers pass (persistent dir), warn (worktree-mode temp dir), and no-warn (temp dir without worktree mode). ## Verification ``` # in cli/ pnpm exec vitest run src/__tests__/database-check.test.ts # 3 passed pnpm exec vitest run src/__tests__/doctor.test.ts # passes (no regression) pnpm exec tsc --noEmit # no new errors in changed file ``` Manual repro of the underlying bug (no warning before this change): ```bash PAPERCLIP_IN_WORKTREE=true PAPERCLIP_HOME="$(mktemp -d)/.paperclip-worktrees" paperclipai run # -> boots an empty DB under /tmp; logins fail with "User not found". # With this change, run/doctor now print a Database WARN pointing at the cause + fix. ``` Note for transparency: one unrelated test (`worktree.test.ts > pauseSeededScheduledRoutines`) fails *locally only* because it shells out to a real `pnpm install` that times out in a sandbox — it does not touch `database-check`. Pre-existing `tsc` errors under `server/src/services/plugin-*` (missing `@paperclipai/plugin-sdk` build artifact) are also unrelated to this change. ## Risks Low risk. Additive, non-fatal `warn` only — no behavior change to startup or existing `pass`/`fail` paths, and gated on `PAPERCLIP_IN_WORKTREE` so it does not fire for intentional ephemeral/CI temp data dirs. A stricter follow-up (refuse to start the primary `run` against a temp-dir data dir unless explicitly opted in) is possible but intentionally out of scope here. ## Model Used - **Provider / model:** Anthropic Claude — Opus 4.8 - **Exact model ID:** `claude-opus-4-8` (1M-context variant) - **Context window:** 1M tokens - **Reasoning mode:** extended thinking enabled - **Capabilities used:** agentic tool use via Claude Code (repository exploration, file edits, local shell, and running the vitest suite locally before pushing) The change was authored by @futhgar with this model as an assistant; the diagnosis, fix, and tests were reviewed and verified locally. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass (new test + affected `doctor` suite; see Verification for one unrelated, environment-only failure) - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (N/A — no UI change) - [x] I have updated relevant documentation to reflect my changes (the warning message + repair hint are self-documenting; no separate docs change needed) - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (pending CI run) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (pending review) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: futhgar <futhgar@users.noreply.github.com> |
||
|
|
14fd8aee36 |
fix(server): flag truncated issue descriptions (#4771)
## Thinking Path > - Paperclip orchestrates AI agents for zero-human companies. > - The issue list API is one of the surfaces API consumers use to synchronize issue metadata. > - The list endpoint intentionally returns a bounded `description` preview so large descriptions do not bloat list responses. > - Before this change, that preview looked like a complete field value because the response did not say whether it had been shortened. > - That made round-trip clients vulnerable to accidentally PATCHing a preview back over the full description. > - This pull request keeps the existing preview behavior but adds an explicit `descriptionTruncated` flag. > - The benefit is backwards-compatible visibility into truncated issue descriptions, so clients can avoid data-loss workflows. ## Linked Issues or Issue Description Fixes #4758. Related PR: #4792 also targets #4758, but it includes unrelated logger changes and currently has separate review/security concerns. This PR keeps the fix scoped to the issue-list description truncation API behavior. ## What Changed - Added `descriptionTruncated` to the issue list projection when `description` exceeds the existing 1200-character preview limit. - Exposed `descriptionTruncated?: boolean` on the shared `Issue` type. - Added service tests for truncated descriptions, exact-limit descriptions, null descriptions, and multibyte-safe preview truncation. ## Verification June 18, 2026 refresh after rebasing onto current `origin/master`: - `pnpm install --frozen-lockfile` - `pnpm exec vitest run server/src/__tests__/issues-service.test.ts` - `pnpm --filter @paperclipai/server typecheck` - `pnpm typecheck` - `git diff --check origin/master...HEAD` - GitHub PR checks are green on head `12e828e6`. Earlier pre-review verification also included `pnpm test`. ## Risks - Low risk. This is an additive API response field; existing clients can ignore it. - The list endpoint still returns the same bounded `description` preview. Clients that need full text should continue fetching the issue detail, but can now detect when that is necessary. - No database migration or UI behavior change. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex based on GPT-5, via Codex desktop on April 29, June 15, and June 18, 2026. Used tool-assisted repository inspection, code editing, local test execution, GitHub CLI workflows, and PR review follow-up. Exact context window size is not surfaced by the tool. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] If this change affects the UI, I have included before/after screenshots (N/A: no UI change) - [x] I have updated relevant documentation to reflect my changes (N/A: additive API field covered by tests) - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Sami Rusani <sr@samirusani> |
||
|
|
c1c46f1e4e |
feat: Claude login on the new-agent page before agent creation (#11347)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The Claude local adapter supports subscription login through a sandbox > - The new-agent page must show login before the user creates an agent > - Test results must not expose raw sandbox diagnostics or secret values > - This pull request adds the login UI to both Test lanes and closes the diagnostic boundary > - The branch also adds durable cleanup recovery for failed sandbox teardown > - Reusable sandboxes must retain both their recorded teardown configuration and a valid lifecycle path until destruction succeeds > - The benefit is a usable login flow with fixed public checks, redacted server logs, and recoverable sandbox cleanup ## Linked Issues or Issue Description Related public work: [#9488](https://github.com/paperclipai/paperclip/pull/9488) adds first-class recognition for `CLAUDE_CODE_OAUTH_TOKEN` in headless and remote runs. Related public issue: [#2681](https://github.com/paperclipai/paperclip/issues/2681) requests Claude Code subscription support. This pull request adds the login transport and new-agent UI flow that those changes do not provide. **Subsystem affected:** Claude local adapter, server login probes, sandbox provider setup, cleanup recovery, and the new-agent UI. **Problem or motivation:** The Test lanes did not show the sandbox login panel in all supported cases. Test results also exposed raw probe diagnostics, and JSON escapes could end secret redaction early. **Proposed solution:** Surface the login capability through the bundled provider manifest. Prepare the same probe runtime in the ACP lane. Send diagnostics only to redacted server logs. Keep Test checks on fixed public messages. Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. Consume JSON escapes during redaction. Preserve failed sandbox cleanup state across retries and restarts, and prevent deletion from severing the lifecycle context of a live reusable sandbox. **Alternatives considered:** Keep raw diagnostics in Test checks or trust login URL text from the sandbox. Both choices increase information exposure. Keep separate probe behavior in the ACP lane. That choice would leave the two Test lanes inconsistent. ## What Changed - Surface the sandbox login panel on both Test lanes. - Reconcile the bundled Daytona plugin manifest so `supportsSetupTokenLogin` reaches the UI capability gate. - Prepare the ACP Test lane with the same probe runtime as the CLI Test lane. - Add the `claude_acp_login_probe_unavailable` warning when the ACP probe cannot run. - Send raw sandbox diagnostics only to redacted server logs. - Keep Test checks on fixed public messages in the ACP, managed-config, and CLI paths. - Normalize login URL hints to allowlisted HTTPS Claude and Anthropic hosts. - Redact JSON and escaped-JSON secret values, including escaped quotes and backslashes. - Preserve orphan cleanup records across provider failures, restarts, and unavailable plugins. - Atomically block environment deletion while a live reusable sandbox lease still depends on it. - Verify pending cleanup destroys plugin sandboxes with the provider configuration recorded on the lease, even after the current environment configuration changes. ## Verification - Head under review: `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241`. - Focused environment route/service/runtime coverage passes: 196 tests across 3 files. - `pnpm -r typecheck` passes. - `pnpm build` passes. - The full Vitest run completed with 4,754 passing and 28 failing tests. All 23 source-test failures reproduce unchanged on parent head `58cfe61a33191ce03d965d65085d26064b4888ba`; the other 5 are duplicate executions from stale `server/dist` output. The failures are unrelated macOS path/listener and scheduler-fixture failures, so there is no new bad commit for bisect to localize. - All required CI checks pass for the current head, including build, typecheck/release registry, all server and workspace shards, serialized server suites, canary, and e2e. - A fresh Greptile review for `506b7fa2d83c36bfa5fd722ee9d95b0c7431c241` reports 5/5, “safe to merge,” with no blocking failure remaining. ## Risks - A probe or redaction change could hide useful server diagnostics. - An allowlist change could reject a valid Claude login URL. - Cleanup recovery changes could affect provider teardown ordering. - An environment with a live reusable sandbox can no longer be deleted until the owning issue or execution workspace completes teardown. - The implementation keeps public Test messages fixed and sends detail to redacted server logs. ## Model Used OpenAI GPT-5 via Codex — exact model ID: GPT-5; tool use and code execution enabled; extended reasoning enabled. The implementation author used AI-assisted development. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and documented the result - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation or confirmed no separate documentation change is needed - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3061ce6901 |
feat(sandbox): stream session output by capability, drop three operator flags (#11557)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Sandboxed agents use provider capabilities to select safe execution paths > - Session output still depends on three operator flags that duplicate capability data > - Duplicate flags can drift from the verified sandbox capability snapshot > - This pull request makes the capability snapshot the only streaming decision and removes the obsolete flags > - The benefit is default streaming with a poll fallback when a capability or stream fails ## Linked Issues or Issue Description **What existing behavior does this improve?** ACP sandbox session-output streaming and sandbox execution configuration. **Subsystem affected** Cross-cutting (multiple of the above): server/, packages/shared/, packages/adapter-utils/, and packages/plugins/. **Current behavior** Session-output streaming requires operator flags in the server and Daytona plugin configuration. Saved configurations can retain a removed key. **Proposed behavior** The verified capability snapshot selects streaming. The Daytona plugin uses persistent sessions by default, keeps bypass commands one-shot, and falls back from the log stream to polling. Removed configuration keys become inert. **Reason and benefit** One capability source prevents configuration drift. The fallback keeps output available when capability resolution or log streaming fails. **Breaking changes** The three operator flags no longer control session-output streaming. Existing saved keys load but have no effect. ## What Changed - Remove `useSessions` and `useLogStream` from the Daytona plugin configuration and manifest. - Remove `streamAgentSessionOutput` from server configuration, shared types, and execution-target plumbing. - Select streaming from `persistentProcessSessions` and `independentControlCommands`. - Keep poll fallback on capability resolution failure and stream failure. - Strip removed keys from strict fake-sandbox and catchall plugin configuration. - Update the sandbox capability documentation and focused tests. ## Verification - `tsc --noEmit` passed in `packages/shared`, `packages/adapter-utils`, `server`, and the Daytona plugin. - Daytona `plugin.test.ts` passed 139 tests. - Server capability, configuration, route, and runtime suites passed 160 tests. - `packages/adapter-utils` `execution-target-sandbox.test.ts` passed 44 tests. - The capability matrix covers stream, poll, and resolution-failure paths. - Removed-key tests cover strict fake-sandbox and catchall plugin schemas. ## Risks - A capability snapshot that lacks either required session capability uses polling. - A log stream failure uses polling and can increase request count. - Existing removed configuration keys no longer change behavior. - The isolated-worktree Daytona Vitest run has a pre-existing missing `packages/adapters/droid-local` reference. CI and standard checkouts use the committed configuration. ## Model Used OpenAI Codex, GPT-5, tool use and code review assistance. The exact runtime context window is managed by the Codex platform. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.817.0-canary.12 |
||
|
|
e71ce9a9d3 |
feat: sandbox provider capability contract with fail-closed effective resolution (#11463)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip runs work through adapters and sandbox providers > - Providers need a clear contract so the server can use only verified capabilities > - A declared capability must not grant a method that the live worker did not verify > - This pull request adds manifest declarations and fail-closed effective capability resolution > - The benefit is safe provider reuse across execution targets and run lifecycles ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting (multiple of the above) **Problem or motivation** Sandbox providers expose different runtime methods. The server needs one safe capability contract that accounts for provider declarations, worker verification, and narrowing configuration. **Proposed solution** Add strict manifest validation for five sandbox capabilities. Resolve effective capabilities as the subset of verified, declared, and narrowed values. Store the result as a frozen execution-target snapshot. **Alternatives considered** Trusting the manifest alone could grant methods that the worker does not support. Trusting only a fixed built-in list would reject valid third-party providers. The intersection rule keeps the verified runtime ceiling and supports both provider types. **Roadmap alignment** This change supports the ACP run lifecycle track and the sandbox provider contract work in the current roadmap. **Additional context** The legacy `supportsReusableLeases` field remains supported. The nested capability validator rejects unknown keys. Missing or unavailable verification resolves all capabilities to `false`. ## What Changed - Add strict `sandboxCapabilities` manifest validation with legacy reusable-lease compatibility. - Carry declarations through the ready-driver projection. - Add fail-closed effective resolution from verified, declared, and narrowed capabilities. - Add narrowing for provider configuration, Kubernetes Job leases, and Daytona sessions. - Add a frozen read-only capability snapshot to execution targets. - Add focused tests and keep existing characterization baselines covered. - Add and update sandbox provider capability documentation. ## Verification - `npx vitest run packages/shared/src/validators/plugin.test.ts` - `npx vitest run server/src/__tests__/plugin-environment-driver-sandbox-capabilities.test.ts` - `npx vitest run server/src/__tests__/sandbox-capability-contract.test.ts` - `npx vitest run server/src/__tests__/environment-execution-target-capabilities.test.ts` - `npx vitest run packages/adapter-utils/src/acpx-engine/startup-characterization.test.ts packages/adapter-utils/src/acpx-engine/turn-characterization.test.ts packages/adapter-utils/src/acpx-engine/settlement-characterization.test.ts packages/adapter-utils/src/acpx-engine/composed-run-characterization.test.ts` - Package typechecks for shared, server, and adapter-utils pass. - Stage-2 security review suites pass with 28 tests. ## Risks The resolver fails closed when verification is absent or unavailable. Providers that rely on undeclared capabilities may see narrower behavior until they expose verified worker methods. The change does not alter the existing native-sync guard. ## Model Used OpenAI Codex, GPT-5, exact runtime model ID `gpt-5`, tool use and code execution. The implementation author used this model to assist with the change. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.817.0-canary.11 |
||
|
|
d2fb05d225 | Fix issue document deep-link routing (#11551) | ||
|
|
1a17cbf232 |
fix(runtime-exposure): mediate leased app/HMR port pairs centrally (#11526)
<!-- Simplified Technical English (ASD-STE100). --> > **Stacked pull request.** This targets #11525, which targets #11524. Merge those first. Review only the last commit, `fix(runtime-exposure): mediate leased app/HMR port pairs centrally`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts managed runtime services for execution workspaces, and #11524 and #11525 make those services reachable over Tailscale HTTPS on a loopback port pair > - An HTTPS lane is only safe if one execution workspace holds its port pair exclusively for the whole life of the lane > - A managed start reused a pair that a stopped but still leased workspace owned. Paperclip reported that workspace stopped and its exposure removed, while the host listeners and the Serve mappings for those ports were live and belonged to an unrelated workspace > - The cause is that ownership was decided in more than one place, and no single place saw persisted reservations, live listeners, and Serve mappings together > - This pull request adds one mediator that owns the decision, and makes every mismatch fail closed while naming the conflicting workspace > - The benefit is that a later start cannot collide with, adopt, or interfere with another issue's service, and cannot produce security evidence attributed to the wrong workspace ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the bug report template. **What happened** A managed HTTPS start reused the loopback port pair of a stopped but still exclusively leased execution workspace. The ports were then held by an unrelated workspace. Paperclip continued to report the first workspace's runtime as stopped and its exposure as removed, while the host listeners and the `tailscale serve` mappings for those exact ports were live and owned by the other workspace. **Expected behavior** An active execution-workspace lease reserves its app and HMR pair until the lease is explicitly released or torn down. A start that finds the pair held by a different workspace fails closed and names the conflict. Paperclip never adopts a process or a Serve mapping across execution-workspace ids. **Root cause** Three separate readers each had an incomplete view: - `deprovisionExposure` replaces the exposure status with a fresh `removed` status whose `listeners` array is empty. A later reader asking "which ports did this row own?" gets no answer, so a stopped row's pair looked free even while the row was leased. - Startup reconciliation adopted a persisted service by `row.port` alone, then terminated the local service when its health check failed. Under the `project_primary` strategy, where workspaces share a working directory, the containment check cannot separate two workspaces, so the sweep could adopt and then kill an unrelated workspace's live service. - Allocation checked live port availability but never checked which pairs active leases still reserve. **Impact** Two workspaces can collide on one lane. A start can adopt or interfere with another issue's service, and evidence about an exposure can be attributed to the wrong workspace. ## What Changed - Add `server/src/services/runtime-exposure/port-reservation.ts`, one mediator that decides allocation and ownership from persisted reservations plus live listener and Serve ownership together. - Reserve a pair for as long as its execution workspace holds an active lease, until the lease is explicitly released or torn down. - Re-derive a row's pair from the `port` column and `deriveViteHmrPort` instead of the status `listeners` array, so a `removed` status no longer hides which ports a leased row still reserves. - Refuse to adopt a process or a Serve mapping across execution-workspace ids. A mismatch fails closed and names the conflicting workspace and issue. - Treat an unattributable holder as a conflict. A Serve mapping that is present but cannot be attributed means the host has something there that could not be named, so it fails closed instead of falling through to "allowed". - Make reconciliation surface a stopped or removed row whose reserved ports are live or mapped by another workspace, instead of reporting success. - Leave manual and unknown Serve mappings alone on release and teardown. ## Verification - `npx vitest run --root server src/services/runtime-exposure/ src/__tests__/workspace-runtime-exposure-reservation.test.ts src/services/workspace-runtime-exposure-backfill.test.ts` — 8 files, **112 tests pass**. - `npx tsc --noEmit -p server/tsconfig.json` — **0 errors** with `@paperclipai/plugin-sdk` built. - `pnpm --filter @paperclipai/db typecheck` — migration numbering and safety checks pass. - `pnpm --filter @paperclipai/tailscale-https-broker test` — 87 tests pass. The five required regressions are covered by `workspace-runtime-exposure-reservation.test.ts` and `port-reservation.test.ts`: 1. Reuse of a stopped-but-leased pair is denied. 2. Cross-execution-workspace process adoption is denied. 3. A Serve mapping ownership mismatch is visible and fails closed. 4. Concurrent allocators return unique pairs. 5. Release and teardown make the pair reusable without harming manual or unknown mappings. Note for reviewers: `server/src/services/workspace-runtime-exposure.test.ts` fails on a development host that already runs an HTTPS canary holding ports 42000, 42001, 52000, and 52001, because that fixture stubs port availability and then allocates into the occupied range. It is unaffected by this change and is expected to pass in CI, where no such listener exists. Please read the CI result rather than a local run on an exposing host. ## Risks - The mediator is now the single decision point for allocation and adoption, so a defect in it affects every managed start. This is deliberate: the incident happened because the decision was spread across three readers, and concentrating it is the fix. - Behavior becomes stricter. A start that previously reused a pair now fails closed with a named conflict. This is the intended change, and it can surface pre-existing collisions that used to pass silently. - The remediation path does not stop an unrelated service that already holds a pair. It reports the conflict instead, so it cannot disturb another issue's running lane. - No migration runs in this pull request. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.817.0-canary.10 |
||
|
|
4c349fe6b7 |
feat(runtime): managed Tailscale HTTPS lifecycle, durable runtime leases, and bounded control recovery (#11525)
<!-- Simplified Technical English (ASD-STE100). --> > **Stacked pull request.** This targets #11524. Merge #11524 first. Review only the second commit, `feat(runtime): managed Tailscale HTTPS lifecycle...`. ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services, so an agent's branch can be previewed while the agent works > - The previous pull request added the host broker, the shared contract, and the database columns, but no code used them > - A managed runtime can only be exposed over HTTPS if it holds a stable loopback port pair for the whole life of the service. The current control path cannot promise this: two controls can race the same execution workspace, a stranded control can stay `running` forever, and a start can adopt a port it does not own > - This pull request adds the HTTPS lifecycle and the control-path hardening that the lifecycle depends on > - The benefit is that a managed preview becomes reachable from another device, and a managed control now always reaches a terminal state ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, workspace operations, the execution workspace routes, and the workspace runtime UI. **Problem or motivation** A managed runtime service is reachable only on loopback, so a preview cannot be opened from a phone or a second computer. Exposing it safely needs an exclusively held port pair. Three existing gaps block that. Overlapping controls can race the same workspace. A control whose owner dies stays `running` and blocks the lane forever. Port allocation does not confirm that the process holding a port is the process Paperclip spawned. **Proposed solution** Add the exposure lifecycle on top of the broker from #11524: reserve before spawn, expose after readiness, validate the public URL, and remove on stop. In the same change, make managed controls mutually exclusive per workspace, give each control a durable issue-owned lease and a terminal state, and verify port ownership before use. **Alternatives considered** - Add HTTPS exposure without the control hardening. This was rejected because a raced or stranded control makes exposure point at the wrong process. - Guard the lane with an in-memory lock only. This was rejected because the lock does not survive a server restart, so the lane can be lost or double-claimed. - Trust the requested bind address. This was rejected because a checkout that predates managed HTTPS overwrites `PAPERCLIP_BIND` from its own `--bind` argument, and then binds the wildcard address. **Roadmap alignment** This completes the managed workspace runtime capability that already exists. It adds no new product surface beyond the HTTPS link. **Additional context** This is the second of three pull requests. The third adds central mediation of leased port pairs. ## What Changed Exposure lifecycle: - Add the server-side broker client and the exposure lifecycle manager. The manager reserves the mapping before spawn, exposes after backend readiness, validates the public URL, and removes the mapping on stop. - Default managed worktree runtimes to `tailscale_https`, read exposure intent from legacy `expose` blocks, and backfill runtimes that are still HTTP-only. - Verify listener ownership for the app port and its Vite HMR companion before the broker is asked to expose anything. An unrelated listener on either port fails the start closed. - Force the loopback bind through argv instead of environment hints. Leave a non-Paperclip service's `--bind` argument alone. - Probe loopback for readiness instead of the public URL, and give Vite HMR its own loopback-bound server in middleware mode. - Preserve operator-declared Serve mappings across the managed lifecycle, so cleanup never removes a mapping that Paperclip did not create. - Name which listener predicate denied an expose, so an operator can act on the message. Control-path hardening: - Make `start`, `stop`, `restart`, and job `run` mutually exclusive per execution workspace. An overlap gets `409 workspace_runtime_control_in_progress`, and authorization is still checked first. - Take a durable exclusivity lease on the execution workspace, owned by the controlling issue. A different issue gets `409 workspace_runtime_lease_conflict` before any operation is recorded. Board and operator actions bypass the lease. - Give every control a terminal state. Each control stamps its owning process and pid, heartbeats while it runs, and has a wall-clock ceiling. Recovery of a stranded control uses a compare-and-swap on `updated_at`, so a live owner is never stolen. - Bound readiness probes, verify allocated port ownership on POSIX and Windows, harden sibling port allocation, and reconcile desired runtimes on server startup. - Surface exposure state and bounded runtime errors in the workspace runtime UI. - Record the new behavior in `doc/DEVELOPING.md`. ## Verification Focused checks, all run on this branch: - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, exactly the count on `master`. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. - `pnpm --filter @paperclipai/ui typecheck` — clean. - Server suites, 177 tests pass across 9 files: `workspace-runtime.test.ts`, `workspace-runtime-leases.test.ts`, `workspace-runtime-control-recovery.test.ts`, `execution-workspace-runtime-control-conflict.test.ts`, `execution-workspace-runtime-lease-route.test.ts`, `workspace-operations-reconciliation.test.ts`, `workspace-runtime-start-terminality.test.ts`, `app-hmr-port.test.ts`, and `workspace-runtime-ready-comment.test.ts`. - Exposure unit suites, 77 tests pass: `src/services/runtime-exposure/` and `workspace-runtime-exposure-backfill.test.ts`. - UI: `WorkspaceRuntimeControls.test.tsx` and `WorkspaceServiceControlBar.test.tsx` — 34 tests pass. **One suite is red on the development host and is expected to be green in CI.** `server/src/services/workspace-runtime-exposure.test.ts` has 10 failures on the machine used to write this branch. The cause is host contamination, not the code. That machine already runs an HTTPS canary that holds ports 42000, 42001, 52000, and 52001 on a tailnet address. The suite allocates from the same range, so the new listener-ownership check correctly reports: ``` listener_ownership_mismatch — port 42000 is bound to 100.123.243.20, 127.0.0.1, fd7a:115c:a1e0:0:0:0:dd3a:f314 ... instead of loopback only ``` A CI runner has no listener on those ports, so the check sees loopback only and the suite passes. Please confirm this from the CI result on this pull request rather than from a local run on a host that already exposes a managed runtime. This is a real weakness of the current test fixture, and the third pull request in the series removes it by allocating the pair through a central mediator instead of a stubbed availability check. `workspace-runtime-https-live-exercise.test.ts` needs a live `tailscale` host and was not run locally. ## Risks - This is the behavior-bearing pull request of the three, so it carries the most risk. - Two new `409` responses appear on managed control routes. A caller that assumed a control always starts must handle a conflict. Board and operator actions are deliberately exempt, so an agent lease cannot lock an operator out. - Managed worktree runtimes now default to `tailscale_https`. If the host has no working broker, the start fails closed and reports the exposure failure instead of silently serving plain HTTP. This is intended, and it is the reason the failure message names the denying predicate. - Startup reconciliation touches persisted runtime rows. It is scoped to desired state and does not resurrect a service that never came up. - The lease has a 30-minute time to live and explicit release paths, so a crashed owner cannot hold a lane forever. - No migration runs in this pull request. The tables and columns land in #11524. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass, with the one host-contaminated suite explained above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.817.0-canary.9 |
||
|
|
bfc19e2ebd |
chore(lockfile): refresh pnpm-lock.yaml (#11531)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com>canary/v2026.817.0-canary.8 |
||
|
|
f4802b1bbc |
feat(runtime-exposure): least-privilege Tailscale HTTPS broker, shared contract, and persisted exposure state (#11524)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Paperclip starts and supervises managed runtime services for a project's execution workspaces, so an agent's branch can be previewed while it works > - Those services only listen on plain loopback HTTP. A person on another device, or on a phone, cannot open the preview > - A Tailscale HTTPS mapping solves this, but `tailscale serve` needs host privileges that the Paperclip server process must not hold > - This pull request adds the foundation only: a separate least-privilege host broker, the shared exposure contract, and the database columns that hold exposure state > - Nothing calls the broker yet, so there is no behavior change. The benefit is that the privileged surface is small, reviewable, and isolated before any lifecycle code depends on it ## Linked Issues or Issue Description No public GitHub issue exists. The change follows the feature request template. **Subsystem affected** Managed workspace runtime services, the shared type and validator package, and the database schema. **Problem or motivation** A managed runtime service binds to loopback only. There is no supported way to reach that preview from another device. Adding HTTPS directly to the server would mean the server process runs `tailscale serve`, which needs privileges far wider than the task requires. A compromised or buggy server could then map any port to the tailnet. **Proposed solution** Split the privileged work into a separate broker process with a narrow protocol, and define one shared contract that the server, the UI, the runtime, and the broker all read. Land this foundation first, with no caller, so the privileged code can be reviewed on its own. **Alternatives considered** - Call `tailscale serve` from the server process. This was rejected because it gives the server unrestricted mapping authority. - Use `sudo` for single `tailscale` commands. This was rejected because the argument list is the only guard, and it is easy to widen by accident. - Use a generic reverse proxy. This was rejected because it does not remove the need for a privileged Tailscale mapping step. **Roadmap alignment** This supports the existing managed workspace runtime capability. It adds no new product surface on its own. **Additional context** The broker is the security boundary of the feature, so it is deliberately the first slice. Three later pull requests build on it: the server exposure lifecycle, the runtime lease and recovery integration, and the leased-port mediator. ## What Changed - Add the `@paperclipai/tailscale-https-broker` workspace package. The broker listens on a unix socket, authorizes each peer with `SO_PEERCRED`, and answers a small request protocol. - Restrict what the broker will map. It accepts only same-number HTTPS-to-loopback pairs inside the Paperclip port range, refuses protected ports, and confirms that the loopback port belongs to a Paperclip-owned listener. - Parse every request with a strict JSON reader that rejects duplicate keys, prototype keys, and unknown fields. - Write an append-only audit record for each broker decision. - Add the shared exposure contract in `@paperclipai/shared`: the `RuntimeExposureConfig`, `RuntimeExposureState`, and `RuntimeExposureStatus` types, their zod validators, the app and HMR port rules, and the loopback-bind helpers. - Persist exposure state on `workspace_runtime_services` with the new `exposure` column, plus the server-private `exposure_handle` and `backend_url` columns that are never serialized to API clients. - Add the `execution_workspace_runtime_leases` table that the later lease slice uses. - Extend the runtime read-model test fixture for the three new columns. ## Verification Focused checks, all run on this branch: - `pnpm --filter @paperclipai/tailscale-https-broker test` — 12 files, 82 tests pass. This covers peer credentials, port policy, protected ports, the serve config writer, the strict JSON reader, argv parsing, and the socket server. - `pnpm --filter @paperclipai/tailscale-https-broker typecheck` — clean. - `npx vitest run --root packages/shared src/runtime-exposure src/validators/runtime-exposure.test.ts` — 3 files, 40 tests pass. - `pnpm --filter @paperclipai/db typecheck` — runs `check:migrations` first. Migration numbering and migration safety both pass. - `pnpm --filter @paperclipai/shared typecheck` — clean. - `pnpm --filter @paperclipai/ui typecheck` — clean. - `npx vitest run --root server src/services/workspace-runtime-read-model.test.ts` — 3 tests pass. - `npx tsc --noEmit -p server/tsconfig.json` — 139 errors, which is exactly the count on `master` before this branch. All 139 come from the unbuilt `@paperclipai/plugin-sdk` package. To confirm the exposure state is inert, start a managed runtime service as usual. The new columns stay null and the service behaves as it does today. ## Risks - Migration risk is low. Both migrations only add a table and three nullable columns. No column is backfilled and no existing column changes. The migration safety check passes. - Behavior risk is low. No code path calls the broker in this pull request, and the shared exposure fields are optional. - The broker is privileged, so it is the real risk surface. It is mitigated by peer-credential authorization, a fixed port range, a protected-port deny list, same-number pair enforcement, listener-ownership checks, strict JSON parsing, and an audit trail. Reviewers should read `packages/tailscale-https-broker/src/authorization.ts` and `src/port-policy.ts` closely. - The broker requires a `tailscale` version floor, which its README records. An older host CLI makes the broker refuse to start rather than map incorrectly. - `pnpm-lock.yaml` changes because a new workspace package is added. The diff is the new importer block, plus one duplicate `tinyexec` entry that pnpm removed. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. ## Model Used Claude Opus 5 (`claude-opus-5`), 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.817.0-canary.7 |
||
|
|
55464204a6 |
test(ui): wait for conditions in the last four fixed-turn loops (#11523)
The remaining instances of the pattern #11499 and #11521 replaced in the routing tests, found by grepping `attempt < N` across the suite. Budgets of 20, 25 and 30 turns rather than 3 and 5, which is why they surfaced far less often - SkillStudio was among the failures seen while verifying the earlier PRs. Three of the four were reimplementations of `vi.waitFor` down to rethrowing the last error, differing only in bounding on turns rather than on time. The fourth was the text variant. Two `flushReact` helpers became dead with the loops that used them and are removed. Verified on the mechanism, since budgets this large pass until the machine is loaded and so prove nothing by passing: a throwaway probe drove a 30-turn loop and `vi.waitFor` against a value landing at turn 60. The loop throws, `vi.waitFor` reaches it. Not a lateral move between arbitrary bounds. This closes one spelling of the pattern, not the class, and the sweep that found these was too narrow. Two other shapes do the same thing and a grep for `attempt < N` cannot see either: a fixed-cycle helper, `flushReact(cycles = 4)` in AgentToolsTab.test.tsx, and fixed-duration sleeps in AgentToolsTab, CompanyContext, AgentConfigForm.render, Artifacts, Search and ImportFromVaultDialog. `AgentToolsTab > autosaves installed apps for the current agent` failed one of three full-suite runs here, holds both shapes, and is untouched by this change. Refs #11484. ui typecheck clean; the four files pass; full suite passes two of three runs, the third failing only on that pre-existing instance. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.6 |
||
|
|
65907aa41f |
test(ui): wait for the route, not three turns, in the cases-routing test (#11521)
The instance #11499 named but did not include. `App.cases-routing.test.tsx` carried the identical fixed-turn loop that PR replaced in its sibling `App.activity-routing.test.tsx` - three macrotasks instead of five, otherwise the same helper - so it fails the same way when the suite runs many workers in parallel and the container has not filled yet. It was one of the failures observed while verifying #11499. Same one-line replacement: `vi.waitFor` retries against a time budget, so a loaded worker gets more turns rather than a failure. The two helpers are identical again. Verified on the mechanism rather than on a green run, because the old loop passes in isolation too - that is what made this a flake and not a failure. A throwaway probe drove both helpers against a container whose text lands after ten macrotasks: the three-turn loop throws, `vi.waitFor` resolves. That is the condition a loaded CI worker creates. The probe was deleted rather than committed; it tests a test helper and had one question to answer. This closes one named instance, not the class. The full ui suite passed three consecutive times with no failure in any file, but the other instances seen during #11499's verification - TaskChatComposer, RequestCollapsedSidebar - simply did not recur, so they are rarer rather than fixed. Refs #11484. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e07d605dfc |
test(ui): wait for conditions, not durations, in three flaky tests (#11499)
Three tests yielded a fixed number of macrotasks before asserting - five in one case, one in another - which is ample on an idle machine and not when the suite runs many workers in parallel. The container was still empty, or the state had not landed, and the assertion failed on behaviour that works. `vi.waitFor` retries against a time budget instead, so a loaded worker gets more turns rather than a failure. `DocumentAnnotationPopover` is a different race and is fixed differently. The popover element is in the DOM as soon as React commits, while the effect that registers the document-level keydown and pointerdown listeners runs afterwards. A test dispatching in that gap loses the event outright, and a lost event cannot be recovered by retrying an assertion - so the render is wrapped in `act` to flush passive effects, and the waits only cover the smaller race that remains. Refs #11484. Verified stable over six consecutive runs of the three files, and the full ui suite passes. Other instances of the same class remain: the full suite still shows an occasional failure in a different unrelated test on each run. `App.cases-routing.test.tsx:104-108` is the clearest one - the identical fixed-turn loop this PR replaced in its sibling `App.activity-routing.test.tsx`, three turns instead of five - and takes the same one-line fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.5 |
||
|
|
870c305410 |
refactor(ui): drop the TZ pin now the fixtures are anchored (#11508)
#11480 pinned `TZ: "UTC"` in the vitest config because several suites asserted
local-time renders from UTC instants. That made the suite green everywhere, but
by suppressing the variable rather than fixing what depended on it: afterwards
no test could observe non-UTC behaviour, and a fixture quietly regaining a
local-time dependency would not be caught.
#11478 anchored those fixtures to the clock under test, which is the real fix,
so the pin now carries only its cost. Removed.
The order was load-bearing and is now satisfied. Measured on master before
#11478 landed, removing the pin failed four date-dependent tests at UTC+9 and
ten at UTC+12 - IssueProperties, IssueThreadInteractionCard, SummarySlotCard
and attention, all of which that PR anchors. Re-measured on master at
canary/v2026.817.0-canary.4
|
||
|
|
6d0adbfb5d |
refactor(ui): retire the shared-cache reasoning the account key made obsolete (#11507)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Three places in the UI turn the company list into an authorization verdict: the invite landing page, the onboarding draft gate, and company auto-selection > - Each grew a defense when one `["companies"]` cache entry answered for every account, and each documented that hazard at length > - #11488 keyed the entry by account, so the hazard those comments describe can no longer happen > - The comments stayed, and a comment that describes a trap that no longer exists is how the next reader removes a mechanism that is still holding something up > - This pull request replaces that reasoning with what the mechanisms actually do now, and removes the one condition that genuinely went dead > - The benefit is that the next person to simplify these gates has accurate reasons to work from ## Linked Issues or Issue Description No public issue exists. Follow-up to #11488, #11430 and #11417. The problem follows. **What happened?** The three gates were written against a shared, account-less company cache. #11488 keyed that entry by account, which made the documented hazard impossible — but the documentation stayed. Each gate now carries a long explanation of a cross-account leak that the key prevents, while the mechanism it explains is in fact still required for a different and unrelated reason. That is a maintenance hazard in a specific direction: a reader who checks the comment against the code concludes the mechanism is obsolete, removes it, and reintroduces a failure the comment never mentioned. **Expected behavior** The reasoning next to each gate describes why the gate is there now. **Steps to reproduce** Read the comment above `ownershipDecidable` in `OnboardingWizard.tsx` against `master`. It justifies `isSuccess` on the grounds that "after an account switch the retained value is the previous account's list", which the account-keyed entry makes impossible. **Paperclip version or commit** `master` at `0817fbad9`. ## What Changed - `ui/src/pages/InviteLanding.tsx` — dropped the `Boolean(sessionQuery.data)` conjunct from `membershipListIsCurrent`; rewrote the comment. - `ui/src/components/OnboardingWizard.tsx` — replaced the shared-cache explanation above the ownership gate with the reason the gate still exists. - `ui/src/hooks/useSignOut.ts` — corrected the sweep's rationale, which cited the company list as its example of data the next account could read. ### The one dead condition `membershipListIsCurrent` tested `Boolean(sessionQuery.data) && companiesQuery.isFetchedAfterMount`. The first term cannot be false when the second is true: the query is `enabled` only while a session exists, so the flag cannot be set without one. The lapsed-session case it looked like it covered is covered by the keying instead — the observer re-keys to the anonymous entry and holds no data to leak. Tests pass with it removed, but that only shows no test distinguishes it, which is why the reasoning above is recorded in the code rather than left for the next reader to redo. ### What is deliberately kept Each gate turned out to be load-bearing for a reason that has nothing to do with accounts: - **InviteLanding** still waits for a list fetched this mount. A pending query reads as an empty list, which reads as "not a member", which auto-accepts an invite the customer may already hold. - **OnboardingWizard** still forces a fetch with `staleTime: 0`. A cached list is the right account's but can be thirty seconds old, so a company created moments ago in another tab is missing from it — and missing reads as "you do not own this", which *deletes* the draft rather than withholding it. - **CompanyProvider** still clears the live selection on an account change. That is component state and does not change key with the query. Removing them as redundant is the mistake the stale comments invited; this change is what makes that argument harder to make by accident. ## Verification - `pnpm tsc -b` in `ui`: clean. - `InviteLanding.test.tsx`, `OnboardingWizard.test.tsx`, `CompanyContext.test.tsx`, `useSignOut.test.tsx`, `companies-query.test.ts`: **67 passed**, run twice. No behaviour change is claimed and none is intended: the only non-comment edit is the removal of a condition that cannot alter the expression's value. **Not done:** no browser run. Nothing here is observable at runtime. ## Risks Low. Comments, plus one condition shown to be unreachable-false. The risk that remains is a documentation risk in the other direction: if the keying is ever reverted or bypassed, these comments will understate what the gates protect against. They name #11488 so that connection is findable. **This does not close the class.** Account-scoped entries other than the list — `["companies", id]`, stats, and the rest — still survive an account change that skips the sign-out button. That is unclaimed work, and larger than this. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell for typecheck and test runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.3 |
||
|
|
40e7add71c |
test(ui): make the suite independent of the machine timezone (#11478)
Several suites asserted local-time renders from UTC instants, or built date fixtures from the real clock, so they only held where local time happened to match. CI runs in UTC and never reported it; a contributor anywhere else saw failures on a clean checkout. Seven files, anchored to the clock under test rather than to the machine's. The set grew twice while being fixed: the two tests the issue named surfaced three more at UTC+9, and those surfaced two more at UTC+14. `StatusCards/format` is the interesting one. `rollupUpdatesToday` filters on the UTC calendar day to match the server token cap, while the fixtures were built on the local day. West of UTC that lands `iso(0)` in the previous UTC day for the stretch between UTC midnight and local midnight - about seven hours a day at UTC-7 - and east of UTC+12 "today at local noon" is already yesterday in UTC outright. Either way the rows it means to count drop out. A run crossing midnight UTC splits the same way. Fixes #11476. Deliberately left: IssueProperties.test.tsx:1515-1517 still pin the minute of three timestamps against a UTC fixture. They pass at every offset tried, including UTC+5:45, and the minute there is load-bearing - it distinguishes Created from Started from Completed - so it wants more care than mechanical anchoring. Full ui suite 4113 pass. The TZ pin added by #11480 is still in place here and is now redundant; #11508 removes it, stacked on this branch so it cannot land without the anchoring it depends on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.2 |
||
|
|
0817fbad92 |
fix(ui): scope invite membership checks to the signed-in account (#11417)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Which companies a person belongs to is an authorization fact the server owns, and the UI caches the answer under a single `["companies"]` key > - That cache entry carries no account identity, and `main.tsx` sets `staleTime: 30_000` for every query, so for thirty seconds after a sign-in the previous account's list is served with no request at all > - The invite landing page reads that list to decide whether the person is already a member of the inviting company > - A list that arrives with no loading state and no error therefore looks authoritative while describing somebody else > - This pull request makes the page trust only a list it fetched itself, for the account signed in now > - The benefit is that a membership decision stops depending on cache freshness, which nothing in the app guarantees ## Linked Issues or Issue Description No public issue exists. Refs #11380, #11382. The problem follows. **What happened?** `InviteLanding` read the shared `["companies"]` cache entry as proof of membership in two places: - The post-sign-in redirect called `fetchQuery(companiesListQueryOptions)`, which returns the cached entry without a request while it is inside the app-wide `staleTime`. - An effect cleared the pending invite token whenever the cached list contained the invited company. Neither checked that the list belonged to the account signed in now. A second account signing in on a warm tab, or a session that lapses server-side, is enough to reach both. **Expected behavior** The page decides membership from a company list fetched for the current session. **Steps to reproduce** 1. On a self-hosted instance in `authenticated` mode, sign in as account A, which belongs to company X. 2. Within thirty seconds, open an invite link for company X and sign in as account B, which does not belong to it. 3. The page reads A's cached list, finds company X, and treats B as already a member. **Paperclip version or commit** `master` at `2a4b4bc63`. ## What Changed - `ui/src/pages/InviteLanding.tsx` — the membership query sets `staleTime: 0` so it revalidates on mount, and the verdict is withheld until that fetch lands, keyed on `isFetchedAfterMount`. The token-clearing effect and the "already a member" branch both read through that gate. - `ui/src/pages/InviteLanding.tsx` — the post-sign-in path cancels anything still in flight for the previous session, then forces a fetch for the new one with `staleTime: 0`. - `ui/src/pages/Auth.tsx` — sign-in resets the companies query instead of invalidating it. Invalidation leaves the previous account's list readable, and its fetch running, until the refetch returns. - `ui/src/pages/InviteLanding.test.tsx` — coverage for the warm-cache case, the token-clearing effect, and the `local_trusted` exemption. ### `local_trusted` is exempt Those instances have no accounts, so the shared list is the only identity there is. `membershipIsAccountScoped` is false there and the gate stays open. ### Rebased onto the account-keyed cache #11488 landed while this was open and keys the company list by account, so the page can no longer reach another account's list at all. Two things changed here as a result: - The post-sign-in read now calls `fetchCompanyListForCurrentAccount`, which replaces the `cancelQueries` plus forced `fetchQuery` this PR originally carried. The helper is strictly stronger: it detaches the in-flight `/companies` request inside the query function, and it resolves the account identity past the session invalidation immediately above rather than trusting the session entry still in the cache. - The observer reads through `useCompanyListQuery`. The mount-scoped `isFetchedAfterMount` gate is **kept**, not removed. Its purpose has narrowed — cross-account leakage is now structurally impossible, so what remains is holding the verdict until this page has a list rather than acting on a pending one. It is still load-bearing: disabling it fails two tests here. Removing a defense in the same change that rebases onto a new foundation is the wrong order; that is a follow-up once the keying has proven itself. `Auth.tsx` can safely reset, because it navigates away on success and `InviteLanding` mounts fresh afterward. Measurements of exactly when that rewind does and does not bite are in [#11380](https://github.com/paperclipai/paperclip/pull/11380#issuecomment-5300984911). ## Verification - `InviteLanding.test.tsx`, `Auth.test.tsx`, `companies-query.test.ts`, `CompanyContext.test.tsx` together: **51 passed**, run twice. - `pnpm tsc -b`: clean. - `InviteLanding.test.tsx` and `Auth.test.tsx` together: 21 passed. Both failures are pre-existing and unrelated. Each reproduces on a tree that does not contain this change, in files this change does not touch: | Failure | Why it fails | | --- | --- | | `IssueProperties.test.tsx` | Timezone-dependent: expects `4:08 PM`, gets `9:08 AM` | | `StatusCards/format.test.ts` | Time-of-day dependent: "only counts updates started today" breaks near midnight | **Not done:** no manual two-account run in a browser. The path needs two accounts on an `authenticated` instance, which a local dev instance cannot exercise. ## Risks Low. The failure direction is a membership verdict withheld for one extra round trip, which resolves itself; the direction it removes is one account's membership granted to another, which does not. It adds one request per invite-page mount, on a query key the app already uses. **This does not close the class.** The shared list is still unscoped for every other consumer. #11380 clears it on sign-out and #11382 handles the onboarding draft gate; all three are needed, because an account can change without passing through any one of those paths. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell for typecheck and test runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.1 |
||
|
|
327f59cac2 |
refactor(ui): key the company list by account (#11488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Which companies a person belongs to is an authorization fact the server owns, and the UI caches the answer for speed > - It cached that answer under one `["companies"]` key with no account attached, while `main.tsx` sets `staleTime: 30_000` for every query > - So for thirty seconds after an account change, one person's list answered questions asked about another, arriving with no loading state and no error > - Three separate consumers each grew their own defense against this, and each was a place to forget one > - This pull request keys the entry by account, so a list belonging to someone else is not distrusted but unreachable > - The benefit is that the protection stops depending on every future consumer remembering to defend itself ## Linked Issues or Issue Description No public issue exists. Refs #11380, #11382, #11417, #11430. The problem follows. **What happened?** The company list lived in a single cache entry, `["companies"]`, carrying no record of which account it was fetched for. Combined with the app-wide 30s `staleTime`, any read within that window after an account change returned the previous account's list — from cache, with no request, no loading state and no error. Every consumer that treats the list as an authorization fact had to know this and defend itself: - `InviteLanding` reads it to decide whether you already belong to the inviting company (#11417). - `OnboardingWizard` reads it to decide whether a saved draft belongs to you (#11382, merged). - `CompanyProvider` reads it to pick and persist your active company (#11430). All three defenses are correct. The problem is structural: the fourth consumer has to invent a fourth one. **Expected behavior** A cached company list can only answer questions about the account it was fetched for. **Steps to reproduce** 1. On a self-hosted instance in `authenticated` mode, sign in as account A, which belongs to company X. 2. Within thirty seconds, have account B become the session in that tab — a second tab signing in, or A's session lapsing server-side. 3. Any consumer reading the company list receives A's list, and nothing in the query result indicates it is not B's. **Paperclip version or commit** `master` at `ac91b7f3b`, which includes #11430. ## What Changed - `ui/src/lib/queryKeys.ts` — `companies.list(userId)` replaces `companies.all` as the list's entry. `companies.all` remains the prefix, so it still matches for invalidation. - `ui/src/api/companies-query.ts` — `companyListQueryOptions(userId)` builds the keyed options; `useCompanyListQuery()` is the only observer entry point and holds until the session settles, because the key cannot be built before then; `fetchCompanyListForCurrentAccount(queryClient)` covers imperative paths; `useAccountIdentity()` exposes the session identity the key is built from. - `ui/src/api/companies-query.ts` — the `/companies` detach moved into the query function. - `ui/src/context/CompanyContext.tsx` — drops the session-watching refetch machinery the key now makes unnecessary (`removeQueries`, the explicit replacement fetch, the awaiting gate). It still clears the live selection on an account change, because that is component state and does not change key with the query. - `ui/src/pages/InviteLanding.tsx`, `ui/src/components/OnboardingWizard.tsx` — read through the account-aware API. - Tests — the account-keyed guarantee, prefix invalidation still reaching the list, the detach inside the query function, and updates where suites seeded the old shared key. ### The existing defenses are deliberately left in place The per-consumer gates in #11382, #11417 and #11430 are now belt and braces. They are also what will catch this refactor if it is wrong somewhere, so removing them in the same change that moves the foundation would be the wrong order. Simplifying them is a follow-up, once this has proven itself. ### Why `retry: 1` appears in CompanyProvider An earlier measurement on #11430 found a retry on the replacement fetch changed no outcome, because `removeQueries` made the observer rebind and issue a second request for free. Keying by account removes that mechanism and the free attempt with it. The retry now carries the property the incidental refetch used to — a single blip during an account change should not leave the customer with no companies until they find "Try again". #11430's test for that property is unchanged and still passes, which is how the gap was caught. ### A regression this went through, kept for the record Gating the query on the session settling meant that while the account was unknown the query was *disabled*, and a disabled query reports `isLoading: false` with no data — which the provider defaults to an empty list and reads as "asked, and owns nothing". That is the destructive branch #11477 had just fixed, reached through a different door: it would have cleared the customer's stored company on every cold boot. #11477's test caught it during the rebase. `useCompanyListQuery` now reports the wait for the account as part of the wait for the list. ### What this does not do It does not scope the rest of the per-account cache. `["companies", id]`, stats, and every other account-scoped entry still survive an account change; that is the cache-lifetime work in #11380. ## Verification - `pnpm vitest run` in `ui`: **4018 passed, 1 failed**. - `pnpm tsc -b` in `ui`: clean. - `companies-query.test.ts`: 6 passed. `CompanyContext.test.tsx`: 17 passed. `OnboardingWizard.test.tsx`: 13 passed. `InviteLanding.test.tsx`: 13 passed. The failure is the pre-existing timezone-dependent `IssueProperties.test.tsx`, fixed by #11478. Two behaviours are asserted rather than assumed, because the refactor is only safe if they hold: that invalidating the `companies` prefix still marks the account-keyed list stale (19 call sites depend on it), and that the query function detaches the in-flight `/companies` request before fetching. **Not done:** no manual two-account run in a browser. The path needs two accounts on an `authenticated` instance, which a local dev instance cannot exercise. ## Risks Moderate, and worth reading before approving. **It touches `InviteLanding.tsx`, which #11417 also modifies**, so one of the two will need a rebase — the conflict is mechanical (both change how the same query is read). This was #11481, which GitHub closed automatically when its base branch (#11430's) was deleted on merge; reopening a pull request whose base branch is gone is not permitted, so it continues here against `master` with the same head and the same review already recorded on #11481. **The list now waits for the session query.** The key cannot be built before the account is known. In the app the session is already fetched at boot by many components, so this is a dependency rather than an extra request, but it does serialize: on a cold boot the list waits for the session to land. Every test that renders a company-list consumer now needs a session in the cache, which is why several suites gained a seed. **A missing mock surfaces as a passing gate rather than an error.** The detach inside the query function meant suites whose `companiesApi` mock lacked `detachInflightList` had their query function throw, which read as "decided" in the onboarding gate and mounted the wizard early. Fixed in the affected suites; worth knowing as a failure mode. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell for typecheck and test runs. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>canary/v2026.817.0-canary.0 |
||
|
|
fd472d02ba |
Show ordered live blocker work in task chat (#11487)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The task view shows operators why work cannot continue > - The redesigned task thread now shows direct and ultimate blockers > - But it does not show the ordered task queue while a blocker chain has live work > - This pull request adds a compact ordered live-work queue to the redesigned thread > - The benefit is that operators can see completed, running, and queued dependencies without opening the larger legacy notice ## Linked Issues or Issue Description **What existing behavior does this improve?** The redesigned task thread blocker summary is improved. The merged predecessor is #11456. **Current behavior** The redesigned task thread shows compact direct and ultimate blocker links. It does not show the ordered queue when the blocker tree has live work. The legacy task view shows this queue in a larger notice. **Proposed behavior** Show a compact blue live-work queue at both ends of the redesigned task thread. Order completed tasks first, then running tasks, then queued tasks. Show a live terminal leaf as `Now running`. Return to the amber blocker links when no live dependency remains. **Reason and benefit** Operators can see the active dependency order without leaving the redesigned task view. The compact presentation preserves the new thread's low-chrome layout. **Breaking changes** None. The change only adds UI for blocker data that the task view already receives. ## What Changed - Shared the live blocker ordering helper between the legacy notice and the redesigned task thread. - Added compact ordered dependency links at the top and bottom of the redesigned thread. - Added a separate `Now running` link for a live terminal blocker leaf. - Preserved the compact amber blocker rows when live work is not present. - Added component tests and a Storybook state for the new presentation. ## Verification - `pnpm exec vitest run ui/src/components/TaskChatThread.test.tsx ui/src/components/IssueBlockedNotice.test.tsx` - `pnpm check:token-gates` - `pnpm -r typecheck` - `pnpm test:run` - `pnpm build` - `pnpm --filter @paperclipai/ui build-storybook` - Captured and reviewed the new Storybook state in a headless browser. ## Risks - Low risk. The queue appears only for blocked tasks whose blocker attention state is `covered` and whose dependency set contains live work. - The API does not provide an explicit queue position. The UI preserves the existing legacy ordering rule: completed, running, queued, then numeric task identifier. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used - OpenAI Codex, GPT-5. The runtime does not expose the exact snapshot or context-window size. Reasoning, code execution, repository tools, and browser automation were enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.816.0-canary.8 |
||
|
|
ac91b7f3b2 |
fix(ui): scope the company selection to the signed-in account (#11430)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every company-scoped screen reads the active company from `CompanyProvider`, which picks one from the `["companies"]` list and remembers it in localStorage > - That cache entry is shared app-wide and carries no account identity, so it survives a change of account in the tab > - The provider therefore auto-selects from whatever list is cached, which can belong to the account that just went away > - This pull request makes the provider watch the account and refuse to derive a selection from a list fetched for a different one > - The benefit is that the app stops pointing at a company the signed-in account may not be able to see ## Linked Issues or Issue Description No public issue exists. Refs #11380, #11382, #11417. The problem follows. **What happened?** `CompanyProvider` auto-selects a company from the shared `["companies"]` cache entry and writes that id to `localStorage`. Nothing ties that entry to an account. When the account changes in the tab, the previous account's list is still served, so the provider can select — and persist — a company belonging to the account that just went away. Company-scoped screens then render against a company the current account may not be able to see. Signing in through `Auth.tsx` invalidates the entry, so the in-app sign-in path is covered. Two paths are not: a session that lapses server-side, and a second account signing in on another tab. The sign-out sweep in #11380 does not cover them either, because neither presses the sign-out button. **Expected behavior** The company selection is derived only from a company list fetched for the account that is signed in now. **Steps to reproduce** 1. On a self-hosted instance in `authenticated` mode, sign in as account A, which belongs to company X. 2. In a second tab, sign in as account B, which does not belong to company X. 3. Return to the first tab. The session query refetches and reports account B, while the company list is still account A's. 4. The provider keeps company X selected and leaves its id in `localStorage`. **Paperclip version or commit** `master` at `6542ad1f4`. ## What Changed - `ui/src/context/CompanyContext.tsx` — the provider observes `queryKeys.auth.session`. On a change of session user it clears the live selection, removes the shared company list, and holds auto-select until a list fetched for the new account lands. The stored id is left alone on purpose: `resolveBootstrapCompanySelection` re-validates it, so an account signing back in keeps its company while an unrelated account cannot inherit it. - `ui/src/context/CompanyContext.tsx` — an errored list is treated as undecided rather than as "no companies". With `retry: false` a single network blip sticks, and the empty-list branch read it as proof the account owns nothing and cleared the stored selection. - `ui/src/api/client.ts` — new `detachInflightGet(path)`. GET coalescing keys on the request path alone, so a `/companies` request issued under the previous session could be joined by the replacement fetch and answer it with the previous account's companies. Detaching leaves that request to settle for its own callers and makes the next call issue a fresh one. - `ui/src/api/companies.ts` — `companiesApi.detachInflightList()` wraps that for the list path. - `ui/src/context/CompanyContext.tsx` — `companyListUnavailable` separates "no usable list because a request failed" from "this account owns nothing", and `retryCompanies` gives a recovery action that fetches. Both are derived from the query rather than tracked beside it; a second copy of "did the last attempt succeed" drifted out of step during review, reporting a failure over a later empty list that was simply the truth. - `ui/src/components/SidebarCompanyMenu.tsx` — renders "Couldn't load companies" and a Try again item in place of "No companies", which is a claim about the account that a failed request cannot support. This is the menu `Sidebar` mounts, so it is the only place a customer can act on the failure. - `ui/src/components/CompanySwitcher.tsx` — the same treatment. The application does not render this component (its only mount is a Storybook story), so it is kept in step rather than relied on. - `ui/src/context/CompanyContext.test.tsx`, `ui/src/components/SidebarCompanyMenu.test.tsx`, `ui/src/api/client.test.ts` — coverage for the account switch, a same-account re-observation not churning, the detached GET, the failed replacement and its recovery, a single blip self-healing, unavailability not outliving the failure, and the sidebar rendering the recovery action for a failure but plain "No companies" for an account that owns nothing. ### No `retry` override on the replacement fetch The obvious fix for a failed replacement is a retry, and it is not load-bearing here. A transient failure already gets a second attempt: the observer rebinds to a fresh query on the render those state updates schedule, and issues its own request — measured as two attempts with or without the option. Retries would only add failed round trips before a real outage is reported, and the outage is what needs a way out, which is what `companyListUnavailable` and `retryCompanies` provide. ### Why `removeQueries` here, and why that does not generalise Removal notifies no observer. What rebinds them at this call site is the render the surrounding state updates schedule; every observer re-binds to a fresh query on the next render. A caller without that guarantee would leave mounted observers serving the previous account's value, so this is not a pattern to lift elsewhere — the sign-out sweep in #11380 must use `resetQueries` instead, and its measurements are at [#11380](https://github.com/paperclipai/paperclip/pull/11380#issuecomment-5300984911). The inverse caveat holds for a local reset under an observer that stays mounted, which is why #11417 and #11382 avoid `resetQueries`. ## Verification - `pnpm vitest run` in `ui`: **4014 passed, 1 failed**. - `pnpm tsc -b` in `ui`: clean. - `CompanyContext.test.tsx`: 16 passed. `SidebarCompanyMenu.test.tsx`: 15 passed. `client.test.ts`: 9 passed. The failure is pre-existing and unrelated: `IssueProperties.test.tsx` expects `4:08 PM` and gets `9:08 AM`, a timezone-dependent assertion. It reproduces on a tree without this change, and #11478 fixes it. Each new test was confirmed to fail against the implementation it covers, by reverting that change and re-running rather than by assuming. The account-switch test fails without the fix (the selection stays on the previous account's company and no refetch is issued); the flag-clearing test fails without its clause (an empty list keeps reading as "couldn't load"). **Not done:** no manual two-account run in a browser. The path needs two accounts on an `authenticated` instance, which a local dev instance cannot exercise. ## Risks Low. The failure direction is a company selection withheld for one extra round trip, which resolves when the list arrives. The direction it removes is one account's company selected and persisted for another. It adds one company-list request per account change, on a query key the app already uses. It adds no request at boot: the session query it observes is already fetched app-wide. **This does not close the class.** Company-scoped entries other than the list — `["companies", id]`, stats, and the rest of the per-account cache — still survive an account change. That is the cache-lifetime work in #11380, not this provider's. ## Model Used Claude Opus 5 (`claude-opus-5`), through Claude Code. Extended thinking enabled. Tool use enabled: file read and edit, shell for typecheck and test runs, and a scratch vitest harness to measure `removeQueries` and `resetQueries` notification behaviour against the installed `@tanstack/query-core` 5.101.4. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>canary/v2026.816.0-canary.7 |
||
|
|
10d0555189 |
fix(interactions): authorize resolvers consistently (#11376)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Issue interactions give agents and people a structured decision record. > - Resolver routes used different authorization rules. > - Some routes blocked valid agents, including task watchdogs with normal issue access. > - The API did not show who could resolve a pending interaction. > - This pull request gives every interaction kind one resolver policy evaluator. > - The benefit is a clear decision path with consistent governance and company isolation. ## Linked Issues or Issue Description Fixes: #8087 Refs: #7403 Related PR: #11082 proposes board-only confirmation rules. This change keeps human-only review as an explicit policy. **What happened?** Agents could create issue interactions. Some resolver routes still required board access. This left valid agent confirmations pending. Task watchdogs could see the same problem without board identity. **Expected behavior** Every interaction kind must use one resolver policy contract. The contract must support `anyone`, `not_creator`, and `human_only`. It must also apply all normal governance controls. **Steps to reproduce** 1. Create a `request_confirmation` interaction as an agent. 2. Resolve it with another authorized agent. 3. Observe the board-only denial. **Paperclip version or commit** The problem exists on `master` before this change. **Deployment mode** Local development with `pnpm dev`. ## What Changed - Add canonical policies for `anyone`, `not_creator`, and `human_only`. - Use one server evaluator for every interaction kind. - Apply named addressees, company limits, review rules, and task watchdog scope. - Charge cross-issue resolutions to the existing per-run action limit. - Return the effective resolver audience in attention and interaction data. - Show the audience, governance choices, and denial reasons in the board UI. - Add telemetry, API documents, product documents, and regression fixtures. - Add migration provenance for safe legacy behavior. - Make migration `0218` safe for complete replays and partial prior runs. ## Product Rules - An interaction records a response. It does not grant authority for the next action. - `anyone` lets any authorized issue participant respond. - `not_creator` requires a responder other than the interaction creator. - `human_only` requires an authorized person. - A named addressee, company policy, or governed action can narrow the audience. - These controls cannot widen the audience. - A task watchdog uses the same rules as an ordinary agent. - A task watchdog does not receive board authority. - An agent resolution on another issue uses the shared cross-issue action limit. - Legacy pending interactions keep their earlier restrictions. - The UI shows the effective audience and a permanent denial reason. ## Verification - `pnpm --filter @paperclipai/db check:migrations` - `pnpm --filter @paperclipai/db typecheck` - `pnpm exec vitest run packages/db/src/issue-thread-interaction-resolver-policy-migration.test.ts` - The focused PostgreSQL test applies migration `0218` twice. - The test also completes a partial prior run and preserves existing provenance. - The latest GitHub head has 29 successful checks. - The opt-in Storybook visual check skipped as expected. - Greptile reports 5/5 with no open comments. ## Risks - New interaction writes use `anyone` by default. - Callers must select `not_creator` or `human_only` when they need stricter review. - Legacy pending interactions keep the old creator and human restrictions. - Migration `0218` fills only missing provenance fields during recovery. - Cross-issue resolutions can reach the existing action limit. - The shared evaluator affects every interaction kind. - Route, service, database, shared contract, and UI tests cover these rules. > This work matches the Agent Reviews and Approvals direction in `ROADMAP.md`. It does not duplicate a planned item. ## Model Used OpenAI Codex, GPT-5. The runtime does not expose the exact deployment ID or context window. The agent used reasoning, repository tools, shell commands, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked public issues or described the issue with the required labels - [x] I have not referenced internal Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open comments - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d6acb48551 |
feat(ui): show a calm in-flight notice when a live run is on the issue (#11423)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - When an agent run ends without a recorded disposition, Paperclip raises a "missing disposition" handoff so the work does not stall silently > - The issue page shows that handoff as an amber alarm: "This task still needs a next step." > - The server tells the UI whether the issue has a live continuation, and an earlier change used that flag to hide the alarm while a correction run is active > - Hiding it removed the false alarm but replaced it with nothing, so a reader cannot tell "nothing is wrong" from "nothing is tracked" > - This pull request puts a quiet informational line where the alarm was, and links the live run > - The benefit is that the page stays honest in both states: it is calm while an agent works, and it is loud only when the issue is really stuck ## Linked Issues or Issue Description No public GitHub issue exists for this gap. Description follows `.github/ISSUE_TEMPLATE/enhancement.yml`: **What existing behavior does this improve?** The missing-disposition handoff notice on the issue page. It is the amber banner that reads "This task still needs a next step." **Subsystem affected** Web UI — `ui/src/components/IssueBlockedNotice.tsx`. **Current behavior** An issue with an outstanding missing-disposition handoff shows nothing at all in the blocked-notice slot while a correction run is live. `IssueBlockedNotice` calls `isSuccessfulRunHandoffRequired()`. That helper returns `false` when `successfulRunHandoff.hasLiveContinuation` is set. The component then renders no handoff content. Two tests asserted the empty render. **Proposed behavior** The page states, quietly, that a correction run is in progress. It also states that the alarm returns if the run stops without choosing a next step. The reader can open the live run from that line. The amber alarm does not change when no run is live. **Reason and benefit** Silence and "healthy" look the same. A user who saw the alarm earlier cannot tell whether the handoff was resolved, whether the alert was withdrawn, or whether an agent is working on it now. One muted line removes that ambiguity. It also keeps the loud state meaningful, because the alarm now appears only when the issue is really stuck. **Breaking changes** None. The change is presentational and adds no API or data-shape change. ## What Changed - Added `SuccessfulRunHandoffInFlightNotice` to `ui/src/components/IssueBlockedNotice.tsx`. It renders a muted row with a pulsing live dot and this copy: "A correction run is in progress — the agent is working. This alert returns if the run stops without choosing a next step." - The notice links the live run when the server sends `liveRunId` and the handoff has an `assigneeAgentId`. It shows the short run id as plain text when no agent id is available, and it shows no run reference when `liveRunId` is absent. - Liveness reads either the server `hasLiveContinuation` flag or the fresher client `liveIssueIds` set. This matches the rule that already suppressed the alarm. - The amber alarm is unchanged when no live continuation exists. The unpromoted scheduled-retry carve-out still shows the alarm, so the "Retry now" control stays reachable. - The calm line also renders above the blocker notice when an issue has blockers and a live run at the same time. - Storybook: added `InFlightNotice` and `LivenessComparison` stories to `ui/storybook/stories/successful-run-handoff.stories.tsx`, and removed a duplicated panel from the overview story. - Tests: the two cases that asserted an empty render now assert the calm line. New cases cover a missing `liveRunId`, a missing agent id, a handoff that is not required, and the two "alarm is unchanged" guards. ## Verification Run the component and helper suites from `ui/`: ``` cd ui && NODE_ENV=test npx vitest run \ src/components/IssueBlockedNotice.test.tsx \ src/components/IssueChatThread.test.tsx \ src/components/IssueChatThreadSystemNotice.test.tsx \ src/lib/successful-run-handoff.test.ts ``` Result: 4 files, 111 tests, all pass. Also run: - `cd ui && npx tsc -b --force` — clean. - `node scripts/check-token-gates.mjs` — all gates clean. Manual check in Storybook (`pnpm --dir ui storybook`), story `Paperclip/Successful Run Handoff → Liveness Comparison`: - The alarm panel keeps its 4 remediation bullets, its amber surface, and its run chips. - The calm panel shows one 39 px muted row, no bullets, and a working link to the live run. - Measured contrast of the calm text against its rendered surface: 4.58:1 in light mode and 6.52:1 in dark mode. Both pass WCAG AA for normal text. ## Risks Low risk. The change is limited to one presentational component and its stories. It adds a render path where the component previously returned `null`, so a surface that expected an empty render now shows one muted row. No server, API, or data-shape change. The amber alarm path and the scheduled-retry carve-out are covered by tests that assert the calm line does not appear. ## Model Used Claude Opus 5 (Anthropic), model id `claude-opus-5[1m]`, 1M context window, extended thinking, with tool use and code execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: ClaudeCoder <claudecoder@paperclip.ing> Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.816.0-canary.6 |