mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 02:07:25 +08:00
master
383
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6f2ce27ca7 |
fix(workspaces): prepare checkouts without a local seed config (#14810)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Task preparation can create an isolated Git worktree and run its setup script. > - The Paperclip repository setup script also prepares a seeded development instance. > - A server configured through environment variables can have no local seed config. > - This stops ordinary task preparation before the agent starts. > - This pull request prepares checkout dependencies when no seed source exists, while preserving errors for invalid sources and existing development instances. > - Tasks can start without creating or claiming a seeded development runtime. ## Linked Issues or Issue Description **What happened?** A task with the Paperclip repository fails during setup when the host has no repository-local or default instance config. The automatic worktree provisioner requires a seed source even when the task only needs the checkout. **Expected behavior** A plain checkout should prepare its dependencies without a local development database. A missing custom source, invalid source path, or existing development instance with a missing source should still fail. Starting a seeded runtime must still require a valid source. **Steps to reproduce** 1. Run an environment-configured Paperclip server without a local instance config. 2. Add the Paperclip repository to a project. 3. Start a task that uses an isolated Git worktree without a custom provision command. 4. Observe the setup error before agent execution. **Paperclip version or commit** Reproduced against `0d3e7bf6ac` with a real script subprocess and workspace realization regression. **Deployment mode** Environment-configured server with external PostgreSQL. **Additional context** Searched open and closed GitHub PRs and issues. Related work: Refs #14795 (seed-source diagnostics) and Refs #11733 (source validation). This change keeps source validation and seed-readiness checks in place. ## What Changed - Permit dependency setup when the default seed config is absent (including the Docker image config path) and the worktree has no development-instance state. - Keep missing custom configs, invalid paths, and lost sources for existing instances as errors. - Create no config, environment file, or seed manifest for a plain checkout. - Keep dependency install failures visible and allow normal instance setup once a source becomes available. - Cover the setup script, seed-runtime refusal, and automatic server worktree realization. - Document the difference between checkout preparation and seeded-runtime readiness. ## Verification - Regression tests failed before the fix for absent-source checkout preparation and dependency setup. - `bash -n scripts/provision-worktree.sh` - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs` — 34 passed; 1 existing flock-dependent test skipped on macOS. - Server regression — 2 passed, covering an unset config and the Docker image default path. - `pnpm build` — passed. - `pnpm -r typecheck` — passed. - All CI checks passed, including the full test shards, build, typecheck, browser tests, and canary dry run. - The first local `pnpm test:run` encountered two chat-test failures because skill discovery selected an unrelated parent directory. Both tests pass at the PR commit in a clean temporary checkout. The full local run was not completed; the redundant clean run was stopped after the complete CI suite passed. - `git diff --check` and added-line secrets/PII scan passed. - Greptile: 5/5, no comments. The branch has no merge conflicts. - No live tenant deployment or task retry was performed. ## Risks - A new checkout with no implicit seed config now completes dependency setup. It has no seeded development instance. A runtime request still fails until a valid source exists. - Existing instances and custom source paths retain their failure behavior. The script does not synthesize a source from environment credentials or copy a live database. - No schema, API, or task-setting changes. Revert the commit to restore the previous setup behavior. ## Model Used OpenAI Codex (GPT-6), with tool-assisted analysis, code edits, and local tests. The runtime did not expose a verified model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4ac374103f |
fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.
## Linked Issues or Issue Description
Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.
**What happened?**
Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.
**Expected behavior**
Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.
**Steps to reproduce**
1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.
**Paperclip version or commit**
Reproduced from
|
||
|
|
0dc8d80eea |
Clarify worktree seed source setup failures (#14795)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Managed worktrees can prepare an isolated Paperclip development instance. > - The built-in provisioner requires a canonical registered seed source config. > - A missing source currently has the same message as a rejected symlink or non-regular file. > - This pull request separates those messages and names the supported setup choices. > - Operators can choose the intended setup without changing the source-validation guards. ## Linked Issues or Issue Description **What happened?** A plain repository checkout can select the control-plane instance as its seed source. If that instance runs with environment-only configuration, the source config file can be unavailable. The provisioner stops with a message that also covers noncanonical files and gives no repair guidance. **Expected behavior** The error should identify the selected source and distinguish an unavailable prerequisite from a rejected file. It should explain that a seeded development instance needs a canonical registered source. It should describe the explicit no-op only for a checkout-only worktree. **Steps to reproduce** 1. Use a plain base checkout with no repository-local config. 2. Leave the control-plane instance config file absent. 3. Run the built-in worktree provisioner against an isolated checkout. **Paperclip version or commit** Base commit `c8f874311c`. **Deployment mode** Managed local worktrees, including servers configured only through environment variables. Related: #11733 adds deeper source-readiness checks. #11735 changes runtime and seed lifecycle handling. This change only improves the existing shell guard's diagnostics. ## What Changed - Distinguish unavailable source configs from symlinks and non-regular files. - Identify whether the selected source belongs to the base workspace or control-plane instance. - Explain seeded-instance prerequisites and the explicit checkout-only setup choice. - Verify failure still precedes target-state creation and CLI invocation. - Document the setup choice and its runtime-readiness limit. ## Verification - `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs`: 21 passed; one platform-gated test skipped because macOS lacks `flock`. - `bash -n scripts/provision-worktree.sh` and `git diff --check`: passed. - `pnpm -r typecheck`: passed. - `pnpm exec vitest run server/src/__tests__/ai-connections.test.ts`: 50 passed after running the installed PostgreSQL package's own symlink hydration script in this worktree. - `pnpm test:run`: attempted, then stopped after unrelated database suites failed at startup. The offline install had omitted PostgreSQL native library symlinks. The focused database rerun above verifies the local repair; the complete suite is delegated to CI. - `pnpm build`: passed. - [Required PR CI](https://github.com/paperclipai/paperclip/actions/runs/36797650741) passed on `0a3ba63e12`: 50 successful checks and two intentional Storybook skips. Greptile scored that exact commit 5/5; there are zero unresolved review threads and no merge conflicts. ## Risks - Diagnostics only. This does not supply a source config or repair an existing blocked task. - The failure predicates and exit status stay unchanged. Symlink and non-regular-file errors do not recommend skipping setup. - The checkout-only no-op requires an explicit policy choice. It does not grant runtime or seed readiness. - No schema, migration, tenant policy, deployment, or Sentry reporting change. ## Model Used OpenAI GPT-6-based Codex, with reasoning, shell tools, and code execution. The exact serving model ID and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
33f2b3a159 |
fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Connectors catalog lets people give agents tools or connect agents to conversations. > - GitHub put these two uses behind one card and an extra choice. > - People should choose the connection they need from the catalog. > - This pull request keeps GitHub for tools and adds GitHub Code Review Bot as a separate card. > - Each card opens its setup directly. Both use the existing connection code. ## Linked Issues or Issue Description **What existing behavior does this improve?** GitHub connector discovery and setup. **Current behavior** With chat connectors enabled, GitHub opens a menu that asks whether to use tools or create a bot. Saved tools and bots share the same catalog entry. **Proposed behavior** GitHub opens tool account access. GitHub Code Review Bot opens agent selection. Saved bots and drafts appear under the bot card. Chat-disabled instances show only GitHub tools. **Reason and benefit** The catalog names the two uses and removes an extra setup choice. The bot keeps the existing GitHub provider, credentials, endpoint IDs, setup steps, and runtime. **Additional context** Related work: https://github.com/paperclipai/paperclip/pull/12843 and https://github.com/paperclipai/paperclip/pull/14594 established GitHub account identity. This change preserves that tool flow. No duplicate catalog split was found. ## What Changed - Split the generated app definitions into GitHub tools and GitHub Code Review Bot. Reuse the existing GitHub logo and channel method. - Open bot setup directly, including old resume and reconnect links. - Put existing bot endpoints and drafts under the bot card. Hide duplicate internal chat applications. - Keep pasted GitHub URLs mapped to the tool connection. - Add seven Storybook states for the catalog, saved connections, disabled chat, both setup paths, mobile, and light mode. - Fix narrow-screen bot rows so the label cannot overlap status and setup actions. - Update catalog, route, browser, and API tests, plus the GitHub connector guide. ## Verification - [Hosted Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog): seven states built from this branch. The deployment passed its public-file verification. - All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`. Two optional Storybook jobs skip under their normal trigger rules; the manual Storybook deployment passes. The branch has no merge conflicts. - Greptile: 5/5 on the current head, with no review comments or unresolved threads. - `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm build-storybook` passed. The final Storybook fixture also passed UI typecheck and the hosted build. - Targeted catalog, URL matching, routing, grouping, brand, and chat UI contract tests passed. - GitHub provider browser tests: 2 passed. These cover direct tool setup and the bot setup and management lifecycle with provider responses mocked. - Embedded-browser test on an isolated local instance: opened both cards, selected an agent, saved a bot draft, and resumed the same endpoint under the bot card after a reload. - Storybook Tool Setup and Bot Setup assertions pass in the published preview. Chat Disabled assertions pass locally. Inspected mobile and light mode, including the draft-row layout and official GitHub marks. - Local full-suite limitation: `pnpm test:run` was not clean. A cross-company route assertion failed in the aggregate run and passed in isolation; a workspace-runtime test reached its 30-second hook timeout. Some isolated database reruns skipped when the embedded-PostgreSQL availability probe failed. The local aggregate was stopped after CI completed. The corresponding full CI suites pass all 360 tool-access tests and all 162 workspace-runtime tests. - No live GitHub authorization or installation was performed. The isolated instance correctly stopped at the cloud enrollment or public HTTPS prerequisites. ## Risks - Low scope: catalog presentation and routing change. There is no database migration or provider credential change. - Existing GitHub bot URLs now open bot setup directly. The tool route remains `/apps/connect?source=github`. - The bot remains behind the existing chat-connectors feature flag. Existing endpoints retain `provider: github`. - Channel applications are represented by endpoint rows. Regression tests cover legacy bot applications, tools, active bots, and drafts together. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, and embedded-browser tools. The exact deployed model ID, context window size, and reasoning setting are not exposed to this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ad55d0a281 |
fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents use Apps through a gateway that checks identity, company access, and action policies. > - Personal pasted credentials can point to company secrets. Setup can show success while the gateway rejects every call. > - Several OAuth methods also omit the scopes needed for their supported write actions. > - This pull request gives setup, health checks, and invocation the same credential rules. Owners repair existing connections by reconnecting. > - New connections request reviewed permissions for their supported actions. Read-only choices remain available under Advanced. > - Agents can use the connections people give them, while existing consent, identity boundaries, and action restrictions remain enforced. ## Linked Issues or Issue Description Refs #14009 and #14008. This addresses the personal-credential defect. The separate GitHub organization-identity selection defect is outside this change. Related work: #13942 fixed part of new personal-key setup. #14200 independently fixes legacy personal reconnect and protects managed-agent profile credentials during removal. This PR covers that ownership invariant across key and secret-URL setup, reconnect, health, discovery, and invocation, and keeps owner reconnect as the repair path. #14059 tracks requested versus provider-asserted OAuth scopes; it remains separate work. I searched open PRs and issues for Zapier, Airtable scopes, connector writes, and personal credential failures. **What happened?** A Zapier secret URL saved through personal setup can become a company secret referenced by a user grant. Health checks bypass the gateway's ownership check, so the connection appears healthy but calls fail with `grant_credential_invalid`. Custom-header paths can also receive a duplicate `credentials.` prefix. Omitted OAuth scopes make write access depend on provider defaults. **Expected behavior** Personal invocation credentials belong to the selected user. Setup, health, and actual calls enforce the same rule. New connections request documented permissions for supported read and write actions. Existing tokens gain no permissions without provider consent. **Steps to reproduce** 1. Connect Zapier or a generic secret URL with the personal identity. 2. Allow an agent to use the connection and complete setup. 3. Invoke a tool through a run-scoped gateway. The legacy layout fails ownership validation despite successful setup. **Paperclip version or commit** The implementation started from `44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master at `94e8dec56`. **Deployment mode** Built from source. Regression tests use isolated PostgreSQL fixtures and controlled MCP transports. ## What Changed - Share credential writing, ownership validation, and canonical paths across initial setup, resume, reconnect, rotation, health, discovery, and gateway calls. Keep OAuth client-registration secrets separate from invocation credentials. - Existing personal connections with company-scoped credentials require owner reconnect with a fresh key or secret URL. Reconnect creates a correctly owned value and updates the existing grant and declarations. There is no automatic ownership backfill or new startup hook. - Preserve PostgreSQL timestamp precision when reconnect checks whether a grant changed. Previously, converting the timestamp to a JavaScript Date could reject reconnect with a false concurrent-change error. - Protect credentials used by other grants, connections, bindings, managed-agent profiles, routine triggers, or secret proposals from connection removal. - Review all 117 tool methods, including 84 OAuth methods. Record explicit scopes or documented provider-default exceptions with official evidence. Add Airtable's seven scopes, Hugging Face repository/job scopes, and other documented MCP permissions. - Prefer available write-capable methods. Put explicit read-only choices under Advanced. Explain pasted-key permissions and offer reconnect for missing OAuth consent. Preserve existing grants, policies, Google availability gates, and curated scope allowlists. - Reconnect generic secret URLs and custom headers using their stored credential fields. Refresh the catalog after setup, correct reconnect feedback and error guidance, and let Cancel exit invalid setup while Save & exit retains draft-saving behavior. - Apply ownership checks to the new GitHub repository/skill connection picker. Align the permission audit with the Google scope reductions merged on master. - Add run-scoped gateway, ownership, owner-reconnect, OAuth URL, insufficient-scope, UI, and catalog-wide regression coverage. Update the connector playbook and permission audit. ## Verification Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all CI/status gates (55 completed check runs, no failures or pending checks) and has a completed Greptile review at **5/5 with no outstanding findings**. GitHub reports the PR as mergeable/CLEAN. - **Embedded browser:** used the actual server and built UI from this worktree, a fresh isolated database, and local HTTP MCP fixtures. Completed personal bearer-key, secret-URL, and custom-header setup; reproduced the legacy ownership failure; reconnected through the owner’s form; and completed writes afterward. Read-back was verified for bearer-key and secret-URL connections. Public organization-wide setup appeared immediately in Browse without reload. Zapier URL validation/Cancel and Google’s enrollment gate were also exercised. - **Persistence and invocation:** verified user ownership, canonical `credentials.authorization` / `remote.url` / `headers.X-Api-Key` declarations, and unchanged connection/grant identity. The old company secrets retain their ownership. Separate HTTP calls through an actual run-scoped gateway session completed a write and read-back. - **Backend coverage:** the final gateway suite passes all 82 cases, including catalog Zapier and generic inline reconnect. It checks company/user isolation, canonical declarations, same-endpoint URL validation, fresh credentials, retained restrictions, and real gateway read/write execution using fixture transport. A timestamp with PostgreSQL microseconds covers the former false reconnect conflict. - **Local checks:** 368 catalog, gateway, repository, and UI tests passed before the final extra Zapier case; 49 GitHub skill access tests also passed. All three Apps browser regressions pass, including reconnect through the actual form and catalog visibility without reload. Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final patch, and token gates passed. Full tool-access service runs hit varying 15-second Google fixture timeouts; both affected cases and the updated reconnect assertion pass in isolation (3 tests). The complete test matrix passes in CI on this head. - **Verification limits:** no live provider account was available for Zapier/Airtable/OAuth consent or account-bound write proof. Public metadata and local fixtures do not establish provider consent. The original development database clone failed on a pre-existing missing `tool_connections_transport_check` constraint; browser acceptance used a fresh isolated database created by the normal CLI onboarding flow. ## Risks - Existing broken personal connections stay unusable until their owner reconnects. Health, discovery, and invocation return an actionable ownership error; startup does not rewrite credential ownership. - Scope changes affect new authorization requests. Providers may still require resource selection, account roles, paid plans, or app verification. Existing consent and action restrictions remain unchanged. - Shared credentials are retained rather than reassigned or revoked. Provider-default exceptions and unavailable live checks are documented in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`. - No new endpoint, database table, lockfile change, or CI workflow change is included. ## Model Used OpenAI Codex, based on GPT-6, with reasoning, code editing, shell execution, web research, and browser tools. The exact deployment model ID and context window were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a36cbffa9e |
fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path > - Paperclip manages agents and the services they need for work. > - Agents use connection search to discover a setup path before they request access. > - Tool-only filtering hid channel and AI methods from this search. > - Requiring every query word to match rejected normal task descriptions. > - A small aggregator index also omitted supported apps such as Circleback. > - This pull request broadens retrieval and returns purpose-specific setup guidance. > - Agents can choose a relevant result while existing access and provider-choice checks still apply. ## Linked Issues or Issue Description **What happened?** A search such as “AgentMail create an email address and manage an agent mailbox” returned no usable result. “Help me find tools for circle back” also missed Composio's supported Circleback toolkit. Queries longer than 200 characters failed validation. **Expected behavior** Return useful native and verified aggregator matches from natural-language queries. Include channel/email methods when Chat connectors is enabled. Identify each method's purpose and the correct setup path. **Steps to reproduce** Use the queries above with `connections_search` from an active task. Enable Chat connectors for the AgentMail case. The regression suite reproduces these misses before the change. Related routing work: #13941. This change does not change the runner failure path or add channel setup to tool-only connection cards. ## What Changed - Rank name and capability matches. Accept extra words, split names, small spelling errors, and queries up to 4,000 characters. - Include tool, channel/email, and AI methods. Return company-prefix setup links for channel and AI flows. - Add a dated snapshot of 1,583 official Composio toolkit names and a refresh script. Merge duplicate MCP variants for search and link each support claim to official evidence. - Find authorized indexed aggregator namespaces within longer queries. Return multiple app matches when the agent needs to choose. - Prefer exact app names over fuzzy matches for other apps; retain existing AI readiness. - Preserve native preference, company and identity boundaries, administrative denials, and saved provider consent. - Add relevance and database regressions, extend native tool-authority coverage, and document search behavior. ## Verification - Red: 15 new assertions failed against the previous implementation; the existing baseline passed. Added red-green regressions for Motion versus fuzzy Notion and existing AI access during review. A further regression covers mixed ready/unconfigured AI results and their per-result setup guidance. - Green: all 103 tests in the eight focused shared, database, runtime-tool, fixture, and route suites pass on the latest commit. - `pnpm -r typecheck` and `pnpm build` passed. - Latest-commit CI passed: 54 successful checks and two skipped checks, including the full test matrix, browser E2E, typecheck, build, and canary dry run. - The long local `pnpm test:run` attempt began before the review fixes and retained transformed pre-fix search code; it also hit an unrelated timing failure. Fresh serial reruns of the affected search suites and three timeout cases passed all 124 tests. Parallel local route shards hit two additional database setup timeouts; both suites passed all 17 tests on a fresh serial rerun. The complete corresponding CI suites also passed. Duplicate broad local runs were stopped after CI completed. The local UI suite independently passed all 7,007 tests. - Greptile: 5/5 on `cc6a0180d`; all review findings resolved. - Browser verification passed in a disposable local instance through real process-agent search requests: AgentMail opened its setup flow with the requester selected; the saved Circleback choice produced the Composio setup card; a paragraph-length Notion query produced its setup card. No provider credentials or external accounts were created. - The browser test caught an invalid UUID-based setup URL. The fix uses the company prefix and has a regression assertion. - A 3,971-character catalog query found Circleback first in a local 10 ms spot check after sharing query preparation across the catalog scan. This is a single measurement, not a performance guarantee. ## Risks - Broader retrieval can return extra candidates. Named services rank first; agents must select the relevant method. - The public support snapshot can age. It proves catalog support, not account authorization or the availability of every requested action. - Channel and AI methods use existing setup links. The tool connection card still accepts tool methods only. - No schema, migration, credential, or runner lifecycle changes. ## Model Used OpenAI Codex (GPT-6). The session does not expose a more specific model identifier or context-window size. Used reasoning, repository search, code execution, tests, and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d72389bee2 |
feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Apps gateway gives agents governed access to external tools. > - Browser Use Cloud can run browser work, but a tool result alone does not let a person watch or take over. > - A task needs a durable browser session, a visible viewer, and recorded costs. > - This pull request adds a Browser Use Cloud v4 connection and interactive browser tabs on tasks. > - People can follow the work, interact with the page, and retain the browser after the agent finishes. ## Linked Issues or Issue Description **Problem or motivation** Agents need governed access to Browser Use Cloud. People need to see and interact with the same browser from the task. A browser must remain available after a run finishes and appear at the correct point in the task feed. **Proposed solution** Add a native REST connection for the v4 API. Bind each session to its company, task, agent, and credential grant. Open its interactive viewer in the task side panel. Record provider costs as financial events. Use `browser-use-cloud` as the app and connector key. Keep its skill with the connector and deliver it only with authorized connection tools. **Alternatives considered** A v3 MCP connection would expose tools without the v4 lifecycle integration. An external viewer link would leave the task. A fixed viewer size would prevent pages from responding to changes in the task pane. **Roadmap alignment** This extends the governed Apps gateway and Connected Apps roadmap. It uses the existing task, grant, secret, approval, and financial records. The work was requested by the maintainer. A search found no duplicate Browser Use connector PR or issue. ## What Changed - Add the Browser Use Cloud app, brand asset, API-key connection, and profile settings under the `browser-use-cloud` key. - Bundle the `browser-use-cloud` skill with the connector. Keep it out of global `skills/` discovery. Deliver it only with authorized task/run connection tools. Remove retired connector skill keys from runtime overlays and preserve unrelated browser skills. - Expose seven v4 tools through the governed gateway and deliver them to native and CLI agents. - Persist sessions, browsers, runs, event and recovery cursors, shutdown leases, and cumulative cost accounting. Recover uncertain paid starts without replaying them. - Enforce task ownership, credential grants, approvals, revoked access, and budget limits. - Add interactive task browser tabs and compact chronological feed entries. Retain the viewer across tab switches and keep visible idle browsers open. - Add debounced automatic viewport fitting, standard size presets, and a viewer ownership lease. - Add lifecycle, authorization, accounting, viewport, UI, and Storybook coverage. - Add an idempotent database migration after the current master migration. Preserve deployed migration hashes. Migrate pre-release Cloud connection and financial keys without replacing grants, credentials, or browser history. - Document provider behavior, live acceptance results, and the lack of documented passkey forwarding. ## Verification - Full workspace typecheck and production build pass on the updated branch. - Token gates, brand asset validation, module boundaries, and migration ordering pass. - Cloud tests verify global skill exclusion, authorized task/run delivery, unassigned agents, disabled connections, revocation, adapter isolation, and secret exclusion. The existing AgentMail connector assignment test also passes. - Migration replay runs twice against existing browser work and financial records. It preserves the records and avoids duplicate costs. - The focused provider, app catalog, OpenAPI, connection gateway, and migration regression suites pass. Recovery coverage includes lost replies, process crashes, provider rejection, and browser arrival acknowledgement. - All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`, including the full test matrix, browser E2E shards, build, typecheck, security, and release canary. Two optional Storybook jobs are skipped. - Greptile is 5/5 on the same commit, with zero unresolved review threads. The corrected review uses the actual master-to-head diff. - The local `pnpm test:run` started and was stopped after the full CI matrix passed. It did not complete locally; the full-suite result above comes from CI. - Earlier live acceptance used an isolated company with a capped provider credential. The agent opened paperclip.ing, the embedded viewer accepted navigation, and the same browser stayed available after completion and tab switches. - The local Storybook build passes. Stories cover the panel, footer, feed entries, settings, lifecycle failures, and viewport modes with an offline viewer fixture. ## Risks - Browser Use charges for hosted work. Provider caps and local budget checks reduce exposure; reported costs can arrive after work completes. - Viewer and CDP URLs grant access to the browser. The server validates and restricts them. They are excluded from agent results and durable event data. - Runtime resizing of v4 agent browsers uses a provider option confirmed by live testing but absent from its published agent schema. Resizing during a click may invalidate coordinates. Fixed presets remain available. - Viewport ownership is process-local and resets on restart. The lifecycle and accounting records remain in the database. - The original intermittent embedded-viewer stall has not been fully diagnosed. A bounded reconnect and active-session recovery cover the observed failure paths. - Live tests did not cover every revocation, approval, rate-limit, or restart case. Deterministic integration tests cover those paths. Passkey forwarding is not claimed. - Unknown create outcomes keep the credential available for cleanup. Run-list absence cannot prove a paid POST was rejected, so recovery stays pending until it can identify provider work. ## Model Used OpenAI Codex, GPT-6. Used reasoning, repository search, code execution, browser interaction, and test tools. The exact serving model ID and context-window size were not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3b4b270650 |
fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The shared ACP adapter engine records agent failures for operators. > - ACP providers can report a failure category, title, and detailed cause. > - Our patch kept only the category in the saved error, so an operator could not diagnose a failure when tracing was off. > - This pull request preserves redacted provider diagnostics in the run error, transcript, and structured run result. > - Operators can now inspect the provider message and any supplied request ID or stack trace after the run ends. ## Linked Issues or Issue Description Refs #13889 (the diagnostic gap; this PR does not update the bundled Claude version). Refs #14484 (related model-refusal classification; this PR retains diagnostics for all terminal failure categories). **What happened?** An ACP turn failed with only `ACP agent reported a terminal service failure.` The provider's title and details were available in memory but absent from the saved error and transcript. **Expected behavior** The run retains useful provider diagnostics even when raw tracing is disabled. Credentials remain redacted. A size limit must report truncation instead of silently removing the cause. **Steps to reproduce** 1. Run an ACP agent that returns an error-severity typed session failure. 2. Include an HTTP error, request ID, and stack text in its title and details. 3. Inspect the failed run with tracing disabled. Before this change, only the category survives. ## What Changed - Both pinned ACPX patches pass complete error text to the in-memory callback, so redaction happens before truncation. - The shared engine retains the sanitized category, title, and details in `resultJson.terminalSessionFailure` and includes the text in the run error and error transcript. - Diagnostics redact configured environment values even under arbitrary names, unknown launch-environment values, connection URL passwords, run credentials, and common credential syntax. Known boolean settings remain readable, while credential values are redacted even when embedded in other text. Diagnostics remove control characters and invalid Unicode. - Title and detail limits keep escaped transcript JSON below the server's chunk limit. Truncated fields include an omission count. The safe run-result projection preserves a byte-bounded diagnostic preview when the result exceeds its byte budget, with an explicit pointer to the full adapter-bounded run error and transcript. - The existing UI and CLI display the error. Diagnostics do not become assistant output. Issue continuation summaries and session-compaction prompts receive only the generic category, preventing provider text from becoming handoff instructions. Existing quota classification, warnings, timeout precedence, and control-channel failure precedence remain in place. - Regression tests cover real ACP child processes with both pinned versions in one-shot and persistent modes, credential redaction, request IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and database retrieval of oversized multibyte diagnostics. ## Verification - Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2 intentionally skipped, no pending or failing checks**. Includes typechecking, build, all Vitest shards, Runner checks, browser E2E, and the canary packaging/public-install dry run. - Greptile: **5/5** on this commit. Superagent security scan passes. All review threads are resolved. - Local verification passed: shared ACP engine suite (395 tests); real Claude ACP child-process and diagnostic regressions across both pinned runtimes and both execution modes; run retrieval and model-handoff regressions (59 tests); ACPX patch packaging (16 tests); full typecheck and build. Affected package typechecks and focused tests were rerun after review fixes. - The broad local `pnpm test:run` was stopped after review edits made its cached imports stale. Fresh targeted runs pass, including both affected server suites. Cold-build import failures were also rerun after dependency builds: chat integration (1,063 tests) and tool access (351 tests) pass. The final commit's complete CI matrix is green. ## Risks - Provider diagnostic text is untrusted. This change retains more of it in company-scoped run records. Redaction and size bounds apply before persistence. - Diagnostics are limited to fields the provider supplies. Old runs cannot recover discarded error text. - No schema migration, recovery-policy change, or new Telemetry or OpenTelemetry export. ## Model Used - OpenAI GPT-6 through Codex, with reasoning, repository inspection, code editing, and test execution. The exact serving model ID and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24beb00575 |
feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification. Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
0be2afcca6 |
feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path > - Paperclip lets operators assign tasks to AI agents and review their work. > - The task composer controls the next message and its assigned agent. > - Operators needed a way to choose that agent's model and effort without leaving the composer. > - The old mode selector, upload button, and input cards made the mobile composer crowded and hid normal messaging during a pending decision. > - Harnesses publish different model and effort capabilities, so the picker must follow the selected agent. > - This pull request adds one responsive composer flow, keeps pending cards visible above it, and protects Codex ACP authentication in the local test path. > - Operators can choose run settings, send a message, and answer a pending card as separate actions. ## Linked Issues or Issue Description **Subsystem affected** Task composer UI, issue thread interactions, Codex ACP credential handling, and Storybook. **Problem or motivation** The composer did not expose model or effort for the selected agent. Mobile actions wrapped poorly. Pending questions and confirmations replaced the composer. A local Codex ACP test could also reuse host authentication after the managed key was removed. **Proposed solution** Put assignee search, model search, exact model IDs, effort, and fast mode in one picker. Use a mobile dialog. Replace the direct-upload plus action and separate mode selector with an Add menu and removable Plan or Ask chips. Place pending interaction cards above the usable composer. Keep these cards pending after an ordinary message unless their creator asks for comment superseding. Replace managed ACP auth files atomically and isolate the test key from host credentials. **Roadmap alignment** ROADMAP.md does not list an overlapping composer milestone. This change improves the existing task and review flows. ## What Changed - Added the combined assignee, model, and effort picker to both task composers. Search matches agent name, role, and harness. The server uses a curated Codex list by default and honors instance-declared models. Manual IDs remain available. - Added an effort slider for known model capabilities, a conditional Codex fast control, and reset. The picker opens in a modal on mobile. - Added the Add menu for files, supported goals, Plan mode, and Ask mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling remains available. - Adjusted mobile spacing, avatars, wrapping, and Send placement. Removed the composer divider. - Moved pending question, confirmation, review, and related cards above the composer. Ordinary comments now leave question and confirmation cards pending by default. The onboarding prompt retains explicit comment superseding. - Updated the Storybook composer group with responsive states and the production picker. Added UI, service, route, and browser regression coverage. - Isolated Codex ACP API-key authentication, skipped subscription auth merge and shared-home copy-back for remote API-key runs, and replaced the managed auth file atomically. ## Verification - `pnpm -r typecheck` — passed on the final local head. - `pnpm check:token-gates` — passed on the final local head. - `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31 tests passed, including role and harness search, declared Codex models, and filtering general OpenAI models. - `pnpm exec vitest run server/src/__tests__/issue-thread-interactions-service.test.ts` — 74 tests passed. - `pnpm exec vitest run packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed, including remote API-key copy-back isolation. - `pnpm test:run` — attempted locally; the embedded PostgreSQL test database could not initialize on macOS. The isolated `heartbeat-run-event-sequencing` suite reproduced that environment failure. GitHub CI runs the full test matrix for this head. - `pnpm build` — passed on the final head. `pnpm build-storybook` passed after the last UI change; only server code, tests, and docs changed afterward. - Live local test drive — Codex ACP ran a task with a managed API key. The test agent was restored to its default ACP configuration afterward. - Review the interactive stories under the top-level Composer group with `pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask, picker, and pending-question states. ## Risks - A pending card stays open when an ordinary comment changes the discussion. Its creator can set `supersedeOnUserComment: true` when a new comment should replace it. - Model and effort overrides persist on the task until reset or changed. An unlisted manual model ID may fail when the provider runs it. - Some harness catalogs do not report effort support. The picker hides effort for those models. - No database migration is required. > For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and discuss it in `#dev` before opening the PR. Feature PRs that overlap with planned core work may need to be redirected — check the roadmap first. See `CONTRIBUTING.md`. ## Model Used OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID or context window to the task. The model used code execution and browser tools. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: OpenAI Codex <codex@openai.com> |
||
|
|
18e8c121d9 |
fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path > - Paperclip manages agents through a shared native runner. > - Built-in harness support should ship with Paperclip's public distribution. > - Grok already speaks ACP; it does not require a new public bridge package. > - Sandbox provisioning owns the native executable and its pinned version. > - The runner must verify that prerequisite without downloading it during npm installation. > - This change separates built-in launcher identity from external runtime identity. > - Clean npm installation and live staging checks verify the distribution boundary. ## Linked Issues or Issue Description Refs #13882, #13973, #13977, #13979. This follow-up now targets master after #13882 was squash-merged. It replaces the private `@paperclipai/grok-acp` workspace package with runner-owned assets. Current master is included so the branch also contains the merged scheduler, complete-event capture, and durable cleanup fixes. ## What Changed - Ship Grok launcher and qualification metadata inside the runner's compiled output and the public server's vendored runner tree. - Remove the separate Grok npm package and all package-manager install hooks for this runtime. - Require the checksum-verified Grok Build 1.0.13 binary at `/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution environment. Provision it explicitly in the Daytona image and CI setup. - Keep native binaries outside the provider pack. Bind the built-in launcher into the pack manifest. - Preserve executable leases, descriptor-backed startup, credential fences, permissions, and exact ACP model admission. - Use `builtin:grok-acp` and `native:grok` as profile identities. Historical package-profile sessions fail closed on resume rather than being silently reinterpreted. - Resolve built-in assets from the authenticated sidecar location, including public server npm layouts. Keep the controller path out of provider environments. - Add clean npm tarball installation verification to the existing trusted canary CI job and the admitted manual EC2 verification path. It stages a unified release version and runs npm lifecycle scripts, then verifies missing-prerequisite rejection and admission after separate provisioning without credentials or inference. - Include the controller-owned provider pack in stamped Cloud images. Unstamped local images omit the pack and remain usable; remote ACPX requires full source provenance. - Correct CLI approval-page metadata for an already authenticated Cloud board user; approval authorization remains unchanged. - Honor explicit native-runner enablement in the Cloud agent picker and direct setup page, keeping the flag disabled by default. - Allow selecting the execution environment before connecting credentials. Include Grok in the existing authenticated hello-probe flow, targeting its pinned native prerequisite for runner setup. - Recover an existing subscription sign-in conflict through an explicit cancel-and-retry action, serialized after cancellation succeeds. - Preserve the selected ACPX harness before normalizing config fields, so new Grok agents use the Grok default model. - Keep the credential-free Cloud provider pack root-owned and readable after runtime UID remapping; verify manifest and referenced asset access under an unrelated unprivileged UID during image builds. - Archive prior failover backups alongside explicitly replaced harness state, preserving evidence while preventing stale backups from blocking a fresh replacement. - Update Daytona image content inputs and contract tests for the built-in assets and explicit provisioner. - Document and regression-test the shared `approve-all` default for Grok setup, saved configuration, and native execution. Explicitly saved restrictions remain unchanged. ## Verification Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05` incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the base PR was squash-merged. All 12 conflicts came from incoming files identical to the tested pre-squash base. The final tree exactly matches a three-way merge using that original base, preserving built-in Grok distribution and removal of the obsolete private package. All 252 focused runner/UI tests, six npm-isolation tests, and token gates pass. Fresh exact-head Greptile review is 5/5 with no outstanding findings; security scans and EC2 native compilation pass. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([run 36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)). The repository owner explicitly authorized bypassing code-owner approval after all checks passed; no CI checks or repository protection settings are bypassed or changed. The only remaining PR was removed from the completed stack metadata to permit native auto-merge. Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1` includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28). Two master advances during verification overlapped the eval catalog; the final merge preserves Grok qualification, completion updates, and bounded API-response reading in all 348 cells. All 77 focused catalog/eval/workflow tests pass. Both native stack layers (#14397) are mergeable, and both exact-head Greptile reviews are 5/5 with successful security scans and no unresolved review threads. All current-head CI is green: 56 successful checks/statuses and four intentional skips ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)). The initial attempt lost two EC2 runners to shutdown signals and stalled a third shard during dependency preparation; all three passed the same-commit failed-job-only retry. Trunk code-owner requirements remain enforced. The review summary’s non-blocking saved-asset offset classification note concerns code already merged in #14301; those runtime files are identical to master and outside this stack’s diff. Historical live evidence below retains its original source revisions. [Final public npm verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764) passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public packages, an executed offline lifecycle sentinel, unchanged consumer lock, built-in launcher, missing-prerequisite rejection, and verified separately provisioned binary/command lease. Provisioning and cleanup require no host privilege elevation; only the positive probe mounts the temporary native binary read-only. The verifier is unchanged by the final master merge. All six isolation tests and an offline npm smoke test pass. The prior head had 56 green CI checks and a 5/5 review after two unchanged tests timed out and passed a failed-job-only retry ([CI attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)). All 56 recovery-display/lineage tests pass; re-review cleared the already-covered missed-retry concern. Earlier EC2 failures remain retained: [npm lockfile rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203), [missing compiler in the slim image](https://github.com/paperclipai/paperclip/actions/runs/36440210984), and the aggregate 15-minute test timeouts in those broad runs. Both broad attempts passed typecheck, token gates, Product E2E type/unit checks and build. The focused EC2 lane preserves the existing trusted-actor and immutable-source gates. Earlier documentation/test checkpoint `ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior unchanged. 154 focused tests pass across configuration building, native provider resolution, permission policy, credentials, UI configuration, and new-agent setup (including both Grok auth modes); token gates pass. All fresh CI is green for this head: 56 successful checks/statuses and two intentional skips ([run 36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)). Greptile is 5/5 with no new findings. Grok already inherits the shared `approve-all` default, so unattended setup requires no manual permission change. Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final staging continuation failure before provider startup: explicit replacement archived the old harness but left its failover backups active, which caused `runner_harness_state_mismatch`. The regression fails before the fix and passes after it; all eight adjacent recovery-safety cases also pass. Old backups remain inspectable inside the continuity archive. All fresh CI is green at this head ([run 36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)), with a 5/5 review. One unrelated Cursor test timed out in the initial server shard; the same-commit failed-job rerun passed, and both attempts are retained. Staging deployment is confirmed healthy on this revision. The controller image is `ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`. The final browser-created staging task passed on this exact revision with API authentication: context read → structured human question → controller restart → answer submission → same native provider session resumed → document saved → task Done. The two turns took approximately 119s and 77s. The actual write receipt was applied, and the saved document has exactly one revision containing the selected answer and requested marker. Usage and cost were not reported. [Controller image build](https://github.com/paperclipai/paperclip/actions/runs/36360889243). - Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`: all CI green (53 successful checks/statuses, two intentional skips), including repository typecheck/build/tests, native Runner tests, browser shards, and canary installation checks. [CI run 36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529). Greptile is 5/5 with no unresolved findings. - Focused checks cover Grok credentials, executable admission, launcher assets, provider-pack paths/permissions, workflow contracts, setup defaults, CLI authorization, and subscription conflict recovery. All 39 protocol definitions validate. Final integration checks pass 124 catalog/evidence/cache tests and nine project-form tests; token gates pass. Some local dependency checks could not load the stale installed dependency tree; the corresponding fresh EC2 checks pass. - Clean public npm installation passed on EC2 at `8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run 36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)): 17 unified-version packages, lifecycle scripts enabled, built-in launcher present, no separate Grok package or npm-downloaded binary, missing prerequisite rejected, separately provisioned native executable and command lease verified. No credentials or inference were used. Subsequent changes preserve this npm asset layout. - The immutable Daytona prerequisite image is `ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`, built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous Cloud controller image was `ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`, built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded by the latest image above. Its EC2 build verified provider-pack access under an unrelated unprivileged UID. - Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed full Grok onboarding with the correct `grok-4.7` model, saved credential delivery, and pinned Daytona execution. A browser-created task read context and asked the structured human question. After a controller restart, answering the persisted question resumed the same native provider session, saved the requested document, and completed the task. Actual tool outcomes and durable state agree: one question and one document revision. The two successful turns took 42.7s and 63.1s; usage and cost were not reported. - Restricted policy returned the expected `approval_required` outcome. Functional staging tests explicitly selected `approve-all`; controller authorization and governed approvals remain enforced. Temporary board CLI access was revoked and verified rejected (HTTP 401), and the disposable onboarding agent was paused. Failures remain retained: the pre-fix continuation failure (its task remains blocked; the passing final task is fresh), the original Cloud provider-pack permission failure, the expected restricted-policy denial, the superseded npm staging failure, and an earlier monolithic CI infrastructure timeout. Browser CI exposed a project alias/form race; the final stack uses master's stronger draft-preservation fix and all browser shards pass. Historical full subscription/API protocol and Product rosters retain their original source revisions and do not qualify this packaging revision. No local Docker or Rust build was used. ## Risks The branch includes master’s draft-preservation fix for project URL aliases. It keeps the same project’s edit form mounted and clears prior data when the project or company changes. Custom sandboxes and local execution hosts must provision the pinned binary before Grok starts. Missing, changed, unsupported-platform, and symlinked executables fail admission. The new builtin profile cannot resume sessions created with the former private-package profile. Existing Claude/Codex npm bridge profiles retain their package pins. Grok restricted modes preserve the selected policy but cannot automatically admit Paperclip calls: ACP permission metadata does not independently bind tool authority, so those calls stop with `approval_required`. New Grok configurations default to `approve-all`, including API configurations that omit the mode. Existing explicitly restricted configurations remain restricted; controller authorization and governed approvals remain enforced. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ee8f1fd6e |
ci: retire recurring public cloud image builds (#13827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Core publishes standard images and source verification for downstream services. > - Managed services can now compose private images from the signed standard image. > - Core still builds a second public cloud image on every master push and release. > - That duplicate producer consumes build capacity and retains an obsolete readiness contract. > - This pull request retires recurring cloud publication while preserving the standard producer and rollback artifacts. ## Linked Issues or Issue Description Refs #13797 and #13789. Related: #12856 changes image dependency packaging; it does not retire this producer. **What existing behavior does this improve?** Core's recurring Docker publication and Cloud readiness workflow. **Current behavior** Master pushes call the legacy cloud publisher from Cloud readiness. Release tags and manual Docker runs call it too. Canary promotion also requires the legacy image. **Proposed behavior** Publish standard Core images and retain `Cloud source verified v1`. Let downstream services build their managed image. Keep explicit commit previews and existing images available. ## What Changed - Remove `docker-cloud.yml`, its master and release callers, and its unused cache selector. - Remove the legacy image/migrator wait and `Cloud deployable v1` job. Keep the full source verification workflow and exact source-proof name. - Make canary promotion inspect and promote the standard image only. - Preserve signed standard-image publication, direct migrator publication, and explicit `release.yml` previews. The preview path still uses the Dockerfile `cloud` target. - Update workflow, preview, build-stamp, and packaging tests. Exercise the promotion shell with mocked registry commands, including missing-image and missing-tag cases. - Document frozen legacy aliases, consumer requirements, preview compatibility, and rollback retention. ## Verification - All 377 workflow tests pass: `node --test .github/scripts/tests/*.test.mjs`. - All 129 release-registry tests pass: `pnpm test:release-registry`. - Focused source-proof, standard-image, preview, and workflow tests pass: 256 tests. - Focused image packaging/build-stamp tests pass: 16 tests. - Actionlint passes on all three changed workflow files. `git diff --check` passes. - Full local `pnpm build` and `pnpm -r typecheck` pass. - The policy follow-up updates an old assertion that required the removed readiness job. All 37 source-proof/release-workflow tests pass locally. - Full local `pnpm test:run` did not complete successfully while the Mac ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of generated Cargo output from this isolated worktree with `cargo clean`. GitHub CI passed on the final head: 52 successful checks and 2 optional skips. - Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`: **5/5**, successful current-head check, zero review threads. - September 23 refresh: the unchanged PR head merges cleanly with current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an isolated temporary worktree, all 377 workflow tests and 29 release/preview tests pass on the combined tree. `git diff --cached --check` passes. - Refreshed Actionlint workflow validation passes with ShellCheck disabled. Full Actionlint reports the same 10 existing ShellCheck diagnostics as master, with no added diagnostics. No source changes or new PR commits were needed. - The full local build/typecheck and current-head Linux CI results above remain the verification for the unchanged PR head. They were not rerun for this metadata-only refresh. No image publication or tenant deployment was initiated for this refresh. ## Risks **Deployment prerequisite satisfied (September 23):** The combined cleanup release is deployed to staging and production, and production Support is verified. Active managed-fleet automation uses standard-image composition. Explicit immutable previews remain supported by the retained preview publisher. This PR is ready for maintainer review; keep auto-merge disabled and wait for explicit merge authorization. - A consumer still selecting `Cloud deployable v1` will stop advancing at the last legacy-ready commit. Confirm active automatic consumers use the standard-image composition contract before merge. - Legacy cloud release-channel aliases stop advancing. Standard self-hosted aliases continue. - This PR deletes no registry images, cache tags, migrators, credentials, or runner infrastructure. Existing immutable releases remain usable for rollback. - Explicit legacy previews remain for commit-specific operator deployments. Retiring that compatibility path requires a separate consumer migration. - These changes affect CI publication, not database schema or application behavior. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used repository inspection, reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db8f8fe5b7 |
fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path > - Paperclip manages agent work through shared runner contracts. > - Direct protocol evals qualify provider behavior against a mock control plane. > - Grok supports API keys and company subscription credentials. > - The hosted protocol workflow selected an API key for every Grok cell. > - Product subscription support did not enable subscription protocol runs. > - This change adds explicit subscription selection and checks the recorded authentication mode. ## Linked Issues or Issue Description Refs #13878, #13882, #12618. The direct Grok protocol roster cannot run with subscription authentication through the trusted default-branch workflow. Add an explicit selector while keeping API-key dispatches compatible. Keep the actor allowlist, protected environment, immutable source revisions, and publication gates. ## What Changed - Add `grok_authentication` with `api_key` and `subscription` choices. Keep `api_key` as the compatibility default. - Deliver the protected `GROK_AUTH_JSON` secret only to a subscription-selected Grok cell. Do not provide an API key to that cell. - Read authentication mode from the pinned eval program's actual roster summary. Retain it in the cell, catalog, campaign roster, and result. - Reject missing or mismatched authentication evidence during aggregation. Preserve cell metadata and an allowlisted failure reason before failing a cell, so malformed evidence cannot hide the retained attempt. - Document credential setup, source separation, and temporary-secret cleanup. ## Verification - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed. - Validated all 39 Grok cells at eval revision `3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests only the subscription credential. Validation made zero provider calls. - Negative coverage rejects invalid selectors and missing or API authentication evidence in an otherwise passing subscription attempt. - `git diff --check` passed. All 53 current-head checks passed; the unchanged callback-drain timing test passed its bounded rerun, and the failed attempt is retained. Greptile reviewed `ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining findings. - No Docker or broad builds ran on the developer machine. CI performs repository checks on the configured fleet. ## Risks Grok runs require an eval revision that records `authenticationMode` in the roster summary. Missing evidence fails closed. The credential contains account access and refresh tokens; an owner must approve its delivery to the protected environment before a live run. The change adds no PR trigger or authorization bypass. Live subscription protocol qualification remains pending this workflow reaching master. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
be6f49a425 |
feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path > - Paperclip runs agents through local adapters and the native runner. > - Both paths must use the same installed provider CLI. > - New models require current harness releases. > - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and OpenCode 1.18.29. > - Changing the image alone would fail the runner's exact version and executable checks. > - This pull request updates those dependencies, integrity checks, controller checks, and image pins together. > - Shared installations can then run the current models without a task-time download. ## Linked Issues or Issue Description Refs #13829, which updates model choices and reasoning controls. Searches found no open PR that updates these runtime pins. **Current behavior** The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run Opus 5.5, which requires 2.1.280. Remote controllers reject provider packs whose versions differ from their declared pins. **Proposed behavior** Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge patches and one shared CLI installation per provider. **Reason and benefit** Current harnesses support the new model IDs while preserving executable verification and remote provider-pack compatibility checks. ## What Changed - Update dependency overrides, the Codex ACP package patch, runtime profiles, and remote controller pins. - Verify the new Claude Linux x64 and macOS arm64/x64 executables and Codex Linux x64 executable against integrity-verified npm archives. - Refresh OpenCode version checks, fixtures, and the runner configuration label. - Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI pins and archive hashes. Hermes remains current at 0.19.0. - Refresh the build-time lock digest from clean pnpm 9.15.4 resolution. Leave lockfile commits to repository automation. - Document model compatibility and the separation between CLI runtimes and patched ACP bridges. ## Verification - `pnpm -r typecheck` and `pnpm build` passed. - Rust workspace release tests passed. - Package/patch and OpenCode binary-materialization contract tests: 11 passed. - Real Codex 0.156.0 startup-ownership and paginated session-resume probes passed with isolated synthetic homes and no model turn. - Codex app-server `thread/start` preserved `gpt-6-sol` and `gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in catalog does not include those account-served entries. - Installed Claude integrity probes passed for `claude-opus-5-5` and `claude-fable-5-1`. - `pnpm --filter @paperclipai/paperclip-runner test:opencode:qualification` passed with the actual OpenCode 1.18.32 executable under Node 24 and Node 25. The loopback provider exercise covers health/version, session creation/read/delete, SSE, and a completed async prompt. - `pnpm check:token-gates` passed. - The targeted runner suite passed 130 tests. Three macOS failures in snapshot module lookup and OpenCode final-message selection also reproduce on the unchanged base; Linux CI will provide the platform check. - [Final Linux CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399): all gates passed. Four jobs needed one retry after their CI workers received shutdown signals. The PR has 55 successful checks, two skipped checks, Greptile 5/5, and no unresolved review threads. - Changed runner configuration UI tests: 5 passed. - Full macOS `pnpm test:run` reached 13,094 passing server tests, 84 skipped, and 18 failures before the wrapper stopped. Failures involved skill-cache publication permissions, missing bundled connector skills in the worktree, and a conversation-reset timing case. The 10 cache permission failures reproduce on the unchanged base; both conversation-reset cases passed on a targeted retry. The wrapper did not reach its later workspace/serialized groups locally; Linux CI covers those groups. - The local Docker daemon did not respond, so no local Docker build was run. No billable model requests were made. ## Risks - Deploy the matching controller and provider pack together. Older controllers enforce their previous exact pins. - Current upstream CLIs can change behavior. Existing protocol tests and isolated real Codex probes cover the integration boundaries; authenticated model inference is not part of these checks. - ACP bridge package versions and executable digests stay unchanged because their executable bytes are unchanged. Only the underlying CLI/SDK dependencies move. - No schema migration. Revert the runtime and image pins together to roll back. ## Model Used OpenAI GPT-6 via Codex, with repository tools, code execution, and web research. The exact serving model ID and context window were not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass for the changed surfaces and real-executable probes; full macOS-suite limitations are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e3d8fb0876 |
feat: attest standard production images at full source commits (#13797)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Operators deploy its standard production container on several CPU architectures. > - Downstream image builders need to identify the exact source of their base image. > - A short commit tag does not provide signed source evidence. > - This pull request adds a full commit tag and signed image digest for canonical master pushes. > - Consumers can verify the source and compose from the immutable digest. ## Linked Issues or Issue Description **What existing behavior does this improve?** Publication of the standard multi-platform production image. **Current behavior** The Docker workflow publishes short commit tags and channel tags. It does not provide a signed standard-image contract tied to the complete master commit. **Proposed behavior** Canonical master pushes also publish `sha-<full-commit>` and attest the exact index digest after platform validation and an immutable-image orphan-reaping check. The signer certificate binds the source repository, commit, workflow and ref. Existing tags and the separate cloud producer remain available. **Reason and benefit** Downstream builders can prove the source of a standard base without adding their dependencies or repository details to the public workflow. No matching open issue or duplicate PR was found. ## What Changed - Add the canonical full-SHA tag without changing existing tag mappings. - Validate amd64 and arm64 descriptors and hash the exact registry response bytes and require its digest header to match. - Verify the immutable image and sign it with GitHub artifact attestations. - Run the contract tests in trusted PR verification and document the consumer contract. ## Verification - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/standard-image-contract.test.mjs`: 37 passed. - A read-only check against an existing published index returned its exact expected digest. - `actionlint -shellcheck='' .github/workflows/docker.yml .github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports only existing `ls` and word-splitting warnings. - `pnpm build`: passed locally with Cargo available. - `pnpm -r typecheck`: passed locally. - Full local Vitest was attempted: 8,260 passed, 14 failed, with 34 failing suites. The failures were missing embedded-PostgreSQL library aliases in this fresh install and existing macOS runtime-skill-cache rename errors. Native aliases are now restored. Rerunning the 33 affected database suites produced 577 passes and two unrelated AgentMail skill-root lookup failures (32 suites passed). The four directly failing database tests also pass independently. This is not a claim that the full local suite passed. - All final-head CI checks pass. One unrelated routine-route mock assertion passed on the single-shard retry; its 15 tests also pass locally. Greptile is 5/5 on this exact head, with no unresolved threads. - Actual signing requires a canonical master push. This draft PR does not publish trusted provenance. ## Risks The new attestation step requires OIDC and attestation write permissions in the merge job. Signing failure leaves the image available but without the new admission proof. Consumers must fail closed when proof is missing. Existing release tags, the legacy producer, and image retention remain unchanged. No database or application behavior changes. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell execution and test tools. The session does not expose a more specific model variant or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
3c2bc4f546 |
ci: halve the isolated native Runner check by seeding it from the public master build cache (#13736)
## What Recurring CI health check (PAP-31): on the most recent fully-green PR run (`cb703ac`, run [35519542997](https://github.com/paperclipai/paperclip/actions/runs/35519542997)), the slowest check was **Compile isolated native Runner** at **452s** — ahead of the largest test shards (379s). Every run recompiled the full Rust dependency tree from zero, even though `docker.yml` already refreshes a **public** `mode=max` BuildKit cache (`ghcr.io/paperclipai/paperclip:buildcache-{amd64,arm64}`) on every master push, containing exactly these layers. This PR seeds **only the baseline build** with that registry cache, via anonymous pull. **Measured on this PR's own CI (which exercises the seeded path): the check completed in 123s, down from 452s — a 73% reduction, ~5.5 minutes saved per run.** ## Thinking Path Cost breakdown of the 452s from the job log: `cargo chef cook` dependency compile 214.7s, local cache export 43.1s, runner-core build 37.0s, metadata proof layer 36.8s, `cargo install cargo-chef` 36.4s, rebuild-verification build ~65s, setup/teardown ~20s. The dependency compile and toolchain layers are identical to what the production `docker.yml` build already caches publicly on every master push, so recompiling them here bought no signal — the check's real assertions live in the *verification* build, not the baseline. A first attempt used `actions/cache` plus a master `push` trigger, but the CI bot's GitHub App lacks `workflows` permission; the registry-cache approach is strictly better anyway (shared across PRs immediately, no 10GB Actions-cache quota pressure, no workflow change). ## What Changed - `scripts/check-docker-runner-cache.sh`: the baseline build now adds `--cache-from type=registry,ref=ghcr.io/paperclipai/paperclip:buildcache-{amd64|arm64}` (selected by host arch). `RUNNER_CHECK_SEED_CACHE` overrides the ref, or set it empty to force the old cold path. The script header documents the anonymous external read. - `.github/workflows/docker-runner-check.yml` (comment-only): the stale "no external cache" note now describes the anonymous GHCR seed and the verification build's local-cache-only isolation. This was pushed in a follow-up commit with workflow-edit permissions; the original CI-bot token could not touch workflow files. No Dockerfile stages or verification assertions changed. ## Verification - This PR's own `Compile isolated native Runner` check runs the seeded path (the script is in the workflow's trigger paths): **passed in 123s** vs the 452s baseline. - The rebuild-verification semantics are untouched: it still runs on a **fresh builder** importing **only the local cache exported by this run's baseline**, so it proves exactly what it proved before — that the runner image rebuilds reproducibly from this run's own exported layers. - Verified `ghcr.io/paperclipai/paperclip:buildcache-amd64` is anonymously readable (unauthenticated manifest pull succeeds), so the check gains no credential or secret dependency. ## Risks - **Stale or missing seed cache:** if the GHCR ref is unreachable, private, or garbage-collected, BuildKit logs a warning and falls back to the pre-PR cold compile — the check gets slower, never wrong. `RUNNER_CHECK_SEED_CACHE=""` restores the cold path explicitly. - **Cache trust:** the seed only accelerates the *baseline* build; the verification build still runs on a fresh builder against only this run's locally exported cache, so a stale or poisoned registry cache cannot make verification pass spuriously. The ref lives under `ghcr.io/paperclipai/*`, written only by repo CI on master pushes. ## Model Used Claude Fable 5 (`claude-fable-5`) via Paperclip agent **Bender (Fable)**, issue PAP-31. --------- Co-authored-by: Bender (Fable) <bender-fable@paperclip.local> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
b82661b561 |
refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path > - Paperclip manages agents and their access to external tools. > - Connectors expose these tools through a governed MCP gateway. > - PR #13755 added a direct Composio MCP connection behind the experimental MCP aggregators flag. > - The old project API-key broker still created toolkit child connections and showed a separate Services tab. > - Keeping both paths leaves obsolete setup and session code in the product. > - This change removes the broker and preserves direct MCP setup, credentials, permissions, and execution. > - Saved legacy records fail closed and remain available for explicit removal. ## Linked Issues or Issue Description Related: #13755. This retirement supersedes the legacy-path fixes proposed in #12630, #12632, #12634, and #12906. It does not close those PRs. **What existing behavior does this improve?** Composio connector setup, management, and runtime dispatch. **Current behavior** Composio offers both direct MCP and a project API-key broker. The broker mints sessions and creates one child connection per toolkit. **Proposed behavior** Offer only direct MCP. Remove the toolkit Services UI, REST routes, API client, and session broker. Block saved legacy parent and child records from discovery, execution, health checks, reconnect, and OAuth. Preserve their records and credentials until the operator removes each connection. **Reason and benefit** The direct MCP connector becomes the single supported Composio workflow. Provider accounts remain managed in Composio. ## What Changed - Remove the API-key catalog method and its generated-source definition. - Delete Composio broker clients, session creation, account synchronization, child lifecycle, and toolkit routes. - Remove the Services tab, service rows, child provenance, and cascade-removal controls. Keep Vercel provenance intact. - Retain a shared retirement guard for stored legacy records. Show Retired status and replacement/removal guidance in the connection list and details; hide obsolete runtime controls. - Preserve the experimental MCP aggregators flag and direct MCP infrastructure. - Replace broker fixtures with retirement tests and extend direct Composio catalog/reconnect coverage. ## Verification - Focused shared, server, and UI tests passed with one worker. Server retirement tests use a name filter; no full local test suite was run, as requested. - Server and UI TypeScript checks passed. - Token gates and UI build passed. - Real browser: opened the saved Composio connection, refreshed all 11 tools, and ran the provider's read-only GitHub account-list operation through the standard Test dialog as an agent. The provider returned success using the existing OAuth credentials. - See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live evidence. - Storybook build passed. A fresh real agent used `COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway audit records confirm both calls succeeded. - Browser retirement check: a credential-free legacy fixture showed the guidance, opened the direct MCP replacement flow, and was removed through the standard confirmation. - Focused regressions for the experimental settings copy and exact OpenAPI route coverage passed. All latest-head CI checks passed (54 successful, two intentionally skipped); Greptile scored 5/5 with no unresolved review threads. The PR has no merge conflicts. ## Risks This intentionally breaks the old Composio project API-key and child-connection workflow. Existing legacy records cannot run, even if their stored status is active. Operators must create a new direct MCP connection and choose access rules; credentials and grants are not migrated. Remove each old record separately to delete its credentials. No schema migration or data deletion runs automatically. Direct MCP connections keep their existing grants and secrets. ## Model Used OpenAI GPT-6 via Codex, with reasoning, code execution, and browser tools. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e8c8ba3c19 |
feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its tool gateway applies company access rules and approval controls to connected apps. > - MCP aggregators expose many apps through one provider endpoint. > - Each aggregator needs its own credential, catalog, grants, and lifecycle in Paperclip. > - This pull request adds independent Zapier, Arcade, Composio Connect, and Executor setup with a common Access → Connect layout. > - A default-off MCP aggregators flag lets operators opt in while we complete provider acceptance tests. > - Agents use the normal Paperclip permissions, Test screen, and gateway after setup. ## Linked Issues or Issue Description **Subsystem affected** Apps, connection setup, shared contracts, and the remote MCP gateway. **Problem or motivation** Aggregator endpoints need clear provider setup and correct MCP sessions. Generic setup does not explain each provider's authentication or broad execution tools. Provider approval must preserve the original execution instead of replaying a write. **Proposed solution** Add four separate connectors behind Settings → Experimental → MCP aggregators. Start with human and agent access, then connect the endpoint and read its tools. Enable tools by default. Use the existing Permissions and Test screens after setup. Keep legacy Composio API-key and child connections intact. **Alternatives considered** A shared connection for all providers would mix credentials and access rules. Separate provider-specific permission and test screens would duplicate existing controls. Vercel Connect is outside this change. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps capability and the Connected Apps roadmap area. This work was requested and reviewed by the maintainer. Related work: #11894, #12630, #12632, #12634, and #12906 concern the legacy Composio broker. #13102 also covers remote MCP pagination. This change preserves the broker path and adds initialized sessions, response matching, and provider resume handling alongside pagination. ## What Changed - Add branded setup and interactive Storybooks for Zapier, Arcade, Composio Connect, and Executor. Use the existing access controls and normal action tests. Do not request a connection name or action choices during setup. - Add the default-off `enableMcpAggregators` flag to settings, managed feature metadata, the catalog, and setup guards. Hidden connections keep running. Legacy Composio connections remain unchanged. - Reuse the vault, grants, policy, and catalog models. Support OAuth discovery, bearer tokens, custom headers, and credential-bearing URLs. Add no database tables or migrations. - Initialize and retain Streamable HTTP sessions by connection and effective credentials. Read paginated catalogs and match streaming responses to request IDs. - Classify unfamiliar aggregator tools as writes despite upstream read-only hints; only exact reviewed read capabilities enter the read-only allowlist. Legacy Composio child behavior is preserved. - Preserve provider authorization links and execution IDs. Support Executor approve/resume, decline, and cancel without automatic replay of uncertain writes. - Preserve Off and Ask first choices during refresh and reconnect. Allow new tools and retire removed tools. Keep agent access updates atomic and preserve an empty agent selection. - Document connector UX rules, provider branding sources, and live acceptance results. - Stabilize the existing Sentry release fixture after its repeated CI failure by reusing one module mock; production Sentry behavior is unchanged. ## Verification - Final head `d11781970`: [CI run](https://github.com/paperclipai/paperclip/actions/runs/35633534900) passed, including broad typecheck, test shards, build, and E2E. All 54 checks pass; 2 optional checks are skipped. Greptile is 5/5, Security Scan passes, and all review threads are resolved. - Passed 27 focused connector Vitest checks and 18 connector-only Storybook browser checks before the flag change. All 85 stories rendered at desktop and narrow widths. - Passed 5 connector lifecycle/server checks and 7 selected flag checks after adding the flag. The latter cover settings, managed defaults, cached catalog visibility, and all four setup routes. - Review fixes passed 13 risk/handoff/lifecycle checks, dedicated session-expiration and transport regressions, 13 selected connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A real Composio connection-list call also succeeded through the refreshed UI on `9ab115f71`. - UI and server TypeScript checks passed. UI build, Storybook build, token gates, and diff whitespace checks passed during implementation. - Real browser and real Paperclip agent tests passed for Arcade, Composio, and Executor. Tested action permissions, denied agent access, reconnect, disconnect, and isolation. Tested Arcade catalog additions/removal and Executor provider approve/resume, decline, and cancel. - Zapier live acceptance is incomplete. Its dedicated provider server is configured, but its credential-copy dialog returned an empty clipboard through browser automation. No live Zapier action is claimed. - The three isolated Sentry release cases pass after the CI fixture fix. - Local verification is deliberately narrow at the maintainer's request. The full local suite, recursive typecheck, and repository-wide build were not run. CI provides the broader checks. ## Risks - Shared MCP transport changes affect other remote MCP servers. Protocol fixtures cover initialized sessions, streaming response matching, pagination, and isolation. - Broad execution tools remain broad permissions. The provider governs actions inside those tools. - Provider handoff links are retained briefly in memory. After a server restart, a one-time link may require reopening the provider dashboard. Paperclip does not replay the original call. - Zapier remains unproven live. Custom-header imports and self-hosted endpoints have fixture coverage rather than a separate live account for every variant. - Turning the experimental flag off hides setup; it does not revoke existing credentials or stop existing connections. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, shell execution, and browser automation. The exact runtime model ID and context-window size are not exposed in this session. A separate Anthropic-backed Paperclip agent performed live gateway acceptance tasks. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
86b7ee992c |
feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agents have a persistent visual identity (#13171): a ClipLab character in one of 17 palettes, rendered as cached PNGs in lists and as a live character in larger placements. > - The onboarding wizard is where a person meets that identity first, and it showed a stock ClipLab expression on the previous engine while the rest of the app would show a different character on a newer one. > - The wizard's steps also cut from one screen to the next, so the arc read as separate pages rather than one walk. > - This pull request puts one character on one engine everywhere, gives the wizard's hero the studio's sleepy → wink → idle sequence on Review, and hands the steps over inside one presence. > - The benefit is that what wakes on Review is exactly what the agent looks like on the dashboard afterwards, and the walk to it reads as one screen changing. ## Linked Issues or Issue Description Refs #13171, now merged into master. This PR contains the onboarding and ClipLab update on top of that foundation. Original feature work by @tonio-alucema; merge preparation preserves the original commits. **Problem or motivation** The onboarding hero and the app's avatars were two different characters on two different ClipLab engines. Steps 1 → 4 of the wizard cut between screens, and the wizard mounted cold when a cloud-managed workspace arrived from Cloud's naming screen. **Proposed solution** Vendor ClipLab v0.2.0 as the shared engine and render one studio-exported character from it in every palette, for every pose and size. Play the export's one-shot wake on Review with the palette fading in over the gray dormant loop. Hand steps over inside one presence so the footer slides instead of jumping, and play the arrival half of that hand-off when the wizard opens directly on the agent step. **Alternatives considered** Exporting mp4/webm loops per size: no cursor following, no clean alpha, and the palette "colour in" is a runtime blend. Minting a `cap-v2` character version: nothing had shipped `cap-v1`, so the artwork is regenerated in place instead of migrated. Keeping the separately vendored runtime bundle for the hero: two engines and two characters in one app. ## What Changed - `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0 (`987b6db0`) with the Paperclip adaptations replayed (optional graphics backend for the Node SVG snapshot path, supersampled live textures, character framing, deterministic SVG id prefixes); new upstream `particles.ts`. - `packages/shared/src/cliplab/character.ts`: the studio export, mirrored from `ui/src/assets/cliplab/onboarding.character.json` by `scripts/sync-cliplab-character.mjs` (drift caught by `check:token-gates`). `characterDefinition` builds every palette from it; the resting portrait is its idle beat. - `OnboardingCharacter`: gray `sleepy` loop through the agent and connect steps; on Review the one-shot sleepy → wink → idle plays on two lock-step canvases while the palette fades in, then the `idle` loop. Body-follows the pointer, page-scoped. 160px in the wizard. - `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one `AnimatePresence` (departing content fades and gives its room back; arriving content opens its room then fills); the hero has a room that opens on the walk into the agent step; opening directly on the agent step plays the arrival half; the self-hosted naming step uses the arc's label and field. - Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`, `ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`). - Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent arc` walkable from the naming step plus `Arrive From Cloud`; the companies fixture answers the wizard's create call with a company. - Uses the shared runtime for onboarding; `doc/agent-personas.md` documents the shared character. - Releases both onboarding canvases after partial startup or transition failure. Registers each canvas before seeking so synchronous render errors can release it. Six component tests cover these failures and palette changes before or during wake. - Refreshes both sleeping canvases when the palette changes, including a palette change in the same render as wake. - Moves choreography values into the CSS token layer and preserves the shared motion catalog drift check across the imported stylesheet. - Repairs the static Storybook avatar route and uses accessible heading names/current button labels in the wizard play functions. - Closes the lazy avatar worker pool during application shutdown. ## Verification - Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm build-storybook`, and `pnpm check:token-gates` pass. The final UI typecheck and 123 focused onboarding, lifecycle, and token catalog tests pass. All 55 checks on final head `b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full sharded test suite, runner verification, and all eight browser shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)). The duplicate monolithic local `pnpm test:run` was stopped after CI completed; it is not claimed as a separate completed local run. - Chromium walkthrough: palette change, wake, return to sleep, WebGL failure fallback, Review step hand-offs and cloud arrival pass with normal and reduced motion; no browser errors. The signoff happy-path browser test also passes against a disposable instance. - The final CI run confirms the catalog fix and a passing signoff browser shard. The earlier signoff failure was a heartbeat-run availability timeout; the focused local reproduction and final CI passed without signoff code changes. - Original author verification: - `pnpm check:token-gates` (includes the new character sync check); shared, server avatar/persona (17) and UI onboarding/persona (137) suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs through the worker pipeline. - Storybook: `Agents / Personas` Sizes, Expressions and Palettes render the studio character at every size and pose; `Onboarding / Character → Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc` walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by ~560ms, footer travel continuous). - The original author walked the agent → connect → review flow and wake after a real sign-in on staging. - Not done here: the Linux Storybook visual baselines (`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for the new engine, hero size and naming-step changes. ## Risks - Every avatar's pixels change (new engine, new character) under the unchanged `cap-v1` name. Stacks that rendered avatars on the previous engine keep those PNGs in their cache (`generated-agent-avatars/cap-v1/...`, served immutable) until cleared; only the two pinned staging stacks ever did. - The one-shot handoff to the idle loop is timed from the sequence's authored duration (the engine reports completion by continuing into idle itself); presentation only, nothing in the wizard's state waits on it. - Reduced motion skips the wake and the hand-offs; jsdom is treated the same way, so the wizard tests see the next step's content immediately. - The committed export differs from the studio by one animation (Loop off, leading idle step removed); a re-export without that fix would play a 5.6s idle before the wake. ## Model Used Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in Claude Code, with shell, browser, and file tools. The original context window was not recorded. Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex, with reasoning, shell execution, file editing, GitHub CLI, and automated tests. The session does not expose an exact runtime model ID or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Dotta <bippadotta@protonmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1ef3b08714 |
feat(ui): integrate agent personas across the app (#13171)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A stable agent persona is useful only when the same identity appears across the app. > - Lists, task messages, selectors, and activity feeds need inexpensive static avatars. > - Onboarding and agent headers need a larger character with expressions and pointer tracking. > - This pull request connects the persona foundation to those existing views and preserves onboarding draft assignments. > - Full-page stories and Linux checks make the placements and performance contract reviewable. ## Linked Issues or Issue Description **Problem or motivation** Agents need a stable visual identity in lists, tasks, onboarding, and configuration. External tools also need an image URL for that identity. **Proposed solution** Assign each agent a permanent palette from a fixed ClipLab character library. Store the assignment on the agent. Render and cache preset PNG URLs on demand. Use static images in dense views and one animated character in larger placements. **Alternatives considered** A generated image bundle requires a separate asset build. A live renderer in every avatar adds unnecessary work in large lists. Arbitrary uploaded images do not provide the requested shared character system. **Roadmap alignment** This improves agent identity across existing control-plane views. It preserves agent permissions, company boundaries, and status labels. ROADMAP.md has no separate ClipLab persona milestone. Related approaches: #2422 adds configurable image URLs and DiceBear generation; #5578 adds optional uploaded avatars. This work uses a fixed, versioned character library and preset URLs. ## What Changed - Replace agent icons with static persona images across lists, the sidebar, org charts, tasks, comments, selectors, activity, and dashboard views. - Put one animated character in the agent header. Let it follow the pointer across the page, with reduced-motion and touch fallbacks. - Add larger padded characters to agent creation. Keep the palette stable across draft refreshes and connection retries, then reveal it after success. - Pass appearance through shared projections rather than fetching each agent separately. - Add real full-page Storybook examples for the agent list, overview, task, dashboard, new-agent dialog, and connection page. - Add Linux screenshot, clipping, density, and 500-avatar performance checks. ## Verification - `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased tree. Persona lifecycle tests pass. - The rebased feature passes 38 Linux screenshot/performance checks, including both display densities, corner pointer positions, and the no-WebGL/no-live-download contract for 500 avatars. - The final Linux persona suite passes all 38 visual, lifecycle, density, and full-page checks using the standard Storybook configuration and real on-demand avatar endpoint. - Final local focused verification: 45 avatar/native-recovery tests pass; UI identity/routine tests, typecheck/build, token gates, and Storybook build pass. - Current-head CI passes: full workspace/server tests, all serialized server groups, typecheck/release checks, build, canary validation, and end-to-end shards. The build passed after retrying a native-runner concurrency-test failure; its three targeted cases also pass locally. - Manual inspection covered stable identities in the app, header placement, full-page mouse tracking, onboarding size, and task/dashboard placements. ### Screenshots Linux captures use synthetic Storybook fixtures. Full-page captures use reduced motion. The live character, mouse tracking, and disposal are checked separately. <details> <summary>Agent overview with the character in its header</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png" width="900" alt="Agent overview with the character in its header" /> </details> <details> <summary>Task messages and assignee identity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png" width="900" alt="Task messages and assignee identity" /> </details> <details> <summary>Larger onboarding character with room for expressions</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png" width="900" alt="Larger onboarding character with room for expressions" /> </details> <details> <summary>Dashboard agent activity</summary> <img src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png" width="900" alt="Dashboard agent activity" /> </details> ## Risks - This PR depends on #13170, the persona foundation. Merge the foundation first, then retarget this PR to master. - Many placements change from icons to character silhouettes. Human avatars and authoritative agent status labels retain their existing behavior. - Only one character can render live per view. Reduced motion, hidden/offscreen content, touch input, and renderer failures use the defined fallbacks. - The full-page stories use fixture data. They do not contact a real company or complete real provider sign-in. ## Model Used OpenAI Codex, GPT-6 family. The exact model identifier and context window are not exposed in this session. Used code editing, shell execution, browser inspection, and Linux visual testing. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Tonio <tonework@gmail.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
d54b750111 |
Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset time when the runtime reduced provider metadata to a generic category error. Inspect terminal metadata in memory and retain only safe recovery labels and a parsed reset timestamp. Preserve the existing handling of other limits. Verified real child processes on both pinned ACPX runtimes, adapter and server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate a pre-existing chat regression from unrelated fixtures’ retry work. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
352153b5ed |
fix: run the cargo-building native-runner CI suite in the Rust-cached vitest lane (#13586)
## Thinking Path Post-merge of #13557, the slowest check on the freshest fully-green PR run ([35246999382](https://github.com/paperclipai/paperclip/actions/runs/35246999382)) was `ci / General tests (server (1/12))` at **339s**. The cause is one suite: `server/src/services/native-runtime/native-codex-runner.integration.test.ts` runs 1 test in **277s of a 291s vitest step (95%)** because its `beforeAll` cargo-builds the Runner release binaries, and the general-server shards carry no Rust cache — every PR run cold-compiles the full third-party crate graph. The other 19 suites in that shard finish in under 70ms each. The obvious fix (a dedicated Rust-cached matrix lane) requires editing workflow files, which the available GitHub App credentials cannot push (`workflows` permission). But the `Verify Paperclip Runner` lanes **already restore the shared `release-runner-v1` Rust cache read-only**, and their commands are `pnpm --filter @paperclipai/paperclip-runner <package script>` — so the suite can move into a Rust-cached lane purely through script changes. ## What Changed - `scripts/run-vitest-stable.mjs`: new `general-server-native-runner` group carrying exactly that suite. Under the PR workflow (`GITHUB_WORKFLOW == "PR"`, inherited from `pr.yml` by the reusable `pr-trusted.yml`) the without-chat server shards exclude it and rebalance to ~211s of tests each. Every other caller — local runs, `release-verify.yml` under the Release / Cloud readiness workflows — keeps the suite in the shards, so a renamed or unknown workflow degrades to today's slower-but-covered behavior instead of dropping coverage. - `packages/paperclip-runner`: `test:typescript:vitest` now routes through `scripts/run-pr-vitest-lane.mjs` — the identical `ensure:eval-build-deps && build:rust && vitest run` chain (shard flags passed through), plus the native-runner group on the **final PR shard only** (`--shard=N/M` with `N == M`, i.e. today's `vitest 2/2`, the 122s lane). With the restored cache the suite's cargo build becomes an incremental rebuild. - `scripts/__tests__/run-vitest-stable-shard.test.mjs`: guards pin the whole contract — PR 12-shard coverage (shards + chat + native-runner = full server group exactly), Release/local 10-shard runs keep the suite, `pr.yml` is named `PR`, the vitest lanes partition with exactly one final shard, the package-script wiring, and the wrapper's shard/workflow gating via its `--dry-run` plan output. No workflow files change. `.github/workflows/*` are untouched. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: **36/36 pass** locally on this branch (includes the new coverage, wiring, and wrapper-gating guards). Both files run in CI's `Test general-server shard partition` / `Test release verify workflow wiring` steps. - Wrapper `--dry-run` plan matrix verified for all six shard/workflow combinations plus malformed-shard rejection (pinned as a guard test). - The executing proof is this PR's own CI: `ci / Verify Paperclip Runner (vitest 2/2)` must go green while running the native-runner suite (its log will show the `general-server-native-runner` group after the package vitest shard), and the 12 `ci / General tests (server (x/12))` shards must go green without it. ## Risks - The exclusion keys on `GITHUB_WORKFLOW == "PR"`. Failure mode of a rename is safe (suite falls back into the server shards, slower but covered) and the guard test on `pr.yml`'s name makes it loud. - `vitest 2/2` grows from ~122s to an expected ~210–260s — still well under the ~306s `vitest 1/2` and ~326s e2e shards, and inside the 20-minute lane timeout. If the cache misses (key drift), the lane pays a cold compile like the server shard does today; a miss is slow, never wrong. - Double-run/coverage-loss combinations are enumerated in the wrapper header and pinned by tests: each caller runs the suite exactly once. ## Model Used Claude (Bender agent, Paperclip) — Fable 5. --- Expected savings once merged: the 339s `server (1/12)` check drops to ~265s-equivalent shard levels (~211s of tests), the slowest `ci /` check becomes the ~326s e2e shard (~13–33s off PR wall time), and every PR run stops paying ~4.5 min of billed cold Rust compile. For the merger (squash): please keep the trailer below in the squash body to preserve authorship. `Co-Authored-By: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>` ## Related PRs Searched the GitHub PR list for prior work on this surface — related groundwork, none duplicate this change: - #13457 — restored master's Rust dependency cache on the PR runner lane (the read-only cache this PR relies on) - #13500 — made that cache key image-toolchain-independent so GitHub-hosted PR runners actually hit it - #13521 — rebalanced PR shards and split the Verify Paperclip Runner lanes this PR extends - #13557 — previous health-check iteration (split the runnerd transport suite); this PR targets the next slowest check ## Checklist - [x] I have searched GitHub for duplicate or related PRs and linked them above Co-authored-by: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fcdb3f2499 |
feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). --> ## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agents that do research work need live information from the web > - Paperclip reaches external systems through governed, catalog-based MCP connections > - The Apps catalog is data-driven: a researched provider with a hosted remote MCP server becomes a connectable app with no runtime code change > - You.com operates a hosted remote MCP server for web search, content extraction, and research tools > - The server supports OAuth 2.1 with dynamic client registration, an API key in a bearer header, and a keyless free profile at a separate endpoint > - This pull request adds You.com to the self-serve MCP research ledger and generates its catalog entry with three connection methods: browser sign-in, API key, and the keyless free profile > - The benefit is that an operator can give agents live web search through the normal connection governance, and the free profile needs no account at all ## Linked Issues or Issue Description No public issue exists for this provider. The problem description follows the new-adapter issue template. **Agent or provider** You.com — web search and research tools over a hosted remote MCP server. **Why this adapter is useful** Agents that do research, monitoring, or fact-finding tasks need current web results. You.com exposes web search (`you-search`), live page extraction (`you-contents`), citation-backed research (`you-research`), and finance research (`you-finance`) as MCP tools. Any Paperclip company can connect it in a few clicks. The free profile offers `you-search` without an account, so a new company can try agent web search at zero cost and zero setup. **How the agent is invoked** Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`. Three supported access paths, verified against the live server on 2026-09-16: - OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate` challenge with RFC 9728 protected-resource metadata and advertises a dynamic client registration endpoint, so Paperclip's automatic DCR path applies. - API key. Sent as an `Authorization: Bearer` header per the provider's official server manifest and docs. Keys come from you.com/platform and unlock higher rate limits plus the full tool set. - Keyless free profile at `https://api.you.com/mcp?profile=free`. Provides a reduced, read-only tool set. Official docs: https://you.com/docs/build-with-agents/mcp-server **Are you willing to implement it?** Yes. Implemented in this pull request. **Additional context** Research evidence collected 2026-09-16, from live protocol probes and official provider sources only: - Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with `WWW-Authenticate: Bearer resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource" scope="Tools offline_access"`. - RFC 9728 metadata lists one authorization server with scopes `Tools` and `offline_access`. - The authorization-server metadata (RFC 8414) publishes authorization, token, and revocation endpoints, and advertises a `registration_endpoint`, so DCR is available. No registration was performed during research, per the runbook's non-registering preflight rule. - The keyless free profile answers `initialize` (server `You.com`, version `4.0.1`), lists the tools `you-search` and `you-discover`, and executed both tools successfully during the probe. - The API-key placement matches the provider's official `server.json` in the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer <key>`. ## What Changed - Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve MCP research ledger in `packages/shared/src/self-serve-mcp-research.json`, and refreshed the ledger verification date. - Added the You.com category (`ai`) and API-key header spec to `scripts/ingest-app-definitions.mjs`. - Added a You.com case to `specialMethodsFor` that emits three methods: browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer header), and keyless free profile (`mcp-free`, no auth). - Regenerated `packages/shared/src/app-definitions/youcom.json` and the generated registry via the ingestion script (`--definitions-only` mode; no unrelated provider churn). - Added the official You.com wordmark artwork (light and dark theme variants, taken from the provider's docs site) under `ui/public/brands/apps/`, with a manifest entry. - Updated `packages/shared/src/app-definitions.test.ts`: ledger counts (47 providers, 44 candidates), store count (48), verification date, and assertions for the three You.com methods and their endpoints. ## Verification - `node scripts/ingest-app-definitions.mjs --definitions-only` — passed. Generated the new definition and registry import only; no other provider JSON changed. - `node scripts/check-app-brand-assets.mjs` — passed (71 identities). - `node --test scripts/app-brand-validation.test.mjs` — passed. - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts ui/src/lib/app-brand-assets.test.ts ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests). - `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts server/src/__tests__/generic-mcp-connection.test.ts server/src/__tests__/tool-connection-removal.test.ts ui/src/pages/apps/AppsConnect.test.tsx ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server suites that require embedded Postgres skipped on this machine by their own environment gate; the gate is unrelated to this change. - `pnpm --filter @paperclipai/shared typecheck` — passed. - `pnpm --filter @paperclipai/ui typecheck` — passed. - `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed; builds `@paperclipai/shared` with the new definition. - `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592 skipped. Every failure is environmental on this container: the embedded-Postgres suites refuse to start because the machine runs as root, the native runtime suites need the Rust runner binary that this container cannot build, and one media suite needs a native HEIC binary. No failure touches the app-catalog, connection, branding, or shared-package surface; those suites pass locally. CI is the authoritative gate for the full suite. - `pnpm --filter @paperclipai/server typecheck` — not completed: the script's `prepare:runner-vendor` prelude builds the Rust runner, which cannot build on this container. A direct `tsc --noEmit` reports only pre-existing errors from the missing vendored runner types; no error touches this change. No server code is changed. - Live You.com proof on 2026-09-16 (keyless free profile, real network calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned results ✓, `you-discover` call returned results ✓. - Live proof NOT run: an authenticated OAuth connect and an API-key call against the full server. This environment has no You.com account or API key. Per the runbook, this proof stays outstanding and must not be assumed from the keyless probe. Both paths match the reviewed `mcp_remote` patterns (DCR and bearer header) used by existing providers. - Browser e2e suites not run: opt-in per `AGENTS.md`, and this change adds catalog data only, with no UI code. ## Risks - Low risk. The change is catalog data plus generated output. It adds no runtime code and touches no existing provider. - The free-profile method is a fixed keyless endpoint. If You.com changes or removes `?profile=free`, that method breaks and the entry needs a ledger update. The OAuth and API-key methods do not depend on it. - The authenticated tool catalog is discovered live at connect time, so provider-side tool changes appear through the normal catalog refresh and quarantine flow, not through this definition. - Rollback is a single revert; no migration and no state are involved. ## Model Used - Provider: Zhipu AI, via OpenRouter - Model: GLM-5.3 (`z-ai/glm-5.3`) - Context window: 200K tokens - Capabilities used: tool use (shell, file edits, live HTTP probes), long-context repository reading - The change was produced with AI assistance and reviewed by a human before submission. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes: #` / `Refs: #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6fe8e30625 |
feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external resources. > - Operators need to inspect Railway services, read logs, deploy code, and run container commands. > - Railway offers hosted MCP with OAuth, but broad remote actions hide their internal operations. > - This PR adds a branded connection and fixed direct operations through the existing gateway. > - Separate SSH keys enable container commands under the same grants and policies. > - Operators can require approval for an action and inspect the resulting audit record. ## Linked Issues or Issue Description **Subsystem affected** Apps catalog, connection setup, gateway execution, and connection documentation. **Problem or motivation** Agents need Railway access through Paperclip. Operators need to grant and revoke that access, inspect available actions, and govern deployment and container operations without giving agents provider credentials. **Proposed solution** Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and the gateway. Probe the actual credential before enabling fixed GraphQL operations. Use a dedicated grant-owned SSH key for bounded container commands. **Alternatives considered** A catalog entry alone cannot execute the missing operations. The hosted general agent has opaque internal effects. An unrestricted CLI runtime can bypass action policy and inherit ambient credentials. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps path and the Connected Apps direction in ROADMAP.md. It does not add a plugin or parallel connection service. Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway. They do not add this outbound Apps connection. The separate shared agent-picker fix is #13414 and is not included here. ## What Changed - Add the generated Railway catalog entry, official marks, provenance, and OAuth setup guidance. - Add fixed service/deployment status, bounded logs, and redeploy/restart/rollback tools. Block source deployment until the provider can atomically bind the approved repository and commit. - Verify API access with an explicit workspace before exposing direct tools. - Add grant-owned SSH key setup and a bounded runner with host verification, target checks, isolated state, and cleanup. - Block the opaque hosted railway-agent and accept-deploy actions. Preserve normal Allowed defaults and Ask-first policies for other actions. - Quarantine new or changed Railway schemas after initial discovery, including reconnect. - Add provider, lifecycle, gateway, SSH, UI, and browser fixtures. Document setup, limitations, and the release checklist. ## Verification - Security follow-up: removed the unsafe source-deployment mutation. Direct calls and old active catalog entries are denied before any upstream request, including normalized aliases. Refresh marks retired entries disabled. All 386 focused Railway, catalog and gateway tests passed, and server TypeScript checking passed. Full [GitHub CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144) passed on |
||
|
|
d08abcba15 |
ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every pull request runs the Trusted PR CI workflow before merge > - The test suites roughly tripled in six weeks, and shard balance did not keep up, so PR runs crept from ~4 to ~17 minutes > - Slow CI delays every merge and every contributor > - This pull request rebalances the shards from fresh measurements, splits the largest test files, reuses the Rust build cache in three more jobs, and takes the policy job off the critical path > - The benefit is a PR wall clock near 6 minutes with the same coverage ## Linked Issues or Issue Description **What existing behavior does this improve?** PR CI wall clock. A typical green run took 16-17 minutes. Two months ago it took about 4 minutes. **Subsystem affected** The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the shard-duration manifests, the vitest shard runner scripts, the `paperclip-runner` package scripts, and the dry-run branch of `release.sh`. **Current behavior** The shard-duration manifests were stale. The general-server manifest had durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29 specs. Stale median weights made shard steps range 417s-806s (server) and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release build. Every test lane waited ~60s for the policy job before it could start. **Proposed behavior** All lanes finish in a narrow ~200-290s band. The manifests carry fresh measured durations for every suite. The three largest test files are split so no single file caps a shard. The Rust cache restore runs in every job that builds the Runner binary. Test lanes start as soon as the gate resolves. **Reason and benefit** Merges stop waiting on CI. The projected wall clock is ~6 minutes for the same test coverage. ## What Changed - Rebuild `scripts/general-server-shard-durations.json` (646 suites) and `scripts/e2e-shard-durations.json` (all specs) from per-suite completion timestamps in runs 35036001734 and 35024948947. - Move the PR server lane to the release-verify shape: `general-server-without-chat` across twelve duration-balanced shards, plus the chat integration suite split by collected test location across three dedicated lanes. - Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and `-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions` and `-projects` specs. Each pair shares fixtures through a `.shared.ts` module. Playwright collects the same test sets (39 and 20 tests). - Raise e2e shards to eight and serialized shards to nine. - Run the runner package's `check:all` as four matrix lanes: `check:static`, `check:runner`, and two native vitest `--shard` halves. The union is exactly `check:all`. - Add the read-only Rust cache restore (toolchain pin, `save-if: false`) to the Canary Dry Run, Build, and Typecheck jobs. - Make release.sh preview publish payloads concurrently in batches of eight during `--dry-run`. The real publish path stays strictly serial. - Drop the policy-job lockfile artifact chain. Each lane installs with `--frozen-lockfile` and falls back to an inline `--resolution-only` regeneration. The policy job stays a required check through the `verify` and `e2e` aggregates. - Update the shard-count mirrors and workflow assertions in the partition and gate tests. ## Verification - `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/e2e-shard.test.mjs` — 30 pass. - `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass. - `node --test scripts/__tests__/release-verify-workflow.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass. - `playwright test --list` collects 39 tests across the chat-adapters split and 20 across the agent-chat split, equal to the original files. - A local vitest collection of the chat suite partitions 995 tests into 498/497 line shards. - Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s x8, serialized ~216s x9. ## Risks - The split spec files reorder tests relative to the original files. Every describe seeds its own company, so the specs stay independent; a hidden cross-describe dependency would surface as a deterministic failure in one shard. - The inline lockfile fallback changes install behavior for manifest-changing and stacked PRs. The policy job still validates resolution as a required check. - `release.sh` changes are confined to the `--dry-run` preview branch. The publish loop is untouched. `bash -n` passes and the release dry-run tests pass. - One PR now schedules ~44 fleet runners. If the RunsOn fleet caps concurrency, queueing may absorb part of the gain; watch the first runs. ## Model Used - Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with tool use (shell, file edits) in Claude Code. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
18989a9e73 |
docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its evaluations test both the Runner and complete product workflows. > - The guides and run histories are in separate places. > - The shared Evalbook viewer can make the test boundary unclear. > - This pull request names the two families and adds a guide, authoring skills, and a public hub. > - Contributors can choose the correct test and inspect its history. ## Linked Issues or Issue Description **Issue type** Missing documentation. **Where is the issue?** Runner and Product E2E evaluation guides, case-authoring procedures, and public result navigation. **What's wrong?** There is no single entry point. A report format can be mistaken for an execution boundary. There are no dedicated case-authoring skills for these two families. **Suggested fix** Add a guide and three skills. Link both existing histories from a public hub. Keep existing campaign URLs and grading unchanged. ## What Changed - Add `doc/evals.md` and links from existing guides. - Add the `paperclip-evals`, `add-runner-eval`, and `add-product-e2e-eval` skills. Install copies in `~/paperclipai/.agents/skills`. - Add a static hub builder that reads the existing public history feeds. - Show a dated snapshot for each family. Label partial campaigns and preserve measurement dates across report refreshes. - Document publication and refresh commands for https://pages.paperclip.ing/evals/. ## Verification - Seven Python summary tests pass: `python3 -m unittest discover -s scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR does not modify package scripts. - All three skills pass the skill-creator `quick_validate.py` check with `/usr/bin/python3`. - Build tested with saved history fixtures and the live public feeds. - Desktop and mobile browser checks pass. The mobile page has no horizontal overflow. - Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200, no page errors, all eight links return HTTP 200, no mobile overflow. - Independent skill exercises found the existing Notion-decline case and a direct Runner permission-denial case. Roster validation with an explicit run ID passes. - Missing refresh measurement date: regression fails before the fix and passes after it. - `git diff --check` passes. - No paid evals were run for this documentation and reporting change. The preceding head passed typecheck, build, server/workspace tests, runner verification, browser E2E, and the canary dry run. Checks for the latest commit are pending. Local repository-wide typecheck, test, and build were not repeated because no product code changed. ## Risks The hub is a dated static snapshot. It can lag behind the linked histories until an operator refreshes it. A changed history schema stops the build. Existing archives and grades are not modified. The published guide link is pinned to the reviewed commit so branch deletion cannot break it. Later builds can use master. ## Model Used OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna for documentation and independent skill checks. Both used repository tools and code execution. Context window sizes are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green (latest commit pending) - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups (preceding head was 5/5; latest commit pending) - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
a8d32e5e61 |
feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Agent work runs in sandboxes that provider plugins supply > - Operators can choose a provider to run agent work > - CreateOS adds another provider with workspace-preserving pause and resume > - This pull request adds a CreateOS provider plugin > - The benefit is that operators can preserve a workspace between runs without keeping its compute active ## Linked Issues or Issue Description Refs #13203 and the earlier closed #13096. This continues the CreateOS contribution from @bhautikchudasama and @ashwaq06. The branch preserves the original implementation commit. Thank you to both contributors. When squash-merging, preserve the original author's credit in the squash commit body: ```text Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com> ``` The original fork rejects maintainer pushes. This branch includes the merge-conflict resolution and review fixes. The request is described below using `adapter_request.yml`. **Agent or provider** CreateOS sandbox API (https://api.sb.createos.sh). **Why this adapter is useful** CreateOS can pause a sandbox and resume it by ID. The workspace survives the pause. This adds a reusable-lease option to the existing sandbox provider system. **How the agent is invoked** Build and install the local plugin as described in its README. Open Instance Settings, then Environments. Select the `createos` driver. Supply an API key and shape. The driver then supplies sandbox leases for agent runs. **Are you willing to implement it?** Yes. This pull request is the implementation. ## What Changed - Adds the `createos` sandbox provider under `packages/plugins/sandbox-providers/createos`. - Calls the CreateOS HTTP API directly. The package adds no vendor SDK. - Implements the environment lifecycle hooks, incremental process output, and binary workspace sync. - Registers the optional bundled provider and its trusted host credential fallback. The fallback is limited to the official API origin; custom endpoints require an explicit key. - Lists the package in the release manifest with `publishFromCi: false` until its first npm publish is bootstrapped. - Waits through delayed pause/resume state updates without duplicate action requests. - Cancels queued API requests promptly while preserving request spacing. - Uses direct CLI invocation in the setup guide so paths and IDs are passed without an extra shell expansion. - Includes current master and retains its existing Git-subfolder containment fix. ## Demo Fresh setup and a run against a CreateOS sandbox. https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110 https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084 ## Verification All 25 jobs in [CI run 34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542) passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including typecheck, build, native runner verification, server and workspace tests, browser tests, and the canary release dry run. Greptile reviewed the same commit at 5/5 with no unresolved review threads. GitHub reports no merge conflicts. The remaining merge gate is code-owner approval for the new `package.json`, as required by `.github/CODEOWNERS` and the `master` ruleset. Reviewers have been requested automatically. Local checks passed: - Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke skipped), and `pnpm build`. - Host: focused credential and bundled-plugin tests (17 passed), plus CLI invocation safety (39 passed). - Release: package manifest check and release policy tests (18 passed). The full local `pnpm test:run` attempt caught the README command issue; its focused rerun now passes. The full local run stopped after its general-server group: 7,804 tests passed, with unrelated embedded PostgreSQL startup failures and 10 failures in unchanged runtime-skill-cache tests (`EACCES` on directory rename on macOS). It did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm build` reach the runner package and stop because this machine has no Rust/Cargo installation. The corresponding CI checks passed on provisioned runners, as linked above. The live CreateOS smoke requires explicit provider credentials and was not run during this review. It is available with `CREATEOS_LIVE_TEST=1 pnpm test` in the provider directory. The author supplied the demo links above. ## Risks The provider is opt-in and is not installed by default. It is available through a local-path install or explicit image inclusion. npm publication remains disabled until a maintainer bootstraps the package and enables publishing. Sandbox creation has no idempotency key. An ambiguous create response can leave a resource that requires provider-account inspection. Process tracking is in memory; durable lease recovery belongs to the host. The provider does not advertise guaranteed expiry, interactive login, snapshots, duplex channels, or ingress. Live native-runner qualification remains outside this PR's tested claims. ## Model Used Original provider implementation: human-authored by @bhautikchudasama, as reported in #13203. The original description reports Claude Opus 5 assistance. Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review, editing, and tool execution. The precise runtime model variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass (focused checks; full-suite environment limits documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com> Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
bfb4ceabbb |
perf(release): wait for npm registry visibility of all packages concurrently (#13495)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - The release system publishes ~33 public npm packages per release across the canary, nightly, beta, and stable channels > - `scripts/release.sh` publishes them strictly sequentially: publish one package, poll npm until that version is registry-visible, then start the next > - npm accepts a publish in seconds, but registry packument propagation can lag minutes per package (the CI budget was raised to 30 minutes per package after two aborted releases), so the total publish time is the sum of every package's lag — about two hours on a bad npm day, paid by every channel run including every canary on every master push > - This pull request keeps the publishes sequential but runs all the registry visibility polls concurrently once every publish is accepted > - The benefit is that the wall-clock cost of npm propagation drops from the sum of all packages' lag to the single slowest package's lag, with every existing safety property preserved ## Linked Issues or Issue Description No existing issue. Description follows the enhancement template: **What existing behavior does this improve?** The npm publish step of `scripts/release.sh` (Step 5), used by every release channel. **Subsystem affected** Release tooling (`scripts/release.sh`, `scripts/release-lib.sh`). **Current behavior** Packages publish one at a time, and after each publish the script polls npm until that package's version is visible in the registry packument before publishing the next. With per-package propagation lag of minutes (observed up to ~15 minutes; per-package CI budget is 30 minutes), the full 33-package set takes up to ~2 hours of mostly idle waiting. **Proposed behavior** Phase 1 publishes every package sequentially exactly as today (a rejected publish still aborts the batch immediately with exact attribution). Phase 2 then polls registry visibility for all packages concurrently. Each package keeps its own `NPM_PUBLISH_VERIFY_ATTEMPTS` × `NPM_PUBLISH_VERIFY_DELAY_SECONDS` budget, and any version that never becomes visible still hard-fails the release, now naming every straggler. **Reason and benefit** Total publish wait becomes the slowest single package's lag instead of the sum of all lags — typically minutes instead of hours. This shortens every canary, nightly, beta, and stable run and reduces exposure to job timeouts during npm slowdowns. **Breaking changes** None. Dry-run output is byte-identical in structure, dist-tags are still applied at publish time (`--tag`), the Sigstore TLOG duplicate-recovery path is untouched, and the later dist-tag integrity check (`wait_for_release_registry_state`) is unchanged. ## What Changed - `scripts/release-lib.sh`: replaced `publish_package_to_npm_and_wait` with `wait_for_npm_package_versions`, which takes the package tuple list and polls every package's visibility in background subshells, each reusing the existing `wait_for_npm_package_version` poll (same per-package budget), then reports per-package success or fails naming all stragglers - `scripts/release.sh` Step 5: the publish loop calls `publish_package_to_npm` only (sequential, fail-fast on a rejected publish), followed by one call to `wait_for_npm_package_versions` for the whole set; Step 6's recap line updated to match - `scripts/release-lib.test.mjs`: the registry-visibility and workflow-budget tests now drive the new function (same assertions on `npm view` counts, virtual sleeps, and the fail-closed message, which now names the straggler); a new cross-visibility test proves concurrency — two fake packages that each become visible only after the other has been polled can only converge when polled in parallel, so the test fails if the waits ever serialize again Safety analysis for the ordering change: nothing in the publish loop resolves sibling packages from the registry. `prepare-bundled-package.mjs` bundles and patches from the local workspace tree, and the TLOG duplicate-recovery path only queries the package it just published. The only consumer of the "visible before next publish" invariant was the release script's own final verification, which still runs against the full set. ## Verification - `node --test scripts/release-lib.test.mjs` — 15 tests pass, including the new concurrency proof and the existing 15-minute-20-second budget tolerance test against the new function - `npm run test:release-registry` — full lane, 140 tests pass - `bash -n` on both scripts; `shellcheck` reports no findings beyond the three pre-existing ones on master (verified by comparing counts against `origin/master` copies) - A real-release exercise happens on the next master push: every canary run executes this exact path ## Risks - Low. The failure mode most worth watching is a release where some packages become visible and others never do: previously the run stopped at the first invisible package with later packages unpublished; now all packages are accepted before visibility is enforced, and the run fails naming every straggler. Recovery is identical in both worlds (the next attempt derives a new version number), and the accepted-but-lagging packages carry the correct dist-tag either way. - Publish jobs run the source commit's copy of `release.sh`, so this change takes effect for a given channel only once its source commit includes this merge — promoted nightlies/betas cut from older commits keep the old sequential behavior until their trains catch up. ## Model Used Claude Fable 5 (`claude-fable-5`), extended thinking with tool use (Claude Code). ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
6cfe4acff7 |
fix: preserve README images in npm package (#13488)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Its CLI is published as the paperclipai npm package, with the root README shown on the package page > - The root README uses repository-relative image paths so images render correctly on GitHub > - npm resolves those paths under the package repository directory, which is cli, so the image requests point to missing cli/doc/assets files > - This pull request prepares the generated npm README by converting only image src and srcset asset paths to stable raw GitHub URLs > - The benefit is that the same source README remains correct on GitHub and the published npm README displays its images ## Linked Issues or Issue Description **Where is the issue?** The issue is in the root README image assets and the npm packaging step in scripts/build-npm.sh. The affected public page is https://www.npmjs.com/package/paperclipai. **What's wrong?** The npm build copies README.md into cli/ before publishing. npm resolves relative image paths beneath the package repository directory, so doc/assets/banner.jpg becomes cli/doc/assets/banner.jpg. Those files do not exist, and the images render as broken on npm. **Suggested fix** Keep the root README paths relative for GitHub. Rewrite repository-relative image paths only in the generated npm README copy to absolute raw.githubusercontent.com URLs. ## What Changed - Added a small npm README preparation script that rewrites relative image src and srcset asset paths. - Updated scripts/build-npm.sh to use the preparation step when generating the npm package README. - Added a regression test for src, srcset, immutable refs, absolute URLs, and non-image Markdown links. - Pinned release-build image URLs to the source commit, while preserving tarball builds by passing their known source refs. ## Verification - Passed: node --test scripts/prepare-npm-readme.test.mjs - Passed: bash -n scripts/build-npm.sh scripts/e2e-install-lifecycle.sh scripts/e2e-update-migrations.sh - Passed: focused README and E2E migration harness tests - Passed: git diff origin/master...HEAD --check - Generated README asset URLs were checked against raw GitHub and all seven returned HTTP 200. ## Risks Low risk. The change affects only the temporary README generated for npm packaging. It does not change the GitHub README or runtime code. The generated npm README depends on the public raw GitHub asset URLs remaining available. ## Model Used OpenAI Codex, GPT-5. Tool-enabled repository inspection, code execution, browser verification, and git/GitHub operations were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with Fixes: # / Closes: # / Refs: # OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs) - [x] My branch name describes the change (e.g. docs/... or fix/...) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4510bf7c9e |
ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Pull request CI uses a trusted workflow on the AWS runner fleet. > - The caller used a fixed SHA that also needed runner-group admission. > - A mainline pin update left CI queued because the group still allowed older SHAs. > - This pull request calls the trusted workflow on master, which requires code-owner review. > - New merged workflow versions can use the existing master runner-group entry. ## Linked Issues or Issue Description Related: #12968. That Dependabot PR proposes another SHA rotation. This change keeps this first-party workflow on master instead. **What happened?** CI run 34975562974 stayed queued because its trusted workflow SHA was absent from the runner-group allowlist. The fleet itself was healthy. **Expected behavior** New reviewed versions of the trusted workflow on master should receive runner access without a separate SHA allowlist update. **Steps to reproduce** Change the caller to a new trusted workflow SHA without adding that SHA to the restricted runner group. Its jobs remain queued. The master reference removes that recurring synchronization step. ## What Changed - Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`. - Exclude this exact first-party workflow from Dependabot updates. - Update the existing E2E shard workflow tests for the master caller contract. - Document the runner-group entry, required code-owner review, and old-reference retention. ## Verification - `actionlint .github/workflows/pr.yml` passed. - `node --test scripts/__tests__/e2e-shard.test.mjs .github/scripts/tests/cloud-runner-routing.test.mjs .github/scripts/tests/pr-runner-rust-cache.test.mjs .github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests. - Parsed Dependabot YAML and checked the exact workflow exclusion. - `git diff --check` passed. - Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS has no errors, and the active master ruleset requires code-owner review. This covers the trusted workflow, caller, and CODEOWNERS itself. - The approved organization setting now allows `paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`. All previous references and other runner-group settings remain intact. - Full application typecheck, tests, and build were not repeated locally for this workflow-only change. This PR's CI and review are pending. ## Risks New versions of the trusted workflow take effect for new callers after merge to master. Keep code-owner review and master protection enabled. Existing administrator pull-request bypasses remain unchanged. Third-party action pins, runner routing, and infrastructure are unchanged. Older callers still use their SHA pins; their allowed refs remain in place. ## Model Used OpenAI GPT-6 through Codex performed implementation and orchestration; its exact runtime variant and context-window size were not exposed. OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions and applied the approved runner-group setting. Both used code and tool access. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4cc387f907 |
fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.
## Linked Issues or Issue Description
Refs: #13455, #13454, #13192
**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.
**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.
**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.
**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.
## What Changed
- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.
## Verification
- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
|
||
|
|
da77a0c28c |
feat(ci): publish immutable cloud migrator artifacts (#13455)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images and a matching database migrator.
> - New migrator versions must currently become available on npm before
cloud can use them.
> - npm can serve package metadata while the named archive still returns
404.
> - This pull request publishes immutable migrator archives and a
complete dependency lockfile through the existing artifact store.
> - Cloud can install these exact packages without waiting for their new
npm versions.
> - This producer change prepares a separate cloud consumer and
readiness cutover.
## Linked Issues or Issue Description
Refs: #13454
**What happened?**
A recent master run built both packages by 06:22:41 UTC on 2026-09-15.
Both archives became downloadable from npm at 06:31:56 UTC. Fresh
metadata requests did not remove the delay.
**What did you expect to happen?**
Cloud should be able to install the verified migrator as soon as its
package build and artifact upload finish.
**Steps to reproduce**
Compare package build completion, npm publication, version metadata
availability, and tarball download availability for a fresh full commit
SHA.
**Version**
Master commit `08adcc70d5ec45b7ced9619a3dc10c1d1bec397d`.
## What Changed
- Add a master-only workflow that builds the DB and shared archives
without publication credentials.
- Resolve the dependency lockfile from local archives, then pin those
archives to content-addressed URLs.
- Publish the complete bundle to a separate prefix in the existing
S3/CloudFront artifact store. Write the commit manifest last and verify
public downloads.
- Add a dedicated OIDC role policy. Only canonical master can assume it.
Writes require `If-None-Match: *`; the role cannot overwrite or delete
objects.
- Add source, integrity, lockfile, publication, and real npm install
tests. Document the format and staged rollout.
- Attest the validated manifest with GitHub/Sigstore before S3
publication. The signature binds every package and lockfile hash to the
exact master workflow and source commit.
## Verification
- `node --test scripts/cloud-migrator-artifacts.test.mjs`: 7 tests pass,
including real `npm ci` with an empty cache and no new-version metadata
lookup.
- `pnpm test:release-registry`: 136 tests pass.
- `actionlint .github/workflows/cloud-migrator-artifacts.yml` and `git
diff --check`: pass.
- Ran the workflow's filtered install and package build against the
exact master source. Built and validated the dependency lockfile from
those real archives.
- Latest-head application tests passed, including reruns of two failures
in unchanged application tests. The final CI aggregate passed. The
application source is unchanged. Common-source local typecheck and build
passed; the full local suite has the same documented macOS
read-only-directory rename limitation as #13454 (13 failures in two
unchanged suites).
- The dedicated role and additive bucket read permission are configured.
IAM simulation allows only conditional writes in the intended prefix;
overwrite without the condition, other prefixes, and deletion are
denied.
- Published the verified master
|
||
|
|
058c55bf8f |
fix(ci): overlap preview migrator publication waits (#13454)
## Thinking Path
> - Paperclip manages AI agents and their work.
> - Cloud deploys an application image with a migrator from the same
source commit.
> - The migrator consists of the DB package and its matching shared
package.
> - npm can accept a package several minutes before readers can see it.
> - The publisher waits for shared visibility before it submits the DB
package.
> - This change submits both validated packages before polling their
visibility.
> - Their availability delays can then overlap while the final gate
still requires both packages.
## Linked Issues or Issue Description
Related: #13334. No duplicate open migrator-publication fix was found.
**What happened?**
The cloud migrator publisher waits up to ten minutes for the shared
package before submitting the DB package. The two registry delays
therefore accumulate. A shared visibility timeout can prevent the DB
package from being submitted at all.
**Expected behavior**
Submit both valid packages, then wait for both to pass the existing
exact-source metadata checks. Report which package remains unavailable.
**Steps to reproduce**
Return 404 for shared metadata after npm accepts the shared archive.
Observe that the old publisher never submits the DB archive until shared
becomes visible. The new regression test requires both submissions
before the first visibility sleep.
**Paperclip version**
Base commit
|
||
|
|
08adcc70d5 |
fix(ci): retry transient GitHub reads while waiting for source verification (#13429)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work
> - Every commit on master is published as an npm canary, and a canary
is the only thing a nightly, a beta and then a stable can be promoted
from
> - A canary only publishes after the release run confirms that Cloud
readiness passed for that exact commit, which it does by polling the
GitHub Actions API for up to 45 minutes
> - That poll used a read that threw on any non-OK response, so one
gateway error ended the wait immediately
> - The release run then failed and the commit published no canary,
although readiness itself had passed
> - A commit with no canary can never be promoted, so this silently
removes commits from the release path
> - This pull request retries the reads that mean "ask again" and leaves
every real failure fast
> - The benefit is that a transient API error costs a few seconds
instead of a release
## Linked Issues or Issue Description
No existing issue. The problem, in the bug report format:
**What happened**
The `Reuse exact-source verification` job failed 12 seconds into a
45-minute wait:
```
Waiting for Cloud source verified v1 for
|
||
|
|
52d120f68d |
fix(release): wait 30 minutes for npm to expose a published version (#13436)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work > - Every release publishes a batch of npm packages and then waits for each one to become visible before continuing > - npm accepts a publish immediately but exposes it later, so that wait exists to keep a release from continuing past a package nobody can install yet > - Today that wait was too short twice in a row, and each timeout aborted the release with the batch half-published > - Version numbers are derived from what is already on npm, so every retry moves to a new number and meets the same lag > - Two canary attempts burned two versions this way and shipped nothing > - This pull request raises the per-package budget from ten minutes to thirty > - The benefit is that ordinary registry lag costs waiting instead of a failed, half-published release ## Linked Issues or Issue Description No existing issue. The problem, in the bug report format: **What happened** `publish_canary` failed twice in a row with the batch half-published: ``` Warning: npm accepted @paperclipai/server@2026.914.0-canary.2, but the version did not become registry-visible. Error: stopping release: npm did not publish and expose @paperclipai/server@2026.914.0-canary.2 ``` The version was accepted at 17:49:09 and became visible at 18:04:29 — about five minutes after the poll gave up. `shared`, `db` and `adapter-utils` published at that version; `server`, `paperclip-runner` and the root package did not. **Expected behavior** Ordinary registry propagation delay costs the release some waiting, not a failure. A package that becomes visible after 15 minutes must not abort the batch because of the old 10-minute wait. Longer registry outages can still leave a partial batch. **Steps to reproduce** 1. Publish any channel while npm is propagating slowly. 2. A package takes longer than `NPM_PUBLISH_VERIFY_ATTEMPTS * NPM_PUBLISH_VERIFY_DELAY_SECONDS` to become visible. 3. The release aborts, that version is half-published, and the retry picks a new version number and meets the same lag. **Paperclip version or commit** Present on master. Observed on 2026-09-14 across canary runs in workflow run 34869494325. ## What Changed - Increase npm visibility checks from 60 to 180, retaining the 10-second delay: about 30 minutes per package. - Increase canary, nightly, beta, and stable publish job timeouts from 90 to 150 minutes. - Add offline regression tests using the workflow's actual settings. They cover the observed 15-minute 20-second delay, immediate visibility, exhausted retries, and job timeout sizing. - Load the access router in test setup so its cold transform does not consume the first permission test's 10-second timeout. The permission assertions are unchanged. ## Verification - `node --test scripts/release-lib.test.mjs`: 14 passed. - `pnpm run test:release-registry`: 129 passed. - `pnpm exec vitest run server/src/__tests__/access-routes-permissions-upgrade.test.ts`: 3 passed. - Regression proof in temporary fixtures: restoring 60 attempts fails the observed-delay test; restoring 90-minute jobs fails the timeout-budget test. - [CI run 34916632804](https://github.com/paperclipai/paperclip/actions/runs/34916632804): all jobs passed, including typecheck/release registry, build, all general and serialized server shards, browser tests, runner verification, and canary dry run. The PR has 31 successful checks and two expected Storybook skips at `a27f5e896`. - [Previously failing serialized shard](https://github.com/paperclipai/paperclip/actions/runs/34916632804/job/104215809616): all three access-route permission tests passed in CI after preloading the router. - Greptile's final review is 5/5 with no outstanding findings. Both review threads are resolved. - Local limits: `pnpm -r typecheck` and `pnpm build` stop at the runner package because this machine has no Rust `cargo` executable. The duplicate full local `pnpm test:run` was interrupted while the complete CI matrix ran. Targeted local results are listed above; full validation is from CI. The previous CI failures were unrelated to npm propagation: - [PR review](https://github.com/paperclipai/paperclip/actions/runs/34891770396) required a test file for this fix. - [Serialized server shard 3](https://github.com/paperclipai/paperclip/actions/runs/34891773736/job/104136392196) timed out in the first access-route permission test at 10 seconds. The other two tests in that file passed. - The canary dry run passed in that same CI run. ## Risks Low risk. Production behavior changes only in release waiting budgets. - Polling exits as soon as npm exposes the version, so healthy publishes do not wait longer. - An unavailable version now takes about 30 minutes to report. Polls remain bounded and still fail the release on exhaustion. - The 150-minute jobs leave roughly 30 minutes for setup/build plus four full polling windows. npm command runtime and later release steps also consume that budget; a broader outage can still interrupt a batch. - The permission-test change moves module loading into a bounded setup hook; it does not relax authorization assertions. - No schema changes or operational migrations. ## Model Used - Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run through Claude Code with tool use and code execution. - OpenAI GPT-6 (Codex), with reasoning, tool use, and code execution, for the CI follow-up. Context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted checks listed above; full validation passed in CI - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes — workflow comments and verification details - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
728f7185f6 |
feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Self-hosted boards need a way to show occasional product announcements. > - An app release should not be required to publish or withdraw a card. > - Native card controls keep publishing consistent; the hero can use a static image or isolated HTML/CSS animation. > - This pull request renders a validated JSON feed with native components. > - It stores dismissals per account on each instance, so a closed card stays closed across companies and browsers. > - Named staging feeds let authors test content before production publication. ## Linked Issues or Issue Description **Subsystem affected** Board application shell, announcement delivery, and user preferences. **Problem or motivation** Operators need a small, optional announcement card. Users need reliable dismissal state. Authors need to test remote content without changing the production feed. **Proposed solution** Add one non-modal AnnouncementWell. Fetch validated JSON and content-addressed media through the instance server. Keep card controls native, with optional sandboxed HTML/CSS animation in the hero. Use stable announcement IDs for dismissal, an explicit empty manifest and quiet 404 handling. Provide a staged publishing helper and isolated test-drive guide. **Alternatives considered** Hosting the entire card as a page would move navigation and dismissal into remote content. This change limits HTML to a scriptless, isolated visual hero and keeps controls native. Browser-only storage would lose dismissals across browsers, so the instance stores account preferences. **Roadmap alignment** ROADMAP.md has no overlapping announcement feature. A GitHub title search found no related announcement pull requests. This work implements a maintainer-requested feature. ## What Changed - Add shared feed types, strict validation of every object, supported routes, expiration and version checks. - Add a board-only current-feed API, constrained media proxy, and idempotent dismissal API. Store the first dismissal and its company audit entry in one transaction. - Cache upstream data for one hour. Use conditional requests, request deduplication, response limits, public destination checks, and a three-second deadline. Treat a remote 404 as an empty feed with a fifteen-minute retry cooldown. - Keep announcement visibility stable when focus moves to browser chrome or another app pane; only tab visibility starts a return check. - Add a responsive native announcement card. Respect onboarding, dialogs and toast placement. Sync pending dismissals across tabs and retry after reconnect or return. - Add idempotent migrations for dismissals and validated publication IDs, design-guide examples, static and animated Storybook examples, and focused tests. The publication registry supports offline retries without accepting caller-invented IDs. - Add HTML/CSS animated heroes with static posters, automatic playback, reduced-motion handling, strict DOMPurify validation, an empty iframe sandbox and CSP that blocks scripts/network resources. - Add validated staging publication, content-addressed assets, an empty production manifest, preview fixtures, and authoring/operator documentation. ## Verification - The preceding implementation passed 98 targeted shared/server/publisher/route/OpenAPI/UI tests and 127 tests including the master rebase. The playback-control removal passes all 21 announcement UI tests, covering the rendered sandbox, fallback, reduced motion, dismissal and slow/stale state lookups. The preceding shared/server tests cover HTML validation and response sandbox headers. - The playback-control removal passes UI typecheck, production UI build, Storybook build and token gates locally. Browser verification confirms the animated card has only its dismiss button and two links, with no page errors. The full canonical CI matrix passed on current head `00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two optional Storybook deployment checks skipped. This run needed no retries. Greptile reviewed this same head at 5/5 with no outstanding findings. - The local canonical general-server run passed 12,063 tests before reporting embedded-PostgreSQL startup failures in an unrelated fixture. All 31 tests in that fixture passed across isolated retries. The UI group passed 6,219 tests and other workspace groups passed 3,201; two CLI database-startup failures also passed individually. Serialized server suites were verified by the full CI matrix rather than repeating them locally. No source changes were needed for these environment failures. - The real S3/CloudFront staging manifest and both media asset headers were verified. Production remains empty/unpublished. The guide distinguishes the preview host's disabled edge cache from production cache requirements. - In the isolated test-drive, the animation visibly moves without playback controls. A 390×844 browser viewport keeps the card above navigation. Reduced motion makes no animation request. Both themes render correctly and browser page errors are empty. Browser fault injection verified that scripts cannot execute and CSS cannot make network requests; a missing animation leaves its poster and controls. - Refresh leaves the animated card visible. Closing it persists after reload and the API returns null. Earlier live checks verified dismissal across browsers, company-relative CTA navigation, modal deferral/restoration, and new-ID eligibility after restarting the same database. - The deployed empty feed and a real remote 404 return HTTP 200 with null from the board API, with a usable dashboard and no announcement popup or browser warnings. - Authoring documentation covers staging, animated HTML constraints, test-drive, withdrawal, ID reuse and cache-refresh steps. ## Risks - Animation supports self-contained visual HTML/CSS and inline SVG, without JavaScript or external resources. A static image is required. Older builds that do not recognize the optional animation field quietly hide that unsupported feed. - The default feed makes an outbound request from an instance when a board is used. Operators can disable it. Requests contain no account IDs, company data, cookies or interaction events. - Feed publication and withdrawal can take about 65 minutes to reach returning users because of CDN and instance caches. Expiration also removes visible cards locally. - Dismissals follow an account within one instance. No-login instances share the existing local-board identity. Separate installations do not share state. - Both tables are additive. A unique key prevents duplicate dismissals; the transaction prevents duplicate first-dismissal audit entries. The publication registry retains only validated IDs. AGENTS.md and the implementation spec document the required exception to company scope for these instance-level records. - Publication was limited to separate public staging prefixes on the existing preview host. Production remains empty/unpublished. No AWS policies or infrastructure were changed. ## Model Used OpenAI GPT-6 through Codex. The exact runtime model ID and context-window size are not exposed in this session. Capabilities used: reasoning, code editing, shell execution, tests, browser interaction, and tool use. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
2d47a8058d |
fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly approves. ## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connector screens need recognizable app artwork. > - Several bundled marks have inconsistent artwork or dark-theme behavior. > - The shared logo component should retain the existing frame. > - This change replaces selected artwork and fixes local theme fallback. > - Users see consistent connector icons across shared component callers. ## Linked Issues or Issue Description **Current behavior** Some connector marks use outdated artwork or unsuitable theme variants. A remote dark logo can override a canonical local mark that works in both themes. **Proposed behavior** Use the selected bundled artwork with the existing gray rounded frame. Use local light artwork in both themes unless a distinct local dark variant exists. **Subsystem affected** Connector artwork, the shared AppLogo resolver, and its validation. **Breaking changes** No connector capability, permission, credential, or catalog activation changes. No duplicate PR was found in the earlier search. ## What Changed - Update 46 artwork files and only the app definitions whose logo paths need to change. - Keep a compact public manifest of identities, paths, visibility, and aliases. - Prefer local artwork in both themes and preserve the existing frame and padding. - Add locally runnable artwork safety checks and a light/dark Storybook gallery; leave PR workflows unchanged. - Document artwork conventions. Source research is kept outside the public manifest. ## Verification - Artwork check: 69 identities pass. - Node artwork validation tests: 13 pass (rechecked September 14). - Focused resolver, component, and catalog tests: 40 pass (rechecked September 14). - Maintainer decision: dedicated icon-validation CI is not required; the workflow remains unchanged. Greptile acknowledged 5/5 with no remaining code concerns on September 14. - UI typecheck, token gates, and Storybook build pass. - Earlier full build and repository typecheck passed. The broad local test run was stopped after workspace-runtime dependency fixture failures outside this change. Current GitHub CI remains the full-suite gate. - Visual approval remains outstanding. Review the canonical icon registry in both themes at 24–48px. ## Risks The SVG checker rejects common active features; it is not a general sanitizer for arbitrary uploads. Optical balance still requires human review. Remote fallback for unknown brands retains existing behavior. This change does not add an asset importer or new connector capabilities. ## Model Used OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and browser tooling. Exact hosted model IDs and context window sizes were not exposed. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [ ] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [ ] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ## Artwork comparison Before and after for every affected identity. Images follow your GitHub light/dark theme and are pinned to the base and PR commits. This compares artwork; the existing gray rounded product frame and padding are unchanged. | Connector | Before | After | |---|:---:|:---:| | AgentMail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg" alt="AgentMail" width="48" height="48"></picture> | | Airtable | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg" alt="Airtable" width="48" height="48"></picture> | | Asana | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg" alt="Asana" width="48" height="48"></picture> | | ClickHouse | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg" alt="ClickHouse" width="48" height="48"></picture> | | Cloudflare | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg" alt="Cloudflare" width="48" height="48"></picture> | | Cloudinary | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg" alt="Cloudinary" width="48" height="48"></picture> | | Discord | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg" alt="Discord" width="48" height="48"></picture> | | GitHub | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg" alt="GitHub" width="48" height="48"></picture> | | Gmail | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg" alt="Gmail" width="48" height="48"></picture> | | Google Calendar | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg" alt="Google Calendar" width="48" height="48"></picture> | | Google Chat | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg" alt="Google Chat" width="48" height="48"></picture> | | Google Docs | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg" alt="Google Docs" width="48" height="48"></picture> | | Google Drive | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg" alt="Google Drive" width="48" height="48"></picture> | | Google People | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg" alt="Google People" width="48" height="48"></picture> | | Google Sheets | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg" alt="Google Sheets" width="48" height="48"></picture> | | Google Slides | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg" alt="Google Slides" width="48" height="48"></picture> | | Google Workspace Search | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg" alt="Google Workspace Search" width="48" height="48"></picture> | | Hugging Face | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg" alt="Hugging Face" width="48" height="48"></picture> | | Jam.dev (library only) | New library entry | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg" alt="Jam.dev" width="48" height="48"></picture> | | Jira | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg" alt="Jira" width="48" height="48"></picture> | | Linear | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg" alt="Linear" width="48" height="48"></picture> | | Manufact | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg" alt="Manufact" width="48" height="48"></picture> | | Microsoft Teams | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg" alt="Microsoft Teams" width="48" height="48"></picture> | | Miro | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg" alt="Miro" width="48" height="48"></picture> | | Mixpanel | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg" alt="Mixpanel" width="48" height="48"></picture> | | Netlify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg" alt="Netlify" width="48" height="48"></picture> | | Notion | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg" alt="Notion" width="48" height="48"></picture> | | PagerDuty | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg" alt="PagerDuty" width="48" height="48"></picture> | | PostHog | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg" alt="PostHog" width="48" height="48"></picture> | | Postman | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg" alt="Postman" width="48" height="48"></picture> | | Shopify | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg" alt="Shopify" width="48" height="48"></picture> | | Slack | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png" alt="Slack" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg" alt="Slack" width="48" height="48"></picture> | | Stripe | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg" alt="Stripe" width="48" height="48"></picture> | | Supabase | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg" alt="Supabase" width="48" height="48"></picture> | | Telegram | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg" alt="Telegram" width="48" height="48"></picture> | | Todoist | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg" alt="Todoist" width="48" height="48"></picture> | | Wix | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg" alt="Wix" width="48" height="48"></picture> | | Zapier | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | <picture><source media="(prefers-color-scheme: dark)" srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg" alt="Zapier" width="48" height="48"></picture> | |
||
|
|
13368c5183 |
fix: unblock clean-machine onboarding for api_key AI connections (nightly smoke) (#13372)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The release pipeline gates each nightly on a Docker onboarding smoke. The smoke proves a clean machine can finish onboarding and hire the first agent. > - #13247, #13248, #13344, and #13351 changed the Connect step. Connect now creates an AI connection that the server verifies live with the provider. > - The managed adoption check also demanded a CLI hello probe. A clean machine has no provider CLI and cannot complete a subscription login. Onboarding dead-ends and the nightly gate fails. > - This pull request lets a live-verified API key adopt on the engine's own verdict, and re-verifies the key with the provider at adoption time. > - It also drives the release smoke through the API-key path against a provider mock that lives inside the test harness. > - The benefit is a green, deterministic release gate with no paid credential in CI, and a working first run for API-key users on clean installs. ## Linked Issues or Issue Description No public issue exists. The failure surfaced in the nightly release gate. Related PRs (no duplicates found): #13247, #13248, #13344, #13351 (the Connect changes), and #12423, #12135, #12151 (earlier release-smoke updates). **What happened?** The nightly Release cut failed its gate: [run 34749840498](https://github.com/paperclipai/paperclip/actions/runs/34749840498), job `smoke_nightly / smoke`, on published canary `2026.913.0-canary.2`. The wizard never left the "Connect a model" step. The subscription path waits for a human to run `claude auth login` on the server. The API-key path saves and live-validates the key, but the environment test then fails with `Command not found in PATH: "claude"` and `ai_connection_validation_incomplete`, and the wizard blocks the hire. **Expected behavior** A clean machine with a provider-accepted API key completes onboarding and hires the lead agent. The release smoke passes without a real paid credential in CI. **Steps to reproduce** 1. Run `scripts/docker-onboard-smoke.sh` with `PAPERCLIPAI_VERSION=2026.913.0-canary.2`. 2. Sign in, complete onboarding to "Connect a model", select "Use API key instead", pick Claude, enter a valid API key, and press Connect. 3. The environment test fails on the missing `claude` CLI and blocks the hire. **Paperclip version or commit** `2026.913.0-canary.2` (nightly candidate `c9e3bb7ca`). ## What Changed - `server/src/routes/agents.ts`: `testManagedEnvironment` no longer forces the CLI-lane hello probe for a resolved `api_key` binding. It re-verifies the key against the provider's live endpoint instead (the same `validateAiApiKey` check the save performed, which needs no CLI). A key the provider rejects fails adoption with `ai_connection_api_key_rejected`. Subscription adoption keeps the strict hello-probe requirement. - `scripts/docker-onboard-smoke.sh`: the harness now serves `api.anthropic.com` itself. A sibling container (the already-built smoke image) runs a small HTTPS mock. The app container gets `--add-host` for that one hostname and trusts the mock's certificate through `NODE_EXTRA_CA_CERTS`. The private key stays mode 600 in the mock container; the app container mounts only the certificate. The mock serves only `GET /v1/models` and returns 404 for every other path. `SMOKE_PROVIDER_MOCK=false` disables it. - `tests/release-smoke/docker-auth-onboarding.spec.ts`: the spec drives the API-key path — switch the credential mode before the source tile (the link hides when the row collapses), enter the key, and Connect. Loopback targets use a placeholder key that the mock accepts. Any other target must set `PAPERCLIP_RELEASE_SMOKE_ANTHROPIC_API_KEY`, and the test fails on arrival without it. - `server/src/__tests__/agent-test-environment-routes.test.ts`: three new route tests cover accepted keys (no CLI probe consulted), provider-rejected keys, and subscriptions that cannot complete a hello probe. ## Verification - `npx vitest run src/__tests__/agent-test-environment-routes.test.ts` — 26/26 pass. - `npx vitest run src/__tests__/ai-connections.test.ts src/__tests__/ai-legacy-compatibility.test.ts` — 42/42 pass. - `tsc --noEmit` reports no errors in the touched files. - Full local harness + suite run against the exact failing canary: the app container reaches the mock (request visible in the mock log), the placeholder key validates, and the connection saves as the default. The flow then stops at the forced CLI hello probe — the exact server check this PR removes, still present in the published canary. The next canary that includes this fix is the end-to-end proof. - Hardening check: from inside the app container, the mock answers with status 200 and `key.pem` is not visible. ## Risks - Behavior shift: `api_key` adoption no longer requires a CLI hello probe. It re-verifies the key with the provider at adoption instead. Subscription adoption is unchanged. - The mock returns 404 for unexpected provider calls, so a future onboarding change that calls a new endpoint fails the smoke loudly instead of passing silently. - Release-smoke runs against non-loopback targets now require an explicit key and fail fast without one. - No database migration. No dependency change. No provider routing change. ## Model Used Claude Fable 5 (`claude-fable-5`) through the Claude Code CLI, with extended thinking and tool use (shell, file edits, Playwright runs, GitHub CLI). No other models were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge |
||
|
|
13bae6fa21 |
fix: reject unsupported REST tool connections without stdio validation (#13346)
## Thinking Path > - Paperclip manages AI agents and their connections. > - Connection checks must use the configured transport. > - The tool service treated every remaining transport as local stdio. > - Anthropic's old REST method therefore failed with a templateId error. A REST connection with a valid stdio template could incorrectly pass. > - Anthropic now has a supported AI-account flow. This pull request removes its obsolete REST setup option and limits stdio checks to stdio connections. > - Users can connect an AI account, and existing unsupported connections receive an accurate error. ## Linked Issues or Issue Description Related: #13248 added the supported AI-account flow. Searches for related REST health and templateId bugs found no duplicate fix. **What happened?** The Anthropic REST API-key connection showed `Local stdio MCP connections must use an approved templateId`. Health checks and catalog discovery both fell through to the local stdio path. A REST connection with an approved template could report success and expose the template's catalog without a REST integration. **Expected behavior** Only local stdio connections use command templates. Unsupported transports return an accurate HTTP 422 error. New Anthropic accounts use the supported runtime authentication flow. **Steps to reproduce** 1. Check out the test-only commit `924e6e85a` in a separate worktree and install dependencies. 2. Run `pnpm exec vitest run packages/shared/src/app-definitions.test.ts server/src/__tests__/tool-access-service.test.ts -t 'unsupported REST|obsolete Anthropic'`. 3. The tests exercise saved Anthropic REST configuration and an unsupported REST connection containing an approved stdio template. They cover health checks and catalog discovery separately. 4. Run the same tests on the fix commit. They pass. The full affected files also pass. **Paperclip version or commit** Reproduced against master `6cef9743c`. **Deployment mode** Server transport handling. Reproduced with an isolated embedded PostgreSQL test database. No provider account or live credentials are required. ## What Changed - Restrict stdio health checks and tool discovery to `local_stdio`. - Return and audit `tool_connection_transport_unsupported` with HTTP 422 for unsupported tool transports. - Remove Anthropic's obsolete REST method from the generated catalog and its durable ingestion source. Keep its subscription and API-key AI methods. - Cover the reported error, false-success case, rejected obsolete setup, connection removal, and the UI's AI-account submission path. - Replace impossible reconnect forms for removed methods with supported setup, while preserving connection removal. - Preserve AI-versus-tool intent isolation for legacy requests and reject new unsupported Anthropic tool requests. - Document recovery for existing unsupported connections. ## Verification - Clean-worktree red/green: the same command failed all six regression cases at `924e6e85a` and passed all six at `4d3de9de0`. The failing run includes the reported templateId error. - Green: all 555 tests across the six affected test files passed. - Recovery UI red/green: three added cases failed before the recovery fix and passed afterward; all 200 tests across setup, detail, and advanced controls passed. - After the recovery UI update, UI typecheck/build and token gates passed again. - `pnpm -r typecheck` — passed. - `pnpm build` — passed. - `pnpm check:token-gates` — passed. - Catalog regeneration — passed with the documented `PAPERCLIP_CONTENT_TEMPLATES` override for the local capture corpus. - Full CI on `69fb31fd4` — passed all general and serialized test shards, browser shards, typecheck, build, runner verification, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/34726975425. - The local serial `pnpm test:run` was stopped after the fixture correction superseded that run; full-suite verification above comes from CI. All 555 affected tests passed locally, including all 17 connection-intent tests after the correction. - Greptile — 5/5, successful check on final commit `69fb31fd4`, no unresolved findings. - No live Anthropic validation was performed. The UI regression uses a fake key and a mocked AI-account response. ## Risks Existing obsolete REST connections remain in needs-attention state. Users must add an account through the supported flow and remove the old connection. Credentials and grants are not transferred automatically. Removal remains covered. The specialized AgentMail and Composio paths keep their existing behavior. There are no schema or permission changes. ## Model Used OpenAI GPT-6 through Codex. The exact serving model ID and context-window capacity are not exposed in this session. Used reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
47ded8bf97 |
feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Agent runs need credentials for a specific provider and sign-in method. > - Connections already owns accounts, grants, and access permissions. > - AI authentication should use those same boundaries. > - This pull request adds the storage, API, adoption, and runtime foundation. > - Legacy agents keep their authentication until they explicitly adopt a managed connection. ## Linked Issues or Issue Description **Problem or motivation** AI credentials are configured separately from Connections. Agents cannot consistently reuse a responsible user's account or a permitted shared account. **Proposed solution** Manage AI accounts with the existing Connections grants and permissions. Keep model and harness selection independent from credential selection. Preserve legacy authentication until validated adoption. **Alternatives considered** A separate credential registry would duplicate ownership and access policy. Automatic fallback would risk using the wrong account. **Roadmap alignment** This extends the shipped Apps, multi-user, secrets, and agent-runtime capabilities. The maintainer requested the feature and reviewed the UI. Related groundwork: #11899 (connection permissions), #10910 (connection wizard), #11692 (Claude subscription profiles), and #11854 (Codex account rotation). ## What Changed - Add AI-purpose/runtime-auth contracts and an additive, idempotent migration. - Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and catalog entries. - Store credentials on grants. Resolve responsible-user defaults or explicit permitted grants. - Isolate managed credentials and provider sessions across accounts. Block missing credentials without ambient fallback. - Keep imported legacy secrets unchanged during reconnect. Use independent local Codex/Grok sign-in attempts for rotating credentials. - Add authorization, migration, concurrent refresh, retry, cancellation, and legacy-compatibility tests. This is part 1 of a two-PR stack. The app UI follows in #13248. Merge the foundation first. ## Verification - Updated against master `04e364236`, preserving upstream provider login and connector workflows. - Full workspace typecheck, production build, Storybook build, and token gates passed on the integrated branch. Final local-login changes passed 59 focused tests; new-agent and inbox regression suites passed 63 tests. - Browser checks verified automatic local Claude account detection, resumable Codex login commands, retry, focus restoration, and desktop/phone layouts. Commands create their isolated directory before invoking the CLI. - All current-head CI checks passed on `2a996560a`, including all server/workspace tests, browser shards, runner verification, typecheck, build, and canary dry run. Greptile reviewed that commit at 5/5 with no unresolved threads. Earlier local full-suite attempts hit the Mac PostgreSQL shared-memory limit; the complete suites passed in CI. - Renumbered the additive AI migration to `0276` after upstream migrations and regenerated its snapshot. Existing legacy agents retain their configuration. - Added local login status checks, owner-scoped retry, managed OpenCode remote homes, credential-aware model discovery, and task connection-repair delivery. ## Risks - Managed credential failures intentionally block execution. They do not restore legacy fallback. - Preview-era copied Codex/Grok subscriptions require independent reconnect. - The integrated branch has live provider acceptance coverage. This update verifies local Claude detection and Codex API-key task repair; it does not add a new subscription authorization/refresh or Daytona stress pass. - Runtime-auth connections must stay excluded from tool and channel handling. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser testing. The exact runtime model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
4d317274ce |
feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Channels connect external conversations to company tasks and agent execution. > - Slack, Discord, and AgentMail already provide durable delivery and access controls. > - People also need to reach an agent from Apple Messages and send photos. > - Photon provides shared Pro DMs, dedicated numbers, and authenticated event recovery. > - This pull request connects Photon to the existing channel services. > - People can message an agent while Paperclip retains task ownership and approval authority. ## Linked Issues or Issue Description **Subsystem affected** Cross-cutting: channel services, shared contracts, database constraints, Apps, and agent Channels UI. **Problem or motivation** Paperclip has no iMessage channel. A person cannot use Apple Messages to start a task, send a photo, or answer an agent's pending question. **Proposed solution** Add experimental **iMessage Photon** with Pro-compatible shared DMs or a dedicated Photon Cloud number per agent channel. Reuse channel admission, identity links, task generations, publication, and interaction continuation. Keep groups disabled for shared allocation. Dedicated lines support groups that an operator explicitly enables. Require a fresh linked message and a published agent response before setup completes. **Alternatives considered** Shared allocation has no owned phone number, so it reserves one project and allows DMs only. Dedicated allocation reserves one stable number. Local Mac access needs a separate deployment model. The upstream Photon Chat SDK adapter does not persist the poll mappings and send receipts required here. This change uses the lower-level SDK without adding another agent runtime. **Roadmap alignment** This extends Connected Apps and agent communication through the existing channel subsystem. It does not add a parallel tool connection or agent loop. GitHub searches for Photon and iMessage found no matching provider implementation. **Additional context** This ships behind the existing experimental channel gate. Dedicated-line release qualification remains incomplete. Real Photon Pro DMs passed task/reply, native poll, text answers, confirmation rejection, media, restart, pause, reconnect, revocation, and removal tests. An operator-supplied iPhone camera HEIC also passed the full round trip. Dedicated groups remain unqualified. See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the implementation plan](doc/plans/2026-09-11-imessage-photon.md). ## What Changed - Add the provider catalog entry, shared setup contracts, and a forward migration. A global partial index reserves the dedicated number or shared project until its endpoint is archived. - Add Cloud project inspection, vaulted project credentials, selected-line token renewal, and a leased receiver. Persist checkpoint updates under the receiver lease. Shared project replay accepts sparse increasing sequences only after a complete recovery barrier. - Connect DMs and enabled groups to existing task generations, sender authorization, ordered delivery, and publication services. Keep each iMessage conversation on its task after completion; only explicit `/new` or `/close` releases the binding. Publish committed inbound comments live and label their human bubbles “Sent from iMessage” in both task-chat renderers. - Persist immutable text/file send identities, upload receipts, poll IDs, option IDs, per-person drafts, and canonical interaction continuation proofs. - Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG previews, and related Live Photo companion video retention. - Add the three-step setup flow and channel management surfaces with official branding. Preserve the experimental gate and existing pause/disconnect behavior. - Add interactive production-component Storybooks for setup, access, recovery, and ongoing conversations. Add provider, integration, catalog, and browser regression coverage. Document setup, recovery, supported boundaries, and qualification gaps. ## Verification - Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and receive native Codex replies in Apple Messages. Unlinked senders cannot start work. - Three real follow-ups each reopened the same completed task. Incoming bubbles appeared on its open page without reload and showed “Sent from iMessage.” The third follow-up ran after restarting the server on `4d7222110`; the agent correctly repeated its previous reply from before the restart. - Native polls after restart, sequential text drafts, required-field correction, explicit submission, approval rejection with a required reason, and native continuation passed against Photon. - PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC passed in both directions. The camera photo produced a 3024×4032 JPEG preview. The native agent described it and returned the received HEIC byte-for-byte. - Pause/resume, reconnect, identity revocation, removal, `/status`, `/new`, `/close`, and stale answers after close passed live. Messages suppressed by pause did not become work on resume. Removal stopped intake and removed credential bindings. - All 304 focused tests passed on `4d7222110`. These cover Photon unit/integration behavior, both task-chat renderers, live comment hydration, completed-task continuity after restart, enabled groups, duplicate delivery, and explicit reset/close. The selected Teams completion-boundary regression also passed. Full workspace typecheck/build and token gates passed for the conversation fix; the final UI changes passed their affected typecheck/build and tests. - All 26 new Photon Storybook Playwright cases passed in light and dark themes, including the complete shared-DM setup journey and 390px mobile follow-ups. UI typecheck and the Storybook build passed. These stories use simulated Photon responses and do not replace the live evidence above. - The full chat-adapters browser suite previously passed all 39 cases. Migration checks passed, and migration 0275 applied to the isolated live instance with the earlier Photon migration already applied. - The local full Vitest run was previously interrupted by the host's embedded-Postgres shared-memory limit; it is not a full-suite pass. All 30 applicable CI checks passed on preceding head `7a5419cac`, with two skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit required-story discovery guard to the 26 passing Storybook cases. Greptile rates this final head 5/5 with no unresolved review threads. All 30 applicable CI checks passed, with two optional checks skipped. - A repeated live send key suppressed the duplicate but returned gRPC 6 / SDK `internalError` without an original receipt. Paperclip keeps unknown delivery unresolved. This provider behavior is covered by a regression test. - See [the verification record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package versions, redacted live evidence, deterministic coverage, and remaining qualification gaps. ## Risks - Dedicated group qualification remains unrun; groups are disabled for the approved Pro scope. Real iPhone camera HEIC passed transport, preview generation, agent inspection, and return. Keep the channel experimental; the dedicated-line release matrix remains incomplete. - Shared recovery and attachment aliases were verified against the live gateway. Duplicate writes currently return an error without the original receipt; unresolved sends require operator resolution. The implementation fails visibly on invalid replay ordering, a reset cursor, or changed identity. - The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF binaries have not been executed in this work. Linux musl has no packaged converter. Unsupported conversion retains the original and reports the missing preview. - The migration adds a global reservation across companies for Photon numbers and shared projects. Paused and revoked endpoints keep that reservation until removal. - Integration touches shared channel services. Existing provider browser coverage passes; broad repository verification is recorded above. - `pnpm-lock.yaml` is intentionally excluded under repository policy. The repository bot owns lockfile updates. The additional Superagent supply-chain scan is neutral/inconclusive because these new dependencies are not yet in the committed lockfile. Its security scan passed; all required CI checks pass. ## Model Used OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code execution, browser testing, and tool use. The exact served model identifier and context-window size are not exposed in this session. No sub-agents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
44f6312cd8 |
fix(ci): reuse one available Cloud registry cache (#13334)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud needs a verified image for each merged source commit. > - Fresh builders restore compiled native dependencies from registry caches. > - The current workflow imports up to eleven historical cache manifests at once. > - Live builds missed native layers that a fresh builder reused from one manifest. > - This PR selects the nearest available cache and tests reuse across fresh builders. ## Linked Issues or Issue Description Refs #13329 and #13330. A search of open cache PRs found no duplicate of this change. **What existing behavior does this improve?** Remote Docker cache reuse on fresh Cloud image builders. **Current behavior** [Cloud run 34714483272](https://github.com/paperclipai/paperclip/actions/runs/34714483272/job/103609096836) imported the previous cache manifest successfully but rebuilt `cargo-chef` and Rust dependencies. The dependency compile took 3m43s. The preceding image build had already exported those layers. A controlled [fresh-builder diagnostic](https://github.com/paperclipai/paperclip/actions/runs/34715336530) used the same source and registry cache. The single-manifest job reused both layers immediately. The multiple-manifest job rebuilt them and failed the cache assertion. Both jobs used GitHub-hosted runners with read-only access. **Proposed behavior** Inspect cache manifests in first-parent order and import only the nearest available one. Keep full-SHA cache exports, the ten-commit search bound, and the legacy fallback. If caches cannot be read, permit a cold build. **Reason and benefit** Avoid the observed cache misses without changing image contents or builder sizes. Expected savings include about four minutes of native tool/dependency compilation when those inputs are unchanged. The final merge-to-deployable gain still needs a post-merge measurement. **Breaking changes** No image, artifact, deployment, or runner-routing contract changes. ## What Changed - Select one available ancestor cache after Docker login and Buildx setup. - Preserve separate writable cache tags for each full source SHA. - Test cache ordering, missing caches, registry errors, and workflow integration. - Add the selector tests to the existing release-registry suite. - Export a local test cache, remove the first builder, and verify a source rebuild on a fresh builder. - Document cache selection and the stronger Docker check. ## Verification - Passed 456 focused workflow, routing, readiness, preview-artifact, and cache-selector tests. - Passed shell syntax, ShellCheck for the changed probe, actionlint workflow validation, and `git diff --check`. actionlint's shell checks were disabled for the workflow validation because unchanged migration-label commands trigger existing SC2012 notes. - The fresh-builder registry diagnostic proves the single-cache behavior. The [permanent two-builder probe passed](https://github.com/paperclipai/paperclip/actions/runs/34715771048/job/103612624090), including a changed real binary and dependency-declaration invalidation. - Passed all 35 latest-head checks (green or intentionally skipped), including full typecheck, test, build, and browser suites in [PR CI run 34715771217](https://github.com/paperclipai/paperclip/actions/runs/34715771217). - The real selector CLI inspected registry metadata and chose the nearest available ancestor cache. - Fresh Greptile review is 5/5 with no open findings. The PR title was corrected to meet the source-change naming rule; the review check passed after that correction. - Local full-suite runs and Docker builds are unavailable because the local Docker daemon is unresponsive after disk exhaustion. CI provides the Linux verification. ## Risks - Missing or unreadable caches cause a slower cold build. The selector logs that condition and preserves image publication. - Inspecting several missing ancestors adds lookup time. Each lookup has a ten-second timeout and the search is bounded. - The Docker test now exports a local cache. It removes the first builder before starting the second to release disk space, then cleans up its builders and files. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7435b2ee9c |
ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys images that contain the native Rust Runner. > - The image already builds that Runner before copying ordinary app source. > - A Rust source change still invalidates its entire compiled dependency layer. > - Compiled dependencies can survive source changes when their recipe is unchanged. > - This PR adds a separate locked dependency build before compiling the real workspace. ## Linked Issues or Issue Description Refs #13195. A search of related Docker and Cargo cache PRs found no duplicate dependency-recipe change. **What existing behavior does this improve?** Docker image build time after Rust source or embedded protocol changes. **Current behavior** The `runner-build` stage compiles dependencies and workspace code in one layer. In Cloud readiness run 34698143548, that stage took about 3m48s when its cache was unavailable. **Proposed behavior** Generate a recipe with pinned cargo-chef 0.1.73. Build locked release dependencies in `runner-deps`, then copy and compile real Rust source and embedded protocol inputs in `runner-build`. Source edits can reuse the dependency layer from the existing registry cache. **Reason and benefit** Reduce dependency recompilation during source changes and merge bursts. Expected savings are roughly 2–4 minutes when the old native layer would miss but dependency layers are available. Full cold builds also pay for the recipe tool installation. Ordinary app-only cache hits gain little from this change. **Breaking changes** None to the shipped application or image tags. The recipe tool and compiled dependencies remain in build stages. ## What Changed - Install a pinned recipe generator with its locked dependencies and the existing package-owned compiler. - Add recipe planning and compiled dependency stages. Use the same release profile, package, binary, and lockfile enforcement as the real native build. - Remove generated source stubs before copying actual source. Preserve protocol inputs, timestamp normalization, binary staging, and application checks. - Add Docker cache wiring regressions and update the Docker cache documentation. - Run a two-build probe in Docker Runner check. It requires dependency reuse, changed real binary metadata after a source edit, and a changed recipe after a dependency declaration edit. It uses a disposable tracked-source context and exports only small metadata files. ## Verification - Passed all five Docker build-stamp and dependency-cache tests with `pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`. - Passed the local ARM64 `docker buildx build --target runner-build --progress plain`. Local Docker then hit storage errors during a runtime probe; cache invalidation verification continues on GitHub-hosted Linux. - Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`, and `git diff --check`. - Passed a [Linux AMD64 cache probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199) against the PR source: dependencies compiled in 3m49s for the baseline and were `CACHED` after a source edit; real source compilation took about 37 seconds. Binary metadata changed and dependency declaration changes altered the recipe. The permanent probe is also running in latest-head Docker Runner check. - Passed latest-head [Docker Runner check](https://github.com/paperclipai/paperclip/actions/runs/34711145160), including the permanent source/dependency invalidation probe. - Passed full [PR verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2): typecheck, all grouped tests, native verification, build, release dry run, and browser checks. One unrelated signoff-policy browser test failed waiting for a heartbeat run on attempt 1; only that failed shard and dependent checks were retried, and passed. - Latest-head Greptile is 5/5 with no unresolved findings. Full local tests/build were limited by local disk exhaustion; Linux CI completed those checks. ## Risks - The two-build CI probe has a 20-minute job limit to cover the cold build and source rebuild. It adds no AWS routing. - A fully cold build must install cargo-chef and populate the dependency layer. Both become reusable registry layers; no Actions cache is added. - The recipe and final build must keep the same compiler, build profile, package, binary, and directory layout. A source-change rebuild probe checks real cache reuse and binary invalidation. - Dependency or compiler changes still require rebuilding dependencies. Existing image verification and full-SHA publication gates remain unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8df3ee2cf5 |
ci: rebalance serialized tests with current suite durations (#13328)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Cloud deployments wait for verified source commits. > - Verification splits serialized server tests across independent runners. > - The shard duration estimates came from August and no longer match current tests. > - Stale estimates put much more work on one runner than the others. > - This PR refreshes the estimates from a complete successful run to balance the existing runners. ## Linked Issues or Issue Description **What existing behavior does this improve?** The time spent waiting for the slowest serialized server-test shard in PR and release verification. **Current behavior** In [Cloud readiness run 34705914878](https://github.com/paperclipai/paperclip/actions/runs/34705914878), the five serialized shards spent 479, 371, 322, 390, and 365 seconds running tests. The recovery suite had a 55-second estimate but now takes about 156 seconds including process overhead. **Proposed behavior** Use fresh per-suite measurements with the existing deterministic duration balancer. Applying the same measured costs to the new assignment gives 385, 385, 386, 385, and 385 seconds. This predicts about 93 seconds less waiting for the slowest shard, before runner/setup overhead. Live CI will confirm the result. **Reason and benefit** Use the existing runners more evenly. No extra runner, test parallelism, cache, timeout, or routing change is needed. **Breaking changes** Suite-to-shard assignments change. The full suite set, assertions, and per-suite process isolation stay the same. **Additional context** Searched related CI and shard PRs. This updates the existing duration manifest, without duplicating a pending sharding implementation. ## What Changed - Refresh all 145 serialized suite weights from the same successful release verification run. - Record source job IDs and the measurement method in the manifest. Durations include process startup, imports, collection, tests, and shutdown. ## Verification - Passed all 30 shard and release-workflow tests: `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Confirmed every measured suite appears exactly once across all five source logs. - Compared old and new assignments using the same measured weights. The maximum fell from 478862ms to 385551ms. - In [PR CI run 34710696242](https://github.com/paperclipai/paperclip/actions/runs/34710696242), all five serialized jobs passed in 7m05s–7m27s including setup. The measured assignment is now balanced in a live run. - The same run passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Local full-suite verification on this base was limited by disk exhaustion; local typecheck and targeted shard tests passed. - Latest-head Greptile is 5/5 with no open findings. Every current-head CI check must be green or intentionally skipped before merge. ## Risks - Individual durations vary with load and future test changes. These estimates affect assignment only; missing or renamed suites receive the existing median weight. - Both PR and release verification read this manifest, so both receive the new assignments. Each suite still runs in its own serialized Vitest process. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
d2e940f4c1 |
ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Paperclip Cloud deploys verified images from merged source commits. > - Cloud readiness waits for every release verification check. > - Runner verification currently runs long TypeScript tests before Rust checks. > - These checks can run on independent runners with their own build directories. > - This PR runs them in parallel while preserving all checks and the shared dependency cache. ## Linked Issues or Issue Description Refs #13194. Related prior work: #13142 and #13259. A search found no duplicate parallel release-check change. **What existing behavior does this improve?** Time from merge to Cloud source verification and deployment readiness. **Current behavior** Recent successful runs take roughly 13 minutes from merge to deployable. In run 34705914878, Runner verification took 11m23s. Protocol tests finished before Rust tests and API authority checks started. **Proposed behavior** Run protocol and Rust verification in two matrix jobs. Cloud readiness still requires both jobs to pass. **Reason and benefit** Remove the serial dependency between independent checks. Expected improvement is about 2–3 minutes on a typical cached run, until the image build or server tests become the longest job. This is an estimate; post-merge timing will confirm it. **Breaking changes** Individual release Runner job names gain a lane suffix. Cloud source and readiness marker names stay the same. PR runner routing is unchanged. ## What Changed - Split release Runner checks into protocol and Rust lanes. Keep every constituent of `check:all` exactly once. - Restore the existing Rust dependency cache in both lanes. Allow only the Rust lane to save it after warming both build profiles. - Add coverage and cache authorization regressions. Document the parallel verification and single cache writer. ## Verification - Passed 477 workflow and source-verification tests with `node --test .github/scripts/tests/*.test.mjs scripts/cloud-source-verification.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`. - Passed `actionlint`, `git diff --check`, and the private AWS routing regression suite. - Passed local `pnpm -r typecheck` and the standalone `check:runner && check:api-authority` lane, including all 1,671 API tests before the protocol lane had built TypeScript output. - The broad local protocol run under Node 25 had four failures. The two affected files passed under CI's Node 24.19.0: 67 passed, 6 platform skips. - Local `pnpm test:run` aborted when disk space ran out; local `pnpm build` could not run afterward. These are local verification limits. [Linux CI run 34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424) passed full typecheck, all grouped tests, native verification, build, release dry run, and browser checks. Native protocol CI passed 1,986 tests, plus 1,671 API tests and the Rust suites. - Latest-head Greptile is 5/5 with no open findings. All 33 current-head checks are successful or intentionally skipped. ## Risks - Uses one additional short-lived verification runner per release verification. The existing AWS exact-master restriction remains in place. - The Rust lane warms debug dependencies so its cache save also serves protocol tests. Both lanes always rebuild workspace code. - A workflow regression could omit a check. The new coverage test compares the matrix checks directly with `check:all`; Cloud readiness depends on the complete reusable workflow. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, and code execution. The exact serving model ID and context window are not exposed by this environment. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7e6d512597 |
fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The first agent helps the board define work and hire other agents. > - That agent can have the general role while its instructions require hiring skills. > - Missing skills and blocked schema discovery make valid requests fail. > - Repeated confirmation and invalid waiting guidance can turn these failures into extra runs. > - This PR supplies the required skills, opens read-only schema discovery, and corrects the guidance. > - The agent can complete an authorized hire while company approval and duplicate checks still apply. ## Linked Issues or Issue Description Refs #13068 — the first-task onboarding flow that this change repairs. Refs #12029 — related drift between the sandbox allowlist and bundled hiring guidance. This PR adds schema access; it does not replace the earlier hiring-route fix. **What happened?** A general-role onboarding chief received hiring instructions without the core hiring skills. Sandbox requests to the documented OpenAPI endpoint failed. The agent then guessed question and hire payloads. The persona required new confirmation after validation errors and described waiting states that agents cannot set. **Expected behavior** A direct request authorizes the requested hire. The chief asks only for material missing details, uses valid API payloads, and completes the task. Formal company approval gates still apply. A saved human-input card gives the task a valid waiting state. **Steps to reproduce** 1. Create an onboarding chief with role `general` through the board. 2. Ask it to hire a friendly robot with a supplied name and responsibilities. 3. Check its assigned skills, schema requests, question cards, hire requests, and final task state. **Paperclip version or commit** Reproduced on the first-task onboarding implementation after #13068. The live local verification used this branch at `112f44610`. **Deployment mode** The original failure used a hosted sandbox with legacy Codex ACP. Live verification used an isolated local instance and real `codex_local` execution. Queue and HTTP/2 transport access is covered by automated tests. ## What Changed - Give board-created onboarding chiefs the existing core skills regardless of role. Preserve explicit skill version pins, including aliases. Keep ordinary general-agent defaults and authorization checks. - Allow exactly `GET /api/openapi.json` through both sandbox bridge transports. - Publish validator-tested question, free-text, hire, and waiting examples. Regenerate the runner API reference and capability inventory. - Clarify direct authorization, material ambiguity, and correction of confirmed pre-creation validation failures. Preserve uncertain-outcome reconciliation, duplicate protection, and company approval gates. - Align disposition instructions with agent permissions and the saved human-input waiting path. ## Verification - After rebasing onto current `master`: 69 targeted server tests, 110 queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These cover core skill defaults, version pins, actor restrictions, schema access, published examples, hire validation, idempotency, and approval gates. Waiting recovery tests and live question flows also passed before the rebase. - `pnpm -r typecheck` and `pnpm build` passed again after the rebase. Frozen dependency installation and both generated capability checks passed. - Ran the full `pnpm test:run` suite. The initial run had 14 failed server files due to local database resource limits, a missing built test fixture, and socket failures. All 14 files passed after fixture repair and isolated retries. UI, CLI, workspace packages, database tests, and all 145 serialized server files passed. - Real one-request hiring replay: one hire, one successful run, task done in 2m16s. No repeated approval or recovery escalation. - Real two-turn browser conversation: start with an unspecified hire, then supply a name and friendly robot responsibilities. One clarification card, one hire, two successful runs, task done in 3m27s of execution. No failed writes, confirmation cards, or recovery actions. - Assigned the hired robot a welcome-message task through the browser. It produced a warm message under 100 words and finished in one successful 66-second run, with no questions or recovery actions. - The two-turn flow still asked an optional preferences question and gave a technical final reply. These are remaining presentation limits. - Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads. Fixed its generator finding and passed 1,655 published-example/runtime API tests plus server typecheck. All latest-head CI checks are green (32 passed; 2 unrelated Storybook checks skipped). The signoff-policy browser test initially timed out while waiting for an approver run. Its shard passed on one rerun without code changes. [CI run](https://github.com/paperclipai/paperclip/actions/runs/34698211049). ## Risks - Onboarding chiefs receive more default skills. Ordinary general agents retain existing defaults, and explicit versions take precedence. - Prompt guidance can affect model behavior. The live replays are examples, not a guarantee that every model follows the guidance. - Retry guidance applies only when validation confirms that nothing was created. Uncertain outcomes still require checking existing agents. - No database migration or new public endpoint. Existing company boundaries, approval gates, and bounded recovery remain in force. ## Model Used OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code editing, and live browser verification. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
ab15aff390 |
feat: add experimental persistent agent chat (#13284)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Conversations must use the same tasks, controls, and execution history. > - Users need an ongoing chat with an agent without managing task properties. > - Agents should clarify and plan work, then hand execution to assigned project tasks. > - This pull request combines the reviewed Agent Chat stack for one squash merge. > - The benefit is persistent conversation with normal task governance and shared UI. ## Linked Issues or Issue Description **Subsystem affected** Task lifecycle, agent runtime tools, shared task UI, and browser/paid runner tests. **Problem or motivation** Users need one persistent conversation with each agent. A separate chat store or renderer would duplicate task behavior and bypass existing controls. **Proposed solution** Use a task-backed chat per company, user, and agent. Reuse the task composer and transcript. Clarify and plan in chat, then create assigned project tasks with the relevant plan. Keep Agent Chat behind its own disabled-by-default experimental setting. **Roadmap alignment** This implements the task-backed direction in [CEO Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat). Related proposals: #2504 and #9693. Related request: #7981. The maintainer requested one squash merge of the complete stack. Consolidates the reviewed runtime [#13281](https://github.com/paperclipai/paperclip/pull/13281), backend [#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI [#13283](https://github.com/paperclipai/paperclip/pull/13283) layers with this PR's E2E coverage. All four layers passed CI and received Greptile 5/5 before consolidation. This PR targets master and includes the complete feature. ## What Changed - Add personal canonical chat tasks with ordinary company visibility, immutable identity, idempotent first sends, and an idle waiting state. - Process `/new` in queue order. Preserve history, release a chat pause, and fence old provider context and delayed writes. - Keep chat lifecycle rules across recovery, finalization, assignment, task lists, and rollups. - Support research and plan revision in chat. Hand plans to ordinary assigned project tasks before execution starts. Reject new chat subtasks. - Add repository-aware project creation and discovery tools, including multiple repository IDs and GitHub URLs, authorization, idempotency, and durable project-created cards. - Reuse task UI components for chat, with starred/recent agent navigation and a separate `enableAgentChat` experimental flag. - Add deterministic browser tests and 24 paid chat cells across four Codex/Claude profiles, with validated reports and screenshots. - Integrate current master recovery, controller lease, queued-message, and task UI changes. Gate chat interruption and deferred promotion on ownership/feature policy. Guarantee lease renewal and active controls are stopped even if teardown fails. - Preserve master's migration 0273 and generate chat migration 0274 with idempotent replay for development databases. ## Verification - Prior exact heads of all four PRs passed Linux CI, including build, typecheck, general/serialized tests, and browser E2E. Each had Greptile 5/5 and no unresolved findings. - Integrated local verification passed: full repository typecheck and production build, Storybook build, token gates, 340 focused UI tests, all 20 deterministic chat browser tests, two migration replay tests, 88 focused chat/queue/native/controller tests, and provider/session regressions including real lease expiry. These include the three lifecycle regressions for the final admission/teardown fixes; server typecheck also passes. Current head `1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no unresolved findings and passing security scans. All final-head CI gates passed: build, full Runner verification, typecheck/release registry, canary, all general/serialized test shards, and all browser E2E shards ([CI run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)). Local PostgreSQL startup contention required serialized retries; skipped fixtures do not count as passing coverage. - The earlier paid campaign passed all 24 chat cells and retained 32 screenshots: [report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat). It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior evidence, not a paid run of this integrated head. - Manual check: enable Agent Chat in Experimental settings, open an agent, clarify and revise a plan, then hand off to an assigned project task. Stop a reply, send `/new`, and verify fresh context with retained history. Disable the setting and verify agent shortcuts/new chat turns are blocked. ## Risks - Queue/session integration can affect retries and delayed writes. Tests cover ownership, cancellation, reset boundaries, idle recovery, and ordinary task behavior. - Migration 0274 adds conversation fields and constraints. Replay is idempotent and preserves existing development chat history. - This combines the previously reviewed stack at the maintainer's request. Agent Chat remains off by default and is separate from Conference Room. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code execution, browser tools, and parallel review. The exact context-window size is not exposed in this session. Codex and Claude also ran as test subjects in the linked paid campaign. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |