383 Commits
Author SHA1 Message Date
Devin FoleyandPaperclip 6f2ce27ca7 fix(workspaces): prepare checkouts without a local seed config (#14810)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Task preparation can create an isolated Git worktree and run its
setup script.
> - The Paperclip repository setup script also prepares a seeded
development instance.
> - A server configured through environment variables can have no local
seed config.
> - This stops ordinary task preparation before the agent starts.
> - This pull request prepares checkout dependencies when no seed source
exists, while preserving errors for invalid sources and existing
development instances.
> - Tasks can start without creating or claiming a seeded development
runtime.

## Linked Issues or Issue Description

**What happened?**

A task with the Paperclip repository fails during setup when the host
has no repository-local or default instance config. The automatic
worktree provisioner requires a seed source even when the task only
needs the checkout.

**Expected behavior**

A plain checkout should prepare its dependencies without a local
development database. A missing custom source, invalid source path, or
existing development instance with a missing source should still fail.
Starting a seeded runtime must still require a valid source.

**Steps to reproduce**

1. Run an environment-configured Paperclip server without a local
instance config.
2. Add the Paperclip repository to a project.
3. Start a task that uses an isolated Git worktree without a custom
provision command.
4. Observe the setup error before agent execution.

**Paperclip version or commit**

Reproduced against `0d3e7bf6ac` with a real script subprocess and
workspace realization regression.

**Deployment mode**

Environment-configured server with external PostgreSQL.

**Additional context**

Searched open and closed GitHub PRs and issues. Related work: Refs
#14795 (seed-source diagnostics) and Refs #11733 (source validation).
This change keeps source validation and seed-readiness checks in place.

## What Changed

- Permit dependency setup when the default seed config is absent
(including the Docker image config path) and the worktree has no
development-instance state.
- Keep missing custom configs, invalid paths, and lost sources for
existing instances as errors.
- Create no config, environment file, or seed manifest for a plain
checkout.
- Keep dependency install failures visible and allow normal instance
setup once a source becomes available.
- Cover the setup script, seed-runtime refusal, and automatic server
worktree realization.
- Document the difference between checkout preparation and
seeded-runtime readiness.

## Verification

- Regression tests failed before the fix for absent-source checkout
preparation and dependency setup.
- `bash -n scripts/provision-worktree.sh`
- `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs`
— 34 passed; 1 existing flock-dependent test skipped on macOS.
- Server regression — 2 passed, covering an unset config and the Docker
image default path.
- `pnpm build` — passed.
- `pnpm -r typecheck` — passed.
- All CI checks passed, including the full test shards, build,
typecheck, browser tests, and canary dry run.
- The first local `pnpm test:run` encountered two chat-test failures
because skill discovery selected an unrelated parent directory. Both
tests pass at the PR commit in a clean temporary checkout. The full
local run was not completed; the redundant clean run was stopped after
the complete CI suite passed.
- `git diff --check` and added-line secrets/PII scan passed.
- Greptile: 5/5, no comments. The branch has no merge conflicts.
- No live tenant deployment or task retry was performed.

## Risks

- A new checkout with no implicit seed config now completes dependency
setup. It has no seeded development instance. A runtime request still
fails until a valid source exists.
- Existing instances and custom source paths retain their failure
behavior. The script does not synthesize a source from environment
credentials or copy a live database.
- No schema, API, or task-setting changes. Revert the commit to restore
the previous setup behavior.

## Model Used

OpenAI Codex (GPT-6), with tool-assisted analysis, code edits, and local
tests. The runtime did not expose a verified model variant or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:00:25 -07:00
DottaandPaperclip 4ac374103f fix(connections): repair Asana MCP and add shared-app sign-in (#14756)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Connections let agents use provider tools through the permission
gateway.
> - Asana provides an official remote MCP server, but its v2 server
requires a registered MCP OAuth app.
> - Setup can discover retired v1 endpoints and send a callback that
differs from the displayed URL.
> - This pull request repairs custom app setup and adds sign-in through
Paperclip's shared app.
> - Users can choose their own app without enrolling with Paperclip
Cloud.
> - Agents can use Asana tools after the user connects their account and
sets action permissions.

## Linked Issues or Issue Description

Related: #14739 supplies the personal credential repair used by resumed
Asana setup. No duplicate Asana authentication PR was found.

**What happened?**

Asana setup failed even with a user-created app. Root discovery metadata
still points at v1. MCP v2 uses the Asana OAuth issuer and requires an
MCP app with a client secret. Local setup also displayed a localhost
callback while an Origin header could make authorization use a numeric
loopback callback.

**Expected behavior**

Sign in with Paperclip's app when its broker profile is available. Keep
custom MCP app setup available without Cloud enrollment. Use the correct
issuer, callback, client credentials, and resource throughout setup.

**Steps to reproduce**

1. Open Asana in the connection catalog.
2. Supply an Asana MCP app's client ID and secret.
3. Start OAuth on a local instance opened with a numeric loopback
address, or resume a draft that cached v1 metadata.
4. Observe the wrong discovery endpoint or callback mismatch.

**Paperclip version or commit**

Reproduced from b54b2dc35c. Rebased onto
master at `0829d94af` after the single-screen setup change in #14811.

**Deployment mode**

Local development from source. The managed path also supports enrolled
self-hosted instances.

## What Changed

- Add the `asana.mcp` managed profile and the default Sign in with Asana
method.
- Use Asana's reviewed v2 protected-resource metadata before cached
endpoints.
- Require a custom MCP app's client secret and retain saved credentials
during setup or reconnect. Repair only the known Asana v1
issuer/resource binding, retaining company and callback checks.
- Expose a boolean for the acting user's saved client secret. The form
offers secret reuse only when that user has an active grant with the
required reference.
- Let users select their own app from the enrollment and
shared-app-unavailable screens, or from Advanced on the single-screen
setup page.
- Canonicalize HTTP loopback callbacks even when the request includes an
Origin header.
- Extend signed broker claims and provider URL validation for Asana.
Require refresh credentials on managed authorization.
- Document setup, distribution, and shared-app rollout requirements.
- Resolve permission-profile name collisions when finishing another
account. The live staging test found this after renaming the first Asana
connection; OAuth succeeded but profile finalization failed.

## Verification

- Final live staging proof used app commit
`c6053157c4e42ac017727117b754ae77fa5c45fa` and the real Cloud broker at
`767b63835170f664542afd0df99a76615e204b62`. In the embedded browser,
default shared sign-in required no client credentials, returned through
the central Cloud callback to the tenant, and discovered 39 actions. Get
me succeeded through the gateway as the selected QA agent (2.1 seconds).
The custom-app connection also returned a real result on this final
build (0.9 seconds).
- Retried the shared draft that failed during the first staging test. It
completed after the profile-name fix, retained the selected agent, and
kept the existing custom connection intact. Two database regressions
reproduced the collision before the fix and passed afterward. The
updated transaction rollback test also passes.
- Shared reconnect returned to the same staging connection with 39
actions. Earlier staging checks verified the custom-app fallback when
the shared profile was unavailable, saved-secret reuse on reconnect, and
Off blocking the action test. Allowed was restored after that check.
- Local live-provider checks also repaired a saved Asana v1
issuer/resource binding without reentering the secret and verified that
numeric-loopback setup uses the displayed localhost callback. Expiring
the local managed access-token timestamp triggered a real Asana refresh
and a successful Get me call. These early local broker tests used
enrollment/authentication and storage fixtures; the final staging proof
used deployed Cloud identity and persistent storage.
- The new production app is registered and configured, but production
sign-in has not been deployed or verified. Live provider revocation was
not run because the existing staging test app is shared with other
connections.

- After rebasing onto the single-screen setup flow, full `pnpm -r
typecheck`, `pnpm build`, and `pnpm check:token-gates` pass. Focused
verification passes 365 service and 36 broker-client tests. Broader
checks pass all 833 shared-package tests and all 530 connector-page
tests. The shared suite uses `TMPDIR=/private/tmp` to avoid macOS
temporary-directory symlinks in its canonical-path tests. The UI tests
verify the shared-app default and switching to a custom app with its
required client secret.
- Embedded-browser smoke on the current rebased build verified the
shared sign-in default, Advanced → custom app (client ID and secret
required), and switching back to Paperclip. Both existing Asana
connections remained connected after restart. No new provider
authorization was performed during this smoke.
- The full local `pnpm test:run` was interrupted when the execution
session restarted. Before interruption, it reported one runtime-slot
restart test failure. That test passed on an isolated retry after
clearing two unused PostgreSQL shared-memory segments. The full local
suite did not complete; CI must pass on the current head before merge.
- All 52 CI and security checks pass on
`9318fd4e9b616cdc3de12f40cdb9bd32d865af4c` (CI run `36882064080`),
including all eight browser shards, nine serialized-server shards,
build, typecheck, and canary dry run. Two optional Storybook checks were
skipped. Greptile review 4 reports 5/5 on this exact commit, with all
review threads resolved. Its updated summary identifies the current SHA;
this comment-triggered review did not publish a separate GitHub check
run.
- Provider revocation is unit-tested in the companion broker. Live
provider revocation was not run because the existing test app is shared
with other connections.

## Risks

- The shared Paperclip Asana MCP app has been registered with its
production callback and Any workspace distribution. Its secret is
provisioned in the production secret store, and the runtime client ID
and secret reference are configured. The production profile is enabled
in the saved deployment configuration. The companion broker has merged
and passed staging deployment; production sign-in still requires a
production deployment and live verification. Custom setup remains
available.
- Asana MCP uses the provider's fixed `default` grant. Paperclip action
policies limit agent tool use; they do not narrow provider consent.
- The callback correction affects HTTP loopback OAuth flows. Public
HTTPS callbacks retain their existing behavior.
- Reviewed discovery URLs now override stale cached endpoints. Tests
cover the Asana v1-to-v2 repair.
- No schema migration. Connection removal retains the existing
local-only revocation behavior.

## Model Used

OpenAI GPT-6 through Codex, with code execution, browser testing, and
GitHub tooling. The exact model variant and context window are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-10-01 10:24:44 -05:00
Devin FoleyandPaperclip 0dc8d80eea Clarify worktree seed source setup failures (#14795)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed worktrees can prepare an isolated Paperclip development
instance.
> - The built-in provisioner requires a canonical registered seed source
config.
> - A missing source currently has the same message as a rejected
symlink or non-regular file.
> - This pull request separates those messages and names the supported
setup choices.
> - Operators can choose the intended setup without changing the
source-validation guards.

## Linked Issues or Issue Description

**What happened?**

A plain repository checkout can select the control-plane instance as its
seed source. If that instance runs with environment-only configuration,
the source config file can be unavailable. The provisioner stops with a
message that also covers noncanonical files and gives no repair
guidance.

**Expected behavior**

The error should identify the selected source and distinguish an
unavailable prerequisite from a rejected file. It should explain that a
seeded development instance needs a canonical registered source. It
should describe the explicit no-op only for a checkout-only worktree.

**Steps to reproduce**

1. Use a plain base checkout with no repository-local config.
2. Leave the control-plane instance config file absent.
3. Run the built-in worktree provisioner against an isolated checkout.

**Paperclip version or commit**

Base commit `c8f874311c`.

**Deployment mode**

Managed local worktrees, including servers configured only through
environment variables.

Related: #11733 adds deeper source-readiness checks. #11735 changes
runtime and seed lifecycle handling. This change only improves the
existing shell guard's diagnostics.

## What Changed

- Distinguish unavailable source configs from symlinks and non-regular
files.
- Identify whether the selected source belongs to the base workspace or
control-plane instance.
- Explain seeded-instance prerequisites and the explicit checkout-only
setup choice.
- Verify failure still precedes target-state creation and CLI
invocation.
- Document the setup choice and its runtime-readiness limit.

## Verification

- `node --test scripts/__tests__/provision-worktree-self-heal.test.mjs`:
21 passed; one platform-gated test skipped because macOS lacks `flock`.
- `bash -n scripts/provision-worktree.sh` and `git diff --check`:
passed.
- `pnpm -r typecheck`: passed.
- `pnpm exec vitest run server/src/__tests__/ai-connections.test.ts`: 50
passed after running the installed PostgreSQL package's own symlink
hydration script in this worktree.
- `pnpm test:run`: attempted, then stopped after unrelated database
suites failed at startup. The offline install had omitted PostgreSQL
native library symlinks. The focused database rerun above verifies the
local repair; the complete suite is delegated to CI.
- `pnpm build`: passed.

- [Required PR
CI](https://github.com/paperclipai/paperclip/actions/runs/36797650741)
passed on `0a3ba63e12`: 50 successful checks and two intentional
Storybook skips. Greptile scored that exact commit 5/5; there are zero
unresolved review threads and no merge conflicts.

## Risks

- Diagnostics only. This does not supply a source config or repair an
existing blocked task.
- The failure predicates and exit status stay unchanged. Symlink and
non-regular-file errors do not recommend skipping setup.
- The checkout-only no-op requires an explicit policy choice. It does
not grant runtime or seed readiness.
- No schema, migration, tenant policy, deployment, or Sentry reporting
change.

## Model Used

OpenAI GPT-6-based Codex, with reasoning, shell tools, and code
execution. The exact serving model ID and context-window size are not
exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 18:35:33 -07:00
DottaandPaperclip 33f2b3a159 fix: separate GitHub tools and code review bot connections (#14750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Connectors catalog lets people give agents tools or connect
agents to conversations.
> - GitHub put these two uses behind one card and an extra choice.
> - People should choose the connection they need from the catalog.
> - This pull request keeps GitHub for tools and adds GitHub Code Review
Bot as a separate card.
> - Each card opens its setup directly. Both use the existing connection
code.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

GitHub connector discovery and setup.

**Current behavior**

With chat connectors enabled, GitHub opens a menu that asks whether to
use tools or create a bot. Saved tools and bots share the same catalog
entry.

**Proposed behavior**

GitHub opens tool account access. GitHub Code Review Bot opens agent
selection. Saved bots and drafts appear under the bot card.
Chat-disabled instances show only GitHub tools.

**Reason and benefit**

The catalog names the two uses and removes an extra setup choice. The
bot keeps the existing GitHub provider, credentials, endpoint IDs, setup
steps, and runtime.

**Additional context**

Related work: https://github.com/paperclipai/paperclip/pull/12843 and
https://github.com/paperclipai/paperclip/pull/14594 established GitHub
account identity. This change preserves that tool flow. No duplicate
catalog split was found.

## What Changed

- Split the generated app definitions into GitHub tools and GitHub Code
Review Bot. Reuse the existing GitHub logo and channel method.
- Open bot setup directly, including old resume and reconnect links.
- Put existing bot endpoints and drafts under the bot card. Hide
duplicate internal chat applications.
- Keep pasted GitHub URLs mapped to the tool connection.
- Add seven Storybook states for the catalog, saved connections,
disabled chat, both setup paths, mobile, and light mode.
- Fix narrow-screen bot rows so the label cannot overlap status and
setup actions.
- Update catalog, route, browser, and API tests, plus the GitHub
connector guide.

## Verification

- [Hosted
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fgithub-review-connection/?path=/story/connections-github-and-code-review-bot--catalog):
seven states built from this branch. The deployment passed its
public-file verification.
- All GitHub checks pass on `d13a2cd53561645bb2a15c6f8e75a61a936d6459`.
Two optional Storybook jobs skip under their normal trigger rules; the
manual Storybook deployment passes. The branch has no merge conflicts.
- Greptile: 5/5 on the current head, with no review comments or
unresolved threads.
- `pnpm -r typecheck`, `pnpm build`, `pnpm check:token-gates`, and `pnpm
build-storybook` passed. The final Storybook fixture also passed UI
typecheck and the hosted build.
- Targeted catalog, URL matching, routing, grouping, brand, and chat UI
contract tests passed.
- GitHub provider browser tests: 2 passed. These cover direct tool setup
and the bot setup and management lifecycle with provider responses
mocked.
- Embedded-browser test on an isolated local instance: opened both
cards, selected an agent, saved a bot draft, and resumed the same
endpoint under the bot card after a reload.
- Storybook Tool Setup and Bot Setup assertions pass in the published
preview. Chat Disabled assertions pass locally. Inspected mobile and
light mode, including the draft-row layout and official GitHub marks.
- Local full-suite limitation: `pnpm test:run` was not clean. A
cross-company route assertion failed in the aggregate run and passed in
isolation; a workspace-runtime test reached its 30-second hook timeout.
Some isolated database reruns skipped when the embedded-PostgreSQL
availability probe failed. The local aggregate was stopped after CI
completed. The corresponding full CI suites pass all 360 tool-access
tests and all 162 workspace-runtime tests.
- No live GitHub authorization or installation was performed. The
isolated instance correctly stopped at the cloud enrollment or public
HTTPS prerequisites.

## Risks

- Low scope: catalog presentation and routing change. There is no
database migration or provider credential change.
- Existing GitHub bot URLs now open bot setup directly. The tool route
remains `/apps/connect?source=github`.
- The bot remains behind the existing chat-connectors feature flag.
Existing endpoints retain `provider: github`.
- Channel applications are represented by endpoint rows. Regression
tests cover legacy bot applications, tools, active bots, and drafts
together.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and
embedded-browser tools. The exact deployed model ID, context window
size, and reasoning setting are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 17:15:02 -05:00
DottaandPaperclip ad55d0a281 fix(connections): repair personal credentials and request write access (#14739)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use Apps through a gateway that checks identity, company
access, and action policies.
> - Personal pasted credentials can point to company secrets. Setup can
show success while the gateway rejects every call.
> - Several OAuth methods also omit the scopes needed for their
supported write actions.
> - This pull request gives setup, health checks, and invocation the
same credential rules. Owners repair existing connections by
reconnecting.
> - New connections request reviewed permissions for their supported
actions. Read-only choices remain available under Advanced.
> - Agents can use the connections people give them, while existing
consent, identity boundaries, and action restrictions remain enforced.

## Linked Issues or Issue Description

Refs #14009 and #14008. This addresses the personal-credential defect.
The separate GitHub organization-identity selection defect is outside
this change.

Related work: #13942 fixed part of new personal-key setup. #14200
independently fixes legacy personal reconnect and protects managed-agent
profile credentials during removal. This PR covers that ownership
invariant across key and secret-URL setup, reconnect, health, discovery,
and invocation, and keeps owner reconnect as the repair path. #14059
tracks requested versus provider-asserted OAuth scopes; it remains
separate work. I searched open PRs and issues for Zapier, Airtable
scopes, connector writes, and personal credential failures.

**What happened?**

A Zapier secret URL saved through personal setup can become a company
secret referenced by a user grant. Health checks bypass the gateway's
ownership check, so the connection appears healthy but calls fail with
`grant_credential_invalid`. Custom-header paths can also receive a
duplicate `credentials.` prefix. Omitted OAuth scopes make write access
depend on provider defaults.

**Expected behavior**

Personal invocation credentials belong to the selected user. Setup,
health, and actual calls enforce the same rule. New connections request
documented permissions for supported read and write actions. Existing
tokens gain no permissions without provider consent.

**Steps to reproduce**

1. Connect Zapier or a generic secret URL with the personal identity.
2. Allow an agent to use the connection and complete setup.
3. Invoke a tool through a run-scoped gateway. The legacy layout fails
ownership validation despite successful setup.

**Paperclip version or commit**

The implementation started from
`44736c9c7c67b7b646ead9d51721db10f5b83835` and was rebased onto master
at `94e8dec56`.

**Deployment mode**

Built from source. Regression tests use isolated PostgreSQL fixtures and
controlled MCP transports.

## What Changed

- Share credential writing, ownership validation, and canonical paths
across initial setup, resume, reconnect, rotation, health, discovery,
and gateway calls. Keep OAuth client-registration secrets separate from
invocation credentials.
- Existing personal connections with company-scoped credentials require
owner reconnect with a fresh key or secret URL. Reconnect creates a
correctly owned value and updates the existing grant and declarations.
There is no automatic ownership backfill or new startup hook.
- Preserve PostgreSQL timestamp precision when reconnect checks whether
a grant changed. Previously, converting the timestamp to a JavaScript
Date could reject reconnect with a false concurrent-change error.
- Protect credentials used by other grants, connections, bindings,
managed-agent profiles, routine triggers, or secret proposals from
connection removal.
- Review all 117 tool methods, including 84 OAuth methods. Record
explicit scopes or documented provider-default exceptions with official
evidence. Add Airtable's seven scopes, Hugging Face repository/job
scopes, and other documented MCP permissions.
- Prefer available write-capable methods. Put explicit read-only choices
under Advanced. Explain pasted-key permissions and offer reconnect for
missing OAuth consent. Preserve existing grants, policies, Google
availability gates, and curated scope allowlists.
- Reconnect generic secret URLs and custom headers using their stored
credential fields. Refresh the catalog after setup, correct reconnect
feedback and error guidance, and let Cancel exit invalid setup while
Save & exit retains draft-saving behavior.
- Apply ownership checks to the new GitHub repository/skill connection
picker. Align the permission audit with the Google scope reductions
merged on master.
- Add run-scoped gateway, ownership, owner-reconnect, OAuth URL,
insufficient-scope, UI, and catalog-wide regression coverage. Update the
connector playbook and permission audit.

## Verification

Latest commit `97bc0b86e0eae0ec892e4ac44beff1a66164b20e` passes all
CI/status gates (55 completed check runs, no failures or pending checks)
and has a completed Greptile review at **5/5 with no outstanding
findings**. GitHub reports the PR as mergeable/CLEAN.

- **Embedded browser:** used the actual server and built UI from this
worktree, a fresh isolated database, and local HTTP MCP fixtures.
Completed personal bearer-key, secret-URL, and custom-header setup;
reproduced the legacy ownership failure; reconnected through the owner’s
form; and completed writes afterward. Read-back was verified for
bearer-key and secret-URL connections. Public organization-wide setup
appeared immediately in Browse without reload. Zapier URL
validation/Cancel and Google’s enrollment gate were also exercised.
- **Persistence and invocation:** verified user ownership, canonical
`credentials.authorization` / `remote.url` / `headers.X-Api-Key`
declarations, and unchanged connection/grant identity. The old company
secrets retain their ownership. Separate HTTP calls through an actual
run-scoped gateway session completed a write and read-back.
- **Backend coverage:** the final gateway suite passes all 82 cases,
including catalog Zapier and generic inline reconnect. It checks
company/user isolation, canonical declarations, same-endpoint URL
validation, fresh credentials, retained restrictions, and real gateway
read/write execution using fixture transport. A timestamp with
PostgreSQL microseconds covers the former false reconnect conflict.
- **Local checks:** 368 catalog, gateway, repository, and UI tests
passed before the final extra Zapier case; 49 GitHub skill access tests
also passed. All three Apps browser regressions pass, including
reconnect through the actual form and catalog visibility without reload.
Full `pnpm -r typecheck`, `pnpm build`, server typecheck after the final
patch, and token gates passed. Full tool-access service runs hit varying
15-second Google fixture timeouts; both affected cases and the updated
reconnect assertion pass in isolation (3 tests). The complete test
matrix passes in CI on this head.
- **Verification limits:** no live provider account was available for
Zapier/Airtable/OAuth consent or account-bound write proof. Public
metadata and local fixtures do not establish provider consent. The
original development database clone failed on a pre-existing missing
`tool_connections_transport_check` constraint; browser acceptance used a
fresh isolated database created by the normal CLI onboarding flow.

## Risks

- Existing broken personal connections stay unusable until their owner
reconnects. Health, discovery, and invocation return an actionable
ownership error; startup does not rewrite credential ownership.
- Scope changes affect new authorization requests. Providers may still
require resource selection, account roles, paid plans, or app
verification. Existing consent and action restrictions remain unchanged.
- Shared credentials are retained rather than reassigned or revoked.
Provider-default exceptions and unavailable live checks are documented
in `doc/connections/CONNECTOR-PERMISSION-AUDIT.md`.
- No new endpoint, database table, lockfile change, or CI workflow
change is included.

## Model Used

OpenAI Codex, based on GPT-6, with reasoning, code editing, shell
execution, web research, and browser tools. The exact deployment model
ID and context window were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 14:32:34 -05:00
DottaandPaperclip a36cbffa9e fix(connections): broaden natural-language and aggregator search (#14725)
## Thinking Path

> - Paperclip manages agents and the services they need for work.
> - Agents use connection search to discover a setup path before they
request access.
> - Tool-only filtering hid channel and AI methods from this search.
> - Requiring every query word to match rejected normal task
descriptions.
> - A small aggregator index also omitted supported apps such as
Circleback.
> - This pull request broadens retrieval and returns purpose-specific
setup guidance.
> - Agents can choose a relevant result while existing access and
provider-choice checks still apply.

## Linked Issues or Issue Description

**What happened?**

A search such as “AgentMail create an email address and manage an agent
mailbox” returned no usable result. “Help me find tools for circle back”
also missed Composio's supported Circleback toolkit. Queries longer than
200 characters failed validation.

**Expected behavior**

Return useful native and verified aggregator matches from
natural-language queries. Include channel/email methods when Chat
connectors is enabled. Identify each method's purpose and the correct
setup path.

**Steps to reproduce**

Use the queries above with `connections_search` from an active task.
Enable Chat connectors for the AgentMail case. The regression suite
reproduces these misses before the change.

Related routing work: #13941. This change does not change the runner
failure path or add channel setup to tool-only connection cards.

## What Changed

- Rank name and capability matches. Accept extra words, split names,
small spelling errors, and queries up to 4,000 characters.
- Include tool, channel/email, and AI methods. Return company-prefix
setup links for channel and AI flows.
- Add a dated snapshot of 1,583 official Composio toolkit names and a
refresh script. Merge duplicate MCP variants for search and link each
support claim to official evidence.
- Find authorized indexed aggregator namespaces within longer queries.
Return multiple app matches when the agent needs to choose.
- Prefer exact app names over fuzzy matches for other apps; retain
existing AI readiness.
- Preserve native preference, company and identity boundaries,
administrative denials, and saved provider consent.
- Add relevance and database regressions, extend native tool-authority
coverage, and document search behavior.

## Verification

- Red: 15 new assertions failed against the previous implementation; the
existing baseline passed. Added red-green regressions for Motion versus
fuzzy Notion and existing AI access during review. A further regression
covers mixed ready/unconfigured AI results and their per-result setup
guidance.
- Green: all 103 tests in the eight focused shared, database,
runtime-tool, fixture, and route suites pass on the latest commit.
- `pnpm -r typecheck` and `pnpm build` passed.
- Latest-commit CI passed: 54 successful checks and two skipped checks,
including the full test matrix, browser E2E, typecheck, build, and
canary dry run.
- The long local `pnpm test:run` attempt began before the review fixes
and retained transformed pre-fix search code; it also hit an unrelated
timing failure. Fresh serial reruns of the affected search suites and
three timeout cases passed all 124 tests. Parallel local route shards
hit two additional database setup timeouts; both suites passed all 17
tests on a fresh serial rerun. The complete corresponding CI suites also
passed. Duplicate broad local runs were stopped after CI completed. The
local UI suite independently passed all 7,007 tests.
- Greptile: 5/5 on `cc6a0180d`; all review findings resolved.
- Browser verification passed in a disposable local instance through
real process-agent search requests: AgentMail opened its setup flow with
the requester selected; the saved Circleback choice produced the
Composio setup card; a paragraph-length Notion query produced its setup
card. No provider credentials or external accounts were created.
- The browser test caught an invalid UUID-based setup URL. The fix uses
the company prefix and has a regression assertion.
- A 3,971-character catalog query found Circleback first in a local 10
ms spot check after sharing query preparation across the catalog scan.
This is a single measurement, not a performance guarantee.

## Risks

- Broader retrieval can return extra candidates. Named services rank
first; agents must select the relevant method.
- The public support snapshot can age. It proves catalog support, not
account authorization or the availability of every requested action.
- Channel and AI methods use existing setup links. The tool connection
card still accepts tool methods only.
- No schema, migration, credential, or runner lifecycle changes.

## Model Used

OpenAI Codex (GPT-6). The session does not expose a more specific model
identifier or context-window size. Used reasoning, repository search,
code execution, tests, and browser tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 10:20:50 -05:00
DottaandPaperclip d72389bee2 feat: add Browser Use Cloud connector and live task browsers (#14627)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The Apps gateway gives agents governed access to external tools.
> - Browser Use Cloud can run browser work, but a tool result alone does
not let a person watch or take over.
> - A task needs a durable browser session, a visible viewer, and
recorded costs.
> - This pull request adds a Browser Use Cloud v4 connection and
interactive browser tabs on tasks.
> - People can follow the work, interact with the page, and retain the
browser after the agent finishes.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need governed access to Browser Use Cloud. People need to see and
interact with the same browser from the task. A browser must remain
available after a run finishes and appear at the correct point in the
task feed.

**Proposed solution**

Add a native REST connection for the v4 API. Bind each session to its
company, task, agent, and credential grant. Open its interactive viewer
in the task side panel. Record provider costs as financial events. Use
`browser-use-cloud` as the app and connector key. Keep its skill with
the connector and deliver it only with authorized connection tools.

**Alternatives considered**

A v3 MCP connection would expose tools without the v4 lifecycle
integration. An external viewer link would leave the task. A fixed
viewer size would prevent pages from responding to changes in the task
pane.

**Roadmap alignment**

This extends the governed Apps gateway and Connected Apps roadmap. It
uses the existing task, grant, secret, approval, and financial records.
The work was requested by the maintainer. A search found no duplicate
Browser Use connector PR or issue.

## What Changed

- Add the Browser Use Cloud app, brand asset, API-key connection, and
profile settings under the `browser-use-cloud` key.
- Bundle the `browser-use-cloud` skill with the connector. Keep it out
of global `skills/` discovery. Deliver it only with authorized task/run
connection tools. Remove retired connector skill keys from runtime
overlays and preserve unrelated browser skills.
- Expose seven v4 tools through the governed gateway and deliver them to
native and CLI agents.
- Persist sessions, browsers, runs, event and recovery cursors, shutdown
leases, and cumulative cost accounting. Recover uncertain paid starts
without replaying them.
- Enforce task ownership, credential grants, approvals, revoked access,
and budget limits.
- Add interactive task browser tabs and compact chronological feed
entries. Retain the viewer across tab switches and keep visible idle
browsers open.
- Add debounced automatic viewport fitting, standard size presets, and a
viewer ownership lease.
- Add lifecycle, authorization, accounting, viewport, UI, and Storybook
coverage.
- Add an idempotent database migration after the current master
migration. Preserve deployed migration hashes. Migrate pre-release Cloud
connection and financial keys without replacing grants, credentials, or
browser history.
- Document provider behavior, live acceptance results, and the lack of
documented passkey forwarding.

## Verification

- Full workspace typecheck and production build pass on the updated
branch.
- Token gates, brand asset validation, module boundaries, and migration
ordering pass.
- Cloud tests verify global skill exclusion, authorized task/run
delivery, unassigned agents, disabled connections, revocation, adapter
isolation, and secret exclusion. The existing AgentMail connector
assignment test also passes.
- Migration replay runs twice against existing browser work and
financial records. It preserves the records and avoids duplicate costs.
- The focused provider, app catalog, OpenAPI, connection gateway, and
migration regression suites pass. Recovery coverage includes lost
replies, process crashes, provider rejection, and browser arrival
acknowledgement.
- All 54 checks pass on `2974b5f03641ad0cea3c941d8c02579316fa8c92`,
including the full test matrix, browser E2E shards, build, typecheck,
security, and release canary. Two optional Storybook jobs are skipped.
- Greptile is 5/5 on the same commit, with zero unresolved review
threads. The corrected review uses the actual master-to-head diff.
- The local `pnpm test:run` started and was stopped after the full CI
matrix passed. It did not complete locally; the full-suite result above
comes from CI.
- Earlier live acceptance used an isolated company with a capped
provider credential. The agent opened paperclip.ing, the embedded viewer
accepted navigation, and the same browser stayed available after
completion and tab switches.
- The local Storybook build passes. Stories cover the panel, footer,
feed entries, settings, lifecycle failures, and viewport modes with an
offline viewer fixture.

## Risks

- Browser Use charges for hosted work. Provider caps and local budget
checks reduce exposure; reported costs can arrive after work completes.
- Viewer and CDP URLs grant access to the browser. The server validates
and restricts them. They are excluded from agent results and durable
event data.
- Runtime resizing of v4 agent browsers uses a provider option confirmed
by live testing but absent from its published agent schema. Resizing
during a click may invalidate coordinates. Fixed presets remain
available.
- Viewport ownership is process-local and resets on restart. The
lifecycle and accounting records remain in the database.
- The original intermittent embedded-viewer stall has not been fully
diagnosed. A bounded reconnect and active-session recovery cover the
observed failure paths.
- Live tests did not cover every revocation, approval, rate-limit, or
restart case. Deterministic integration tests cover those paths. Passkey
forwarding is not claimed.
- Unknown create outcomes keep the credential available for cleanup.
Run-list absence cannot prove a paid POST was rejected, so recovery
stays pending until it can identify provider work.

## Model Used

OpenAI Codex, GPT-6. Used reasoning, repository search, code execution,
browser interaction, and test tools. The exact serving model ID and
context-window size were not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-30 07:13:16 -05:00
DottaandPaperclip 3b4b270650 fix(adapters): preserve ACP terminal failure diagnostics (#14573)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The shared ACP adapter engine records agent failures for operators.
> - ACP providers can report a failure category, title, and detailed
cause.
> - Our patch kept only the category in the saved error, so an operator
could not diagnose a failure when tracing was off.
> - This pull request preserves redacted provider diagnostics in the run
error, transcript, and structured run result.
> - Operators can now inspect the provider message and any supplied
request ID or stack trace after the run ends.

## Linked Issues or Issue Description

Refs #13889 (the diagnostic gap; this PR does not update the bundled
Claude version).
Refs #14484 (related model-refusal classification; this PR retains
diagnostics for all terminal failure categories).

**What happened?**
An ACP turn failed with only `ACP agent reported a terminal service
failure.` The provider's title and details were available in memory but
absent from the saved error and transcript.

**Expected behavior**
The run retains useful provider diagnostics even when raw tracing is
disabled. Credentials remain redacted. A size limit must report
truncation instead of silently removing the cause.

**Steps to reproduce**
1. Run an ACP agent that returns an error-severity typed session
failure.
2. Include an HTTP error, request ID, and stack text in its title and
details.
3. Inspect the failed run with tracing disabled. Before this change,
only the category survives.

## What Changed

- Both pinned ACPX patches pass complete error text to the in-memory
callback, so redaction happens before truncation.
- The shared engine retains the sanitized category, title, and details
in `resultJson.terminalSessionFailure` and includes the text in the run
error and error transcript.
- Diagnostics redact configured environment values even under arbitrary
names, unknown launch-environment values, connection URL passwords, run
credentials, and common credential syntax. Known boolean settings remain
readable, while credential values are redacted even when embedded in
other text. Diagnostics remove control characters and invalid Unicode.
- Title and detail limits keep escaped transcript JSON below the
server's chunk limit. Truncated fields include an omission count. The
safe run-result projection preserves a byte-bounded diagnostic preview
when the result exceeds its byte budget, with an explicit pointer to the
full adapter-bounded run error and transcript.
- The existing UI and CLI display the error. Diagnostics do not become
assistant output. Issue continuation summaries and session-compaction
prompts receive only the generic category, preventing provider text from
becoming handoff instructions. Existing quota classification, warnings,
timeout precedence, and control-channel failure precedence remain in
place.
- Regression tests cover real ACP child processes with both pinned
versions in one-shot and persistent modes, credential redaction, request
IDs after the old 4 KiB cutoff, transcript parsing, storage bounds, and
database retrieval of oversized multibyte diagnostics.

## Verification

- Full CI on `20ad4f5f1f66c46d2c260e6ad0339cbea607b4cf`: **54 passed, 2
intentionally skipped, no pending or failing checks**. Includes
typechecking, build, all Vitest shards, Runner checks, browser E2E, and
the canary packaging/public-install dry run.
- Greptile: **5/5** on this commit. Superagent security scan passes. All
review threads are resolved.
- Local verification passed: shared ACP engine suite (395 tests); real
Claude ACP child-process and diagnostic regressions across both pinned
runtimes and both execution modes; run retrieval and model-handoff
regressions (59 tests); ACPX patch packaging (16 tests); full typecheck
and build. Affected package typechecks and focused tests were rerun
after review fixes.
- The broad local `pnpm test:run` was stopped after review edits made
its cached imports stale. Fresh targeted runs pass, including both
affected server suites. Cold-build import failures were also rerun after
dependency builds: chat integration (1,063 tests) and tool access (351
tests) pass. The final commit's complete CI matrix is green.

## Risks

- Provider diagnostic text is untrusted. This change retains more of it
in company-scoped run records. Redaction and size bounds apply before
persistence.
- Diagnostics are limited to fields the provider supplies. Old runs
cannot recover discarded error text.
- No schema migration, recovery-policy change, or new Telemetry or
OpenTelemetry export.

## Model Used

- OpenAI GPT-6 through Codex, with reasoning, repository inspection,
code editing, and test execution. The exact serving model ID and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 10:18:30 -05:00
DottaandPaperclip 24beb00575 feat(runner): add rich ACP transport and durable interaction foundation (#14430)
Add shared rich ACP transport, durable questions and permissions, verified provider packaging, and bounded activity and plan presentation. Keep Cursor, Copilot, and Pi pending their separate provider qualification.

Persist interaction settlement before publication, fence failed writes until fresh recovery, and preserve owned-process cleanup. Incorporate reviewed mainline integration with extended harness coverage.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-29 08:56:21 -05:00
0be2afcca6 feat(ui): improve task composer controls and pending input (#14322)
## Thinking Path

> - Paperclip lets operators assign tasks to AI agents and review their
work.
> - The task composer controls the next message and its assigned agent.
> - Operators needed a way to choose that agent's model and effort
without leaving the composer.
> - The old mode selector, upload button, and input cards made the
mobile composer crowded and hid normal messaging during a pending
decision.
> - Harnesses publish different model and effort capabilities, so the
picker must follow the selected agent.
> - This pull request adds one responsive composer flow, keeps pending
cards visible above it, and protects Codex ACP authentication in the
local test path.
> - Operators can choose run settings, send a message, and answer a
pending card as separate actions.

## Linked Issues or Issue Description

**Subsystem affected**

Task composer UI, issue thread interactions, Codex ACP credential
handling, and Storybook.

**Problem or motivation**

The composer did not expose model or effort for the selected agent.
Mobile actions wrapped poorly. Pending questions and confirmations
replaced the composer. A local Codex ACP test could also reuse host
authentication after the managed key was removed.

**Proposed solution**

Put assignee search, model search, exact model IDs, effort, and fast
mode in one picker. Use a mobile dialog. Replace the direct-upload plus
action and separate mode selector with an Add menu and removable Plan or
Ask chips. Place pending interaction cards above the usable composer.
Keep these cards pending after an ordinary message unless their creator
asks for comment superseding. Replace managed ACP auth files atomically
and isolate the test key from host credentials.

**Roadmap alignment**

ROADMAP.md does not list an overlapping composer milestone. This change
improves the existing task and review flows.

## What Changed

- Added the combined assignee, model, and effort picker to both task
composers. Search matches agent name, role, and harness. The server uses
a curated Codex list by default and honors instance-declared models.
Manual IDs remain available.
- Added an effort slider for known model capabilities, a conditional
Codex fast control, and reset. The picker opens in a modal on mobile.
- Added the Add menu for files, supported goals, Plan mode, and Ask
mode. Plan and Ask are exclusive removable chips. Keyboard mode cycling
remains available.
- Adjusted mobile spacing, avatars, wrapping, and Send placement.
Removed the composer divider.
- Moved pending question, confirmation, review, and related cards above
the composer. Ordinary comments now leave question and confirmation
cards pending by default. The onboarding prompt retains explicit comment
superseding.
- Updated the Storybook composer group with responsive states and the
production picker. Added UI, service, route, and browser regression
coverage.
- Isolated Codex ACP API-key authentication, skipped subscription auth
merge and shared-home copy-back for remote API-key runs, and replaced
the managed auth file atomically.

## Verification

- `pnpm -r typecheck` — passed on the final local head.
- `pnpm check:token-gates` — passed on the final local head.
- `pnpm exec vitest run server/src/__tests__/adapter-models.test.ts
ui/src/components/task-chat/ComposerRunSettingsPicker.test.tsx` — 31
tests passed, including role and harness search, declared Codex models,
and filtering general OpenAI models.
- `pnpm exec vitest run
server/src/__tests__/issue-thread-interactions-service.test.ts` — 74
tests passed.
- `pnpm exec vitest run
packages/adapters/codex-local/src/server/acp.test.ts` — 42 tests passed,
including remote API-key copy-back isolation.
- `pnpm test:run` — attempted locally; the embedded PostgreSQL test
database could not initialize on macOS. The isolated
`heartbeat-run-event-sequencing` suite reproduced that environment
failure. GitHub CI runs the full test matrix for this head.
- `pnpm build` — passed on the final head. `pnpm build-storybook` passed
after the last UI change; only server code, tests, and docs changed
afterward.
- Live local test drive — Codex ACP ran a task with a managed API key.
The test agent was restored to its default ACP configuration afterward.
- Review the interactive stories under the top-level Composer group with
`pnpm storybook`. Check a narrow desktop width and mobile Plan, Ask,
picker, and pending-question states.

## Risks

- A pending card stays open when an ordinary comment changes the
discussion. Its creator can set `supersedeOnUserComment: true` when a
new comment should replace it.
- Model and effort overrides persist on the task until reset or changed.
An unlisted manual model ID may fail when the provider runs it.
- Some harness catalogs do not report effort support. The picker hides
effort for those models.
- No database migration is required.

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

OpenAI GPT-6 via Codex. This runtime does not expose the exact model ID
or context window to the task. The model used code execution and browser
tools.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI Codex <codex@openai.com>
2026-09-28 23:02:39 +00:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip 18dac1e1ef feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Connections give agents governed access to external tools.
> - Agents need durable memory across tasks and execution environments.
> - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory
tools.
> - This pull request adds their setup flows behind an experimental
toggle.
> - It also delivers assigned MCP tools through the native remote Codex
runner.
> - Operators can connect a provider once and use the same governed
tools locally or in Daytona.

## Linked Issues or Issue Description

**Problem or motivation**

The Apps catalog lacks a complete set of memory providers. Remote native
Codex agents also need access to assigned managed MCP tools without
receiving provider credentials.

**Proposed solution**

Add five memory connectors behind the disabled-by-default Experimental
memory connectors setting. Use the existing connection setup and
permissions UI. Default their tools to Allowed. Preserve company
boundaries, operator permission changes, provider scopes, and audit
attribution.

**Alternatives considered**

Direct provider credentials in each sandbox would duplicate setup and
bypass the managed gateway. The remote runner instead uses its existing
protocol channel to call the gateway on the server.

**Roadmap alignment**

This maintainer-requested experiment supports the Memory / Knowledge and
Connected Apps roadmap areas. It adds provider connections without
introducing a separate memory UI. Related connector authoring
documentation is tracked in #13692; no duplicate memory-provider
implementation was found.

## What Changed

- Add provider definitions, official branding, and the experimental
setting for all five providers.
- Use OAuth for Zep and Supermemory, API credentials for Mem0 and
Honcho, and a bundled Cloud API bridge for Cognee with no runtime
downloads or subprocesses.
- Default memory tools to Allowed and classify destructive actions
explicitly.
- Fix personal remote credential resolution and propagate provider tool
errors.
- Relay assigned managed MCP tools to remote native Codex through the
runner protocol. Recheck current authority for each call and rotate
stale tool contracts.
- Add 21 Storybook states and complete OAuth walkthrough fixtures.
- Document provider research, sanitized tool inventory, and live local
and Daytona proof.

## Verification

- Passed workspace typecheck: `pnpm -r typecheck`.
- Passed production build: `pnpm build`.
- Full local `pnpm test:run`: 13,273 passed, with failures from
process/readiness timeouts under parallel load. Reran all 16 affected
suites with one worker: 456 passed, leaving two macOS `/var` versus
`/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`.
No test failures remain unverified. Latest-head remote CI passes all 54
checks (two optional Storybook jobs skipped). Greptile is 5/5 with no
unresolved findings.
- Latest Cognee gateway regression: 71 passed, including public
deployment without a runtime host and immediate recovery after a
provider error. Bundled bridge tests: 25 passed.
- Browser setup and real agent tasks exercised all five providers. Mem0,
Cognee, Zep, and Honcho have successful store/retrieve proof.
- Supermemory now has scoped read/write consent. Local storage and real
Daytona write, document read, and semantic recall passed; indexing
completion was verified before claiming success.
- All five providers were exercised through a real Daytona sandbox and
its native runner MCP relay. Provider credentials remained on the
server. The final bundled Cognee bridge also passed a fresh Daytona
store/recall run. All disposable sandboxes were removed and verified
absent after testing.
- Storybook is rebuilt and contains the experimental toggle, catalog,
setup, permissions, and error states. The Zep and Supermemory access
steps advance correctly.

## Risks

- Provider OAuth scopes and plan limits remain independent of Paperclip
tool permissions. An Allowed tool can still be rejected by the provider.
- Providers can queue memory indexing; save acceptance does not prove
that semantic recall is ready.
- Remote tool contracts must stay synchronized with current connection
authority. Regression tests cover revocation and stale contracts.
- This change has no database migration. Existing connections remain
usable when the experimental catalog toggle is disabled.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository
edits, code execution, and browser automation. The exact context window
size is not exposed in this session. Live acceptance agents used
`gpt-5.6-sol`.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-24 15:54:38 -05:00
Devin FoleyandPaperclip 8ee8f1fd6e ci: retire recurring public cloud image builds (#13827)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Core publishes standard images and source verification for
downstream services.
> - Managed services can now compose private images from the signed
standard image.
> - Core still builds a second public cloud image on every master push
and release.
> - That duplicate producer consumes build capacity and retains an
obsolete readiness contract.
> - This pull request retires recurring cloud publication while
preserving the standard producer and rollback artifacts.

## Linked Issues or Issue Description

Refs #13797 and #13789. Related: #12856 changes image dependency
packaging; it does not retire this producer.

**What existing behavior does this improve?**

Core's recurring Docker publication and Cloud readiness workflow.

**Current behavior**

Master pushes call the legacy cloud publisher from Cloud readiness.
Release tags and manual Docker runs call it too. Canary promotion also
requires the legacy image.

**Proposed behavior**

Publish standard Core images and retain `Cloud source verified v1`. Let
downstream services build their managed image. Keep explicit commit
previews and existing images available.

## What Changed

- Remove `docker-cloud.yml`, its master and release callers, and its
unused cache selector.
- Remove the legacy image/migrator wait and `Cloud deployable v1` job.
Keep the full source verification workflow and exact source-proof name.
- Make canary promotion inspect and promote the standard image only.
- Preserve signed standard-image publication, direct migrator
publication, and explicit `release.yml` previews. The preview path still
uses the Dockerfile `cloud` target.
- Update workflow, preview, build-stamp, and packaging tests. Exercise
the promotion shell with mocked registry commands, including
missing-image and missing-tag cases.
- Document frozen legacy aliases, consumer requirements, preview
compatibility, and rollback retention.

## Verification

- All 377 workflow tests pass: `node --test
.github/scripts/tests/*.test.mjs`.
- All 129 release-registry tests pass: `pnpm test:release-registry`.
- Focused source-proof, standard-image, preview, and workflow tests
pass: 256 tests.
- Focused image packaging/build-stamp tests pass: 16 tests.
- Actionlint passes on all three changed workflow files. `git diff
--check` passes.
- Full local `pnpm build` and `pnpm -r typecheck` pass.
- The policy follow-up updates an old assertion that required the
removed readiness job. All 37 source-proof/release-workflow tests pass
locally.
- Full local `pnpm test:run` did not complete successfully while the Mac
ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of
generated Cargo output from this isolated worktree with `cargo clean`.
GitHub CI passed on the final head: 52 successful checks and 2 optional
skips.
- Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`:
**5/5**, successful current-head check, zero review threads.
- September 23 refresh: the unchanged PR head merges cleanly with
current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an
isolated temporary worktree, all 377 workflow tests and 29
release/preview tests pass on the combined tree. `git diff --cached
--check` passes.
- Refreshed Actionlint workflow validation passes with ShellCheck
disabled. Full Actionlint reports the same 10 existing ShellCheck
diagnostics as master, with no added diagnostics. No source changes or
new PR commits were needed.
- The full local build/typecheck and current-head Linux CI results above
remain the verification for the unchanged PR head. They were not rerun
for this metadata-only refresh. No image publication or tenant
deployment was initiated for this refresh.

## Risks

**Deployment prerequisite satisfied (September 23):** The combined
cleanup release is deployed to staging and production, and production
Support is verified. Active managed-fleet automation uses standard-image
composition. Explicit immutable previews remain supported by the
retained preview publisher. This PR is ready for maintainer review; keep
auto-merge disabled and wait for explicit merge authorization.

- A consumer still selecting `Cloud deployable v1` will stop advancing
at the last legacy-ready commit. Confirm active automatic consumers use
the standard-image composition contract before merge.
- Legacy cloud release-channel aliases stop advancing. Standard
self-hosted aliases continue.
- This PR deletes no registry images, cache tags, migrators,
credentials, or runner infrastructure. Existing immutable releases
remain usable for rollback.
- Explicit legacy previews remain for commit-specific operator
deployments. Retiring that compatibility path requires a separate
consumer migration.
- These changes affect CI publication, not database schema or
application behavior.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used repository inspection,
reasoning, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 19:35:19 -07:00
DottaandPaperclip db8f8fe5b7 fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path

> - Paperclip manages agent work through shared runner contracts.
> - Direct protocol evals qualify provider behavior against a mock
control plane.
> - Grok supports API keys and company subscription credentials.
> - The hosted protocol workflow selected an API key for every Grok
cell.
> - Product subscription support did not enable subscription protocol
runs.
> - This change adds explicit subscription selection and checks the
recorded authentication mode.

## Linked Issues or Issue Description

Refs #13878, #13882, #12618.

The direct Grok protocol roster cannot run with subscription
authentication through the trusted default-branch workflow. Add an
explicit selector while keeping API-key dispatches compatible. Keep the
actor allowlist, protected environment, immutable source revisions, and
publication gates.

## What Changed

- Add `grok_authentication` with `api_key` and `subscription` choices.
Keep `api_key` as the compatibility default.
- Deliver the protected `GROK_AUTH_JSON` secret only to a
subscription-selected Grok cell. Do not provide an API key to that cell.
- Read authentication mode from the pinned eval program's actual roster
summary. Retain it in the cell, catalog, campaign roster, and result.
- Reject missing or mismatched authentication evidence during
aggregation. Preserve cell metadata and an allowlisted failure reason
before failing a cell, so malformed evidence cannot hide the retained
attempt.
- Document credential setup, source separation, and temporary-secret
cleanup.

## Verification

- `node --test
packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs
packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed.
- Validated all 39 Grok cells at eval revision
`3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests
only the subscription credential. Validation made zero provider calls.
- Negative coverage rejects invalid selectors and missing or API
authentication evidence in an otherwise passing subscription attempt.
- `git diff --check` passed. All 53 current-head checks passed; the
unchanged callback-drain timing test passed its bounded rerun, and the
failed attempt is retained. Greptile reviewed
`ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining
findings.
- No Docker or broad builds ran on the developer machine. CI performs
repository checks on the configured fleet.

## Risks

Grok runs require an eval revision that records `authenticationMode` in
the roster summary. Missing evidence fails closed. The credential
contains account access and refresh tokens; an owner must approve its
delivery to the protected environment before a live run. The change adds
no PR trigger or authorization bypass. Live subscription protocol
qualification remains pending this workflow reaching master.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 18:43:21 -05:00
DottaandPaperclip 24429024e7 feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external tools through stored
credentials.
> - Routines start work when an external service sends an event.
> - Fireflies provides meeting transcripts and summaries through an
official hosted MCP server.
> - This PR adds that connection and accepts signed meeting events
through the shared app webhook flow.
> - Agents can review completed meetings with the same permissions and
audit records as other work.

## Linked Issues or Issue Description

**Problem or motivation**

Operators need agents to read Fireflies meetings and start follow-up
work when a summary is ready. The Apps catalog lacks Fireflies. The
shared app webhook flow needs to accept its signed deliveries.

**Proposed solution**

Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer
API key. Extend the existing Another app or script flow with signed
webhook support. Verify the raw-body signature and pass the JSON payload
as external data. Select Meeting Summarized in Fireflies. Deduplicate
identical signed deliveries, including setup deliveries.

**Alternatives considered**

A separate REST connector would duplicate the governed MCP path.
Polling, legacy V1 payloads, and automatic provider-side webhook
registration are outside this change.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps and Scheduled Routines
surfaces. It adds a provider to those systems. It does not introduce a
second integration framework.

**Additional context**

A GitHub search found no existing Fireflies issues or PRs. Provider
references and verification limits are in
`doc/connections/FIREFLIES.md`.

## What Changed

- Add the official Fireflies catalog definition, generated registry,
provider evidence, and branded artwork.
- Reuse Access → Connect, dynamic discovery, Permissions, vault storage,
policy, and audit behavior.
- Classify Fireflies sharing, movement, and access revocation as writes.
- Preserve Off and Ask first restrictions during OAuth reauthorization
and API-key replacement. New actions retain normal defaults.
- Add `app_webhook` authentication to the shared Another app or script
flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier
`fireflies_hmac` triggers and revision snapshots for compatibility.
Existing text columns need no migration.
- Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact
request body. Preserve generic event payloads and deduplicate identical
signed requests.
- Keep the routine wizard generic. Show one webhook URL and secret in
Another app or script. Keep all new app webhook event names
provider-neutral. Keep provider setup instructions in the connector
documentation.
- Pass generic webhook JSON to the task in an explicit external-data
block, capped at 16,384 characters. Keep strict meeting validation for
existing legacy Fireflies triggers.

## Verification

- Feature implementation commit `0882dc8a1`: all 54 CI checks passed;
two conditional Storybook checks skipped. This includes full tests,
typecheck, build, browser E2E, canary dry run, and security checks.
Greptile rated this commit 5/5; all review threads are resolved.
- Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed
on the final code. Targeted connector, gateway, webhook, revision, and
UI suites passed during implementation. After the provider-neutral
follow-up, all 84 app-webhook and routine-service tests passed; the
final payload-to-task assertion also passed in the 72-test routine suite
and a clean-config rerun.
- The long local `pnpm test:run` invocation started before the final
edits and was stopped after the final-commit CI suites passed. It
reported one generic webhook test failure while those files were
changing; that test and the entire routine suite passed on the final
source, including a clean-config reproduction. The interrupted local run
is not counted as a full-suite pass.
- In the embedded browser, completed official OAuth consent and
discovered 20 live actions. Real meeting listing, transcript retrieval,
and summary/action-item retrieval succeeded as the selected agent.
Turning a live read Off blocked its test; catalog refresh preserved the
restriction.
- Embedded-browser Another app or script setup, back/save/resume, narrow
layout, and a signed synthetic Fireflies delivery succeeded. The UI
reported authentication passed without creating a task. Fixtures cover
signature tampering, malformed requests, ordinary app event names,
duplicate/setup deliveries, rotation, revisions, pause/archive, and
company isolation.
- Existing MCP browser suite: 8 passed and 2 provider-dependent cases
skipped. Branding checks passed; connector artwork and webhook setup
were checked at desktop/mobile widths and in light/dark modes.
- An unauthenticated POST to a correctly formatted public webhook URL
reached the staging tenant verifier through the existing Cloud gateway.
- A real Fireflies webhook delivery remains unverified. A staging
callback is available for the operator walkthrough. Live API-key
authorization, credential expiry, and a new meeting's summary completion
were not tested against the provider. Fixtures cover these protocol and
lifecycle paths where applicable.

- Storybook follow-up `c54174faa`: 27 production-component stories cover
every UI change, with a source-to-story map in the connector
documentation. Static Storybook build, UI typecheck, token gates, and
Playwright checks for all stories and the mobile footer pass. All PR
checks passed for this Storybook follow-up; Greptile reviewed
`c54174faa` at 5/5.

## Risks

- Fireflies may change its hosted MCP tools or OAuth behavior. Tool
discovery stays dynamic. Experimental search/fetch tools are not
required.
- Public webhook setup requires HTTPS and a separate signing secret.
Fireflies normally emits events for meetings owned by the configuring
account.
- Reauthorization touches shared MCP permission code. Regression tests
cover existing restrictions, new actions, connection removal, and other
gateway callers.
- Webhook receipt grants no tool access. The routine agent still needs
an authorized Fireflies connection.

- New generic triggers rely on provider event subscriptions. Without a
sender-supplied idempotency key, changed request bytes count as a new
event. Existing legacy Fireflies triggers retain summary-only filtering
and per-meeting deduplication.

## Model Used

OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing,
code execution, and embedded-browser testing. The runtime did not expose
a context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:11:16 -05:00
DottaandPaperclip b648d8cdda fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path

> - Paperclip manages AI agents and their provider connections.
> - Product E2E checks real tasks through the browser, server, and
runner.
> - Grok qualification needs separate API-key and subscription evidence.
> - Product subscription tests and direct Grok protocol evals need
explicit credential delivery.
> - This change supplies each credential only to its selected profile
and prepares the pinned binary.
> - Maintainer authorization and protected-environment gates remain
required.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13882.

The Grok feature branch has a manual subscription qualification profile.
The trusted master workflow must admit its selected credential and
prepare the same verified binary and artifact verifier as the API
profile. Direct protocol evals also need the selected xAI key and pinned
Grok binary. These prerequisites do not register or schedule the new
profiles on master.

## What Changed

- Deliver `GROK_AUTH_JSON` from the protected paid environment only when
the selected profile requests that credential.
- Install the checksum-verified Grok binary for the local subscription
profile.
- Prepare the pinned artifact verifier for the manual subscription
suite.
- Extend workflow security assertions to cover the new credential and
profile.
- Add the ACPX Grok credential mapping to the trusted-master catalog,
then deliver only the selected `XAI_API_KEY` to direct protocol cells
and install the target’s checksum-verified Grok binary before packaging.
- Allow a direct-protocol concurrency override from two cases up to the
existing configured ceiling; it can only lower concurrency.
- Document the Grok protocol workflow and its API-only credential
boundary.
- Render missing LLM usage and cost as Unavailable, and label partial
observations with coverage. Preserve raw records, grades, and actual
zero costs.
- Preserve measured campaign source metadata during report regeneration
instead of inheriting the renderer checkout or CI event; skip empty
legacy source records when recovering older provenance.

## Verification

- Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54
reported checks successful, two intentional skips, Greptile 5/5, and
zero unresolved review threads. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35890978288).

- After merging current master, all 17 workflow security/image tests and
23 catalog/workflow policy tests passed. The trusted catalog also
generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per
shard. The new policy tests execute the concurrency guard against valid,
out-of-range, and malformed values.
- The Grok branch separately passed 450 Product harness unit tests,
including private company credential staging, cleanup, and
token-fragment redaction.
- The fresh-login native subscription smoke passed three repetitions of
tool execution, session resume, restrictive permissions, and cleanup.
These are setup evidence; full subscription Product qualification
remains pending.
- All 72 focused report/billing/history/catalog tests and the Product
harness typecheck passed for the report-display change. The initial
sandbox run could not open the tsx IPC socket; the permitted rerun
passed. A zero-provider-call replay of the actual 16-cell Grok campaign
preserved all result records, grades, timing, and source provenance
while correcting missing usage labels.
- Review the thirteen-file diff. Provider credentials still enter only
the selected paid-test step; default-branch, numeric-actor, and
environment restrictions are unchanged.

## Risks

This admits a refreshable subscription credential to explicitly selected
trusted tests. Store it only in `runner-e2e-paid`, use a test login, and
remove it after qualification. Unselected profiles receive an empty
value. Pull requests cannot trigger the paid workflow. This PR changes
no fleet admission or actor allowlist.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:05:52 -05:00
Devin FoleyandPaperclip be6f49a425 feat(runner): refresh shared coding harness runtimes (#13838)
## Thinking Path

> - Paperclip runs agents through local adapters and the native runner.
> - Both paths must use the same installed provider CLI.
> - New models require current harness releases.
> - The runner still pins Codex 0.153.4, Claude SDK 0.3.263, and
OpenCode 1.18.29.
> - Changing the image alone would fail the runner's exact version and
executable checks.
> - This pull request updates those dependencies, integrity checks,
controller checks, and image pins together.
> - Shared installations can then run the current models without a
task-time download.

## Linked Issues or Issue Description

Refs #13829, which updates model choices and reasoning controls.
Searches found no open PR that updates these runtime pins.

**Current behavior**

The shared provider pack ships old CLIs. Claude Code 2.1.263 cannot run
Opus 5.5, which requires 2.1.280. Remote controllers reject provider
packs whose versions differ from their declared pins.

**Proposed behavior**

Use Codex 0.156.0, Claude Agent SDK 0.3.280 / Claude Code 2.1.280, and
OpenCode 1.18.32 throughout the runner. Keep the reviewed ACP bridge
patches and one shared CLI installation per provider.

**Reason and benefit**

Current harnesses support the new model IDs while preserving executable
verification and remote provider-pack compatibility checks.

## What Changed

- Update dependency overrides, the Codex ACP package patch, runtime
profiles, and remote controller pins.
- Verify the new Claude Linux x64 and macOS arm64/x64 executables and
Codex Linux x64 executable against integrity-verified npm archives.
- Refresh OpenCode version checks, fixtures, and the runner
configuration label.
- Refresh the eval image's Grok, Gemini, Kimi, Cursor, and GitHub CLI
pins and archive hashes. Hermes remains current at 0.19.0.
- Refresh the build-time lock digest from clean pnpm 9.15.4 resolution.
Leave lockfile commits to repository automation.
- Document model compatibility and the separation between CLI runtimes
and patched ACP bridges.

## Verification

- `pnpm -r typecheck` and `pnpm build` passed.
- Rust workspace release tests passed.
- Package/patch and OpenCode binary-materialization contract tests: 11
passed.
- Real Codex 0.156.0 startup-ownership and paginated session-resume
probes passed with isolated synthetic homes and no model turn.
- Codex app-server `thread/start` preserved `gpt-6-sol` and
`gpt-6-luna`; no `turn/start` was sent. An unauthenticated built-in
catalog does not include those account-served entries.
- Installed Claude integrity probes passed for `claude-opus-5-5` and
`claude-fable-5-1`.
- `pnpm --filter @paperclipai/paperclip-runner
test:opencode:qualification` passed with the actual OpenCode 1.18.32
executable under Node 24 and Node 25. The loopback provider exercise
covers health/version, session creation/read/delete, SSE, and a
completed async prompt.
- `pnpm check:token-gates` passed.
- The targeted runner suite passed 130 tests. Three macOS failures in
snapshot module lookup and OpenCode final-message selection also
reproduce on the unchanged base; Linux CI will provide the platform
check.
- [Final Linux
CI](https://github.com/paperclipai/paperclip/actions/runs/35798076399):
all gates passed. Four jobs needed one retry after their CI workers
received shutdown signals. The PR has 55 successful checks, two skipped
checks, Greptile 5/5, and no unresolved review threads.
- Changed runner configuration UI tests: 5 passed.
- Full macOS `pnpm test:run` reached 13,094 passing server tests, 84
skipped, and 18 failures before the wrapper stopped. Failures involved
skill-cache publication permissions, missing bundled connector skills in
the worktree, and a conversation-reset timing case. The 10 cache
permission failures reproduce on the unchanged base; both
conversation-reset cases passed on a targeted retry. The wrapper did not
reach its later workspace/serialized groups locally; Linux CI covers
those groups.
- The local Docker daemon did not respond, so no local Docker build was
run. No billable model requests were made.

## Risks

- Deploy the matching controller and provider pack together. Older
controllers enforce their previous exact pins.
- Current upstream CLIs can change behavior. Existing protocol tests and
isolated real Codex probes cover the integration boundaries;
authenticated model inference is not part of these checks.
- ACP bridge package versions and executable digests stay unchanged
because their executable bytes are unchanged. Only the underlying
CLI/SDK dependencies move.
- No schema migration. Revert the runtime and image pins together to
roll back.

## Model Used

OpenAI GPT-6 via Codex, with repository tools, code execution, and web
research. The exact serving model ID and context window were not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass for the changed surfaces
and real-executable probes; full macOS-suite limitations are listed
above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 17:02:29 -07:00
Devin FoleyandPaperclip e3d8fb0876 feat: attest standard production images at full source commits (#13797)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators deploy its standard production container on several CPU
architectures.
> - Downstream image builders need to identify the exact source of their
base image.
> - A short commit tag does not provide signed source evidence.
> - This pull request adds a full commit tag and signed image digest for
canonical master pushes.
> - Consumers can verify the source and compose from the immutable
digest.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Publication of the standard multi-platform production image.

**Current behavior**

The Docker workflow publishes short commit tags and channel tags. It
does not provide a signed standard-image contract tied to the complete
master commit.

**Proposed behavior**

Canonical master pushes also publish `sha-<full-commit>` and attest the
exact index digest after platform validation and an immutable-image
orphan-reaping check. The signer certificate binds the source
repository, commit, workflow and ref. Existing tags and the separate
cloud producer remain available.

**Reason and benefit**

Downstream builders can prove the source of a standard base without
adding their dependencies or repository details to the public workflow.
No matching open issue or duplicate PR was found.

## What Changed

- Add the canonical full-SHA tag without changing existing tag mappings.
- Validate amd64 and arm64 descriptors and hash the exact registry
response bytes and require its digest header to match.
- Verify the immutable image and sign it with GitHub artifact
attestations.
- Run the contract tests in trusted PR verification and document the
consumer contract.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/standard-image-contract.test.mjs`: 37 passed.
- A read-only check against an existing published index returned its
exact expected digest.
- `actionlint -shellcheck='' .github/workflows/docker.yml
.github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports
only existing `ls` and word-splitting warnings.
- `pnpm build`: passed locally with Cargo available.
- `pnpm -r typecheck`: passed locally.
- Full local Vitest was attempted: 8,260 passed, 14 failed, with 34
failing suites. The failures were missing embedded-PostgreSQL library
aliases in this fresh install and existing macOS runtime-skill-cache
rename errors. Native aliases are now restored. Rerunning the 33
affected database suites produced 577 passes and two unrelated AgentMail
skill-root lookup failures (32 suites passed). The four directly failing
database tests also pass independently. This is not a claim that the
full local suite passed.
- All final-head CI checks pass. One unrelated routine-route mock
assertion passed on the single-shard retry; its 15 tests also pass
locally. Greptile is 5/5 on this exact head, with no unresolved threads.
- Actual signing requires a canonical master push. This draft PR does
not publish trusted provenance.

## Risks

The new attestation step requires OIDC and attestation write permissions
in the merge job. Signing failure leaves the image available but without
the new admission proof. Consumers must fail closed when proof is
missing. Existing release tags, the legacy producer, and image retention
remain unchanged. No database or application behavior changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell
execution and test tools. The session does not expose a more specific
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 07:33:30 -07:00
3c2bc4f546 ci: halve the isolated native Runner check by seeding it from the public master build cache (#13736)
## What

Recurring CI health check (PAP-31): on the most recent fully-green PR
run (`cb703ac`, run
[35519542997](https://github.com/paperclipai/paperclip/actions/runs/35519542997)),
the slowest check was **Compile isolated native Runner** at **452s** —
ahead of the largest test shards (379s). Every run recompiled the full
Rust dependency tree from zero, even though `docker.yml` already
refreshes a **public** `mode=max` BuildKit cache
(`ghcr.io/paperclipai/paperclip:buildcache-{amd64,arm64}`) on every
master push, containing exactly these layers. This PR seeds **only the
baseline build** with that registry cache, via anonymous pull.

**Measured on this PR's own CI (which exercises the seeded path): the
check completed in 123s, down from 452s — a 73% reduction, ~5.5 minutes
saved per run.**

## Thinking Path

Cost breakdown of the 452s from the job log: `cargo chef cook`
dependency compile 214.7s, local cache export 43.1s, runner-core build
37.0s, metadata proof layer 36.8s, `cargo install cargo-chef` 36.4s,
rebuild-verification build ~65s, setup/teardown ~20s. The dependency
compile and toolchain layers are identical to what the production
`docker.yml` build already caches publicly on every master push, so
recompiling them here bought no signal — the check's real assertions
live in the *verification* build, not the baseline. A first attempt used
`actions/cache` plus a master `push` trigger, but the CI bot's GitHub
App lacks `workflows` permission; the registry-cache approach is
strictly better anyway (shared across PRs immediately, no 10GB
Actions-cache quota pressure, no workflow change).

## What Changed

- `scripts/check-docker-runner-cache.sh`: the baseline build now adds
`--cache-from
type=registry,ref=ghcr.io/paperclipai/paperclip:buildcache-{amd64|arm64}`
(selected by host arch). `RUNNER_CHECK_SEED_CACHE` overrides the ref, or
set it empty to force the old cold path. The script header documents the
anonymous external read.
- `.github/workflows/docker-runner-check.yml` (comment-only): the stale
"no external cache" note now describes the anonymous GHCR seed and the
verification build's local-cache-only isolation. This was pushed in a
follow-up commit with workflow-edit permissions; the original CI-bot
token could not touch workflow files.

No Dockerfile stages or verification assertions changed.

## Verification

- This PR's own `Compile isolated native Runner` check runs the seeded
path (the script is in the workflow's trigger paths): **passed in 123s**
vs the 452s baseline.
- The rebuild-verification semantics are untouched: it still runs on a
**fresh builder** importing **only the local cache exported by this
run's baseline**, so it proves exactly what it proved before — that the
runner image rebuilds reproducibly from this run's own exported layers.
- Verified `ghcr.io/paperclipai/paperclip:buildcache-amd64` is
anonymously readable (unauthenticated manifest pull succeeds), so the
check gains no credential or secret dependency.

## Risks

- **Stale or missing seed cache:** if the GHCR ref is unreachable,
private, or garbage-collected, BuildKit logs a warning and falls back to
the pre-PR cold compile — the check gets slower, never wrong.
`RUNNER_CHECK_SEED_CACHE=""` restores the cold path explicitly.
- **Cache trust:** the seed only accelerates the *baseline* build; the
verification build still runs on a fresh builder against only this run's
locally exported cache, so a stale or poisoned registry cache cannot
make verification pass spuriously. The ref lives under
`ghcr.io/paperclipai/*`, written only by repo CI on master pushes.

## Model Used

Claude Fable 5 (`claude-fable-5`) via Paperclip agent **Bender
(Fable)**, issue PAP-31.

---------

Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-21 20:22:48 -07:00
DottaandPaperclip b82661b561 refactor(connections): retire the legacy Composio broker (#13758)
## Thinking Path

> - Paperclip manages agents and their access to external tools.
> - Connectors expose these tools through a governed MCP gateway.
> - PR #13755 added a direct Composio MCP connection behind the
experimental MCP aggregators flag.
> - The old project API-key broker still created toolkit child
connections and showed a separate Services tab.
> - Keeping both paths leaves obsolete setup and session code in the
product.
> - This change removes the broker and preserves direct MCP setup,
credentials, permissions, and execution.
> - Saved legacy records fail closed and remain available for explicit
removal.

## Linked Issues or Issue Description

Related: #13755. This retirement supersedes the legacy-path fixes
proposed in #12630, #12632, #12634, and #12906. It does not close those
PRs.

**What existing behavior does this improve?**

Composio connector setup, management, and runtime dispatch.

**Current behavior**

Composio offers both direct MCP and a project API-key broker. The broker
mints sessions and creates one child connection per toolkit.

**Proposed behavior**

Offer only direct MCP. Remove the toolkit Services UI, REST routes, API
client, and session broker. Block saved legacy parent and child records
from discovery, execution, health checks, reconnect, and OAuth. Preserve
their records and credentials until the operator removes each
connection.

**Reason and benefit**

The direct MCP connector becomes the single supported Composio workflow.
Provider accounts remain managed in Composio.

## What Changed

- Remove the API-key catalog method and its generated-source definition.
- Delete Composio broker clients, session creation, account
synchronization, child lifecycle, and toolkit routes.
- Remove the Services tab, service rows, child provenance, and
cascade-removal controls. Keep Vercel provenance intact.
- Retain a shared retirement guard for stored legacy records. Show
Retired status and replacement/removal guidance in the connection list
and details; hide obsolete runtime controls.
- Preserve the experimental MCP aggregators flag and direct MCP
infrastructure.
- Replace broker fixtures with retirement tests and extend direct
Composio catalog/reconnect coverage.

## Verification

- Focused shared, server, and UI tests passed with one worker. Server
retirement tests use a name filter; no full local test suite was run, as
requested.
- Server and UI TypeScript checks passed.
- Token gates and UI build passed.
- Real browser: opened the saved Composio connection, refreshed all 11
tools, and ran the provider's read-only GitHub account-list operation
through the standard Test dialog as an agent. The provider returned
success using the existing OAuth credentials.
- See `doc/connections/COMPOSIO-BROKER-RETIREMENT.md` for scope and live
evidence.
- Storybook build passed. A fresh real agent used
`COMPOSIO_SEARCH_TOOLS` and `COMPOSIO_MULTI_EXECUTE_TOOL` to return the
actual Paperclip DeepWiki hierarchy: one success, zero errors. Gateway
audit records confirm both calls succeeded.
- Browser retirement check: a credential-free legacy fixture showed the
guidance, opened the direct MCP replacement flow, and was removed
through the standard confirmation.
- Focused regressions for the experimental settings copy and exact
OpenAPI route coverage passed. All latest-head CI checks passed (54
successful, two intentionally skipped); Greptile scored 5/5 with no
unresolved review threads. The PR has no merge conflicts.

## Risks

This intentionally breaks the old Composio project API-key and
child-connection workflow. Existing legacy records cannot run, even if
their stored status is active. Operators must create a new direct MCP
connection and choose access rules; credentials and grants are not
migrated. Remove each old record separately to delete its credentials.
No schema migration or data deletion runs automatically. Direct MCP
connections keep their existing grants and secrets.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, code execution, and browser
tools. The exact runtime variant and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 13:57:35 -05:00
DottaandPaperclip e8c8ba3c19 feat(apps): add experimental MCP aggregator connectors (#13755)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its tool gateway applies company access rules and approval controls
to connected apps.
> - MCP aggregators expose many apps through one provider endpoint.
> - Each aggregator needs its own credential, catalog, grants, and
lifecycle in Paperclip.
> - This pull request adds independent Zapier, Arcade, Composio Connect,
and Executor setup with a common Access → Connect layout.
> - A default-off MCP aggregators flag lets operators opt in while we
complete provider acceptance tests.
> - Agents use the normal Paperclip permissions, Test screen, and
gateway after setup.

## Linked Issues or Issue Description

**Subsystem affected**

Apps, connection setup, shared contracts, and the remote MCP gateway.

**Problem or motivation**

Aggregator endpoints need clear provider setup and correct MCP sessions.
Generic setup does not explain each provider's authentication or broad
execution tools. Provider approval must preserve the original execution
instead of replaying a write.

**Proposed solution**

Add four separate connectors behind Settings → Experimental → MCP
aggregators. Start with human and agent access, then connect the
endpoint and read its tools. Enable tools by default. Use the existing
Permissions and Test screens after setup. Keep legacy Composio API-key
and child connections intact.

**Alternatives considered**

A shared connection for all providers would mix credentials and access
rules. Separate provider-specific permission and test screens would
duplicate existing controls. Vercel Connect is outside this change.

**Roadmap alignment**

Extends the existing MCP Tool Gateway & Apps capability and the
Connected Apps roadmap area. This work was requested and reviewed by the
maintainer.

Related work: #11894, #12630, #12632, #12634, and #12906 concern the
legacy Composio broker. #13102 also covers remote MCP pagination. This
change preserves the broker path and adds initialized sessions, response
matching, and provider resume handling alongside pagination.

## What Changed

- Add branded setup and interactive Storybooks for Zapier, Arcade,
Composio Connect, and Executor. Use the existing access controls and
normal action tests. Do not request a connection name or action choices
during setup.
- Add the default-off `enableMcpAggregators` flag to settings, managed
feature metadata, the catalog, and setup guards. Hidden connections keep
running. Legacy Composio connections remain unchanged.
- Reuse the vault, grants, policy, and catalog models. Support OAuth
discovery, bearer tokens, custom headers, and credential-bearing URLs.
Add no database tables or migrations.
- Initialize and retain Streamable HTTP sessions by connection and
effective credentials. Read paginated catalogs and match streaming
responses to request IDs.
- Classify unfamiliar aggregator tools as writes despite upstream
read-only hints; only exact reviewed read capabilities enter the
read-only allowlist. Legacy Composio child behavior is preserved.
- Preserve provider authorization links and execution IDs. Support
Executor approve/resume, decline, and cancel without automatic replay of
uncertain writes.
- Preserve Off and Ask first choices during refresh and reconnect. Allow
new tools and retire removed tools. Keep agent access updates atomic and
preserve an empty agent selection.
- Document connector UX rules, provider branding sources, and live
acceptance results.
- Stabilize the existing Sentry release fixture after its repeated CI
failure by reusing one module mock; production Sentry behavior is
unchanged.

## Verification

- Final head `d11781970`: [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35633534900)
passed, including broad typecheck, test shards, build, and E2E. All 54
checks pass; 2 optional checks are skipped. Greptile is 5/5, Security
Scan passes, and all review threads are resolved.

- Passed 27 focused connector Vitest checks and 18 connector-only
Storybook browser checks before the flag change. All 85 stories rendered
at desktop and narrow widths.
- Passed 5 connector lifecycle/server checks and 7 selected flag checks
after adding the flag. The latter cover settings, managed defaults,
cached catalog visibility, and all four setup routes.
- Review fixes passed 13 risk/handoff/lifecycle checks, dedicated
session-expiration and transport regressions, 13 selected
connector/gateway CI cases, and 10 selected setup/reconnect UI cases. A
real Composio connection-list call also succeeded through the refreshed
UI on `9ab115f71`.
- UI and server TypeScript checks passed. UI build, Storybook build,
token gates, and diff whitespace checks passed during implementation.
- Real browser and real Paperclip agent tests passed for Arcade,
Composio, and Executor. Tested action permissions, denied agent access,
reconnect, disconnect, and isolation. Tested Arcade catalog
additions/removal and Executor provider approve/resume, decline, and
cancel.
- Zapier live acceptance is incomplete. Its dedicated provider server is
configured, but its credential-copy dialog returned an empty clipboard
through browser automation. No live Zapier action is claimed.
- The three isolated Sentry release cases pass after the CI fixture fix.
- Local verification is deliberately narrow at the maintainer's request.
The full local suite, recursive typecheck, and repository-wide build
were not run. CI provides the broader checks.

## Risks

- Shared MCP transport changes affect other remote MCP servers. Protocol
fixtures cover initialized sessions, streaming response matching,
pagination, and isolation.
- Broad execution tools remain broad permissions. The provider governs
actions inside those tools.
- Provider handoff links are retained briefly in memory. After a server
restart, a one-time link may require reopening the provider dashboard.
Paperclip does not replay the original call.
- Zapier remains unproven live. Custom-header imports and self-hosted
endpoints have fixture coverage rather than a separate live account for
every variant.
- Turning the experimental flag off hides setup; it does not revoke
existing credentials or stop existing connections.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, shell
execution, and browser automation. The exact runtime model ID and
context-window size are not exposed in this session. A separate
Anthropic-backed Paperclip agent performed live gateway acceptance
tasks.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-21 12:53:11 -05:00
86b7ee992c feat(onboarding): ClipLab sleepy-to-wake hero and step hand-offs (#13629)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents have a persistent visual identity (#13171): a ClipLab
character in one of 17 palettes, rendered as cached PNGs in lists and as
a live character in larger placements.
> - The onboarding wizard is where a person meets that identity first,
and it showed a stock ClipLab expression on the previous engine while
the rest of the app would show a different character on a newer one.
> - The wizard's steps also cut from one screen to the next, so the arc
read as separate pages rather than one walk.
> - This pull request puts one character on one engine everywhere, gives
the wizard's hero the studio's sleepy → wink → idle sequence on Review,
and hands the steps over inside one presence.
> - The benefit is that what wakes on Review is exactly what the agent
looks like on the dashboard afterwards, and the walk to it reads as one
screen changing.

## Linked Issues or Issue Description

Refs #13171, now merged into master. This PR contains the onboarding and
ClipLab update on top of that foundation. Original feature work by
@tonio-alucema; merge preparation preserves the original commits.

**Problem or motivation**

The onboarding hero and the app's avatars were two different characters
on two different ClipLab engines. Steps 1 → 4 of the wizard cut between
screens, and the wizard mounted cold when a cloud-managed workspace
arrived from Cloud's naming screen.

**Proposed solution**

Vendor ClipLab v0.2.0 as the shared engine and render one
studio-exported character from it in every palette, for every pose and
size. Play the export's one-shot wake on Review with the palette fading
in over the gray dormant loop. Hand steps over inside one presence so
the footer slides instead of jumping, and play the arrival half of that
hand-off when the wizard opens directly on the agent step.

**Alternatives considered**

Exporting mp4/webm loops per size: no cursor following, no clean alpha,
and the palette "colour in" is a runtime blend. Minting a `cap-v2`
character version: nothing had shipped `cap-v1`, so the artwork is
regenerated in place instead of migrated. Keeping the separately
vendored runtime bundle for the hero: two engines and two characters in
one app.

## What Changed

- `packages/shared/src/cliplab`: re-vendored from ClipLab v0.2.0
(`987b6db0`) with the Paperclip adaptations replayed (optional graphics
backend for the Node SVG snapshot path, supersampled live textures,
character framing, deterministic SVG id prefixes); new upstream
`particles.ts`.
- `packages/shared/src/cliplab/character.ts`: the studio export,
mirrored from `ui/src/assets/cliplab/onboarding.character.json` by
`scripts/sync-cliplab-character.mjs` (drift caught by
`check:token-gates`). `characterDefinition` builds every palette from
it; the resting portrait is its idle beat.
- `OnboardingCharacter`: gray `sleepy` loop through the agent and
connect steps; on Review the one-shot sleepy → wink → idle plays on two
lock-step canvases while the palette fades in, then the `idle` loop.
Body-follows the pointer, page-scoped. 160px in the wizard.
- `OnboardingWizard`: steps 1 → 2 → 3 → 4 hand over inside one
`AnimatePresence` (departing content fades and gives its room back;
arriving content opens its room then fills); the hero has a room that
opens on the walk into the agent step; opening directly on the agent
step plays the arrival half; the self-hosted naming step uses the arc's
label and field.
- Motion vocabulary in `onboarding-motion.ts` (`stepContentMotion`,
`ledeMotion`, `heroRoomMotion`, `heroRoomArrival`, `titleSwapMotion`).
- Storybook: `Onboarding / Character` (Wake Up), `Onboarding / Agent
arc` walkable from the naming step plus `Arrive From Cloud`; the
companies fixture answers the wizard's create call with a company.
- Uses the shared runtime for onboarding; `doc/agent-personas.md`
documents the shared character.
- Releases both onboarding canvases after partial startup or transition
failure. Registers each canvas before seeking so synchronous render
errors can release it. Six component tests cover these failures and
palette changes before or during wake.
- Refreshes both sleeping canvases when the palette changes, including a
palette change in the same render as wake.
- Moves choreography values into the CSS token layer and preserves the
shared motion catalog drift check across the imported stylesheet.
- Repairs the static Storybook avatar route and uses accessible heading
names/current button labels in the wizard play functions.
- Closes the lazy avatar worker pool during application shutdown.

## Verification

- Merge-preparation checks: `pnpm -r typecheck`, `pnpm build`, `pnpm
build-storybook`, and `pnpm check:token-gates` pass. The final UI
typecheck and 123 focused onboarding, lifecycle, and token catalog tests
pass. All 55 checks on final head
`b4f5e201a1564083d163abc6f93f5b3da06ccefd` pass, including the full
sharded test suite, runner verification, and all eight browser shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/35445430535)).
The duplicate monolithic local `pnpm test:run` was stopped after CI
completed; it is not claimed as a separate completed local run.
- Chromium walkthrough: palette change, wake, return to sleep, WebGL
failure fallback, Review step hand-offs and cloud arrival pass with
normal and reduced motion; no browser errors. The signoff happy-path
browser test also passes against a disposable instance.
- The final CI run confirms the catalog fix and a passing signoff
browser shard. The earlier signoff failure was a heartbeat-run
availability timeout; the focused local reproduction and final CI passed
without signoff code changes.
- Original author verification:
- `pnpm check:token-gates` (includes the new character sync check);
shared, server avatar/persona (17) and UI onboarding/persona (137)
suites pass; `pnpm build-storybook` packages all 3,564 avatar PNGs
through the worker pipeline.
- Storybook: `Agents / Personas` Sizes, Expressions and Palettes render
the studio character at every size and pose; `Onboarding / Character →
Wake Up` plays the wake on the shared engine; `Onboarding / Agent arc`
walks 1 → 4 with the hand-offs, and `Arrive From Cloud` plays the
arrival (measured: content room 6 → 65px over 320ms, fade to 1.0 by
~560ms, footer travel continuous).
- The original author walked the agent → connect → review flow and wake
after a real sign-in on staging.
- Not done here: the Linux Storybook visual baselines
(`tests/storybook-visual/agent-personas.spec.ts`) need re-baselining for
the new engine, hero size and naming-step changes.

## Risks

- Every avatar's pixels change (new engine, new character) under the
unchanged `cap-v1` name. Stacks that rendered avatars on the previous
engine keep those PNGs in their cache
(`generated-agent-avatars/cap-v1/...`, served immutable) until cleared;
only the two pinned staging stacks ever did.
- The one-shot handoff to the idle loop is timed from the sequence's
authored duration (the engine reports completion by continuing into idle
itself); presentation only, nothing in the wizard's state waits on it.
- Reduced motion skips the wake and the hand-offs; jsdom is treated the
same way, so the wizard tests see the next step's content immediately.
- The committed export differs from the studio by one animation (Loop
off, leading idle step removed); a re-export without that fix would play
a 5.6s idle before the wake.

## Model Used

Original feature: Anthropic Claude Fable 5.1 (`claude-fable-5-1`) in
Claude Code, with shell, browser, and file tools. The original context
window was not recorded.

Merge preparation and lifecycle regression fixes: OpenAI GPT-6 in Codex,
with reasoning, shell execution, file editing, GitHub CLI, and automated
tests. The session does not expose an exact runtime model ID or
context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Dotta <bippadotta@protonmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 08:30:14 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
Devin FoleyandPaperclip d54b750111 Preserve Claude ACP quota classification and reset time (#13651)
Typed Claude ACP quota failures lost their recovery classification and reset
time when the runtime reduced provider metadata to a generic category error.
Inspect terminal metadata in memory and retain only safe recovery labels and
a parsed reset timestamp. Preserve the existing handling of other limits.

Verified real child processes on both pinned ACPX runtimes, adapter and
server recovery regressions, all PR CI gates, and Greptile 5/5. Also isolate
a pre-existing chat regression from unrelated fixtures’ retry work.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-18 19:42:28 -07:00
352153b5ed fix: run the cargo-building native-runner CI suite in the Rust-cached vitest lane (#13586)
## Thinking Path

Post-merge of #13557, the slowest check on the freshest fully-green PR
run
([35246999382](https://github.com/paperclipai/paperclip/actions/runs/35246999382))
was `ci / General tests (server (1/12))` at **339s**. The cause is one
suite:
`server/src/services/native-runtime/native-codex-runner.integration.test.ts`
runs 1 test in **277s of a 291s vitest step (95%)** because its
`beforeAll` cargo-builds the Runner release binaries, and the
general-server shards carry no Rust cache — every PR run cold-compiles
the full third-party crate graph. The other 19 suites in that shard
finish in under 70ms each.

The obvious fix (a dedicated Rust-cached matrix lane) requires editing
workflow files, which the available GitHub App credentials cannot push
(`workflows` permission). But the `Verify Paperclip Runner` lanes
**already restore the shared `release-runner-v1` Rust cache read-only**,
and their commands are `pnpm --filter @paperclipai/paperclip-runner
<package script>` — so the suite can move into a Rust-cached lane purely
through script changes.

## What Changed

- `scripts/run-vitest-stable.mjs`: new `general-server-native-runner`
group carrying exactly that suite. Under the PR workflow
(`GITHUB_WORKFLOW == "PR"`, inherited from `pr.yml` by the reusable
`pr-trusted.yml`) the without-chat server shards exclude it and
rebalance to ~211s of tests each. Every other caller — local runs,
`release-verify.yml` under the Release / Cloud readiness workflows —
keeps the suite in the shards, so a renamed or unknown workflow degrades
to today's slower-but-covered behavior instead of dropping coverage.
- `packages/paperclip-runner`: `test:typescript:vitest` now routes
through `scripts/run-pr-vitest-lane.mjs` — the identical
`ensure:eval-build-deps && build:rust && vitest run` chain (shard flags
passed through), plus the native-runner group on the **final PR shard
only** (`--shard=N/M` with `N == M`, i.e. today's `vitest 2/2`, the 122s
lane). With the restored cache the suite's cargo build becomes an
incremental rebuild.
- `scripts/__tests__/run-vitest-stable-shard.test.mjs`: guards pin the
whole contract — PR 12-shard coverage (shards + chat + native-runner =
full server group exactly), Release/local 10-shard runs keep the suite,
`pr.yml` is named `PR`, the vitest lanes partition with exactly one
final shard, the package-script wiring, and the wrapper's shard/workflow
gating via its `--dry-run` plan output.

No workflow files change. `.github/workflows/*` are untouched.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: **36/36 pass**
locally on this branch (includes the new coverage, wiring, and
wrapper-gating guards). Both files run in CI's `Test general-server
shard partition` / `Test release verify workflow wiring` steps.
- Wrapper `--dry-run` plan matrix verified for all six shard/workflow
combinations plus malformed-shard rejection (pinned as a guard test).
- The executing proof is this PR's own CI: `ci / Verify Paperclip Runner
(vitest 2/2)` must go green while running the native-runner suite (its
log will show the `general-server-native-runner` group after the package
vitest shard), and the 12 `ci / General tests (server (x/12))` shards
must go green without it.

## Risks

- The exclusion keys on `GITHUB_WORKFLOW == "PR"`. Failure mode of a
rename is safe (suite falls back into the server shards, slower but
covered) and the guard test on `pr.yml`'s name makes it loud.
- `vitest 2/2` grows from ~122s to an expected ~210–260s — still well
under the ~306s `vitest 1/2` and ~326s e2e shards, and inside the
20-minute lane timeout. If the cache misses (key drift), the lane pays a
cold compile like the server shard does today; a miss is slow, never
wrong.
- Double-run/coverage-loss combinations are enumerated in the wrapper
header and pinned by tests: each caller runs the suite exactly once.

## Model Used

Claude (Bender agent, Paperclip) — Fable 5.

---
Expected savings once merged: the 339s `server (1/12)` check drops to
~265s-equivalent shard levels (~211s of tests), the slowest `ci /` check
becomes the ~326s e2e shard (~13–33s off PR wall time), and every PR run
stops paying ~4.5 min of billed cold Rust compile.

For the merger (squash): please keep the trailer below in the squash
body to preserve authorship.

`Co-Authored-By: Bender (Fable)
<Paperclip-Paperclip@users.noreply.github.com>`

## Related PRs

Searched the GitHub PR list for prior work on this surface — related
groundwork, none duplicate this change:
- #13457 — restored master's Rust dependency cache on the PR runner lane
(the read-only cache this PR relies on)
- #13500 — made that cache key image-toolchain-independent so
GitHub-hosted PR runners actually hit it
- #13521 — rebalanced PR shards and split the Verify Paperclip Runner
lanes this PR extends
- #13557 — previous health-check iteration (split the runnerd transport
suite); this PR targets the next slowest check

## Checklist

- [x] I have searched GitHub for duplicate or related PRs and linked
them above

Co-authored-by: Bender (Fable) <Paperclip-Paperclip@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-18 07:22:12 -07:00
mouse-value-add fcdb3f2499 feat: add optional you.com search integration (#13555)
<!-- Simplified Technical English (ASD-STE100). -->

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agents that do research work need live information from the web
> - Paperclip reaches external systems through governed, catalog-based
MCP connections
> - The Apps catalog is data-driven: a researched provider with a hosted
remote MCP server becomes a connectable app with no runtime code change
> - You.com operates a hosted remote MCP server for web search, content
extraction, and research tools
> - The server supports OAuth 2.1 with dynamic client registration, an
API key in a bearer header, and a keyless free profile at a separate
endpoint
> - This pull request adds You.com to the self-serve MCP research ledger
and generates its catalog entry with three connection methods: browser
sign-in, API key, and the keyless free profile
> - The benefit is that an operator can give agents live web search
through the normal connection governance, and the free profile needs no
account at all

## Linked Issues or Issue Description

No public issue exists for this provider. The problem description
follows the new-adapter issue template.

**Agent or provider**

You.com — web search and research tools over a hosted remote MCP server.

**Why this adapter is useful**

Agents that do research, monitoring, or fact-finding tasks need current
web results. You.com exposes web search (`you-search`), live page
extraction (`you-contents`), citation-backed research (`you-research`),
and finance research (`you-finance`) as MCP tools. Any Paperclip company
can connect it in a few clicks. The free profile offers `you-search`
without an account, so a new company can try agent web search at zero
cost and zero setup.

**How the agent is invoked**

Hosted remote MCP server (Streamable HTTP) at `https://api.you.com/mcp`.
Three supported access paths, verified against the live server on
2026-09-16:

- OAuth 2.1 browser sign-in. The server returns a `WWW-Authenticate`
challenge with RFC 9728 protected-resource metadata and advertises a
dynamic client registration endpoint, so Paperclip's automatic DCR path
applies.
- API key. Sent as an `Authorization: Bearer` header per the provider's
official server manifest and docs. Keys come from you.com/platform and
unlock higher rate limits plus the full tool set.
- Keyless free profile at `https://api.you.com/mcp?profile=free`.
Provides a reduced, read-only tool set.

Official docs: https://you.com/docs/build-with-agents/mcp-server

**Are you willing to implement it?**

Yes. Implemented in this pull request.

**Additional context**

Research evidence collected 2026-09-16, from live protocol probes and
official provider sources only:

- Unauthenticated `POST https://api.you.com/mcp` returns HTTP 401 with
`WWW-Authenticate: Bearer
resource_metadata="https://api.you.com/mcp/.well-known/oauth-protected-resource"
scope="Tools offline_access"`.
- RFC 9728 metadata lists one authorization server with scopes `Tools`
and `offline_access`.
- The authorization-server metadata (RFC 8414) publishes authorization,
token, and revocation endpoints, and advertises a
`registration_endpoint`, so DCR is available. No registration was
performed during research, per the runbook's non-registering preflight
rule.
- The keyless free profile answers `initialize` (server `You.com`,
version `4.0.1`), lists the tools `you-search` and `you-discover`, and
executed both tools successfully during the probe.
- The API-key placement matches the provider's official `server.json` in
the youdotcom-oss/mcp repository: header `Authorization`, value `Bearer
<key>`.

## What Changed

- Added You.com (slug `youcom`, wave 4, risk tier S2) to the self-serve
MCP research ledger in
`packages/shared/src/self-serve-mcp-research.json`, and refreshed the
ledger verification date.
- Added the You.com category (`ai`) and API-key header spec to
`scripts/ingest-app-definitions.mjs`.
- Added a You.com case to `specialMethodsFor` that emits three methods:
browser sign-in (`mcp-oauth`, DCR), API key (`mcp-api-key`, bearer
header), and keyless free profile (`mcp-free`, no auth).
- Regenerated `packages/shared/src/app-definitions/youcom.json` and the
generated registry via the ingestion script (`--definitions-only` mode;
no unrelated provider churn).
- Added the official You.com wordmark artwork (light and dark theme
variants, taken from the provider's docs site) under
`ui/public/brands/apps/`, with a manifest entry.
- Updated `packages/shared/src/app-definitions.test.ts`: ledger counts
(47 providers, 44 candidates), store count (48), verification date, and
assertions for the three You.com methods and their endpoints.

## Verification

- `node scripts/ingest-app-definitions.mjs --definitions-only` — passed.
Generated the new definition and registry import only; no other provider
JSON changed.
- `node scripts/check-app-brand-assets.mjs` — passed (71 identities).
- `node --test scripts/app-brand-validation.test.mjs` — passed.
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
ui/src/lib/app-brand-assets.test.ts
ui/src/pages/apps/AppLogo.brand-assets.test.tsx` — passed (39 tests).
- `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts
server/src/__tests__/generic-mcp-connection.test.ts
server/src/__tests__/tool-connection-removal.test.ts
ui/src/pages/apps/AppsConnect.test.tsx
ui/src/pages/apps/Browse.test.tsx` — passed (181 tests). Two server
suites that require embedded Postgres skipped on this machine by their
own environment gate; the gate is unrelated to this change.
- `pnpm --filter @paperclipai/shared typecheck` — passed.
- `pnpm --filter @paperclipai/ui typecheck` — passed.
- `pnpm --filter @paperclipai/plugin-sdk ensure-build-deps` — passed;
builds `@paperclipai/shared` with the new definition.
- `pnpm test:run` (full Vitest suite) — 7,930 passed, 18 failed, 4,592
skipped. Every failure is environmental on this container: the
embedded-Postgres suites refuse to start because the machine runs as
root, the native runtime suites need the Rust runner binary that this
container cannot build, and one media suite needs a native HEIC binary.
No failure touches the app-catalog, connection, branding, or
shared-package surface; those suites pass locally. CI is the
authoritative gate for the full suite.
- `pnpm --filter @paperclipai/server typecheck` — not completed: the
script's `prepare:runner-vendor` prelude builds the Rust runner, which
cannot build on this container. A direct `tsc --noEmit` reports only
pre-existing errors from the missing vendored runner types; no error
touches this change. No server code is changed.
- Live You.com proof on 2026-09-16 (keyless free profile, real network
calls): preflight 401 challenge with RFC 9728/8414 metadata and DCR
endpoint ✓, `initialize` ✓, `tools/list` ✓, `you-search` call returned
results ✓, `you-discover` call returned results ✓.
- Live proof NOT run: an authenticated OAuth connect and an API-key call
against the full server. This environment has no You.com account or API
key. Per the runbook, this proof stays outstanding and must not be
assumed from the keyless probe. Both paths match the reviewed
`mcp_remote` patterns (DCR and bearer header) used by existing
providers.
- Browser e2e suites not run: opt-in per `AGENTS.md`, and this change
adds catalog data only, with no UI code.

## Risks

- Low risk. The change is catalog data plus generated output. It adds no
runtime code and touches no existing provider.
- The free-profile method is a fixed keyless endpoint. If You.com
changes or removes `?profile=free`, that method breaks and the entry
needs a ledger update. The OAuth and API-key methods do not depend on
it.
- The authenticated tool catalog is discovered live at connect time, so
provider-side tool changes appear through the normal catalog refresh and
quarantine flow, not through this definition.
- Rollback is a single revert; no migration and no state are involved.

## Model Used

- Provider: Zhipu AI, via OpenRouter
- Model: GLM-5.3 (`z-ai/glm-5.3`)
- Context window: 200K tokens
- Capabilities used: tool use (shell, file edits, live HTTP probes),
long-context repository reading
- The change was produced with AI assistance and reviewed by a human
before submission.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` /
`Closes: #` / `Refs: #` OR (b) described the issue in-PR following the
relevant issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-17 10:01:49 -07:00
Devin FoleyandPaperclip 6fe8e30625 feat(apps): add Railway connection and governed deployment tools (#13415)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Apps gives agents governed access to external resources.
> - Operators need to inspect Railway services, read logs, deploy code,
and run container commands.
> - Railway offers hosted MCP with OAuth, but broad remote actions hide
their internal operations.
> - This PR adds a branded connection and fixed direct operations
through the existing gateway.
> - Separate SSH keys enable container commands under the same grants
and policies.
> - Operators can require approval for an action and inspect the
resulting audit record.

## Linked Issues or Issue Description

**Subsystem affected**

Apps catalog, connection setup, gateway execution, and connection
documentation.

**Problem or motivation**

Agents need Railway access through Paperclip. Operators need to grant
and revoke that access, inspect available actions, and govern deployment
and container operations without giving agents provider credentials.

**Proposed solution**

Reuse hosted MCP OAuth, vault storage, catalog discovery, grants, and
the gateway. Probe the actual credential before enabling fixed GraphQL
operations. Use a dedicated grant-owned SSH key for bounded container
commands.

**Alternatives considered**

A catalog entry alone cannot execute the missing operations. The hosted
general agent has opaque internal effects. An unrestricted CLI runtime
can bypass action policy and inherit ambient credentials.

**Roadmap alignment**

This extends the existing MCP Tool Gateway & Apps path and the Connected
Apps direction in ROADMAP.md. It does not add a plugin or parallel
connection service.

Related PRs #311, #939, and #7861 concern hosting Paperclip on Railway.
They do not add this outbound Apps connection. The separate shared
agent-picker fix is #13414 and is not included here.

## What Changed

- Add the generated Railway catalog entry, official marks, provenance,
and OAuth setup guidance.
- Add fixed service/deployment status, bounded logs, and
redeploy/restart/rollback tools. Block source deployment until the
provider can atomically bind the approved repository and commit.
- Verify API access with an explicit workspace before exposing direct
tools.
- Add grant-owned SSH key setup and a bounded runner with host
verification, target checks, isolated state, and cleanup.
- Block the opaque hosted railway-agent and accept-deploy actions.
Preserve normal Allowed defaults and Ask-first policies for other
actions.
- Quarantine new or changed Railway schemas after initial discovery,
including reconnect.
- Add provider, lifecycle, gateway, SSH, UI, and browser fixtures.
Document setup, limitations, and the release checklist.

## Verification

- Security follow-up: removed the unsafe source-deployment mutation.
Direct calls and old active catalog entries are denied before any
upstream request, including normalized aliases. Refresh marks retired
entries disabled. All 386 focused Railway, catalog and gateway tests
passed, and server TypeScript checking passed. Full [GitHub
CI](https://github.com/paperclipai/paperclip/actions/runs/35139421144)
passed on d86530ab9, including typecheck, build, all tests, runner
checks, and browser tests. Superagent passed and confirmed the P2 fix.
Greptile reviewed the same commit at 5/5 with no findings.

- CI follow-up: fixed the missing Railway SSH operation in the OpenAPI
document, including its request schema, operator-only authentication,
and error responses. The failure reproduced locally before the fix; all
403 selected API, Railway, catalog, and artwork tests passed after it.
Synced current master and resolved the catalog/artwork conflicts.

- After rebase: 440 focused provider, lifecycle, gateway, catalog, and
container-panel tests passed. AppDetail and AppsConnect passed another
196 tests.
- Full typecheck, build, token gates, and the gallery browser check
passed after rebase.
- During implementation, full build and the gallery browser check
passed. Shared generic-MCP fixtures covered OAuth callback/state/issuer
binding and failure paths.
- Local live consent and tools/list succeeded. There were 44 active
hosted actions and two blocked actions. A workspace-bound API probe and
direct project/service/environment reads succeeded. The inspected
project had no deployed services. No provider mutation ran.
- Full GitHub CI passed on commit 303340f19, including all
server/workspace test groups, typecheck, build, runtime verification,
release dry run, and browser tests. The original local full-run attempt
was incomplete; the complete automated suite is now verified in CI.

Manual review: connect Railway, review the actual actions, install for
an agent, and run a resource read through the gateway. Choose Ask first
before testing a deployment mutation. Configure a dedicated key only
when container access is needed.

**Release qualification is still open.** Live agent gateway reads/logs,
rejected and approved deployment calls, refresh/revoke, public HTTPS
consent, and SSH enrollment/commands/cleanup need an authorized
disposable service. The passing API diagnostic does not replace those
tests. See doc/connections/RAILWAY.md and RAILWAY-REVIEW.md.

## Risks

Overall risk is medium. New runtime behavior is gated to Railway
connections, but the PR changes shared catalog, credential lifecycle,
and gateway code. A regression in those paths can affect other Apps
connections. The highest-impact operations are Railway deployments and
container commands.

- Provider consent can authorize an entire workspace. Catalog labels are
not local resource allowlists. Direct tools check target membership, and
provider permissions still apply.
- Shell commands have broad internal authority. Action policy cannot
approve each internal shell step. Timeouts close the local connection
but cannot guarantee remote child-process termination.
- Log and command output may contain application secrets that pattern
redaction cannot recognize.
- Source deployment is unavailable until the provider supports atomic
repository/commit binding. Existing deployments can still be redeployed,
restarted or rolled back.
- No database migration is required. Rollback can remove promotion and
direct dispatch while preserving connection data and the generic MCP
path.
- Live Railway qualification must still pass before release acceptance.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and browser testing.
An independent read-only security agent reviewed the local
implementation. The exact serving model ID and context window were not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 13:54:44 -07:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
DottaandPaperclip 18989a9e73 docs: add eval guide, authoring skills, and public history hub (#13535)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Its evaluations test both the Runner and complete product workflows.
> - The guides and run histories are in separate places.
> - The shared Evalbook viewer can make the test boundary unclear.
> - This pull request names the two families and adds a guide, authoring
skills, and a public hub.
> - Contributors can choose the correct test and inspect its history.

## Linked Issues or Issue Description

**Issue type**

Missing documentation.

**Where is the issue?**

Runner and Product E2E evaluation guides, case-authoring procedures, and
public result navigation.

**What's wrong?**

There is no single entry point. A report format can be mistaken for an
execution boundary. There are no dedicated case-authoring skills for
these two families.

**Suggested fix**

Add a guide and three skills. Link both existing histories from a public
hub. Keep existing campaign URLs and grading unchanged.

## What Changed

- Add `doc/evals.md` and links from existing guides.
- Add the `paperclip-evals`, `add-runner-eval`, and
`add-product-e2e-eval` skills. Install copies in
`~/paperclipai/.agents/skills`.
- Add a static hub builder that reads the existing public history feeds.
- Show a dated snapshot for each family. Label partial campaigns and
preserve measurement dates across report refreshes.
- Document publication and refresh commands for
https://pages.paperclip.ing/evals/.

## Verification

- Seven Python summary tests pass: `python3 -m unittest discover -s
scripts/evals-hub -p 'test_*.py'`. Run these checks directly; this PR
does not modify package scripts.
- All three skills pass the skill-creator `quick_validate.py` check with
`/usr/bin/python3`.
- Build tested with saved history fixtures and the live public feeds.
- Desktop and mobile browser checks pass. The mobile page has no
horizontal overflow.
- Published https://pages.paperclip.ing/evals/. Browser check: HTTP 200,
no page errors, all eight links return HTTP 200, no mobile overflow.
- Independent skill exercises found the existing Notion-decline case and
a direct Runner permission-denial case. Roster validation with an
explicit run ID passes.
- Missing refresh measurement date: regression fails before the fix and
passes after it.
- `git diff --check` passes.
- No paid evals were run for this documentation and reporting change.
The preceding head passed typecheck, build, server/workspace tests,
runner verification, browser E2E, and the canary dry run. Checks for the
latest commit are pending. Local repository-wide typecheck, test, and
build were not repeated because no product code changed.

## Risks

The hub is a dated static snapshot. It can lag behind the linked
histories until an operator refreshes it. A changed history schema stops
the build. Existing archives and grades are not modified. The published
guide link is pinned to the reviewed commit so branch deletion cannot
break it. Later builds can use master.

## Model Used

OpenAI gpt-6-astra for implementation and review. OpenAI gpt-5.6-luna
for documentation and independent skill checks. Both used repository
tools and code execution. Context window sizes are not exposed by this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green (latest commit pending)
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
(preceding head was 5/5; latest commit pending)
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-16 08:41:06 -05:00
a8d32e5e61 feat(sandbox-providers): add CreateOS sandbox provider (#13434)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent work runs in sandboxes that provider plugins supply
> - Operators can choose a provider to run agent work
> - CreateOS adds another provider with workspace-preserving pause and
resume
> - This pull request adds a CreateOS provider plugin
> - The benefit is that operators can preserve a workspace between runs
without keeping its compute active

## Linked Issues or Issue Description

Refs #13203 and the earlier closed #13096.

This continues the CreateOS contribution from @bhautikchudasama and
@ashwaq06. The branch preserves the original implementation commit.
Thank you to both contributors.

When squash-merging, preserve the original author's credit in the squash
commit body:

```text
Co-Authored-By: bhautikchudasama <BhautikChudasama@users.noreply.github.com>
```

The original fork rejects maintainer pushes. This branch includes the
merge-conflict resolution and review fixes. The request is described
below using `adapter_request.yml`.

**Agent or provider**

CreateOS sandbox API (https://api.sb.createos.sh).

**Why this adapter is useful**

CreateOS can pause a sandbox and resume it by ID. The workspace survives
the pause. This adds a reusable-lease option to the existing sandbox
provider system.

**How the agent is invoked**

Build and install the local plugin as described in its README. Open
Instance Settings, then Environments. Select the `createos` driver.
Supply an API key and shape. The driver then supplies sandbox leases for
agent runs.

**Are you willing to implement it?**

Yes. This pull request is the implementation.

## What Changed

- Adds the `createos` sandbox provider under
`packages/plugins/sandbox-providers/createos`.
- Calls the CreateOS HTTP API directly. The package adds no vendor SDK.
- Implements the environment lifecycle hooks, incremental process
output, and binary workspace sync.
- Registers the optional bundled provider and its trusted host
credential fallback. The fallback is limited to the official API origin;
custom endpoints require an explicit key.
- Lists the package in the release manifest with `publishFromCi: false`
until its first npm publish is bootstrapped.
- Waits through delayed pause/resume state updates without duplicate
action requests.
- Cancels queued API requests promptly while preserving request spacing.
- Uses direct CLI invocation in the setup guide so paths and IDs are
passed without an extra shell expansion.
- Includes current master and retains its existing Git-subfolder
containment fix.

## Demo

Fresh setup and a run against a CreateOS sandbox.


https://github.com/user-attachments/assets/e71b9e06-c006-4fb9-b847-52dfd68f6110


https://github.com/user-attachments/assets/43b5ac75-66bd-4f76-8563-67e4c7759084

## Verification

All 25 jobs in [CI run
34884260542](https://github.com/paperclipai/paperclip/actions/runs/34884260542)
passed at commit `f8d0997677024b784fdadf9d44a84c01cb4e813c`, including
typecheck, build, native runner verification, server and workspace
tests, browser tests, and the canary release dry run. Greptile reviewed
the same commit at 5/5 with no unresolved review threads.

GitHub reports no merge conflicts. The remaining merge gate is
code-owner approval for the new `package.json`, as required by
`.github/CODEOWNERS` and the `master` ruleset. Reviewers have been
requested automatically.

Local checks passed:

- Provider: `pnpm typecheck`, `pnpm test` (52 passed, one live smoke
skipped), and `pnpm build`.
- Host: focused credential and bundled-plugin tests (17 passed), plus
CLI invocation safety (39 passed).
- Release: package manifest check and release policy tests (18 passed).

The full local `pnpm test:run` attempt caught the README command issue;
its focused rerun now passes. The full local run stopped after its
general-server group: 7,804 tests passed, with unrelated embedded
PostgreSQL startup failures and 10 failures in unchanged
runtime-skill-cache tests (`EACCES` on directory rename on macOS). It
did not reach the later test groups. Local `pnpm -r typecheck` and `pnpm
build` reach the runner package and stop because this machine has no
Rust/Cargo installation. The corresponding CI checks passed on
provisioned runners, as linked above.

The live CreateOS smoke requires explicit provider credentials and was
not run during this review. It is available with `CREATEOS_LIVE_TEST=1
pnpm test` in the provider directory. The author supplied the demo links
above.

## Risks

The provider is opt-in and is not installed by default. It is available
through a local-path install or explicit image inclusion. npm
publication remains disabled until a maintainer bootstraps the package
and enables publishing.

Sandbox creation has no idempotency key. An ambiguous create response
can leave a resource that requires provider-account inspection. Process
tracking is in memory; durable lease recovery belongs to the host. The
provider does not advertise guaranteed expiry, interactive login,
snapshots, duplex channels, or ingress. Live native-runner qualification
remains outside this PR's tested claims.

## Model Used

Original provider implementation: human-authored by @bhautikchudasama,
as reported in #13203. The original description reports Claude Opus 5
assistance.

Review and follow-up fixes: OpenAI GPT-6 via Codex, with code review,
editing, and tool execution. The precise runtime model variant and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full-suite
environment limits documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: bhautikchudasama <bhautikrchudasama@gmail.com>
Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 17:09:43 -07:00
Devin Foley bfb4ceabbb perf(release): wait for npm registry visibility of all packages concurrently (#13495)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The release system publishes ~33 public npm packages per release
across the canary, nightly, beta, and stable channels
> - `scripts/release.sh` publishes them strictly sequentially: publish
one package, poll npm until that version is registry-visible, then start
the next
> - npm accepts a publish in seconds, but registry packument propagation
can lag minutes per package (the CI budget was raised to 30 minutes per
package after two aborted releases), so the total publish time is the
sum of every package's lag — about two hours on a bad npm day, paid by
every channel run including every canary on every master push
> - This pull request keeps the publishes sequential but runs all the
registry visibility polls concurrently once every publish is accepted
> - The benefit is that the wall-clock cost of npm propagation drops
from the sum of all packages' lag to the single slowest package's lag,
with every existing safety property preserved

## Linked Issues or Issue Description

No existing issue. Description follows the enhancement template:

**What existing behavior does this improve?**
The npm publish step of `scripts/release.sh` (Step 5), used by every
release channel.

**Subsystem affected**
Release tooling (`scripts/release.sh`, `scripts/release-lib.sh`).

**Current behavior**
Packages publish one at a time, and after each publish the script polls
npm until that package's version is visible in the registry packument
before publishing the next. With per-package propagation lag of minutes
(observed up to ~15 minutes; per-package CI budget is 30 minutes), the
full 33-package set takes up to ~2 hours of mostly idle waiting.

**Proposed behavior**
Phase 1 publishes every package sequentially exactly as today (a
rejected publish still aborts the batch immediately with exact
attribution). Phase 2 then polls registry visibility for all packages
concurrently. Each package keeps its own `NPM_PUBLISH_VERIFY_ATTEMPTS` ×
`NPM_PUBLISH_VERIFY_DELAY_SECONDS` budget, and any version that never
becomes visible still hard-fails the release, now naming every
straggler.

**Reason and benefit**
Total publish wait becomes the slowest single package's lag instead of
the sum of all lags — typically minutes instead of hours. This shortens
every canary, nightly, beta, and stable run and reduces exposure to job
timeouts during npm slowdowns.

**Breaking changes**
None. Dry-run output is byte-identical in structure, dist-tags are still
applied at publish time (`--tag`), the Sigstore TLOG duplicate-recovery
path is untouched, and the later dist-tag integrity check
(`wait_for_release_registry_state`) is unchanged.

## What Changed

- `scripts/release-lib.sh`: replaced `publish_package_to_npm_and_wait`
with `wait_for_npm_package_versions`, which takes the package tuple list
and polls every package's visibility in background subshells, each
reusing the existing `wait_for_npm_package_version` poll (same
per-package budget), then reports per-package success or fails naming
all stragglers
- `scripts/release.sh` Step 5: the publish loop calls
`publish_package_to_npm` only (sequential, fail-fast on a rejected
publish), followed by one call to `wait_for_npm_package_versions` for
the whole set; Step 6's recap line updated to match
- `scripts/release-lib.test.mjs`: the registry-visibility and
workflow-budget tests now drive the new function (same assertions on
`npm view` counts, virtual sleeps, and the fail-closed message, which
now names the straggler); a new cross-visibility test proves concurrency
— two fake packages that each become visible only after the other has
been polled can only converge when polled in parallel, so the test fails
if the waits ever serialize again

Safety analysis for the ordering change: nothing in the publish loop
resolves sibling packages from the registry.
`prepare-bundled-package.mjs` bundles and patches from the local
workspace tree, and the TLOG duplicate-recovery path only queries the
package it just published. The only consumer of the "visible before next
publish" invariant was the release script's own final verification,
which still runs against the full set.

## Verification

- `node --test scripts/release-lib.test.mjs` — 15 tests pass, including
the new concurrency proof and the existing 15-minute-20-second budget
tolerance test against the new function
- `npm run test:release-registry` — full lane, 140 tests pass
- `bash -n` on both scripts; `shellcheck` reports no findings beyond the
three pre-existing ones on master (verified by comparing counts against
`origin/master` copies)
- A real-release exercise happens on the next master push: every canary
run executes this exact path

## Risks

- Low. The failure mode most worth watching is a release where some
packages become visible and others never do: previously the run stopped
at the first invisible package with later packages unpublished; now all
packages are accepted before visibility is enforced, and the run fails
naming every straggler. Recovery is identical in both worlds (the next
attempt derives a new version number), and the accepted-but-lagging
packages carry the correct dist-tag either way.
- Publish jobs run the source commit's copy of `release.sh`, so this
change takes effect for a given channel only once its source commit
includes this merge — promoted nightlies/betas cut from older commits
keep the old sequential behavior until their trains catch up.

## Model Used

Claude Fable 5 (`claude-fable-5`), extended thinking with tool use
(Claude Code).

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-15 16:09:51 -07:00
DottaandPaperclip 6cfe4acff7 fix: preserve README images in npm package (#13488)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Its CLI is published as the paperclipai npm package, with the root
README shown on the package page
> - The root README uses repository-relative image paths so images
render correctly on GitHub
> - npm resolves those paths under the package repository directory,
which is cli, so the image requests point to missing cli/doc/assets
files
> - This pull request prepares the generated npm README by converting
only image src and srcset asset paths to stable raw GitHub URLs
> - The benefit is that the same source README remains correct on GitHub
and the published npm README displays its images

## Linked Issues or Issue Description

**Where is the issue?**

The issue is in the root README image assets and the npm packaging step
in scripts/build-npm.sh. The affected public page is
https://www.npmjs.com/package/paperclipai.

**What's wrong?**

The npm build copies README.md into cli/ before publishing. npm resolves
relative image paths beneath the package repository directory, so
doc/assets/banner.jpg becomes cli/doc/assets/banner.jpg. Those files do
not exist, and the images render as broken on npm.

**Suggested fix**

Keep the root README paths relative for GitHub. Rewrite
repository-relative image paths only in the generated npm README copy to
absolute raw.githubusercontent.com URLs.

## What Changed

- Added a small npm README preparation script that rewrites relative
image src and srcset asset paths.
- Updated scripts/build-npm.sh to use the preparation step when
generating the npm package README.
- Added a regression test for src, srcset, immutable refs, absolute
URLs, and non-image Markdown links.
- Pinned release-build image URLs to the source commit, while preserving
tarball builds by passing their known source refs.

## Verification

- Passed: node --test scripts/prepare-npm-readme.test.mjs
- Passed: bash -n scripts/build-npm.sh scripts/e2e-install-lifecycle.sh
scripts/e2e-update-migrations.sh
- Passed: focused README and E2E migration harness tests
- Passed: git diff origin/master...HEAD --check
- Generated README asset URLs were checked against raw GitHub and all
seven returned HTTP 200.

## Risks

Low risk. The change affects only the temporary README generated for npm
packaging. It does not change the GitHub README or runtime code. The
generated npm README depends on the public raw GitHub asset URLs
remaining available.

## Model Used

OpenAI Codex, GPT-5. Tool-enabled repository inspection, code execution,
browser verification, and git/GitHub operations were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with Fixes: # / Closes: #
/ Refs: # OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub #NNN / github.com/paperclipai/paperclip URLs)
- [x] My branch name describes the change (e.g. docs/... or fix/...) and
contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 13:34:30 -05:00
DottaandPaperclip 4510bf7c9e ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Pull request CI uses a trusted workflow on the AWS runner fleet.
> - The caller used a fixed SHA that also needed runner-group admission.
> - A mainline pin update left CI queued because the group still allowed
older SHAs.
> - This pull request calls the trusted workflow on master, which
requires code-owner review.
> - New merged workflow versions can use the existing master
runner-group entry.

## Linked Issues or Issue Description

Related: #12968. That Dependabot PR proposes another SHA rotation. This
change keeps this first-party workflow on master instead.

**What happened?**

CI run 34975562974 stayed queued because its trusted workflow SHA was
absent from the runner-group allowlist. The fleet itself was healthy.

**Expected behavior**

New reviewed versions of the trusted workflow on master should receive
runner access without a separate SHA allowlist update.

**Steps to reproduce**

Change the caller to a new trusted workflow SHA without adding that SHA
to the restricted runner group. Its jobs remain queued. The master
reference removes that recurring synchronization step.

## What Changed

- Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`.
- Exclude this exact first-party workflow from Dependabot updates.
- Update the existing E2E shard workflow tests for the master caller
contract.
- Document the runner-group entry, required code-owner review, and
old-reference retention.

## Verification

- `actionlint .github/workflows/pr.yml` passed.
- `node --test scripts/__tests__/e2e-shard.test.mjs
.github/scripts/tests/cloud-runner-routing.test.mjs
.github/scripts/tests/pr-runner-rust-cache.test.mjs
.github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests.
- Parsed Dependabot YAML and checked the exact workflow exclusion.
- `git diff --check` passed.
- Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS
has no errors, and the active master ruleset requires code-owner review.
This covers the trusted workflow, caller, and CODEOWNERS itself.
- The approved organization setting now allows
`paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`.
All previous references and other runner-group settings remain intact.
- Full application typecheck, tests, and build were not repeated locally
for this workflow-only change. This PR's CI and review are pending.

## Risks

New versions of the trusted workflow take effect for new callers after
merge to master. Keep code-owner review and master protection enabled.
Existing administrator pull-request bypasses remain unchanged.
Third-party action pins, runner routing, and infrastructure are
unchanged. Older callers still use their SHA pins; their allowed refs
remain in place.

## Model Used

OpenAI GPT-6 through Codex performed implementation and orchestration;
its exact runtime variant and context-window size were not exposed.
OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions
and applied the approved runner-group setting. Both used code and tool
access.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 09:43:24 -05:00
Devin FoleyandPaperclip 4cc387f907 fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.

## Linked Issues or Issue Description

Refs: #13455, #13454, #13192

**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.

**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.

**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.

**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.

## What Changed

- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.

## Verification

- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
da77a0c28c passed in 6.52 seconds,
including the real GitHub CLI signature policy and all artifact
downloads.
- Cloud consumer resolution with the certificate encoding fix passed in
5.45 seconds with zero npm metadata requests or npm processes. The
consumer is deployed and enabled in staging and production; their live
resolution APIs passed in 2.34 and 2.38 seconds. Both report the
expected fixed harness commit. A fresh tenant deployment follows this
cutover merge.

## Risks

- `Cloud deployable v1` no longer promises npm preview availability.
Enable the cloud direct-artifact consumer in staging and production
before merging this change.
- Artifact storage and GitHub attestations become required services for
new direct releases. Missing or invalid evidence fails explicitly.
- Restore the old npm dispatcher and readiness gate together before
disabling the consumer switch. Retain artifacts referenced by existing
releases.
- This change does not expand AWS runner access. The producer and
readiness bookkeeping use GitHub-hosted runners. Existing PR allowlists
and source verification gates remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full
application host limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 01:14:43 -07:00
Devin FoleyandPaperclip da77a0c28c feat(ci): publish immutable cloud migrator artifacts (#13455)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images and a matching database migrator.
> - New migrator versions must currently become available on npm before
cloud can use them.
> - npm can serve package metadata while the named archive still returns
404.
> - This pull request publishes immutable migrator archives and a
complete dependency lockfile through the existing artifact store.
> - Cloud can install these exact packages without waiting for their new
npm versions.
> - This producer change prepares a separate cloud consumer and
readiness cutover.

## Linked Issues or Issue Description

Refs: #13454

**What happened?**
A recent master run built both packages by 06:22:41 UTC on 2026-09-15.
Both archives became downloadable from npm at 06:31:56 UTC. Fresh
metadata requests did not remove the delay.

**What did you expect to happen?**
Cloud should be able to install the verified migrator as soon as its
package build and artifact upload finish.

**Steps to reproduce**
Compare package build completion, npm publication, version metadata
availability, and tarball download availability for a fresh full commit
SHA.

**Version**
Master commit `08adcc70d5ec45b7ced9619a3dc10c1d1bec397d`.

## What Changed

- Add a master-only workflow that builds the DB and shared archives
without publication credentials.
- Resolve the dependency lockfile from local archives, then pin those
archives to content-addressed URLs.
- Publish the complete bundle to a separate prefix in the existing
S3/CloudFront artifact store. Write the commit manifest last and verify
public downloads.
- Add a dedicated OIDC role policy. Only canonical master can assume it.
Writes require `If-None-Match: *`; the role cannot overwrite or delete
objects.
- Add source, integrity, lockfile, publication, and real npm install
tests. Document the format and staged rollout.
- Attest the validated manifest with GitHub/Sigstore before S3
publication. The signature binds every package and lockfile hash to the
exact master workflow and source commit.

## Verification

- `node --test scripts/cloud-migrator-artifacts.test.mjs`: 7 tests pass,
including real `npm ci` with an empty cache and no new-version metadata
lookup.
- `pnpm test:release-registry`: 136 tests pass.
- `actionlint .github/workflows/cloud-migrator-artifacts.yml` and `git
diff --check`: pass.
- Ran the workflow's filtered install and package build against the
exact master source. Built and validated the dependency lockfile from
those real archives.
- Latest-head application tests passed, including reruns of two failures
in unchanged application tests. The final CI aggregate passed. The
application source is unchanged. Common-source local typecheck and build
passed; the full local suite has the same documented macOS
read-only-directory rename limitation as #13454 (13 failures in two
unchanged suites).
- The dedicated role and additive bucket read permission are configured.
IAM simulation allows only conditional writes in the intended prefix;
overwrite without the condition, other prefixes, and deletion are
denied.
- Published the verified master 08adcc70d5
bundle with the operator session and verified all public downloads.
GitHub OIDC publication is still pending the master workflow run.
- Cloud resolved the real bundle and checked all 278 SQL migrations in
2.9 seconds with zero npm metadata requests or npm processes. The
existing migration runner applied it to a disposable local PostgreSQL
database and succeeded again on repeat.
- The producer now requires an empty-cache smoke install of the actual
package archives and their full dependency graph before upload. That
check and imports of both installed packages passed locally.

## Risks

- This is an additive producer rollout. It does not yet change the cloud
resolver or the deployable marker.
- The dedicated role and bucket read statement must be installed before
the workflow can publish. Existing bucket policy statements and
public-access blocks must be preserved.
- Referenced artifacts must be retained for rollback. No expiry rule
applies to this prefix.
- Existing external dependencies still download from npm, with SHA-512
pins. New DB and shared versions do not require npm metadata.
- The workflow uses GitHub-hosted runners and has no PR trigger. It adds
no AWS compute routing or PR access.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused release and
real-artifact tests; full-suite host limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 00:46:05 -07:00
Devin FoleyandPaperclip 058c55bf8f fix(ci): overlap preview migrator publication waits (#13454)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Cloud deploys an application image with a migrator from the same
source commit.
> - The migrator consists of the DB package and its matching shared
package.
> - npm can accept a package several minutes before readers can see it.
> - The publisher waits for shared visibility before it submits the DB
package.
> - This change submits both validated packages before polling their
visibility.
> - Their availability delays can then overlap while the final gate
still requires both packages.

## Linked Issues or Issue Description

Related: #13334. No duplicate open migrator-publication fix was found.

**What happened?**

The cloud migrator publisher waits up to ten minutes for the shared
package before submitting the DB package. The two registry delays
therefore accumulate. A shared visibility timeout can prevent the DB
package from being submitted at all.

**Expected behavior**

Submit both valid packages, then wait for both to pass the existing
exact-source metadata checks. Report which package remains unavailable.

**Steps to reproduce**

Return 404 for shared metadata after npm accepts the shared archive.
Observe that the old publisher never submits the DB archive until shared
becomes visible. The new regression test requires both submissions
before the first visibility sleep.

**Paperclip version**

Base commit 08adcc70d5. GitHub Actions
cloud-migrator publication.

## What Changed

- Validate both package archives before either publication starts.
- Submit missing packages before polling their visibility.
- Report each visible package and name packages missing at timeout.
- Cover delayed visibility, partial reuse, and invalid package pairs.
- Document the publication order.

## Verification

- Focused preview tests: 19 passed.
- Release registry tests: 132 passed.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- All 33 latest-head GitHub checks are green or intentionally skipped.
Greptile is 5/5 with no unresolved findings.
- The local full Vitest run found three failures in unchanged
company-skills cache tests. A standalone filesystem test reproduces this
macOS host rejecting a read-only directory rename with EACCES. Those
tests pass in Linux CI. The local full run continues.

## Risks

The DB package can become visible before its shared dependency. The
publisher still requires both to pass before succeeding. This removes
avoidable serialization; it does not eliminate npm propagation delays or
prove archive downloadability. Those remain separate parts of the
migrator availability work.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused release tests; the
full-suite macOS limitation is documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 23:57:12 -07:00
Devin Foley 08adcc70d5 fix(ci): retry transient GitHub reads while waiting for source verification (#13429)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every commit on master is published as an npm canary, and a canary
is the only thing a nightly, a beta and then a stable can be promoted
from
> - A canary only publishes after the release run confirms that Cloud
readiness passed for that exact commit, which it does by polling the
GitHub Actions API for up to 45 minutes
> - That poll used a read that threw on any non-OK response, so one
gateway error ended the wait immediately
> - The release run then failed and the commit published no canary,
although readiness itself had passed
> - A commit with no canary can never be promoted, so this silently
removes commits from the release path
> - This pull request retries the reads that mean "ask again" and leaves
every real failure fast
> - The benefit is that a transient API error costs a few seconds
instead of a release

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
The `Reuse exact-source verification` job failed 12 seconds into a
45-minute wait:

```
Waiting for Cloud source verified v1 for 5054c9ef9b.
GitHub Actions read failed (HTTP 502).
##[error]Process completed with exit code 1.
```

`publish_canary` was skipped, so the commit published no canary. Cloud
readiness for that same commit had already completed successfully.

**Expected behavior**
A gateway error from the GitHub API is a reason to ask again, not a
verdict on the commit. The wait should continue and the canary should
publish.

**Steps to reproduce**
1. Push to master and let Cloud readiness pass for that commit.
2. Have the GitHub Actions API return 502 for any single read the
verification poll makes.
3. The release run fails and no canary is published for that commit.

**Paperclip version or commit**
Present on master. Observed on 2026-09-14 in a release run for
`5054c9ef9`.

## What Changed

- `scripts/cloud-source-verification.mjs` gains `createActionsReader`,
the transport the CLI entry point now uses. It retries transient
transport failures — network errors and HTTP 408, 425, 429, 500, 502,
503 and 504 — with four bounded attempts and a growing backoff.
- Every other non-OK status still throws on the first response. 401, 403
and 404 mean the token or the target is wrong, and waiting those out
would only delay the failure.
- The verification logic itself is untouched. A readiness run that
genuinely failed still stops the release immediately, and an ambiguous
or mismatched run still throws.

## Verification

```
node --test scripts/cloud-source-verification.test.mjs   # 16 passed
node scripts/cloud-source-verification.mjs               # still exits cleanly with the token message
```

New tests cover a retried gateway error, each transient status with its
growing backoff, a retried network failure and the message that survives
exhaustion, no retry for authorization failures, and the attempt budget.
The pre-existing verification tests are unchanged and still pass.

## Risks

Low risk, and limited to one CI script.

- A genuinely unreachable API now takes four attempts before failing,
which adds a few seconds to a run that was going to fail anyway.
- The retry cannot mask a failed readiness run: that path throws a
different error, from the verification logic rather than the transport.
- No product code and no workflow changes.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [ ] I have updated relevant documentation to reflect my changes — no
documented behavior changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-14 23:20:15 -07:00
Devin FoleyandPaperclip 52d120f68d fix(release): wait 30 minutes for npm to expose a published version (#13436)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every release publishes a batch of npm packages and then waits for
each one to become visible before continuing
> - npm accepts a publish immediately but exposes it later, so that wait
exists to keep a release from continuing past a package nobody can
install yet
> - Today that wait was too short twice in a row, and each timeout
aborted the release with the batch half-published
> - Version numbers are derived from what is already on npm, so every
retry moves to a new number and meets the same lag
> - Two canary attempts burned two versions this way and shipped nothing
> - This pull request raises the per-package budget from ten minutes to
thirty
> - The benefit is that ordinary registry lag costs waiting instead of a
failed, half-published release

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
`publish_canary` failed twice in a row with the batch half-published:

```
Warning: npm accepted @paperclipai/server@2026.914.0-canary.2, but the version did not become registry-visible.
Error: stopping release: npm did not publish and expose @paperclipai/server@2026.914.0-canary.2
```

The version was accepted at 17:49:09 and became visible at 18:04:29 —
about five minutes after the poll gave up. `shared`, `db` and
`adapter-utils` published at that version; `server`, `paperclip-runner`
and the root package did not.

**Expected behavior**
Ordinary registry propagation delay costs the release some waiting, not
a failure. A package that becomes visible after 15 minutes must not
abort the batch because of the old 10-minute wait. Longer registry
outages can still leave a partial batch.

**Steps to reproduce**
1. Publish any channel while npm is propagating slowly.
2. A package takes longer than `NPM_PUBLISH_VERIFY_ATTEMPTS *
NPM_PUBLISH_VERIFY_DELAY_SECONDS` to become visible.
3. The release aborts, that version is half-published, and the retry
picks a new version number and meets the same lag.

**Paperclip version or commit**
Present on master. Observed on 2026-09-14 across canary runs in workflow
run 34869494325.

## What Changed

- Increase npm visibility checks from 60 to 180, retaining the 10-second
delay: about 30 minutes per package.
- Increase canary, nightly, beta, and stable publish job timeouts from
90 to 150 minutes.
- Add offline regression tests using the workflow's actual settings.
They cover the observed 15-minute 20-second delay, immediate visibility,
exhausted retries, and job timeout sizing.
- Load the access router in test setup so its cold transform does not
consume the first permission test's 10-second timeout. The permission
assertions are unchanged.

## Verification

- `node --test scripts/release-lib.test.mjs`: 14 passed.
- `pnpm run test:release-registry`: 129 passed.
- `pnpm exec vitest run
server/src/__tests__/access-routes-permissions-upgrade.test.ts`: 3
passed.
- Regression proof in temporary fixtures: restoring 60 attempts fails
the observed-delay test; restoring 90-minute jobs fails the
timeout-budget test.
- [CI run
34916632804](https://github.com/paperclipai/paperclip/actions/runs/34916632804):
all jobs passed, including typecheck/release registry, build, all
general and serialized server shards, browser tests, runner
verification, and canary dry run. The PR has 31 successful checks and
two expected Storybook skips at `a27f5e896`.
- [Previously failing serialized
shard](https://github.com/paperclipai/paperclip/actions/runs/34916632804/job/104215809616):
all three access-route permission tests passed in CI after preloading
the router.
- Greptile's final review is 5/5 with no outstanding findings. Both
review threads are resolved.
- Local limits: `pnpm -r typecheck` and `pnpm build` stop at the runner
package because this machine has no Rust `cargo` executable. The
duplicate full local `pnpm test:run` was interrupted while the complete
CI matrix ran. Targeted local results are listed above; full validation
is from CI.

The previous CI failures were unrelated to npm propagation:

- [PR
review](https://github.com/paperclipai/paperclip/actions/runs/34891770396)
required a test file for this fix.
- [Serialized server shard
3](https://github.com/paperclipai/paperclip/actions/runs/34891773736/job/104136392196)
timed out in the first access-route permission test at 10 seconds. The
other two tests in that file passed.
- The canary dry run passed in that same CI run.

## Risks

Low risk. Production behavior changes only in release waiting budgets.

- Polling exits as soon as npm exposes the version, so healthy publishes
do not wait longer.
- An unavailable version now takes about 30 minutes to report. Polls
remain bounded and still fail the release on exhaustion.
- The 150-minute jobs leave roughly 30 minutes for setup/build plus four
full polling windows. npm command runtime and later release steps also
consume that budget; a broader outage can still interrupt a batch.
- The permission-test change moves module loading into a bounded setup
hook; it does not relax authorization assertions.
- No schema changes or operational migrations.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

- OpenAI GPT-6 (Codex), with reasoning, tool use, and code execution,
for the CI follow-up. Context-window size is not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted checks listed
above; full validation passed in CI
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes —
workflow comments and verification details
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 20:16:29 -07:00
DottaandPaperclip 728f7185f6 feat: add native in-app announcements with persistent dismissal (#13403)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Self-hosted boards need a way to show occasional product
announcements.
> - An app release should not be required to publish or withdraw a card.
> - Native card controls keep publishing consistent; the hero can use a
static image or isolated HTML/CSS animation.
> - This pull request renders a validated JSON feed with native
components.
> - It stores dismissals per account on each instance, so a closed card
stays closed across companies and browsers.
> - Named staging feeds let authors test content before production
publication.

## Linked Issues or Issue Description

**Subsystem affected**

Board application shell, announcement delivery, and user preferences.

**Problem or motivation**

Operators need a small, optional announcement card. Users need reliable
dismissal state. Authors need to test remote content without changing
the production feed.

**Proposed solution**

Add one non-modal AnnouncementWell. Fetch validated JSON and
content-addressed media through the instance server. Keep card controls
native, with optional sandboxed HTML/CSS animation in the hero. Use
stable announcement IDs for dismissal, an explicit empty manifest and
quiet 404 handling. Provide a staged publishing helper and isolated
test-drive guide.

**Alternatives considered**

Hosting the entire card as a page would move navigation and dismissal
into remote content. This change limits HTML to a scriptless, isolated
visual hero and keeps controls native. Browser-only storage would lose
dismissals across browsers, so the instance stores account preferences.

**Roadmap alignment**

ROADMAP.md has no overlapping announcement feature. A GitHub title
search found no related announcement pull requests. This work implements
a maintainer-requested feature.

## What Changed

- Add shared feed types, strict validation of every object, supported
routes, expiration and version checks.
- Add a board-only current-feed API, constrained media proxy, and
idempotent dismissal API. Store the first dismissal and its company
audit entry in one transaction.
- Cache upstream data for one hour. Use conditional requests, request
deduplication, response limits, public destination checks, and a
three-second deadline. Treat a remote 404 as an empty feed with a
fifteen-minute retry cooldown.
- Keep announcement visibility stable when focus moves to browser chrome
or another app pane; only tab visibility starts a return check.
- Add a responsive native announcement card. Respect onboarding, dialogs
and toast placement. Sync pending dismissals across tabs and retry after
reconnect or return.
- Add idempotent migrations for dismissals and validated publication
IDs, design-guide examples, static and animated Storybook examples, and
focused tests. The publication registry supports offline retries without
accepting caller-invented IDs.
- Add HTML/CSS animated heroes with static posters, automatic playback,
reduced-motion handling, strict DOMPurify validation, an empty iframe
sandbox and CSP that blocks scripts/network resources.
- Add validated staging publication, content-addressed assets, an empty
production manifest, preview fixtures, and authoring/operator
documentation.

## Verification

- The preceding implementation passed 98 targeted
shared/server/publisher/route/OpenAPI/UI tests and 127 tests including
the master rebase. The playback-control removal passes all 21
announcement UI tests, covering the rendered sandbox, fallback, reduced
motion, dismissal and slow/stale state lookups. The preceding
shared/server tests cover HTML validation and response sandbox headers.
- The playback-control removal passes UI typecheck, production UI build,
Storybook build and token gates locally. Browser verification confirms
the animated card has only its dismiss button and two links, with no
page errors. The full canonical CI matrix passed on current head
`00e416431edb610861599d50490270bbd0f3c6b6`: 32 successful checks and two
optional Storybook deployment checks skipped. This run needed no
retries. Greptile reviewed this same head at 5/5 with no outstanding
findings.
- The local canonical general-server run passed 12,063 tests before
reporting embedded-PostgreSQL startup failures in an unrelated fixture.
All 31 tests in that fixture passed across isolated retries. The UI
group passed 6,219 tests and other workspace groups passed 3,201; two
CLI database-startup failures also passed individually. Serialized
server suites were verified by the full CI matrix rather than repeating
them locally. No source changes were needed for these environment
failures.
- The real S3/CloudFront staging manifest and both media asset headers
were verified. Production remains empty/unpublished. The guide
distinguishes the preview host's disabled edge cache from production
cache requirements.
- In the isolated test-drive, the animation visibly moves without
playback controls. A 390×844 browser viewport keeps the card above
navigation. Reduced motion makes no animation request. Both themes
render correctly and browser page errors are empty. Browser fault
injection verified that scripts cannot execute and CSS cannot make
network requests; a missing animation leaves its poster and controls.
- Refresh leaves the animated card visible. Closing it persists after
reload and the API returns null. Earlier live checks verified dismissal
across browsers, company-relative CTA navigation, modal
deferral/restoration, and new-ID eligibility after restarting the same
database.
- The deployed empty feed and a real remote 404 return HTTP 200 with
null from the board API, with a usable dashboard and no announcement
popup or browser warnings.
- Authoring documentation covers staging, animated HTML constraints,
test-drive, withdrawal, ID reuse and cache-refresh steps.

## Risks

- Animation supports self-contained visual HTML/CSS and inline SVG,
without JavaScript or external resources. A static image is required.
Older builds that do not recognize the optional animation field quietly
hide that unsupported feed.
- The default feed makes an outbound request from an instance when a
board is used. Operators can disable it. Requests contain no account
IDs, company data, cookies or interaction events.
- Feed publication and withdrawal can take about 65 minutes to reach
returning users because of CDN and instance caches. Expiration also
removes visible cards locally.
- Dismissals follow an account within one instance. No-login instances
share the existing local-board identity. Separate installations do not
share state.
- Both tables are additive. A unique key prevents duplicate dismissals;
the transaction prevents duplicate first-dismissal audit entries. The
publication registry retains only validated IDs. AGENTS.md and the
implementation spec document the required exception to company scope for
these instance-level records.
- Publication was limited to separate public staging prefixes on the
existing preview host. Production remains empty/unpublished. No AWS
policies or infrastructure were changed.

## Model Used

OpenAI GPT-6 through Codex. The exact runtime model ID and
context-window size are not exposed in this session. Capabilities used:
reasoning, code editing, shell execution, tests, browser interaction,
and tool use.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 10:19:54 -05:00
Michael Nguyen 2d47a8058d fix(apps): update connector artwork and theme fallback (#13361)
> Awaiting author review. Do not merge until the author explicitly
approves.

## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Connector screens need recognizable app artwork.
> - Several bundled marks have inconsistent artwork or dark-theme
behavior.
> - The shared logo component should retain the existing frame.
> - This change replaces selected artwork and fixes local theme
fallback.
> - Users see consistent connector icons across shared component
callers.

## Linked Issues or Issue Description

**Current behavior**

Some connector marks use outdated artwork or unsuitable theme variants.
A remote dark logo can override a canonical local mark that works in
both themes.

**Proposed behavior**

Use the selected bundled artwork with the existing gray rounded frame.
Use local light artwork in both themes unless a distinct local dark
variant exists.

**Subsystem affected**

Connector artwork, the shared AppLogo resolver, and its validation.

**Breaking changes**

No connector capability, permission, credential, or catalog activation
changes. No duplicate PR was found in the earlier search.

## What Changed

- Update 46 artwork files and only the app definitions whose logo paths
need to change.
- Keep a compact public manifest of identities, paths, visibility, and
aliases.
- Prefer local artwork in both themes and preserve the existing frame
and padding.
- Add locally runnable artwork safety checks and a light/dark Storybook
gallery; leave PR workflows unchanged.
- Document artwork conventions. Source research is kept outside the
public manifest.

## Verification

- Artwork check: 69 identities pass.
- Node artwork validation tests: 13 pass (rechecked September 14).
- Focused resolver, component, and catalog tests: 40 pass (rechecked
September 14).
- Maintainer decision: dedicated icon-validation CI is not required; the
workflow remains unchanged. Greptile acknowledged 5/5 with no remaining
code concerns on September 14.
- UI typecheck, token gates, and Storybook build pass.
- Earlier full build and repository typecheck passed. The broad local
test run was stopped after workspace-runtime dependency fixture failures
outside this change. Current GitHub CI remains the full-suite gate.
- Visual approval remains outstanding. Review the canonical icon
registry in both themes at 24–48px.

## Risks

The SVG checker rejects common active features; it is not a general
sanitizer for arbitrary uploads. Optical balance still requires human
review. Remote fallback for unknown brands retains existing behavior.
This change does not add an asset importer or new connector
capabilities.

## Model Used

OpenAI Codex, assisted by GPT-5 and GPT-6 with code execution and
browser tooling. Exact hosted model IDs and context window sizes were
not exposed.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [ ] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge


## Artwork comparison

Before and after for every affected identity. Images follow your GitHub
light/dark theme and are pinned to the base and PR commits. This
compares artwork; the existing gray rounded product frame and padding
are unchanged.

| Connector | Before | After |
|---|:---:|:---:|
| AgentMail | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/agentmail.svg"
alt="AgentMail" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/agentmail.svg"
alt="AgentMail" width="48" height="48"></picture> |
| Airtable | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/airtable.svg"
alt="Airtable" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/airtable.svg"
alt="Airtable" width="48" height="48"></picture> |
| Asana | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/asana.svg"
alt="Asana" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/asana.svg"
alt="Asana" width="48" height="48"></picture> |
| ClickHouse | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/clickhouse.svg"
alt="ClickHouse" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/clickhouse.svg"
alt="ClickHouse" width="48" height="48"></picture> |
| Cloudflare | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudflare.svg"
alt="Cloudflare" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudflare.svg"
alt="Cloudflare" width="48" height="48"></picture> |
| Cloudinary | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/cloudinary.svg"
alt="Cloudinary" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/cloudinary.svg"
alt="Cloudinary" width="48" height="48"></picture> |
| Discord | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/discord.svg"
alt="Discord" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/discord.svg"
alt="Discord" width="48" height="48"></picture> |
| GitHub | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/github.svg"
alt="GitHub" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/github.svg"
alt="GitHub" width="48" height="48"></picture> |
| Gmail | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/gmail.svg"
alt="Gmail" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/gmail.svg"
alt="Gmail" width="48" height="48"></picture> |
| Google Calendar | <picture><source media="(prefers-color-scheme:
dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-calendar.svg"
alt="Google Calendar" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-calendar.svg"
alt="Google Calendar" width="48" height="48"></picture> |
| Google Chat | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-chat.svg"
alt="Google Chat" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-chat.svg"
alt="Google Chat" width="48" height="48"></picture> |
| Google Docs | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-docs.svg"
alt="Google Docs" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-docs.svg"
alt="Google Docs" width="48" height="48"></picture> |
| Google Drive | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-drive.svg"
alt="Google Drive" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-drive.svg"
alt="Google Drive" width="48" height="48"></picture> |
| Google People | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-people.svg"
alt="Google People" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-people.svg"
alt="Google People" width="48" height="48"></picture> |
| Google Sheets | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-sheets.svg"
alt="Google Sheets" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-sheets.svg"
alt="Google Sheets" width="48" height="48"></picture> |
| Google Slides | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-slides.svg"
alt="Google Slides" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-slides.svg"
alt="Google Slides" width="48" height="48"></picture> |
| Google Workspace Search | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/google-workspace-search.svg"
alt="Google Workspace Search" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/google-workspace-search.svg"
alt="Google Workspace Search" width="48" height="48"></picture> |
| Hugging Face | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/hugging-face.svg"
alt="Hugging Face" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/hugging-face.svg"
alt="Hugging Face" width="48" height="48"></picture> |
| Jam.dev (library only) | New library entry | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jam-dev.svg"
alt="Jam.dev" width="48" height="48"></picture> |
| Jira | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/jira.svg"
alt="Jira" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/jira.svg"
alt="Jira" width="48" height="48"></picture> |
| Linear | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/linear.svg"
alt="Linear" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/linear.svg"
alt="Linear" width="48" height="48"></picture> |
| Manufact | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/manufact.svg"
alt="Manufact" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/manufact.svg"
alt="Manufact" width="48" height="48"></picture> |
| Microsoft Teams | <picture><source media="(prefers-color-scheme:
dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/microsoft-teams.svg"
alt="Microsoft Teams" width="48" height="48"></picture> |
<picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/microsoft-teams.svg"
alt="Microsoft Teams" width="48" height="48"></picture> |
| Miro | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/miro.svg"
alt="Miro" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/miro.svg"
alt="Miro" width="48" height="48"></picture> |
| Mixpanel | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/mixpanel.svg"
alt="Mixpanel" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/mixpanel.svg"
alt="Mixpanel" width="48" height="48"></picture> |
| Netlify | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/netlify.svg"
alt="Netlify" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/netlify.svg"
alt="Netlify" width="48" height="48"></picture> |
| Notion | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/notion.svg"
alt="Notion" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/notion.svg"
alt="Notion" width="48" height="48"></picture> |
| PagerDuty | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/pagerduty.svg"
alt="PagerDuty" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/pagerduty.svg"
alt="PagerDuty" width="48" height="48"></picture> |
| PostHog | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/posthog.svg"
alt="PostHog" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/posthog.svg"
alt="PostHog" width="48" height="48"></picture> |
| Postman | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/postman.svg"
alt="Postman" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/postman.svg"
alt="Postman" width="48" height="48"></picture> |
| Shopify | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/shopify.svg"
alt="Shopify" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/shopify.svg"
alt="Shopify" width="48" height="48"></picture> |
| Slack | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/slack.png"
alt="Slack" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/slack.svg"
alt="Slack" width="48" height="48"></picture> |
| Stripe | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/stripe.svg"
alt="Stripe" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/stripe.svg"
alt="Stripe" width="48" height="48"></picture> |
| Supabase | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/supabase.svg"
alt="Supabase" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/supabase.svg"
alt="Supabase" width="48" height="48"></picture> |
| Telegram | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/telegram.svg"
alt="Telegram" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/telegram.svg"
alt="Telegram" width="48" height="48"></picture> |
| Todoist | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/todoist.svg"
alt="Todoist" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/todoist.svg"
alt="Todoist" width="48" height="48"></picture> |
| Wix | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/wix.svg"
alt="Wix" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix-dark.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/wix.svg"
alt="Wix" width="48" height="48"></picture> |
| Zapier | <picture><source media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/c9e3bb7ca40160b2ff80958ec1a8c0254638ad42/ui/public/brands/apps/zapier.svg"
alt="Zapier" width="48" height="48"></picture> | <picture><source
media="(prefers-color-scheme: dark)"
srcset="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"><img
src="https://raw.githubusercontent.com/paperclipai/paperclip/ea053c60d9d2c5c3470275e2f4dfde5f8fe5db53/ui/public/brands/apps/zapier.svg"
alt="Zapier" width="48" height="48"></picture> |
2026-09-14 05:05:27 -10:00
Devin Foley 13368c5183 fix: unblock clean-machine onboarding for api_key AI connections (nightly smoke) (#13372)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The release pipeline gates each nightly on a Docker onboarding
smoke. The smoke proves a clean machine can finish onboarding and hire
the first agent.
> - #13247, #13248, #13344, and #13351 changed the Connect step. Connect
now creates an AI connection that the server verifies live with the
provider.
> - The managed adoption check also demanded a CLI hello probe. A clean
machine has no provider CLI and cannot complete a subscription login.
Onboarding dead-ends and the nightly gate fails.
> - This pull request lets a live-verified API key adopt on the engine's
own verdict, and re-verifies the key with the provider at adoption time.
> - It also drives the release smoke through the API-key path against a
provider mock that lives inside the test harness.
> - The benefit is a green, deterministic release gate with no paid
credential in CI, and a working first run for API-key users on clean
installs.

## Linked Issues or Issue Description

No public issue exists. The failure surfaced in the nightly release
gate. Related PRs (no duplicates found): #13247, #13248, #13344, #13351
(the Connect changes), and #12423, #12135, #12151 (earlier release-smoke
updates).

**What happened?**

The nightly Release cut failed its gate: [run
34749840498](https://github.com/paperclipai/paperclip/actions/runs/34749840498),
job `smoke_nightly / smoke`, on published canary `2026.913.0-canary.2`.
The wizard never left the "Connect a model" step. The subscription path
waits for a human to run `claude auth login` on the server. The API-key
path saves and live-validates the key, but the environment test then
fails with `Command not found in PATH: "claude"` and
`ai_connection_validation_incomplete`, and the wizard blocks the hire.

**Expected behavior**

A clean machine with a provider-accepted API key completes onboarding
and hires the lead agent. The release smoke passes without a real paid
credential in CI.

**Steps to reproduce**

1. Run `scripts/docker-onboard-smoke.sh` with
`PAPERCLIPAI_VERSION=2026.913.0-canary.2`.
2. Sign in, complete onboarding to "Connect a model", select "Use API
key instead", pick Claude, enter a valid API key, and press Connect.
3. The environment test fails on the missing `claude` CLI and blocks the
hire.

**Paperclip version or commit**

`2026.913.0-canary.2` (nightly candidate `c9e3bb7ca`).

## What Changed

- `server/src/routes/agents.ts`: `testManagedEnvironment` no longer
forces the CLI-lane hello probe for a resolved `api_key` binding. It
re-verifies the key against the provider's live endpoint instead (the
same `validateAiApiKey` check the save performed, which needs no CLI). A
key the provider rejects fails adoption with
`ai_connection_api_key_rejected`. Subscription adoption keeps the strict
hello-probe requirement.
- `scripts/docker-onboard-smoke.sh`: the harness now serves
`api.anthropic.com` itself. A sibling container (the already-built smoke
image) runs a small HTTPS mock. The app container gets `--add-host` for
that one hostname and trusts the mock's certificate through
`NODE_EXTRA_CA_CERTS`. The private key stays mode 600 in the mock
container; the app container mounts only the certificate. The mock
serves only `GET /v1/models` and returns 404 for every other path.
`SMOKE_PROVIDER_MOCK=false` disables it.
- `tests/release-smoke/docker-auth-onboarding.spec.ts`: the spec drives
the API-key path — switch the credential mode before the source tile
(the link hides when the row collapses), enter the key, and Connect.
Loopback targets use a placeholder key that the mock accepts. Any other
target must set `PAPERCLIP_RELEASE_SMOKE_ANTHROPIC_API_KEY`, and the
test fails on arrival without it.
- `server/src/__tests__/agent-test-environment-routes.test.ts`: three
new route tests cover accepted keys (no CLI probe consulted),
provider-rejected keys, and subscriptions that cannot complete a hello
probe.

## Verification

- `npx vitest run src/__tests__/agent-test-environment-routes.test.ts` —
26/26 pass.
- `npx vitest run src/__tests__/ai-connections.test.ts
src/__tests__/ai-legacy-compatibility.test.ts` — 42/42 pass.
- `tsc --noEmit` reports no errors in the touched files.
- Full local harness + suite run against the exact failing canary: the
app container reaches the mock (request visible in the mock log), the
placeholder key validates, and the connection saves as the default. The
flow then stops at the forced CLI hello probe — the exact server check
this PR removes, still present in the published canary. The next canary
that includes this fix is the end-to-end proof.
- Hardening check: from inside the app container, the mock answers with
status 200 and `key.pem` is not visible.

## Risks

- Behavior shift: `api_key` adoption no longer requires a CLI hello
probe. It re-verifies the key with the provider at adoption instead.
Subscription adoption is unchanged.
- The mock returns 404 for unexpected provider calls, so a future
onboarding change that calls a new endpoint fails the smoke loudly
instead of passing silently.
- Release-smoke runs against non-loopback targets now require an
explicit key and fail fast without one.
- No database migration. No dependency change. No provider routing
change.

## Model Used

Claude Fable 5 (`claude-fable-5`) through the Claude Code CLI, with
extended thinking and tool use (shell, file edits, Playwright runs,
GitHub CLI). No other models were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-13 17:02:02 -07:00
DottaandPaperclip 13bae6fa21 fix: reject unsupported REST tool connections without stdio validation (#13346)
## Thinking Path

> - Paperclip manages AI agents and their connections.
> - Connection checks must use the configured transport.
> - The tool service treated every remaining transport as local stdio.
> - Anthropic's old REST method therefore failed with a templateId
error. A REST connection with a valid stdio template could incorrectly
pass.
> - Anthropic now has a supported AI-account flow. This pull request
removes its obsolete REST setup option and limits stdio checks to stdio
connections.
> - Users can connect an AI account, and existing unsupported
connections receive an accurate error.

## Linked Issues or Issue Description

Related: #13248 added the supported AI-account flow. Searches for
related REST health and templateId bugs found no duplicate fix.

**What happened?**

The Anthropic REST API-key connection showed `Local stdio MCP
connections must use an approved templateId`. Health checks and catalog
discovery both fell through to the local stdio path. A REST connection
with an approved template could report success and expose the template's
catalog without a REST integration.

**Expected behavior**

Only local stdio connections use command templates. Unsupported
transports return an accurate HTTP 422 error. New Anthropic accounts use
the supported runtime authentication flow.

**Steps to reproduce**

1. Check out the test-only commit `924e6e85a` in a separate worktree and
install dependencies.
2. Run `pnpm exec vitest run packages/shared/src/app-definitions.test.ts
server/src/__tests__/tool-access-service.test.ts -t 'unsupported
REST|obsolete Anthropic'`.
3. The tests exercise saved Anthropic REST configuration and an
unsupported REST connection containing an approved stdio template. They
cover health checks and catalog discovery separately.
4. Run the same tests on the fix commit. They pass. The full affected
files also pass.

**Paperclip version or commit**

Reproduced against master `6cef9743c`.

**Deployment mode**

Server transport handling. Reproduced with an isolated embedded
PostgreSQL test database. No provider account or live credentials are
required.

## What Changed

- Restrict stdio health checks and tool discovery to `local_stdio`.
- Return and audit `tool_connection_transport_unsupported` with HTTP 422
for unsupported tool transports.
- Remove Anthropic's obsolete REST method from the generated catalog and
its durable ingestion source. Keep its subscription and API-key AI
methods.
- Cover the reported error, false-success case, rejected obsolete setup,
connection removal, and the UI's AI-account submission path.
- Replace impossible reconnect forms for removed methods with supported
setup, while preserving connection removal.
- Preserve AI-versus-tool intent isolation for legacy requests and
reject new unsupported Anthropic tool requests.
- Document recovery for existing unsupported connections.

## Verification

- Clean-worktree red/green: the same command failed all six regression
cases at `924e6e85a` and passed all six at `4d3de9de0`. The failing run
includes the reported templateId error.
- Green: all 555 tests across the six affected test files passed.
- Recovery UI red/green: three added cases failed before the recovery
fix and passed afterward; all 200 tests across setup, detail, and
advanced controls passed.
- After the recovery UI update, UI typecheck/build and token gates
passed again.
- `pnpm -r typecheck` — passed.
- `pnpm build` — passed.
- `pnpm check:token-gates` — passed.
- Catalog regeneration — passed with the documented
`PAPERCLIP_CONTENT_TEMPLATES` override for the local capture corpus.
- Full CI on `69fb31fd4` — passed all general and serialized test
shards, browser shards, typecheck, build, runner verification, and
canary dry run:
https://github.com/paperclipai/paperclip/actions/runs/34726975425.
- The local serial `pnpm test:run` was stopped after the fixture
correction superseded that run; full-suite verification above comes from
CI. All 555 affected tests passed locally, including all 17
connection-intent tests after the correction.
- Greptile — 5/5, successful check on final commit `69fb31fd4`, no
unresolved findings.
- No live Anthropic validation was performed. The UI regression uses a
fake key and a mocked AI-account response.

## Risks

Existing obsolete REST connections remain in needs-attention state.
Users must add an account through the supported flow and remove the old
connection. Credentials and grants are not transferred automatically.
Removal remains covered. The specialized AgentMail and Composio paths
keep their existing behavior. There are no schema or permission changes.

## Model Used

OpenAI GPT-6 through Codex. The exact serving model ID and
context-window capacity are not exposed in this session. Used reasoning,
code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 19:26:58 -05:00
DottaandPaperclip 47ded8bf97 feat: manage AI runtime credentials through Connections (#13247)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agent runs need credentials for a specific provider and sign-in
method.
> - Connections already owns accounts, grants, and access permissions.
> - AI authentication should use those same boundaries.
> - This pull request adds the storage, API, adoption, and runtime
foundation.
> - Legacy agents keep their authentication until they explicitly adopt
a managed connection.

## Linked Issues or Issue Description

**Problem or motivation**
AI credentials are configured separately from Connections. Agents cannot
consistently reuse a responsible user's account or a permitted shared
account.

**Proposed solution**
Manage AI accounts with the existing Connections grants and permissions.
Keep model and harness selection independent from credential selection.
Preserve legacy authentication until validated adoption.

**Alternatives considered**
A separate credential registry would duplicate ownership and access
policy. Automatic fallback would risk using the wrong account.

**Roadmap alignment**
This extends the shipped Apps, multi-user, secrets, and agent-runtime
capabilities. The maintainer requested the feature and reviewed the UI.
Related groundwork: #11899 (connection permissions), #10910 (connection
wizard), #11692 (Claude subscription profiles), and #11854 (Codex
account rotation).

## What Changed

- Add AI-purpose/runtime-auth contracts and an additive, idempotent
migration.
- Add Claude, OpenAI, OpenRouter, and Grok provider capabilities and
catalog entries.
- Store credentials on grants. Resolve responsible-user defaults or
explicit permitted grants.
- Isolate managed credentials and provider sessions across accounts.
Block missing credentials without ambient fallback.
- Keep imported legacy secrets unchanged during reconnect. Use
independent local Codex/Grok sign-in attempts for rotating credentials.
- Add authorization, migration, concurrent refresh, retry, cancellation,
and legacy-compatibility tests.

This is part 1 of a two-PR stack. The app UI follows in #13248. Merge
the foundation first.

## Verification

- Updated against master `04e364236`, preserving upstream provider login
and connector workflows.
- Full workspace typecheck, production build, Storybook build, and token
gates passed on the integrated branch. Final local-login changes passed
59 focused tests; new-agent and inbox regression suites passed 63 tests.
- Browser checks verified automatic local Claude account detection,
resumable Codex login commands, retry, focus restoration, and
desktop/phone layouts. Commands create their isolated directory before
invoking the CLI.
- All current-head CI checks passed on `2a996560a`, including all
server/workspace tests, browser shards, runner verification, typecheck,
build, and canary dry run. Greptile reviewed that commit at 5/5 with no
unresolved threads. Earlier local full-suite attempts hit the Mac
PostgreSQL shared-memory limit; the complete suites passed in CI.
- Renumbered the additive AI migration to `0276` after upstream
migrations and regenerated its snapshot. Existing legacy agents retain
their configuration.
- Added local login status checks, owner-scoped retry, managed OpenCode
remote homes, credential-aware model discovery, and task
connection-repair delivery.

## Risks

- Managed credential failures intentionally block execution. They do not
restore legacy fallback.
- Preview-era copied Codex/Grok subscriptions require independent
reconnect.
- The integrated branch has live provider acceptance coverage. This
update verifies local Claude detection and Codex API-key task repair; it
does not add a new subscription authorization/refresh or Daytona stress
pass.
- Runtime-auth connections must stay excluded from tool and channel
handling.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, code
execution, and browser testing. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 16:30:10 -05:00
DottaandPaperclip 4d317274ce feat(channels): add experimental iMessage Photon (#13299)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Channels connect external conversations to company tasks and agent
execution.
> - Slack, Discord, and AgentMail already provide durable delivery and
access controls.
> - People also need to reach an agent from Apple Messages and send
photos.
> - Photon provides shared Pro DMs, dedicated numbers, and authenticated
event recovery.
> - This pull request connects Photon to the existing channel services.
> - People can message an agent while Paperclip retains task ownership
and approval authority.

## Linked Issues or Issue Description

**Subsystem affected**

Cross-cutting: channel services, shared contracts, database constraints,
Apps, and agent Channels UI.

**Problem or motivation**

Paperclip has no iMessage channel. A person cannot use Apple Messages to
start a task, send a photo, or answer an agent's pending question.

**Proposed solution**

Add experimental **iMessage Photon** with Pro-compatible shared DMs or a
dedicated Photon Cloud number per agent channel. Reuse channel
admission, identity links, task generations, publication, and
interaction continuation. Keep groups disabled for shared allocation.
Dedicated lines support groups that an operator explicitly enables.
Require a fresh linked message and a published agent response before
setup completes.

**Alternatives considered**

Shared allocation has no owned phone number, so it reserves one project
and allows DMs only. Dedicated allocation reserves one stable number.
Local Mac access needs a separate deployment model. The upstream Photon
Chat SDK adapter does not persist the poll mappings and send receipts
required here. This change uses the lower-level SDK without adding
another agent runtime.

**Roadmap alignment**

This extends Connected Apps and agent communication through the existing
channel subsystem. It does not add a parallel tool connection or agent
loop. GitHub searches for Photon and iMessage found no matching provider
implementation.

**Additional context**

This ships behind the existing experimental channel gate. Dedicated-line
release qualification remains incomplete. Real Photon Pro DMs passed
task/reply, native poll, text answers, confirmation rejection, media,
restart, pause, reconnect, revocation, and removal tests. An
operator-supplied iPhone camera HEIC also passed the full round trip.
Dedicated groups remain unqualified. See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) and [the
implementation plan](doc/plans/2026-09-11-imessage-photon.md).

## What Changed

- Add the provider catalog entry, shared setup contracts, and a forward
migration. A global partial index reserves the dedicated number or
shared project until its endpoint is archived.
- Add Cloud project inspection, vaulted project credentials,
selected-line token renewal, and a leased receiver. Persist checkpoint
updates under the receiver lease. Shared project replay accepts sparse
increasing sequences only after a complete recovery barrier.
- Connect DMs and enabled groups to existing task generations, sender
authorization, ordered delivery, and publication services. Keep each
iMessage conversation on its task after completion; only explicit `/new`
or `/close` releases the binding. Publish committed inbound comments
live and label their human bubbles “Sent from iMessage” in both
task-chat renderers.
- Persist immutable text/file send identities, upload receipts, poll
IDs, option IDs, per-person drafts, and canonical interaction
continuation proofs.
- Add source-bound file recovery, bounded HEIC/HEIF conversion, JPEG
previews, and related Live Photo companion video retention.
- Add the three-step setup flow and channel management surfaces with
official branding. Preserve the experimental gate and existing
pause/disconnect behavior.
- Add interactive production-component Storybooks for setup, access,
recovery, and ongoing conversations. Add provider, integration, catalog,
and browser regression coverage. Document setup, recovery, supported
boundaries, and qualification gaps.

## Verification

- Live Photon Pro, SDK 2.1.0: linked iPhone messages create a task and
receive native Codex replies in Apple Messages. Unlinked senders cannot
start work.
- Three real follow-ups each reopened the same completed task. Incoming
bubbles appeared on its open page without reload and showed “Sent from
iMessage.” The third follow-up ran after restarting the server on
`4d7222110`; the agent correctly repeated its previous reply from before
the restart.
- Native polls after restart, sequential text drafts, required-field
correction, explicit submission, approval rejection with a required
reason, and native continuation passed against Photon.
- PNG, text documents, synthetic HEIC, and a real iPhone camera HEIC
passed in both directions. The camera photo produced a 3024×4032 JPEG
preview. The native agent described it and returned the received HEIC
byte-for-byte.
- Pause/resume, reconnect, identity revocation, removal, `/status`,
`/new`, `/close`, and stale answers after close passed live. Messages
suppressed by pause did not become work on resume. Removal stopped
intake and removed credential bindings.
- All 304 focused tests passed on `4d7222110`. These cover Photon
unit/integration behavior, both task-chat renderers, live comment
hydration, completed-task continuity after restart, enabled groups,
duplicate delivery, and explicit reset/close. The selected Teams
completion-boundary regression also passed. Full workspace
typecheck/build and token gates passed for the conversation fix; the
final UI changes passed their affected typecheck/build and tests.
- All 26 new Photon Storybook Playwright cases passed in light and dark
themes, including the complete shared-DM setup journey and 390px mobile
follow-ups. UI typecheck and the Storybook build passed. These stories
use simulated Photon responses and do not replace the live evidence
above.
- The full chat-adapters browser suite previously passed all 39 cases.
Migration checks passed, and migration 0275 applied to the isolated live
instance with the earlier Photon migration already applied.
- The local full Vitest run was previously interrupted by the host's
embedded-Postgres shared-memory limit; it is not a full-suite pass. All
30 applicable CI checks passed on preceding head `7a5419cac`, with two
skipped checks and Greptile 5/5. Head `24f8e1aae` adds an explicit
required-story discovery guard to the 26 passing Storybook cases.
Greptile rates this final head 5/5 with no unresolved review threads.
All 30 applicable CI checks passed, with two optional checks skipped.
- A repeated live send key suppressed the duplicate but returned gRPC 6
/ SDK `internalError` without an original receipt. Paperclip keeps
unknown delivery unresolved. This provider behavior is covered by a
regression test.
- See [the verification
record](doc/connections/IMESSAGE-PHOTON-VERIFICATION.md) for package
versions, redacted live evidence, deterministic coverage, and remaining
qualification gaps.

## Risks

- Dedicated group qualification remains unrun; groups are disabled for
the approved Pro scope. Real iPhone camera HEIC passed transport,
preview generation, agent inspection, and return. Keep the channel
experimental; the dedicated-line release matrix remains incomplete.
- Shared recovery and attachment aliases were verified against the live
gateway. Duplicate writes currently return an error without the original
receipt; unresolved sends require operator resolution. The
implementation fails visibly on invalid replay ordering, a reset cursor,
or changed identity.
- The HEIF converter passed on macOS arm64 and in Linux CI. Windows HEIF
binaries have not been executed in this work. Linux musl has no packaged
converter. Unsupported conversion retains the original and reports the
missing preview.
- The migration adds a global reservation across companies for Photon
numbers and shared projects. Paused and revoked endpoints keep that
reservation until removal.
- Integration touches shared channel services. Existing provider browser
coverage passes; broad repository verification is recorded above.
- `pnpm-lock.yaml` is intentionally excluded under repository policy.
The repository bot owns lockfile updates. The additional Superagent
supply-chain scan is neutral/inconclusive because these new dependencies
are not yet in the committed lockfile. Its security scan passed; all
required CI checks pass.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository inspection, code
execution, browser testing, and tool use. The exact served model
identifier and context-window size are not exposed in this session. No
sub-agents were used.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 15:23:50 -05:00
Devin FoleyandPaperclip 44f6312cd8 fix(ci): reuse one available Cloud registry cache (#13334)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image for each merged source
commit.
> - Fresh builders restore compiled native dependencies from registry
caches.
> - The current workflow imports up to eleven historical cache manifests
at once.
> - Live builds missed native layers that a fresh builder reused from
one manifest.
> - This PR selects the nearest available cache and tests reuse across
fresh builders.

## Linked Issues or Issue Description

Refs #13329 and #13330. A search of open cache PRs found no duplicate of
this change.

**What existing behavior does this improve?**

Remote Docker cache reuse on fresh Cloud image builders.

**Current behavior**

[Cloud run
34714483272](https://github.com/paperclipai/paperclip/actions/runs/34714483272/job/103609096836)
imported the previous cache manifest successfully but rebuilt
`cargo-chef` and Rust dependencies. The dependency compile took 3m43s.
The preceding image build had already exported those layers.

A controlled [fresh-builder
diagnostic](https://github.com/paperclipai/paperclip/actions/runs/34715336530)
used the same source and registry cache. The single-manifest job reused
both layers immediately. The multiple-manifest job rebuilt them and
failed the cache assertion. Both jobs used GitHub-hosted runners with
read-only access.

**Proposed behavior**

Inspect cache manifests in first-parent order and import only the
nearest available one. Keep full-SHA cache exports, the ten-commit
search bound, and the legacy fallback. If caches cannot be read, permit
a cold build.

**Reason and benefit**

Avoid the observed cache misses without changing image contents or
builder sizes. Expected savings include about four minutes of native
tool/dependency compilation when those inputs are unchanged. The final
merge-to-deployable gain still needs a post-merge measurement.

**Breaking changes**

No image, artifact, deployment, or runner-routing contract changes.

## What Changed

- Select one available ancestor cache after Docker login and Buildx
setup.
- Preserve separate writable cache tags for each full source SHA.
- Test cache ordering, missing caches, registry errors, and workflow
integration.
- Add the selector tests to the existing release-registry suite.
- Export a local test cache, remove the first builder, and verify a
source rebuild on a fresh builder.
- Document cache selection and the stronger Docker check.

## Verification

- Passed 456 focused workflow, routing, readiness, preview-artifact, and
cache-selector tests.
- Passed shell syntax, ShellCheck for the changed probe, actionlint
workflow validation, and `git diff --check`. actionlint's shell checks
were disabled for the workflow validation because unchanged
migration-label commands trigger existing SC2012 notes.
- The fresh-builder registry diagnostic proves the single-cache
behavior. The [permanent two-builder probe
passed](https://github.com/paperclipai/paperclip/actions/runs/34715771048/job/103612624090),
including a changed real binary and dependency-declaration invalidation.
- Passed all 35 latest-head checks (green or intentionally skipped),
including full typecheck, test, build, and browser suites in [PR CI run
34715771217](https://github.com/paperclipai/paperclip/actions/runs/34715771217).
- The real selector CLI inspected registry metadata and chose the
nearest available ancestor cache.
- Fresh Greptile review is 5/5 with no open findings. The PR title was
corrected to meet the source-change naming rule; the review check passed
after that correction.
- Local full-suite runs and Docker builds are unavailable because the
local Docker daemon is unresponsive after disk exhaustion. CI provides
the Linux verification.

## Risks

- Missing or unreadable caches cause a slower cold build. The selector
logs that condition and preserves image publication.
- Inspecting several missing ancestors adds lookup time. Each lookup has
a ten-second timeout and the search is bounded.
- The Docker test now exports a local cache. It removes the first
builder before starting the second to release disk space, then cleans up
its builders and files.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 13:20:18 -07:00
Devin FoleyandPaperclip 7435b2ee9c ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images that contain the native Rust Runner.
> - The image already builds that Runner before copying ordinary app
source.
> - A Rust source change still invalidates its entire compiled
dependency layer.
> - Compiled dependencies can survive source changes when their recipe
is unchanged.
> - This PR adds a separate locked dependency build before compiling the
real workspace.

## Linked Issues or Issue Description

Refs #13195. A search of related Docker and Cargo cache PRs found no
duplicate dependency-recipe change.

**What existing behavior does this improve?**

Docker image build time after Rust source or embedded protocol changes.

**Current behavior**

The `runner-build` stage compiles dependencies and workspace code in one
layer. In Cloud readiness run 34698143548, that stage took about 3m48s
when its cache was unavailable.

**Proposed behavior**

Generate a recipe with pinned cargo-chef 0.1.73. Build locked release
dependencies in `runner-deps`, then copy and compile real Rust source
and embedded protocol inputs in `runner-build`. Source edits can reuse
the dependency layer from the existing registry cache.

**Reason and benefit**

Reduce dependency recompilation during source changes and merge bursts.
Expected savings are roughly 2–4 minutes when the old native layer would
miss but dependency layers are available. Full cold builds also pay for
the recipe tool installation. Ordinary app-only cache hits gain little
from this change.

**Breaking changes**

None to the shipped application or image tags. The recipe tool and
compiled dependencies remain in build stages.

## What Changed

- Install a pinned recipe generator with its locked dependencies and the
existing package-owned compiler.
- Add recipe planning and compiled dependency stages. Use the same
release profile, package, binary, and lockfile enforcement as the real
native build.
- Remove generated source stubs before copying actual source. Preserve
protocol inputs, timestamp normalization, binary staging, and
application checks.
- Add Docker cache wiring regressions and update the Docker cache
documentation.
- Run a two-build probe in Docker Runner check. It requires dependency
reuse, changed real binary metadata after a source edit, and a changed
recipe after a dependency declaration edit. It uses a disposable
tracked-source context and exports only small metadata files.

## Verification

- Passed all five Docker build-stamp and dependency-cache tests with
`pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`.
- Passed the local ARM64 `docker buildx build --target runner-build
--progress plain`. Local Docker then hit storage errors during a runtime
probe; cache invalidation verification continues on GitHub-hosted Linux.
- Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`,
and `git diff --check`.
- Passed a [Linux AMD64 cache
probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199)
against the PR source: dependencies compiled in 3m49s for the baseline
and were `CACHED` after a source edit; real source compilation took
about 37 seconds. Binary metadata changed and dependency declaration
changes altered the recipe. The permanent probe is also running in
latest-head Docker Runner check.
- Passed latest-head [Docker Runner
check](https://github.com/paperclipai/paperclip/actions/runs/34711145160),
including the permanent source/dependency invalidation probe.
- Passed full [PR
verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2):
typecheck, all grouped tests, native verification, build, release dry
run, and browser checks. One unrelated signoff-policy browser test
failed waiting for a heartbeat run on attempt 1; only that failed shard
and dependent checks were retried, and passed.
- Latest-head Greptile is 5/5 with no unresolved findings. Full local
tests/build were limited by local disk exhaustion; Linux CI completed
those checks.

## Risks

- The two-build CI probe has a 20-minute job limit to cover the cold
build and source rebuild. It adds no AWS routing.
- A fully cold build must install cargo-chef and populate the dependency
layer. Both become reusable registry layers; no Actions cache is added.
- The recipe and final build must keep the same compiler, build profile,
package, binary, and directory layout. A source-change rebuild probe
checks real cache reuse and binary invalidation.
- Dependency or compiler changes still require rebuilding dependencies.
Existing image verification and full-SHA publication gates remain
unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:57:32 -07:00
Devin FoleyandPaperclip 8df3ee2cf5 ci: rebalance serialized tests with current suite durations (#13328)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments wait for verified source commits.
> - Verification splits serialized server tests across independent
runners.
> - The shard duration estimates came from August and no longer match
current tests.
> - Stale estimates put much more work on one runner than the others.
> - This PR refreshes the estimates from a complete successful run to
balance the existing runners.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The time spent waiting for the slowest serialized server-test shard in
PR and release verification.

**Current behavior**

In [Cloud readiness run
34705914878](https://github.com/paperclipai/paperclip/actions/runs/34705914878),
the five serialized shards spent 479, 371, 322, 390, and 365 seconds
running tests. The recovery suite had a 55-second estimate but now takes
about 156 seconds including process overhead.

**Proposed behavior**

Use fresh per-suite measurements with the existing deterministic
duration balancer. Applying the same measured costs to the new
assignment gives 385, 385, 386, 385, and 385 seconds. This predicts
about 93 seconds less waiting for the slowest shard, before runner/setup
overhead. Live CI will confirm the result.

**Reason and benefit**

Use the existing runners more evenly. No extra runner, test parallelism,
cache, timeout, or routing change is needed.

**Breaking changes**

Suite-to-shard assignments change. The full suite set, assertions, and
per-suite process isolation stay the same.

**Additional context**

Searched related CI and shard PRs. This updates the existing duration
manifest, without duplicating a pending sharding implementation.

## What Changed

- Refresh all 145 serialized suite weights from the same successful
release verification run.
- Record source job IDs and the measurement method in the manifest.
Durations include process startup, imports, collection, tests, and
shutdown.

## Verification

- Passed all 30 shard and release-workflow tests: `node --test
scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- Confirmed every measured suite appears exactly once across all five
source logs.
- Compared old and new assignments using the same measured weights. The
maximum fell from 478862ms to 385551ms.
- In [PR CI run
34710696242](https://github.com/paperclipai/paperclip/actions/runs/34710696242),
all five serialized jobs passed in 7m05s–7m27s including setup. The
measured assignment is now balanced in a live run.
- The same run passed full typecheck, all grouped tests, native
verification, build, release dry run, and browser checks. Local
full-suite verification on this base was limited by disk exhaustion;
local typecheck and targeted shard tests passed.
- Latest-head Greptile is 5/5 with no open findings. Every current-head
CI check must be green or intentionally skipped before merge.

## Risks

- Individual durations vary with load and future test changes. These
estimates affect assignment only; missing or renamed suites receive the
existing median weight.
- Both PR and release verification read this manifest, so both receive
the new assignments. Each suite still runs in its own serialized Vitest
process.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:36:20 -07:00
Devin FoleyandPaperclip d2e940f4c1 ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys verified images from merged source commits.
> - Cloud readiness waits for every release verification check.
> - Runner verification currently runs long TypeScript tests before Rust
checks.
> - These checks can run on independent runners with their own build
directories.
> - This PR runs them in parallel while preserving all checks and the
shared dependency cache.

## Linked Issues or Issue Description

Refs #13194. Related prior work: #13142 and #13259. A search found no
duplicate parallel release-check change.

**What existing behavior does this improve?**

Time from merge to Cloud source verification and deployment readiness.

**Current behavior**

Recent successful runs take roughly 13 minutes from merge to deployable.
In run 34705914878, Runner verification took 11m23s. Protocol tests
finished before Rust tests and API authority checks started.

**Proposed behavior**

Run protocol and Rust verification in two matrix jobs. Cloud readiness
still requires both jobs to pass.

**Reason and benefit**

Remove the serial dependency between independent checks. Expected
improvement is about 2–3 minutes on a typical cached run, until the
image build or server tests become the longest job. This is an estimate;
post-merge timing will confirm it.

**Breaking changes**

Individual release Runner job names gain a lane suffix. Cloud source and
readiness marker names stay the same. PR runner routing is unchanged.

## What Changed

- Split release Runner checks into protocol and Rust lanes. Keep every
constituent of `check:all` exactly once.
- Restore the existing Rust dependency cache in both lanes. Allow only
the Rust lane to save it after warming both build profiles.
- Add coverage and cache authorization regressions. Document the
parallel verification and single cache writer.

## Verification

- Passed 477 workflow and source-verification tests with `node --test
.github/scripts/tests/*.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- Passed `actionlint`, `git diff --check`, and the private AWS routing
regression suite.
- Passed local `pnpm -r typecheck` and the standalone `check:runner &&
check:api-authority` lane, including all 1,671 API tests before the
protocol lane had built TypeScript output.
- The broad local protocol run under Node 25 had four failures. The two
affected files passed under CI's Node 24.19.0: 67 passed, 6 platform
skips.
- Local `pnpm test:run` aborted when disk space ran out; local `pnpm
build` could not run afterward. These are local verification limits.
[Linux CI run
34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424)
passed full typecheck, all grouped tests, native verification, build,
release dry run, and browser checks. Native protocol CI passed 1,986
tests, plus 1,671 API tests and the Rust suites.
- Latest-head Greptile is 5/5 with no open findings. All 33 current-head
checks are successful or intentionally skipped.

## Risks

- Uses one additional short-lived verification runner per release
verification. The existing AWS exact-master restriction remains in
place.
- The Rust lane warms debug dependencies so its cache save also serves
protocol tests. Both lanes always rebuild workspace code.
- A workflow regression could omit a check. The new coverage test
compares the matrix checks directly with `check:all`; Cloud readiness
depends on the complete reusable workflow.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:33:20 -07:00
DottaandPaperclip 7e6d512597 fix(onboarding): make chief-of-staff hiring reliable (#13317)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The first agent helps the board define work and hire other agents.
> - That agent can have the general role while its instructions require
hiring skills.
> - Missing skills and blocked schema discovery make valid requests
fail.
> - Repeated confirmation and invalid waiting guidance can turn these
failures into extra runs.
> - This PR supplies the required skills, opens read-only schema
discovery, and corrects the guidance.
> - The agent can complete an authorized hire while company approval and
duplicate checks still apply.

## Linked Issues or Issue Description

Refs #13068 — the first-task onboarding flow that this change repairs.
Refs #12029 — related drift between the sandbox allowlist and bundled
hiring guidance. This PR adds schema access; it does not replace the
earlier hiring-route fix.

**What happened?**

A general-role onboarding chief received hiring instructions without the
core hiring skills. Sandbox requests to the documented OpenAPI endpoint
failed. The agent then guessed question and hire payloads. The persona
required new confirmation after validation errors and described waiting
states that agents cannot set.

**Expected behavior**

A direct request authorizes the requested hire. The chief asks only for
material missing details, uses valid API payloads, and completes the
task. Formal company approval gates still apply. A saved human-input
card gives the task a valid waiting state.

**Steps to reproduce**

1. Create an onboarding chief with role `general` through the board.
2. Ask it to hire a friendly robot with a supplied name and
responsibilities.
3. Check its assigned skills, schema requests, question cards, hire
requests, and final task state.

**Paperclip version or commit**

Reproduced on the first-task onboarding implementation after #13068. The
live local verification used this branch at `112f44610`.

**Deployment mode**

The original failure used a hosted sandbox with legacy Codex ACP. Live
verification used an isolated local instance and real `codex_local`
execution. Queue and HTTP/2 transport access is covered by automated
tests.

## What Changed

- Give board-created onboarding chiefs the existing core skills
regardless of role. Preserve explicit skill version pins, including
aliases. Keep ordinary general-agent defaults and authorization checks.
- Allow exactly `GET /api/openapi.json` through both sandbox bridge
transports.
- Publish validator-tested question, free-text, hire, and waiting
examples. Regenerate the runner API reference and capability inventory.
- Clarify direct authorization, material ambiguity, and correction of
confirmed pre-creation validation failures. Preserve uncertain-outcome
reconciliation, duplicate protection, and company approval gates.
- Align disposition instructions with agent permissions and the saved
human-input waiting path.

## Verification

- After rebasing onto current `master`: 69 targeted server tests, 110
queue/HTTP2 bridge tests, and 4 capability inventory tests passed. These
cover core skill defaults, version pins, actor restrictions, schema
access, published examples, hire validation, idempotency, and approval
gates. Waiting recovery tests and live question flows also passed before
the rebase.
- `pnpm -r typecheck` and `pnpm build` passed again after the rebase.
Frozen dependency installation and both generated capability checks
passed.
- Ran the full `pnpm test:run` suite. The initial run had 14 failed
server files due to local database resource limits, a missing built test
fixture, and socket failures. All 14 files passed after fixture repair
and isolated retries. UI, CLI, workspace packages, database tests, and
all 145 serialized server files passed.
- Real one-request hiring replay: one hire, one successful run, task
done in 2m16s. No repeated approval or recovery escalation.
- Real two-turn browser conversation: start with an unspecified hire,
then supply a name and friendly robot responsibilities. One
clarification card, one hire, two successful runs, task done in 3m27s of
execution. No failed writes, confirmation cards, or recovery actions.
- Assigned the hired robot a welcome-message task through the browser.
It produced a warm message under 100 words and finished in one
successful 66-second run, with no questions or recovery actions.
- The two-turn flow still asked an optional preferences question and
gave a technical final reply. These are remaining presentation limits.
- Greptile: 5/5 on `b71f83ba2`, with zero unresolved review threads.
Fixed its generator finding and passed 1,655 published-example/runtime
API tests plus server typecheck. All latest-head CI checks are green (32
passed; 2 unrelated Storybook checks skipped). The signoff-policy
browser test initially timed out while waiting for an approver run. Its
shard passed on one rerun without code changes. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/34698211049).

## Risks

- Onboarding chiefs receive more default skills. Ordinary general agents
retain existing defaults, and explicit versions take precedence.
- Prompt guidance can affect model behavior. The live replays are
examples, not a guarantee that every model follows the guidance.
- Retry guidance applies only when validation confirms that nothing was
created. Uncertain outcomes still require checking existing agents.
- No database migration or new public endpoint. Existing company
boundaries, approval gates, and bounded recovery remain in force.

## Model Used

OpenAI Codex, model `gpt-6-astra`, with reasoning, tool use, code
editing, and live browser verification. The exact context-window size is
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 12:59:42 -05:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00