mirror of
https://github.com/paperclipai/paperclip.git
synced 2026-10-02 10:16:20 +08:00
5ee9e751fbb38787c055ee793c3412ecfc89da2e
4598
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5ee9e751fb |
fix(runner): discover assigned tools when direct catalogs exceed limits (#14218)
Keep assigned app tools accessible when the combined Runner catalog exceeds its operation or byte limits. Reserve task and completion tools, then expose bounded discovery and call tools for large catalogs. Fetch oversized schemas in reauthorized chunks without blocking later search results. Retain task ownership, work-mode restrictions, pinned assignments, current gateway authorization, approvals, and audit. Small catalogs stay direct. Validation: 46 focused server tests, two Runner capacity tests, local typecheck/build, 52 passing CI checks, and Greptile 5/5 with no open findings. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.927.0-canary.3 |
||
|
|
640dee1802 |
fix(runner): acquire reusable leases when switching to warm sessions (#14187)
## Thinking Path > - Paperclip coordinates agent work across task turns. > - Warm runners need reusable sandbox leases. > - Lease acquisition read only the environment reuse setting. > - Startup then read the agent's new warm setting and rejected the ephemeral lease. > - This fix requests reuse before acquisition, with provider checks intact. ## Linked Issues or Issue Description **What happened?** Switching an existing native runner task from per-turn to warm fails with `runner_warm_environment_requires_reusable_lease`. Environment `reuseLease` defaults to false. Acquisition therefore creates an ephemeral lease, and capability narrowing correctly denies reuse. **Expected behavior** Changing lifecycle between turns requests a compatible lease without replacing the task workspace or changing the shared environment. **Steps to reproduce** 1. Start a native runner task with inherited lifecycle and environment reuse disabled. 2. Change the agent from per-turn to warm. 3. Continue the task on a reusable-capable sandbox provider. **Paperclip version or commit** Reproduced on `a6c4e7a`. Related: #12904 established warm workspace continuity; this fixes the earlier lease-selection mismatch. ## What Changed - Pass agent settings into acquisition and derive a run-scoped reuse request. - Respect explicit environment lifecycle overrides; leave other adapters unchanged. - Keep provider capability, ownership, cleanup and restore gates intact. - Add orchestration and database-backed transition tests; document the contract. ## Verification - Red: three new assertions failed before the fix. - Green: 156 tests pass in environment-run-orchestrator, environment-runtime, native-sandbox-lifecycle, and environment-execution-target-capabilities. - Database-backed regression proves ephemeral → reusable → resumed lease, same workspace, and unchanged stored environment. Provider RPCs are mocked, not live Daytona. - `git diff --check` passes. - Latest-head CI passes typecheck, build, tests, runner checks, E2E, and canary dry run. Greptile: 5/5; no open threads. - Recovery uses the persisted lifecycle. Direct regression: red before, six cases green after. CI runs the added Vitest cases. - Full local checks were unavailable (missing dependencies and earlier memory limits). ## Risks Warm mode now requests retained resources despite an environment's default `reuseLease:false`. Existing idle timeout and cleanup still apply. Unsupported providers remain denied. No migration, credential [REDACTED], or shared-environment mutation. ## Model Used OpenAI Codex agent; reasoning, code editing and test execution. Exact model ID and context size were not exposed to this run. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details available) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal Paperclip ticket [REDACTED] or instance-derived details - [x] I have run targeted tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f2ed0b65c4 |
fix(runner): enable API tools by default (#14186)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner exposes tools for company tasks.
> - API search and call tools cover operations without a dedicated tool.
> - The current default hides these tools unless an operator sets an
environment variable.
> - This pull request enables the tools when that variable is absent.
> - Operators can still disable the tools or restrict them to selected
companies.
## Linked Issues or Issue Description
Refs #13003, which added the guarded API tools.
**What happened?**
The native runner does not advertise `search_api` or `call_api` with the
default server configuration.
**Expected behavior**
The tools are available without a special environment variable. Existing
authorization checks still apply.
**Steps to reproduce**
Remove `PAPERCLIP_RUNNER_API_TOOLS_ENABLED` and
`PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS`. Create a normal runner
authority. Inspect its tool definitions.
**Paperclip version or commit**
canary/v2026.927.0-canary.2
|
||
|
|
a6c4e7a8d1 |
fix(heartbeat): cancel obsolete execution continuations before dispatch (#13761)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A comment can queue execution before the task changes state. > - The task may be completed, cancelled, deleted, or reassigned before setup checks ownership. > - The ownership guard must prevent that obsolete execution from starting. > - This pull request records that guard outcome as cancellation and settles the wake request. > - Missing history and authorization errors remain failures that operators can investigate. ## Linked Issues or Issue Description **What happened?** A comment can queue a run just before a user marks its task Done. The ownership guard stops setup before adapter dispatch, but the run becomes a setup failure and can put the agent into an error state. **Expected behavior** Cancel obsolete work when the typed task ownership guard rejects it. Keep genuine setup errors visible. Do not change the task's terminal state or current owner. **Steps to reproduce** Queue a task continuation, then mark the task Done or Cancelled before continuation setup. The run should settle as cancelled without invoking the adapter. Deleting or reassigning the task must also prevent dispatch. Missing source history or explicit user authorization must still fail. **Paperclip version or commit** Refreshed against master `b2e9e82f053af14777fd68e8834fff6bb64c84c4`. Related: #13546 covers lost issue-lock claims, and #13888 covers persisted continuation decisions and retry budgets. This change handles the final continuation ownership guard during setup. Thanks to @MrBlackTongue for the original typed cancellation implementation and regression coverage. This update preserves the contributor's commits, resolves the master conflict, and narrows cancellation to task ownership invalidation. ## What Changed - Add a typed `StaleExecutionContinuationError` for `continuation_task_ownership_changed`. - Use existing cancellation settlement for that typed guard, including the run, wake request, issue execution ownership, and agent state. Suppress immediate recovery of obsolete work. - Preserve failure classification for missing source context, missing user authorization, and untyped errors, including an untyped error with identical text. - Cover Done and Cancelled tasks with real database checks. Retain company and assignment guard coverage. Verify cancellation performs no adapter dispatch or automatic replay. - Document the cancellation event and its error classification in the run-log guide. ## Verification Current head: `b875486e81`. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm exec vitest run server/src/__tests__/heartbeat-process-recovery.test.ts server/src/services/execution-continuation.test.ts --maxWorkers=1 --no-file-parallelism -t 'authorized continuation context|obsolete continuation setup|untyped continuation setup failures'`: 17 passed; 316 unrelated tests excluded by the name filter. - `pnpm test:run`: local full run is still finishing. It is not a clean pass: the observed failures are three company-skills cache checks (`EACCES` while renaming read-only cache directories on macOS), one tool-access assertion, and one comment-wake timeout. The three cache failures reproduce in isolation in unchanged code. The tool-access check passes in isolation, and the entire comment-wake file passes (29 tests), including explicit feedback after completion. Final full-run totals will be added when available. - [CI](https://github.com/paperclipai/paperclip/actions/runs/36282532652): all gates green on this head. An unchanged Runner workspace-diff test initially returned no diff; an isolated check passed, and the failed job plus its dependent gate passed on retry with no code change. Original failed job: 108517010989. - Greptile: 5/5 on the full current head, with no unresolved review threads. Branch is current with master and conflict-free. - Diff check and local secret/PII scan passed. No customer data or private deployment identifiers are included. ## Risks The classification change is limited to the typed ownership guard. A task with missing history remains a failure rather than being treated as an expected cancellation. Existing company, task-owner, terminal-state, and authorization checks still prevent dispatch. Cancellation retains its specific reason in the run log and does not grant a retry. No schema, dependency, or API changes; no migration is required. ## Model Used OpenAI GPT-6 via Codex assisted investigation, implementation, code review, and test execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the bug issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the focused regression tests locally and they pass; full-suite results are tracked above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation - [x] I have considered and documented the risks - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before merge Co-authored-by: Paperclip <noreply@paperclip.ing> Co-authored-by: Devin Foley <devin@paperclip.ing>canary/v2026.927.0-canary.1 |
||
|
|
b2e9e82f05 |
fix: stop remote Grok runs before continuing queued messages (#14100)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The execution service owns each run and saves messages sent while it runs. > - Interrupt must stop the current executor before it delivers those messages. > - Remote Grok commands did not register the host cancellation control. > - A cancelled task run could still write Done and prevent queue recovery. > - This pull request connects remote cancellation and revokes cancelled run writes. > - Saved input can use the existing queue admission rules after verified cleanup. ## Linked Issues or Issue Description **What happened?** Interrupting a queued message marked a remote Grok run cancelled before its sandbox stopped. The old run could still post a reply and mark the task Done. Its saved follow-up remained deferred behind execution recovery. **Expected behavior** Stop revokes run write authority and waits for verified termination. Saved messages remain durable and enter one successor through normal admission after cleanup. **Steps to reproduce** 1. Run a task with `grok_local` in a remote sandbox. 2. Send a follow-up and use Interrupt while the command runs. 3. Let the old command attempt a task status update after cancellation. 4. Observe the task disposition and the saved message queue. **Paperclip version or commit** The gap is present in master at `d3e0f0a238`. **Deployment mode** Authenticated server with a Daytona sandbox. Related work: #14028 and #14046 handle bounded continuation. #13291 covers infrastructure interruption and verified remote cleanup. #13332 addresses atomic recovery holds. This change handles direct Grok operator cancellation and stale task writes. ## What Changed - Register remote Grok cancellation before preparation. Keep command ownership until the host confirms sandbox termination. - Reuse the sandbox cancellation boundary for the direct CLI invocation. Reject fresh attempts after cancellation and preserve workspace restore failure evidence. - Reject writes from cancelled task JWTs and runs with a pending stop. Preserve diagnostic reads and existing conversation error codes. - Recheck run authority under a database lock before task updates and interaction responses commit. - Preserve authorized handoffs that stop their own run. Only the server-issued stop receipt for that request permits the final task update. - Add tests for hung commands, unverified stops, early cancellation, copy-back failures, late Done, late interaction responses, authorized handoffs, exact lease receipts, and one queue successor across concurrent restart sweeps. - Document the cancellation and write-authority contract. ## Verification - Targeted adapter, cancellation-boundary, authentication, queued-message, interaction-service, and activity-route tests passed. The expanded run passed 214 tests; one new test had an incomplete fixture. After correcting the fixture, all 8 selected follow-up cases passed. - `pnpm -r typecheck`: passed on `179c86caf1bf0d89914a503d46e24af7e4b8c557`. - `pnpm build`: passed on the same commit. - `pnpm test:run`: the general-server group completed with 13,521 passed, 99 skipped, and 18 failed tests. It then stopped, so the remaining local groups did not run. Five Slack, email, and wake-batching failures passed on focused reruns after correcting the local environment. The remaining 13 failures reproduce as `EACCES` on rename in unchanged skill-cache code on macOS. Two custom-image suite setup hooks also failed to start embedded PostgreSQL after the machine exhausted shared-memory slots; all 31 tests in that file passed on rerun after the local resource issue was resolved. CI covers all test groups. - CI: 53 checks passed and 2 were skipped on the latest commit, including the aggregate verification gate. The last server shard passed on its single rerun after a preview-server startup timeout. The affected file also passed locally with 28 passed and 3 skipped. - Greptile: 5/5 on the latest commit. Both review threads are resolved. - No live deployment or staging task mutation has been performed. ## Risks - Stopping the sandbox can prevent file copy-back. The result preserves workspace restore failure evidence; termination does not imply restored files. - If provider termination fails, the adapter keeps ownership of its outstanding command and does not acknowledge Stop. - The write restriction now applies to ordinary cancelled tasks. Reads remain allowed. Task and interaction checks add a shared run-row lock to agent mutations. An exact server-issued receipt permits the task request that stopped its own run to complete its handoff. - Existing terminal tasks are not reopened automatically. An operator must correct a historical late Done before its saved queue can continue. - No schema migration or UI change. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository inspection, code execution, and test tools. The precise backend revision and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.927.0-canary.0 |
||
|
|
01d9a12185 |
fix: make keyboard shortcut enablement a personal preference (#14141)
Store keyboard shortcut enablement per user and expose it in Profile settings. Co-Authored-By: Codie <Codie@users.noreply.github.com> |
||
|
|
d3e0f0a238 |
fix(server): validate CLI auth challenge IDs (#14095)
Reject malformed CLI challenge UUIDs before database access while preserving secret and authentication precedence. Document the HTTP responses and pin the supervisor test fixture to the CI-selected Node executable. Validated by 21 focused tests, root typecheck/build, and full CI. Mac broad-suite baseline limitations are documented in the PR. Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.926.0-canary.4 nightly/v2026.926.0-nightly.0 |
||
|
|
1e3a148f46 |
fix(connections): retain safe broker rejection diagnostics (#14098)
Preserve fixed broker rejection reason codes within strict size/time limits while retaining public error codes and status. Unknown bodies remain generic; no raw response or credential material enters the error. Validated by 33 consumer tests, a synthetic producer HTTP contract fixture, root typecheck/build, and full CI. Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
ffa32373bc |
fix(runtime): honor provider acquisition timeout defaults (#14097)
Declare provider acquisition budgets so slow Daytona creation does not hit the host’s 30-second fallback. Bound creation and setup to one deadline and preserve scoped cleanup ownership after timeout. Legacy drivers retain their original call shape. Validated by 238 provider/manifest tests, focused database and heartbeat regressions, root and standalone provider typecheck/build, and full CI. Local broad tests also expose recorded Mac baseline limitations. Greptile 5/5. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.926.0-canary.3 |
||
|
|
7f3c06dac4 |
refactor(ui): remove the legacy Cloud organization switcher (#14061)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its sidebar provides navigation between organizations. > - The generic organization switcher slot lets installed plugins own that navigation. > - Both Core UI shells still contain the old Cloud portfolio menu. > - Cloud now uses its private Account plugin for this menu. > - This pull request removes the duplicate Cloud menu and keeps the built-in company menu. > - This reduces Cloud-specific code without changing the plugin host contract. ## Linked Issues or Issue Description Refs #13832 and #13854. Related #14060 changes the account popup, not this organization switcher. **What existing behavior does this improve?** Organization navigation in both sidebar shells. **Current behavior** Core retains Cloud portfolio fetching, stack rows and stack-entry links behind the plugin replacement slot. **Proposed behavior** The installed switcher plugin owns Cloud navigation. Core lists local companies when no usable replacement exists. The Members-page Cloud invitation action keeps its existing portfolio API; it is a live caller, not an old-image fallback. ## What Changed - Remove Cloud portfolio queries, stack rendering and Cloud creation/entry branches from both built-in menus. - Remove the unused stack-entry URL helper. - Keep company selection, ordering, invitations, logout and plugin error handling. Hide local company creation on managed hosts, where the server forbids it. - Update switcher tests and the navigation contract. ## Verification - `pnpm -r typecheck` passed, including Rust checks with the installed Cargo toolchain on PATH. - `pnpm exec vitest run --project @paperclipai/ui`: 637 files and 6,727 tests passed. - Focused switcher, plugin host and Cloud link tests: 26 passed after the managed-host creation guard. - `pnpm check:token-gates` passed. - Full `pnpm test:run` was attempted, then stopped after failures in unchanged server tests. Targeted reproduction found an ancestor skills-directory collision for Slack and macOS EACCES errors renaming the company skills cache. Other local failures appeared in email connector skill setup and a process-turn test. This is not a local full-suite pass; clean Linux CI covers the complete suite. - Full `pnpm build`, UI production build and Storybook build passed. All [latest-head CI checks](https://github.com/paperclipai/paperclip/actions/runs/36201814444) passed, including the full Linux test matrix, browser tests, typecheck, build and release checks. Greptile is 5/5 with no unresolved comments. ## Risks - A managed host without a usable switcher plugin now gets the ordinary company menu. It no longer gets the old Cloud portfolio menu, and local company creation remains unavailable there. - The Cloud portfolio endpoint remains required by the Members-page invitation action. This PR does not remove that live endpoint or change its authorization. - No database, authentication, plugin protocol or deployment changes. ## Model Used OpenAI Codex, GPT-6, with reasoning, repository inspection and code execution. The exact deployment identifier and context-window size are not exposed by this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.926.0-canary.2 |
||
|
|
4ca404b49a |
fix: record safe sandbox restore failure diagnostics (#14064)
Record a bounded diagnostic for failed workspace and staged-asset restores. Preserve the original error, retry policy, and archive safety checks. Never copy raw provider messages, credentials, paths, or asset names into the log. Nested failures log once; safe fields survive throwing property getters. Verified 129 focused restore/Claude tests, typecheck/build, and green full PR CI. Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.926.0-canary.1 |
||
|
|
c3ddb288b1 |
fix: validate work product execution workspace references (#14063)
Reject invalid and cross-company execution workspace references with a useful 422 before changing the work product. Hold the validated reference through the transaction so concurrent deletion cannot turn validation into a 500. Verified 13 focused tests, full typecheck/build, and green PR CI. Greptile 5/5 with no unresolved comments. The known UI copy-toast flake passed on an unchanged-commit retry and in a focused local run. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.926.0-canary.0 |
||
|
|
96bf004a79 |
fix: use persisted state for lifecycle continuation and retry budgets (#13888)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.6 |
||
|
|
6bc830b62b |
fix(recovery): escalate an issue whose interrupted run has spent its retry budget (#14046)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The stranded-issue sweeper is what puts an assigned issue back on a
live path when its run dies.
> - A failed or interrupted run gets a bounded transient retry; when
that budget is spent, the retry scheduler queues nothing and reports the
exhaustion.
> - The sweeper treated that "nothing queued" like every other "nothing
queued" and skipped the issue, on every tick, forever: `in_progress`, no
run, no path, no notice.
> - Three server restarts in a row (a deploy storm) are enough to spend
the budget, so this is reachable in ordinary operation.
> - This pull request makes the sweeper escalate a spent budget to a
board-owned recovery action, the same visible `blocked` state its other
dead ends use.
> - The benefit is that no issue can sit assigned and silent after its
retries run out.
## Linked Issues or Issue Description
No existing issue found. Related work: #14028 (this branch's
predecessor: queue drain-time wakes, retry interrupted corrective runs)
fixed two neighbouring gaps but not this one.
**What happened?**
A routine-created issue's run was interrupted by a graceful server
shutdown, and both bounded transient retries were interrupted by the
next two shutdowns. The retry scheduler logged "Bounded retry exhausted
after 2 scheduled attempts; no further automatic retry will be queued"
and stopped. On every tick after that, `reconcileStrandedAssignedIssues`
reached the generic continuation lane, `enqueueStrandedIssueRecovery`
delegated to `scheduleRecoveryRetry`, which returned null (exhausted),
and the sweeper counted the issue as `skipped`. The issue stayed
`in_progress` with no run, no recovery action, and no comment for hours
until a person woke the agent by hand. The same hole exists in the
assigned-`todo` dispatch lane.
**Expected behavior**
When the transient retry budget for an interrupted or failed run is
spent, the sweeper should treat it as the dead end it is: escalate the
issue to `blocked` with a board-owned recovery action and a notice,
exactly as it does when a continuation retry chain or an assignment
retry chain is exhausted. Leaving the budget spent across restarts is
correct and unchanged; leaving the issue silent is not.
**Steps to reproduce**
1. Assign an agent using a conversation adapter (its interrupted runs
carry `conversationContinuation`, so legacy reconciliation does not
terminalize them) an issue and let a run start.
2. Interrupt the run with a graceful shutdown, let the transient retry
start, interrupt it, and repeat once more so `scheduledRetryAttempt`
reaches 2.
3. Run the stranded-issue sweep. Before this change: `skipped` every
tick, issue `in_progress`, no run, no recovery action. After:
`escalated`, issue `blocked`, one active board-owned
`issue_recovery_actions` row, a "No live execution path" notice.
**Paperclip version or commit**
master at
canary/v2026.925.0-canary.5
|
||
|
|
bd6caf51bb |
fix: preserve restore failure results and stop unsafe retries (#14035)
Preserve agent output and earlier execution errors when workspace restore fails. Report the restore phase and confirmed saved-plan links. Require verified repair before retrying unsafe archives, while preserving approval states and the retry budget. Verified with full CI, 506 focused regression tests, and Greptile 5/5 with all review threads resolved. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.4 |
||
|
|
5b09d66183 |
fix(heartbeat): keep drain-time wakes queued and retry interrupted disposition handoffs (#14028)
## Thinking Path
> - Paperclip is the open source app people use to manage AI agents for
work.
> - The server's heartbeat scheduler queues, admits, and recovers agent
runs.
> - A control plane arms an instance task drain before it restarts the
process, so no new run starts mid-restart.
> - Two recovery paths treated that transient hold as a permanent
verdict: a wake that arrived during a drain was written as `skipped` and
never replayed, and a corrective disposition run that the restart
interrupted counted as an exhausted attempt, so the issue went to
`blocked`.
> - Both outcomes leave an assigned issue with no run and no path until
a person notices.
> - This pull request keeps drain-time wakes in the durable queue and
gives an interrupted corrective run the same bounded transient retry any
interrupted run gets.
> - The benefit is that a graceful restart never strands or blocks an
issue by itself.
## Linked Issues or Issue Description
No existing issue found. I searched open PRs that touch the task drain
(#13527 adds opt-in termination of active runs; #13965 handles deferred
wakes after a stale cancellation). Neither replays drain-time wakes or
changes the corrective-handoff escalation.
**What happened?**
1. An operator accepted a plan confirmation while the instance was in a
task drain (a control plane armed it before a deploy restart).
`enqueueWakeup` wrote the assignee's continuation wake with `status:
"skipped"` and `heartbeatSkip.reason: "task_drain"`. The drain is
process-local and cleared on the restart seconds later, but nothing
replays skipped wakes. The issue stayed in `todo` with no run for 20
hours until a person retried it by hand. The stranded-issue sweeper did
not help: a `todo` issue whose latest run succeeded is treated as
deliberately handed back.
2. On another issue, the watchdog correctly raised
`successful_run_missing_state` and queued the single corrective handoff
run. A graceful server shutdown (the same deploy restart) interrupted
that run with `server_shutdown_interrupted`. On the next boot,
`reconcileStrandedAssignedIssues` saw a corrective run at
`handoffAttempt >= maxHandoffAttempts` and escalated the issue to
`blocked` with a board-owned `missing_disposition` recovery action. The
agent never got to finish one corrective attempt.
**Expected behavior**
A task drain holds admission, not the request. A wake that arrives
during a drain must run once the drain lifts or the process restarts. An
interrupted corrective run is not evidence that the agent could not
choose a disposition; it should be retried like any interrupted run, and
escalate only when that retry budget is spent or a finished attempt
still leaves no disposition.
**Steps to reproduce**
1. Assign an agent an issue and `POST /api/instance/task-drain` with a
Cloud control assertion (or call `startTaskDrain({})` in a test).
2. Trigger any wake for that agent (assignment, comment, or an accepted
`request_confirmation`).
3. Observe the `agent_wakeup_requests` row: `status = skipped`, `reason
= heartbeat.scheduling_suppressed`, `payload.heartbeatSkip.reason =
task_drain`. Stop the drain or restart: no run is ever created for that
wake.
4. For the second case: let a `finish_successful_run_handoff` run be
interrupted by SIGTERM during a graceful shutdown, then boot. The issue
moves to `blocked` on the startup sweep without any retry.
**Paperclip version or commit**
master at
canary/v2026.925.0-canary.3
|
||
|
|
bd20309323 |
fix(ui): keep mobile task pickers above the keyboard (#14022)
## Thinking Path > - Paperclip helps people manage AI agents and their tasks. > - The New Task dialog lets a person select an assignee and a project. > - A mobile browser reduces the visible viewport when the software keyboard opens. > - The dialog accepted invalid viewport data, and the pickers stayed inside a transformed container. > - This behavior could collapse the dialog or move a picker search field above the visible area. > - This pull request validates viewport data and puts mobile pickers in the visible viewport. > - The benefit is that a mobile user can see and use each picker while the keyboard is open. ## Linked Issues or Issue Description **What happened?** On a mobile device, the Assignee and Project pickers in the New Task dialog could move above the visible viewport. A short invalid viewport value could also collapse the dialog to a line. **Expected behavior** The dialog and each open picker must stay in the visible viewport while the software keyboard is open. **Steps to reproduce** 1. Open the New Task dialog in a mobile browser. 2. Open the Assignee picker or the Project picker. 3. Focus the picker search field so that the software keyboard opens. 4. Observe that the picker can move above the visible viewport. **Paperclip version or commit** `efce9356b5` **Deployment mode** Local dev and hosted browser UI. ## What Changed - Ignore zero, negative, and non-finite Visual Viewport measurements. - Keep the last safe dialog geometry until the browser gives a valid measurement. - Put entity pickers outside the transformed dialog container. - Size and position the mobile picker from the dialog Visual Viewport values. - Add unit tests for invalid viewport recovery. - Add Chromium tests for the Assignee and Project picker states. ## Verification - `pnpm exec vitest run ui/src/components/NewIssueDialog.test.tsx` passed with 34 tests. - The Chromium viewport test passed 35 of 35 runs with five repeats and no retries. - `pnpm --filter @paperclipai/ui typecheck` passed. - `pnpm check:token-gates` passed. - `pnpm build-storybook` passed. - `pnpm -r typecheck` passed. - `pnpm build` passed. - A full local Vitest run reached restricted workspace-runtime tests that require sibling worktree and runtime writes. Remote CI will run the supported test environment. ## Risks - Risk is low because the new picker layout applies only to mobile widths. - The layout depends on Visual Viewport data when the browser supplies valid values. - Unit and browser tests cover invalid data, mobile pickers, tablet layout, and desktop layout. > This change fixes a focused UI bug. It does not duplicate a core feature in `ROADMAP.md`. ## Model Used - OpenAI Codex with GPT-5. The agent used reasoning, repository tools, code execution, and browser automation. The runtime did not expose the context window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.2 |
||
|
|
e2f1a66aa7 |
feat: support default-hidden experimental settings (#13980)
## Thinking Path > - Paperclip is the open source control plane for AI-agent companies. > - Operators can hide settings that their users must not change. > - The server shares the effective restrictions with the UI and settings API. > - An explicit list must change whenever a new experimental flag is added. > - This pull request adds a wildcard with named exceptions to the existing setting. > - Core applies the policy to its current catalog, so new flags stay hidden automatically. ## Linked Issues or Issue Description Related: #11823 introduced settings visibility. #13907 added workspace isolation visibility. I searched related PRs and issues and found no duplicate wildcard implementation. **What existing behavior does this improve?** Operator control of experimental setting visibility through `PAPERCLIP_HIDDEN_SETTINGS`. **Subsystem affected** Shared settings policy and its existing server health and mutation consumers. **Current behavior** Operators must name every hidden experimental toggle. A new Core flag can become visible until the operator updates that list. **Proposed behavior** `instance.experimental.*` hides current and future experimental toggles. Entries such as `!instance.experimental.enableEnvironments` leave named controls available. Explicit hidden keys and the hidden parent page take precedence over exceptions. **Reason and benefit** Operators can maintain a short list of allowed controls instead of a second copy of Core's full feature catalog. **Breaking changes** Existing explicit lists and unset configuration keep their behavior. The new syntax is opt-in. Older images ignore it, so operators must retain explicit restrictions until those images are upgraded. Visibility does not change feature values. ## What Changed - Expand the wildcard into concrete catalog keys in the shared parser. - Limit exceptions to known experimental controls and preserve explicit restrictions in either input order. - Test a synthetic future catalog addition, duplicate and invalid entries, API rejection, same-value echoes, and the effective health payload. - Document the syntax and the transition for deployments with mixed image versions. ## Verification - Targeted parser, future-catalog, health, and settings-route tests pass: 95 tests across four files. - `pnpm -r typecheck` passes, including Rust checks, with the installed Cargo directory on PATH. - `pnpm build` passes. - All current-head CI gates pass, including the full test shards, Rust, build, browser E2E, and canary dry run: https://github.com/paperclipai/paperclip/actions/runs/36083292578. One unchanged runtime-exposure cold-start test passed on its first retry. - The full local `pnpm test:run` did not pass on macOS/Node 25: the first server group reported 13,356 passed, 18 failed, and 99 skipped, with six failed files (including two failed suite setups). Failures were in unchanged runtime/company skill cache, chat/email connector fixtures, embedded-Postgres setup, and workspace cleanup tests. A standalone filesystem probe reproduced the read-only-directory rename permission failure. Missing connector fixture paths, database startup failures, and two integration assertions also occurred; the remaining local groups were not reached after this group failed. The corresponding CI lanes all pass. These local failures are not claimed as fixed by this PR. - Browser suites were not run because this changes the shared policy, not UI rendering or browser workflows. Health payload and route tests cover the shared UI/API contract. ## Risks A malformed exception remains hidden and is reported as unknown. Exceptions cannot override an explicit hidden toggle or parent page. Older images ignore wildcard syntax; keep their explicit list during a mixed-version rollout. No schema or feature-value changes are included. ## Model Used OpenAI GPT-6 through Codex, with reasoning, tool use, and code execution. The exact runtime variant and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
1dbffb4c04 |
fix(daytona): clean up sandbox allocation after create failure (#13979)
## Thinking Path > - Paperclip manages AI agents and their execution environments. > - The Daytona provider creates sandboxes before it returns a lease. > - Daytona can allocate a sandbox, then time out while waiting for it to start. > - The SDK throws without returning the allocated sandbox handle. > - Paperclip previously lost the resource identity and could leave a running sandbox behind. > - This pull request keeps an exact creation identity and deletes a matching failed allocation. ## Linked Issues or Issue Description Related: #13882. Searches for Daytona create cleanup and orphan fixes found no duplicate for this path. **What happened?** In [Grok qualification campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743), Daytona sandbox creation exceeded 300 seconds. No native runner or provider session started. The SDK threw, but an ownership-filtered provider lookup found the allocated sandbox still running. The disposable resource was then deleted separately. The original attempt and its failure remain retained. **Expected behavior** A failed create must retain enough identity to clean up an allocated sandbox. Cleanup must never delete another attempt's resource. If the provider cannot confirm cleanup during sandbox lease acquisition, the host must retain a durable cleanup record and retry the exact owned allocation after restart. **Steps to reproduce** 1. Have Daytona allocate a sandbox for an environment acquisition. 2. Make the SDK throw while waiting for startup, before it returns the sandbox handle. 3. Before this fix, acquisition fails without deleting the allocated sandbox. 4. The regression tests reproduce this failure and verify exact-owner deletion. **Paperclip version or commit** Observed at `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`. The same creation path exists on master; this fix starts from `efce9356b`. ## What Changed - Assign each cold create a unique provider name and creation-attempt label. - On create failure, look up that exact name and verify every ownership label before deletion. - Wait for provider deletion with bounded lookup and deletion calls. - Preserve the original create error after successful cleanup. Report both errors and the provider name when cleanup cannot be confirmed. - Treat a missing lookup after an uncertain create as unconfirmed cleanup: a delayed provider request might still create the sandbox. - Transfer only validated ownership fields through the acquire-lease RPC error; provider exceptions and arbitrary error data are not serialized. - Persist a pending-cleanup lease before the host retries deletion. Reuse the existing cleanup sweep and durable spool, retaining the original environment scope after its row is deleted. - Fence retries by company, environment, run, provider account, unique creation name, and every ownership label. Missing provisional-name lookups remain unresolved. - Journal an ownership-verified provider ID before host-driven deletion. Its absence then confirms cleanup after a lost deletion reply or final database update; failed journal writes block deletion. - Keep ownership labels separate from the mutable input Daytona modifies during create. - Report confirmed host-side deletion accurately while still rejecting the failed acquisition. - Add protocol, malformed-evidence, cross-scope denial, controller-restart, and deleted-environment regressions alongside the original bounded cleanup tests. ## Verification - SDK suite: 92 passed. Daytona plugin suite: 277 passed; six opt-in live tests skipped. Environment-runtime suite: 103 passed, then all four focused observation/recovery variants passed after adding a foreign-observation case. SDK and provider TypeScript checks and `git diff --check` pass. - Regressions failed before the respective fixes: missing durable ownership, misleading successful-cleanup message, and SDK mutation of the ownership labels. - Controlled real-Daytona proof at `d2ac1c95b32ca64daf039c0427137657f351c0e0`: create a real sandbox; inject an SDK failure and a failed immediate lookup; persist the validated ownership envelope; have a child journal the observed provider ID before deleting; deliberately drop its successful deletion receipt; reconcile absence from another fresh process and confirm it with a separate provider read. Passed, with no manual cleanup and zero model calls. The live proof uses an owner-only envelope file; actual host database/spool recovery is tested separately in the environment-runtime suite. - The preceding controlled attempt failed because the real SDK mutated the labels object. That failure is retained. The harness deleted its exact owned sandbox, and a fresh lookup confirmed absence. The fix copies the SDK input labels separately from the ownership snapshot. - [Repository CI](https://github.com/paperclipai/paperclip/actions/runs/36095925399) passes on this exact head, with a clean 5/5 review. The unchanged Cursor remote-command test passed its one bounded rerun after a 10-second timeout; the same test had already passed on the combined tree. The original failure is retained. Earlier `199f0852` CI and immediate-deletion proof are retained, without being promoted to the new source. - Combined-source verification uses temporary PR #13990 against the Grok feature branch. No Docker or Rust build runs on the developer laptop. ## Risks Creation failures now add up to ten seconds for lookup and fifteen seconds for deletion confirmation. The SDK still owns the initial creation timeout. If a request materializes only after the failed lookup, its durable pending record remains eligible for later cleanup; a never-observed creation name is never reported deleted from a 404. A previously observed provider ID can be reconciled as deleted. The host and SDK must be deployed together. Durable handoff applies to plugin-backed sandbox lease acquisition used by heartbeat and runner login. Probe and custom-image interactive setup retain immediate cleanup and explicit failure reporting, but do not use this lease-recovery path. A worker crash before delivering the failure envelope still cannot be recovered through this mechanism. A cleanup error never returns a lease. Successful creation behavior is unchanged except for the provider name and ownership label. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
f56aee5423 |
fix(evals): capture complete durable run event streams (#13977)
## Thinking Path > - Paperclip manages AI agents and records their durable work outcomes. > - Product E2E checks those outcomes through the browser and public API. > - A successful long run can emit more than 1,000 durable events. > - The harness read one page and missed the later completion evidence. > - This pull request reads every page before it checks runtime invariants. > - Invalid or incomplete capture still fails. A missing page cannot produce a pass. ## Linked Issues or Issue Description Related: #13882 (Grok qualification). A duplicate search found no existing event-pagination fix. **What happened?** The local structured-question case in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537) completed the task and passed its six outcome matchers. It failed native runtime invariants because the capture contained exactly 1,000 events. The last captured event preceded the run's completion by more than a minute. The API caps each response at 1,000 rows. The harness did not request the next page. **Expected behavior** Read the complete durable event stream through the public API before checking semantic-result and terminal-event counts. Reject incomplete or malformed evidence. **Steps to reproduce** 1. Complete a native task that emits more than 1,000 durable events. 2. Place the semantic-result and terminal events after row 1,000. 3. Capture the run with the Product E2E harness. 4. Before this fix, the invariant checker sees only the first page. **Paperclip version or commit** Observed at `4196a4cd76db434854b679035e4146c7f69689ce`. The same single-page capture exists on master. The original failed result remains unchanged; missing historical tail evidence is not reconstructed or graded as a pass. ## What Changed - Add a bounded event collector that advances through the public `afterSeq` cursor. - Use it for task success/failure evidence and shared chat run evidence. - Reject invalid pages, missing or non-increasing sequence numbers, repeated cursors, failed later requests, and an exhausted page limit. - Test completion events beyond the first page, exact page boundaries, and malformed evidence. - Correct the existing Everyday catalog test from 38 to the maintained 47 cells. The suite stays explicit-only. - Document the complete-capture requirement and its bound. ## Verification - Product harness typecheck passes. - The full credential-free harness suite passed 477 tests. After adding the chat integration regression, all 34 chat evidence tests pass. - Seventeen pagination tests cover the valid tail and malformed-evidence cases. - `git diff --check` passes. - [Full repository CI](https://github.com/paperclipai/paperclip/actions/runs/36080683422) passes at `aed6f79c089226be79e75dcf390969561ec4f787`: 52 successful checks and two intentional skips, including Rust, typecheck, build, server tests, and browser shards. Greptile gives this exact head 5/5 with no findings. No local Docker, Rust build, browser suite, or model invocation was used for this change. - [Live Grok requalification](https://github.com/paperclipai/paperclip/actions/runs/36080870743) is running on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, with the runner, scheduler, and evidence fixes. Its full credential-free harness passes all 493 tests and typecheck. Live results are pending; the original campaign remains a failed measurement. ## Risks Long runs need more read-only API requests and larger private evidence files. Capture stops with an explicit error after 100 full pages. This changes neither production APIs nor provider behavior. It does not relax an invariant or change a historical grade. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c96c751f39 |
fix(heartbeat): claim task ownership with queued runs (#13973)
## Thinking Path > - Paperclip manages agents and their tasks. > - The scheduler can start several runs for one agent. > - Each task still needs one execution owner. > - Task creation and recovery can queue assignments for the same task. > - The scheduler previously marked both runs running before starting either executor. > - This change claims the run and task ownership in one transaction. > - Tasks with different IDs retain concurrent execution. ## Linked Issues or Issue Description Refs #13882. Related queue work: #11830 and #13965; those address different admission and cancellation paths. **What happened?** The Grok subscription Product campaign exposed two assignment runs for one task. Task creation and periodic recovery started runs within 200 ms. Both acquired Daytona sandboxes. One failed before provider work with `paperclip_runner_attachment_staging_not_authorized` because the other owned the task. Fixture cleanup then found an active session. The original failed result is retained in [campaign 36071063537](https://github.com/paperclipai/paperclip/actions/runs/36071063537). **Expected behavior** Only the task execution owner may start provider setup. Competing queued work must wait. Different tasks may use the agent's available slots. **Steps to reproduce** 1. Set an agent's concurrent run limit to two. 2. Queue task-creation and recovery assignment runs for the same task. 3. Resume the queue while holding adapter execution open. 4. Before this fix, both runs become running. Only one has the task execution lock. **Paperclip version or commit** Reproduced on master `8781f06a8` and in the Grok campaign at `4196a4cd`. ## What Changed - Use the same company-scoped task ownership gate for assignments, direct comments, and queued comments. - Leave competing work queued while another live run owns the task. Transfer a terminal pointer only after the tracked executor and durable environment leases/finalization settle, including cleanup on another controller. - Commit the running state and task execution owner together. - Preserve the dedicated review path and the native replacement checkout guard. - Add sixteen database regressions for assignment/comment orderings, unrelated concurrent tasks, and cross-controller cleanup fences. ## Verification - Live subscription Product E2E planning passed on its first attempt in both Daytona (351,197 ms) and local execution (268,883 ms), with 6/6 matchers and cleanup passing in each, on combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e` ([campaign](https://github.com/paperclipai/paperclip/actions/runs/36080870743)). The local screenshots and persisted state show the same plan revised, the first approval rejected, the revised approval accepted, and one completion. The full campaign remains unqualified because of a separate startup retry and EC2 Spot interruption. - The new same-task regression failed before the fix: two running rows instead of one. - All seven focused concurrency cases passed. Three comment cases failed before the shared gate was added. Five cross-controller cleanup cases failed before the durable gate was added. A warm-retention regression also failed before its successful release receipt was admitted; missing/failed receipts and a different retention policy stay blocked. - `node ../node_modules/vitest/vitest.mjs run src/__tests__/heartbeat-stale-queue-invalidation.test.ts` from `server`: 48 passed. Native cleanup admission adds one passing test; task-drain release adds two. Total focused checks: 51 passed. - `git diff --check`: passed. - Current-head [CI](https://github.com/paperclipai/paperclip/actions/runs/36078495577) passes repository typecheck, tests, Rust checks, build, and browser suites. The unrelated repository-form browser shard passed one bounded rerun without assertion or code changes; its original navigation failure remains retained. No local Docker or Rust build was used. - The failed test's exact two sandboxes were checked: one was already absent; the other was deleted after its company, environment, and run labels matched. ## Risks The claim transaction adds a task row lock. It follows task-before-run lock order. A competing run stays queued until the current owner releases the task. The change does not alter attachment authorization, provider permissions, schemas, or public APIs. Review is 5/5 with all threads resolved on `23c0d13b6`. All CI checks pass on that exact head. Live Grok requalification will use the combined runner and scheduler source. ## Model Used OpenAI GPT-6 through Codex, with repository tools and code execution. The exact serving identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge ### Live regression verification The Daytona plan/revise/accept lifecycle in [Grok campaign 36080870743](https://github.com/paperclipai/paperclip/actions/runs/36080870743) passed on its first attempt at combined source `1b0551bb7c8de3c54f4bee64dbe2c88328b3645e`, in 351,197 ms, with all six outcome matchers and cleanup passing. The earlier competing-run failure remains retained. This is one live planning result; the complete Grok matrix and repeated qualification are separate gates, and this campaign also has separately retained infrastructure failures. --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.925.0-canary.1 |
||
|
|
efce9356b5 |
fix(ui): offer recovery when the app fails before React starts (#13970)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The browser must load its JavaScript before React can render a task. > - A failed import can stop that process before the React error boundary exists. > - The HTML entry then leaves an empty page with no recovery action. > - This PR adds a small recovery screen that works without React. > - The user can retry the same page and return to saved task content. ## Linked Issues or Issue Description Refs #13824 and #13895. This is a follow-up to their browser startup investigation. **What happened?** Interrupting the app bundle or a required import leaves an empty React root. A startup exception has the same effect. A React error boundary cannot handle these failures because React has not started. **Expected behavior** The page must explain the startup failure and offer a manual retry. A late successful load must dismiss the recovery message without a reload. **Steps to reproduce** 1. Open a saved task in the browser. 2. Abort the application bundle request, or make a required module return HTTP 503. 3. Observe the empty page before this change. With this change, use Reload page after the fault clears and verify the saved task and comment. **Paperclip version or commit** The failing regression baseline used master at `8781f06a8`. **Deployment mode** Local source build and compiled UI. Tests cover both initial navigation and a page controlled by the production service worker. The exact cause of the older intermittent Vite stall remains unconfirmed. Forty app loads and thirty replays of retained responses did not reproduce it. This PR fixes the missing recovery path; it does not claim to remove that historical cause. A normal HTTP 304 response is not a failure. ## What Changed - Add an inline startup guard and recovery screen in the HTML entry. It does not depend on the app module graph. - Show a manual reload action after a startup error or after 30 seconds without rendered root content. - Remove the notice, timer, observer, and error listeners when the app starts. Never reload automatically. - Keep the recovery screen outside the React root so it cannot satisfy app-readiness checks. - Add browser tests for interrupted imports, a stalled import, an evaluation error, service-worker-controlled retry, repeated offline retry, and cleanup after successful startup. - Return a static, uncached HTML retry screen when a service-worker-controlled navigation fails offline. It contains no task content. - Add a full-app test that retries an interrupted compiled bundle and checks the saved task, comment, composer, route, and absence of agent runs. - Document the coverage and the limits of the historical diagnosis. ## Verification - Red baseline: four recovery cases failed; the normal-startup case passed. After the change, all five recovery cases passed. The review found an offline retry gap; that additional case failed before the worker fix and passed afterward. - Full provider-free browser-support suite: 16 passed. - Compiled-app browser tests: four passed, including saved-task reload, interrupted-bundle recovery, slow-CPU service-worker reload, and sidebar navigation. - Expanded service-worker, offline response, PWA, and worker build-ID unit tests: 37 passed. The two old plain-text offline expectations were reproduced as failures and updated for the HTML retry contract. - UI production build, full local repository typecheck (`pnpm -r typecheck`), runner-E2E typecheck, and design token checks passed. - Manual browser check: a temporary server failed the compiled bundle once. The recovery screen appeared. Clicking Reload page restored the same saved task, comment, and composer. - Full local `pnpm build` passed. - Full local `pnpm test:run` was attempted with a bounded deadline and stopped after it timed out. Workspace runtime/cleanup tests reported timeouts on this host. The monolithic local run is not a pass. The focused tests above and the complete Linux CI run provide the successful verification. - Final-head [CI run](https://github.com/paperclipai/paperclip/actions/runs/36072201966) passed. All 53 check runs succeeded; the two Storybook jobs were intentionally skipped. The legacy security status also passed. - Greptile reviewed `f83e0f51fb760541d83353f2c1df4e182f3948f9`: 5/5. Both review findings are fixed and resolved. ## Risks - The guard only handles startup before React renders root content. Existing React boundaries handle later rendering errors. - A slow startup can show the message after 30 seconds. A later successful render removes it; the page does not reload by itself. - The fallback uses native HTML when the app stylesheet is unavailable. - The worker changes only its offline navigation response. It returns static HTML with a reload button and `Cache-Control: no-store`. Its cache allowlist, private-response protections, task state, provider prompts, and grading rules stay unchanged. - This does not establish or fix the unknown cause of the historical intermittent Vite stall. ## Model Used OpenAI GPT-6 through Codex. The session exposes the GPT-6 family but not an exact served model ID or context window size. Used reasoning, code editing, shell tools, and browser testing. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>nightly/v2026.925.0-nightly.0 canary/v2026.925.0-canary.0 |
||
|
|
8781f06a87 |
feat(connections): enable MCP aggregators by default (#13964)
## Thinking Path > - Paperclip helps people manage AI agents for work. > - Connections let those agents use external services with explicit access rules. > - Zapier, Arcade, Composio Connect, and Executor already have setup and runtime support. > - Their experimental switch still blocks discovery and setup by default. > - This pull request removes those gates and the Settings toggle. > - Users can connect these providers without enabling an experiment. ## Linked Issues or Issue Description Refs #13755. Refs #13941. **What existing behavior does this improve?** Apps browsing, inline setup, and agent connection search for the four MCP aggregators. **Current behavior** An instance must enable the MCP aggregators experiment before users or agents can start setup. **Proposed behavior** All four providers are available by default on local and managed instances. Old stored and managed values still parse but cannot disable them. ## What Changed - Remove the aggregator gates from Apps, inline setup, server setup, and agent search. - Remove the Settings toggle and its UI hook. - Retain the old setting key only for upgrade compatibility. Normalize it to true and ignore managed overrides, as Apps already does. - Replace opt-in fixtures with default-on coverage. Test old false values, all four setup flows, provider choice, and the removed toggle. - Update current connector guidance and remove the opt-in from the runner acceptance fixture. ## Verification - 306 focused tests passed across eight files: shared remote MCP contracts; server remote MCP lifecycle, aggregator fallback, settings normalization, and managed overlay; UI Apps browsing, setup, and experimental settings. - Server and UI TypeScript checks passed. - UI token gates and `git diff --check` passed. - The full local suite was not run, per the maintainer's instruction. All 54 CI checks passed; two checks were skipped. One unrelated workspace-preview readiness timeout passed on one failed-shard retry. - The setup fixtures use simulated MCP responses. This change does not claim new live provider acceptance. ## Risks - Existing instances now show all four providers, even if the old flag was false. This is intentional. - External provider choice, credentials, company isolation, agent grants, and tool policies still apply. Showing a connector does not authorize an external account. - No data migration is required. The compatibility key keeps old managed configuration documents valid. - Historical Zapier live acceptance remains incomplete in the existing evidence report. The maintainer explicitly requested the default-on rollout for all four existing providers; the report records that scoped exception. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, repository tools, and test execution. The context window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.5 |
||
|
|
aa8fc86331 |
feat(connections): prefer native apps and ask users to choose external providers (#13941)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external services. > - Native connections should remain the first choice for a supported app. > - Other apps may be available through an external MCP provider. > - The user must know which external provider handles the connection and choose it before setup. > - This pull request adds ranked alternatives and server-authored instructions to connection search. > - Agents can follow the returned instructions while Paperclip validates saved choices and access. ## Linked Issues or Issue Description Related: #13879, which fixed inline MCP provider setup. This PR adds discovery and provider selection on top of that work. **Subsystem affected** Cross-cutting: shared connection contracts, server search and intent services, native runtime, CLI, inline setup UI, and evals. **Problem or motivation** An agent cannot offer a clear external-provider choice when Paperclip has no native connection for an app. Adding provider-specific branches to the core prompt would make those instructions harder to maintain. **Proposed solution** Prefer a native connection. Otherwise return verified alternatives in Composio, Arcade, Executor, Zapier order. Include an external-service disclosure, a question with None, and the next instruction in the search result. Validate the saved human choice before creating a selected fallback setup card. Reuse existing provider accounts and verify underlying app access separately. **Alternatives considered** Do not silently choose a provider. Do not claim that broad execution tools prove support for every app. Reuse existing questions and connection intents rather than add another connection model. **Roadmap alignment** Extends the existing MCP Tool Gateway & Apps and Agent evals & feedback capabilities. The MCP aggregators experiment remains the gate. No duplicate provider-routing PR was found in the public search. ## What Changed - Add a dated support index and authorized cached-tool evidence for external routes. - Return provider questions and next-step instructions from `connections_search`. - Preserve pending choices and declines across continuation. Validate company, task, agent, human, app, and current route eligibility. - Carry the selected app into new setup and account reuse, validate explicit provider requests against persisted human messages, and distinguish provider readiness from app authorization. - Sync native, MCP, REST, and CLI contracts. Keep core agent instructions provider-neutral. - Add production-component Storybooks, focused database tests, and three real-agent browser eval cases. - Record the plan, observed failures, fixes, passing evidence, and acceptance limits. ## Verification - Latest head `586f0e6cd`: 54 checks passed, 2 skipped; Greptile 5/5 and all review threads resolved. - After rebasing on master `18dac1e1e`: 64 focused shared, validator, route-contract, and database tests passed; server typecheck passed. - Embedded-browser test drive on the rebased head: native Jira card, HubSpot external-provider question, Arcade account reuse, one actual MCP read against a local synthetic fixture, reload persistence, and None preventing further calls. A real OpenAI-backed agent performed discovery and continuation. - UX observation: the agent initially combined mutually exclusive request fields; the server rejected it and the agent recovered without changing access. This extra retry remains visible in the transcript. - After rebase: 23 focused eval grader/catalog tests, affected TypeScript checks, token gates, production UI build, and Storybook build passed. Full local tests are intentionally excluded at the maintainer's request. - Before rebase: four browser/real-agent attempts passed: native Jira, None, and reuse of the second provider on two Codex profiles. - Browser evals used an isolated deterministic MCP fixture through the real Paperclip gateway. They do not prove production compatibility with all four providers. - Review `Apps / Connections / Provider choice` in Storybook. Choose Arcade, continue through Access, and verify the app name, external-service disclosure, and URL configuration. - The detailed verification report is `doc/connections/2026-09-23-aggregator-routing-verification.md`. ## Risks - The public support index is finite and can age. Account capability and app authorization still require verification after selection. - Existing installed-tool permissions remain in effect. Provider choice is not a new execution permission boundary. An early Mini attempt skipped search; clearer provider-neutral instructions made the targeted rerun pass. This is not a measured reliability rate. - Explicit requests skip provider confirmation only when a clear persisted human message or saved provider choice supports them. Other phrasing falls back to confirmation; the agent query alone is not consent. These routes do not add tool permissions. - No database migration or legacy Composio broker is added. Real-provider acceptance remains separate from fixture proof. ## Model Used OpenAI GPT-6 through Codex, with reasoning, repository tools, code execution, and browser tools. The exact deployment variant and context-window size are not exposed in this session. Product evals separately used the repository's primary Codex and Codex Mini profiles; those agents supplied test behavior, not independent provider compatibility proof. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.4 |
||
|
|
18dac1e1ef |
feat(connections): add experimental memory providers and remote MCP access (#13942)
## Thinking Path > - Paperclip manages AI agents and their work. > - Connections give agents governed access to external tools. > - Agents need durable memory across tasks and execution environments. > - Mem0, Zep, Supermemory, Cognee, and Honcho provide hosted memory tools. > - This pull request adds their setup flows behind an experimental toggle. > - It also delivers assigned MCP tools through the native remote Codex runner. > - Operators can connect a provider once and use the same governed tools locally or in Daytona. ## Linked Issues or Issue Description **Problem or motivation** The Apps catalog lacks a complete set of memory providers. Remote native Codex agents also need access to assigned managed MCP tools without receiving provider credentials. **Proposed solution** Add five memory connectors behind the disabled-by-default Experimental memory connectors setting. Use the existing connection setup and permissions UI. Default their tools to Allowed. Preserve company boundaries, operator permission changes, provider scopes, and audit attribution. **Alternatives considered** Direct provider credentials in each sandbox would duplicate setup and bypass the managed gateway. The remote runner instead uses its existing protocol channel to call the gateway on the server. **Roadmap alignment** This maintainer-requested experiment supports the Memory / Knowledge and Connected Apps roadmap areas. It adds provider connections without introducing a separate memory UI. Related connector authoring documentation is tracked in #13692; no duplicate memory-provider implementation was found. ## What Changed - Add provider definitions, official branding, and the experimental setting for all five providers. - Use OAuth for Zep and Supermemory, API credentials for Mem0 and Honcho, and a bundled Cloud API bridge for Cognee with no runtime downloads or subprocesses. - Default memory tools to Allowed and classify destructive actions explicitly. - Fix personal remote credential resolution and propagate provider tool errors. - Relay assigned managed MCP tools to remote native Codex through the runner protocol. Recheck current authority for each call and rotate stale tool contracts. - Add 21 Storybook states and complete OAuth walkthrough fixtures. - Document provider research, sanitized tool inventory, and live local and Daytona proof. ## Verification - Passed workspace typecheck: `pnpm -r typecheck`. - Passed production build: `pnpm build`. - Full local `pnpm test:run`: 13,273 passed, with failures from process/readiness timeouts under parallel load. Reran all 16 affected suites with one worker: 456 passed, leaving two macOS `/var` versus `/private/var` path assertions. Both passed with `TMPDIR=/private/tmp`. No test failures remain unverified. Latest-head remote CI passes all 54 checks (two optional Storybook jobs skipped). Greptile is 5/5 with no unresolved findings. - Latest Cognee gateway regression: 71 passed, including public deployment without a runtime host and immediate recovery after a provider error. Bundled bridge tests: 25 passed. - Browser setup and real agent tasks exercised all five providers. Mem0, Cognee, Zep, and Honcho have successful store/retrieve proof. - Supermemory now has scoped read/write consent. Local storage and real Daytona write, document read, and semantic recall passed; indexing completion was verified before claiming success. - All five providers were exercised through a real Daytona sandbox and its native runner MCP relay. Provider credentials remained on the server. The final bundled Cognee bridge also passed a fresh Daytona store/recall run. All disposable sandboxes were removed and verified absent after testing. - Storybook is rebuilt and contains the experimental toggle, catalog, setup, permissions, and error states. The Zep and Supermemory access steps advance correctly. ## Risks - Provider OAuth scopes and plan limits remain independent of Paperclip tool permissions. An Allowed tool can still be rejected by the provider. - Providers can queue memory indexing; save acceptance does not prove that semantic recall is ready. - Remote tool contracts must stay synchronized with current connection authority. Regression tests cover revocation and stale contracts. - This change has no database migration. Existing connections remain usable when the experimental catalog toggle is disabled. ## Model Used OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, repository edits, code execution, and browser automation. The exact context window size is not exposed in this session. Live acceptance agents used `gpt-5.6-sol`. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
e006c18f20 |
test(paperclip-runner): stabilize the assigned-skills ACPX runtime host test (#13453)
Release sealed skill homes before test cleanup and reuse one prepared sandbox across the assigned-skills test opens. Preserve the current skill-refresh and prompt assertions without increasing test timeouts. Co-Authored-By: Devin Foley <devinfoley@users.noreply.github.com> Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.3 |
||
|
|
2914501c08 |
test(chat): retire fixture leases before recovery sweeps (#13952)
Retire verifying chat fixtures and their synthetic endpoint leases after service shutdown. Add a regression that recovers a stale receipt while preserving another company's lease. Co-Authored-By: Paperclip <noreply@paperclip.ing> |
||
|
|
f3a214fe77 |
Fix relative symlinks in secondary sandbox repositories (#13953)
## Thinking Path > - Paperclip manages agents and their task workspaces. > - A task can use several independent Git repositories. > - Sandbox staging copies secondary repositories from temporary clones. > - The copy changed relative symlinks into absolute host paths. > - Those links broke skill discovery and made workspace restore fail. > - This change preserves link targets while keeping extraction checks intact. ## Linked Issues or Issue Description **What happened?** Staging a secondary repository rewrites a link such as `.claude/skills/demo -> ../../skills/demo` to an absolute path in a temporary Git clone. That clone is then removed. The link is broken in the sandbox, and Daytona refuses the outbound archive during workspace restore. An agent can finish its turn but still have its run fail during restore. **Expected behavior** Repository links keep their original targets after staging. Links within a repository remain usable, and changes return to the local checkout. Unsafe outbound archive links still fail before extraction. **Steps to reproduce** 1. Create a project with a primary repository and a secondary repository. 2. Commit a relative skill directory link in the secondary repository. 3. Stage the workspace for sandbox execution and inspect the copied link. 4. Restore that repository through Daytona. Before this fix, the link points at a removed host temporary directory and restore rejects it. **Paperclip version/commit** Reproduced on `0f8750627f11d855552abce9a837d7f3b67c9ddf` in the multi-repository sandbox path. Related: #13442 introduced multi-repository provisioning. #13882 adds native Grok but keeps legacy adapters; #12991 addresses Grok instruction isolation and leaves skill staging unchanged. Searches of open/closed PRs and open issues found no direct fix for this copy behavior. ## What Changed - Set `verbatimSymlinks: true` when copying secondary Git clones. This [Node option](https://nodejs.org/api/fs.html#fspromisescpsrc-dest-options) preserves the stored link target instead of resolving it against the temporary source. - Test directory, file, chained and dangling links after temporary-clone cleanup. Assert the copied Git checkout remains clean. - Cover skill-link reads, edits and restore in fresh, warm-adoption and durable-seed workspace modes. - Extend Daytona checks for valid relative directory links and rejected absolute targets. - Document the staging behavior. ## Verification - Four regression cases fail without the source fix: one clone test and three staging modes. - `pnpm exec vitest run packages/adapter-utils/src/git-workspace-sync.test.ts packages/adapter-utils/src/sandbox-managed-runtime.test.ts packages/plugins/sandbox-providers/daytona/src/plugin.test.ts`: 301 passed. - `pnpm -r typecheck` and `pnpm build`: passed. - `pnpm test:run` was started locally, then stopped after the full Linux CI suite passed. It did not finish locally; this is not a full local-suite pass. - Full PR CI: 53 checks passed, two conditional skips, on `07b30ff298baab327a4e60708ade899935688a3f`. Three server shards were interrupted by runner shutdowns; the unchanged mobile repository test timed out waiting for a disabled Save changes button. One same-commit failed-job rerun passed. Original attempts remain in [run 36036538369](https://github.com/paperclipai/paperclip/actions/runs/36036538369). - Greptile: 5/5 on the same head, no review threads or actionable findings. - Diff scanned for secrets and private identifiers; no matches. ## Risks Low risk: the production change is one copy option. It preserves symlinks instead of following or materializing their targets. Daytona extraction guards, workspace exclusions, authentication and database behavior do not change. This prevents corruption in newly staged snapshots. It does not rewrite an already corrupted warm workspace or durable seed; those need fresh staging from the source checkout. No live provider run or customer-task replay was performed. The tests use real Git, filesystem and tar operations with mocked provider transport. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.2 |
||
|
|
0f8750627f |
fix(codex): preserve typed ACP quota classification and reset time (#13945)
## Thinking Path > - Paperclip runs agents and recovers failed tasks. > - Codex ACP can report usage exhaustion as a typed terminal failure. > - The shared engine removes provider text before it stores the result. > - Codex did not use the existing terminal classifier hook, so a quota failure became a generic turn failure. > - This change classifies explicit usage exhaustion before that text is removed. > - Recovery can then use quota backoff and the supported reset clock. ## Linked Issues or Issue Description **What happened?** A Codex ACP `limit` failure with explicit usage-exhaustion text produced `acpx_turn_failed`. Recovery could schedule the ordinary short retries because the quota classification and reset time were lost. **What did you expect to happen?** Keep the failure visible and use the existing provider-quota wait. Preserve a supported reset timestamp without storing provider text. **Steps to reproduce** Return a terminal ACP session failure with category `limit` and title `You've hit your usage limit for GPT-5. Switch to another model now, or try again at 4:30 PM (America/Chicago).` The real child-process regression tests exercise both pinned ACPX versions in persistent and oneshot modes. **Paperclip version** Base commit: `32573876d4`. Related work: #13651 and #13831 provide the shared hook and Claude classification. #11854 includes quota handling as part of optional credential rotation, but reads the already-sanitized result; this patch handles the typed terminal boundary without adding rotation. #13549 reads recovery text after this boundary and cannot recover discarded provider text. #9011 concerns the Codex CLI backoff. This change leaves those other mechanisms in place. ## What Changed - Register a Codex terminal-failure classifier with the existing ACP engine hook. - Recognize explicit usage exhaustion only in a typed `limit` failure. - Reuse the Codex reset-time parser and existing provider-quota recovery fields. - Leave context, turn, rate, budget, storage-capacity, and unknown failures on their existing paths. - Test real ACP children, both dependency patches, privacy, and recovery classification. Document the boundary. ## Verification - 88 focused tests passed across Codex ACP, parsing, and server recovery classification. - An initial cross-adapter run passed 53 tests, including the existing Claude quota suite. - Removing only the classifier registration makes five new integration tests fail; restoring it passes all 19 new tests. - `pnpm -r typecheck` and `pnpm build` passed. - Full Linux CI passed on `b10d60e002`, including all unit/integration shards, browser shards, typecheck/build, runner checks, and release/package gates. - The duplicate local `pnpm test:run` reported three skill-cache failures and two runner-suite failures. It is not claimed as a full local pass. - The three skill-cache failures reproduce in a clean worktree at the unchanged base commit (3 failed, 107 passed, 3 skipped across the skills and runner files). The cache rename reports `EACCES` on macOS. - An isolated runner-suite run reports two embedded PostgreSQL startup failures before its assertions (37 tests pass). No runner source is changed; the Linux CI runner and server suites passed. - Greptile reviewed the current head at 5/5 with no actionable findings or unresolved threads. ## Risks - Explicit usage-exhaustion failures now wait for quota recovery instead of short generic retries. - Unknown wording retains the existing behavior. The generic historical terminal-limit message alone cannot establish quota exhaustion. - Reset parsing keeps the existing Codex clock formats. A missing or unsupported reset uses the existing quota backoff. - No schema, credential, UI, dependency, or deployment changes. Provider text stays in memory and is absent from results and logs. ## Model Used OpenAI GPT-6 via Codex. Exact runtime model variant and context-window size were not exposed. Used reasoning, repository inspection, editing, and local tests. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.924.0-canary.1 |
||
|
|
32573876d4 |
chore(lockfile): refresh pnpm-lock.yaml (#13840)
Auto-generated lockfile refresh after dependencies changed on master. This PR only updates pnpm-lock.yaml. Co-authored-by: lockfile-bot <lockfile-bot@users.noreply.github.com> |
||
|
|
b0155a681a |
feat(slack): connect Paperclip conversations and scheduled messages (#13920)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Slack conversations use the same tasks and agents as the Paperclip board. > - A board reply must reach that Slack conversation and let the agent continue the work. > - An assigned agent also needs its Slack tools during normal tasks and scheduled routines. > - Both paths must keep the linked user's authority, delivery rules, and conversation history. > - This pull request adds those paths and reduces setup friction for Slack bots. ## Linked Issues or Issue Description **Subsystem affected** Server orchestration, Slack connector tools, shared contracts, and chat setup UI. **Problem or motivation** Replies entered in Paperclip did not provide a complete round trip to the linked Slack thread. Slack and board wakeups could select different model sessions for the same task. Agents also lacked their assigned Slack tools outside Slack-origin work, which prevented a routine from sending its responsible user a briefing. Inviting a bot could leave the new channel disabled. **Proposed solution** Mirror human board messages with author attribution and route the agent result to the same thread. Use the same session key across both entry points. Supply Slack tools to the connection's assigned agent in normal tasks and routines, using the current responsible user's verified link. Enable newly invited channels while preserving explicit disabled choices. Add a browser-agent setup prompt to the Slack wizard. **Alternatives considered** A separate Slack scheduler or task dispatcher would duplicate existing Paperclip workflows. Reusing the connection owner's identity would grant the wrong authority. Replaying old channel history could start unintended work. This change uses ordinary task wakeups, routine dispatch, and Slack's original invitation mention event instead. **Roadmap alignment** Extends the shipped Scheduled Routines and governed Apps capabilities. It does not add a separate task lifecycle. Related work: #13828 and #13809. Related test stabilization: #13877. The existing plugin Slack-control proposals are separate from this built-in connector change. ## What Changed - Queue human Paperclip messages for the original Slack thread with display-name attribution and stable delivery identities. Require the author’s current linked Slack identity and recheck access before delivering messages or agent replies. - Apply pause, dependency, cancellation, and closed-workspace guards before explicit Board sends request work and again when the durable outbox dispatches it. - Route agent results back to Slack and preserve model-session continuity, including replies that reopen completed tasks. - Resolve assigned Slack connections for normal agent tasks and routines. Recheck the responsible user's link, membership, and permissions at execution. - Add `slack_open_dm` for the responsible user's bot DM and request the `im:write` scope. - Enable newly discovered invited channels. Keep explicit OFF choices and normal admission and deduplication rules. - Add a copyable Slack setup prompt for a computer-use agent, with Storybook coverage. Share the prompt-button component with GitHub. - Update Slack tool documentation and runtime instructions. - Stabilize the mobile project browser test by waiting for the final canonical route before editing, preserving all persistence assertions. ## Verification - Live staging: invited the bot after the first mention. The channel became enabled and the bot answered that original mention. - Live staging: a normal Paperclip reply appeared in Slack with author attribution. The agent completed the calculation and replied once in the original thread and in Paperclip. - Live staging: a codeword entered in Slack was recalled from Paperclip. A following Slack calculation used the result from the Paperclip turn. Run metadata confirmed the same model session for both entry points. - Live staging: a scheduled routine used `slack_open_dm` and `slack_post_message` to deliver one DM. The existing app was reinstalled with `im:write`. The test routine was paused after verification. - Before the master merge: 397 focused feature tests passed. The continuity fix passed all 76 issue comment/update route tests and six focused route/integration cases. Typecheck, build, and token gates passed. - Review fixes: 47 focused integration cases passed, covering link revocation/replacement, private membership removal, guarded outbox dispatch, concurrent workers, lost scheduler responses, a real one-connection pool, exact reply provenance, and attachment retries. All 99 issue-comment route tests and the Slack catalog browser test passed. - Full local typecheck, production build, token gates, and module-boundary checks passed. The full local test command passed 25,915 tests before a 15-second timeout in `issue-thread-interaction-routes.test.ts`; that entire suite passed on isolated rerun (81 tests). Remaining serialized coverage is provided by the current-head CI shards. - An unchanged Cursor adapter test hit its 10-second limit in CI; all five tests in that file passed on a local rerun in 3.11 seconds, and the failed CI shard passed on its single retry. - The preview-server readiness test passed a local rerun (28 tests). The mobile-project readiness fix passed three repetitions of both browser tests (6/6). - Final commit `64ac0d9897f4353375996f1b1b38e5040bdeb0a0`: all CI gates passed, including all eight browser shards, all server/chat suites, typecheck, build, runner checks, and security checks. Greptile reviewed this exact commit at 5/5; all review threads are resolved. ## Risks - Human messages on a Slack-linked task now publish to its Slack thread. The task banner states this behavior. Incoming Slack messages and internal agent bookkeeping must not echo back. - Normal tasks and routines can now use the assigned bot. Authority remains bound to the current responsible user's link; it does not fall back to the connection owner. Revocation, private-context limits, and queued-write checks still apply. - Existing Slack apps need `im:write` and a reinstall to open DMs. Other existing capabilities remain available without that scope. - New invited channels default to enabled. Explicit disabled choices remain disabled. Channels created by bot tools still require a person to enable responses. - No database migration or new provider credentials are required. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code execution, GitHub CLI, and browser tools. The runtime does not expose a more specific model build identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
c341588bdb |
fix(ui): add Cloud invitations to the Members page (#13922)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - People manage collaborators from the Members page. > - Cloud manages invitations outside the tenant's local invitation system. > - The local Invites tab is hidden on Cloud, so this page has no way to invite a person. > - This pull request adds an Invite people action for the current Cloud stack's owner or admin. > - The action opens the existing Cloud People settings for that stack. ## Linked Issues or Issue Description **What happened?** A Cloud owner opens Organization Settings → Members and finds no invitation action. The tenant-local Invites tab is hidden, and the page does not link to Cloud's invitation flow. **Expected behavior** Cloud owners and admins can start an invitation from Members. **Steps to reproduce** 1. Sign in to a Cloud-managed instance as the current stack's owner or admin. 2. Open Organization Settings → Members with `company.invites` hidden. 3. Look for an invitation action beside the page heading. **Paperclip version or commit** `7b7c4d4172d6aac14919e2682b702ae87bc17653`. **Deployment mode** Cloud-managed, authenticated. Related search: #2388 proposes broader member-management UI. This change only connects the existing Members page to Cloud invitations. No duplicate Cloud invitation action PR was found. ## What Changed - Add **Invite people** beside the Members heading for the current Cloud stack's owner/admin. - Read the role from the authenticated Cloud portfolio. Ownership of another stack does not enable the action. - Navigate to the current stack's People settings on the configured Cloud origin. Keep the local Invites tab hidden when configured. - Cover allowed roles, denied roles, loading, failed refresh, missing configuration, current-stack selection, and self-hosted behavior. Document the navigation contract. ## Verification - Focused Members and Cloud link tests: 23 passed. - Full UI suite: 6,669 passed across 634 files. - UI typecheck and `pnpm check:token-gates`: passed. - `pnpm build`: passed. - `pnpm -r typecheck`: passed. - All PR CI checks passed, including the full general, serialized, browser, and runner test jobs. The duplicate local repository-wide `pnpm test:run` was stopped after CI completed; it is not reported as a local pass. - Manual acceptance after tenant rollout: an owner/admin opens Members, selects **Invite people**, and reaches the same stack's Cloud People settings. A member does not see the action. ## Risks - The action needs a tenant app update before it appears on an existing stack. - The portfolio request must identify the current stack and its role. The action stays hidden when that information is unavailable or the request fails. - Cloud rechecks invitation authorization at the destination. No schema, API, or invitation-acceptance behavior changes. ## Model Used - OpenAI Codex, GPT-6, with reasoning, code execution, and repository tools. The session does not expose a more specific model identifier or context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
7b7c4d4172 |
fix(ui): recover an archived Cloud stack from its own health probe (#13913)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - On Paperclip Cloud, a stack whose organization is archived is served through the Cloud harness router. > - An archived stack answers every tenant request — including the SPA's own health probe — with a 423 archived status page. > - If the SPA is already loaded (a stale tab, or a load that raced the stack's archival), that 423 surfaces as a dead "Failed to load health" screen with no way out. > - This pull request recovers it the same way an expired tenant session already does: a top-level reload re-enters the Cloud harness, which redirects a fresh browser navigation on an archived host to the portfolio. > - The benefit is that an archived stack never traps the user on a broken health screen. ## Linked Issues or Issue Description No public issue exists. Description follows the bug template: **What happened?** A Cloud stack that becomes archived while its SPA is open (or is opened in a stale tab) shows a "Failed to load health" screen. The SPA's health probe receives the harness's 423 archived status page and has no recovery path, so the browser is stuck on a dead screen. **Expected behavior** The browser should leave the archived stack for the Cloud portfolio, as it already does for an expired tenant session — no dead "Failed to load health" screen. **Steps to reproduce** 1. On a Cloud-managed instance, open an organization's SPA in the browser. 2. Archive that organization (or open a stale tab for one that was archived). 3. The SPA's health probe returns 423 (archived) and the page shows "Failed to load health" with no way out. **Paperclip version or commit** `master` (Paperclip Cloud managed instances). **Deployment mode** Cloud-managed (`authenticated`). Self-hosted instances are unaffected: the 423 archived status page only originates from the Cloud harness router that fronts managed stacks; a self-hosted server serves its own health. **Subsystem affected** UI: tenant document recovery (`ui/src/lib/tenant-session-recovery.ts`), which the health/client/heartbeats/audit API surfaces already route error responses through. ## What Changed - `isArchivedStackRecoveryError(status, body)`: true for a 423 whose body carries a `statusPage.code === "archived"`. - `isTenantDocumentRecoveryError` unifies the existing 401 tenant-session codes with the new archived case; the recovery coordinator now fires on either. - The recovery action is unchanged — a single top-level reload — so an archived stack re-enters the harness and is redirected to the portfolio. The existing single-reload guard prevents any loop, and only the exact `archived` code triggers it (suspended/deleted/other 423s pass through as normal errors). ## Verification - `pnpm exec vitest run src/lib/tenant-session-recovery.test.ts src/api/auth.test.ts src/api/audit.test.ts` — pass. - New tests: the predicate accepts a 423 archived body and rejects suspended/deleted/error-shaped/503/null; the coordinator triggers the same single top-level reload for an archived body and stays inert for unrelated responses. - `pnpm exec tsc --noEmit` in `ui/` is clean. - Pairs with the harness-side redirect (paperclipai/paperclip-cloud#545): the harness redirects fresh navigations on an archived host to `/orgs`; this makes an already-loaded SPA trigger that navigation on its own health 423. ## Risks - Low. The recovery path and its single-reload guard are unchanged; only the set of conditions that trigger it grows by one exact code. Non-archived 423s and all other statuses behave as before. ## Model Used - Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergenightly/v2026.924.0-nightly.0 canary/v2026.924.0-canary.0 |
||
|
|
94a0aa7726 |
fix(ui): hide workspace isolation controls for managed hosts (#13907)
## Thinking Path > - Paperclip manages AI agents and their work. > - Workspace isolation keeps task checkouts separate. > - Managed hosts can enable isolation and hide its experimental toggles. > - Project, task, and routine forms still expose choices that override that policy. > - This pull request adds an operator visibility key for those controls. > - Workspace access stays available, and execution keeps its existing policy. ## Linked Issues or Issue Description **What existing behavior does this improve?** Operator control over workspace isolation settings in the UI. This follows the settings list cleanup in #13905. **Current behavior** Hiding the experimental isolation toggles leaves project policy editors, task selectors, routine and pipeline overrides, recovery actions, and workspace configuration visible. **Proposed behavior** Set `PAPERCLIP_HIDDEN_SETTINGS=workspaces.isolation` to hide these controls. Keep workspace navigation, status, files, and runtime access. Hide experimental toggles separately. Instances that do not set this key keep their controls. **Reason and benefit** Users on managed hosts should use the host's isolation default. A hidden form must not submit a stale draft that overrides it. ## What Changed - Add the UI-only `workspaces.isolation` key to the shared visibility registry. - Hide project workspace policy, task and subtask selectors, routine and pipeline overrides, and isolated re-issue actions. - Hide the workspace Configuration tab and redirect direct links to workspace issues. - Omit hidden new-task and routine overrides. Keep explicit task/subtask workspace launch context, saved policies, and automatic branch values for workspace routine runs. - Wait for the health visibility policy before showing controls. Keep workspace access and all execution APIs available. - Document the key and test visibility, form payloads, deep links, and unchanged workspace access. ## Verification - All 52 CI checks pass on `50cc770d33` (two expected skips). The branch is mergeable. Greptile is 5/5 with no unresolved comments. - `pnpm -r typecheck` passed. - `pnpm build` passed. - `pnpm check:token-gates` passed. - Targeted UI checks passed: 346 tests across 14 suites, including hidden project/task controls, stale task drafts, routine branch defaults, recovery actions, configuration deep links, and workspace access. - Shared settings-visibility tests passed: 11 tests. - `pnpm test:run` was run and stopped after reproducing five failures in unchanged server tests: two `chat-channels.integration` cases (linked-request provenance and direct external-chat finals) and three `company-skills-service` cases (runtime refresh, concurrent download, and explicit update). The earlier local run for #13905 showed the same failures. The full local suite is not claimed green. Targeted UI/shared checks pass. All PR CI shards, including the affected chat and skills suites, pass. - Reviewed the diff for secrets, private links, and run artifacts. ## Risks This key changes UI visibility only. It does not reject API calls or change feature values. Operators must enable isolation and its default through their existing policy mechanism. Older app versions ignore the new key until upgraded. Removing the key restores the controls. No schema changes. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either linked existing issues or described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run tests locally; targeted checks pass (full-suite limitation documented above) - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
8ee8f1fd6e |
ci: retire recurring public cloud image builds (#13827)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Core publishes standard images and source verification for downstream services. > - Managed services can now compose private images from the signed standard image. > - Core still builds a second public cloud image on every master push and release. > - That duplicate producer consumes build capacity and retains an obsolete readiness contract. > - This pull request retires recurring cloud publication while preserving the standard producer and rollback artifacts. ## Linked Issues or Issue Description Refs #13797 and #13789. Related: #12856 changes image dependency packaging; it does not retire this producer. **What existing behavior does this improve?** Core's recurring Docker publication and Cloud readiness workflow. **Current behavior** Master pushes call the legacy cloud publisher from Cloud readiness. Release tags and manual Docker runs call it too. Canary promotion also requires the legacy image. **Proposed behavior** Publish standard Core images and retain `Cloud source verified v1`. Let downstream services build their managed image. Keep explicit commit previews and existing images available. ## What Changed - Remove `docker-cloud.yml`, its master and release callers, and its unused cache selector. - Remove the legacy image/migrator wait and `Cloud deployable v1` job. Keep the full source verification workflow and exact source-proof name. - Make canary promotion inspect and promote the standard image only. - Preserve signed standard-image publication, direct migrator publication, and explicit `release.yml` previews. The preview path still uses the Dockerfile `cloud` target. - Update workflow, preview, build-stamp, and packaging tests. Exercise the promotion shell with mocked registry commands, including missing-image and missing-tag cases. - Document frozen legacy aliases, consumer requirements, preview compatibility, and rollback retention. ## Verification - All 377 workflow tests pass: `node --test .github/scripts/tests/*.test.mjs`. - All 129 release-registry tests pass: `pnpm test:release-registry`. - Focused source-proof, standard-image, preview, and workflow tests pass: 256 tests. - Focused image packaging/build-stamp tests pass: 16 tests. - Actionlint passes on all three changed workflow files. `git diff --check` passes. - Full local `pnpm build` and `pnpm -r typecheck` pass. - The policy follow-up updates an old assertion that required the removed readiness job. All 37 source-proof/release-workflow tests pass locally. - Full local `pnpm test:run` did not complete successfully while the Mac ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of generated Cargo output from this isolated worktree with `cargo clean`. GitHub CI passed on the final head: 52 successful checks and 2 optional skips. - Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`: **5/5**, successful current-head check, zero review threads. - September 23 refresh: the unchanged PR head merges cleanly with current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an isolated temporary worktree, all 377 workflow tests and 29 release/preview tests pass on the combined tree. `git diff --cached --check` passes. - Refreshed Actionlint workflow validation passes with ShellCheck disabled. Full Actionlint reports the same 10 existing ShellCheck diagnostics as master, with no added diagnostics. No source changes or new PR commits were needed. - The full local build/typecheck and current-head Linux CI results above remain the verification for the unchanged PR head. They were not rerun for this metadata-only refresh. No image publication or tenant deployment was initiated for this refresh. ## Risks **Deployment prerequisite satisfied (September 23):** The combined cleanup release is deployed to staging and production, and production Support is verified. Active managed-fleet automation uses standard-image composition. Explicit immutable previews remain supported by the retained preview publisher. This PR is ready for maintainer review; keep auto-merge disabled and wait for explicit merge authorization. - A consumer still selecting `Cloud deployable v1` will stop advancing at the last legacy-ready commit. Confirm active automatic consumers use the standard-image composition contract before merge. - Legacy cloud release-channel aliases stop advancing. Standard self-hosted aliases continue. - This PR deletes no registry images, cache tags, migrators, credentials, or runner infrastructure. Existing immutable releases remain usable for rollback. - Explicit legacy previews remain for commit-specific operator deployments. Retiring that compatibility path requires a separate consumer migration. - These changes affect CI publication, not database schema or application behavior. ## Model Used OpenAI Codex, GPT-6. The runtime does not expose a more specific model identifier or context-window size. Used repository inspection, reasoning, code editing, shell tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
db8f8fe5b7 |
fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path > - Paperclip manages agent work through shared runner contracts. > - Direct protocol evals qualify provider behavior against a mock control plane. > - Grok supports API keys and company subscription credentials. > - The hosted protocol workflow selected an API key for every Grok cell. > - Product subscription support did not enable subscription protocol runs. > - This change adds explicit subscription selection and checks the recorded authentication mode. ## Linked Issues or Issue Description Refs #13878, #13882, #12618. The direct Grok protocol roster cannot run with subscription authentication through the trusted default-branch workflow. Add an explicit selector while keeping API-key dispatches compatible. Keep the actor allowlist, protected environment, immutable source revisions, and publication gates. ## What Changed - Add `grok_authentication` with `api_key` and `subscription` choices. Keep `api_key` as the compatibility default. - Deliver the protected `GROK_AUTH_JSON` secret only to a subscription-selected Grok cell. Do not provide an API key to that cell. - Read authentication mode from the pinned eval program's actual roster summary. Retain it in the cell, catalog, campaign roster, and result. - Reject missing or mismatched authentication evidence during aggregation. Preserve cell metadata and an allowlisted failure reason before failing a cell, so malformed evidence cannot hide the retained attempt. - Document credential setup, source separation, and temporary-secret cleanup. ## Verification - `node --test packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed. - Validated all 39 Grok cells at eval revision `3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests only the subscription credential. Validation made zero provider calls. - Negative coverage rejects invalid selectors and missing or API authentication evidence in an otherwise passing subscription attempt. - `git diff --check` passed. All 53 current-head checks passed; the unchanged callback-drain timing test passed its bounded rerun, and the failed attempt is retained. Greptile reviewed `ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining findings. - No Docker or broad builds ran on the developer machine. CI performs repository checks on the configured fleet. ## Risks Grok runs require an eval revision that records `authenticationMode` in the roster summary. Missing evidence fails closed. The credential contains account access and refresh tokens; an owner must approve its delivery to the protected environment before a live run. The change adds no PR trigger or authorization bypass. Live subscription protocol qualification remains pending this workflow reaching master. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
794b09f834 |
fix(ui): alphabetize experimental settings and hide empty groups (#13905)
## Thinking Path Paperclip operators use Experimental settings to find and manage optional controls. New cards have accumulated outside alphabetical order, and operator-hidden cards leave empty developer and legacy sections. Sort the displayed controls within their existing groups and remove groups with no visible controls. ## Related Issues **What existing behavior does this improve?** The Instance Settings → Experimental page. **Current behavior** Agent Chat follows Cases, MCP aggregators precedes External Objects, and hidden developer/legacy controls leave empty headings. The worktree execution card also ignores its operator visibility key. **Proposed behavior** Cards appear alphabetically by their displayed title within each section. Empty developer and legacy sections disappear. The worktree execution card follows the same operator visibility policy as other experimental controls. **Reason and benefit** Operators can scan the list predictably. Hosted installations show only the controls their operator permits, without empty sections or a worktree-only exception. No matching sorting PR was found in the duplicate search. This is a small improvement to an existing settings page and does not add a roadmap feature. ## What Changed - Reorder existing cards without changing their values or mutation handlers. - Hide empty developer/legacy sections and respect the hidden worktree-execution key. - Cover alphabetical ordering, conditional isolated-workspace controls, a restricted three-control policy, and partly visible sections. - Document sorting and operator visibility. ## Verification - `pnpm --dir ui exec vitest run src/pages/InstanceExperimentalSettings.test.tsx`: 45 tests passed. - `pnpm check:token-gates`: passed. - `pnpm -r typecheck`: passed. - `pnpm build`: passed. - `pnpm --filter @paperclipai/ui typecheck`: passed. - All PR CI checks passed on `88011049be`, including the chat integration shards, full typecheck, build, browser tests, and policy checks. Greptile is 5/5 with no review threads. - Local `pnpm test:run` reported eight failures in unchanged chat, company-skills, email-channel, and workspace exposure suites. Stopped the remaining local run after CI completed successfully. Isolated chat rechecks were skipped by the host database support gate and do not count as passes. All matching CI shards passed; the full local suite is not claimed as green. - Frozen installation is blocked on the base branch by existing overrides/patch configuration drift from the lockfile. Local validation uses `pnpm@9.15.4 install --no-frozen-lockfile`; the original lockfile and manifests are unchanged in this PR. - Diff reviewed for secrets, private references, and run artifacts. ## Risks Settings only change position or visibility. Existing values, managed locks, API contracts, and feature dependencies are unchanged. Each section keeps its own alphabetical list. Reverting this change restores the previous presentation. ## Model Used OpenAI GPT-6 (Codex), with reasoning, repository tools, and test execution. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have described the issue in-PR following the enhancement template - [x] I have not referenced internal/instance-local Paperclip issues or links - [x] My branch name describes the change and contains no internal ticket id - [x] I have run the relevant tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
fff410dfe7 |
fix(ui): retain bounded context for browser render errors (#13904)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its browser error boundaries recover from failed renders and report the exception when error monitoring is enabled. > - The boundaries already receive the React component stack, but discard it before reporting the error. > - A minified DOM insertion error can therefore lack enough context to identify the affected component. > - Browser translation can replace text nodes that React still uses as insertion anchors. > - This pull request preserves bounded component names and browser state for the next failure. > - Maintainers can locate the failed component without collecting page content or customer URLs. ## Linked Issues or Issue Description Related: #13719 attributes browser errors to the loaded release. #13784 supplies the browser environment. This change adds context to error-boundary reports. **What happened?** A browser error boundary reports a DOM `NotFoundError` with only a minified JavaScript stack. The React component trace is logged to the console, which production builds remove. Browser translation is a plausible cause, but the report cannot identify the component or confirm the translation marker. **Expected behavior** An error-boundary report includes a bounded component trace and limited browser state. It excludes component props, text, HTML, element IDs, arbitrary CSS classes, page URLs, and query strings. **Steps to reproduce** 1. Render a component with conditional content before a text node. 2. Replace that text node with a translation element outside React. 3. Enable the preceding conditional content. React tries to insert before the detached text node and throws `NotFoundError`. 4. The boundary shows its recovery UI, but the old report loses the component trace. **Paperclip version or commit** Based on `6681c71b40`. The regression test reproduces the DOM mutation with the installed React version. **Deployment mode** Built browser UI with optional Sentry monitoring enabled for the signed-in session. ## What Changed - Pass the component stack and boundary kind from both error boundaries. - Keep at most 40 component names from at most 16 KiB of stack input. Drop locations and unrecognized lines. - Snapshot document readiness, visibility, and the browser translation root-class marker before the asynchronous reporting queue runs. - Attach diagnostics to that event only. Preserve the existing monitoring gate, sign-out behavior, and original exception if diagnostics fail. - Preserve function names in production bundles. Test the actual Vite production pipeline. - Document the fields, privacy limits, and translation-marker limitations. ## Verification - Focused diagnostics, boundary, real Sentry SDK, and production-build tests: 49 passed. - The translation DOM-mutation test reproduces `NotFoundError`, verifies the failed component trace, and keeps the recovery UI usable. - The real SDK test checks emitted events, private fixture exclusion, event isolation, and no capture after sign-out. Its transport stays in-process. - `pnpm check:token-gates`: passed. - [Greptile review](https://github.com/paperclipai/paperclip/pull/13904#issuecomment-5804220318): 5/5 on `92a2a23d9c`, with no review threads. - `pnpm -r typecheck` and `pnpm build`: passed. - Full CI unit and integration test matrix: passed, including all server, chat, workspace, runner, and serialized suites. - Local `pnpm test:run` was started, then stopped after the equivalent full CI matrix passed. The unsharded local run was not completed and is not claimed as a pass. - `pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/project-repositories.spec.ts --repeat-each=3 --trace=on --reporter=line`: six tests passed against the production build. - Initial CI browser shard 8 timed out after the repository form replaced an enabled Save button with a disabled one before the click. The retained page snapshot shows the original repository selection. No error-boundary fallback appeared. [CI attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35929990870/attempts/2) passed the failed jobs on the same commit. The full PR check matrix is green. This confirms an intermittent failure, but does not establish the cause of the first failure. - Full browser builds with and without name preservation passed. The initial JavaScript chunk grows from 1,571,919 to 1,668,453 gzip bytes (+6.1%). Total JavaScript across all chunks grows by 219,610 gzip bytes (+5.3%). ## Risks - This adds diagnostics. It reproduces a translation failure mode but does not identify or repair the specific application component from a past report. - The root-class marker is a hint. Other translation tools may omit it, and its presence does not prove causation. - Name preservation increases bundle size as measured above. No source maps are published by this change. - The error is still reported. DOM operations, browser translation, and recovery behavior are unchanged. ## Model Used OpenAI GPT-6 through Codex, with reasoning, code editing, terminal tools, and test execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.14 |
||
|
|
a6448cd060 |
fix(runner): tolerate Codex account notifications during active turns (#13902)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - The Rust runner connects task runs to Codex app-server sessions. > - Codex sends account notifications on the connection during authentication updates. > - These notifications have no task, thread, or turn identity. > - The runner treated a valid account update as an invalid task event and stopped the run. > - This pull request classifies account updates as connection information while preserving task identity checks. > - Sandbox images can also contain an older runner with compatible metadata. The server must compare its bytes before reuse. > - Slack conversations use the deployed runner and can complete when Codex refreshes account state. ## Linked Issues or Issue Description Related: #13853. That change fixes the supported Codex version range. This fixes a separate notification failure after the version check succeeds. **What happened?** An active Codex task stopped with `thread_binding_mismatch` when the provider emitted `account/updated` without a thread ID. Slack showed that the agent stopped before completing its turn. **Expected behavior** Connection-level account updates must not stop a task or acquire task authority. Account details must not appear in task output. **Steps to reproduce** 1. Start a task with the native Codex runner. 2. Emit `account/updated` during the turn, with `authMode` and `planType` but no thread ID. 3. Observe that the old runner rejects the notification and fails the turn. **Paperclip version or commit** Reproduced after #13853. The fix is based on `b648d8cdd`. **Deployment mode** Cloud staging with a remote Codex runner and a Slack chat connection. ## What Changed - Classify `account/updated` and `account/login/completed` as connection information. - Reject account notifications that contain execution identity fields. - Reuse the existing bounded diagnostic path without publishing account payloads. - Test both notification types during two consecutive turns. Preserve existing identity rejection tests. - Compare preinstalled sandbox runner bytes with the controller artifact. Stage the deployed binary after a mismatch, failed checksum, or timeout. Keep exact retained artifact reuse. ## Verification - `cargo test --release --locked -p paperclip-runner-core`: passed, 582 test executions; two existing ignored cases. - `cargo test --release --locked -p paperclip-runner-core --test codex_provider`: 89 passed; two existing ignored cases. - `cargo fmt --all` and `git diff --check`: passed. - Repository-wide `pnpm -r typecheck`: passed. Server typecheck also passed after the artifact-selection change. - Server executor suite: 456 tests passed, including preinstalled digest match, mismatch, checksum failure, and timeout. - Repository-wide Vitest and build are running. - Deployed `5648d90d5dd762c5a8b697face269f6eea9cd59d` to one staging stack and verified its serving SHA. - Retried the failed conversation through the Slack UI. The agent replied and the task completed. - Sent a fresh Slack mention asking which bots belong to the channel. Verified successful `slack_members` and `slack_user` calls, a successful run, and the delivered Slack answer. - Sent a follow-up in the same Slack thread without another mention. The agent returned the requested names from context, with one final reply. ## Risks - Only two known connection-level notification methods change behavior. Unknown authoritative methods and malformed or conflicting identities still fail closed. - An older sandbox runner now requires one upload when its bytes differ from the controller artifact. Remote OS and architecture checks still apply. - No schema, credential, permission, Codex version-range, or UI changes. The Codex minimum remains 0.149.0. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code editing, terminal tools, and browser verification. The context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.13 |
||
|
|
30182c704c |
feat(telemetry): propose connector telemetry events behind proposal markers (#13582)
Adds connection.created / connection.updated / connection.invoked as proposed-telemetry wrappers (unregistered; the TelemetryClient cannot queue or send them) plus the server-side emitter wiring and tests. No telemetry is transmitted by this change. Proposal: #13578. Coordinated release plan and review package: PAP-24 (Step A.1).canary/v2026.923.0-canary.12 |
||
|
|
6681c71b40 |
fix(ci): stabilize chat startup and close the initial live-update gap (#13895)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Browser tests check chat state across navigation, reload, and agent runs. > - Runtime tests check that a service is ready before Paperclip publishes its address. > - CI for #13891 failed in these paths, then passed attempt 3 with the same code. > - The browser traces stopped during development asset startup. A separate sidebar assertion used an unstable focus path through the rich editor. > - This PR gives browser tests fresh built assets and a direct keyboard path to the star button. It adds service-worker reload coverage under CPU throttling. > - A later browser failure exposed a real reload race: comments can change between the first query and the first live subscription. The UI now refreshes active queries when that subscription opens. > - Runtime fixtures now have separate registry state and better failure evidence. The readiness deadline remains unchanged. > - A later CI run exposed a wall-clock backoff assertion and three authorization cases sharing one test lifecycle. The PR anchors the assertion to transport time and separates the cases. ## Linked Issues or Issue Description Refs #13891. Related runtime ownership and cleanup work: #11791, #11389, #11278. Evidence: [original CI run, attempt 3](https://github.com/paperclipai/paperclip/actions/runs/35910335089/attempts/3). Attempt 3 passed both affected shards without code changes. This PR is separate from the wake-payload change, which has since merged. | Failure | Diagnosis and classification | | --- | --- | | Sidebar blank page and retry text missing after reload in CI | Both traces show a blank document before React startup. Vite connects, but the failing page makes no application API requests. The service worker forwards the unbundled development module graph. The retry response is already stored and its process-adapter run succeeded before reload. This places the failure in browser bootstrap, not wake payload or reply persistence. The exact reason the development module graph stopped is not established by the retained trace. The harness now serves built assets, and the new test covers startup and controlled reload at 4x CPU throttling. | | Star opacity remains zero on macOS | Reproduced on unchanged master `f55759942b`. The old test clicked the rich editor, focused the star, then used Tab and Shift+Tab. An instrumented baseline run captured the sequence: Tab moved from the star to the next sidebar link, then the editor bundle called `focus()` on its contenteditable before Shift+Tab. That key reached the composer, where it is a work-mode shortcut. This confirms a pending editor selection update stole focus; it was not the browser skipping sidebar buttons. The assertion did not prove the star retained focus. The test now moves the pointer away, focuses the preceding sidebar link, presses Tab once, and asserts both actual focus and opacity. This is test synchronization and keyboard traversal, not a demonstrated CSS defect. | | Runtime readiness exceeds 10 seconds | The old error only says `fetch failed`. There is no child startup output in that failure, so it cannot distinguish slow process startup from a refused or stalled probe. It passed unchanged locally and took 3.416 seconds in attempt 3. Resource contention is plausible but unproved. This PR does not claim a proven historical runtime root cause: it isolates fixture registry/log files, checks ports before spawn and after stop, checks live backends at publication, and preserves transport errors, probe count, elapsed time, and fixture startup timestamps for the next occurrence. | | Later CI: recovery reply disappears after reload | Product synchronization bug, distinct from the blank-page bootstrap failure. Reproduced locally with a trace: the comments request started at `21:56:07.443`, the server saved the reply at `.520`, and the first live subscription started at `.541`. The reload fetched a successful run and its complete log, but missed the comment event. The provider refreshed after reconnects only. It now refreshes active queries on the first connection too. A deterministic regression test fails before the fix. No browser assertion was changed. | | Later CI: Slack backoff and authorization tests | The 30-second backoff assertion required more than 25 seconds to remain when it read the saved action. CI spent 10.311 seconds in the test, exceeding that five-second allowance. A local six-second read delay reproduces the failure; the new transport-anchored lower and upper bounds pass the same fault injection. The neighbouring authorization test ran three independent fixtures in one test and timed out at 15 seconds. Each case now has its own fixture cleanup and the normal per-test deadline, so earlier cases do not remain active during later global worker sweeps. No specific production slow call was established. | | Self-hosted runner loses communication or shuts down | The original lost-communication failure has no assertion. On final-head [attempt 1](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/1), Build, Runner Vitest 1/2, and chat 2/3 ran on three separate fleet instances. All received a runner shutdown signal at `22:04:55 UTC`, within 35 milliseconds, then cancellation. Server shard 6/12 received the same shutdown signal one minute later. All four were Spot `m7i-flex.xlarge` instances in `us-east-1a`. No test assertion or build error preceded those stops. This is infrastructure interruption; the reason the fleet stopped the runners is not available in job logs. | ## What Changed - Build the browser fixture UI into the static server's preferred directory, `server/ui-dist`, and disable Vite middleware. Ignore these generated assets. Start the source CLI directly from the repository root, as required by the CLI invocation safety contract. - Keep all existing chat assertions. Add first takeover and three service-worker-controlled reloads under CPU throttling. Assert that the page uses built module assets and that opening it creates no chat task. - Use forward keyboard traversal from the Zeta link to its star. Check focus before checking the reveal style. - Give each runtime exposure test a temporary Paperclip home and restore environment state after process cleanup. - Check that reserved ports are free before spawn, serve the fixture response before exposure, and become free after stop. - Include the nested fetch error, probe count, and elapsed time in readiness failures. Add a unit test for this diagnostic contract. - Refresh active queries when the first live connection opens, covering events missed during initial page loading. Keep reconnect toast suppression unchanged. - Measure Slack retry timing from the transport attempt and recovery completion. This checks the full provider-requested delay without spending a small wall-clock allowance on unrelated processing. - Run each Slack authorization-revocation scenario as a separate test, with cleanup between cases. All assertions remain. - Document the browser fixture's build and serving mode. No configured assertion deadline, readiness deadline, retry count, or skip was added. ## Verification - Final-head [Linux CI, attempt 2](https://github.com/paperclipai/paperclip/actions/runs/35925901613/attempts/2): **green**. All 52 check runs passed; two conditional checks were skipped. The legacy Snyk status also passed. No pending or failed checks remain. - `pnpm -r typecheck` passed. - `pnpm build` passed again after the final UI fix. - Focused runtime suites: 38 passed, 3 existing platform skips. - CLI invocation safety suite: 39 passed. - Slack timing negative control: inserting a six-second delay before reading the saved action fails the old assertion. All four revised authorization/backoff cases pass with that same delay. The diagnostic delay is not committed. - Live update suites: 92 passed, including the new regression. UI typecheck and token gates passed. - Recovery browser negative control: the unchanged tests reproduced the missing reply (1 failed, 9 passed). After the first-connection fix, all six recovery paths passed twice (12 passed). Browser assertions and deadlines are unchanged. - Server typecheck passed after the Slack test adjustment. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-chat --shard-index 0 --shard-count 3`: 342 passed. The other 682 tests belong to the remaining shards; collection verified exact coverage. - Final direct-source CLI launch: all five chat session tests passed. - `PAPERCLIP_E2E_PORT=32993 PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome pnpm exec playwright test --config tests/e2e/playwright.config.ts tests/e2e/agent-chat-sessions.spec.ts --repeat-each=3`: 15 passed, including nine controlled reloads at 4x CPU throttling. - `GITHUB_WORKFLOW=PR pnpm test:run:general -- --group general-server-without-chat --shard-index 10 --shard-count 12`: 55 files passed, 867 tests passed, 21 existing skips. An earlier run could not initialize PostgreSQL because this Mac exhausted its System V shared-memory slots. After reclaiming the orphaned segment from this task's stopped browser server, the full shard passed. - On final commit `1f1fafc08d`, [browser 4/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400837947) passed all 16 tests; [browser 8/8](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838155) passed all 23 tests; [server 11/12](https://github.com/paperclipai/paperclip/actions/runs/35925901613/job/107400838328) passed 887 tests with one existing skip. The readiness lifecycle case took 1.842 seconds. Chat 1/3 also passed. All four jobs interrupted by runner shutdowns passed unchanged on their single rerun. - Greptile reviewed `1f1fafc08d`: **5/5**, with no unresolved comments. - Full local `pnpm test:run` was attempted and stopped after confirming failures outside this patch: a sibling `skills` directory shadows bundled Slack/AgentMail skills; macOS rejects rename of read-only skill-cache directories (`EACCES`, also reproduced in an isolated filesystem probe); and the host exhausts PostgreSQL System V shared-memory slots. Focused reruns confirmed these limits. The complete Linux CI run is the repository-wide verification; the full local run is not green. - Baseline evidence: the original sidebar test failed on unchanged master; a development-mode run at 4x CPU throttling passed six selected cases, so CPU pressure alone did not reproduce the CI bootstrap stall. ## Risks - Default browser tests now exercise the shipped static UI. They no longer implicitly cover Vite middleware or HMR; use the development server for those checks. - The new browser startup test uses Chromium CDP, matching the only configured browser project. - Opening a live subscription now causes one active-query refresh to close the initial event gap. This adds startup API reads but no recurring poll. - Runtime behavior and deadlines are unchanged except for error details. The historical readiness stall remains unconfirmed; a green rerun alone cannot establish its cause. - Local verification runs on macOS. The runtime lifecycle checks also passed on Linux CI. ## Model Used - OpenAI GPT-6 through Codex. The session identifies the model family as GPT-6; an exact served model ID and context window size are not exposed. Used reasoning, repository inspection, shell tools, code editing, and test execution. No subagents were used. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #123` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass — targeted suites passed; full local host limits are listed above - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
24429024e7 |
feat: add Fireflies connector and summary-ready routines (#13890)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Apps gives agents governed access to external tools through stored credentials. > - Routines start work when an external service sends an event. > - Fireflies provides meeting transcripts and summaries through an official hosted MCP server. > - This PR adds that connection and accepts signed meeting events through the shared app webhook flow. > - Agents can review completed meetings with the same permissions and audit records as other work. ## Linked Issues or Issue Description **Problem or motivation** Operators need agents to read Fireflies meetings and start follow-up work when a summary is ready. The Apps catalog lacks Fireflies. The shared app webhook flow needs to accept its signed deliveries. **Proposed solution** Use the official Fireflies MCP endpoint with OAuth or a vaulted bearer API key. Extend the existing Another app or script flow with signed webhook support. Verify the raw-body signature and pass the JSON payload as external data. Select Meeting Summarized in Fireflies. Deduplicate identical signed deliveries, including setup deliveries. **Alternatives considered** A separate REST connector would duplicate the governed MCP path. Polling, legacy V1 payloads, and automatic provider-side webhook registration are outside this change. **Roadmap alignment** This extends the existing MCP Tool Gateway & Apps and Scheduled Routines surfaces. It adds a provider to those systems. It does not introduce a second integration framework. **Additional context** A GitHub search found no existing Fireflies issues or PRs. Provider references and verification limits are in `doc/connections/FIREFLIES.md`. ## What Changed - Add the official Fireflies catalog definition, generated registry, provider evidence, and branded artwork. - Reuse Access → Connect, dynamic discovery, Permissions, vault storage, policy, and audit behavior. - Classify Fireflies sharing, movement, and access revocation as writes. - Preserve Off and Ask first restrictions during OAuth reauthorization and API-key replacement. New actions retain normal defaults. - Add `app_webhook` authentication to the shared Another app or script flow. Accept bearer tokens or raw-body HMAC-SHA256. Preserve earlier `fireflies_hmac` triggers and revision snapshots for compatibility. Existing text columns need no migration. - Verify `X-Hub-Signature` or `X-Hub-Signature-256` against the exact request body. Preserve generic event payloads and deduplicate identical signed requests. - Keep the routine wizard generic. Show one webhook URL and secret in Another app or script. Keep all new app webhook event names provider-neutral. Keep provider setup instructions in the connector documentation. - Pass generic webhook JSON to the task in an explicit external-data block, capped at 16,384 characters. Keep strict meeting validation for existing legacy Fireflies triggers. ## Verification - Feature implementation commit `0882dc8a1`: all 54 CI checks passed; two conditional Storybook checks skipped. This includes full tests, typecheck, build, browser E2E, canary dry run, and security checks. Greptile rated this commit 5/5; all review threads are resolved. - Full local `pnpm -r typecheck`, `pnpm build`, and token gates passed on the final code. Targeted connector, gateway, webhook, revision, and UI suites passed during implementation. After the provider-neutral follow-up, all 84 app-webhook and routine-service tests passed; the final payload-to-task assertion also passed in the 72-test routine suite and a clean-config rerun. - The long local `pnpm test:run` invocation started before the final edits and was stopped after the final-commit CI suites passed. It reported one generic webhook test failure while those files were changing; that test and the entire routine suite passed on the final source, including a clean-config reproduction. The interrupted local run is not counted as a full-suite pass. - In the embedded browser, completed official OAuth consent and discovered 20 live actions. Real meeting listing, transcript retrieval, and summary/action-item retrieval succeeded as the selected agent. Turning a live read Off blocked its test; catalog refresh preserved the restriction. - Embedded-browser Another app or script setup, back/save/resume, narrow layout, and a signed synthetic Fireflies delivery succeeded. The UI reported authentication passed without creating a task. Fixtures cover signature tampering, malformed requests, ordinary app event names, duplicate/setup deliveries, rotation, revisions, pause/archive, and company isolation. - Existing MCP browser suite: 8 passed and 2 provider-dependent cases skipped. Branding checks passed; connector artwork and webhook setup were checked at desktop/mobile widths and in light/dark modes. - An unauthenticated POST to a correctly formatted public webhook URL reached the staging tenant verifier through the existing Cloud gateway. - A real Fireflies webhook delivery remains unverified. A staging callback is available for the operator walkthrough. Live API-key authorization, credential expiry, and a new meeting's summary completion were not tested against the provider. Fixtures cover these protocol and lifecycle paths where applicable. - Storybook follow-up `c54174faa`: 27 production-component stories cover every UI change, with a source-to-story map in the connector documentation. Static Storybook build, UI typecheck, token gates, and Playwright checks for all stories and the mobile footer pass. All PR checks passed for this Storybook follow-up; Greptile reviewed `c54174faa` at 5/5. ## Risks - Fireflies may change its hosted MCP tools or OAuth behavior. Tool discovery stays dynamic. Experimental search/fetch tools are not required. - Public webhook setup requires HTTPS and a separate signing secret. Fireflies normally emits events for meetings owned by the configuring account. - Reauthorization touches shared MCP permission code. Regression tests cover existing restrictions, new actions, connection removal, and other gateway callers. - Webhook receipt grants no tool access. The routine agent still needs an authorized Fireflies connection. - New generic triggers rely on provider event subscriptions. Without a sender-supplied idempotency key, changed request bytes count as a new event. Existing legacy Fireflies triggers retain summary-only filtering and per-meeting deduplication. ## Model Used OpenAI Codex, model `gpt-6-astra`. Used reasoning, repository editing, code execution, and embedded-browser testing. The runtime did not expose a context-window size. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing> |
||
|
|
b648d8cdda |
fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path > - Paperclip manages AI agents and their provider connections. > - Product E2E checks real tasks through the browser, server, and runner. > - Grok qualification needs separate API-key and subscription evidence. > - Product subscription tests and direct Grok protocol evals need explicit credential delivery. > - This change supplies each credential only to its selected profile and prepares the pinned binary. > - Maintainer authorization and protected-environment gates remain required. ## Linked Issues or Issue Description Refs #13845, #13847, #13850, #13882. The Grok feature branch has a manual subscription qualification profile. The trusted master workflow must admit its selected credential and prepare the same verified binary and artifact verifier as the API profile. Direct protocol evals also need the selected xAI key and pinned Grok binary. These prerequisites do not register or schedule the new profiles on master. ## What Changed - Deliver `GROK_AUTH_JSON` from the protected paid environment only when the selected profile requests that credential. - Install the checksum-verified Grok binary for the local subscription profile. - Prepare the pinned artifact verifier for the manual subscription suite. - Extend workflow security assertions to cover the new credential and profile. - Add the ACPX Grok credential mapping to the trusted-master catalog, then deliver only the selected `XAI_API_KEY` to direct protocol cells and install the target’s checksum-verified Grok binary before packaging. - Allow a direct-protocol concurrency override from two cases up to the existing configured ceiling; it can only lower concurrency. - Document the Grok protocol workflow and its API-only credential boundary. - Render missing LLM usage and cost as Unavailable, and label partial observations with coverage. Preserve raw records, grades, and actual zero costs. - Preserve measured campaign source metadata during report regeneration instead of inheriting the renderer checkout or CI event; skip empty legacy source records when recovering older provenance. ## Verification - Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54 reported checks successful, two intentional skips, Greptile 5/5, and zero unresolved review threads. [CI run](https://github.com/paperclipai/paperclip/actions/runs/35890978288). - After merging current master, all 17 workflow security/image tests and 23 catalog/workflow policy tests passed. The trusted catalog also generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per shard. The new policy tests execute the concurrency guard against valid, out-of-range, and malformed values. - The Grok branch separately passed 450 Product harness unit tests, including private company credential staging, cleanup, and token-fragment redaction. - The fresh-login native subscription smoke passed three repetitions of tool execution, session resume, restrictive permissions, and cleanup. These are setup evidence; full subscription Product qualification remains pending. - All 72 focused report/billing/history/catalog tests and the Product harness typecheck passed for the report-display change. The initial sandbox run could not open the tsx IPC socket; the permitted rerun passed. A zero-provider-call replay of the actual 16-cell Grok campaign preserved all result records, grades, timing, and source provenance while correcting missing usage labels. - Review the thirteen-file diff. Provider credentials still enter only the selected paid-test step; default-branch, numeric-actor, and environment restrictions are unchanged. ## Risks This admits a refreshable subscription credential to explicitly selected trusted tests. Store it only in `runner-e2e-paid`, use a test login, and remove it after qualification. Unselected profiles receive an empty value. Pull requests cannot trigger the paid workflow. This PR changes no fleet admission or actor allowlist. ## Model Used OpenAI GPT-6 through Codex, with tool use and code execution. The exact serving model identifier and context-window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.11 |
||
|
|
4b8ec588f3 |
Stop duplicating wake context in adapter environments (#13891)
## Thinking Path
> - Paperclip manages agent work and preserves task context.
> - Built-in adapters already include wake context in the agent prompt.
> - They also copy the full wake JSON into a process environment
variable.
> - A large environment entry can prevent the agent from starting with
`spawn E2BIG`.
> - This change removes the duplicate environment entry and uses the
existing prompt delivery.
> - The agent keeps its context without extra file transport or new
history limits.
## Linked Issues or Issue Description
Refs #13144, #13860, #13872, #13793.
Large wake payloads can exceed the operating system limit for one
environment entry. The launch-envelope fix in #13793 handles the outer
transport but leaves that child environment entry intact.
Credit to @nickyleach for the prompt-only approach in #13144. This PR
applies that part on current master. It does not include that PR's
30-item history limits or recovery-history endpoint. Those behavior
changes can be reviewed separately from the process launch fix.
## What Changed
- Stop exporting `PAPERCLIP_WAKE_PAYLOAD_JSON` in the shared ACP engine
and all ten built-in adapter writers.
- Ignore configured values of the retired variable so saved adapter
settings cannot restore the oversized entry. Also drop inherited copies
in Hermes, which builds its environment directly.
- Keep scalar runtime variables, existing prompt rendering, continuation
history, resume deltas, gateway bodies, and Hermes JSON template
variables.
- Document the prompt delivery contract and the migration for custom
instructions that read the retired variable.
- Test large local and sandbox child-process launches, fresh and resumed
ACP turns, SDK delivery, and configured-variable filtering.
## Verification
- `pnpm -r typecheck` passed.
- Focused adapter utility, ACP, Codex child-process, and Cursor Cloud
suites: 340 tests passed.
- Hermes execution and prompt tests: 18 tests passed using its package
Vitest configuration.
- The child-process tests deliver over 128 KB of context through stdin
and check the complete text. The ACP test retains 50 complete messages
and 50 completed actions, then checks the resumed delta.
- `pnpm build` passed.
- `pnpm test:run` was attempted, then stopped after it reproduced ten
macOS runtime-skill-cache permission failures (also reproduced on
unchanged master) and one HTTPS backfill test failure. The HTTPS test
passed when rerun unchanged on this branch and master. The complete
local suite was not completed; Linux CI provides the full-suite gate.
- CI is green on
canary/v2026.923.0-canary.10
|
||
|
|
f55759942b |
fix(runner): accept compatible Codex releases from 0.149.0 (#13853)
Keep the minimum fixed at 0.149.0 until a deliberate maintainer change. Accept stable versions below 0.157.0 separately from the reproducible install pin. Co-Authored-By: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.9 |
||
|
|
b41ccf097f |
fix(apps): configure MCP aggregators from inline task cards (#13879)
## Thinking Path > - Paperclip lets people manage AI agents and their work. > - Agents request connections through cards in task threads. > - MCP aggregators need provider-specific URLs and authentication. > - The task dialog used the generic setup form and omitted these fields. > - This pull request uses the same provider setup controller in tasks and Apps. > - Users can configure a connection without leaving the task. ## Linked Issues or Issue Description Related: #13755, #13855. **What happened?** An inline Executor request opened a very wide dialog with an empty credential step. Connect failed because the MCP URL was missing. The other aggregator cards also bypassed their provider setup. **Expected behavior** Each card shows its provider instructions, URL field, and authentication options in a bounded dialog. Completing setup grants access only to the requesting agent. **Steps to reproduce** 1. Enable experimental MCP aggregators. 2. Have an agent request Zapier, Arcade, Composio, or Executor from a task. 3. Open the card and continue past Access. ## What Changed - Route page and task setup through the same provider controller. - Bound the task dialog width and preserve the requesting agent's access scope. - Support existing accounts, saved drafts, URL/token setup, and task-bound OAuth. - Keep a sign-in link available when the browser cannot open a popup. Verify completion through the existing durable callback path. - Add inline Access, configuration, and narrow Storybooks for all four providers. - Document the shared setup requirement and correct Executor's URL instructions. ## Verification - Focused Vitest selection: 24 passed. Covers all four inline forms, requester access, existing accounts, saved drafts, OAuth retry, callback validation, popup cleanup, and generic reconnect endpoint preservation. - UI typecheck, UI build, design-token gates, and Storybook build passed. - Live local browser: new Executor, Arcade, and Composio connections completed provider consent from task cards. Each appeared Connected with the requester selected. - Real Test calls returned Executor output `4`, an Arcade public GitHub star count, and Composio tool-discovery results. An ungranted agent was denied access. Real Paperclip process-agent runs discovered each provider catalog with only the requester’s connection installed. - Deployed implementation commit `5d722e89c` to the isolated staging tenant. The original failing Executor card now completes, discovers seven actions, limits access to its requesting agent, and resumes that agent. Its continuation completed real Executor calls and the provider resume flow, then returned an upstream Airtable authorization link. A real staging Test call returned `4` in 1.5 seconds. Later PR commits add regression coverage and popup-unmount cleanup. - Zapier fresh-token browser test remains pending a provider clipboard handoff. Its URL and token flow passes focused tests. - Latest-head CI (`d14f4c73c`): 53 checks passed; optional Storybook deployment and visual regression jobs skipped. Greptile 5/5, both review threads resolved. Canceled runners and unrelated chat timeouts passed the single retry on unchanged code. - The full local test suite was not run, as requested by the maintainer. CI runs the repository gates. ## Risks - OAuth popup behavior differs by browser. The explicit sign-in link and durable server completion checks provide recovery. - Saved task drafts store only a connection ID in browser storage. Credentials remain in the existing server vault. - No database or server protocol changes. ## Model Used OpenAI GPT-6 (`gpt-6-astra`) through Codex, with reasoning, code execution, and browser tools. The exact context-window size is not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.8 |
||
|
|
60c7c9cd1a |
fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Product E2E campaigns test the native runner in local and Daytona environments. > - Each campaign resolves one target lockfile and verifies its downloaded artifact. > - The Daytona image job did not pass that artifact digest to Docker. > - Docker used an older default digest and stopped before any selected task ran. > - This pull request passes and validates the campaign digest at the image build boundary. > - The image keeps its checksum check and frozen package installation. ## Linked Issues or Issue Description **What happened?** The merged-master [qualification campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409) stopped in the Daytona image build. The resolved target lock digest was `e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`. Docker used its default digest, `57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The checksum check rejected the mismatch. All 12 selected model cells were skipped. This follows the image provenance work in #13814. **Expected behavior** The image build must check the same lockfile artifact that the campaign restored and verified. A changed lockfile must still fail the checksum check. **Steps to reproduce** 1. Start a Product E2E campaign with a Daytona cell on master `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. 2. Resolve a target lockfile whose digest differs from the Dockerfile default. 3. Observe the provider-pack image stage reject the lockfile before model execution. **Paperclip version or commit** `7944ed3d976d1a7cc26a2d0cee51f227f3542084`. **Deployment mode** GitHub Actions Product E2E campaign with a Daytona image build. **Install method** Built from source with the campaign lockfile artifact. **Agent adapter(s) involved** Native Codex and ACPX Claude cells were selected. No model cell ran in this failed campaign. **Database mode** Not involved. The failure occurs during image creation. **Access context** The authorized default-branch paid workflow. The build receives no provider credentials. ## What Changed - Read the image checksum from the existing target-lock job output. - Require a 64-character lowercase hexadecimal digest before image inspection or build. - Pass the digest as the existing Docker build argument. - Add regression checks and document the campaign checksum handoff. ## Verification - The Daytona image regression fails with the original workflow and passes with the fix. - All six Daytona image contract tests pass. - All 450 Product E2E unit tests pass. - Product E2E typecheck passes. - Actionlint passes for the changed workflow. - A context-shaped resolution probe preserves the downloaded lockfile bytes and digest. - All latest-head CI checks passed on `67d41fd9439b2a9a809ddb05765f8617585072c5` ([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)). - Greptile gave 5/5 on this head; its test-scoping comment is addressed and resolved. - A hosted Daytona image rebuild and the three remote qualification cells remain pending after merge. ## Risks The campaign digest comes from the existing trusted target-lock job. The restored artifact checks, Docker checksum check, frozen install, content identity, image signing, and verification remain in place. The standalone Docker default remains available. This change does not alter task behavior, prompts, credentials, or dependency versions. ## Model Used OpenAI `gpt-6-astra` through Codex performed diagnosis and review with code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation, implementation, and verification. Context window limits are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.7 |
||
|
|
4721f55803 |
fix(apps): recover MCP OAuth setup after consent errors (#13855)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Connections give agents governed access to external tools. > - MCP aggregator setup can return from provider consent to a saved draft. > - The branded setup did not explain failed or cancelled authorization. > - It also offered identity changes that the server does not apply when a saved connection resumes. > - This pull request explains OAuth return outcomes and keeps the displayed identity consistent with the saved policy. > - Users can understand the outcome and retry the same connection. ## Linked Issues or Issue Description Related: #13755 introduced the MCP aggregators. #13758 retired the legacy Composio broker. #13584 proposes changes to the provider handoff window; this fix retains the current handoff behavior. **What happened?** Cancelling Composio consent returned to the setup form with no explanation. The same controller ignored failed OAuth callback outcomes. Returning to Access on a saved draft also offered personal/shared choices, although the server retains the saved identity. This could make the OAuth request disagree with that identity. **Expected behavior** Explain cancellation or failure, preserve the draft, and offer Try again. Display the retained credential identity and start OAuth with the policy returned by the server. **Steps to reproduce** 1. Enable MCP aggregators and start a Composio connection. 2. Continue to provider consent and cancel it. 3. Observe the return screen. Before this change, it showed the form without cancellation feedback. 4. Go back to Access. Before this change, the form offered personal/shared choices even though resume retains the original identity. **Paperclip version or commit** Observed before the fix on `8c6cc7dccf91523e0720bd86f95487e66b4b0e63`. The original report of successful consent leaving setup unfinished did not reproduce. This PR addresses the recovery defects observed during that investigation. **Deployment mode** Isolated local development instance and authenticated staging deployment, with real Composio consent and provider calls. ## What Changed - Read the OAuth callback outcome in branded MCP setup and show cancellation or failure feedback. - Retry the same saved draft without displaying untrusted callback error text. - Keep saved personal/shared identity fixed in Access and select OAuth identity from the returned credential policy. - Add six focused regression cases and authorization-failure Storybook states for Arcade, Composio, and Executor. - Document return-screen recovery and retained identity in the connector playbook. ## Verification - `pnpm exec vitest run ui/src/pages/apps/AppsConnect.test.tsx -t 'OAuth return' --maxWorkers=1`: 6 passed, 143 skipped, including after integration with current master. - `pnpm --filter @paperclipai/ui exec tsc --noEmit`: passed. - `pnpm check:token-gates`: passed. - `pnpm --filter @paperclipai/ui build`: passed. - `pnpm --filter @paperclipai/ui build-storybook`: passed for the implementation commit. - Browser recovery: cancelled real Composio consent, observed the new feedback, returned to Access, retried the same personal draft, completed consent, and ran a real tool call. - Staging on implementation commit `84bb40aa70662e0c8955bf692c5714661b4bea93`: fresh shared and personal connections each completed on the first consent attempt and loaded 11 tools. Real discovery, execution, and schema calls succeeded. Both connections stayed Connected after reload. A real agent used the shared connection through the Paperclip gateway and returned the public repository documentation hierarchy with one success and zero errors. - All current-head CI gates passed on `662f84a67e867a52a2e5526026adbed00f6b59bf`. The Cursor execution and agent-chat browser shards each had an initial timeout; both passed on one targeted rerun without code changes. Greptile reviewed this exact head at 5/5, with no open review threads. - No full local suite was run, as requested. CI provides the broader checks. The PR adds a master merge and documentation after the live-tested implementation commit. ## Risks - This shared setup controller also serves Arcade and Executor. Their callback rendering and retry behavior have focused test coverage; this investigation used Composio for live provider testing. - Saved identity remains fixed during resume. A different identity requires a new connection, consistent with server behavior. - No database, protocol, credential storage, or gateway policy changes. ## Model Used OpenAI GPT-6 via Codex, with code editing, shell tools, and browser testing. The exact runtime model ID and context window size are not exposed in this session. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.6 |
||
|
|
7944ed3d97 |
fix(runner): preserve hire runtime safety and first-activity timing (#13852)
## Thinking Path > - Paperclip is the open source control plane for companies of AI agents. > - Native runner agents need governed tools, durable runtime state, and useful execution evidence. > - A first activity trace waited 53.467 seconds even though tool activity took 6.274 seconds; provider input arrived before the server API call executed. > - Native agents also need a safe way to hire teammates without asking the model to rebuild runtime configuration. > - This pull request separates the observed ACP input-stream window from the actual server `tool.execute` span and adds a server-owned native hire contract. > - The benefit is clearer latency evidence and safer native teammates with existing approval, auth, and company boundaries preserved. ## Linked Issues or Issue Description Related Daytona provenance work is in [#13814](https://github.com/paperclipai/paperclip/pull/13814). No duplicate public PR was found for this combined timing and native-hire change. **What existing behavior does this improve?** Native runner agents can use governed tools and request hires. The server did not expose a safe native hire operation that reused the caller's validated runtime settings. First-activity traces also mixed provider input timing with server tool execution timing. **Current behavior** A native hire must construct a separate runner configuration. Full configuration copying could expose paths, instructions, secrets, or sessions. Timing evidence could make a provider or MCP identity join appear proven when the trace did not contain that join. **Proposed behavior** The native `hire_agent` operation accepts identity and persona inputs. The server sends `adapterType: "paperclip_runner"` with `inheritRuntimeFrom: "caller"`, then copies only validated provider, model, permission, lifecycle, and bounded execution settings. It inherits and validates the default environment, derives the managed AI binding through existing normalization, preserves approval and permissions, and creates fresh child instructions. Caller secrets, paths, prompts, and sessions are excluded. Provider events now include the optional boolean `inputUpdated`, with Rust forwarding support. Timing evidence separately records the ACP input-stream window and the actual server `tool.execute` activity. It does not claim a provider or MCP join without matching evidence. **Reason and benefit** Native agents can hire teammates that start with the caller's approved execution policy. Operators retain company boundaries, auth rules, approval gates, and requalification. Reviewers can distinguish provider streaming time from server API execution time when diagnosing first-activity delays. **Breaking changes** None for existing hires or tool calls. `inheritRuntimeFrom` is optional and only applies to same-company native agent callers. Conflicting explicit runtime settings are rejected. The provider event field is optional for existing producers. ## What Changed - Added the native `hire_agent` protocol action, catalog entry, API contract, and runner authority checks. - Added `inheritRuntimeFrom: "caller"` validation and a closed native runtime inheritance allowlist. - Preserved managed AI binding normalization, default-environment validation, approval snapshots, permissions, requalification, and fresh child instructions. - Added provider `inputUpdated` schema support and Rust forwarding. - Added first-activity and server tool timing evidence with conservative identity-join handling. - Added route, authority, provider-event, sidecar, API, catalog, and Rust-focused tests. - Kept private Honeycomb links, raw traces, and local result paths out of this description. ## Verification Focused checks passed: - 458 timing/session checks. - 61 native hire inheritance checks. - 20 hire authority checks. - 1,741 API checks. - 106 catalog checks. - 54 provider sidecar checks. - 12 Rust provider checks. Live R2 and R3 each passed 45 checks across 6 runs (361,135 ms for R2). R1 stopped at missing Docker image setup. The final trace is available at https://ui.honeycomb.io/paperclip/environments/test/datasets/paperclip/result/BiMypLNvmiB?tab=traces. Latest-head CI passed all required build, typecheck, Rust, static, Vitest, serialized-server, workspace, chat, and E2E jobs. The focused local checks listed above passed; the broad local suite was not run before the live evaluation, while CI provides the full repository verification. ## Risks - Timing fields describe separate observed windows. They do not prove a provider or MCP owner without a valid trace join. - The inheritance allowlist must stay synchronized with native runner configuration fields. - Approval snapshots include resolved safe inherited settings and should be reviewed when native configuration fields change. - The focused local suite is narrower than the full repository suite; latest-head CI covers the broader repository checks. > Roadmap review: `ROADMAP.md` places this work within Paperclip's bring-your-own-agent direction. It extends existing native runner hiring and observability behavior. ## Model Used OpenAI GPT-6 (exact serving model ID is not exposed), with extended reasoning and repository tool use; GPT-5.6 Luna assisted with focused implementation and verification work. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>canary/v2026.923.0-canary.5 nightly/v2026.923.0-nightly.0 |
||
|
|
64faf0ae90 |
fix(ui): leave the archived company's settings after archiving it (#13846)
## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - A person can archive a company from that company's settings page. > - After the archive, the page stays on the archived company. The only visible change is the button text "Already archived". > - It is unclear that anything happened, and the user has no next step on screen. > - This pull request navigates away after a successful archive: to another active company, to the Cloud portfolio when no active company remains on a managed instance, or to the companies list when self-hosted. > - The benefit is a clear outcome: the user sees where they are now and a toast that names what happened. ## Linked Issues or Issue Description No public issue exists. Description follows the enhancement template: **What existing behavior does this improve?** The archive action on the company settings page. The mutation works, but the view stays on the archived company and gives no feedback beyond a disabled button. **Subsystem affected** UI: company settings page, company selection helpers, cloud links. **Current behavior** Archive succeeds. The button changes to "Already archived". The user stays on the archived company's settings. On a single-company instance nothing else changes. **Proposed behavior** After a successful archive, the app departs: another active company's dashboard with a toast; the Cloud portfolio's manage view when no active company remains on a managed instance; the companies list when self-hosted with nothing active, because that list shows the archived state and owns the unarchive action. **Reason and benefit** Staying on the archived company reads as "nothing happened". Leaving to a live surface makes the outcome clear and gives the user their next step. ## What Changed - `ui/src/lib/company-selection.ts`: new pure `resolveCompanyArchiveDeparture` helper. Another active company wins; the just-archived company is excluded by id, so a stale cached status cannot select it. Cloud portfolio is the fallback on managed instances; the companies list is the final fallback. - `ui/src/lib/cloudLinks.ts`: new `cloudPortfolioManageUrl` (`/orgs?manage=1`). The manage view matters: the plain launchpad auto-forwards a solo user back into their one openable stack — the page this navigation is escaping. - `ui/src/pages/CompanySettings.tsx`: the archive mutation now departs on success — toast + `navigate` for in-app destinations, a top-level navigation for the Cloud portfolio — instead of only setting the selected company id. - `ui/src/pages/CompanySettingsRenameHint.test.tsx`: render harness wraps the page in a `MemoryRouter` for the new router dependency. ## Verification - `pnpm exec vitest run src/lib/company-selection.test.ts src/lib/cloudLinks.test.ts src/pages/CompanySettingsRenameHint.test.tsx src/pages/CompanySettings.test.tsx` — 22 tests pass. - New unit tests cover: active sibling wins (also on cloud), the just-archived company is never the destination even with a stale cached status, cloud portfolio fallback, self-hosted companies-list fallback, and the portfolio URL helper (null base included). - `pnpm exec tsc --noEmit` in `ui/` is clean after the plugin SDK build. ## Risks - Low risk. The archive API call is unchanged; only post-success navigation is new. The toast uses the optional actions hook, so surfaces without a ToastProvider stay safe. - On Cloud, the portfolio navigation is a full top-level navigation, so the query-cache invalidation for the departed document is skipped intentionally. ## Model Used - Claude Fable 5 (`claude-fable-5`), via Claude Code CLI, extended thinking and tool use enabled. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change (e.g. `docs/...`, `fix/...`) and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting mergecanary/v2026.923.0-canary.4 |