## Thinking Path > - Paperclip is the open source app people use to manage AI agents for work. > - Its control plane decides when a task can continue, wait, stop, or complete. > - Legacy continuation could change when an agent changed its wording without changing task state. > - Shared attempt counts also let repair and infrastructure retries affect each other's limits. > - This pull request uses persisted state and separate, bounded allowances for these decisions. > - If automatic repair stops, the task explains what happened and offers a guarded retry. > - Paired tests and real-provider evaluations verify that Stop, approvals, ownership, and spending limits remain authoritative. ## Linked Issues or Issue Description Related work: Refs #13761, Refs #11126, Refs #13610. These cover obsolete continuation dispatch and retry storms. Open and closed issues and PRs were searched for related lifecycle, continuation, and retry work. **What happened?** Legacy continuation depended on English wording and progress heuristics. Repair, failure retry, and productive continuation could consume shared counts. When bounded repair stopped, the task showed a technical recovery message without a clear next action. **Expected behavior** Persisted disposition and owned execution paths determine the next action. Missing disposition prompts bounded agent repair. Explicit work mode determines planning mode. Narrative changes and raw activity counts cannot replenish allowances. An exhausted repair shows a readable notice. An explicit retry checks current controls and preserves the assigned agent. **Steps to reproduce** Run `pnpm test:lifecycle-baseline`. The paired probes keep structured state constant while varying completion, planning, blocker, and progress prose. Run the explicit `lifecycle-baseline` and `continuation-accounting` Product E2E suites for real-provider coverage. In Storybook, open **Design previews / Recovery notice** to inspect the production component's normal, pending, acknowledged, unavailable, failure, and mobile states. ## What Changed - Hide the image attachment button, icon, and drop/paste hint in answer composers. Image paste and drop support remains available. - Merge current master and retain both browser regression sets. Use a production-stamped service worker in the offline recovery browser fixture. - Share one state-based legacy continuation decision across immediate, delayed, and recovered dispatch. Bind bounded repairs to their source run and episode. - Remove title and description wording from work-mode authority. Agents can still write requested plans in execution mode. - Persist separate failure-retry and productive-continuation counters. Disposition repair and resource waits cannot consume or reset those allowances. - Validate delayed repair identity, then recheck current gates before provider dispatch. Fence native startup cancellation. - Show **Agent needs attention**, a plain-language explanation, **Retry agent**, and expandable details in both task interfaces. Report request progress, acknowledgement, and errors inline. - Store typed recovery notice metadata. Recognize older active notices only through exact stored action and run IDs. Notice text never grants retry authority. - Use the existing recovery-action endpoint for retry. Recheck current action, status, owner, agent availability, dependencies, active runs, pending questions and confirmations, approvals, pause controls, and budget. Duplicate requests do not wake twice. - Add component, page, route, database, contract, and Storybook coverage. Keep the scenario inventory and executable evals here. Historical reports and snapshots live in the [commit-pinned paperclip-evals archive](https://github.com/paperclipai/paperclip-evals/blob/ce3e5afcd4a1184650f586a2b5b8be5874c66c8b/experiments/2026-09-lifecycle-authority/README.md). - Preserve unsaved project fields while the same project URL changes to its canonical alias. Do not reuse data across projects or companies. This separate fix addresses the repeated repository-editor browser failure without changing the browser test. - Keep the development service worker from intercepting Vite module reloads. Update the connection-intent browser fixture to record progress and completion through the agent API. ## Verification Merge preparation on September 25, commit `c1e8e4b7ddd9fbc4913ed55ce21b8e12906c2f97`: - Merged master `bd2030932` and resolved the browser test-list conflict by keeping both sets of regressions. - Deterministic lifecycle baseline: 1,090/1,090 assertions passed; no failures, skips, or missing selected evidence. Unit 423, runner 184, database integration 397, grading 86. - Browser support: 17/17 passed. The offline recovery test first failed with an unstamped development worker, then passed with the production stamp. Its assertions are unchanged. - Focused interaction UI and offline fallback tests: 19/19 passed. Verified the custom-answer composer in Storybook: no attachment controls or hint; entering an answer enables Next. - Recursive typecheck, production build, token gates, and diff checks passed. The worktree is clean. No new real-provider campaign was run. - Current CI and review: [Current PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36166011243): 55 successful checks and two optional Storybook skips. Greptile scored this exact commit 5/5. Hiding the question attachment controls is an intentional UI change; paste/drop remains available. Earlier recovery UI verification, commit `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: - Recursive typecheck, production build, token gates, and diff checks passed. - Focused UI coverage: 338 tests passed across six suites (336 before the interaction guard, with the two affected suites rerun at 149 passed after it). Covers both task interfaces, the real page mutation, pending/error acknowledgement, stale state, and unavailable controls. - Recovery database integration: 352 tests passed before the interaction guard. The complete recovery-action and mutation-route suites passed 181 tests after it. The two new pending question/confirmation regressions failed before the fix and passed afterward, including resolved-interaction controls. Shared validator suite: 31 passed. E2E catalog suites: 34 passed. - Browser inspection passed for light/dark themes, mobile layout, expandable details, pending retry, acknowledgement, failure, and disabled retry. Storybook renders the production component; its request is simulated. - The broad local run hit two chat callback-order wait failures and was stopped after all CI unit/database/runner shards passed. Both local failures passed when rerun without the competing full-suite process. - CI exposed a repeated project-repository draft-loss race during canonical redirects. A new unit regression failed before the fix; all nine project-page tests now pass, including controls for other projects and companies. Both unchanged repository browser tests passed against a fresh local server. UI typecheck, production UI build, and token gates passed after this fix. - [Earlier PR CI passed](https://github.com/paperclipai/paperclip/actions/runs/36072486798) on `21be0fec0e90e86b6d662b8ee4831847cd041cdb`: 55 successful checks, two optional Storybook skips, and no failed or pending checks. The repository browser shard passed with the production fix. Greptile is 5/5 on this exact commit with no unresolved review threads. The PR is mergeable. Historical, source-qualified lifecycle evidence: - Lifecycle baseline: 1,074 assertions. Native session coverage: 447 tests. Product E2E support: 515 tests. Browser support: 11 tests. Full earlier verification is retained in the archive. - [Real-provider campaign: 8/8 passed, zero retries](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35881382080-1/index.html), source `e88d210417280140b44a36449027290adcb1aeaa`. Evidence and cleanup checks passed. This includes deliberately exhausted repair cases that correctly remain blocked; it does not mean every task finished Done. This campaign predates the recovery UI change. - Archive migration verified all 16 original JSON files byte-for-byte and all 24 checksum entries. App tests do not need private archive access. [Archive PR #27](https://github.com/paperclipai/paperclip-evals/pull/27) is merged. ## Risks - Agents that omit durable disposition receive at most two repair attempts by default. Prose-only completion exposes missing state rather than silently changing scheduling. - A retry is an explicit board action. The server rechecks current controls. A successful response confirms the task returned to To do; it does not claim that the provider has already started. - Existing notice metadata remains valid. Only older active notices with matching structured evidence receive the new UI. Historical notices without that evidence keep their existing rendering. No schema migration is required. - Old run records require conservative retry accounting. Tests cover old counters, alternating retry lanes, restarts, and exhausted repairs. - Historical snapshots require private `paperclip-evals` access. The app index retains public campaign links. Live campaigns qualify specific sources and scenarios; no new real-provider campaign has run for the recovery UI commit. > This fixes existing lifecycle and recovery behavior and does not duplicate planned core work. ## Model Used OpenAI GPT-6 through Codex assisted implementation, reasoning, code execution, and review. The exact serving model ID and context window are not exposed in this task. Historical real-provider evaluations used Codex model `gpt-5.6-sol`, separately from the implementation assistant. ## Checklist - [x] I have included a thinking path that traces from project context to this change - [x] I have specified the model used (with version and capability details) - [x] I have checked ROADMAP.md and confirmed this PR does not duplicate planned core work - [x] I have searched GitHub for duplicate or related PRs and linked them above - [x] I have either (a) linked existing issues with `Fixes: #` / `Closes #` / `Refs #` OR (b) described the issue in-PR following the relevant issue template - [x] I have not referenced internal/instance-local Paperclip issues or links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip` URLs) - [x] My branch name describes the change and contains no internal Paperclip ticket id or instance-derived details - [x] I have run tests locally and they pass - [x] I have added or updated tests where applicable - [x] I have updated relevant documentation to reflect my changes - [x] I have considered and documented any risks above - [x] All Paperclip CI gates are green - [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups - [x] I will address all Greptile and reviewer comments before requesting merge --------- Co-authored-by: Paperclip <noreply@paperclip.ing>
Lifecycle behavior baseline
This suite establishes a measurement before removing narrative/regex authority from lifecycle decisions. The original baseline changes no production policy. Assertions describe intended behavior; observed failures are retained rather than blessed as expected outcomes. Subsequent fixes and fresh measurements are recorded separately.
Recorded results moved to paperclip-evals
The 16 saved result JSON files and seven dated measurement reports now live in
the lifecycle authority archive
in the private paperclip-evals repository. Links below pin the archive commit.
The migration manifest
records original paths and checksums. JSON measurements are unchanged, including
failed and partial attempts; report edits only repair links to application files.
Executable tests, fixtures, graders, run commands, and the scenario inventory
remain here. These tests do not require the archive or access to the private
repository. New runs still write ignored local output under .lifecycle-baseline/;
archive retained measurements in paperclip-evals, with their source revisions
and coverage, instead of committing result snapshots to the app repository.
| Measurement | Archived report (private) | Published Product E2E report |
|---|---|---|
| September 21 — deterministic baseline | Report | Deterministic tests; run locally below |
| September 21 — initial live baseline | Report | Campaign 35672810261 |
| September 21 — cancellation and fixture fixes | Report | Campaign 35680906634 |
| September 22 — continuation authority | Report | Campaign 35747200170 |
| September 22 — explicit work mode | Report | Campaign 35806360797 |
| September 22 — accounting baseline | Report | Campaign 35813099816 |
| September 23 — accounting fixes and PR verification | Report | Campaign 35881382080 |
The public reports remain available without private-repository access. Each campaign measures its recorded source and selected cells; this index does not combine them into one score or qualify later revisions. Large logs, traces, videos, and browser reports stay in existing campaign artifact storage.
Run and inspect
From an installed Paperclip checkout:
pnpm test:lifecycle-baseline --list
pnpm test:lifecycle-baseline unit
pnpm test:lifecycle-baseline runner
pnpm test:lifecycle-baseline integration
pnpm test:lifecycle-baseline grading
pnpm test:lifecycle-baseline
pnpm test:lifecycle-baseline:support
pnpm exec tsc -p tests/lifecycle-baseline/tsconfig.json
All four lanes are credential-free. The integration lane uses disposable embedded Postgres and scripted providers. No command above invokes a model, starts a paid campaign, or changes an existing Paperclip instance. Tests are outside default server/workspace discovery; Product E2E matcher calibration remains in its normal opt-in support suite. The baseline command returns nonzero on failed assertions, missing evidence, or unavailable selected coverage. Ordinary CI is unaffected.
Each invocation retains an independent directory under .lifecycle-baseline/:
baseline.mdandbaseline.json: scenario inventory joined to actual assertions;- per-lane Vitest JSON: complete assertion results, durations, and failures;
- observation JSONL: fixture inputs/classifications and actual authority effects, including both sides of narrative pairs and persisted heartbeat outcomes;
source.diff: tracked implementation delta from the recorded commit.
The commit plus fingerprint covers tracked differences and untracked authored
files. Keep the worktree or commit its tests with retained measurements when
comparing revisions. The scripted report marks live coverage not_run; use the separate live Actions
record for provider measurements. Authored definitions and passing grader
calibration are not proof of real provider behavior. A test
suite that cannot load is an evidence/harness failure, not a product finding.
Skips are unavailable coverage, never a pass. Failed assertions require triage;
raw runner/provider text is not uploaded by this command.
Scenario inventory
inventory.mjs is the executable mapping. References to existing suites reuse
their actual assertions rather than duplicating test bodies or counting catalog
entries as executed tests. The report records which matching assertions ran.
| ID | Scenario | Contract |
|---|---|---|
| LCA-01 | Ordinary completion | Valid completion, delivered response, no extra execution |
| LCA-02 | Productive multiple turns | Continue appropriate work without repeating completed operations |
| LCA-03 | Human question | Durable question, matching response, one causal continuation |
| LCA-04 | Approval/decline | Correct actor and approval; other admission gates still apply |
| LCA-05 | Planning/revision | Work mode and exact accepted revision authorize execution |
| LCA-06 | Dependencies | Satisfied dependency condition wakes the parent once |
| LCA-07 | External monitor | Durable eligible wait, one-shot wake, bounded expiry/retry |
| LCA-08 | Conversation | Deliver the reply without forcing task completion or repair loops |
| LCA-09 | Missing disposition | Bounded explicit repair; comments/restarts do not reset attempts |
| LCA-10 | Stop/pause/budget | Preserve distinct semantics; no unauthorized continuation |
| LCA-11 | Terminal ordering | Completion report, provider terminal, and cleanup remain distinct |
| LCA-12 | Replay/restart/ownership | Preserve receipts and budgets; fence stale finalizers |
| LCA-13 | Review | Concrete reviewer/owner and authoritative review outcome |
Troublesome combinations included in the inventory and reused suites:
- Completion followed by failure/cancellation or stream closure without terminal.
- Approval while paused/over budget; response arriving during input handoff/cleanup.
- Reassignment/closure before late finalization; restart between commit and delivery.
- Exhaustion followed by commentary versus an authorized new user request.
- Stale plan revision approval; duplicate question/dependency/reconciler events.
- Wrong company/task or unauthorized resolution; productive versus failure retries.
Narrative pairs and positive controls
authority.test.ts records the legacy classifier for diagnostic comparison and
executes the production structured continuation decision. Assertions compare
scheduling decisions rather than requiring diagnostic labels to be identical.
The persisted legacy authority tests cover replay, restart, exhaustion and
dispatch gates; they also run in the ordinary server suite. Native pairs exercise the actual
status arbiter. The heartbeat cases cross the real persistence/finalization
boundary for both runtimes, including explicit native continuation and the legacy
missing-disposition path. Observations precede cleanup; a second queue/drain pass
checks for additional dispatch. This is not proof about arbitrary future timers:
the existing monitor/retry/reconciler suites separately exercise due-time policy.
Vary summaries/results, existing comment bodies, continuation summaries, stdout, stderr, titles, and descriptions. Variants cover negation, historical quotation, Spanish, optional next steps, unsupported completion claims, encouraging prose, and empty narrative. Commentary-only evidence must not manufacture progress. Keep authority identical for each pair. Positive controls change real status, approval, continuation, budget, or ownership with identical prose. New authenticated user messages are separate causal events, not interchangeable text.
Scripted runner tests validate actual session/event behavior and structured-result contracts. They do not prove OS termination or real provider compliance. Native replacement tests inject verifier evidence at their documented boundary. Route unit tests mock services; database suites test persisted application behavior. These proof boundaries must remain visible when interpreting the baseline.
Live evals
The dedicated live lifecycle suite adds
40 explicitly selected real-LLM/browser cells on both runtime generations.
It was authored after the initial scripted baseline and has now run on GitHub
Actions; see the separate live record for measured outcomes. Use
pnpm test:e2e:runner -- --list --suite lifecycle-baseline to inspect it.
Product E2E uses the existing continuation, agent-chat, and
everyday-workflows catalogs. Continuation now retains an explicit lifecycle
snapshot at each browser checkpoint and grades pending question identity, plan
revision binding, original run receipts, and the absence of execution/recovery
paths after completion. The grader has valid/wrong/missing-evidence calibration.
The continuation definition version advances so measurements are distinguishable.
Validate/discover without spending:
pnpm test:e2e:runner:typecheck
pnpm test:e2e:runner:unit
pnpm test:e2e:runner -- --list --suite continuation
pnpm test:e2e:runner -- --list --suite agent-chat
pnpm test:e2e:runner -- --list --suite everyday-workflows
The sibling paperclip-evals change adds the opt-in
rosters/live-lifecycle-narrative-baseline.json: ordinary finish/block controls, two
misleading-summary variants, and question/review/dependency/wake cases. It does
not expand the maintained paid campaign. Both cases and the roster are validated
without providers; successful, wrong-state, missing-tool, and extra-wake evidence
calibrate their actual grader.
For subsequent live execution use existing explicit selectors and record exact
App/Evals revisions, profile/environment, retries, usage/cost, and artifact IDs.
See doc/evals.md. Do not combine mock-authority Runner Eval scores with Product
E2E scores, or claim full qualification from a partial selection.
September 22 follow-up: legacy continuation implementation and verification, including preserved failed campaigns and the remaining backlog.
The continuation accounting matrix adds ACCT-01 through ACCT-04 for separate allowances, false progress, late gates and restart/replay. Its real-provider companion is the explicit-only continuation-accounting Product E2E suite.
The September 22 measurement records the
enabled failures and preserves both initial and corrected live campaigns.
The September 23 fixes and fresh verification
retain the original measurements and cover separate persisted allowances, delayed
repair promotion and the current native question/response continuation contract.
The inexpensive browser regressions use real Chromium without a provider or Paperclip instance. They check screenshot readiness and development service-worker module revalidation across repeated reloads:
pnpm exec playwright test --config tests/runner-e2e/playwright-support.config.ts
Set PAPERCLIP_PLAYWRIGHT_CHANNEL=chrome to use an installed Chrome browser.