fix(chat): resolve approvals and preserve unanswered questions (#14613)

## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents ask for decisions and optional details through cards in chat.
> - A clear approval in a message can leave the matching card pending.
> - An unanswered question can also block an unrelated later reply.
> - Decisions need a saved source message, while optional questions need
to remain answerable in history.
> - This pull request records conversational decisions and lets users
move on from questions and answer them later.

## Linked Issues or Issue Description

**What happened?**

Native Claude and Codex could act on approval in chat while the original
approval card stayed pending. Pending question forms stayed above the
composer, were absent from history, and could suppress later chat
replies. A late native question answer could wait for a finished run to
reconnect.

**Expected behavior**

The active agent records a clear approval or refusal against the exact
card and user message. Ambiguous replies do not grant consent. Users can
send another message without answering a question. The question remains
pending in history and can be reopened and answered later. The saved
answer reaches the agent.

**Steps to reproduce**

1. Ask an agent to propose work with a confirmation card, then approve
it in chat.
2. Check that the original card records that approval before work
starts.
3. Ask an interactive question, send an unrelated message, and reload.
4. Open the unanswered question from history and submit an answer.

Related work: #14408 added completion delivery. #14607 tests completion
reporting turns. Neither records conversational answers on approval
cards.

## What Changed

- Add a confirmation endpoint backed by a user comment, with schema
validation, OpenAPI discovery, and native Plan-mode access. Ask mode
remains read-only.
- Check company, active run, actor, current session, message provenance,
revision, and resolver policy. Save the decision and audit in one
transaction. Retries do not repeat effects. Emit resolution telemetry
after commit.
- Give fresh and resumed chat turns the actual pending confirmation
identities. Teach agents to save clear conversational decisions before
acting and to clarify ambiguity.
- Keep unanswered Agent Chat questions as compact history entries. A
newer user message closes the old form. Question cards never contribute
to composer pending counts or navigation, including after dismissing a
fresh form. The history card is the sole reminder; clicking it restores
that exact form and draft.
- Preserve Agent Chat questions when later messages or questions arrive.
Historical ordinary inputs no longer gate later chat replies.
Current-run requests, task execution, and governed approvals keep their
gates. Remove the special acknowledgement-publication proof helpers that
this rule replaces.
- Route answers to finished native runs through durable fresh-wake
delivery, with existing idempotency and source-question context. Settle
late replies against contiguous completed conversation turns and freeze
their history replay; failed, unhandled, and newly arriving messages
remain actionable.
- Add real-component Storybook scenarios, database and UI regressions,
and a three-turn native Claude/Codex E2E case. Capture distinct,
UI-ready screenshots and report the individual assertions.

## Verification

- Focused decision/publication/UI regressions after merging master: 288
passed; subsequent UI draft, failed-send, and conversation checks: 199
passed.
- Native question and durable delivery regressions: 106 passed,
including all four terminal run states and exactly-once late delivery.
Seven targeted regressions fail against the original implementation and
pass with the fix.
- Latest conversation/decision/native-delivery regressions after the
master merge: 121 passed. Covers completed progress, missing or failed
intervening turns, new messages during a late reply, stale sessions, and
frozen retry/replay boundaries. Four new assertions fail before the
ordering fix.
- E2E support suite after the master merge: 792 passed. Negative
controls reject expired cards, wrong questions/answers, stale or missing
replies, unrelated clarification forms, and unexpected tasks.
- The embedded-browser walkthrough caught one additional defect:
dismissing a fresh question still showed a composer badge. Both Cancel
and close-button regressions failed before the fix. The fix at
`65f2ade12` passes 170 chat-thread tests and 792 E2E support tests.
After merging master, 232 chat-thread/confirmation tests, server/UI
typechecks, and token gates pass. The preview and two-provider live E2E
pass at `e5512a206`; Greptile is 5/5 with zero unresolved threads at
that commit. All 55 checks are now successful at `e5512a206` (four
conditional checks skipped), including the aggregate verification gate
and clean-install canary test. The first attempt was interrupted by
simultaneous CI worker shutdowns; one failed-job rerun passed without
code changes.
- [Published
Storybook](https://d1p6rlowie26tp.cloudfront.net/storybook/branches/codex~2Fchat-approval-resolution/?path=/story/chat-comments-agent-chat-unanswered-questions--moved-on):
nine real-component scenarios. Manually exercised move on, reopen,
preserve draft, answer later, answer one of multiple questions, and a
custom mobile answer in the embedded browser. Retested fresh Cancel and
close-button dismissal in the updated build, then reopened and submitted
the preserved Green selection and inspected its answered receipt. Static
preview has no live model/backend; its callbacks are fixture responses.
- [First live
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36714504406-1/)
reproduced the late-answer completion-state defect on both providers
despite correct saved answers and acknowledgements. It also exposed a
valid imperative clarification rejected by the old oracle. Both issues
are fixed with regression controls; this failing run is retained as
evidence.
- [Four-cell
qualification](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36717804064-1/)
passed 4/4 at `2bf8a1009`: unanswered-question return and ambiguous
confirmation, each on native Claude and Codex. Inspected saved state,
source-message decisions, visible cards, and agent replies. Both
late-answer chats settled to waiting; no unrequested tasks were created.
[Final branch
rerun](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36719666238-1/)
passed 2/2 at `142630720`: the same unanswered-question journey after
merging master, plus an additional screenshot and browser assertion for
the actual late-answer acknowledgement.
- [Composer-reminder
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36727006818-1/)
passed 2/2 at `5b62c52d9`: native Claude and Codex, three turns each,
with explicit no-badge assertions before and after reload. Inspected
saved pending/answered state, both screenshots with a clear composer,
and actual Blue acknowledgements; all five behavioral matchers passed
per provider and neither created tasks. Cost coverage is partial; this
is bounded workflow qualification.
- [Fresh-dismissal
E2E](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36742773318-1/)
passed 2/2 at `e5512a206`: native Claude and Codex, including fresh
Cancel, clear composer, reopen, unrelated message, reload, late Blue
answer, and actual agent acknowledgement. All five behavioral matchers
pass per provider. Inspected the fresh-dismissal screenshots and saved
pending/answered identity; neither created tasks. Cost coverage is
partial (4/6 runs).
- Prior evidence remains available in [the earlier
campaign](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36642252725-1/).
Its early loading screenshot and overwritten final capture prompted the
UI-ready, distinct screenshot fixes.

## Risks

- The model interprets intent. The server verifies permission and
provenance; it does not infer consent from text. Ambiguous and unrelated
replies are not approvals.
- Historical questions can accumulate. They remain visible, pending, and
answerable; no automatic answer or expiry is invented.
- The change to completion gates is scoped to Agent Chat and ordinary
historical inputs. Current-turn and governed approvals retain their
existing controls.
- Live qualification is limited to the selected stories. Broader native
onboarding finalization remains separate work.
- No database migration. Telemetry adds no fields or values; the
contract and README document the commit boundary. Privacy review was
requested on the PR.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, repository tools, code
execution, and browser-test orchestration. The exact model ID and
context-window size are not exposed to this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
This commit is contained in:
Dotta
2026-09-30 11:46:46 -05:00
committed by GitHub
co-authored by Paperclip
parent 0e5830887b
commit 3c561642b4
59 changed files with 1897 additions and 181 deletions
+2
View File
@@ -108,3 +108,5 @@ tokenize motion. Principles — reasoning only; values live in `ui/src/index.css
- **Reduced motion is honored at the token layer.** A `prefers-reduced-motion: reduce`
block collapses the duration/stagger tokens to zero, cascading to every scoped token,
in addition to each animation's own component-level guard.
Agent Chat keeps pending questions as compact “Unanswered question” entries at their original position in history. A newer user message dismisses the old question form without resolving it. Opening the history entry restores the original form and its draft; submitting later uses the same durable question response path. Actual permission reviews retain their permission checks.
+5
View File
@@ -366,3 +366,8 @@ Copilot and Pi ACP profiles on local and Daytona. See the
[fixture admission, credentials and budget contract](../tests/runner-e2e/README.md#extended-acp-harnesses-explicit-only).
The private Runner Evals campaign of the same name provides complementary
semantic protocol cases; catalog membership is not live qualification.
The explicit-only Product E2E `confirmation-replies` suite tests conversational
approval and rejection, persisted message provenance, approval before execution,
ambiguous proposals, and the existing card-click path with native Claude/Codex.
See the [suite contract](../tests/runner-e2e/README.md#conversational-confirmation-replies-explicit-only).
@@ -8,9 +8,9 @@ The skill/reference inventory and eval cases are the only normative behavior sou
## Baseline Counts
- Skill/reference headings: 155
- Skill/reference headings: 157
- Eval cases: 106 across 16 groups
- Total normative rows: 261
- Total normative rows: 263
- Legacy MCP aliases folded into normative rows: 42
| Eval group | Cases |
@@ -71,6 +71,7 @@ The skill/reference inventory and eval cases are the only normative behavior sou
| skill:skills/paperclip/SKILL.md:key-endpoints-hot-routes:679 | optional_agent_tool | skills/paperclip/SKILL.md:679 |
| skill:skills/paperclip/SKILL.md:searching-issues:708 | optional_agent_tool | skills/paperclip/SKILL.md:708 |
| skill:skills/paperclip/SKILL.md:full-reference:718 | optional_agent_tool | skills/paperclip/SKILL.md:718 |
| skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:722 | always_agent_tool | skills/paperclip/SKILL.md:722 |
| skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1 | always_agent_tool | skills/paperclip/references/artifacts.md:1 |
| skill:skills/paperclip/references/artifacts.md:workspace-only-file-references:15 | optional_agent_tool | skills/paperclip/references/artifacts.md:15 |
| skill:skills/paperclip/references/cases.md:cases:1 | optional_agent_tool | skills/paperclip/references/cases.md:1 |
@@ -180,22 +181,23 @@ The skill/reference inventory and eval cases are the only normative behavior sou
| skill:skills/paperclip/references/api-reference.md:ceo-strategy-approval:912 | optional_agent_tool | skills/paperclip/references/api-reference.md:912 |
| skill:skills/paperclip/references/api-reference.md:questions-and-waiting-for-human-input:921 | always_agent_tool | skills/paperclip/references/api-reference.md:921 |
| skill:skills/paperclip/references/api-reference.md:issue-thread-confirmations:1028 | always_agent_tool | skills/paperclip/references/api-reference.md:1028 |
| skill:skills/paperclip/references/api-reference.md:checkbox-confirmations:1086 | always_agent_tool | skills/paperclip/references/api-reference.md:1086 |
| skill:skills/paperclip/references/api-reference.md:item-verdict-requests:1201 | optional_agent_tool | skills/paperclip/references/api-reference.md:1201 |
| skill:skills/paperclip/references/api-reference.md:checking-approval-status:1311 | optional_agent_tool | skills/paperclip/references/api-reference.md:1311 |
| skill:skills/paperclip/references/api-reference.md:approval-follow-up-requesting-agent:1317 | always_agent_tool | skills/paperclip/references/api-reference.md:1317 |
| skill:skills/paperclip/references/api-reference.md:issue-lifecycle:1335 | always_agent_tool | skills/paperclip/references/api-reference.md:1335 |
| skill:skills/paperclip/references/api-reference.md:error-handling:1365 | control_plane_owned | skills/paperclip/references/api-reference.md:1365 |
| skill:skills/paperclip/references/api-reference.md:full-api-reference:1379 | optional_agent_tool | skills/paperclip/references/api-reference.md:1379 |
| skill:skills/paperclip/references/api-reference.md:agents:1381 | optional_agent_tool | skills/paperclip/references/api-reference.md:1381 |
| skill:skills/paperclip/references/api-reference.md:issues-tasks:1402 | optional_agent_tool | skills/paperclip/references/api-reference.md:1402 |
| skill:skills/paperclip/references/api-reference.md:companies-projects-goals:1442 | optional_agent_tool | skills/paperclip/references/api-reference.md:1442 |
| skill:skills/paperclip/references/api-reference.md:routines:1466 | optional_agent_tool | skills/paperclip/references/api-reference.md:1466 |
| skill:skills/paperclip/references/api-reference.md:approvals-costs-activity-dashboard:1482 | optional_agent_tool | skills/paperclip/references/api-reference.md:1482 |
| skill:skills/paperclip/references/api-reference.md:secrets:1504 | optional_agent_tool | skills/paperclip/references/api-reference.md:1504 |
| skill:skills/paperclip/references/api-reference.md:agent-secret-proposals:1517 | optional_agent_tool | skills/paperclip/references/api-reference.md:1517 |
| skill:skills/paperclip/references/api-reference.md:agent-secret-access:1617 | optional_agent_tool | skills/paperclip/references/api-reference.md:1617 |
| skill:skills/paperclip/references/api-reference.md:common-mistakes:1657 | optional_agent_tool | skills/paperclip/references/api-reference.md:1657 |
| skill:skills/paperclip/references/api-reference.md:conversational-confirmation-answers:1086 | always_agent_tool | skills/paperclip/references/api-reference.md:1086 |
| skill:skills/paperclip/references/api-reference.md:checkbox-confirmations:1092 | always_agent_tool | skills/paperclip/references/api-reference.md:1092 |
| skill:skills/paperclip/references/api-reference.md:item-verdict-requests:1207 | optional_agent_tool | skills/paperclip/references/api-reference.md:1207 |
| skill:skills/paperclip/references/api-reference.md:checking-approval-status:1317 | optional_agent_tool | skills/paperclip/references/api-reference.md:1317 |
| skill:skills/paperclip/references/api-reference.md:approval-follow-up-requesting-agent:1323 | always_agent_tool | skills/paperclip/references/api-reference.md:1323 |
| skill:skills/paperclip/references/api-reference.md:issue-lifecycle:1341 | always_agent_tool | skills/paperclip/references/api-reference.md:1341 |
| skill:skills/paperclip/references/api-reference.md:error-handling:1371 | control_plane_owned | skills/paperclip/references/api-reference.md:1371 |
| skill:skills/paperclip/references/api-reference.md:full-api-reference:1385 | optional_agent_tool | skills/paperclip/references/api-reference.md:1385 |
| skill:skills/paperclip/references/api-reference.md:agents:1387 | optional_agent_tool | skills/paperclip/references/api-reference.md:1387 |
| skill:skills/paperclip/references/api-reference.md:issues-tasks:1408 | optional_agent_tool | skills/paperclip/references/api-reference.md:1408 |
| skill:skills/paperclip/references/api-reference.md:companies-projects-goals:1449 | optional_agent_tool | skills/paperclip/references/api-reference.md:1449 |
| skill:skills/paperclip/references/api-reference.md:routines:1473 | optional_agent_tool | skills/paperclip/references/api-reference.md:1473 |
| skill:skills/paperclip/references/api-reference.md:approvals-costs-activity-dashboard:1489 | optional_agent_tool | skills/paperclip/references/api-reference.md:1489 |
| skill:skills/paperclip/references/api-reference.md:secrets:1511 | optional_agent_tool | skills/paperclip/references/api-reference.md:1511 |
| skill:skills/paperclip/references/api-reference.md:agent-secret-proposals:1524 | optional_agent_tool | skills/paperclip/references/api-reference.md:1524 |
| skill:skills/paperclip/references/api-reference.md:agent-secret-access:1624 | optional_agent_tool | skills/paperclip/references/api-reference.md:1624 |
| skill:skills/paperclip/references/api-reference.md:common-mistakes:1664 | optional_agent_tool | skills/paperclip/references/api-reference.md:1664 |
## Legacy MCP Alias Index
@@ -271,6 +271,15 @@
"semanticOperation": "runtime_reconciliation",
"expectedMockState": "runtime_decision_record"
},
{
"id": "skill:skills/paperclip/SKILL.md:722",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/SKILL.md#L722:conversational-confirmation-answers",
"heading": "Conversational confirmation answers",
"primaryDisposition": "always_agent_tool",
"semanticOperation": "call_api",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1",
"kind": "skill_heading",
@@ -751,151 +760,160 @@
{
"id": "skill:skills/paperclip/references/api-reference.md:1086",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1086:checkbox-confirmations",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1086:conversational-confirmation-answers",
"heading": "Conversational confirmation answers",
"primaryDisposition": "always_agent_tool",
"semanticOperation": "call_api",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1092",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1092:checkbox-confirmations",
"heading": "Checkbox confirmations",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1201",
"id": "skill:skills/paperclip/references/api-reference.md:1207",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1201:item-verdict-requests",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1207:item-verdict-requests",
"heading": "Item verdict requests",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1311",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1311:checking-approval-status",
"heading": "Checking approval status",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1317",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1317:approval-follow-up-requesting-agent",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1317:checking-approval-status",
"heading": "Checking approval status",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1323",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1323:approval-follow-up-requesting-agent",
"heading": "Approval follow-up (requesting agent)",
"primaryDisposition": "control_plane_owned",
"semanticOperation": "runtime_reconciliation",
"expectedMockState": "runtime_decision_record"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1335",
"id": "skill:skills/paperclip/references/api-reference.md:1341",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1335:issue-lifecycle",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1341:issue-lifecycle",
"heading": "Issue Lifecycle",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1365",
"id": "skill:skills/paperclip/references/api-reference.md:1371",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1365:error-handling",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1371:error-handling",
"heading": "Error Handling",
"primaryDisposition": "control_plane_owned",
"semanticOperation": "runtime_reconciliation",
"expectedMockState": "runtime_decision_record"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1379",
"id": "skill:skills/paperclip/references/api-reference.md:1385",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1379:full-api-reference",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1385:full-api-reference",
"heading": "Full API Reference",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1381",
"id": "skill:skills/paperclip/references/api-reference.md:1387",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1381:agents",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1387:agents",
"heading": "Agents",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1402",
"id": "skill:skills/paperclip/references/api-reference.md:1408",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1402:issues-tasks",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1408:issues-tasks",
"heading": "Issues (Tasks)",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1442",
"id": "skill:skills/paperclip/references/api-reference.md:1449",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1442:companies-projects-goals",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1449:companies-projects-goals",
"heading": "Companies, Projects, Goals",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1466",
"id": "skill:skills/paperclip/references/api-reference.md:1473",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1466:routines",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1473:routines",
"heading": "Routines",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1482",
"id": "skill:skills/paperclip/references/api-reference.md:1489",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1482:approvals-costs-activity-dashboard",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1489:approvals-costs-activity-dashboard",
"heading": "Approvals, Costs, Activity, Dashboard",
"primaryDisposition": "control_plane_owned",
"semanticOperation": "runtime_reconciliation",
"expectedMockState": "runtime_decision_record"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1504",
"id": "skill:skills/paperclip/references/api-reference.md:1511",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1504:secrets",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1511:secrets",
"heading": "Secrets",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1517",
"id": "skill:skills/paperclip/references/api-reference.md:1524",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1517:agent-secret-proposals",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1524:agent-secret-proposals",
"heading": "Agent secret proposals",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1570",
"id": "skill:skills/paperclip/references/api-reference.md:1577",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1570:re-bind-an-existing-secret-under-a-new-path-no-secret-id",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1577:re-bind-an-existing-secret-under-a-new-path-no-secret-id",
"heading": "Re-bind an existing secret under a new path (no secret ID)",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1617",
"id": "skill:skills/paperclip/references/api-reference.md:1624",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1617:agent-secret-access",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1624:agent-secret-access",
"heading": "Agent secret access",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
"expectedMockState": "operation_result"
},
{
"id": "skill:skills/paperclip/references/api-reference.md:1657",
"id": "skill:skills/paperclip/references/api-reference.md:1664",
"kind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1657:common-mistakes",
"sourceAnchor": "skills/paperclip/references/api-reference.md#L1664:common-mistakes",
"heading": "Common Mistakes",
"primaryDisposition": "optional_agent_tool",
"semanticOperation": "scoped_discovery",
@@ -2,9 +2,9 @@
Generated by `scripts/generate-capability-contract.mjs`; do not edit generated files.
- Skill/reference headings: 156
- Skill/reference headings: 158
- Legacy MCP tools: 42
- Eval cases: 106 across 16 groups
- Deterministic content SHA-256: `fa6e347b59c519dce414714c4c99f8742b30ca92b3e12df042e4c969ffc481de`
- Deterministic content SHA-256: `60255e688c2b172a02979b4bd305ec089c9ad93b6d4eb4ffae59051a60170d7e`
Every row has exactly one primary disposition, a source anchor, a semantic operation, and a mock-state expectation.
File diff suppressed because one or more lines are too long
@@ -24,7 +24,7 @@
"prpVersion": 1,
"nativeExecutionVersion": 1,
"catalogVersion": 1,
"catalogSha256": "sha256:61c6ba99926c14285ca6a89a05fe316848b9e5bc38840d43dd19b0990dc482d0",
"catalogSha256": "sha256:994cf391ef851c37a4815db38368aa97bd3830def50218b2ffe9428966fc994f",
"driverContractVersion": 1,
"driverKind": "paperclip-deterministic",
"driverVersion": "1.0.0"
@@ -160,7 +160,7 @@
},
{
"path": "fixtures/evals/native-execution-seeded.json",
"sha256": "c9c56cc11b420302be2d0044aeee30302e43da74c076e13a31addfbcaacd0458",
"sha256": "076427778f1ecced927c0680ab3d59e52e06d3e680a610bee3af6fd0795b6a8c",
"expectation": "accept",
"compatibilityCase": "canonical"
},
@@ -42,7 +42,7 @@ function validInventories() {
schemaVersion: 2,
inventoryRole: "normative",
generatedFrom: ["skills/paperclip/SKILL.md"],
rows: Array.from({ length: 155 }, (_, index) => row(`capability-${index}`)),
rows: Array.from({ length: 157 }, (_, index) => row(`capability-${index}`)),
},
evaluations: {
schemaVersion: 2,
@@ -25,6 +25,7 @@ function sourceAnchor(path, line, heading) {
function classifyHeading(path, heading) {
const normalized = heading.toLowerCase();
if (normalized === "conversational confirmation answers") return "always_agent_tool";
if (/(authentication|identity|checkout|budget|error|wake|heartbeat|approval follow-up|activity|audit|release|terminology)/.test(normalized)) {
return "control_plane_owned";
}
@@ -39,6 +40,7 @@ function classifyHeading(path, heading) {
function semanticOperation(disposition, heading) {
const normalized = heading.toLowerCase();
if (normalized === "conversational confirmation answers") return "call_api";
if (disposition === "control_plane_owned") return "runtime_reconciliation";
if (normalized.includes("document") || normalized.includes("plan")) return "write_document";
if (normalized.includes("comment") || normalized.includes("report")) return "report_progress";
@@ -247,7 +247,7 @@ export async function buildMcpInventory(repoRoot) {
export function validateInventories(inventories) {
const errors = [];
const expectedCounts = { capabilities: 155, evaluations: 106, legacyMcpAliases: 42 };
const expectedCounts = { capabilities: 157, evaluations: 106, legacyMcpAliases: 42 };
const normativeNames = ["capabilities", "evaluations"];
const normativeRows = new Map();
const globalNormativeIds = new Set();
@@ -463,6 +463,21 @@
"skill:skills/paperclip/SKILL.md:718"
]
},
{
"id": "skill:skills/paperclip/SKILL.md:conversational-confirmation-answers:722",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/SKILL.md:722",
"title": "Conversational confirmation answers",
"expectedSemantics": "Skill guidance headed “Conversational confirmation answers”.",
"primaryDisposition": "always_agent_tool",
"requiredGrants": [],
"assertionClasses": [
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/SKILL.md:722"
]
},
{
"id": "skill:skills/paperclip/references/artifacts.md:generated-artifacts-and-work-products:1",
"sourceKind": "skill_heading",
@@ -2099,11 +2114,11 @@
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:checkbox-confirmations:1086",
"id": "skill:skills/paperclip/references/api-reference.md:conversational-confirmation-answers:1086",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1086",
"title": "Checkbox confirmations",
"expectedSemantics": "Skill guidance headed “Checkbox confirmations”.",
"title": "Conversational confirmation answers",
"expectedSemantics": "Skill guidance headed “Conversational confirmation answers”.",
"primaryDisposition": "always_agent_tool",
"requiredGrants": [],
"assertionClasses": [
@@ -2114,9 +2129,24 @@
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:item-verdict-requests:1201",
"id": "skill:skills/paperclip/references/api-reference.md:checkbox-confirmations:1092",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1201",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1092",
"title": "Checkbox confirmations",
"expectedSemantics": "Skill guidance headed “Checkbox confirmations”.",
"primaryDisposition": "always_agent_tool",
"requiredGrants": [],
"assertionClasses": [
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1092"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:item-verdict-requests:1207",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1207",
"title": "Item verdict requests",
"expectedSemantics": "Skill guidance headed “Item verdict requests”.",
"primaryDisposition": "optional_agent_tool",
@@ -2125,13 +2155,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1201"
"skill:skills/paperclip/references/api-reference.md:1207"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:checking-approval-status:1311",
"id": "skill:skills/paperclip/references/api-reference.md:checking-approval-status:1317",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1311",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1317",
"title": "Checking approval status",
"expectedSemantics": "Skill guidance headed “Checking approval status”.",
"primaryDisposition": "optional_agent_tool",
@@ -2139,29 +2169,29 @@
"assertionClasses": [
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1311"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:approval-follow-up-requesting-agent:1317",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1317",
"title": "Approval follow-up (requesting agent)",
"expectedSemantics": "Skill guidance headed “Approval follow-up (requesting agent)”.",
"primaryDisposition": "always_agent_tool",
"requiredGrants": [],
"assertionClasses": [
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1317"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:issue-lifecycle:1335",
"id": "skill:skills/paperclip/references/api-reference.md:approval-follow-up-requesting-agent:1323",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1335",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1323",
"title": "Approval follow-up (requesting agent)",
"expectedSemantics": "Skill guidance headed “Approval follow-up (requesting agent)”.",
"primaryDisposition": "always_agent_tool",
"requiredGrants": [],
"assertionClasses": [
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1323"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:issue-lifecycle:1341",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1341",
"title": "Issue Lifecycle",
"expectedSemantics": "Skill guidance headed “Issue Lifecycle”.",
"primaryDisposition": "always_agent_tool",
@@ -2170,13 +2200,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1335"
"skill:skills/paperclip/references/api-reference.md:1341"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:error-handling:1365",
"id": "skill:skills/paperclip/references/api-reference.md:error-handling:1371",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1365",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1371",
"title": "Error Handling",
"expectedSemantics": "Skill guidance headed “Error Handling”.",
"primaryDisposition": "control_plane_owned",
@@ -2185,13 +2215,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1365"
"skill:skills/paperclip/references/api-reference.md:1371"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:full-api-reference:1379",
"id": "skill:skills/paperclip/references/api-reference.md:full-api-reference:1385",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1379",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1385",
"title": "Full API Reference",
"expectedSemantics": "Skill guidance headed “Full API Reference”.",
"primaryDisposition": "optional_agent_tool",
@@ -2200,13 +2230,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1379"
"skill:skills/paperclip/references/api-reference.md:1385"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:agents:1381",
"id": "skill:skills/paperclip/references/api-reference.md:agents:1387",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1381",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1387",
"title": "Agents",
"expectedSemantics": "Skill guidance headed “Agents”.",
"primaryDisposition": "optional_agent_tool",
@@ -2215,13 +2245,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1381"
"skill:skills/paperclip/references/api-reference.md:1387"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:issues-tasks:1402",
"id": "skill:skills/paperclip/references/api-reference.md:issues-tasks:1408",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1402",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1408",
"title": "Issues (Tasks)",
"expectedSemantics": "Skill guidance headed “Issues (Tasks)”.",
"primaryDisposition": "optional_agent_tool",
@@ -2230,13 +2260,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1402"
"skill:skills/paperclip/references/api-reference.md:1408"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:companies-projects-goals:1442",
"id": "skill:skills/paperclip/references/api-reference.md:companies-projects-goals:1449",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1442",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1449",
"title": "Companies, Projects, Goals",
"expectedSemantics": "Skill guidance headed “Companies, Projects, Goals”.",
"primaryDisposition": "optional_agent_tool",
@@ -2245,13 +2275,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1442"
"skill:skills/paperclip/references/api-reference.md:1449"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:routines:1466",
"id": "skill:skills/paperclip/references/api-reference.md:routines:1473",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1466",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1473",
"title": "Routines",
"expectedSemantics": "Skill guidance headed “Routines”.",
"primaryDisposition": "optional_agent_tool",
@@ -2260,13 +2290,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1466"
"skill:skills/paperclip/references/api-reference.md:1473"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:approvals-costs-activity-dashboard:1482",
"id": "skill:skills/paperclip/references/api-reference.md:approvals-costs-activity-dashboard:1489",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1482",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1489",
"title": "Approvals, Costs, Activity, Dashboard",
"expectedSemantics": "Skill guidance headed “Approvals, Costs, Activity, Dashboard”.",
"primaryDisposition": "optional_agent_tool",
@@ -2275,13 +2305,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1482"
"skill:skills/paperclip/references/api-reference.md:1489"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:secrets:1504",
"id": "skill:skills/paperclip/references/api-reference.md:secrets:1511",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1504",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1511",
"title": "Secrets",
"expectedSemantics": "Skill guidance headed “Secrets”.",
"primaryDisposition": "optional_agent_tool",
@@ -2290,13 +2320,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1504"
"skill:skills/paperclip/references/api-reference.md:1511"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:agent-secret-proposals:1517",
"id": "skill:skills/paperclip/references/api-reference.md:agent-secret-proposals:1524",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1517",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1524",
"title": "Agent secret proposals",
"expectedSemantics": "Skill guidance headed “Agent secret proposals”.",
"primaryDisposition": "optional_agent_tool",
@@ -2305,13 +2335,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1517"
"skill:skills/paperclip/references/api-reference.md:1524"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:agent-secret-access:1617",
"id": "skill:skills/paperclip/references/api-reference.md:agent-secret-access:1624",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1617",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1624",
"title": "Agent secret access",
"expectedSemantics": "Skill guidance headed “Agent secret access”.",
"primaryDisposition": "optional_agent_tool",
@@ -2320,13 +2350,13 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1617"
"skill:skills/paperclip/references/api-reference.md:1624"
]
},
{
"id": "skill:skills/paperclip/references/api-reference.md:common-mistakes:1657",
"id": "skill:skills/paperclip/references/api-reference.md:common-mistakes:1664",
"sourceKind": "skill_heading",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1657",
"sourceAnchor": "skills/paperclip/references/api-reference.md:1664",
"title": "Common Mistakes",
"expectedSemantics": "Skill guidance headed “Common Mistakes”.",
"primaryDisposition": "optional_agent_tool",
@@ -2335,7 +2365,7 @@
"control_plane_invariant"
],
"evidenceIds": [
"skill:skills/paperclip/references/api-reference.md:1657"
"skill:skills/paperclip/references/api-reference.md:1664"
]
}
]
@@ -6,9 +6,9 @@ export type CapabilityPrimaryDisposition =
| "optional_agent_tool";
export const capabilityInventoryCounts = {
"skillReferenceCapabilities": 155,
"skillReferenceCapabilities": 157,
"evalCases": 106,
"normativeRows": 261,
"normativeRows": 263,
"legacyMcpAliases": 42
} as const;
@@ -6,6 +6,14 @@ import { PAPERCLIP_CORE_PROTOCOL_ACTIONS } from "./core.js";
describe("core Paperclip protocol action contracts", () => {
const ajv = new Ajv2020({ allErrors: true, allowUnionTypes: true, strict: false });
it("keeps the live input tool's guidance consistent with conversational confirmation recording", () => {
const action = PAPERCLIP_CORE_PROTOCOL_ACTIONS.find(action => action.id === "request_human_input")!;
expect(action.live!.descriptor.description).toBe(action.documentation.description);
expect(action.live!.descriptor.description).toContain("resolve-from-comment");
expect(action.live!.descriptor.description).toContain("resolver permissions");
expect(action.live!.descriptor.description).not.toContain("Never infer answers, answer your own card");
});
it.each(PAPERCLIP_CORE_PROTOCOL_ACTIONS.map((action) => [action.id, action] as const))(
"%s has immutable metadata and schema-valid examples",
(operationId, action) => {
@@ -27,7 +27,7 @@ export const requestHumanInputAction = {
},
"documentation": {
"title": "Request structured human input",
"description": "Create a durable human question or approval card on the current Paperclip task bound to this run; Paperclip renders it and authenticates the response. Use questions with continuationPolicy='wake_assignee' when an answer is needed, including otherwise tool-free chat turns. Supply a stable idempotencyKey and reuse it on retries. For one question at a time, ask only the next unanswered question and wait for its real answer. Never infer answers, answer your own card, or treat clarification as approval. Preserve existing review gates. Call this tool before claiming a question was asked; if creation fails, report the failure. Do not fabricate answer links or Markdown buttons, post duplicate cards, or substitute call_api. Use payload.questions for choices and payload.questionSet for text fields; see the payload schema for formats.",
"description": "Create a durable human question or approval card on the current Paperclip task bound to this run; Paperclip renders it and authenticates the response. Use questions with continuationPolicy='wake_assignee' when an answer is needed, including otherwise tool-free chat turns. Supply a stable idempotencyKey and reuse it on retries. For one question at a time, ask only the next unanswered question and wait for its real answer. Never fabricate answers or treat ambiguous clarification as approval. For an ordinary confirmation or checkbox card, record a clear user chat answer with call_api POST /api/issues/{id}/interactions/{interactionId}/resolve-from-comment, using the source commentId and decision (accept/reject), plus explicit selectedOptionIds for checkbox acceptance. Existing resolver permissions still apply; governed tool, secret, and connection approvals are excluded. Question forms retain their dedicated answer workflow. Preserve existing review gates. Call this tool before claiming a question was asked; if creation fails, report the failure. Do not fabricate answer links or Markdown buttons, post duplicate cards, or use call_api to create the card. Use payload.questions for choices and payload.questionSet for text fields; see the payload schema for formats.",
"note": null
},
"examples": {
@@ -74,7 +74,7 @@ export const requestHumanInputAction = {
"operationId": "request_human_input",
"version": 1,
"title": "Request structured human input",
"description": "Create a durable human question or approval card on the current Paperclip task bound to this run; Paperclip renders it and authenticates the response. Use questions with continuationPolicy='wake_assignee' when an answer is needed, including otherwise tool-free chat turns. Supply a stable idempotencyKey and reuse it on retries. For one question at a time, ask only the next unanswered question and wait for its real answer. Never infer answers, answer your own card, or treat clarification as approval. Preserve existing review gates. Call this tool before claiming a question was asked; if creation fails, report the failure. Do not fabricate answer links or Markdown buttons, post duplicate cards, or substitute call_api. Use payload.questions for choices and payload.questionSet for text fields; see the payload schema for formats.",
"description": "Create a durable human question or approval card on the current Paperclip task bound to this run; Paperclip renders it and authenticates the response. Use questions with continuationPolicy='wake_assignee' when an answer is needed, including otherwise tool-free chat turns. Supply a stable idempotencyKey and reuse it on retries. For one question at a time, ask only the next unanswered question and wait for its real answer. Never fabricate answers or treat ambiguous clarification as approval. For an ordinary confirmation or checkbox card, record a clear user chat answer with call_api POST /api/issues/{id}/interactions/{interactionId}/resolve-from-comment, using the source commentId and decision (accept/reject), plus explicit selectedOptionIds for checkbox acceptance. Existing resolver permissions still apply; governed tool, secret, and connection approvals are excluded. Question forms retain their dedicated answer workflow. Preserve existing review gates. Call this tool before claiming a question was asked; if creation fails, report the failure. Do not fabricate answer links or Markdown buttons, post duplicate cards, or use call_api to create the card. Use payload.questions for choices and payload.questionSet for text fields; see the payload schema for formats.",
"exposure": "always",
"requiredClaims": [],
"allowedModes": [
@@ -2,6 +2,7 @@ import assert from "node:assert/strict";
import { readFile } from "node:fs/promises";
import test from "node:test";
import { resolve } from "node:path";
import { decodeInventory } from "../scripts/lib/capability-inventory.mjs";
import { validateRows } from "../scripts/generate-capability-contract.mjs";
const phaseDirectory = resolve(import.meta.dirname, "../generated/capability");
@@ -17,7 +18,7 @@ test("generated Capability inventory has full source coverage", async () => {
readRows("eval-traceability.yaml"),
]);
assert.equal(capabilities.length, 155);
assert.equal(capabilities.length, 158);
assert.equal(tools.length, 42);
assert.equal(evals.length, 106);
assert.equal(new Set(evals.map((row) => row.group)).size, 16);
@@ -39,3 +40,14 @@ test("contract validation rejects missing, duplicate, and unclassified entries",
assert.throws(() => validateRows([row, { ...row, id: "example:2" }], "fixture"), /duplicate source anchor/);
assert.throws(() => validateRows([{ ...row, primaryDisposition: "unclassified" }], "fixture"), /valid primary disposition/);
});
test("conversational answer guidance has the same agent-operation classification in both inventories", async () => {
const generated = (await readRows("capabilities.yaml")).filter(row => row.heading === "Conversational confirmation answers");
const spec = decodeInventory(await readFile(resolve(import.meta.dirname, "../spec/capability/capabilities.yaml"), "utf8")).rows
.filter(row => row.title === "Conversational confirmation answers");
assert.equal(generated.length, 2);
assert.equal(spec.length, 2);
for (const row of [...generated, ...spec]) assert.equal(row.primaryDisposition, "always_agent_tool");
for (const row of generated) assert.equal(row.semanticOperation, "call_api");
});
+2
View File
@@ -2002,6 +2002,8 @@ export {
requestItemVerdictsResultSchema,
createIssueThreadInteractionSchema,
acceptIssueThreadInteractionSchema,
resolveConfirmationFromCommentSchema,
type ResolveConfirmationFromComment,
rejectIssueThreadInteractionSchema,
cancelIssueThreadInteractionSchema,
skipIssueThreadInteractionSchema,
+6
View File
@@ -78,6 +78,12 @@ in the generated contract. Its `legacy_inherited_restriction` dimension is
restriction. It is `false` for canonical new writes. This dimension describes
policy provenance. It does not contain user content or an identifier.
Emit `interaction.resolved` only after the complete decision transaction commits.
Conversational answers include the card outcome, source-message reference, and
activity audit in that transaction. A rollback or matching retry must not emit
a resolution event. This changes emission timing only: message text, comment IDs,
and user IDs remain in the instance database and are not added to telemetry.
Use `trackInteractionCreated()` and `trackInteractionResolved()` from
`events.ts` to emit these events. The generated contract remains the authority
for their exact dimensions and optionality.
@@ -58,6 +58,7 @@ interaction_kind: ("suggest_tasks" | "ask_user_questions" | "request_confirmatio
used_deprecated_resolver_policy_alias: boolean
}
/** Emit only after the complete resolution transaction commits, including response provenance. */
export interface PaperclipInteractionResolvedDimensions {
interaction_kind: ("suggest_tasks" | "ask_user_questions" | "request_confirmation" | "request_checkbox_confirmation" | "request_item_verdicts" | "other")
status: ("accepted" | "rejected" | "answered" | "cancelled" | "expired" | "failed" | "other")
+2
View File
@@ -481,6 +481,8 @@ export {
requestItemVerdictsResultSchema,
createIssueThreadInteractionSchema,
acceptIssueThreadInteractionSchema,
resolveConfirmationFromCommentSchema,
type ResolveConfirmationFromComment,
rejectIssueThreadInteractionSchema,
cancelIssueThreadInteractionSchema,
skipIssueThreadInteractionSchema,
+17
View File
@@ -2039,6 +2039,23 @@ export type AcceptIssueThreadInteraction = z.infer<
typeof acceptIssueThreadInteractionSchema
>;
/** Records an agent's interpretation of a real user reply without widening resolver permissions. */
export const resolveConfirmationFromCommentSchema = z.object({
commentId: z.string().guid(),
decision: z.enum(["accept", "reject"]),
selectedOptionIds: z.array(z.string().trim().min(1).max(120))
.max(REQUEST_CHECKBOX_CONFIRMATION_OPTION_LIMIT).optional(),
reason: z.string().trim().max(4000).optional(),
}).strict().superRefine((value, ctx) => {
if (value.decision === "reject" && value.selectedOptionIds !== undefined) {
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["selectedOptionIds"], message: "Selections apply only to acceptance" });
}
if (value.selectedOptionIds && new Set(value.selectedOptionIds).size !== value.selectedOptionIds.length) {
ctx.addIssue({ code: z.ZodIssueCode.custom, path: ["selectedOptionIds"], message: "Selections must be unique" });
}
});
export type ResolveConfirmationFromComment = z.infer<typeof resolveConfirmationFromCommentSchema>;
export const rejectIssueThreadInteractionSchema = z.object({
reason: z.string().trim().max(4000).optional(),
});
@@ -892,6 +892,54 @@ const support = await getEmbeddedPostgresTestSupport();
if (!scenario.ordinary) expect(after.conversationSessionGeneration).toBe(scenario.generation);
});
it.each(["answered", "new-message", "response-less", "failed", "stale-session", "failed-gap", "missing-gap"])("settles a historical question using completed conversation progress (%s)", async scenario => {
const chat = await create();
const svc = issueService(db);
const first = await svc.addComment(chat.id, "Which color?", { userId: "local-board" });
const original = await runFor(chat.id, first.id);
const [question] = await db.insert(issueThreadInteractions).values({ companyId, issueId: chat.id,
kind: "ask_user_questions", status: "answered", continuationPolicy: "wake_assignee", sourceRunId: original.id,
createdByAgentId: agentId, resolvedByUserId: "local-board", payload: { version: 1, questions: [] }, result: { version: 1, answers: [] } }).returning();
if (scenario === "failed-gap" || scenario === "missing-gap") {
const skipped = await svc.addComment(chat.id, "This request still needs an answer.", { userId: "local-board" });
if (scenario === "failed-gap") {
const failed = await runFor(chat.id, skipped.id, { conversationSessionGeneration: 0 });
await db.update(heartbeatRuns).set({ status: "failed" }).where(eq(heartbeatRuns.id, failed.id));
}
}
const newer = await svc.addComment(chat.id, "What is the capital of France?", { userId: "local-board" });
const replyRun = await runFor(chat.id, newer.id, { conversationSessionGeneration: 0 });
if (scenario !== "response-less") await svc.addComment(chat.id, "Paris.", { agentId, runId: replyRun.id });
await db.update(heartbeatRuns).set({ status: scenario === "failed" ? "failed" : "succeeded",
...(scenario === "stale-session" ? { contextSnapshot: { ...replyRun.contextSnapshot, conversationSessionGeneration: 99 } } : {}) }).where(eq(heartbeatRuns.id, replyRun.id));
const answerRun = await runFor(chat.id, first.id, { source: "issue.interaction.respond", interactionId: question.id,
interactionKind: "ask_user_questions", interactionStatus: "answered" });
const prepared = await prepareConversationTurn(db, answerRun);
expect(prepared.context.wakeCommentId).toBe(first.id);
expect(prepared.context.conversationReplyBoundaryCommentId).toBe(["answered", "new-message"].includes(scenario) ? newer.id : first.id);
// A retry reuses the frozen boundary, even if another message has arrived.
if (scenario === "new-message") await svc.addComment(chat.id, "New instructions while you answer", { userId: "local-board" });
const retry = await prepareConversationTurn(db, { ...answerRun, contextSnapshot: prepared.context });
expect(retry.context.conversationReplyBoundaryCommentId).toBe(prepared.context.conversationReplyBoundaryCommentId);
await svc.addComment(chat.id, "You chose Blue.", { agentId, runId: answerRun.id });
expect(await settleConversationTurn(db, { ...answerRun, status: "succeeded", contextSnapshot: prepared.context })).toBe(true);
const [settled] = await db.select().from(issues).where(eq(issues.id, chat.id));
expect(settled.conversationState).toBe(scenario === "answered" ? "waiting" : "active");
expect(retry.context.conversationReplayThroughCommentId).toBe(prepared.context.conversationReplayThroughCommentId);
const replay = await conversationReplay(db, companyId, chat.id, first.id,
prepared.context.conversationReplayThroughCommentId as string);
if (["answered", "new-message"].includes(scenario)) {
expect(replay).toContain("What is the capital of France?");
expect(replay).toContain("Paris.");
} else {
expect(replay).not.toContain("What is the capital of France?");
expect(replay).not.toContain("Paris.");
}
expect(replay).not.toContain("This request still needs an answer.");
expect(replay).not.toContain("New instructions while you answer");
expect(replay).not.toContain("You chose Blue.");
});
it("only parks answered turns and preserves idle across recovery classification", async () => {
const issue = await create();
const message = await issueService(db).addComment(
@@ -0,0 +1,260 @@
import { randomUUID } from "node:crypto";
import { eq, sql } from "drizzle-orm";
import { afterAll, beforeAll, describe, expect, it, vi } from "vitest";
import { activityLog, agents, companies, createDb, documents, documentRevisions, heartbeatRuns,
issueComments, issueDocuments, issues, issueThreadInteractions } from "@paperclipai/db";
import { getEmbeddedPostgresTestSupport, startEmbeddedPostgresTestDatabase } from "./helpers/embedded-postgres.js";
import { resolveConfirmationFromComment } from "../services/confirmation-comment-resolution.js";
import { issueThreadInteractionService } from "../services/issue-thread-interactions.js";
import { pendingNativeGovernance } from "../services/native-runtime/native-run-finalizer.js";
import { getConversationConfirmationContext } from "../services/conversation-confirmation-context.js";
const { resolvedTelemetry } = vi.hoisted(() => ({ resolvedTelemetry: vi.fn() }));
vi.mock("@paperclipai/shared/telemetry", async importOriginal => ({
...await importOriginal<Record<string, unknown>>(), trackInteractionResolved: resolvedTelemetry, trackInteractionCreated: vi.fn(),
}));
vi.mock("../telemetry.js", async importOriginal => ({
...await importOriginal<Record<string, unknown>>(), getTelemetryClient: () => ({}),
}));
const support = await getEmbeddedPostgresTestSupport();
(support.supported ? describe : describe.skip)("conversational confirmation resolution", () => {
let temporary: Awaited<ReturnType<typeof startEmbeddedPostgresTestDatabase>>;
let db: ReturnType<typeof createDb>;
beforeAll(async () => { temporary = await startEmbeddedPostgresTestDatabase("confirmation-replies-"); db = createDb(temporary.connectionString); }, 30_000);
afterAll(async () => { await temporary?.cleanup(); });
async function seed(checkbox = false, conversation = false) {
const companyId = randomUUID(), agentId = randomUUID(), issueId = randomUUID(), runId = randomUUID();
await db.insert(companies).values({ id: companyId, name: "Replies", issuePrefix: companyId.slice(0, 8) });
await db.insert(agents).values({ id: agentId, companyId, name: "Planner", adapterType: "paperclip_runner" });
const [issue] = await db.insert(issues).values({ id: issueId, companyId, title: "Proposal", status: "in_progress", assigneeAgentId: agentId,
...(conversation ? { conversationAgentId: agentId, conversationUserId: "operator", conversationState: "active", conversationSessionGeneration: 1 } : {}) }).returning();
await db.insert(heartbeatRuns).values({ id: runId, companyId, agentId, nativeIssueId: issueId, status: "running", runtimeMode: "native", contextSnapshot: { issueId, conversationSessionGeneration: 1 } });
await db.update(issues).set({ executionRunId: runId }).where(eq(issues.id, issueId));
const svc = issueThreadInteractionService(db);
const card = await svc.create(issue, checkbox ? {
kind: "request_checkbox_confirmation", payload: { version: 1, prompt: "Approve selected work?", options: [{ id: "note", label: "Welcome note" }, { id: "poster", label: "Poster" }], minSelected: 1, defaultSelectedOptionIds: ["poster"] },
} : { kind: "request_confirmation", payload: { version: 1, prompt: "Approve the welcome note?" } }, { agentId });
const [comment] = await db.insert(issueComments).values({ companyId, issueId, authorUserId: "operator", authorType: "user", body: "Yes, write the welcome note." }).returning();
const args = { companyId, issueId, interactionId: card.id, actor: { agentId, runId }, input: { commentId: comment.id, decision: "accept" as const, ...(checkbox ? { selectedOptionIds: ["note"] } : {}) } };
return { args, card, issue, comment, svc, companyId, issueId, agentId, runId };
}
const readCard = (id: string) => db.select().from(issueThreadInteractions).where(eq(issueThreadInteractions.id, id)).then(rows => rows[0]!);
const audit = (id: string) => db.select().from(activityLog).where(eq(activityLog.entityId, id)).then(rows => rows.filter(row => row.details?.source === "conversation_reply"));
it.each(["ask_user_questions", "request_confirmation", "request_checkbox_confirmation"])("an older chat %s stays pending without gating a later reply", async kind => {
const f = await seed(false, true);
await db.update(issueThreadInteractions).set({ kind, sourceRunId: null }).where(eq(issueThreadInteractions.id, f.card.id));
expect(await pendingNativeGovernance({ db, ...f, executionState: null })).toBeNull();
expect((await readCard(f.card.id)).status).toBe("pending");
await db.update(issueThreadInteractions).set({ sourceRunId: f.runId }).where(eq(issueThreadInteractions.id, f.card.id));
expect(await pendingNativeGovernance({ db, ...f, executionState: null })).toEqual({ kind: "interaction", id: f.card.id });
});
it.each(["ordinary-task", "human_only", "toolAction", "secretProposal", "connectionAuthorization"])("preserves the existing %s completion gate", async kind => {
const f = await seed(false, kind !== "ordinary-task");
if (kind === "human_only") await db.update(issueThreadInteractions).set({ effectiveResolverPolicy: "human_only" }).where(eq(issueThreadInteractions.id, f.card.id));
else if (kind !== "ordinary-task") await db.update(issueThreadInteractions).set({ payload: { ...f.card.payload, [kind]: {} } }).where(eq(issueThreadInteractions.id, f.card.id));
expect(await pendingNativeGovernance({ db, ...f, executionState: null })).toEqual({ kind: "interaction", id: f.card.id });
expect(await pendingNativeGovernance({ db, ...f, executionState: { status: "pending" } })).toEqual({ kind: "execution_stage", id: f.runId });
expect(await pendingNativeGovernance({ db, ...f, companyId: randomUUID(), executionState: null })).toBeNull();
});
it.each([true, false])("preserves historical questions only in agent chat (conversation=%s)", async conversation => {
const f = await seed(false, conversation);
const question = (prompt: string) => ({ kind: "ask_user_questions" as const, payload: { version: 1 as const,
supersedeOnUserComment: true, questions: [{ id: "color", prompt, selectionMode: "single" as const, required: true,
options: [{ id: "blue", label: "Blue" }, { id: "green", label: "Green" }] }] } });
const first = await f.svc.create(f.issue, question("Which color?"), { agentId: f.agentId });
await f.svc.expireRequestConfirmationsSupersededByComment(f.issue,
{ id: f.comment.id, authorUserId: "operator", createdAt: new Date(Date.now() + 1000) }, { userId: "operator" });
expect((await readCard(first.id)).status).toBe(conversation ? "pending" : "expired");
const second = await f.svc.create(f.issue, question("Which shade?"), { agentId: f.agentId });
await f.svc.create(f.issue, question("Which finish?"), { agentId: f.agentId });
expect((await readCard(second.id)).status).toBe(conversation ? "pending" : "expired");
if (conversation) {
const answered = await f.svc.answerQuestions(f.issue, first.id, { answers: [{ questionId: "color", optionIds: ["blue"] }] }, { userId: "operator" });
expect(answered).toMatchObject({ status: "answered", result: { answers: [{ questionId: "color", optionIds: ["blue"] }] } });
expect((await readCard(second.id)).status).toBe("pending");
await expect(f.svc.answerQuestions(f.issue, first.id, { answers: [{ questionId: "color", optionIds: ["green"] }] }, { userId: "operator" })).rejects.toThrow();
}
});
it("supplies actual pending card identities and explicit choices, then refreshes after resolution", async () => {
const f = await seed(true, true);
const snapshot = await getConversationConfirmationContext({ db, ...f });
expect(snapshot).toMatchObject({ truncated: false, cards: [{ id: f.card.id, status: "pending", resolverPolicy: "anyone",
options: [{ id: "note", label: "Welcome note" }, { id: "poster", label: "Poster" }] }] });
expect(JSON.stringify(snapshot)).not.toContain("defaultSelectedOptionIds");
await resolveConfirmationFromComment(db, f.args);
expect(await getConversationConfirmationContext({ db, ...f })).toEqual({ truncated: false, cards: [] });
});
it.each(["ordinary-task", "other-company", "other-agent"])("does not supply confirmation context for %s", async kind => {
const f = await seed(false, kind !== "ordinary-task");
expect(await getConversationConfirmationContext({ db, ...f,
...(kind === "other-company" ? { companyId: randomUUID() } : {}),
...(kind === "other-agent" ? { agentId: randomUUID() } : {}),
})).toBeNull();
});
it.each(["toolAction", "secretProposal", "connectionAuthorization", "questions", "expired"])("excludes %s from ordinary confirmation context", async kind => {
const f = await seed(false, true);
await db.update(issueThreadInteractions).set(kind === "questions" ? { kind: "ask_user_questions" }
: kind === "expired" ? { status: "expired" }
: { payload: { ...f.card.payload, [kind]: { value: "must-not-leak" } } }).where(eq(issueThreadInteractions.id, f.card.id));
expect(await getConversationConfirmationContext({ db, ...f })).toEqual({ truncated: false, cards: [] });
});
it("excludes prior-session cards and preserves the current card's human-only policy", async () => {
const f = await seed(false, true);
const boundaryAt = new Date(Date.now() + 1000);
const [boundary] = await db.insert(issueComments).values({ companyId: f.companyId, issueId: f.issueId,
authorType: "user", authorUserId: "operator", body: "/new", createdAt: boundaryAt }).returning();
await db.update(issues).set({ conversationBoundaryCommentId: boundary.id }).where(eq(issues.id, f.issueId));
const [current] = await db.insert(issueThreadInteractions).values({ companyId: f.companyId, issueId: f.issueId,
kind: "request_confirmation", effectiveResolverPolicy: "human_only", payload: { version: 1, prompt: "Current proposal" },
createdAt: new Date(boundaryAt.getTime() + 1000) }).returning();
expect(await getConversationConfirmationContext({ db, ...f })).toMatchObject({ cards: [{ id: current.id, resolverPolicy: "human_only" }] });
});
it("bounds card and proposal data and declares truncation instead of hiding it", async () => {
const f = await seed(false, true);
await db.update(issueThreadInteractions).set({ payload: { version: 1, prompt: "x".repeat(2500) } }).where(eq(issueThreadInteractions.id, f.card.id));
await db.insert(issueThreadInteractions).values(Array.from({ length: 13 }, () => ({ companyId: f.companyId, issueId: f.issueId,
kind: "request_confirmation", payload: { version: 1 as const, prompt: "Another proposal" }, createdAt: new Date(Date.now() + 1000) })));
const snapshot = await getConversationConfirmationContext({ db, ...f });
expect(snapshot?.truncated).toBe(true);
expect(snapshot?.cards).toHaveLength(12);
expect(snapshot?.cards[0]?.prompt).toHaveLength(2000);
expect(snapshot?.cards[0]?.promptTruncated).toBe(true);
});
it.each([false, true])("persists acceptance, explicit selection and source-message audit (checkbox=%s)", async checkbox => {
const f = await seed(checkbox, true);
const answer = await resolveConfirmationFromComment(db, f.args);
expect(answer).toMatchObject({ deduplicated: false, interaction: { status: "accepted", resolvedByAgentId: f.agentId, resolvedByRunId: f.runId, resolvedByUserId: null, result: { outcome: "accepted", commentId: f.comment.id, ...(checkbox ? { selectedOptionIds: ["note"] } : {}) } } });
expect(await audit(f.issueId)).toMatchObject([{ action: "issue.thread_interaction_accepted", details: { responseCommentId: f.comment.id, responseUserId: "operator" } }]);
});
it.each(["accept", "reject"] as const)("does not emit a %s resolution when the provenance audit rolls back", async decision => {
const f = await seed();
await db.execute(sql.raw(`CREATE FUNCTION fail_confirmation_audit() RETURNS trigger LANGUAGE plpgsql AS $$ BEGIN IF NEW.entity_id = '${f.issueId}' THEN RAISE EXCEPTION 'audit write failed'; END IF; RETURN NEW; END $$`));
await db.execute(sql.raw("CREATE TRIGGER fail_confirmation_audit BEFORE INSERT ON activity_log FOR EACH ROW EXECUTE FUNCTION fail_confirmation_audit()"));
resolvedTelemetry.mockClear();
try {
await expect(resolveConfirmationFromComment(db, { ...f.args, input: { commentId: f.comment.id, decision } })).rejects.toThrow();
expect((await readCard(f.card.id)).status).toBe("pending");
expect(await audit(f.issueId)).toHaveLength(0);
expect(resolvedTelemetry).not.toHaveBeenCalled();
} finally {
await db.execute(sql.raw("DROP TRIGGER fail_confirmation_audit ON activity_log"));
await db.execute(sql.raw("DROP FUNCTION fail_confirmation_audit()"));
}
await resolveConfirmationFromComment(db, { ...f.args, input: { commentId: f.comment.id, decision } });
expect(resolvedTelemetry).toHaveBeenCalledExactlyOnceWith(expect.anything(), expect.objectContaining({ status: decision === "accept" ? "accepted" : "rejected" }));
await resolveConfirmationFromComment(db, { ...f.args, input: { commentId: f.comment.id, decision } });
expect(resolvedTelemetry).toHaveBeenCalledTimes(1);
});
it("records refusal and its reason on the card", async () => {
const f = await seed();
const answer = await resolveConfirmationFromComment(db, { ...f.args, input: { commentId: f.comment.id, decision: "reject", reason: "Not needed" } });
expect(answer.interaction).toMatchObject({ status: "rejected", result: { outcome: "rejected", reason: "Not needed", commentId: f.comment.id } });
});
it("serializes simultaneous retries into one decision and one audit record", async () => {
const f = await seed(true);
const answers = await Promise.all([resolveConfirmationFromComment(db, f.args), resolveConfirmationFromComment(db, f.args)]);
expect(answers.map(x => x.deduplicated).sort()).toEqual([false, true]);
expect(await audit(f.issueId)).toHaveLength(1);
const before = await readCard(f.card.id);
await resolveConfirmationFromComment(db, f.args);
expect(await readCard(f.card.id)).toEqual(before);
});
it("does not rewrite an opposite decision or a different checkbox selection", async () => {
const f = await seed(true);
await resolveConfirmationFromComment(db, f.args);
await expect(resolveConfirmationFromComment(db, { ...f.args, input: { commentId: f.comment.id, decision: "reject" } })).rejects.toThrow("different decision");
await expect(resolveConfirmationFromComment(db, { ...f.args, input: { ...f.args.input, selectedOptionIds: ["poster"] } })).rejects.toThrow("different decision");
expect((await readCard(f.card.id)).result).toMatchObject({ selectedOptionIds: ["note"] });
});
it.each(["deleted", "agent", "system", "agent-attributed", "quarantined", "before-card", "other-task", "other-company", "newer-message", "wrong-user", "reset-command"])("rejects %s answer evidence without effects", async kind => {
const f = await seed(false, true);
const change: Record<string, unknown> = kind === "deleted" ? { deletedAt: new Date() }
: kind === "agent" ? { authorType: "agent", authorAgentId: f.agentId }
: kind === "system" ? { authorType: "system" }
: kind === "agent-attributed" ? { createdByRunId: f.runId }
: kind === "quarantined" ? { sourceTrust: { preset: "low_trust", disposition: "quarantined" } }
: kind === "before-card" ? { createdAt: new Date(0) }
: kind === "wrong-user" ? { authorUserId: "somebody-else" }
: kind === "reset-command" ? { body: "/new" } : {};
if (kind === "other-task") {
const [otherIssue] = await db.insert(issues).values({ companyId: f.companyId, title: "Other proposal", status: "in_progress" }).returning();
const [otherComment] = await db.insert(issueComments).values({ companyId: f.companyId, issueId: otherIssue.id, authorUserId: "operator", authorType: "user", body: "Yes, go ahead." }).returning();
f.args.input.commentId = otherComment.id;
} else if (kind === "other-company") {
const other = await seed();
f.args.input.commentId = other.comment.id;
} else if (kind === "newer-message") {
await db.insert(issueComments).values({ companyId: f.companyId, issueId: f.issueId, authorUserId: "operator", authorType: "user", body: "Wait, do not proceed.", createdAt: new Date(Date.now() + 1000) });
} else await db.update(issueComments).set(change).where(eq(issueComments.id, f.comment.id));
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow();
expect((await readCard(f.card.id)).status).toBe("pending");
expect(await audit(f.issueId)).toHaveLength(0);
});
it.each(["cancelled", "finished", "lost-owner", "wrong-run-task", "reset-generation", "old-card"])("fences %s execution", async kind => {
const f = await seed(false, true);
if (kind === "cancelled" || kind === "finished") await db.update(heartbeatRuns).set({ status: kind === "cancelled" ? "cancelled" : "succeeded" }).where(eq(heartbeatRuns.id, f.runId));
if (kind === "lost-owner") await db.update(issues).set({ executionRunId: null }).where(eq(issues.id, f.issueId));
if (kind === "wrong-run-task") await db.update(heartbeatRuns).set({ nativeIssueId: null, contextSnapshot: {} }).where(eq(heartbeatRuns.id, f.runId));
if (kind === "reset-generation") await db.update(issues).set({ conversationSessionGeneration: 2 }).where(eq(issues.id, f.issueId));
if (kind === "old-card") {
const [boundary] = await db.insert(issueComments).values({ companyId: f.companyId, issueId: f.issueId, authorType: "user", authorUserId: "operator", body: "/new" }).returning();
await db.update(issues).set({ conversationBoundaryCommentId: boundary.id }).where(eq(issues.id, f.issueId));
}
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow();
expect((await readCard(f.card.id)).status).toBe("pending");
});
it.each(["human_only", "not_creator", "addressee", "toolAction", "secretProposal", "connectionAuthorization", "questions", "closed"])("preserves %s restrictions", async kind => {
const f = await seed();
if (kind === "closed") await db.update(issues).set({ status: "done" }).where(eq(issues.id, f.issueId));
else if (kind === "human_only" || kind === "not_creator") await db.update(issueThreadInteractions).set({ effectiveResolverPolicy: kind }).where(eq(issueThreadInteractions.id, f.card.id));
else if (kind === "addressee") await db.update(issueThreadInteractions).set({ addresseeUserId: "another-user" }).where(eq(issueThreadInteractions.id, f.card.id));
else if (kind === "questions") await db.update(issueThreadInteractions).set({ kind: "ask_user_questions" }).where(eq(issueThreadInteractions.id, f.card.id));
else await db.update(issueThreadInteractions).set({ payload: { ...f.card.payload, [kind]: {} } }).where(eq(issueThreadInteractions.id, f.card.id));
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow();
expect((await readCard(f.card.id)).status).toBe("pending");
});
it.each(["human_only", "not_creator"] as const)("does not promote a user's yes through an additional %s review restriction", async policy => {
const f = await seed();
await expect(resolveConfirmationFromComment(db, { ...f.args, actor: { ...f.args.actor, resolverPolicyRestriction: policy } })).rejects.toThrow();
expect((await readCard(f.card.id)).status).toBe("pending");
expect(await audit(f.issueId)).toHaveLength(0);
});
it("rechecks narrowed permissions on an otherwise matching retry", async () => {
const f = await seed();
await resolveConfirmationFromComment(db, f.args);
await db.update(issueThreadInteractions).set({ effectiveResolverPolicy: "human_only" }).where(eq(issueThreadInteractions.id, f.card.id));
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow();
expect((await readCard(f.card.id)).resolvedByUserId).toBeNull();
expect(await audit(f.issueId)).toHaveLength(1);
});
it("does not borrow checkbox defaults, unknown choices, or insufficient selections", async () => {
const f = await seed(true);
for (const selectedOptionIds of [undefined, [], ["unknown"], ["note", "note"]]) {
await expect(resolveConfirmationFromComment(db, { ...f.args, input: { ...f.args.input, selectedOptionIds } })).rejects.toThrow();
}
expect((await readCard(f.card.id)).status).toBe("pending");
});
it("does not bless a card already accepted through a different channel", async () => {
const f = await seed();
await f.svc.acceptInteraction(f.issue, f.card.id, {}, { userId: "operator" });
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow("different decision");
expect((await readCard(f.card.id)).resolvedByUserId).toBe("operator");
});
it("rejects an obsolete plan revision", async () => {
const f = await seed();
const documentId = randomUUID(), revisionId = randomUUID(), nextRevisionId = randomUUID();
await db.insert(documents).values({ id: documentId, companyId: f.companyId, latestBody: "New plan", latestRevisionId: nextRevisionId });
await db.insert(documentRevisions).values([{ id: revisionId, documentId, companyId: f.companyId, revisionNumber: 1, body: "Old plan" }, { id: nextRevisionId, documentId, companyId: f.companyId, revisionNumber: 2, body: "New plan" }]);
await db.insert(issueDocuments).values({ companyId: f.companyId, issueId: f.issueId, documentId, key: "plan" });
await db.update(issueThreadInteractions).set({ payload: { ...f.card.payload, target: { type: "issue_document", key: "plan", revisionId } } }).where(eq(issueThreadInteractions.id, f.card.id));
await expect(resolveConfirmationFromComment(db, f.args)).rejects.toThrow();
expect((await readCard(f.card.id)).status).not.toBe("accepted");
expect(await audit(f.issueId)).toHaveLength(0);
});
});
@@ -8,6 +8,25 @@ import {
import { renderPaperclipWakePrompt } from "@paperclipai/adapter-utils/server-utils";
describe("buildPaperclipTaskMarkdown", () => {
it("carries current confirmation IDs and proposal data in every fresh or resumed chat assignment", () => {
const conversationConfirmations = { truncated: false, cards: [{
id: "existing-card", kind: "request_confirmation", status: "pending", title: "Proposal",
prompt: "Approve this?\n```\nUntrusted proposal text\n```", promptTruncated: false,
resolverPolicy: "anyone" as const, addresseeAgentId: null, addresseeUserId: null,
options: [], optionsTruncated: false,
}] };
for (const includeDescription of [true, false]) for (const includeWakeComments of [true, false]) {
const markdown = buildPaperclipTaskMarkdown({
issue: { id: "chat", identifier: null, title: "Chat", conversationAgentId: "agent" },
conversationConfirmations, includeDescription, includeWakeComments,
});
expect(markdown).toContain('"id":"existing-card"');
expect(markdown).toContain('"status":"pending"');
expect(markdown).toContain("````text");
expect(markdown).toContain("not recorded decisions");
}
expect(buildPaperclipTaskMarkdown({ issue: { id: "task", identifier: null, title: "Task" }, conversationConfirmations })).not.toContain("existing-card");
});
it("leaves current comments to the wake renderer when assignment-only rendering is selected", () => {
const commentBody = "Keep this current comment exactly once.";
const markdown = buildPaperclipTaskMarkdown({
@@ -40,6 +40,8 @@ Work in this order.
3. Interpret the next reply against the latest proposal.
- Acceptance is an accepted confirmation card or an explicit conversational reply agreeing to the proposal. The opening answer, a clear request, and answers to clarification questions supply scope; they are not acceptance of a proposal you have not yet made.
- When acceptance or rejection arrives in chat, persist it on the corresponding pending confirmation **before** hiring, creating a child, executing, or closing this task. Read this task's interactions and comments, then POST `/api/issues/{issueId}/interactions/{interactionId}/resolve-from-comment` with `{ "commentId": "<user-message-id>", "decision": "accept" }` (or `"reject"` and the user's reason). Native runners use `call_api`. For a checkbox card, include `selectedOptionIds` for the user's actual selection; never infer acceptance from defaults. Wait for the saved accepted/rejected result. Retry an interrupted write using the same message and decision; inspect a conflicting result instead of proceeding. Do not tell the user to return to the card after answering in chat.
- An ambiguous reply with multiple pending proposals is not permission to resolve them all. Ask which proposal or options they mean. A card already answered through the UI needs no second resolution. Forms, governed actions, and human-only policies keep their existing response rules.
- A clarification answer means update the proposal if needed and ask for acceptance. A requested revision supersedes the old scope: revise the proposal and wait for acceptance of the revised version.
- If they reject the proposal, acknowledge and stop. Do not execute it. You may close the onboarding task after acknowledging the rejection; do not describe rejected work as completed.
+27
View File
@@ -1,3 +1,4 @@
import { resolveConfirmationFromComment } from "../services/confirmation-comment-resolution.js";
import { createIssueReadTiming } from "../services/issue-read-timing.js";
import { isNativeWorkspaceExportRepairCause } from "@paperclipai/shared";
import { retryNativeWorkspaceExport } from "../services/native-runtime/native-workspace-export-retry.js";
@@ -63,6 +64,7 @@ import {
import {
addIssueCommentSchema,
acceptIssueThreadInteractionSchema,
resolveConfirmationFromCommentSchema,
attachmentArtifactWorkProductMetadataSchema,
cancelIssueThreadInteractionSchema,
skipIssueThreadInteractionSchema,
@@ -15889,6 +15891,31 @@ export function issueRoutes(
},
);
router.post(
"/issues/:id/interactions/:interactionId/resolve-from-comment",
validate(resolveConfirmationFromCommentSchema),
async (req, res) => {
const issue = await getAccessibleResource(req, res, svc.getById(req.params.id as string), "Issue not found");
if (!issue) return;
if (req.actor.type !== "agent") throw forbidden("Conversational resolution requires the responding agent run");
const authorization = await getIssueThreadInteractionResolutionAuthorization(
req, res, issue, req.params.interactionId as string,
);
if (!authorization) return;
const actor = getActorInfo(req);
const result = await resolveConfirmationFromComment(db, {
companyId: issue.companyId, issueId: issue.id,
interactionId: req.params.interactionId as string,
input: req.body,
actor: { agentId: actor.agentId!, runId: actor.runId!,
resolverPolicyRestriction: authorization.resolutionAuthorization.resolverPolicyRestriction },
});
// The authenticated responding run already owns this turn. Do not enqueue
// another self-wake or restart its session after it records the answer.
res.json(result);
},
);
router.post(
"/issues/:id/interactions/:interactionId/accept",
validate(acceptIssueThreadInteractionSchema),
+15
View File
@@ -173,6 +173,7 @@ import {
createIssueThreadInteractionSchema,
createChildIssueSchema,
acceptIssueThreadInteractionSchema,
resolveConfirmationFromCommentSchema,
rejectIssueThreadInteractionSchema,
respondIssueThreadInteractionSchema,
skipIssueThreadInteractionSchema,
@@ -7363,6 +7364,20 @@ registry.registerPath({
responses: { 200: r.ok(), 400: r.badRequest, 401: r.unauthorized },
});
registry.registerPath({
method: "post",
path: "/api/issues/{id}/interactions/{interactionId}/resolve-from-comment",
tags: ["issues"],
summary: "Record a user's conversational confirmation answer",
description: "An eligible active agent run resolves a confirmation on its own task using the latest user comment. Resolver permissions, target staleness and conversation reset boundaries remain enforced. Matching retries are idempotent. No new wake is scheduled.",
request: {
params: z.object({ id: z.string(), interactionId: z.string() }),
body: jsonBody(resolveConfirmationFromCommentSchema),
},
responses: { 200: r.ok(), 400: r.badRequest, 401: r.unauthorized, 403: r.forbidden, 404: r.notFound,
409: { description: "Stale or conflicting decision" }, 422: { description: "Invalid answer evidence or selection" } },
});
registry.registerPath({
method: "post",
path: "/api/issues/{id}/interactions/{interactionId}/accept",
+61 -6
View File
@@ -83,6 +83,8 @@ Research, clarify, and develop full plans here using the conversation's plan doc
When the user asks to approve a plan before handoff, publish the plan and create a revision-bound approval card before ending the turn. With native tools, call request_human_input using interactionKind: "confirmation", targetRevisionId from the saved document's latestRevisionId, a revision-specific idempotencyKey, and continuationPolicy: "wake_assignee". Set payload.target to { type: "issue_document", key: "plan", revisionId: latestRevisionId }. Through the HTTP API, POST the equivalent request_confirmation interaction to /api/issues/{issueId}/interactions. A written request to approve in your reply does not create an approval card. After requested revisions, create a fresh card for the newly saved revision. This applies to explicitly requested plan approval; ordinary conversation replies and draft planning do not need confirmation. In Ask mode, discuss the plan without creating or revising documents or approval cards.
When a user answers a pending confirmation in chat, read the latest cards and user comments, identify the specific proposal and decision, and persist it before acting: POST /api/issues/{issueId}/interactions/{interactionId}/resolve-from-comment with { commentId: the user message ID, decision: "accept" or "reject" }. Use call_api on native runners. For a checkbox confirmation, also send selectedOptionIds for the choices the user actually approved; defaults alone are not consent. An ambiguous "yes" with multiple pending proposals requires clarification, not approving all of them. A requested revision is not acceptance. Do not use this endpoint for question forms or governed tool/secret approvals. Respect resolver-policy denials. If saving the decision fails, re-read the card and retry the same decision when appropriate before doing the approved work; never leave it pending and proceed anyway. If a card is already resolved, use its recorded outcome. Do not ask the user to clear a card after recording their answer.
Before handing off work, inspect available projects and repositories. Every task you create from this chat must belong to a suitable project. Reuse an appropriate existing project; otherwise use create_project. Consider all relevant available repositories and pass repositoryIds for one or multiple repositories when the work spans them. For existing GitHub repositories you can access that are absent from the catalog, pass their HTTPS repositoryUrls; this registers them with the project without creating remote GitHub repositories. You may combine known IDs and URLs and attach multiple repositories. The direct HTTP equivalent is POST /api/companies/{companyId}/projects with name, repositoryIds and/or repositoryUrls arrays, and an idempotencyKey. Include all selected repositories in that creation; do not combine these arrays with workspace. Never invent repository IDs or substitute inaccessible repositories. Ask when the choice is materially ambiguous or required access is missing. Non-code projects may need no repository.
Create ordinary assigned tasks, never subtasks of this conversation. Give each task a clear outcome, context, acceptance criteria, project, and appropriate assignee. Use create_task with initialPlan to copy the relevant plan into the new task before execution starts. If using the HTTP API directly, POST /api/companies/{companyId}/issues with projectId, assigneeAgentId, status: "todo", initialPlan containing the relevant plan Markdown, and an idempotencyKey; omit parentId. Putting a plan in description does not create the task's plan document. Verify the new task's plan document before claiming the handoff is complete. Preserve the original plan here. Carry forward only the remaining execution steps, not completed planning, approval, project creation, or task creation steps. Include the source conversation ID, approved plan revision ID, and accepted interaction ID so the worker can verify the recorded approval. State the approved scope and what is already done; do not claim the new plan document has its own approval. The worker should execute that authorized scope, and ask again only if scope changes or another applicable gate requires it. When splitting work, include the relevant part of the plan in each task. Create and link each task before claiming it exists.
@@ -202,6 +204,53 @@ export async function prepareConversationTurn(
),
);
}
// An answer to a historical question references its original user message.
// Keep that provenance, but do not mistake already answered later messages
// for new work. Only advance through a successful, durably answered turn in
// this same session; queued, failed, or response-less turns do not qualify.
if (context.interactionKind === "ask_user_questions" && context.interactionStatus === "answered"
&& typeof context.interactionId === "string" && typeof context.conversationReplyBoundaryCommentId !== "string") {
const [answered] = await tx.select({ id: issueThreadInteractions.id }).from(issueThreadInteractions).where(and(
eq(issueThreadInteractions.id, context.interactionId), eq(issueThreadInteractions.companyId, run.companyId),
eq(issueThreadInteractions.issueId, issueId), eq(issueThreadInteractions.kind, "ask_user_questions"),
eq(issueThreadInteractions.status, "answered"),
));
if (answered) {
const progress = await tx.select({
id: issueComments.id,
handled: sql<boolean>`exists (select 1 from ${heartbeatRuns} completed
where completed.company_id = ${run.companyId}::uuid and completed.agent_id = ${issue.conversationAgentId}::uuid
and completed.status = 'succeeded' and completed.context_snapshot->>'issueId' = ${issueId}
and completed.context_snapshot->>'conversationSessionGeneration' = ${String(generation)}
and coalesce(completed.context_snapshot->>'wakeCommentId', completed.context_snapshot->>'commentId') = issue_comments.id::text
and (exists (select 1 from issue_comments reply where reply.company_id = ${run.companyId}::uuid
and reply.issue_id = ${issueId}::uuid and reply.created_by_run_id = completed.id
and reply.author_agent_id = ${issue.conversationAgentId}::uuid and reply.deleted_at is null)
or exists (select 1 from issue_thread_interactions question where question.company_id = ${run.companyId}::uuid
and question.issue_id = ${issueId}::uuid and question.source_run_id = completed.id
and question.created_by_agent_id = ${issue.conversationAgentId}::uuid)))`,
}).from(issueComments).where(and(
eq(issueComments.companyId, run.companyId), eq(issueComments.issueId, issueId),
isNull(issueComments.deletedAt), isNull(issueComments.createdByRunId), isNull(issueComments.authorAgentId),
sql`${issueComments.authorUserId} is not null`,
comment ? sql`(${issueComments.createdAt}, ${issueComments.id}) > (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${comment.id}::uuid)` : sql`false`,
)).orderBy(issueComments.createdAt, issueComments.id);
let replyBoundary = commentId;
let firstUnhandled: string | null = null;
for (const candidate of progress) {
if (!candidate.handled) { firstUnhandled = candidate.id; break; }
replyBoundary = candidate.id;
}
context.conversationReplyBoundaryCommentId = replyBoundary;
// Freeze the visible history too. Include prior replies, but never pull
// an unhandled or newly arriving message into this historical answer.
const [replayThrough] = await tx.select({ id: issueComments.id }).from(issueComments).where(and(
eq(issueComments.companyId, run.companyId), eq(issueComments.issueId, issueId), isNull(issueComments.deletedAt),
firstUnhandled ? sql`(${issueComments.createdAt}, ${issueComments.id}) < (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${firstUnhandled}::uuid)` : undefined,
)).orderBy(desc(issueComments.createdAt), desc(issueComments.id)).limit(1);
context.conversationReplayThroughCommentId = replayThrough?.id ?? commentId;
}
}
await tx
.update(issues)
.set({
@@ -281,9 +330,11 @@ export async function settleConversationTurn(
// Messages arriving during the reply remain actionable, including the
// crash window between their comment commit and wake enqueue.
const wakeId =
typeof context.wakeCommentId === "string"
? context.wakeCommentId
: context.commentId;
typeof context.conversationReplyBoundaryCommentId === "string"
? context.conversationReplyBoundaryCommentId
: typeof context.wakeCommentId === "string"
? context.wakeCommentId
: context.commentId;
const [wake] =
typeof wakeId === "string"
? await tx
@@ -356,6 +407,7 @@ export async function conversationReplay(
companyId: string,
issueId: string,
wakeCommentId: string | null,
throughCommentId?: string,
) {
const [issue] = await db
.select()
@@ -368,13 +420,14 @@ export async function conversationReplay(
.from(issueComments)
.where(eq(issueComments.id, issue.conversationBoundaryCommentId))
: [];
const [wake] = wakeCommentId
const replayCutoffId = throughCommentId ?? wakeCommentId;
const [wake] = replayCutoffId
? await db
.select()
.from(issueComments)
.where(
and(
eq(issueComments.id, wakeCommentId),
eq(issueComments.id, replayCutoffId),
eq(issueComments.issueId, issueId),
),
)
@@ -391,7 +444,9 @@ export async function conversationReplay(
? sql`(${issueComments.createdAt}, ${issueComments.id}) > (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${boundary.id}::uuid)`
: undefined,
wake
? sql`(${issueComments.createdAt}, ${issueComments.id}) < (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${wake.id}::uuid)`
? throughCommentId
? sql`(${issueComments.createdAt}, ${issueComments.id}) <= (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${wake.id}::uuid)`
: sql`(${issueComments.createdAt}, ${issueComments.id}) < (select cursor.created_at, cursor.id from issue_comments cursor where cursor.id = ${wake.id}::uuid)`
: undefined,
),
)
@@ -0,0 +1,139 @@
import { and, desc, eq, isNull } from "drizzle-orm";
import { heartbeatRuns, issueComments, issueThreadInteractions, issues, type Db } from "@paperclipai/db";
import { resolveConfirmationFromCommentSchema, type ResolveConfirmationFromComment } from "@paperclipai/shared";
import { assertAgentRunWriteAllowed } from "../agent-run-cancellation.js";
import { conflict, forbidden, notFound, unprocessable } from "../errors.js";
import { persistActivity, publishActivity, type ActivityPublication } from "./activity-log.js";
import { assertIssueThreadInteractionResolverAudience } from "./issue-thread-interaction-resolution.js";
import { issueThreadInteractionService } from "./issue-thread-interactions.js";
type ResolutionActor = Parameters<ReturnType<typeof issueThreadInteractionService>["acceptInteraction"]>[3];
/**
* Record the agent's interpretation of a reply, using the agent's own authority.
* The message is provenance, not a credential or server-verified proof of consent.
* Never promote its author to the resolver: human-only and independent-review
* gates must still reject this agent, even when the referenced text says yes.
*/
export async function resolveConfirmationFromComment(db: Db, args: {
companyId: string;
issueId: string;
interactionId: string;
input: ResolveConfirmationFromComment;
actor: ResolutionActor & { agentId: string; runId: string };
}) {
const input = resolveConfirmationFromCommentSchema.parse(args.input);
let publication: ActivityPublication | null = null;
const commitEffects: Array<(committedDb: Db) => Promise<void>> = [];
const result = await db.transaction(async (transaction) => {
const tx = transaction as unknown as Db;
// Match task mutation/Stop/reset lock ordering, including replay checks.
const [issue] = await tx.select().from(issues).where(and(
eq(issues.id, args.issueId), eq(issues.companyId, args.companyId),
)).for("update");
if (!issue) throw notFound("Issue not found");
await assertAgentRunWriteAllowed(tx, args.companyId, args.actor);
const [run] = await tx.select().from(heartbeatRuns).where(and(
eq(heartbeatRuns.id, args.actor.runId), eq(heartbeatRuns.companyId, args.companyId),
eq(heartbeatRuns.agentId, args.actor.agentId),
)).for("share");
if (!run || run.status !== "running" || issue.executionRunId !== run.id
|| issue.assigneeAgentId !== args.actor.agentId
|| (run.nativeIssueId ?? run.contextSnapshot?.issueId ?? run.contextSnapshot?.taskId) !== issue.id) {
throw forbidden("Only the active responding run can record a conversational decision");
}
const [card] = await tx.select().from(issueThreadInteractions).where(and(
eq(issueThreadInteractions.id, args.interactionId), eq(issueThreadInteractions.companyId, args.companyId),
eq(issueThreadInteractions.issueId, issue.id),
)).for("update");
if (!card) throw notFound("Interaction not found");
if (card.kind !== "request_confirmation" && card.kind !== "request_checkbox_confirmation") {
throw unprocessable("Only confirmation cards can be answered from a conversation");
}
const payload = card.payload as unknown as Record<string, unknown>;
if (payload.toolAction || payload.secretProposal || payload.connectionAuthorization) {
throw forbidden("Governed actions must use their dedicated approval controls");
}
const [comment] = await tx.select().from(issueComments).where(and(
eq(issueComments.id, input.commentId), eq(issueComments.companyId, args.companyId),
eq(issueComments.issueId, issue.id),
)).for("share");
if (!comment || comment.deletedAt || comment.authorType !== "user"
|| !comment.authorUserId || comment.authorAgentId || comment.createdByRunId
|| comment.sourceTrust?.disposition === "quarantined"
|| comment.createdAt < card.createdAt || comment.body.trim() === "/new") {
throw unprocessable("The answer must be a user message posted after this confirmation on the same task");
}
if ((issue.conversationUserId && comment.authorUserId !== issue.conversationUserId)
|| (run.responsibleUserId && comment.authorUserId !== run.responsibleUserId)) {
throw forbidden("The answer is not from this conversation's user");
}
if (issue.conversationAgentId) {
if (run.contextSnapshot?.conversationSessionGeneration !== issue.conversationSessionGeneration) {
throw conflict("Conversation session changed");
}
if (issue.conversationBoundaryCommentId) {
const [boundary] = await tx.select().from(issueComments).where(and(
eq(issueComments.id, issue.conversationBoundaryCommentId), eq(issueComments.companyId, args.companyId),
eq(issueComments.issueId, issue.id),
));
if (!boundary || card.createdAt <= boundary.createdAt || comment.createdAt <= boundary.createdAt) {
throw conflict("A new conversation cannot answer a previous session's confirmation");
}
}
}
const [latest] = await tx.select({ id: issueComments.id }).from(issueComments).where(and(
eq(issueComments.companyId, args.companyId), eq(issueComments.issueId, issue.id),
eq(issueComments.authorType, "user"), isNull(issueComments.deletedAt),
)).orderBy(desc(issueComments.createdAt), desc(issueComments.id)).limit(1);
if (latest?.id !== comment.id) throw conflict("A newer user message must be considered before resolving this confirmation");
const svc = issueThreadInteractionService(tx);
const current = await svc.getForIssue(issue, card.id);
assertIssueThreadInteractionResolverAudience({
actor: { type: "agent", agentId: args.actor.agentId, runId: args.actor.runId },
interaction: card, additionalRestriction: args.actor.resolverPolicyRestriction,
});
const expected = input.decision === "accept" ? "accepted" : "rejected";
const selected = input.selectedOptionIds;
if (card.kind === "request_checkbox_confirmation" && input.decision === "accept" && selected === undefined) {
throw unprocessable("Record the user's explicit selection; checkbox defaults are not an answer");
}
if (card.kind === "request_confirmation" && selected !== undefined) {
throw unprocessable("Selections apply only to checkbox confirmations");
}
if (card.status !== "pending") {
const prior = current.result as { commentId?: string; selectedOptionIds?: string[]; reason?: string } | null;
const sameSelection = JSON.stringify([...(prior?.selectedOptionIds ?? [])].sort()) === JSON.stringify([...(selected ?? [])].sort());
if (card.status !== expected || prior?.commentId !== comment.id || !sameSelection
|| (input.decision === "reject" && (prior?.reason ?? "") !== (input.reason ?? ""))) {
throw conflict("This confirmation already has a different decision");
}
return { interaction: current, deduplicated: true };
}
// These service methods recheck resolver audience, company review policy,
// target revision, selection bounds and issue lifecycle under the same lock.
const mutationOptions = { deferConfirmationCommitEffects: (effect: (committedDb: Db) => Promise<void>) => { commitEffects.push(effect); } };
const interaction = input.decision === "accept"
? (await svc.acceptInteraction(issue, card.id, { selectedOptionIds: selected }, args.actor, mutationOptions)).interaction
: await svc.rejectInteraction(issue, card.id, { reason: input.reason }, args.actor, mutationOptions);
if ((interaction.kind !== "request_confirmation" && interaction.kind !== "request_checkbox_confirmation") || !interaction.result) {
throw conflict("Confirmation resolution did not persist a result");
}
await tx.update(issueThreadInteractions).set({
result: { ...interaction.result, commentId: comment.id },
}).where(and(eq(issueThreadInteractions.id, card.id), eq(issueThreadInteractions.companyId, args.companyId)));
publication = (await persistActivity(tx, {
companyId: args.companyId, actorType: "agent", actorId: args.actor.agentId,
agentId: args.actor.agentId, runId: args.actor.runId,
action: `issue.thread_interaction_${expected}`, entityType: "issue", entityId: issue.id,
details: { interactionId: card.id, interactionKind: card.kind, interactionStatus: expected,
responseCommentId: comment.id, responseUserId: comment.authorUserId, source: "conversation_reply" },
})).publication;
return { interaction: await svc.getForIssue(issue, card.id), deduplicated: false };
});
if (publication) publishActivity(publication);
for (const effect of commitEffects) await effect(db);
return result;
}
@@ -0,0 +1,45 @@
import { and, asc, eq, gt, inArray, isNotNull, sql } from "drizzle-orm";
import { issueComments, issues, issueThreadInteractions, type Db } from "@paperclipai/db";
/** Supply current card identities even when they were created outside the provider session. */
export async function getConversationConfirmationContext(input: {
db: Db; companyId: string; issueId: string; agentId: string;
}) {
const { db, companyId, issueId, agentId } = input;
const [issue] = await db.select({ boundaryId: issues.conversationBoundaryCommentId }).from(issues).where(and(
eq(issues.id, issueId), eq(issues.companyId, companyId),
eq(issues.conversationAgentId, agentId), eq(issues.assigneeAgentId, agentId), isNotNull(issues.conversationUserId),
));
if (!issue) return null;
const [boundary] = issue.boundaryId ? await db.select({ createdAt: issueComments.createdAt }).from(issueComments).where(and(
eq(issueComments.id, issue.boundaryId), eq(issueComments.companyId, companyId), eq(issueComments.issueId, issueId),
)) : [];
if (issue.boundaryId && !boundary) return null;
const limit = 12;
const rows = await db.select().from(issueThreadInteractions).where(and(
eq(issueThreadInteractions.companyId, companyId), eq(issueThreadInteractions.issueId, issueId),
eq(issueThreadInteractions.status, "pending"),
inArray(issueThreadInteractions.kind, ["request_confirmation", "request_checkbox_confirmation"]),
// Dedicated approval payloads are not conversational confirmations.
sql`not (${issueThreadInteractions.payload} ?| array['toolAction', 'secretProposal', 'connectionAuthorization'])`,
boundary ? gt(issueThreadInteractions.createdAt, boundary.createdAt) : undefined,
)).orderBy(asc(issueThreadInteractions.createdAt), asc(issueThreadInteractions.id)).limit(limit + 1);
return {
truncated: rows.length > limit,
cards: rows.slice(0, limit).map(card => {
const payload = card.payload as unknown as Record<string, unknown>;
const options = Array.isArray(payload.options) ? payload.options as Array<{ id: string; label: string }> : [];
const prompt = typeof payload.prompt === "string" ? payload.prompt : "";
return {
id: card.id, kind: card.kind, status: card.status, title: card.title?.slice(0, 256) ?? null,
prompt: prompt.slice(0, 2000), promptTruncated: prompt.length > 2000,
resolverPolicy: card.effectiveResolverPolicy,
addresseeAgentId: card.addresseeAgentId, addresseeUserId: card.addresseeUserId,
options: options.slice(0, 20).map(option => ({ id: option.id, label: option.label.slice(0, 256) })),
optionsTruncated: options.length > 20,
};
}),
};
}
export type ConversationConfirmationContext = Awaited<ReturnType<typeof getConversationConfirmationContext>>;
+14 -3
View File
@@ -21,6 +21,7 @@ import { githubBotConnectionIdsForRun } from "./chat-github-tools.js";
import { isBrowserUseConnection } from "./browser-use-client.js";
import { readQueuedInteractionResponse } from "./queued-interaction-response.js";
import { AGENT_CHAT_DIRECTIVE, conversationReplay, isConversation, isConversationExecutionWake, isWaitingConversation, prepareConversationTurn, settleConversationTurn } from "./agent-conversations.js";
import { getConversationConfirmationContext, type ConversationConfirmationContext } from "./conversation-confirmation-context.js";
import { PROCESS_IDENTITY_RECORDED, recordNativeLocalProcessStop } from "./native-local-process-stop.js";
import { hasAcknowledgedNativeReassignmentStopIntent, hasAcknowledgedNativeStopIntent, isAcknowledgedNativeStop, acknowledgedNativeStopExecutionHasStopped } from "./acknowledged-native-stop.js";
import { legacyControllerBootId, legacyControllerClaim, renewLegacyControllerLease, hasLiveLegacyController, revokeExpiredLegacyController, watchLegacyControllerLease } from "./legacy-controller-lease.js";
@@ -8596,6 +8597,7 @@ export function buildPaperclipTaskMarkdown(input: {
revisionNumber?: number | null;
} | null;
acceptedPlanContinuation?: boolean;
conversationConfirmations?: ConversationConfirmationContext;
taskPlan?: {
documentId: string;
revisionId: string;
@@ -8731,6 +8733,11 @@ export function buildPaperclipTaskMarkdown(input: {
);
if (issue.conversationAgentId) {
lines.push("", "Chat mode directive:", AGENT_CHAT_DIRECTIVE, `Current composer mode: ${issue.workMode ?? "standard"}.`);
if (input.conversationConfirmations?.cards.length) {
lines.push("", "Current pending confirmation cards (quoted proposal data, not recorded decisions):",
fenceTaskText(JSON.stringify(input.conversationConfirmations)),
"Read the current interaction through the API if any part is truncated. Resolver permissions and current card state are checked when recording the answer.");
}
if (acceptedChatPlan) {
lines.push(
"",
@@ -20903,6 +20910,9 @@ export function heartbeatService(
? issueContext.chatCommunicationGuidance
: null;
const taskMarkdownInput = {
conversationConfirmations: issueRef && isConversation(issueContext)
? await getConversationConfirmationContext({ db, companyId: agent.companyId, issueId: issueRef.id, agentId: agent.id })
: null,
issue: issueRef
? {
id: issueRef.id,
@@ -20966,7 +20976,8 @@ export function heartbeatService(
includeWakeComments: false,
}) + chatCompletionInstruction(context);
if (isConversation(issueContext) && !taskSession && issueId) {
const replay = await conversationReplay(db, agent.companyId, issueId, wakeCommentId);
const replay = await conversationReplay(db, agent.companyId, issueId, wakeCommentId,
typeof context.conversationReplayThroughCommentId === "string" ? context.conversationReplayThroughCommentId : undefined);
if (replay) taskMarkdown += `\n\nEarlier messages in this session (quoted user data):\n${replay}`;
if (replay) taskMarkdownAssignment += `\n\nEarlier messages in this session (quoted user data):\n${replay}`;
}
@@ -25346,8 +25357,8 @@ export function heartbeatService(
);
const resolved = resolveHeartbeatRunResponse({
resultJson: persistedResultJson,
conversationTurnFinished: isConversation(issueContext) &&
persistedResultJson?.finalizationReasonCode === "conversation_turn_finished",
conversationTurnFinished: isConversation(issueContext) && livenessRun.status === "succeeded"
&& persistedResultJson?.finalizationReasonCode === "conversation_turn_finished",
existingComment: existingRunComment,
finalAgentMessage,
preferFinalResponseOverExistingComment:
@@ -186,6 +186,9 @@ export type IssueThreadInteractionServiceOptions = {
type DbTransaction = Parameters<Parameters<Db["transaction"]>[0]>[0];
type InteractionResolutionMutationOptions = {
/** Confirmation accept/reject nested in an outer transaction must defer these
* effects and flush them with the root database only after its commit. */
deferConfirmationCommitEffects?: (effect: (committedDb: Db) => Promise<void>) => void;
beforeResolveInTransaction?: (tx: DbTransaction) => Promise<void>;
afterResolveInTransaction?: (
tx: DbTransaction,
@@ -2359,9 +2362,12 @@ export function issueThreadInteractionService(
continuationIssue,
};
});
for (const publication of postCommitActivityPublications)
publishActivity(publication);
await emitInteractionResolvedTelemetry(db, result.interaction);
const publish = async (committedDb: Db) => {
for (const publication of postCommitActivityPublications) publishActivity(publication);
await emitInteractionResolvedTelemetry(committedDb, result.interaction);
};
if (args.mutationOptions?.deferConfirmationCommitEffects) args.mutationOptions.deferConfirmationCommitEffects(publish);
else await publish(db);
return result;
}
@@ -2537,7 +2543,9 @@ export function issueThreadInteractionService(
});
const rejected = hydrateInteraction(updated);
await emitInteractionResolvedTelemetry(db, rejected);
const publish = (committedDb: Db) => emitInteractionResolvedTelemetry(committedDb, rejected);
if (args.mutationOptions?.deferConfirmationCommitEffects) args.mutationOptions.deferConfirmationCommitEffects(publish);
else await publish(db);
return rejected;
}
@@ -3506,7 +3514,7 @@ export function issueThreadInteractionService(
const result = await db.transaction(async (tx) => {
await assertInteractionRunWriteAllowed(tx as unknown as Db, issue, actor);
const [issueRow] = await tx
.select({ status: issues.status })
.select({ status: issues.status, conversationAgentId: issues.conversationAgentId, conversationUserId: issues.conversationUserId })
.from(issues)
.where(
and(
@@ -3597,16 +3605,15 @@ export function issueThreadInteractionService(
// An agent replacing its own still-pending card supersedes the older
// one so the thread never accumulates stale sibling cards. This covers
// request_confirmation drafts and ask_user_questions (PAP-437: probe
// question cards that agents never withdrew). Each kind keeps its own
// result shape. Scoped strictly to the same agent + issue + kind, so
// other agents' or other kinds' pending cards are untouched.
// request_confirmation drafts and ordinary task questions. Agent Chat
// questions remain answerable in history even when another is asked.
// Scoped to the same agent + issue + kind; other actors are untouched.
const canSupersedeSiblingCards =
options.supersedePendingSiblingInteractions !== false &&
((data.kind === "request_confirmation" &&
data.payload.toolAction === undefined &&
data.payload.secretProposal === undefined) ||
data.kind === "ask_user_questions");
(data.kind === "ask_user_questions" && (!issueRow.conversationAgentId || !issueRow.conversationUserId)));
if (!actor.agentId || !canSupersedeSiblingCards) {
await enqueueIssueInteractionChatPublications(
tx as unknown as Db,
@@ -3638,6 +3645,7 @@ export function issueThreadInteractionService(
eq(issueThreadInteractions.createdByAgentId, actor.agentId),
eq(issueThreadInteractions.status, "pending"),
ne(issueThreadInteractions.id, row.id),
),
)
.returning();
@@ -4226,6 +4234,9 @@ export function issueThreadInteractionService(
// machine; createdByRunId can. Only genuine human comments (no run context) supersede.
if (comment.createdByRunId) return [];
const [scope] = await db.select({ conversationAgentId: issues.conversationAgentId, conversationUserId: issues.conversationUserId })
.from(issues).where(and(eq(issues.id, issue.id), eq(issues.companyId, issue.companyId)));
const rows = await db
.select()
.from(issueThreadInteractions)
@@ -4241,6 +4252,7 @@ export function issueThreadInteractionService(
);
const superseded = rows.filter((row) => {
if (row.kind === "ask_user_questions" && scope?.conversationAgentId && scope.conversationUserId) return false;
if (!isUserCommentSupersedableKind(row.kind)) return false;
const interaction = hydrateInteraction(
row,
@@ -0,0 +1,61 @@
import { createIssueThreadInteractionSchema } from "@paperclipai/shared";
import { randomUUID } from "node:crypto";
import { eq } from "drizzle-orm";
import { afterAll, beforeAll, describe, expect, it } from "vitest";
import { issueComments, issueThreadInteractions, issues } from "@paperclipai/db";
import { createLocalAgentJwt } from "../../agent-auth-jwt.js";
import { startRunnerApiTestServer } from "../../__tests__/helpers/runner-api-server.js";
import { issueThreadInteractionService } from "../issue-thread-interactions.js";
import { runnerApiMutationRestriction } from "./runner-api-policy.js";
describe("confirmation reply through the native API tool", () => {
let server: Awaited<ReturnType<typeof startRunnerApiTestServer>>;
const oldSecret = process.env.PAPERCLIP_AGENT_JWT_SECRET;
beforeAll(async () => { process.env.PAPERCLIP_AGENT_JWT_SECRET = randomUUID(); server = await startRunnerApiTestServer(); }, 60_000);
afterAll(async () => { await server?.close(); if (oldSecret === undefined) delete process.env.PAPERCLIP_AGENT_JWT_SECRET; else process.env.PAPERCLIP_AGENT_JWT_SECRET = oldSecret; });
async function seed(mode: "standard" | "planning" | "ask" = "standard") {
const f = await server.fixture({ conversation: true, disableWakeOnDemand: true, mode });
const [issue] = await server.db.select().from(issues).where(eq(issues.id, f.issueId));
const card = await issueThreadInteractionService(server.db).create(issue!, createIssueThreadInteractionSchema.parse({ kind: "request_confirmation", payload: { version: 1, prompt: "Write the welcome note?" } }), { agentId: f.agentId });
const [comment] = await server.db.insert(issueComments).values({ companyId: f.companyId, issueId: f.issueId, authorType: "user", authorUserId: f.userId, body: "Yes, write it." }).returning();
const call = { tool: "call_api", callId: randomUUID(), arguments: { operationId: "POST /api/issues/{id}/interactions/{interactionId}/resolve-from-comment", pathParams: { id: f.issueId, interactionId: card.id }, body: { commentId: comment!.id, decision: "accept" } } };
return { f, card, comment: comment!, call };
}
it("advertises and executes the real authenticated route, including retry", async () => {
const { f, card, comment, call } = await seed();
const first = await f.authority.execute(call);
expect(first).toMatchObject({ status: 200, data: { deduplicated: false, interaction: { id: card.id, status: "accepted", result: { commentId: comment.id } } } });
const retry = await f.authority.execute({ ...call, callId: randomUUID() });
expect(retry).toMatchObject({ status: 200, data: { deduplicated: true } });
const snapshot = await f.snapshot();
expect(snapshot.activity.filter(row => row.action === "issue.thread_interaction_accepted")).toHaveLength(1);
});
it("denies human-only resolution and foreign-message provenance through the same tool", async () => {
const { f, card, call, comment } = await seed();
await server.db.update(issueThreadInteractions).set({ effectiveResolverPolicy: "human_only" }).where(eq(issueThreadInteractions.id, card.id));
expect(await f.authority.execute(call)).toMatchObject({ status: 403 });
const other = await seed();
const foreignAnswer = { ...other.call, arguments: { ...other.call.arguments, body: { ...other.call.arguments.body, commentId: comment.id } } };
expect(await other.f.authority.execute(foreignAnswer)).toMatchObject({ status: 422 });
});
it("records a conversational decision in Plan mode without opening writes in Ask mode", async () => {
const planning = await seed("planning");
expect(await planning.f.authority.execute(planning.call)).toMatchObject({ status: 200, data: { interaction: { status: "accepted" } } });
const ask = await seed("ask");
await expect(ask.f.authority.execute(ask.call)).rejects.toThrow("only reads");
expect((await server.db.select().from(issueThreadInteractions).where(eq(issueThreadInteractions.id, ask.card.id)))[0]?.status).toBe("pending");
});
it("rejects unauthenticated HTTP requests", async () => {
const { f, card, call, comment } = await seed();
const url = `${server.apiUrl}/api/issues/${f.issueId}/interactions/${card.id}/resolve-from-comment`;
expect((await fetch(url, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify(call.arguments.body) })).status).toBe(404);
const token = createLocalAgentJwt(f.agentId, f.companyId, "paperclip_runner", f.runId, f.userId);
expect((await fetch(url, { method: "POST", headers: { "Content-Type": "application/json", Authorization: `Bearer ${token}` }, body: JSON.stringify({ ...call.arguments.body, actorUserId: f.userId }) })).status).toBe(400);
});
it("does not open the ordinary interaction or governance mutation routes", () => {
for (const operation of ["accept", "reject", "respond", "verdicts", "withdraw"]) {
expect(runnerApiMutationRestriction(`/api/issues/{id}/interactions/{interactionId}/${operation}`)).not.toBeNull();
}
expect(runnerApiMutationRestriction("/api/issues/{id}/interactions/{interactionId}/resolve-from-comment")).toBeNull();
});
});
@@ -42,6 +42,7 @@ import {
issueService,
type IssuePostCommitAction,
} from "../issues.js";
import { questionResponseDeliveryService } from "../question-response-delivery.js";
import { heartbeatService } from "../heartbeat.js";
import { DurablePrpControlPlane } from "../../vendor/paperclip-runner/index.js";
import { PaperclipControlPlanePort } from "./paperclip-control-plane-port.js";
@@ -403,6 +404,32 @@ describeEmbeddedPostgres("native question bridge", () => {
release();
});
it.each(["succeeded", "failed", "cancelled", "timed_out"])("delivers a historical answer exactly once after a %s native run", async status => {
await seed();
await db.update(issues).set({ conversationAgentId: agentId, conversationUserId: "operator-1", conversationState: "active", conversationSessionGeneration: 1, executionRunId: null }).where(eq(issues.id, issueId));
const interaction = await projectNativeRuntimeRequest({ db, binding: binding(), event: runtimeRequestEvent() });
await db.update(heartbeatRuns).set({ status, finishedAt: new Date() }).where(eq(heartbeatRuns.id, runId));
await issueThreadInteractionService(db).answerQuestions({ id: issueId, companyId, status: "in_progress" }, interaction!.id,
{ answers: [{ questionId: "color", optionIds: ["green"] }] }, { userId: "operator-1" });
const queueCommand = vi.fn();
const release = registerNativeQuestionCommandTarget({ binding: { companyId, issueId, runId, agentId }, queueCommand });
const targetRunId = randomUUID();
const wakeup = vi.fn(async () => db.insert(heartbeatRuns).values({ id: targetRunId, companyId, agentId, status: "queued",
contextSnapshot: { issueId } }).returning().then(rows => rows[0]!));
const service = questionResponseDeliveryService(db, { heartbeat: { wakeup } as never,
resolveNativeQuestion: candidate => deliverNativeQuestionResponse(db, candidate) });
expect(await service.deliver(interaction!.id)).toMatchObject({ status: "fallback_queued", targetRunId });
expect((await service.deliver(interaction!.id))?.duplicate).toBe(true);
expect(wakeup).toHaveBeenCalledTimes(1);
expect(JSON.stringify(wakeup.mock.calls)).toContain(interaction!.id);
const [saved] = await db.select().from(issueThreadInteractions).where(eq(issueThreadInteractions.id, interaction!.id));
expect(saved).toMatchObject({ status: "answered", result: { answers: [{ questionId: "color", optionIds: ["green"] }] } });
const [delivery] = await db.select().from(issueQuestionResponseDeliveries).where(eq(issueQuestionResponseDeliveries.interactionId, interaction!.id));
expect(delivery).toMatchObject({ targetRunId, sourceRunId: runId, deliveryMode: "wake_fallback", status: "fallback_queued" });
expect(queueCommand).not.toHaveBeenCalled();
release();
});
it("binds projection to the persisted native run and ignores legacy delivery", async () => {
await seed();
const mismatched = runtimeRequestEvent();
@@ -336,7 +336,10 @@ export async function deliverNativeQuestionResponse(
return "not_native";
}
const run = await authorizedNativeRun(db, interaction);
if (!run) return "not_native";
// A historical question can be answered after its provider turn has ended.
// Fall through to durable fresh-wake delivery instead of waiting forever for
// a command target that cannot return for this terminal run.
if (!run || ["succeeded", "failed", "cancelled", "timed_out"].includes(run.status)) return "not_native";
const response = canonicalResponse(interaction.payload.questionSet, interaction.result.answers);
const target = activeTargets.get(run.id);
if (
@@ -135,6 +135,8 @@ export async function pendingNativeGovernance(input: {
if (record(input.executionState).status === "pending") {
return { kind: "execution_stage", id: input.runId };
}
const [issue] = await input.db.select({ conversationAgentId: issues.conversationAgentId, conversationUserId: issues.conversationUserId })
.from(issues).where(and(eq(issues.id, input.issueId), eq(issues.companyId, input.companyId)));
const [pendingInteraction, pendingApproval] = await Promise.all([
input.db
.select({ id: issueThreadInteractions.id })
@@ -144,6 +146,18 @@ export async function pendingNativeGovernance(input: {
eq(issueThreadInteractions.companyId, input.companyId),
eq(issueThreadInteractions.issueId, input.issueId),
eq(issueThreadInteractions.status, "pending"),
// A previous chat turn's ordinary input remains answerable in history;
// it does not own the lifecycle of every subsequent reply. Current-turn
// requests, task execution, and governed approvals keep their gates.
isConversation(issue) ? sql`(
${issueThreadInteractions.sourceRunId} is not distinct from ${input.runId}
or not (
${issueThreadInteractions.kind} = 'ask_user_questions'
or (${issueThreadInteractions.kind} in ('request_confirmation', 'request_checkbox_confirmation')
and ${issueThreadInteractions.effectiveResolverPolicy} = 'anyone'
and not (${issueThreadInteractions.payload} ?| array['toolAction', 'secretProposal', 'connectionAuthorization']))
)
)` : undefined,
),
)
.limit(1)
@@ -102,7 +102,9 @@ describe("PaperclipRunnerToolAuthority", () => {
expect(authority.definitions()).toHaveLength(34);
const questions = authority.definitions().find(tool => tool.name === "request_human_input")!;
expect(questions.description).toContain("ask only the next unanswered question");
expect(questions.description).toContain("Never infer answers");
expect(questions.description).toContain("Never fabricate answers");
expect(questions.description).toContain("resolve-from-comment");
expect(questions.description).toContain("Existing resolver permissions still apply");
expect(questions.description).toContain("Do not fabricate answer links");
expect(JSON.stringify(questions.inputSchema)).toContain("at least two distinct meaningful options");
expect(JSON.stringify(questions.inputSchema)).toContain("answerMode:'text'");
@@ -81,6 +81,8 @@ export function buildRunnerApiCatalog(document: Json = buildOpenApiDocument()):
if (!METHODS.has(verb)) continue;
const method = verb.toUpperCase();
const restriction = runnerApiRestriction(method, path);
const conversationalConfirmation = method === "POST"
&& /^\/api\/issues\/\{[^}]+\}\/interactions\/\{[^}]+\}\/resolve-from-comment$/.test(path);
const skillReference = runnerApiReference[`${method} ${path.replace(/\{[^}]+\}/g, "{}")}`];
const protocol = !path.startsWith("/api/") || /\/(oauth|auth|runtime-tools|mcp|ws)(\/|$)/.test(path)
|| /\/(claude-login|login-sessions|start-authorization|finalize-oauth-access)(\/|$)/.test(path)
@@ -100,11 +102,15 @@ export function buildRunnerApiCatalog(document: Json = buildOpenApiDocument()):
const descriptor = CAPABILITY_SEMANTIC_TOOL_CATALOG.find(tool => tool.operationId === name);
return descriptor ? [{ name, description: descriptor.description, supportedParameters: Object.keys((descriptor.inputSchema as Json).properties ?? {}) }] : [];
}),
runnerRestrictions: [...(restriction ? [restriction] : []), "Active run and assignment must remain authorized.", "Cannot replace checkout, completion, task status/ownership changes, approval decisions, or runner execution control. Dedicated tools retain their existing permissions."],
runnerRestrictions: [...(restriction ? [restriction] : []), "Active run and assignment must remain authorized.",
conversationalConfirmation
? "This records an ordinary conversational confirmation under existing resolver permissions. It cannot decide governed approvals or impersonate the user."
: "Cannot replace checkout, completion, task status/ownership changes, governed approval decisions, or runner execution control. Dedicated tools retain their existing permissions."],
dedicatedToolGuidance: dedicatedTools(method, path).includes("hire_agent")
? "Use hire_agent for native teammate identity and persona fields. It fixes the reportsTo and source task context and inherits the caller's native runtime; never use call_api to supply adapter, environment, or credential configuration."
: "Use an available dedicated tool for its supported fields. Inspect that tool's advertised schema; call_api may be used for additional API fields, subject to lifecycle restrictions.",
allowedModes: method === "GET" || method === "HEAD" ? ["standard", "ask", "planning", "skill_test"] : ["standard", "skill_test"],
allowedModes: method === "GET" || method === "HEAD" ? ["standard", "ask", "planning", "skill_test"]
: conversationalConfirmation ? ["standard", "planning", "skill_test"] : ["standard", "skill_test"],
...(skillReference ? { skillReference } : {}),
});
}
@@ -30,6 +30,10 @@ export function runnerApiRestriction(
/** Runner-owned transitions cannot be reached by the generic HTTP escape hatch. */
export function runnerApiMutationRestriction(path: string): string | null {
// This endpoint only records an eligible responding agent's interpretation
// of a same-task user message. It checks the live run, resolver policy and
// session under lock, and cannot execute governed approvals or enqueue work.
if (/^\/api\/issues\/\{[^}]+\}\/interactions\/\{[^}]+\}\/resolve-from-comment$/.test(path)) return null;
const issueRoute = /^\/api\/issues\/\{[^}]+\}/.test(path);
const routineAnnotation =
/^\/api\/routines\/\{[^}]+\}\/description\/annotations(?:\/\{[^}]+\}(?:\/comments)?)?$/.test(
@@ -375,6 +375,10 @@ export const runnerApiReference: Record<string, { section: string; description?:
}
]
},
"POST /api/issues/{}/interactions/{}/resolve-from-comment": {
"section": "Issues (Tasks)",
"description": "Resolve a confirmation from the latest user reply; body: commentId, decision (accept/reject), selectedOptionIds for checkbox acceptance, optional reason"
},
"POST /api/issues/{}/interactions/{}/accept": {
"section": "Issues (Tasks)",
"description": "Accept suggested tasks or confirmation (body: `selectedClientKeys` for `suggest_tasks`; `selectedOptionIds` for `request_checkbox_confirmation`)",
@@ -58,6 +58,15 @@ describe("bounded response capture receipts", () => {
});
describe("runner API catalog", () => {
it("advertises the conversational recording exception without general approval authority", () => {
const operationId = "POST /api/issues/{id}/interactions/{interactionId}/resolve-from-comment";
const operation = runnerApiOperation(operationId);
expect(operation.allowedModes).toContain("planning");
expect(operation.allowedModes).not.toContain("ask");
expect(operation.runnerRestrictions?.join(" ")).toContain("conversational confirmation");
expect(runnerApiOperation("POST /api/issues/{id}/interactions/{interactionId}/accept").callPolicy).toBe("restricted");
expect(runnerApiOperation(createProject).allowedModes).not.toContain("planning");
});
it("accounts for unique operations with resolved request contracts", () => {
const catalog = runnerApiCatalog();
expect(catalog.length).toBeGreaterThan(400);
+6
View File
@@ -718,3 +718,9 @@ Results are ranked by relevance: title matches first, then identifier, descripti
## Full Reference
For detailed API tables, JSON response schemas, worked examples (IC and Manager heartbeats), governance/approvals, cross-team delegation rules, error codes, issue lifecycle diagram, and the common mistakes table, read: `skills/paperclip/references/api-reference.md`
### Conversational confirmation answers
When the user answers a pending confirmation in a message, record the answer before acting. Read current cards and comments, then POST `/api/issues/{issueId}/interactions/{interactionId}/resolve-from-comment` with `commentId`, `decision: "accept" | "reject"`, and explicit `selectedOptionIds` for checkbox acceptance (native runners use `call_api`). Ambiguous replies among proposals require clarification. Revisions are not acceptance. Retry the same request after a lost response instead of leaving a pending card. Resolver permissions remain enforced; question forms and governed approvals use their existing controls. See the API reference for scope and retry rules.
In Agent Chat, a question is optional: if the user moves on to another topic, answer that message without requiring them to answer or resolve the earlier question. Leave its card unanswered so they can reopen it later. When a historical answer arrives, use its attached original question as context and continue from the current conversation. Unrelated messages are never approval.
@@ -1083,6 +1083,12 @@ Rules:
- A pending interaction is an explicit waiting path. Before ending the heartbeat, update the source issue into a visible waiting posture, normally `in_review`, and leave a comment that names the response needed and the effective audience.
- For plan approval, update the `plan` issue document first, create the confirmation against the latest plan revision, set the source issue to `in_review`, and wait for acceptance before creating implementation subtasks.
### Conversational confirmation answers
To record a user's conversational answer, an eligible agent responding on this task may POST `/api/issues/{issueId}/interactions/{interactionId}/resolve-from-comment` with `{ "commentId": "<latest-user-message-id>", "decision": "accept" }` (or `"reject"` and `reason`). Native runners use `call_api`. Checkbox acceptance must include explicit `selectedOptionIds`; defaults alone are not consent. The result is `{ interaction, deduplicated }`, with the user message retained in `interaction.result.commentId` and the activity audit. Resolver attribution remains the responding agent/run. This does not widen permissions: `human_only`, independent-review restrictions, named addressees, and governed-action controls still apply. Only confirmation and checkbox cards are supported, not forms, secret/tool approvals, or connection authorizations.
Read current cards and comments before interpreting the reply. Resolve the specific proposal before performing the approved work. Ask for clarification when a reply is ambiguous among multiple proposals or checkbox choices; do not approve all of them. Requested revisions are not acceptance. If the write is interrupted, retry the same card/message/decision: matching retries return `deduplicated: true` without another wake. Conflicting, stale, deleted, superseded, wrong-user, and previous-session answers fail. Do not ask the user to clear a card after their decision is saved.
### Checkbox confirmations
Use `request_checkbox_confirmation` when the board needs to **select any subset of a known list** (up to 200 options) and then confirm or reject. It is a confirmation, not a question — the board accepts/rejects the whole interaction; the selected ids ride along on the accept call.
@@ -1420,6 +1426,7 @@ Terminal states: `done`, `cancelled`
| DELETE | `/api/issues/:issueId/inbox-archive` | Reverse inbox archive; same target and policy rules |
| GET | `/api/issues/:issueId/interactions` | List issue-thread interactions |
| POST | `/api/issues/:issueId/interactions` | Create issue-thread interaction (`suggest_tasks`, `ask_user_questions`, `request_confirmation`, `request_checkbox_confirmation`, `request_item_verdicts`) |
| POST | `/api/issues/:issueId/interactions/:interactionId/resolve-from-comment` | Resolve a confirmation from the latest user reply; body: commentId, decision (accept/reject), selectedOptionIds for checkbox acceptance, optional reason |
| POST | `/api/issues/:issueId/interactions/:interactionId/accept` | Accept suggested tasks or confirmation (body: `selectedClientKeys` for `suggest_tasks`; `selectedOptionIds` for `request_checkbox_confirmation`) |
| POST | `/api/issues/:issueId/interactions/:interactionId/reject` | Reject suggested tasks or confirmation |
| POST | `/api/issues/:issueId/interactions/:interactionId/respond` | Respond to structured questions |
+45
View File
@@ -33,6 +33,50 @@ full traces and checks that Vite module loads do not create worker fetches or
leave the page empty. This isolates browser loading; it does not create a
Paperclip task, run an agent, or replace a Product E2E result.
## Conversational confirmation replies (explicit only)
`--suite confirmation-replies` selects ten local native Claude/Codex cells:
conversational single-task approval, saved-plan approval, rejection, card-click
acceptance as a control, and ambiguous approval with two independent pending
proposals. The onboarding cells use the production wizard, runtime switch and
persona. The ambiguity fixture creates ordinary board cards through the public
API, sends "Yes, go ahead" through the browser, requires both to stay pending
with a clarification reply and no execution, then approves only one and rejects
the other through separate browser messages.
The board-created cards also check that fresh and resumed chat turns receive
the current confirmation identities, including cards outside provider memory.
Only ordinary, current-session pending confirmations enter this bounded context;
the resolution endpoint still rechecks live state and permissions.
Clarification may be a fresh chat reply or a source-bound question card, including
a native question-set description and its proposal choices. A generic question
about tone or deadlines is not evidence that the ambiguous approval was clarified.
Unauthorized state changes fail immediately. Clarification wording is graded
after capturing later decisions and reload receipts, so a new wording variant
does not discard the rest of the paid journey's evidence. A wording failure
still fails the case; any later offline regrade must be reported separately.
The independent oracle requires the exact original card to hold the decision,
source user-comment ID and resolving agent/run. Acceptance must precede child
creation. Expiring/hiding the card, reporting acceptance only in prose, or
finishing work with a pending card fails. Browser reload verifies the displayed
accepted/rejected receipt. Existing card-click behavior remains unchanged.
Accepted onboarding cases also use the 120-second completion/result-access
probe and retain its semantic evidence. Inspect final prose for obsolete
requests to clear the approval card; mechanical success alone does not establish
prose quality. Provider-scoped Claude jobs require separate retained-probe
judging, as described below.
No model-generated outcome or direct database mutation supplies a pass. Source
SHA, definition hash, model, run evidence, screenshots, failures, cleanup and
billing use the existing report pipeline. This suite is opt-in and excluded
from `--all`; it does not qualify governed tool approvals, human-only policies,
question-form extraction, or remote execution.
```sh
pnpm test:e2e:runner -- --list --suite confirmation-replies
pnpm test:e2e:runner -- --suite confirmation-replies
```
## Completion-update probes (explicit only)
`--suite completion-updates` selects ten local Product E2E cells: native Codex
@@ -1318,6 +1362,7 @@ The separate Runner Evals `extended-harnesses` campaign lives in the private
`paperclip-evals` repository and grades semantic protocol behavior against the
mock control plane. Neither suite substitutes for the other.
The explicit-only `confirmation-replies` suite also includes `unanswered-question-return` for native Claude and Codex (three provider turns). The browser asks a saved color question, dismisses and reopens the fresh form, sends an unrelated message, verifies the reply while the original stays pending, reloads, reopens the history entry, submits Blue, and verifies the saved answer plus a later agent acknowledgement. After dismissing the fresh form and before and after reload, the history card is the only pending-question reminder; the composer has no duplicate pending-input badge. It checks that no tasks were created. Unique, UI-ready screenshots show each checkpoint; individual checks are included in the report. This is a bounded mechanical workflow check, not broader semantic answer-quality qualification.
## Direct blocker guidance
`blocker-guidance` is an explicit-only Product E2E suite for the production
+3 -3
View File
@@ -113,10 +113,10 @@ describe("runner E2E catalog", () => {
expect(localIntegrityTasks).toHaveLength(2);
expect(openRouterBreadthTasks).toHaveLength(3);
expect(runnerSuites.map((suite) => suite.expectedMatrixSize)).toEqual([
6, 30, 3, 16, 16, 2, 8, 46, 23, 47, 20, 52, 28, 18, 6, 6, 10, 48, 16, 10, 2, 1, 1,
6, 30, 3, 16, 16, 2, 8, 46, 23, 47, 20, 52, 28, 18, 6, 6, 12, 10, 48, 16, 10, 2, 1, 1,
]);
expect(validateRunnerCatalog()).toHaveLength(415);
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(415);
expect(validateRunnerCatalog()).toHaveLength(427);
expect(new Set(runnerMatrix.map((entry) => entry.id)).size).toBe(427);
expect(
runnerMatrix.filter((entry) => entry.suite.id === "core-compatibility"),
).toHaveLength(48);
+12
View File
@@ -1,3 +1,4 @@
import { chatConfirmationTasks } from "./chat-cases.js";
import { instructionPersistenceTask } from "./instruction-persistence.js";
import { apiResponseReadingTask } from "./api-response-reading.js";
import { blockerTasks, blockerProfile } from "./blocker-cases.js";
@@ -1215,6 +1216,17 @@ export const runnerSuites: readonly RunnerSuiteFixture[] = [
environments: [localEnvironment], tasks: chatQualificationTasks, expectedMatrixSize: 6,
definitionMetadata: { version: 10, instructionSetup: "read-before-write-base-hash", permissions: "production-defaults", instructions: "production", crashBoundary: "verified-native-worker-pid-at-file-wait", recovery: "new-user-message-after-verified-cleanup", answerGrading: "exact-grounded-propositions-plus-separate-semantic-review", scheduling: "explicit-only" },
},
{
id: "confirmation-replies", label: "Conversational Approval Cards", manualOnly: true,
description: "Persist approval/refusal from chat before execution, preserve card clicks, clarify ambiguous proposals, and answer historical questions after moving on.",
groups: ["chat", "native"],
profiles: runnerProfiles.filter(profile => ["runner-codex", "runner-acpx-claude"].includes(profile.id))
.map(profile => productionStoryProfile(defaultPermissionProfile(profile))),
environments: [localEnvironment],
tasks: [...firstTaskTasks.filter(task => ["task-reply-accept", "interview-plan-accept", "reject-no-execution", "task-card-accept"].includes(task.id)), ...chatConfirmationTasks],
expectedMatrixSize: 12,
definitionMetadata: { version: 6, instructions: "production", grading: "card-message-provenance-before-child-creation", unansweredQuestionComposer: "history-card-only-including-fresh-dismissal", completionObservation: "120-seconds", scheduling: "explicit-only" },
},
{
id: "completion-updates", label: "Delegated Completion Updates", manualOnly: true,
description: "Qualify completion delivery in onboarding and idle, busy, multiple-task, and restart Agent Chat handoffs.",
+2
View File
@@ -73,3 +73,5 @@ export const chatCompletionTasks = buildChatTasks([
["handoff-completion-multiple", "Report multiple delegated results as they finish", 7],
["handoff-completion-restart", "Recover pending completion delivery across a server restart", 5],
]).map(task => ({ ...task, minimumExpectedRunCount: 2 }));
export const chatConfirmationTasks = buildChatTasks([["confirmation-ambiguous", "Clarify ambiguous approval, then resolve only the chosen cards", 4], ["unanswered-question-return", "Move on, reopen a historical question, and deliver the late answer", 3]]);
+7 -1
View File
@@ -1,3 +1,4 @@
import { runAmbiguousConfirmationReply, runUnansweredQuestionReturn } from "./confirmation-replies.js";
import { expect, type Page } from "@playwright/test";
import { pollUntil, type RunnerApi } from "./api.js";
import type {
@@ -344,6 +345,7 @@ export interface ChatFlowInput {
observe: (issue: ChatIssue, runs: ChatRun[]) => void;
capture: (id: string, label: string, file: string) => Promise<void>;
evidence: (name: string, data: unknown) => Promise<void>;
check?: (id: string, passed: boolean, detail: string) => void;
}
export async function runChatFlow(input: ChatFlowInput) {
const { page, api, fixtures: f, execution, nonce } = input;
@@ -430,7 +432,11 @@ export async function runChatFlow(input: ChatFlowInput) {
expect(await api.get(chatPath)).toBeNull();
expect(await allRuns()).toHaveLength(0);
if (caseId.startsWith("handoff-completion-")) {
if (caseId === "confirmation-ambiguous") {
await runAmbiguousConfirmationReply({ input, issue: () => issue!, idle, comments });
} else if (caseId === "unanswered-question-return") {
await runUnansweredQuestionReturn({ input, issue: () => issue!, idle, comments });
} else if (caseId.startsWith("handoff-completion-")) {
await runChatCompletionUpdate({ input, marker, allRuns, issue: () => issue!,
refreshIssue: async () => { issue = await api.get<ChatIssue>(chatPath); if (issue) input.observe(issue, await allRuns()); } });
} else if (execution.suite.id === "agent-chat-qualification") {
+15 -1
View File
@@ -1,9 +1,23 @@
import { describe, expect, it } from "vitest";
import { COMPLETION_QUALITY_CONFIG, completionQualityControls, completionQualityStatus, completionQualityRequest, judgeCompletionQuality, reserveCompletionQuality, validateCompletionQuality } from "./completion-quality.js";
import { runsCompletionUpdateProbe, COMPLETION_QUALITY_CONFIG, completionQualityControls, completionQualityStatus, completionQualityRequest, judgeCompletionQuality, reserveCompletionQuality, validateCompletionQuality } from "./completion-quality.js";
import { runnerMatrix } from "./catalog.js";
const observation = { sourceId: "chat", marker: "REF", worker: { id: "task", title: "Welcome", status: "done", completedAt: "2026-09-01T00:00:00Z" }, documents: [{ id: "doc", issueId: "task", key: "welcome", body: "Welcome to the garden. Meet at 10:30." }], comments: [{ id: "reply", issueId: "chat", authorAgentId: "agent", createdAt: "2026-09-01T00:01:00Z", body: "The note is ready at /issues/task", createdByRunId: "run" }], runs: [] };
const criteria = Object.keys(COMPLETION_QUALITY_CONFIG.rubric).map(id => ({ id, passed: true, rationale: "Supported by the saved note and reply", evidenceIds: ["reply", "doc"] }));
const reports = [{ replyId: "reply", rationale: "Reports the completed task", completedTaskIdsReferenced: ["task"], resultAccessTaskIds: ["task"], correctsReplyIds: [] as string[] }];
describe("completion semantic qualification", () => {
it.each(["runner-codex", "runner-acpx-claude"])("requires a completion judge only for accepted work on %s", profile => {
const cells = runnerMatrix.filter(e => e.suite.id === "confirmation-replies" && e.profile.id === profile);
expect(cells).toHaveLength(6);
expect(cells.filter(runsCompletionUpdateProbe).map(e => e.task.id).sort()).toEqual([
"interview-plan-accept", "task-card-accept", "task-reply-accept",
]);
});
it("retains the completion suite's gate and excludes ordinary chat journeys", () => {
const completion = runnerMatrix.filter(e => e.suite.id === "completion-updates");
expect(completion).toHaveLength(10);
expect(completion.every(runsCompletionUpdateProbe)).toBe(true);
expect(runnerMatrix.filter(e => !["confirmation-replies", "completion-updates"].includes(e.suite.id)).some(runsCompletionUpdateProbe)).toBe(false);
});
it("uses a pinned no-tool judge and separates untrusted evidence from instructions", () => {
const request = completionQualityRequest(observation);
expect(request.model).toBe(COMPLETION_QUALITY_CONFIG.model); expect(request).not.toHaveProperty("tools");
+7
View File
@@ -4,6 +4,13 @@ import { createHash } from "node:crypto";
import { FIRST_TASK_JUDGE_CONFIG } from "./first-task-quality.js";
import type { CompletionObservation } from "./completion-updates.js";
/** Only journeys that finish delegated work can require a completion report. */
export function runsCompletionUpdateProbe(execution: { suite: { id: string }; task: { id: string; flow: string } }) {
return execution.suite.id === "completion-updates"
|| (execution.suite.id === "confirmation-replies" && execution.task.flow === "first_task"
&& execution.task.id !== "reject-no-execution");
}
export const COMPLETION_QUALITY_CONFIG = {
version: 14, resultAccessEvidence: "observed-rendered-task-links", duplicateRule: "per-reply-completed-task-references-new-access-or-correction", model: FIRST_TASK_JUDGE_CONFIG.model, temperature: 0, maxOutputTokens: 1600,
correctionRule: "Grade the final corrected position of the conversation. If a later reply explicitly corrects an earlier stale or inaccurate statement and provides the result without a new user request, the corrected statement replaces the earlier statement for ALL three criteria. Do not fail a criterion solely because the corrected earlier reply failed it. Uncorrected false claims still fail.",
@@ -0,0 +1,171 @@
import { gradeUnansweredQuestion, type UnansweredQuestionEvidence } from "./confirmation-replies.js";
import { describe, expect, it, vi } from "vitest";
import { gradeConfirmationReply, assertAmbiguousReplyUnresolved, ambiguousConfirmationFixtures } from "./confirmation-replies.js";
import { firstTaskScenario } from "./first-task-cases.js";
import type { FirstTaskEvidence } from "./first-task-scoring.js";
import { provisionFirstTaskFixtures } from "./first-task-fixtures.js";
import { runnerMatrix } from "./catalog.js";
import { parseRunnerSelectors, selectRunnerExecutions } from "./selectors.js";
function recording(rejected = false): FirstTaskEvidence {
const caseId = rejected ? "reject-no-execution" : "task-reply-accept";
const scenario = firstTaskScenario(caseId, "oracle");
const card = { id: "proposal", kind: "request_confirmation", status: "pending" };
const base = { issueId: "onboard", tasks: [{ id: "onboard" }], agents: [], comments: [], interactions: [card], documents: [], runs: [] };
const message = { id: "user-answer", authorUserId: "user", body: rejected ? scenario.rejection : scenario.acceptance, createdAt: "2026-09-29T00:01:00Z" };
return { caseId, nonce: "oracle", onboardingIssueId: "onboard", agentId: "planner", initialTaskIds: ["onboard"], instructions: [], configuredModel: null, observedModels: [], checks: [], checkpoints: [
{ ...base, id: "before", phase: "response", at: "2026-09-29T00:00:00Z" },
{ ...base, id: "decision", phase: rejected ? "rejected" : "accepted", at: message.createdAt, comments: [message] },
{ ...base, id: "after", phase: "finished", at: "2026-09-29T00:03:00Z", comments: [message], runs: [{ id: "resolver", agentId: "planner" }],
tasks: [{ id: "onboard" }, ...rejected ? [] : [{ id: "child", createdAt: "2026-09-29T00:02:01Z" }]],
interactions: [{ ...card, status: rejected ? "rejected" : "accepted", resolvedAt: "2026-09-29T00:02:00Z", resolvedByAgentId: "planner", resolvedByRunId: "resolver", result: { outcome: rejected ? "rejected" : "accepted", commentId: message.id } }] },
] };
}
describe("confirmation-reply independent oracle", () => {
it("uses valid public confirmation fixtures without treating checkbox defaults as consent", () => {
expect(ambiguousConfirmationFixtures).toHaveLength(2);
const checkbox = ambiguousConfirmationFixtures[1]!;
expect(checkbox.kind).toBe("request_checkbox_confirmation");
expect(checkbox.payload).toMatchObject({ minSelected: 1, defaultSelectedOptionIds: ["poster"] });
});
it.each([false, true])("accepts the recorded decision with provenance (rejection=%s)", rejected => {
expect(gradeConfirmationReply(recording(rejected)).every(c => c.passed)).toBe(true);
});
it.each(["pending", "expired", "wrong-message", "wrong-agent", "no-run", "wrong-run-agent", "missing-time", "no-card", "too-late", "multiple-pending", "opposite"])("rejects plausible wrong outcome: %s", kind => {
const e = recording(), last = e.checkpoints.at(-1)!, card = last.interactions[0]!;
if (kind === "pending" || kind === "expired") card.status = kind;
if (kind === "wrong-message") card.result.commentId = "another-message";
if (kind === "wrong-agent") card.resolvedByAgentId = "someone-else";
if (kind === "no-run") last.runs = [];
if (kind === "wrong-run-agent") last.runs[0]!.agentId = "someone-else";
if (kind === "missing-time") delete card.resolvedAt;
if (kind === "no-card") e.checkpoints[0]!.interactions = [];
if (kind === "too-late") card.resolvedAt = "2026-09-29T00:02:02Z";
if (kind === "multiple-pending") e.checkpoints[0]!.interactions.push({ id: "another-proposal", kind: "request_confirmation", status: "pending" });
if (kind === "opposite") card.result.outcome = "rejected";
expect(gradeConfirmationReply(e).some(c => !c.passed)).toBe(true);
});
it("does not treat checkbox defaults as a persisted selection", () => {
const e = recording(), card = e.checkpoints.at(-1)!.interactions[0]!;
card.kind = "request_checkbox_confirmation";
expect(gradeConfirmationReply(e).every(c => c.passed)).toBe(false);
card.result.selectedOptionIds = ["welcome-note"];
expect(gradeConfirmationReply(e).every(c => c.passed)).toBe(true);
});
it("requires clarification and no effects for ambiguous approval", () => {
const e = { cards: [{ id: "a", status: "pending" }, { id: "b", status: "pending" }], originalIds: ["a", "b"], tasks: [], reply: "Which proposal do you mean: the note or the poster?" };
expect(() => assertAmbiguousReplyUnresolved(e)).not.toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, cards: [{ id: "a", status: "accepted" }, e.cards[1]!] })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, tasks: [{ id: "unauthorized-task" }] })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "Both proposals are approved." })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "What is the deadline?" })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "Which tone should the welcome note and poster use?" })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "Would you like the welcome note and poster to be formal?" })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "Do you want the welcome note and poster by Friday?" })).toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply: "Do you mean the welcome note or the poster?" })).not.toThrow();
for (const reply of ["Which proposal do you mean?\n- Welcome note\n- Poster", "Which proposal do you mean?\n\n* Welcome note\n* Poster", "Which one\nshould I proceed with: the welcome note or poster?",
"Should I proceed with the welcome note or the poster?", "Do you mean the note or the poster, or both?", "Should we start the poster or the welcome note, or both?", "Which garden-club item should we move forward with: the note or poster?",
"Which pending proposal does yes approve?\n- Welcome note\n- Poster",
"Which project(s) should I start on now?\n- Welcome note only\n- Poster only\n- Both",
'There are two pending proposals: the **welcome note** and the **poster**. Your yes does not specify which. Could you clarify:\n\n- Just the welcome note?\n- Just the poster?\n- Both?\n\nI will record your decision.',
]) {
expect(() => assertAmbiguousReplyUnresolved({ ...e, reply })).not.toThrow();
}
});
it("accepts a current structured clarification and rejects stale or unrelated question cards", () => {
const question = { id: "question", kind: "ask_user_questions", status: "pending", createdByAgentId: "planner",
originCommentIds: ["ambiguous-answer"], payload: { questions: [{ prompt: "Which item(s) does go ahead authorize me to plan?", options: [{ label: "Welcome note only" }, { label: "Poster only" }] }] } };
const evidence = { cards: [{ id: "a", status: "pending" }, { id: "b", status: "pending" }, question], originalIds: ["a", "b"], tasks: [], reply: "",
agentId: "planner", answerId: "ambiguous-answer" };
expect(() => assertAmbiguousReplyUnresolved(evidence)).not.toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...evidence, cards: [...evidence.cards.slice(0, 2), { ...question,
payload: { questions: [{ prompt: "Which garden club item should I move forward with?", options: [{ label: "Welcome note" }, { label: "Poster" }] }] },
}] })).not.toThrow();
const native = { ...question, payload: { questionSet: { description: "Which item(s) would you like me to start developing a plan for?",
questions: [{ prompt: "Scope", options: [{ label: "Welcome note only" }, { label: "Poster only" }, { label: "Both" }] }] } } };
expect(() => assertAmbiguousReplyUnresolved({ ...evidence, cards: [...evidence.cards.slice(0, 2), native] })).not.toThrow();
expect(() => assertAmbiguousReplyUnresolved({ ...evidence, cards: [...evidence.cards.slice(0, 2), { ...native,
payload: { questionSet: { ...native.payload.questionSet, description: "What is the deadline?" } },
}] })).toThrow();
for (const patch of [{ originCommentIds: ["old-answer"] }, { createdByAgentId: "other-agent" }, { status: "answered" }, { payload: { questions: [] } }, { payload: { questions: [{ prompt: "What is the deadline?" }] } }, { payload: { questions: [{ prompt: "Which tone should the welcome note and poster use?", options: [{ label: "Warm" }, { label: "Formal" }] }] } }]) {
expect(() => assertAmbiguousReplyUnresolved({ ...evidence, cards: [...evidence.cards.slice(0, 2), { ...question, ...patch }] })).toThrow();
}
});
it.each([
"Which part of the welcome note or poster should I revise?",
"Which font should I choose for the welcome note or poster?",
"Which color option should the welcome note or poster use?",
"Which of the fonts should I use for the welcome note or poster?",
"Should I select a deadline for the welcome note or poster?",
"Pick a font for the welcome note or poster?",
"Would you like the welcome note or poster to be formal?",
"Do you want the welcome note or poster by Friday?",
"Could you clarify:\n- What font for the welcome note?\n- What font for the poster?",
])("does not mistake a detail question for proposal selection: %s", prompt => {
const cards = [{ id: "a", status: "pending" }, { id: "b", status: "pending" }];
const evidence = { cards, originalIds: ["a", "b"], tasks: [], reply: prompt, agentId: "planner", answerId: "answer" };
expect(() => assertAmbiguousReplyUnresolved(evidence)).toThrow();
// Named alternatives do not turn an unrelated field into a scope question.
expect(() => assertAmbiguousReplyUnresolved({ ...evidence, reply: "", cards: [...cards, {
id: "question", kind: "ask_user_questions", status: "pending", createdByAgentId: "planner", originCommentIds: ["answer"],
payload: { questions: [{ prompt, options: [{ label: "Welcome note" }, { label: "Poster" }] }] },
}] })).toThrow();
});
it.each(runnerMatrix.filter(e => e.suite.id === "confirmation-replies" && e.task.flow === "first_task"))("provisions the real onboarding fixture contract for $id", async execution => {
const get = vi.fn().mockResolvedValue([{ id: "local", driver: "local" }]);
const postSensitive = vi.fn().mockResolvedValue({ id: "secret" });
const credential = execution.profile.credential;
const fixtures = await provisionFirstTaskFixtures({ api: { get, postSensitive }, execution, nonce: "fixture",
company: { id: "company", name: "Garden" }, credentials: { [credential]: "test-credential" } });
expect(get).toHaveBeenCalledExactlyOnceWith("/api/companies/company/environments?driver=local");
expect(postSensitive).toHaveBeenCalledExactlyOnceWith("/api/companies/company/secrets", expect.objectContaining({ key: credential }));
expect(fixtures.agent.id).toBe(""); // The real wizard must still create the agent.
expect(JSON.stringify(fixtures)).not.toContain("test-credential");
});
it.each(["Select every proposal you want to approve.", "Choose the proposals to approve."])("accepts explicit proposal selection in an imperative form: %s", prompt => {
expect(() => assertAmbiguousReplyUnresolved({ cards: [{ id: "a", status: "pending" }, { id: "b", status: "pending" },
{ id: "question", kind: "ask_user_questions", status: "pending", createdByAgentId: "agent", originCommentIds: ["reply"],
payload: { questions: [{ prompt, options: [{ label: "Welcome note" }, { label: "Poster" }] }] } }],
originalIds: ["a", "b"], tasks: [], reply: "", agentId: "agent", answerId: "reply" })).not.toThrow();
});
it("rejects an imperative about fonts even with named proposal options", () => {
expect(() => assertAmbiguousReplyUnresolved({ cards: [{ id: "a", status: "pending" }, { id: "b", status: "pending" },
{ id: "question", kind: "ask_user_questions", status: "pending", createdByAgentId: "agent", originCommentIds: ["reply"],
payload: { questions: [{ prompt: "Select every font you want to approve.", options: [{ label: "Welcome note" }, { label: "Poster" }] }] } }],
originalIds: ["a", "b"], tasks: [], reply: "", agentId: "agent", answerId: "reply" })).toThrow();
});
it("selects exactly twelve explicit-only cases using production native profiles", () => {
const cells = runnerMatrix.filter(e => e.suite.id === "confirmation-replies");
expect(cells).toHaveLength(12);
expect(new Set(cells.map(c => c.profile.id))).toEqual(new Set(["runner-codex", "runner-acpx-claude"]));
expect(new Set(cells.map(c => c.task.id)).size).toBe(6);
expect(selectRunnerExecutions(parseRunnerSelectors(["--suite", "confirmation-replies"]))).toHaveLength(12);
expect(selectRunnerExecutions(parseRunnerSelectors(["--all"])).some(c => c.suite.id === "confirmation-replies")).toBe(false);
});
});
describe("unanswered question workflow oracle", () => {
const good = () => ({
original: { id: "color", status: "pending", payload: { questions: [{ options: [{ id: "blue", label: "Blue" }, { id: "green", label: "Green" }] }] } },
afterMove: { id: "color", status: "pending", result: null, resolvedAt: null },
afterAnswer: { id: "color", status: "answered", resolvedByUserId: "user", resolvedAt: "2026-09-01T12:02:00Z", result: { answers: [{ optionIds: ["blue"] }] } },
unrelatedComment: { id: "unrelated-comment", authorUserId: "user", createdAt: "2026-09-01T12:00:00Z" },
unrelatedReply: { id: "unrelated-reply", authorAgentId: "agent", body: "Paris.", createdAt: "2026-09-01T12:01:00Z" },
lateReply: { id: "late-reply", authorAgentId: "agent", createdByRunId: "later-run", body: "Blue it is.", createdAt: "2026-09-01T12:03:00Z" }, taskCount: 0,
});
it("accepts the independently saved workflow", () => expect(gradeUnansweredQuestion(good()).every(check => check.passed)).toBe(true));
it.each(["expired", "wrong-question", "missing-reply", "stale-reply", "wrong-answer", "missing-late-reply", "old-acknowledgement", "invented-work"])("rejects %s", kind => {
const evidence: UnansweredQuestionEvidence = good();
if (kind === "expired") evidence.afterMove.status = "expired";
if (kind === "wrong-question") evidence.afterAnswer.id = "other-question";
if (kind === "missing-reply") evidence.unrelatedReply = undefined;
if (kind === "stale-reply") evidence.unrelatedReply!.createdAt = "2026-08-01T12:00:00Z";
if (kind === "wrong-answer") evidence.afterAnswer.result.answers[0].optionIds = ["green"];
if (kind === "missing-late-reply") evidence.lateReply = undefined;
if (kind === "old-acknowledgement") evidence.lateReply!.createdAt = "2026-08-01T12:00:00Z";
if (kind === "invented-work") evidence.taskCount = 1;
expect(gradeUnansweredQuestion(evidence).some(check => !check.passed)).toBe(true);
});
});
+260
View File
@@ -0,0 +1,260 @@
import { expect, type Page } from "@playwright/test";
import { createIssueThreadInteractionSchema } from "../../packages/shared/src/validators/issue.js";
import type { FirstTaskEvidence, FirstTaskCheck, Row } from "./first-task-scoring.js";
import { firstTaskScenario } from "./first-task-cases.js";
import { sendChatMessage, type ChatFlowInput, type ChatIssue } from "./chat-flow.js";
// Validate public fixture requests before spending provider turns. The selected
// checkbox default deliberately proves that a default is not user consent.
export const ambiguousConfirmationFixtures = [
createIssueThreadInteractionSchema.parse({ kind: "request_confirmation", title: "Welcome note proposal", continuationPolicy: "none", payload: { version: 1, prompt: "Approve writing the welcome note?" } }),
createIssueThreadInteractionSchema.parse({ kind: "request_checkbox_confirmation", title: "Poster proposal", continuationPolicy: "none", payload: { version: 1, prompt: "Approve the poster?", options: [{ id: "poster", label: "Create a garden club poster" }], defaultSelectedOptionIds: ["poster"], minSelected: 1 } }),
];
/** Check the saved user-facing receipt, including its original proposal. */
export async function assertConfirmationReceipt(page: Page, card: Row) {
expect(["accepted", "rejected"]).toContain(card.status);
const receipt = page.getByTestId("task-chat-interaction-receipt").filter({ hasText: card.payload.prompt });
const summary = card.status === "accepted"
? card.kind === "request_confirmation" ? "Confirmed request"
: `Confirmed ${card.result.selectedOptionIds.length} of ${card.payload.options.length} options`
: card.kind === "request_checkbox_confirmation" ? "Declined selection"
: card.payload.rejectLabel?.trim() ? `Selected “${card.payload.rejectLabel.trim()}”` : "Declined request";
await expect(receipt.locator("summary")).toHaveText(summary);
await receipt.locator("summary").click();
await expect(receipt.getByText(card.payload.prompt, { exact: true })).toBeVisible();
if (card.kind === "request_checkbox_confirmation" && card.status === "accepted") {
for (const option of card.payload.options.filter((option: Row) => card.result.selectedOptionIds.includes(option.id))) {
await expect(receipt).toContainText(option.label);
}
}
}
/** The independent oracle checks recorded card state, provenance and ordering, not agent claims. */
export function gradeConfirmationReply(e: FirstTaskEvidence): FirstTaskCheck[] {
const expected = e.caseId === "reject-no-execution" ? "rejected" : "accepted";
const decision = e.checkpoints.find(c => c.phase === expected);
const last = e.checkpoints.at(-1);
const scenario = firstTaskScenario(e.caseId, e.nonce);
const reply = decision?.comments.find(c => !c.authorAgentId && c.authorUserId && c.body === (expected === "accepted" ? scenario.acceptance : scenario.rejection));
const before = e.checkpoints.filter(c => ["response", "clarified", "revised"].includes(c.phase)).at(-1);
const pending = before?.interactions.filter(c => c.status === "pending" && ["request_confirmation", "request_checkbox_confirmation"].includes(c.kind)) ?? [];
const card = pending.length === 1 ? last?.interactions.find(c => c.id === pending[0]!.id) : undefined;
const resolvedAt = Date.parse(card?.resolvedAt ?? "");
const run = last?.runs.find(r => r.id === card?.resolvedByRunId);
const valid = Boolean(reply && card && card.status === expected && card.result?.outcome === expected
&& card.result?.commentId === reply.id && card.resolvedByAgentId === e.agentId && run?.agentId === e.agentId
&& Number.isFinite(resolvedAt) && resolvedAt >= Date.parse(reply.createdAt ?? decision!.at)
&& (card.kind !== "request_checkbox_confirmation" || expected === "rejected" || Array.isArray(card.result?.selectedOptionIds)));
const children = last?.tasks.filter(t => !e.initialTaskIds.includes(t.id) && t.id !== e.onboardingIssueId) ?? [];
return [{ id: "conversation-resolves-confirmation", passed: valid,
evidence: [before?.id, decision?.id, last?.id].filter((v): v is string => Boolean(v)),
detail: "The exact pending proposal records the user's reply and the resolving agent run as an accepted/rejected outcome" },
{ id: "resolution-before-execution", passed: valid && children.every(t => Date.parse(t.createdAt ?? "") >= resolvedAt),
evidence: last ? [last.id] : [], detail: "Structured approval is persisted before any child task is created" }];
}
// This fixture asks the user to choose between two named proposals. A question
// about tone, deadline, or another detail does not disambiguate that approval.
function asksWhichProposal(body: string, options: string[] = []): boolean {
const text = body.replace(/^(\s*)\*\s/gm, "$1- ").replace(/[*_`]/g, "");
// Keep the interrogative's object explicit. Arbitrary intervening words (or
// bare "of") also match "which part of" and "which font option", which ask
// about details rather than choosing a proposal. "Garden club" is this
// fixture's named scope, not a wildcard for any modifier.
const scopedChoice = /\bwhich\s+(?:(?:pending|current|proposed|available|separate|two)\s+|garden[\s-]+club\s+)?(?:one|ones|item|items|proposal|proposals|project|projects|task|tasks|option|options)\b/i;
// Structured forms can ask imperatively, without a question mark. Require
// both named alternatives plus an explicit proposal-selection instruction.
const namedOptions = options.some(option => /\b(?:welcome\s+)?note\b/i.test(option))
&& options.some(option => /\bposter\b/i.test(option));
if (namedOptions && /\b(?:select|choose|pick)\s+(?:(?:all|every|each|one|the|any)(?:\s+of\s+the)?\s+)?proposals?\s+(?:you\s+(?:want|wish)\s+to\s+|to\s+)?(?:approve|authorize)\b/i.test(text)) return true;
const proposalName = "(?:the\\s+)?(?:(?:welcome\\s+)?note|poster)(?:\\s+proposal)?";
const choiceIntent = "(?:(?:do|did|would)\\s+you\\s+(?:mean|want|prefer|like)|should\\s+(?:i|we)\\s+(?:start|proceed\\s+with))";
const directChoice = new RegExp(`\\b${choiceIntent}\\s+${proposalName}\\s+or\\s+${proposalName}(?:,?\\s+or\\s+both)?\\s*\\?`, "i");
const clarificationList = text.match(/\b(?:could|can|would)\s+you\s+(?:please\s+)?clarify\s*:\s*\n((?:\s*[-+]\s+(?:just\s+)?(?:the\s+)?(?:(?:welcome\s+)?note|poster|both)(?:\s+only)?\??[ \t]*(?:\n|$)){2,})/i)?.[1];
if (clarificationList && /\bnote\b/i.test(clarificationList) && /\bposter\b/i.test(clarificationList)) return true;
return [...text.matchAll(/[^?]*\?/g)].some(match => {
const question = match[0];
const listedOptions: string[] = [];
for (const line of text.slice(match.index! + question.length).trimStart().split("\n")) {
const item = line.match(/^\s*(?:[-+]|\d+[.)])\s+(.+)$/);
if (!item) break;
listedOptions.push(item[1]!);
}
const choices = [...options, ...listedOptions];
const alternatives = [question, ...choices].join(" ");
const namesBoth = /\b(?:welcome\s+)?note\b/i.test(alternatives) && /\bposter\b/i.test(alternatives);
const choosesNamedScope = scopedChoice.test(question);
return namesBoth && (choosesNamedScope || directChoice.test(question));
});
}
type AmbiguousReplyEvidence = { cards: Row[]; originalIds: string[]; tasks: Row[]; reply: string; agentId?: string; answerId?: string };
function assertAmbiguousStateUnchanged(input: AmbiguousReplyEvidence) {
expect(input.originalIds).toHaveLength(2);
expect(input.cards.filter(c => input.originalIds.includes(c.id)).map(c => c.status)).toEqual(["pending", "pending"]);
expect(input.tasks).toHaveLength(0);
}
export function assertAmbiguousReplyUnresolved(input: AmbiguousReplyEvidence) {
assertAmbiguousStateUnchanged(input);
const questionCard = input.cards.some(card => card.kind === "ask_user_questions" && card.status === "pending"
&& input.agentId && card.createdByAgentId === input.agentId && input.answerId && card.originCommentIds?.includes(input.answerId)
&& (card.payload?.questionSet?.questions ?? card.payload?.questions ?? []).some((question: Row) => asksWhichProposal(
[card.payload?.questionSet?.description, question.prompt].filter(Boolean).join("\n"),
(question.options ?? []).map((option: Row) => option.label ?? ""))));
expect(asksWhichProposal(input.reply) || questionCard, "Ask which proposal the ambiguous reply refers to").toBe(true);
}
export async function runAmbiguousConfirmationReply(context: {
input: ChatFlowInput; issue(): ChatIssue; idle(count: number): Promise<void>; comments(): Promise<Row[]>;
}) {
const { input } = context;
const { api, page } = input;
await sendChatMessage(page, "I am considering a welcome note and a poster for our garden club. Neither is approved. Just acknowledge for now; do not create tasks or start either one.");
await context.idle(1);
const issue = context.issue();
const cardsPath = `/api/issues/${issue.id}/interactions`;
// Public board-created cards of different kinds retain independent pending
// decisions, rather than depending on a model producing two simultaneous tools.
const note = await api.post<Row>(cardsPath, ambiguousConfirmationFixtures[0]);
const poster = await api.post<Row>(cardsPath, ambiguousConfirmationFixtures[1]);
const before = await api.get<Row[]>(cardsPath);
expect(before.filter(c => c.status === "pending").map(c => c.id).sort()).toEqual([note.id, poster.id].sort());
await input.evidence("confirmation-before-reply.json", { issueId: issue.id, cards: before });
await page.reload({ waitUntil: "domcontentloaded" });
await expect(page.getByTestId("task-chat-composer-takeover")).toBeVisible();
await expect(page.getByTestId("task-chat-composer-takeover")).toContainText("Approve");
await input.capture("confirmation-pending", "Two independent pending proposals", "confirmation-pending.png");
await sendChatMessage(page, "Yes, go ahead.");
await context.idle(2);
const afterReply = await context.comments();
const ambiguousAnswer = afterReply.findLast(c => !c.authorAgentId && c.authorUserId && c.body === "Yes, go ahead.");
expect(ambiguousAnswer, "Persist the exact ambiguous user reply").toBeTruthy();
const e = { cards: await api.get<Row[]>(cardsPath), originalIds: [note.id, poster.id],
tasks: await api.get<Row[]>(`/api/companies/${input.fixtures.company.id}/issues`),
agentId: input.fixtures.agent.id, answerId: ambiguousAnswer!.id,
reply: afterReply.filter(c => c.authorAgentId === input.fixtures.agent.id && c.createdAt >= ambiguousAnswer!.createdAt).at(-1)?.body ?? "" };
await input.evidence("confirmation-ambiguous.json", e);
// Stop immediately for unauthorized effects. Retain the complete later
// decision/reload evidence before grading clarification wording: a new valid
// wording must not force another paid run just to observe those later steps.
assertAmbiguousStateUnchanged(e);
await sendChatMessage(page, "I approve only the welcome note proposal. Record that decision, but do not start execution or create a task yet. Leave the poster proposal pending.");
await context.idle(3);
let cards = await api.get<Row[]>(cardsPath);
const yes = (await context.comments()).filter(c => !c.authorAgentId && c.body.includes("I approve only")).at(-1)!;
expect(cards.find(c => c.id === note.id)).toMatchObject({ status: "accepted", result: { commentId: yes.id } });
expect(cards.find(c => c.id === poster.id)?.status).toBe("pending");
await sendChatMessage(page, "No, do not proceed with the poster. Reject that proposal. We are still not starting any work.");
await context.idle(4);
cards = await api.get<Row[]>(cardsPath);
const no = (await context.comments()).filter(c => !c.authorAgentId && c.body.includes("No, do not proceed")).at(-1)!;
expect(cards.find(c => c.id === poster.id)).toMatchObject({ status: "rejected", result: { commentId: no.id } });
expect(await api.get<Row[]>(`/api/companies/${input.fixtures.company.id}/issues`)).toHaveLength(0);
await page.reload({ waitUntil: "domcontentloaded" });
for (const id of [note.id, poster.id]) {
await assertConfirmationReceipt(page, cards.find(card => card.id === id)!);
}
await input.evidence("confirmation-decisions.json", { cards, comments: await context.comments(), activity: await api.get(`/api/issues/${issue.id}/activity`) });
await input.capture("confirmation-decisions", "Conversational approval and rejection persisted", "confirmation-decisions.png");
input.check?.("confirmation-decisions-persisted", true, "Exact user comments resolve the chosen cards; reloaded receipts show approval and rejection, with no tasks created");
assertAmbiguousReplyUnresolved(e);
input.check?.("ambiguous-approval-not-assumed", true, "The agent asks which proposal is intended and leaves both decisions pending until explicitly answered");
}
export interface UnansweredQuestionEvidence {
original: Row;
afterMove: Row;
afterAnswer: Row;
unrelatedComment: Row | undefined;
unrelatedReply: Row | undefined;
lateReply: Row | undefined;
taskCount: number;
}
export function gradeUnansweredQuestion(e: UnansweredQuestionEvidence) {
const originalOptions: Row[] = (e.original.payload?.questionSet?.questions ?? e.original.payload?.questions ?? [])[0]?.options ?? [];
const blueId = originalOptions.find(option => option.label.toLowerCase() === "blue")?.id;
const answer = e.afterAnswer.result?.answers?.[0];
return [
{ id: "unanswered-preserved", passed: Boolean(e.original.id && e.afterMove.id === e.original.id && e.original.status === "pending"
&& e.afterMove.status === "pending" && !e.afterMove.result && !e.afterMove.resolvedAt), detail: "Moving on preserves the exact unanswered question without inventing a resolution" },
{ id: "unrelated-turn-completed", passed: Boolean(e.unrelatedComment?.authorUserId && e.unrelatedReply?.authorAgentId
&& Date.parse(e.unrelatedReply.createdAt) >= Date.parse(e.unrelatedComment.createdAt)
&& /\bParis\b/i.test(e.unrelatedReply.body)), detail: "The agent answers the new message while its earlier question remains pending" },
{ id: "historical-answer-recorded", passed: Boolean(e.afterAnswer.id === e.original.id && e.afterAnswer.status === "answered"
&& e.afterAnswer.resolvedByUserId && blueId && answer?.optionIds?.length === 1 && answer.optionIds[0] === blueId), detail: "The reopened original question records the user's Blue selection" },
{ id: "historical-answer-delivered", passed: Boolean(e.lateReply?.authorAgentId && e.lateReply.createdByRunId
&& Date.parse(e.lateReply.createdAt) >= Date.parse(e.afterAnswer.resolvedAt ?? "") && /\bblue\b/i.test(e.lateReply.body)), detail: "A later agent turn acknowledges the saved answer after the original turn ended" },
{ id: "no-unrequested-work", passed: e.taskCount === 0, detail: "No tasks are created by the question, unrelated reply, or late answer" },
];
}
export async function runUnansweredQuestionReturn(context: {
input: ChatFlowInput; issue(): ChatIssue; idle(count: number): Promise<void>; comments(): Promise<Row[]>;
}) {
const { input } = context;
const { api, page } = input;
await sendChatMessage(page, "Help me choose a color for a garden club welcome note. Ask me one interactive question using Paperclip's question card: Which color should the welcome note use? Offer Blue and Green. When I eventually answer, acknowledge my chosen color in chat. For now only ask; do not create tasks or write the note.");
await context.idle(1);
const path = `/api/issues/${context.issue().id}/interactions`;
const original = (await api.get<Row[]>(path)).filter(card => card.kind === "ask_user_questions" && card.status === "pending").at(-1);
expect(original, "Agent creates an actual saved question").toBeTruthy();
const questionRow = () => page.getByTestId("task-chat-unanswered-question").filter({ hasText: /color/i });
await expect(page.getByTestId("task-chat-composer-takeover")).toBeVisible();
await expect(questionRow()).toBeVisible();
await input.capture("question-asked", "Original question awaiting an answer", "question-asked.png");
await page.getByTestId("task-chat-composer-takeover").getByRole("button", { name: "Cancel", exact: true }).click();
await expect(questionRow()).toBeVisible();
await expect(page.getByTestId("task-chat-composer-takeover")).toHaveCount(0);
await expect(page.getByTestId("task-chat-pending-input-indicator")).toHaveCount(0);
await input.capture("question-dismissed", "Dismissing a fresh question leaves only its history card", "question-dismissed.png");
await questionRow().click();
await expect(page.getByTestId("task-chat-composer-takeover")).toBeVisible();
const unrelated = "Leave that color question unanswered for now. What is the capital of France? Answer that in chat; do not create tasks.";
await sendChatMessage(page, unrelated);
await context.idle(2);
const afterMove = (await api.get<Row[]>(path)).find(card => card.id === original!.id)!;
const moveComments = await context.comments();
const unrelatedComment = moveComments.findLast(comment => comment.authorUserId && comment.body === unrelated);
const unrelatedReply = moveComments.findLast(comment => comment.authorAgentId === input.fixtures.agent.id
&& unrelatedComment && comment.createdAt >= unrelatedComment.createdAt);
expect(afterMove).toMatchObject({ status: "pending", resolvedAt: null, result: null });
expect(unrelatedReply?.body).toMatch(/\bParis\b/i);
await expect(page.getByTestId("task-chat-pending-input-indicator")).toHaveCount(0);
await page.reload({ waitUntil: "domcontentloaded" });
await expect(questionRow()).toBeVisible();
await expect(page.getByTestId("task-chat-composer-input")).toBeVisible();
await expect(page.getByTestId("task-chat-composer-takeover")).toHaveCount(0);
await expect(page.getByTestId("task-chat-pending-input-indicator")).toHaveCount(0);
await input.capture("question-left-unanswered", "After reload: question stays in history without a duplicate composer reminder", "question-left-unanswered.png");
await questionRow().click();
const form = page.getByTestId("task-chat-composer-takeover");
await expect(form).toContainText(/color/i);
await input.capture("question-reopened", "The original question can be reopened from history", "question-reopened.png");
await form.getByRole("radio", { name: "Blue", exact: true }).click();
await form.getByRole("button", { name: /^(Send|Submit) answers$/ }).click();
await context.idle(3);
const afterAnswer = (await api.get<Row[]>(path)).find(card => card.id === original!.id)!;
const lateReply = (await context.comments()).findLast(comment => comment.authorAgentId === input.fixtures.agent.id
&& comment.createdAt >= afterAnswer.resolvedAt);
const evidence = { original: original!, afterMove, afterAnswer, unrelatedComment, unrelatedReply, lateReply,
taskCount: (await api.get<Row[]>(`/api/companies/${input.fixtures.company.id}/issues`)).length };
await input.evidence("unanswered-question.json", evidence);
for (const check of gradeUnansweredQuestion(evidence)) {
input.check?.(check.id, check.passed, check.detail);
expect(check.passed, check.detail).toBe(true);
}
await page.reload({ waitUntil: "domcontentloaded" });
const receipt = page.getByTestId("task-chat-answered-questions-receipt");
await expect(receipt).toBeVisible();
await expect(page.getByTestId("task-chat-unanswered-question")).toHaveCount(0);
await receipt.locator("summary").click();
await expect(receipt).toContainText("Blue");
await input.capture("question-answered-later", "The historical question stores the submitted Blue answer", "question-answered-later.png");
const acknowledgement = page.locator(`[id="comment-${lateReply!.id}"]`);
await expect(acknowledgement).toContainText(/blue/i);
await acknowledgement.scrollIntoViewIfNeeded();
await input.capture("question-answer-acknowledged", "The agent acknowledges the late answer in a new chat turn", "question-answer-acknowledged.png");
}
+1 -1
View File
@@ -14,7 +14,7 @@ export async function provisionFirstTaskFixtures(input: {
}): Promise<LiveFixtureValues> {
const { api, execution, nonce, company, credentials } = input;
if (
!["first-task", "completion-updates"].includes(execution.suite.id) || execution.task.flow !== "first_task" ||
!["first-task", "completion-updates", "confirmation-replies"].includes(execution.suite.id) || execution.task.flow !== "first_task" ||
execution.environment.id !== "local" ||
![
"legacy-codex",
+15 -2
View File
@@ -1,3 +1,5 @@
import { runsCompletionUpdateProbe } from "./completion-quality.js";
import { gradeConfirmationReply, assertConfirmationReceipt } from "./confirmation-replies.js";
import { firstTaskUserRequest } from "./first-task-transcript.js";
import { isBlockedUnstartedWake, isTerminalUnstartedWake } from "./non-execution-wake.js";
import { firstTaskRejectionReplyRecorded, isFirstTaskRejectionCancellation } from "./first-task-rejection.js";
@@ -256,6 +258,7 @@ export async function runFirstTaskFlow(input: {
};
e.checkpoints.push(checkpoint);
e.checks = gradeFirstTask(e);
if (execution.suite.id === "confirmation-replies" && scenario.id !== "task-card-accept") e.checks.push(...gradeConfirmationReply(e));
input.observe(issue, runs, e);
await input.evidence("first-task.json", e);
await input.evidence("api-state.json", checkpoint);
@@ -343,7 +346,7 @@ export async function runFirstTaskFlow(input: {
const agent = await api.get<Row>(`/api/agents/${fixtures.agent.id}`);
e.configuredModel = agent.adapterConfig?.model ?? null;
e.runtimeSettings = {
completionDeliveryProbe: execution.suite.id === "completion-updates",
completionDeliveryProbe: runsCompletionUpdateProbe(execution),
onboardingRuntime: fixtures.onboardingRuntime,
adapterType: agent.adapterType,
adapterConfig: agent.adapterConfig,
@@ -598,7 +601,16 @@ export async function runFirstTaskFlow(input: {
scenario.id !== "interview-plan-accept",
);
}
if (execution.suite.id === "completion-updates") {
if (execution.suite.id === "confirmation-replies" && scenario.id !== "task-card-accept") {
const last = e.checkpoints.at(-1)!;
const decisionCard = last.interactions.find(card => card.result?.commentId && ["accepted", "rejected"].includes(card.status));
if (decisionCard) {
await page.reload({ waitUntil: "domcontentloaded" });
await assertConfirmationReceipt(page, decisionCard);
}
await input.evidence("confirmation-audit.json", await api.get(`/api/issues/${issue.id}/activity`));
}
if (runsCompletionUpdateProbe(execution)) {
const children = (await api.get<Row[]>(tasksPath)).filter(t => t.parentId === issue.id);
expect(children).toHaveLength(1);
const completion = await observeCompletionUpdate({ ...input, sourceId: issue.id, workerId: children[0]!.id,
@@ -607,6 +619,7 @@ export async function runFirstTaskFlow(input: {
await snapshot("finished");
}
e.checks = gradeFirstTask(e);
if (execution.suite.id === "confirmation-replies" && scenario.id !== "task-card-accept") e.checks.push(...gradeConfirmationReply(e));
await input.evidence("first-task.json", e);
await input.capture(
"final-state",
+5 -4
View File
@@ -1,5 +1,5 @@
import { gitFinalizationEvidence, gitStreamingEvidence, setupGitStreamingWorkspace } from "./daytona-git-streaming.js";
import { completionQualityControls, completionQualityStatus, judgeCompletionQuality, reserveCompletionQuality, type CompletionQualityRecord } from "./completion-quality.js";
import { runsCompletionUpdateProbe, completionQualityControls, completionQualityStatus, judgeCompletionQuality, reserveCompletionQuality, type CompletionQualityRecord } from "./completion-quality.js";
import { completionDelivery, type CompletionObservation } from "./completion-updates.js";
import { runInstructionPersistenceFlow } from "./instruction-persistence.js";
import { gradeApiResponsePaging, readResponseProof, responseEvidenceDescription } from "./api-response-reading.js";
@@ -567,7 +567,7 @@ for (const execution of executions) {
const completionQuality: CompletionQualityRecord[] = [];
const completionEvidence = async (name: string, data: unknown) => {
await writeSanitizedJson(snapshotsDir, name, data, secrets);
if (execution.suite.id !== "completion-updates" || !name.endsWith("completion-update.json")) return;
if (!["completion-updates", "confirmation-replies"].includes(execution.suite.id) || !name.endsWith("completion-update.json")) return;
const probe = data as { observation?: CompletionObservation };
if (!probe.observation || !completionDelivery(probe.observation).checks.every(c => c.passed)) return;
if (!credentials.OPENAI_API_KEY) {
@@ -959,9 +959,10 @@ for (const execution of executions) {
observe: (chatIssue, chatRuns) => { issue = chatIssue; selectedRuns = chatRuns; },
capture: captureScreenshot,
evidence: completionEvidence,
check: (id, passed, detail) => matcherResults.push({ matcher: { kind: "json_path", path: `chat.${id}`, expected: true }, passed, detail }),
});
issue = chat.issue; selectedRuns = chat.runs;
matcherResults = [{ matcher: { kind: "issue_status", expected: "in_review" }, passed: true, detail: "Chat workflow and durable handoff/session assertions passed" }];
if (matcherResults.length === 0) matcherResults = [{ matcher: { kind: "issue_status", expected: "in_review" }, passed: true, detail: "Chat workflow and durable handoff/session assertions passed" }];
} else if (execution.task.flow === "first_task") {
const firstTask = await runFirstTaskFlow({
page, api, fixtures, execution, nonce, secrets,
@@ -2588,7 +2589,7 @@ for (const execution of executions) {
);
}
}
if (execution.suite.id === "completion-updates" && credentials.OPENAI_API_KEY) {
if (runsCompletionUpdateProbe(execution) && credentials.OPENAI_API_KEY) {
const qualification = completionQualityStatus(completionQuality);
if (qualification === "unqualified") {
failureClassOverride = "permanent_infrastructure";
+128 -1
View File
@@ -42,6 +42,7 @@ const DIRECT_ADAPTER_TYPES = [
"http",
"custom_plugin",
] as const;
const editorCallbacks = vi.hoisted(() => ({ change: (_value: string) => {} }));
const streamlinedState = vi.hoisted(() => ({ enabled: true }));
vi.mock("@/components/transcript/useLiveRunTranscripts", () => ({
@@ -92,13 +93,14 @@ vi.mock("@/lib/router", () => ({
}));
vi.mock("@/components/MarkdownEditor", () => ({
MarkdownEditor: forwardRef(function MockMarkdownEditor(
{ value }: { value: string },
{ value, onChange }: { value: string; onChange: (value: string) => void },
ref: ForwardedRef<unknown>,
) {
useImperativeHandle(ref, () => ({
insertMarkdown: () => {},
focus: () => {},
}));
editorCallbacks.change = onChange;
return <div data-testid="mock-editor">{value}</div>;
}),
}));
@@ -2749,6 +2751,131 @@ describe("TaskChatThread runtime transcript selection", () => {
);
});
describe("Agent Chat unanswered question history", () => {
const old = questionInteraction("old", "Which color?", "2026-08-15T12:00:01Z");
const newer = questionInteraction("new", "Which tone?", "2026-08-15T12:05:00Z");
const movedOn = createLongThreadComments();
const props = { conversationMode: true, issueId: "issue-1", onAdd: async () => {}, onSubmitInteractionAnswers: vi.fn() };
const takeover = () => container.querySelector('[data-testid="task-chat-composer-takeover"]');
const pendingIndicator = () => container.querySelector<HTMLButtonElement>('[data-testid="task-chat-pending-input-indicator"]');
const dismiss = async () => {
await act(async () => takeover()!.querySelector<HTMLButtonElement>('[aria-label^="Dismiss "]')!.click());
};
const click = async (text: string) => {
const button = Array.from(container.querySelectorAll<HTMLButtonElement>("button")).find(button => button.textContent?.trim() === text);
expect(button).toBeTruthy();
await act(async () => button!.click());
};
it("leaves a compact, reopenable question on reload after moving on", async () => {
render(<TaskChatThread {...props} comments={movedOn} interactions={[old]} />);
expect(takeover()).toBeNull();
expect(pendingIndicator()).toBeNull();
const row = container.querySelector<HTMLButtonElement>('[data-testid="task-chat-unanswered-question"]');
expect(row?.textContent).toContain("Which color?");
await act(async () => row!.click());
expect(takeover()?.textContent).toContain("Which color?");
await click("Yes");
await click("Submit answers");
expect(props.onSubmitInteractionAnswers).toHaveBeenCalledWith(old, [{ questionId: "old-question", optionIds: ["yes"] }]);
});
it("collapses a displayed question when a new user message arrives", async () => {
render(<TaskChatThread {...props} comments={[]} interactions={[old]} />);
expect(takeover()).not.toBeNull();
await act(async () => render(<TaskChatThread {...props} comments={movedOn} interactions={[old]} />));
expect(takeover()).toBeNull();
expect(pendingIndicator()).toBeNull();
expect(container.querySelector('[data-testid="task-chat-unanswered-question"]')).not.toBeNull();
});
it.each(["Cancel", "dismiss"])("leaves only the history card after %s on a fresh question", async action => {
render(<TaskChatThread {...props} comments={[]} interactions={[old]} />);
await click("Yes");
if (action === "Cancel") await click("Cancel");
else await dismiss();
expect(takeover()).toBeNull();
expect(pendingIndicator()).toBeNull();
await act(async () => container.querySelector<HTMLButtonElement>('[aria-label="Answer question: Which color?"]')!.click());
expect(takeover()?.querySelector('[role="radio"][aria-checked="true"]')?.textContent).toContain("Yes");
});
it("keeps multiple historical questions out of the composer after reopening and dismissing one", async () => {
const anotherOld = questionInteraction("another-old", "Which format?", "2026-08-15T12:00:02Z");
render(<TaskChatThread {...props} comments={movedOn} interactions={[old, anotherOld]} />);
expect(container.querySelectorAll('[data-testid="task-chat-unanswered-question"]')).toHaveLength(2);
expect(pendingIndicator()).toBeNull();
await act(async () => container.querySelector<HTMLButtonElement>('[aria-label="Answer question: Which color?"]')!.click());
expect(takeover()?.textContent).toContain("Which color?");
expect(takeover()?.textContent).not.toContain("2 pending");
await dismiss();
expect(takeover()).toBeNull();
expect(pendingIndicator()).toBeNull();
});
it("counts and cycles approvals without including fresh or historical question cards", async () => {
const approval = { ...planReviewInteraction(), id: "approval", title: "Approve the outline?", payload: { version: 1, prompt: "Approve the outline?" } } as IssueThreadInteraction;
const nextApproval = { ...approval, id: "next-approval", title: "Approve the note?", createdAt: new Date("2026-08-23T10:02:00Z"), payload: { version: 1, prompt: "Approve the note?" } } as IssueThreadInteraction;
render(<TaskChatThread {...props} comments={movedOn} interactions={[old, newer, approval, nextApproval]} />);
expect(takeover()?.textContent).toContain("Approve the note?");
await click("2 pending");
expect(takeover()?.textContent).toContain("Approve the outline?");
await click("2 pending");
expect(takeover()?.textContent).toContain("Approve the note?");
await act(async () => container.querySelector<HTMLButtonElement>('[aria-label="Answer question: Which color?"]')!.click());
expect(takeover()?.textContent).toContain("Which color?");
expect(takeover()?.textContent).toContain("2 pending");
expect(takeover()?.textContent).not.toContain("3 pending");
await dismiss();
expect(pendingIndicator()?.textContent).toContain("2 pending inputs");
await act(async () => pendingIndicator()!.click());
expect(takeover()?.textContent).toContain("Approve the note?");
});
it("opens the new question but can reopen the exact historical native question", async () => {
const nativeOld = { ...old, sourceRunId: "old-run", payload: { ...old.payload, runtimeRequestId: "old-request" } } as IssueThreadInteraction;
render(<TaskChatThread {...props} comments={movedOn} interactions={[nativeOld, newer]} />);
expect(takeover()?.textContent).toContain("Which tone?");
const oldRow = container.querySelector<HTMLButtonElement>('[aria-label="Answer question: Which color?"]');
await act(async () => oldRow!.click());
expect(takeover()?.textContent).toContain("Which color?");
expect(takeover()?.textContent).not.toContain("2 pending");
render(<TaskChatThread {...props} comments={[...movedOn]} interactions={[nativeOld, newer]} />);
expect(takeover()?.textContent).toContain("Which color?");
await dismiss();
expect(pendingIndicator()).toBeNull();
await act(async () => container.querySelector<HTMLButtonElement>('[aria-label="Answer question: Which tone?"]')!.click());
expect(takeover()?.textContent).toContain("Which tone?");
});
it.each([true, false])("only collapses the question after a successful send (success=%s)", async success => {
const onAdd = success ? vi.fn().mockResolvedValue(undefined) : vi.fn().mockRejectedValue(new Error("Message not sent"));
render(<TaskChatThread {...props} comments={[]} interactions={[old]} onAdd={onAdd} />);
await click("Yes");
await act(async () => editorCallbacks.change("Let's talk about tasks instead."));
const send = container.querySelector<HTMLButtonElement>('[data-testid="task-chat-composer-send"]');
expect(send?.disabled).toBe(false);
await act(async () => send!.click());
expect(onAdd).toHaveBeenCalledTimes(1);
if (success) {
expect(takeover()).toBeNull();
expect(pendingIndicator()).toBeNull();
await act(async () => container.querySelector<HTMLButtonElement>('[data-testid="task-chat-unanswered-question"]')!.click());
}
expect(takeover()).not.toBeNull();
expect(takeover()?.querySelector('[role="radio"][aria-checked="true"]')?.textContent).toContain("Yes");
expect(container.querySelector('[data-testid="task-chat-unanswered-question"]')).not.toBeNull();
});
it("does not change ordinary task question behavior", async () => {
render(<TaskChatThread {...props} conversationMode={false} comments={movedOn} interactions={[old]} />);
expect(takeover()?.textContent).toContain("Which color?");
expect(container.querySelector('[data-testid="task-chat-unanswered-question"]')).toBeNull();
await dismiss();
expect(pendingIndicator()?.textContent).toContain("1 pending input");
});
});
describe("TaskChatThread composer alignment", () => {
it("matches the thread width at every breakpoint", () => {
render(<TaskChatThread comments={[]} onAdd={async () => {}} />);
+47 -17
View File
@@ -2441,46 +2441,70 @@ export function TaskChatThread(props: TaskChatThreadProps) {
}
return result;
}, [interactions, pendingRuntimeRequest]);
// Agent Chat questions already have an answerable history card, including
// when the user dismisses a fresh form without sending another message.
const pendingReminderInputs = useMemo(() => pendingComposerInputs.filter(input =>
!(conversationMode && input.kind === "durable" && input.interaction.kind === "ask_user_questions")),
[pendingComposerInputs, conversationMode]);
const latestUserComment = useMemo(() => comments
.filter(comment => !comment.deletedAt && comment.authorUserId && !comment.authorAgentId && !comment.createdByRunId)
.reduce<(typeof comments)[number] | null>((latest, comment) =>
!latest || toMs(comment.createdAt) >= toMs(latest.createdAt) ? comment : latest, null), [comments]);
const latestUserCommentId = latestUserComment?.id ?? null;
// A later message leaves a question answerable in history, without reopening
// its form on every render or reload. Explicit selection can still reopen it.
const currentPendingInputs = useMemo(() => pendingComposerInputs.filter(input =>
!(conversationMode && input.kind === "durable"
&& input.interaction.kind === "ask_user_questions" && latestUserComment
&& toMs(input.interaction.createdAt) < toMs(latestUserComment.createdAt))),
[pendingComposerInputs, conversationMode, latestUserComment]);
const currentPendingKeys = useMemo(() => new Set(currentPendingInputs.map(input => input.key)), [currentPendingInputs]);
const [takeoverMode, setTakeoverMode] = useState<"open" | "normal">("open");
const [selectedPendingKey, setSelectedPendingKey] = useState<string | null>(
null,
);
const knownPendingKeysRef = useRef<Set<string>>(new Set());
const lastInputUserCommentRef = useRef(latestUserCommentId);
useEffect(() => {
const currentKeys = new Set(
pendingComposerInputs.map((input) => input.key),
);
const hasNewInput = pendingComposerInputs.some(
(input) => !knownPendingKeysRef.current.has(input.key),
(input) => currentPendingKeys.has(input.key) && !knownPendingKeysRef.current.has(input.key),
);
const movedOn = lastInputUserCommentRef.current !== latestUserCommentId;
lastInputUserCommentRef.current = latestUserCommentId;
knownPendingKeysRef.current = currentKeys;
setSelectedPendingKey((current) =>
current && currentKeys.has(current)
current && currentKeys.has(current) && !(movedOn && !currentPendingKeys.has(current))
? current
: (pendingComposerInputs[0]?.key ?? null),
: (currentPendingInputs[0]?.key ?? null),
);
if (hasNewInput) setTakeoverMode("open");
}, [pendingComposerInputs]);
}, [pendingComposerInputs, latestUserCommentId, currentPendingInputs, currentPendingKeys]);
const selectedPendingInput =
pendingComposerInputs.find((input) => input.key === selectedPendingKey) ??
pendingComposerInputs[0] ??
currentPendingInputs[0] ??
null;
const interactionDraftKey = selectedPendingInput
? `paperclip:task-input:${issueId ?? "unknown"}:${selectedPendingInput.key}`
: undefined;
const openPendingTakeover = useCallback(() => {
if (!pendingReminderInputs.some(input => input.key === selectedPendingInput?.key)) {
setSelectedPendingKey(pendingReminderInputs[0]?.key ?? null);
}
setTakeoverMode("open");
}, []);
}, [pendingReminderInputs, selectedPendingInput]);
const showNextPendingInput = useCallback(() => {
if (pendingComposerInputs.length < 2) return;
const currentIndex = pendingComposerInputs.findIndex(
if (pendingReminderInputs.length < 2) return;
const currentIndex = pendingReminderInputs.findIndex(
(input) => input.key === selectedPendingInput?.key,
);
setSelectedPendingKey(
pendingComposerInputs[(currentIndex + 1) % pendingComposerInputs.length]
pendingReminderInputs[(currentIndex + 1) % pendingReminderInputs.length]
?.key ?? null,
);
}, [pendingComposerInputs, selectedPendingInput?.key]);
}, [pendingReminderInputs, selectedPendingInput?.key]);
const skipPendingInput = useCallback(
async (input: PendingComposerInput) => {
if (input.kind === "runtime") {
@@ -2516,12 +2540,16 @@ export function TaskChatThread(props: TaskChatThreadProps) {
if (assigneeUsesPaperclipRunner) setRunnerSubmissionPending(true);
try {
await onAdd(...args);
if (conversationMode && selectedPendingInput?.kind === "durable" && selectedPendingInput.interaction.kind === "ask_user_questions") {
setSelectedPendingKey(null);
setTakeoverMode("normal");
}
} catch (error) {
setRunnerSubmissionPending(false);
throw error;
}
},
[assigneeUsesPaperclipRunner, onAdd],
[assigneeUsesPaperclipRunner, onAdd, conversationMode, selectedPendingInput],
);
const optimisticRunnerStartup =
assigneeUsesPaperclipRunner &&
@@ -2599,15 +2627,16 @@ export function TaskChatThread(props: TaskChatThreadProps) {
);
const reopenToolReview = useCallback((interactionId: string) => {
setSelectedPendingKey(`interaction:${interactionId}`);
setSelectedPendingKey(pendingComposerInputs.find(input => input.kind === "durable" && input.interaction.id === interactionId)?.key ?? `interaction:${interactionId}`);
setTakeoverMode("open");
}, []);
}, [pendingComposerInputs]);
const renderInteraction = useCallback(
(item: TaskChatInteractionItem) => (
<TaskChatInteractionCard
item={item}
onReviewRequest={reopenToolReview}
showUnansweredQuestion={conversationMode}
planDocument={planDocument}
showPlanPreview={
!threadOwnsPlanPreview(
@@ -2648,6 +2677,7 @@ export function TaskChatThread(props: TaskChatThreadProps) {
settledRunIds,
tailRunId,
reopenToolReview,
conversationMode,
],
);
@@ -2704,7 +2734,7 @@ export function TaskChatThread(props: TaskChatThreadProps) {
selectedPendingInput.kind === "durable" &&
selectedPendingInput.interaction.kind === "request_confirmation" &&
Boolean(selectedPendingInput.interaction.payload.toolAction),
pendingCount: pendingComposerInputs.length,
pendingCount: Math.max(1, pendingReminderInputs.length),
content: takeoverContent,
onDismiss: () => setTakeoverMode("normal"),
onSkip: () => skipPendingInput(selectedPendingInput),
@@ -3122,10 +3152,10 @@ export function TaskChatThread(props: TaskChatThreadProps) {
onRunnerGoalCommand={runnerGoal.executeComposerCommand}
onRunnerGoalReassign={reassignForRunnerGoal}
pendingTakeover={
pendingComposerInputs.length > 0
pendingReminderInputs.length > 0
? {
count: pendingComposerInputs.length,
label: `${pendingComposerInputs.length} pending input${pendingComposerInputs.length === 1 ? "" : "s"}`,
count: pendingReminderInputs.length,
label: `${pendingReminderInputs.length} pending input${pendingReminderInputs.length === 1 ? "" : "s"}`,
onOpen: openPendingTakeover,
}
: null
@@ -5,6 +5,7 @@ import { IssueThreadInteractionCard } from "@/components/IssueThreadInteractionC
import { AppLogo } from "@/pages/apps/AppLogo";
import { MarkdownBody } from "@/components/MarkdownBody";
import { Button } from "@/components/ui/button";
import { CircleHelp, ChevronRight } from "lucide-react";
import { shouldHideInteractionCard } from "@/lib/issue-thread-interactions";
import { TaskChatCompactInteractionCard } from "./TaskChatCompactInteractionCard";
import { TaskChatPlanPreviewCard } from "./TaskChatPlanPreviewCard";
@@ -18,6 +19,7 @@ type InteractionCardProps = Omit<
export interface TaskChatInteractionCardProps extends InteractionCardProps {
item: TaskChatInteractionItem;
onReviewRequest?: (interactionId: string) => void;
showUnansweredQuestion?: boolean;
planDocument?: IssueDocument | null;
showPlanPreview?: boolean;
presentation?: "timeline" | "takeover";
@@ -35,6 +37,7 @@ export interface TaskChatInteractionCardProps extends InteractionCardProps {
export function TaskChatInteractionCard({
item,
onReviewRequest,
showUnansweredQuestion = false,
planDocument,
showPlanPreview = true,
presentation = "timeline",
@@ -42,6 +45,24 @@ export function TaskChatInteractionCard({
...cardProps
}: TaskChatInteractionCardProps) {
const interaction = item.interaction;
if (presentation === "timeline" && showUnansweredQuestion && !shouldHideInteractionCard(interaction) && interaction.kind === "ask_user_questions" && interaction.status === "pending") {
const prompt = interaction.payload.questionSet?.questions[0]?.prompt ?? interaction.payload.questions[0]?.prompt ?? interaction.title ?? "Question";
return (
<button
type="button"
id={`interaction-${interaction.id}`}
data-testid="task-chat-unanswered-question"
aria-label={`Answer question: ${prompt}`}
disabled={!onReviewRequest}
onClick={() => onReviewRequest?.(interaction.id)}
className="group flex w-full items-center gap-2 rounded-sm px-1 py-1.5 text-left text-sm text-muted-foreground transition-colors hover:text-foreground focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-ring"
>
<CircleHelp aria-hidden className="h-3.5 w-3.5 shrink-0" />
<span className="min-w-0 flex-1"><span className="block text-xs">Unanswered question</span><span className="block truncate">{prompt}</span></span>
<ChevronRight aria-hidden className="h-3 w-3 shrink-0" />
</button>
);
}
if (interaction.kind === "request_confirmation" && interaction.payload.toolAction) {
const action = interaction.payload.toolAction;
if (presentation === "takeover") {
@@ -0,0 +1,124 @@
import { useState } from "react";
import type { Meta, StoryObj } from "@storybook/react-vite";
import { expect, userEvent, within, waitFor } from "storybook/test";
import type { AskUserQuestionsInteraction, IssueComment } from "@paperclipai/shared";
import { QueryClient, QueryClientProvider } from "@tanstack/react-query";
import { queryKeys } from "@/lib/queryKeys";
import { TaskChatThread } from "@/components/TaskChatThread";
import { pendingAskUserQuestionsInteraction } from "@/fixtures/issueThreadInteractionFixtures";
import { storybookAgentMap } from "../fixtures/paperclipData";
const question: AskUserQuestionsInteraction = {
...pendingAskUserQuestionsInteraction,
id: "unanswered-color", title: "Welcome note preference", sourceRunId: "original-run",
createdAt: new Date("2026-04-01T12:01:00Z"),
payload: { version: 1, runtimeRequestId: "color-request", questions: [{ id: "color", prompt: "Which color should the welcome note use?",
selectionMode: "single", required: true, options: [{ id: "blue", label: "Blue" }, { id: "green", label: "Green" }] }] },
};
const otherQuestion: AskUserQuestionsInteraction = {
...question, id: "unanswered-tone", sourceRunId: "second-run", title: "Welcome note tone",
createdAt: new Date("2026-04-01T12:03:00Z"),
payload: { version: 1, questions: [{ id: "tone", prompt: "Which tone should the welcome note use?",
selectionMode: "single", required: true, options: [{ id: "friendly", label: "Friendly" }, { id: "formal", label: "Formal" }] }] },
};
function comment(id: string, body: string, at: string, agent = false): IssueComment {
return { id, companyId: question.companyId, issueId: question.issueId, body, authorType: agent ? "agent" : "user",
authorAgentId: agent ? question.createdByAgentId ?? null : null, authorUserId: agent ? null : "user-board",
presentation: null, metadata: null, createdAt: new Date(at), updatedAt: new Date(at) };
}
function QuestionChat({ movedOn = false, multiple = false, answered = false }: { movedOn?: boolean; multiple?: boolean; answered?: boolean }) {
const [queryClient] = useState(() => {
const client = new QueryClient({ defaultOptions: { queries: { retry: false, staleTime: Infinity } } });
// This conversation has no plan. Keep the real thread's plan query local
// to the fixture rather than requesting an API from the static publisher.
client.setQueryData([...queryKeys.issues.documents(question.issueId), "plan"], null);
return client;
});
const [interactions, setInteractions] = useState<AskUserQuestionsInteraction[]>([
answered ? { ...question, status: "answered", resolvedAt: new Date("2026-04-01T12:06:00Z"), result: { version: 1, answers: [{ questionId: "color", optionIds: ["blue"] }] } } : question,
...(multiple ? [otherQuestion] : []),
]);
const [comments, setComments] = useState<IssueComment[]>([
comment("start", "Help me plan a welcome note for our garden club.", "2026-04-01T12:00:00Z"),
...(movedOn || multiple || answered ? [
comment("move-on", "Let's leave those choices for later. What can Paperclip tasks track?", "2026-04-01T12:04:00Z"),
comment("reply", "Tasks track ownership, progress, and the work needed to reach a goal.", "2026-04-01T12:05:00Z", true),
] : []),
...(answered ? [comment("late-answer", "Blue it is. I'll use that preference when we return to the welcome note.", "2026-04-01T12:07:00Z", true)] : []),
]);
return <QueryClientProvider client={queryClient}><div className="flex h-screen flex-col bg-background text-foreground">
<TaskChatThread conversationMode comments={comments} interactions={interactions} issueId={question.issueId}
issueStatus="in_review" currentUserId="user-board" agentMap={storybookAgentMap} enableLiveTranscriptPolling={false}
threadHeader={<div className="p-4"><h1 className="text-xl font-semibold">Garden club chat</h1><p className="text-sm text-muted-foreground">Questions can wait. Open an unanswered question in history whenever you're ready.</p></div>}
onAdd={async body => {
setComments(rows => [...rows, comment(`user-${rows.length}`, body, new Date().toISOString()),
comment(`agent-${rows.length}`, "We can come back to that question later. What would you like to work on next?", new Date(Date.now() + 1).toISOString(), true)]);
}}
onSubmitInteractionAnswers={async (interaction, answers) => {
setInteractions(rows => rows.map(row => row.id === interaction.id ? { ...row, status: "answered", resolvedAt: new Date(), result: { version: 1, answers } } : row));
setComments(rows => [...rows, comment(`answer-${rows.length}`, "Thanks, I've received your answer to the earlier question.", new Date().toISOString(), true)]);
}}
/>
</div></QueryClientProvider>;
}
const meta = { title: "Chat & Comments/Agent Chat Unanswered Questions", parameters: { layout: "fullscreen" }, component: QuestionChat,
beforeEach: () => {
for (const key of Object.keys(localStorage)) {
if (key.includes(`paperclip:task-input:${question.issueId}:`)) localStorage.removeItem(key);
}
},
} satisfies Meta<typeof QuestionChat>;
export default meta;
type Story = StoryObj<typeof meta>;
export const JustAsked: Story = { args: {} };
export const DismissFreshQuestion: Story = { args: {}, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getByRole("radio", { name: "Green" })).toBeVisible());
await userEvent.click(canvas.getByRole("radio", { name: "Green" }));
await userEvent.click(canvas.getByRole("button", { name: /^Cancel$/ }));
await expect(canvas.getByTestId("task-chat-unanswered-question")).toBeVisible();
await expect(canvas.queryByTestId("task-chat-composer-takeover")).not.toBeInTheDocument();
await expect(canvas.queryByTestId("task-chat-pending-input-indicator")).not.toBeInTheDocument();
} };
export const MovedOn: Story = { args: { movedOn: true }, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getByTestId("task-chat-unanswered-question")).toBeVisible());
await expect(canvas.queryByTestId("task-chat-composer-takeover")).not.toBeInTheDocument();
await expect(canvas.queryByTestId("task-chat-pending-input-indicator")).not.toBeInTheDocument();
} };
export const Reopened: Story = { args: { movedOn: true }, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getByTestId("task-chat-unanswered-question")).toBeVisible());
await userEvent.click(canvas.getByRole("button", { name: "Answer question: Which color should the welcome note use?" }));
await waitFor(() => expect(canvas.getByTestId("task-chat-composer-takeover")).toBeVisible());
await expect(canvas.getByRole("radio", { name: "Blue" })).toBeVisible();
} };
export const AnswerLater: Story = { args: { movedOn: true }, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getByTestId("task-chat-unanswered-question")).toBeVisible());
await userEvent.click(canvas.getByTestId("task-chat-unanswered-question"));
await userEvent.click(canvas.getByRole("radio", { name: "Blue" }));
await userEvent.click(canvas.getByRole("button", { name: /^(Send|Submit) answers$/ }));
await expect(canvas.queryByTestId("task-chat-unanswered-question")).not.toBeInTheDocument();
await waitFor(() => expect(canvas.getByTestId("task-chat-answered-questions-receipt")).toBeVisible());
} };
export const MultipleUnanswered: Story = { args: { multiple: true }, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getAllByTestId("task-chat-unanswered-question")).toHaveLength(2));
await expect(canvas.queryByTestId("task-chat-composer-takeover")).not.toBeInTheDocument();
await expect(canvas.queryByTestId("task-chat-pending-input-indicator")).not.toBeInTheDocument();
} };
export const AnsweredHistory: Story = { args: { answered: true } };
export const Mobile: Story = { args: { movedOn: true }, globals: { viewport: { value: "mobile", isRotated: false } } };
export const MoveOnWithoutAnswering: Story = { args: {}, play: async ({ canvasElement }) => {
const canvas = within(canvasElement);
await waitFor(() => expect(canvas.getByTestId("task-chat-unanswered-question")).toBeVisible());
await userEvent.click(canvas.getByRole("radio", { name: "Green" }));
await userEvent.type(canvas.getByRole("textbox", { name: "editable markdown" }), "Leave that for later. Tell me about tasks.");
await userEvent.click(canvas.getByRole("button", { name: /^Send$/ }));
await expect(canvas.queryByTestId("task-chat-composer-takeover")).not.toBeInTheDocument();
await expect(canvas.queryByTestId("task-chat-pending-input-indicator")).not.toBeInTheDocument();
await userEvent.click(canvas.getByTestId("task-chat-unanswered-question"));
await expect(canvas.getByRole("radio", { name: "Green" })).toBeChecked();
} };