256 Commits
Author SHA1 Message Date
1d23cb6962 ci: pin a checkout-independent Rust cache path so PR lanes hit master's cache (#14394)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Paperclip ships a native Runner binary, written in Rust, and seven
CI lanes build it on every pull request
> - `Canary Dry Run` is the slowest check on every green PR run, and
most of its time is `cargo build --release` on third-party crates
> - Master saves a Rust dependency cache for these lanes, but every PR
lane logs `No cache found` and compiles every crate from zero
> - The cache key matches, but GitHub also compares a hash of the
absolute cache paths, and the master writer (RunsOn fleet,
`/home/runner/_work/...`) and the PR readers (GitHub-hosted,
`/home/runner/work/...`) hash different paths
> - This pull request gives both sides a checkout-independent workspace
path, so the hashes match and the PR lanes restore master's cache
> - The benefit is about 2.5 minutes less wall clock per PR run and
about 18 fewer runner-minutes per run

## Linked Issues or Issue Description

No public issue exists for this problem. The description below follows
the enhancement template.

Related prior PRs on the same cache: Refs #13194, Refs #13259, Refs
#13457, Refs #13459, Refs #13500, Refs #13586. None of them pins the
workspace path, so none of them fixes this miss.

**What existing behavior does this improve?**

The `Swatinem/rust-cache` restore step in the PR workflow lanes that
build the Runner: `Canary Dry Run`, `Build`, `Typecheck + Release
Registry`, and the four `Verify Paperclip Runner` lanes.

**Subsystem affected**

CI workflows under `.github/workflows/`, their guard tests under
`.github/scripts/tests/`, and `doc/RELEASE-AUTOMATION-SETUP.md`.

**Current behavior**

Every PR lane logs `No cache found` although master holds an entry with
the exact key. Run 36424309181 computed
`v0-rust-release-runner-v1-Linux-x64-c3a3ca66-a95b0328`, and master
holds a 678 MB entry with that key. GitHub matches a cache entry on the
key and on a version hash of the absolute paths in the cache. The master
writer runs on the RunsOn fleet, where the checkout is
`/home/runner/_work/paperclip/paperclip`. The PR readers run on
GitHub-hosted `ubuntu-latest`, where the checkout is
`/home/runner/work/paperclip/paperclip`. The stored version `5c40870d…`
is the sha256 of the `_work` paths plus `zstd-without-long|1.0`. The
`work` paths hash to `1656e9ee…`. The key can never match, so each lane
compiles every third-party crate again.

**Proposed behavior**

The writer and the readers pass the same checkout-independent path to
`rust-cache`. Both runner layouts then produce the same version hash,
and the PR lanes restore master's cache.

**Reason and benefit**

`Canary Dry Run` takes 533s on a green run. 251s of that is dependency
compilation that a warm cache removes. Seven lanes pay this cost in
every PR run.

**Breaking changes**

None. This change affects CI only.

## What Changed

- Add a `Pin the Runner Rust workspace path` step before `rust-cache` in
the master writer (`release-verify.yml`, typecheck and runner lanes) and
in all four PR readers (`pr-trusted.yml`). The step creates the symlink
`$HOME/paperclip-runner-rust` →
`$GITHUB_WORKSPACE/packages/paperclip-runner/runner` and passes that
path to `rust-cache` as `workspaces: <path> -> target`. `rust-cache`
resolves the input with `path.resolve`, which does not follow symlinks,
so both runner layouts now produce the same cache paths and the same
version hash. `$HOME` is `/home/runner` on both images, which is why the
`~/.cargo` paths already agreed.
- Bump the shared keys `release-runner-v1` → `release-runner-v2` and
`release-typecheck-v1` → `release-typecheck-v2`. The old, unreachable
entries are then visibly orphaned instead of sharing a key with the new
ones.
- Extend the guard tests `pr-runner-rust-cache`, `release-runner-cache`,
and `typecheck-rust-cache`. They now require the pin step in both
workflows with identical text, placed before the cache step, and they
reject a `workspaces:` value that resolves under the checkout. The
`pr-runner-rust-cache` test checks all four PR reader jobs and fails if
a `rust-cache` step appears in a PR job that is not in its reader list.
- Update `doc/RELEASE-AUTOMATION-SETUP.md` to name the
`release-runner-v2` key and to explain the pinned workspace path.

### Expected savings once merged

Measured from run 36424309181. "Removed" is the dependency-compile time
that a warm restore removes, minus about 18s to restore the 680 MB
entry. The fleet writer's own restore shows this cost.

| Lane | Today | Removed | Expected |
|---|---|---|---|
| Canary Dry Run | 533s | ~150s | ~380s |
| Typecheck + Release Registry | 462s | ~155s | ~305s |
| Verify Paperclip Runner (vitest 2/2) | 453s | ~245s | ~210s |
| Verify Paperclip Runner (rust) | 400s | ~175s | ~225s |
| Build | 348s | ~130s | ~220s |
| Verify Paperclip Runner (static checks) | 321s | ~170s | ~150s |
| Verify Paperclip Runner (vitest 1/2) | 346s | ~70s | ~275s |

- Wall clock per PR run: about 533s → about 385s. That is about 2.5
minutes faster to a green check set. `Canary Dry Run` stays the longest
check. The rest is the non-cargo work in `release.sh` (standalone
package builds ~30s, publish-payload preview ~73s).
- Runner time: about 18 runner-minutes saved per PR run across the seven
lanes.
- The first master push after merge compiles from zero once in the fleet
writer (about 4 extra minutes on that one run) and saves the v2 entry.
Later PRs hit it. When a PR changes `Cargo.lock`, the prefix restore key
still gives a partial hit, as before.

## Verification

- Run the guard tests for the three cache lanes:
`node --test .github/scripts/tests/pr-runner-rust-cache.test.mjs
.github/scripts/tests/release-runner-cache.test.mjs
.github/scripts/tests/typecheck-rust-cache.test.mjs`
  Result: 21 pass, 0 fail.
- Run the full guard suite: `node --test
'.github/scripts/tests/*.test.mjs'`. Result: 376 pass, 3 fail. The 3
failures are in `docker-canary-promotion.test.mjs`. They hit a sandbox
temp-file ENOENT and fail the same way on the unmodified branch.
- Run `node --test scripts/__tests__/release-verify-workflow.test.mjs`.
Result: 14 pass.
- Local archive test: create a tar from the `_work` layout through the
symlink (relative `../../../paperclip-runner-rust/target` entries, `tar
-P -C $GITHUB_WORKSPACE`, the same way `@actions/cache` does). Extract
it on the `work` layout. The files land in the real target directory and
the symlink stays intact.
- After merge, open any GitHub-hosted PR run and confirm that the seven
Rust lanes log `Restored from cache key ...release-runner-v2...` in
place of `No cache found`.

## Risks

- Low risk. The change touches CI workflows, their tests, and one doc
page. No product code changes.
- If the pin step fails, `rust-cache` reports a miss and the lane
compiles from zero, as it does today. The build does not break.
- Both runner layouts sit four levels under `/home/runner`, so the
relative `../../../` archive entries line up. The existing
`~/.cargo/registry` and `~/.cargo/git` cache paths already rely on this
property. A future runner image with a different `$HOME` depth would
miss the cache but would not fail the job.
- `rm -rf "$pinned"` acts on the symlink itself (no trailing slash),
never on the checkout behind it. It only matters on a reused runner.
- Squash-merge note: the branch carries commits by `Bender (Fable)`. Add
`Co-Authored-By: Bender (Fable) <bender-fable@paperclip.local>` to the
squash body to keep that authorship.

## Model Used

- Anthropic Claude Fable 5.1 (`claude-fable-5-1`), run through Claude
Code inside a Paperclip agent heartbeat. Extended thinking was on. Tool
use: shell, GitHub CLI, and the GitHub REST API for workflow logs, cache
listings, and PR operations.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change and contains no internal
ticket id. The agent execution workspace fixed this branch name, so I
cannot rename it. Squash-merge drops the branch name.
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: devinfoley <139239+devinfoley@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
2026-09-30 16:04:23 -07:00
Devin FoleyandPaperclip e5bf9d49a5 fix(sentry): retain recorded process exit details (#14575)
## Thinking Path

> - Paperclip manages agents and records their task runs.
> - Operators can enable Sentry reports for terminal run failures.
> - A failed adapter can leave only the generic message “Adapter
failed.”
> - The run already records a process exit code and signal, but the
report omits them.
> - This pull request carries those two values through a strict capture
boundary.
> - Operators can distinguish a nonzero exit from signal termination
when the message is generic.

## Linked Issues or Issue Description

**What happened?**
A failed run can store `exitCode: 1` or `signal: "SIGTERM"` while its
Sentry event contains only `adapter_failed` and “Adapter failed.” The
existing reporter drops both recorded fields. This occurs on the current
master reporting path.

**Expected behavior**
The opt-in report preserves bounded process exit evidence without
exporting adapter output or changing run behavior.

**Steps to reproduce**
Enable the backend Sentry DSN and report a failed run whose message is
“Adapter failed” and whose stored signal is `SIGTERM`. Before this
change, the event has no signal field. After this change,
`run_failure.signal` is `SIGTERM` and the existing fingerprint stays the
same.

Related: #12105 and #8222 describe missing adapter/HTTP failure details.
#13152 changes terminal-result cleanup classification, and #12886 adds
process-failure classification and runtime URL checks. None forwards
these stored fields through the Sentry reporter. This change does not
resolve those broader issues.

## What Changed

- Forward the stored exit code and signal from the terminal run
reporter.
- Accept only signed 32-bit integer exit codes; use `null` for missing
or malformed values.
- Accept only the reporting host's Node signal constants; use `null` for
missing values and `unknown` for unrecognized values.
- Keep the added fields in event-local context, outside tags and
fingerprints.
- Cover database-backed reporting, malformed input, privacy, and
isolation through the real Sentry SDK.
- Document the fields and their limits.
- Give the dedicated Sentry job the normal PR dependency-resolution
fallback, with lifecycle scripts disabled on every install and the
required real-SDK test retained.

## Verification

- Before the change: 16 report-shape/exit-field assertions failed in the
focused capture suite.
- After the change: 91 focused Sentry, DSN, and database-backed
reporting tests passed. The real SDK uses an in-memory transport.
- Final real-SDK test also passed with malformed metadata; it verifies
that arbitrary signal text is absent from captured events and unrelated
errors inherit no run context.
- `pnpm -r typecheck`: passed.
- `pnpm build`: passed.
- `pnpm test:run`: exited nonzero after 718 general-server files: 14,085
tests passed, 14 failed, 86 skipped. Thirteen skill-service/cache
failures reproduce on the unchanged base commit on this macOS host. One
comment-wake test timed out; the complete 29-test suite passes on the
unchanged base and in final-head Linux CI. Local isolated rechecks
skipped because embedded PostgreSQL could not start; these are not
counted as passes. The remaining local workspace/serialized lanes did
not run after the failing first lane; all CI lanes passed.
- Initial Sentry CI failed before tests with
`ERR_PNPM_LOCKFILE_CONFIG_MISMATCH`. Its install step lacked the normal
PR fallback. The repaired real-SDK job passed. Security review then
requested disabling lifecycle scripts for resolved dependencies; every
install now uses `--ignore-scripts`. A fresh isolated checkout passed
the exact script-disabled fallback and real-SDK contract. Final-head
real-SDK CI and the security scan passed.

- GitHub CI: all 54 checks passed on
`37e0a836e50660f7753d367bcf5a4959eaf89b90`, including required `ci /
verify` and `ci / e2e`; two unrelated checks skipped.
- Greptile: 5/5 on that commit. No unresolved review comments.
- Merge status: conflict-free; required CODEOWNER approval for the
workflow change is still pending.

## Risks

Low risk: this only adds two validated fields to existing opt-in error
reports. It changes no database schema, run status, retry, fingerprint,
or suppression rule. Process output and adapter result payloads remain
excluded. A recorded signal does not identify its sender or prove an
out-of-memory kill. Missing exit evidence stays unknown; this change
does not establish the cause of a historical generic adapter failure.

## Model Used

OpenAI GPT-6 (Codex), with reasoning, terminal tools, and code
execution. The context window size is not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes:` / `Closes`
/ `Refs` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
ticket id or instance-derived details
- [x] I have run focused tests locally and they pass; broader validation
is recorded above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-29 12:29:17 -05:00
DottaandPaperclip 18e8c121d9 fix(runner): include Grok support in public installs with sandbox prerequisites (#14024)
## Thinking Path

> - Paperclip manages agents through a shared native runner.
> - Built-in harness support should ship with Paperclip's public
distribution.
> - Grok already speaks ACP; it does not require a new public bridge
package.
> - Sandbox provisioning owns the native executable and its pinned
version.
> - The runner must verify that prerequisite without downloading it
during npm installation.
> - This change separates built-in launcher identity from external
runtime identity.
> - Clean npm installation and live staging checks verify the
distribution boundary.

## Linked Issues or Issue Description

Refs #13882, #13973, #13977, #13979.

This follow-up now targets master after #13882 was squash-merged. It
replaces the private `@paperclipai/grok-acp` workspace package with
runner-owned assets. Current master is included so the branch also
contains the merged scheduler, complete-event capture, and durable
cleanup fixes.

## What Changed

- Ship Grok launcher and qualification metadata inside the runner's
compiled output and the public server's vendored runner tree.
- Remove the separate Grok npm package and all package-manager install
hooks for this runtime.
- Require the checksum-verified Grok Build 1.0.13 binary at
`/opt/paperclip/providers/grok/1.0.13/grok` in the selected execution
environment. Provision it explicitly in the Daytona image and CI setup.
- Keep native binaries outside the provider pack. Bind the built-in
launcher into the pack manifest.
- Preserve executable leases, descriptor-backed startup, credential
fences, permissions, and exact ACP model admission.
- Use `builtin:grok-acp` and `native:grok` as profile identities.
Historical package-profile sessions fail closed on resume rather than
being silently reinterpreted.
- Resolve built-in assets from the authenticated sidecar location,
including public server npm layouts. Keep the controller path out of
provider environments.
- Add clean npm tarball installation verification to the existing
trusted canary CI job and the admitted manual EC2 verification path. It
stages a unified release version and runs npm lifecycle scripts, then
verifies missing-prerequisite rejection and admission after separate
provisioning without credentials or inference.
- Include the controller-owned provider pack in stamped Cloud images.
Unstamped local images omit the pack and remain usable; remote ACPX
requires full source provenance.
- Correct CLI approval-page metadata for an already authenticated Cloud
board user; approval authorization remains unchanged.
- Honor explicit native-runner enablement in the Cloud agent picker and
direct setup page, keeping the flag disabled by default.
- Allow selecting the execution environment before connecting
credentials. Include Grok in the existing authenticated hello-probe
flow, targeting its pinned native prerequisite for runner setup.
- Recover an existing subscription sign-in conflict through an explicit
cancel-and-retry action, serialized after cancellation succeeds.
- Preserve the selected ACPX harness before normalizing config fields,
so new Grok agents use the Grok default model.
- Keep the credential-free Cloud provider pack root-owned and readable
after runtime UID remapping; verify manifest and referenced asset access
under an unrelated unprivileged UID during image builds.
- Archive prior failover backups alongside explicitly replaced harness
state, preserving evidence while preventing stale backups from blocking
a fresh replacement.
- Update Daytona image content inputs and contract tests for the
built-in assets and explicit provisioner.
- Document and regression-test the shared `approve-all` default for Grok
setup, saved configuration, and native execution. Explicitly saved
restrictions remain unchanged.

## Verification

Current merge-repair head `df09eb3e1a619430ad8419a0ee9aedd486689b05`
incorporates master `f1a394bd30cb56fb9e479f98b9f50176fe921858` after the
base PR was squash-merged. All 12 conflicts came from incoming files
identical to the tested pre-squash base. The final tree exactly matches
a three-way merge using that original base, preserving built-in Grok
distribution and removal of the obsolete private package. All 252
focused runner/UI tests, six npm-isolation tests, and token gates pass.
Fresh exact-head Greptile review is 5/5 with no outstanding findings;
security scans and EC2 native compilation pass. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([run
36468768035](https://github.com/paperclipai/paperclip/actions/runs/36468768035)).
The repository owner explicitly authorized bypassing code-owner approval
after all checks passed; no CI checks or repository protection settings
are bypassed or changed. The only remaining PR was removed from the
completed stack metadata to permit native auto-merge.

Earlier integration head `78cb306ecc41b5c96577c26c1d89153b0ef865a1`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36447124691)).
The initial attempt lost two EC2 runners to shutdown signals and stalled
a third shard during dependency preparation; all three passed the
same-commit failed-job-only retry. Trunk code-owner requirements remain
enforced. The review summary’s non-blocking saved-asset offset
classification note concerns code already merged in #14301; those
runtime files are identical to master and outside this stack’s diff.
Historical live evidence below retains its original source revisions.
[Final public npm
verification](https://github.com/paperclipai/paperclip/actions/runs/36445542764)
passed on `76ea70cd4d13786a042af9df82f0fd7a8c85ae30`: 17 public
packages, an executed offline lifecycle sentinel, unchanged consumer
lock, built-in launcher, missing-prerequisite rejection, and verified
separately provisioned binary/command lease. Provisioning and cleanup
require no host privilege elevation; only the positive probe mounts the
temporary native binary read-only. The verifier is unchanged by the
final master merge. All six isolation tests and an offline npm smoke
test pass. The prior head had 56 green CI checks and a 5/5 review after
two unchanged tests timed out and passed a failed-job-only retry ([CI
attempts](https://github.com/paperclipai/paperclip/actions/runs/36444597313)).
All 56 recovery-display/lineage tests pass; re-review cleared the
already-covered missed-retry concern. Earlier EC2 failures remain
retained: [npm lockfile
rejection](https://github.com/paperclipai/paperclip/actions/runs/36436311203),
[missing compiler in the slim
image](https://github.com/paperclipai/paperclip/actions/runs/36440210984),
and the aggregate 15-minute test timeouts in those broad runs. Both
broad attempts passed typecheck, token gates, Product E2E type/unit
checks and build. The focused EC2 lane preserves the existing
trusted-actor and immutable-source gates.


Earlier documentation/test checkpoint
`ff244c4fd78a7ede5a3e00efe09f475f133ef33e` leaves runtime behavior
unchanged. 154 focused tests pass across configuration building, native
provider resolution, permission policy, credentials, UI configuration,
and new-agent setup (including both Grok auth modes); token gates pass.
All fresh CI is green for this head: 56 successful checks/statuses and
two intentional skips ([run
36367065119](https://github.com/paperclipai/paperclip/actions/runs/36367065119)).
Greptile is 5/5 with no new findings. Grok already inherits the shared
`approve-all` default, so unattended setup requires no manual permission
change.

Runtime head `bb5a9307991f1ac567b781970ef11b39d518e19b` fixes a final
staging continuation failure before provider startup: explicit
replacement archived the old harness but left its failover backups
active, which caused `runner_harness_state_mismatch`. The regression
fails before the fix and passes after it; all eight adjacent
recovery-safety cases also pass. Old backups remain inspectable inside
the continuity archive. All fresh CI is green at this head ([run
36360839248](https://github.com/paperclipai/paperclip/actions/runs/36360839248)),
with a 5/5 review. One unrelated Cursor test timed out in the initial
server shard; the same-commit failed-job rerun passed, and both attempts
are retained. Staging deployment is confirmed healthy on this revision.
The controller image is
`ghcr.io/paperclipai/paperclip@sha256:6ad91c487910ccd2596ff7aed0a3a3ea5233d12b51b83cd6e1402237749b9673`.
The final browser-created staging task passed on this exact revision
with API authentication: context read → structured human question →
controller restart → answer submission → same native provider session
resumed → document saved → task Done. The two turns took approximately
119s and 77s. The actual write receipt was applied, and the saved
document has exactly one revision containing the selected answer and
requested marker. Usage and cost were not reported. [Controller image
build](https://github.com/paperclipai/paperclip/actions/runs/36360889243).

- Previous integration head `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`:
all CI green (53 successful checks/statuses, two intentional skips),
including repository typecheck/build/tests, native Runner tests, browser
shards, and canary installation checks. [CI run
36358672529](https://github.com/paperclipai/paperclip/actions/runs/36358672529).
Greptile is 5/5 with no unresolved findings.
- Focused checks cover Grok credentials, executable admission, launcher
assets, provider-pack paths/permissions, workflow contracts, setup
defaults, CLI authorization, and subscription conflict recovery. All 39
protocol definitions validate. Final integration checks pass 124
catalog/evidence/cache tests and nine project-form tests; token gates
pass. Some local dependency checks could not load the stale installed
dependency tree; the corresponding fresh EC2 checks pass.
- Clean public npm installation passed on EC2 at
`8b172ebcf8e02e30662d830c00f3961e3bd459ec` ([run
36164964900](https://github.com/paperclipai/paperclip/actions/runs/36164964900)):
17 unified-version packages, lifecycle scripts enabled, built-in
launcher present, no separate Grok package or npm-downloaded binary,
missing prerequisite rejected, separately provisioned native executable
and command lease verified. No credentials or inference were used.
Subsequent changes preserve this npm asset layout.
- The immutable Daytona prerequisite image is
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:98957d5be0ac774d086b6402b5849e8e6356fec70fb8c09fca6eb4ed6de918e0`,
built from `5a2db471f3ddabe77f9f80e76ed27f996cb97fba`. The previous
Cloud controller image was
`ghcr.io/paperclipai/paperclip@sha256:fd914e1ab1e45f741e8e078ff452d16f082d7ac05f9b4b3506d3a3c64150d204`,
built from `a44f7dbb6b6f77cd9ed893756ca453307f281e5f`; it is superseded
by the latest image above. Its EC2 build verified provider-pack access
under an unrelated unprivileged UID.
- Browser staging at `40f898bc4cba73c1dff4e6344a3983ba0fb247ef` passed
full Grok onboarding with the correct `grok-4.7` model, saved credential
delivery, and pinned Daytona execution. A browser-created task read
context and asked the structured human question. After a controller
restart, answering the persisted question resumed the same native
provider session, saved the requested document, and completed the task.
Actual tool outcomes and durable state agree: one question and one
document revision. The two successful turns took 42.7s and 63.1s; usage
and cost were not reported.
- Restricted policy returned the expected `approval_required` outcome.
Functional staging tests explicitly selected `approve-all`; controller
authorization and governed approvals remain enforced. Temporary board
CLI access was revoked and verified rejected (HTTP 401), and the
disposable onboarding agent was paused.

Failures remain retained: the pre-fix continuation failure (its task
remains blocked; the passing final task is fresh), the original Cloud
provider-pack permission failure, the expected restricted-policy denial,
the superseded npm staging failure, and an earlier monolithic CI
infrastructure timeout. Browser CI exposed a project alias/form race;
the final stack uses master's stronger draft-preservation fix and all
browser shards pass. Historical full subscription/API protocol and
Product rosters retain their original source revisions and do not
qualify this packaging revision. No local Docker or Rust build was used.

## Risks

The branch includes master’s draft-preservation fix for project URL
aliases. It keeps the same project’s edit form mounted and clears prior
data when the project or company changes.

Custom sandboxes and local execution hosts must provision the pinned
binary before Grok starts. Missing, changed, unsupported-platform, and
symlinked executables fail admission. The new builtin profile cannot
resume sessions created with the former private-package profile.
Existing Claude/Codex npm bridge profiles retain their package pins.
Grok restricted modes preserve the selected policy but cannot
automatically admit Paperclip calls: ACP permission metadata does not
independently bind tool authority, so those calls stop with
`approval_required`. New Grok configurations default to `approve-all`,
including API configurations that omit the mode. Existing explicitly
restricted configurations remain restricted; controller authorization
and governed approvals remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 14:54:44 -05:00
DottaandPaperclip f1a394bd30 feat(runner): add Grok Build through native ACP (#13882)
## Thinking Path

> - Paperclip manages AI agents and governs their work.
> - Its native runner uses structured provider protocols for sessions
and tools.
> - Grok Build supports ACP over stdio, but the runner did not expose
it.
> - Native execution requires company-scoped credentials, verified
identities, and permission gates.
> - This change adds Grok through ACPX for local and Daytona execution.
> - Subscription login and explicit API-key execution have separate
credential paths.
> - Qualification grades real tool outcomes, durable state, and browser
workflows.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13878, #13901, #13973, #13977, #13979.

Add **Grok Build** to `paperclip_runner` with `provider: "acpx"`,
`acpxAgent: "grok"`, and model `grok-4.7`. Existing legacy Grok agents
keep their adapter. Merge the three companion fixes (#13973, #13977,
#13979) before treating the integrated Product qualification as deployed
behavior.

## What Changed

- Synchronize shared, TypeScript, Rust, server, validation, and UI
provider contracts.
- Run Grok native ACP stdio through ACPX and the authenticated Paperclip
MCP bridge. Verify the pinned executable and exact ACP model identity.
- Prefer company subscription login. Support an explicit company-secret
API key without automatic paid fallback. Fence refresh and copyback to
the same account and remove private runtime credentials after
containment.
- Preserve selected permissions, cancellation, durable session identity,
resume, and restart recovery. Keep unsupported steering and goals
unavailable. Preserve missing usage and cost as unknown.
- Package checksum-verified Grok Build 1.0.13 for Daytona with an
immutable, signed image built on EC2.
- Add deterministic admission, protocol, permissions, identity,
credential, failure, and cleanup checks. Add the maintained 39-case
protocol roster and separate subscription/API Product profiles.
- Fix live-test findings in reasoning events, reloads, idle-owner
retirement, credential-home cleanup, expired-login model discovery,
launcher pinning, and rerun evidence selection.
- Align control-plane state readers with the transport's 64 MiB bound
while retaining identity, ownership, lifecycle, and size rejection
checks.
- Stabilize two asynchronous CI assertions while retaining actual
outcome and filesystem-evidence checks.

## Verification

Current integration head `f114948376056fe0b6b34c1496ae8667b59daa63`
includes master `3447609d2247e75e55d91493dda91a608364f672` (2026-09-28).
Two master advances during verification overlapped the eval catalog; the
final merge preserves Grok qualification, completion updates, and
bounded API-response reading in all 348 cells. All 77 focused
catalog/eval/workflow tests pass. Both native stack layers (#14397) are
mergeable, and both exact-head Greptile reviews are 5/5 with successful
security scans and no unresolved review threads. All current-head CI is
green: 56 successful checks/statuses and four intentional skips ([CI
run](https://github.com/paperclipai/paperclip/actions/runs/36447097232)).
Trunk code-owner requirements remain enforced. The review summary’s
non-blocking saved-asset offset classification note concerns code
already merged in #14301; those runtime files are identical to master
and outside this stack’s diff. Historical live evidence below retains
its original source revisions.


Earlier integration checkpoint:
`24fc9b94ca0afb21ccdc8d26dbb2e4b258ad72cb`. Refreshed against master
`0f14d2612`, preserving Grok qualification alongside the new accounting
and lifecycle suites. All 124 focused catalog, evidence, and
service-worker checks pass. The current base workflow includes the
explicitly selected public-install verification lane; follow-up #14024
supplies its verifier script. CI at that earlier checkpoint was green
(56 successful checks/statuses, four intentional skips), and the review
is 5/5 with no unresolved findings. Prior feature CI at
`fd73f0a9b1ecdf4094685054028df71739ddc3e1` passed ([run
36148259902](https://github.com/paperclipai/paperclip/actions/runs/36148259902));
that is historical evidence, not a current-head result.

Paid Product measurements use frozen integrated source
`2d939a92b21dcaf5c77c88b54d96784d2ddd0699`, which combines the feature
with #13973, #13977, and #13979. That source passed all 52 CI checks and
clean 5/5 review. Later master syncs incorporate upstream changes. Their
checks remain separate from these pinned live measurements.

| Check | Result and source-pinned report |
| --- | --- |
| Subscription protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-36046839612-1/index.html),
runtime `bc6833f7`, evals `92bb4b8c` |
| API protocol roster | [39/39 first attempts; 206
assertions](https://d1p6rlowie26tp.cloudfront.net/runner-protocol-evals/campaigns/gha-35926577007-1/index.html),
runtime `4a1061c8`, evals `3213dbec` |
| Subscription full Product matrix | [16/16 first attempts; 144
assertions; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36096908572-1/index.html),
source `2d939a92` |
| Subscription core repetitions | 18/18: tool use, planning approval,
and Stop/resume each passed three times in local and Daytona profiles.
The full matrix contains repetition one; [repeat
two](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36104551060-1/index.html)
and [repeat
three](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36108100404-1/index.html)
each passed 6/6. Total: 28 unique subscription attempts at `2d939a92`. |
| API smoke and question continuation | [4/4 first attempts; cleanup
passed](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-36147315401-1/index.html),
both environments at `2d939a92` |
| Historical API Product coverage | [16/16 full
matrix](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-35875144860-1/index.html)
and 18/18 core repetitions at `4a1061c8`; retained as measurements of
that revision |
| Native Daytona proof | Three subscription and three API
MCP/permissions/resume runs passed at `bc6833f7`. Three expired-login
admission and fenced refresh checks passed without inference. All test
sandboxes were removed. |
| Inspectable artifacts and UI | Current-source screenshots verify
planning approval, direct Ask completion, question continuation after
controller restart, and two downloadable project revisions. The project
downloads pass 12 and 18 tests; all 40 independent artifact oracle
checks pass. |
| Provider-free checks | 116 eval-validator tests, 39 Grok definitions,
and 359 enabled/external campaign cells pass. Continuation regressions
above 2 MiB and 16 MiB failed before their fixes; 32 focused
recovery/ownership/size checks pass. |

The 32 unique current-source Product attempts have no failures, retries,
or skipped cells, and all cleanup checks pass. Whole-workflow timing,
model identity, image and provider-pack provenance, attempts, and
accounting coverage are retained in the canonical reports. The report
publisher's conservative `complete=false` flag is preserved; independent
audits verify the exact selected source catalog and immutable result
rows.

Pins: Grok Build `1.0.13 (5e9a58528b76)`, ACPX `0.13.1`, ACP model
`grok-4.7`. Linux binary SHA-256:
`edf79521581bb5e6b95abef848491a6a742e860da3e237ebe86a280d30dce4c1`.
Launcher SHA-256:
`f0b698395a3704ed2ffaf84ea19bdb20c36c8a0a70b7c629c7b6ffe144e59e55`.
Image:
`ghcr.io/paperclipai/paperclip-daytona-runner@sha256:76b24edfd850219e949418b19e4ceba690e84d51d199ade426e484953329b5e9`.
Image build source is `4196a4cd`, recorded separately from application
source `2d939a92`; each campaign verifies the image signature and
provider pack.

Original failed campaigns remain available: [continuation
bound](https://github.com/paperclipai/paperclip/actions/runs/36057718059),
[scheduler/event
capture](https://github.com/paperclipai/paperclip/actions/runs/36071063537),
and [startup cleanup plus EC2
interruption](https://github.com/paperclipai/paperclip/actions/runs/36080870743).
They retain their original grades. No Docker or Rust builds ran on the
developer laptop for these follow-ups.

## Risks

Merge packaging follow-up #14024 with this base before public release.
The follow-up replaces the private Grok bridge package with a built-in
launcher and makes the native binary an explicit sandbox prerequisite.

Three separate, reviewed fixes are part of the tested integrated
behavior: #13973 serializes task-run admission; #13977 captures complete
event evidence; #13979 durably reconciles failed Daytona creation. Each
has green CI and clean 5/5 review. Failed-create recovery has 277 plugin
tests, 92 SDK tests, host-runtime recovery tests, and a real Daytona
lost-deletion-receipt proof. The live proof uses a private file for
journal persistence; database durability is covered by host tests.
Worker death before delivery of a failure envelope remains outside that
recovery mechanism.

Subscription fixtures stage an authorized company login; interactive
browser sign-in is not qualified. Local Product profiles ran on EC2
Linux. The temporary subscription credential was removed from the
protected GitHub environment after all subscription audits, with absence
verified. Runtime homes and refresh copyback remain ownership-fenced.

Protocol results remain pinned to their original revisions; they are not
relabeled as tests of the latest feature commit. New binary/model
versions require qualification. Missing token usage and model cost
remain unknown; runtime estimates do not establish a full bill.
Automatic paid Grok scheduling remains disabled pending separate
reviewed enablement. The 64 MiB bound can increase memory use for
verbose sessions, and larger files still fail closed. No automatic
legacy-agent migration occurs.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 13:54:36 -05:00
Devin FoleyandPaperclip 8ee8f1fd6e ci: retire recurring public cloud image builds (#13827)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Core publishes standard images and source verification for
downstream services.
> - Managed services can now compose private images from the signed
standard image.
> - Core still builds a second public cloud image on every master push
and release.
> - That duplicate producer consumes build capacity and retains an
obsolete readiness contract.
> - This pull request retires recurring cloud publication while
preserving the standard producer and rollback artifacts.

## Linked Issues or Issue Description

Refs #13797 and #13789. Related: #12856 changes image dependency
packaging; it does not retire this producer.

**What existing behavior does this improve?**

Core's recurring Docker publication and Cloud readiness workflow.

**Current behavior**

Master pushes call the legacy cloud publisher from Cloud readiness.
Release tags and manual Docker runs call it too. Canary promotion also
requires the legacy image.

**Proposed behavior**

Publish standard Core images and retain `Cloud source verified v1`. Let
downstream services build their managed image. Keep explicit commit
previews and existing images available.

## What Changed

- Remove `docker-cloud.yml`, its master and release callers, and its
unused cache selector.
- Remove the legacy image/migrator wait and `Cloud deployable v1` job.
Keep the full source verification workflow and exact source-proof name.
- Make canary promotion inspect and promote the standard image only.
- Preserve signed standard-image publication, direct migrator
publication, and explicit `release.yml` previews. The preview path still
uses the Dockerfile `cloud` target.
- Update workflow, preview, build-stamp, and packaging tests. Exercise
the promotion shell with mocked registry commands, including
missing-image and missing-tag cases.
- Document frozen legacy aliases, consumer requirements, preview
compatibility, and rollback retention.

## Verification

- All 377 workflow tests pass: `node --test
.github/scripts/tests/*.test.mjs`.
- All 129 release-registry tests pass: `pnpm test:release-registry`.
- Focused source-proof, standard-image, preview, and workflow tests
pass: 256 tests.
- Focused image packaging/build-stamp tests pass: 16 tests.
- Actionlint passes on all three changed workflow files. `git diff
--check` passes.
- Full local `pnpm build` and `pnpm -r typecheck` pass.
- The policy follow-up updates an old assertion that required the
removed readiness job. All 37 source-proof/release-workflow tests pass
locally.
- Full local `pnpm test:run` did not complete successfully while the Mac
ran out of disk space. No full-suite pass is claimed. Removed 1.2 GiB of
generated Cargo output from this isolated worktree with `cargo clean`.
GitHub CI passed on the final head: 52 successful checks and 2 optional
skips.
- Fresh Greptile review for `4f5fe1951f0bd7f7739cf6655d395ff78f1ed944`:
**5/5**, successful current-head check, zero review threads.
- September 23 refresh: the unchanged PR head merges cleanly with
current master `db8f8fe5b73a2697684a30261b0d306a9c631aba`. In an
isolated temporary worktree, all 377 workflow tests and 29
release/preview tests pass on the combined tree. `git diff --cached
--check` passes.
- Refreshed Actionlint workflow validation passes with ShellCheck
disabled. Full Actionlint reports the same 10 existing ShellCheck
diagnostics as master, with no added diagnostics. No source changes or
new PR commits were needed.
- The full local build/typecheck and current-head Linux CI results above
remain the verification for the unchanged PR head. They were not rerun
for this metadata-only refresh. No image publication or tenant
deployment was initiated for this refresh.

## Risks

**Deployment prerequisite satisfied (September 23):** The combined
cleanup release is deployed to staging and production, and production
Support is verified. Active managed-fleet automation uses standard-image
composition. Explicit immutable previews remain supported by the
retained preview publisher. This PR is ready for maintainer review; keep
auto-merge disabled and wait for explicit merge authorization.

- A consumer still selecting `Cloud deployable v1` will stop advancing
at the last legacy-ready commit. Confirm active automatic consumers use
the standard-image composition contract before merge.
- Legacy cloud release-channel aliases stop advancing. Standard
self-hosted aliases continue.
- This PR deletes no registry images, cache tags, migrators,
credentials, or runner infrastructure. Existing immutable releases
remain usable for rollback.
- Explicit legacy previews remain for commit-specific operator
deployments. Retiring that compatibility path requires a separate
consumer migration.
- These changes affect CI publication, not database schema or
application behavior.

## Model Used

OpenAI Codex, GPT-6. The runtime does not expose a more specific model
identifier or context-window size. Used repository inspection,
reasoning, code editing, shell tools, and test execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 19:35:19 -07:00
DottaandPaperclip db8f8fe5b7 fix(evals): select Grok subscription protocol credentials explicitly (#13901)
## Thinking Path

> - Paperclip manages agent work through shared runner contracts.
> - Direct protocol evals qualify provider behavior against a mock
control plane.
> - Grok supports API keys and company subscription credentials.
> - The hosted protocol workflow selected an API key for every Grok
cell.
> - Product subscription support did not enable subscription protocol
runs.
> - This change adds explicit subscription selection and checks the
recorded authentication mode.

## Linked Issues or Issue Description

Refs #13878, #13882, #12618.

The direct Grok protocol roster cannot run with subscription
authentication through the trusted default-branch workflow. Add an
explicit selector while keeping API-key dispatches compatible. Keep the
actor allowlist, protected environment, immutable source revisions, and
publication gates.

## What Changed

- Add `grok_authentication` with `api_key` and `subscription` choices.
Keep `api_key` as the compatibility default.
- Deliver the protected `GROK_AUTH_JSON` secret only to a
subscription-selected Grok cell. Do not provide an API key to that cell.
- Read authentication mode from the pinned eval program's actual roster
summary. Retain it in the cell, catalog, campaign roster, and result.
- Reject missing or mismatched authentication evidence during
aggregation. Preserve cell metadata and an allowlisted failure reason
before failing a cell, so malformed evidence cannot hide the retained
attempt.
- Document credential setup, source separation, and temporary-secret
cleanup.

## Verification

- `node --test
packages/paperclip-runner/scripts/runner-protocol-eval-campaign.test.mjs
packages/paperclip-runner/scripts/runner-protocol-eval-workflow-security.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`: 34 tests passed.
- Validated all 39 Grok cells at eval revision
`3213dbec7e8ca1865ea95e6db7e7d34b095eb47a`; every selected cell requests
only the subscription credential. Validation made zero provider calls.
- Negative coverage rejects invalid selectors and missing or API
authentication evidence in an otherwise passing subscription attempt.
- `git diff --check` passed. All 53 current-head checks passed; the
unchanged callback-drain timing test passed its bounded rerun, and the
failed attempt is retained. Greptile reviewed
`ecd3dcc0e998f07cf56fcb1f087946f50f388bec` at 5/5 with no remaining
findings.
- No Docker or broad builds ran on the developer machine. CI performs
repository checks on the configured fleet.

## Risks

Grok runs require an eval revision that records `authenticationMode` in
the roster summary. Missing evidence fails closed. The credential
contains account access and refresh tokens; an owner must approve its
delivery to the protected environment before a live run. The change adds
no PR trigger or authorization bypass. Live subscription protocol
qualification remains pending this workflow reaching master.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 18:43:21 -05:00
DottaandPaperclip b648d8cdda fix(evals): support explicit Grok qualification workflows (#13878)
## Thinking Path

> - Paperclip manages AI agents and their provider connections.
> - Product E2E checks real tasks through the browser, server, and
runner.
> - Grok qualification needs separate API-key and subscription evidence.
> - Product subscription tests and direct Grok protocol evals need
explicit credential delivery.
> - This change supplies each credential only to its selected profile
and prepares the pinned binary.
> - Maintainer authorization and protected-environment gates remain
required.

## Linked Issues or Issue Description

Refs #13845, #13847, #13850, #13882.

The Grok feature branch has a manual subscription qualification profile.
The trusted master workflow must admit its selected credential and
prepare the same verified binary and artifact verifier as the API
profile. Direct protocol evals also need the selected xAI key and pinned
Grok binary. These prerequisites do not register or schedule the new
profiles on master.

## What Changed

- Deliver `GROK_AUTH_JSON` from the protected paid environment only when
the selected profile requests that credential.
- Install the checksum-verified Grok binary for the local subscription
profile.
- Prepare the pinned artifact verifier for the manual subscription
suite.
- Extend workflow security assertions to cover the new credential and
profile.
- Add the ACPX Grok credential mapping to the trusted-master catalog,
then deliver only the selected `XAI_API_KEY` to direct protocol cells
and install the target’s checksum-verified Grok binary before packaging.
- Allow a direct-protocol concurrency override from two cases up to the
existing configured ceiling; it can only lower concurrency.
- Document the Grok protocol workflow and its API-only credential
boundary.
- Render missing LLM usage and cost as Unavailable, and label partial
observations with coverage. Preserve raw records, grades, and actual
zero costs.
- Preserve measured campaign source metadata during report regeneration
instead of inheriting the renderer checkout or CI event; skip empty
legacy source records when recovering older provenance.

## Verification

- Latest commit `05d05801477104c8155977bbbe3e119a5241f960`: all 54
reported checks successful, two intentional skips, Greptile 5/5, and
zero unresolved review threads. [CI
run](https://github.com/paperclipai/paperclip/actions/runs/35890978288).

- After merging current master, all 17 workflow security/image tests and
23 catalog/workflow policy tests passed. The trusted catalog also
generated all 39 pinned Grok cells with `XAI_API_KEY` and one case per
shard. The new policy tests execute the concurrency guard against valid,
out-of-range, and malformed values.
- The Grok branch separately passed 450 Product harness unit tests,
including private company credential staging, cleanup, and
token-fragment redaction.
- The fresh-login native subscription smoke passed three repetitions of
tool execution, session resume, restrictive permissions, and cleanup.
These are setup evidence; full subscription Product qualification
remains pending.
- All 72 focused report/billing/history/catalog tests and the Product
harness typecheck passed for the report-display change. The initial
sandbox run could not open the tsx IPC socket; the permitted rerun
passed. A zero-provider-call replay of the actual 16-cell Grok campaign
preserved all result records, grades, timing, and source provenance
while correcting missing usage labels.
- Review the thirteen-file diff. Provider credentials still enter only
the selected paid-test step; default-branch, numeric-actor, and
environment restrictions are unchanged.

## Risks

This admits a refreshable subscription credential to explicitly selected
trusted tests. Store it only in `runner-e2e-paid`, use a test login, and
remove it after qualification. Unselected profiles receive an empty
value. Pull requests cannot trigger the paid workflow. This PR changes
no fleet admission or actor allowlist.

## Model Used

OpenAI GPT-6 through Codex, with tool use and code execution. The exact
serving model identifier and context-window size are not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 17:05:52 -05:00
DottaandPaperclip 60c7c9cd1a fix(runner-e2e): pass verified lock digest to Daytona image build (#13876)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Product E2E campaigns test the native runner in local and Daytona
environments.
> - Each campaign resolves one target lockfile and verifies its
downloaded artifact.
> - The Daytona image job did not pass that artifact digest to Docker.
> - Docker used an older default digest and stopped before any selected
task ran.
> - This pull request passes and validates the campaign digest at the
image build boundary.
> - The image keeps its checksum check and frozen package installation.

## Linked Issues or Issue Description

**What happened?**

The merged-master [qualification
campaign](https://github.com/paperclipai/paperclip/actions/runs/35863582409)
stopped in the Daytona image build. The resolved target lock digest was
`e0c928a494f90ddad3c00791e83f09315ee8c82df2a0418a809dcf93649a8ab3`.
Docker used its default digest,
`57b298aceebc48bb94ea0593347348256475da7b2fddb77025d9e57cc8759420`. The
checksum check rejected the mismatch. All 12 selected model cells were
skipped. This follows the image provenance work in #13814.

**Expected behavior**

The image build must check the same lockfile artifact that the campaign
restored and verified. A changed lockfile must still fail the checksum
check.

**Steps to reproduce**

1. Start a Product E2E campaign with a Daytona cell on master
`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.
2. Resolve a target lockfile whose digest differs from the Dockerfile
default.
3. Observe the provider-pack image stage reject the lockfile before
model execution.

**Paperclip version or commit**

`7944ed3d976d1a7cc26a2d0cee51f227f3542084`.

**Deployment mode**

GitHub Actions Product E2E campaign with a Daytona image build.

**Install method**

Built from source with the campaign lockfile artifact.

**Agent adapter(s) involved**

Native Codex and ACPX Claude cells were selected. No model cell ran in
this failed campaign.

**Database mode**

Not involved. The failure occurs during image creation.

**Access context**

The authorized default-branch paid workflow. The build receives no
provider credentials.

## What Changed

- Read the image checksum from the existing target-lock job output.
- Require a 64-character lowercase hexadecimal digest before image
inspection or build.
- Pass the digest as the existing Docker build argument.
- Add regression checks and document the campaign checksum handoff.

## Verification

- The Daytona image regression fails with the original workflow and
passes with the fix.
- All six Daytona image contract tests pass.
- All 450 Product E2E unit tests pass.
- Product E2E typecheck passes.
- Actionlint passes for the changed workflow.
- A context-shaped resolution probe preserves the downloaded lockfile
bytes and digest.
- All latest-head CI checks passed on
`67d41fd9439b2a9a809ddb05765f8617585072c5`
([run](https://github.com/paperclipai/paperclip/actions/runs/35865739359)).
- Greptile gave 5/5 on this head; its test-scoping comment is addressed
and resolved.
- A hosted Daytona image rebuild and the three remote qualification
cells remain pending after merge.

## Risks

The campaign digest comes from the existing trusted target-lock job. The
restored artifact checks, Docker checksum check, frozen install, content
identity, image signing, and verification remain in place. The
standalone Docker default remains available. This change does not alter
task behavior, prompts, credentials, or dependency versions.

## Model Used

OpenAI `gpt-6-astra` through Codex performed diagnosis and review with
code execution tools. OpenAI `gpt-5.6-luna` assisted with investigation,
implementation, and verification. Context window limits are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-23 08:45:52 -05:00
DottaandPaperclip ac854a9b59 ci: build Daytona eval images on the authorized fleet (#13847)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Product E2E campaigns test the browser, server, and runner together.
> - The trusted workflow selects an authorized execution fleet.
> - Daytona image builds still use a fixed GitHub-hosted runner.
> - This pull request applies the existing fleet selection to image
builds.
> - Campaign builds and tests then use the configured EC2 fleet.

## Linked Issues or Issue Description

Refs: https://github.com/paperclipai/paperclip/pull/13845

**What existing behavior does this improve?**

The location of Daytona Product E2E image builds.

**Current behavior**

The image job uses ubuntu-latest even when RUNNER_E2E_AWS_ENABLED
selects EC2 for the rest of the campaign.

**Proposed behavior**

Use the existing authorized runner output for the image job. Preserve
the GitHub-hosted fallback when the EC2 switch is off.

## What Changed

- Route Daytona image builds through the existing authorized runner
selection.
- Assert that the image job depends on authorization and receives no
provider credentials.
- Document the build routing and credential boundary.

## Verification

- Product E2E workflow security: 11 tests passed.
- git diff --check passed.
- Reviewed actor checks, workflow triggers, target commit selection,
signing, and package permissions. They are unchanged.
- Required CI checks pass on the latest head. The first CI attempt had
failures in unchanged chat tests; one diagnostic rerun passed. Both
attempts remain in Actions.
- Greptile reviewed the latest head at 5/5 with no inline findings.
- Live EC2 image verification follows after this trusted workflow change
is merged. No local Docker execution.

## Risks

EC2 image builds depend on the fleet having working Docker and
sufficient disk space. Existing authorization, image digest checks,
signing, and the GitHub-hosted fallback remain in place. No provider
secrets are added to the image job. No application or database changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool
execution. The exact deployment model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 21:30:09 -05:00
DottaandPaperclip 950ccb8eef ci: enable Grok qualification in the trusted paid workflow (#13845)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Product E2E tests verify tasks through the browser, server, and
runner.
> - These paid tests use a trusted workflow from master and an isolated
target commit.
> - The Grok target branch selects XAI_API_KEY, but the trusted workflow
does not deliver that credential.
> - Its artifact test also needs the pinned Python verifier before
execution.
> - This pull request adds both bindings inside the existing paid
boundary.
> - The tests can then run on the configured EC2 fleet without laptop
Docker.

## Linked Issues or Issue Description

Related evaluation infrastructure:
https://github.com/paperclipai/paperclip/pull/11297. No duplicate Grok
paid-workflow change was found.

**What existing behavior does this improve?**

Branch-targeted Grok Product E2E qualification in the existing paid
workflow.

**Current behavior**

Grok cells cannot receive their selected API credential. The Grok
build-revise case also misses the artifact-verifier setup step.

**Proposed behavior**

Deliver XAI_API_KEY only when the selected matrix credential is
XAI_API_KEY. Prepare the existing pinned verifier for the Grok
qualification suite.

**Reason and benefit**

Run the controller, browser, runner and artifact checks on the EC2
fleet. Preserve default-branch workflow authorization and protected
environment secret access.

## What Changed

- Bind the selected XAI credential only in the paid test step.
- Install the checksum-verified Grok binary for local cells before
provider access.
- Include Grok qualification in the existing pinned artifact-verifier
preparation.
- Add an optional max_parallel input that can only lower the configured
campaign concurrency. Use 1 for the Grok test key.
- Extend security assertions and document setup.

## Verification

- Ran the Product E2E workflow-security tests: 11 passed.
- Checked the diff for whitespace errors.
- Reviewed credential selection, setup ordering, numeric actor gates,
target commit pinning, and trusted report checkout.
- Live Grok execution follows after this workflow is available on
master. This PR does not claim completed Grok qualification.

## Risks

The paid test step can use the selected XAI credential and incur
provider charges. The credential remains in runner-e2e-paid and is
absent from setup, build, and reporting jobs. The default-branch gate
and existing environment restrictions remain in place. No database
migration or product behavior changes.

## Model Used

OpenAI Codex, GPT-6 family, with reasoning, code editing, and tool
execution. The exact deployment model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 20:17:25 -05:00
Devin FoleyandPaperclip 92d4868e79 fix(server): isolate run errors and redact runtime capability headers (#13826)
## Thinking Path

> - Paperclip manages AI agents and their work.
> - Operators use Sentry to investigate failed runs and server errors.
> - Run reports attach a task ID, run ID, error code, and adapter
fingerprint.
> - The server skips Sentry's OpenTelemetry setup to preserve its
separate tracing and privacy settings.
> - Without an async context manager, a scope mutation can attach old
run data to later errors.
> - The HTTP logger also retains a runtime credential capability header.
> - This change isolates run metadata and redacts that header so
diagnostics identify failures without leaking credentials.

## Linked Issues or Issue Description

Refs #13446 and #13719.

**What happened?**

After a terminal run failure, an unrelated server exception can inherit
that run's tags, context, and fingerprint. Sentry then groups a database
error with an earlier adapter failure. The real SDK reproduces this with
the application's `skipOpenTelemetrySetup: true` setting. HTTP request
logs also retain the `x-paperclip-github-capability` header, which must
be treated as a credential.

**Expected behavior**

Run metadata belongs to the terminal run event. Later exceptions must
not inherit it. Every genuine error must still be captured. Runtime
capability headers must be redacted on success and failure logs.

**Steps to reproduce**

1. Initialize the optional Sentry SDK with the application's options and
an in-memory transport.
2. Capture a terminal run failure.
3. Capture an unrelated exception.
4. Inspect the second event. Before this fix, it contains the first
run's identity and fingerprint.
5. Send a request with a fixture runtime GitHub capability header.
Before this fix, HTTP logs retain the fixture value.

## What Changed

- Pass tags, context, and fingerprint directly to `captureException`
instead of mutating the ambient scope.
- Preserve the existing run fields, grouping keys, ordinary exception
capture, and privacy settings.
- Test two run identities interleaved with unrelated exceptions against
the real optional SDK.
- Update the capture contract tests and document event-local run
metadata.
- Redact the runtime GitHub capability header through the existing HTTP
logger policy. Test successful, denied, and failed requests.
- Add a dedicated GitHub-hosted CI check that installs the exact
optional SDK version declared in `server/package.json`. It fails if the
real-SDK regression would be skipped. The SDK stays outside the
workspace and production dependency graph.

## Verification

- The real-SDK regression failed before the fix because the unrelated
event contained `contexts.run_failure`.
- Five focused suites passed: 123 tests, including all optional SDK
tests. Suites: `run-failure-sentry-real-sdk.test.ts`,
`run-failure-sentry.test.ts`, `sentry.test.ts`,
`run-failure-report.test.ts`, and `http-log-redaction.test.ts`. A custom
in-memory transport prevented outbound Sentry delivery.
- All three new header-redaction cases failed before the policy fix and
passed afterward.
- The dedicated CI command passed locally with
`PAPERCLIP_REQUIRE_SENTRY_TEST_SDK=1` and the audited SDK available
through `NODE_PATH`.
- Server TypeScript check passed with a scratch configuration that
resolves this checkout's workspace packages. The existing dependency
links point to another checkout.
- `node scripts/check-module-boundaries.mjs` and `git diff --check`
passed.
- Gitleaks and a separate private-data scan passed before push.
- Full local workspace typecheck, test, and build were not run. The
machine has less than 2 GiB free and those commands include Rust builds.
Full PR CI must pass before merge.
- The dedicated real-SDK GitHub check passed with 1 test executed and no
skips: https://github.com/paperclipai/paperclip/actions/runs/35774449002
- Greptile reviewed c9db03bcab at 5/5. Its only thread is resolved. Full
PR CI passed on that same head:
https://github.com/paperclipai/paperclip/actions/runs/35774449020

## Risks

Small change to error attribution. Unrelated errors may now form their
correct Sentry groups instead of reopening a prior run group. No errors
are filtered or suppressed. No tracing is enabled and no new event
fields are added. No schema or runtime-execution changes. HTTP logs
retain their request and status diagnostics while masking the capability
value. The new SDK job has read-only permissions, no secrets, and an
in-memory Sentry transport.

## Model Used

OpenAI GPT-6 (Codex), with tool use and code execution.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` OR (b) described the issue in-PR following the relevant issue
template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub references)
- [x] My branch name describes the change and contains no internal
Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 12:53:31 -07:00
Devin FoleyandPaperclip e3d8fb0876 feat: attest standard production images at full source commits (#13797)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Operators deploy its standard production container on several CPU
architectures.
> - Downstream image builders need to identify the exact source of their
base image.
> - A short commit tag does not provide signed source evidence.
> - This pull request adds a full commit tag and signed image digest for
canonical master pushes.
> - Consumers can verify the source and compose from the immutable
digest.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

Publication of the standard multi-platform production image.

**Current behavior**

The Docker workflow publishes short commit tags and channel tags. It
does not provide a signed standard-image contract tied to the complete
master commit.

**Proposed behavior**

Canonical master pushes also publish `sha-<full-commit>` and attest the
exact index digest after platform validation and an immutable-image
orphan-reaping check. The signer certificate binds the source
repository, commit, workflow and ref. Existing tags and the separate
cloud producer remain available.

**Reason and benefit**

Downstream builders can prove the source of a standard base without
adding their dependencies or repository details to the public workflow.
No matching open issue or duplicate PR was found.

## What Changed

- Add the canonical full-SHA tag without changing existing tag mappings.
- Validate amd64 and arm64 descriptors and hash the exact registry
response bytes and require its digest header to match.
- Verify the immutable image and sign it with GitHub artifact
attestations.
- Run the contract tests in trusted PR verification and document the
consumer contract.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/standard-image-contract.test.mjs`: 37 passed.
- A read-only check against an existing published index returned its
exact expected digest.
- `actionlint -shellcheck='' .github/workflows/docker.yml
.github/workflows/pr-trusted.yml`: passed. Normal ShellCheck reports
only existing `ls` and word-splitting warnings.
- `pnpm build`: passed locally with Cargo available.
- `pnpm -r typecheck`: passed locally.
- Full local Vitest was attempted: 8,260 passed, 14 failed, with 34
failing suites. The failures were missing embedded-PostgreSQL library
aliases in this fresh install and existing macOS runtime-skill-cache
rename errors. Native aliases are now restored. Rerunning the 33
affected database suites produced 577 passes and two unrelated AgentMail
skill-root lookup failures (32 suites passed). The four directly failing
database tests also pass independently. This is not a claim that the
full local suite passed.
- All final-head CI checks pass. One unrelated routine-route mock
assertion passed on the single-shard retry; its 15 tests also pass
locally. Greptile is 5/5 on this exact head, with no unresolved threads.
- Actual signing requires a canonical master push. This draft PR does
not publish trusted provenance.

## Risks

The new attestation step requires OIDC and attestation write permissions
in the merge job. Signing failure leaves the image available but without
the new admission proof. Consumers must fail closed when proof is
missing. Existing release tags, the legacy producer, and image retention
remain unchanged. No database or application behavior changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository inspection, shell
execution and test tools. The session does not expose a more specific
model variant or context-window size.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-22 07:33:30 -07:00
3c2bc4f546 ci: halve the isolated native Runner check by seeding it from the public master build cache (#13736)
## What

Recurring CI health check (PAP-31): on the most recent fully-green PR
run (`cb703ac`, run
[35519542997](https://github.com/paperclipai/paperclip/actions/runs/35519542997)),
the slowest check was **Compile isolated native Runner** at **452s** —
ahead of the largest test shards (379s). Every run recompiled the full
Rust dependency tree from zero, even though `docker.yml` already
refreshes a **public** `mode=max` BuildKit cache
(`ghcr.io/paperclipai/paperclip:buildcache-{amd64,arm64}`) on every
master push, containing exactly these layers. This PR seeds **only the
baseline build** with that registry cache, via anonymous pull.

**Measured on this PR's own CI (which exercises the seeded path): the
check completed in 123s, down from 452s — a 73% reduction, ~5.5 minutes
saved per run.**

## Thinking Path

Cost breakdown of the 452s from the job log: `cargo chef cook`
dependency compile 214.7s, local cache export 43.1s, runner-core build
37.0s, metadata proof layer 36.8s, `cargo install cargo-chef` 36.4s,
rebuild-verification build ~65s, setup/teardown ~20s. The dependency
compile and toolchain layers are identical to what the production
`docker.yml` build already caches publicly on every master push, so
recompiling them here bought no signal — the check's real assertions
live in the *verification* build, not the baseline. A first attempt used
`actions/cache` plus a master `push` trigger, but the CI bot's GitHub
App lacks `workflows` permission; the registry-cache approach is
strictly better anyway (shared across PRs immediately, no 10GB
Actions-cache quota pressure, no workflow change).

## What Changed

- `scripts/check-docker-runner-cache.sh`: the baseline build now adds
`--cache-from
type=registry,ref=ghcr.io/paperclipai/paperclip:buildcache-{amd64|arm64}`
(selected by host arch). `RUNNER_CHECK_SEED_CACHE` overrides the ref, or
set it empty to force the old cold path. The script header documents the
anonymous external read.
- `.github/workflows/docker-runner-check.yml` (comment-only): the stale
"no external cache" note now describes the anonymous GHCR seed and the
verification build's local-cache-only isolation. This was pushed in a
follow-up commit with workflow-edit permissions; the original CI-bot
token could not touch workflow files.

No Dockerfile stages or verification assertions changed.

## Verification

- This PR's own `Compile isolated native Runner` check runs the seeded
path (the script is in the workflow's trigger paths): **passed in 123s**
vs the 452s baseline.
- The rebuild-verification semantics are untouched: it still runs on a
**fresh builder** importing **only the local cache exported by this
run's baseline**, so it proves exactly what it proved before — that the
runner image rebuilds reproducibly from this run's own exported layers.
- Verified `ghcr.io/paperclipai/paperclip:buildcache-amd64` is
anonymously readable (unauthenticated manifest pull succeeds), so the
check gains no credential or secret dependency.

## Risks

- **Stale or missing seed cache:** if the GHCR ref is unreachable,
private, or garbage-collected, BuildKit logs a warning and falls back to
the pre-PR cold compile — the check gets slower, never wrong.
`RUNNER_CHECK_SEED_CACHE=""` restores the cold path explicitly.
- **Cache trust:** the seed only accelerates the *baseline* build; the
verification build still runs on a fresh builder against only this run's
locally exported cache, so a stale or poisoned registry cache cannot
make verification pass spuriously. The ref lives under
`ghcr.io/paperclipai/*`, written only by repo CI on master pushes.

## Model Used

Claude Fable 5 (`claude-fable-5`) via Paperclip agent **Bender
(Fable)**, issue PAP-31.

---------

Co-authored-by: Bender (Fable) <bender-fable@paperclip.local>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-09-21 20:22:48 -07:00
3790ca2f13 fix(runner): repair approval and Stop races and eval infrastructure (#13750)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Runner tasks must continue after approval and stop when the user
presses Stop.
> - Live evals found races at approval delivery and provider startup.
> - Browser readiness and CI setup errors also hid the actual task
results.
> - This pull request fixes those races and the related test
infrastructure.
> - Regression tests and saved live reports show which cases now pass.

## Linked Issues or Issue Description

Companion eval definitions PR:
https://github.com/paperclipai/paperclip-evals/pull/25 (AgentCore paused
and provider/environment infrastructure).

Related: #13741 now supplies the late-startup Stop fence and
warm-attachment recovery; this PR retains that fence and extends startup
tracking and regression coverage to both native backend paths. #13539
introduced queued approvals during active runs. #13738 fixes child
assignment, task replies, and warm process continuity and is already in
the base. #13291 concerns automatic continuation of interrupted legacy
sandbox runs; this PR fixes native startup cancellation and does not
change that recovery policy.

**What happened?**
An accepted service approval could wait after its source run stopped.
Stop could return success before the provider handle existed. Work could
then start after Stop, or a cancelled run could be recorded as failed.
Some E2E tests also failed on unloaded browser content or irrelevant
reply wording. Runner CI could fail before model work because of
dependency or sandbox setup.

**Expected behavior**
Deliver each settled approval once after its source run stops. Do not
start work after an acknowledged Stop. Preserve the audited
cancellation. Test the intended product behavior with a ready browser
and verified runtime dependencies.

**Steps to reproduce**
1. Approve a service request while its source run is active. Let the run
finish. Check that its result starts one continuation.
2. Delay provider startup. Press Stop before its handle is available.
Check cancellation, then submit `/new`.
3. Run the browser, warm-workspace, and Stop-and-redirect cases from the
linked report.

**Paperclip version or commit**
The branch includes master at `9d19f98b5`. The report records the
original source for each focused attempt.

**Deployment mode**
Isolated local development instances and disposable Daytona sandboxes.

## What Changed

- Deliver settled tool-action results for the exact company and source
run during final cleanup. Keep the existing idempotent receipt and
periodic recovery sweep.
- Wait for startup to hand off its provider handle before acknowledging
Stop. Reject first-turn admission after cancellation. Preserve a
matching audited pending or acknowledged cancellation.
- Wait for mounted task history and connector controls in browser tests.
Record failure evidence. Grade workspace contents and process continuity
separately from exact reply wording. Require each warm-turn marker once
and in order, allowing surrounding prose.
- Stop-and-redirect now checks that the source file exists and work is
active before Stop.
- Resolve target dependency locks in an uncredentialed CI job. Verify
the lock artifact hash. Keep orchestration and publication on the
trusted workflow revision.
- Materialize the pinned OpenCode executable and configure the exact
Codex executable's user-namespace profile before provider credentials
are available.
- Compress Daytona directory uploads with gzip. Preserve files,
executable modes, symlinks, empty directories, and confinement checks.
- Classify file-transfer RPC deadlines as infrastructure. Keep unrelated
runner RPC failures visible.

## Verification

- [Focused live report with screenshots and original
attempts](https://pages.paperclip.ing/runner-reliability-20260921/): 14
of 15 selected Product E2E cases pass across the recorded revisions.
Claude and Codex Stop → `/new`, Claude service approval, delegation,
both hiring/reuse cases, and native Daytona warm continuity pass.
- Two credentialed Runner smoke cases pass. These are not full protocol
coverage.
- E2E harness after the master merge: 429 tests pass. E2E and server
TypeScript checks pass.
- Daytona plugin: 239 tests pass, 6 skipped. Plugin TypeScript build
passes. The compression test fails against the old code and passes with
the change.
- Runner backend/runtime regression group: 161 tests pass.
Cancellation/startup selection: 26 tests pass. Approval delivery: 34
real-database tests pass.
- Workflow security: 7 tests pass. Both edited workflows pass
actionlint. Runner TypeScript and Rust builds pass.
- After merging master, all 389 native executor tests pass, including
both native backend paths and late startup after the Stop deadline.
- Post-merge `pnpm -r typecheck` and `pnpm build` pass. The monolithic
local `pnpm test:run` was interrupted to integrate master and is
inconclusive. The [hosted CI test
partitions](https://github.com/paperclipai/paperclip/actions/runs/35620461738)
pass on `50a3e43822bcba1e0d07b1b45b0be91cbf9312da`. An unchanged sandbox
callback schema test initially received HTTP 503. It passed five
isolated local runs, its full local test file, and one failed-job CI
retry. No assertion was weakened.

## Risks

- Stop can wait for the bounded startup handoff. If it cannot settle,
the existing pending-recovery state remains instead of a false
acknowledgement.
- Immediate approval delivery must remain idempotent across cleanup and
recovery sweeps. Tests cover duplicate delivery and company/run
boundaries.
- The workflow changes still need hosted Linux verification. They retain
the trusted workflow and credential boundaries.
- Gzip reduces the observed provider upload from about 1.8 GB to 663 MB.
It does not yet fix the remaining Claude Daytona transfer timeout. That
recovery test never reached Claude, so recovery remains unverified. Use
a matching image with the verified provider package preinstalled for the
next recovery test; retain cold-upload coverage separately.
- The report preserves diagnostic runs with missing source metadata and
marks them as such. It does not claim a new full-suite pass.
- This PR adds no new prompt policy or historical status reconciliation.

## Model Used

OpenAI GPT-6 through Codex performed the primary implementation and
review. The exact primary backend model ID is not exposed in this
session. OpenAI `gpt-5.6-luna` assisted with bounded infrastructure work
and verification. The agents used repository tools, code execution, and
browser tests. The exact backend revision and context-window size are
not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: OpenAI GPT-6 <noreply@openai.com>
2026-09-21 12:50:50 -05:00
DottaandPaperclip 4cfc0e3f6c fix(runner-e2e): provision complete worker prerequisites (#13674)
## Thinking Path

> - Paperclip manages work performed by AI agents.
> - Runner E2E tests verify that work in local and remote environments.
> - The full catalog exposed setup failures before agents could execute
tasks.
> - ACPX Codex missed sandbox provisioning, and new artifact cases
missed Docker preparation.
> - A host Claude version probe also stopped native Claude stories that
use a packaged provider.
> - This pull request repairs trusted worker setup and keeps its
selectors covered by catalog tests.
> - The resulting reruns can measure behavior instead of missing
prerequisites.

## Linked Issues or Issue Description

Follow-up to #13655. Full-catalog campaign:
https://github.com/paperclipai/paperclip/actions/runs/35417932353.

**What happened?**

ACPX Codex could not start its sandbox. New artifact stories missed
Docker preparation. Native Claude stories tried to spawn an unrelated
host CLI. The report job failed on trusted lockfile drift.

**What did you expect to happen?**

Prepare each worker's required capabilities before paid execution and
publish the retained results from trusted code.

**Steps to reproduce**

Run local ACPX Codex, everyday agent-review-handoff, or native Claude
everyday cells in the full-stack workflow at 43acbcc39.

## What Changed

- Include local ACPX Codex in exact-binary sandbox provisioning and
preflight.
- Prepare the pinned artifact oracle for agent-review-handoff. Do not
require it for skill creation.
- Remove host CLI version probes from native everyday stories.
- Retry GitHub actor lookup up to three times, with a deadline, while
retaining all authorization checks.
- Resolve reporter dependencies from the trusted checkout before its
frozen install.
- Test worker selections against the catalog and document the
default-branch requirement.

## Verification

- `pnpm test:e2e:runner:unit`: 383 tests pass.
- `pnpm test:e2e:runner:typecheck`: passes.
- PR checks: 54 passed, two intentionally skipped. Greptile: 5/5, no
inline findings.
- The paid workflow reads trusted setup from master. These setup changes
need to land before the affected Linux cells can verify them. Runtime
fixes and other affected reruns are on a separate branch.

## Risks

The setup selectors decide which workers receive sandbox policy and
Docker preparation. Catalog coverage checks their scope. Actor lookup
still fails closed. Reporter dependency resolution uses only the trusted
checkout; target branch code does not receive publication credentials.
No production prompt changes.

## Model Used

OpenAI GPT-6 via Codex, with repository inspection, code editing, and
test execution. The exact API model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 09:27:15 -05:00
1ef3b08714 feat(ui): integrate agent personas across the app (#13171)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - A stable agent persona is useful only when the same identity appears
across the app.
> - Lists, task messages, selectors, and activity feeds need inexpensive
static avatars.
> - Onboarding and agent headers need a larger character with
expressions and pointer tracking.
> - This pull request connects the persona foundation to those existing
views and preserves onboarding draft assignments.
> - Full-page stories and Linux checks make the placements and
performance contract reviewable.

## Linked Issues or Issue Description

**Problem or motivation**

Agents need a stable visual identity in lists, tasks, onboarding, and
configuration. External tools also need an image URL for that identity.

**Proposed solution**

Assign each agent a permanent palette from a fixed ClipLab character
library. Store the assignment on the agent. Render and cache preset PNG
URLs on demand. Use static images in dense views and one animated
character in larger placements.

**Alternatives considered**

A generated image bundle requires a separate asset build. A live
renderer in every avatar adds unnecessary work in large lists. Arbitrary
uploaded images do not provide the requested shared character system.

**Roadmap alignment**

This improves agent identity across existing control-plane views. It
preserves agent permissions, company boundaries, and status labels.
ROADMAP.md has no separate ClipLab persona milestone.

Related approaches: #2422 adds configurable image URLs and DiceBear
generation; #5578 adds optional uploaded avatars. This work uses a
fixed, versioned character library and preset URLs.

## What Changed

- Replace agent icons with static persona images across lists, the
sidebar, org charts, tasks, comments, selectors, activity, and dashboard
views.
- Put one animated character in the agent header. Let it follow the
pointer across the page, with reduced-motion and touch fallbacks.
- Add larger padded characters to agent creation. Keep the palette
stable across draft refreshes and connection retries, then reveal it
after success.
- Pass appearance through shared projections rather than fetching each
agent separately.
- Add real full-page Storybook examples for the agent list, overview,
task, dashboard, new-agent dialog, and connection page.
- Add Linux screenshot, clipping, density, and 500-avatar performance
checks.

## Verification

- `pnpm -r typecheck`, `pnpm build`, and token gates pass on the rebased
tree. Persona lifecycle tests pass.
- The rebased feature passes 38 Linux screenshot/performance checks,
including both display densities, corner pointer positions, and the
no-WebGL/no-live-download contract for 500 avatars.
- The final Linux persona suite passes all 38 visual, lifecycle,
density, and full-page checks using the standard Storybook configuration
and real on-demand avatar endpoint.
- Final local focused verification: 45 avatar/native-recovery tests
pass; UI identity/routine tests, typecheck/build, token gates, and
Storybook build pass.
- Current-head CI passes: full workspace/server tests, all serialized
server groups, typecheck/release checks, build, canary validation, and
end-to-end shards. The build passed after retrying a native-runner
concurrency-test failure; its three targeted cases also pass locally.
- Manual inspection covered stable identities in the app, header
placement, full-page mouse tracking, onboarding size, and task/dashboard
placements.


### Screenshots

Linux captures use synthetic Storybook fixtures. Full-page captures use
reduced motion. The live character, mouse tracking, and disposal are
checked separately.

<details>
<summary>Agent overview with the character in its header</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-agent-overview.png"
width="900" alt="Agent overview with the character in its header" />

</details>
<details>
<summary>Task messages and assignee identity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-task.png"
width="900" alt="Task messages and assignee identity" />

</details>
<details>
<summary>Larger onboarding character with room for expressions</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-meet-your-next-agent.png"
width="900" alt="Larger onboarding character with room for expressions"
/>

</details>
<details>
<summary>Dashboard agent activity</summary>

<img
src="https://raw.githubusercontent.com/paperclipai/paperclip/8c68f42b268ada22b67d79d0fe1bb0a2f84ec25c/screenshots/full-page-company-dashboard.png"
width="900" alt="Dashboard agent activity" />

</details>

## Risks

- This PR depends on #13170, the persona foundation. Merge the
foundation first, then retarget this PR to master.
- Many placements change from icons to character silhouettes. Human
avatars and authoritative agent status labels retain their existing
behavior.
- Only one character can render live per view. Reduced motion,
hidden/offscreen content, touch input, and renderer failures use the
defined fallbacks.
- The full-page stories use fixture data. They do not contact a real
company or complete real provider sign-in.

## Model Used

OpenAI Codex, GPT-6 family. The exact model identifier and context
window are not exposed in this session. Used code editing, shell
execution, browser inspection, and Linux visual testing.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Tonio <tonework@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-19 07:57:53 -05:00
DottaandPaperclip 43acbcc398 fix(runner): preserve sessions and complete question and approval continuations (#13655)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - The native runner connects task state to provider sessions.
> - Follow-up turns must retain provider memory and carry new user
direction.
> - Lost session IDs caused repeated context and extra input tokens.
> - Native question answers and approval races could leave valid work
blocked.
> - This pull request repairs those paths and adds regression coverage.
> - Agents can continue accepted work without repeating the conversation
or losing the user's answer.

## Linked Issues or Issue Description

Refs #13574. That merged PR shortened continuation prompts and moved
question instructions into tool documentation. This change preserves
sessions and fixes failures exposed by broader testing. Related runtime
work: #13408 and #13410.

**What happened?**

Native follow-up turns could lose the provider session ID. Completion
guidance could replace the original task with its latest comment. Claude
native questions could remain pending after the user answered. Approval
during a running tool call could suspend the run before the tool
response arrived. Onboarding and chat handoff instructions also caused
repeated planning or missing plan documents.

**Expected behavior**

Reuse a valid provider session. Send only new events when that session
already has the history. Preserve the task requirements and apply later
user direction. Store the question answer and deliver it to the waiting
run. Finish governed tool responses before suspending. Execute the
accepted plan without asking for the same approval again.

**Steps to reproduce**

Run the continuation, local-session-integrity, first-task, and
agent-chat suites with native Codex and Claude. Include
provider-question-bridge, accept-while-running, and plan-handoff.

**Paperclip version or commit**

This branch is based on master d54b75011. The active full catalog run
tests 4e75881db. Later review fixes have separate regression coverage.

**Deployment mode**

Isolated local instances and Daytona sandboxes in the existing Runner
full-stack E2E harness.

## What Changed

- Retain provider session identity across turns and late usage
snapshots. Send new continuation events on session reuse, with full
context available for a fresh session.
- Preserve task requirements and later direction in completion guidance.
Return the current contract revision after a stale completion
submission.
- Bridge native Claude questions to saved Paperclip cards. Submit
answers through the saved card and resume the same run.
- Delay governed suspension until tool results settle. Add a
deterministic test barrier for approval during an active run.
- Clarify free-text question examples, explicit onboarding plans, and
execution of accepted chat plans.
- Fix continuation readiness, verified output evidence, and declared
screenshot collection.
- Qualify the legacy Claude test CLI at 2.1.277. The old 2.1.19 CLI did
not discover mounted skills. Update the existing workflow pin and
isolated launcher together.
- Refresh the Daytona image lockfile integrity pin after reviewing
master patch updates.
- Carry continuation mode as runtime metadata instead of inferring it
from user-visible text. Install the test Claude CLI without lifecycle
scripts.

## Verification

- Targeted paid verification: 20/20 cases passed across
local-environment campaigns before the rebase.
[Report](https://pages.paperclip.ing/runner-e2e-seven-fixes-35397904249/).
- Harness checks: 379 unit tests passed; harness typecheck passed.
- Latest-head PR checks: 55 passed, two intentionally skipped. Greptile
is 5/5; the security scan passes.
- Review regressions: 350 executor tests and 204 session/driver tests
passed. A script-free Claude install was verified with the actual CLI.
- Full catalog, including the explicit-only everyday suite: [run
35417932353](https://github.com/paperclipai/paperclip/actions/runs/35417932353).
Completed: **164/205 passed; 41 failed**. [Full dashboard and failure
investigation](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/).
Includes 204 case artifacts and one pre-case GitHub authorization
timeout; missing evidence is not scored as a pass. The full run tested
`4e75881db`; Final-head metadata/CLI smoke cases both passed. In the
separate [six infrastructure
retries](https://github.com/paperclipai/paperclip/actions/runs/35419769343),
the GitHub timeout case passed and all five Docker preflight failures
repeated. [Follow-up
dashboard](https://pages.paperclip.ing/runner-e2e-full-catalog-35417932353/follow-up/).
- Full local typecheck and build passed on the rebased branch. The full
local unit run completed with 657 passing files, two test timeouts and
one suite setup timeout. All three affected files passed when rerun in
isolation (84 tests). The first full local run was not clean.
- Focused regression coverage includes the live question bridge,
same-run response delivery, UI routing, stale revisions, approval
overlap, and session reuse.

## Full-catalog follow-ups

- Test infrastructure: 14 Claude everyday cells probe an absent host
CLI; six cells failed pre-task GitHub/Docker qualification (GitHub
passes on retry; all five Docker cases repeat; the workflow preflight
allowlist omits their case IDs); five ACPX Codex cells cannot create
sandbox namespaces.
- Runtime: four OpenCode completion-criteria mismatches masked by
shutdown errors, one service-approval suspension failure; three Daytona
recovery failures encounter existing skill files; one duplicate
completion wake.
- Confirmed test defects: question pagination and a noncanonical plan
document key.
- Product/behavior: mismatched visible/required question sets, an
attachment instead of the requested task document, one lone-option
onboarding question, early completion instead of review, and a Codex
Mini completion-schema failure.
- The report job itself fails on trusted master’s stale patch/lock
configuration. The linked report is rebuilt with the shared renderer
from original cell results and public fixture screenshots; it excludes
private snapshots, logs and traces.

These are investigated follow-ups, not silently regraded passes.
First-task passed 51/52. The PR checks are green independently of the
broader catalog’s behavioral/infrastructure failures.

## Risks

- Session reuse depends on a valid provider identity and context
coverage. Fresh-session fallback and reset tests cover this boundary.
- Native question delivery spans saved interaction state and a live
provider run. Tests cover duplicate events, closed runs, and same-run
answers.
- Provider behavior varies. The full paid catalog may expose failures
beyond these targeted fixes; those results will be reported without
relaxing valid approval or output checks.
- The legacy Claude version update is limited to test infrastructure. No
database migration is included.

## Model Used

OpenAI Codex, GPT-6 family, with repository inspection, code execution,
and browser/E2E tools. The exact runtime model identifier and
context-window size are not exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (targeted checks and all
three timeout-file reruns pass; full-run timeout caveat above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-19 07:42:57 -05:00
Devin Foley d08abcba15 ci: cut PR wall clock from ~16 to ~6 minutes (#13521)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every pull request runs the Trusted PR CI workflow before merge
> - The test suites roughly tripled in six weeks, and shard balance did
not keep up, so PR runs crept from ~4 to ~17 minutes
> - Slow CI delays every merge and every contributor
> - This pull request rebalances the shards from fresh measurements,
splits the largest test files, reuses the Rust build cache in three more
jobs, and takes the policy job off the critical path
> - The benefit is a PR wall clock near 6 minutes with the same coverage

## Linked Issues or Issue Description

**What existing behavior does this improve?**

PR CI wall clock. A typical green run took 16-17 minutes. Two months ago
it took about 4 minutes.

**Subsystem affected**

The Trusted PR CI workflow (`.github/workflows/pr-trusted.yml`), the
shard-duration manifests, the vitest shard runner scripts, the
`paperclip-runner` package scripts, and the dry-run branch of
`release.sh`.

**Current behavior**

The shard-duration manifests were stale. The general-server manifest had
durations for ~400 of 649 suites. The e2e manifest was missing 14 of 29
specs. Stale median weights made shard steps range 417s-806s (server)
and 277s-745s (e2e). Three jobs each paid a ~3m40s cold cargo release
build. Every test lane waited ~60s for the policy job before it could
start.

**Proposed behavior**

All lanes finish in a narrow ~200-290s band. The manifests carry fresh
measured durations for every suite. The three largest test files are
split so no single file caps a shard. The Rust cache restore runs in
every job that builds the Runner binary. Test lanes start as soon as the
gate resolves.

**Reason and benefit**

Merges stop waiting on CI. The projected wall clock is ~6 minutes for
the same test coverage.

## What Changed

- Rebuild `scripts/general-server-shard-durations.json` (646 suites) and
`scripts/e2e-shard-durations.json` (all specs) from per-suite completion
timestamps in runs 35036001734 and 35024948947.
- Move the PR server lane to the release-verify shape:
`general-server-without-chat` across twelve duration-balanced shards,
plus the chat integration suite split by collected test location across
three dedicated lanes.
- Split `tests/e2e/chat-adapters-ui.spec.ts` into `-providers` and
`-messaging` specs, and `tests/e2e/agent-chat.spec.ts` into `-sessions`
and `-projects` specs. Each pair shares fixtures through a `.shared.ts`
module. Playwright collects the same test sets (39 and 20 tests).
- Raise e2e shards to eight and serialized shards to nine.
- Run the runner package's `check:all` as four matrix lanes:
`check:static`, `check:runner`, and two native vitest `--shard` halves.
The union is exactly `check:all`.
- Add the read-only Rust cache restore (toolchain pin, `save-if: false`)
to the Canary Dry Run, Build, and Typecheck jobs.
- Make release.sh preview publish payloads concurrently in batches of
eight during `--dry-run`. The real publish path stays strictly serial.
- Drop the policy-job lockfile artifact chain. Each lane installs with
`--frozen-lockfile` and falls back to an inline `--resolution-only`
regeneration. The policy job stays a required check through the `verify`
and `e2e` aggregates.
- Update the shard-count mirrors and workflow assertions in the
partition and gate tests.

## Verification

- `node --test scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/e2e-shard.test.mjs` — 30 pass.
- `node --test '.github/scripts/tests/'*.test.mjs` — 410 pass.
- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-dry-run-notes.test.mjs` — 42 pass.
- `playwright test --list` collects 39 tests across the chat-adapters
split and 20 across the agent-chat split, equal to the original files.
- A local vitest collection of the chat suite partitions 995 tests into
498/497 line shards.
- Projected shard weights: server 230s x12, chat ~143s x3, e2e 207-242s
x8, serialized ~216s x9.

## Risks

- The split spec files reorder tests relative to the original files.
Every describe seeds its own company, so the specs stay independent; a
hidden cross-describe dependency would surface as a deterministic
failure in one shard.
- The inline lockfile fallback changes install behavior for
manifest-changing and stacked PRs. The policy job still validates
resolution as a required check.
- `release.sh` changes are confined to the `--dry-run` preview branch.
The publish loop is untouched. `bash -n` passes and the release dry-run
tests pass.
- One PR now schedules ~44 fleet runners. If the RunsOn fleet caps
concurrency, queueing may absorb part of the gain; watch the first runs.

## Model Used

- Claude Fable 5 (`claude-fable-5`, Anthropic), extended thinking, with
tool use (shell, file edits) in Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge
2026-09-16 11:45:14 -07:00
4577d10029 fix: prepare everyday artifact and Codex sandbox prerequisites in CI (#13516)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - The runner E2E workflow executes paid everyday workflow stories on
disposable CI hosts
> - The everyday artifact oracle requires a pinned Python image and
fails closed when it is absent
> - Fresh CI hosts did not prepare this image before the paid cell, so
project stories failed during preflight
> - This pull request prepares and verifies the pinned image before the
affected everyday cells
> - The benefit is reliable artifact isolation checks on fresh trusted
CI hosts

## Linked Issues or Issue Description

**What happened?**

Fresh trusted CI runners did not have the pinned Python artifact oracle
image.

**Expected behavior:**

The workflow prepares and verifies the pinned image before an everyday
project story starts.

**Steps to reproduce:**

Run an everyday project story on a fresh CI host without the image
cached. The `everyday-artifact.py --preflight` check fails before task
creation.

**Paperclip version or commit:**

`master` at `bd51f157e`.

**Deployment mode:**

Other: GitHub Actions trusted paid workflow.

## What Changed

- Add a matrix-gated CI step for everyday project and recovery cells.
- Check Docker, pull the fixed digest with bounded timeouts, and verify
the exact repo digest.
- Apply the existing Codex sandbox preparation to both native Codex
profiles, including the mini profile.
- Add workflow security assertions for ordering, condition, digest,
timeouts, and secret isolation.
- Document that CI prepares the pinned oracle image.
- Check provisioning eligibility against every catalog cell, and scope
the Daytona registry inspection assertion to the Daytona image job.

## Verification

- `pnpm test:e2e:runner:unit` — 24 files and 222 tests passed.
- `pnpm test:e2e:runner:typecheck` — passed.
- `pnpm exec vitest run --config tests/runner-e2e/vitest.config.ts
workflow-security.test.ts` — 10 tests passed.
- Python artifact oracle calibration — 12/12 passed.
- `git diff --check` — passed.

## Risks

Low risk. The image step runs only for everyday cells that execute the
artifact preflight. The Codex setup now covers both native Codex
profiles. It uses a fixed public image digest and has no provider
credentials.

## Model Used

OpenAI gpt-5.6-luna (implementation subagent) and gpt-6-astra (review
fixes and orchestration), using code execution and repository tools.
Context window size is not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
Co-authored-by: Paperclip <paperclip@paperclip.ing>
2026-09-16 07:49:48 -05:00
Nicky LeachandClaude Opus 5 bd51f157e9 fix(ci): make the Runner Rust cache key independent of the image toolchain (#13500)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change goes through the pull request CI workflow, and its
`Verify Paperclip Runner` lane builds and tests the native Rust runner
> - https://github.com/paperclipai/paperclip/pull/13457 made that lane
restore master's prebuilt Rust dependency cache instead of recompiling
313 crates
> - Measured over 26 runs since it went live, the cache hits on the
RunsOn fleet and misses on every GitHub-hosted runner
> - The cause is the cache key: `rust-cache` hashes every installed
toolchain, and each runner image ships a different stable Rust next to
the pinned one
> - This pull request removes the extra toolchains before the key is
computed, in the reader and the writer
> - The benefit is that the saving the cache already delivers, 4.8
minutes per run, reaches the 73% of runs that currently miss it

## Linked Issues or Issue Description

No public GitHub issue exists for this. The problem follows the
enhancement issue template below.

- Refs https://github.com/paperclipai/paperclip/pull/13457 — added the
cache this pull request repairs
- Refs https://github.com/paperclipai/paperclip/pull/13459 — activated
it
- Refs https://github.com/paperclipai/paperclip/pull/13194 — created the
`release-runner-v1` entry on master

I searched this repository for other pull requests touching this cache
and found no duplicate and nothing in flight.

**What existing behavior does this improve?**

The Rust dependency cache added in #13457 misses on GitHub-hosted
runners, so most pull requests still recompile the whole dependency
tree.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

The cache works, but only on one runner class. Across 26 successful
`Verify Paperclip Runner` jobs since #13459 merged:

| Runner | Runs | Before | After | Change | Cache |
|---|---|---|---|---|---|
| RunsOn fleet | 7 | 16.4m | 11.6m | −4.8m | 4 of 4 full hit |
| ubuntu-latest | 19 | 14.9m | 14.7m | −0.2m | 0 of 6 hit |
| All | 26 | 15.0m | 13.9m | −1.1m | |

Only 27% of runs reach the fleet, so the fleet-wide saving is 1.1
minutes rather than the 4.8 minutes the cache delivers where it lands.

The two runners compute different keys:

```
fleet:         v0-rust-release-runner-v1-Linux-x64-4bb3b8ea-a95b0328
ubuntu-latest: v0-rust-release-runner-v1-Linux-x64-9fdc73e3-a95b0328
```

The lockfile half agrees. The environment half does not. `rust-cache`
logs why, under `Environment considered`:

| Runner | Toolchains it found |
|---|---|
| fleet | 1.97.1 and **1.98.0** |
| ubuntu-latest | 1.97.1 and **1.98.1** |

`rust-cache` hashes every installed toolchain, not only the active one.
Both images carry the pinned 1.97.1. Each also ships its own stable
Rust, and those differ by a patch version. The post-merge writer runs on
a RunsOn image, so the fleet agrees with it and GitHub-hosted runners
cannot.

Pinning `RUSTUP_TOOLCHAIN` in #13457 was necessary but not sufficient.
It fixes which toolchain builds the code. It does not change which
toolchains exist on the image.

**Proposed behavior**

Remove every toolchain except the pin, before the cache step, in both
the reader and the `release-runner-v1` writer. The key then depends on
the pinned compiler and the lockfile alone, not on what the image
happens to carry.

**Reason and benefit**

The cache already proves its value where it lands: release compile drops
from 5m07s to 1m30s, debug from 1m48s to 13s, and the job from 16.4m to
11.6m. This change extends that to the other 73% of runs.

Expected fleet-wide mean: about 11.5m, against 13.9m today and 15.0m
before #13457.

It also removes a standing fragility. The fleet hits today only because
two RunsOn images happen to agree. If either image updates its stable
Rust on its own, the hit rate drops to zero with no code change.

**Breaking changes**

None. The change only affects cache key computation. A miss reproduces
the current behavior.

**Additional context**

The typecheck writer in `release-verify.yml` keeps its current step on
purpose. It restores and saves on the same post-merge image, so its key
never disagrees with itself.

## What Changed

- Added a toolchain normalization block to `Select the pinned Runner
Rust toolchain` in `.github/workflows/pr-trusted.yml`, before the cache
restore. It keeps the pinned toolchain and uninstalls the rest.
- Added the identical block to the same step in the
`verify_paperclip_runner` job of `.github/workflows/release-verify.yml`,
which writes `release-runner-v1`. The reader and the writer must agree,
or the key matches nothing.
- Made the block tolerant. If a toolchain cannot be removed it prints a
notice and continues, so a pull request loses the cache rather than the
run.
- Extended `.github/scripts/tests/pr-runner-rust-cache.test.mjs` with
two tests: the block is byte-identical in both workflows, and it runs
before the cache step in each.

## Verification

Run the workflow shape tests:

```bash
node --test '.github/scripts/tests/*.test.mjs' ./scripts/__tests__/e2e-shard.test.mjs ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs ./scripts/cloud-source-verification.test.mjs
```

Result: 471 pass, 0 fail.

I ran the step body against a stub `rustup` to confirm the logic, rather
than only checking syntax. Three paths, all exit 0:

| Case | Result |
|---|---|
| Pin plus an extra stable toolchain | Uninstalls only the extra,
exports `RUSTUP_TOOLCHAIN=1.97.1-x86_64-unknown-linux-gnu` |
| Pin only, already normalized | No uninstall calls, no error |
| No `rustup` on `PATH` | Prints the notice and continues |

I also mutation-tested the new parity assertion. Each mutation edits
only the writer, then the reader and writer disagree:

| Mutation to `release-verify.yml` | Result |
|---|---|
| One word changed in the shared comment | Caught |
| `uninstall` changed to `remove` | Caught |
| Trailing `rustup toolchain list` deleted | Caught |

After this merges, confirm the repair in CI. Take a `Verify Paperclip
Runner` job that ran on `ubuntu-latest` and check the restore step for
`full match: true`. Under `Environment considered`, `Rust Versions` must
list only 1.97.1. The job should finish near 11.5m rather than 14.7m.

## Risks

- **Master must republish the cache once.** This changes the key, so the
existing `release-runner-v1` entry no longer matches.
`cloud-readiness.yml` runs `release-verify.yml` on every push to master,
so the first push after this lands writes the new entry. Pull requests
merged before that miss the cache, which is exactly what most of them do
today. No run breaks.
- **The reader and the writer must stay in step.** If they diverge, no
run hits the cache. The new parity test fails on any difference in the
block, including a comment.
- **Uninstalling the image toolchain is deliberate.** Everything in this
job builds through the pinned 1.97.1, resolved from
`packages/paperclip-runner/rust-toolchain.toml` and `RUSTUP_TOOLCHAIN`.
Nothing in the job uses the image default.
- **Low blast radius.** The change only affects cache key inputs. A miss
compiles from scratch, as today.
- **Image drift is now handled.** The key no longer depends on the
image's own Rust version, so a future image update cannot silently
disable the cache.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
the `gh` CLI and the GitHub API to pull 26 post-merge job records and 10
full job logs, log parsing in Python to isolate the per-phase timings
and the two cache keys, a stub `rustup` on `PATH` to exercise the new
step, and local `node --test` runs to verify and mutation-test the
guards. Run through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The CI and Greptile boxes stay unchecked
until those checks finish. On documentation: no document describes the
Runner cache keys, so there is nothing to update.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 17:24:22 -07:00
DottaandPaperclip 4510bf7c9e ci: use code-owner-reviewed master for trusted PR workflow (#13470)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Pull request CI uses a trusted workflow on the AWS runner fleet.
> - The caller used a fixed SHA that also needed runner-group admission.
> - A mainline pin update left CI queued because the group still allowed
older SHAs.
> - This pull request calls the trusted workflow on master, which
requires code-owner review.
> - New merged workflow versions can use the existing master
runner-group entry.

## Linked Issues or Issue Description

Related: #12968. That Dependabot PR proposes another SHA rotation. This
change keeps this first-party workflow on master instead.

**What happened?**

CI run 34975562974 stayed queued because its trusted workflow SHA was
absent from the runner-group allowlist. The fleet itself was healthy.

**Expected behavior**

New reviewed versions of the trusted workflow on master should receive
runner access without a separate SHA allowlist update.

**Steps to reproduce**

Change the caller to a new trusted workflow SHA without adding that SHA
to the restricted runner group. Its jobs remain queued. The master
reference removes that recurring synchronization step.

## What Changed

- Call `paperclipai/paperclip/.github/workflows/pr-trusted.yml@master`.
- Exclude this exact first-party workflow from Dependabot updates.
- Update the existing E2E shard workflow tests for the master caller
contract.
- Document the runner-group entry, required code-owner review, and
old-reference retention.

## Verification

- `actionlint .github/workflows/pr.yml` passed.
- `node --test scripts/__tests__/e2e-shard.test.mjs
.github/scripts/tests/cloud-runner-routing.test.mjs
.github/scripts/tests/pr-runner-rust-cache.test.mjs
.github/scripts/tests/pr-dependency-cache.test.mjs` passed: 35 tests.
- Parsed Dependabot YAML and checked the exact workflow exclusion.
- `git diff --check` passed.
- Live GitHub checks confirmed `.github/**` has code owners, CODEOWNERS
has no errors, and the active master ruleset requires code-owner review.
This covers the trusted workflow, caller, and CODEOWNERS itself.
- The approved organization setting now allows
`paperclipai/paperclip/.github/workflows/pr-trusted.yml@refs/heads/master`.
All previous references and other runner-group settings remain intact.
- Full application typecheck, tests, and build were not repeated locally
for this workflow-only change. This PR's CI and review are pending.

## Risks

New versions of the trusted workflow take effect for new callers after
merge to master. Keep code-owner review and master protection enabled.
Existing administrator pull-request bypasses remain unchanged.
Third-party action pins, runner routing, and infrastructure are
unchanged. Older callers still use their SHA pins; their allowed refs
remain in place.

## Model Used

OpenAI GPT-6 through Codex performed implementation and orchestration;
its exact runtime variant and context-window size were not exposed.
OpenAI `gpt-5.6-luna` with high reasoning inspected workflow assumptions
and applied the approved runner-group setting. Both used code and tool
access.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 09:43:24 -05:00
Nicky Leach 5b913e7943 ci: activate restore-only Rust dependency cache (#13459)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change to Paperclip goes through the pull request CI workflow
before it merges
> - That workflow is split in two on purpose: `pr.yml` is the caller,
and it pins `pr-trusted.yml` at an immutable SHA
> - The pin means a change to `pr-trusted.yml` on master does nothing
until someone advances the pin
> - https://github.com/paperclipai/paperclip/pull/13457 added a
read-only Rust dependency cache to the `Verify Paperclip Runner` lane,
and it is inert for that reason
> - This pull request advances the pin, which is the second step of that
rollout
> - The benefit is that the saving measured in #13457 starts to apply,
about 3.9 minutes per run and about 4.7 compute-hours each day

## Linked Issues or Issue Description

This pull request is the activation half of a two-step rollout. #13457
merged on 2026-09-15, so this pull request now targets master directly.

- Refs https://github.com/paperclipai/paperclip/pull/13457 — added the
cache step this pull request activates. Merged as `f97a3f886`.
- Refs https://github.com/paperclipai/paperclip/pull/13302 — the
previous activation, and the change that introduced the `# Pin:` comment
convention this pull request follows
- Refs https://github.com/paperclipai/paperclip/pull/13300 — the change
#13302 activated, and the current pin target

I searched this repository for other pull requests that move this pin.
One is open:

- Refs https://github.com/paperclipai/paperclip/pull/12968 — an
automated bump of the same pin. See Risks.

**What existing behavior does this improve?**

The pull request CI lane still recompiles the full Rust dependency tree
on every run, because the cache step added in #13457 is not yet part of
the active CI definition.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

`.github/workflows/pr.yml` pins `pr-trusted.yml` at `44dde2de`, the
squashed commit of #13300. GitHub reads `pr.yml` from the pull request
and takes every job from `pr-trusted.yml` at that SHA. A change to
`pr-trusted.yml` on master therefore has no effect on any pull request
until the pin advances.

#13457 is the only change to `pr-trusted.yml` since that pin, and it is
currently inert.

**Proposed behavior**

Advance the pin to the commit that carries the cache step, and update
the `# Pin:` comment to name the pull request it activates.

**Reason and benefit**

The saving measured in #13457 begins to apply. Master's own warm-cache
lanes run the same checks in 3.4 minutes against 7.1 minutes cold. The
net saving is about 3.9 minutes per run after the 20 second restore,
across about 73 runs each day.

**Breaking changes**

None. The activated change only adds a cache restore. A cache miss
reproduces today's behavior exactly.

**Additional context**

The last five activations all landed on the same day as the change they
activated: #13302, #12860, #12810, #12509, and #12464. This pull request
follows that convention. #13457 merged today.

## What Changed

- Advanced the `uses:` pin in `.github/workflows/pr.yml` from `44dde2de`
(#13300) to `f97a3f886`, the squashed merge of #13457.
- Updated the `# Pin:` comment to name #13457 and the capability it
activates, matching the convention #13302 introduced.

The diff is the same two lines every previous activation changed.

## Verification

Run the workflow and pin tests:

```bash
node --test ./scripts/__tests__/e2e-shard.test.mjs ./scripts/__tests__/run-vitest-stable-shard.test.mjs ./scripts/__tests__/release-verify-workflow.test.mjs ./scripts/cloud-source-verification.test.mjs '.github/scripts/tests/*.test.mjs'
```

Result: 469 pass, 0 fail. This includes `pr.yml calls the trusted PR
workflow at an immutable SHA`, which reads the pinned workflow out of
git and asserts on its content.

Confirm the pin resolves to a workflow that contains the cache step:

```bash
git show $(grep -oE '[0-9a-f]{40}' .github/workflows/pr.yml):.github/workflows/pr-trusted.yml | grep -c "Restore Runner Rust dependencies (read only)"
```

This prints `1`.

This pull request also verifies itself. GitHub uses the pull request's
own `pr.yml` for `pull_request` events, so this run takes its jobs from
the newly pinned workflow. The `Verify Paperclip Runner` job in this run
is therefore the cached version, running the exact definition this pull
request makes active. Check its log for `Cache restored from key:
v0-rust-release-runner-v1-Linux-x64-...`, confirm cargo prints no
`Compiling` lines for third-party crates, and compare the job duration
against the 15.0 minute baseline recorded in #13457.

## Risks

- **An automated pin bump is open and may race this.** #12968 moves the
same pin. It is a no-op today, because `pr-trusted.yml` is identical
between the two SHAs. If it rebases after #13457 lands, its target moves
to a commit that contains the cache step, and merging it would activate
the change with a stale `# Pin:` comment that still names #13300.
Closing #12968 before merging this avoids the ambiguity.
- **The activated change itself is low risk.** It only adds a read-only
cache restore. A miss reproduces today's behavior. #13457 records the
full risk list.
- **The rollback is one commit.** Restoring the previous pin value
returns CI to the current definition without touching `pr-trusted.yml`.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
`git` to confirm the merge strategy, the pin history, and the squashed
merge SHA, the `gh` CLI and the GitHub API to read the repository merge
settings and to find the open automated bump, and local `node --test`
runs to verify the pin resolves and the guard tests pass. Run through
Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The CI and Greptile boxes stay unchecked
until those checks finish on this pull request. On tests: the existing
pin guard in `scripts/__tests__/e2e-shard.test.mjs` already covers this
change, so this pull request adds no new test. On documentation: no
document describes the pull request CI pin, so there is nothing to
update.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-15 01:19:14 -07:00
Nicky Leach f97a3f886e perf(ci): restore master's Rust dependency cache on the PR runner lane (#13457)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every change to Paperclip goes through the pull request CI workflow
before it merges
> - One lane of that workflow, `Verify Paperclip Runner`, builds and
tests the native Rust runner
> - That lane compiles all 313 third-party crates from scratch on every
pull request, in two profiles
> - Master already builds and stores exactly those compiled crates, but
the pull request lane never reads them
> - This pull request restores that existing cache on the pull request
lane, read-only
> - The benefit is about 3.9 minutes less compute per run, or about 4.7
compute-hours each day, with no change to what CI checks

## Linked Issues or Issue Description

No public GitHub issue exists for this. The problem follows the
enhancement issue template below.

Related merged pull requests, found by searching this repository:

- Refs https://github.com/paperclipai/paperclip/pull/13194 — created the
`release-runner-v1` cache that this pull request reads
- Refs https://github.com/paperclipai/paperclip/pull/13300 — established
the restore-only dependency cache pattern that this pull request follows
- Refs https://github.com/paperclipai/paperclip/pull/13326 — split the
post-merge runner job into the `protocol` and `rust` lanes used as the
warm-cache baseline below
- Refs https://github.com/paperclipai/paperclip/pull/13259 — applied the
same caching idea to the post-merge typecheck job

I found no open pull request that duplicates this work.

**What existing behavior does this improve?**

The `Verify Paperclip Runner` job in `.github/workflows/pr-trusted.yml`
recompiles the full Rust dependency tree on every pull request.

**Subsystem affected**

Cross-cutting (multiple of the above). The change touches CI workflow
configuration only. It does not change product code.

**Current behavior**

The job has no Rust cache. Its only cache step restores the pnpm store.
Each run therefore downloads about 280 crates and compiles all 313
third-party crates twice, once in the `dev` profile and once in the
`release` profile.

Measured over 12 successful runs on 2026-09-15, the job takes 15.0
minutes on average. The range is 12.1 to 17.4 minutes.

| Phase | Mean | Share |
|---|---|---|
| `check:eval-kernel` | 0.0m | — |
| `typecheck:typescript` and protocol manifest | 0.1m | 1% |
| `build:rust` — cargo `dev` profile | 1.8m | 12% |
| TypeScript tests (`node --test` and vitest, 143 files) | 5.8m | 39% |
| `check:replay-goldens` | 0.1m | 1% |
| `typecheck:rust` — `cargo fmt` and `cargo check` | 0.9m | 6% |
| `test:rust` compile — cargo `release` profile | 4.9m | 33% |
| `test:rust` run | 0.8m | 5% |
| Parity checks | 0.0m | — |
| `check:api-authority` | 0.5m | 3% |

Cargo reports the two compiles directly. The `dev` profile takes 1m24s
to 1m53s. The `release` profile takes 3m54s to 5m36s.

**Proposed behavior**

The job restores master's existing Rust dependency cache before it runs
the checks.

Master already writes this cache. The post-merge `rust` lane in
`.github/workflows/release-verify.yml` writes `release-runner-v1`. It
also warms both profiles into that entry. The entry is 677MB and lives
on `refs/heads/master`. Pull request branches are allowed to read caches
from the default branch.

The pull request lane restores that entry read-only. Master stays the
only writer. A pull request never saves a branch-scoped copy. A pull
request never evicts the shared entry. This matches the rule the pnpm
store in the same file already follows.

**Reason and benefit**

Master runs these same checks with a warm cache. That gives a direct
measurement of the saving.

| Check set | Cold (pull request today) | Warm (master) | Change |
|---|---|---|---|
| `check:runner` and `check:api-authority` | 7.1m | 3.4m | −3.7m |
| `check:eval-kernel` and `check:protocol` | 7.8m | 7.3m | −0.5m |
| Cache restore step | — | 20–21s | +0.35m |

The net saving is about 3.9 minutes per run, or about 26%. About 73 runs
execute this job each day. The daily saving is therefore about 4.7
compute-hours.

The protocol side changes very little. Vitest dominates that side, not
compilation.

The cache should almost always hit.
`packages/paperclip-runner/runner/Cargo.lock` changed in 7 of the last
1697 commits on master. A miss costs nothing more than today's behavior.

**Breaking changes**

None. The change adds two steps to one CI job. It does not change any
check, any test, or any product code.

**Additional context**

No documentation covers pull request CI caching, so this pull request
updates no documents.

## What Changed

- Added a `Select the pinned Runner Rust toolchain` step to the
`verify_paperclip_runner` job in `.github/workflows/pr-trusted.yml`. The
step exports `RUSTUP_TOOLCHAIN`. `rust-cache` hashes `rustc -vV` into
the cache key, so the key needs the pinned compiler. This lane had no
rustup step before, so `rustc` resolved to each runner image's default
instead of the pinned 1.97.1.
- Made that step tolerate a missing `rustup`. `release-verify.yml` runs
on one post-merge fleet image. The gate in this workflow routes to
either `ubuntu-latest` or the public pull request fleet. A missing
`rustup` now costs the cache. It does not fail the pull request.
- Added a `Restore Runner Rust dependencies (read only)` step that uses
`Swatinem/rust-cache` with `save-if: false`.
- Mirrored every cache key input from the master writer in
`release-verify.yml`: the same action SHA, `workspaces`, `shared-key`,
`cache-workspace-crates`, and `cache-bin`. Neither side sets
`prefix-key`. Any drift causes a silent miss and a full recompile.
- Added `.github/scripts/tests/pr-runner-rust-cache.test.mjs`. It
asserts key-input parity across the two workflow files, the step order,
and the restore-only contract.

## Verification

Run the workflow shape tests:

```bash
node --test '.github/scripts/tests/*.test.mjs'
```

Result: 413 pass, 0 fail.

I also mutation-tested the new assertions. I applied each mutation to
the workflow, ran the new test file, then restored the file. All 7
mutations fail the suite:

| Mutation | Result |
|---|---|
| `shared-key` changed to `pr-runner-v1` | Caught |
| `save-if` changed to `true` | Caught |
| `cache-bin` changed to `true` | Caught |
| `rustup show active-toolchain` line deleted | Caught |
| `RUSTUP_TOOLCHAIN` export line deleted | Caught |
| `working-directory` line deleted | Caught |
| `command -v rustup` guard deleted | Caught |

To confirm the cache key inputs match the master writer, parse both
workflows and compare:

```bash
ruby -ryaml -e 'pr=YAML.safe_load(File.read(".github/workflows/pr-trusted.yml"), aliases: true); rv=YAML.safe_load(File.read(".github/workflows/release-verify.yml"), aliases: true); a=pr["jobs"]["verify_paperclip_runner"]["steps"].find{|s| s["uses"].to_s.include?("rust-cache")}; b=rv["jobs"]["verify_paperclip_runner"]["steps"].find{|s| s["uses"].to_s.include?("rust-cache")}; puts a["uses"]==b["uses"]; %w[workspaces shared-key cache-workspace-crates cache-bin].each{|k| puts "#{k}: #{a["with"][k]==b["with"][k]}"}'
```

Every line prints `true`.

After the pin advances (see Risks), confirm the cache works in CI. The
restore step log must show `Cache restored from key:
v0-rust-release-runner-v1-Linux-x64-...`. Cargo must stop printing
`Compiling` lines for third-party crates. The job should drop from about
15.0 minutes to about 11.0 minutes.

## Risks

Low risk overall. A cache miss produces exactly today's behavior, so the
worst case is no improvement.

- **This change does nothing until a second pull request lands.**
`.github/workflows/pr.yml` pins this reusable workflow by SHA. The file
header describes this two-step rollout. A follow-up pull request must
bump that pin. That follow-up validates itself, because GitHub uses the
pull request's own `pr.yml` for `pull_request` events.
- **A key mismatch would silently remove the benefit.** The runner
images here may ship a different default `rustc` than the post-merge
fleet. The toolchain step pins the compiler to prevent this. The new
test guards the remaining key inputs. If the first runs still miss,
compare `rustup show` output against a master run.
- **A missing `rustup` degrades quietly.** The step prints a GitHub
notice and continues. CI stays green and the run is simply uncached.
- **No cache poisoning path.** Pull requests only read. `save-if: false`
stops any write. GitHub also isolates pull request cache writes from the
default branch.
- **The cache entry can expire.** GitHub evicts unused entries after 7
days and enforces a repository size limit. Master pushes are frequent,
so the entry should stay warm. Eviction only causes a miss.

## Model Used

Claude Opus 5, provider Anthropic, exact model ID `claude-opus-5`, 1M
context window. Adaptive thinking was on. I used tool use throughout:
the `gh` CLI to read 12 job logs and step timings from recent CI runs,
the GitHub Actions cache API to read cache entry keys and sizes, `git
log` to measure `Cargo.lock` churn, and local `node --test` runs to
verify and mutation-test the change. Run through Claude Code.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [ ] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Note on the unchecked boxes. The branch name
`claude/paperclip-runner-conditional-81a300` carries a tool-generated
suffix, so it does not meet the branch naming rule. I can rename it if
you want. The CI and Greptile boxes stay unchecked until those checks
finish on this pull request.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-09-15 01:16:34 -07:00
Devin FoleyandPaperclip 4cc387f907 fix(ci): remove npm propagation from cloud readiness (#13456)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image and matching database
migrator before it can deploy a merge.
> - New npm package versions can take minutes to become downloadable
after the package build finishes.
> - The direct producer now publishes signed archives and a complete
dependency lockfile for each master commit.
> - This pull request makes readiness verify those artifacts and removes
the duplicate automatic npm migrator run.
> - Deployment still requires all source checks, exact image identity,
migration compatibility, and pinned dependencies.

## Linked Issues or Issue Description

Refs: #13455, #13454, #13192

**What existing behavior does this improve?**
The time from a master merge to the `Cloud deployable v1` signal.

**Current behavior**
Readiness polls npm metadata for the new DB and shared versions. An
automatic dispatcher also starts a separate npm-only migrator workflow.
A measured source built its packages at 06:22:41 UTC on 2026-09-15, but
both npm archives were not downloadable until 06:31:56 UTC.

**Proposed behavior**
Wait for the successful exact-source direct producer, verify its signed
manifest and all pinned downloads, and publish readiness only after the
existing source and image jobs pass. Keep manual npm migrators and
branch previews available.

**Reason and benefit**
Remove new-version npm propagation from merge-to-deployable time. The
gain depends on whether image building or source verification finishes
later; it is not a fixed subtraction from every run.

## What Changed

- Require a successful producer from the canonical repository, exact
commit, master ref, expected workflow, and approved event.
- Verify the manifest's GitHub attestation with the hosted GitHub CLI.
Enforce the exact source SHA, master workflow identity, and hosted
runner.
- Download and validate both archives and the complete dependency
lockfile after publication succeeds. Reject invalid signatures,
inaccessible objects, corrupt bytes, and source mismatches.
- Remove automatic npm-only migrator dispatch. Retain manual release and
branch-preview publication.
- Document the cloud feature-switch prerequisite and coordinated
rollback.

## Verification

- `node --test .github/scripts/tests/*.test.mjs`: 405 pass.
- Focused readiness, routing, preview, and artifact tests: 249 pass.
- Workflow lint and `git diff --check`: pass.
- `pnpm test:release-registry`: 139 pass after installing this
worktree's dependencies.
- All latest-head GitHub CI checks passed. Greptile is 5/5 with no
unresolved comments.
- Application source is unchanged. Common-source local typecheck and
build passed. The full local application suite has the documented macOS
read-only-directory rename limitation from #13454 (13 failures in two
unchanged suites); Linux CI is the final application gate.
- Live readiness verification of master
da77a0c28c passed in 6.52 seconds,
including the real GitHub CLI signature policy and all artifact
downloads.
- Cloud consumer resolution with the certificate encoding fix passed in
5.45 seconds with zero npm metadata requests or npm processes. The
consumer is deployed and enabled in staging and production; their live
resolution APIs passed in 2.34 and 2.38 seconds. Both report the
expected fixed harness commit. A fresh tenant deployment follows this
cutover merge.

## Risks

- `Cloud deployable v1` no longer promises npm preview availability.
Enable the cloud direct-artifact consumer in staging and production
before merging this change.
- Artifact storage and GitHub attestations become required services for
new direct releases. Missing or invalid evidence fails explicitly.
- Restore the old npm dispatcher and readiness gate together before
disabling the consumer switch. Retain artifacts referenced by existing
releases.
- This change does not expand AWS runner access. The producer and
readiness bookkeeping use GitHub-hosted runners. Existing PR allowlists
and source verification gates remain enforced.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused checks; full
application host limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 01:14:43 -07:00
Devin FoleyandPaperclip da77a0c28c feat(ci): publish immutable cloud migrator artifacts (#13455)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images and a matching database migrator.
> - New migrator versions must currently become available on npm before
cloud can use them.
> - npm can serve package metadata while the named archive still returns
404.
> - This pull request publishes immutable migrator archives and a
complete dependency lockfile through the existing artifact store.
> - Cloud can install these exact packages without waiting for their new
npm versions.
> - This producer change prepares a separate cloud consumer and
readiness cutover.

## Linked Issues or Issue Description

Refs: #13454

**What happened?**
A recent master run built both packages by 06:22:41 UTC on 2026-09-15.
Both archives became downloadable from npm at 06:31:56 UTC. Fresh
metadata requests did not remove the delay.

**What did you expect to happen?**
Cloud should be able to install the verified migrator as soon as its
package build and artifact upload finish.

**Steps to reproduce**
Compare package build completion, npm publication, version metadata
availability, and tarball download availability for a fresh full commit
SHA.

**Version**
Master commit `08adcc70d5ec45b7ced9619a3dc10c1d1bec397d`.

## What Changed

- Add a master-only workflow that builds the DB and shared archives
without publication credentials.
- Resolve the dependency lockfile from local archives, then pin those
archives to content-addressed URLs.
- Publish the complete bundle to a separate prefix in the existing
S3/CloudFront artifact store. Write the commit manifest last and verify
public downloads.
- Add a dedicated OIDC role policy. Only canonical master can assume it.
Writes require `If-None-Match: *`; the role cannot overwrite or delete
objects.
- Add source, integrity, lockfile, publication, and real npm install
tests. Document the format and staged rollout.
- Attest the validated manifest with GitHub/Sigstore before S3
publication. The signature binds every package and lockfile hash to the
exact master workflow and source commit.

## Verification

- `node --test scripts/cloud-migrator-artifacts.test.mjs`: 7 tests pass,
including real `npm ci` with an empty cache and no new-version metadata
lookup.
- `pnpm test:release-registry`: 136 tests pass.
- `actionlint .github/workflows/cloud-migrator-artifacts.yml` and `git
diff --check`: pass.
- Ran the workflow's filtered install and package build against the
exact master source. Built and validated the dependency lockfile from
those real archives.
- Latest-head application tests passed, including reruns of two failures
in unchanged application tests. The final CI aggregate passed. The
application source is unchanged. Common-source local typecheck and build
passed; the full local suite has the same documented macOS
read-only-directory rename limitation as #13454 (13 failures in two
unchanged suites).
- The dedicated role and additive bucket read permission are configured.
IAM simulation allows only conditional writes in the intended prefix;
overwrite without the condition, other prefixes, and deletion are
denied.
- Published the verified master 08adcc70d5
bundle with the operator session and verified all public downloads.
GitHub OIDC publication is still pending the master workflow run.
- Cloud resolved the real bundle and checked all 278 SQL migrations in
2.9 seconds with zero npm metadata requests or npm processes. The
existing migration runner applied it to a disposable local PostgreSQL
database and succeeded again on repeat.
- The producer now requires an empty-cache smoke install of the actual
package archives and their full dependency graph before upload. That
check and imports of both installed packages passed locally.

## Risks

- This is an additive producer rollout. It does not yet change the cloud
resolver or the deployable marker.
- The dedicated role and bucket read statement must be installed before
the workflow can publish. Existing bucket policy statements and
public-access blocks must be preserved.
- Referenced artifacts must be retained for rollback. No expiry rule
applies to this prefix.
- Existing external dependencies still download from npm, with SHA-512
pins. New DB and shared versions do not require npm metadata.
- The workflow uses GitHub-hosted runners and has no PR trigger. It adds
no AWS compute routing or PR access.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused release and
real-artifact tests; full-suite host limitation documented above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-15 00:46:05 -07:00
Devin FoleyandPaperclip 52d120f68d fix(release): wait 30 minutes for npm to expose a published version (#13436)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Every release publishes a batch of npm packages and then waits for
each one to become visible before continuing
> - npm accepts a publish immediately but exposes it later, so that wait
exists to keep a release from continuing past a package nobody can
install yet
> - Today that wait was too short twice in a row, and each timeout
aborted the release with the batch half-published
> - Version numbers are derived from what is already on npm, so every
retry moves to a new number and meets the same lag
> - Two canary attempts burned two versions this way and shipped nothing
> - This pull request raises the per-package budget from ten minutes to
thirty
> - The benefit is that ordinary registry lag costs waiting instead of a
failed, half-published release

## Linked Issues or Issue Description

No existing issue. The problem, in the bug report format:

**What happened**
`publish_canary` failed twice in a row with the batch half-published:

```
Warning: npm accepted @paperclipai/server@2026.914.0-canary.2, but the version did not become registry-visible.
Error: stopping release: npm did not publish and expose @paperclipai/server@2026.914.0-canary.2
```

The version was accepted at 17:49:09 and became visible at 18:04:29 —
about five minutes after the poll gave up. `shared`, `db` and
`adapter-utils` published at that version; `server`, `paperclip-runner`
and the root package did not.

**Expected behavior**
Ordinary registry propagation delay costs the release some waiting, not
a failure. A package that becomes visible after 15 minutes must not
abort the batch because of the old 10-minute wait. Longer registry
outages can still leave a partial batch.

**Steps to reproduce**
1. Publish any channel while npm is propagating slowly.
2. A package takes longer than `NPM_PUBLISH_VERIFY_ATTEMPTS *
NPM_PUBLISH_VERIFY_DELAY_SECONDS` to become visible.
3. The release aborts, that version is half-published, and the retry
picks a new version number and meets the same lag.

**Paperclip version or commit**
Present on master. Observed on 2026-09-14 across canary runs in workflow
run 34869494325.

## What Changed

- Increase npm visibility checks from 60 to 180, retaining the 10-second
delay: about 30 minutes per package.
- Increase canary, nightly, beta, and stable publish job timeouts from
90 to 150 minutes.
- Add offline regression tests using the workflow's actual settings.
They cover the observed 15-minute 20-second delay, immediate visibility,
exhausted retries, and job timeout sizing.
- Load the access router in test setup so its cold transform does not
consume the first permission test's 10-second timeout. The permission
assertions are unchanged.

## Verification

- `node --test scripts/release-lib.test.mjs`: 14 passed.
- `pnpm run test:release-registry`: 129 passed.
- `pnpm exec vitest run
server/src/__tests__/access-routes-permissions-upgrade.test.ts`: 3
passed.
- Regression proof in temporary fixtures: restoring 60 attempts fails
the observed-delay test; restoring 90-minute jobs fails the
timeout-budget test.
- [CI run
34916632804](https://github.com/paperclipai/paperclip/actions/runs/34916632804):
all jobs passed, including typecheck/release registry, build, all
general and serialized server shards, browser tests, runner
verification, and canary dry run. The PR has 31 successful checks and
two expected Storybook skips at `a27f5e896`.
- [Previously failing serialized
shard](https://github.com/paperclipai/paperclip/actions/runs/34916632804/job/104215809616):
all three access-route permission tests passed in CI after preloading
the router.
- Greptile's final review is 5/5 with no outstanding findings. Both
review threads are resolved.
- Local limits: `pnpm -r typecheck` and `pnpm build` stop at the runner
package because this machine has no Rust `cargo` executable. The
duplicate full local `pnpm test:run` was interrupted while the complete
CI matrix ran. Targeted local results are listed above; full validation
is from CI.

The previous CI failures were unrelated to npm propagation:

- [PR
review](https://github.com/paperclipai/paperclip/actions/runs/34891770396)
required a test file for this fix.
- [Serialized server shard
3](https://github.com/paperclipai/paperclip/actions/runs/34891773736/job/104136392196)
timed out in the first access-route permission test at 10 seconds. The
other two tests in that file passed.
- The canary dry run passed in that same CI run.

## Risks

Low risk. Production behavior changes only in release waiting budgets.

- Polling exits as soon as npm exposes the version, so healthy publishes
do not wait longer.
- An unavailable version now takes about 30 minutes to report. Polls
remain bounded and still fail the release on exhaustion.
- The 150-minute jobs leave roughly 30 minutes for setup/build plus four
full polling windows. npm command runtime and later release steps also
consume that budget; a broader outage can still interrupt a batch.
- The permission-test change moves module loading into a bounded setup
hook; it does not relax authorization assertions.
- No schema changes or operational migrations.

## Model Used

- Claude Fable 5 (`claude-fable-5`), 1M context, extended thinking, run
through Claude Code with tool use and code execution.

- OpenAI GPT-6 (Codex), with reasoning, tool use, and code execution,
for the CI follow-up. Context-window size is not exposed in this
session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — targeted checks listed
above; full validation passed in CI
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes —
workflow comments and verification details
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-14 20:16:29 -07:00
Devin FoleyandPaperclip 44f6312cd8 fix(ci): reuse one available Cloud registry cache (#13334)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud needs a verified image for each merged source
commit.
> - Fresh builders restore compiled native dependencies from registry
caches.
> - The current workflow imports up to eleven historical cache manifests
at once.
> - Live builds missed native layers that a fresh builder reused from
one manifest.
> - This PR selects the nearest available cache and tests reuse across
fresh builders.

## Linked Issues or Issue Description

Refs #13329 and #13330. A search of open cache PRs found no duplicate of
this change.

**What existing behavior does this improve?**

Remote Docker cache reuse on fresh Cloud image builders.

**Current behavior**

[Cloud run
34714483272](https://github.com/paperclipai/paperclip/actions/runs/34714483272/job/103609096836)
imported the previous cache manifest successfully but rebuilt
`cargo-chef` and Rust dependencies. The dependency compile took 3m43s.
The preceding image build had already exported those layers.

A controlled [fresh-builder
diagnostic](https://github.com/paperclipai/paperclip/actions/runs/34715336530)
used the same source and registry cache. The single-manifest job reused
both layers immediately. The multiple-manifest job rebuilt them and
failed the cache assertion. Both jobs used GitHub-hosted runners with
read-only access.

**Proposed behavior**

Inspect cache manifests in first-parent order and import only the
nearest available one. Keep full-SHA cache exports, the ten-commit
search bound, and the legacy fallback. If caches cannot be read, permit
a cold build.

**Reason and benefit**

Avoid the observed cache misses without changing image contents or
builder sizes. Expected savings include about four minutes of native
tool/dependency compilation when those inputs are unchanged. The final
merge-to-deployable gain still needs a post-merge measurement.

**Breaking changes**

No image, artifact, deployment, or runner-routing contract changes.

## What Changed

- Select one available ancestor cache after Docker login and Buildx
setup.
- Preserve separate writable cache tags for each full source SHA.
- Test cache ordering, missing caches, registry errors, and workflow
integration.
- Add the selector tests to the existing release-registry suite.
- Export a local test cache, remove the first builder, and verify a
source rebuild on a fresh builder.
- Document cache selection and the stronger Docker check.

## Verification

- Passed 456 focused workflow, routing, readiness, preview-artifact, and
cache-selector tests.
- Passed shell syntax, ShellCheck for the changed probe, actionlint
workflow validation, and `git diff --check`. actionlint's shell checks
were disabled for the workflow validation because unchanged
migration-label commands trigger existing SC2012 notes.
- The fresh-builder registry diagnostic proves the single-cache
behavior. The [permanent two-builder probe
passed](https://github.com/paperclipai/paperclip/actions/runs/34715771048/job/103612624090),
including a changed real binary and dependency-declaration invalidation.
- Passed all 35 latest-head checks (green or intentionally skipped),
including full typecheck, test, build, and browser suites in [PR CI run
34715771217](https://github.com/paperclipai/paperclip/actions/runs/34715771217).
- The real selector CLI inspected registry metadata and chose the
nearest available ancestor cache.
- Fresh Greptile review is 5/5 with no open findings. The PR title was
corrected to meet the source-change naming rule; the review check passed
after that correction.
- Local full-suite runs and Docker builds are unavailable because the
local Docker daemon is unresponsive after disk exhaustion. CI provides
the Linux verification.

## Risks

- Missing or unreadable caches cause a slower cold build. The selector
logs that condition and preserves image publication.
- Inspecting several missing ancestors adds lookup time. Each lookup has
a ten-second timeout and the search is bounded.
- The Docker test now exports a local cache. It removes the first
builder before starting the second to release disk space, then cleans up
its builders and files.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 13:20:18 -07:00
Devin FoleyandPaperclip 0ce7df2648 ci: keep Cloud readiness markers out of the builder queue (#13330)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud consumes versioned readiness markers for each merged
source commit.
> - Those markers can be published only after the required checks and
artifacts pass.
> - The marker jobs currently wait for the same AWS runner capacity as
builds and tests.
> - A busy builder pool can delay readiness after all required work has
finished.
> - This PR moves the small readiness jobs to GitHub-hosted runners
while retaining every dependency gate.

## Linked Issues or Issue Description

Refs #13326 and #13328.

**What existing behavior does this improve?**

Time from completed Cloud verification to a deployable marker, and AWS
capacity occupied by artifact polling.

**Current behavior**

In [Cloud readiness run
34711557083](https://github.com/paperclipai/paperclip/actions/runs/34711557083),
all builds and tests finished at 18:43:05 UTC. The source marker did not
start until 18:44:21, and the deployable marker did not start until
18:44:37. Merge-to-deployable was 11m33s, although the prerequisite work
finished in 9m44s.

**Proposed behavior**

Run the artifact wait and both versioned marker jobs on `ubuntu-latest`.
Keep the compute jobs on the approved post-merge AWS fleet.

**Reason and benefit**

Avoid builder-pool queue delays after verification finishes. This also
removes the long artifact-wait job from AWS capacity. Expected savings
depend on queue depth: the observed run had over 90 seconds of avoidable
marker waiting. The marker commands themselves take only seconds.

**Breaking changes**

Runner placement changes for three bookkeeping jobs. Marker names,
exact-source artifact checks, required verification, and image
verification stay the same.

**Additional context**

Searched the related runner and Cloud readiness work. This addresses
queue time observed after the parallel verification change.

## What Changed

- Place the artifact wait, source-verification marker, and deployable
marker on GitHub-hosted runners.
- Keep all existing job dependencies, source guards, permissions, and
commands.
- Extend routing regressions to enforce this placement and retain
fail-closed readiness gates.
- Document why readiness bookkeeping uses separate runner capacity.

## Verification

- Passed 433 workflow, routing, source-verification, and Cloud readiness
tests with `node --test .github/scripts/tests/*.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/cloud-readiness.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- Passed `actionlint` and `git diff --check`.
- Passed all latest-head CI gates in [run 34712624340, attempt
2](https://github.com/paperclipai/paperclip/actions/runs/34712624340),
including typecheck, build, browser, Runner, and all general/serialized
tests.
- Attempt 1 had one localhost readiness timeout in an unchanged test. A
single targeted retry passed all 166 files (3,225 tests passed, one
existing skip), including all seven tests in that file. No timeout,
assertion, or application source was changed; the retry is documented in
the PR comment.
- Local full-suite verification is limited by local disk exhaustion; the
focused checks above pass.
- Fresh Greptile review is 5/5 with no unresolved findings.

## Risks

- GitHub-hosted capacity can also queue, but these jobs no longer
compete with AWS build/test demand. The change does not reserve
instances or change box sizes.
- Readiness must still fail if any prerequisite fails. The existing
`needs` relationships and success-only execution are preserved and
tested.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 12:32:46 -07:00
Devin FoleyandPaperclip 7435b2ee9c ci: cache compiled Docker Rust dependencies separately from source (#13329)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images that contain the native Rust Runner.
> - The image already builds that Runner before copying ordinary app
source.
> - A Rust source change still invalidates its entire compiled
dependency layer.
> - Compiled dependencies can survive source changes when their recipe
is unchanged.
> - This PR adds a separate locked dependency build before compiling the
real workspace.

## Linked Issues or Issue Description

Refs #13195. A search of related Docker and Cargo cache PRs found no
duplicate dependency-recipe change.

**What existing behavior does this improve?**

Docker image build time after Rust source or embedded protocol changes.

**Current behavior**

The `runner-build` stage compiles dependencies and workspace code in one
layer. In Cloud readiness run 34698143548, that stage took about 3m48s
when its cache was unavailable.

**Proposed behavior**

Generate a recipe with pinned cargo-chef 0.1.73. Build locked release
dependencies in `runner-deps`, then copy and compile real Rust source
and embedded protocol inputs in `runner-build`. Source edits can reuse
the dependency layer from the existing registry cache.

**Reason and benefit**

Reduce dependency recompilation during source changes and merge bursts.
Expected savings are roughly 2–4 minutes when the old native layer would
miss but dependency layers are available. Full cold builds also pay for
the recipe tool installation. Ordinary app-only cache hits gain little
from this change.

**Breaking changes**

None to the shipped application or image tags. The recipe tool and
compiled dependencies remain in build stages.

## What Changed

- Install a pinned recipe generator with its locked dependencies and the
existing package-owned compiler.
- Add recipe planning and compiled dependency stages. Use the same
release profile, package, binary, and lockfile enforcement as the real
native build.
- Remove generated source stubs before copying actual source. Preserve
protocol inputs, timestamp normalization, binary staging, and
application checks.
- Add Docker cache wiring regressions and update the Docker cache
documentation.
- Run a two-build probe in Docker Runner check. It requires dependency
reuse, changed real binary metadata after a source edit, and a changed
recipe after a dependency declaration edit. It uses a disposable
tracked-source context and exports only small metadata files.

## Verification

- Passed all five Docker build-stamp and dependency-cache tests with
`pnpm exec vitest run server/src/__tests__/docker-build-stamp.test.ts`.
- Passed the local ARM64 `docker buildx build --target runner-build
--progress plain`. Local Docker then hit storage errors during a runtime
probe; cache invalidation verification continues on GitHub-hosted Linux.
- Passed `bash -n scripts/check-docker-runner-cache.sh`, `actionlint`,
and `git diff --check`.
- Passed a [Linux AMD64 cache
probe](https://github.com/paperclipai/paperclip/actions/runs/34711042199)
against the PR source: dependencies compiled in 3m49s for the baseline
and were `CACHED` after a source edit; real source compilation took
about 37 seconds. Binary metadata changed and dependency declaration
changes altered the recipe. The permanent probe is also running in
latest-head Docker Runner check.
- Passed latest-head [Docker Runner
check](https://github.com/paperclipai/paperclip/actions/runs/34711145160),
including the permanent source/dependency invalidation probe.
- Passed full [PR
verification](https://github.com/paperclipai/paperclip/actions/runs/34711145352/attempts/2):
typecheck, all grouped tests, native verification, build, release dry
run, and browser checks. One unrelated signoff-policy browser test
failed waiting for a heartbeat run on attempt 1; only that failed shard
and dependent checks were retried, and passed.
- Latest-head Greptile is 5/5 with no unresolved findings. Full local
tests/build were limited by local disk exhaustion; Linux CI completed
those checks.

## Risks

- The two-build CI probe has a 20-minute job limit to cover the cold
build and source rebuild. It adds no AWS routing.
- A fully cold build must install cargo-chef and populate the dependency
layer. Both become reusable registry layers; no Actions cache is added.
- The recipe and final build must keep the same compiler, build profile,
package, binary, and directory layout. A source-change rebuild probe
checks real cache reuse and binary invalidation.
- Dependency or compiler changes still require rebuilding dependencies.
Existing image verification and full-SHA publication gates remain
unchanged.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:57:32 -07:00
Devin FoleyandPaperclip d2e940f4c1 ci: run release Runner protocol and Rust checks in parallel (#13326)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys verified images from merged source commits.
> - Cloud readiness waits for every release verification check.
> - Runner verification currently runs long TypeScript tests before Rust
checks.
> - These checks can run on independent runners with their own build
directories.
> - This PR runs them in parallel while preserving all checks and the
shared dependency cache.

## Linked Issues or Issue Description

Refs #13194. Related prior work: #13142 and #13259. A search found no
duplicate parallel release-check change.

**What existing behavior does this improve?**

Time from merge to Cloud source verification and deployment readiness.

**Current behavior**

Recent successful runs take roughly 13 minutes from merge to deployable.
In run 34705914878, Runner verification took 11m23s. Protocol tests
finished before Rust tests and API authority checks started.

**Proposed behavior**

Run protocol and Rust verification in two matrix jobs. Cloud readiness
still requires both jobs to pass.

**Reason and benefit**

Remove the serial dependency between independent checks. Expected
improvement is about 2–3 minutes on a typical cached run, until the
image build or server tests become the longest job. This is an estimate;
post-merge timing will confirm it.

**Breaking changes**

Individual release Runner job names gain a lane suffix. Cloud source and
readiness marker names stay the same. PR runner routing is unchanged.

## What Changed

- Split release Runner checks into protocol and Rust lanes. Keep every
constituent of `check:all` exactly once.
- Restore the existing Rust dependency cache in both lanes. Allow only
the Rust lane to save it after warming both build profiles.
- Add coverage and cache authorization regressions. Document the
parallel verification and single cache writer.

## Verification

- Passed 477 workflow and source-verification tests with `node --test
.github/scripts/tests/*.test.mjs
scripts/cloud-source-verification.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs`.
- Passed `actionlint`, `git diff --check`, and the private AWS routing
regression suite.
- Passed local `pnpm -r typecheck` and the standalone `check:runner &&
check:api-authority` lane, including all 1,671 API tests before the
protocol lane had built TypeScript output.
- The broad local protocol run under Node 25 had four failures. The two
affected files passed under CI's Node 24.19.0: 67 passed, 6 platform
skips.
- Local `pnpm test:run` aborted when disk space ran out; local `pnpm
build` could not run afterward. These are local verification limits.
[Linux CI run
34710421424](https://github.com/paperclipai/paperclip/actions/runs/34710421424)
passed full typecheck, all grouped tests, native verification, build,
release dry run, and browser checks. Native protocol CI passed 1,986
tests, plus 1,671 API tests and the Rust suites.
- Latest-head Greptile is 5/5 with no open findings. All 33 current-head
checks are successful or intentionally skipped.

## Risks

- Uses one additional short-lived verification runner per release
verification. The existing AWS exact-master restriction remains in
place.
- The Rust lane warms debug dependencies so its cache save also serves
protocol tests. Both lanes always rebuild workspace code.
- A workflow regression could omit a check. The new coverage test
compares the matrix checks directly with `check:all`; Cloud readiness
depends on the complete reusable workflow.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 11:33:20 -07:00
DottaandPaperclip ab15aff390 feat: add experimental persistent agent chat (#13284)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Conversations must use the same tasks, controls, and execution
history.
> - Users need an ongoing chat with an agent without managing task
properties.
> - Agents should clarify and plan work, then hand execution to assigned
project tasks.
> - This pull request combines the reviewed Agent Chat stack for one
squash merge.
> - The benefit is persistent conversation with normal task governance
and shared UI.

## Linked Issues or Issue Description

**Subsystem affected**

Task lifecycle, agent runtime tools, shared task UI, and browser/paid
runner tests.

**Problem or motivation**

Users need one persistent conversation with each agent. A separate chat
store or renderer would duplicate task behavior and bypass existing
controls.

**Proposed solution**

Use a task-backed chat per company, user, and agent. Reuse the task
composer and transcript. Clarify and plan in chat, then create assigned
project tasks with the relevant plan. Keep Agent Chat behind its own
disabled-by-default experimental setting.

**Roadmap alignment**

This implements the task-backed direction in [CEO
Chat](https://github.com/paperclipai/paperclip/blob/master/ROADMAP.md#-ceo-chat).
Related proposals: #2504 and #9693. Related request: #7981. The
maintainer requested one squash merge of the complete stack.

Consolidates the reviewed runtime
[#13281](https://github.com/paperclipai/paperclip/pull/13281), backend
[#13282](https://github.com/paperclipai/paperclip/pull/13282), and UI
[#13283](https://github.com/paperclipai/paperclip/pull/13283) layers
with this PR's E2E coverage. All four layers passed CI and received
Greptile 5/5 before consolidation. This PR targets master and includes
the complete feature.

## What Changed

- Add personal canonical chat tasks with ordinary company visibility,
immutable identity, idempotent first sends, and an idle waiting state.
- Process `/new` in queue order. Preserve history, release a chat pause,
and fence old provider context and delayed writes.
- Keep chat lifecycle rules across recovery, finalization, assignment,
task lists, and rollups.
- Support research and plan revision in chat. Hand plans to ordinary
assigned project tasks before execution starts. Reject new chat
subtasks.
- Add repository-aware project creation and discovery tools, including
multiple repository IDs and GitHub URLs, authorization, idempotency, and
durable project-created cards.
- Reuse task UI components for chat, with starred/recent agent
navigation and a separate `enableAgentChat` experimental flag.
- Add deterministic browser tests and 24 paid chat cells across four
Codex/Claude profiles, with validated reports and screenshots.
- Integrate current master recovery, controller lease, queued-message,
and task UI changes. Gate chat interruption and deferred promotion on
ownership/feature policy. Guarantee lease renewal and active controls
are stopped even if teardown fails.
- Preserve master's migration 0273 and generate chat migration 0274 with
idempotent replay for development databases.

## Verification

- Prior exact heads of all four PRs passed Linux CI, including build,
typecheck, general/serialized tests, and browser E2E. Each had Greptile
5/5 and no unresolved findings.
- Integrated local verification passed: full repository typecheck and
production build, Storybook build, token gates, 340 focused UI tests,
all 20 deterministic chat browser tests, two migration replay tests, 88
focused chat/queue/native/controller tests, and provider/session
regressions including real lease expiry. These include the three
lifecycle regressions for the final admission/teardown fixes; server
typecheck also passes. Current head
`1268eda16cc2af892055917e7292f068820be135` has Greptile 5/5 with no
unresolved findings and passing security scans. All final-head CI gates
passed: build, full Runner verification, typecheck/release registry,
canary, all general/serialized test shards, and all browser E2E shards
([CI
run](https://github.com/paperclipai/paperclip/actions/runs/34696739927)).
Local PostgreSQL startup contention required serialized retries; skipped
fixtures do not count as passing coverage.
- The earlier paid campaign passed all 24 chat cells and retained 32
screenshots:
[report](https://d1p6rlowie26tp.cloudfront.net/runner-e2e/campaigns/gha-34648511170-1/index.html?report=agent-chat#suite-agent-chat).
It tested `abacbdfd2f660709ec37312cdb758284c8399d04`; it is prior
evidence, not a paid run of this integrated head.
- Manual check: enable Agent Chat in Experimental settings, open an
agent, clarify and revise a plan, then hand off to an assigned project
task. Stop a reply, send `/new`, and verify fresh context with retained
history. Disable the setting and verify agent shortcuts/new chat turns
are blocked.

## Risks

- Queue/session integration can affect retries and delayed writes. Tests
cover ownership, cancellation, reset boundaries, idle recovery, and
ordinary task behavior.
- Migration 0274 adds conversation fields and constraints. Replay is
idempotent and preserves existing development chat history.
- This combines the previously reviewed stack at the maintainer's
request. Agent Chat remains off by default and is separate from
Conference Room.

## Model Used

OpenAI Codex, GPT-6 Astra (`gpt-6-astra`), with reasoning, code
execution, browser tools, and parallel review. The exact context-window
size is not exposed in this session. Codex and Claude also ran as test
subjects in the linked paid campaign.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-12 08:56:04 -05:00
Devin FoleyandPaperclip 09e208c54f ci: activate shared PR dependency cache restores (#13302)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - PR checks share a repository cache budget with post-merge cloud
verification.
> - Per-PR dependency store copies consume about 700 MB each and
displace useful build caches.
> - The reviewed workflow now restores shared dependency caches without
saving PR copies.
> - The public entry point must pin an approved immutable workflow
before it can use AWS runners.
> - This PR activates the merged workflow after the runner group
accepted its exact SHA.
> - Native verification and the app build also run in parallel, as
defined in the merged workflow.

## Linked Issues or Issue Description

Refs #13300. Refs #13301.

The cache definition is merged, and its exact SHA is authorized in the
restricted runner group. This PR activates it. Its base also contains
the two synthetic fixture fixes from #13301. No duplicate activation PR
was found.

## What Changed

- Advance the `pr-trusted.yml` pin from
`03609aa6ecc9a047ed53d6b6469d8be554fbc46d` to merged commit
`44dde2dec42a22746a2f36b595acacc9ccfa1df6`.
- Update the adjacent pin comment to identify #13300 and the activated
behavior.
- Activate restore-only pnpm caching and removal of redundant Node setup
steps from #13300.
- Activate the already-merged separate native verification job, required
by the aggregate `verify` check. The app build no longer waits for
native verification inside the same job.
- Activate the merged module-boundary and source-verification checks and
explicit Runner Evalbook viewer build.

## Verification

- Passed all 505 workflow, routing, cache, sharding, and
source-verification tests.
- Passed `actionlint` for both workflow files and `git diff --check`.
- Passed the runner infrastructure's workflow routing regression tests.
- Confirmed that the gate is unchanged between the old and new workflow
definitions. The restricted runner group permits the exact merged SHA
and retains its existing workflow pins.
- Full workspace typecheck and build passed on the timeout-fix branch.
Its 107 targeted runner tests and all Linux CI groups passed. The broad
local Mac test command had 15 unrelated permission and
missing-skill-path failures, documented in #13301; it stopped before
later groups.
- [Current-head CI run
34662664879](https://github.com/paperclipai/paperclip/actions/runs/34662664879)
passed. All 33 checks are green or intentionally skipped at
`ded4f1d093644b9279cdf56457f69a196fcf8cc2`. Greptile is 5/5, with no
unresolved findings.
- The live build restored the master pnpm cache, downloaded the resolved
lockfile artifact, and completed a frozen install. It had no
dependency-cache upload step. GitHub reports zero cache entries for this
PR. Native verification and the app build started together and both
passed.
- [Post-merge cloud readiness for
#13301](https://github.com/paperclipai/paperclip/actions/runs/34662389233)
passed all 30 jobs in 12m42s from merge. Native verification restored
the master Rust cache and passed the formerly failing fixtures.

## Risks

A new PR-only dependency can require a download until a master cache
contains it. The new pin also activates the merged workflow changes
listed above, so the live PR must pass all checks. The repository
storage cap is still 10 GB; this PR does not increase it. No allowlist
or runner permission logic changes.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 18:05:36 -07:00
Devin FoleyandPaperclip 44dde2dec4 ci: reuse dependency caches without per-PR uploads (#13300)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud releases wait for source verification before deployment.
> - That verification reuses compiled Rust dependencies to finish
sooner.
> - PR jobs save large pnpm stores under separate merge refs and
different lockfile keys.
> - Those copies compete with master build caches for the repository's
10 GB cache limit.
> - This PR makes PR dependency caches restore-only and reuses
master-compatible keys.
> - A separate pin update will activate the reviewed workflow.

## Linked Issues or Issue Description

**What happened?**
PR merge refs accumulated roughly 700 MB copies of the same pnpm store.
Master Rust caches disappeared, and Cloud readiness run
[34656098157](https://github.com/paperclipai/paperclip/actions/runs/34656098157)
rebuilt dependencies after cache misses. The repository currently has a
10 GB limit. GitHub rejected a request for 50 GB; that setting needs
separate organization/billing access.

**Expected behavior**
PR jobs should reuse downloaded packages without evicting post-merge
compilation caches through duplicate uploads.

**Steps to reproduce**
1. Run several PRs while the checked-in lockfile needs policy
regeneration.
2. Compare the setup-node keys in PR jobs and master jobs.
3. List Actions caches by ref, key, and archive size. The PR keys repeat
across merge refs.

**Paperclip version or commit**
f12b647ae, before this change.

**Deployment mode**
GitHub Actions, with GitHub-hosted and allowlisted AWS PR runners.

Refs #13267 (empty pnpm store prevention). Searched open issues and PRs
for pnpm cache duplication and found no duplicate implementation. This
change leaves the paused capacity documentation PR #13280 alone.

## What Changed

- Replace setup-node cache writes with pinned `actions/cache/restore` in
all seven PR install job definitions.
- Restore against the checked-in lockfile before downloading the
regenerated policy artifact. Keep every install frozen against that
artifact.
- Allow an OS/architecture-specific pnpm fallback and disable automatic
setup-node caching.
- Remove dependency-store caching from the resolution-only policy job.
- Add eight regression tests, update the existing stacked-lockfile cache
assertion, and document cache behavior and storage settings.

## Verification

- Passed 505 workflow, routing, cache, and source-verification tests:
`node --test '.github/scripts/tests/*.test.mjs'
scripts/__tests__/e2e-shard.test.mjs
scripts/__tests__/run-vitest-stable-shard.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs
scripts/cloud-source-verification.test.mjs`.
- Passed `actionlint .github/workflows/pr-trusted.yml` and `git diff
--check`.
- The AWS routing gate is unchanged. Author, event sender, and rerun
actor must still be allowlisted.
- This definition PR does not change the active `pr.yml` pin. After
review and merge, authorize its immutable merge SHA additively and
activate it in a separate PR. Verify a populated restore and no cache
uploads in an allowlisted PR.
- All 32 checks passed or were intentionally skipped on
`44b31eca590f61b75cae646de43c491b6c4deae7`, including full native Runner
verification, application build, server/workspace tests, and browser
shards. Current-head Greptile is 5/5 with no findings. No application
code changes in this PR.

## Risks

- New dependencies present only in a PR may download again on each run
until master saves a cache containing them. Frozen installation remains
the source of dependency resolution.
- Missing or expired stores fall back to normal package downloads.
- Existing PR copies remain until expiry or a separate targeted cleanup.
No cache entries are deleted here.
- The workflow only takes effect after the separate immutable pin
rotation. Storage billing settings are not changed by this PR.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 17:31:15 -07:00
Devin FoleyandPaperclip 37d7dfb0e3 ci: allow dependency changes in cloud eval verification (#13286)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployment requires source verification for the exact merged
commit.
> - Contributor PRs leave lockfile updates to a separate bot PR.
> - Most release checks can refresh an outdated lockfile while
installing dependencies.
> - Two Runner checks still require a frozen lockfile and fail after
dependency changes.
> - This PR gives those checks the same install policy as the other
release checks.
> - A valid dependency change can become deployable without waiting for
another merge.

## Linked Issues or Issue Description

Refs #13257. The dependency change in #13256 exposed this gap. The
separate lockfile update is #13279. Related #12115 addresses the bot PR
check trigger; this PR fixes exact-source cloud verification itself.

**What happened?**

[Cloud readiness for
2083bf6](https://github.com/paperclipai/paperclip/actions/runs/34651761811)
failed in the Runner scorer and chaos jobs with
`ERR_PNPM_OUTDATED_LOCKFILE`. The commit added `svix` to server
dependencies. The tracked lockfile still describes the previous
manifest. The other release checks install with `--no-frozen-lockfile`.

**Expected behavior**

Every source check installs and tests the same checked-out commit. A
pending bot lockfile PR must not block cloud readiness.

**Steps to reproduce**

1. Check out master commit 250deab, which retains the manifest/lockfile
mismatch.
2. Run `pnpm install --ignore-scripts --frozen-lockfile`. It fails with
the same outdated-lockfile error.
3. Run `pnpm install --ignore-scripts --no-frozen-lockfile
--resolution-only`. It succeeds.
4. Restore the generated lockfile. This PR does not commit it.

## What Changed

- Use `--no-frozen-lockfile` in the release Runner scorer job.
- Use the same option in the reusable Runner chaos workflow.
- Document why cloud source checks allow a job-local lockfile refresh.
- Update the existing Runner scorer workflow assertion to match its
install policy.

## Verification

- All 457 workflow tests pass across `.github/scripts/tests/*.test.mjs`
and `scripts/__tests__/release-verify-workflow.test.mjs`.
- `actionlint` passes for both changed workflows.
- Reproduced the frozen install failure against the real tracked
manifest and lockfile. The refresh command passes in 4.6 seconds.
- `git diff --check` passes. No lockfile changes remain.
- No application source changes. Full local application typecheck,
build, and test commands were not rerun in this dependency-free workflow
worktree. Current-head GitHub CI must pass before merge.
- After merge, verify both affected jobs pass on the exact master source
even if the lockfile bot PR remains pending.

## Risks

pnpm can resolve allowed dependency ranges when a manifest outgrows the
tracked lockfile. This matches the existing release install policy. The
resulting lockfile stays in the job workspace. Verification commands and
runner routing are unchanged. The security reviewer explicitly accepted
this existing dependency-policy tradeoff for both jobs after reviewing
repository policy and the source/authorization checks. A future shared
immutable dependency artifact would improve reproducibility across jobs.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] Local verification passes: all 457 workflow tests, actionlint, and
the stale-lockfile reproduction described above.
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 15:56:21 -07:00
Devin FoleyandPaperclip 19c76bfc3f fix(ci): avoid empty pnpm caches from lockfile refresh (#13267)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - CI installs dependencies before it verifies and builds cloud
artifacts.
> - Install jobs share a pnpm package-store cache with lockfile refresh.
> - Lockfile refresh resolves versions without downloading packages.
> - That job saved an empty cache before full install jobs could save
theirs.
> - This PR prevents lockfile refresh from publishing that empty entry.
> - Full install jobs can then populate the cache and reuse
dependencies.

## Linked Issues or Issue Description

Refs #13259 for the related cloud verification cache work. No duplicate
empty-cache fix was found.

**What happened?**

Refresh Lockfile run 34517514932 saved a 216-byte default-branch pnpm
cache at 18:58:08 UTC on September 10. Full install jobs still restore
that empty entry. The cache API reports 216 bytes for master and about
703 MB for populated entries with the same key and cache version in PR
scopes.

**Expected behavior**

A job that installs dependencies should populate the shared
package-store cache.

**Steps to reproduce**

1. Run lockfile refresh with a new lockfile cache key.
2. Its resolution-only command leaves the package store empty.
3. The Node action saves the empty archive before a full install
finishes.
4. Later jobs report a cache hit but download packages again.

**Paperclip version or commit**

Observed on master 6728e133f8 and still
present at a23ae894a5.

**Deployment mode**

GitHub Actions cloud verification and release workflows.

## What Changed

- Disable package-manager caching in Refresh Lockfile.
- Document how to remove the existing empty default-branch entry and
verify a populated replacement.
- Add regression coverage for explicit and automatic package-manager
cache selection in a resolution-only job.

## Verification

- actionlint and git diff checks pass.
- All 175 existing workflow-script tests pass. Both new regression cases
pass and fail against the original workflow, covering the explicit pnpm
cache and automatic npm cache paths. This change adds no application
behavior.
- [The cache creator
job](https://github.com/paperclipai/paperclip/actions/runs/34517514932/job/103006542158)
logs a 216-byte upload under the same key still used by cloud
verification.
- The batch-wide local full typecheck and build passed. The local full
test run reported 10,600 passed, 65 skipped, and 13 permission failures
in unchanged runtime-skill suites. These checks were not repeated in
this dependency-free worktree. Current-head Linux CI passes. The
unchanged chat and browser suites passed on their single retry; all
final checks are green. Greptile is 5/5 with all threads resolved.
- After merge, delete only the existing empty master cache entry. Verify
that a master install saves a populated archive and subsequent jobs
reuse packages. Measure the net install-time change before claiming a
latency gain.

## Risks

- Lockfile resolution can require fresh registry metadata. It does not
need a cached package store.
- The existing empty cache must be removed once; this change prevents
its recreation by this workflow.
- Cache benefits vary with download speed and archive extraction time.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — focused workflow checks
pass; the batch-wide local test limitation is disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 14:35:28 -07:00
Devin FoleyandPaperclip bc68312327 ci: use reserved AWS capacity for post-merge cloud verification (#13257)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments consume a verified image and exact-source
migrator.
> - An image alone is not deployable until source checks and artifact
checks pass.
> - GitHub-hosted queues delayed those checks and the final readiness
signal.
> - This PR gives trusted master work a separate concurrency allowance
on existing AWS runners.
> - Community PRs and arbitrary source inputs keep the GitHub-hosted
fallback.

## Linked Issues or Issue Description

Refs #13243.

**What existing behavior does this improve?**
Time from a master merge to the Cloud deployable v1 signal.

**Current behavior**
For merge d0b7ba4, the image was available after 7m 42s, but readiness
took 16m 05s. Typecheck queued for 6m 40s and the final readiness job
queued for 1m 46s.

**Proposed behavior**
Allow up to 36 concurrent post-merge verification and migrator jobs on
the existing four-vCPU, 16-GiB AWS runners. Workers launch on demand and
terminate after their job; no always-on worker pool or AWS Reserved
Instance purchase is introduced. Keep the combined runner ceiling
unchanged. A separate operator switch enables this route only after the
restricted runner group and Fleet exist.

**Reason and benefit**
Remove GitHub-hosted queue delays from the cloud deployment path. The
gain depends on queue pressure and which remaining job finishes last;
the observed queues are not additive savings.

**Breaking changes**
None to source verification or readiness contracts. Paid routing is
limited to canonical master push/manual events, with exact source checks
on reusable and migrator jobs.

## What Changed

- Route source verification, artifact waiting, dispatch, and readiness
jobs to the separate post-merge Fleet when enabled.
- Require source inputs to match the event's master SHA. Preview inputs
and raced older migrator dispatches stay GitHub-hosted.
- Keep npm publication on GitHub-hosted runners for trusted publishing.
- Bound AWS job timeouts below the 45-minute instance lifetime.
- Document activation, capacity reservation, and rollback.
- Exercise each actual runner selector against allowed and rejected
event/source combinations.

## Verification

- 268 routing and timeout cases pass, including unapproved PR, fork,
branch/tag, arbitrary ref, and disabled-switch cases.
- All 461 focused workflow, preview, and readiness tests pass. The 284
routing/preview cases also pass after the review fixes.
- actionlint passes for changed workflows with the existing
SC2012/SC2016/SC2129 warnings excluded.
- Full local typecheck and build pass (167s and 206s). `pnpm test:run`
completed: 10,600 passed, 65 skipped, and 13 failed in the unchanged
company-skills-service/runtime-skill-cache suites with local filesystem
permission errors. Linux CI is the required test gate; this is not a
claim of a fully passing local suite. Current-head Linux CI is green,
Greptile is 5/5, and all findings are resolved. The final Build retry
passed on a verified 60 GiB AWS runner after correcting the earlier
disk-capacity failure.
- After activation, verify a master run selects the separate group and
all readiness prerequisites pass.

## Risks

- A missing or incorrectly restricted runner group can leave eligible
jobs queued. Enable the switch only after Fleet and group verification.
- PR bursts have 64 slots after reserving 36 for post-merge work. The
image Fleet retains eight, for the same 108-runner total.
- A migrator dispatch racing a newer merge uses GitHub-hosted runners.
This preserves source trust but can retain some queue delay.
- Roll back placement by disabling AWS_POST_MERGE_CI_ENABLED and
rerunning the whole workflow.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. The exact serving model ID and context window are not exposed
by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — focused change tests
pass; full-suite local permission failures are disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 14:13:01 -07:00
Devin FoleyandPaperclip a23ae894a5 ci: cache Rust dependencies used by post-merge typecheck (#13259)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud images become deployable only after source verification
passes.
> - The typecheck job builds the native Runner binary through the server
package.
> - Fresh runners repeatedly compile Rust dependencies for that binary.
> - This PR caches those dependencies for exact-source master
verification.
> - Workspace code and every typecheck still rebuild or run as before.

## Linked Issues or Issue Description

Refs #13243 and #13257.

**What existing behavior does this improve?**
The typecheck portion of post-merge cloud source verification.

**Current behavior**
The typecheck job has no Rust dependency cache. An observed release
build in this job took 4m 13s, including dependency compilation.

**Proposed behavior**
Restore dependency build outputs for canonical master pushes with the
exact source SHA. Use a separate cache key from the Runner verification
job, which builds other profiles.

**Reason and benefit**
A warm cache should remove roughly 2–3 minutes of dependency compilation
from this job. Overall deployment gains depend on the remaining critical
path. The first cache population still compiles from scratch.

**Breaking changes**
None. All checks remain enabled. Non-master callers compile without
restoring or saving this cache.

## What Changed

- Select the pinned Rust toolchain before the typecheck cache lookup.
- Reuse the existing pinned Rust cache action with a typecheck-specific
key.
- Exclude workspace crates and installed cargo executables.
- Test restore/save trust boundaries and document cache behavior.

## Verification

- All workflow script tests pass locally, including nine new cache
trust/contract cases.
- actionlint passes for release-verify.yml.
- Full local typecheck and build pass on the same source base (167s and
206s); `pnpm test:run` is still running and is recorded with #13257.
This PR changes only the workflow, cache guard tests, and documentation.
- All 32 current-head checks are successful or intentionally skipped,
including the complete Linux test matrix, build, and Greptile 5/5 with
no unresolved findings. Verify cache population and subsequent restore
on actual master runs.

## Risks

- The first run and any toolchain/dependency invalidation compile from
scratch.
- Cache restore/save overhead reduces the benefit for small dependency
graphs.
- Disable the cache step to roll back; the existing uncached build
remains valid.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, repository tools, and code
execution. Exact serving model ID and context window are not exposed by
this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass — focused change tests
pass; full-suite local permission failures are disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 13:15:06 -07:00
Devin FoleyandPaperclip d0b7ba4194 ci: route approved master cloud builds to AWS (#13243)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Paperclip Cloud deploys images built from master commits.
> - Cloud image builds share GitHub-hosted capacity with other
workflows.
> - The organization already operates AWS runners through RunsOn Fleet.
> - This pull request allows approved master builds to use a dedicated
cloud Fleet.
> - The benefit is separate build capacity with a quick operator
rollback.

## Linked Issues or Issue Description

Refs #13189, #13192.

**What existing behavior does this improve?**

Placement of the Docker cloud build after a master merge.

**Current behavior**

Every Docker cloud build uses a GitHub-hosted runner. Busy periods delay
the job.

**Proposed behavior**

An operator variable enables the approved cloud Fleet for canonical
master pushes and manual master builds. Other events, refs, and
repositories use GitHub-hosted runners.

## What Changed

- Add a guarded AWS runner selector to the Docker cloud job.
- Keep the existing image cache, verification, and publication steps.
- Test the selector against master, branch, tag, PR, fork, and disabled
contexts.
- Document provisioning requirements, placement checks, and rollback.

## Verification

- 29 focused Node tests pass for routing, readiness, and disk handling.
- The full workflow-script Node suite passes.
- `pnpm -r typecheck` passes locally.
- The pinned PR routing regression suite passes. The first live PR run
assigned 21 jobs to the approved AWS PR group. AWS then reclaimed 16
Spot instances. The failed run is being repeated on GitHub-hosted
runners while the Fleet moves to On-Demand.
- Actionlint passes with existing shellcheck findings excluded (SC2012,
SC2016, SC2129).
- `git diff --check` passes.
- Greptile reports 5/5 on commit
`764d505a41dd2023751c3f361906fa9ea35bf0c6`, with no review threads.
- All 30 current-head CI checks pass, including typecheck, build, all
server/workspace test shards, Runner verification, and browser tests.
Two Storybook checks are intentionally skipped for this change. Run:
https://github.com/paperclipai/paperclip/actions/runs/34630550799
- The broader local test/build sequence is still running. This Mac has
reported failures in unchanged application suites; their complete Linux
CI shards pass. Local targeted workflow tests and typecheck pass.
- Both On-Demand Fleets are deployed and healthy. Live master
cloud-build verification follows the merge.

## Risks

- Missing Fleet capacity or runner-group authorization can leave an AWS
job queued. Disable `AWS_CLOUD_BUILDS_ENABLED` and rerun the workflow to
use GitHub-hosted capacity.
- The runner group must restrict access to this repository and the
master version of `docker-cloud.yml`.
- Docker needs more disk space than the PR Fleet. Provision 120 GiB
disks and retain the free-space check.
- This changes image build placement only. Source verification and
migrator publication remain separate prerequisites.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, tool use, and code
execution. The exact serving model identifier and context-window size
are not exposed by this environment.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 11:15:55 -07:00
Devin FoleyandPaperclip 4fde92107e fix(ci): reuse cloud source verification for npm canaries (#13233)
Reuse the exact master source-verification result before npm canary publication, removing a duplicate verification matrix while preserving fail-closed release checks.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-11 09:26:26 -07:00
DottaandPaperclip 96bba78fba feat: add readable Storybook branch bookmarks (#13231)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Maintainers use Storybook previews to review the board UI.
> - Branch previews need stable bookmarks that people can read.
> - The current publisher only provides a hashed branch path.
> - This pull request adds a readable branch bookmark after each
successful upload.
> - Existing branch and build links keep working.

## Linked Issues or Issue Description

Refs #13226.

**What existing behavior does this improve?**

Manual Storybook publication for repository branches.

**Current behavior**

The stable branch path contains a hash. The expected
`/storybook/branches/master/` URL does not exist.

**Proposed behavior**

Each publication updates a readable bookmark. The action summary and
Markdown artifact link it. Master uses `/storybook/branches/master/`.
Other names use a safe path segment that preserves case and escapes
special characters.

**Reason and benefit**

Maintainers can save and share a readable URL that opens the latest
published branch build.

**Breaking changes**

None. Existing hashed branch entries still update. Existing build URLs
remain valid.

**Additional context**

This follows the publisher in #13226. A duplicate search found no
related bookmark change. It does not overlap planned core work in
ROADMAP.md.

## What Changed

- Generate readable branch bookmarks without collisions with existing
build directories.
- Upload the bookmark only after the full build and compatibility entry
uploads succeed.
- Link the bookmark in the existing summary and Markdown artifact.
- Document branch-name escaping and test path isolation, stable links,
and upload order.

## Verification

- `node --test scripts/__tests__/storybook-deploy.test.mjs`: 20 tests
pass.
- `actionlint .github/workflows/storybook-deploy.yml
.github/workflows/storybook-visual.yml`: passes.
- `git diff --check`: passes.
- [Master bookmark
publication](https://github.com/paperclipai/paperclip/actions/runs/34613344758):
passed. Opened `/storybook/branches/master/` in the browser and
confirmed a story renders. Downloaded the Markdown report and verified
its bookmark link.
- [Feature branch bookmark
publication](https://github.com/paperclipai/paperclip/actions/runs/34613449034):
passed. Its separate bookmark uses `codex~2Fstorybook-bookmarks`.
- Greptile: 5/5 on `dccaf10413ecf447cb34e622b6b3c505791abb51`, with no
unresolved review threads. All current-head Paperclip CI gates pass,
including typecheck, tests, build, browser suites, and the canary dry
run.
- Full local repository checks were not repeated for this focused
publisher change. The preceding run passed typecheck but encountered
unrelated native-session test failures.

## Risks

- Special characters in branch names use `~HH` byte escapes. For
example, `feature/foo` becomes `feature~2Ffoo`.
- Names that could overlap an existing hashed build directory escape the
final hyphen. Very long names retain a hash suffix.
- The two branch entries update separately. If the final upload fails,
the workflow fails and a rerun can repair the bookmark.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, shell tools, and live deployment
verification. The exact runtime model ID and context-window size are not
exposed in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 10:19:36 -05:00
Devin FoleyandPaperclip 974949a39b ci: spread cloud server verification across ten runners (#13227)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud waits for source verification before deploying a new image.
> - The slowest server verification job spends about ten minutes running
tests.
> - Each job uses one test worker to preserve test isolation.
> - This pull request distributes those suites across ten standard
hosted runners.
> - The benefit is a shorter verification path with the same test
coverage.

## Linked Issues or Issue Description

**Current behavior**

In [readiness run
34572340764](https://github.com/paperclipai/paperclip/actions/runs/34572340764),
the slowest server job ran for 638 seconds. Test execution used 594
seconds. This held readiness behind the image job.

**Proposed behavior**

Use ten general server jobs in the reusable release verification
workflow. Keep the three chat jobs and every existing prerequisite. The
complete partition test verifies that no server suite is omitted or
duplicated.

**Reason and benefit**

Reduce merge-to-deployable time on the existing runner type. The next
longest prerequisite was Runner verification at 526 seconds, so the
initial expected total gain is about two minutes rather than a halving
of readiness time. Measure actual queue and execution time before
claiming a result.

Related: #13198 introduced the separate chat lane. #12577 refreshes
duration estimates; this change leaves that manifest alone.

## What Changed

- Increase the general server matrix from five jobs to ten.
- Verify the ten-way partition covers the complete server suite when
combined with the chat lane.
- Document runner demand and the unchanged local and PR grouping.

## Verification

- `node --test scripts/__tests__/release-verify-workflow.test.mjs
scripts/__tests__/run-vitest-stable-shard.test.mjs`: 29 passed.
- `actionlint .github/workflows/release-verify.yml`: passed.
- Full local `pnpm -r typecheck` and `pnpm build`: passed.
- All latest-head GitHub CI checks passed, including the complete Linux
test partition, build, typecheck, and browser gates. Greptile: 5/5 with
zero open findings.
- [Ten-shard timing
probe](https://github.com/paperclipai/paperclip/actions/runs/34606772388):
all 16 jobs passed; slowest server job 6m 23s versus 10m 38s in the
earlier five-shard sample. This compares the server lane, not total
readiness, and is not a controlled same-source A/B.
- The full local `pnpm test:run` is also running. It has reproduced
previously observed macOS-only failures in unchanged skill-cache and
native-session suites; the corresponding Linux CI suites passed. Final
local results will be attached separately. No affected-workflow test
failed.

## Risks

Five additional concurrent jobs per release verification run increase
runner demand and repeated setup work. Queueing can offset the gain.
Test workers, timeouts, permissions, and readiness requirements stay
unchanged. Revert the matrix and its partition test to restore the
previous split.

## Model Used

OpenAI GPT-6 / Codex, with reasoning, tool use, and code execution. The
exact serving model identifier and context-window size are not exposed
by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run the affected workflow tests locally and they pass;
full-suite macOS limitations are disclosed above
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 07:22:03 -07:00
DottaandPaperclip a20ecce409 feat: publish CODEOWNER-approved Storybook branch previews (#13226)
## Thinking Path

> - Paperclip helps people manage AI agents for work.
> - Maintainers use Storybook to review the board UI.
> - Reviews need public previews of selected repository branches.
> - Each branch needs its own URL so previews do not replace each other.
> - This pull request adds manual, CODEOWNER-controlled publishing to S3
and CloudFront.
> - The action returns stable branch links and permanent build links in
its summary and a Markdown artifact.

## Linked Issues or Issue Description

**What existing behavior does this improve?**

The existing Storybook build and manual visual-review workflow.

**Current behavior**

The repository has no manual branch-preview publisher. A single GitHub
Pages site cannot support independent publishers without combining their
output.

**Proposed behavior**

A CODEOWNER selects a source branch and approves publication. Each
branch has a stable CloudFront URL. A completed build becomes the branch
target only after its upload succeeds. The action attaches
`storybook-deployment.md` with the preview links and source commit.

**Reason and benefit**

Maintainers can share multiple branch previews at the same time. Branch
builds have no repository token permissions or AWS credentials.
Dependency caching and install hooks are disabled. The publisher cannot
write runner dashboard files or delete objects.

**Breaking changes**

None. Normal visual checks keep their existing behavior. This does not
change application code or GitHub Pages settings.

**Additional context**

Searched public issues and PRs for Storybook deployment work. No
duplicate deployment proposal was found. This is maintainer
infrastructure, not a roadmap-level core feature.

## What Changed

- Add `Storybook Deploy` with a source-branch input and a manual entry
through `Storybook Visual`.
- Check the original actor and rerunner against default-branch
CODEOWNERS. Require a protected deployment environment with CODEOWNER
reviewers.
- Separate public-source builds with no repository permissions from an
OIDC publisher restricted to the Storybook S3 prefix.
- Publish distinct branch URLs and retain build URLs. Preserve Storybook
deep links across the branch redirect.
- Add the run summary, a downloadable Markdown deployment report,
focused tests, and operator setup docs and IAM policies.

## Verification

- `node --test scripts/__tests__/storybook-deploy.test.mjs`: 19 tests
pass.
- `actionlint .github/workflows/storybook-deploy.yml
.github/workflows/storybook-visual.yml`: passes.
- [Feature branch live publication and deployment-only
rerun](https://github.com/paperclipai/paperclip/actions/runs/34533202273):
passed.
- [Master branch live
publication](https://github.com/paperclipai/paperclip/actions/runs/34533204743):
passed.
- Both public branch URLs render a component story without browser
errors. A deployment-only rerun updates only the selected branch entry
and preserves the previous build URL.
- AWS policy simulation allows Storybook uploads and denies dashboard
writes and object deletion.
- Full local typechecking passes. Full local tests, build, and
current-head PR checks are running.
- [Revised build and Markdown artifact
validation](https://github.com/paperclipai/paperclip/actions/runs/34605623088):
passed. Downloaded the report and verified its branch URL, build URL,
and source commit.
- The public verifier also checks that the stable branch URL points to
this build and rejects stale targets.

## Risks

- Storybook previews are public. Maintainers must publish only public UI
fixtures.
- Retained builds accumulate until an operator prunes them.
- Environment reviewers must stay synchronized with CODEOWNERS. The
workflow fails closed if its environment loses required protection.
- The existing CloudFront distribution is shared with runner reports.
Separate S3 prefixes and a dedicated role prevent the publisher from
overwriting those reports.

## Model Used

OpenAI GPT-6 via Codex, with reasoning, shell tools, and browser
verification. The exact runtime model ID and context-window size are not
exposed in this session.

## Checklist


- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 09:13:55 -05:00
Devin FoleyandPaperclip 932c8bec56 fix(ci): bake the managed runtime identity into cloud images (#13210)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Managed deployments start from the image built by the Cloud
workflow.
> - The managed runtime requests user and group 1001.
> - The image currently builds the node user as 1000.
> - Startup must remap that user, which can walk a large mounted home
directory.
> - This pull request uses the existing Docker build arguments to bake
user and group 1001 into Cloud images.
> - Matching the runtime identity removes that startup work and helps
avoid health-check retries.

## Linked Issues or Issue Description

Refs #13208, #1923, and #7861. Searched open and closed PRs for the
Cloud UID change. The older #7861 addresses build context and volume
ownership repair. This change uses the existing identity arguments in
the Cloud workflow and preserves ownership repair.

**What happened?**

A measured rollout had a container log `Updating node UID to 1001` after
startup. The container stayed at this step for at least 2 minutes 55
seconds before rollback stopped it. The baked node identity was 1000,
while the managed runtime requested 1001. A health check timed out and
the target required a second deployment attempt.

**Expected behavior**

Cloud images should already have the managed runtime identity. A
matching image should skip user and group remapping. Fresh or mismatched
volumes must still receive ownership repair.

**Steps to reproduce**

1. Build the current Cloud image with its default build arguments.
2. Start it with `USER_UID=1001`, `USER_GID=1001`, and a populated home
volume.
3. Observe the startup user remap before the application starts.

**Paperclip version or commit**

`fc06f7f05f42c675be71ff0927b6334405d520ed`

**Deployment mode**

Docker on managed hosts.

## What Changed

- Pass `USER_UID=1001` and `USER_GID=1001` to the Cloud image build.
- Check the pushed digest's baked identity before the entrypoint can
repair it. Then check the normal entrypoint's effective identity and
writable home before publishing the verified full-SHA tag.
- Add a workflow regression and two entrypoint cases for a matching
Cloud identity, including a mismatched volume.
- Document the runtime identity and the first-build cache cost.

## Verification

- Focused workflow and artifact tests: 27 passed.
- Entrypoint tests: 11 passed. Actionlint passed. Full local `pnpm -r
typecheck` passed. Full local `pnpm build` passed. The manual [Cloud
image
build](https://github.com/paperclipai/paperclip/actions/runs/34575473213)
passed on the exact PR head. It checked Sentry, baked and effective
identity, writable home, orphan reaping, and full-SHA publication. The
new identity check took one second. All 30 PR checks passed; the
Storybook workflow was intentionally skipped. Greptile reviewed commit
`114d408f637a0b53e2e2b1339c263779b1e4ae54` at 5/5 with no findings or
open threads.
- The full local suite for the same application source was already run
in #13205. Its macOS general-server phase had 10,471 passes and 70
failures in seven unchanged files. Those failures included missing
Runner fixtures, filesystem errors, timeouts, a port conflict, and a
load-count mismatch. After configuring Cargo and rebuilding fixtures, 37
of 38 native tests passed; one unchanged native-resume assertion still
failed. Linux PR CI passed. This change adds entrypoint tests and does
not change application code.

## Risks

- The first build must rebuild layers that depend on the base image
identity. Later builds can reuse them.
- A future managed runtime identity change must update these build
arguments and checks together.
- The Dockerfile's self-hosted defaults remain 1000. Runtime overrides
and mounted-volume ownership repair remain supported.
- The observed startup delay supports this change, but fleet timing also
includes provider startup, image pull, canary order, and retries. No
fixed end-to-end gain is claimed before a live rollout.

## Model Used

OpenAI GPT-6 through Codex, with reasoning, code execution, and tool
use. The exact serving model ID and context-window size are not exposed
in this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (focused workflow tests;
full-suite limitations are listed above)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-11 00:58:27 -07:00
Devin FoleyandPaperclip fc06f7f05f fix(ci): isolate chaos verification by caller workflow (#13208)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Cloud deployments require verified artifacts for the merged source
commit.
> - Cloud readiness and the npm release independently run the same
source checks.
> - Their shared chaos workflow used only the source ref as its
concurrency key.
> - One caller could cancel the other caller's required job for the same
commit.
> - This pull request scopes that key to the caller workflow and source
ref.
> - Both callers can finish their checks without blocking deployment
readiness.

## Linked Issues or Issue Description

Refs #13192 and #13205. Searched for related open issues and PRs; no
duplicate fix was found.

**What happened?**

The master push for `398d304e15739d1ee6105633bd8a0e42c929d33f` started
Cloud readiness and Release together. GitHub cancelled the Cloud
readiness chaos job before it acquired a runner. Its annotation reported
a higher-priority waiting request for the same concurrency group. The
required readiness gate cannot pass after that cancellation.

**Expected behavior**

Cloud readiness and Release must each finish source verification for the
same SHA. Standalone chaos evals must also have a separate group.

**Steps to reproduce**

Merge a commit to master while the npm release queue is empty. Both
callers reach the reusable chaos workflow with the same source SHA. See
[the cancelled
job](https://github.com/paperclipai/paperclip/actions/runs/34569569760/job/103168603926).

**Paperclip version or commit**

`398d304e15739d1ee6105633bd8a0e42c929d33f`.

**Deployment mode**

GitHub Actions on master.

## What Changed

- Add the caller workflow name to the chaos workflow concurrency group.
Retain source isolation and cancellation of duplicate calls within the
same workflow.
- Add a regression test that evaluates the group for Cloud readiness,
Release, and standalone evals at the same source SHA.
- Document the concurrency boundary in the readiness runbook.

## Verification

- `node --test scripts/preview-artifacts.test.mjs
scripts/__tests__/release-verify-workflow.test.mjs` passed: 26 tests.
- The new regression test fails against the previous concurrency key and
passes with this fix.
- `actionlint -shellcheck= -pyflakes=
.github/workflows/runner-chaos-evals.yml
.github/workflows/release-verify.yml
.github/workflows/cloud-readiness.yml` passed.
- `git diff --check` passed.
- The full local typecheck passed for the same application source in
#13205. Its macOS general-server test phase had 10,471 passes and 70
failures in seven unchanged application test files: missing Cargo/Runner
test binaries, filesystem permissions, timeouts, a port conflict, and a
load-test count mismatch. Linux CI test checks passed. The full local
build passed with Cargo on PATH. This PR changes workflow configuration,
its test, and documentation only.
- All CI checks pass on the final head, including typecheck, tests,
browser suites, build, and canary dry run. Greptile is 5/5 with no open
findings. After merge, verify both callers' chaos jobs complete for the
same master SHA and record the resulting readiness time.

## Risks

- Two callers may now run chaos tests at the same time. This uses two
existing GitHub runners, which is the intended cost of independent
verification.
- Renaming a caller changes its concurrency group. The fixed prefix
keeps this child group separate from caller-level concurrency groups.
- The readiness gate continues to require every verification
prerequisite. No gate is bypassed.

## Model Used

- OpenAI GPT-6 / Codex, with reasoning, repository editing, and
command/API tools. Exact serving model ID and context-window size are
not exposed by this session.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass (26 focused
workflow/artifact tests)
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-10 23:59:44 -07:00
Devin FoleyandPaperclip d56be3f3fc fix(ci): verify deployable cloud artifacts independently (#13192)
Verify source, build the cloud image, and wait for exact-source migrator packages concurrently. Emit Cloud deployable v1 only when every prerequisite succeeds for the merged full SHA.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:08:29 -07:00
Devin FoleyandPaperclip 5cc51fad06 fix(release): publish exact-source cloud migrators on merge (#13188)
Publish exact-source shared and database migrator packages for each master merge through the existing trusted Release workflow, independently of the full release and image build.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:07:38 -07:00
Devin FoleyandPaperclip 6a7025ebe3 ci: skip cloud runner cleanup when disk headroom is ample (#13191)
Skip cloud runner disk cleanup when both the Docker and workspace filesystems have at least 64 GiB free. Preserve the existing cleanup for low, unavailable, or invalid measurements and verify the actual shell behavior across eight scenarios.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 21:02:42 -07:00
Devin FoleyandPaperclip 5c660a32f3 ci: build cloud images independently for each merge (#13189)
Build cloud images independently for each master commit through a reusable workflow. Preserve production release dependencies and image runtime checks, and write cloud registry caches per commit with bounded ancestor imports to prevent overlapping builds from replacing each other's cache.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:42:18 -07:00
Devin FoleyandPaperclip 59d74b68b2 ci: cache the native Runner in a separate Docker stage (#13195)
Compile the native Runner from its complete Cargo and protocol inputs in a separate cached Docker stage. Preserve Cargo validation and generated-contract checks during the normal application build, normalize input timestamps across checkouts, and compile the isolated target in PR CI.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:36:55 -07:00
Devin FoleyandPaperclip 42961b6ef1 fix(ci): split release chat verification into test shards (#13198)
Split release chat verification into three validated test-line shards and balance other server suites across five runners using the measured native Runner integration cost. Retire each chat case's fixtures after assertions, preserve complete test coverage, and exercise the real shard CLI in PR tests.

Co-Authored-By: Paperclip <noreply@paperclip.ing>
2026-09-10 20:30:30 -07:00