Files
paperclip/doc/runner-api-tools.md
DottaandPaperclip 3447609d22 fix(runner): stream and page large API responses within capture budgets (#14301)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work.
> - Agents use governed API tools to inspect task evidence.
> - Large API results become saved assets with short previews.
> - Reading an asset through the same tool used to create another asset,
so the agent could not reach the rest of the evidence.
> - The 10 MiB response cap also blocked useful large results. Removing
all bounds allowed excessive disk use.
> - This pull request streams responses up to 1 GiB and makes saved text
readable in bounded pages. It adds durable run budgets and capture
admission limits.
> - Agents can inspect complete evidence while tool results, memory use,
and capture work stay bounded.

## Linked Issues or Issue Description

**What happened?**
A large response became an asset. Reading that asset returned another
asset and the same preview. Responses above 10 MiB failed before the
agent could read any page.

**Expected behavior**
The agent can fetch a large response and read its saved text to EOF.
Each page stays bounded. New snapshots have a generous finite limit and
a durable run budget. Existing larger assets remain readable through
byte ranges.

**Steps to reproduce**
1. Call a GET operation that returns more than 10 MiB of text or JSON.
2. Before the fix, the tool returns `api_transport_failure`.
3. With this change, responses up to 1 GiB become streamed snapshots
with artifact references.
4. Read `GET /api/assets/{assetId}/content` with `responseText:
{offsetBytes: 0, limitBytes: 8192}`. Follow `nextOffsetBytes` until
null.

Related work: #14186 added the API fallback tools. #14218 bounded API
discovery.

## What Changed

- Add authenticated UTF-8 text windows to `call_api`, with byte offsets
and total size. Keep each page at or below 24 KiB.
- Stream new responses above 24 KiB through private temporary files into
company-owned assets. Bound each capture to 1 GiB of decoded bytes.
Reject oversized declared lengths before reading and count streamed
bytes before writing.
- Reserve capture budget in the run record before spilling. Allow 4 GiB
per run. Settle successful captures to their actual size. Failed or
interrupted captures retain their full 1 GiB reservation. Run restarts
do not reset the budget.
- Enforce a 20 GiB company snapshot quota with database reservations.
Count legacy snapshots and unfinished storage work across runs and
processes. Asset deletion frees quota.
- Limit large captures to two per company and four per server process.
Hold slots through storage upload and temporary-file cleanup. Use a
10-minute download deadline and 30-second connection/idle-read timeouts.
- Return explicit size, budget, busy, and timeout errors. Preserve
unknown outcomes for mutations whose response cannot be captured.
- Read saved assets through authenticated storage ranges, with at most
two extra bytes for UTF-8 and EOF handling. Unpaged reads return the
existing asset and digest with a bounded preview. Reads create no copies
and do not consume capture budget.
- Keep existing assets above 1 GiB readable in pages. Use safe integer
offsets and PostgreSQL `bigint` asset sizes.
- Stream large S3 uploads through ordered multipart requests. Abort
failed uploads and remove partial local files.
- Revalidate run authority during downloads. Keep company authorization,
GET-only text paging, redirect denial, and mutation replay receipts.
- Document the separate 10 MiB upload limits. This PR does not raise
memory-buffered attachment ingestion limits. Future large video uploads
need streamed ingestion and storage quotas.

## Verification

- Full workspace `pnpm -r typecheck` and `pnpm build` pass after
rebasing on master.
- Focused API and response tests: 1,761 pass. Cover declared and chunked
oversize responses, incorrect Content-Length, exact-limit success,
active-stream deadline, cancellation, cleanup, concurrency admission,
and mutation outcome handling.
- Real HTTP integration: 28 tests pass, including runnerd → PRP →
authority → HTTP, a 12 MiB snapshot, final-page/EOF reads, cross-company
denial, a persisted 3 GiB sparse asset, and large mutation receipt
replay.
- The HTTP suite verifies durable run-budget accounting, simultaneous
runs competing for company quota, legacy snapshot accounting, deletion
refunds, failed-storage reservations, cleaned-failure refunds,
metadata-rollback cleanup refunds, preservation after a lost commit
acknowledgement, and small/saved reads after capture-budget exhaustion.
- A standalone proof streams exactly 1 GiB through the production
capture helper, verifies the final bytes, and removes its temporary
file. It uses repeated 256 KiB chunks and records a peak process RSS of
191 MiB.
- Earlier storage verification covers exact S3 multipart boundaries,
cleanup/abort failures, and a 17 MiB transfer through the real AWS SDK
to a local HTTP S3 endpoint. No cloud S3 qualification was run for this
follow-up.
- The local full test run was interrupted for the company-quota changes.
A later targeted run hit exhausted macOS shared-memory slots before
tests started; two unattached PostgreSQL segments with dead owners were
reclaimed before retrying. All 55 current-head checks pass at
`aebb80ceeeee77d5a56b67bfffd835f2f846878c`, including the full CI test
suite, typecheck, build, browser suites, security scan, and Greptile
(5/5). There are no unresolved review threads. The combined rebased test
catalog also passes (48 tests).
- Earlier paging acceptance passed Daytona and separate staging at
`7739879e9`. Those runs predate the streaming and budget changes.

## Risks

- The 1 GiB response cap and 10-minute active-download deadline are
intentional product limits. Larger live results must use endpoint
pagination or a direct file workflow. Existing larger assets remain
readable through bounded ranges.
- A durable 20 GiB company snapshot quota counts stored runner-api
assets and active/orphan reservations across runs and processes. The
operator can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a
finite value of at least 1 GiB. Deleting snapshots frees capacity;
possible orphan storage must be reconciled before releasing its
reservation.
- A failed capture uses its full reservation. A new large capture needs
a full 1 GiB available, even if it later completes at a smaller size.
Small reads and existing asset pages remain available.
- Concurrency limits apply per server process. The run byte budget is
shared through the database.
- The `integer` to `bigint` migration rewrites asset metadata and takes
an exclusive table lock. File bytes stay in storage.
- A live endpoint is fetched once before returning its snapshot.
Continue reading the saved artifact for stable pages. Mutations may
commit before any size or transport error; inspect state before
retrying.
- Attachment uploads and native file handoffs still default to 10 MiB.
Raising buffered ingestion paths to GiB sizes is separate work.

## Model Used

OpenAI Codex, based on GPT-6, with code execution and repository tools.
The runtime does not expose an exact serving model variant or
context-window size. The earlier paging work also used browser testing
and subagents.

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [x] All Paperclip CI gates are green
- [x] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-09-28 10:36:32 -05:00

15 KiB
Raw Permalink Blame History

Runner API escape hatch

search_api and call_api extend the native runner when an available dedicated operation cannot express the requested work. Existing tools remain preferred; agents do not have to search before using them. Only two tool definitions are advertised. The API catalog is returned on demand, never injected into the initial prompt.

Default availability and operator controls

The escape hatch is enabled by default. No environment variable is required. The native hire_agent tool shares this availability policy. Set PAPERCLIP_RUNNER_API_TOOLS_ENABLED=false on the server to disable these tools. An explicit true also enables them; other explicit values fail closed.

Operators can restrict availability with PAPERCLIP_RUNNER_API_TOOLS_COMPANY_IDS, a comma-separated list of company UUIDs. An unset list allows every company; an explicitly empty list allows none. IDs must match exactly. This restriction also applies when the enabled flag is unset.

A server-owned binding can disable these tools for a baseline eval but cannot override an operator restriction. The server checks the policy when advertising tools, when accepting a call, and immediately before HTTP dispatch after preparing any files. Existing dedicated tools remain available. Operators must update the environment of each server process and restart it for deployment-level changes; this environment switch is not a live settings API.

All company, run, work-mode, credential, and lifecycle checks below still apply.

Discovery and requests

{"query":"create project","limit":5}

Search is deterministic lexical ranking over OpenAPI paths, summaries and the old skill reference. It supports task/issue and other terminology, exact METHOD /api/path/{parameter} lookup, and opaque query/catalog-bound pagination. Results include resolved request schemas, response descriptions, authorization metadata, work modes, examples where available, and relevant dedicated tools with their supported parameters. limit defaults to five and is capped at ten.

{"operationId":"PATCH /api/projects/{id}","pathParams":{"id":"PROJECT_UUID"},"body":{"description":"Updated project description"}}

The catalog determines method and path. companyId is filled from the active binding. Scalars and arrays are accepted in query. body defaults to JSON; contentType supports text and raw uploads. files accepts entries containing exactly one authorized artifactId or task-workspace path, and an optional multipart field. No arbitrary URL, headers, authentication, or remote file URL can be supplied. Routes still validate payloads and enforce permissions.

Requests have a 16 KiB URL limit and 10 MiB request/upload limit. New response captures are limited to 1 GiB of decoded bytes. This is separate from the 24 KiB inline/page limit. The receiver rejects an oversized Content-Length before reading and counts actual bytes before writing, including chunked or compressed responses. Oversized responses return api_response_too_large; narrow the query or use the endpoint's own pagination. A mutation may already have committed, so inspect its state rather than retrying it to obtain a smaller response.

Connection setup and stalled response reads time out after 30 seconds. An active capture has a 10-minute total download deadline. Responses above 24 KiB stream into a private temporary file, then into a company-owned asset. Memory stays bounded by the inline prefix and stream buffers. Binary responses also become assets; text previews are limited to 2,000 bytes.

Large captures reserve 1 GiB against a durable 4 GiB per-run capture budget before creating a file. A completed capture settles to its actual byte count; failed/interrupted captures retain the full reservation to bound retry loops. A process restart does not reset that budget. Small inline responses and reads of existing assets need no reservation. Concurrent large captures are limited to two per company and four per server process, including the storage upload and temporary-file cleanup. Budget exhaustion returns api_response_capture_limit; concurrency exhaustion returns api_response_capture_busy without queuing more large transfers. A 20 GiB company-wide quota counts all stored runner-api snapshots plus unattached reservations. Admission uses a company database lock, so runs and server processes share the same quota. Existing snapshots from before this change count too. Attaching an asset converts its reservation to actual stored bytes; deleting the asset frees that capacity. Small binary snapshots also need storage admission. Ordinary inline text/JSON and existing-asset pages do not.

Operators can set PAPERCLIP_RUNNER_API_COMPANY_CAPTURE_MAX_BYTES to a positive safe integer of at least 1 GiB. Invalid values fall back to 20 GiB. No zero or unlimited setting is accepted. This quota covers API snapshots, not all company attachments. Storage capacity and backend limits still apply.

Handled pre-storage failures release the company reservation after temporary file cleanup, but retain the run's charge against retry loops. After a metadata transaction fails, a locking read must prove the asset was not committed before the uploaded object is removed. Confirmed cleanup refunds the company quota. A crash, failed cleanup, or ambiguous storage write keeps an unattached reservation. Operators must reconcile possible orphan files/objects before deleting that reservation from runner_api_response_reservations; no automatic expiry silently refunds possibly occupied space. Linked rows are removed with their asset. Deleting a run does not release its unattached company reservations.

Temporary files are removed on success or failure. Long captures revalidate the active run at least every MiB or at the next chunk after one second, and again before returning the snapshot. Stopping the run stops its download.

To inspect saved text without creating another artifact, call its authorized content operation with responseText:

{"operationId":"GET /api/assets/{assetId}/content","pathParams":{"assetId":"RETURNED_ARTIFACT_ID"},"responseText":{"offsetBytes":0,"limitBytes":8192}}

The result contains data as text (including JSON), plus responseText with offsetBytes, nextOffsetBytes, and totalBytes. Continue at nextOffsetBytes until it is null. Windows end at UTF-8 boundaries. The limit defaults to 24 KiB and accepts 4–24,576 bytes. The server rejects binary content, invalid UTF-8, and offsets inside a code point or beyond the response. This option only works with GET; never repeat a mutation to retrieve another part of its response. Read the saved artifact for a stable snapshot instead of paging a changing live response. A text-window call against a live response above 24 KiB also returns that snapshot's artifact reference; continue on its content operation. A saved asset page never creates another asset. An unpaged asset read also fetches only a bounded preview and returns the existing reference instead of copying the file. Offsets and total sizes use safe integer byte counts, including values above 2 GiB. Existing assets larger than the 1 GiB capture limit remain readable because each request transfers only a bounded range. Request/upload limits and all route authorization remain. Saved asset pages use authenticated HTTP byte ranges. The storage provider reads only the requested window, with at most two extra bytes for UTF-8/EOF handling. The client validates Content-Range, the total size, and the received byte count; it rejects unsupported or inconsistent ranges instead of downloading the whole asset for every page. Other GET routes are fetched once in full to create the snapshot, so use the returned asset for subsequent pages. S3 uses streamed multipart uploads for large snapshots. Asset sizes are stored as PostgreSQL bigint, preserving the existing numeric API shape.

Media files and upload limits

The 1 GiB limit covers new snapshots returned by call_api, such as large JSON exports or binary API downloads. It does not raise attachment upload limits or limit files an agent creates and edits inside its workspace. Saved asset downloads stream from storage and support byte ranges, including video seeking.

PAPERCLIP_ATTACHMENT_MAX_BYTES separately defaults to 10 MiB for uploads and native file handoffs. call_api uploads also have their own 10 MiB limit. Several upload and handoff paths buffer complete files in memory; raising those defaults to GiB sizes requires streaming ingestion and corresponding admission/budget controls first. For a future video attachment workflow, 2 GiB per streamed file is a reasonable default, with an operator override and storage quotas. Do not claim that this response-paging change enables GiB attachment uploads.

Tool responses identify the HTTP route with apiOperationId. The native protocol reserves operationId and callId for semantic tool-call identity; API metadata must not masquerade as that envelope. Saved mutation receipts are normalized at the tool boundary as well, without repeating their HTTP request. All redirects are refused. Interrupted mutation responses have an unknown outcome, requiring inspection before another mutation. Mutation responses with HTTP 5xx, HTTP 408, redirects, or malformed JSON also retain an unknown outcome. A server may have committed the write before it failed to return a valid response.

Authority and replay

The server revalidates the active native run, assigned task and actor, then creates a server-held agent JWT bound to that company and run. Requests go through the actual HTTP router with its authorization, validation and domain audit behavior. An additional runner.api_called receipt attributes mutations to the run even where older route audit events omit that field. The run and work mode are checked again after asynchronous file preparation, so a stopped run cannot dispatch an upload prepared under its earlier binding.

Ask and pre-acceptance Plan permit reads through the escape hatch. Existing dedicated-tool exceptions are unchanged. Runner-owned checkout, completion, status/assignment transitions, approval decisions and execution-control actions cannot be bypassed through generic calls. Routine creation, schedule/trigger changes and manual/public routine execution require the existing scheduling clients. Direct workspace runtime commands, runtime-slot stop/restart, case automation retries and skill test-run controls also require their existing execution clients. Gateway session credentials cannot enter generic results. Routine metadata remains readable; annotation threads, comments and thread resolution remain available through the fallback. API-only ordinary fields, such as a task's billingCode, remain accessible even when a dedicated tool covers other fields on that endpoint.

Mutation call IDs reserve a durable receipt in the run's existing resultJson before dispatch. Replays return the recorded result. Reusing an ID with different arguments is rejected. A crash after reservation leaves an unknown outcome and never automatically resends the mutation. The limit is 512 mutation receipts per run. No database migration is needed.

Workspace uploads use the existing workspace resource containment checks, no-symlink file opens covering every path component, and bounded descriptor reads. Local uploads require Linux or macOS; authorized artifacts work on other hosts. Lifecycle-sensitive endpoints require an inline JSON object, so a raw uploaded JSON file cannot hide protected fields from policy checks. Artifacts must belong to the bound company. Secret-value access, credential management, secret proposals and company exports require their existing secure clients. Search describes these operations as restricted. call_api rejects them before creating a replay receipt or making an HTTP request. Safe secret metadata listing remains available. Agent credentials are never returned to the model. Streaming, WebSocket, MCP and authentication handshakes are documented as protocol operations requiring their existing clients.

Catalog maintenance

runner-api-catalog.ts builds from the server OpenAPI registry. Experimental pipeline, Cases and smoke-lab routes now share their validators with discovery. Seven Cases/pipeline route shapes are multiplexed by resource identity: the Cases router intentionally forwards unknown resources to the pipeline router. Their separate catalog entries explain which resource identifier is required. Registry authorization descriptions are documentation; actual route checks are authoritative.

Regenerate old-skill enrichment after editing its API reference:

node scripts/generate-runner-api-reference.mjs
node scripts/generate-runner-api-reference.mjs --check
node scripts/generate-runner-experimental-api-metadata.mjs
node scripts/generate-runner-experimental-api-metadata.mjs --check

Mounted-route coverage tests include experimental routes. Three WebSocket mounts are explicitly classified in the catalog. Shared protocol-action catalogs, provider projections and generated compatibility checks include both tools.

Verification and paid evals

The companion paperclip-evals worktree contains evals/runner-api-tools. Its README documents explicit case/model selectors, the cumulative budget ledger, fixture reset, progressive batches, and Evalbook generation. No command defaults to running the entire paid suite. Capability, forced operation contracts and paired common-operation regressions are reported separately.

Provider-free integration tests exercise real runnerd → PRP → authority → HTTP, route validation and audit, stale bindings, Ask/Plan restrictions, identity spoofing, file containment, uncertain mutation receipts and fixture isolation.

The Evalbook viewer uses the existing shared viewer and stylesheet on master. The report retains actual persisted-state summaries for private local inspection; public replay continues to withhold company-state details.

The ACPX sidecar includes the upstream terminal-usage accounting correction from origin/codex/evalbook-default-chat-sept6. Its qualified Claude executable requires Linux x64. The first macOS stage records a zero-cost ACPX admission failure. A later user-authorized OpenCode/OpenRouter Sonnet profile reached a real HTTP read, but the attempt failed on a missing harness completion contract and incomplete terminal accounting. The harness contract is corrected. The missing fourth request was subsequently recovered from the matching OpenRouter session and generation billing record; the original failed attempt remains immutable. New attempts retain an append-only, flushed event journal and bounded provider trace outside disposable runtime files. Provider-free startup succeeds for OpenRouter Sonnet and DeepSeek. See doc/plans/2026-09-07-runner-api-production-readiness.md for remaining release gates.