Files
paperclip/docs/agents-runtime.md
Devin FoleyandPaperclip cc17f29e7e fix(timeouts): raise sandbox wall-clock backstop to 4h and make acpx_local timeouts self-describing (#9232)
## Thinking Path

> - Paperclip is the open source app people use to manage AI agents for
work
> - Agent runs execute through adapters (e.g. `acpx_local`), which can
run locally, over SSH, or inside sandbox execution targets, each with a
wall-clock execution timeout
> - Sandbox-backed runs defaulted to a 30-minute wall-clock backstop
(`DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC = 1800`), which kills
healthy long agent runs that are still making progress — long before the
recovery watchdog's 4h critical threshold would even consider them stuck
> - On top of that, `acpx_local` resolved its timeout directly from
`adapterConfig.timeoutSec` instead of the shared execution-target
resolver, and its timeout failures surfaced as a bare `Timed out after
Ns` — giving operators no clue which timer fired or which knob raises it
> - This pull request raises the sandbox backstop to 4h (aligned with
the recovery watchdog), routes `acpx_local` through the shared timeout
resolver, logs the effective timeout and its source at run start, and
makes every timeout error message self-describing
> - The benefit is that long-running sandbox agent runs no longer die at
30 minutes, and when a wall-clock timeout does fire, the run log states
exactly which timer fired and how to configure it

## Linked Issues or Issue Description

Refs #4535 (related: wall-clock execution timeouts killing agent runs
that are still making progress — that issue covers a different hardcoded
600s timer, but the operator pain is the same).

No exact public issue exists for this one, so describing it in-PR:

**Bug:** A long sandbox-backed `acpx_local` agent run was killed with a
bare `Timed out after 1800s` even though the agent was actively working.

- **What happened:** The run hit the 30-minute
`DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` backstop. `acpx_local`
never consulted the shared execution-target timeout resolution (it read
`adapterConfig.timeoutSec` directly, default 0), so on sandbox targets
the sandbox-provider default applied with no adapter-level say. The
resulting error named neither the timer that fired nor the knob that
controls it.
- **Expected:** Healthy long runs should not be killed by a 30-minute
wall-clock backstop when the recovery watchdog only treats runs as
critically stuck after 4h of output silence; and any timeout error
should say which timeout fired and how to raise it.
- **Impact:** Long, legitimate agent runs in sandboxes fail mid-work;
operators waste time reverse-engineering which of several timers
produced "Timed out after Ns".

## What Changed

- `packages/adapter-utils/src/execution-target.ts`
- `DEFAULT_REMOTE_SANDBOX_ADAPTER_TIMEOUT_SEC` raised from `1_800` to
`14_400` (4h), with a comment explaining it intentionally matches the
recovery watchdog's `ACTIVE_RUN_OUTPUT_CRITICAL_THRESHOLD_MS` (4h) so
the adapter backstop never fires before the watchdog path.
Output-inactivity monitors remain the primary hang detectors.
- New `resolveAdapterExecutionTargetTimeout(target,
configuredTimeoutSec)` returns `{ timeoutSec, source }` where `source`
is `configured` / `sandbox_default` / `unlimited`. The existing
`resolveAdapterExecutionTargetTimeoutSec` is preserved as a thin
wrapper, so current callers are unaffected.
- New `formatAdapterExecutionTimeoutErrorMessage(resolution)` and
`formatAdapterExecutionTimeoutStartLogLine(resolution)` produce
self-describing messages that name the timer that fired and the
`adapterConfig.timeoutSec` knob that controls it.
- `packages/adapters/acpx-local/src/server/execute.ts`
- `buildRuntime` now resolves the wall-clock timeout through the shared
resolver: sandbox targets default to the 4h backstop, local/SSH keep the
historical "0 = no adapter timeout", and a configured
`adapterConfig.timeoutSec` always wins.
- The executor logs the effective timeout and its source at run start
(`[paperclip] Adapter execution timeout: …`), so a later timeout is
diagnosable from the run log alone.
- All three bare timeout messages (timer cancel reason, turn result
`errorMessage`, catch-path `messageOverride`) now use the
self-describing format.
- `packages/adapters/acpx-local/src/index.ts` — the adapter
configuration doc for `timeoutSec` states the sandbox default and that
the output-inactivity monitor remains the primary hang detector.
- Tests: `packages/adapter-utils/src/execution-target-sandbox.test.ts`
and `packages/adapters/acpx-local/src/server/execute.test.ts` (see
Verification).

## Verification

- `pnpm --filter @paperclipai/adapter-utils typecheck` — passes
- `pnpm --filter @paperclipai/adapter-acpx-local typecheck` — passes
- `npx vitest run
packages/adapter-utils/src/execution-target-sandbox.test.ts
packages/adapters/acpx-local/src/server/execute.test.ts` — 2 files, 40
tests, all pass
- New/updated test coverage:
- sandbox default resolves to 4h (and the constant is asserted to be `4
* 60 * 60`)
- `resolveAdapterExecutionTargetTimeout` reports `configured` /
`sandbox_default` / `unlimited` sources with the correct precedence
(configured > sandbox default; local/SSH stay unlimited)
- exact wording of the self-describing error message and the
start-of-run log line
- `acpx_local` runtime picks up the sandbox default into `timeoutMs`,
keeps the unlimited local default, honors configured-over-default
precedence, emits the start-of-run log line, and surfaces the
self-describing `errorMessage`/cancel reason when the wall-clock timer
kills a turn

## Risks

- **Behavioral shift:** sandbox-backed adapter runs that previously hit
the 30-minute backstop now run up to 4h before the adapter kills them.
Genuinely hung runs are still caught much earlier by the adapters'
output-inactivity monitors and by the recovery watchdog; the wall-clock
timer is a last-resort kill switch. Operators who relied on the
30-minute default can restore it explicitly via
`adapterConfig.timeoutSec`.
- **Error-message consumers:** any tooling that pattern-matched the
exact `Timed out after Ns` string from `acpx_local` will see the new
self-describing message instead.
- No API or schema changes; `resolveAdapterExecutionTargetTimeoutSec`
keeps its exact signature and behavior (modulo the raised sandbox
default).

> For core feature work, check [`ROADMAP.md`](ROADMAP.md) first and
discuss it in `#dev` before opening the PR. Feature PRs that overlap
with planned core work may need to be redirected — check the roadmap
first. See `CONTRIBUTING.md`.

## Model Used

- Claude (Anthropic) via Claude Code CLI — model ID `claude-fable-5`,
extended thinking enabled, agentic tool use (file edits, shell, test
execution)

## Checklist

- [x] I have included a thinking path that traces from project context
to this change
- [x] I have specified the model used (with version and capability
details)
- [x] I have checked ROADMAP.md and confirmed this PR does not duplicate
planned core work
- [x] I have searched GitHub for duplicate or related PRs and linked
them above
- [x] I have either (a) linked existing issues with `Fixes: #` / `Closes
#` / `Refs #` OR (b) described the issue in-PR following the relevant
issue template
- [x] I have not referenced internal/instance-local Paperclip issues or
links (only public GitHub `#NNN` / `github.com/paperclipai/paperclip`
URLs)
- [x] My branch name describes the change (e.g. `docs/...`, `fix/...`)
and contains no internal Paperclip ticket id or instance-derived details
- [x] I have run tests locally and they pass
- [x] I have added or updated tests where applicable
- [x] I have updated relevant documentation to reflect my changes
- [x] I have considered and documented any risks above
- [ ] All Paperclip CI gates are green
- [ ] Greptile is 5/5 with no open P2s, recommendations, or follow-ups
- [x] I will address all Greptile and reviewer comments before
requesting merge

---------

Co-authored-by: Paperclip <noreply@paperclip.ing>
2026-07-08 13:50:00 -07:00

7.0 KiB

Agent Runtime Guide

Status: User-facing guide Last updated: 2026-03-26 Audience: Operators setting up and running agents in Paperclip

1. What this system does

Agents in Paperclip do not run continuously.
They run in heartbeats: short execution windows triggered by a wakeup.

Each heartbeat:

  1. Starts the configured agent adapter (for example, Claude CLI or Codex CLI)
  2. Gives it the current prompt/context
  3. Lets it work until it exits, times out, or is cancelled
  4. Stores results (status, token usage, errors, logs)
  5. Updates the UI live

2. When an agent wakes up

An agent can be woken up in four ways:

  • timer: scheduled interval (for example every 5 minutes)
  • assignment: when work is assigned/checked out to that agent
  • on_demand: manual wakeup (button/API)
  • automation: system-triggered wakeup for future automations

If an agent is already running, new wakeups are merged (coalesced) instead of launching duplicate runs.

3. What to configure per agent

3.1 Adapter choice

Built-in adapters:

  • claude_local: runs your local claude CLI
  • codex_local: runs your local codex CLI
  • opencode_local: runs your local opencode CLI
  • cursor: runs Cursor in background mode
  • pi_local: runs an embedded Pi agent locally
  • hermes_local: starts your local hermes CLI through @paperclipai/hermes-paperclip-adapter
  • hermes_gateway: calls an already-running Hermes API server through @paperclipai/hermes-paperclip-adapter/gateway
  • openclaw_gateway: connects to an OpenClaw gateway endpoint
  • process: generic shell command adapter
  • http: calls an external HTTP endpoint

External plugin adapters (install via the adapter manager or API):

  • droid_local: runs your local Factory Droid CLI (@henkey/droid-paperclip-adapter)

For local CLI adapters (claude_local, codex_local, opencode_local, hermes_local, droid_local), Paperclip assumes the CLI is already installed and authenticated on the host machine. For hermes_gateway, Paperclip assumes the Hermes API server is already running, reachable from the Paperclip server, and configured with an API key. The older @paperclipai/adapter-hermes-gateway npm package is only a deprecated compatibility shim; the adapter type remains hermes_gateway.

3.2 Runtime behavior

In agent runtime settings, configure heartbeat policy:

  • enabled: allow scheduled heartbeats
  • intervalSec: timer interval (0 = disabled)
  • wakeOnAssignment: wake when assigned work
  • wakeOnOnDemand: allow ping-style on-demand wakeups
  • wakeOnAutomation: allow system automation wakeups

3.3 Working directory and execution limits

For local adapters, set:

  • cwd (working directory)
  • timeoutSec (max runtime per heartbeat; 0 uses the target default — no adapter timeout on local/SSH, a 4-hour backstop on sandbox targets — and a negative value disables the adapter timeout everywhere, including sandboxes)
  • graceSec (time before force-kill after timeout/cancel)
  • optional env vars and extra CLI args
  • use Test environment in agent configuration to run adapter-specific diagnostics before saving

3.4 Prompt templates

You can set:

  • promptTemplate: used for every run (first run and resumed sessions)

Templates support variables like {{agent.id}}, {{agent.name}}, and run context values.

Note: bootstrapPromptTemplate is deprecated and should not be used for new agents. Existing configs that use it will continue to work but should be migrated to the managed instructions bundle system.

4. Session resume behavior

Paperclip stores session IDs for resumable adapters.

  • Next heartbeat reuses the saved session automatically.
  • This gives continuity across heartbeats.
  • You can reset a session if context gets stale or confused.

Use session reset when:

  • you significantly changed prompt strategy
  • the agent is stuck in a bad loop
  • you want a clean restart

5. Logs, status, and run history

For each heartbeat run you get:

  • run status (queued, running, succeeded, failed, timed_out, cancelled)
  • error text and stderr/stdout excerpts
  • token usage/cost when available from the adapter
  • full logs (stored outside core run rows, optimized for large output)

In local/dev setups, full logs are stored on disk under the configured run-log path.

6. Live updates in the UI

Paperclip pushes runtime/activity updates to the browser in real time.

You should see live changes for:

  • agent status
  • heartbeat run status
  • task/activity updates caused by agent work
  • dashboard/cost/activity panels as relevant

If the connection drops, the UI reconnects automatically.

7. Common operating patterns

7.1 Simple autonomous loop

  1. Enable timer wakeups (for example every 300s)
  2. Keep assignment wakeups on
  3. Use a focused prompt template that tells agents to act in the same heartbeat, leave durable progress, and mark blocked work with an owner/action
  4. Watch run logs and adjust prompt/config over time

7.2 Event-driven loop (less constant polling)

  1. Disable timer or set a long interval
  2. Keep wake-on-assignment enabled
  3. Use child issues, comments, and on-demand wakeups for handoffs instead of loops that poll agents, sessions, or processes

7.3 Safety-first loop

  1. Short timeout
  2. Conservative prompt
  3. Monitor errors + cancel quickly when needed
  4. Reset sessions when drift appears

8. Troubleshooting

If runs fail repeatedly:

  1. Check adapter command availability (e.g. claude/codex/opencode/hermes installed and logged in).
  2. Verify cwd exists and is accessible.
  3. Inspect run error + stderr excerpt, then full log.
  4. Confirm timeout is not too low.
  5. Reset session and retry.
  6. Pause agent if it is causing repeated bad updates.

Typical failure causes:

  • CLI not installed/authenticated
  • bad working directory
  • malformed adapter args/env
  • prompt too broad or missing constraints
  • process timeout

Claude-specific note:

  • If ANTHROPIC_API_KEY is set in adapter env or host environment, Claude uses API-key auth instead of subscription login. Paperclip surfaces this as a warning in environment tests, not a hard error.

9. Security and risk notes

Local CLI adapters run unsandboxed on the host machine.

That means:

  • prompt instructions matter
  • configured credentials/env vars are sensitive
  • working directory permissions matter

Start with least privilege where possible, and avoid exposing secrets in broad reusable prompts unless intentionally required.

10. Minimal setup checklist

  1. Choose adapter (e.g. claude_local, codex_local, opencode_local, hermes_local, hermes_gateway, cursor, or openclaw_gateway). External plugins like droid_local are also available via the adapter manager.
  2. Set cwd to the target workspace (for local adapters).
  3. Optionally add a prompt template (promptTemplate) or use the managed instructions bundle.
  4. Configure heartbeat policy (timer and/or assignment wakeups).
  5. Trigger a manual wakeup.
  6. Confirm run succeeds and session/token usage is recorded.
  7. Watch live updates and iterate prompt/config.