2869 Commits
Author SHA1 Message Date
Wu Shuwen 6f7d6a4ee4 fix(daemon): relaunch after stopping the last seat (#325) 2026-10-01 08:51:21 -07:00
Wu Shuwen dc6ab9eed5 docs(cli): clarify context trace seat value (#323) 2026-10-01 08:03:13 -07:00
Wu Shuwen e5681fc4c8 docs(cli): clarify broadcast pod scope (#322) 2026-10-01 08:02:51 -07:00
Shravan Sumanthanan 065df99929 fix(security): prevent temporary file permission leaks, DNS rebinding on terminal WS, and path traversal in bundle unpacking (#292)
* fix(security): enforce mode 0600 and exclusive creation on sendText temporary files

* fix(security): prevent DNS rebinding terminal hijack and add OPENRIG_ALLOWED_ORIGINS support in terminal WebSocket

* fix(security): reject backslash path traversal and Windows drive absolute paths in bundle unpacking

* fix(security): tighten loopback IP and URL origin matching in terminal-ws

* fix(security): only unlink temporary file in sendText if created by this call
2026-10-01 07:40:25 -07:00
KorallisandClaude Opus 5.5 51676aa204 perf(daemon): capture every seat's pane in a few tmux calls per structural sweep (#308) (#309)
* perf(daemon): capture every seat's pane in a few tmux calls per structural sweep (#308)

The structural activity sweep forked one `tmux capture-pane` per running seat
every tick; each fork blocks the event loop ~12-20 ms at ~1 GB RSS, and a
73-seat host spent about a third of its main thread in spawn (#308).

TmuxAdapter.capturePanesContent: one `list-sessions`, then the live sessions
in chained calls of at most 24 (`display-message` marker unique to the call,
then `capture-pane`, exact `=<name>:` targets), each pane's text identical to
capturePaneContent's. A missing session maps to null without a fork. tmux
stops a chain at its first failing command, so a failing chunk is split in
half and retried: an output over the exec buffer fits, and a session gone
since the listing is isolated and left to the per-seat read. No listing ->
null -> the sweep reads per seat as before.

The sweep uses it once per tick: 1 + ceil(live/24) calls instead of one per
seat (72 live seats: 4 instead of 72). pollSeat takes the prefetched text and
otherwise captures per seat; MF1 (a null capture invalidates) and MF2
(single-flight) are unchanged. Exact targets also resolve a session name with
a ".", which `-t <name>` reads as window.pane.

Tests: the adapter against a fake tmux with real chain semantics in both the
shell-string and argv modes (listing + one call, missing -> null, chunks of
24, a vanished session isolated, an over-buffer chunk split, no listing ->
null, marker nonce), and the sweep (one batch, MF1, per-seat fallback).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(daemon): doc comments on the batched capture helpers (review)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(daemon): a batched capture keeps its own time; a slow later chunk can't renew it (review)

The sweep waited for every batch and then stamped each pane's text as observed
"now", so a slow later chunk renewed older captures as fresh: an idle pane's
5.2 s-old active text was projected as running where main lets it expire.

capturePanesContent returns { text, capturedAt } per target, capturedAt being
when that target's own chunk was read (missing sessions: the listing time);
it takes the service's clock. pollSeat stores that time as observedAt, so the
freshness window ages each capture from when it was actually read.

Tests: a slow second chunk gets its own later capturedAt while the first
chunk keeps its earlier one (both exec modes); after a 5.2 s sweep an early
capture is past the 5 s window and is not served, while a fresh one is.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 07:36:51 -07:00
KorallisandClaude Opus 5.5 9dc0c3f1e3 fix(queue): a human-closed row's resolved notice opens no delivery episode (never-posted after a posted ALERT) (#296)
* fix(queue): a human-decision-resolved notice on a closed row opens no delivery episode

Rows a human closed by a direct Slack reply read deliveryOutcome=never-posted although their ALERT was posted. The close
writes a human-decision-resolved owner notification, which the Slack outbound never posts (listHumanAlerts lists active
rows only). The ledger took it as the current episode, found no receipt for its key, and reported never-posted after the
post window. On a row that is no longer active, the episode now falls back to the latest outbound notice
(human-required / human-update). Active rows are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(queue): judge a resolved notice by its own transition, and resolve the closing human (review)

Maintainer review: the fallback used the row's CURRENT state, so a resolved
NOTICE that was really posted (or transport-failed) while the row was active
lost its outcome once an agent later closed the row. The fallback now applies
only when the resolved notice's own transition closed the row, i.e. a human's
direct reply closing it; a notice written while the row stayed active keeps
its own outcome whatever happens to the row later.

CodeRabbit: after the fallback the recipient came from the row's current
destination, so an agent's row that had been parked on the human (its
destination the agent, blocked_on cleared by the close) read inactive. The
recipient is now the human who closed it (the resolved notice's actor).

Tests: "notice posted while active, then the agent closes the row" for posted
and transport-failed receipts; an agent's row parked on the human, ALERT
posted, closed by the human's reply -> posted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 07:22:42 -07:00
Mike SchwarzandOpenRig contributors 451e373165 fix(cli): replace proof artifacts without writing through links (#314)
* fix(cli): replace proof artifacts without writing through links

`rig proof add --replace` opened the existing artifact name with the "w"
flag. That follows a symlink and truncates a shared inode, so replacing
an artifact name that was a symlink or hard link overwrote the other
path's bytes, including files outside the slice and proof media.

--replace now writes a uniquely named staging file next to the artifact,
created exclusively, and renames it over the name. The old link or inode
is never opened for writing, so other paths keep their bytes. A failed
write or rename leaves the target as it was and removes only the staging
file this call created; the original error is reported. A self-sourced
replace still works because the source is read before the write. The
non-replace path keeps its exclusive-create refusal.

The replaced artifact is a new inode with default permissions. There is
no fsync, so this is a directory-entry swap, not a crash-durability
guarantee.

* fix(cli): keep artifact permissions and the first failure on proof replace

An explicit --replace of an existing regular artifact now carries its
permission bits onto the staging file before the swap, so a replacement
never broadens access (a 0600 artifact stays 0600; a read-only artifact
is replaced and stays read-only). A symlinked or new name still gets
default permissions.

The write and close failures are collected separately: the first one is
reported and a later close failure is a warning. A close failure on its
own still refuses the rename, and the target keeps its bytes either way.

* fix(cli): create the proof replace staging file no wider than the artifact

The exclusive staging open now passes the replaced artifact's permission
bits (or the default 0666 for a symlinked or new name) as its creation
mode, so the staging file is never created wider than the artifact it
replaces. The fchmod still sets the exact bits regardless of umask.

---------

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-10-01 06:23:13 -07:00
6a2b804c23 fix(daemon): Codex 0.157/0.158 idle composer (#100) plus a far Working row reads active (#295)
* fix(daemon): recognise the Codex 0.157 empty-composer placeholder as idle

Codex 0.157 renders a fixed placeholder in the empty composer
(`› Ask Codex to do anything`, codex-rs/tui/src/chatwidget.rs `PLACEHOLDER`)
and its footer (`← for agents · ? for shortcuts`) no longer carries the
`· Context [` status bar. Once the activity hook ages out
(AGENT_ACTIVITY_FRESHNESS_MS), send readiness falls back to the pane probe,
which classifies an idle Codex 0.157 seat as `unknown`; with --wait-for-idle
every send to such a seat is refused as `target_activity_unknown`.

The placeholder and footer look the same idle and mid-turn. Mid-turn the
status row (`Working … esc to interrupt`) normally sits above the composer,
so the pattern only counts through MID_WORK_PATTERNS. Codex hides that row
while it streams assistant output, so classifySendReadiness also refuses to
let a placeholder-only idle verdict override a display-fresh (<5min)
running/needs_input runtime hook: a running turn stays running until its hook
ages out. An `unknown` hook (e.g. SessionStart) does not block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(daemon): a Codex placeholder composer with a Working row anywhere above it is active (0.158)

Codex 0.158 keeps the 0.157 placeholder (`› Ask Codex to do anything`); its
footer is a status line plus `? for shortcuts`, so the placeholder pattern
covers it. But its status row (`Working … esc to interrupt`) can sit more than
8 non-blank lines above the composer when queued or incoming message blocks
come in between. The mid-work check only looks at the last 8 lines, and an idle
prompt line then falls through to idle_prompt, so a busy seat read idle, and
with its hook aged out a --wait-for-idle send landed mid-turn.

Under the placeholder, a mid-work line anywhere in the capture now gives
agent_active / mid_work_pattern with that line as evidence. A bare prompt keeps
today's stale-scrollback behaviour. Tests: 0.158 idle and busy pane tails
(redacted from live seats), the far Working row (classifier and a
--wait-for-idle send with an aged-out hook), and a bare prompt with stale
Working text beyond the window still idle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(daemon): only Codex's turn-status row counts as busy beyond the 8-line window (review)

The long-range scan under the placeholder matched the generic Working pattern
anywhere in the capture, so completed output such as "• Working directory:
/tmp/project" or "• Working tree is clean." more than eight lines up made an
idle composer read busy, and a --wait-for-idle send timed out. At that range
only the turn-status row now counts: a bullet, a header (Working, or the
reasoning summary Codex shows in its place), then "(<elapsed> • esc to
interrupt)". The 8-line generic check is unchanged.

Tests: both completed-prose cases read idle; a far status row under a
reasoning-summary header reads active; a send with "Working tree is clean."
far above the placeholder is delivered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Oskar Stępniak <oskar.stepniak52@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-01 05:42:45 -07:00
KorallisandClaude Opus 5.5 394613e95e fix(slack): read a local attachment through one descriptor (follow-up to #298) (#305)
CodeRabbit on #298: defaultReadLocalImage resolved the path twice, statSync
for the checks and readFileSync for the bytes, so the file checked need not be
the file sent. Now the path is opened once (O_RDONLY | O_NONBLOCK, so a FIFO
can't block the open; it is refused as not a regular file), fstat'ed, and read
from the same descriptor, bounded by the size just checked; the descriptor is
always closed.

Tests: the reader opens the path once and never calls statSync/readFileSync on
it; a FIFO with an attachment extension returns "not a regular file" without
waiting for a writer.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 05:39:56 -07:00
Mike Schwarzandv-openrig-build 3fba553a2f docs: describe OpenRig teams as persistent in the npm and Context7 descriptions (#304)
Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 05:11:25 -07:00
Sahil 7ca2724101 feat(slack): clickable questions on human decisions (#193) (#195)
* feat(slack): structured questions with clickable options on human decisions (#193)

- `rig queue create --human-questions-file`: 1-4 questions, 2-4 options each
  (one may be recommended), validated at create; stored by migration 091.
- Slack renders each question as a button row with a complete text fallback;
  a typed thread reply still answers ("Other").
- A block_actions click records that answer on the item; once every question
  is answered the answers are final, one reply row lands on the seat and the
  decision resolves like a typed reply. Failed hand-backs are dead-lettered
  and retried; each click is confirmed in the thread.
- App manifest turns on interactivity (Socket Mode, no request URL); setup
  doc and messaging-the-human skill updated.

* fix(slack): own-key answer lookup; require a human destination for questions(#193)

* fix(slack): redact secrets in answer acknowledgements (#193)
2026-10-01 04:51:28 -07:00
Mike Schwarzandv-openrig-build 29a72ad324 fix(gateway): redact secrets in the Slack file-upload title (#302)
The file title sent with files.completeUploadExternal was the queue summary
(or the file name) as is, while the message text carrying the same summary
already went through redactSecrets. A secret in a row summary was therefore
hidden in the text but shown as the uploaded file's title. The title now
goes through the same redaction. Slack does not parse file titles as mrkdwn,
so the text's control-syntax escaping is not applied to it.

Fixes #300

Claude-Session: https://claude.ai/code/session_01V8RzdrMhGkYLV4rn96kNMN

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 04:37:17 -07:00
Mike SchwarzandOpenRig contributors d0623e06cb feat(daemon): turn the web UI and its terminal WebSocket off by default (#301)
The daemon no longer serves the web UI's pages, assets and SPA routes, or
registers the web-terminal WebSocket, unless `ui.enabled` is true. Turn them
on with `rig config set ui.enabled true`, then stop and start the daemon.
While the UI is off, a page request answers 404 with those steps, and
`rig ui open` asks the running daemon and prints the same steps instead of
opening the page.

The /api routes, /healthz and the HTTP terminal routes (open, views, preview,
status) answer as before. The guard that ran in front of the HTTP terminal
routes through the WebSocket route still runs when the WebSocket is absent.

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-10-01 04:34:43 -07:00
KorallisandClaude Opus 5.5 9648c59fe5 fix(slack): form-encode the external-upload calls; attach local video and PDF (#298)
An update row's local evidenceRef screenshot delivered as text only: the daemon
logged "ATTACHMENT upload-url FAILED ...: invalid_arguments (text delivered;
attachment missing)". files.getUploadURLExternal was sent as a JSON POST.
Slack's reference lists JSON for it, but the method reads form fields only and
answers a JSON body with invalid_arguments ("missing required field: length /
filename").

- callWebApi gets a "form-post" shape: url-encoded fields, objects/arrays as
  JSON strings, the way Slack's own Web API client sends every call. Both
  files.getUploadURLExternal and files.completeUploadExternal use it; the
  JSON-POST family (auth.test, apps.connections.open, chat.postMessage) is
  unchanged.
- Local attachments (LOCAL_ATTACHMENT_EXT): the image types plus .mp4, .webm,
  .mov and .pdf, at most 50 MiB (stat before read). LOCAL_IMAGE_EXT still gates
  https Block Kit image blocks (#47), so a video URL never becomes one.
- The byte upload's default timeout grows with the size (15 s + 1 s / 512 KiB).
- An attachment that can't be sent (over the cap, missing, not a regular file)
  logs "ATTACHMENT skipped ... (text delivered; attachment missing)" instead of
  being silently dropped; a non-attachment ref (e.g. PROOF.md) stays silent.

Tests: the slack-images fake now answers a JSON getUploadURLExternal with
invalid_arguments, as live; form shapes for both upload methods; video through
the three legs; the reader's allowlist, cap boundary and skip reasons.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 04:29:30 -07:00
Mike Schwarzandv-openrig-build 9480978770 docs: describe OpenRig as a network of agents in the README, npm and Context7 (#297)
The README opening now says what OpenRig is in one sentence (open-source
software for building and running your own network of agents) and closes
with the line about the AI civilization experiments. The npm package
description and the Context7 description use the same short summary:
"Build your own network of agents from Claude Code, Codex and Pi:
long-running teams with roles, shared context and owned work."

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 04:15:12 -07:00
Sahil 5fe41efa7c feat(slack): let an update post into an earlier item's thread (--reply-to) (#96) (#155)
* feat(slack): let an update post into an earlier item's thread (--reply-to) (#96)

`rig queue create --human-intent update --reply-to <qitemId>` posts the
update under the earlier item's Slack thread instead of a new top-level
post. Decisions keep their own threads.

- Refused with named errors: without --human-intent update
  (reply_to_requires_update), unknown item (reply_to_not_found), or while
  the earlier item still waits on the human (reply_to_open_decision).
- The thread choice is recorded once, before the first post, so retries
  and --verify agree. Only daemon-written records are trusted.
- A missing, closed, other-channel or other-human root falls back to a
  new post; --verify reports threaded: false with the reason.
- If an item is parked on the human again after an update shared its root,
  the new decision gets a fresh root, and a reply in the older root no
  longer answers it.

* fix(slack): address #155 review — seat check, thread_ts ordering, multi-thread tests

- --reply-to falls back with root-other-seat when the root belongs to a
  different seat, so a human reply can't reach the wrong agent. The
  messaging-the-human skill now says to send the update from the seat
  that owns the thread.
- The newest root of a conversation is Slack's latest thread_ts, not
  opened_at, which a rebuild resets. This applies to all outbound thread
  reuse, not only --reply-to.
- Pin that a reply in an older thread of a conversation no longer
  answers its current decision, without --reply-to.
- Pin that the resolved-decision notice never takes the root away from a
  later --reply-to update (Slack-answered decision and resolved park).

* fix(slack): post a --reply-to update top-level instead of refusing it (#155 review)

An update whose --reply-to item still waits on the human is no longer
refused at create time (reply_to_open_decision is gone). Delivery already
falls back: it posts the update as a new top-level message and records
why on the row, so --verify reports threaded: false with
reference-has-live-gate. A missing thread already fell back the same way
(root-missing).

- Still refused: --reply-to without --human-intent update
  (reply_to_requires_update), and an id naming no qitem on this host
  (reply_to_not_found).
- Tests: create accepts a reference with an open decision or a row parked
  on the human; delivery posts both top-level with the fallback reason.
- CLI help and the messaging-the-human skill describe the fallback;
  skill edge digests regenerated.
2026-10-01 04:12:53 -07:00
KorallisandClaude Opus 5.5 c40dbac0c1 perf(daemon): one tmux call per sweep for seat window activity and pane identity (#161) (#293)
The 1 Hz seat-activity sweep ran one `tmux display-message` per seat, and the
identity sweep a getPanePid + getPaneCommand per seat (plus getPanePid again per
native-process sample). At 87 seats that was ~85 tmux spawns/s from the
activity sweep alone and most of the daemon's event-loop time.

- TmuxAdapter.readAllSessionWindowActivity(): one `list-windows -a`, each
  session's CURRENT window activity: the value `display-message -t <session>`
  reads.
- TmuxAdapter.readAllPaneProcesses(): one `list-panes -a`, pane pid + current
  command.
Both use the adapter's argv seam and TMUX_FIELD_SEPARATOR, with the free-text
field (session name, command) last so a separator inside it stays intact.

The sweeps read once and look each seat up, falling back to the per-seat read
on a miss or a failed batch. Direct pollSeat calls still read per target.
Cadence, debounce, arbitration and #169's generation checks are unchanged; each
native-process sample takes its own list-panes, so samples stay independent.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 03:58:23 -07:00
Rudy Mizrahi Celekli 3e1c3b0108 fix(scope): reject self dependencies before scaffolding (#286)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-10-01 03:20:16 -07:00
Mike Schwarzanddev-driver bb5d17223a fix(daemon): retain bound Claude auto identity across restore (#294)
Co-authored-by: dev-driver <dev-driver@openrig-build>
2026-10-01 03:01:43 -07:00
Mike Schwarzandv-openrig-build 8c4b30bee8 docs(onboarding): tell both doghouse failures and hand over intent, not instructions (#289)
Onboarding piece 01 told only half of the doghouse story: the build that
over-grows (now a light that needs power, not a lock) and then asked "How
big is the dog?", which belongs to the other half. It now tells both: the
moon base, stopped by "does the thing actually need this?", and the tidy
doghouse whose door the dog can't fit through, stopped by "how big is the
dog?". A short paragraph adds that shared understanding of the goal, not
tighter instructions, is what prevents both.

Piece 02's ending becomes "When you need more": the capability map is read
when needed rather than by every seat, a world pack is for planning and
routing work, and the world template is pointed to for people setting one up.

The delegating-work skill gains "Hand over intent, not instructions".

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 02:28:22 -07:00
Rudy Mizrahi Celekli 82dbc01443 fix(daemon): Select the latest inserted same-second restore snapshot (#206)
* fix(daemon): rank same-second restore snapshots by insertion order

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(daemon): handle a restore source pruned between selection reads

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

---------

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-10-01 02:20:39 -07:00
Rudy Mizrahi Celekli 5f33183055 fix(config): report filesystem failures when resetting configuration (#284)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-10-01 02:11:25 -07:00
Rudy Mizrahi Celekli 22622b8981 fix: Preserve running TUI control sockets during a second launch (#205)
* fix(tui): preserve live control sockets when launching

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* test(tui): keep collision sockets inside isolated temporary root

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(tui): serialize stale control socket recovery

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(tui): keep standalone terminals usable on socket collisions

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* test(tui): build child-process entrypoints before package tests

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(tui): recover reservations abandoned by killed launchers

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(tui): expose the complete selected control endpoint for copying

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* test(tui): verify recovered endpoint without assuming inode allocation

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* test(tui): size long socket fixture from byte budget

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

---------

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-10-01 01:51:34 -07:00
Mike SchwarzandOpenRig contributors 809c338843 fix(cli): report a mutating handover's timeout as an unknown outcome (#198)
A mutating `rig seat handover` (and `rig handover`, which shares it) waits up to
120 s for the daemon (#260). When the client reaches that bound the daemon may
still be working, so the handover's outcome is unknown, but the CLI surfaced only
a raw transport timeout (reported in korallis/agent-stack#32).

The CLI now prints "handover outcome unknown" in text and JSON, exits 1, and
makes no second request. It tells the user to inspect the seat before any new
handover, and that a result shown there may belong to an earlier attempt.
A completed or refused handover, a connection failure, and dry-run behave as
before.

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-10-01 01:26:01 -07:00
Rudy Mizrahi Celekli d70c5995a6 fix(bundle): retain recovery exports for inaccessible declared skill paths (#274)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-10-01 00:58:15 -07:00
Mike Schwarzandv-openrig-build 792ef42e5f docs: add contributor maps and a developing-openrig skill (#282)
* docs: add contributor maps and a developing-openrig skill

Add ARCHITECTURE.md (packages, request path, counts with the commands that
produce them, where to add a command, route, migration, adapter, skill,
context pack or scenario), docs/as-built/arteries.md (high-impact areas,
their dependents and past regressions) and docs/as-built/test-layers.md
(what to run before a pull request and what each layer proves).

Add a repo-local developing-openrig skill for Claude Code (.claude/skills)
and Codex (.agents/skills) that routes to these maps, with narrow .gitignore
exceptions so only that skill is tracked. Point CONTRIBUTING.md at the maps
and correct its skill-mirror instructions. Replace the private-host paragraph
in the shipped openrig-skills router with a pointer to the repository skill.

* docs: correct six factual points from review

Newest as-built markers; skills may have one or several copies; the host
scenario runner drops rather than refuses TMUX and daemon variables;
scenario-10 declares its own normaliser; the managed Claude launch applies
only with an explicit permission mode; Dockerfile.scenarios is layered by
run-pr-scenarios.sh. Also tighten several wordings the review flagged.

---------

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 00:49:21 -07:00
Leon.C 94a68a3d77 fix(daemon): render non-image https evidenceRef as a plain link, not an image block (#89)
* fix(daemon): render non-image https evidenceRef as a plain link, not an image block

* fix(daemon): validate Slack evidence rendering boundaries
2026-10-01 00:39:34 -07:00
Shravan Sumanthanan 274e05fdc2 fix(security): prevent SSRF in notifications, drive-by CSRF on API, and pairing queue flooding (#149)
* fix(security): origin guard on /api/*, target URL validation, pairing hygiene, and UTC test fix

- API Origin protection: add apiOriginProtection middleware rejecting unauthorized browser Origin headers on /api/* routes to prevent drive-by web pages posting to the local daemon.
- Notification target URL validation: refuse non-HTTP(S) schemes and embedded URL credentials with clear startup errors and URL credential redaction, set redirect: 'error' on webhook delivery, and preserve self-hosted notifier support on localhost/LAN.
- Pairing hygiene: synchronously reserve capacity before awaited calls to prevent race conditions exceeding MAX_PENDING_PAIRS, prune requests only after a 1-hour grace window past PAIR_TTL_MS so waiting clients receive status 'expired', and sanitize requester strings.
- Timezone test fix: use parseSqliteUtcMs in effective-model-occupant-identity.test.ts to eliminate non-UTC test failure.

* fix(notifications): log warning and disable on invalid target URL, drop redirect error

* fix(notifications): support basic auth credentials in webhook and ntfy URLs

* fix(notifications): safely decode URI components in target URLs falling back on URIError
2026-10-01 00:30:29 -07:00
Mike Schwarzandv-openrig-build 07dab0761d fix(daemon): shipped startup files and resources follow the running install (#261) (#269) (#281)
* fix(daemon): deliver built-in startup files from the running install (#261)

Built-in startup files (CULTURE-default.md, openrig-start.md and the two
onboarding files) are stored per seat as absolute paths inside the install
that created the seat. After an upgrade removes that install (mise and other
versioned managers), every fresh launch rolled back with ENOENT, and restore
blocked on the missing required file. While the old install remained, seats
kept receiving the old version's shipped guidance.

Fresh launch (SeatLifecycleService.readStartupContext) and restore
(pre-validation and replay) now re-anchor a stored file to the running
daemon's assets when it is a recognized built-in: one of the four logical
names, stored at exactly <ownerRoot>/<known relative path>, with ownerRoot a
daemon/assets directory. Custom rig and agent files, including a same-named
CULTURE-default.md, are delivered exactly as stored. Required/appliesOn/
delivery metadata is unchanged, a missing current built-in still fails
honestly, and a same-native resume still replays nothing.

Refs #261

* fix(daemon): project shipped-spec resources from the running install (#261)

Kernel and library agents ship inside the install, so their stored projection
entries (agent guidance, shared runtime settings/MCP fragments) also point
into the install that created the rig. After that install is removed, a fresh
launch still rolled back in projection, before startup-file delivery.

Fresh launch and restore (pre-validation drift checks and replay) now
re-anchor a stored projection entry to the running install's daemon/specs when
both its sourcePath and its resource path lie under the same OpenRig install's
daemon/specs (packaged @openrig/cli/daemon/specs or a packages/daemon/specs
checkout). A resource stored elsewhere, such as the openrig-core plugin under
~/.openrig/plugins, user specs, or a folder merely named daemon/specs, is
projected exactly as stored. Identifiers, category, target, merge behavior and
the existing skill exclusion are unchanged; a resource missing from the
running version keeps the existing failure or warning.

Refs #261

* fix(daemon): own-key built-in lookup; restore-check judges the running paths (#261)

The built-in name lookup used a plain object, so a custom startup file named
constructor, toString or __proto__ matched an inherited key and fresh launch
threw on a valid seat (review finding on be599a64). Use a Map.

RestoreCheckService.checkStartupContext checked and reported the stored
paths directly, so a seat whose old install was removed showed missing
required inputs that replay no longer uses, and named stale paths in its
evidence and remediation. It now applies the same resolvers to the inspected
context and uses the resolved paths in both the predicates and the reported
text. The red/yellow rule, genuinely missing current or custom files, and
missing/malformed/probe-error handling are unchanged. The route's context
reader passes ownerRoot and sourcePath through so the resolvers can apply.

Refs #261

* fix(daemon): shipped-spec startup files follow the running install (#261)

Kernel and library seats also store their rig culture and agent startup files
(culture/CULTURE.md, guidance/role.md, startup/context.md) as absolute paths
inside the install's daemon/specs. With the projection fixed, a kernel created
by a removed install still rolled back on the culture file, and its post-launch
role/context would have followed.

The startup-file resolver now also re-anchors a stored startup file whose
ownerRoot and absolutePath both lie under the same recognized install's
daemon/specs, using the same mapping as projection entries. It applies through
fresh launch (pre- and post-launch delivery), restore replay and
pre-validation, and restore-check. Custom, external and plugin paths, and all
required/applicability/delivery metadata, are unchanged; a file missing from
the running version still fails or warns honestly.

Refs #261

* fix(daemon): re-anchor dev-checkout paths only when missing (#261)

A daemon running from one install remapped every recognized shipped path,
including ones stored from a developer checkout (packages/daemon/...). Those
files can be newer than, or absent from, the running install, so a present
file was remapped to a missing running path and the launch rolled back
(review finding, confirmed by both reviewers).

Packaged installs (@openrig/cli/daemon/...) still re-anchor even while the
stored file exists, so seats never keep an old version's shipped content.
Dev-checkout files re-anchor only when the stored file is missing. Recognition
is limited to those two layouts for built-in assets too. Each consumer passes
its own existence check (restore fsOps, restore-check deps; fresh launch uses
the filesystem).

Refs #261

* fix(daemon): leave legacy persisted startup shapes as stored (#261)

Snapshots and contexts persisted by older versions can lack ownerRoot on
startup files or sourcePath on projection entries. The #261 resolvers called
path.resolve on those fields, which threw, so restore reported restore_error
instead of its real outcome; restore pre-validation walks projection entries
even for an exact resume (CI: restore-honesty-d1-d6 D6a, 4 failures).

Entries without a string root or path are now returned exactly as stored,
which is their pre-#261 behavior.

Refs #261

---------

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-10-01 00:06:13 -07:00
Mike Schwarzanddev-driver d1ba4ee71e fix(cli): honor the recorded daemon endpoint for local commands (#266) (#279)
* fix(cli): use current home daemon endpoint for local clients

* fix(cli): retain recorded endpoint after daemon exit

---------


(cherry picked from commit cdf5fff724)

Co-authored-by: dev-driver <dev-driver@openrig-build>
2026-09-30 23:37:14 -07:00
Mike SchwarzandOpenRig contributors 6963421604 fix: give the managed Claude capability query room before refusing a launch (#278)
A managed Claude launch with an explicit permission mode first runs
`claude --help` to read the supported modes. The query was killed after one
second, so a valid but slow help refused the launch with "capability query
failed". The daemon now allows five seconds; a hung query still ends at that
bound with the same refusal. `rig seat set-permissions` waits up to 10 seconds
and a mutating `rig seat handover` gets the existing 120-second launch window,
so the daemon's answer arrives before the client deadline.

Port of #271 from release/0.6.4 (squash bfa66821). Original author:
dev50-driver; one-line S03 test finisher: dev60-driver.

Refs #260

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-09-30 23:32:01 -07:00
Mike SchwarzandOpenRig contributors b24ef4f6f6 fix(daemon): keep binary files byte-identical in rig bundles (#276)
Bundle create read every vendored agent-package file, and the declared
culture, docs and startup files, as UTF-8 text and wrote it back as UTF-8.
Any byte sequence that is not valid UTF-8 became U+FFFD, and the integrity
manifest was hashed from the changed copy, so inspect and install verified
the corrupted bytes as intact.

Copy those files as bytes: the pod assembler reads them with readFileBuffer
and writes the same bytes, and the install-side skills and workflow-spec
routers copy with fs.copyFileSync. Rewriting rig.yaml and agent.yaml import
refs stays a text operation.

Port of #257 from release/0.6.4 (squash 2174b11e). The async test change in
that squash is already on main via #268 and is not repeated here.

Fixes #245

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-09-30 23:31:44 -07:00
Mike SchwarzandOpenRig contributors 15763c46bb fix(daemon): relaunch a sole seat after fresh --stop ends its tmux server (#277)
Stopping the server's last session ends tmux's server, so launchFresh's
later probes saw transport_unavailable and refused with tmux_probe_failed
after the occupant was already stopped. After its own successful stop,
launchFresh now restores an empty server with the existing startServer()
(a no-op while the server is up; no session is invented), so the
classified probes get a positive answer. Transport, permission,
collision, claimed-seat and wrong-pane refusals are unchanged.

Port of #267 from release/0.6.4 (squash fa770114), without its RC-only
CHANGELOG notes.

Refs #265

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-09-30 23:31:26 -07:00
MelonSmasher 217b5282f8 fix: report identity hook tokenPersisted from stored state (#40) 2026-09-30 23:04:20 -07:00
Mike Schwarzandv-openrig-build 700e6a05be fix(cli): include the repository README in the npm package (#272)
The npm page for @openrig/cli shows "No README data found" because the
published package root has no README.md. The CLI build already copies the
root LICENSE into the package; it now copies the root README.md the same
way, and the generated copy is ignored like LICENSE. The packaging test
now also checks that npm pack carries README.md byte-for-byte.

Co-authored-by: v-openrig-build <v-openrig-build@users.noreply.github.com>
2026-09-30 23:02:03 -07:00
Rudy Mizrahi Celekli 11b14d0150 fix(bundle): warn about unresolved skills in recovery exports (#259)
* fix(bundle): refuse missing declared agent skills

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(bundle): preserve recovery exports with unresolved skill warnings

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

---------

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 22:44:37 -07:00
Rudy Mizrahi Celekli 76ef4d9b13 fix(daemon): compare chat history cutoffs as UTC instants (#208)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 21:37:59 -07:00
Rudy Mizrahi Celekli 02a3ca7bf2 fix(tui): cancel activity streams opened after shutdown (#204)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 21:37:43 -07:00
Rudy Mizrahi Celekli e1139e2276 fix(daemon): probe bracketed IPv6 service targets (#215)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 21:37:09 -07:00
Rudy Mizrahi Celekli 25ee889495 fix(daemon): reject failed Pi startup RPC responses (#213)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 21:29:43 -07:00
Mike Schwarzanddev-driver 964b904fd2 fix(daemon): prove managed Claude identity through shell wrappers (#264)
* fix(daemon): prove managed Claude identity during restore

* fix(daemon): sample Claude lineage only for shell labels

* test(daemon): distinguish bare shells from unusable Claude panes

* fix(identity): retain unavailable final Claude command observation

---------

Co-authored-by: dev-driver <dev-driver@openrig-build>
2026-09-30 21:28:33 -07:00
Rudy Mizrahi Celekli 662288d791 fix(tmux): preserve delimiters in pane working directories (#216)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 21:19:35 -07:00
Mike SchwarzandOpenRig contributors 030227c5d1 test(daemon): assert async ps samplers stay pending through a microtask drain (#268)
The three anti-sync discriminators in list-processes-async.test.ts raced a
0ms timer against the sampler's result. Timer order is not a valid guarantee
of async execution: in a local probe, a 70ms synchronous stall after spawning
a single-PID `ps eww` let the result settle before the 0ms timer in most
trials. The CI failure of this test is consistent with that race; the CI
timing itself was not measured.

Assert instead that the returned promise is still pending after draining 20
microtask ticks: an async execFile cannot settle without an event-loop turn,
while a synchronous ps inside the call settles within a few ticks. The real
rows and HOME assertions are unchanged.

Refs #245

Co-authored-by: OpenRig contributors <noreply@openrig.dev>
2026-09-30 21:08:42 -07:00
Rudy Mizrahi Celekli e0b135ca43 fix(daemon): tolerate unusable snapshots during restore selection (#207)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 20:54:31 -07:00
Rudy Mizrahi Celekli bc8c608506 fix(cli): fail bootstrap and requirements on client errors (#202)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 20:46:15 -07:00
Rudy Mizrahi Celekli 650b96bb4e fix(daemon): keep truncated ntfy titles ASCII (#217)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 20:45:38 -07:00
Rudy Mizrahi Celekli ed9dc62f5a fix(daemon): require every Compose replica to be healthy (#218)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 20:45:17 -07:00
Rudy Mizrahi Celekli b4a05d279a fix(scope): warn about dependencies left behind by slice moves (#263)
Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 19:49:38 -07:00
Ege 1347d825c3 fix(daemon): create parent directory when packing bundle to missing path (#258) 2026-09-30 19:21:09 -07:00
Rudy Mizrahi Celekli 3691134c5b fix: Render terminal-open HTTP errors before consuming open results (#229)
* fix(cli): render terminal open HTTP errors before results

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

* fix(cli): select terminal result printer by response shape

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>

---------

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
2026-09-30 18:38:25 -07:00