The desktop app killed its `$SHELL -ilc` probe after 5s and kept launchd's
minimal PATH for the whole session, so local runs could not find tools like
`uv`. Startup still waits 5s, but the probe now keeps running (capped at
60s) and a late answer supplies PATH only; the directory variables stay
inherited because storage is already using them.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
On Linux, a long-running `orx up` whose binary was replaced on disk sees
current_exe() as `<path> (deleted)`, so the supervisor spawn failed with
os error 2 and hook/MCP configs recorded a path that no longer exists.
Move updates.rs's relaunch_target into paths::spawnable_exe and use it at
every daemon-executed site that spawns or writes the orx path.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Add generated release notes and issue links
* Run release note check in CI
* Grant release job read access to linked issues
* Preserve release tags and keep issue lookup optional
* Fix version check release notes link
On Linux, a long-running `orx up` whose binary was replaced on disk sees
current_exe() as `<path> (deleted)`, so spawning `orx supervise` for
chat-launched runs failed with os error 2 and left runs in `starting`.
Strip the marker with the existing relaunch_target helper.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Fix SSH run cancellation when wrapper is defunct
* Finish SSH cancellation and avoid full scans for live wrappers
* Bound SSH cancellation for TERM-ignoring payloads
* Cancel SSH payload process groups including descendants
* Finish cooperative SSH cancellation promptly
* Keep SSH cancellation escalation when ps is unavailable
* Require public reproduction inputs in agent feedback
* Clarify exact errors and conditional identifiers in feedback
* Keep feedback filing compatible with agent permission checks
* Make restricted public inputs reconstructible in feedback
* Generalize feedback detail guidance around sensitivity
* Bump OpenResearch CLI to 0.2.13
* Link the README's downloads to the Windows installer and Linux AppImages
The Windows button now downloads OpenResearch-Setup.exe instead of the CLI
zip, and Get started links the x86_64 and ARM64 AppImages, with the SmartScreen
prompt and the Linux requirements alongside.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Apply suggestion from @greptile-apps[bot]
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Make the README's Linux button a Download for Linux link
Match the macOS and Windows buttons: a left-aligned "Download for Linux" image
that downloads the x86_64 AppImage, with ARM64 linked from Get started.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
* Land file:line citations on the cited line
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Hold line jumps for stale editor buffers; keep schemes out of citations
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Reseed clean buffers that lag the loaded version
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Test the citation sanitizer carve-out; reseed on path changes
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Rebuild ui/dist
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Cite extensionless files; keep cited hrefs out of live links
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Unit-test the citation sanitizer
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Rebuild ui/dist
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
parse_paper_id treated arXiv:1706.03762 as the id, dropped the archive
from arXiv:hep-th/9711200, and turned a trailing slash into an empty id.
Strip the citation prefix, a trailing [cs.CL] tag, trailing slashes, and
.html so pasted abs-page lines and URLs resolve.
* Make the Linux app's single-instance claim race-free and private
Take a lock file before binding the focus socket, so two launches at once
can't both become the app, and fall back to a per-user, per-host directory
instead of the shared temp dir when XDG_RUNTIME_DIR is unset. Clear an
interrupted update's staged AppImage before staging another.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Trust only a private runtime dir, and let a waiting launch take a freed lock
Use XDG_RUNTIME_DIR only when the user owns it and no one else can write it,
retry the lock alongside the socket so a launch waiting on an app that gave
the lock up takes over, say when no private directory exists, and test the
handoff in a scratch directory.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Open the macOS app's dashboard in its own window
The OpenResearch.app now hosts the dashboard in a native window (tao + wry
WKWebView) instead of handing it to the user's browser. It adds a standard
menu bar, save panels for downloads, a native window.confirm panel, and a
Cmd+Q that flushes workspace state and shuts the server down cleanly.
Pop-ups open in the system browser (http/https/mailto only).
The app now prefers port 4792 so the window's localStorage survives
relaunches. The Dock-click tab-focus script, its Apple-events entitlement,
and the SSE client counter it relied on are removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Open the dashboard in its own window on Windows
Adds the Windows desktop app. A small GUI-subsystem OpenResearch.exe
(windows/launcher) starts the orx.exe beside it as `orx app` in a hidden
console, so orx and the git, shell, and agent processes it runs share one
invisible console instead of each flashing a window.
`orx app` shares the macOS window code in src/commands/app.rs: WebView2
window, pop-ups to the system browser, save panel, page-load reveal, and
port 4792. On Windows, closing the window quits, a second launch focuses
the running window (named mutex + event), and the taskbar groups the
window with the Start menu shortcut. An update restart relaunches as the
app on the same port. Quit on both platforms now goes through
up::request_shutdown instead of a self-sent SIGTERM.
An Inno Setup script builds a per-user OpenResearch-Setup.exe with a Start
menu entry and a WebView2 bootstrap. CI builds it as an artifact, and
release-windows-app.yml attaches it to releases once WINDOWS_APP_ENABLED
is set. The icon is embedded via embed-resource.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Fix the Windows build and address review of the Windows app
- The focus listener captured its bare HANDLE (edition 2021 disjoint
capture) instead of the Send wrapper, which failed to compile on Windows.
- An orx.exe away from the CLI installer's prefix, like the app's, is now
Portable even when the installer's receipt exists, so it can update itself.
- Single instance creates its event before the mutex; the launcher hands its
foreground rights to orx.exe.
- Focus requests and server readiness are ignored while quitting, and focus
waits for the first load. The save dialog is owned by the window.
- Telemetry counts an app start only after the instance claim.
- The installer reports a failed WebView2 bootstrap and shows progress.
- Icon embedding now fails the build if no resource compiler is found.
- CI format-checks the launcher; the release job drops an unpinned action.
- Docs: maintainer notes for the release gate, two known gaps, and wording.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tighten the Windows app after a second review
- Keep the macOS save panel free-floating: only Windows gets the window as
its owner, since a parent makes rfd show a sheet on macOS.
- Treat the WebView2 bootstrap as failed unless the runtime is then present.
- Move the UTF-16 helper out of the single-instance module, and bind the
kernel object names before the mutex call that GetLastError follows.
- Docs and a dead_code reason.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Add the Linux desktop app as a self-updating AppImage
A `desktop` cargo feature builds the windowed app on Linux (glibc, WebKitGTK),
leaving the static musl CLI builds untouched. build.rs sets cfg(desktop_app)
for macOS, Windows, and Linux with the feature, replacing the macOS-or-Windows
gates.
The AppImage's AppRun starts `orx app`, sharing the Windows entry point:
the webview is built into the GTK window's box (X11 and Wayland), closing
quits, a Unix socket in XDG_RUNTIME_DIR keeps one instance and focuses it,
and the login shell's PATH is adopted as on macOS.
scripts/build-linux-appimage.sh bundles WebKitGTK with pinned linuxdeploy,
its GTK plugin, and appimagetool, copying WebKit's helper processes and
rewriting libwebkit2gtk's /usr paths to ././ so they resolve inside the
image. CI builds x86_64 and aarch64 AppImages on Ubuntu 22.04 and
smoke-tests each under Xvfb; release-linux-app.yml attaches them with
linux-app.json once LINUX_APP_ENABLED is set.
The AppImage updates itself (InstallChannel::AppImage, updates/linux_app.rs):
a sha256-checked download renamed over the file, production builds only,
and Restart execs the new AppImage on the same port.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Fix the Linux app's downloads, bundled WebKit, and host environment
From review of the Linux desktop app:
- Downloads froze the app: rfd's GTK backend runs its dialog on a thread
of its own, which waits forever on the main context tao holds. The save
dialog is now a gtk::FileChooserDialog run on the main thread.
- The bundled WebKit helpers found their libraries only beside themselves,
so they loaded the host's WebKit or none. They now get an RPATH to the
image's usr/lib, and linuxdeploy no longer makes stray copies of them.
- The GTK hook's variables reached every host program orx opens. AppRun
now saves the session's values and orx hands them back to xdg-open, the
folder picker, error dialogs, and agents.
- The smoke test now runs after WebKitGTK is removed from the runner and
checks that each WebKit helper runs, and loads libwebkit2gtk, from the
image; cleanup kills the app's whole process group.
- Startup failures now reach stderr and a zenity or kdialog dialog.
- The tools and the AppImage runtime are pinned by sha256, and only the
release job that publishes can write to the release.
- Smaller fixes: the shell probe runs after the single-instance claim and
falls back to /bin/sh; APPDIR is canonicalized and must not be empty or
root; relative project paths resolve against home inside the image.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Keep the session's GTK settings across Linux app restarts and terminals
Restore the host GTK variables before an update relaunches the AppImage, so
the new AppRun doesn't save the old image's values as the session's, and in
the dashboard's terminals. Share one APPDIR containment check between the
update channel, relative project paths, and the folder picker, which now keeps
a terminal `orx up`'s working directory.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Relaunch the AppImage only from inside it, and tidy after review
Pick the relaunch target with the same APPDIR check as the update channel, so
an `orx app` started from the app's terminal restarts itself. Hand the shell
probe the session's GTK settings, harden the smoke test's helper checks and
cleanup, and bring the docs up to date.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Run the AppImage detection test only on Unix
Its mount paths aren't absolute on Windows, which has no AppImage anyway.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Import the AppImage test's helper inside the Unix-only test
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Show the app window even when the dashboard never finishes loading
On a UTM VM the Linux app ran with no window: it stays hidden until the page
finishes loading, and that never happened. Show it 15 seconds after the
server is up regardless, say so on stderr, and have the smoke test fail if
that fallback fires or no window appears.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Give the Linux app its name and icon in the dock, and keep GIO off host modules
GNOME shows an unmatched window as a gear named after its class, so name the
program OpenResearch and have each launch install a desktop entry and icon
pointing at the AppImage (TryExec hides it once the file is gone).
Bundle GLib's TLS module and set GIO_MODULE_DIR to the image's, so the bundled
GLib stops loading the host's modules, built against a newer one. A second
launch now shows the window even before the page loads, so a stuck instance
can't swallow it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tidy the Linux app after the fourth review
Write the desktop entry only from the AppImage's own orx (shared with the
update relaunch), keep a session's GIO_EXTRA_MODULES off the bundled GLib,
mark a Dock-reopened macOS window as shown so the load fallback can't reopen
it, close a race in the smoke test, and note the Fedora and openSUSE TLS gap.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Show a stalled app window, and close the app before the installer runs
If the dashboard never finishes loading, show the window 15 seconds after the
server is up and say so on stderr; a relaunch (or a Dock click on macOS) also
shows it before the page loads. The installer and uninstaller now check the
app's single-instance mutex and ask the user to close it, rather than
replacing files under a running app and skipping its shutdown path.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Keep a replaced download's original until it lands, and don't offer a launch without WebView2
Move the file a download replaces aside and restore it if the download fails
or is cancelled (WebView2's flyout can cancel), rather than deleting it up
front. The installer no longer offers to start OpenResearch when the WebView2
Runtime is still missing, since the window could not open.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Add Astra to Codex fallback model catalog
* Preserve custom model IDs in composer
* Use CLI model catalogs without static fallbacks
* Allow CLI default after manual Codex model
* Show terminal model error once in chat
* Preserve model discovery errors and distinct turn failures
* Recover legacy Codex threads when selecting CLI default
* Clarify model ID entry in picker
---------
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Open the macOS app's dashboard in its own window
The OpenResearch.app now hosts the dashboard in a native window (tao + wry
WKWebView) instead of handing it to the user's browser. It adds a standard
menu bar, save panels for downloads, a native window.confirm panel, and a
Cmd+Q that flushes workspace state and shuts the server down cleanly.
Pop-ups open in the system browser (http/https/mailto only).
The app now prefers port 4792 so the window's localStorage survives
relaunches. The Dock-click tab-focus script, its Apple-events entitlement,
and the SSE client counter it relied on are removed.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Open the dashboard in its own window on Windows
Adds the Windows desktop app. A small GUI-subsystem OpenResearch.exe
(windows/launcher) starts the orx.exe beside it as `orx app` in a hidden
console, so orx and the git, shell, and agent processes it runs share one
invisible console instead of each flashing a window.
`orx app` shares the macOS window code in src/commands/app.rs: WebView2
window, pop-ups to the system browser, save panel, page-load reveal, and
port 4792. On Windows, closing the window quits, a second launch focuses
the running window (named mutex + event), and the taskbar groups the
window with the Start menu shortcut. An update restart relaunches as the
app on the same port. Quit on both platforms now goes through
up::request_shutdown instead of a self-sent SIGTERM.
An Inno Setup script builds a per-user OpenResearch-Setup.exe with a Start
menu entry and a WebView2 bootstrap. CI builds it as an artifact, and
release-windows-app.yml attaches it to releases once WINDOWS_APP_ENABLED
is set. The icon is embedded via embed-resource.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Fix the Windows build and address review of the Windows app
- The focus listener captured its bare HANDLE (edition 2021 disjoint
capture) instead of the Send wrapper, which failed to compile on Windows.
- An orx.exe away from the CLI installer's prefix, like the app's, is now
Portable even when the installer's receipt exists, so it can update itself.
- Single instance creates its event before the mutex; the launcher hands its
foreground rights to orx.exe.
- Focus requests and server readiness are ignored while quitting, and focus
waits for the first load. The save dialog is owned by the window.
- Telemetry counts an app start only after the instance claim.
- The installer reports a failed WebView2 bootstrap and shows progress.
- Icon embedding now fails the build if no resource compiler is found.
- CI format-checks the launcher; the release job drops an unpinned action.
- Docs: maintainer notes for the release gate, two known gaps, and wording.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Tighten the Windows app after a second review
- Keep the macOS save panel free-floating: only Windows gets the window as
its owner, since a parent makes rfd show a sheet on macOS.
- Treat the WebView2 bootstrap as failed unless the runtime is then present.
- Move the UTF-16 helper out of the single-instance module, and bind the
kernel object names before the mutex call that GetLastError follows.
- Docs and a dead_code reason.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Show a stalled app window, and close the app before the installer runs
If the dashboard never finishes loading, show the window 15 seconds after the
server is up and say so on stderr; a relaunch (or a Dock click on macOS) also
shows it before the page loads. The installer and uninstaller now check the
app's single-instance mutex and ask the user to close it, rather than
replacing files under a running app and skipping its shutdown path.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Keep a replaced download's original until it lands, and don't offer a launch without WebView2
Move the file a download replaces aside and restore it if the download fails
or is cancelled (WebView2's flyout can cancel), rather than deleting it up
front. The installer no longer offers to start OpenResearch when the WebView2
Runtime is still missing, since the window could not open.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Finder's window bounds include the 28pt title bar, so the 640x320 frame
left a 292pt content area: the background's bottom was cropped and the
labels sat closer to the bottom edge than the icons to the top. Size the
frame for a full 320pt content area, and park dot-folders off-window so
Finder builds that show hidden files don't shift the icons down.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The OpenResearch.app now hosts the dashboard in a native window (tao + wry
WKWebView) instead of handing it to the user's browser. It adds a standard
menu bar, save panels for downloads, a native window.confirm panel, and a
Cmd+Q that flushes workspace state and shuts the server down cleanly.
Pop-ups open in the system browser (http/https/mailto only).
The app now prefers port 4792 so the window's localStorage survives
relaunches. The Dock-click tab-focus script, its Apple-events entitlement,
and the SSE client counter it relied on are removed.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Users who connect to orx via the SSH remote gateway (e.g. a laptop
accessing a workstation that orchestrates compute nodes) could not use
the SSH or Slurm compute settings pages — the Connect button, status
badge, terminal, and execution settings were all hidden.
The original intent was to prevent a remote node from re-opening SSH
connections to itself, but this incorrectly penalises the legitimate
3-tier topology: laptop → remote gateway → workstation → compute nodes.
Changes:
- Backend: remove /api/settings/ssh/{master,preflight,connect} from
remote_route_forbidden(); these routes are safe to proxy through the
gateway to the remote orx instance.
- Frontend: remove the remote guards from SshSection and SlurmSection so
the Connect button, status badge, execution settings, and interactive
SSH terminal are always visible regardless of session mode.
Co-authored-by: jie.yuan <yuanjielovejesus@gmail.com>
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Trust the OS root store alongside the bundled webpki roots
reqwest and tokio-tungstenite were built with rustls' compiled-in Mozilla
root list only. On a network where a proxy terminates and re-signs TLS
(Zscaler, Netskope, Palo Alto), the re-signing CA is installed in the OS
trust store but invisible to rustls, so every outbound HTTPS request
failed at connect:
invalid peer certificate: UnknownIssuer
That takes out `discover`, `paper`, update checks, telemetry, and the
Overleaf websocket — `orx` has no working network path on such a network,
while `curl` to the same URL succeeds.
Enable `rustls-tls-native-roots` in addition to `rustls-tls-webpki-roots`
on both dependencies. reqwest merges both sets into one root store, so
this is additive: the bundled roots still apply where there is no OS
store (scratch containers), and the OS store is consulted on top. reqwest
already skips unparseable native certificates, which matters because
system bundles routinely carry a few.
No API or configuration change; corporate users need no flag or env var.
`SSL_CERT_FILE` / `SSL_CERT_DIR` are also honored now, since rustls reads
them when loading native roots.
The musl static-link rationale in the removed comment still holds — this
does not reintroduce OpenSSL.
* Clarify direct websocket trust scope
---------
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Surface the cause chain on alphaXiv transport errors
`main` prints an error's `Display`, so the five alphaXiv request sites
reported only reqwest's own wrapper:
Could not reach alphaXiv at https://api.alphaxiv.org: error sending
request for url (...)
The actionable detail always lives in `source()` — a refused connection,
a DNS failure, or `invalid peer certificate: UnknownIssuer` behind a
TLS-inspecting proxy. None of it reached the user, so a network failure
could not be diagnosed without attaching a debugger or writing a probe
binary against the same client.
Add `transport_error`, which walks `source()` and appends each distinct
cause to the message. An `anyhow` context layer alone would not work
here: that detail is only rendered by `{:#}`, and the entry point uses
`{}`.
Message prefixes are unchanged, so existing output is a strict subset of
the new output.
* Make transport error regression test proxy independent
---------
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Default `orx logs` to a compact path/size/preview summary.
Add `--full` for streaming whole logs and regression tests for raw selectors, UTF-8 previews, and output routing; update orx-evidence guidance for targeted inspection.
* Read a bounded log suffix for default preview instead of the whole file.
Tighten logs_cli coverage for 64KiB-window vs 500-char preview, UTF-8
boundary straddling, --full newline/footer, and --head selector precedence.
* fix(logs): honor captured size during concurrent writes
* docs: describe compact logs default in help
* Simplify run log access to local file preview
* Clarify log file guidance across agent prompts
---------
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* OR-317: limit experiment nodes to measured work
* OR-317: tie experiment nodes to measured hypotheses
* OR-317: narrow experiment node guidance to relevant code
PDF page count, zoom, find, and a new download button now sit in the file
header row instead of a separate toolbar, and image, audio, and video
previews get a header download button in place of the bottom
"Download {name}" strip. Hosts expose a MediaToolbarSlot that previews
portal into; lower-priority controls hide in narrow headers.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The dashboard's PDF preview used <object type="application/pdf">, which
depends on the webview having a PDF viewer. WebKitGTK, which the Linux
desktop app will use, has none. Previews now render with a lazily loaded
PDF.js viewer on every platform: page count, zoom, fit to width, and
in-document search, with links routed like the rest of the dashboard.
- Uses the legacy PDF.js build, whose polyfills cover 2024-era Safari,
WebKitGTK, and Chromium; older engines fall back to the download link.
- Adds a ReadableStream async-iterator polyfill, without which PDF.js's
text extraction (and so search) fails silently in WebKit.
- Ships PDF.js's cmaps, standard fonts, and wasm decoders under a
versioned /pdfjs/<version>/ path, since the server caches unhashed
assets as immutable.
- Scopes the light theme's color-scheme to :root[data-theme="light"] so
PDF.js's stylesheet can't override it.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
The dashboard reports its Paraglide locale to orx on load and on change;
orx persists it in settings.json and sends it as context.locale on every
analytics event. Declined consent events omit it.
Requires the openresearch.sh ingest contract to accept context.locale
before release.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Sign and notarize the macOS CLI binaries in releases
The install.sh archives for macOS shipped unsigned, so device-management
policies on work Macs blocked them. A release-signing-gated job now signs
and notarizes each darwin archive between dist's local and global builds,
rewriting its checksums so the installers and sha256.sum match.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* Harden CLI signing from review
Keep the called workflow from reporting skipped on dry runs, narrow its
token to read, pin the Developer ID requirement the updater checks, keep
notarization logs on failure, and correct the allow-dirty docs.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Work computers block the unsigned CLI that install.sh downloads, while
the DMG app is signed and notarized. Tell those users to install the
app and link its orx into their terminal.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Add a `linked issue` check that fails PRs from forks unless the
description links an issue in this repository. PRs from branches in the
repo are exempt. It runs on pull_request_target so a fork cannot edit its
own check, and never checks out PR code.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Let agents report OpenResearch feedback with orx feedback (OR-308)
Adds the orx feedback subcommand, which posts a bug, feature request, or
frustration report to the openresearch.sh /feedback endpoint (token only when
logged in, no-op when telemetry is opted out), and an orx-feedback agent skill
that sets a high bar for filing, requires redacting research details, and keeps
the report silent. feedback is allow-listed in the Claude plan gate.
* Gate orx feedback like analytics: no reports from development builds
orx feedback now uses telemetry's effective_disabled_reason, so development
builds, ORX_TELEMETRY_ENV test runs, and opted-out users send nothing.
* chore: bump version to 0.2.10
The four pulsing skeleton cards gave no clue that orx was reading the
project, so they looked broken. Surface the existing localized
"Reading the project to suggest where to start…" message with a spinner
below the skeletons, and make that row the status region instead of an
aria-label on the decorative grid.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
`orx discover pubmed` searches PubMed through NCBI E-utilities (esearch for
Best Match PMIDs, efetch XML for records), and `orx paper` reads a PMID,
`pmid:` id, or PubMed URL. PubMed joins the Settings/composer source toggles,
chat activity rendering, and the orx-lit-review skill alongside alphaXiv,
OpenAlex, and bioRxiv.
- Date windows map to `datetype=pdat` with both bounds filled; recency and
historical rerank a wider relevance pool like OpenAlex. No citation counts,
so `popular` keeps relevance order.
- Book records (StatPearls, GeneReviews) are parsed alongside journal articles.
- 429s are retried per NCBI's Retry-After with jitter, since keyless
E-utilities allows 3 req/s and agents issue parallel calls.
- Adds roxmltree because efetch serves abstracts only as XML.
- The PubMed logo is a placeholder mark.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Add "Reveal in file manager" action to the file viewer
Binary and unrecognized files can only be downloaded from the viewer,
which on most systems saves a scary extensionless file instead of
showing where it lives. Add a reveal action next to "Open in default
editor" that opens the OS file manager with the file selected (Finder
on macOS, Explorer on Windows) or the containing folder on Linux.
Mirrors open_in_default_app / open_project_file with the same path
confinement, and is blocked for remote sessions like the open route.
Closes#147
* Address review: dedupe file confinement, cover new locales, fix fallout
- Extract confined_checkout_file so the reveal route shares open's
path validation instead of carrying a seventh copy
- Split reveal_in_file_manager into a testable reveal_command +
shared spawn_detached; the old test popped a real Finder window on
every cargo test run and asserted nothing
- Quote the explorer /select path (commas are legal in NTFS names) and
build it as an OsString so non-UTF-16 names survive
- Add the missing ar/es/hi translations and sort the new key; without
them lint:i18n fails and ui/dist cannot be regenerated
- Exempt /file/reveal from write invalidation like /file/open - a
read-only OS action should not refetch the whole file tree
- Cover the new route in the remote-forbidden test and share the
busy/error wrapper between the two file-viewer buttons
* Use raw_arg for explorer /select on Windows
explorer reparses its own command line rather than consuming argv, so
Command::arg's " escaping of the quoted /select path made the raw
command line unparseable and Explorer silently opened Documents. Emit
/select,"<path>" verbatim instead; NTFS forbids '"' in names so the
quoting can't break out. Also restore doc comments displaced by the
confinement helper extraction.
* Trim reveal comments to the non-obvious why
* Show open/reveal disabled with a reason when the file isn't on disk
* Give missing files their own reason, expose it to screen readers, fix Arabic wording
---------
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Retry ETXTBSY on detection spawns to fix flaky main CI
Two main-branch CI runs failed within hours when a test exec'd a
just-written shell script: execve returned ETXTBSY because a concurrent
fork elsewhere in the test binary briefly held an inherited write fd on
the target until its own exec. probe_bin mapped that transient spawn
error to Unknown, failing assertions that expect a deterministic Broken.
spawn_with_permit is the funnel for every detection child, so retry
ExecutableFileBusy there — the window is milliseconds. Also covers the
production equivalent: probing a CLI while its installer replaces it.
* Escalate the ETXTBSY retry and reuse it for the OpenCode server spawn
3x25ms only tolerated ~75ms of sibling fork-to-exec slowness; the budget
now escalates 25->400ms (~0.8s total), still far inside every caller
deadline, and only paid when ETXTBSY actually fires. Extracted as
spawn_retrying_busy so the same guard covers the one other site where
the transient reaches a user: start_server spawning an opencode binary
its installer is mid-replacing.
* Drop the dead backoff cap, document Unix-only ETXTBSY, test the retry
Review nits: with five retries the escalating backoff peaks at 400ms on
its own, so the .min() clamp never binds; ExecutableFileBusy is a Unix
phenomenon; and the FnMut closure makes the retry policy checkable with
injected errors — no process fixture needed.
* Add composer commands that work the same on every coding agent
The composer's `/` menu now offers /new (alias /clear), /resume, /model,
/plan, /copy and /export on every harness. The dashboard runs each one
itself, so behavior no longer depends on what the selected agent's own CLI
supports — none of them act on slash text in the mode orx runs them in
except Claude Code, and only partially.
`planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the
chips and the send path all read from. A skill whose bare name collides
with a command is dropped from the menu: the composer inserts, and the
server resolves, skills by bare name, so it would otherwise shadow.
Only /plan composes with a prompt. The rest run as a whole message, so
prose that merely mentions `/export` still reaches the agent, and a
command is chipped only where it would actually run.
/resume lists chats across every project: GET /api/chat/sessions now takes
`scope=all` (400 without it or a projectId) and returns the 500 newest.
Picking a chat in another project switches to it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add /compact, which every coding agent can run
The composer's /compact (alias /summarize) shortens a chat's context on
every harness. Claude Code and Codex compact their own context in place
and OpenCode summarizes through its endpoint, so the session survives;
Cursor and the legacy `codex exec` path have nothing to call, so orx
summarizes the transcript itself and reseeds a fresh native session with
it through `bootstrap_context`, which is injected exactly while no native
session exists.
An agent that is merely not running is never traded for a summary: a
resumable session reports "send a message first" instead, since a summary
keeps only the newest 32 KiB of the chat.
Compaction runs as a turn. It holds the session's turn slot, reports
itself as a transcript row that shimmers while it works, and every write
is gated on still owning that slot — Stop cancels it, and a cancelled row
settles as failed rather than sitting at running. The row is the progress:
it survives a reload, reaches every open client, and carries the reason
when it fails. Codex waits for the compaction turn it started rather than
the RPC ack, filtering by that turn's own id; Claude retires its child if
its /compact fails, so an abandoned result cannot land in the next turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add /goal, a standing objective every turn is reminded of
`/goal <text>` gives the chat an objective, `/goal clear` drops it, and a
bare `/goal` reports it. A pill beside Plan shows an active goal and
clears it on click, so it is never invisible state.
The goal is orx's own, stored on the session and re-injected into every
turn. Claude Code's and Codex's native goals are deliberately unused:
theirs mean different things (work until a condition is met, versus a
long-lived objective), and one command cannot mean three things. Being
orx-owned also makes it outlive compaction, a lost native session and a
resume — the agent has to be told its goal every turn regardless.
Goal is the second command that reads the rest of the message, so it is
anchored: it must lead, and it is matched before Plan's unanchored token
so a `/plan` mentioned inside a goal stays part of the goal. An empty
composer gets a session first rather than losing the goal typed into it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Let /resume adopt chats from the agents' own CLIs (#410)
The picker now lists, under "From your terminal", the chats the user had
in Claude Code or Codex outside orx. Adopting one binds an orx session to
the agent's own id and backfills a readable copy of the transcript, so it
does not open blank; the agent itself resumes from its own session, which
is why a marker row says new turns continue in that history.
orx could already resume such a chat — `native_store` resolves an id in
the user's real agent home and the turn runs there. What was missing was
finding them: Claude keeps one JSONL per session (named by its own title
records, else by the first prompt), Codex a SQLite index read `mode=ro` so
its newest threads are not hidden behind the WAL.
Every backfilled row records the native session, so a rewind or a branch
switch resumes the chat rather than silently unbinding it. Adoption is
idempotent under one immediate transaction: two sessions driving one
native chat would split its history. orx's own throwaway children are
filtered out by comparing their working directory against the temp
directory canonically — macOS reports one as `/private/var`, the other as
`/var`, and without that 39 auto-titler runs sat at the top of the picker.
Cursor and OpenCode are deliberately left out: Cursor keys chats by a hash
of the working directory and stores bodies in an encrypted blob, and
OpenCode's live database would have to be copied rather than read.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add composer commands that work the same on every coding agent
The composer's `/` menu now offers /new (alias /clear), /resume, /model,
/plan, /copy and /export on every harness. The dashboard runs each one
itself, so behavior no longer depends on what the selected agent's own CLI
supports — none of them act on slash text in the mode orx runs them in
except Claude Code, and only partially.
`planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the
chips and the send path all read from. A skill whose bare name collides
with a command is dropped from the menu: the composer inserts, and the
server resolves, skills by bare name, so it would otherwise shadow.
Only /plan composes with a prompt. The rest run as a whole message, so
prose that merely mentions `/export` still reaches the agent, and a
command is chipped only where it would actually run.
/resume lists chats across every project: GET /api/chat/sessions now takes
`scope=all` (400 without it or a projectId) and returns the 500 newest.
Picking a chat in another project switches to it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add /compact, which every coding agent can run
The composer's /compact (alias /summarize) shortens a chat's context on
every harness. Claude Code and Codex compact their own context in place
and OpenCode summarizes through its endpoint, so the session survives;
Cursor and the legacy `codex exec` path have nothing to call, so orx
summarizes the transcript itself and reseeds a fresh native session with
it through `bootstrap_context`, which is injected exactly while no native
session exists.
An agent that is merely not running is never traded for a summary: a
resumable session reports "send a message first" instead, since a summary
keeps only the newest 32 KiB of the chat.
Compaction runs as a turn. It holds the session's turn slot, reports
itself as a transcript row that shimmers while it works, and every write
is gated on still owning that slot — Stop cancels it, and a cancelled row
settles as failed rather than sitting at running. The row is the progress:
it survives a reload, reaches every open client, and carries the reason
when it fails. Codex waits for the compaction turn it started rather than
the RPC ack, filtering by that turn's own id; Claude retires its child if
its /compact fails, so an abandoned result cannot land in the next turn.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add /goal, a standing objective every turn is reminded of (#385)
`/goal <text>` gives the chat an objective, `/goal clear` drops it, and a
bare `/goal` reports it. A pill beside Plan shows an active goal and
clears it on click, so it is never invisible state.
The goal is orx's own, stored on the session and re-injected into every
turn. Claude Code's and Codex's native goals are deliberately unused:
theirs mean different things (work until a condition is met, versus a
long-lived objective), and one command cannot mean three things. Being
orx-owned also makes it outlive compaction, a lost native session and a
resume — the agent has to be told its goal every turn regardless.
Goal is the second command that reads the rest of the message, so it is
anchored: it must lead, and it is matched before Plan's unanchored token
so a `/plan` mentioned inside a goal stays part of the goal. An empty
composer gets a session first rather than losing the goal typed into it.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* Add composer commands that work the same on every coding agent
The composer's `/` menu now offers /new (alias /clear), /resume, /model,
/plan, /copy and /export on every harness. The dashboard runs each one
itself, so behavior no longer depends on what the selected agent's own CLI
supports — none of them act on slash text in the mode orx runs them in
except Claude Code, and only partially.
`planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the
chips and the send path all read from. A skill whose bare name collides
with a command is dropped from the menu: the composer inserts, and the
server resolves, skills by bare name, so it would otherwise shadow.
Only /plan composes with a prompt. The rest run as a whole message, so
prose that merely mentions `/export` still reaches the agent, and a
command is chipped only where it would actually run.
/resume lists chats across every project: GET /api/chat/sessions now takes
`scope=all` (400 without it or a projectId) and returns the 500 newest.
Picking a chat in another project switches to it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add /compact, which every coding agent can run (#384)
The composer's /compact (alias /summarize) shortens a chat's context on
every harness. Claude Code and Codex compact their own context in place
and OpenCode summarizes through its endpoint, so the session survives;
Cursor and the legacy `codex exec` path have nothing to call, so orx
summarizes the transcript itself and reseeds a fresh native session with
it through `bootstrap_context`, which is injected exactly while no native
session exists.
An agent that is merely not running is never traded for a summary: a
resumable session reports "send a message first" instead, since a summary
keeps only the newest 32 KiB of the chat.
Compaction runs as a turn. It holds the session's turn slot, reports
itself as a transcript row that shimmers while it works, and every write
is gated on still owning that slot — Stop cancels it, and a cancelled row
settles as failed rather than sitting at running. The row is the progress:
it survives a reload, reaches every open client, and carries the reason
when it fails. Codex waits for the compaction turn it started rather than
the RPC ack, filtering by that turn's own id; Claude retires its child if
its /compact fails, so an abandoned result cannot land in the next turn.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
The composer's `/` menu now offers /new (alias /clear), /resume, /model,
/plan, /copy and /export on every harness. The dashboard runs each one
itself, so behavior no longer depends on what the selected agent's own CLI
supports — none of them act on slash text in the mode orx runs them in
except Claude Code, and only partially.
`planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the
chips and the send path all read from. A skill whose bare name collides
with a command is dropped from the menu: the composer inserts, and the
server resolves, skills by bare name, so it would otherwise shadow.
Only /plan composes with a prompt. The rest run as a whole message, so
prose that merely mentions `/export` still reaches the agent, and a
command is chipped only where it would actually run.
/resume lists chats across every project: GET /api/chat/sessions now takes
`scope=all` (400 without it or a projectId) and returns the 500 newest.
Picking a chat in another project switches to it.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
OpenCode 2.0.6 removed /api/status in favor of /api/info, so orx probed
two dead routes until the 30s timeout and reported "Could not read
OpenCode configuration". The V2 check now negotiates /api/info ->
/api/status -> /api/health and requires a JSON body carrying `version`,
since the SPA fallback can answer unknown routes with HTML 200s.
/api/health answering still marks the legacy (<=2.0.3) route layout for
export/generate paths.
On WSL, xdg-open often has nothing to launch, so the dashboard URL never
reached the Windows browser. On WSL the URL now goes to the Windows host
through rundll32.exe url.dll,FileProtocolHandler (by PATH, then by
absolute System32 path for appendWindowsPath=false), then wslview, then
xdg-open. rundll32 is tried first because some wslview builds re-parse
the URL through cmd and truncate at '&', which breaks the login
callback's state parameter. Native Windows switches from cmd /C start to
the same argv-safe rundll32 call.
The compat fixture learns the same health negotiation, treats the 2.0.6+
synthetic idle message correctly when waiting for a reply, and looks for
export under /api/experimental first.
Refs #341
* Fix first-install onboarding stall via two-phase harness detection
The welcome page gated Continue on GET /api/harnesses, which ran a full
synchronous detection sweep: sequential --version probes per binary
candidate, auth probes, model-catalog calls, and OpenCode V2 server
startup — ~3s on Linux, ~6s on macOS, and 16-25s on Windows where each
subprocess costs 1.5-3s. Startup preflight ran the same sweep again and
discarded it, and the auth monitor cleared the cache on every generation
bump so even warm calls re-detected. OR-305: ~15% of users churned here.
Serve a spawn-free snapshot instead: filesystem-existence install checks
plus file/env auth evidence, marked catalogPending and clamped to
agentReady=false so no consumer treats provisional data as verified. A
single-flight background fill runs the full sweep, swaps the cache, and
emits harness.catalog; the UI polls at 1 Hz while pending so a missed
SSE event can't strand the provisional payload. refresh=1 still runs the
full sweep inline, outside the cache lock.
Also: preflight shares the cache path (no duplicate sweep), candidate
--version probes run in parallel, demo repo install prewarms in the
background off the onboarding click path, and Claude credential presence
(not expiry) indicates refreshable login so the auth overlay no longer
flaps the cache. Pending harnesses render "checking" in Onboarding,
Settings, ModelPicker, ChatPanel, and LocalModelSetup rather than a
false "unavailable"/sign-in prompt.
Measured on fresh VMs: cold harnesses 16-25s -> ~350ms on Windows,
~6.1s -> 0ms on macOS, ~3.2s -> ~1ms on Linux; warm ~0-4ms; onboarding
complete ~4.5s -> ~720ms on Windows. Catalog fill (~2-13s) is off-path.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Cut detection spawns, commit fills progressively, instrument timing
Second OR-305 pass on the Windows first-install profile. CreateProcess
blocks the runtime worker while AV scans each CLI, so detection spawns
now run on the blocking pool; spawn count drops from ~10 to ~4 via
claude's PE version resource + Credential Manager sign-out check, codex
version from the install path, opencode's install manifest, cursor's
bundled node, a merged claude auth/ultracode probe, and canonicalized
path dedupe. The catalog fill commits each harness entry as its probes
land instead of batching behind the slowest, so the first ready agent
unblocks onboarding at ~3.2s rather than the ~8s all-clear. Demo prewarm
git children run at idle priority and escalate if the user clicks in.
Every timed probe records into a bounded buffer that the fill drains and
reports as harness_detect / harness_detect_probe analytics events on a
shared fillId — per-probe and per-pass latency with harness/OS metadata,
so detection hangs are visible in production. Delivery is batched into
one payload per fill and stays nonblocking and opt-out-aware.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Scope probe timings per fill, hold spawn permits across process creation
Review found the shared probe-timing buffer misattributed concurrent
detects to the in-flight fill, semaphore acquisition spent probe
deadlines, a cancelled spawn released its lane while CreateProcess was
still running, and progressive claude entries were overlaid back to
Unknown until fill end. Probe timings are now scoped to a task-local
fill id (spawned probes re-enter it explicitly), the permit travels into
the blocking spawn and returns with a kill_on_drop child, every
deadline acquires its lane before the clock starts, and each committed
claude entry adopts its own auth verdict. first_ready_ms records only
published, overlay-stable readiness.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Scope fill timings to a per-fill sink, bound OpenCode probes, fix JSONC
Review follow-ups: probe timings now collect into a per-fill Arc<Mutex>
sink carried by the task-local instead of a process-global buffer — no
fill ids on records, no drain partitioning, no stale-residue clearing.
OpenCode's full pass falls back to `debug config --pure` when a config
source can't strict-parse (`.jsonc`, comments), so local providers no
longer silently vanish, and `run_models` puts process creation inside
its deadline so a wedged CreateProcess can't hold the fill. The final
first_ready_ms fallback only fires after cache ownership is confirmed,
and ready/installed counts come from the reconciled payload.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Address review majors: telemetry off the lock, timing on cancelled probes
- capture_detect runs via spawn_blocking so settings+outbox IO never sits
on the lock-held /api/harnesses path or a runtime worker
- probe timings record through a drop guard holding the sink, so aborted
speculative probes still report; the spec-miss rerun is timed too
- codex: an explicitly too-old version vetoes steering even when the
model/list handshake answered — catalog liveness isn't turn/steer proof
- opencode manifest pin: reject binary/manifest mtime skew in either
direction, not just a newer binary
- opencode config: snapshot_config returns an "unresolved" flag covering
.jsonc, comments, OPENCODE_CONFIG_CONTENT, and OPENCODE_CONFIG_DIR, so
the full pass knows when to pay for `debug config --pure`
- credman: ERROR_NOT_FOUND means signed-out; other enumeration failures
err toward the live probe
- pe_version reads VS_FIXEDFILEINFO by name instead of u32 offsets
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Await aborted speculative probes before a pass can drain timings
abort() only schedules teardown — the probe's timing guard lands in the
fill's sink when the task is dropped, which could trail the fill's drain
and lose the row. Awaiting the handle makes teardown deterministic.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* chore: bump version to 0.2.8
---------
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Fix orx paper for old-style arXiv ids
parse_paper_id kept only the last path segment, so hep-th/9711200 (the id
orx discover returns for a pre-2007 paper) became 9711200, which alphaXiv
does not know. Keep the archive when the last segment is an old-style
YYMMNNN number.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Review follow-ups: honest is_archive doc, negative-path tests
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Daniel Kim <sox8502@gmail.com>
* Keep unsent composer content scoped to its chat
The composer's draft, attachments, and annotations were a single shared
state across every chat session and the "new chat" screen, so text typed
in one chat appeared in the next. Stash composer content per session
scope on the way out and restore it on return; clear before the first
await on send so the scope-change cleanup can never stash sent text.
Fixes OR-301.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
* Anchor the demo prefill to one chat
The seeded experiment prompt re-seeded on every scope change, so on a
fresh database it appeared in every session and in New Chat — the exact
symptom OR-301 reports. Latch the offer to the first real session scope
instead, clear the latch when that session is forgotten, and keep the
stash cleanup's suppression tied to the scope where the prefill was
offered.
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
---------
Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>