mirror of
https://github.com/alphaXiv/OpenResearch.git
synced 2026-10-02 01:34:34 +08:00
main
96
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a87bb4ce5a |
OR-346: Route Windows local runs' python around the Store alias stub (#483)
* OR-346: Route Windows local runs' python around the Store alias stub Git Bash resolved bare `python` to the Microsoft Store App Execution Alias, which exits 9009 (49 to bash) and writes nothing to the run log. run.sh now puts a real interpreter first on PATH (next on PATH, else `py`'s), forwards a stub `python3` to `python`, and otherwise fails with a clear message. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test that Git Bash executes the generated python shims Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8f2fc4f514 |
OR-347: Kill the whole Codex app-server process group on teardown (#484)
* Kill the whole Codex app-server process group on teardown (OR-347) npm installs run `node codex.js app-server`, which spawns the native app-server in the same process group. orx only SIGKILLed the wrapper, so the native server survived holding the thread writer lock, and a replacement server's thread/resume failed with "already has an active writer". Teardown now signals the whole group, once, with a bounded reap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Detect the killed child by stdout EOF in the teardown test A pid probe succeeds on an unreaped zombie, so the test could fail when the new parent reaps slowly; EOF proves every writer in the group exited. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
1483698805 |
OR-327 Add Agentic/Copilot autonomy control to the model picker (#488)
* Add an Agentic/Copilot autonomy control to the model picker Copilot adds a per-turn <orx-autonomy> block telling the agent to explain its plan and ask before launching runs or changing direction. Agentic is the default and sends nothing, so today's behavior is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address review: fix autonomy test, type the API field, tighten UI saves Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Address review round 2: typed store setters, always roll back failed chat saves Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Serialize autonomy preference saves so a quick second change always wins Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
395ec53748 |
OR-349: Retry rate-limited literature requests with a clear wait message (#481)
* OR-349: Retry rate-limited literature requests with a clear wait message alphaXiv, OpenAlex, bioRxiv and arXiv requests failed on the first 429, so parallel agent calls to `orx paper`/`orx discover` gave up mid-analysis. PubMed's existing retry becomes a shared `public_get` used by every keyless literature request: up to 3 attempts honoring Retry-After (1s/2s backoff without it), failing fast with "retry in Ns" when the host asks for longer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Honor HTTP-date Retry-After and randomize retry jitter Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f336b12152 |
OR-338: Restore vanished chat worktrees at their last checkout (#478)
* OR-338: Restore vanished chat worktrees at their last checkout Session worktrees have been reported disappearing mid-turn (Codex, 0.2.10) with no identified cause. Make recovery lossless for committed work: - When a session worktree is missing but still registered, recreate it on its last branch (or commit) instead of the baseline, and log it. - Replace the global `git worktree prune` with a targeted `worktree remove --force <dir>`, so one session's recovery or deletion no longer erases other vanished sessions' checkouts; prune stays as a last resort when the baseline add still fails. - Demo seeding keeps landing on its exact commit. - Refuse enabling GitHub sync (409) while a busy chat is still on the legacy worktree layout, whose migration would `git worktree move` it mid-turn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * OR-338: Make legacy migration atomic with turn starts; drop fallback prune Addresses review on #478: - Run the legacy worktree migration inside ChatHost::while_idle, which holds the turn map and refuses if an affected session has a turn or durable lease, so no turn can start between the check and `git worktree move`. Replaces the racy handler-level guard. - Remove the last-resort global `git worktree prune`: an unrelated add failure could erase other vanished sessions' restorable checkouts. Remove an empty leftover dir up front instead, the one case the prune was covering. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
ddb603173b |
OR-341: Alert waiting agents and the dashboard when Slurm monitoring is lost (#475)
* Surface lost Slurm monitoring to waiting agents and the dashboard Keep supervisor stderr in run-logs/<id>.supervisor.log, alert a session's pending wake-up once per monitoring outage, and show monitoringError on live runs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Debounce Slurm monitoring errors and combine per-session alerts Report monitoringError only after 60s of failed polls, send one alert per session listing every newly unmonitored run, dedupe repeated ssh log-stream errors, and word the alert without promising a wake-up that needs monitoring back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep monitoring tests in their temp store and log each poll failure Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Log monitoring recovery and pin alert formatting in tests Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Exercise the Slurm monitoring grace period in the supervisor integration test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Drop stray bytecode and anchor the grace-period check on the first poll Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep remote ssh text out of monitoring alerts and bound the supervisor log Name unmonitored runs by local ids and point the agent at orx exp status, show a warning from any live run on the dashboard, and restart the supervisor log past 1 MiB. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Show every live run's monitoring reason and roll the supervisor log over orx exp status now prints the monitoring error of older live runs, the dashboard takes the newest live run's error through one helper, and a full supervisor log is renamed rather than truncated under a running supervisor. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3f2d41264a |
OR-340: Refuse agent spawns on a harness orx up cannot find (#474)
* OR-340: Refuse agent spawns on a harness orx up cannot find
`orx agent spawn` only queues the helper, so a `--harness` the resident
server could not launch reported success and failed afterwards with
"codex not found on PATH". Spawn now asks `orx up` (new
`GET /api/harnesses/{id}/snapshot`, the server's own PATH) whether a
switched-to harness is installed and refuses up front if not. Any failure
to ask falls back to queueing as before.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
* OR-340: Accept the callback token on the spawn preflight
On a persistent remote `orx up` host, agents authenticate with the
callback token, which only covered run submission and cancellation, so the
harness snapshot answered 401 and the preflight silently fell open.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
4c3e1b84a1 |
OR-339: Adopt the shell PATH when the app's env probe answers late (#471)
The desktop app killed its `$SHELL -ilc` probe after 5s and kept launchd's minimal PATH for the whole session, so local runs could not find tools like `uv`. Startup still waits 5s, but the probe now keeps running (capped at 60s) and a late answer supplies PATH only; the directory variables stay inherited because storage is already using them. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f9b0579308 |
Spawn the replaced orx binary, not its "(deleted)" path (#470)
On Linux, a long-running `orx up` whose binary was replaced on disk sees current_exe() as `<path> (deleted)`, so the supervisor spawn failed with os error 2 and hook/MCP configs recorded a path that no longer exists. Move updates.rs's relaunch_target into paths::spawnable_exe and use it at every daemon-executed site that spawns or writes the orx path. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b421d24e41 |
Spawn supervisors from the replaced orx binary (OR-334) (#469)
On Linux, a long-running `orx up` whose binary was replaced on disk sees current_exe() as `<path> (deleted)`, so spawning `orx supervise` for chat-launched runs failed with os error 2 and left runs in `starting`. Strip the marker with the existing relaunch_target helper. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
96e444b08e |
Link the README's downloads to the Windows installer and Linux AppImages (#460)
* Link the README's downloads to the Windows installer and Linux AppImages The Windows button now downloads OpenResearch-Setup.exe instead of the CLI zip, and Get started links the x86_64 and ARM64 AppImages, with the SmartScreen prompt and the Linux requirements alongside. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Apply suggestion from @greptile-apps[bot] Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> * Make the README's Linux button a Download for Linux link Match the macOS and Windows buttons: a left-aligned "Download for Linux" image that downloads the x86_64 AppImage, with ARM64 linked from Get started. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com> |
||
|
|
882f13c338 |
chore: bump version to 0.2.12 (#459)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
d52635a980 |
Open file:line citations at the cited line (#456)
* Land file:line citations on the cited line Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Hold line jumps for stale editor buffers; keep schemes out of citations Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Reseed clean buffers that lag the loaded version Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test the citation sanitizer carve-out; reseed on path changes Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rebuild ui/dist Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Cite extensionless files; keep cited hrefs out of live links Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Unit-test the citation sanitizer Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Rebuild ui/dist Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
89f28490aa |
Make the Linux app's single-instance claim race-free and private (#455)
* Make the Linux app's single-instance claim race-free and private Take a lock file before binding the focus socket, so two launches at once can't both become the app, and fall back to a per-user, per-host directory instead of the shared temp dir when XDG_RUNTIME_DIR is unset. Clear an interrupted update's staged AppImage before staging another. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Trust only a private runtime dir, and let a waiting launch take a freed lock Use XDG_RUNTIME_DIR only when the user owns it and no one else can write it, retry the lock alongside the socket so a launch waiting on an app that gave the lock up takes over, say when no private directory exists, and test the handoff in a scratch directory. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c95be22a35 |
OR-326 Add the Linux desktop app as a self-updating AppImage (#445)
* Open the macOS app's dashboard in its own window The OpenResearch.app now hosts the dashboard in a native window (tao + wry WKWebView) instead of handing it to the user's browser. It adds a standard menu bar, save panels for downloads, a native window.confirm panel, and a Cmd+Q that flushes workspace state and shuts the server down cleanly. Pop-ups open in the system browser (http/https/mailto only). The app now prefers port 4792 so the window's localStorage survives relaunches. The Dock-click tab-focus script, its Apple-events entitlement, and the SSE client counter it relied on are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Open the dashboard in its own window on Windows Adds the Windows desktop app. A small GUI-subsystem OpenResearch.exe (windows/launcher) starts the orx.exe beside it as `orx app` in a hidden console, so orx and the git, shell, and agent processes it runs share one invisible console instead of each flashing a window. `orx app` shares the macOS window code in src/commands/app.rs: WebView2 window, pop-ups to the system browser, save panel, page-load reveal, and port 4792. On Windows, closing the window quits, a second launch focuses the running window (named mutex + event), and the taskbar groups the window with the Start menu shortcut. An update restart relaunches as the app on the same port. Quit on both platforms now goes through up::request_shutdown instead of a self-sent SIGTERM. An Inno Setup script builds a per-user OpenResearch-Setup.exe with a Start menu entry and a WebView2 bootstrap. CI builds it as an artifact, and release-windows-app.yml attaches it to releases once WINDOWS_APP_ENABLED is set. The icon is embedded via embed-resource. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the Windows build and address review of the Windows app - The focus listener captured its bare HANDLE (edition 2021 disjoint capture) instead of the Send wrapper, which failed to compile on Windows. - An orx.exe away from the CLI installer's prefix, like the app's, is now Portable even when the installer's receipt exists, so it can update itself. - Single instance creates its event before the mutex; the launcher hands its foreground rights to orx.exe. - Focus requests and server readiness are ignored while quitting, and focus waits for the first load. The save dialog is owned by the window. - Telemetry counts an app start only after the instance claim. - The installer reports a failed WebView2 bootstrap and shows progress. - Icon embedding now fails the build if no resource compiler is found. - CI format-checks the launcher; the release job drops an unpinned action. - Docs: maintainer notes for the release gate, two known gaps, and wording. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tighten the Windows app after a second review - Keep the macOS save panel free-floating: only Windows gets the window as its owner, since a parent makes rfd show a sheet on macOS. - Treat the WebView2 bootstrap as failed unless the runtime is then present. - Move the UTF-16 helper out of the single-instance module, and bind the kernel object names before the mutex call that GetLastError follows. - Docs and a dead_code reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Add the Linux desktop app as a self-updating AppImage A `desktop` cargo feature builds the windowed app on Linux (glibc, WebKitGTK), leaving the static musl CLI builds untouched. build.rs sets cfg(desktop_app) for macOS, Windows, and Linux with the feature, replacing the macOS-or-Windows gates. The AppImage's AppRun starts `orx app`, sharing the Windows entry point: the webview is built into the GTK window's box (X11 and Wayland), closing quits, a Unix socket in XDG_RUNTIME_DIR keeps one instance and focuses it, and the login shell's PATH is adopted as on macOS. scripts/build-linux-appimage.sh bundles WebKitGTK with pinned linuxdeploy, its GTK plugin, and appimagetool, copying WebKit's helper processes and rewriting libwebkit2gtk's /usr paths to ././ so they resolve inside the image. CI builds x86_64 and aarch64 AppImages on Ubuntu 22.04 and smoke-tests each under Xvfb; release-linux-app.yml attaches them with linux-app.json once LINUX_APP_ENABLED is set. The AppImage updates itself (InstallChannel::AppImage, updates/linux_app.rs): a sha256-checked download renamed over the file, production builds only, and Restart execs the new AppImage on the same port. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the Linux app's downloads, bundled WebKit, and host environment From review of the Linux desktop app: - Downloads froze the app: rfd's GTK backend runs its dialog on a thread of its own, which waits forever on the main context tao holds. The save dialog is now a gtk::FileChooserDialog run on the main thread. - The bundled WebKit helpers found their libraries only beside themselves, so they loaded the host's WebKit or none. They now get an RPATH to the image's usr/lib, and linuxdeploy no longer makes stray copies of them. - The GTK hook's variables reached every host program orx opens. AppRun now saves the session's values and orx hands them back to xdg-open, the folder picker, error dialogs, and agents. - The smoke test now runs after WebKitGTK is removed from the runner and checks that each WebKit helper runs, and loads libwebkit2gtk, from the image; cleanup kills the app's whole process group. - Startup failures now reach stderr and a zenity or kdialog dialog. - The tools and the AppImage runtime are pinned by sha256, and only the release job that publishes can write to the release. - Smaller fixes: the shell probe runs after the single-instance claim and falls back to /bin/sh; APPDIR is canonicalized and must not be empty or root; relative project paths resolve against home inside the image. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep the session's GTK settings across Linux app restarts and terminals Restore the host GTK variables before an update relaunches the AppImage, so the new AppRun doesn't save the old image's values as the session's, and in the dashboard's terminals. Share one APPDIR containment check between the update channel, relative project paths, and the folder picker, which now keeps a terminal `orx up`'s working directory. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Relaunch the AppImage only from inside it, and tidy after review Pick the relaunch target with the same APPDIR check as the update channel, so an `orx app` started from the app's terminal restarts itself. Hand the shell probe the session's GTK settings, harden the smoke test's helper checks and cleanup, and bring the docs up to date. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Run the AppImage detection test only on Unix Its mount paths aren't absolute on Windows, which has no AppImage anyway. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Import the AppImage test's helper inside the Unix-only test Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Show the app window even when the dashboard never finishes loading On a UTM VM the Linux app ran with no window: it stays hidden until the page finishes loading, and that never happened. Show it 15 seconds after the server is up regardless, say so on stderr, and have the smoke test fail if that fallback fires or no window appears. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Give the Linux app its name and icon in the dock, and keep GIO off host modules GNOME shows an unmatched window as a gear named after its class, so name the program OpenResearch and have each launch install a desktop entry and icon pointing at the AppImage (TryExec hides it once the file is gone). Bundle GLib's TLS module and set GIO_MODULE_DIR to the image's, so the bundled GLib stops loading the host's modules, built against a newer one. A second launch now shows the window even before the page loads, so a stuck instance can't swallow it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tidy the Linux app after the fourth review Write the desktop entry only from the AppImage's own orx (shared with the update relaunch), keep a session's GIO_EXTRA_MODULES off the bundled GLib, mark a Dock-reopened macOS window as shown so the load fallback can't reopen it, close a race in the smoke test, and note the Fedora and openSUSE TLS gap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Show a stalled app window, and close the app before the installer runs If the dashboard never finishes loading, show the window 15 seconds after the server is up and say so on stderr; a relaunch (or a Dock click on macOS) also shows it before the page loads. The installer and uninstaller now check the app's single-instance mutex and ask the user to close it, rather than replacing files under a running app and skipping its shutdown path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep a replaced download's original until it lands, and don't offer a launch without WebView2 Move the file a download replaces aside and restore it if the download fails or is cancelled (WebView2's flyout can cancel), rather than deleting it up front. The installer no longer offers to start OpenResearch when the WebView2 Runtime is still missing, since the window could not open. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
2c63f548c2 |
OR-326 Open the dashboard in its own window on Windows (#439)
* Open the macOS app's dashboard in its own window The OpenResearch.app now hosts the dashboard in a native window (tao + wry WKWebView) instead of handing it to the user's browser. It adds a standard menu bar, save panels for downloads, a native window.confirm panel, and a Cmd+Q that flushes workspace state and shuts the server down cleanly. Pop-ups open in the system browser (http/https/mailto only). The app now prefers port 4792 so the window's localStorage survives relaunches. The Dock-click tab-focus script, its Apple-events entitlement, and the SSE client counter it relied on are removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Open the dashboard in its own window on Windows Adds the Windows desktop app. A small GUI-subsystem OpenResearch.exe (windows/launcher) starts the orx.exe beside it as `orx app` in a hidden console, so orx and the git, shell, and agent processes it runs share one invisible console instead of each flashing a window. `orx app` shares the macOS window code in src/commands/app.rs: WebView2 window, pop-ups to the system browser, save panel, page-load reveal, and port 4792. On Windows, closing the window quits, a second launch focuses the running window (named mutex + event), and the taskbar groups the window with the Start menu shortcut. An update restart relaunches as the app on the same port. Quit on both platforms now goes through up::request_shutdown instead of a self-sent SIGTERM. An Inno Setup script builds a per-user OpenResearch-Setup.exe with a Start menu entry and a WebView2 bootstrap. CI builds it as an artifact, and release-windows-app.yml attaches it to releases once WINDOWS_APP_ENABLED is set. The icon is embedded via embed-resource. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Fix the Windows build and address review of the Windows app - The focus listener captured its bare HANDLE (edition 2021 disjoint capture) instead of the Send wrapper, which failed to compile on Windows. - An orx.exe away from the CLI installer's prefix, like the app's, is now Portable even when the installer's receipt exists, so it can update itself. - Single instance creates its event before the mutex; the launcher hands its foreground rights to orx.exe. - Focus requests and server readiness are ignored while quitting, and focus waits for the first load. The save dialog is owned by the window. - Telemetry counts an app start only after the instance claim. - The installer reports a failed WebView2 bootstrap and shows progress. - Icon embedding now fails the build if no resource compiler is found. - CI format-checks the launcher; the release job drops an unpinned action. - Docs: maintainer notes for the release gate, two known gaps, and wording. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Tighten the Windows app after a second review - Keep the macOS save panel free-floating: only Windows gets the window as its owner, since a parent makes rfd show a sheet on macOS. - Treat the WebView2 bootstrap as failed unless the runtime is then present. - Move the UTF-16 helper out of the single-instance module, and bind the kernel object names before the mutex call that GetLastError follows. - Docs and a dead_code reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Show a stalled app window, and close the app before the installer runs If the dashboard never finishes loading, show the window 15 seconds after the server is up and say so on stderr; a relaunch (or a Dock click on macOS) also shows it before the page loads. The installer and uninstaller now check the app's single-instance mutex and ask the user to close it, rather than replacing files under a running app and skipping its shutdown path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Keep a replaced download's original until it lands, and don't offer a launch without WebView2 Move the file a download replaces aside and restore it if the download fails or is cancelled (WebView2's flyout can cancel), rather than deleting it up front. The installer no longer offers to start OpenResearch when the WebView2 Runtime is still missing, since the window could not open. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
b05312eb8c |
Even out the DMG installer window's vertical padding (#454)
Finder's window bounds include the 28pt title bar, so the 640x320 frame left a 292pt content area: the background's bottom was cropped and the labels sat closer to the bottom edge than the icons to the top. Size the frame for a full 320pt content area, and park dot-folders off-window so Finder builds that show hidden files don't shift the icons down. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
7159b1f134 |
Open the macOS app's dashboard in its own window (#438)
The OpenResearch.app now hosts the dashboard in a native window (tao + wry WKWebView) instead of handing it to the user's browser. It adds a standard menu bar, save panels for downloads, a native window.confirm panel, and a Cmd+Q that flushes workspace state and shuts the server down cleanly. Pop-ups open in the system browser (http/https/mailto only). The app now prefers port 4792 so the window's localStorage survives relaunches. The Dock-click tab-focus script, its Apple-events entitlement, and the SSE client counter it relied on are removed. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
f2c6a315c7 |
Move media preview controls into the file header (#446)
PDF page count, zoom, find, and a new download button now sit in the file
header row instead of a separate toolbar, and image, audio, and video
previews get a header download button in place of the bottom
"Download {name}" strip. Hosts expose a MediaToolbarSlot that previews
portal into; lower-priority controls hide in narrow headers.
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
|
||
|
|
5048460120 |
Render PDF previews with PDF.js instead of the browser's viewer (#444)
The dashboard's PDF preview used <object type="application/pdf">, which depends on the webview having a PDF viewer. WebKitGTK, which the Linux desktop app will use, has none. Previews now render with a lazily loaded PDF.js viewer on every platform: page count, zoom, fit to width, and in-document search, with links routed like the rest of the dashboard. - Uses the legacy PDF.js build, whose polyfills cover 2024-era Safari, WebKitGTK, and Chromium; older engines fall back to the download link. - Adds a ReadableStream async-iterator polyfill, without which PDF.js's text extraction (and so search) fails silently in WebKit. - Ships PDF.js's cmaps, standard fonts, and wasm decoders under a versioned /pdfjs/<version>/ path, since the server caches unhashed assets as immutable. - Scopes the light theme's color-scheme to :root[data-theme="light"] so PDF.js's stylesheet can't override it. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
3037f727e8 |
Track the dashboard locale in usage analytics (#437)
The dashboard reports its Paraglide locale to orx on load and on change; orx persists it in settings.json and sends it as context.locale on every analytics event. Declined consent events omit it. Requires the openresearch.sh ingest contract to accept context.locale before release. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
de469a3f44 |
OR-319 Sign and notarize the macOS CLI binaries in releases (#427)
* Sign and notarize the macOS CLI binaries in releases The install.sh archives for macOS shipped unsigned, so device-management policies on work Macs blocked them. A release-signing-gated job now signs and notarizes each darwin archive between dist's local and global builds, rewriting its checksums so the installers and sha256.sum match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Harden CLI signing from review Keep the called workflow from reporting skipped on dry runs, narrow its token to read, pin the Developer ID requirement the updater checks, keep notarization logs on failure, and correct the allow-dirty docs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
c81d46ca99 |
Point managed Macs at the signed app in the README (#426)
Work computers block the unsigned CLI that install.sh downloads, while the DMG app is signed and notarized. Tell those users to install the app and link its orx into their terminal. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
13d7c1fd44 |
Require fork PRs to link an issue (#425)
Add a `linked issue` check that fails PRs from forks unless the description links an issue in this repository. PRs from branches in the repo are exempt. It runs on pull_request_target so a fork cannot edit its own check, and never checks out PR code. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
8de34d2dc7 |
Show a visible hint while starter prompts load (#416)
The four pulsing skeleton cards gave no clue that orx was reading the project, so they looked broken. Surface the existing localized "Reading the project to suggest where to start…" message with a spinner below the skeletons, and make that row the status region instead of an aria-label on the decorative grid. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
678f477982 |
chore: bump version to 0.2.9 (#420)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
22496add7c |
Add PubMed as a literature source (#419)
`orx discover pubmed` searches PubMed through NCBI E-utilities (esearch for Best Match PMIDs, efetch XML for records), and `orx paper` reads a PMID, `pmid:` id, or PubMed URL. PubMed joins the Settings/composer source toggles, chat activity rendering, and the orx-lit-review skill alongside alphaXiv, OpenAlex, and bioRxiv. - Date windows map to `datetype=pdat` with both bounds filled; recency and historical rerank a wider relevance pool like OpenAlex. No citation counts, so `popular` keeps relevance order. - Book records (StatPearls, GeneReviews) are parsed alongside journal articles. - 429s are retried per NCBI's Retry-After with jitter, since keyless E-utilities allows 3 req/s and agents issue parallel calls. - Adds roxmltree because efetch serves abstracts only as XML. - The PubMed logo is a placeholder mark. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
9291fb95e6 |
Let /resume adopt chats from the agents' own CLIs (#413)
* Add composer commands that work the same on every coding agent The composer's `/` menu now offers /new (alias /clear), /resume, /model, /plan, /copy and /export on every harness. The dashboard runs each one itself, so behavior no longer depends on what the selected agent's own CLI supports — none of them act on slash text in the mode orx runs them in except Claude Code, and only partially. `planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the chips and the send path all read from. A skill whose bare name collides with a command is dropped from the menu: the composer inserts, and the server resolves, skills by bare name, so it would otherwise shadow. Only /plan composes with a prompt. The rest run as a whole message, so prose that merely mentions `/export` still reaches the agent, and a command is chipped only where it would actually run. /resume lists chats across every project: GET /api/chat/sessions now takes `scope=all` (400 without it or a projectId) and returns the 500 newest. Picking a chat in another project switches to it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Add /compact, which every coding agent can run The composer's /compact (alias /summarize) shortens a chat's context on every harness. Claude Code and Codex compact their own context in place and OpenCode summarizes through its endpoint, so the session survives; Cursor and the legacy `codex exec` path have nothing to call, so orx summarizes the transcript itself and reseeds a fresh native session with it through `bootstrap_context`, which is injected exactly while no native session exists. An agent that is merely not running is never traded for a summary: a resumable session reports "send a message first" instead, since a summary keeps only the newest 32 KiB of the chat. Compaction runs as a turn. It holds the session's turn slot, reports itself as a transcript row that shimmers while it works, and every write is gated on still owning that slot — Stop cancels it, and a cancelled row settles as failed rather than sitting at running. The row is the progress: it survives a reload, reaches every open client, and carries the reason when it fails. Codex waits for the compaction turn it started rather than the RPC ack, filtering by that turn's own id; Claude retires its child if its /compact fails, so an abandoned result cannot land in the next turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Add /goal, a standing objective every turn is reminded of `/goal <text>` gives the chat an objective, `/goal clear` drops it, and a bare `/goal` reports it. A pill beside Plan shows an active goal and clears it on click, so it is never invisible state. The goal is orx's own, stored on the session and re-injected into every turn. Claude Code's and Codex's native goals are deliberately unused: theirs mean different things (work until a condition is met, versus a long-lived objective), and one command cannot mean three things. Being orx-owned also makes it outlive compaction, a lost native session and a resume — the agent has to be told its goal every turn regardless. Goal is the second command that reads the rest of the message, so it is anchored: it must lead, and it is matched before Plan's unanchored token so a `/plan` mentioned inside a goal stays part of the goal. An empty composer gets a session first rather than losing the goal typed into it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Let /resume adopt chats from the agents' own CLIs (#410) The picker now lists, under "From your terminal", the chats the user had in Claude Code or Codex outside orx. Adopting one binds an orx session to the agent's own id and backfills a readable copy of the transcript, so it does not open blank; the agent itself resumes from its own session, which is why a marker row says new turns continue in that history. orx could already resume such a chat — `native_store` resolves an id in the user's real agent home and the turn runs there. What was missing was finding them: Claude keeps one JSONL per session (named by its own title records, else by the first prompt), Codex a SQLite index read `mode=ro` so its newest threads are not hidden behind the WAL. Every backfilled row records the native session, so a rewind or a branch switch resumes the chat rather than silently unbinding it. Adoption is idempotent under one immediate transaction: two sessions driving one native chat would split its history. orx's own throwaway children are filtered out by comparing their working directory against the temp directory canonically — macOS reports one as `/private/var`, the other as `/var`, and without that 39 auto-titler runs sat at the top of the picker. Cursor and OpenCode are deliberately left out: Cursor keys chats by a hash of the working directory and stores bodies in an encrypted blob, and OpenCode's live database would have to be copied rather than read. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
4cb5b90b25 |
Add /goal, a standing objective every turn is reminded of (#412)
* Add composer commands that work the same on every coding agent The composer's `/` menu now offers /new (alias /clear), /resume, /model, /plan, /copy and /export on every harness. The dashboard runs each one itself, so behavior no longer depends on what the selected agent's own CLI supports — none of them act on slash text in the mode orx runs them in except Claude Code, and only partially. `planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the chips and the send path all read from. A skill whose bare name collides with a command is dropped from the menu: the composer inserts, and the server resolves, skills by bare name, so it would otherwise shadow. Only /plan composes with a prompt. The rest run as a whole message, so prose that merely mentions `/export` still reaches the agent, and a command is chipped only where it would actually run. /resume lists chats across every project: GET /api/chat/sessions now takes `scope=all` (400 without it or a projectId) and returns the 500 newest. Picking a chat in another project switches to it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Add /compact, which every coding agent can run The composer's /compact (alias /summarize) shortens a chat's context on every harness. Claude Code and Codex compact their own context in place and OpenCode summarizes through its endpoint, so the session survives; Cursor and the legacy `codex exec` path have nothing to call, so orx summarizes the transcript itself and reseeds a fresh native session with it through `bootstrap_context`, which is injected exactly while no native session exists. An agent that is merely not running is never traded for a summary: a resumable session reports "send a message first" instead, since a summary keeps only the newest 32 KiB of the chat. Compaction runs as a turn. It holds the session's turn slot, reports itself as a transcript row that shimmers while it works, and every write is gated on still owning that slot — Stop cancels it, and a cancelled row settles as failed rather than sitting at running. The row is the progress: it survives a reload, reaches every open client, and carries the reason when it fails. Codex waits for the compaction turn it started rather than the RPC ack, filtering by that turn's own id; Claude retires its child if its /compact fails, so an abandoned result cannot land in the next turn. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Add /goal, a standing objective every turn is reminded of (#385) `/goal <text>` gives the chat an objective, `/goal clear` drops it, and a bare `/goal` reports it. A pill beside Plan shows an active goal and clears it on click, so it is never invisible state. The goal is orx's own, stored on the session and re-injected into every turn. Claude Code's and Codex's native goals are deliberately unused: theirs mean different things (work until a condition is met, versus a long-lived objective), and one command cannot mean three things. Being orx-owned also makes it outlive compaction, a lost native session and a resume — the agent has to be told its goal every turn regardless. Goal is the second command that reads the rest of the message, so it is anchored: it must lead, and it is matched before Plan's unanchored token so a `/plan` mentioned inside a goal stays part of the goal. An empty composer gets a session first rather than losing the goal typed into it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a15b3a61a4 |
Add /compact, which every coding agent can run (#411)
* Add composer commands that work the same on every coding agent The composer's `/` menu now offers /new (alias /clear), /resume, /model, /plan, /copy and /export on every harness. The dashboard runs each one itself, so behavior no longer depends on what the selected agent's own CLI supports — none of them act on slash text in the mode orx runs them in except Claude Code, and only partially. `planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the chips and the send path all read from. A skill whose bare name collides with a command is dropped from the menu: the composer inserts, and the server resolves, skills by bare name, so it would otherwise shadow. Only /plan composes with a prompt. The rest run as a whole message, so prose that merely mentions `/export` still reaches the agent, and a command is chipped only where it would actually run. /resume lists chats across every project: GET /api/chat/sessions now takes `scope=all` (400 without it or a projectId) and returns the 500 newest. Picking a chat in another project switches to it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Add /compact, which every coding agent can run (#384) The composer's /compact (alias /summarize) shortens a chat's context on every harness. Claude Code and Codex compact their own context in place and OpenCode summarizes through its endpoint, so the session survives; Cursor and the legacy `codex exec` path have nothing to call, so orx summarizes the transcript itself and reseeds a fresh native session with it through `bootstrap_context`, which is injected exactly while no native session exists. An agent that is merely not running is never traded for a summary: a resumable session reports "send a message first" instead, since a summary keeps only the newest 32 KiB of the chat. Compaction runs as a turn. It holds the session's turn slot, reports itself as a transcript row that shimmers while it works, and every write is gated on still owning that slot — Stop cancels it, and a cancelled row settles as failed rather than sitting at running. The row is the progress: it survives a reload, reaches every open client, and carries the reason when it fails. Codex waits for the compaction turn it started rather than the RPC ack, filtering by that turn's own id; Claude retires its child if its /compact fails, so an abandoned result cannot land in the next turn. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1a8c6ad40e |
Add composer commands that work the same on every coding agent (#376)
The composer's `/` menu now offers /new (alias /clear), /resume, /model, /plan, /copy and /export on every harness. The dashboard runs each one itself, so behavior no longer depends on what the selected agent's own CLI supports — none of them act on slash text in the mode orx runs them in except Claude Code, and only partially. `planCommand.ts` becomes `composerCommands.ts`, a registry the menu, the chips and the send path all read from. A skill whose bare name collides with a command is dropped from the menu: the composer inserts, and the server resolves, skills by bare name, so it would otherwise shadow. Only /plan composes with a prompt. The rest run as a whole message, so prose that merely mentions `/export` still reaches the agent, and a command is chipped only where it would actually run. /resume lists chats across every project: GET /api/chat/sessions now takes `scope=all` (400 without it or a projectId) and returns the 500 newest. Picking a chat in another project switches to it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6af7c10273 |
Add a workspace Terminal tab with an interactive shell (#360)
The Workspace tools gain a Terminal button that opens an interactive shell
tab in the right panel. A new /api/projects/{id}/terminal WebSocket spawns
the user's login shell in the session worktree (or the project clone) over
the existing PTY relay. The tab is a persisted home view that stays mounted
behind other tabs so the shell survives tab switches, spawns lazily on first
selection, and can be restarted after it exits.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
8cc8fe313f |
OR-217 Run coding-agent sign-in and setup commands in an embedded settings terminal (#361)
* Run coding-agent sign-in and setup commands in an embedded settings terminal Settings notes that mention a terminal command now carry a play button that runs it in place: harness sign-ins (claude auth login, codex login, agent login, opencode auth login), gh auth login, hf auth login, and the Cursor installer. The terminal follows the app theme, stays open after the command exits, and continues as the user's login shell for follow-up commands. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address review: keep the settings terminal alive, size and env for the follow-up shell - Hoist the run state out of the note so a successful sign-in (which removes the note) no longer unmounts the terminal mid-session. - PTY children get the imported shell PATH/config env and a TERM. - The follow-up shell opens at the client's last reported size and only after the command actually started. - Light-theme ANSI palette for the app-themed terminal. - Non-zero exits under the shell no longer raise a toast. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address round-2 review: report server failures, share the shell hand-off, drop opencode models - A command failure under the follow-up shell now writes the message to the terminal and marks the header failed; only transport failures toast. - One continue_in_shell helper serves both routes and carries OpenCode's isolated store into the shell after opencode auth login. - opencode models is no longer runnable: it would run in a different configuration than detection and mislead. - Drop the dead on_success parameter; origin check first everywhere; setup route uses the shared origin guard; tests for resize tracking and the shell query flag. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address round-3 review nits Keep the success state after the shell exits, write a late shell-start error to the terminal, report sessionEnded accurately for non-shell terminals, and fix comment anchors and wording. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
0ceb9606e4 |
Show pre-written starter prompts for blank projects (#349)
A blank project has no paper and no files, so a model call over its brief could only produce generic prompts. The backend now flags such projects as blank and skips the harness one-shot on creation, on the form's pre-warm, and on the empty-chat fetch; the UI fills the four boxes from localized, pre-written prompts instead. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b38c8136e6 |
OR-280 Nudge new installs to run the demo experiment (#338)
* Nudge new installs to run the demo experiment Analytics for Sep 9-12 installs showed the largest funnel leak between "onboarding complete" and "first message sent" (43.8%). The demo welcome steered users to read the recorded sessions rather than send anything. - Rename the welcome's primary button to "Run the demo experiment" and hand focus to the composer, which already holds the run-experiment prompt. - Keep that prefill on offer until the user has sent anything in the demo, detected by a recorded session's active leaf moving off its seeded message or a new session appearing, instead of only until the welcome closes. - Show a dismissible hint above the composer explaining the 200-step probe, with a ring on the send button while the prefilled prompt is untouched. - Open only the experiments tab on the first demo visit. - Default the right pane to a quarter of the window (max 480px) on installs with no saved layout, and share that default with the projects home so it is not overwritten before the demo opens. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Refine the demo nudge and add a running-experiment hint - Relabel the welcome button "Try the demo experiment": the click only focuses the prefilled prompt, so "Run" over-promised. - Style the composer hint as a neutral surface bubble instead of a tint. - When a demo run starts, show a second dismissible hint above the composer with a Logs button that opens that run's logs; it clears when the run leaves the running state. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Address review on the running-experiment hint - Seed the running-run hint from the runs baseline so it survives a reload or a return to the demo project mid-run. - Announce the hint to assistive tech via a polite live region and return focus to the composer on dismiss, matching the prompt hint. - Share the hint bubble classes, drop the redundant Logs tooltip key, and use the idiomatic Arabic phrasing for the tap instruction. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
72c781311f |
chore: bump version to 0.2.2 (#335)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
4a646d9b4c |
Add self-update on Windows (#334)
Windows installs could never update themselves: the PowerShell installer's receipt lives under %LOCALAPPDATA%, which orx did not look in, and the installer cannot overwrite a running exe. orx now stages the release via the installer beside orx.exe and swaps it in with two renames, parks the old binary under a unique .old name, and rewrites the receipt version itself. A zip-extracted orx.exe outside package-manager and cargo target paths is a new "portable" channel that updates in place with no receipt. `orx up` relaunches on Windows by spawning the new binary, which waits on the parent's process handle before binding the port. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
35417ef1b8 |
chore: bump version to 0.1.123 (#310)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9b05f9771c |
OR-199 Windows support for the CLI and dashboard (#299)
* Make orx build and resolve its tools on Windows Block A of the Windows beta: the fixes that need no Windows machine to write, so the branch is ready to clone the moment one is set up. - Extend a bare binary name with PATHEXT when searching PATH. Every harness is looked up by bare name, so on Windows the picker reported Claude Code, Codex and OpenCode all absent and the dashboard could do nothing. - Compose a child's PATH with `;` rather than a hardcoded `:`. - Find bash through `git` rather than PATH. The `bash.exe` usually on PATH is the WSL launcher in System32, which sees a different filesystem than the run directory it would be handed. - Give git an empty path it understands. `/dev/null` reads as a missing relative path on Windows, silently re-enabling repository hooks. - Check a local run's liveness with a zero-timeout wait; there is no `ps` to shell out to. - Link through native_store::create_symlink, which already has both branches, instead of calling the unix symlink directly. - Add a windows-latest CI job so none of this regresses. The Windows branches here have not been compiled yet — that is the next step, on the machine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Allow CI to be dispatched against a branch * Compile the remaining unix-only code paths out on Windows Round one of the Windows compile, from the first CI run: 25 diagnostics, seven root causes. - Give remote_host the `#[cfg(not(unix))]` twins its uid/mode helpers were missing, matching what `directory_writable` already does. Windows keeps only the symlink and directory checks in `ensure_private_dir`, which is a weaker guarantee than the unix path makes. - Gate the control channel itself. It is a Unix domain socket, so `--remote-host` is now refused up front on Windows rather than starting a host nothing could attach to. Local `orx up` is untouched. - Annotate `signal`, whose type came only from the cfg(unix) block. - Gate the EXDEV rename fallback, the inode and mode assertions, and the shell-hook tests, all of which assert on things Windows does not have. - Silence the parameters that are only read on unix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Gate the remote-host command surface off on Windows Round two, from the cascade the unresolved socket import had been suppressing: every caller of the six control-channel functions. The module splits cleanly in two. `orx up` uses its types and its dashboard locking — HostDescriptor, DashboardLock, canonical_data_dir — and none of that needs a socket. The `orx remote-host` command surface is socket-backed all the way down, so it is gated as a unit and the command now refuses on Windows instead of failing somewhere further in. Also: the two imports only the gated half reached, and `signal`, which loses its one assignment on Windows and with it the need for `mut`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Keep the compiled-out remote-host half rather than gate it item by item Round three was 24 dead-code and unused-import lints and no real compile errors — the tail of removing a feature's entry points while its machinery stays behind. One scoped allow on the module says that plainly. Gating each constant, helper and request type would be bulkier than the code it guards, would have to be unpicked the moment the control channel gets a Windows transport, and would say nothing the module comment does not. updates.rs is the same shape at smaller scale: relaunch_target and relaunch_args keep their tests on every platform and lose only their caller in the binary. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Stop handing git a verbatim path, and let test cleanup fail The Windows suite ran for the first time: 662 passed, 46 failed. Two causes account for 27 of them. `repository_git_dir` canonicalizes, and Windows answers canonicalize with a `\\?\C:\…` verbatim path. Git rejects those outright — "not a git repository" — so importing any existing repository failed. This is a real bug, not a test artifact; the prefix now comes off for drive paths, which are the only ones with a plain form. The other 21 were `remove_dir_all(root).unwrap()` in test teardown, which Windows refuses while any handle into the tree is still open. The codebase already prefers a tolerant cleanup in 146 other places; these 86 now match. All of them are inside `#[cfg(test)]`. Also: the probe fixtures asserted on POSIX PATHs, which are not absolute on Windows and so were rejected by the parser under test. They are parameterized rather than gated — the parsing they cover is not platform-specific, only the fixture was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Check out LF on every platform Git for Windows defaults to core.autocrlf=true, and SYSTEM_PROMPT.md and the agent-skills SKILL.md files are include_str!'d at build time. A Windows build therefore embedded CRLF and then failed to parse it: playbook_md() splits the template on `-->\n\n` to drop its leading HTML comment, that split stopped matching, and `unwrap_or` kept the comment — so every agent session would have been handed OpenResearch's internal notes as the opening of its system prompt. Two tests caught it, one of them by scanning the rendered playbook for unresolved `{token}` placeholders rather than a fixed list. That is the only reason this surfaced as a test failure and not as a shipped binary. Nothing renormalizes: no tracked text file has CRLF in the index today, so this only governs what lands in a working tree. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix what Windows found: exec bits, remote paths, cancellation Round two of the Windows suite: 691 passed, 17 failed. Four of those were real bugs rather than test artifacts. Demo onboarding rejected the demo it had just built. NTFS has no executable bit, so runcpu.sh committed as 100644, the tree hash moved, and the commit ids drifted off the BASELINE_SHA/EXPERIMENT_SHA the demo validates itself against. The mode now reaches the tree through the index, where it does not depend on the filesystem. Remote paths were validated with local rules. `storage_root` asked `Path::is_absolute` about `/home/me/.cargo/bin/orx`, which is false under Windows, so a Windows client rejected every legitimate remote path — and the PathBuf it built would have reached a Linux shell separated by backslashes. Remote paths are POSIX whatever the client runs, so they are handled as strings now. `manage_local_file` took `reports/result.md` from the dashboard and answered `reports\summary.md`, which the next request would have carried straight back. Relative paths cross the API `/`-separated. Cancelling a local run did nothing: there is no process group to TERM. `taskkill /T` takes the tree instead, which is what TERMing the group was for. The rest were fixtures asserting on POSIX shapes, a filename using a character Windows reserves, and two LaTeX tests that need a symlink and a shebang. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Let the index own the demo's exec bit Setting the mode through `update-index` fixed the commit ids but left the index at 100755 against a working tree NTFS always reports as 0644. With `core.filemode true` git believed the filesystem, read runcpu.sh as modified, and refused to check main back out. Windows trusts the index instead, which is where the mode came from. The two telemetry failures are not diagnosed. HTTP delivery itself works on Windows — `sender_posts_the_first_party_endpoint` asserts an Acknowledged post and passes — so what fails is the removal that follows it, and reading has not told me why. Both assertions now report what a second removal says, which separates a blocked delete from a delete that succeeded while the name lingered. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Keep the demo's origin path native, and judge cancel by liveness Three down to two. `join("demo-repos/nanochat.git")` embeds a separator the platform did not choose, so on Windows the bare origin came out as `...\data\demo-repos/nanochat.git` — mixed, and textually unequal to what git echoes back for the remote it was given. Filesystem calls did not care; comparing the recorded URL to git's answer did. Split into two joins, in product code as well as the tests that caught it. `taskkill /T` returns non-zero when any descendant has already exited on its own, which says nothing about whether the run stopped. Cancellation now reads the leader's liveness, which is the question actually being asked. The telemetry pair is one step from diagnosed. The retry in the previous round answered Ok(()), so the file was there and freely deletable and the removal was never reached — delivery must have come back Retryable. The assertion now names the outcome, which separates a request that never landed from a response classified unexpectedly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Read the request before answering it in the telemetry mock servers The last two Windows failures, and they were the tests' fault rather than delivery's. Two of the three mock servers accepted a connection and wrote their response without ever reading the request. Closing a socket that still holds unread data aborts the connection on Windows instead of closing it gracefully, so the client saw a reset rather than the reply it had already been sent, and classified a 400 as Retryable — which left the outbox file the assertion expected to be gone. The third server always read, which is why it passed and misled the search toward the removal. The read loop it already had is now shared by all three. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Canonicalize into paths a child process can actually use An audit of the `orx up` paths no test exercises. Every bug this branch has fixed so far was in code a test happened to touch; the dashboard's own request handling was the largest area where that was not true. Windows `canonicalize` answers with a verbatim `\\?\C:\…` path, and `CreateProcessW` will not take one as a working directory. The dashboard canonicalizes a session's checkout root and then runs git in it — twelve call sites behind `resolve_checkout_root`, covering the code tree, file reads and writes, and diffs — so none of that could have worked. The same root reaches `run_shell_command` as the cwd for a `!` command, and Codex receives canonicalized paths as its sandbox writable roots. Consistency matters as much as the spelling: containment checks compare a canonicalized child against a canonicalized root, and a mix of the two forms denies access to paths that are genuinely inside. So every canonicalization in the crate now goes through one helper rather than the handful that happened to be reachable from a failing test — that is what makes the property hold rather than the individual fixes. The earlier one-off in `repository_git_dir`, which the import failure found, is folded into it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Look for the harness binaries under the names they actually have Second audit pass, over the spawn path. `find_claude` and `find_opencode` fall back to their installers' drop locations when PATH does not answer, and both spell them as bare names — `~/.local/bin/claude`, `~/.opencode/bin/opencode` — which match no file on Windows. `~/.local/bin` is exactly where Claude Code's own Windows installer puts claude.exe, so the fallback was dead precisely where it was needed. Both now look for the same names PATH does. The rest of the spawn path holds up: config directories travel as environment variables, where a raw path is correct; Codex's config.toml is written through serde_json, whose backslash escape is the one TOML wants; and the binaries themselves come from find_on_path, which learned about PATHEXT earlier on this branch. One gap left deliberately: the PATH guard writes zsh and bash startup hooks and needs ZDOTDIR or HOME to place them, neither of which Windows sets, so it returns early and the guard is simply absent there. It fails closed — a session gets no hook rather than a broken one — and building a Windows equivalent is worth more than a beta needs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round one Six blockers, all real. `orx_bin_dir` dropped any directory whose name contains a colon, which on Windows is every absolute path, so agents were never handed the running orx on PATH — and the separator constant added earlier sat in a branch that never ran. It filters on that constant now. A dashboard shell command that hit its timeout killed nothing on Windows and then waited on the child with no timeout, hanging the turn forever. The non-unix arm terminates the tree, the shape latex already used. BIBINPUTS/BSTINPUTS joined with a colon, so bibtex on Windows lost both the paper's own directory and the system defaults. `restore_local_repository` never pinned core.autocrlf, so a restored demo clone was written CRLF against LF blobs and then failed the cleanliness check that reads it with global config disabled — telling the user their demo "is not the OpenResearch nanochat demo". The remove_dir_all sweep had converted four setup removals along with the teardown, letting tests pass while testing nothing. The Windows uid/mode stubs were unreachable: their whole chain is rooted in cfg(unix) entry points. Gated instead, stubs deleted, and with them a comment claiming a weaker-but-working private runtime directory that does not exist. Also from review: prefix stripping now matches the parsed prefix, so a non-UTF-8 name cannot keep `\\?\` while its root loses it and fail its own containment check; UNC verbatim paths are decoded rather than passed through. PATHEXT set-but-empty no longer reports every tool missing. The bash fallback no longer resolves back to the WSL launcher it rejects. pdfPath goes through the same API path rule as every other relative path. The telemetry signature widened for a test message is reverted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round two Two blockers, both introduced by round one's own fixes. Deleting the Windows uid/mode stubs was wrong: `open_lock` is ungated and calls all three, and `DashboardLock::acquire` reaches it on every `orx up`. That is a Windows compile error on the one platform this branch exists for, and macOS cannot see it. The ownership and mode checks are gated inside `open_lock` now, which keeps the file-type check everywhere and states plainly what Windows does not enforce. Teaching `orx_bin_dir` about the platform separator also handed `ORX_BIN_DIR` to PATH_GUARD, a POSIX snippet that joins `PATH` on colons. Under Git Bash a drive letter split it in two and the guard re-prepended on every startup. The stamp is skipped off unix; `prepare_env` still fronts the directory, which was the point. Also: `cancel_job`'s cfg'd `return` is the function's last statement once stripped, which the Windows clippy job would have rejected; the restore path pins `core.eol` alongside autocrlf so a global `core.eol=crlf` cannot reproduce the same dirty checkout; the verbatim UNC head is built from `OsString` like the tail it is joined to; and the null-device comment no longer claims git falls back to the repository's own hooks, which it does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Polish from review round three No blockers in either round-three review. Taking the nits: `open_lock` fstats once instead of twice; the Windows branch of the `ORX_BIN_DIR` stamp is a `#[cfg]` rather than a runtime-false `cfg!`, so it stops resolving the executable only to discard it; two doc comments that described one platform now describe both; and the ownership comment names the case it does not cover — a lock under a shared `ORX_DATA_DIR` has only its ACL on Windows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Keep the windows job out of the release gate release.yml calls ci.yml and requires it to pass, so the windows job added here would have held up a macOS or Linux release over a platform those releases ship nothing for. `plan` is passed only by release.yml, so skipping on it draws the line where it belongs: Windows still gates pull requests, which is what keeps it from rotting, and never gates a release. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Upload a release orx.exe, not a debug one The artifact is what a tester is handed, and a debug build is slow enough to colour their impression of the dashboard rather than the code's. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Turn SSH multiplexing off explicitly on Windows Win32-OpenSSH implements no connection multiplexing: it accepts ControlMaster and ControlPath and then never creates a master. Omitting them is not enough, because the user's own ssh_config can still set a ControlPath, and one it honours fails the connection outright with "getsockname failed: Not a socket". Passing ControlMaster=no and ControlPath=none is what overrides that. The socket machinery goes with it — the control directory, its path hash and its preparation are unix's, and `master_is_running` can only answer false where no master is ever created. Nothing changes on unix, where the shared master still covers later status, log and job commands. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review Two tests asserted the unix option set unconditionally and would have failed on Windows — the very platform this change is for, and one CI never builds. Both derive the expectation from the platform now. `PathBuf`'s only uses became unix-gated, so its import would have failed the Windows clippy job under -D warnings. git shells out to ssh with its own command line, which never carried the override, so a user's ssh_config could still break git-over-SSH on Windows for exactly the reason this change exists. Routed through one helper. The rest: a doc comment my new test had come between and its rightful test; four comments still promising multiplexing unconditionally, the `interactive_args` one worst because it states the invariant that is false on Windows; the new test pinned to the exact vector the way its neighbour already is; and BatchMode coverage split out, since it holds on every platform and was about to be gated away with the ControlPath half. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round two No blockers in either round-two review. Taking the nits. Deriving the option count from `multiplexing_opts` had quietly removed the only assertion pinning the unix vector — ControlPersist could have been deleted and nothing in the tree would have failed. Pinned alongside the Windows one. The rest: a test whose name promised reuse on a platform that reuses nothing; `git_ssh_command` says which command it builds; two more doc comments of the class round one fixed, in the two other files that carry them; and the Windows sentence dropped from `ssh_opts`, where it was the third statement of the same fact within 150 lines. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Point git's config paths at an empty file, not the null device Onboarding died on Windows at the demo's first `git init`: fatal: unable to access 'NUL': Invalid argument `NUL` is right for `diff --no-index`, which recognizes it before touching the path, and wrong everywhere a path is actually read. Git skips a config file that reports ENOENT and fatals on any other error, and Windows cannot `access()` its null device by name at all. A real empty file behaves the same on every platform, and this file already had that idiom in `clone_public_repository`. Shared now, and pointing `hooksPath` at it still disables hooks, since git finds no hook inside a file. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Give the spawned bash its own toolchain on PATH A local run died at `run.sh: line 6: mkdir: command not found`. bash was running; it just had no coreutils. Git for Windows keeps those in `usr\bin` and puts only `cmd` on the Windows PATH, so a bash started from a Windows process inherits a PATH with no Unix tools at all. Fronted for the two bash spawns only — a local run's launcher and the composer's `!` command. Doing it for every child would shadow Windows' own `find` and `sort` with the MSYS ones, which is a worse bug than the one being fixed. prepare_env's PATH is now a function the bash spawn can extend, so the running orx still comes first and the toolchain sits ahead of both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fold the toolchain into the PATH run.sh exports The previous commit set the toolchain PATH on the spawn, which run.sh then overwrote on its second line: localrun exports the shell's PATH into the script so a user's own command finds python and uv, and that export wins over anything the process was started with. Caught by the agent inside the dashboard, which read the generated run.sh rather than the launcher and saw the hardcoded Windows PATH for what it was. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Let the demo probe run under Git Bash Two things differ there, and the probe hit both. uv ships no shell installer for Windows, so the setup uses its PowerShell one; and uv lays the venv out as `Scripts/` rather than `bin/`, so both activations are tried — either platform can have either layout depending on how the venv was made. These are Rust constants rather than files in the demo repository, so the commit ids the demo validates itself against do not move. `runs/runcpu.sh` carries the same POSIX assumption and is committed, so making the project's own run command portable is a separate change that regenerates those ids. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Stop writing a Windows PATH into a bash script The run script exported the process PATH so a user's command could find python and uv. On Windows that value is semicolon-separated, and bash splits PATH on `:` — so it collapsed at the first drive colon and nothing resolved, `mkdir` included. Adding the Git toolchain to it, as the last two commits did, only added more entries to the same unusable string. Git Bash converts the PATH it inherits into POSIX form for itself, so the fix is to leave it alone: the launcher sets the process PATH, and the script no longer overwrites it. The `cd` on the next line had the same shape — bash reads a quoted `C:\…` literally — so the run directory is forward-slashed now. Found by the agent running inside the dashboard, which read the generated run.sh, spotted the separator mismatch, and verified a candidate PATH against both splitters before proposing anything. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Hand the shell paths it can read, not Windows ones `tar: Cannot connect to C: resolve failed` — GNU tar reads a colon before the first slash as a `host:path` remote spec, so the snapshot archive's Windows path made it try to resolve `C:` as a hostname. Forward slashes would not have helped; the colon is what triggers it. One conversion now covers both places a local path reaches the shell: the archive tar unpacks, and the `cd` the launcher writes. `/c/Users/…` has no colon at all. The HF backend passes a Linux container path and is deliberately left alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Give runs a UTF-8 stdout A run died mid-training on `UnicodeEncodeError`: CPython encodes stdout in the console's codepage, cp1252 on a default Windows install, and the script printed a banner outside Latin-1. Nothing to do with the experiment — it had already trained a tokenizer by then. `PYTHONIOENCODING=utf-8` joins `PYTHONUNBUFFERED` as a default the author can override, through the same helper every backend already calls, and open-coded once more for kubernetes, whose env is a JSON array. Set on every platform rather than only Windows: a Linux container with `LANG=C` has the same ASCII default and would fail the same way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Let git write paths past Windows' 260-character cap Creating a session worktree for a deep repository failed partway through with "Filename too long". Windows refuses paths over 260 characters unless a program opts into the wide API, and a session worktree spends about a hundred of them on `worktrees/<project>/chat_<session>` before the repository's own paths begin. `core.longpaths=true` on the invocations that write a working tree: the shared helper every worktree and checkout goes through, the authenticated fetch, and the anonymous clone. Read-only plumbing does not need it. This was graded Minor in the original survey. It is a blocker for any repository with a deep tree, which is most of them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Link the agent's config without asking Windows for a privilege Windows grants symlinks only to an elevated process or one in Developer Mode, so preparing the isolated agent home failed onboarding outright with os error 1314. Fall back to the links that need no privilege: a junction for a directory, which reads back exactly like a symlink, and a hard link for a file, which does not — so the reconcile now recognizes one instead of filing a conflict against it on every launch. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Stop asking Windows' ssh to multiplex Windows' OpenSSH cannot create the AF_UNIX socket ControlMaster needs, and rather than decline the option it kills the connection outright with "getsockname failed: Not a socket" — so every ssh call orx made from Windows failed, whatever the destination. Drop the Control* options there and pay a handshake per call. Same option vector, same platform: `UserKnownHostsFile=/dev/null` is a name Windows' ssh takes literally, creating a `\dev\null` on the current drive. It now gets a scratch file of orx's own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Count the Windows ssh opts the Windows way Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Revert the duplicate multiplexing fix PR #302 already turns Windows' multiplexing off, and does it by setting ControlMaster=no explicitly rather than by omission — which is what overrides a ControlPath in the user's own ssh_config — and covers git's ssh besides. Take that one instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Give Windows a real file for the host keys it must forget `UserKnownHostsFile=/dev/null` is a name Windows' ssh takes literally, creating a `\dev\null` on whichever drive is current. Point it at a scratch file of orx's own instead; StrictHostKeyChecking=no is what makes a recycled provider key acceptable, so the behaviour is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Ship a Windows binary and a way to install it The demo's own run command reached for a POSIX uv installer and `.venv/bin/activate`, neither of which exists on Windows, so the nanochat demo could not run there; both now branch at runtime. That changes the experiment tree, hence the new EXPERIMENT_SHA. Releases built no Windows artifact at all, which is what the CI job's non-blocking rationale rested on. Add the target and a PowerShell installer so a beta tester has something to install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Tell Windows testers what they need before they need it Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Let Windows hold up a release now that a release ships it The job skipped release runs because releases carried no Windows artifact. They do now, so a regression on main would ship broken rather than fail loudly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Say what signing will and will not fix Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Say Git for Windows is missing at startup, not at the first run Without it there is no usable bash, so experiments and the demo die on their first command with a spawn error naming a path the user has never heard of. The agents check next to it already warns this way. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Make a second click reach the dashboard, not die on its port Double-clicking orx.exe runs `orx up`. Doing it while orx is already running failed the bind, printed an error into a console Explorer then closed, and looked to the user like nothing happened at all. Open the running dashboard instead, and hold the console open for whatever else a startup failure has to say. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Make a failed start say so where a console cannot The console Explorer opens for a double-clicked orx.exe dies with the process, so both ways this program exits with something to say — a returned error and a panic — reached a window that was already gone. Put the message in a dialog instead, and only when no terminal is attached to read the printed one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Start the dashboard when orx.exe is double-clicked A bare `orx` prints usage and exits, which is right at a prompt and useless from Explorer: the console it printed into closes with the process, so the gesture looked like nothing happening. Owning the console is what tells the two apart, and it now starts `up` — the same trade the macOS .app makes for the same gesture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Keep demo installs seeded before the Windows change working Moving EXPERIMENT_SHA left every existing mac and linux install pointing at a commit its repository does not contain. Each agent turn resolves the session worktree from that commit, so all three demo chats would have failed on their next message after upgrading. An install now starts its sessions from whichever known experiment its demo branch descends from. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round four - Release verify job reads the Windows .zip and orx.exe; without it no release would have published for any platform. - "Already running" reuse only for a double-click, and only of a dashboard speaking this build's protocol. On mac and linux a second `orx up` fails on the port again, as it did before this branch. - An old demo install whose cache was wiped but origin kept restores its own experiment rather than building a sibling it cannot push. - Same-file probe errors fall through to the old reconcile; the demo probe only puts ~/.local/bin on PATH after installing uv; the panic hook is Windows-only. - Docs name the assets cargo-dist actually publishes. - Comments cut to the one-or-two-line house rule; three doc comments moved back onto the functions they describe. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round five - A double-click keeps the flags it was given (`--no-telemetry` from a shortcut no longer turns telemetry back on). - A kept demo origin is validated before a clone is restored from it, so a foreign or moved origin is named as the problem instead of the cache. - The ssh opt-vector test reads the Windows known-hosts path once; it follows XDG_CONFIG_HOME, which telemetry tests mutate in parallel. - Every comment this branch adds is now one or two lines. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Address review round six - Remote storage paths accept an interior `.` again. `Path::components` skipped it on main, so `ORX_DATA_DIR=./orx` on a remote box — probed as `/home/u/./orx/orx.db` — broke `orx up --remote` on this branch. - A foreign kept demo origin is rejected before anything is restored from it, now with a test; `install_repository` keeps main's shape, since the restore helper creates its own parent. - Comments trimmed last round get back the why they lost, and the PYTHONIOENCODING one no longer claims the console codepage: redirected output uses the ANSI one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ea7fe8af8c |
chore: bump version to 0.1.121 (#298)
* fix: rebuild ui/dist for the telemetry UI changes in #297
|
||
|
|
66d2cd88fa |
feat: emit onboarding, demo, starter and first-action telemetry (#297)
Emits the four new product events the API now ingests: an onboarding step per screen, a demo experiment (curated conversation or real run), a starter-prompt click, and the first action taken on the demo or on a new project. Three of the four originate in the UI, which previously had no telemetry path beyond consent, so this adds POST /api/telemetry/event. Every field is matched against a fixed allowlist before it reaches capture(), so the local endpoint cannot emit arbitrary telemetry. first_action is claimed once per surface and persisted, so relaunching cannot promote a later action into the first slot. The claim happens only when telemetry is enabled — claiming first would burn the slot for a later opt-in. Starter clicks report slot position because the prompts are model-generated; the accepted range carries headroom so a fifth box cannot 400 a whole batch. Requires the matching API contract to be deployed first. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6ab3a7a44f |
Date the nanochat demo relative to onboarding (#296)
The bundled demo seeded every project, experiment, run, session, and message with fixed August-2026 epoch literals, so a new install opened on a demo whose history was dated weeks or months in the past. Compute the timestamps from now_ms() at seed time instead, laying the seeded history out over the ~4 hours before onboarding: the pipeline conversation reads ~4h old, the two analysis conversations ~1h and ~1h35, and the follow-up probes a few minutes old. The run keeps a duration that covers the wall-clock span of the log it ships with (10:51:52 -> 14:33:21), and the baseline assistant message now lands after the run it reports on rather than three hours before it. Timestamps are only recomputed on a fresh seed; create_demo_snapshot's INSERT OR IGNORE leaves an installed demo untouched. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bcc05dacf0 |
OR-262 Live Overleaf sync over the editor channel (#295)
* feat: live Overleaf sync over the editor channel A linked paper now follows Overleaf as people type, and sends saves straight back, instead of syncing over the git bridge every 30s. The bridge snapshots on demand and rate-limits polling, so it can never be immediate; the editor's own socket.io channel can. - `local::overleaf_live` speaks the channel (sharejs text OT, UTF-16 positions, one in-flight op) for the paper's `.tex`/`.bib`/class files. Each doc keeps the confirmed server text as a three-way base, persisted per paper folder and seeded only from the git sync's agreed hashes, so a reconnect can tell a local edit from a remote one. Where both sides moved the same span the file is held for the git sync's conflict prompt rather than guessed at. - The git sync stays for figures, conflicts and fallback, pausing the channel while it runs and re-seeding it afterwards. It also runs on a slow timer while live, so figures move without being asked. - The session cookie can be imported from a signed-in browser (Chrome/Edge/Brave/Arc/Vivaldi/Chromium/Firefox, macOS) instead of pasted. Only the Overleaf cookie is read, and the Keychain is asked for a browser's key only when that browser actually holds one. - Saves carry a hash of the bytes the editor loaded, so a save cannot overwrite a collaborator's edit that landed first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: address review of the live Overleaf sync - A channel stuck reconnecting no longer suppresses the git sync for ever; only one that is actually up carries the text. - Whether a sync has agreed a base is held per paper, not derived from the channel, so the effect's own teardown can no longer swallow a retry. - A remote edit keeps the file's CRLF line endings instead of rewriting every line as LF. - Reads refuse a symlink, as writes already did, so a link out of the paper cannot have its target pushed into the Overleaf document. - The git sync releases the channel through a guard, so a folder cannot be left paused if the sync unwinds. - A cookie refusal says so on the status, rather than the dashboard reading the prose of an error message. - Firefox profiles that cannot be read are skipped like Chromium's, and the WAL sidecars are copied so a cookie set moments ago is visible. - The sync button only claims text is live when it is, the paste and import buttons no longer share a label, and the live line is a status region. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
fa6d82571e |
OR-264 feat: "Restart now" button on the update banner (#285)
* feat: restart button on the "Updated, restart to use it" banner The banner (and Settings → Updates) now offer "Restart now". The server exposes POST /api/update/restart, which answers and then relaunches into the copy the updater installed: a terminal `orx up` execs itself on the same port with --no-browser, and the macOS app exits and reopens the bundle with the previous port handed over so the open tab reconnects. The dashboard polls until a different server instance answers, then reloads so the UI is the new build too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Show the update banner on the zero-projects screen too Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * Fix unused warnings on non-macOS targets in the relaunch path Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
b66af71a5b |
fix: keep starter prompts when switching harness or model (#284)
* fix: keep starter prompts when switching harness or model The four starter prompts are cached per project brief and locale only, and the UI keys them by project, so changing the composer's harness or model no longer regenerates them or flashes the loading skeleton. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore: rebuild ui/dist Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: retry starter prompts only for a harness that failed A failed generation is remembered per agent, not per project, so switching from a broken harness to a working one still produces prompts. The route answers with an empty list when there is nothing to offer, so projects with experiments never show the loading skeleton on a switch. Rebuilds ui/dist. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
56ed86d466 |
feat: ! bash mode in the chat composer (#283)
A draft that opens with `!` is a shell command, as in Claude Code: the
composer turns amber with a Bash chip, Enter (or the Run button) runs it
in the session's worktree through POST /api/chat/sessions/{id}/shell,
and the exchange lands on the transcript as a user-side card. Commands
run since the last real message are folded into the next turn's prompt
as untrusted context. Output is capped, the process group is killed on
timeout, and a backgrounded child cannot hold the request open.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
7d6da83157 |
fix: accept the renamed alphaXiv/OpenResearch repository in release builds (#281)
The GitHub repository was renamed from openresearch-cli to OpenResearch, so build.rs rejected ORX_OFFICIAL_RELEASE_BUILD=1 and the v0.1.119 release failed. Update the guard and the release/download URLs to the new name. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f9cfcc883f |
chore: bump version to 0.1.119 (#280)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
cc496956c4 |
feat: simplify the nanochat demo and seed runnable follow-up experiments (#279)
Collapse the seeded demo transcript's 51 per-checkpoint validation messages into three milestones and the eight SFT updates into two. Seed two idle 200-step probe experiments under the completed baseline, each with its own branch at the baseline commit and a self-contained run command, and tell the demo session's agent to launch them directly with `orx exp run`. On first demo open, lead the right pane with the experiments tab and prefill the composer with a prompt to run one of them. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
e867a4a92d |
feat: starter prompts follow-ups — composer model, OpenCode dead-key detection, box outlines (#277)
* feat: run starter-prompt one-shots on the composer's model Codex and OpenCode children take the model the user picked (the stored preference for warm-ups) instead of the CLI's configured default, so an OpenCode setup whose default provider has no working key still works with a free opencode/* model. Cache keys include the model. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * feat(opencode): detect a rejected provider key Probe each signed-in OpenCode provider once per auth.json version with a tiny read-only request; on an authentication rejection hide that provider's models, mark it in the account line, and explain the fix. Settings shows harness notes whenever one is set. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ui: tint each starter prompt box with its own accent Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ui: outline starter prompt boxes instead of tinting them Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ui: soften the starter box outlines Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * review: follow-up nits for model plumbing and dead-key detection Share one opencode child helper between the one-shot and the key probe, cache only definitive probe verdicts, drop the bare 401 marker, mark OpenCode not ready when no model is left, reconcile a saved composer model the harness no longer lists, seed the pre-warm model like the composer does, ignore the model in cache keys for harnesses that ignore it, and render the Settings note. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
dccc6a2ce0 |
OR-249 feat: starter prompts for a new project's empty chat (#276)
* feat: starter prompts for a new project's empty chat Four clickable boxes under "What should we research?" whose prompts are written by a headless one-shot call on the user's chat harness from a brief of the project (paper overview, README, manifests, entry points, file tree). Cached by content fingerprint, warmed on project creation and pre-warmed from the new-project form. The per-harness title children are generalized into a reusable one_shot request. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * review: starter prompts round-two fixes Build the brief off the runtime, cap the git listing, remember failures briefly, bound concurrent model children, clip model output, only pre-warm briefs that will match the created project, and fold the per-harness title overrides into the trait default. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * review: starter prompts round-three nits Keep shallow paths when capping the listing, pre-warm folders only when git is ready, and tidy names, docs, and comments. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c5b49fe37f |
OR-247 feat: simplify the skills experience (#272)
* feat: simplify the skills experience The Customize tab was five sections deep in choices that no longer earned their place: a Global/This project scope picker on every card, a separate "import from your agent" list, and two split skill lists. - Skills installed in a coding agent are now mirrored automatically — read live from its skills dir on every listing and every session write, so a skill edited in Claude Code is the one the next session runs. The import step and its two endpoints are gone. A session hosted by the agent a skill came from is not handed a copy it already loads. - Claude Code's installed plugins are mirrored too, discovered through installed_plugins.json so only real installs count. - Every skill and LaTeX template is global. Project scope is removed from the store, the API, and the UI; anything already saved under it is migrated into the single store on first access. - Skill frontmatter is parsed the way skills are actually written: folded and literal blocks, and values wrapped over indented lines. A skill whose frontmatter we cannot read is a skill the user never sees. - ~/.agents/skills is read whenever it exists, rather than only when ~/.codex does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: harden the skills mirror after review Review of the mirroring change turned up defects worth fixing before it ships: - A plugin's skills are registered namespaced (`runpod:flash`), so the bare `/name` only resolves from a copy. They are no longer treated as natively loaded by the agent that installed them. - Uploaded skills resolve through their folder name again, not the frontmatter `name` a hand-edit can change out from under them. - The hosting agent's own skills are dropped only after they have won their `/name`, so a same-named skill from another agent can't take their place in the worktree while the dashboard shows the first. - Session skill dirs are replaced and pruned only when the manifest says we wrote them, so a `.claude/skills` the project itself commits is left alone; manifest names are re-validated before any removal. - A folder whose source is unchanged is not re-copied, folders over the upload budget are skipped, and neither the copy nor the size walk follows symlinks. - The retired per-project store is emptied, never deleted: anything that can't move stays put as the user's only copy. - A `#` comment no longer folds into the value above it, and a long description truncates instead of dropping the skill. - The Customize tab distinguishes a failed skills fetch from an empty one, and ignores drops while an upload is in flight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep one answer for what a skill name resolves to Second review round: - `source_dirs` resolves uploads through the same listing the dashboard reads, so a folder whose SKILL.md won't parse can't win a name in the worktree while the tab shows the mirrored skill it shadowed. - A folder tally now carries a digest over every (path, size) pair, so a rename or a move inside a skill brings the session copy forward; each source and destination is walked once per turn instead of three times. - A destination whose content already matches its source is adopted as ours, so a lost manifest heals instead of freezing that skill forever. - A mirrored folder over the upload budget is left out of the listing too, rather than offering a `/name` that never reaches the worktree. - LaTeX templates get the same skip-if-unchanged treatment. - Archive junk under the retired per-project store no longer keeps it alive on every call. The freshness test now watches the inode: `fs::copy` carries the mtime across on macOS, so the timestamp it asserted on could never have failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: never let an adopted skill dir become prunable Round three: adopting a destination whose content already matches its source healed a lost manifest, but it also recorded a directory orx had never written — so once the source went away, the prune deleted a skill the project itself commits. Adoption is now limited to the case it was for: no manifest at all. While a manifest exists it stays the whole truth about what we own. Also from that round: the upload budget moved into `source_dirs`, so the menu, the hover preview and the session write agree on which `/name` exists; a stray `.DS_Store` beside the retired per-project store no longer keeps it alive; and the uncapped walks stopped pretending they can fail. The fingerprint test never reached the SKILL.md byte comparison it was supposed to cover, and the migration test never asserted the retired tree was gone — both fixed, and both verified by reverting the fix under them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: one budget answer for uploads too The listing kept sizing uploads with an uncapped walk while the resolver had started budgeting them, so a hand-edited store could list a skill no session would be given. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
369d4df7fa |
feat(ui): render HTML files in the right-hand panel (#273)
An .html file opened in the right pane now renders as the page it is — a generated plot or report — instead of showing its source, matching how markdown already behaves. The existing header toggle switches to source (and, for a checkout file, the editor). The document renders in an iframe sandboxed without allow-same-origin, so its scripts run in an opaque origin and cannot reach the loopback API, the dashboard's storage, or the parent page. That origin is also refused any loopback subresource, so project-local assets a document references are fetched by the parent and inlined as data: URIs rather than rewritten. Inlining is gated on the response Content-Type matching the element it would fill, which requires the server to type a response by the path it resolved to. It did not: content type came from the requested name while the bytes came from the canonicalized target, so an in-repo symlink (logo.png -> .env) served secrets as image/png. The three raw handlers now pass the resolved path, disk_response's parameter is renamed type_path to say so, and chat_attachment is made uniform. A srcdoc document inherits the embedder's base URL, so a report's #section link navigated the frame into the dashboard SPA. A base of about:srcdoc is injected unless the document names an origin of its own, and external anchors are given target="_blank". Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3c8c8dbcbe |
OR-241 fix(ui): tell the dashboard when the backend is gone (#268)
* feat(ui): tell the dashboard when the backend is gone The dashboard's one EventSource was also its only evidence the local server still exists, and a dropped stream was treated as a transient reconnect. Quitting the macOS app therefore left a tab that looked live but was frozen, on an ephemeral port nothing will ever answer again. Track the stream's state and surface it as a banner. The stream is now re-openable too: a CLOSED EventSource is the browser giving up for good, which no retry of its own will recover from. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ui): address review on the offline banner Copy is descriptive rather than prescriptive: the macOS app binds a fresh ephemeral port per launch, so telling the user to reopen it promised this tab a recovery it will never get. The live region is now permanently mounted with swapping text, since one inserted together with its content is missed by VoiceOver and NVDA, and an error arriving after teardown can no longer arm a timer nothing is left to clear. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d69717f42e |
OR-244 feat(ui): name the open project in the browser tab title (#269)
* feat(ui): name the open project in the browser tab title
The tab now reads "OpenResearch - {Project Name}" while a project is open,
so multiple dashboard tabs are distinguishable.
The title falls back to plain "OpenResearch" on every screen that renders
instead of a project — projects home, the startup-error page, and the
loading spinner — since a partial startup failure or an SSE project event
can leave `projects` populated behind those. The project name is wrapped in
`autoDir` so an RTL name cannot reorder the brand prefix, matching how
project names are interpolated elsewhere.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(ui): lead the tab title with the project name
Browser tabs truncate from the right, so the invariant "OpenResearch"
prefix survived while the project name — the part that distinguishes one
tab from another — was clipped first. The title now reads
"{Project Name} - OpenResearch".
The autoDir isolate around the name becomes load-bearing in a second way
here: as the leading run it would otherwise supply the title's first strong
character, so an RTL project name would flip the whole title's base
direction. FSI is skipped by the base-direction scan, keeping it LTR.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
5130caea0a |
fix(ui): trim filler copy from settings and the model picker (#264)
Remove subtitles that restate their own heading, and rewrite the model-picker harness note as two sentences instead of an em-dash splice. Drops the "General" subtitle and four message keys whose UI was already removed in #262 and #258, leaving the catalogs with dead entries. Restores the Environment subtitle wording that #262 replaced. The Persian model-picker string deliberately ends without a period: it renders in a bare div, and index.html pins dir="ltr" for every locale, so a trailing neutral would paint at the wrong end of the RTL line. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
72f550c193 |
Keep an agent's orx the one running its session (#266)
`prepare_env` fronts this build's directory on a harness child's PATH, but the child's tool shell then sources the user's own startup files, and a profile prepending its bin dir — or macOS `path_helper` rebuilding PATH wholesale — pushes it back down. Agents shelling out to plain `orx` reached whatever copy was installed instead: one session ran `orx agent spawn` against a build predating the subcommand, while the skill documenting it came from the binary serving that session. Export the running orx's bin dir and re-front it from the per-session shell hooks, which are the last thing orx gets to run after the user's files. The guard rides every zsh startup file and the bash BASH_ENV hook; a harness that snapshots the user's shell captures the guarded order, since the capture itself passes through these files. `orx_bin_dir` now degrades to the un-canonicalized `current_exe` path rather than giving up, which also stops app-mode children from inheriting launchd's bare PATH after a failed canonicalize, and drops a relative or colon-bearing directory that could not be spelled in a PATH entry. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a9af660f9e |
OR-216 Add an orx-figures skill for publication-quality figures (#260)
* Add an orx-figures skill for publication-quality figures Figures were left to matplotlib defaults: screen-sized, titled where a caption belongs, and rasterized where the document wants vector. This adds a bundled skill that routes by the question a figure answers and ships the style layer rather than describing it. `orx-figures` follows the `orx-compute` shape — one SKILL.md carrying the non-negotiables, then exactly one of six references (curves, scaling, comparison, pareto, matrix, diagram). Two assets install beside them: `orx_figstyle.py`, vendored into the project so a figure stays reproducible after the session ends, and a TikZ scaffold sharing its palette. `save()` audits every figure and prints the result: printed width against the known column widths, Type 42 font embedding, stray axes titles, missing axis labels, sub-5pt text, overlapping text, and text running off the canvas. Prose the agent can skip is how the first unusable figures shipped; these checks run whether or not the guidance was read. The playbook and the `orx skill` overview both direct figure work here — a session with the skill installed still plotted raw matplotlib until they did. Content is calibrated against 39 popular recent alphaXiv papers (283 captions) and 10 arXiv e-print sources: 87% of real figure PDFs are built wider than any single-column page, so the width check is the load-bearing one; multi-panel is 27% of figures; and captions run a median of 28 words. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix review findings in the figures skill Correctness: - `label_ends` computed the reserved x-room in the pre-`set_xlim` scale, so the labels overhung the spine by need^2/(width+need) — about a quarter of a long label. Solve for the limit at which the original span occupies (axes width - label width) pixels instead. - The audit's categorical-axis exemption keyed off whether tick text parsed as a number, but log formatters emit mathtext, so every log axis was classified categorical and escaped the missing-axis-label check. Decide by scale. - `save()` audited before writing, and the audit needs an Agg-family renderer that is not guaranteed in the very case its first check reports. Write the files first; the audit can no longer be the reason a figure is lost. - `curves.md` interpolated every seed onto the first seed's grid, and `np.interp` clamps, so a run that stopped early gained a fabricated flat tail. Clip to the range every seed reached. - `comparison.md` hardcoded `ylim(0, 90)`: any score above it drew as a bar clipped at the axes edge, which is the dishonesty that reference is about. - `scaling.md` zipped families against three markers, silently dropping a fourth family from the main panel while still plotting its residuals. - Guards for degenerate input: zero-variance `diff_ci`, all-NaN heatmaps, empty confusion rows, resamples that fail to fit, empty label lists. Consistency: - `diagram.md` still taught the retired natural-size include, and its inline path dropped `\familydefault`, which would render a serif diagram beside sans plots. Both contradicted SKILL.md. - `orx-paper` and `orx-reports` examples named a rasterized plot. - The audit emits a font-fallback finding the docs never listed, and it is the one finding an agent should hand over rather than fix. Tests now assert the rendered playbook rather than the raw template, and pin the audit's problem strings rather than text that also appears in docstrings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix round-two review findings The round-one TikZ fix was worse than the bug: it told the agent to move `\renewcommand{\familydefault}{\sfdefault}` into the paper's preamble, which is document-wide and would have set the body text, headings, and captions of a submission to sans. The face now rides in the `orx`, `edgelbl`, and `stagelbl` styles, so it travels with the picture and touches nothing else. Also from review: - `diagram.md` still carried the broken relative `cp assets/...`, which does not resolve from a session's cwd. - The new "write a report figure to the artifacts directory" note contradicted "keep the generating script beside its output"; both destinations now keep the pair together. - A `WIDE` figure only gets a 6.75in `\linewidth` inside `figure*`; in a plain `figure` it silently halves, which non-negotiable #1 now says. - Skipping all text on an axis-off axes also exempted the user's own content, which on such an axes is the figure. Exclude only the undrawn axis artists. - `save()` leaked the figure when the audit raised, and did not create the stem's parent directory. - A single seed drew a zero-height band that reads as certainty; `mean_ci` and `diff_ci` now warn, matching what the references demand of the caption. - Guards for disjoint seed ranges and all-zero scores. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Close the round-three review findings None were blockers, but three were real defects on secondary paths: - The inline `\input` route omitted `\usepackage{tikz}` and the `\usetikzlibrary` line, so a paper following it literally would not compile: `trapezium`, `Stealth`, `fit=`, and `on background layer` each need one. - Vendoring was hardcoded to `figs/`, so a report script written under the artifacts directory could not import the style module beside it. - Removing `\familydefault` left no documented way back to a serif diagram while `use_style(family="serif")` still exists for plots. The colorbar guard now accepts `ax.get_label() == "<colorbar>"` as well as the private `_colorbar`; verified against matplotlib 3.11 that both hold today, so a rename can no longer make `clean` unreachable for every heatmap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5c365c6e07 |
Never chat with a coding agent that cannot start (#247)
A user's `codex` npm install had lost its vendored platform binary, so the node wrapper on PATH ran and exited 1 on every invocation. orx reported the harness ready anyway — `installed` came from PATH presence alone, and codex reads its credentials off disk — so the composer let them send, and the CLI's own crash landed in the transcript. The failed turn was then classified as possibly-delivered, so the card offered a "may have been accepted" Continue that re-ran the same doomed spawn. A `--version` probe that fails conclusively now marks the install broken: never ready, with a reinstall note in place of a sign-in one. Conclusively means the CLI printed no version and either exited non-zero or could not be executed for a stable reason — a timeout or fork pressure stays inconclusive, because falsely locking a working harness out of chat is worse than the bug. Claude needed two more guards: its auth overlay was overwriting the reinstall note with `claude auth status`, and a claude that broke mid-session had its cached ready entry restored and re-cached, permanently. For the turn itself, a codex exec that exits having emitted no event never reached the model, so it reports as undelivered and the card offers Retry. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8663e2137f |
Show that chat file chips open a right-pane tab (#252)
File and run chips looked like inline code, so nothing said a click opens them in the right pane. Each chip now ends with a panel glyph and carries `.tool-target`'s underline on its label, and the tooltip names the destination. Non-interactive chips drop both cues. The rules live unlayered in tailwind.css beside `.md .file-chip` so the wrapper's `.run-chip svg` utilities cannot tint the trailing glyph. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0260bcb624 |
OR-193 Keep a paper in step with an Overleaf project (#250)
A `.tex` in the checkout can be linked to an Overleaf project and synced with it in both directions over Overleaf's git bridge: edits made here are pushed on save, edits made in Overleaf are pulled while the tab is open. Accounts without the bridge — it is a paid Overleaf feature, and Overleaf publishes no way to ask whether one has it — can upload the paper as a new project instead, which any account can do. Each sync records what both sides agreed on, path to content hash. That agreement is what gives a later difference a direction: one side moved, or both did. A file both sides changed is left untouched on both sides and reported, because resolving it either way would discard somebody's writing; the panel asks which copy to keep. The agreement is scoped to the checkout it was taken from, since a session worktree and the hub clone hold different copies of the same path. A sync clones the project fresh, works out the whole exchange without writing anything, pushes, and only then touches the checkout — so a rejected push cannot leave the disk ahead of what was recorded, and a co-author's edits can never be clobbered by a stale local copy. The token is stored outside `~/.openresearch/env`: that file is fanned out to every compute backend and into the agent's environment, and nothing off this machine pushes to Overleaf. It reaches git only through a credential helper scoped to the linked project's host. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
6f1da0a915 |
Chip a slash command where it was typed (#235)
The composer hoisted a picked command into React state and pinned its chip to the textarea's first line, so a command chosen mid-message jumped to the front and the message had to be reassembled on send. The `/name` token now stays in the draft exactly where it was typed and a mirror of the textarea paints the chip behind the real text — the draft is the wire form, with nothing to reassemble. The mirror copies the textarea's metrics at runtime and its pill bleeds via box-shadow spread, so it adds no layout width and every glyph after it stays put. Backspace just behind a chip deletes the whole command, a command typed with no args shows its arg hint as ghost text (and to screen readers, which the placeholder used to carry), and the transcript chips the tokens where they were sent. Only a leading command expands server-side, so only that one is lowercased for the backend's exact match. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2025865fa1 |
Tint the chip-targeted line blue instead of red (#233)
Clicking a `file:line` code chip highlighted the target line with `--primary`, a deep red that read as an error marker. Use the existing `--accent-blue-subtle` band with an accent-blue rail instead. The dist rebuild also prunes three stale bundles and repairs index.html's dangling stylesheet href, which release builds embed. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f27c858c78 |
Simplify the DMG installer window to a plain drag-to-install (#228)
Drop the brand mark from the DMG background so the window is just the app icon, an arrow, and the /Applications alias. With the mark gone the window shrinks 640x400 -> 640x320 and the icon row recenters (86px above the icons, 86px below the labels). The arrow now sits level with the icon centers (ICON_Y) and is centered on the icons' artwork gap rather than their 128px cells, since the two icons inset their art by different amounts. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0526961e4e |
Invoke a slash skill from anywhere in the composer draft (#229)
The chat composer only offered skills while `/name` was the first token; inline slashes were narrowed to `/plan`. A `/name` under the caret now opens the menu wherever it sits, and picking it chips the command with the surrounding text as its args. Picking mid-sentence takes a deliberate Tab or click: Enter accepts only where the command opens the draft or one of its lines, so an ordinary sentence mentioning a `/token` still sends as typed. Auto-convert on a trailing space stays limited to a draft that is nothing but the command. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
40d57170d1 |
Keep the code editor's selection translucent (#230)
The editor layers a transparent-text textarea over the highlighted code, so the opaque global ::selection background painted blank rectangles over the tokens underneath — selecting a line hid the code you were selecting. Give the edit surface its own translucent selection tint instead. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a77c2218dc |
Start a blank paper project when a paper has no public repo (#231)
A paper with no linked GitHub repository could not become a project at all: the API rejected it and the form hid every control below the paper picker. Such a paper now creates a project seeded with its PDF, committed as the repository's initial commit, and the form explains why there is no repository to clone. The destination has to be an empty folder of its own, outside any Git repository, because the paper is committed into the repository the project initializes — seeding an existing repository would leave the paper untracked, and a subdirectory would put it outside the project root. A failed PDF download fails the request rather than quietly producing an empty project, since the PDF is the only content such a project has. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0bfeb1ebbe |
Find the LaTeX toolchain on the user's PATH, and recommend Tectonic (#227)
* Give shell_env the one PATH search every lookup uses `find_on_path` sat in `harness::detect` as `pub(super)`, so anything outside the harness registry that needed a shell-PATH lookup wrote the loop again — `find_opencode` and `resolves_on_path` both had a copy, the former commented "Mirrors `find_on_path`". Move it to `shell_env`, which already owns `search_path()` and whose module doc is about which PATH orx searches, and point all four call sites at it. The shared lookup now also drops *relative* PATH entries, not just empty ones. A relative entry names no fixed directory: it resolved against orx's own cwd, which for a Finder-launched bundle is never where a user's tool lives, and the child that later ran it would resolve the same name somewhere else again. `other_orx_on_path` in install_cli.rs keeps its own loop — it needs every candidate and canonicalizes them, so it is a different search. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Find the LaTeX toolchain on the user's PATH, not launchd's A `.tex` file opened in the macOS app reported "No LaTeX toolchain found on PATH" on a machine with tectonic installed. The bundle is started by launchd with `PATH=/usr/bin:/bin:/usr/sbin:/sbin`, where no TeX lives, and `latex.rs` was the one lookup still resolving against the process environment — the harnesses had gone through `shell_env`'s probed shell PATH for exactly this reason. Verified against a bundle launched with launchd's environment: `/api/latex/engine` returns `tectonic` where the installed build returns null. Finding the tool is not enough on its own: latexmk finds the engine, biber and bibtex on PATH itself, so the child is handed the same PATH we found it on. The tool is spawned by its absolute path but symlinks are left unresolved — TeX picks its format file from the name it was invoked as, so canonicalizing `pdflatex` to `pdftex` would silently run plain TeX. `on_path` no longer spawns `<binary> --version` to decide. It costs a probe that a present-but-unrunnable binary would have caught, and buys back five subprocesses on every `.tex` tab open. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Point a user with no LaTeX at Tectonic rather than MacTeX The no-engine hint led with MacTeX and offered `brew install --cask mactex` as the command to copy — a multi-gigabyte install for someone who just wants to see their paper. Tectonic is one self-contained binary that fetches each document's packages, so it is the install a user can actually finish, and it is what the copyable command now installs. The distribution stays named second, because Tectonic runs only XeTeX. The hint says so: it is the last screen that can, since the hint disappears the moment an engine exists and the amber substitution note only fires for a document that names its engine explicitly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
51e510f1a2 |
Compile LaTeX papers to PDF in the file viewer (#225)
A `.tex` file in the live checkout now compiles on open and shows the real PDF, with a toggle to a live source editor that recompiles on save. Compilation targets Overleaf's behaviour: `latexmk` drives whichever engine the document asks for via `% !TeX program`, running biber or bibtex and repeating passes until references settle. Without latexmk the engine is driven directly and that orchestration is done here; with neither, `tectonic` stands in and says so, since it is XeTeX and cannot honour a LuaLaTeX document. Users can upload their own templates — a conference class, a house style — as a `.tex` or a `.zip` carrying its `.cls`/`.sty`/`.bst`. Templates are global or scoped to one project, and are written into each session worktree so the agent can start a paper from one. Uploading and managing them lives in the renamed Customize tab. The `orx-paper` skill and the `/write-paper` command teach the agent to write papers as `.tex` in the working tree, follow an uploaded template when one exists, and ground every reported number in a run rather than in memory. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a0c30df38a |
OR-129 Let an agent delegate a task to a second agent session (#221)
* Let an agent delegate a task to a second agent session An agent could only do work itself or queue it behind the current turn. `orx agent spawn "<task>"` now starts a helper in its own top-level session — visible in Recents, its own transcript, its own worktree — and resumes this chat with the helper's closing reply when it finishes. The CLI only writes rows: the child's `chat_sessions` row (carrying a new `parent_session_id`) and a `chat_spawns` record. The resident `orx up` watcher starts the child's first turn and later reports back, the same store-and-watcher split as `orx exp wake` and for the same reason — the spawning `orx` is a short-lived subprocess with no harness to run a turn on. It is a CLI verb rather than an MCP tool because the mcp-gate bridge is Claude-only, while all three harnesses already shell out to `orx` with `ORX_CHAT_SESSION_ID` exported. Spawn rows walk pending → starting → running → notifying → done. The two claimed states carry a token so a watcher that dies mid-step leaves work reclaimable rather than lost, and a helper whose parent was deleted retires quietly instead of stranding the row. Spawned agents cannot spawn their own. A helper that can delegate turns one request into an unbounded tree of paid sessions, and nothing downstream bounds it. Settings carry over only when parent and child share a harness: a model or permission-mode id from one CLI means nothing to another. Plan never carries over — a helper is spawned to do the task, not to plan it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Harden spawned agents against losing the delegated task Review of the spawn feature found several ways a delegated task could be accepted and then quietly not happen. A helper's completion was read from `ChatHost::is_busy`, a map this process owns. A second `orx up` — or the same one after a restart — sees an empty map, so a helper working under it was reported "finished" within 3s and the delegation was abandoned. The gate now consults the durable turn lease, and `finish_turn` stamps the spawn row, so an idle helper with no stamp is reported as interrupted rather than as a false completion carrying half-written text. A failed start returned the row to Pending and retried every 3s forever. The brief is persisted before `send_message_showing` can report the turn didn't start, so each retry appended it again. Retries now skip the already-recorded brief, stop after three attempts, and wake the parent with "could not be started" — the CLI had promised it a wake-up. Claude activates Plan through its permission mode, not the plan axis, so clearing `plan_mode` still handed a planning parent's helper a mode that only ever produces a plan. The plan permission id is now dropped when inheriting. The closing reply was read from the last assistant message. An answered prompt card rides its own text-less message and becomes the branch tip, so an opencode helper that hit one card reported "no written reply" and its actual answer was lost. A turn that died on a harness error reported the same thing, reading as "nothing to say" rather than "it failed". The report now walks back to the newest message with text and distinguishes a reply, a failure, and genuine silence. Depth was capped at one level; breadth was not, so a looping parent could create as many paid sessions and worktrees as it liked. Five in flight. Also: the transcript's squash key ignored the spawned ids, collapsing two adjacent spawns into one row and stranding the second helper; following a spawn card set the one session filter that hides archived rows; the delegation playbook told agents to hand helpers a node this session owns, which cardinal rule 1 and one-branch-one-owner both forbid, and never told them to bound a helper's compute. The playbook placeholder test checked a hardcoded token list, so it could not catch a newly added token — it now scans the rendered prompt, and fails on the `{max_spawns}` this change introduced if its substitution is removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Never retire a delegation before its report is delivered Round-two review found four more ways a spawned task could be lost. The give-up notification was fire-and-forget: it returned silently when the parent's turn slot was taken, then retired the row anyway. The parent is busy in exactly the expected case — it spawned from inside a turn, and the third failed attempt lands ~9s later — so the promised "could not be started" usually never arrived. Both wake-ups now go through one `deliver_wake_up`, which retires the row only once the message actually started a turn and otherwise leaves it for the next tick. `--no-wake` settled to Done the instant the turn started, and the in-flight cap counts only rows that are not Done. A looping parent could spawn five, wait one tick, and spawn five more forever. Fire-and-forget rows now stay Running until the helper is genuinely idle, so "in flight" means the same thing for both kinds. `interrupted` was read from the listing snapshot while the liveness check was live. A helper finishing cleanly while an earlier iteration awaited a whole turn was reported as interrupted and its reply thrown away — the same loss the stamp was added to prevent, inverted. It is re-read after the liveness check. A user pressing Stop reaches the same `finish_turn`, so a manually stopped helper was stamped finished and its half-written text handed back as a closing reply. `spawn_outcome` now recognizes the interrupt marker. The wake-up also claimed the helper's commits were "on its own branch". Session worktrees start detached, so a helper only has a branch if it ran `orx create-experiment`, and for the delegations the playbook recommends there are usually no commits at all. Restored the accurate wording, in the report and the playbook, and the clause is omitted entirely when the worktree isn't there. Also: a reclaimed `starting` claim now counts as an attempt, so a watcher that dies inside the send cannot retry forever; `ChatSpawn` no longer carries a `state` no reader used, and the INSERT hardcodes `pending`; branch-local dev slots get ALTERs for the renamed columns. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2f9f01dea2 |
Fork chat turns only from user messages (#218)
Drop the assistant-side "Try another response" control and its version pager, leaving "Edit and re-send" on the user message as the single way to fork a turn. Fork positions are now computed only for persisted user messages, and ForkControls collapses to that one purpose. The server's fork and branch APIs are unchanged, so a session that already holds re-sampled replies keeps them; only the transcript's way back to them is gone. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
38e9943df6 |
Wrap long code lines instead of scrolling sideways (#219)
Code files in the viewer scrolled horizontally for any line wider than
the pane. Wrap them instead, in both the read-only view and the editor.
Wrapping breaks the old gutter's assumption that source line N is visual
row N, so each line becomes its own row carrying its own number:
highlightLines() splits the refractor tree per source line, reopening
tokens that straddle a newline.
In the editor the overlay's numbers are absolutely positioned out of the
line box — an in-flow number is an atomic inline offering a break
opportunity the textarea lacks, which desynced the two layers on long
URLs and base64. Every code element also spells out its font: Preflight
isn't imported, so the UA `pre, code { font-family: monospace }` beats an
inherited font and would give the layers different fonts and `ch` units.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
d16359ce19 |
Open chat-chip tabs in preview mode (#217)
A tab opened from a chat chip (file, run, experiment, plan, sub-agent) now renders italic and is replaced by the next chip click, VS Code style. Double-clicking the tab, pressing Enter on it, typing in an editable file, or following a chip out of it commits the tab so it stays open. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c82bf91a41 |
OR-185 Steer a running turn instead of only queueing behind it (#211)
* Steer a running turn instead of only queueing behind it Typing while a turn ran parked the message: the queue held it until the turn ended, then drained it into a NEW turn. So "no, use the other dataset" arrived only after the work it was meant to redirect had finished. Both chat harnesses can already take input mid-turn — claude keeps stdin open on its resident `--input-format stream-json` child, and the codex app-server exposes `turn/steer` against the active turn. A running turn now registers a steering sink, and Enter hands the message straight to it; Cmd/Alt+Enter still parks it. Harnesses that can't steer are unchanged, gated on a `supports_steering` capability the UI reads. The delivered text rides the in-flight assistant message as a `steer` part, so it renders where the agent actually received it — a separate user row would carry a later timestamp and sort after the whole turn, misdescribing when it landed. Steering never silently drops a message. Attachments and changed composer settings still queue (both only take effect at a turn boundary), and every delivery failure — dead child, older app-server, review or compaction turn, a turn that ended underneath the send — parks the text as an ordinary chip. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Harden steering against losing a message Review of the steering turn found four ways a steered message could be accepted and then quietly do the wrong thing. A send that raced the turn's end was accepted into a channel with no reader: the sink stayed published through the whole epilogue — flush, usage write, finish_turn, drain_queue — while the harness had stopped reading the moment run_turn returned. The task now deregisters, closes, and parks whatever is buffered before any of that. Settings changed mid-turn were compared against the session row, which the composer updates *before* the message that carries them: picking a permission mode or toggling Plan and then typing steered a read-only instruction into a child still holding write access. The comparison now runs against a snapshot of what the running turn is actually under, and the composer always sends its current state — including the plan axis, which went missing entirely once a toggle had persisted. codex advertised steering on the legacy exec path, which cannot read the channel. One predicate now decides both dispatch and the reported capability, and the trait is a ceiling a `detect` may narrow but never widen. `turn/steer` rode the shared 150s request timeout while awaited inside the event loop, so a slow app-server froze the transcript and hid any card it raised; and a timeout was treated as a rejection, re-running text codex may already have applied. It now has its own short bound, parks only on a definitive rejection, and records the text either way. Also: absolute watchdog deadlines, so repeated steering cannot postpone the wedged-child detector; a steer no longer flattens the streaming caret or the live tool row; and a retry fork carries the steered instruction instead of silently re-asking the pre-steer prompt. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Apply rustfmt to the merge resolution The merge commit staged the conflict resolution before `cargo fmt` ran, so a line rustfmt joins went in split and CI's format gate caught it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
45eb162e19 |
Focus the existing dashboard tab on a Dock click (#212)
The macOS app's Dock-icon handler called `open <url>` on every click. Once the dashboard SPA has navigated off `/` the browser has no exact-URL match, so the click opened a new tab instead of returning to the dashboard already open. The handler now tries, best first: the exact tab via a JXA script over the scriptable browsers; the browser window itself, when a tab still holds its `/api/events` stream (the only way to infer an open dashboard in Firefox, which exposes no tab API); and finally the plain open it did before. A repeat click within five seconds takes the plain open, so a raise that did not surface the dashboard is always recoverable — the app's port is ephemeral, and the user has no other route back to it. JXA rather than AppleScript because it resolves each browser's terminology at run time: an AppleScript `using terms from` block fails to compile on a Mac without Chrome installed. Signing now passes macos/entitlements.plist, since the hardened runtime denies the Apple event outright without it — a failure that appears only in signed builds, never in a local one. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b4c0c966b8 |
Open an agent-cited artifact that lives in the checkout (#210)
Clicking a file chip an agent cited as `artifacts/<name>` showed "Artifact not found in the project." whenever the agent had written that file into its session checkout rather than the project's artifacts dir — the file was plainly there under Files, but the viewer only ever asked the artifacts endpoint. The checkout side has always fallen back to artifacts; make the mirror true. An `artifacts/…` tab that misses the store now probes the checkout, trying the cited path first (`artifacts/<rel>`, the layout the publish playbook itself recommends) and then the prefix-stripped name, taking the first hit. A probe that throws — unknown session, directory name — just means "not here" and can't replace the not-found copy with a raw error. Whatever answers becomes the file the tab operates on: `filePath` drives the header, the raw-bytes URL, saves, and open-in-editor, so a hit on the prefixed candidate can never read one path and write another. Such a file is editable like any checkout file, which is why the provenance note sits outside the scroll body — `CodeEditor` is `h-full`, so an in-body note overflowed it by its own height and could be scrolled out of view, taking the only signal that this is not the artifact with it. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
c3a56a5833 |
Fork chat turns to sample multiple responses (#206)
Retry re-samples a reply as a sibling of the one already there, and editing a sent message re-asks it as a sibling prompt. A `‹ 2/3 ›` pager switches between the forks of a turn. Transcript messages become a tree: `chat_messages.parent_id` plus `chat_sessions.active_leaf_id`, which picks the branch on screen. Harness CLIs mint a new session id per turn, so each user message records the id its turn started from (`base_native_session_id`) and each reply the id it ended at (`result_native_session_id`) — re-sampling resumes from the recorded start and branches the harness's own history instead of extending a sibling's. A `user_version`-guarded backfill chains existing transcripts and hands each session's newest message its live harness id, so a turn recorded before this change keeps its conversation when forked. Forks share the session's git worktree, so a re-sampled turn starts from whatever files the previous one left behind. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e2447aabb1 |
chore: release 0.1.107 (#204)
Puts the CLI and macOS app self-update (#201) into users' hands. It merged after v0.1.106 was cut, so no released build contains it yet. For the macOS app this is the changeover release: 0.1.106 and earlier have no updater compiled in, so nothing can fetch an update for them. Everyone on the app today needs one manual install of this build; from then on it self-updates. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
967defd7b5 |
Release only when a push changes the version (#203)
A duplicate release was dispatched on 2026-08-17: #202 merged with a version bump and started a release build; #201 merged eight minutes later, and because the tag is only created when that build *publishes* (~10 minutes in, per dispatch-releases mode), the "does the tag exist" guard saw nothing and dispatched a second release for v0.1.106. It failed with "a release with the same tag name already exists". The failure was the visible half. The two runs were dispatched from different commits, so whichever finished first decided what v0.1.106 contained — the released content for a version was a race. Guard on the version actually changing in the push instead, comparing Cargo.toml at github.event.before. A push that didn't touch the version can no longer release, whatever the tag state. The tag check stays as a second line of defense and as the fallback when the previous head is unreadable. Drop the concurrency group. It existed to stop two quick merges double-dispatching, which the version guard now does per-run and correctly. Keeping it would have been actively harmful: queueing a run cancels the pending one in its group, so a third push could cancel the bump's run, and with no later push re-dispatching, the release would vanish with nothing red to show for it — trading a loud duplicate for a silent miss. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
be25acc3f1 |
Self-update the CLI and the macOS app (#201)
Neither install shape reliably stayed current. The CLI had the plumbing to check and to apply an update, but nothing ever fired it — the user had to notice a stderr warning and run `orx update`. The macOS app had no update mechanism at all: the only way forward was re-downloading the DMG by hand. Both now update themselves in the background. When a check finds a newer release a detached updater installs it, so the CLI's next invocation and the app's next launch are already current; the dashboard offers a restart in the meantime. Nothing blocks the command the user actually ran, and a failing update backs off instead of retrying on every invocation. The macOS updater is the new part: it downloads the DMG, verifies its digest *and* requires a notarized Developer ID signature from our team before swapping the bundle by directory rename (writing into a running bundle invalidates its code signature and gets the process killed). It keys on a new `macos-app.json` release asset rather than the CLI's `dist-manifest.json`, because the DMG is attached by a separate workflow and legitimately lags — reading the CLI's version would advertise a build the app cannot install. Also here: `orx install-cli` links the app's binary onto PATH so an app user's CLI is the same build the app updates; a Settings → Updates card; and a guard on `orx delete --cli`, which through that new link would have resolved to the app's executable and gutted the bundle. Installs orx doesn't own — cargo, Homebrew, Nix — are never modified. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
64781c9716 |
chore: release 0.1.102 (#197)
Merging this triggers a release: `release-on-bump.yml` sees the new version on main and dispatches `release.yml`, which publishes `v0.1.102` and creates the tag. Cuts a tag containing the macOS app fixes from #195, which merged without a version bump — so the released `v0.1.101` predates them and there is currently no tag the DMG pipeline can build the fixed app from. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
46d359ff5e |
OR-173 Make the macOS app find harnesses, and stop it fighting a CLI install (#195)
* fix(app): detect harnesses in the macOS app by adopting the shell PATH Finder launches OpenResearch.app through launchd, so the bundle starts with PATH=/usr/bin:/bin:/usr/sbin:/sbin and no shell rc ever sourced. Harness detection resolves `claude`/`codex`/`opencode` off that PATH, so a DMG user with a Homebrew, nvm, or npm-global install saw every harness missing -- including an already-signed-in Claude Code session -- while a terminal `orx up` on the same machine detected them all. App mode now probes the user's shell once at startup and installs the result via `local::search_path`, which harness lookup and harness children consult instead of the process PATH. The probe is `$SHELL -ilc`: interactive, because zsh reads .zshrc only for interactive shells and that is where PATH edits live -- a login-only probe missed the ~/.local/bin holding `claude`. It runs an inner `sh -c` so fish reports a colon-separated PATH rather than its own list-valued $PATH, fences the value between per-probe nonce markers so rc-file chatter cannot forge one, and requires an absolute directory in the result. Best-effort throughout: a 5s cap, and every outcome logged. The five detection probes that spawned a harness binary with no env prep (`bin_version`, `claude_accepts_ultracode`, `claude_list_models`, `codex_model_list`, `run_models`) now go through `prepare_env` like `probe_auth` already did. Without that, a node-shebang install resolves via the hydrated PATH but its `--version` fails for want of `node`, and `gate_oauth_version` turns the missing version into `Unknown` -- telling a signed-in user to sign in. Verified end to end: a bundle run under launchd's PATH with an isolated ORX_DATA_DIR reports Claude Code agentReady, and finds a `codex` reachable only through the hydrated PATH -- codex has no fallback drop-location, so that resolution can only have come from the override. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): stop the bundle and a CLI install from conflicting Three ways the DMG app and a curl-installed `orx` stepped on each other. `orx delete` takes an exclusive lifecycle lock and refuses when another OpenResearch process holds it -- "Close `orx up`, `orx serve`, and any active runs before trying again." CLI `orx up` holds the matching read lock via `dispatch`, but app mode returns from `main` before `dispatch` runs, so it held nothing: `orx delete` from a CLI install would wipe the store out from under a live app. App mode now takes the same read lock for the life of the process. Agents shell out to `orx`, and `chat::prepare_env` prepends the running executable's directory so they get THIS build. In the bundle that directory held only `OpenResearch`, so `orx` fell through to whatever CLI the user had -- a different version, or nothing at all for a DMG-only user. The bundle now ships `Contents/MacOS/orx` as a symlink to the executable. Verified the symlink survives `codesign --options runtime` + `--verify --strict`. That alias would otherwise open a second dashboard, since a bare `orx` reaches app mode's no-arguments condition. `launched_as_app_bundle` now also requires that the process was invoked under the bundle executable's own name, which keeps the alias a plain CLI while leaving a direct run of the executable path -- the documented way to watch the app's logs -- in GUI mode. macOS `current_exe` reports the launch path symlink and all, so it is canonicalized before the comparison. Verified safe already, and left alone: the app binds an ephemeral port rather than 4791; the store is WAL with a 5s busy timeout; and run supervisors hold a per-run exclusive lock, so a second server recovering the same active run exits instead of double-driving it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): address review of the PATH and coexistence work Blocker: `launched_as_app_bundle` had grown real logic -- four distinct outcomes turning on a subtle platform fact -- with no coverage, and being macOS-gated it could not get any, since CI runs ubuntu only. The decision is now the pure `is_bundle_exe_launch(exe, argv0)`, compiled everywhere and table-tested for all four cases; only the process-state read stays gated. Same seam `extract_path` uses. The lifecycle lock moves ahead of the PATH probe. Taking it afterwards left the store unprotected for however long the probe ran -- up to 5s -- which is exactly when an `orx delete` racing app startup would land. A failed `read()` is now logged too; it was the one outcome in this module that could pass silently, and silently is how the app would then run unlocked. `orx delete` blocks on a running app now, so its message says so instead of naming three things the user isn't running. Probe diagnostics split the timeout from the spawn failure and name the shell, since "which shell, and did it hang or fail to start" is the whole question a report needs to answer. `bin_version` gets the `NO_COLOR` guard its siblings already carry: it feeds `parse_version`, an escape-laden version line parses as `None`, and `gate_oauth_version` turns that into `Unknown` -- reproducing the bug this branch exists to fix, via the synced env the probe now inherits. `find_opencode` skips empty PATH components, matching `find_on_path`. Packaging asserts the `orx` symlink survived staging. Every step preserves it today, but a dereferencing copy would fail silently -- the app builds, signs, notarizes and runs, and agents just quietly fall back to the user's own CLI. Docs: corrected the same overstatement in `search_path`'s module doc that DISTRIBUTION.md already fixed, and recorded the known gap that the probe imports only PATH, so `ORX_DATA_DIR`/`XDG_*` from a shell rc reach the CLI but not a Finder-launched app. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): import the shell's data and config dirs, not just PATH `ORX_DATA_DIR`, `XDG_DATA_HOME`, and `XDG_CONFIG_HOME` exported from a shell rc reached the CLI but never a Finder-launched app, so the two resolved different directories: the dashboard opened the default database while `orx` on the same machine used the user's, and the lifecycle lock added last commit guarded a file neither of them shared. Same root cause as the PATH bug, one layer down. `search_path` generalizes to `shell_env`: an allowlist (`IMPORTED`) the startup probe reads, NUL-separated so a directory can contain anything but NUL, behind `var()` -- process environment when no override is installed, so nothing changes for a terminal `orx`. `store::env_path` and `config::config_dir` are the only two places that resolve these, and both now go through it. The allowlist stays short on purpose: these are the variables whose divergence makes the app and the CLI act like separate installs, while credentials already reach harness children via `prepare_env`. This moves the lifecycle lock back after the probe, undoing last commit's reordering. The reason is now structural rather than incidental: the lock path is derived from `config_dir()`, so taking it before the probe would lock the default path while the CLI locks the user's -- protecting nothing. A correct lock a few hundred milliseconds late beats an immediate lock on the wrong file. Verified with the app's process environment stripped of both variables while the probe shell exported them: `orx.db` lands in the shell's data dir and `orx.lifecycle.lock` in the shell's config dir, and the HOME-derived defaults are never created. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): carry the imported environment into orx child processes Importing the shell's directories fixed the app's own resolution but not its children's, which made things worse rather than better for anyone exporting `ORX_DATA_DIR`. `spawn_detached_supervise` and the harness adapters spawn plain CLI processes; those never probe, so they re-resolved to the default store while the app read the user's. A run launched from the app would have been supervised against a different database than the dashboard was showing — before this branch both landed on the default and at least agreed. `shell_env::export_to` hands the imported variables to a child. `prepare_env` applies it, so an agent's `orx exp run` shares the dashboard's store, and `spawn_detached_supervise` applies it plus PATH, which also reaches the run payload it launches (`localbox` spawns `bash run.sh` with an inherited env, so it was losing the user's PATH entirely — no `python`, no `uv`, no `conda`). `opencode::spawn_agent` had hand-rolled a copy of `prepare_env`; it now calls it rather than growing a third copy of the same fix. Verified with the app's process environment stripped of both variables while the probe shell exported them: a harness child is handed ORX_DATA_DIR and XDG_CONFIG_HOME, and an `orx` run with that environment creates no default store. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cli): print a runnable command in the hints the bundle breaks A DMG user has no `orx` anywhere on PATH, so every hint telling them to run `orx exp wait …` names a command their shell cannot find. Agents are already covered -- `prepare_env` puts the bundle's `orx` first on their PATH -- but a human driving the bundle's CLI is not. `invocation::orx()` resolves the spelling once: plain `orx` when it is on PATH, which is every CLI install and every agent, otherwise this binary's own path, shell-quoted when it needs it. Nothing changes for anyone who installed the CLI. Applied to the three hints the report named, each of which had been copied per backend: the post-launch "Follow it with …" line (nine copies), the missing run-command error (eight), and the outdated-version warning. Both hint bodies move into `invocation`, so the eight backends now share one line each -- a net deletion, and one place to fix the next time the wording moves. `warning_for` takes the spelling as an argument rather than reading process state, keeping its tests deterministic. Not covered: the long tail of `orx login`-style hints in error strings, which still print a literal `orx`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): close the environment gaps review found A dashboard-synced `ORX_DATA_DIR`/`XDG_*` silently beat the shell-imported one for harness children: `prepare_env` exported the imported vars, then the synced loop overwrote any whose key was absent from the *process* env -- which in app mode is exactly these. The app read one store while claude and opencode wrote another. The guard now consults `shell_env::var`, so the imported value counts as present. codex escaped this only because it re-pins `ORX_DATA_DIR` after `prepare_env`; claude and opencode have no such pin. Local run payloads were never actually fixed. The previous commit claimed the supervisor passed its environment to `bash run.sh`, but the supervisor does not spawn it -- `localbox::run_job` does, in the launching process, from `LocalJobSpec::env`. A run started from the app still executed the user's script under launchd's PATH, with no python, uv, or conda. The env now carries the imported PATH and directories, and the supervisor comment drops the wrong claim. Three more readers of these same variables were still going to the process env, and two of them made harness detection itself diverge between app and terminal: `harness::xdg_config_home` (whose doc claims to mirror `config::config_dir`) and `opencode_auth_path`, which would report opencode signed out in the app while the CLI saw it signed in. `invocation::resolves_on_path` was the third -- it decides whether a hint can say plain `orx`, and the reader pastes that into the shell whose PATH we imported. Also: `orx login` -- the first hint a fresh DMG user meets -- now resolves like the rest; `export_to` iterates in `IMPORTED` order so the startup log is diffable; and the doc drift is repaired, including the octal-escape invariant that makes the nonce's leading underscore load-bearing, the reopened `orx delete` window, and the scope note listing what still inherits the process environment (`git`, `gh`, `kubectl`, `ssh`, `publish-branch`). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(app): allow dead_code for the launch check off macOS `is_bundle_exe_launch` is deliberately un-gated so its tests run on CI's Linux runner, but its only non-test caller is `#[cfg(target_os = "macos")]`, so the Linux bin target sees it as dead. `mod local` carries a blanket `#[allow(dead_code)]`, which is why `shell_env::parse_probe` -- same shape -- did not trip; `mod commands` does not. Verified by rebuilding with every `target_os = "macos"` in `src/` swapped to `"linux"`, which reproduces exactly what the runner cfgs out. Clean, and no other item on this branch is dead off macOS. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f460a71af6 |
OR-162 Inline code editing, open-in-editor, and editable file browsers (#194)
* feat(ui): inline code editing, open-in-editor, and editable file browsers Make right-pane code files editable: click into the code and type, save on ⌘S or blur. Adds a live-checkout write endpoint, editor detection + open-in- editor, and a highlighted-overlay editor. The experiment code browser opens its live worktree (editable) when checked out on the branch, and always labels the correct branch. Fixes clipped end-aligned tooltips. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): address code-review blockers - write/open endpoints reject paths under .git and a stale session worktree that fell back to the clone; open rejects directories - editors.rs opens via explorer.exe on Windows (no cmd.exe metachar reparse) - preserve original CRLF line endings on save; compare/seed the buffer in LF - don't clobber keystrokes typed during an in-flight save (skip the optimistic reseed); clone-fallback session files stay read-only - CodeEditor re-navigates on file:line chips into an already-open file - use `instanceof Error` instead of `as Error` assertions Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(review): case-insensitive .git guard; trim comments Round-2 nits: match .git case-insensitively so .GIT can't slip past on case-insensitive filesystems, and trim two comments to the 1-2 line convention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
5412465c11 |
OR-170 Fix onboarding paper titles and the reproduce-paper starter (#184)
* Fix onboarding paper titles and reproduce-paper starter - Onboarding: resolve the canonical paper title after adding, instead of storing the Google-scraped fast-search hit (truncated / id-prefixed / reworded). Falls back to the cleaned search title if resolve fails. - Reproduce-paper starter: drop the dangling "on " and the redundant pre-filled paper id. The linked paper is already injected into every harness playbook and compute defaults to the configured target, so the armed skill chip + placeholder is enough. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Rebuild ui/dist Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: reconcile placeholder, prefer linked paper in no-args - ChatPanel: placeholder + comments now state compute is optional (was contradicting the starter change); tighten the onClick comment. - skills.rs: bare /reproduce-paper and /paper-to-marimo now prefer the linked paper from the playbook before repo inference — the reproduce starter relies on this no-args path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8c17722d88 |
feat(macos): styled drag-to-Applications DMG installer window (#188)
Opening OpenResearch.dmg now shows OpenResearch.app and an /Applications alias to drag it onto, over a branded background — instead of a bare app in the DMG root. - scripts/generate-dmg-background.mjs: pure-Node PNG background (no image deps, mirroring generate-icon.mjs) with the brand mark + a drag arrow, rendered at 1x and 2x and combined into a HiDPI background.tiff. - scripts/package-macos-app.sh: stage app + Applications symlink + background, lay out the Finder window (size, icon positions, background) via AppleScript on a read-write image, then flatten to the compressed DMG. Styling is best-effort: if the Finder step fails (e.g. headless), a functional but unstyled DMG still ships, and the temp/mount cleanup runs regardless. - release-macos-app.yml: bound the 10x-billed macOS job with a timeout. - DISTRIBUTION.md: document the installer window and its constants. Verified locally: styled DMG builds with the app + Applications alias + branded background, 640x400 window; happy-path cleanup leaves no leaked temp dirs or mounts. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4f0fb0bf12 |
fix(release): attach macOS DMG to releases (repo var + PAT cascade + manual dispatch) (#186)
The signed macOS DMG never made it onto releases, so releases/latest/download/OpenResearch.dmg 404s. Two root causes: 1. The cheap gate reads `vars.MACOS_SIGNING_ENABLED`, which only sees repo/org variables, but the variable was created at environment scope. Fixed out-of-band by adding the repo-level variable. 2. `release-on-bump` dispatches `release.yml` with GITHUB_TOKEN, so GitHub's anti-recursion rule suppresses the `workflow_run` cascade and the macOS attach never runs for real releases (only PR dry-runs, which correctly skip). The workflow's comment claiming `workflow_run` fires regardless of the upstream author was wrong. Changes: - release-on-bump.yml: dispatch release.yml with a PAT (RELEASE_DISPATCH_TOKEN) so the Release run is owned by a real token and its completion cascades to the attach. Falls back to GITHUB_TOKEN when the PAT is missing or invalid so a bad PAT never blocks a release (the DMG is then a manual dispatch). - release-macos-app.yml: add a `workflow_dispatch` (tag input) trigger as the manual/recovery path; the gate and checkout handle both trigger types. This is how releases cut before the pipeline worked get their DMG. - DISTRIBUTION.md: document the repo-variable scope, the PAT secret and rotation, the manual re-attach command, and the required-reviewer gate. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
dccc37c4e8 |
feat: downloadable macOS app (OpenResearch.app) + signed release hosting (#180)
* feat: scaffold downloadable macOS app (OpenResearch.app) Adds a native macOS `.app` bundle whose executable IS the `orx` binary. Launched from the bundle (double-click, no args) it enters GUI "app mode": an AppKit NSApplication + delegate run loop on the main thread — Dock icon and "OpenResearch" name from the bundle's Info.plist + .icns, and a Dock-icon click that reopens the dashboard in the browser — while the `orx up` server runs on background tokio worker threads. App mode is entered only when the executable lives in `.app/Contents/MacOS` AND argv is empty, so the bundled binary is still usable as a CLI. - src/commands/app.rs: bundle detection + AppKit delegate + background server - macos/Info.plist: bundle metadata (name, icon, identifier, version) - scripts/generate-icon.mjs: transparent-PNG rasterizer (from favicon.svg) - scripts/build-macos-app.sh: builds release orx, generates .icns, assembles OpenResearch.app into dist/ - objc2/objc2-app-kit gated under cfg(target_os = "macos") so musl Linux is untouched The bundle is unsigned; code-signing + notarization is a follow-up. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat: sign, notarize, and host OpenResearch.app via GitHub Releases Automates distribution of the macOS app as a signed, notarized DMG attached to each GitHub Release: - scripts/build-macos-app.sh: ORX_APP_UNIVERSAL=1 builds a universal (arm64 + x86_64) binary via lipo for distribution. - scripts/package-macos-app.sh: codesigns (hardened runtime, inside-out), notarizes + staples the .app (so it launches offline once dragged out of the DMG), packages a DMG, then notarizes + staples the DMG. Env-gated: runs unsigned locally, fully signed in CI. - .github/workflows/release-macos-app.yml: on the "Release" workflow completing (workflow_run — a GITHUB_TOKEN-created release: published event can't trigger workflows), a cheap ubuntu job gates on a real dispatched release + all signing secrets + the release existing, then a macOS job builds/signs/notarizes and uploads OpenResearch.dmg to that release. - macos/DISTRIBUTION.md: the one-time Apple setup + the six repo secrets. Inert until the signing secrets are configured, so it is safe to merge before the Apple Developer account exists. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: gate macOS signing behind CODEOWNERS + a required-reviewer environment Hardens the release-signing pipeline for a public repo: - .github/CODEOWNERS: marks the release workflows and macOS signing scripts as owned so they can't change unreviewed (with branch protection's "Require review from Code Owners"). - release-macos-app.yml: the cert-using job now runs in the `release-signing` environment, so adding required reviewers to it pauses signing for human approval — the Developer ID cert is never used by an unreviewed change. - DISTRIBUTION.md: documents creating the environment + branch protection. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: add @sox8502 to CODEOWNERS Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * security: environment-scoped signing secrets + SHA-pinned actions Hardens the release-signing pipeline: - The 6 signing secrets move from repo secrets to the `release-signing` environment, so only the reviewed, environment-gated macos-app job can read them — the cert secrets never exist in the cheap ubuntu gate. - The gate now keys on a non-secret repo variable MACOS_SIGNING_ENABLED instead of reading the secrets to detect configuration. - actions/checkout and actions/setup-node are pinned to commit SHAs so a moved tag can't inject code into the job that holds the Developer ID cert. - DISTRIBUTION.md updated: create the environment (required reviewers + main-only deployment branches), add environment secrets, set the variable. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * docs: trim DISTRIBUTION.md to repo-specific config Drop the generic Apple-portal walkthrough (enrolment, cert creation, export click-by-click) — Apple documents that. Keep only what's specific to this repo: the environment/secrets/variable, the local build+sign commands, and the download URL. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a94c399458 |
OR-166 Queue chat messages sent while a turn is running (mid-run steering) (#178)
* Queue messages sent while a turn is running
Sending a chat message while a turn is in flight previously errored with
"session is busy — interrupt it first". Instead, park the message and run
it automatically when the current turn finishes — Claude-desktop-style
mid-run steering. Lives in ChatHost above the harnesses, so Claude, Codex,
and OpenCode are all covered.
Backend: a per-session in-memory queue on ChatHost; send_message enqueues
on a busy claim and emits chat.queued; a natural turn completion drains the
next message (FIFO); user Stop and session delete clear the queue; a new
DELETE .../queue/{itemId} cancels one parked message; the messages endpoint
returns the queued snapshot for reload recovery.
Frontend: send() parks instead of bailing while busy; queued messages render
as chips above the composer with a cancel ✕; chat.queued drives the reducer;
the seed snapshot restores chips after a reload.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Address review: drop vs re-park on drain failure, clear queue on project delete
- drain_queue re-parks a message only on a genuine busy-race (is_busy after
the failed send); genuine setup errors and deleting sessions now drop+log
instead of stranding a chip that re-fails on every future drain.
- delete_project clears each session's queue (parity with delete_session), so
no parked entries orphan a deleted project.
- Image/file-only parked messages get an "N attachment(s)" chip label instead
of a blank one.
- Import ordering + comment trims from review.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* Coalesce queued messages into one turn
When several messages are parked behind a running turn, drain them all into a
single combined turn (texts joined by blank lines, attachments concatenated,
most-recent composer overrides win) instead of running a full turn per message
— so successive steering prompts run together as soon as the turn ends.
Re-park on a busy-race restores the individual messages (front, arrival order)
so they re-coalesce on the next drain.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
4738715e94 |
OR-134 Add Skills tab: upload, import, and /-invoke agent skills (#175)
* Add Skills tab: upload, /-invoke, and import agent skills Adds a left-rail Skills section where users manage custom agent skills: - Upload a SKILL.md or a .zip of a skill folder (Global or per-project). - Import skills already installed in the user's coding agents. - Uploaded skills are written into each session worktree beside the built-in orx-* skills and surfaced in the composer's / menu. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Rebuild ui/dist for the Skills tab Regenerated committed assets (`pnpm build`) after the Skills tab UI and the merge with main's Tailwind migration, per AGENTS.md. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address code review on the Skills tab - Prune deleted/renamed user skills from session worktrees via a manifest, and replace (not merge) each skill dir so shadows and re-uploads are clean. - Bound zip entry reads to the extract budget (deflate-bomb guard). - Accept a quoted SKILL.md `name:`; import keys strictly off the frontmatter name so folder==name always holds. - Dedupe the composer `/` menu (project shadows global); drop dead helpers. - Surface delete errors, guard re-entrant uploads, a11y + timeAgo tweaks. - Add tests for prune, clean shadow, zip caps, and quoted names. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Rebuild ui/dist after review fixes Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
e53a84d3be |
OR-165 Consolidate composer controls behind a settings switch icon (#173)
* Consolidate composer controls behind a settings switch icon Move Mode, Model, Reasoning, and Literature sources out of the composer bar and into a single popover opened by the switch icon. The bar now shows just the switch, attach, context meter, and send. - Nested pickers open as right-side flyouts so they never clip or overlap the panel's own rows. - Literature sources render inline (no extra click) and are cached across opens so reopening the panel doesn't flash "Loading…". - Unify hover styles across all composer controls: a shared surface-fill that fades in over 150ms with matching radius; the attach icon is the same text color as the others and no longer shifts color on hover. - Keep the model label on one line (no wrap) and align every row's chevron. Rebuilds the embedded ui/dist. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Move model picker to the composer bar with a harness logo Take the model picker out of the settings popover and place it back on the bar's right side, next to the context-window meter. Its menu opens right-aligned again so it doesn't run off the edge. - Show the in-use harness's brand mark on the model-picker trigger. The onboarding agent-card logos are extracted into a shared HarnessLogo component (single source of truth) that both call sites use. - Rename the settings "Mode" row and its dropdown header to "Permissions". Rebuilds the embedded ui/dist. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Re-implement composer settings popover in Tailwind Rebuild the composer redesign against the new inline-Tailwind UI (the prior CSS-file implementation was obsoleted by the styles.css → Tailwind migration merged from main). - Permissions, Reasoning, and Literature sources live behind the switch icon in a settings popover; nested pickers open as right-side flyouts via arbitrary-variant overrides on the panel; sources render inline. - Model picker stays on the bar's right and shows the in-use harness's brand mark; the onboarding agent logos are extracted into a shared HarnessLogo component. - Unify hover across all composer controls (switch, attach, pills, bare, context meter): a surface fill that fades in over 150ms; the attach icon matches the others' color and no longer shifts color on hover. Rebuilds the embedded ui/dist. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9851cf548a |
OR-144, OR-164 Better evidence-based chats, summaries of long-horizon research (#171)
* Add evidence chips linking chat claims to their source Chat markdown now renders `<run id label/>` and `<file path lines exp/>` tags as clickable evidence chips so a claim links to what backs it: run chips open the run's logs, file chips open the source scrolled to and highlighting the cited line, and the code tab labels where the file came from — the cited experiment, or "baseline". - Md.tsx: unify `<file>`/`<run>` mention parsing; RunChip; FileChip forwards line + exp - CodeView.tsx: scroll-to-line + highlight band - FileViewer.tsx / App.tsx: source badge via fileSourceLabel (experiment title vs "baseline") - ChatPanel.tsx / App.tsx: thread onOpenRun and the widened onOpenFile - SYSTEM_PROMPT.md: guidance to cite evidence for claims Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Show the branch name on code tabs instead of experiment/baseline The code-tab header now always names the git branch its contents came from — an experiment's branch (orx/…) or the baseline (main) — replacing the experiment-title-vs-"baseline" badge that read inconsistently. A cited experiment still pins the file to its committed branch; only the label is simplified to the branch itself. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Prompt: close experiment turns with a summary of runs Tell the agent to end any turn that ran or changed experiments with a short per-node summary (what it tested, status, headline result) backed by the evidence chips, so the user — who didn't watch the runs — can reorient instead of getting lost. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: stable openFileTab, subagent run chips, cleanup - openFileTab reads experiments through a ref (like runsRef) instead of a dep, so it stays referentially stable and doesn't churn the memoized chat transcript on experiment polls - thread onOpenRun into sub-agent transcripts so <run> chips work there - delete the now-orphaned .file-view-ref CSS (superseded by .file-view-branch) - note why CodeView's highlight effect depends on `rendered` Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
b4ee41ee39 |
OR-150 Add copy button to onboarding command text (#162)
* Add copy button to onboarding command text Commands shown to the user in onboarding (e.g. `claude auth login` in an agent card's `agentNote`) are now copyable: renderNote emits a code pill with a copy button, shared across onboarding, the model picker, and the chat-panel harness warning. An unready agent card is rendered as a plain <div> instead of a disabled <button> — a disabled button suppresses pointer events on its children, so the nested copy button could not be clicked; interactive cursor/hover styles are scoped to button.onb-agent-choice so the inert card stays inert. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Bump version to 0.1.95 Merging this release-triggers a v0.1.95 tag/release. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
9508aed337 |
Add System/Light/Dark theme toggle to Settings (#161)
The app previously followed the OS via `prefers-color-scheme` only, with no way to override or persist a theme. Add an Appearance section to Settings with a System/Light/Dark segmented control. - theme.ts: singleton store (useSyncExternalStore) that persists the preference to localStorage `orx:theme`, stamps `data-theme` on <html>, and keeps System mode live-syncing to the OS app-wide. - index.html: pre-paint inline script resolves the theme before first render to avoid a flash of the wrong palette. - styles.css: convert the two `@media (prefers-color-scheme: dark)` blocks to `:root[data-theme="dark"]` so JS drives the palette; add the segmented-control styles. - SettingsPage.tsx: AppearanceTab with an accessible radiogroup (roving tabindex + arrow-key navigation, focus follows selection). Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
899c3b68ac |
Add PDF/file upload to the chat composer (#159)
* Add PDF/file upload to the chat composer Add a paperclip button (and paste/drag-drop) that attaches PDFs and images to a chat message. Files ride the existing base64 attachment pipeline, are saved to disk preserving the original name, and their paths are injected into the turn as <attached-files> so the agent Reads them (Claude Code handles PDFs natively). Raise the API request-body limit to 64 MB (axum's 2 MB default rejected any real paper, surfacing as a client NetworkError) and add a 30 MB per-file client guard with an inline error. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Rebuild UI bundle for PDF/file upload Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Address review: aggregate size guard, a11y, theme color, comments Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * cargo fmt Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
a80f1c97ae |
Better paper integration + lit review chat view (#155)
* Add bioRxiv and OpenAlex sources to orx lit / orx paper `orx lit` and `orx paper` covered only alphaXiv (arXiv corpus, not biomed). Add OpenAlex (general scholarly graph) and bioRxiv (biology preprints) so lit reviews reach beyond arXiv. - `orx lit --source alphaxiv|openalex|biorxiv` (default alphaxiv, so existing behavior is unchanged). bioRxiv has no search API, so `--source biorxiv` searches OpenAlex filtered to bioRxiv's corpus (S4306402567). - `orx paper <id>` auto-detects the source from the id (override with `--source`): arXiv id -> alphaXiv report/--full; 10.1101/... DOI -> bioRxiv; any other DOI or a W... id -> OpenAlex. A DOI is recognized only when it carries the mandatory '/', so October arXiv ids (e.g. 1810.04805) are not mistaken for DOIs. - `orx lit --json` now emits a uniform LitHit shape across all sources (adds `source`/`citations`, preserves alphaXiv `votes`/`snippets`). - OpenAlex/bioRxiv print a title/authors/date/citations + abstract card with DOI and PDF links; they have no extracted full text, so `--full` points at the PDF. All three sources are public (no login). New config hosts: OPENALEX_API_URL, BIORXIV_API_URL, OPENALEX_MAILTO. Docs (orx-lit skill, SKILL/SYSTEM_PROMPT/README, lit-review template) updated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Add Settings toggles to enable/disable literature sources Adds a "Literature sources" section to the dashboard Settings with on/off toggles for alphaXiv, OpenAlex, and bioRxiv. The choice is hard-enforced: `orx lit` and `orx paper` refuse a disabled source. - Persisted in settings.json as `disabledLitSources` (empty = all enabled, so a source added later defaults on). New `telemetry` getter/setter via the locked `mutate_settings` RMW; `config::disabled_lit_sources` re-export. - `GET`/`POST /api/settings/lit-sources` returns/accepts `{alphaxiv, openalex, biorxiv}` booleans, mirroring the profile settings handlers. - UI `LiteratureSourcesTab` (built from the ProjectDefaultsTab template) in the Settings stack; `getLitSources`/`setLitSources` in api.ts. Rebuilt ui/dist. - Enforcement: `orx lit --source <disabled>` errors; bare `orx lit` falls back to the first enabled source (noting the swap on stderr) and errors if all are off; `orx paper <id>` on a disabled source errors (no fallback — the id is fixed). `LitArgs.source` is now `Option<LitSource>` to distinguish an explicit choice from the default. `LitSource::as_str` centralizes the wire name. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Render orx lit / orx paper chat rows as a real search In chat, a literature search rendered as a generic "Ran orx lit …" shell line. Now `orx lit` / `orx paper` tool calls read as a real search — the source's official logo plus natural language ("Searching OpenAlex for "…"", "Reading 2401.12345 on bioRxiv") — so a lit review feels first-class. - New `LitSourceLogo.tsx`: the three official brand SVGs (inlined at build via `?raw`, shown in a small white tile so the solid-black OpenAlex/bioRxiv marks stay visible in dark mode) plus `parseOrxLit`, which recognizes an `orx lit`/`orx paper` command and pulls out the source + query/id. `detectPaperSource` mirrors the Rust `detect_source` so a bare `orx paper <id>` shows the right source. - `ChatPanel` gains `toolSummary`: Bash rows matching an orx literature command render the logo + sentence; everything else falls back to the plain `toolLine`. The row still expands to the exact command + output — nothing is hidden. - The text ellipsizes like other tool rows; the logo is aria-hidden (the source name is in the adjacent text). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Make orx paper chat rows link to the paper's source page Clicking a fetched paper in chat now opens it on its own source: an arXiv id → alphaXiv (alphaxiv.org/abs/<id>), a bioRxiv DOI → doi.org (resolves to bioRxiv), and OpenAlex → the DOI or the OpenAlex work page. A small external-link arrow marks the row as clickable; the row still expands to the raw command + output. - `paperUrl(source, id)` in LitSourceLogo builds the per-source URL; `doiFrom` extracts a bare DOI from an id or URL and strips a bioRxiv content-URL version suffix (`v1.full`) so doi.org resolves it. - `toolSummary` renders `orx paper` rows (with an id) as an `<a target=_blank rel=noopener>`; search rows stay plain text. The href scheme/origin are fixed literals, so the agent-supplied id can only ever be a path segment. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Polish lit-review UI: drop redundant source name, logo the settings toggles - Chat: `orx lit`/`orx paper` rows now read "Searching for "…"" / "Reading <id>" — the logo already names the source, so the text no longer repeats it. The logo carries the source name for screen readers (aria-label) when no adjacent text does; a new `decorative` prop keeps it aria-hidden where text names it. - Settings → Literature sources: each toggle row now shows the source's official logo next to its name. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Move literature-source toggles from Settings to the composer The source on/off toggles were tucked away in Settings. Surface them where a lit review happens: a switch icon at the bottom-left of the chat composer opens a small popover to toggle alphaXiv / OpenAlex / bioRxiv. State still lives in settings.json via /api/settings/lit-sources, so `orx lit`/`orx paper` enforcement is unchanged. - New `LitSourcesPicker` mirrors the composer's OptionPicker pattern (usePopover + option-menu); rows are `role="switch"` with the source logo + name. - Remove the Literature sources section (and its dead styles) from SettingsPage. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Drop redundant per-row status dots inside tool groups Rows expanded under "Used N tools" already sit inside the group's status dot and indent rule, so the leading per-row dot was just noise. Hide it for grouped rows only; the group summary dot and single-row dots are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
601a201d0e |
Onboarding: research-background step; drop telemetry step (#152)
* Remove usage-analytics onboarding step The first-run walkthrough is now two steps — coding agents → GitHub — instead of three. GitHub becomes the final step and finishes onboarding directly via onDone; the data-location disclosure moves onto it. Telemetry stays opt-out and is controlled from the CLI (`orx telemetry off` / `--no-telemetry`); it no longer prompts during onboarding. Removes the now-dead telemetry UI API client exports and rebuilds the embedded ui/dist bundle. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Add research-background onboarding step Third onboarding step (optional): a free-text background blurb plus alphaXiv papers linked via title search (reusing the paper-search infra). Prefills from any saved profile and saves best-effort on finish, so an empty or failed save never blocks completing onboarding. Persisted locally in settings.json via a new GET/POST /api/settings/profile (`background` + `linkedPapers`), sharing the same locked read-modify-write as the other settings so it can't clobber install_id/telemetry_disabled. The blurb is user free-text kept on local disk — never sent in telemetry and outside the snapshotted data dir. Not yet consumed by the agent. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Myles Anderson <mylesanderson@Myless-MacBook-Pro.local> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |