Myles AndersonandClaude Opus 5 a9af660f9e OR-216 Add an orx-figures skill for publication-quality figures (#260)
* Add an orx-figures skill for publication-quality figures

Figures were left to matplotlib defaults: screen-sized, titled where a
caption belongs, and rasterized where the document wants vector. This adds
a bundled skill that routes by the question a figure answers and ships the
style layer rather than describing it.

`orx-figures` follows the `orx-compute` shape — one SKILL.md carrying the
non-negotiables, then exactly one of six references (curves, scaling,
comparison, pareto, matrix, diagram). Two assets install beside them:
`orx_figstyle.py`, vendored into the project so a figure stays reproducible
after the session ends, and a TikZ scaffold sharing its palette.

`save()` audits every figure and prints the result: printed width against
the known column widths, Type 42 font embedding, stray axes titles, missing
axis labels, sub-5pt text, overlapping text, and text running off the
canvas. Prose the agent can skip is how the first unusable figures shipped;
these checks run whether or not the guidance was read.

The playbook and the `orx skill` overview both direct figure work here — a
session with the skill installed still plotted raw matplotlib until they
did.

Content is calibrated against 39 popular recent alphaXiv papers (283
captions) and 10 arXiv e-print sources: 87% of real figure PDFs are built
wider than any single-column page, so the width check is the load-bearing
one; multi-panel is 27% of figures; and captions run a median of 28 words.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix review findings in the figures skill

Correctness:
- `label_ends` computed the reserved x-room in the pre-`set_xlim` scale, so
  the labels overhung the spine by need^2/(width+need) — about a quarter of a
  long label. Solve for the limit at which the original span occupies
  (axes width - label width) pixels instead.
- The audit's categorical-axis exemption keyed off whether tick text parsed as
  a number, but log formatters emit mathtext, so every log axis was classified
  categorical and escaped the missing-axis-label check. Decide by scale.
- `save()` audited before writing, and the audit needs an Agg-family renderer
  that is not guaranteed in the very case its first check reports. Write the
  files first; the audit can no longer be the reason a figure is lost.
- `curves.md` interpolated every seed onto the first seed's grid, and
  `np.interp` clamps, so a run that stopped early gained a fabricated flat
  tail. Clip to the range every seed reached.
- `comparison.md` hardcoded `ylim(0, 90)`: any score above it drew as a bar
  clipped at the axes edge, which is the dishonesty that reference is about.
- `scaling.md` zipped families against three markers, silently dropping a
  fourth family from the main panel while still plotting its residuals.
- Guards for degenerate input: zero-variance `diff_ci`, all-NaN heatmaps,
  empty confusion rows, resamples that fail to fit, empty label lists.

Consistency:
- `diagram.md` still taught the retired natural-size include, and its inline
  path dropped `\familydefault`, which would render a serif diagram beside
  sans plots. Both contradicted SKILL.md.
- `orx-paper` and `orx-reports` examples named a rasterized plot.
- The audit emits a font-fallback finding the docs never listed, and it is the
  one finding an agent should hand over rather than fix.

Tests now assert the rendered playbook rather than the raw template, and pin
the audit's problem strings rather than text that also appears in docstrings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Fix round-two review findings

The round-one TikZ fix was worse than the bug: it told the agent to move
`\renewcommand{\familydefault}{\sfdefault}` into the paper's preamble, which
is document-wide and would have set the body text, headings, and captions of
a submission to sans. The face now rides in the `orx`, `edgelbl`, and
`stagelbl` styles, so it travels with the picture and touches nothing else.

Also from review:
- `diagram.md` still carried the broken relative `cp assets/...`, which does
  not resolve from a session's cwd.
- The new "write a report figure to the artifacts directory" note contradicted
  "keep the generating script beside its output"; both destinations now keep
  the pair together.
- A `WIDE` figure only gets a 6.75in `\linewidth` inside `figure*`; in a plain
  `figure` it silently halves, which non-negotiable #1 now says.
- Skipping all text on an axis-off axes also exempted the user's own content,
  which on such an axes is the figure. Exclude only the undrawn axis artists.
- `save()` leaked the figure when the audit raised, and did not create the
  stem's parent directory.
- A single seed drew a zero-height band that reads as certainty; `mean_ci` and
  `diff_ci` now warn, matching what the references demand of the caption.
- Guards for disjoint seed ranges and all-zero scores.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Close the round-three review findings

None were blockers, but three were real defects on secondary paths:

- The inline `\input` route omitted `\usepackage{tikz}` and the
  `\usetikzlibrary` line, so a paper following it literally would not compile:
  `trapezium`, `Stealth`, `fit=`, and `on background layer` each need one.
- Vendoring was hardcoded to `figs/`, so a report script written under the
  artifacts directory could not import the style module beside it.
- Removing `\familydefault` left no documented way back to a serif diagram
  while `use_style(family="serif")` still exists for plots.

The colorbar guard now accepts `ax.get_label() == "<colorbar>"` as well as the
private `_colorbar`; verified against matplotlib 3.11 that both hold today, so
a rename can no longer make `clean` unreachable for every heatmap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 18:19:47 -07:00
…
…
…
2026-08-24 19:19:08 -07:00
…

OpenResearch

The local-first workspace for research agents and autoresearch.

Turn Claude Code, Codex, or OpenCode into research agents that can review literature, develop hypotheses, run experiments, and produce research artifacts.

Download the desktop app · Documentation · Releases

Get started

Download the local desktop app from openresearch.sh/download, or install the CLI on macOS or Linux:

curl -LsSf https://openresearch.sh/install.sh | sh
orx up

orx up opens the local dashboard at http://127.0.0.1:4791.

Create an account at openresearch.sh to receive email updates and use managed OpenResearch compute.

Built for research agents

OpenResearch gives you
Parallel exploration Give each research direction an independent agent session and isolated git worktree.
Reproducible experiments Track variants in a git-native experiment tree; every run receives an immutable archive of its recorded commit.
Evidence in context Keep logs, diffs, files, results, and artifacts tied to the work that produced them.
Your choice of agent Use Claude Code, Codex, or OpenCode, with the harness and model selected per session.
Your choice of compute Run locally, on your own infrastructure, or with managed OpenResearch compute.
Local ownership Keep projects, conversations, experiments, runs, logs, code, and artifacts on your machine.

Autoresearch

OpenResearch can run the full loop autonomously: propose an idea, change the code, launch an experiment, inspect the evidence, and decide what to try next. Multiple agents can explore different directions in parallel while the experiment tree preserves their lineage.

Run anywhere

The same committed source snapshot can run locally, over SSH, or on Slurm, Kubernetes, Ray, Hugging Face Jobs, Modal, Tinker, and managed OpenResearch compute. Publishing the repository is not required.

Run the workspace next to remote GPUs while using the browser on your laptop:

orx up --remote user@host

SSH config aliases and custom ports are supported. The remote service binds to loopback and has no application-level authentication, so other users on that host can reach it.

CLI and agent integration

Install the OpenResearch skill into supported coding agents:

orx install-skills

Common commands:

orx projects
orx project view <project-id>
orx runs <project-id>
orx logs <run-id>
orx exp run <experiment-id>
orx discover keyword <query>
orx paper <arxiv-id-or-doi>

Run orx --help or orx <command> --help for the complete interface.

Local by default

OpenResearch runs on 127.0.0.1 with a local SQLite store. Creating a project or launching a run does not publish your code. An openresearch.sh account is only used for service-owned capabilities such as organizations and managed compute.

Usage analytics

Official release builds send opt-out, coarse usage events tied to a random installation ID. They do not include code, prompts, file contents or paths, repository names, tokens, emails, or project and experiment identifiers.

orx telemetry off
orx telemetry status
orx <command> --no-telemetry

Source and development builds do not send analytics.

S
Description
GitHub Trending: alphaXiv/OpenResearch
Readme MIT
124 MiB
Languages
Rust 70.4%
TypeScript 24.2%
JavaScript 3.5%
Python 0.9%
Shell 0.4%
Other 0.3%