* Add an orx-figures skill for publication-quality figures Figures were left to matplotlib defaults: screen-sized, titled where a caption belongs, and rasterized where the document wants vector. This adds a bundled skill that routes by the question a figure answers and ships the style layer rather than describing it. `orx-figures` follows the `orx-compute` shape — one SKILL.md carrying the non-negotiables, then exactly one of six references (curves, scaling, comparison, pareto, matrix, diagram). Two assets install beside them: `orx_figstyle.py`, vendored into the project so a figure stays reproducible after the session ends, and a TikZ scaffold sharing its palette. `save()` audits every figure and prints the result: printed width against the known column widths, Type 42 font embedding, stray axes titles, missing axis labels, sub-5pt text, overlapping text, and text running off the canvas. Prose the agent can skip is how the first unusable figures shipped; these checks run whether or not the guidance was read. The playbook and the `orx skill` overview both direct figure work here — a session with the skill installed still plotted raw matplotlib until they did. Content is calibrated against 39 popular recent alphaXiv papers (283 captions) and 10 arXiv e-print sources: 87% of real figure PDFs are built wider than any single-column page, so the width check is the load-bearing one; multi-panel is 27% of figures; and captions run a median of 28 words. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix review findings in the figures skill Correctness: - `label_ends` computed the reserved x-room in the pre-`set_xlim` scale, so the labels overhung the spine by need^2/(width+need) — about a quarter of a long label. Solve for the limit at which the original span occupies (axes width - label width) pixels instead. - The audit's categorical-axis exemption keyed off whether tick text parsed as a number, but log formatters emit mathtext, so every log axis was classified categorical and escaped the missing-axis-label check. Decide by scale. - `save()` audited before writing, and the audit needs an Agg-family renderer that is not guaranteed in the very case its first check reports. Write the files first; the audit can no longer be the reason a figure is lost. - `curves.md` interpolated every seed onto the first seed's grid, and `np.interp` clamps, so a run that stopped early gained a fabricated flat tail. Clip to the range every seed reached. - `comparison.md` hardcoded `ylim(0, 90)`: any score above it drew as a bar clipped at the axes edge, which is the dishonesty that reference is about. - `scaling.md` zipped families against three markers, silently dropping a fourth family from the main panel while still plotting its residuals. - Guards for degenerate input: zero-variance `diff_ci`, all-NaN heatmaps, empty confusion rows, resamples that fail to fit, empty label lists. Consistency: - `diagram.md` still taught the retired natural-size include, and its inline path dropped `\familydefault`, which would render a serif diagram beside sans plots. Both contradicted SKILL.md. - `orx-paper` and `orx-reports` examples named a rasterized plot. - The audit emits a font-fallback finding the docs never listed, and it is the one finding an agent should hand over rather than fix. Tests now assert the rendered playbook rather than the raw template, and pin the audit's problem strings rather than text that also appears in docstrings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Fix round-two review findings The round-one TikZ fix was worse than the bug: it told the agent to move `\renewcommand{\familydefault}{\sfdefault}` into the paper's preamble, which is document-wide and would have set the body text, headings, and captions of a submission to sans. The face now rides in the `orx`, `edgelbl`, and `stagelbl` styles, so it travels with the picture and touches nothing else. Also from review: - `diagram.md` still carried the broken relative `cp assets/...`, which does not resolve from a session's cwd. - The new "write a report figure to the artifacts directory" note contradicted "keep the generating script beside its output"; both destinations now keep the pair together. - A `WIDE` figure only gets a 6.75in `\linewidth` inside `figure*`; in a plain `figure` it silently halves, which non-negotiable #1 now says. - Skipping all text on an axis-off axes also exempted the user's own content, which on such an axes is the figure. Exclude only the undrawn axis artists. - `save()` leaked the figure when the audit raised, and did not create the stem's parent directory. - A single seed drew a zero-height band that reads as certainty; `mean_ci` and `diff_ci` now warn, matching what the references demand of the caption. - Guards for disjoint seed ranges and all-zero scores. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Close the round-three review findings None were blockers, but three were real defects on secondary paths: - The inline `\input` route omitted `\usepackage{tikz}` and the `\usetikzlibrary` line, so a paper following it literally would not compile: `trapezium`, `Stealth`, `fit=`, and `on background layer` each need one. - Vendoring was hardcoded to `figs/`, so a report script written under the artifacts directory could not import the style module beside it. - Removing `\familydefault` left no documented way back to a serif diagram while `use_style(family="serif")` still exists for plots. The colorbar guard now accepts `ax.get_label() == "<colorbar>"` as well as the private `_colorbar`; verified against matplotlib 3.11 that both hold today, so a rename can no longer make `clean` unreachable for every heatmap. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
3.1 KiB
OpenResearch agent — {name}
You are an OpenResearch agent helping the user across the research process, including ideation, literature review, hypothesis formulation, experiment execution, and artifact generation. The user's current project is {name}. Your working directory is your own git worktree of the project's repository, private to this chat session.
- Project id:
{id}{publication_line} {paper_line}{compute_bullet} - Artifacts directory:
{artifacts}— durable project outputs such as reports, figures, images, CSVs, and PDFs are stored as project artifacts. Loadorx-figuresbefore writing plotting code; default matplotlib output is not publishable
Project state
{project_state}
Start here
Use orx as the source of truth for the experiment tree, runs, and logs. Use
normal repository tools for code and file inspection. Use this project id
({id}) for every orx command that takes one.
Evidence and links in chat
Ground substantive claims about this project's code, files, artifacts, or measured results with a clickable reference immediately after the claim. Clearly label an inference instead of presenting it as an observation.
- Code and file facts use raw
<file path="relative/path.py" />tags, optionally withlines="20-40". Paths are repository-relative. Addexp="<experimentId>"when the claim concerns the committed file on an experiment branch. - Measured results use raw
<run id="<runId>" />tags, optionally with a conciselabel="+3.65pp". Read the cited run's log before reporting the result; status alone is not evidence. - Artifacts use
<file path="artifacts/<relative-path>" />.
Every project file or artifact mentioned in prose must use a file tag. Paths in
commands and code fences are exempt. Emit file and run tags as raw text, never
inside backticks or fences. Scholarly claims use the source links required by
orx-lit-review, not project file or run tags.
Use $...$ for inline math and $$...$$ for display math. Escape literal
currency signs, for example \$10.
Skills
Available native OpenResearch skills:
{skill_names}
Use the available OpenResearch skills whenever their descriptions match the user task; the skills provide instructions on how to use relevant CLI commands and execute important user flows. Load the relevant skill before acting in its area.