Show original paper figures inline and improve literature review (#367)

* Render local images inline in chat and guide literature-review figures

* Open inline figures in the viewer and separate contextual captions

* Polish inline figure viewer and streamline grounded literature review

* Restore literature retrieval procedure and prioritize comparative visual evidence

* Keep skill routing generic and simplify visual guidance

Preserve canonical descriptions in the generic skill list because crowded native catalogs can abbreviate them. This intentionally trades a small prompt increase for complete routing context. Keep the full retrieval procedure and clarify source PDF extraction.
This commit is contained in:
Daniel Kim
2026-09-20 22:40:32 -07:00
committed by GitHub
parent ecc531ef7a
commit 38a880f4be
24 changed files with 707 additions and 300 deletions
+27 -2
View File
@@ -47,6 +47,12 @@ normal repository tools for code and file inspection. Use this project id
- Follow user instructions and established dependency tooling; inspect project
setup first. `pyproject.toml` alone does not imply uv. Otherwise prefer uv
when available on the execution host.
- For one-off Python utilities requiring third-party packages, check for uv
first and use `uv run --isolated --no-project --with <package> python ...`.
For PDF extraction, the package is `pymupdf` and the import is `pymupdf`.
Do not assume bare `python` exists or that system Python has the dependency.
If uv is unavailable, use a verified environment or a venv with the required
package installed. Keep one-off utilities out of project dependency files.
- For uv projects, use `uv run --locked`, preserve configuration, and commit
dependency declarations and locks together. Fix stale locks rather than
bypassing them.
@@ -78,8 +84,27 @@ label an inference instead of presenting it as an observation.
alone is not evidence.
- Artifacts use `<file path="artifacts/<relative-path>" />`.
Every project file or artifact mentioned in prose must use a file tag. Paths in
commands and code fences are exempt. Emit file and run tags as raw text, never
Display images inline with Markdown: `![Description](path/to/figure.png)`.
Use a session-relative path, `artifacts/<relative-path>`, or an absolute local
path, not a `file://` URL (forward slashes on Windows). For any path containing
spaces, use `![Description](<path with spaces/figure.png>)`; percent-encode a
literal `%` as `%25`. Keep the file available for later
reads of the conversation. Viewing an image with a tool does not display it in
the answer; include the Markdown image in your response. Use file tags when
linking a file, not when showing an image.
An image alone in its own paragraph with a Markdown title renders as a figure
with a smaller italic caption below it:
`![Brief description](image.png "Caption text")`
Captions support inline Markdown and links. Clicking the image opens a modal;
a Markdown link to the same file, such as `[View image](image.png)`, opens it
in the right pane. Use these local links for image references instead of file tags.
Show each underlying image file inline only once per conversation; use a local
link for later references. Different crops or edits may be shown separately.
Other project files or artifacts mentioned in prose must use a file tag.
Paths in commands and code fences are exempt. Emit file and run tags as raw text, never
inside backticks or fences. Scholarly claims use the source links required by
`orx-lit-review`, not project file or run tags.
+102 -23
View File
@@ -1,13 +1,33 @@
---
name: orx-lit-review
description: "Search and read research papers. The main agent calls alphaXiv, OpenAlex, and bioRxiv discovery primitives, ranks the combined candidates, and chooses sources for focused follow-ups. Use for literature reviews, related work, prior art, papers, authors, methods, benchmarks, or research claims; never delegate the retrieval loop to a sub-agent."
description: "Explain and compare scientific or technical concepts using original research evidence. Use before answering conceptual or architectural questions, research claims, literature reviews, or related-work requests, even when no paper, citation, or search is requested. Retrieve with relevant alphaXiv, OpenAlex, and bioRxiv connectors; scale retrieval to the question."
---
# Literature retrieval
You are the retrieval ranker. Call the alphaXiv, OpenAlex, and bioRxiv
primitives yourself, inspect the returned candidates, and decide which sources
are useful for each focused follow-up. Never delegate this loop to a sub-agent.
Use this skill for scientific explanations and comparisons, even without a named
paper. Never delegate retrieval to a sub-agent.
Use enabled literature connectors appropriate to the topic. General web search
is a fallback only when relevant connectors provide no useful evidence. Do not
supplement successful retrieval with a web search for a familiar or preferred
paper; use a focused connector query or `orx paper`. Opening a selected original
PDF to extract evidence is source access; opening its abstract first is unnecessary.
## Explain with visual evidence
- Prefer explaining through original figures, tables, and diagrams. For
comparisons, choose visuals that cover the relevant alternatives; let the
reader see the architecture, relationship, or result being explained.
- Answer as a guided reading of those visuals, presenting them early. Keep
prose brief: what to notice, why it matters, and the caveats, grounded in
contextual author quotations.
Visual evidence should replace standalone tutorials, generated comparison
tables, and recitations of data or structure. Choose visuals for explanatory
value, without a fixed image count. Text is the fallback when relevant sources
contain no useful visual or extraction remains blocked; explain that gap.
One unavailable figure is not a reason to omit other accessible visuals.
Each command performs exactly one public endpoint request and emits its
structured JSON result. No login is required:
@@ -21,8 +41,7 @@ orx discover biorxiv "<biology preprint query>"
- `keyword` searches title, abstract, and full text. Results include the match
snippets that explain why each paper was retrieved. Use short exact terms:
method names, acronyms, benchmarks, authors, or title phrases. Use only terms
stated by the user or observed in results; never invent an acronym expansion.
method names, acronyms, benchmarks, authors, or title phrases.
- `embedding` searches titles and abstracts semantically, then reranks by
similarity and the requested priority. Use the user's actual question or a
concise description of a genuinely missing facet.
@@ -67,8 +86,6 @@ orx discover biorxiv "<query>" --limit 20
## Main-agent retrieval loop
You are the low-latency retrieval ranker. Run the loop below yourself.
### Set up the retrieval query
1. If using keyword retrieval, build focused terms using only wording from the
@@ -151,25 +168,15 @@ papers before retrieval has produced its ranked 5–15 candidate set.
Do not compare alphaXiv votes numerically with OpenAlex citations; they measure
different things. Topical fit is the cross-source ranking signal.
In the final answer, link every alphaXiv/arXiv paper title or paper ID to
`https://www.alphaxiv.org/abs/<versionless-paperId>`. Never return an
`arxiv.org` link for those papers. Link a DOI result to `https://doi.org/<doi>`
and a bare OpenAlex `W…` id to `https://openalex.org/<id>`.
For claim-level synthesis, place the supporting source link immediately after
each substantive scholarly claim, and use a paper as claim-level support only
after reading it. A discovery-only result list may link candidate titles, but
must not imply that their methods or findings were verified from snippets alone.
## Reading selected papers
`orx paper` auto-detects an arXiv id/URL, bioRxiv DOI, other DOI, or OpenAlex
`W…` id. For alphaXiv it returns a compact structured report; use `--full` only
when you explicitly need raw text even if a report exists. Without `--full`, a
missing report automatically falls back to extracted full text in the same
`W…` id. For alphaXiv it returns a compact structured report; use `--full` for
exact wording and surrounding context, even when a report exists. Without
`--full`, a missing report automatically falls back to extracted full text in the same
command. `--full` skips the report entirely rather than acting as a superset of
the default. If extracted text is also unavailable, use the alphaXiv paper link
it returns.
the default. If extracted text is also unavailable, locate the original PDF
through the returned paper link.
`orx paper` prints the alphaXiv link before the content. When alphaXiv has an
associated repository, it then prints `GitHub: <url>`. This is the most-starred
@@ -178,3 +185,75 @@ so sanity-check it before treating it as the implementation.
All discovery and paper commands honor the user's disabled literature-source
settings; do not work around an error saying a source is disabled.
Read a paper before using it as claim-level support. Discovery lists may link
candidate titles, but must not imply that methods or findings were verified
from snippets alone.
## Original visuals
Crop directly from the verified original PDF, using alphaXiv's linked PDF for
alphaXiv papers. For a verified arXiv ID, download the source PDF from
`https://arxiv.org/pdf/<id>`. The alphaXiv `/pdf/` URL serves a viewer page,
not the PDF asset; keep user-facing citations pointed at that viewer.
Preserve panel titles, axes, legends, and table headings; exclude the printed
caption and surrounding prose. Render legibly and inspect
the crop. Do not substitute thumbnails, redraw results, or generate lookalikes.
If extraction fails, inspect the error and try available PDF tooling; report
an unresolved obstacle rather than silently omitting the figure.
Save crops durably in the session working tree. Use the figure component with
brief accessible alt text and a contextual caption in the Markdown title:
```markdown
![Architecture overview](paper/figure1.png "Figure 1. What this shows and why it matters. [p. 6](https://www.alphaxiv.org/pdf/PAPER_ID?page=6)")
```
Replace example values with verified paths, IDs, and pages. Base your caption
on the original, tailor it to the question, and preserve important qualifications.
It is your explanation, not an author quote. Do not bake it into the image or
repeat it below the component.
Embed each underlying file once per conversation. Later references use
`[Figure 1](paper/figure1.png)` or `[Table 1](paper/table1.png)` to open the same
local file in the right pane. Different crops/edits may be embedded; renaming an
unchanged image does not make it new. Keep paper-provenance links separate.
## Quotes and citations
Ground claims in original evidence and distinguish your interpretation.
Use direct quotations to ground authors' reasoning, methods, assumptions, and
limitations instead of paraphrasing them all. Choose complete sentences or
self-contained passages. Read surrounding context; isolated
numbers and clipped phrases are not sufficient. Do not quote values already
clear in a displayed visual. Cite immediately after a quote.
Quote only original text you actually read, including full-text snippets with
sufficient context—not generated reports or summaries. Preserve wording and
qualifications; mark omissions and never join separate snippets into a continuous
quote. Respect quotation limits by selecting fewer complete passages, not by
clipping context. If exact evidence is unavailable, say so; never fabricate it
or present a paraphrase as a quotation.
Prefer `https://www.alphaxiv.org/pdf/<paper-id>?page=N` when alphaXiv contains
the cited version and evidence. N is the verified one-based PDF page index;
omit it when unknown. Do not substitute different preprint results for a journal
version. Use another verified paper viewer when alphaXiv lacks that evidence
or is disabled. Raw PDFs are for extraction, not user-facing citations.
Do not cite abstract pages or invent exact-passage highlighting URL parameters.
For discovery results without an alphaXiv representation, link a DOI to
`https://doi.org/<doi>` or a bare OpenAlex `W…` ID to
`https://openalex.org/<id>`. Never substitute an arXiv link for an available
alphaXiv representation.
| Reference | Label | Destination |
| --- | --- | --- |
| One paper, page known | `p. N` | Paper viewer at that PDF page |
| Multiple papers | `Short title, p. N` | Corresponding paper/page |
| Page unknown | `Paper` or consistent short title | Paper viewer without page |
| Embedded visual | `Figure N` / `Table N` | Existing local image file |
Use these labels consistently, not vague labels such as "Source" or descriptions
of the claim. Disambiguate figures by paper when needed. Introductory paper-title links can
retain their titles. Discovery and figure-provenance links provide navigation;
they do not imply every claim in a paper has been verified.
+1 -1
View File
@@ -174,7 +174,7 @@ const S_AGENT_DELEGATION: AgentSkill = AgentSkill {
};
const S_LIT: AgentSkill = AgentSkill {
name: "orx-lit-review",
description: "Search and read research papers. The main agent calls alphaXiv, OpenAlex, and bioRxiv discovery primitives, ranks the combined candidates, and chooses sources for focused follow-ups. Use for literature reviews, related work, prior art, papers, authors, methods, benchmarks, or research claims; never delegate the retrieval loop to a sub-agent.",
description: "Explain and compare scientific or technical concepts using original research evidence. Use before answering conceptual or architectural questions, research claims, literature reviews, or related-work requests, even when no paper, citation, or search is requested. Retrieve with relevant alphaXiv, OpenAlex, and bioRxiv connectors; scale retrieval to the question.",
content: LIT,
resources: &[],
};
+3 -5
View File
@@ -239,7 +239,7 @@ fn playbook_md(project: &LocalProject, state: &ProjectState) -> String {
let project_state = project_state_md(project, state);
let skill_names = super::agent_skills::skills(super::agent_skills::SkillSet::Local)
.iter()
.map(|skill| format!("- `{}`", skill.name))
.map(|skill| format!("- `{}`: {}", skill.name, skill.description))
.collect::<Vec<_>>()
.join("\n");
let template = SYSTEM_PROMPT
@@ -793,14 +793,12 @@ mod tests {
"template comment not stripped"
);
assert!(!md.contains("<!--"), "HTML comment leaked into the prompt");
// Sanity: skill routing names every installed native skill without
// duplicating the descriptions already surfaced by the harness.
assert!(md.contains("Use the available OpenResearch skills"));
assert!(md.contains("execute important user flows"));
assert!(!md.contains("orx skill <name>"));
// Native catalogs may abbreviate descriptions; preserve routing context here.
for skill in agent_skills::skills(SkillSet::Local) {
assert!(md.contains(&format!("- `{}`", skill.name)));
assert!(!md.contains(skill.description));
assert!(md.contains(&format!("- `{}`: {}", skill.name, skill.description)));
}
assert!(md.contains("orx-compute"));
assert!(md.contains("helping the user across the research process"));
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+2 -2
View File
@@ -49,8 +49,8 @@
html { background: #ffffff; }
html[data-theme="dark"] { background: #0e0c0c; }
</style>
<script type="module" crossorigin src="/assets/index-DdaRJyBa.js"></script>
<link rel="stylesheet" crossorigin href="/assets/index-D5arMJwk.css">
<script type="module" crossorigin src="/assets/index-BayGR9Qm.js"></script>
<link rel="stylesheet" crossorigin href="/assets/index-B0BAnUeV.css">
</head>
<body>
<div
+4 -1
View File
@@ -1184,5 +1184,8 @@
"onboarding_loading_account": "جارٍ تحميل تفاصيل الحساب…",
"onboarding_opencode_not_ready": "تعذّر التحقق من جاهزية OpenCode. أعد المحاولة للتحقق من التثبيت والنماذج.",
"settings_free_models_no_sign_in": "نماذج مجانية · لا يلزم تسجيل الدخول",
"onboarding_choose_another_agent": "اختيار وكيل برمجة آخر"
"onboarding_choose_another_agent": "اختيار وكيل برمجة آخر",
"image_zoom_out": "تصغير",
"image_zoom_in": "تكبير",
"image_zoom_reset": "إعادة ضبط التكبير للملاءمة"
}
+4 -1
View File
@@ -1184,5 +1184,8 @@
"harness_setup_install_preview": "Approve the command below to download and install {agent} on your computer.",
"harness_setup_update_preview": "Approve the command below to update {agent} on your computer.",
"harness_setup_approve_run": "Approve and run",
"onboarding_choose_another_agent": "Choose another coding agent"
"onboarding_choose_another_agent": "Choose another coding agent",
"image_zoom_out": "Zoom out",
"image_zoom_in": "Zoom in",
"image_zoom_reset": "Reset zoom to fit"
}
+4 -1
View File
@@ -1184,5 +1184,8 @@
"harness_setup_install_preview": "Aprueba el siguiente comando para descargar e instalar {agent} en tu equipo.",
"harness_setup_update_preview": "Aprueba el siguiente comando para actualizar {agent} en tu equipo.",
"harness_setup_approve_run": "Aprobar y ejecutar",
"onboarding_choose_another_agent": "Elegir otro agente de programación"
"onboarding_choose_another_agent": "Elegir otro agente de programación",
"image_zoom_out": "Alejar",
"image_zoom_in": "Acercar",
"image_zoom_reset": "Restablecer zoom para ajustar"
}
+4 -1
View File
@@ -1184,5 +1184,8 @@
"harness_setup_install_preview": "برای دانلود و نصب {agent} روی رایانه خود، دستور زیر را تأیید کنید.",
"harness_setup_update_preview": "برای به‌روزرسانی {agent} روی رایانه خود، دستور زیر را تأیید کنید.",
"harness_setup_approve_run": "تأیید و اجرا",
"onboarding_choose_another_agent": "انتخاب یک عامل برنامه‌نویسی دیگر"
"onboarding_choose_another_agent": "انتخاب یک عامل برنامه‌نویسی دیگر",
"image_zoom_out": "کوچک‌نمایی",
"image_zoom_in": "بزرگ‌نمایی",
"image_zoom_reset": "بازنشانی بزرگ‌نمایی برای جا شدن"
}
+4 -1
View File
@@ -1184,5 +1184,8 @@
"harness_setup_install_preview": "अपने कंप्यूटर पर {agent} डाउनलोड और इंस्टॉल करने के लिए नीचे दी गई कमांड को स्वीकार करें।",
"harness_setup_update_preview": "अपने कंप्यूटर पर {agent} अपडेट करने के लिए नीचे दी गई कमांड को स्वीकार करें।",
"harness_setup_approve_run": "स्वीकार करें और चलाएं",
"onboarding_choose_another_agent": "दूसरा कोडिंग एजेंट चुनें"
"onboarding_choose_another_agent": "दूसरा कोडिंग एजेंट चुनें",
"image_zoom_out": "ज़ूम आउट",
"image_zoom_in": "ज़ूम इन",
"image_zoom_reset": "फ़िट करने के लिए ज़ूम रीसेट करें"
}
+4 -1
View File
@@ -1184,5 +1184,8 @@
"harness_setup_install_preview": "批准以下命令以在您的电脑上下载并安装 {agent}。",
"harness_setup_update_preview": "批准以下命令以在您的电脑上更新 {agent}。",
"harness_setup_approve_run": "批准并运行",
"onboarding_choose_another_agent": "选择其他编程智能体"
"onboarding_choose_another_agent": "选择其他编程智能体",
"image_zoom_out": "缩小",
"image_zoom_in": "放大",
"image_zoom_reset": "重置为适合窗口"
}
+3 -1
View File
@@ -97,6 +97,7 @@ import {
} from "./api";
import { WorkspaceTools } from "./components/WorkspaceTools";
import { ProjectTerminal } from "./components/ProjectTerminal";
import { isWindowsDrivePath } from "./markdownTarget";
import { ChatPanel, findPartById, spawnRowTitle } from "./components/ChatPanel";
import { usePopover } from "./components/ModelPicker";
import { SubagentTab } from "./components/SubagentTab";
@@ -162,7 +163,7 @@ function parseFilePath(
}
// A home-anchored path (`~` or `~/…`) is disk, never a repo file — the backend
// expands the `~`, so hand it over verbatim.
if (path === "~" || path.startsWith("~/")) return { path, source: "abs" };
if (path === "~" || path.startsWith("~/") || isWindowsDrivePath(path)) return { path, source: "abs" };
// `path` relative to `base` (`""` when equal), else null. macOS symlinks
// `/tmp`→`/private/tmp` and `/var`→`/private/var`, so an agent-inlined path
// and the stored dir can differ only by that prefix — strip it on both sides.
@@ -1987,6 +1988,7 @@ export default function App({ runtime, projectId, pane }: { runtime: RuntimeInfo
</TabBody>
) : subagentTab ? (
<SubagentTab
projectId={projectId}
// Remount per spawn part so the seed + subscription reset cleanly.
key={subagentTab.spawnPartId}
sessionId={subagentTab.sessionId}
+28 -26
View File
@@ -144,7 +144,7 @@ import {
} from "../orxCommand";
import { LitSourceLogo, parseOrxLit, paperUrl } from "./LitSourceLogo";
import { LitSourcesList } from "./LitSourcesPicker";
import { Md } from "./Md";
import { ChatImageScope, Md } from "./Md";
import { PlanStrip } from "./PlanStrip";
import { SETTINGS_NAV, type SettingsTab } from "./SettingsPage";
import { SkillMenu } from "./SkillMenu";
@@ -6072,31 +6072,33 @@ export function ChatPanel({
}}
>
<div className="chat-thread-inner max-w-readable my-0 mx-auto pt-4 px-4 pb-8 flex flex-col gap-4" ref={threadInnerRef}>
<Transcript
key={activeId}
scrollRef={threadRef}
scrollToEndRef={scrollToEndRef}
stickToBottom={stickToBottom}
onPinToBottom={pinTranscriptToBottom}
messages={messages}
allMessages={allMessages}
canFork={canFork}
onFork={forkTurn}
onSelectFork={selectBranch}
busy={busy}
onOpenFile={openFileInSession}
onOpenRun={onOpenRun}
onOpenSpawnedSession={openSpawnedSession}
runExperimentName={runExperimentName}
onOpenExperiment={onOpenExperiment}
experimentName={experimentName}
onRespond={respond}
onOpenPlan={openPlan}
onOpenSubagent={openSubagent}
recoveringTurnId={recoveringTurnId}
onRecover={recoverFailedTurn}
skills={commands}
/>
<ChatImageScope projectId={projectId} sessionId={activeId}>
<Transcript
key={activeId}
scrollRef={threadRef}
scrollToEndRef={scrollToEndRef}
stickToBottom={stickToBottom}
onPinToBottom={pinTranscriptToBottom}
messages={messages}
allMessages={allMessages}
canFork={canFork}
onFork={forkTurn}
onSelectFork={selectBranch}
busy={busy}
onOpenFile={openFileInSession}
onOpenRun={onOpenRun}
onOpenSpawnedSession={openSpawnedSession}
runExperimentName={runExperimentName}
onOpenExperiment={onOpenExperiment}
experimentName={experimentName}
onRespond={respond}
onOpenPlan={openPlan}
onOpenSubagent={openSubagent}
recoveringTurnId={recoveringTurnId}
onRecover={recoverFailedTurn}
skills={commands}
/>
</ChatImageScope>
{busy && awaitingInput && (
<div className="flex items-center gap-2 text-subtext text-sm pt-0.5 px-0 pb-2 italic">{m.chat_panel_waiting_for_your_input()}</div>
)}
File diff suppressed because one or more lines are too long
+14 -9
View File
@@ -7,6 +7,7 @@ import { type ChatPart } from "../api";
import { findPartById, SubagentTranscript } from "./ChatPanel";
import type { TabOpenIntent } from "../tabPreview";
import { ChatImageScope } from "./Md";
import { TabBody } from "./layout/TabBody";
const PANE_CONTENT_CLASS_NAME = [
@@ -22,6 +23,7 @@ const PANE_CONTENT_CLASS_NAME = [
* the same source the inline block renders from, so it stays in sync as the
* sub-agent works. No dedicated fetch endpoint needed. */
export function SubagentTab({
projectId,
sessionId,
spawnPartId,
onOpenFile,
@@ -31,6 +33,7 @@ export function SubagentTab({
experimentName,
onOpenSubagent,
}: {
projectId: string;
sessionId: string;
spawnPartId: string;
onOpenFile?: (
@@ -112,15 +115,17 @@ export function SubagentTab({
>
<div ref={innerRef}>
{spawn ? (
<SubagentTranscript
spawn={spawn}
onOpenFile={onOpenFile}
onOpenRun={onOpenRun}
runExperimentName={runExperimentName}
onOpenExperiment={onOpenExperiment}
experimentName={experimentName}
onOpenSubagent={onOpenSubagent}
/>
<ChatImageScope projectId={projectId} sessionId={sessionId}>
<SubagentTranscript
spawn={spawn}
onOpenFile={onOpenFile}
onOpenRun={onOpenRun}
runExperimentName={runExperimentName}
onOpenExperiment={onOpenExperiment}
experimentName={experimentName}
onOpenSubagent={onOpenSubagent}
/>
</ChatImageScope>
) : (
<div className="subagent-empty py-[3px] px-1 text-sm text-muted">{m.subagent_tab_this_sub_agent_is_no_longer_available()}</div>
)}
+8 -6
View File
@@ -22,29 +22,31 @@ const SIZES: Record<IconButtonSize, string> = {
small: "h-7 w-7 rounded-sm",
};
function classes(variant: IconButtonVariant, size: IconButtonSize, active: boolean, className?: string) {
return cn(BASE, VARIANTS[variant], SIZES[size], active && "active", className);
function classes(variant: IconButtonVariant, size: IconButtonSize, active: boolean, shape: "rounded" | "circle", className?: string) {
return cn(BASE, VARIANTS[variant], SIZES[size], shape === "circle" && "rounded-full", active && "active", className);
}
export type IconButtonProps = ButtonHTMLAttributes<HTMLButtonElement> & {
active?: boolean;
shape?: "rounded" | "circle";
size?: IconButtonSize;
variant?: IconButtonVariant;
};
export const IconButton = forwardRef<HTMLButtonElement, IconButtonProps>(function IconButton(
{ active = false, size = "default", variant = "default", className, ...props },
{ active = false, shape = "rounded", size = "default", variant = "default", className, ...props },
ref,
) {
return <button ref={ref} className={classes(variant, size, active, className)} {...props} />;
return <button ref={ref} className={classes(variant, size, active, shape, className)} {...props} />;
});
export type IconButtonLinkProps = AnchorHTMLAttributes<HTMLAnchorElement> & {
active?: boolean;
shape?: "rounded" | "circle";
size?: IconButtonSize;
variant?: IconButtonVariant;
};
export function IconButtonLink({ active = false, size = "default", variant = "default", className, ...props }: IconButtonLinkProps) {
return <a className={classes(variant, size, active, className)} {...props} />;
export function IconButtonLink({ active = false, shape = "rounded", size = "default", variant = "default", className, ...props }: IconButtonLinkProps) {
return <a className={classes(variant, size, active, shape, className)} {...props} />;
}
+20
View File
@@ -51,3 +51,23 @@ export function resolveMarkdownTarget(
export function markdownTargetUrl(url: string, target: MarkdownTarget): string {
return `${url}${target.query ? `&${target.query}` : ""}${target.hash}`;
}
export function isWindowsDrivePath(src: string): boolean {
return /^[a-z]:[\\/]/i.test(src);
}
export function chatImageTarget(src: string): (Pick<MarkdownTarget, "path" | "hash"> & { source: "absolute" | "artifact" | "checkout" }) | null {
const windows = isWindowsDrivePath(src);
if (!windows && isExternalMarkdownTarget(src)) return null;
const resolved = resolveMarkdownTarget("", windows ? src.replaceAll("\\", "/") : src, true);
if (!resolved) return null;
// Local-image queries must not override the selected file or session.
const target = { path: resolved.path, hash: resolved.hash };
if (target.path.startsWith("/") || target.path.startsWith("~/") || windows) {
return { ...target, source: "absolute" };
}
if (target.path.startsWith("artifacts/")) {
return { ...target, path: target.path.slice("artifacts/".length), source: "artifact" };
}
return { ...target, source: "checkout" };
}
+61
View File
@@ -0,0 +1,61 @@
import { unified } from "unified";
import remarkParse from "remark-parse";
interface FigureNode {
type: string;
title?: string | null;
alt?: string | null;
url?: string;
children?: FigureNode[];
data?: { hName?: string; hProperties?: Record<string, unknown> };
value?: string;
}
const captionParser = unified().use(remarkParse);
function figureKey(text: string): string | undefined {
const match = text.trim().match(/^(?:original\s+)?(figure|fig\.?|table)\s+(\d+(?:\.\d+)*[a-z]?)(?=$|[\s:.,])/i);
return match ? `${match[1].toLowerCase().startsWith("fig") ? "figure" : "table"} ${match[2].toLowerCase()}` : undefined;
}
function nodeText(node: FigureNode): string {
return node.value ?? node.children?.map(nodeText).join("") ?? "";
}
export function remarkFigures() {
return function transform(tree: FigureNode) {
const images = new Map<string, string | null>();
function collect(node: FigureNode) {
if (node.type === "image" && node.url) {
const key = figureKey(node.title ?? "") ?? figureKey(node.alt ?? "");
if (key) images.set(key, images.has(key) && images.get(key) !== node.url ? null : node.url);
}
node.children?.forEach(collect);
}
collect(tree);
function visit(parent: FigureNode) {
for (const node of parent.children ?? []) {
const key = node.type === "link" ? figureKey(nodeText(node)) : undefined;
const src = key ? images.get(key) : undefined;
if (src) node.data = { ...node.data, hProperties: { ...node.data?.hProperties, "data-figure-src": src } };
const image = node.type === "paragraph" && node.children?.length === 1
? node.children[0] : undefined;
if (image?.type === "image" && image.title?.trim()) {
const caption = captionParser.parse(image.title);
const paragraph = caption.children.length === 1 && caption.children[0]?.type === "paragraph"
? caption.children[0] : undefined;
node.type = "paperFigure";
node.data = { hName: "figure" };
node.children = [image, {
type: "figureCaption",
data: { hName: "figcaption" },
children: paragraph?.children ?? [{ type: "text", value: image.title }],
}];
image.title = null;
}
visit(node);
}
}
visit(tree);
};
}
+1
View File
@@ -61,6 +61,7 @@
--color-danger-hover: color-mix(in oklab, var(--accent-red) 8%, transparent);
--color-danger-active: color-mix(in oklab, var(--accent-red) 14%, transparent);
--color-plan-caret: color-mix(in oklab, var(--base) 35%, var(--text));
--color-image-backdrop: rgb(29 27 26 / 85%);
--color-modal-backdrop: rgb(29 27 26 / 42%);
--color-modal-backdrop-light: rgb(29 27 26 / 40%);
--color-diff-selection: color-mix(in oklab, var(--surface) 76%, var(--primary));
+27
View File
@@ -2,6 +2,7 @@ import assert from "node:assert/strict";
import test from "node:test";
import {
chatImageTarget,
isExternalMarkdownTarget,
markdownTargetUrl,
resolveMarkdownTarget,
@@ -53,3 +54,29 @@ test("resolved image URLs preserve query parameters and fragments", () => {
"/api/file/raw?path=docs%2Fimage.png&raw=1#preview",
);
});
test("chat images resolve local paths without crossing the session root", () => {
assert.deepEqual(chatImageTarget("paper/figures/plot%20one.png"), {
path: "paper/figures/plot one.png", hash: "", source: "checkout",
});
assert.equal(chatImageTarget("/tmp/figure.png").source, "absolute");
assert.equal(chatImageTarget("~/figures/figure.png").source, "absolute");
assert.deepEqual(chatImageTarget("C:/papers/figure.png"), {
path: "C:/papers/figure.png", hash: "", source: "absolute",
});
assert.deepEqual(chatImageTarget("artifacts/paper/figure.png"), {
path: "paper/figure.png", hash: "", source: "artifact",
});
for (const src of ["../secret.png", "%E0%A4%A", "javascript:alert(1)", "data:text/html,test", "https://example.com/image.png"]) {
assert.equal(chatImageTarget(src), null, src);
}
});
test("chat image URLs discard injected query parameters and preserve SVG fragments", () => {
assert.deepEqual(chatImageTarget("figure.svg?path=/secret&sessionId=other#diagram"), {
path: "figure.svg", hash: "#diagram", source: "checkout",
});
assert.equal(chatImageTarget("figures/100%25.png").path, "figures/100%.png");
assert.equal(chatImageTarget(String.raw`C:\papers\figure.png`).path, "C:/papers/figure.png");
});
+45
View File
@@ -0,0 +1,45 @@
import assert from "node:assert/strict";
import test from "node:test";
import { unified } from "unified";
import remarkParse from "remark-parse";
import remarkRehype from "remark-rehype";
import { remarkFigures } from "../src/remarkFigures.ts";
const processor = unified().use(remarkParse).use(remarkFigures).use(remarkRehype);
const render = (text) => processor.runSync(processor.parse(text));
test("a standalone captioned image becomes a figure with a separate linked caption", () => {
const tree = render('![Performance plot](<figures/plot one.png> "Figure 1. Fewer samples. [Source, p. 6](https://www.alphaxiv.org/abs/1234)")');
const figure = tree.children[0];
assert.equal(figure.tagName, "figure");
const [image, caption] = figure.children;
assert.equal(image.tagName, "img");
assert.equal(image.properties.src, "figures/plot%20one.png");
assert.equal(image.properties.alt, "Performance plot");
assert.equal(image.properties.title, undefined);
assert.equal(caption.tagName, "figcaption");
assert.equal(caption.children[0].value, "Figure 1. Fewer samples. ");
assert.equal(caption.children[1].tagName, "a");
assert.equal(caption.children[1].properties.href, "https://www.alphaxiv.org/abs/1234");
});
test("ordinary images and unrelated paragraphs remain ordinary Markdown", () => {
for (const text of ['![Plot](plot.png)\n\nFigure 1: body text', 'Text ![Plot](plot.png "tooltip")']) {
const tree = render(text);
assert.equal(tree.children[0].tagName, "p");
assert.equal(tree.children.some((node) => node.tagName === "figure"), false);
}
assert.equal(render('```md\n![Plot](plot.png "Caption")\n```').children[0].tagName, "pre");
});
test("numbered references open the matching inline image while paper citations stay external", () => {
const tree = render('[**Table 1**](https://www.alphaxiv.org/abs/1234) and [paper](https://www.alphaxiv.org/abs/1234)\n\n![Results](table1.png "Table 1. Results. [Source](https://www.alphaxiv.org/abs/1234)")');
const [reference, , paper] = tree.children[0].children;
assert.equal(reference.properties["data-figure-src"], "table1.png");
assert.equal(paper.properties["data-figure-src"], undefined);
const local = render('[Fig. 2](fig2.png)').children[0].children[0];
assert.equal(local.properties["data-figure-src"], undefined);
assert.equal(local.properties.href, "fig2.png");
const ambiguous = render('[Table 1](https://example.com/paper)\n\n![Table 1](a.png)\n\n![Table 1](b.png)');
assert.equal(ambiguous.children[0].children[0].properties["data-figure-src"], undefined);
});