refactor(whiteboard): file-based prompts + geometry conflict detection (#485)

* feat(prompts): declare whiteboard role prompt IDs + whiteboard-reference snippet

* feat(prompts): add whiteboard role template skeletons + reference snippet stub

* refactor(orchestration): move buildWhiteboardGuidelines to role-specific markdown templates

* docs(prompts): whiteboard guidelines now in markdown templates

* feat(prompts): whiteboard-reference canvas + JSON output sections

* feat(prompts): whiteboard-reference action schemas for text/shape/line

* feat(prompts): whiteboard-reference action schemas for latex/chart/table

* feat(prompts): whiteboard-reference action schemas for code/delete/clear/close

* feat(prompts): whiteboard-reference LaTeX JSON escape section

* feat(prompts): whiteboard-reference bounds + overlap section

* feat(prompts): whiteboard-reference font + latex height tables

* feat(prompts): whiteboard-reference pre-output checklist

* feat(prompts): expand teacher whiteboard template

* feat(prompts): expand assistant whiteboard template

* feat(prompts): expand student whiteboard template

* test(prompts): assert whiteboard-reference sections reach every role prompt

* tune(prompts): cap teacher element budget at 4/turn, encourage wb_clear on crowded board

* tune(prompts): simplify teacher template — single conservative rule + state-awareness emphasis

* tune(prompts): add LaTeX-height × text-fontSize pairing table for visual weight consistency

* feat(orchestration): preserve image parts from UIMessage to AI SDK ModelMessage

* feat(eval): attach prior-turn whiteboard screenshot as user message image part

Gated behind EVAL_ATTACH_PRIOR_SCREENSHOT=1 env. When on, captures a
screenshot after every turn (not just checkpoints) and attaches it as a
file part with mediaType:'image/png' on the next user message. Teacher
prompt gets a new 'Prior-state image' section telling the agent to
treat it as direct visual feedback.

* fix(orchestration): strip data: prefix before passing image to AI SDK

Vercel AI SDK's streamText/generateText treats ImagePart.image as a URL
to fetch when it's a string. data: URLs fail the http/https scheme check
and throw AI_DownloadError. Strip the data URL prefix and pass raw base64
with mediaType separately — AI SDK treats base64 strings as data content.

* refactor(whiteboard): add conflict summarizer, drop prior-screenshot feature

Eval across flash (image-on vs off × repeat-3) and pro (same) showed the
prior-screenshot image-input feature gives ~0 net benefit: +0.4 overall on
flash and −0.6 on pro. The image feedback loop's theoretical upside
(agent self-correcting from visual) didn't materialize — weak models
couldn't act on the image and strong models "found things to fix" and
over-corrected. Meanwhile programmatic geometry detection on the raw JSON
cleanly lifted overlap from 6.3 → 8.1 on flash.

So: remove the image-input plumbing, keep a programmatic equivalent.

## Removed (reverts 0e07e9e, b5675bc, image half of 097ec2f)

- lib/orchestration/ai-sdk-adapter.ts: revert multimodal content passthrough
- lib/orchestration/summarizers/message-converter.ts: revert image part
  extraction; user content is back to string-only
- lib/orchestration/summarizers/conversation-summary.ts: revert multimodal
  content flattening
- lib/orchestration/director-graph.ts: revert HumanMessage multimodal cast
- tests/orchestration/message-converter.test.ts: deleted (was purely
  image-conversion tests)
- eval/whiteboard-layout/runner.ts: drop image attach, vision gate,
  per-turn screenshot capture; revert to checkpoint-only capture
- Role templates (wb-teacher / wb-assistant / wb-student): strip
  "Prior-state image" sections

## Added

- lib/orchestration/summarizers/whiteboard-conflicts.ts: pure geometry
  detector. Computes bbox IoU, line segment vs bbox intersection, and
  canvas edge clipping on the current whiteboard state. Renders a text
  block for inclusion in the system prompt.
- lib/orchestration/summarizers/state-context.ts: wires the conflict
  block in after the whiteboard element listing.
- Role templates: replace image-reading guidance with concise
  "Layout conflicts" sections pointing at the computed list, with a
  strong anti-action anchor so agents don't self-trigger wb_clear when
  the list is empty.

## Eval instrumentation (orthogonal)

- EVAL_ENABLE_THINKING=1 opt-in to enable model thinking per-request.
  app/api/chat/route.ts reads body.thinking; runner.ts forwards it.
  Default chat path stays on enabled:false for latency.
- Per-turn wall-clock timing: eval/whiteboard-layout/{runner,reporter,
  types}.ts capture and report mean/p50/p95/total turn latency.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(whiteboard): address code-review feedback

W1: drop stale "prior-state image" reference from conflict-block prose
    (lib/orchestration/summarizers/whiteboard-conflicts.ts) — the image
    feature was removed; the text now says "real visible problem on the
    current board".

W3: type `thinking` on StatelessChatRequest instead of reading it through
    a cast. `app/api/chat/route.ts` now accesses `body.thinking` directly;
    frontend callers can discover the field via TS completion.

W4: unify canvas height constant to 563 (matching the actual rendered
    pixel count from capture.ts). The geometry detector's
    CANVAS_HEIGHT, the snippet's Dimensions / coordinate system /
    layout-guide text, and examples all now agree.

N6: add tests/orchestration/whiteboard-conflicts.test.ts — 17 cases
    covering empty input, bbox overlap threshold (30%, 50%, 100%, plus
    10% sub-threshold), line-crosses-bbox (through, endpoint-inside,
    path-above), canvas clipping on all 4 edges, exact-edge placement
    (not reported), malformed elements (skipped not crashed), and
    output format.

All 43 tests pass; tsc --noEmit clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(eval): finish canvas-height 563 unification

Second-pass code review caught one residual 562.5 in app/eval/whiteboard/page.tsx
that the first cleanup (commit 7a006bb) missed. The 0.5px disagreement was
harmless in practice (Playwright captures rounded to 563), but inconsistent
with the geometry detector and the prompt snippet now agreeing on 563.

All 43 tests pass; tsc clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: prettier format whiteboard-conflicts.test + eval/runner

CI's `pnpm check` (prettier --check) caught two unformatted files:
- eval/whiteboard-layout/runner.ts
- tests/orchestration/whiteboard-conflicts.test.ts

Auto-fixed via `pnpm prettier --write`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
This commit is contained in:
wyuc
2026-04-25 15:18:41 +08:00
committed by GitHub
co-authored by Claude Opus 4.7 杨慎
parent 4754bbae6d
commit fc6b186929
18 changed files with 999 additions and 116 deletions
+24
View File
@@ -75,6 +75,30 @@ export function generateReport(
}
lines.push('');
// Timing summary across all turns in all scenario runs
const allTurnDurations: number[] = [];
for (const scenario of report.scenarios) {
if (scenario.turnDurationsMs) {
for (const ms of scenario.turnDurationsMs) allTurnDurations.push(ms);
}
}
if (allTurnDurations.length > 0) {
const sorted = [...allTurnDurations].sort((a, b) => a - b);
const p50 = sorted[Math.floor(sorted.length * 0.5)];
const p95 = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * 0.95))];
const meanMs = mean(allTurnDurations);
const totalS = allTurnDurations.reduce((a, b) => a + b, 0) / 1000;
lines.push('## Turn latency');
lines.push('| Metric | Value |');
lines.push('|--------|-------|');
lines.push(`| Turns measured | ${allTurnDurations.length} |`);
lines.push(`| Mean | ${(meanMs / 1000).toFixed(2)}s |`);
lines.push(`| p50 | ${(p50 / 1000).toFixed(2)}s |`);
lines.push(`| p95 | ${(p95 / 1000).toFixed(2)}s |`);
lines.push(`| Total across all turns | ${totalS.toFixed(1)}s |`);
lines.push('');
}
lines.push('## Scenarios');
for (const scenario of report.scenarios) {
const lastCp = scenario.checkpoints[scenario.checkpoints.length - 1];
+26 -7
View File
@@ -35,6 +35,8 @@ const { values: args } = parseArgs({
const BASE_URL = args['base-url']!;
const CHAT_MODEL_RAW = process.env.EVAL_CHAT_MODEL || process.env.DEFAULT_MODEL;
const SCORER_MODEL_RAW = process.env.EVAL_SCORER_MODEL;
const ENABLE_THINKING =
process.env.EVAL_ENABLE_THINKING === '1' || process.env.EVAL_ENABLE_THINKING === 'true';
if (!CHAT_MODEL_RAW) {
console.error(
'Error: EVAL_CHAT_MODEL (or DEFAULT_MODEL) must be set. Example: EVAL_CHAT_MODEL=openai:gpt-4.1',
@@ -98,12 +100,15 @@ async function runScenario(
metadata?: unknown;
}> = [];
// Per-turn wall-clock latency around runAgentLoop. Used to compare cost
// when toggling EVAL_ENABLE_THINKING.
const turnDurationsMs: number[] = [];
try {
for (let turnIdx = 0; turnIdx < scenario.turns.length; turnIdx++) {
const turn = scenario.turns[turnIdx];
console.log(` Turn ${turnIdx + 1}: "${turn.userMessage.slice(0, 50)}..."`);
// Add user message
messages.push({
role: 'user',
content: turn.userMessage,
@@ -127,6 +132,7 @@ async function runScenario(
// Use the shared agent loop — same logic as frontend
const controller = new AbortController();
const turnStartMs = Date.now();
await runAgentLoop(
{
config: scenario.config,
@@ -147,10 +153,17 @@ async function runScenario(
iterResult = null;
actionChain = Promise.resolve();
// Inject thinking config when EVAL_ENABLE_THINKING is set.
// The chat route defaults to disabled; this opt-in lets us
// measure latency / quality tradeoff without changing prod.
const bodyWithThinking = ENABLE_THINKING
? { ...body, thinking: { enabled: true } }
: body;
return fetch(`${BASE_URL}/api/chat`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body),
body: JSON.stringify(bodyWithThinking),
signal,
});
},
@@ -240,14 +253,20 @@ async function runScenario(
controller.signal,
MAX_AGENT_TURNS,
);
const turnDurationMs = Date.now() - turnStartMs;
turnDurationsMs.push(turnDurationMs);
console.log(
` [timing] turn ${turnIdx + 1} ran in ${(turnDurationMs / 1000).toFixed(1)}s`,
);
// Checkpoint: capture + score
const isLastTurn = turnIdx === scenario.turns.length - 1;
if (turn.checkpoint || isLastTurn) {
const isCheckpoint = turn.checkpoint || isLastTurn;
if (isCheckpoint) {
const elements = stateManager.getWhiteboardElements();
const screenshotFilename = `run${runIndex}_turn${turnIdx}.png`;
const screenshotPath = await captureWhiteboard(elements, scenarioDir, screenshotFilename);
console.log(` Captured: ${screenshotFilename} (${elements.length} elements)`);
try {
@@ -257,7 +276,6 @@ async function runScenario(
} catch (scoreErr) {
const msg = scoreErr instanceof Error ? scoreErr.message : String(scoreErr);
console.error(` Score error (continuing): ${msg.slice(0, 120)}`);
// Preserve screenshot with null score so the report can still include it
checkpoints.push({ turnIndex: turnIdx, screenshotPath, score: null, elements });
}
}
@@ -265,12 +283,12 @@ async function runScenario(
} catch (error) {
const msg = error instanceof Error ? error.message : String(error);
console.error(` Error: ${msg}`);
return { scenarioId: scenario.id, runIndex, model, checkpoints, error: msg };
return { scenarioId: scenario.id, runIndex, model, checkpoints, turnDurationsMs, error: msg };
} finally {
stateManager.dispose();
}
return { scenarioId: scenario.id, runIndex, model, checkpoints };
return { scenarioId: scenario.id, runIndex, model, checkpoints, turnDurationsMs };
}
// ==================== Rescore Mode ====================
@@ -331,6 +349,7 @@ async function main() {
console.log('=== Whiteboard Layout Eval ===');
console.log(`Chat: ${CHAT_MODEL} | Scorer: ${SCORER_MODEL} | Repeats: ${REPEAT}`);
console.log(`Thinking: ${ENABLE_THINKING ? 'ON' : 'OFF'}`);
console.log('');
const scenarios = loadScenarios();
+2
View File
@@ -60,6 +60,8 @@ export interface ScenarioRunResult {
runIndex: number;
model: string;
checkpoints: CheckpointResult[];
/** Per-turn wall-clock latency (ms) from runAgentLoop start to end. */
turnDurationsMs?: number[];
error?: string;
}