mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 01:15:18 +08:00
refactor(whiteboard): file-based prompts + geometry conflict detection (#485)
* feat(prompts): declare whiteboard role prompt IDs + whiteboard-reference snippet * feat(prompts): add whiteboard role template skeletons + reference snippet stub * refactor(orchestration): move buildWhiteboardGuidelines to role-specific markdown templates * docs(prompts): whiteboard guidelines now in markdown templates * feat(prompts): whiteboard-reference canvas + JSON output sections * feat(prompts): whiteboard-reference action schemas for text/shape/line * feat(prompts): whiteboard-reference action schemas for latex/chart/table * feat(prompts): whiteboard-reference action schemas for code/delete/clear/close * feat(prompts): whiteboard-reference LaTeX JSON escape section * feat(prompts): whiteboard-reference bounds + overlap section * feat(prompts): whiteboard-reference font + latex height tables * feat(prompts): whiteboard-reference pre-output checklist * feat(prompts): expand teacher whiteboard template * feat(prompts): expand assistant whiteboard template * feat(prompts): expand student whiteboard template * test(prompts): assert whiteboard-reference sections reach every role prompt * tune(prompts): cap teacher element budget at 4/turn, encourage wb_clear on crowded board * tune(prompts): simplify teacher template — single conservative rule + state-awareness emphasis * tune(prompts): add LaTeX-height × text-fontSize pairing table for visual weight consistency * feat(orchestration): preserve image parts from UIMessage to AI SDK ModelMessage * feat(eval): attach prior-turn whiteboard screenshot as user message image part Gated behind EVAL_ATTACH_PRIOR_SCREENSHOT=1 env. When on, captures a screenshot after every turn (not just checkpoints) and attaches it as a file part with mediaType:'image/png' on the next user message. Teacher prompt gets a new 'Prior-state image' section telling the agent to treat it as direct visual feedback. * fix(orchestration): strip data: prefix before passing image to AI SDK Vercel AI SDK's streamText/generateText treats ImagePart.image as a URL to fetch when it's a string. data: URLs fail the http/https scheme check and throw AI_DownloadError. Strip the data URL prefix and pass raw base64 with mediaType separately — AI SDK treats base64 strings as data content. * refactor(whiteboard): add conflict summarizer, drop prior-screenshot feature Eval across flash (image-on vs off × repeat-3) and pro (same) showed the prior-screenshot image-input feature gives ~0 net benefit: +0.4 overall on flash and −0.6 on pro. The image feedback loop's theoretical upside (agent self-correcting from visual) didn't materialize — weak models couldn't act on the image and strong models "found things to fix" and over-corrected. Meanwhile programmatic geometry detection on the raw JSON cleanly lifted overlap from 6.3 → 8.1 on flash. So: remove the image-input plumbing, keep a programmatic equivalent. ## Removed (reverts0e07e9e,b5675bc, image half of097ec2f) - lib/orchestration/ai-sdk-adapter.ts: revert multimodal content passthrough - lib/orchestration/summarizers/message-converter.ts: revert image part extraction; user content is back to string-only - lib/orchestration/summarizers/conversation-summary.ts: revert multimodal content flattening - lib/orchestration/director-graph.ts: revert HumanMessage multimodal cast - tests/orchestration/message-converter.test.ts: deleted (was purely image-conversion tests) - eval/whiteboard-layout/runner.ts: drop image attach, vision gate, per-turn screenshot capture; revert to checkpoint-only capture - Role templates (wb-teacher / wb-assistant / wb-student): strip "Prior-state image" sections ## Added - lib/orchestration/summarizers/whiteboard-conflicts.ts: pure geometry detector. Computes bbox IoU, line segment vs bbox intersection, and canvas edge clipping on the current whiteboard state. Renders a text block for inclusion in the system prompt. - lib/orchestration/summarizers/state-context.ts: wires the conflict block in after the whiteboard element listing. - Role templates: replace image-reading guidance with concise "Layout conflicts" sections pointing at the computed list, with a strong anti-action anchor so agents don't self-trigger wb_clear when the list is empty. ## Eval instrumentation (orthogonal) - EVAL_ENABLE_THINKING=1 opt-in to enable model thinking per-request. app/api/chat/route.ts reads body.thinking; runner.ts forwards it. Default chat path stays on enabled:false for latency. - Per-turn wall-clock timing: eval/whiteboard-layout/{runner,reporter, types}.ts capture and report mean/p50/p95/total turn latency. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(whiteboard): address code-review feedback W1: drop stale "prior-state image" reference from conflict-block prose (lib/orchestration/summarizers/whiteboard-conflicts.ts) — the image feature was removed; the text now says "real visible problem on the current board". W3: type `thinking` on StatelessChatRequest instead of reading it through a cast. `app/api/chat/route.ts` now accesses `body.thinking` directly; frontend callers can discover the field via TS completion. W4: unify canvas height constant to 563 (matching the actual rendered pixel count from capture.ts). The geometry detector's CANVAS_HEIGHT, the snippet's Dimensions / coordinate system / layout-guide text, and examples all now agree. N6: add tests/orchestration/whiteboard-conflicts.test.ts — 17 cases covering empty input, bbox overlap threshold (30%, 50%, 100%, plus 10% sub-threshold), line-crosses-bbox (through, endpoint-inside, path-above), canvas clipping on all 4 edges, exact-edge placement (not reported), malformed elements (skipped not crashed), and output format. All 43 tests pass; tsc --noEmit clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(eval): finish canvas-height 563 unification Second-pass code review caught one residual 562.5 in app/eval/whiteboard/page.tsx that the first cleanup (commit7a006bb) missed. The 0.5px disagreement was harmless in practice (Playwright captures rounded to 563), but inconsistent with the geometry detector and the prompt snippet now agreeing on 563. All 43 tests pass; tsc clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore: prettier format whiteboard-conflicts.test + eval/runner CI's `pnpm check` (prettier --check) caught two unformatted files: - eval/whiteboard-layout/runner.ts - tests/orchestration/whiteboard-conflicts.test.ts Auto-fixed via `pnpm prettier --write`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
This commit is contained in:
co-authored by
Claude Opus 4.7
杨慎
parent
4754bbae6d
commit
fc6b186929
@@ -75,6 +75,30 @@ export function generateReport(
|
||||
}
|
||||
lines.push('');
|
||||
|
||||
// Timing summary across all turns in all scenario runs
|
||||
const allTurnDurations: number[] = [];
|
||||
for (const scenario of report.scenarios) {
|
||||
if (scenario.turnDurationsMs) {
|
||||
for (const ms of scenario.turnDurationsMs) allTurnDurations.push(ms);
|
||||
}
|
||||
}
|
||||
if (allTurnDurations.length > 0) {
|
||||
const sorted = [...allTurnDurations].sort((a, b) => a - b);
|
||||
const p50 = sorted[Math.floor(sorted.length * 0.5)];
|
||||
const p95 = sorted[Math.min(sorted.length - 1, Math.floor(sorted.length * 0.95))];
|
||||
const meanMs = mean(allTurnDurations);
|
||||
const totalS = allTurnDurations.reduce((a, b) => a + b, 0) / 1000;
|
||||
lines.push('## Turn latency');
|
||||
lines.push('| Metric | Value |');
|
||||
lines.push('|--------|-------|');
|
||||
lines.push(`| Turns measured | ${allTurnDurations.length} |`);
|
||||
lines.push(`| Mean | ${(meanMs / 1000).toFixed(2)}s |`);
|
||||
lines.push(`| p50 | ${(p50 / 1000).toFixed(2)}s |`);
|
||||
lines.push(`| p95 | ${(p95 / 1000).toFixed(2)}s |`);
|
||||
lines.push(`| Total across all turns | ${totalS.toFixed(1)}s |`);
|
||||
lines.push('');
|
||||
}
|
||||
|
||||
lines.push('## Scenarios');
|
||||
for (const scenario of report.scenarios) {
|
||||
const lastCp = scenario.checkpoints[scenario.checkpoints.length - 1];
|
||||
|
||||
@@ -35,6 +35,8 @@ const { values: args } = parseArgs({
|
||||
const BASE_URL = args['base-url']!;
|
||||
const CHAT_MODEL_RAW = process.env.EVAL_CHAT_MODEL || process.env.DEFAULT_MODEL;
|
||||
const SCORER_MODEL_RAW = process.env.EVAL_SCORER_MODEL;
|
||||
const ENABLE_THINKING =
|
||||
process.env.EVAL_ENABLE_THINKING === '1' || process.env.EVAL_ENABLE_THINKING === 'true';
|
||||
if (!CHAT_MODEL_RAW) {
|
||||
console.error(
|
||||
'Error: EVAL_CHAT_MODEL (or DEFAULT_MODEL) must be set. Example: EVAL_CHAT_MODEL=openai:gpt-4.1',
|
||||
@@ -98,12 +100,15 @@ async function runScenario(
|
||||
metadata?: unknown;
|
||||
}> = [];
|
||||
|
||||
// Per-turn wall-clock latency around runAgentLoop. Used to compare cost
|
||||
// when toggling EVAL_ENABLE_THINKING.
|
||||
const turnDurationsMs: number[] = [];
|
||||
|
||||
try {
|
||||
for (let turnIdx = 0; turnIdx < scenario.turns.length; turnIdx++) {
|
||||
const turn = scenario.turns[turnIdx];
|
||||
console.log(` Turn ${turnIdx + 1}: "${turn.userMessage.slice(0, 50)}..."`);
|
||||
|
||||
// Add user message
|
||||
messages.push({
|
||||
role: 'user',
|
||||
content: turn.userMessage,
|
||||
@@ -127,6 +132,7 @@ async function runScenario(
|
||||
|
||||
// Use the shared agent loop — same logic as frontend
|
||||
const controller = new AbortController();
|
||||
const turnStartMs = Date.now();
|
||||
await runAgentLoop(
|
||||
{
|
||||
config: scenario.config,
|
||||
@@ -147,10 +153,17 @@ async function runScenario(
|
||||
iterResult = null;
|
||||
actionChain = Promise.resolve();
|
||||
|
||||
// Inject thinking config when EVAL_ENABLE_THINKING is set.
|
||||
// The chat route defaults to disabled; this opt-in lets us
|
||||
// measure latency / quality tradeoff without changing prod.
|
||||
const bodyWithThinking = ENABLE_THINKING
|
||||
? { ...body, thinking: { enabled: true } }
|
||||
: body;
|
||||
|
||||
return fetch(`${BASE_URL}/api/chat`, {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify(body),
|
||||
body: JSON.stringify(bodyWithThinking),
|
||||
signal,
|
||||
});
|
||||
},
|
||||
@@ -240,14 +253,20 @@ async function runScenario(
|
||||
controller.signal,
|
||||
MAX_AGENT_TURNS,
|
||||
);
|
||||
const turnDurationMs = Date.now() - turnStartMs;
|
||||
turnDurationsMs.push(turnDurationMs);
|
||||
console.log(
|
||||
` [timing] turn ${turnIdx + 1} ran in ${(turnDurationMs / 1000).toFixed(1)}s`,
|
||||
);
|
||||
|
||||
// Checkpoint: capture + score
|
||||
const isLastTurn = turnIdx === scenario.turns.length - 1;
|
||||
if (turn.checkpoint || isLastTurn) {
|
||||
const isCheckpoint = turn.checkpoint || isLastTurn;
|
||||
|
||||
if (isCheckpoint) {
|
||||
const elements = stateManager.getWhiteboardElements();
|
||||
const screenshotFilename = `run${runIndex}_turn${turnIdx}.png`;
|
||||
const screenshotPath = await captureWhiteboard(elements, scenarioDir, screenshotFilename);
|
||||
|
||||
console.log(` Captured: ${screenshotFilename} (${elements.length} elements)`);
|
||||
|
||||
try {
|
||||
@@ -257,7 +276,6 @@ async function runScenario(
|
||||
} catch (scoreErr) {
|
||||
const msg = scoreErr instanceof Error ? scoreErr.message : String(scoreErr);
|
||||
console.error(` Score error (continuing): ${msg.slice(0, 120)}`);
|
||||
// Preserve screenshot with null score so the report can still include it
|
||||
checkpoints.push({ turnIndex: turnIdx, screenshotPath, score: null, elements });
|
||||
}
|
||||
}
|
||||
@@ -265,12 +283,12 @@ async function runScenario(
|
||||
} catch (error) {
|
||||
const msg = error instanceof Error ? error.message : String(error);
|
||||
console.error(` Error: ${msg}`);
|
||||
return { scenarioId: scenario.id, runIndex, model, checkpoints, error: msg };
|
||||
return { scenarioId: scenario.id, runIndex, model, checkpoints, turnDurationsMs, error: msg };
|
||||
} finally {
|
||||
stateManager.dispose();
|
||||
}
|
||||
|
||||
return { scenarioId: scenario.id, runIndex, model, checkpoints };
|
||||
return { scenarioId: scenario.id, runIndex, model, checkpoints, turnDurationsMs };
|
||||
}
|
||||
|
||||
// ==================== Rescore Mode ====================
|
||||
@@ -331,6 +349,7 @@ async function main() {
|
||||
|
||||
console.log('=== Whiteboard Layout Eval ===');
|
||||
console.log(`Chat: ${CHAT_MODEL} | Scorer: ${SCORER_MODEL} | Repeats: ${REPEAT}`);
|
||||
console.log(`Thinking: ${ENABLE_THINKING ? 'ON' : 'OFF'}`);
|
||||
console.log('');
|
||||
|
||||
const scenarios = loadScenarios();
|
||||
|
||||
@@ -60,6 +60,8 @@ export interface ScenarioRunResult {
|
||||
runIndex: number;
|
||||
model: string;
|
||||
checkpoints: CheckpointResult[];
|
||||
/** Per-turn wall-clock latency (ms) from runAgentLoop start to end. */
|
||||
turnDurationsMs?: number[];
|
||||
error?: string;
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user