feat: whiteboard layout quality eval harness (#425)

* feat(eval): add state manager bridging ActionEngine for eval

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): add shared types for whiteboard layout eval harness

* feat(eval): add SSE chat client for whiteboard eval

* feat(eval): add Playwright capture module for whiteboard screenshots

* feat(eval): add VLM scorer for whiteboard layout evaluation

* feat(eval): add report generator for whiteboard eval results

* feat(eval): add 8 constructed scenarios for whiteboard layout eval

* feat(eval): add minimal whiteboard render page for Playwright screenshots

Creates app/eval/whiteboard/page.tsx — a headless client page that
seeds the stageStore with a synthetic slide scene, exposes
window.__setElements() for Playwright to inject PPTElement[], and
renders them via ScreenElement inside a 1000×562.5px white canvas.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): add main runner for whiteboard layout eval

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): add eval:whiteboard script, install tsx, gitignore results

* fix(eval): fix TS errors, lint, and prettier formatting

* fix(eval): address code review — add cue_user/empty turn guards, validate VLM output, fix empty array crash

* refactor(eval): replace synthetic scenarios with realistic ones

- Replace 8 generic scenarios with 6 that match real usage patterns
- Multi-agent discussion with short user replies (嗯, 明白了, 继续)
- Include real slide scene data as initialStoreState
- Generated agent configs with Chinese names and proper roles
- Cover: physics, math, finance, primary school, economics, medical

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: extract shared agent loop from use-chat-sessions

Extract the core agent loop logic into lib/chat/agent-loop.ts as a pure
async function with callback injection. Both the frontend React hook and
the eval harness now share the same loop — SSE parsing, exit conditions
(END/cue_user/empty turns/max turns), and director state accumulation.

The frontend wires StreamBuffer callbacks for UI pacing; the eval wires
ActionEngine + message accumulation for headless execution. If loop logic
changes in the shared module, both consumers automatically stay in sync.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(eval): use project LLM infrastructure, fix eval page and model config

- Rewrite scorer to use resolveModel() + generateText() from AI SDK
  instead of raw fetch — supports all providers (OpenAI, Google, Anthropic)
- Model config via env vars (EVAL_CHAT_MODEL, EVAL_SCORER_MODEL),
  matching the pattern from outline-language eval
- Fix eval page: bootstrap store before SceneProvider mounts
- Fix __dirname for tsx CJS mode
- Remove --api-key/--scorer-model CLI args (use env vars instead)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): organize results by model/timestamp

* fix: remove double turnCount increment in shared agent loop

The extracted agent-loop.ts had both `turnCount = directorState?.turnCount ?? turnCount + 1`
(line 190) and a redundant `turnCount++` (line 215), causing multi-agent scenarios to hit
maxTurns at half the expected number of iterations.

Also removes unused processSSEStream import from use-chat-sessions.ts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): organize screenshots by scenario subdirectory

Results structure: results/<model>/<timestamp>/<scenario>/run0_turn1.png
Report files stay at the timestamp level.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): revise scorer rubric and add rescore mode

- Replace space_utilization with rendering_correctness and content_completeness
- Rubric now evaluates from a teacher's perspective (empty space is normal)
- readability emphasizes font size consistency
- Add --rescore flag to re-score existing screenshots without re-running chat
- Increase maxOutputTokens to 2000, add JSON parse error recovery
- Score errors no longer abort the entire scenario

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(eval): sharpen scorer rubric with teacher-perspective examples

The rubric now catches specific classroom whiteboard failure modes:
- overlap now explicitly penalizes writing over existing content when empty space is available (spatial planning failure)
- rendering_correctness calls out diagram accuracy (e.g., parabola drawn as V-shape), raw subscripts (G_x), Chinese inside LaTeX math mode
- content_completeness emphasizes canvas edge clipping and bare unlabeled diagrams
- readability penalizes text styled as UI components (gray card backgrounds)
- overall instructed to weight overlap and rendering_correctness more heavily
- explicit note to ignore the "N" page UI element

Also increases maxOutputTokens to 3000 since longer rubric produces longer justifications.
Reporter now guards against null scores (scorer failures no longer crash report generation).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address code review findings

Critical fixes:
- use-chat-sessions: restore agent_end handling, currentMessageId fallback
  for text_delta/action events with missing messageId, and re-throw on
  SSE error events (previously silently pushed to buffer only).
- eval runner: serialize ActionEngine executions via promise chain.
  void-fire-and-forget raced with ensureWhiteboardOpen's 2s delay and
  could insert elements out of order or before the first element was
  committed to the store.

Important fixes:
- CHAT_MODEL default: 'openai/gpt-4o-mini' -> 'openai:gpt-4o-mini'
  (parseModelString splits on ':', not '/').
- CheckpointResult.score is now VlmScore | null; removed the
  'as unknown as' cast that hid the null contract from consumers.
- Delete dead code: eval/whiteboard-layout/chat-client.ts and
  components/chat/process-sse-stream.ts (both unused after the
  shared agent loop refactor).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
This commit is contained in:
wyuc
2026-04-18 15:29:54 +08:00
committed by GitHub
co-authored by Claude Opus 4.6 杨慎
parent 2f542cf33f
commit ea0e8126a5
19 changed files with 2498 additions and 250 deletions
+66
View File
@@ -0,0 +1,66 @@
import { chromium, type Browser, type Page } from '@playwright/test';
import type { PPTElement } from '@/lib/types/slides';
import { mkdirSync } from 'fs';
import { join } from 'path';
const VIEWPORT = { width: 1000, height: 563 };
let browser: Browser | null = null;
let page: Page | null = null;
/**
* Initialize Playwright browser (reused across captures).
*/
export async function initCapture(baseUrl: string): Promise<void> {
browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: VIEWPORT });
page = await context.newPage();
await page.goto(`${baseUrl}/eval/whiteboard`);
// Wait for the page to signal readiness
await page.waitForFunction(
() => (window as unknown as Record<string, unknown>).__evalReady === true,
);
}
/**
* Capture a screenshot of the whiteboard with the given elements.
* Returns the path to the saved screenshot.
*/
export async function captureWhiteboard(
elements: PPTElement[],
outputDir: string,
filename: string,
): Promise<string> {
if (!page) throw new Error('Capture not initialized. Call initCapture() first.');
// Inject elements into the page
await page.evaluate(
(els: unknown[]) => {
const setter = (window as unknown as Record<string, (els: unknown[]) => void>).__setElements;
setter(els);
},
elements as unknown as unknown[],
);
// Wait for rendering to stabilize (fonts, KaTeX, images)
await page.waitForTimeout(1500);
mkdirSync(outputDir, { recursive: true });
const filepath = join(outputDir, filename);
await page.screenshot({ path: filepath, clip: { x: 0, y: 0, width: 1000, height: 563 } });
return filepath;
}
/**
* Close the browser.
*/
export async function closeCapture(): Promise<void> {
if (browser) {
await browser.close();
browser = null;
page = null;
}
}
+103
View File
@@ -0,0 +1,103 @@
import { writeFileSync, mkdirSync } from 'fs';
import { join } from 'path';
import type { EvalReport, VlmScore } from './types';
function mean(nums: number[]): number {
if (nums.length === 0) return 0;
return nums.reduce((a, b) => a + b, 0) / nums.length;
}
function formatNum(n: number): string {
return n.toFixed(1);
}
/**
* Generate JSON + Markdown reports from eval results.
*/
export function generateReport(
report: EvalReport,
outputDir: string,
): { json: string; md: string } {
mkdirSync(outputDir, { recursive: true });
// Collect all scores across all checkpoints
const allScores: VlmScore[] = [];
for (const scenario of report.scenarios) {
for (const cp of scenario.checkpoints) {
if (cp.score) allScores.push(cp.score);
}
}
const dimensions = [
'readability',
'overlap',
'rendering_correctness',
'content_completeness',
'layout_logic',
] as const;
// Build summary stats (guard against empty arrays)
const summary: Record<string, { mean: number; min: number; max: number }> = {};
if (allScores.length > 0) {
for (const dim of dimensions) {
const vals = allScores.map((s) => s[dim]?.score).filter((v): v is number => v != null);
if (vals.length === 0) continue;
summary[dim] = {
mean: mean(vals),
min: Math.min(...vals),
max: Math.max(...vals),
};
}
const overallVals = allScores.map((s) => s.overall);
summary['overall'] = {
mean: mean(overallVals),
min: Math.min(...overallVals),
max: Math.max(...overallVals),
};
}
// Write JSON
const jsonPath = join(outputDir, 'report.json');
writeFileSync(jsonPath, JSON.stringify(report, null, 2));
// Build Markdown
const lines: string[] = [];
lines.push('# Whiteboard Layout Eval Report');
lines.push(
`Run: ${report.timestamp} | Model: ${report.model} | Scenarios: ${report.scenarios.length}`,
);
lines.push('');
lines.push('## Summary');
lines.push('| Metric | Mean | Min | Max |');
lines.push('|--------|------|-----|-----|');
for (const [key, stats] of Object.entries(summary)) {
lines.push(`| ${key} | ${formatNum(stats.mean)} | ${stats.min} | ${stats.max} |`);
}
lines.push('');
lines.push('## Scenarios');
for (const scenario of report.scenarios) {
const lastCp = scenario.checkpoints[scenario.checkpoints.length - 1];
lines.push(`### ${scenario.scenarioId} (run ${scenario.runIndex + 1})`);
if (scenario.error) {
lines.push(`- Error: ${scenario.error}`);
} else if (lastCp) {
if (lastCp.score) {
lines.push(`- Overall: ${lastCp.score.overall}`);
lines.push(`- Overlap: ${lastCp.score.overlap.score} — ${lastCp.score.overlap.reason}`);
if (lastCp.score.issues.length > 0) {
lines.push(`- Issues: ${lastCp.score.issues.join('; ')}`);
}
} else {
lines.push(`- Score: (scoring failed)`);
}
lines.push(`- Screenshot: ${lastCp.screenshotPath}`);
}
lines.push('');
}
const mdPath = join(outputDir, 'report.md');
writeFileSync(mdPath, lines.join('\n'));
return { json: jsonPath, md: mdPath };
}
+366
View File
@@ -0,0 +1,366 @@
import { readFileSync, readdirSync, mkdirSync } from 'fs';
import { join, dirname } from 'path';
import { fileURLToPath } from 'url';
import { parseArgs } from 'util';
import type { EvalScenario, ScenarioRunResult, CheckpointResult, EvalReport } from './types';
import type { Action } from '@/lib/types/action';
import { runAgentLoop, type AgentLoopIterationResult } from '@/lib/chat/agent-loop';
import { EvalStateManager } from './state-manager';
import { initCapture, captureWhiteboard, closeCapture } from './capture';
import { scoreScreenshot } from './scorer';
import { generateReport } from './reporter';
// ==================== CLI Args ====================
//
// Model configuration follows the same pattern as the outline-language eval:
// EVAL_CHAT_MODEL Model for chat generation (default: DEFAULT_MODEL or gpt-4o-mini)
// EVAL_SCORER_MODEL Model for VLM scoring (default: openai:gpt-4o)
//
// Usage:
// EVAL_CHAT_MODEL=google:gemini-3.1-pro-preview \
// EVAL_SCORER_MODEL=google:gemini-2.0-flash \
// pnpm eval:whiteboard --scenario physics-force-decomposition
const { values: args } = parseArgs({
options: {
scenario: { type: 'string' },
repeat: { type: 'string', default: '1' },
'base-url': { type: 'string', default: 'http://localhost:3000' },
'output-dir': { type: 'string', default: 'eval/whiteboard-layout/results' },
rescore: { type: 'string' }, // Path to existing run dir — rescore only, no chat
},
});
const BASE_URL = args['base-url']!;
const CHAT_MODEL = process.env.EVAL_CHAT_MODEL || process.env.DEFAULT_MODEL || 'openai:gpt-4o-mini';
const SCORER_MODEL = process.env.EVAL_SCORER_MODEL || 'openai:gpt-4o';
const REPEAT = parseInt(args.repeat || '1', 10);
const OUTPUT_DIR = args['output-dir']!;
const SCENARIO_FILTER = args.scenario;
const MAX_AGENT_TURNS = 10;
// ==================== Scenario Loading ====================
function loadScenarios(): EvalScenario[] {
const currentDir =
typeof __dirname !== 'undefined' ? __dirname : dirname(fileURLToPath(import.meta.url));
const scenarioDir = join(currentDir, 'scenarios');
const files = readdirSync(scenarioDir).filter((f) => f.endsWith('.json'));
const scenarios: EvalScenario[] = [];
for (const file of files) {
const scenario: EvalScenario = JSON.parse(readFileSync(join(scenarioDir, file), 'utf-8'));
if (SCENARIO_FILTER && scenario.id !== SCENARIO_FILTER && !file.includes(SCENARIO_FILTER)) {
continue;
}
scenarios.push(scenario);
}
return scenarios;
}
// ==================== Single Scenario Run ====================
async function runScenario(
scenario: EvalScenario,
runIndex: number,
runDir: string,
): Promise<ScenarioRunResult> {
const model = scenario.model || CHAT_MODEL;
const checkpoints: CheckpointResult[] = [];
console.log(` [run ${runIndex + 1}] Starting...`);
// Per-scenario sub-directory: runDir/<scenario-id>/
const scenarioDir = join(runDir, scenario.id);
mkdirSync(scenarioDir, { recursive: true });
const stateManager = new EvalStateManager(scenario.initialStoreState);
const messages: Array<{
role: string;
content: string;
parts?: unknown[];
metadata?: unknown;
}> = [];
try {
for (let turnIdx = 0; turnIdx < scenario.turns.length; turnIdx++) {
const turn = scenario.turns[turnIdx];
console.log(` Turn ${turnIdx + 1}: "${turn.userMessage.slice(0, 50)}..."`);
// Add user message
messages.push({
role: 'user',
content: turn.userMessage,
parts: [{ type: 'text', text: turn.userMessage }],
metadata: { createdAt: Date.now() },
});
// Per-iteration state for the eval callbacks
let iterResult: AgentLoopIterationResult | null = null;
let currentAgentId: string | null = null;
let currentMessageId: string | null = null;
const textParts: string[] = [];
const actionParts: Array<{ type: string; actionName: string; params: unknown }> = [];
let cueUserReceived = false;
// Serial action queue: `wb_*` actions must apply in emission order because
// ActionEngine.ensureWhiteboardOpen() awaits an internal delay on first
// call, which would let later actions race ahead and insert elements
// out of order. We chain each execute() onto the previous one and await
// the tail in onIterationEnd before the screenshot.
let actionChain: Promise<void> = Promise.resolve();
// Use the shared agent loop — same logic as frontend
const controller = new AbortController();
await runAgentLoop(
{
config: scenario.config,
apiKey: '', // Server resolves API key from env/YAML
model,
},
{
getStoreState: () => stateManager.getStoreState(),
getMessages: () => messages,
fetchChat: async (body, signal) => {
// Reset per-iteration accumulators
currentAgentId = null;
currentMessageId = null;
textParts.length = 0;
actionParts.length = 0;
cueUserReceived = false;
iterResult = null;
actionChain = Promise.resolve();
return fetch(`${BASE_URL}/api/chat`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify(body),
signal,
});
},
onEvent: (event) => {
switch (event.type) {
case 'agent_start':
currentAgentId = event.data.agentId;
currentMessageId = event.data.messageId;
break;
case 'text_delta':
textParts.push(event.data.content);
break;
case 'action': {
const action: Action = {
id: event.data.actionId,
type: event.data.actionName,
...event.data.params,
} as Action;
// Serialize execution: chain each action onto the previous
// one so they apply in emission order. We await `actionChain`
// in onIterationEnd before screenshotting.
actionChain = actionChain.then(() => stateManager.executeAction(action));
actionParts.push({
type: `action-${event.data.actionName}`,
actionName: event.data.actionName,
params: event.data.params,
});
break;
}
case 'cue_user':
cueUserReceived = true;
break;
case 'done':
iterResult = {
directorState: event.data.directorState,
totalAgents: event.data.totalAgents,
agentHadContent: event.data.agentHadContent ?? true,
cueUserReceived,
};
break;
case 'error':
throw new Error(`API error: ${event.data.message}`);
}
},
onIterationEnd: async () => {
// Wait for all queued actions to apply to the store before we
// use its state (message construction, screenshot capture).
try {
await actionChain;
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
console.error(` Action execution error: ${msg.slice(0, 120)}`);
}
// Build assistant message for conversation history
if (currentMessageId && (textParts.length > 0 || actionParts.length > 0)) {
const parts: unknown[] = [];
if (textParts.length > 0) {
parts.push({ type: 'text', text: textParts.join('') });
}
for (const ap of actionParts) {
parts.push({ ...ap, state: 'result', output: { success: true } });
}
messages.push({
role: 'assistant',
content: textParts.join(''),
parts,
metadata: {
senderName: currentAgentId || 'agent',
originalRole: 'agent',
agentId: currentAgentId,
createdAt: Date.now(),
},
});
}
return iterResult;
},
},
controller.signal,
MAX_AGENT_TURNS,
);
// Checkpoint: capture + score
const isLastTurn = turnIdx === scenario.turns.length - 1;
if (turn.checkpoint || isLastTurn) {
const elements = stateManager.getWhiteboardElements();
const screenshotFilename = `run${runIndex}_turn${turnIdx}.png`;
const screenshotPath = await captureWhiteboard(elements, scenarioDir, screenshotFilename);
console.log(` Captured: ${screenshotFilename} (${elements.length} elements)`);
try {
const score = await scoreScreenshot(screenshotPath, SCORER_MODEL);
console.log(` Score: overall=${score.overall}, overlap=${score.overlap.score}`);
checkpoints.push({ turnIndex: turnIdx, screenshotPath, score, elements });
} catch (scoreErr) {
const msg = scoreErr instanceof Error ? scoreErr.message : String(scoreErr);
console.error(` Score error (continuing): ${msg.slice(0, 120)}`);
// Preserve screenshot with null score so the report can still include it
checkpoints.push({ turnIndex: turnIdx, screenshotPath, score: null, elements });
}
}
}
} catch (error) {
const msg = error instanceof Error ? error.message : String(error);
console.error(` Error: ${msg}`);
return { scenarioId: scenario.id, runIndex, model, checkpoints, error: msg };
} finally {
stateManager.dispose();
}
return { scenarioId: scenario.id, runIndex, model, checkpoints };
}
// ==================== Rescore Mode ====================
async function rescoreRun(runDir: string) {
console.log('=== Rescore Mode ===');
console.log(`Scorer: ${SCORER_MODEL}`);
console.log(`Run dir: ${runDir}`);
// Read the existing report to get scenario metadata
const reportPath = join(runDir, 'report.json');
const oldReport: EvalReport = JSON.parse(readFileSync(reportPath, 'utf-8'));
const allResults: ScenarioRunResult[] = [];
for (const oldResult of oldReport.scenarios) {
console.log(`\nScenario: ${oldResult.scenarioId} (run ${oldResult.runIndex + 1})`);
const checkpoints: CheckpointResult[] = [];
for (const oldCp of oldResult.checkpoints) {
const pngPath = oldCp.screenshotPath;
console.log(` Rescoring: ${pngPath}`);
try {
const score = await scoreScreenshot(pngPath, SCORER_MODEL);
console.log(` Score: overall=${score.overall}, overlap=${score.overlap.score}`);
checkpoints.push({ ...oldCp, score });
} catch (scoreErr) {
const msg = scoreErr instanceof Error ? scoreErr.message : String(scoreErr);
console.error(` Score error: ${msg.slice(0, 120)}`);
checkpoints.push(oldCp); // Keep old score
}
}
allResults.push({ ...oldResult, checkpoints });
}
const report: EvalReport = {
timestamp: new Date().toISOString(),
model: oldReport.model,
scenarios: allResults,
};
const { json, md } = generateReport(report, runDir);
console.log(`\nReport saved:`);
console.log(` JSON: ${json}`);
console.log(` Markdown: ${md}`);
}
// ==================== Main ====================
async function main() {
// Rescore mode: only re-score existing screenshots
if (args.rescore) {
await rescoreRun(args.rescore);
return;
}
console.log('=== Whiteboard Layout Eval ===');
console.log(`Chat: ${CHAT_MODEL} | Scorer: ${SCORER_MODEL} | Repeats: ${REPEAT}`);
console.log('');
const scenarios = loadScenarios();
if (scenarios.length === 0) {
console.error('No scenarios found. Check eval/whiteboard-layout/scenarios/');
process.exit(1);
}
console.log(`Loaded ${scenarios.length} scenario(s)`);
// Create run directory: results/<model>/<timestamp>/
const sanitizedModel = CHAT_MODEL.replace(/[:/]/g, '-');
const timestamp = new Date().toISOString().replace(/[:.]/g, '-').slice(0, 19);
const runDir = join(OUTPUT_DIR, sanitizedModel, timestamp);
mkdirSync(runDir, { recursive: true });
console.log(`Output: ${runDir}`);
await initCapture(BASE_URL);
const allResults: ScenarioRunResult[] = [];
for (const scenario of scenarios) {
console.log(`\nScenario: ${scenario.name} (${scenario.id})`);
const repeats = scenario.repeat ?? REPEAT;
for (let r = 0; r < repeats; r++) {
const result = await runScenario(scenario, r, runDir);
allResults.push(result);
}
}
await closeCapture();
const report: EvalReport = {
timestamp: new Date().toISOString(),
model: CHAT_MODEL,
scenarios: allResults,
};
const { json, md } = generateReport(report, runDir);
console.log(`\nReport saved:`);
console.log(` JSON: ${json}`);
console.log(` Markdown: ${md}`);
}
main().catch((err) => {
console.error('Fatal error:', err);
process.exit(1);
});
@@ -0,0 +1,92 @@
{
"id": "econ-tech-innovation",
"name": "Development Economics — Technology & Innovation",
"description": "qa模式,英文课程,chart+table并排布局测试",
"tags": ["economics", "qa", "single-agent", "en-US", "chart", "table"],
"initialStoreState": {
"stage": {
"id": "eval-econ-innovation",
"name": "Development Economics",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "en-US"
},
"scenes": [
{
"id": "sc-econ-1",
"stageId": "eval-econ-innovation",
"type": "slide",
"title": "Technology and Innovation",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-5",
"content": "<p style=\"font-size: 32px;\">Technology Progress & Innovation</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "sub-5",
"content": "<p style=\"font-size: 18px;\">Schumpeter's Creative Destruction Theory</p>",
"left": 80,
"top": 130,
"width": 500,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "image",
"id": "img-econ",
"src": "https://placehold.co/400x300",
"left": 540,
"top": 120,
"width": 400,
"height": 280,
"rotate": 0,
"fixedRatio": true
}
]
}
}
}
],
"currentSceneId": "sc-econ-1"
},
"config": {
"agentIds": ["default-1"],
"sessionType": "qa"
},
"turns": [
{
"userMessage": "Can you compare R&D intensity vs capital returns on the whiteboard?"
},
{
"userMessage": "Add a table with specific examples",
"checkpoint": true
},
{
"userMessage": "Now show the Silicon Valley innovation formula"
}
]
}
@@ -0,0 +1,197 @@
{
"id": "finance-tax-architecture",
"name": "企业财务 — 三层架构税务筹划",
"description": "qa模式,多agent讨论,表格+公式+形状混合白板",
"tags": ["finance", "qa", "multi-agent", "zh-CN", "table", "latex"],
"initialStoreState": {
"stage": {
"id": "eval-finance-tax",
"name": "企业财务战略",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "zh-CN"
},
"scenes": [
{
"id": "sc-fin-1",
"stageId": "eval-finance-tax",
"type": "slide",
"title": "企业架构与税务优化",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-3",
"content": "<p style=\"font-size: 28px;\">家族公司+持股公司+业务子公司 三层架构</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "shape",
"id": "box-1",
"viewBox": [1000, 1000],
"path": "M 0 0 L 1000 0 L 1000 1000 L 0 1000 Z",
"left": 60,
"top": 130,
"width": 280,
"height": 120,
"rotate": 0,
"fill": "#E3F2FD",
"fixedRatio": false
},
{
"type": "text",
"id": "label-1",
"content": "<p style=\"font-size: 20px;\">家族公司</p>",
"left": 100,
"top": 170,
"width": 200,
"height": 40,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "shape",
"id": "box-2",
"viewBox": [1000, 1000],
"path": "M 0 0 L 1000 0 L 1000 1000 L 0 1000 Z",
"left": 360,
"top": 130,
"width": 280,
"height": 120,
"rotate": 0,
"fill": "#FFF3E0",
"fixedRatio": false
},
{
"type": "text",
"id": "label-2",
"content": "<p style=\"font-size: 20px;\">持股公司</p>",
"left": 400,
"top": 170,
"width": 200,
"height": 40,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "shape",
"id": "box-3",
"viewBox": [1000, 1000],
"path": "M 0 0 L 1000 0 L 1000 1000 L 0 1000 Z",
"left": 660,
"top": 130,
"width": 280,
"height": 120,
"rotate": 0,
"fill": "#E8F5E9",
"fixedRatio": false
},
{
"type": "text",
"id": "label-3",
"content": "<p style=\"font-size: 20px;\">业务子公司</p>",
"left": 700,
"top": 170,
"width": 200,
"height": 40,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
}
]
}
}
}
],
"currentSceneId": "sc-fin-1"
},
"config": {
"agentIds": ["gen-teacher-01", "gen-assistant-01"],
"sessionType": "qa",
"agentConfigs": [
{
"id": "gen-teacher-01",
"name": "林教授",
"role": "teacher",
"persona": "严谨认真的林教授,善于用白板辅助讲解。",
"avatar": "👨‍🏫",
"color": "#4A90D9",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line",
"spotlight",
"laser"
],
"priority": 10
},
{
"id": "gen-assistant-01",
"name": "小雅",
"role": "assistant",
"persona": "热情活泼的小雅,负责补充老师遗漏的要点。",
"avatar": "🧑‍💼",
"color": "#E8913A",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line"
],
"priority": 7
}
]
},
"turns": [
{
"userMessage": "工资和分红在税务上有什么区别?"
},
{
"userMessage": "发奖金也是工资薪金吧,分红是分红",
"checkpoint": true
},
{
"userMessage": "那家族公司到底怎么省税的"
},
{
"userMessage": "确实心疼",
"checkpoint": true
},
{
"userMessage": "搞明白了,那IPO有什么影响"
}
]
}
@@ -0,0 +1,100 @@
{
"id": "math-quadratic-inequality",
"name": "高中数学 — 二次函数与不等式",
"description": "qa模式,单agent,用户追问驱动公式推导和图表绘制",
"tags": ["math", "qa", "single-agent", "zh-CN", "latex"],
"initialStoreState": {
"stage": {
"id": "eval-math-quadratic",
"name": "高中数学函数",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "zh-CN"
},
"scenes": [
{
"id": "sc-math-1",
"stageId": "eval-math-quadratic",
"type": "slide",
"title": "二次函数与一元二次不等式",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-2",
"content": "<p style=\"font-size: 32px;\">二次函数与一元二次不等式</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "def-1",
"content": "<p style=\"font-size: 18px;\">一元二次不等式 ax²+bx+c>0 的解集</p>",
"left": 80,
"top": 140,
"width": 500,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "def-2",
"content": "<p style=\"font-size: 18px;\">与二次函数 y=ax²+bx+c 的图像关系</p>",
"left": 80,
"top": 200,
"width": 500,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
}
]
}
}
}
],
"currentSceneId": "sc-math-1"
},
"config": {
"agentIds": ["default-1"],
"sessionType": "qa"
},
"turns": [
{
"userMessage": "能在白板上推导一下 x²-5x+6>0 怎么解吗"
},
{
"userMessage": "嗯,然后呢",
"checkpoint": true
},
{
"userMessage": "那如果是小于零呢"
},
{
"userMessage": "画个图看看",
"checkpoint": true
},
{
"userMessage": "韦达定理也写一下"
}
]
}
@@ -0,0 +1,150 @@
{
"id": "med-gcp-compliance",
"name": "临床医学 — GCP合规与风险监查",
"description": "discussion模式,紧凑递进式白板布局",
"tags": ["medical", "discussion", "multi-agent", "zh-CN"],
"initialStoreState": {
"stage": {
"id": "eval-med-gcp",
"name": "临床试验GCP",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "zh-CN"
},
"scenes": [
{
"id": "sc-med-1",
"stageId": "eval-med-gcp",
"type": "slide",
"title": "GCP合规要点",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-6",
"content": "<p style=\"font-size: 28px;\">ICH-GCP 药物临床试验质量管理</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "p-1",
"content": "<p style=\"font-size: 18px;\">传统核查 (SDV) vs 基于风险的监查 (RBM)</p>",
"left": 80,
"top": 140,
"width": 600,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "p-2",
"content": "<p style=\"font-size: 18px;\">知情同意的电子化转型</p>",
"left": 80,
"top": 200,
"width": 600,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
}
]
}
}
}
],
"currentSceneId": "sc-med-1"
},
"config": {
"agentIds": ["gen-teacher-01", "gen-assistant-01", "gen-student-张强"],
"sessionType": "discussion",
"triggerAgentId": "gen-student-张强",
"agentConfigs": [
{
"id": "gen-teacher-01",
"name": "林教授",
"role": "teacher",
"persona": "严谨认真的林教授,善于用白板辅助讲解。",
"avatar": "👨‍🏫",
"color": "#4A90D9",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line",
"spotlight",
"laser"
],
"priority": 10
},
{
"id": "gen-assistant-01",
"name": "苏助手",
"role": "assistant",
"persona": "热情活泼的苏助手,负责补充老师遗漏的要点。",
"avatar": "🧑‍💼",
"color": "#E8913A",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line"
],
"priority": 7
},
{
"id": "gen-student-张强",
"name": "张强",
"role": "student",
"persona": "好奇心强的学生张强。临床医学专业",
"avatar": "🧑‍🎓",
"color": "#66BB6A",
"allowedActions": ["wb_open", "wb_draw_text", "wb_draw_latex"],
"priority": 3
}
]
},
"turns": [
{
"userMessage": "SDV和RBM到底有什么区别?"
},
{
"userMessage": "嗯,那博弈点在哪",
"checkpoint": true
},
{
"userMessage": "动态合规怎么理解"
}
]
}
@@ -0,0 +1,191 @@
{
"id": "physics-force-decomposition",
"name": "初中物理 — 力的分解",
"description": "discussion模式,4个agent,用户短回复驱动多轮白板绘制",
"tags": ["physics", "discussion", "multi-agent", "zh-CN"],
"initialStoreState": {
"stage": {
"id": "eval-physics-forces",
"name": "初中物理力学",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "zh-CN"
},
"scenes": [
{
"id": "sc-phys-1",
"stageId": "eval-physics-forces",
"type": "slide",
"title": "力的合成与分解",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-1",
"content": "<p style=\"font-size: 32px;\">力的合成与分解</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "shape",
"id": "bg-1",
"viewBox": [1000, 1000],
"path": "M 0 0 L 1000 0 L 1000 1000 L 0 1000 Z",
"left": 60,
"top": 120,
"width": 880,
"height": 3,
"rotate": 0,
"fill": "#cccccc",
"fixedRatio": false
},
{
"type": "text",
"id": "point-1",
"content": "<p style=\"font-size: 18px;\">合力与分力的关系</p>",
"left": 80,
"top": 150,
"width": 400,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "text",
"id": "point-2",
"content": "<p style=\"font-size: 18px;\">平行四边形定则</p>",
"left": 80,
"top": 210,
"width": 400,
"height": 50,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "image",
"id": "img-1",
"src": "https://placehold.co/400x300",
"left": 540,
"top": 140,
"width": 380,
"height": 280,
"rotate": 0,
"fixedRatio": true
}
]
}
}
}
],
"currentSceneId": "sc-phys-1"
},
"config": {
"agentIds": ["gen-teacher-01", "gen-assistant-01", "gen-student-小明", "gen-student-小红"],
"sessionType": "discussion",
"triggerAgentId": "gen-teacher-01",
"agentConfigs": [
{
"id": "gen-teacher-01",
"name": "张老师",
"role": "teacher",
"persona": "严谨认真的张老师,善于用白板辅助讲解。",
"avatar": "👨‍🏫",
"color": "#4A90D9",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line",
"spotlight",
"laser"
],
"priority": 10
},
{
"id": "gen-assistant-01",
"name": "小助手",
"role": "assistant",
"persona": "热情活泼的小助手,负责补充老师遗漏的要点。",
"avatar": "🧑‍💼",
"color": "#E8913A",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line"
],
"priority": 7
},
{
"id": "gen-student-小明",
"name": "小明",
"role": "student",
"persona": "好奇心强的学生小明。",
"avatar": "🧑‍🎓",
"color": "#66BB6A",
"allowedActions": ["wb_open", "wb_draw_text", "wb_draw_latex"],
"priority": 3
},
{
"id": "gen-student-小红",
"name": "小红",
"role": "student",
"persona": "好奇心强的学生小红。喜欢提问",
"avatar": "🧑‍🎓",
"color": "#66BB6A",
"allowedActions": ["wb_open", "wb_draw_text", "wb_draw_latex"],
"priority": 3
}
]
},
"turns": [
{
"userMessage": "怎么把一个力分成两个力啊?"
},
{
"userMessage": "嗯。",
"checkpoint": true
},
{
"userMessage": "那个平行四边形怎么画?"
},
{
"userMessage": "明白了。",
"checkpoint": true
},
{
"userMessage": "斜面上的物体怎么分解?"
}
]
}
@@ -0,0 +1,144 @@
{
"id": "primary-math-rotation",
"name": "小学数学 — 图形旋转",
"description": "discussion模式,大量shape组合表示复杂图形,多次wb_clear",
"tags": ["math", "discussion", "multi-agent", "zh-CN", "shapes"],
"initialStoreState": {
"stage": {
"id": "eval-math-rotation",
"name": "小学数学图形",
"createdAt": 1700000000,
"updatedAt": 1700000000,
"languageDirective": "zh-CN"
},
"scenes": [
{
"id": "sc-rot-1",
"stageId": "eval-math-rotation",
"type": "slide",
"title": "图形的旋转",
"order": 0,
"content": {
"type": "slide",
"canvas": {
"id": "slide-0",
"viewportSize": 1000,
"viewportRatio": 0.5625,
"theme": {
"backgroundColor": "#ffffff",
"themeColors": ["#5b9bd5", "#ed7d31", "#a5a5a5", "#ffc000", "#4472c4"],
"fontColor": "#333333",
"fontName": "Microsoft YaHei"
},
"elements": [
{
"type": "text",
"id": "title-4",
"content": "<p style=\"font-size: 32px;\">图形的旋转与对称</p>",
"left": 60,
"top": 40,
"width": 880,
"height": 70,
"rotate": 0,
"defaultFontName": "Microsoft YaHei",
"defaultColor": "#333333"
},
{
"type": "image",
"id": "img-rot",
"src": "https://placehold.co/400x300",
"left": 300,
"top": 140,
"width": 400,
"height": 300,
"rotate": 0,
"fixedRatio": true
}
]
}
}
}
],
"currentSceneId": "sc-rot-1"
},
"config": {
"agentIds": ["gen-teacher-01", "gen-assistant-01", "gen-student-乐乐"],
"sessionType": "discussion",
"triggerAgentId": "gen-teacher-01",
"agentConfigs": [
{
"id": "gen-teacher-01",
"name": "高老师",
"role": "teacher",
"persona": "严谨认真的高老师,善于用白板辅助讲解。",
"avatar": "👨‍🏫",
"color": "#4A90D9",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line",
"spotlight",
"laser"
],
"priority": 10
},
{
"id": "gen-assistant-01",
"name": "方块姐姐",
"role": "assistant",
"persona": "热情活泼的方块姐姐,负责补充老师遗漏的要点。",
"avatar": "🧑‍💼",
"color": "#E8913A",
"allowedActions": [
"wb_open",
"wb_close",
"wb_clear",
"wb_delete",
"wb_draw_text",
"wb_draw_shape",
"wb_draw_chart",
"wb_draw_latex",
"wb_draw_table",
"wb_draw_line"
],
"priority": 7
},
{
"id": "gen-student-乐乐",
"name": "乐乐",
"role": "student",
"persona": "好奇心强的学生乐乐。活泼好动",
"avatar": "🧑‍🎓",
"color": "#66BB6A",
"allowedActions": ["wb_open", "wb_draw_text", "wb_draw_latex"],
"priority": 3
}
]
},
"turns": [
{
"userMessage": "门的旋转中心在哪里?"
},
{
"userMessage": "嗯",
"checkpoint": true
},
{
"userMessage": "360度"
},
{
"userMessage": "嗯嗯,对",
"checkpoint": true
},
{
"userMessage": "左转两次等于右转两次吗"
}
]
}
+145
View File
@@ -0,0 +1,145 @@
/**
* VLM Scorer for whiteboard layout quality.
*
* Uses the project's LLM infrastructure (resolveModel + generateText from AI SDK)
* so model configuration follows the same `provider:model` convention as the rest
* of the codebase. Supports all providers (OpenAI, Google, Anthropic, etc.).
*
* Environment variable: EVAL_SCORER_MODEL (default: openai:gpt-4o)
*/
import { readFileSync } from 'fs';
import { generateText } from 'ai';
import { resolveModel } from '@/lib/server/resolve-model';
import type { VlmScore } from './types';
const SCORER_MODEL_DEFAULT = 'openai:gpt-4o';
const RUBRIC_PROMPT = `You are evaluating a classroom whiteboard screenshot from an AI teaching assistant. Score like a teacher reviewing their own board work for a student's benefit.
Context: This is a real-time teaching whiteboard, NOT a poster or infographic.
- Empty space is NORMAL and NOT a problem — teachers write in one area at a time.
- What matters: would a student be confused, misled, or unable to read the content?
- Ignore the small dark circle "N" in the corner — it is a page UI element, not whiteboard content.
Score each dimension from 1 to 10 (10 = perfect, 1 = broken):
1. readability — Can a student read every element easily?
- Font size CONSISTENCY is critical: penalize heavily if some text is 2x+ larger than other text on the same board (e.g., one giant title + tiny formulas).
- Are characters crisp? Any Chinese rendered as boxes or missing glyphs?
- Penalize text styled like UI components (gray boxes, card backgrounds) that don't match handwritten whiteboard feel.
2. overlap — Are elements clear of each other, AND does new content respect existing content?
- Penalize any occlusion (shapes over text, text stacked on text, arrows piercing labels).
- CRITICAL: penalize "writing over existing content" — if a new formula is placed directly on top of an existing table row when empty space was available nearby, that is a layout failure, not just overlap.
- 10 = everything distinct; 1 = multiple elements unreadable due to occlusion.
3. rendering_correctness — Are formulas, shapes, and symbols drawn correctly?
- LaTeX must render: raw source like "\\\\frac", "\\\\theta", or garbled chunks like "0ext", "Gsinheta", "heta" = major penalty.
- Subscripts/superscripts must render: "G_x" shown as raw underscore (not Gₓ) = penalty.
- Chinese inside LaTeX math mode (e.g., "口诀(当 a > 0 ext 时)") = penalty.
- Diagram ACCURACY matters: a parabola drawn as V-shape straight lines, a circle drawn as ellipse-when-should-be-circle, an angle labeled wrong = penalty.
- 10 = all math/shapes render correctly and match the concept; 1 = multiple broken renders OR fundamentally wrong diagrams.
4. content_completeness — Is the content whole, bounded, and annotated?
- Edge clipping: any element cut off at canvas edge (formula missing its left character, table column cut, arrow head beyond edge) = major penalty.
- Unexpected clearing: if previous turns' content has vanished in a later turn with no reason, penalize.
- Bare diagrams with no labels (a circle with no annotation of what it represents) = penalty.
- 10 = all content fully visible and annotated; 1 = significant content lost, truncated, or unlabeled.
5. layout_logic — Does the arrangement support teaching flow?
- Related elements grouped (a diagram with its labels/formulas together)?
- Natural reading order for the concept (cause → effect, equation → graph → solution)?
- Spatial planning: does new content go to sensibly-chosen empty areas rather than crammed near or over existing elements?
overall: 1–10 holistic teaching-quality score. Weight overlap and rendering_correctness more heavily since they directly block comprehension.
issues: 1-5 short concrete problem descriptions a teacher would call out.
Output ONLY a JSON object with this exact structure (no markdown, no code fences):
{"readability":{"score":N,"reason":"..."},"overlap":{"score":N,"reason":"..."},"rendering_correctness":{"score":N,"reason":"..."},"content_completeness":{"score":N,"reason":"..."},"layout_logic":{"score":N,"reason":"..."},"overall":N,"issues":["..."]}`;
/**
* Score a whiteboard screenshot using a VLM.
*
* Model is resolved via EVAL_SCORER_MODEL env var or the provided modelString,
* using the same resolveModel() infrastructure as the rest of the project.
*/
export async function scoreScreenshot(
screenshotPath: string,
modelString?: string,
): Promise<VlmScore> {
const imageBuffer = readFileSync(screenshotPath);
const { model } = await resolveModel({
modelString: modelString || process.env.EVAL_SCORER_MODEL || SCORER_MODEL_DEFAULT,
});
const result = await generateText({
model,
messages: [
{
role: 'user',
content: [
{ type: 'text', text: RUBRIC_PROMPT },
{ type: 'image', image: imageBuffer },
],
},
],
temperature: 0,
maxOutputTokens: 3000,
});
const content = result.text;
// Extract JSON from response (may be wrapped in markdown code fences)
const jsonMatch = content.match(/\{[\s\S]*\}/);
if (!jsonMatch) {
throw new Error(`VLM returned non-JSON response: ${content.slice(0, 200)}`);
}
// eslint-disable-next-line @typescript-eslint/no-explicit-any
let raw: any;
try {
raw = JSON.parse(jsonMatch[0]);
} catch {
// VLM sometimes produces unescaped quotes or trailing content — attempt cleanup
const cleaned = jsonMatch[0]
.replace(/,\s*}/g, '}') // trailing commas
.replace(/,\s*]/g, ']');
try {
raw = JSON.parse(cleaned);
} catch (e2) {
throw new Error(
`VLM returned invalid JSON: ${(e2 as Error).message}\n${jsonMatch[0].slice(0, 300)}`,
);
}
}
const dimensions = [
'readability',
'overlap',
'rendering_correctness',
'content_completeness',
'layout_logic',
] as const;
for (const dim of dimensions) {
if (!raw[dim] || typeof raw[dim].score !== 'number') {
throw new Error(`VLM response missing or invalid dimension: ${dim}`);
}
}
if (typeof raw.overall !== 'number') {
throw new Error('VLM response missing overall score');
}
const score: VlmScore = {
readability: raw.readability,
overlap: raw.overlap,
rendering_correctness: raw.rendering_correctness,
content_completeness: raw.content_completeness,
layout_logic: raw.layout_logic,
overall: raw.overall,
issues: Array.isArray(raw.issues) ? raw.issues : [],
};
return score;
}
+100
View File
@@ -0,0 +1,100 @@
import { useStageStore } from '@/lib/store/stage';
import { useCanvasStore } from '@/lib/store/canvas';
import { useWhiteboardHistoryStore } from '@/lib/store/whiteboard-history';
import { ActionEngine } from '@/lib/action/engine';
import type { Action } from '@/lib/types/action';
import type { PPTElement } from '@/lib/types/slides';
import type { Stage, Scene } from '@/lib/types/stage';
interface InitialState {
stage: Stage | null;
scenes: Scene[];
currentSceneId: string | null;
whiteboardElements?: PPTElement[];
}
/**
* Manages headless Zustand stores + ActionEngine for eval.
*
* Zustand stores are singletons (module-level). We reset them
* for each scenario via setState(). ActionEngine reads/writes
* these same stores — no simulation drift.
*/
export class EvalStateManager {
private actionEngine: ActionEngine;
constructor(initial: InitialState) {
// Reset stores to clean state
useCanvasStore.setState({
whiteboardOpen: false,
whiteboardClearing: false,
});
useWhiteboardHistoryStore.setState({ snapshots: [] });
// Build stage with optional pre-existing whiteboard elements
const now = Date.now();
const stage: Stage = initial.stage ?? {
id: 'eval-stage',
name: 'Eval Stage',
languageDirective: 'en-US',
createdAt: now,
updatedAt: now,
};
// If pre-existing whiteboard elements provided, seed the whiteboard
if (initial.whiteboardElements && initial.whiteboardElements.length > 0) {
stage.whiteboard = [
{
id: 'eval-whiteboard',
viewportSize: 1000,
viewportRatio: 16 / 9,
elements: initial.whiteboardElements,
background: { type: 'solid', color: '#ffffff' },
animations: [],
},
];
}
useStageStore.setState({
stage,
scenes: initial.scenes,
currentSceneId: initial.currentSceneId,
mode: 'autonomous',
});
// ActionEngine takes the store module as its StageStore argument
this.actionEngine = new ActionEngine(useStageStore);
}
async executeAction(action: Action): Promise<void> {
await this.actionEngine.execute(action);
}
getStoreState(): {
stage: Stage | null;
scenes: Scene[];
currentSceneId: string | null;
mode: string;
whiteboardOpen: boolean;
} {
const s = useStageStore.getState();
return {
stage: s.stage,
scenes: s.scenes,
currentSceneId: s.currentSceneId,
mode: s.mode,
whiteboardOpen: useCanvasStore.getState().whiteboardOpen,
};
}
getWhiteboardElements(): PPTElement[] {
const stage = useStageStore.getState().stage;
if (!stage?.whiteboard || stage.whiteboard.length === 0) return [];
const lastWb = stage.whiteboard[stage.whiteboard.length - 1];
return lastWb.elements ?? [];
}
dispose(): void {
this.actionEngine.dispose();
}
}
+70
View File
@@ -0,0 +1,70 @@
import type { PPTElement } from '@/lib/types/slides';
import type { Stage, Scene } from '@/lib/types/stage';
// ==================== Scenario ====================
export interface EvalTurn {
userMessage: string;
checkpoint?: boolean;
}
export interface EvalScenario {
id: string;
name: string;
description: string;
tags: string[];
initialStoreState: {
stage: Stage | null;
scenes: Scene[];
currentSceneId: string | null;
whiteboardElements?: PPTElement[];
};
config: {
agentIds: string[];
sessionType: 'qa' | 'discussion';
};
turns: EvalTurn[];
model?: string;
repeat?: number;
}
// ==================== Scoring ====================
export interface DimensionScore {
score: number;
reason: string;
}
export interface VlmScore {
readability: DimensionScore;
overlap: DimensionScore;
rendering_correctness: DimensionScore;
content_completeness: DimensionScore;
layout_logic: DimensionScore;
overall: number;
issues: string[];
}
// ==================== Results ====================
export interface CheckpointResult {
turnIndex: number;
screenshotPath: string;
/** null when VLM scoring failed — screenshot is still preserved. */
score: VlmScore | null;
elements: PPTElement[];
}
export interface ScenarioRunResult {
scenarioId: string;
runIndex: number;
model: string;
checkpoints: CheckpointResult[];
error?: string;
}
export interface EvalReport {
timestamp: string;
model: string;
scenarios: ScenarioRunResult[];
}