Commit Graph
288 Commits
Author SHA1 Message Date
jackefnandYanpeng Wang 3b7a0ca0c0 feat(document): add multi-format course material upload (#741)
* feat(document): add multi-format course material upload

* fix(document): support office uploads with mineru cloud

* fix(document): handle mineru cloud office edge cases

* fix(document): align mineru cloud artifact metadata

---------

Co-authored-by: Yanpeng Wang <yanpg.wang@gmail.com>
2026-07-02 17:58:58 +08:00
cd5f997dd2 docs: document the dev-server OOM workaround for large generations (#808)
Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-02 05:44:20 -04:00
Dustin Persek 9ffdd72f77 fix(quiz): render formulas in quiz text (#833)
* fix(quiz): render formulas in quiz text

* fix(quiz): tighten formula rendering heuristics
2026-07-01 23:20:14 -04:00
9b4746efe9 feat(dsl): activate the migration registry + runner (#787 Part B-2) (#825)
* feat(dsl): activate the migration registry + runner (#787 Part B-2)

`version.ts` was a stub (`DSL_MIGRATIONS = []`, no runner). Fill in the
document-level migration mechanism, keeping @openmaic/dsl zero-runtime-dep:

- `migrate(doc)` — walks a document from its written `dslVersion` (absent =>
  legacy) up the `DSL_MIGRATIONS` ladder to `DSL_VERSION`, stamping the result.
  Idempotent, forward-compatible (never downgrades a newer doc), fail-loud on a
  broken ladder. Mirrors the app's `migrateSlideContent` philosophy.
- `dslVersionOf` / `needsMigration` readers; `DslVersioned` envelope type;
  `UNVERSIONED_DSL_VERSION` / `DSL_VERSION_KEY` constants.
- First ladder entry: a no-op transform stamping legacy docs to the current
  `DSL_VERSION` (0.1.0) — no serialized shape changed in #811/#817, so this
  wires the pipeline end to end without faking a shape bump.
- `DSL_VERSION` stays 0.1.0 (serialized-contract axis); package minor-bumps to
  0.2.0 for the new API surface. README documents the two independent version
  axes (+ the orthogonal app-side `SlideContent.schemaVersion`).
- test/version.test.ts: ladder invariants, stamp/idempotence/purity,
  forward-compat, fail-loud.

Closes the migration-registry acceptance box of #787; the app-side wiring of a
normalized store onto this pipeline remains future work (per the issue's open
question).

* fix(dsl): pin migration endpoints to literals, not the moving DSL_VERSION

Cross-review (codex, P2): the first ladder entry used `to: DSL_VERSION`. Once a
future shape change bumps DSL_VERSION and appends a real step, that entry's `to`
would move too, so legacy docs would be stamped straight to the new version and
skip the appended transform. Migration endpoints must be immutable.

- add `INITIAL_DSL_VERSION` (pinned '0.1.0' literal); first entry targets it
  instead of the moving `DSL_VERSION` constant.
- test asserts `DSL_MIGRATIONS[0].to === INITIAL_DSL_VERSION` to guard the intent.

* fix(dsl): validate version stamps + align needsMigration/migrate on non-objects

Addresses cosarah's review on #825:

- Malformed version stamps no longer silently bypass migration. `dslVersionOf`
  now fails loud on a present-but-malformed stamp (e.g. "1", "0.1",
  "0.1.0-beta"), instead of letting `parseVersion` coerce it into a comparable
  value that reads as >= DSL_VERSION. Absent stamp / non-object still map to the
  unversioned baseline. Adds an `isValidVersion` (`x.y.z`) guard.
- `needsMigration` and `migrate` now agree on non-objects: both treat a
  non-object as "nothing to migrate" (needsMigration -> false, migrate -> input
  unchanged), so `while (needsMigration(x)) x = migrate(x)` can't loop forever.

Tests cover malformed-stamp fail-loud and the non-object agreement invariant.

---------

Co-authored-by: wyuc <zdq1204@gmail.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-07-01 18:53:12 +08:00
3f851bbcce fix(ai): close PROVIDERS/THINKING_CAPABILITIES metadata drift with a guard (#809)
* fix: close PROVIDERS/THINKING_CAPABILITIES metadata drift with a guard

* fix: use toggle-budget thinking capability for SiliconFlow DeepSeek-V3.2

---------

Co-authored-by: Matt Van Horn <455140+mvanhorn@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 05:20:29 -04:00
6d1fd2babb fix(export): compute SVG path bounding box via getBounds() (#656)
getSvgPathRange rolled its own bbox by reading x/y off every command and
defaulting missing coordinates to 0. That:
- injected a spurious (0,0) for Z/close (no x/y), stretching any glyph that
  doesn't touch the origin;
- injected a 0 on the missing axis for H/V commands;
- treated relative-command deltas as absolute coordinates (the parser does
  not normalise to absolute), producing a completely wrong box;
- ignored arc bulge, giving e.g. a semicircle zero extent on its bulge axis.

These bounds feed the viewBox of the LaTeX -> SVG-image fallback in the PPTX
export (use-export-pptx.ts), so affected formulas were shifted, clipped, or
collapsed. Delegate to svg-pathdata's getBounds(), which handles all of the
above, keeping the empty/malformed -> {0,0,0,0} contract.

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 04:51:26 -04:00
6fb87e7d0b fix(tts): respect string context when splitting the Doubao stream (#677)
Doubao streams a run of concatenated JSON objects with no delimiter. The
splitter counted `{`/`}` braces without tracking whether it was inside a
string literal, so a brace in a string value — e.g. an error
`{"message":"bad {input}"}` — mis-aligned the object boundaries, dropping
the chunk (lost audio, or a swallowed error).

Extract a string-aware `splitConcatenatedJsonObjects` helper (mirrors the
scanner in json-repair.ts: tracks inString + escapes) and use it. Behaviour
is otherwise unchanged: parse failures are skipped, code 20000000 ends the
stream, rate-limit/error codes still throw.

Closes #676

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 04:29:38 -04:00
eca6811c62 fix(export): keep sibling attributes when style is empty (#683)
formatAttributes is a reduce; the empty-style branch returned '' instead of
the accumulator, so every attribute emitted before an empty `style=""`
attribute was dropped on stringify (e.g. `<div class="x" style=""> -> <div>`).
Return `attrs` to omit the empty style while keeping the rest.

Closes #682

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 03:38:31 -04:00
74ff4ada20 fix(web-search): match Brave's current result-title markup (#688)
* fix(web-search): match Brave's current result-title markup

Brave moved the web-result title from `<span class="search-snippet-title">`
to `<div class="… search-snippet-title …">`, so parseBraveSearchHtml hit
`if (!title) continue` for every snippet and returned 0 results against the
live page. The existing test stayed green because its fixture still used the
old <span> markup (drifted from reality).

Accept either <span> or <div> for the title, and update the fixtures to the
current markup (keeping one legacy <span> case). Verified end-to-end against a
real search.brave.com scrape: 0 results before, real results after.

Closes #687

* fix(web-search): pair Brave title open/close tags via backreference

Use a backreference (`<(span|div)…></\1>`) so the title's closing tag must
match its opening tag, per review feedback. Prevents a malformed
`<span …>…</div>` from being mis-parsed as a title. Title text moves to
capture group 2; existing matched-tag markup (div and legacy span) is
unaffected. Adds a regression test for the mismatched-tag case.

---------

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 03:25:39 -04:00
Yanpeng Wangandwyuc b6fe2814d2 fix(quiz): stop leaking questions on entry and pass results to chat agent (#823)
* fix(quiz): stop leaking questions on entry and pass results to chat agent

Fixes #822.

Two related defects on quiz scenes:

1. The `quiz-actions` opening monologue could enumerate question stems,
   compare options, and trigger a `discussion` action that walked agents
   through the quiz BEFORE the learner had answered anything.

2. After the learner submitted, the chat/discussion agent never saw
   their answers, correctness, the canonical `analysis`, or the
   `aiComment` from `/api/quiz-grade` — so post-quiz feedback was
   generic and often re-explained the whole quiz from scratch.

Changes:

- `lib/prompts/templates/quiz-actions/{system,user}.md`: collapse the
  generator down to a brief 1-2 segment opening, forbid `discussion`
  actions in scene.actions, and add explicit "no preview, no answer,
  no detailed concept re-teach" safety rules.
- `lib/types/chat.ts` + `lib/chat/agent-loop.ts`: add optional
  `quizResults` to `StatelessChatRequest.storeState` /
  `AgentLoopStoreState` (sceneId + answers + per-question
  status/earned/aiComment).
- `components/chat/use-chat-sessions.ts`: hydrate `quizResults` from
  `readSubmittedState(sceneId)` once per agent-loop iteration (and on
  initial dispatch) so retries are picked up.
- `lib/orchestration/summarizers/state-context.ts`: when results are
  present, surface per-question student answer / correct answer /
  verdict / analysis / aiComment with a "address THIS student's
  mistakes" instruction; when unsubmitted, keep question visibility
  for clarifying questions but pin strict no-leak rules so a learner
  saying "I'm done" cannot bait the agent into reciting the quiz.

Discussion / multi-agent round-robin and the canonical session flow
are untouched.

* style: apply prettier formatting to state-context

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-01 14:50:15 +08:00
wyucandClaude Opus 4.8 6538e718ff feat(editor): in-editor authoring of classroom agents (Stage-level roster) (#816)
* feat(editor): agent roster edit operations

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(editor): dedupe agent.delete lookup + cover history cap/color/teacherCount

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(editor): materialize agent roster from presets

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* refactor(editor): cycle prepended-teacher index + document branch-1 as-is

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(editor): stage store viewMode + agent roster persistence

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(editor): useAgentRoster controller hook

Implements the AgentRosterController hook that bridges the store, history
stack, and agent-ops layer for the agents authoring view.

- Materializes roster from stage on mount via materializeRoster + resolvePreset
- Holds AgentRosterHistory in React state; exposes add/update/remove/reorder
- Guards LAST_TEACHER errors (caught, no-op) in update + remove
- Exposes SurfaceHistory-shaped history object for undo/redo wiring
- Syncs history.present to setStageAgents via useEffect on every change

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(editor): Agents view UI components

Adds AgentsView, AgentRosterGrid, AgentCard, AgentInspector, AvatarPicker.

Layout: full-height panel with a scrollable grid of agent cards on the left
and a fixed inspector panel on the right (opens when a card is selected).

- AgentCard: avatar + name + role badge + persona excerpt; ◀/▶ reorder
  buttons; delete button disabled + titled when canRemove=false (last teacher)
- AgentRosterGrid: responsive grid with trailing [+ 添加角色] add tile
- AgentInspector: name input, role select (teacher/assistant/student),
  AvatarPicker over AGENT_DEFAULT_AVATARS, persona textarea capped 2000 chars
- AvatarPicker: grid of AGENT_DEFAULT_AVATARS with violet selection ring
- AgentsView: composes the hook + grid + inspector; renders its own
  CommandBar (title="Agents", history=hook.history) with leading/trailing slots

Styling follows quiz surface conventions (rounded-2xl cards, zinc palette,
violet accent on selection, same FOCUS ring treatment for inputs).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(editor): [Slides]/[Agents] view toggle wired into Pro-mode chrome

- CommandBar: add optional `leading?: ReactNode` center slot for the toggle
- EditShell/Frame: thread new `commandLeading` prop down to CommandBar
- EditChromeRoot: read viewMode/setViewMode from useStageStore; build a
  ViewModeToggle ([幻灯片] / [角色] segmented control) passed as commandLeading
  to both EditShell (slides mode) and AgentsView (agents mode)
- When viewMode==='agents', render AgentsView instead of EditShell; leftRail
  and bottomRail are naturally absent since AgentsView owns its own layout
- When viewMode==='slides', existing EditShell path is unchanged
- CommandBar history in agents mode binds to useAgentRoster().history;
  title = "Agents"

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(editor): functional state updates in useAgentRoster + last-teacher role feedback

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(editor): break infinite render loop in Agents view persistence effect

Remove `stage` from the useEffect dep array in useAgentRoster — setStageAgents
mutates stage (new object ref), which was triggering re-render → effect → loop
(React error #185). setStageAgents already no-ops when stage is null
(lib/store/stage.ts:287), so the `if (stage)` guard is also removed.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(editor): persist agent roster edits to registry + stage snapshot

- Add generatedAgentConfigs to StageRecord (database.ts) so Dexie can
  round-trip it through db.stages.
- Include generatedAgentConfigs in saveStageData whitelist (stage-storage.ts)
  so editor reloads restore the user's roster edits.
- In saveToStorage (stage.ts), call saveGeneratedAgents(stageId, configs)
  after the stage snapshot write; this is already debounced (500 ms) via
  debouncedSave, so db.generatedAgents + the in-memory registry are synced
  without writing on every keystroke. saveGeneratedAgents clears-then-
  bulk-inserts, so deleted agents are fully removed from the registry.
- Add 3 timer-driven tests that verify saveStageData carries
  generatedAgentConfigs and saveGeneratedAgents is called with the
  correct payload after the debounce fires.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(editor): scope agent-registry sync to real roster edits (cross-review)

- FIX A (P1): remove saveGeneratedAgents from shared saveToStorage; add
  dedicated debouncedSaveAgents called only from setStageAgents — scene
  advances (setCurrentSceneId etc.) no longer churn db.generatedAgents
- FIX B (P2.1): add isDirtyRef guard in useAgentRoster so opening the
  Agents tab without editing does not persist/rewrite the roster; ref is
  set in applyOp, add, undo, redo before any setHistState mutation
- FIX C (P2.2): coalesce consecutive agent.update ops on the same
  agent+fields into one history entry so undo steps over a whole edit
  rather than each keystroke
- Tests: add regression guards asserting saveGeneratedAgents is NOT
  called on setCurrentSceneId or plain saveToStorage, IS called after
  setStageAgents (new dedicated debounce path)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(editor): clone preset agents to fresh ids on materialization (cross-review P1)

* refactor(editor): move agent roster into right-rail tab beside Edit with AI

The [幻灯片][角色] top toggle + full-canvas AgentsView swap is replaced
by a tabbed right rail (RightRailTabs) with "Edit with AI" and "角色"
tabs. The slide canvas stays visible in all modes; agents are edited in
the 角色 tab of the right rail.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(editor): restyle agent panel to Design B accordion (课堂阵容)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(editor): cross-review — persona draft sync, scoped preset re-id, AI-tab gating, dead-code cleanup

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(editor): avatar click expands card, editable-name affordance, add-button copy

* fix(editor): cross-review — drop blur-level undo coalescing, sync playback selection to edited roster

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(editor): prettier format agent panel files (CI fix)

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-01 02:04:28 -04:00
wyucandClaude Opus 4.8 7c19ab8f1e feat(dsl): JSON Schema artifacts + pure validators (#787 Part B-1) (#817)
* feat(dsl): add ACTION_TYPES set + isActionType guard

A frozen set of every valid ActionType plus a pure membership guard,
mirroring the PPT_ELEMENT_TYPES / isPPTElementType idiom in guards.ts.
Used by the new validators; useful standalone for narrowing untrusted
action.type values.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(dsl): build-time JSON Schema codegen for stage/scene/action

TS types stay the single source of truth; ts-json-schema-generator (a
devDependency, build-only) emits dist/schema/{stage,scene,action}.schema.json,
shipped under the new ./schema/* export. Serves non-TS / bring-your-own-validator
consumers without adding a runtime dependency.

JSON Schema is not generic, so schema-roots.ts provides a concrete
Scene<Action, SceneContent> entry point (kept internal, not re-exported).
schema.test.ts compiles the generated schema with ajv (devDep) to prove the
artifact is real and self-contained on a clean checkout.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(dsl): add pure structural validators (stage/scene/action)

validateStage / validateScene / validateAction: hand-written, zero-dependency,
fail-loud structural checks layered on the existing guards. They verify
discriminants, required fields, and known nested discriminants — a gate for
untrusted input (LLM/agent output) and persistence boundaries. Exhaustive
per-field validation is delegated to the shipped JSON Schema. Error-collecting
ValidationResult reports every issue with a JSON-pointer-ish path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(dsl): document the runtime schema + validator layer

Add a Runtime layer section covering the two complementary consumption modes
(schema-as-data via @openmaic/dsl/schema/* and the zero-dep validate* gate),
the validate.ts module row, the build:schema step, and tick the JSON Schema
roadmap box.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(dsl): harden validators per cross-review (exhaustiveness, discriminant agreement)

Addresses cross-review findings on the validator layer:

- Add compile-time exhaustiveness assertions for ACTION_TYPES and SCENE_TYPES.
  `satisfies` only proved each entry is valid, not that the tuple covers the
  whole union — a newly-added Action/Scene variant could silently be rejected
  by the validators. The assertion now fails the build on drift.
- validateScene cross-checks that scene.type agrees with content.type; a
  slide-typed scene carrying quiz content (or vice versa) is now rejected.
- Move SCENE_TYPES + add an isSceneType guard to stage.ts, next to SceneType,
  matching the PPT_ELEMENT_TYPES / isPPTElementType idiom in guards.ts; validate.ts
  reuses them instead of re-encoding the union locally.
- Make the validators' scope honest in the doc comment + README: they check the
  structural envelope, known discriminants, and scene/content agreement; they do
  NOT exhaustively check per-variant fields, and app-side interactive/pbl content
  is validated only at the envelope level (exhaustive checks → the JSON Schema).
- Note the contract-owned content scope in schema-roots.ts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* perf(dsl): reuse one schema generator + harden direct-invocation guard

Addresses cross-review findings on the codegen path:

- gen-schema.mjs built a fresh ts-json-schema-generator (full TS program parse)
  per root type; build one generator over the tsconfig program and reuse it for
  every createSchema call — parses once, identical per-type output.
- Harden the `invoked directly` check to compare realpaths, so a symlinked
  invocation (bin shim / pnpm link) still runs main() instead of silently
  emitting no schema.
- schema.test.ts generates all roots once at module scope, then only compiles
  per case — no repeated heavy codegen.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(dsl): bind scene schema discriminants + pin node-20-compatible generator

Addresses round-2 cross-review (codex) findings:

- scene.schema.json now ties scene `type` to the matching `content.type`.
  SerializedScene became a discriminated union (SlideScene | QuizScene via
  `Omit<Scene<…>, 'type'> & { type }`), so the generator emits a const-narrowed
  discriminant per branch. Previously the schema accepted e.g.
  { type: 'quiz', content: { type: 'slide', … } } even though validateScene
  rejects it — the schema is now consistent with the validator. Locked with a test.
- Pin ts-json-schema-generator to ~2.4.0. ^2.3.0 resolved to 2.9.0, which
  declares `node >=22`, conflicting with the repo's `engines.node >=20.9.0`;
  2.4.x supports node >=18, keeping the build runnable on every supported Node.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(dsl): prettier formatting + drop unused ts-expect-error for green CI

- Apply repo Prettier (`pnpm check`) formatting to the new/edited files.
- Remove the `@ts-expect-error` on the gen-schema.mjs import: the root
  `tsc --noEmit` (which compiles test files; allowJs infers the .mjs exports)
  flagged it as an unused directive (TS2578). Plain comment retained.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(dsl): align validators + schema to the public Scene type (review)

Addresses cosarah's review: the validator, the JSON Schema, and the public TS
type had drifted into three different notions of a valid scene. Realign them so
the TS types stay the single source of truth, and position the two runtime
layers honestly.

- schema-roots.ts: revert SerializedScene to the plain default Scene<Action,
  SceneContent>. The schema no longer encodes a stricter type/content binding
  that the public Scene type doesn't express (the constraint had effectively
  moved into a hidden internal type).
- validate.ts: drop the scene.type<->content.type agreement check (the public
  Scene does not bind them, so neither does the validator); restrict content to
  the contract-owned kinds (slide/quiz), matching SceneContent and the schema,
  so validator and schema no longer disagree on interactive/pbl content.
- Reposition docs (validate.ts header + README): the JSON Schema is the
  authoritative, exhaustive per-field validator (it checks variant fields like
  an action's elementId) for trust boundaries; validate* are a cheap,
  zero-dep structural pre-check and a strict subset of the schema — not a
  separate or stricter gate.
- Tests updated to lock the realigned behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(dsl): bind scene type<->content in the public Scene contract (review)

cosarah's review showed the scene type<->content agreement is a real, load-bearing
invariant: consumers branch on `scene.type` and then read `scene.content` as the
matching shape (e.g. `stage-api-canvas` casts `content as SlideContent` after a
`scene.type === 'slide'` check; `complete-summary` counts by `scene.type` and reads
`content as QuizContent`). A validator/schema that blesses mismatched pairs is unsafe.
So tighten the *exported* contract rather than loosening the runtime checks to match
a hole — and make the pure validators the authoritative, relied-upon boundary.

DSL contract:
- `Scene` becomes a distributive discriminated type binding `type` to `content['type']`
  (`SceneCore<TAction> & { type: TContent['type']; content: TContent }` distributed over
  TContent). TS itself now rejects mismatched scenes; generic widening still works.
  Extract `SceneCore` for the kind-independent fields.
- `validateScene` re-establishes the type<->content agreement check; scene `type` is
  the contract's slide/quiz; `scene.schema.json` (from the spelled-out discriminated
  `SerializedScene`) binds `type` to content per branch. Type, validator, and schema
  now describe the same thing.
- `validateAction` checks each variant's required fields (e.g. spotlight.elementId,
  discussion.topic) — closes the false-positive gap. A test pins the hand map to the
  generated schema (the TS-derived source of truth) so it cannot drift.
- Reposition docs: `validate*` is the zero-dep in-process boundary producers rely on;
  the JSON Schema is the cross-language mirror + exhaustive value checker.

App consumers (the sites the discriminated Scene exposes):
- Add `ScenePatch` (non-distributive partial) + `makeScene(core, content)` helper to
  lib/types/stage.ts — one isolated cast, type derived from content. Patch-typed
  signatures (`updateScene`, `applyScenePatchInSync`, regenerate plan) and scene
  rebuilds (`stage-api-scene`, store `updateScene`, `migrateScene`, stage-storage
  deserialize) route through these. Closes two latent desync foot-guns.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(dsl): validate action field types + migrate scene.update; bidirectional drift-guard (review)

Round-2 review of the contract-tightening change (codex + Claude):

- validateAction now checks each variant-required field's TYPE, not just presence
  (`{ type:'spotlight', elementId: 123 }` is rejected — elementId must be a string).
  ACTION_REQUIRED_FIELDS carries a runtime kind per field.
- The schema lockstep test is now bidirectional and type-aware: it asserts a
  fully-typed action is accepted (catches the map OVER-listing an optional field),
  every required field flags when missing (under-listing), and flags when present
  but mis-typed — all keyed off the generated schema's field names + types.
- Migrate the public `stageApi.scene.update()` to ScenePatch + makeScene (the
  create() path was done last commit; update() still did a raw spread that could
  desync type<->content — masked from tsc only because StageStore.setState is typed
  `any`). Now both public scene-construction paths rebind type to content.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(api): keep scene.create's params.type authoritative over content.type (review)

cosarah's review: `CreateSceneParams` exposes an independent `type` plus
`content?: Partial<SceneContent>`, and after switching create() to makeScene()
(which derives the scene kind from content.type), a call like
`create({ type: 'slide', content: { type: 'interactive', ... } })` would silently
become an interactive scene — params.type ignored — and mix slide defaults into it.

Reject a `content.type` that disagrees with `params.type`, and pin the merged
content's `type` to `params.type` so a partial content override can't flip the
scene's discriminant. params.type stays authoritative. Adds a test both ways.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 01:04:34 +08:00
wyuc 5daf3e5a6e fix(editor): keep emptied/zero-action scenes playable, bind outline by stable id, surface incomplete content (#814)
Editor empty-state robustness: dwell on zero-action scenes (engine + PlaybackChromeRoot) so emptied/blank slides stay playable; bind a scene's outline by stable outlineId (reorder/insert/duplicate-safe, allOutlines in scene order); clear the cue preview on glyph unmount; surface incomplete content calmly (amber dashed frame on clips, amber dot on blank outline titles + rail, neutral N/M-ready generate gate). Cross-reviewed (Codex + Claude), 5 rounds to a clean pass.
2026-06-30 05:57:26 -04:00
b516427d27 perf(generation): index assigned images by id in fixElementDefaults (#701)
fixElementDefaults looked up each image element's metadata with
assignedImages.find inside the per-element map, making the pass
O(elements × images). Build a Map<id, image> once and use an O(1)
Map.get lookup instead. Behavior is identical (Map.get returns the same
entry as the find), so it's purely a lookup-mechanism swap.

Closes #700

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-30 05:15:01 -04:00
19ff0cef42 fix(export): convert PPTX shadow offset from px to pt (#679)
getShadowOption converted the shadow blur from px to pt but left the offset
in px. pptxgenjs expects both in points (ratioPx2Pt = 96/72 × viewportSize/960),
so shadows exported ~33% too far from their element at the default viewport
while the blur radius was correct. Divide the offset by ratioPx2Pt like blur.

Closes #678

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-30 04:59:00 -04:00
a3f88d53e5 fix(mathml2omml): call includes() instead of indexing it (#681)
parse.js used `textContainerNames.includes[arr[level].name]`, which reads a
property of the includes function (always undefined) instead of calling it,
so the "trailing text node" branch never ran and trailing text inside a
MathML text container (mtext/mi/mn/mo/ms) was dropped from the OMML. Line 48
of the same file already uses the correct `includes(...)` form.

dist/ is gitignored and rebuilt by postinstall, so the runtime bundle picks
up the fix on install.

Closes #680

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-30 04:50:07 -04:00
wyuc b122ca30d3 feat(dsl): bring the Action playback verbs into @openmaic/dsl (#787)
Promote the Action contract (Action union + variants, ActionType, the action-category lists, PercentageGeometry) into @openmaic/dsl/src/action.ts, zero-runtime-dep. Widget interaction actions graduate into the contract; Scene<TAction> defaults to the standard Action union; lib/types/action.ts becomes a re-export shim. Part A of #787 (Phase 2 under #720).
2026-06-29 23:38:28 -04:00
xuyuanwei678 261fddb062 fix(packages): add openmaic repository metadata (#813)
* fix(packages): add openmaic repository metadata

* ci(packages): publish on openmaic manifest changes
2026-06-29 16:52:03 +08:00
wyuc 0bcab10b83 fix(lecture-notes): render interactive-webpage widget actions in notes area (#810)
Interactive-webpage scenes (SceneType 'interactive') emit four widget_*
action types — widget_highlight, widget_setState, widget_annotation,
widget_reveal — each typically followed by a speech action. These execute
correctly on stage during playback but were dropped from the Lecture Notes
timeline due to two compounding gaps:

1. chat-area.tsx built lectureNotes with an allowlist that only admitted
   speech/spotlight/laser/play_video/discussion, filtering out all widget_*
   types so they never entered the notes data.
2. lecture-notes-view.tsx ACTION_ICON_ONLY had no entries for widget_*, and
   unknown types render as null, so even with data they would be invisible.

Add the four widget_* types to the allowlist and extend ACTION_ICON_ONLY with
icon/style entries (Highlighter/SlidersHorizontal/StickyNote/Eye), reusing the
existing inline icon-only badge pattern already used by spotlight/laser.
2026-06-29 04:33:22 -04:00
xuyuanwei678andwyuc 72e3ef04d7 chore(packages): set openmaic package versions to 0.0.2 (#812)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-29 16:22:13 +08:00
wyucandClaude Opus 4.8 b93434ae2c feat(agent-edit): multi-session conversation history for the AI editor (#801)
* feat(agent-edit): multi-session conversation history for the AI editor

The "Edit with AI" panel kept only one thread per stage in localStorage, and
"New conversation" deleted it, so previous conversations were unrecoverable.

Add a per-stage session list backed by IndexedDB (Dexie v12 `agentEditSessions`):

- "New conversation" now archives the current session and starts a fresh one
  instead of deleting it; a history popover in the panel header lists past
  sessions (auto-titled from the first user message) to switch back to or delete.
- One-time migration of the existing single-thread localStorage entry.
- The active session id is remembered per stage so a refresh restores it,
  including a just-cleared empty session (clean slate survives reload).
- saveSession preserves the original createdAt and tombstones deleted ids so an
  in-flight save can't resurrect a deleted session; switching stages drops the
  previous thread synchronously so it can't leak into the new stage.
- Sessions are pruned to a soft cap per stage and removed when a course is
  deleted.
- i18n for all 8 locales; unit tests for title derivation and the store.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(types): drop redundant Window.SpeechRecognition augmentation conflicting with @assistant-ui/core

@assistant-ui/core's speech adapter declares Window.SpeechRecognition and
webkitSpeechRecognition globally as SpeechRecognitionConstructor. Three files
(prompt-input, use-audio-recorder, use-browser-asr) re-declared them as `any`,
which conflicts (TS2717) once that adapter's types are pulled into the program,
and the constructed instance then lacks the event handlers the code sets
(TS2339/2551). Rely on the global declaration and cast the instance to the local
rich shape. Takes `tsc --noEmit` back to zero.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 23:08:23 -04:00
wyuc da0b394b81 release: v0.3.0 v0.3.0 2026-06-28 22:34:18 -04:00
杨慎 a93f875044 fix(agent): keep edit runtime alive across read-only scenes 2026-06-28 11:23:18 -04:00
cec3c6aaea Refactor widget actions into the scene action pipeline (#796)
* refactor: align widget actions with scene pipeline

* feat(edit): migrate legacy teacherActions off interactive content on load

The widget-actions pipeline no longer persists a teacherActions authoring
layer; playback reads only scene.actions. Add migrateInteractiveContent to
strip the dead field from existing documents as they load through migrateScene
(wired into all store/classroom load paths). Satisfies the #794 forward-migration
acceptance item.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(generation): default widget_setState.state to {} when the LLM omits it

The removed TeacherAction->Action converter always defaulted state to {}.
Restore that guard in the structured-output parser so an action missing
state can't forward state: undefined into the widget iframe, where
SET_WIDGET_STATE handlers dereference it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(edit): also strip nested widgetConfig.teacherActions in migration

Legacy WidgetConfig variants each declared their own teacherActions, so existing
documents can carry the dead field nested inside content.widgetConfig as well as
at the top level. Extend migrateInteractiveContent to drop both, preserving all
other widgetConfig fields.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(prompts): document stable widget-action selectors for diagram/game/viz3d

The interactive-actions prompt documented stable target selectors only for
procedural-skill and simulation widgets. Add the canonical selectors the
diagram / game / visualization3d content prompts actually emit (node IDs via
revealOrder; #zoom-in-btn / #canvas-container / #speed-slider / ...; #start-btn),
plus a rule to use widget_setState or a speech-only beat instead of guessing a
target when no stable selector is known (e.g. free-form code widgets). Hardens
against widget actions silently no-oping on non-existent selectors.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-06-28 21:13:21 +08:00
Yizuki_Ame 243e9f3a48 feat: auto-retry transient failures during scene generation (#788)
Adds automatic retry with exponential backoff for scene generation (content/actions/TTS): shared withGenerationRetry helper at the fetch boundary, smart error classification (retries 429/5xx/timeouts/network, fails fast on 401/403/404), AI SDK RetryError unwrapping, abort-aware backoff, per-route thinking config preserved, and upstream HTTP status surfaced for client retry classification. Fixes #755.
2026-06-27 07:20:22 -04:00
wyuc a88ee3d119 feat(editor): add discussion authoring to the script timeline (#798)
Add an append-only 'Add discussion' control + inline editor (topic/prompt/initiating agent) to the script timeline, respecting the action-parser invariant that a discussion is the scene's last action and at most one per scene. Yellow terminal-anchor styling; agent picker (with avatars) sourced from the selected agents the playback engine gates on. Runtime unchanged.
2026-06-26 04:21:35 -04:00
5226f9589f fix: tighten PBL v2 planner eval harness (#805)
Co-authored-by: mzb25 <mzb25@mails.tsinghua.edu.cn>
Co-authored-by: wu-yx25 <wu-yx25@mails.tsinghua.edu.cn>
Co-authored-by: 805813606 <805813606@qq.com>
Co-authored-by: zlnn23 <zlnn23@mails.tsinghua.edu.cn>
2026-06-26 14:22:39 +08:00
杨慎andzlnn23 8f77d61fa4 test: add PBL v2 planner eval harness (#803)
Co-authored-by: zlnn23 <zlnn23@mails.tsinghua.edu.cn>
2026-06-26 13:56:40 +08:00
杨慎andwyuc 0f44e9a46f Add PBL v2 runtime APIs and classroom UI (#799)
* feat: add PBL v2 runtime UI

* fix: ship PBL v2 instructor avatar

* fix: ship PBL v2 workspace logo

* fix: prevent scenario PBL legacy fallback

* chore: restore package lint comments

* test: cover ordinary PBL legacy fallback

* fix: front-load PBL single-call validation

* fix: preserve PBL v2 content when assembling scenes

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-26 11:56:02 +08:00
xuyuanwei678andwyuc 8e39e7e665 fix(importer): port PPTX shape restoration hotfixes (#789)
* fix(importer): port pptx shape hotfixes

* chore: format pptx importer hotfix

* test(importer): satisfy strict typecheck

* fix(renderer): constrain slide element hit target

* test(renderer): keep hit target regression in package

* fix(renderer): preserve thumbnail click-through

* fix(importer): correct warp height and blank date labels

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-25 23:35:28 -04:00
杨慎 09863af421 Add PBL v2 core schema and generation path (#795)
* feat: add PBL v2 core planner

* chore: restore unrelated lint comments

* fix: document PBL v2 scenario outline flags

* fix: expose PBL scenario subtype in outlines

* chore: clear PBL v2 completion stats lint warning

* chore: address PBL instructor code quality comments

* fix: address PBL v2 review findings

* chore: restore unrelated lint comments
2026-06-25 18:44:46 +08:00
wyucandClaude Opus 4.8 8d74639c19 fix(generation): show full key-point text in outline editor (#782)
Key-point chips in the pre-generation outline editor used a single-line
`truncate`, so longer points were cut off with an ellipsis and users could
not read the full text while reviewing the outline. Let the text wrap within
the chip's max width and add a native title tooltip as a fallback.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 04:43:05 -04:00
wyucandClaude Opus 4.8 81dae36280 fix(editor): stop dumping raw tool-failure blob in AI edit tool cards (#785)
A failed tool card rendered the tool result's full text inline. For
framework schema-validation failures that text is a large blob — the
validation message plus the entire received-arguments JSON, including
the full interactive-page HTML — which floods the Edit-with-AI panel.

Remove the inline failure body. The card is non-expandable by design (a
single status row), so failure surfaces through the amber ✗ status mark
and its hover tooltip; the model also restates the actionable reason in
its own reply. Drops the now-unused `failText` prop from ToolCard and its
wiring at the three call sites, and refreshes the stale "surfaced inline"
doc comments.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 15:14:00 +08:00
wyuc 7cb129183d chore(packages): publish the @openmaic/* SDK family to npm (#778) (#780)
* chore(packages): publish the @openmaic/* SDK family to npm (#778)

Prepares the @openmaic/{dsl,renderer,importer} family for its first npm
publish, and moves the SDK packages onto the @openmaic scope.

Why the scope move: the @maic org name is unavailable on npm (an unscoped
`maic` package already holds the name), so @maic/* is not claimable. @openmaic
matches the project name, the scope is free, and the repo already ships an
@openmaic/docs package — so the SDK family now lines up with that convention.

- rename @maic/{dsl,renderer,importer} -> @openmaic/* across packages, the
  workspace glob, the package dir, and all import sites; lockfile regenerated
- renderer: add publishConfig (public, registry.npmjs.org) — was missing, so
  a scoped publish would default to the wrong registry / restricted access
- importer: add a files allowlist (dist, README, LICENSE) and drop the
  fragile .npmignore blacklist that shipped src; add an exports map so ESM
  consumers resolve dist/index.js instead of falling back to the .cjs main
- all three: add a prepublishOnly build (+ test/typecheck) guard so a publish
  can never ship a stale or empty dist
- add a tag-triggered publish workflow with npm provenance, pinned by name to
  the three @openmaic packages so the vendored forks (mathml2omml, pptxgenjs)
  are never published

Refs #778, #720 (Phase 1).

* fix(packages): address cross-review on the @openmaic publish prep

Cross-review (Claude /code-review + codex) on this PR surfaced:

- renderer's advertised CJS entry was broken: it keeps @openmaic/dsl external
  and imports a runtime enum from it, but dsl is ESM-only (no `require`
  condition), so `require('@openmaic/renderer')` would throw
  ERR_PACKAGE_PATH_NOT_EXPORTED. Make renderer ESM-only: drop the `.cjs`
  rollup output, `main` now points at the ESM build, and the `require`
  conditions are removed from `exports`. (importer is unaffected — it bundles
  dsl, so its CJS build still works.)
- prepublishOnly re-ran the test suite during `pnpm -r publish`, so a flaky
  test after dsl had already published gave a non-atomic partial release.
  Reduce prepublishOnly to a build-only guard (never ship stale/empty dist)
  and move the real test/typecheck gate into the workflow, before any publish.
- document that an @openmaic/* tag publishes the whole family via `pnpm -r`
  (pnpm skips already-published versions); the tag is a release marker, not a
  per-package gate.

Verified: dsl + renderer + importer build; renderer emits ESM only (0 .cjs),
all exports entries resolve; `npm pack` ships dist + README + LICENSE with no
src leak; frozen-lockfile passes.

Refs #778.

* style: reflow @openmaic/dsl type imports past print-width after rename

The @maic -> @openmaic rename lengthened two single-line type imports past
prettier's 100-col width; prettier --check flagged them. Pure formatting.

Refs #778.

* docs(importer): mark @openmaic/importer browser-only (cr-loop accepted limitation)

codex cross-review flagged that the published @openmaic/importer throws
`XMLHttpRequest is not a constructor` when loaded in a pure Node process —
its rollup build is browser-targeted (`nodeResolve({browser:true})` + a
browser pdf.js build). The app only consumes it client-side ('use client'),
so this is by design. Document it as an accepted limitation: prominent
browser-only note in the README and a `browser` field in the manifest.

Refs #778.
2026-06-24 15:12:23 +08:00
wyuc fb4ce5341c fix(619): keep-alive e2e aligns with #777 edit-mode visibility (+ pi-ai server-external) (#781)
* fix(agent): mark pi-ai/pi-agent-core server-external to fix #619 e2e

The editor-agent packages `@earendil-works/pi-ai` and `pi-agent-core` lazily
load node builtins via a computed `import(specifier)` (deliberately, to avoid
breaking browser/Vite builds). webpack cannot statically analyze that, so once
these run on the server (the Pro-mode "Edit with AI" path) the bundled form
throws `Cannot find module as expression is too dynamic` as an unhandled
rejection. That broke the #619 interactive-iframe keep-alive e2e: entering Pro
mode no longer settled, so the keep-alive iframe stayed visible
(`interactive-iframe-keepalive-619.spec.ts:189` expected hidden).

Add both packages to `serverExternalPackages` so Next loads them natively and
their dynamic import resolves as a real Node call.

Refs #619, #777.

* test(619): interactive iframe stays visible in edit mode since #777

#777 intentionally dropped the `mode !== 'edit'` guard in InteractiveIframeHost
so the interactive iframe stays visible during Pro-mode editing — the editor
agent ("Edit with AI") fixes interactive HTML, so the teacher must see the live
page while editing. The keep-alive e2e was still asserting the old "hidden in
edit mode" behavior (toBeHidden at line 189) and so failed on every run since
#777 landed.

Update Trigger A to the new contract: after the Pro-mode toggle the iframe is
visible and its in-iframe counter state is preserved (the real keep-alive
proof — neither unmounted nor reloaded). Trigger B (scene switch to a slide)
still hides it via ownership release, unchanged.

Refs #619, #777.
2026-06-24 14:39:21 +08:00
wyuc 42ca1407ef feat(ai): add GLM-5.2 and Kimi K2.7 Code (#774) 2026-06-24 13:14:59 +08:00
wyucandClaude Opus 4.8 dcf8a63333 fix(generation): make outline type changes take effect (interactive/PBL no longer downgraded to slide) (#772)
* feat(generation): changeOutlineType total constructor for valid type switches

Switching a scene outline's type in the editor only flipped `type` and left
the new type without its required config, so applyOutlineFallbacks silently
downgraded interactive/pbl back to slide. changeOutlineType returns a new
outline that strips foreign per-type config and seeds the target type's
required config from the shared fields, valid by construction.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* i18n(generation): keys for interactive widget kind and PBL project config

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(editor): make outline type changes take effect with config forms

TypePill now builds a valid outline via changeOutlineType (replace, not
partial-merge) so a type switch drops stale config and seeds the new type's
required config — no longer silently downgraded to slide by applyOutlineFallbacks.
Adds InteractiveConfigDisclosure (widget kind + concept) and PblConfigDisclosure
(topic / description / target skills), mirroring QuizConfigDisclosure, so users
can configure interactive and PBL scenes that previously had no input UI.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(editor): move type-specific config under the type pill

The quiz/interactive/PBL config disclosure rendered as a detached chip at the
bottom-left of the card, reading like a stray key point. Pin it to the
bottom-right of the key-points row instead, so it sits directly under the
TypePill it configures (type top-right -> its config bottom-right, same edge),
with key points flowing on the left.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(editor): cascade type config as a connected segment left of the type pill

Replace the standalone config chip with a cascading control in the header: the
type-specific config (quiz count / interactive widget kind / PBL project) renders
as a themed segment joined to the left of the TypePill, the two clipped into one
rounded pill split by a hairline divider. Reads as a single 'config › type'
cascade and groups the config with the selector it refines.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(editor): address cross-review — preserve widget config, drop stale fields

- changeOutlineType: preserve an existing procedural-skill widget config
  instead of replacing it with a seeded simulation (codex P2 — task-engine
  fields were silently dropped on Interactive re-select).
- InteractiveConfigDisclosure.setWidgetType: reset widgetOutline to the shared
  concept only, so switching widget kind drops the previous kind's fields
  (claude P2 — language/gameType/etc. leaked across kinds).
- i18n locale-coverage test now also checks ko-KR and pt-BR.
- Drop a redundant WidgetType cast; add tests for procedural-skill preservation
  and pbl targetSkills dedupe/cap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(editor): round 2 cross-review nits — procedural-skill label + skill a11y

- InteractiveConfigDisclosure: when an outline carries a preserved
  procedural-skill widget, surface it as the current (selected) kind with its
  own label instead of mislabeling it 'Simulation', and make a same-kind
  re-select a no-op so its task-engine fields are never silently clobbered.
- PBL target-skill remove button uses a dedicated removeSkill aria-label.
- Add widgetProceduralSkill / removeSkill i18n keys (8 locales) + test coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(generation): make re-selecting the current outline type a no-op

changeOutlineType rebuilt the outline on every type-menu click, so re-selecting
the already-current type dropped fields it doesn't re-seed — legacy
interactiveConfig, or a partial pblConfig whose projectTopic is empty — turning a
harmless click into data loss. Short-circuit same-type selections to return the
outline untouched (codex P2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(format): prettier --write files left unformatted on main

These 14 agent/edit files (not part of this PR's feature) fail the repo-wide
`prettier . --check` in the required Lint job — they landed on main unformatted,
turning main's CI (and therefore this PR after a branch update) red. Formatting
them here unblocks the required check; pure formatting, no logic changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(lint): suppress intentional set-state-in-effect in reasoning timer

The reasoning timer does a deliberate one-shot setState in its effect to show
the right elapsed value before the first 1s interval tick. main's CI fails the
required Lint job on the new react-hooks/set-state-in-effect rule here; suppress
it with a justification (no behavior change). Unblocks the required check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 17:48:07 +08:00
wyuc 1d1ce80e04 feat(maic-agent): editor-agent line (v0) — Pro-mode "Edit with AI" (#777)
Promote the editor-agent line to main: read-then-act editor agent (read_scene_content, regenerate_scene, regenerate_scene_actions, edit_interactive_html), a routable maic-agent model stage, the agent sidebar UI, interactive-scene runtime-error capture, and the timeline/script editor. Cross-reviewed and smoke-verified; see #777.
2026-06-23 16:56:37 +08:00
wyucandClaude Opus 4.8 13b42206f6 fix(generation): don't regenerate a deleted slide on a finished deck (#769)
Deleting a slide in Pro/Edit mode removed it from `scenes` but left its
entry in the persisted `outlines` plan. On the next classroom mount the
resume-generation pass treats the now-missing `order` as a pending
outline and regenerates the slide, appending it as the last page.

Rather than prune outlines by `order` (which is rebalanced by Pro-mode
insert/duplicate/reorder, so it is not a stable key and would mis-target
or break delete-undo), gate resume on a persisted `generationComplete`
flag:

- The flag is recorded (alongside outlines, in the stageOutlines record)
  when generation finishes — both generateRemaining completion paths, the
  retry-to-completion path (markGenerationCompleteIfDone), and the
  already-materialized-at-mount path — and restored on load. It is written
  only after a verified scene flush, so it never outruns the deck's scenes.
- A finished deck is frozen for editing: loadFromStorage no longer derives
  pending placeholders from orphaned outlines, the classroom resume effect
  skips generateRemaining, media resume is filtered to materialized
  outlines, and isCourseComplete honours the flag — so a deleted slide
  neither regenerates nor leaves a phantom pending page.
- Decks generated before the flag existed self-heal on load (every outline
  has a scene ⇒ inferred complete). deleteScene additionally records
  completion when deleting from a deck that is complete right now, so the
  "Course complete" end page and resume-suppression survive even when the
  flag was never persisted (e.g. edited without a reload).

`outlines` are never pruned, so order-rebalancing and delete-undo are
unaffected. Genuinely interrupted generations (not complete) still resume.

Fixes #768

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 12:23:17 +08:00
wyucandwyuc 267965d6ed feat(generation): infer concise courseTitle from outlines for readable course names (#756)
* feat(generation): infer concise courseTitle from outlines for readable course names

Generated courses were named after the raw user prompt (truncated to 500 chars
client-side, or 50 chars / first outline title server-side), which is long and
unreadable. The outline-generation step already understands the full course
semantics, so have it emit a concise `courseTitle` alongside `languageDirective`
and `outlines`, and use it as the stage name.

- Prompts (requirements-to-outlines, interactive-outlines): top-level JSON gains
  a required `courseTitle` (≤30 chars, in the teaching language, noun phrase).
- Streaming route: extract `courseTitle` head-bound like `languageDirective`,
  emit a `courseTitle` SSE event, and include it in the `done` event.
- Non-streaming outline-generator: parse and return optional `courseTitle`
  (defensive trim + 120-char cap).
- Naming: stage.name = courseTitle || outlines[0]?.title || requirement fallback
  (client preview path backfills stage.name after outlines resolve; server
  classroom-generation uses it directly).
- Session: persist `courseTitle` on GenerationSessionState for reload safety.
- Tests: add courseTitle parsing coverage; e2e mock done event carries it.

Fully optional in the pipeline — every consumer falls back to the previous
behavior when the field is absent, so legacy/LLM-skipped cases are unaffected.

* fix(generation): make courseTitle propagation robust across all outline paths

Address review feedback on the courseTitle threading.

Prompts — align every outline-prompt schema statement to the 3-key shape.
The user-prompt "final reminder" / top-level shape blocks for
requirements-to-outlines, interactive-outlines, and task-engine-outlines still
demanded the old two-key {languageDirective, outlines} object; as the last
instruction the model reads, that caused it to omit courseTitle and the stage
name to fall back to the raw requirement. Task Engine prompts now emit
courseTitle as well.

Streaming route — normalize the streamed courseTitle (trim, ignore
whitespace-only, cap length) to match the non-streaming parser, and add a
full-buffer fallback so a title emitted after the outlines array or beyond the
head-scan window is still recovered before the done event.

Client — reset latched languageDirective/title on outline retry so a
succeeding attempt that omits them falls back instead of inheriting a failed
attempt's stale value; also carry courseTitle in the stream-end resolve path
that fires when the stream closes without an explicit done event.

courseTitle remains optional on every consumer path.

---------

Co-authored-by: wyuc <zdq1204@gmail.com>
2026-06-18 18:31:38 +08:00
wyucandClaude Opus 4.8 25cf58d19f feat(model): optional per-stage LLM model routing (#745) (#747)
* feat(model): optional per-stage LLM model routing (#745)

Add a config-only stage -> model map consulted during model resolution,
falling back to DEFAULT_MODEL when unset (zero behavior change unless opted in).

- lib/server/model-routes.ts: parse/validate/cache MODEL_ROUTES env JSON;
  LLM_STAGES registry of routable stages; getStageModel(stage).
- resolveModel: resolution order x-model > stage route > DEFAULT_MODEL > builtin;
  thread optional `stage` through resolveModelFromHeaders/FromRequest.
- Wire each route's resolveModel call site to its stage; classroom-generation
  resolves generate-classroom and web-search-query-rewrite independently.
- .env.example: document MODEL_ROUTES; unit tests for both modules.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): address cross-review findings for per-stage routing (#745)

- classroom-generation: resolve the web-search-query-rewrite model lazily,
  only when explicitly routed and inside the web-search branch, with try/catch
  fallback to the classroom model. Fixes a regression where a misconfigured
  optional route (keyless/invalid provider) threw and aborted ALL classroom
  generation even with web search disabled; also avoids wasted resolution on the
  common no-search path and preserves classroom-model inheritance when unrouted.
- resolve-model: type the stage param as LlmStage across resolveModel/
  resolveModelFromHeaders/resolveModelFromRequest so a mistyped stage literal is
  a compile error instead of silently falling through to DEFAULT_MODEL.
- model-routes: simplify the cache to a process singleton (matching
  provider-config), dropping the raw-string-keyed cache and redundant STAGE_SET.
- .env.example: clarify x-model precedence and that a route to an unconfigured
  provider fails at request time (no startup validation), like a bad DEFAULT_MODEL.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): stage route takes precedence over client x-model (#745)

Cross-review (codex P2 + Claude tracer) found that the browser UI always sends
its saved model as x-model, so the previous x-model > stage order made
MODEL_ROUTES inert for normal UI traffic on exactly the heavy stages the RFC
targets (scene-content, quiz, pbl-chat, chat). Flip resolution to
stage route > x-model > DEFAULT_MODEL: a configured route is the operator's
deliberate per-stage choice and wins, while unrouted stages still honor the
client's x-model. Update docs and tests accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(model): prettier-format changed files for #745 CI

Run the repo formatter (pnpm check / prettier) on the changed routes, the new
model-routes module, and the new tests so `prettier . --check` (CI) passes;
also correct a stale precedence comment in the resolve-model test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): isolate routed model from client connection params (#745)

When a stage route overrides the client's x-model and points at a different
provider, the client-sent apiKey/baseUrl/providerType (for the client's model)
must not bleed onto the routed provider — otherwise a routed Anthropic model
would be constructed with the client's OpenAI providerType/key and fail. A
routed model now resolves its connection params purely from server config, as
if no x-model was sent. Unrouted stages still honor the client params. Adds
tests covering both the drop (routed) and keep (unrouted) cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(model): correct graceful-degradation comment in classroom (#745)

resolveModel does not throw on a missing key; clarify that a keyless route
surfaces later in callLLM (degraded by the outer try/catch), while only a
resolution-time failure is caught at the rewrite re-resolve. Comment-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(model): route scene-content per scene type (#745)

Extend MODEL_ROUTES with composite keys scene-content:<type> for the four core
scene types (slide/quiz/interactive/pbl). getStageModel now resolves composite
`a:b` stages most-specific-first, trimming `:` segments, so scene-content:<type>
falls back to the base scene-content route, then x-model, then DEFAULT_MODEL.
The scene-content route derives its stage from outline.type. Backward
compatible: with no composite key configured, behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(model): per-stage thinking effort + drop hardcoded gpt fallback (#745)

MODEL_ROUTES values may now be {model, effort} (not just a model string) to pin
the thinking effort per stage. Arbitration mirrors model routing: a routed stage
with effort set wins over the client's thinking (effort "none" disables); routed
with no effort uses the model's default and drops client thinking; unrouted
stages keep the client thinking. resolveModel is the single arbiter — body
thinking is threaded in via resolveModelFromRequest. classroom-generation passes
the resolved thinking into its callLLM calls so generate-classroom and
web-search-query-rewrite honor route effort too.

Also: removed the hardcoded `|| 'gpt-5.4-mini'` fallback in resolveModel — if no
model resolves (no route / x-model / DEFAULT_MODEL) it now throws instead of
silently picking a vendor default.

model-routes: getStageRoute returns {model, effort?}; getStageModel delegates to
it. Docs + tests updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(model): route value carries full ThinkingConfig (#745)

Per-stage routing accepts the full unified ThinkingConfig in the route value
({model, thinking:{mode,effort,level,enabled,budgetTokens,excludeReasoningOutput}})
instead of just an effort string. The route's thinking is passed through
resolveModel and normalized per the model's capability by callLLM, so
budgetTokens (qwen), level (Gemini), enabled/mode, etc. all work.

(qwen3.7-plus/max thinking capability now comes from upstream #753, so this no
longer registers them.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): honor routed thinking for chat-adapter; doc fixes (#745)

Cross-review (codex P2 + Claude) found chat-adapter was the only routable stage
that ignored the route's thinking: it took the routed model but built its own
thinkingConfig from the request body, so operator-pinned thinking never applied
and client thinking still leaked onto a routed model. Now pass the client
thinking into resolveModel and use the resolved thinkingConfig (route-pinned for
a routed stage, client otherwise; defaults to disabled for low-latency chat).

Also: clarify DEFAULT_MODEL is now required for server-side stages (the
hardcoded gpt-5.4-mini fallback was removed → resolveModel throws if nothing
resolves), and fix a stale {model,effort}→{model,thinking} doc comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 22:26:37 +08:00
wyucandClaude Opus 4.8 6dd9390192 feat(ai): add Qwen3.7 Plus and Qwen3.7 Max to the model registry (#753)
Adds the Qwen3.7 series to the built-in qwen provider, mirroring the
existing Qwen3.6 model-card shape:

- qwen3.7-plus: 1M context, 64K output, vision, tools, toggleable
  thinking + budget (default on)
- qwen3.7-max: 1M context, 64K output, text-only, tools, toggleable
  thinking + budget (default on)

Both run over the already-configured DashScope OpenAI-compatible
endpoint; no transport changes. Thinking is wired through the existing
qwen request adapter (qwenBudgetEnabled, 0-81920 budget). Output ceiling
and thinking-budget max reuse the Qwen3.6 conventions, since the public
docs don't spell these out per-model.

Closes #752

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 21:54:22 +08:00
wyucandClaude Opus 4.8 beac93ab3d chore: relicense from AGPL-3.0 to MIT
Switch the OpenMAIC root and the in-house @maic/* SDK packages
(@maic/dsl, @maic/importer, @maic/renderer) from AGPL-3.0 to the MIT
License. Updates LICENSE files, package.json license fields, README
badges and license sections (EN/ZH), CONTRIBUTING, and renderer FONTS
note.

Third-party vendored packages are left untouched: packages/mathml2omml
remains LGPL-3.0, packages/pptxgenjs remains MIT.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 04:00:00 -04:00
2f53fdbdc7 fix(renderer): make code entrance animation type line-by-line (#531) (#724)
On first render animStates was an empty Map, populated only via a
setTimeout(0) effect. Every CodeLineRow therefore mounted with
animState undefined, latched mounted=true from its initial props, and
took the non-animated branch. When the states arrived one tick later,
every already-mounted row fired setTyping(true) in the same tick, so
all lines typed simultaneously - the delayed-mount stagger path was
unreachable on first paint.

Compute the first-render typing states synchronously in the useState
initializer so rows see their animState and typingDelay on the very
first render: line 0 types immediately, each later row stays unmounted
until its cumulative delay. The effect now only handles subsequent
line edits (inserted/replaced), with prevLinesRef seeded to the
initial lines.

Behavior note: previously, toggling animate from false to true marked
every line "inserted" (prevLinesRef started empty); now the ref is
seeded so a toggle does not replay the entrance - arguably the less
surprising behavior.

Fixes #531

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-16 23:20:28 +08:00
0989bd5d15 fix(proxy): bypass proxy for loopback hosts and honor NO_PROXY (#718)
proxyFetch routed every request through the configured HTTP(S) proxy,
including loopback targets. A deployment with HTTP_PROXY set would send
a http://localhost:8888 SearXNG (or any self-hosted service) request to
the external proxy, where localhost resolves to the proxy machine and
the request fails.

- Always bypass loopback hosts (localhost, *.localhost, 127.0.0.0/8, ::1).
- Honor the standard no_proxy / NO_PROXY env var: comma-separated hosts,
  `*` wildcard, optional `:port` (with protocol-default ports), and
  curl-style domain-suffix matching (`example.com` matches
  `api.example.com`; a leading dot is equivalent).

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-16 18:15:14 +08:00
wyucand杨慎 b1e5beee0f feat(dsl): promote Stage/Scene lesson skeleton into @maic/dsl (#740) (#743)
* feat(dsl): promote Stage/Scene lesson skeleton into @maic/dsl (#740)

Phase 2 of the @maic SDK consumption epic (#720): move the universal lesson
skeleton into the contract package, with Scene generic so app-side feature
surfaces (Action, widgets, PBL) plug in via type parameters.

DSL (packages/@maic/dsl):
- New src/stage.ts: Stage, SceneType, StageMode, Whiteboard, VideoManifest,
  GeneratedAgentConfig, MultiAgentConfig, SlideContent, QuizContent, and a
  narrowed SceneContent (SlideContent | QuizContent).
- Scene<TAction = never, TContent extends { type: SceneType } = SlideContent |
  QuizContent> — the contract owns only the structure + universal content
  kinds; TContent is constrained structurally so an app can pass a wider
  content union. Defaults give renderers/importers a feature-free skeleton.
- Pure guards isSlideContent / isQuizContent.
- Registered in index.ts barrel; README roadmap item checked off + 'Stage /
  Scene split' section added.
- Added vitest config + test script + stage.test.ts (9 contract tests).

App (lib/types/stage.ts):
- Rewritten as a shim: re-exports the skeleton types from @maic/dsl, keeps
  InteractiveContent / PBLContent (app feature surfaces), and aliases Scene =
  Scene<Action, AppSceneContent> so all ~55 existing import sites keep their
  semantics. The generation re-export block is dropped (verified zero
  consumers import those five types via @/lib/types/stage).

Scope B from #740 (lesson container in DSL, features plug in). Decision:
generic Scene<TAction, TContent> over a registry/extension seam — explicit,
isolatedModules-friendly, no declaration-merge pitfalls.

* fix(dsl): address review — guard genericity, standalone tests, type pin

1. isSlideContent / isQuizContent: relax <T extends SceneContent> to
   <T extends { type: SceneType }> so the guards accept an app-widened
   content union (interactive / pbl) — the generic case Scene<TAction,
   TContent> is meant to support. Add a regression test exercising the
   widened union.

2. tests: resolve @maic/dsl to ./src/index.ts via resolve.alias in
   vitest.config.ts, so 'pnpm --filter @maic/dsl test' is standalone on a
   clean checkout (no dist build required). Verified by deleting dist and
   re-running.

3. Pin the TAction-defaults-to-never type test to .toEqualTypeOf<never[] |
   undefined> instead of the trivially-true .toMatchTypeOf<unknown[] | …>.

* style: prettier-format lib/types/stage.ts (CI prettier --check)

* fix(stage): value-export isSlideContent/isQuizContent from the compat shim

The two discriminant guards are runtime functions but were sitting inside the
shim's `export type { … }` block, which erases them: `import { isSlideContent }
from '@/lib/types/stage'` would resolve to `undefined` at runtime and TS-error
'cannot be used as a value'. No call sites import the guards through the shim
today (latent), but the shim's job is back-compat re-export, so move them to a
plain `export { … }`.

---------

Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-06-16 16:11:15 +08:00
612a1471e7 fix(agent): respond to the user's turn before lecturing (+ answer-content eval) (#699)
* test(orchestration): add answer-content eval for agent first-reply answering

The existing director question-answering eval only checks that the director
ROUTES to the teacher; it never generates the dispatched agent's reply, so it
cannot catch the case where routing is correct but the agent's first sentence
drifts (greets / opens a lecture / pivots) and only a later turn answers.

This adds an answer-content eval that runs the real agent-generate inputs
(buildStructuredPrompt + convertMessagesToOpenAI), parses the structured output
with the app's runtime parser, and uses an LLM judge for two signals:
leads_with_answer (first sentence addresses the literal ask) and
answered_anywhere (addressed at all). Their gap quantifies "drift-then-answer".

A/B mirrors the answering-runner rule-13 strip: baseline removes the
"Answering the User's Question" section from agent-system, with_rule is as-shipped.
Scenarios are synthetic/anonymized (opening-lecture override, ignored
format/capability/navigation request, ignored correction, frustration re-ask,
adjacent pivot, vague-clarify) plus clean controls; no real user data.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(orchestration): give answer-content scenarios live-shaped classroom context

The minimal store state (empty scenes, null stage, no profile) made the
assembled agent prompt far thinner than the live agent-generate path and
suppressed the slide-narration "lecture reflex" that drives opener drift.

Each scenario now carries a topical slide scene (title + body + latex elements
with ids), a stage (name/description/languageDirective), and a student profile,
so buildStructuredPrompt emits a "# Current State", "# Student Profile" and
"# Language" of the same shape and bulk as production. The runner reads this
context from the scenario instead of a hardcoded empty state.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(agent): respond to the user's turn before continuing the lecture

Generalize the agent-system "Answering the User's Question" section into
"Responding to the User's Turn": the user's latest message always takes priority
over continuing the planned lecture. Beyond factual questions it now explicitly
covers the interaction requests agents were ignoring — navigation/pacing
("skip to next page"), format/language ("explain in Chinese"), capability
requests it cannot fulfill ("make a video" -> say so + offer an alternative),
corrections, frustration re-asks, and vague asks (ask one clarifying question).

Strengthen the answer-content eval that measures this:
- the navigation case now uses a 3-slide deck so "skip to next page" is well-posed;
  runner supports multi-scene state and strips the renamed section for the A/B;
- judge is fairer to legitimate behavior: explaining primarily in the requested
  language while keeping technical terms standard counts; a verbal transition to
  the next slide counts as honoring navigation (no page-turn action exists).

Result (gemini-3-flash-preview, n=10): mean leads-with-answer rises from 42%
(section stripped) to 93% (as-shipped); all 12 scenarios pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(agent): don't over-respond to bare acknowledgements; guard empty eval set

Code-review follow-ups on the "Responding to the User's Turn" change:

- Prompt: the absolute "every user message needs a direct response / treat
  continuing the lecture as a failure" wording had no carve-out for bare
  acknowledgements ("ok" / "嗯" / "good" / "thanks"), risking the agent halting
  the lesson to manufacture a reply, and conflicted with agent-initiated
  discussions where no user turn exists. Add a bare-acknowledgement carve-out,
  allow a brief "好的"/"Sure" before the answer, and scope the override to actual
  requests (and to "if there is a user turn to respond to").

- Eval: answer-content-runner exited 0 with a PASS verdict and NaN% aggregates
  when the scenario set was empty (e.g. an EVAL_SCENARIO typo). Guard it: error
  and exit 1 when no scenario matches.

Re-run unchanged otherwise (gemini-3-flash-preview, n=10): mean leads-with-answer
42% (stripped) -> 94% (as-shipped), 11/12 scenarios pass; the chem opener is
variance-prone around the 70% bar and not over-tuned.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(eval): cr-loop cleanups on answer-content-runner

- ScenarioAgentSpec extends the shared ScenarioAgent (./types) instead of
  re-declaring id/name/role/priority.
- Drop the dead expectedPreFix plumbing from ScenarioResult (it was copied in
  but never read; it stays as an informational annotation in the scenarios).
- extractTexts: replace the hand-rolled JSON fallback with the app's
  finalizeParser, which recovers plain-prose / unclosed-array output the same
  way the runtime does.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: tighten acknowledgement carve-out; don't surface action-JSON as text (cr round 2)

Round-2 review follow-ups.

Prompt: the bare-acknowledgement carve-out was too loose — it could swallow
ack+request hybrids ("嗯,那为什么X?") as acks (re-introducing drift), double-listed
"好的" as both an allowed lead-in and a no-response ack, and listed pacing words
("继续"/"go on") that collide with the navigation bullet. Restrict it to a
STANDALONE acknowledgement with nothing else attached; route an embedded
question/request to that request; route pacing words to the navigation bullet.

Eval: extractTexts gated finalizeParser recovery on !state.jsonStarted, so an
actions-only / unclosed array yields "no text" (the judge's intended empty case)
instead of finalizeParser's raw-buffer fallback surfacing action JSON as fake speech.

Re-run (gemini-3-flash-preview, n=5): no regression — 11/12 pass, meta-requests
100/100/80/100; the chem opener remains variance-prone around the 70% bar.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address review — honest slide-nav, strict judge booleans, errors fail (cr)

Review follow-ups from @cosarah:

- Navigation (P2): agents have no slide-navigation action, so "do it" + a verbal
  transition let "skip to next page" pass the eval while the UI stayed on the
  same slide. Reframe prompt, judge, and scenario so a correct reply ACKNOWLEDGES
  the request and is HONEST that it cannot change the slide itself (continue
  verbally / tell the user how to navigate); pretending it flipped the slide no
  longer counts.
- Judge booleans (P2): answer-content-judge coerced fields with Boolean(...), so a
  string "false" became true. Parse strictly (real bool or "true"/"false");
  malformed output is flagged as an error.
- Error accounting (P3): aggregate excluded errored samples from the denominator,
  letting a scenario pass on one good sample. Errors now count as failures
  (denominator = all requested samples).

Re-run (gemini-3-flash-preview, n=10): navigation 0% -> 100% under the honest
framing; aggregate leads-with-answer ~93%, no regression. Opener scenarios
(math/chem) remain variance-prone near the 70% bar.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(eval): judge first-sentence lead; use canonical teacher actions (cr round 2)

Second review round from @cosarah:

- leads_with_answer was judged against texts[0] (the whole first JSON text
  block), so "Today we'll discuss parabolas. The formula is x=-b/(2a)." counted
  as leading even though the first sentence drifted. Now extract the first ~2
  sentences as the "opening" and have the judge make the call: a brief
  acknowledgement of the user before the answer counts as leading, but a greeting
  / self-intro / topic preamble before the answer does not. (Two sentences +
  semantic judge handles "Sorry, let me clarify. m is inertial mass." vs
  "Welcome! …" without a brittle keyword list.)

- The runner hardcoded TEACHER_ACTIONS, omitting canonical teacher actions
  (play_video, wb_draw_chart). Use getActionsForRole('teacher') so the available-
  actions section matches production.

Re-run (gemini-3-flash-preview agent, n=10): 11/12 pass, aggregate leads ~90%.
The chem opener now correctly fails (~10%): its "Welcome! Today we…" persona
greets first, so at sentence granularity it is genuine first-sentence drift
(answered-anywhere ~90%) — exactly what the stricter lead measurement should
catch.

Note: the judge model is deepseek-v4-pro in this run because the anthropic
gateway is currently returning Bad Request in our environment; the eval is
judge-model-agnostic via EVAL_JUDGE_MODEL.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(eval): judge the opening across all text blocks, not just texts[0] (cr round 3)

Review follow-up from @cosarah: leadFromTexts only looked inside the first
text block, so a valid multi-block reply like
[{"type":"text","content":"Sure."},{"type":"text","content":"The derivative is 2x."}]
was judged on "Sure." alone and failed leads_with_answer despite following the
allowed acknowledgement-then-answer pattern. Build the opening from the first ~2
speech sentences across ALL ordered text chunks.

Re-run (gemini-3-flash-preview, n=10, deepseek judge): 11/12 pass, aggregate
leads ~93%. The chem opener stays sub-threshold (~50%) — its "Welcome! Today
we…" persona greets before answering about half the time, a genuine
first-sentence drift.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-06-15 23:55:30 +08:00
Sebastion d89ef645e1 fix(renderer): remove allow-same-origin from interactive srcDoc iframe sandbox (#726)
Combining allow-scripts with allow-same-origin on a srcDoc iframe negates the sandbox; since the embedded HTML can come from LLM output or imported data, the iframe now runs in a null origin. postMessage control (targetOrigin='*') is unaffected. Adds a regression test. Fixes #725.
2026-06-15 16:45:41 +08:00
wyucandClaude Opus 4.8 4dce1f7414 refactor(types): import the slide DSL directly from @maic/dsl; drop the shim (#738)
#707 left lib/types/slides.ts as a thin `export * from '@maic/dsl'` shim so the
~100 existing `@/lib/types/slides` import sites kept working unchanged. Point
them at `@maic/dsl` directly and delete the shim, so the app consumes the
package contract with no indirection and the legacy module path is gone.

Pure module-specifier migration — every symbol imported from
`@/lib/types/slides` was already a re-export of the same `@maic/dsl` symbol, so
behavior is identical. Also redirects the two `lib/types` relative importers
(`./slides`), the e2e fixture's relative import, and the dynamic
`import('../types/slides').Slide` type alias in stage-storage; merges the now
same-module import pair in slide-edit-elements.

Verified: tsc --noEmit clean, eslint clean, vitest 804/804, next build OK,
e2e (Playwright) 19/19.

Part of #720 (Phase 1).

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-15 15:04:06 +08:00
254ec743db chore(phase-1): drop dead ThumbnailSlide + fix @maic/dsl Node ESM resolution (#736)
* refactor(slide-renderer): remove dead legacy ThumbnailSlide/ThumbnailElement

These were the in-app read-only thumbnail renderers replaced by SlideThumbnail
(→ @maic/renderer) in #707. They have zero importers now — the editor nav rail
renders through SceneThumbnailContent → SlideThumbnail; the only remaining
"ThumbnailSlide" mentions are an unrelated local type alias / util name and
historical doc comments. Delete the dead code (~215 LOC).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(dsl): emit explicit .js extensions so @maic/dsl resolves under Node ESM

The build used `moduleResolution: bundler` + extensionless re-exports, so the
emitted `dist/index.js` did `export * from './slides'` — which fails bare-node
ESM resolution (`ERR_MODULE_NOT_FOUND`), even though Next's bundler tolerated
it. Switch the package tsconfig to `NodeNext` and add `.js` extensions to the
relative specifiers so tsc emits resolvable ESM. `node -e import('@maic/dsl')`
now loads; NodeNext also enforces extensions going forward.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-06-15 14:05:03 +08:00