135 Commits
Author SHA1 Message Date
Xinyu Mu 1635b16899 docs: add multilingual vocational task engine guide (#1749) 2026-10-01 21:01:09 +08:00
fa47efe1b3 feat(persistence)!: server-backed persistence only, with a one-way legacy browser import (#1710)
* ci: run CI for the persistence-default integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(docker): start PostgreSQL by default and add a single-user owner

* feat(docker): start PostgreSQL by default and add a single-user owner

`docker compose up` now runs a server-backed app against a bundled
PostgreSQL, started only once PostgreSQL is healthy and published on
127.0.0.1 only, with a new built-in singleUser owner auth method: every
request resolves to one fixed owner that may publish, and no anonymous
cookie is minted.

Single-user mode (OWNER_SINGLE_USER=true, OWNER_SINGLE_USER_ID default
"local") runs with or without ACCESS_CODE. Without one, the server logs
one prominent startup warning that anyone who can reach it shares, edits
and can delete the single library; it never inspects request peers or
forwarding headers. It excludes PERSISTENCE_SHARED_OWNER_ID and follows
the sharedTeam registration rules. Its principal gets a claim candidate,
so a browser's earlier anonymous work is claimed into the single owner
(OWNER_CLAIM_TRIGGER=auto in the Compose defaults).

Compose defaults live in docker-compose.defaults.env, read before
.env.local so they can be overridden. The app warns when its declared
published address (OPENMAIC_PUBLISH_ADDRESS, passed by Compose) is not
loopback while PostgreSQL uses the default password.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(docker): keep claims explicit and let .env.local set DATABASE_URL

- Compose no longer defaults OWNER_CLAIM_TRIGGER to auto: an irreversible
  merge of every browser's anonymous library into the single owner must
  not happen on upgrade. The single-user principal still gets a claim
  candidate, so POST /api/identity/claim works; the docs explain it and
  warn about auto on a deployment several people used.
- The bundled DATABASE_URL moves to docker-compose.defaults.env, read
  before .env.local, so an external database or a rotated password set
  there keeps working. The password is inserted unencoded, so it must be
  letters and digits (documented).
- Upgrade docs no longer claim browser-stored courses are copied as they
  are opened: they stay in the browser, are not deleted, and are moved by
  the automatic browser-to-server migration shipped with this work.
- Single-user mode without ACCESS_CODE now logs one warning instead of
  two; the docs note that the single owner id is permanent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence)!: server-backed persistence is the only persistence (#1707)

* feat(persistence)!: make server-backed persistence the only persistence

Courses, folders and folder membership, chat history and learner runtime,
and generated media are now always stored on the server through the
embedded /api/persistence endpoint. The browser keeps no durable data of
its own.

- Remove the build-time NEXT_PUBLIC_PERSISTENCE switch: the client
  bootstrap always configures the HTTP document, runtime and asset seams,
  and the Dockerfile, docker-compose.yml, the Pi route and the whiteboard
  runtime no longer read it. Unconfigured seams refuse to resolve instead
  of falling back to IndexedDB, and the learner key is always the
  server-derived one.
- Remove the browser write paths: the browser document, runtime and asset
  stores as backends, the browser-mode branches of stage, folder, chat,
  media, narration and import storage, the Dexie backup export/import and
  the whole-database clear. The library lists through /api/stages, folders
  through /api/folders, and membership now goes through
  /api/folders/members. Regeneration always forks to a fresh asset.
- Device-local state moves to its own IndexedDB database
  (lib/device-storage, maic-device-cache): the local media and narration
  cache and refused-bytes retention, staged PDF images, undo history,
  browser voice profiles (carried over once from the old database) and
  the auto-voice cache. "Clear Local Cache" clears exactly that and no
  longer claims to delete classrooms or chat history.
- Keep the pre-server data readable for the one-way importer: the Dexie
  schema and read-only accessors for its tables and for the browser
  document, runtime and asset databases live in lib/legacy-browser-storage,
  which nothing on the regular paths reads and which has no write path.
- Require a database: the server exits at boot without DATABASE_URL, with
  a message naming `pnpm db:up` and `docker compose up`. `pnpm db:up` /
  `pnpm db:down` start and stop the Compose postgres service alone,
  published on 127.0.0.1 through docker-compose.db.yml; .env.example
  carries the matching local DATABASE_URL.
- E2E specs seed their courses through the persistence endpoint, each
  under a fresh id, instead of writing IndexedDB.
- Guard tests pin the boot requirement, the absence of any
  NEXT_PUBLIC_PERSISTENCE read, and the legacy-storage boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: run the E2E suite against PostgreSQL

The app refuses to start without DATABASE_URL, so the E2E job gets a
postgres:16 service and a job-level DATABASE_URL. The unit, storage and
render-service jobs are unchanged and need no database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document always-on server persistence and the database requirement

README (English and Chinese), the deployment guide in every locale, the
extension cookbook and the changelog now describe PostgreSQL as required:
`pnpm db:up` for local development, `docker compose up` for a deployment,
and an external database for Vercel and other serverless hosts (the
Deploy button prompts for DATABASE_URL). They list what stays in the
browser, drop the NEXT_PUBLIC_PERSISTENCE instructions, and say that
courses an earlier browser-only build stored stay in the browser until the
one-way importer moves them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pbl): use the server-derived learner key for PBL v2 runtime

PBL v2 synchronization, hydration and drain resolved the learner key from
their default watermark KV store, which reads or mints a device key in
localStorage instead of asking runtime configuration. The runtime store is
the server one, which refuses any key but the owner's, so saving a course
with a PBL v2 scene failed before the document write, and the minted key
landed in the slot the one-way importer reads as the pre-server runtime
partition.

The learner key now comes from getLearnerKey(args.kv): the configured
server-derived key, or an explicitly injected store. The default KV store
still holds the drain watermarks and nothing else.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): pin that clearing the cache keeps legacy databases

- Seed the four pre-server databases (MAIC-Database, maic-documents,
  maic-runtime, maic-asset-pool) with real rows, run Clear Local Cache, and
  assert every database and row survives; assert the device database name
  differs from every legacy name.
- The legacy-module import rule now also catches relative static,
  side-effect, dynamic and require imports at any depth.
- New rule: the runtime learner key is never read from a KV store the
  caller made up, only from configuration or an injected store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): run the development database as its own Compose project

`pnpm db:up` / `db:down` shared the checkout's default Compose project with
`docker compose up`, so they recreated or stopped a running stack's
database. docker-compose.db.yml now declares its own project
(`openmaic-dev-db`), which gives the development database its own
container and data volume; it no longer touches the stack, and the two do
not share data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: start the development database before `pnpm dev`

CONTRIBUTING, the getting-started guide in every locale and the startup
modes reference now run `pnpm db:up` and set DATABASE_URL before
`pnpm dev`. README, the deployment guides, the cookbook and the changelog
describe `pnpm db:up` as a separate development database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): sharpen the legacy-storage and learner-key guards

- The Clear Local Cache test closes its seeding connections to
  maic-documents and maic-runtime, so a regression deleting either fails
  on the assertion that names the database instead of by timeout.
- The legacy-module and Dexie import rules accept comments inside
  import() / require(), so a webpackIgnore hint no longer hides one.
- The learner-key rule also flags an injected-looking name that is bound
  to a default local store in the same file; its remaining limit is
  documented, with the behavioural PBL test as the guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): pin the development database's Compose project

`pnpm db:up` / `db:down` now pass `-p openmaic-dev-db`, which takes
precedence over COMPOSE_PROJECT_NAME from the shell or a .env file; the
file's `name:` alone could be overridden and point the scripts at a
running stack's database. The compose test pins the flag. The docs (all
locales) and the file header say the development database is shared by
every checkout on the machine, and how to run a separate one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): import legacy browser data into the server, once (#1708)

* refactor(persistence): prepare the seams the legacy browser importer uses

- Legacy module: read-only accessors for the auto-voice cache, the course
  ids the old media and narration tables name, the pre-server learner key,
  and asset bytes from the browser asset pool.
- Narration adoption exports its row ownership rule so the importer applies
  the same one to rows of the old database.
- The quiz legacy migration is split into a non-deleting core
  (importLegacyQuizSnapshot) and the regular path, which still deletes the
  keys it migrated.
- The lazy document migration exports its snapshot canonicalizer.
- Clear Local Cache keeps the importer's per-owner ledgers, as it keeps the
  pre-server learner key.
- A library-changed window event lets an open library list courses that
  arrive in the background.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(persistence): import legacy browser data into the server, once

On the first load after the upgrade, the client moves what earlier builds
kept in this browser to the server, for the owner the server resolves:
courses from the browser document store and the original tables (with
chat, learner runtime, playback position, roster, folders and membership,
pre-runtime quiz state), their media bytes through commitToPool and the
existing write-back funnels, the bytes of server courses from earlier
opt-in server builds that only the old tables hold, and device-only rows
(failure records, refused bytes, auto-voice clips) into the device cache.

It is automatic and silent, runs after the page is idle, never writes to a
legacy store, and records every step in a per-owner ledger so it resumes
after a crash and never imports twice. Runs are serialized across tabs
with Web Locks. A course the owner already has stays authoritative; one
whose id another owner holds is imported under a derived fresh id; one the
owner deleted on the server stays deleted. Transient failures retry on a
later load with backoff; permanent ones are recorded per item.

The module is temporary and self-contained; its README lists the removal
steps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin that each owner keeps its own import ledger

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): run the legacy import against the app routes and a browser

A PostgreSQL suite routes every request the importer and the app seams
send to the real persistence, library and folder handlers, owner resolved
from the anonymous cookie: a full import, a course another owner holds, a
course the owner deleted, and an existing course whose media is filled in.
An end-to-end spec seeds an old pre-server database in Chromium, loads the
app, and checks that the course reaches the library with its narration
uploaded, stays in the browser untouched, and is visible from a fresh
browser context with the same owner cookie.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: describe the one-way import of legacy browser data

The README (and its Chinese version), the deployment guide in every locale
and the changelog now say what the importer does instead of announcing it:
courses an earlier browser-only build kept in the browser move to the
server automatically on the first load after the upgrade, and the browser
copy stays untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): probe server assets through the shared asset-URL owner

The importer asked the pool directly whether an id exists, which the
asset-URL ownership boundary forbids outside use-asset-url. It now uses
assetRefExists, the existing metadata-only probe.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): guard the importer's pool probe like every other entry point

Only a reference the pool could have issued is probed, as the lease guard
requires; the importer is added to the list of guarded entry points.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import legacy browser data once per browser

Review round 1 of the legacy browser importer.

- Once per browser, not once per owner: the data belongs to whoever used
  the browser before the upgrade, so the first owner the import runs for
  claims it and any other owner gets nothing. The ledger is one key,
  maic:legacy-import:v2, recording the owner only as a SHA-256 digest
  (an anonymous owner id is a bearer credential) and a random salt.
- Handoff: when the claiming owner is retired (403 OWNER_RETIRED), or the
  current owner holds courses or folders the importer created (a claim
  moved them), the current owner continues the unfinished items; courses
  already done are never imported again, so a claim neither duplicates nor
  resurrects them, and a half-copied course keeps its runtime.
- Fresh ids come from the salt and the legacy id through SHA-256, not from
  the owner: unlinkable, unpredictable, and the same in every tab.
- 401 pauses with backoff instead of closing the import; 403
  FORBIDDEN_LEARNER stops the run and leaves the item pending.
- Without Web Locks: ledger writes merge with the stored copy, and the
  library is listed again right before a course is placed, so a second tab
  does not import a course the first one just created.
- Media the server refuses for good gets the app's failed-media record
  instead of a dangling reference; the legacy bytes stay.
- A legacy whiteboard or PBL session is not created when the server
  already has an active one of that kind.
- A folder that is gone leaves the course unfiled instead of failed, and a
  membership naming a folder the old database lacks counts as unfiled.
- Narration rows from before the course column are adopted when exactly
  one legacy course names their key; unconvertible chat rows are noted; an
  unreadable document-store copy falls back to the original tables; PBL v2
  documents go through the save-path strip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the importer's handoff, tabs, refusals and review gaps

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: say the legacy import runs once per browser

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the first owner to load
the upgraded app in a browser claims its legacy data, what the handoff to
an account does, that the ledger holds no owner id, and how refused media
and the new pause and skip cases behave.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): note runtime the importer cannot find without the old learner key

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): let an owner confirm it absorbed a given owner through a claim

GET /api/identity/merged-from?salt=<hex>&digest=<hex> answers whether a
claim merged into the requesting owner an owner whose id hashes to
SHA-256(salt, id). It reads only the requester's own owner_merges rows,
never lists or names an owner, resolves the owner like every other
owner-scoped route, and is uncacheable. The one-way import of pre-server
browser data uses it to decide whether a different owner may continue an
import the first owner started.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): hand the legacy import to another owner only on a confirmed claim

Review round 2 of the legacy browser importer.

- The only way the browser's legacy data moves from the owner that
  claimed it to a different owner is the server confirming, through
  GET /api/identity/merged-from, that the current owner absorbed that owner
  through a claim. The inferred signals (an owner that never wrote, a
  library listing that holds imported courses) and the retired flag are
  gone; a 403 OWNER_RETIRED only stops the run and leaves items pending.
- The first owner is bound only once the run's first authenticated request
  of its own (the library listing) succeeded, so a run that dies before
  that leaves the ledger unclaimed.
- The owner is recorded as SHA-256 of the per-browser salt and the owner
  id; the server computes the same value from the salt the browser sends.
- A "not absorbed" answer is reused for an hour instead of asking on every
  load, and the claiming owner stored first survives another tab's save
  unless that tab handed the import over; completion survives too.
- Run-level failures (401, OWNER_RETIRED, FORBIDDEN_LEARNER, OWNER_BUSY)
  from any call stop the run and leave the course pending, including the
  "is the id taken" read and the library re-list, which could previously
  mark the course failed and close the import.
- Refused media without a generation request (the user's own) gets a new
  non-retryable ASSET_REFUSED code, so it shows as failed without a Retry
  that could only fail; a skipped singleton runtime session leaves a note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin the importer ledger's owner merge and the singleton-skip note

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the legacy import moves to another owner only on a confirmed claim

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the import runs once per
browser and moves to another owner only when that owner claimed the
original one, confirmed by GET /api/identity/merged-from; that the owner
is recorded as a salted digest; and how refused user media shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): never let a transient failure settle a legacy import item

Review round 3 of the legacy browser importer.

- A read of the old browser stores that fails in storage (an aborted
  IndexedDB transaction, a closed database) now pauses the run and leaves
  the course pending. Only a record that cannot be migrated, parsed or
  validated is skipped. This covers the course read (which no longer falls
  back to the older table copy on a storage failure), and the quiz-scene
  and speech indexes, which used to read a failure as "no scenes" or "no
  holder" and drop quiz state or narration for good.
- After a session-id collision, a failed read of the session propagates
  and is retried; only a successful read that shows the id taken skips it.
- Without Web Locks, two tabs of different owners can both find the
  ledger unbound. Every ledger write keeps the owner stored first, and the
  run now re-reads the stored ledger after binding, before folders, before
  each course and before each course document is created, and stops if
  another owner holds it.
- The "not yours" answer is reused for ten minutes instead of an hour, a
  stored wait further ahead than the longest the importer sets (a clock
  that ran ahead) counts as passed, and expired answers are pruned.
- Generated media refused without an error code keeps its Retry; only
  media nothing can regenerate gets ASSET_REFUSED. A poster that fails
  transiently leaves the video pending instead of being dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover transient reads, racing owners, the not-yours cache and the claim client

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): say what the ledger's salt protects, and the new pauses

The salt defeats precomputed and cross-browser tables, not a targeted
guess against one browser's ledger; the module README and the digest's
comment now say so. The README also covers storage read failures, which
pause instead of skipping, the two-owner race without Web Locks, and the
ten-minute recheck.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): bind a browser's legacy data to one owner on the server

The one-way import of pre-server browser data now asks the server whose
data it is, instead of deciding in the browser.

- legacy_import_bindings (browser_id primary key, owner_id, timestamps),
  provisioned with owner_merges on fresh and upgraded databases.
- POST /api/identity/legacy-import-binding {browserId}: one atomic insert
  (ON CONFLICT DO NOTHING) and a read-back; answers only whether the
  requesting owner holds the browser. Same-origin JSON, resolved and
  write-fenced like every owner write, 400 for a malformed id.
- A claim participant re-keys the claimed owner's bindings to the account
  inside the claim transaction, so the account simply holds the browser.
- Owner resolution refuses any request carrying X-OpenMAIC-Legacy-Import
  with 409 LEGACY_IMPORT_NOT_BOUND unless the owner it resolves to holds
  that browser: one central fence every owner-scoped route goes through.
  Requests without the header are unaffected.
- GET /api/identity/merged-from is removed; the binding replaces it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import through fenced clients bound by the server

Review round 4 of the legacy browser importer.

- The ledger holds a random browser id and no owner information; the
  owner digest, the "not yours" cache and the client-side handoff are
  gone. A run first asks the server to bind the browser and imports only
  when the requesting owner holds it.
- Every importer request goes through its own clients that carry the
  browser id, so the server refuses any write under an owner that does not
  hold the browser: after a cookie switch in another tab, from a page whose
  memoized owner is stale, or from a tab that lost the binding race. The
  runtime learner key is asked for during each run, never taken from the
  page. 409 LEGACY_IMPORT_NOT_BOUND stops the run with items pending.
- The write-back funnels, commitToPool and the PBL save-path strip take an
  optional store, put or runtime, so the importer passes its fenced ones.
- A failed read of one legacy course holds up only what it could affect:
  the quiz-scene and speech indexes leave the course out as unindexed, and
  only quiz state or narration it could own waits. Read failures are
  counted per course; after five failing runs over at least a day, a
  course whose document-store copy cannot be read falls back to the table
  copy, or is skipped with the reason when there is none.
- Runtime ids are rewritten by replacing only their course segment, so a
  short course id cannot corrupt the prefix or the learner segment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the server binding, its fence, bounded reads and short ids

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the server binds a browser's legacy data to its first owner

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say: once per browser; the server
binds the browser to the first owner; a claim carries the binding to the
account; the importer's requests are refused for any other owner. They
also describe the bounded pause for a course the old storage cannot read,
and the removal steps for the server side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): keep a dropped asset request pending in the legacy import

The asset client reports a fetch that got no answer (a dropped request, an
aborted upload, the existence probe's deadline) as status 0 with
HTTP_REQUEST_FAILED, which the importer read as a final refusal: the media
was marked failed and the course could complete without it. Read that
code as transient, keep the client's local validation failures final, and
treat a 2xx/3xx answer a client could not use as transient too. Unit tests
drive every client the importer uses with a rejected fetch and a non-JSON
answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): ask a refused binding again only every ten minutes

A browser whose legacy data another owner holds sent the binding request
(a write) on every page load. Record a ten-minute recheck in the ledger
(clamped like every other deferral), so a claim that moves the binding is
still picked up, and open the old databases only once the browser is
bound. The read budget now counts consecutive failing runs (a run that
reads the course resets it) and counts a first failure dated in the
future from the next one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): answer 503 with the cookie when the import fence cannot read

The legacy import fence now reads the header name from the shared
constant, and a database error while reading the binding answers the
usual JSON error (503 PERSISTENCE_UNAVAILABLE) with the resolution's
Set-Cookie values instead of escaping as a bare 500. Route tests cover
that, a bind by an owner a claim retired (403 OWNER_RETIRED, no row), and
one header name on both sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover dropped asset requests, the recheck and the read budget

With the real asset client's error for an upload and an existence probe,
the media stays pending and a later run imports it. A non-holder's five
loads send one binding request and never open the old databases, and a
claim still hands the import to the account after the delay. The read
budget's run count and time span each hold on their own, a clock that ran
ahead does not hold a course open, and a good read resets the count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): complete the legacy import removal steps

List every file the removal touches (the claim scenarios test, the claim
participant table and list, every doc), keep the bindings table for one
release after the code goes because of rolling deploys and give the DROP
statement for that release. Add the claim's eighth participant to the
README lists, fix the fresh-id wording in the changelog, and describe the
recheck delay, the transient status-0 asset failures and the consecutive
read budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): keep legacy quiz state through Clear Local Cache until the import completes

Pre-runtime quiz drafts, answers, results and attempt ids exist only in
localStorage, and the one-way importer copies them to the server with their
course. Clear Local Cache deleted them even while the import was still
pending (it waits for an idle page and can stay pending after a failure), so
that quiz progress was lost for good.

Clear Local Cache now keeps the four quiz key families unless the importer's
ledger records the import as complete, read through a new
`legacyImportIsComplete` helper in the ledger module. The ledger key moves
there too, so the two modules do not import each other. Once the import is
complete, the keys are cleared as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: show DATABASE_URL set in the local development setup

The getting-started guide showed the required DATABASE_URL line commented
out, so copying the block left the server unable to start. Each locale now
shows `pnpm db:up` and, separately, the `.env.local` line to set.

The `.env.example` header said every variable is optional; it now says that
provider variables are, and that DATABASE_URL is required except under
docker-compose.yml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): drop two upgrade notes the final code contradicts

The Compose entry said its default DATABASE_URL overrides .env.local; the
defaults file is read first, so a value in .env.local wins, as the same entry
says further on. The runtime-session entry offered turning server
persistence off as an upgrade path; that switch no longer exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(env): restore the ACCESS_CODE warning notes in .env.example

A later change to the example file dropped the lines saying that an unset
ACCESS_CODE fails open, that the server then logs a one-time startup warning
and GET /api/health reports accessCodeConfigured: false. They are back, and
say that single-user mode logs its own warning in place of the generic one,
as the startup code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* ci: run every app PostgreSQL suite in the contract job

The app-domain step named its suites one by one, and seven had been left
off, among them the legacy importer's route suite. No other job gives the
app suites a database, so they skipped everywhere. The step now lists every
app `*.pg.test.ts`; all of them work in a schema or database of their own,
or truncate only their own tables, and pass before and after the storage
package suite in the same database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): establish the anonymous owner before the first API request (#1721)

* fix(identity): establish the anonymous owner on the page response

On a browser's first load the page sent several API requests without an
owner cookie; each minted its own anonymous owner and set its own cookie,
and the last answer to arrive won. Work done in between (the legacy import
binding, a first course write) then belonged to an owner the browser no
longer presented.

The middleware now mints the anonymous cookie on document navigations that
carry no valid one, with the route handlers' own value format and
attributes (the cookie module is Edge-safe now: Web Crypto instead of
node:crypto), and forwards it on the request so the page render resolves
the same owner. API, RSC, prefetch and Server Action requests never mint
there, and a valid cookie is never replaced. It mints only when neither
OWNER_SINGLE_USER nor PERSISTENCE_SHARED_OWNER_ID is set; host
registrations are invisible to an Edge middleware, so the README tells
hosts where to skip it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(legacy-import): bind only an owner the browser already presents

A binding request whose owner was minted by that very request answers
409 OWNER_NOT_ESTABLISHED with the minted cookie and binds nothing: other
cookieless requests may still be minting owners of their own, and a
binding to this one could be left with an owner nobody presents. The
importer treats it as transient and retries on a later load. It also says
plainly when a request after its own successful bind is refused as
another owner (the owner cookie changed mid-run); items stay pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): cover a first visit whose first answers arrive out of order

Against the real routes and PostgreSQL: a browser without a cookie loads a
page, sends two requests at once, and the first answer reaches it only
after the importer bound the browser; the import completes under the one
owner the page established. A second case binds nothing for an owner the
bind itself minted and imports on the next run. The E2E drives the same
ordering in Chromium by holding the first owner-scoped answer until the
binding answered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): drop an unused initial assignment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): note the page response now sets the anonymous owner cookie

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): let hosts turn page-response minting off

The middleware cannot see owner auth methods a host registers, so a host
with anonymousFallback: false still got an anonymous cookie on a cookieless
page load. Beside the host credential it is a claim candidate; with
OWNER_CLAIM_TRIGGER=auto it is claimed and cleared on the next request,
again after every such page load.

OWNER_ANONYMOUS_PREMINT (default on; true/1/false/0, anything else fails
startup) is read by the middleware before minting. At startup the server
warns once when a registration turns the anonymous fallback off while
pre-minting is still on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): document OWNER_ANONYMOUS_PREMINT

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): keep the anonymous identity for 400 days and renew it in use

The anonymous cookie had a fixed 30-day lifetime from its mint and was
never renewed, so every anonymous visitor lost their whole library 30 days
after the first visit however active they were; with persistence on the
server only, that affects every anonymous deployment.

It now lasts 400 days (the browser cap, one shared constant), and every
route handler and Server Action response that resolves to a valid
anonymous cookie re-sends the same value with a fresh Max-Age (best-effort
where next/headers refuses writes). Page responses never renew, so pages
stay cacheable, and a host principal never renews the anonymous cookie
beside it.

A response that clears the cookie (a claim, a retired owner) must not
renew it, whatever order a route merged its headers in:
withRequestOwner, and the owner-events retired answer, keep only the
clearing value for a cookie a value clears. A renewal that lands after a
claim cleared the cookie is pinned by a test: the next write with it,
alone or beside the account, writes nothing, and its answer clears it
again without renewing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): 400-day renewed anonymous cookie; when to turn pre-minting off

Hosts set OWNER_ANONYMOUS_PREMINT=false only when their registration sets
anonymousFallback: false; hosts that keep anonymous visitors need
pre-minting and skip it per request in middleware for requests their
methods authenticate. The CHANGELOG notes the new lifetime and renewal,
and that a lost cookie means a new owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): forward owner cookies from routes that resolve the owner directly

The Pi chat and whiteboard-visibility routes and asset-id document
extraction resolved the owner themselves and never sent its Set-Cookie
values back, so an anonymous identity used only through them was never
renewed (and a minted one never stored).

attachOwnerCookies attaches a resolution's cookies to a response a route
built itself, with the clear-wins rule, before a stream's body starts. Both
Pi routes answer through it, success and error alike; resolveServerAsset
hands the cookies back with every answer and the extraction route attaches
them.

A guard test scans for every module outside the identity seam that calls
the owner resolution (or the asset helper) directly and requires it to
forward the cookies; the two callers that cannot are listed with the
reason. Behaviour tests cover the three paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: wyuc <dhq1204@yahoo.com>
2026-09-29 17:51:03 +08:00
Ethereal49 877bf792d7 fix(generation): preserve template literals in interactive HTML (#1706) 2026-09-29 15:58:28 +08:00
wyucandClaude Opus 5.5 a686481e73 feat(auth): owner identity seam — composable auth methods, owner-keyed persistence, host hooks, claims
* ci: run CI for the owner identity seam integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 1 — pluggable authenticator with today's behavior

* feat(auth): add the owner identity seam with built-in authenticators

Introduce lib/server/identity/: an OwnerAuthenticator contract
(SubjectKind, OwnerPrincipal, AuthOutcome), a single-shot, boot-validated
configureOwnerAuthenticator() registry, and the anonymousCookie and
sharedTeam built-ins that reproduce today's owner resolution exactly
(same anonymous_id cookie, owner id strings, statuses and env semantics).

Every owner-scoped route handler and the workspace Server Action now
resolve their owner through one memoized per-request resolution instead
of the three previous copies (agent-runtime/owner.ts, with-owner.ts and
the Server Action re-implementation). Publish and unpublish decide from
the principal's course:publish role instead of an anon: id prefix.

An invalid credential from an authenticator is answered with 401 on
every surface and never falls back to an anonymous owner. A source scan
keeps the anonymous owner cookie and anon: prefix checks out of code
outside the identity module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): describe the owner identity seam and host registration

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): tighten the identity boundary guard and Server Action cookies

- Scan workspace package sources too, flag imports of the concrete
  built-in authenticators outside lib/server/identity, and police the
  slice/substring/indexOf/includes spellings of an anon: id-shape check.
  Stop re-exporting the built-ins from the host-facing index.
- Refuse setCookies returned by authenticateFromContext, as the
  authenticate() fallback already did, and document that it must write
  cookies itself through next/headers.
- Let the route-test owner stub carry explicit kind/roles, and cover the
  403 forbidden branch of publish and unpublish for a signed-in owner
  without course:publish.
- README: registration conflicts fail the boot; an invalid principal is
  rejected per request with a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

Runtime sessions and assets on /api/persistence now take their identity
from the same memoized owner resolution as documents, instead of a
client-chosen x-learner-key behind a development bearer token that
shipped in the public bundle.

- Runtime: the storage handler's learner key is principal.ownerId. A
  path or body naming another learner answers 403 FORBIDDEN_LEARNER and
  another learner's session answers 404, whatever x-learner-key says; the
  header is no longer read. The browser learns its key from the new
  GET /api/persistence/learner-key and sends no credential of its own.
  The Pi whiteboard routes derive the same key from the owner.
  PERSISTENCE_DEV_TOKEN, NEXT_PUBLIC_PERSISTENCE_TOKEN and
  PERSISTENCE_ALLOW_INSECURE_DEV_AUTH are gone from the server path and
  the client, and server-auth.ts is deleted. Learner merge stays refused.
- Assets: allocations land in a per-owner partition (owner:<ownerId>),
  so quota is per owner and replace/delete reach only the owner's
  entries. Reads stay capability-by-id where a course viewer needs them:
  another owner's committed entry is readable while a live course
  references it. Legacy entries in the old shared partition stay
  readable by id, and may be replaced or deleted only by an owner who
  owns every course referencing them. Server-generated media is stored
  under the run owner's partition; asset-id extraction and vision image
  resolution use the same rule.
- Tombstones: runtime of a deleted course reads as absent (404, empty
  lists) and takes no new sessions or records, through a guard in the
  app's runtime composition; no package schema change.

The identity boundary scan now also fails on any source that reads
x-learner-key or the retired development token variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): describe owner-keyed runtime and assets, drop the dev token

Remove the development persistence token from the server-persistence
recipe, .env.example, the Docker build arguments and the deployment
docs; describe runtime learner keys, per-owner asset partitions and
quota, the legacy shared-partition rule and the tombstone behavior; and
record the breaking change for existing server-persistence runtime data
under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(storage): scope asset references to principals, typed stage refusal, uniform create collision

@openmaic/storage 0.32.0.

- PgDocumentStore takes assetReferencePrincipals: with reference
  tracking on, a write references and commits only entries held by the
  listed principals. An id naming another principal's entry records no
  row and changes no lifecycle column, like an unknown id, and the write
  is never refused. Without the option any held entry is referenced, as
  before. AssetCollector takes a matching per-document-owner function so
  the one-time backfill scopes each document the same way. This keeps a
  document naming a leaked id from committing, exposing or pinning
  another owner's allocation.
- RuntimeStageNotFoundError: a store may refuse createSession for a
  stage its host considers absent; the runtime handler answers
  404 STAGE_NOT_FOUND, recognizing the error by class or by code.
- A session create over a taken id answers 409 SESSION_ALREADY_EXISTS
  whoever holds it. It used to answer 403 for another learner's session
  and 409 for one's own, which told a caller whose session an id was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): close owner-seam boundary gaps in references, tombstones and legacy assets

- Document writes reference and commit only the writer's own asset
  entries and legacy shared ones (owner-bound document store and the
  collector backfill), so another owner's course naming a pending id
  can no longer commit, expose or pin it.
- The foreign-read rule additionally requires the referencing live
  course to belong to the entry's owner, as defense in depth for
  reference rows that did not come from a scoped write.
- The tombstone refusal for new runtime sessions moved from a raw path
  comparison in the route into the guarded store's createSession
  (RuntimeStageNotFoundError, 404 STAGE_NOT_FOUND), so no encoded
  spelling of the path routes around it.
- Replacing or deleting a legacy shared entry checks reference
  ownership and mutates in one transaction under the entry row lock,
  which a concurrent reference insert must wait for.
- Tests build the uncommitted-but-referenced, tombstoned-with-rows and
  foreign-reference states directly so each half of the read rule is
  observable, and a PostgreSQL suite holds a competing reference open
  across a legacy delete.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): upgrade note for dev-token-gated deployments, owner-scoped references

Tell operators whose only access gate was the development token to put
the deployment behind an access code or gateway, register an owner
authenticator, or turn server persistence off before upgrading, and
describe that a course references only its owner's media.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): per-owner reference principals and a typed taken-id conflict

- PgDocumentStore's assetReferencePrincipals is now a function of the
  store's bound owner, evaluated per write, instead of a fixed list. A
  store re-bound with forOwner therefore scopes to the new owner rather
  than carrying the previous owner's principals onto its writes. It has
  the same shape as AssetCollector's backfill option, so one function
  serves both.
- PgRuntimeStore and BrowserRuntimeStore raise RuntimeSessionExistsError
  for a taken session id, and the runtime handler classifies it as
  409 SESSION_ALREADY_EXISTS even when its re-read cannot see the holder
  (for example a session a host-side wrapper hides). Recognized by class
  or by code.

Still the unreleased 0.32.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): guard the Pi whiteboard runtime and answer hidden collisions with 409

- Every server-side runtime writer now uses the tombstone-guarded store:
  the Pi native whiteboard service gets it through the same helper as
  /api/persistence/runtime/*, so a deleted course takes no new
  whiteboard runtime either.
- Creating a session whose id is held by a session a tombstone hides
  answers the uniform 409 SESSION_ALREADY_EXISTS instead of a 500,
  covered on PGlite and on real PostgreSQL.
- The owner-bound document store passes its reference-principal
  function, so the scoping follows the bound owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 3 — accounts through an identity gateway (trustedProxyHeader)

* feat(auth): owner identity seam, phase 3 — trusted-proxy-header authenticator

Add the trustedProxyHeader built-in: real accounts through an identity
gateway (oauth2-proxy, Authelia, a Keycloak-based proxy, an institutional
reverse proxy) that signs the user in and forwards the verified user in a
request header. OpenMAIC never handles passwords or tokens.

- Selected with OWNER_AUTHENTICATOR=trusted-proxy and validated in
  instrumentation register(). TRUSTED_PROXY_SECRET (32-1024 printable
  ASCII characters) is required; header names default to
  x-openmaic-proxy-secret and x-forwarded-user, and an optional groups
  header with TRUSTED_PROXY_ADMIN_GROUPS grants admin. An unknown
  selector, a TRUSTED_PROXY_* variable without the selector, a shared
  owner id or a host-registered authenticator alongside it, and malformed,
  reserved or repeated header names fail the boot.
- Trust boundary: Next.js exposes no TCP peer address to route handlers,
  middleware or Server Actions, and fills x-forwarded-for from the socket
  only when the client sent none, so no address allowlist is offered. The
  gateway secret is compared in constant time (hashed, timingSafeEqual).
- Per request: a missing or wrong secret, a missing or blank user, a user
  header with a comma (duplicate lines arrive comma-joined), or a user
  whose proxy:<user> id fails the owner id guard is 401
  INVALID_CREDENTIAL, never an anonymous owner. The guard failure is a
  401 rather than a 500 because the value comes with the request. The
  user is trimmed and keeps its case. Every gateway user holds
  course:publish; kind user, assurance verified, channel proxy; no
  cookie. Route handlers and Server Actions share one code path.
- The unset-ACCESS_CODE boot warning is skipped in this mode; the gateway
  is the access gate and ACCESS_CODE stays independent.
- The identity boundary scan now also fails on gateway identity headers
  (x-forwarded-user/-groups/-email, x-auth-request-*, remote-user, the
  secret header) and TRUSTED_PROXY_* read outside lib/server/identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): document accounts through an identity gateway

Describe the trustedProxyHeader built-in in the README owner identity
section (and its Chinese counterpart): configuration, request semantics,
the trust boundary and why it is a shared secret, the interaction with
ACCESS_CODE, boot-time validation, and an oauth2-proxy example that
injects the user, groups and secret headers. Add the .env.example block,
a SECURITY.md note and an Unreleased changelog entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): harden the trusted-proxy authenticator after review

- Reserve the header names HTTP, Next.js and forwarding proxies set
  themselves: configuring the user, groups or secret header to rsc,
  next-action, forwarded, x-forwarded-for/-host/-proto/-port, x-real-ip,
  or anything starting with x-middleware-, x-invoke-, x-nextjs- or
  next-router- now fails the boot.
- Log a one-time boot warning when TRUSTED_PROXY_ADMIN_GROUPS is set: the
  admin role is granted from the groups header, which is trusted on the
  strength of the shared secret alone, so the gateway must overwrite or
  strip it.
- POST /api/chat/pi resolves the owner up front and answers 401 when the
  authenticator rejects the credential, like every other owner-resolving
  route, instead of continuing without an owner. The anonymous default
  always resolves, so it is unaffected.
- The identity boundary scan also covers the Authentik, Azure App Service
  authentication, WebAuth, Cloudflare Access, Google IAP and AWS ALB OIDC
  header families.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): admin-groups warning and a corrected oauth2-proxy example

- Warn next to TRUSTED_PROXY_ADMIN_GROUPS (README, README-zh,
  .env.example) that the groups header is trusted on the strength of the
  shared secret alone, and list the reserved header names.
- The oauth2-proxy example now states it targets v7.14+ (nested
  claimSource/secretSource), uses the canonical bindAddress, sets
  preserveRequestValue: false on all three headers and
  insecureSkipNonce: false, forwards the subject with `claim: user`, and
  notes that the secret variable holds the literal secret. It adds the
  legacy-option migration caveat, a warning not to exempt app routes via
  skip-auth or trusted-ip options, and notes that the IdP must release the
  groups claim and group names must not contain commas.
- Changelog: the new boot checks and the /api/chat/pi 401.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 4 — host extension hooks

* feat(persistence): owner identity seam, phase 4 — host extension hooks

Let a host add product behavior at four points without forking a route,
registered once from instrumentation.ts register() like the owner
authenticator and sealed on first use. Nothing changes when no hook is
registered.

- configurePersistenceHooks({ name, authorizeCreate, onCreate, library,
  beforeAssetAllocate }):
  - authorizeCreate / onCreate run once per created course, inside the
    create transaction, in the transaction that inserts the course's
    ownership row. Saves of a course the owner already holds, and a create
    that loses a race to a concurrent create of the same id, are updates and
    run no create hook. A refusal answers 403 CREATE_REFUSED; a throw rolls
    the course, its ownership row and the host's writes back together. The
    hooks receive the request principal, or only the owner id for a
    background agent run.
  - library chooses which stage ids GET /api/stages lists; the route builds
    the usual summaries in the provider's order and drops every id the read
    path would refuse (deleted or unclaimed). folderId is shown only on the
    principal's own courses.
  - beforeAssetAllocate runs inside the storage handler's asset
    authorization step for POST /assets, after the owner is resolved and
    before the body is read; a returned Response is sent instead, with
    nothing stored and no quota counted.
- configureAssetByteStore({ name, create, signsReadUrls }) replaces the
  ASSET_S3_BUCKET switch for both the persistence provider and the asset
  collector. A host store must keep bytes outside the registry database.
  Under ASSET_BYTE_EGRESS=redirect a store that does not declare
  signsReadUrls stops the server at boot; the built-in layers keep their
  behavior.
- @openmaic/storage 0.33.0: DocumentWriteRefusedError lets a store refuse a
  write as policy; the document HTTP handler answers 403 with its code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): harden the host extension hooks after review

Correctness:
- Registration reads every hook once and stores it bound to the object the
  host passed, so hooks defined on a class prototype or as non-enumerable
  properties are kept. A spread copy dropped them silently, including the
  authorizeCreate and beforeAssetAllocate gates. Same for
  configureAssetByteStore. The unknown-key check now applies to plain
  objects only, since class instances carry their own state fields.
- The owner-bound document store carries each call's operation in its own
  async context (AsyncLocalStorage) instead of a field on the instance. One
  store is shared by every tool of an agent run; a concurrent call could
  overwrite the field before the transaction read it, so a create was gated
  as a read (no ownership row, no create hooks) or claimed another stage.
  create_stage also runs sequentially within a tool batch.

Hardening:
- isDocumentWriteRefusedError recognizes a refusal from another copy of the
  storage package by name and shape (not by code alone). The document
  handler and POST /api/stages use it.
- beforeAssetAllocate also gates PUT /assets/{id}/content. The request
  carries operation ('create' | 'replace') and the decoded assetId from the
  handler's own routing.
- Agent runs: a refusal on a background write carries a fixed message, and
  create_stage answers it with a stable tool result. The host's message never
  reaches the model.
- A byte store declaring signsReadUrls without a signer degrades to direct
  bytes with a warning instead of failing reads, and the collector no longer
  requires signing.
- DocumentActor is discriminated by source ('request' with the principal,
  'background' without one).
- Library ids the read path cannot address are dropped, and a provider
  answer is capped at 5000 ids.
- The boot-time ownership backfill logs how many courses it adopted without
  hooks.
- The server persistence provider no longer exposes an unscoped document
  store.
- The writesOutsideRegistryDatabase trust boundary and the scope of upload
  admission (HTTP uploads, not server-generated media) are documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): refuse misspelled hooks on host classes too

A class instance skipped the unknown-key check, so a misspelled hook such
as authorizeCreat registered cleanly and never ran, leaving creation
ungated. Both configurePersistenceHooks and configureAssetByteStore now
reject any function-valued property of a class instance (own or on its
prototype chain up to Object.prototype, constructor excluded) that is not a
known key, and any unknown key of a plain object. The error names the key
and suggests the known key within edit distance 2. Host classes keep their
helpers private (#helper) or register a plain object; the README says so.

Also:
- isDocumentWriteRefusedError requires a string message from a cross-copy
  candidate, and documents that the name is the effective discriminator.
- CHANGELOG: ServerPersistenceProvider no longer exposes an unscoped
  documentStore; use the owner-bound store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

Course ownership was recorded twice: in stage_meta, which already answers
every access decision, and in document_stages.owner_id, which the storage
package scoped by. Two tenancy records double what a later claim has to
re-key, and a host that keeps ownership in its own table had to rewrite the
package's SQL to strip the column.

@openmaic/storage 0.34.0:
- PgDocumentStore never reads or writes document_stages.owner_id. An
  owner-bound store scopes listings, writes, deletes, the freshness manifest
  and folder membership through the host's ownership relation, named with
  the new documentOwnership option ({ table, stageIdColumn, ownerIdColumn,
  tombstoneColumn, claimOnCreate }, or false for no document scoping).
  Identifiers are validated, never quoted into SQL from arbitrary input.
- forOwner / ownerId without documentOwnership throws at construction,
  so a direct consumer cannot upgrade into a store that silently scopes
  nothing. An unbound store is tenant-agnostic.
- claimOnCreate inserts the ownership row in the create transaction; a
  concurrent create of the same id by another owner rolls back.
- folders: false lets a host whose document_stages has no folder_id use
  the store; folder ids are unique per owner, so membership goes through
  the ownership relation too.
- AssetCollector reads each document's owner from documentOwnership for
  assetReferencePrincipals, and refuses the function without it.
- Schema: fresh installs get no owner_id column and none of its indexes;
  an existing column is kept for one release, made nullable with its
  default dropped (each ALTER guarded by a catalog check, so a replay takes
  no lock), and its two indexes are dropped. document_stages_folder_idx
  replaces the owner/folder index.

Application:
- The owner-bound store binds the package to stage_meta
  (STAGE_META_OWNERSHIP) and lists through it in one query; the asset
  collector schedule passes the same relation.
- The boot backfill copies column-only owners into stage_meta before any
  store is built, only when the legacy column exists, idempotently, and
  logs how many it adopted and how many disagree (stage_meta stands).

Tests cover a fresh install and an upgrade from the previous schema on
PGlite and PostgreSQL 16, a host table with neither ownership nor folder
columns, folder isolation across owners sharing a folder id, and the
concurrent-create race.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): serialize schema bootstrap; guard unscoped owner-bound stores

Review follow-ups for the tenancy change.

- Concurrent starts: two instances starting against one database raced on
  the schema DDL (a duplicate pg_class / pg_type entry, or "tuple
  concurrently updated"), failing one instance's first request. Every
  schema bootstrap the application runs -- the persistence provider's
  sequence (runtime, documents, stage_meta and its ownership backfill,
  owner materials, assets), the agent-session, session-material and
  user-skill stores, and the asset collector's -- now runs under one
  PostgreSQL advisory lock held on a dedicated connection and released in
  finally. It lives in the application because what needs serializing is
  the whole sequence, including tables the package does not own. A
  PostgreSQL test boots eight instances at once, four rounds each, on a
  fresh and on an upgraded database; without the lock it fails every run.
- documentOwnership: false on an owner-bound PgDocumentStore now also
  requires allowCrossOwnerDocumentAccess: true, so binding an owner for
  folders or asset principals cannot silently expose every owner's
  documents.
- Documented that an ownership relation must cascade with the document
  rows (a leftover row keeps the id reserved for its owner, which is also
  how a host keeps retired ids from being reused), with tests for both.
  Leftover rows are deliberately not made claimable: a host cannot be
  told apart from one that keeps them as retirement markers.
- Upgrade notes: the one-time table locks on the first start, and the
  out-of-band CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY path
  that leaves the start with nothing to build or drop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): leave the schema bootstrap lock out of the held-lock check

The create-hooks suite asserts that its host lock is released with the
create transaction by counting granted advisory locks database-wide. A suite
booting a provider against the same database at that moment now holds the
schema bootstrap lock, which is not the lock under test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

* feat(storage): owner merge primitives for claiming anonymous work (0.35.0)

Add the per-store pieces a host needs to move one owner's data to another
inside a single transaction of its own:

- reassignDocumentFolders(tx, { fromOwnerId, toOwnerId, documentOwnership })
  moves folders and filing: a same-named folder merges into the target's, a
  colliding id is renumbered, and filing is rewritten in one statement from
  the old ids.
- PgAssetStore.reassignPrincipal(fromKey, toKey) moves a principal's entries
  under both principals' write locks, in key order; quota is not re-checked.
- PgUserSkillStore.mergeOwner(from, to) moves skills, renaming a live handle
  the target already uses with the first free numeric suffix.
- PgUserSkillStore takes a resolveFinalOwner hook, run as the create
  transaction's first statement, like the agent-session store's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

A visitor who works anonymously and then signs in can move that work to their
account in one transaction, and the anonymous id is retired afterwards.

- Principal: pendingClaim { fromOwnerId, assurance }. The trusted-proxy
  built-in sets it when a gateway request also carries a valid anonymous
  cookie; the anonymous and shared-team built-ins never do. Authenticators may
  implement describeStoredOwner, canonicalize and clearPendingClaim;
  principalFromStoredOwner describes a stored id for work without a request.
- Claims: claimOwner / claimPendingOwner and registerClaimParticipant (sealed
  on first claim). One transaction, both owners' identity locks taken
  exclusively in key order, participants in a fixed order (folders, courses,
  materials, agent sessions, skills, runtime, assets), recorded in the new
  owner_merges table. Idempotent per pair; refuses a non-anonymous source, an
  anonymous target, a source already claimed elsewhere, and chains. Same-named
  folders merge, colliding folder ids are renumbered, colliding skill handles
  get a suffix, and quotas are not applied to what moves.
- Identity lock: every owner write (documents, folders, asset allocations,
  material registration, agent sessions, skills) takes the owner's advisory
  lock in shared mode first, so a write racing a claim is moved or refused,
  never orphaned or deadlocked.
- Retired ids: request writes are refused with 403 OWNER_RETIRED; agent runs
  and media generation that started before the claim follow the id to the
  account (forwarding document store, canonicalized stage probes, forwarded
  generated assets and skills).
- Trigger: POST /api/identity/claim (same-origin JSON only; clears the
  anonymous cookie), or OWNER_CLAIM_TRIGGER=auto on the first request carrying
  a pending claim. The runtime contract's learner merge now performs the same
  claim, allowed only for the request's own pending claim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): count only the suite's own advisory locks after concurrent creates

The check that a host's transaction-scoped lock is released counted every
granted advisory lock in the database. Suites that run against the same
database at the same time hold advisory locks of their own (a schema
bootstrap, the owner identity locks a claim takes), which made it fail
depending on timing. Count only locks held by this suite's connections.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): use a neutral example table for host claim participants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): owner merge fixes and store policy errors over HTTP (review round 1)

- PgUserSkillStore.mergeOwner reserves every handle either owner holds before
  renaming: a claimed skill whose handle the target uses could be renamed into
  a handle another claimed skill still held, which violated the per-owner
  unique index and made that merge fail every time.
- PgUserSkillStore.list orders ties by id.
- PgRuntimeStore takes resolveFinalLearner(tx, learnerKey), run as the
  create transaction's first statement, so a host can fence session creates
  against its owner merges; reassignLearner(from, to) re-keys sessions without
  re-validating them, so one session a newer version wrote cannot block a
  merge.
- Every HTTP handler (runtime, documents, assets) answers a
  DocumentWriteRefusedError as 403 with its code, and the new StorageBusyError
  (code, message, retryAfterSeconds) as 503 with Retry-After.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 1 — fenced runtime creates, bounded locks, retired-cookie recovery

- Runtime session creates take the owner's identity lock as the create
  transaction's first statement and refuse a retired owner, closing the window
  in which a create racing a claim was left under the retired id.
- Lock waits are bounded: a claim waits OWNER_CLAIM_LOCK_WAIT_MS (default 5 s)
  for its identity locks, a write OWNER_WRITE_LOCK_WAIT_MS (default 30 s);
  running out, or losing a deadlock or serialization race, answers 503
  OWNER_BUSY with Retry-After on the claim routes and on every fenced write.
- Every 403 OWNER_RETIRED carries the Set-Cookie values that drop the retired
  anonymous credential (the anonymous built-in now implements
  clearPendingClaim), so a browser whose claim response was lost recovers.
  Writes by id to rows that moved (skill delete, session messages) answer
  OWNER_RETIRED instead of not-found or a bare 403.
- OWNER_CLAIM_TRIGGER=auto no longer claims ahead of the explicit claim
  routes, which now report their own claim instead of a false refusal.
- The owner-events stream tells a retired owner only that it moved, never the
  claiming account's id.
- Creates of one course id take turns (a per-id advisory lock before the
  ownership probe), so a concurrent second create saves as an update instead
  of failing as reserved-document.
- A claim's source must be described as anonymous by the authenticator; the
  write fences skip the retirement lookup for every other owner. The host
  canonicalize hook is removed: forwarding is core's owner_merges alone.
- The folders participant locks the source's stage_meta rows first; runtime
  sessions are re-keyed without re-validation; the owner asset store fences
  every allocation and refuses to allocate without transactions; background
  skill reads and patches follow a mid-run claim; the forwarding document
  store forwards methods only.
- Docs: when a pending claim can arise with the shipped authenticators, the
  anonymous cookie as a bearer credential, bounded waits, the exact scope of
  OWNER_RETIRED, and claims narrowed to what the tests show.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 2 — merge-record guard, boot-checked lock waits, stream credential drop

- owner_merges records claims of anonymous owners only, since the write
  fences enforce retirement only for ids the authenticator describes as
  anonymous. Reading a row that retires any other owner now fails loudly
  instead of being followed but not enforced. The comment that pointed hosts
  at claimOwner for their own account merges is corrected: a host merging two
  signed-in accounts moves the rows itself and refuses the merged-away account
  in its authenticator. describeStoredOwner must classify ids stably.
- OWNER_WRITE_LOCK_WAIT_MS and OWNER_CLAIM_LOCK_WAIT_MS are validated at boot
  with the same parser the lock paths use, so a malformed value stops the
  server instead of failing every write.
- The owner-events stream answers a connect by an already retired identity
  with one owner_moved event and the Set-Cookie values that drop the retired
  credential, so a stale tab's reconnect loop ends after one round trip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* chore(auth): final polish before landing the owner identity seam

* chore(auth): final polish before landing the owner identity seam

- Pi routes refuse a rejected credential with the seam's INVALID_CREDENTIAL body.
- Drop the unused LegacyAssetMutations alias and resolveLockWaitMs re-export.
- Scope the document_stages.owner_id deprecation wording: the store no longer
  reads or writes it; the boot backfill reads it and a claim mirrors ownership
  into it for rollback.
- Document OWNER_CLAIM_TRIGGER and the two lock-wait variables in .env.example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(storage): 0.35.1 for the README wording fix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(identity): one identity entry point with composable auth methods

* feat(identity): compose owner auth methods behind one entry point

Owner resolution now asks an ordered list of owner auth methods instead
of one configured authenticator. Each method answers `authenticated`,
`not-applicable` (no credential of its kind) or `invalid` (its credential
is present but bad). Core keeps the first `authenticated` answer, refuses
any `invalid` answer with 401 INVALID_CREDENTIAL without asking later
methods, and when no method applies falls back to the anonymous cookie
owner, or answers 401 when the host turned the fallback off. Route
handlers and Server Actions share the same order and per-request
memoization.

- configureOwnerAuthentication({ methods, anonymousFallback? }) replaces
  configureOwnerAuthenticator: single-shot, sealed on first use, and
  validated at registration and at boot.
- anonymousCookie becomes the built-in fallback method; sharedTeam
  becomes a method, selected by PERSISTENCE_SHARED_OWNER_ID as before
  when nothing is registered, and kept beside host methods only when the
  host lists sharedTeamAuthMethod() last. The variable set beside a
  registration that leaves it out fails the boot.
- Core attaches pendingClaim itself: a host method's non-anonymous
  principal on a request that also carries a valid anonymous cookie. A
  method may not set one. Claim cookie clearing and OWNER_RETIRED
  recovery go through the anonymous method's clearCredential.
- principalFromStoredOwner consults the anonymous method first, then the
  configured methods' describeStoredOwner in order.
- Remove the built-in trusted-proxy header authenticator with
  OWNER_AUTHENTICATOR and every TRUSTED_PROXY_* variable. The boundary
  test now also keeps the resolution internals inside
  lib/server/identity/ and asserts no source reads the removed variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): document owner auth methods and a gateway JWT recipe

Rewrite the owner identity section around the ordered auth methods: the
three answers and what resolution does with each, the anonymous
fallback, registration and its boot checks, and the sharedTeam rules.
Replace the removed built-in gateway authenticator's documentation with
a short host-owned recipe that verifies a gateway-forwarded signed JWT
(oauth2-proxy, Cloudflare Access, IAP) against the identity provider's
keys with the jose library. Update the claim section for the core-owned
claim candidate, and the Chinese README, SECURITY.md, CHANGELOG and the
deployment docs in every locale to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(identity): tighten retired-owner cookies, retired config and the boundary

- On 403 OWNER_RETIRED, core now calls clearCredential only on the
  anonymous fallback and on methods that declare
  `issuesAnonymousOwners: true`, so an account method's session cookie is
  never cleared there. A clearCredential that throws is logged and
  skipped instead of turning the refusal into a 500.
- Setting any retired OWNER_AUTHENTICATOR / TRUSTED_PROXY_* variable fails
  the boot with a pointer to the auth-methods docs, instead of silently
  serving every request anonymously. Only the names are matched, in one
  module the boundary test checks never reads a value.
- Add lib/server/identity/host/ for host auth methods. The boundary test
  now scans core identity files for gateway identity headers too, polices
  reads of an incoming Authorization header everywhere but the host
  directory, and keeps resolution internals out of host code.
- Document on OwnerAuthMethodResult that a present but unusable
  credential must answer `invalid`, never `not-applicable`, and that a
  method that cannot decide throws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): fix the gateway JWT recipe's error split and add notes

- The recipe now answers `invalid` only for token failures (expired,
  claim validation, malformed JWS/JWT, bad signature, no or several
  matching keys, disallowed or unsupported algorithm) and rethrows
  everything else: jose reports a key endpoint that is down, answers
  non-200 or returns unparsable data as a generic JOSEError, which the
  previous version turned into a 401 for every user.
- The recipe file goes in lib/server/identity/host/, with an accurate
  description of what the boundary test enforces.
- Document invalid versus not-applicable for malformed credentials,
  issuesAnonymousOwners, the boot failure on retired variables, and that
  the claim candidate is an unsigned bearer cookie a sibling subdomain can
  shadow (serve on a registrable domain of its own; cookie hardening is
  deferred). Mirror the notes in the Chinese README, SECURITY.md and the
  changelog.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 11:54:26 +08:00
Frank ZhuandFrank-zhu0404 56322a5e06 fix(generation): reject unusable interactive scripts and surface runtime errors (#1649)
Reject classic inline interactive scripts that fail to parse at generation time (extracted with parse5, checked with node:vm Script without executing), and surface iframe runtime errors on the active interactive scene.

Addresses #1622 (partial: no recovery / regeneration UI).

Co-authored-by: Frank-zhu0404 <Frank-zhu0404@users.noreply.github.com>
2026-09-23 13:50:18 +08:00
杨慎andwyuc d04a75cbc6 feat(ai): add Xiaomi MiMo V2.6 models [AI-assisted] (#1655)
* feat(ai): add Xiaomi MiMo V2.6 models

* docs(ai): preserve Xiaomi provider references

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 13:11:49 +08:00
Frank Zhuandwyuc fbc51fcbd6 fix(generation): reject invalid quiz option shapes and constrain values at the source (Closes #1375) (#1651)
* fix(generation): enforce quiz option value/label contract

When a model puts choice content in value and a bare A-Z letter in label,
swap those fields so persisted quiz options keep a letter as the selection
identity and the content as the label. Correct objects, plain strings, and
answer-key normalization stay as they were.

Closes #1375

* fix(generation): rewrite swapped quiz letter keys to option values

When a bare label is moved onto the option value, a raw answer token that
exactly equals that original label is rewritten to the uppercased letter
before exact alignment. The stored key is then the letter QuizView submits,
so that submission grades correct. A stored "a" on an already-correct
value "A" is left unresolved.

Closes #1375

* fix(generation): reject invalid quiz option shapes instead of swapping

Stop rewriting quiz options when a model puts content in value and a
letter in label. The quiz prompt now requires value to be one ASCII
letter A-Z and label to be the option content. A choice question that
still breaks that contract, or whose answer does not name an option
value, is invalid model output and regenerates.

Closes #1375

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 12:46:40 +08:00
Yizuki_Ameandwyuc 283a3096aa fix(editor): share line bounds for alignment and dragging (#1631)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 15:39:56 +08:00
LING DUANandwyuc 7be21f2d89 fix(chat): pin Pi routing and preserve provider errors (#1637)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 15:18:54 +08:00
AbelWangYaBoandwyuc 7e8a065ecb fix(importer): bound pptx zip inflation by default (#1588)
ZipParseLimits was optional field by field and both import paths called
parseZip with no argument, so every bound the parser already implemented was
inert: a .pptx could inflate without limit.

- apply the defaults per field, so a caller overriding one bound no longer
  silently drops the others, and Number.POSITIVE_INFINITY switches a single
  bound off
- add maxCompressionRatio for text parts, defaulting to 200:1. It applies to
  slides, layouts, themes and chart XML only. Uncompressed bitmaps and silent
  PCM are ordinary deck content sitting at the top of DEFLATE's range — a
  solid-colour 24-bit BMP measures ~1027:1, 30s of silent stereo PCM ~1016:1,
  and a 1920x1080 white screenshot with a grid ~297:1 — so a fixed ratio there
  rejects real decks. Media and embedded objects are bounded by
  maxEntryUncompressedBytes and maxMediaBytes instead, which for an honest
  archive are checked before anything is inflated. Text parts have no such
  problem: real decks peak around 19:1, so 200 leaves ample headroom.
- reject a supplied NaN and a non-integer maxEntries rather than letting them
  compare false against every bound and disable it silently
- export the defaults and the limit type so callers can extend them
- add the first tests for this parser: the regression itself (a call with no
  limits must still apply the bounds), the per-field defaulting invariant, each
  bound pinned, the binary-part exemption pinned with a solid-colour BMP, a
  silent WAV and an embedded object, and the repo's own regression deck

The bounds read the sizes the archive declares about itself, which an archive
is free to understate. One that declares small and inflates large passes all of
them and is stopped only by JSZip's own size check, after that entry has been
inflated — peak RSS in the reviewed case was ~1.19 GB, and with maxConcurrency 8
several such entries inflate in parallel. Bounding the allocation itself needs a
byte-budgeted inflate inside the read path, which is a larger change than
activating the bounds that already existed here.

Version to 0.3.0: ^0.2.6 admits 0.2.7, and this changes default behaviour for
existing callers. Consumers pinned to ^0.2.x need to move deliberately.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 14:24:58 +08:00
LING DUAN 2a77a8a476 feat(chat): make Pi classroom runtime the default (#1628)
* feat(chat): make Pi classroom runtime the default

* chore(chat): align docs and E2E with Pi default
2026-09-22 11:41:13 +08:00
Yizuki_Ameandwyuc df16d7e322 fix(export): include active line geometry in bounds (#1626)
Share corrected line bounds across renderer and React editing paths, preserve double-elbow routing, and translate PPTX points into the shape bounding box.

Refs #674 and the prior implementation/review in #675. AI-assisted.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-21 22:22:35 +08:00
d7b31aa5e7 fix(generation): add browser-safe package entry (#1609)
Co-authored-by: SY <sy@SYdeMacBook-Pro.local>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 16:10:52 +08:00
8de7ec2870 fix(importer): include all color attributes in style cache key (#702) (#1571)
Co-authored-by: cham <2577781125@qq.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-19 22:21:30 +08:00
4d2e2bab82 fix(generation): keep narration speech TTS-readable — no formulas or LaTeX (#1586)
* fix(generation): keep narration speech TTS-readable — no formulas or LaTeX

Narration speech authored by the four *-actions prompts could carry raw
formula notation (a^2 x / y) or LaTeX (\frac{a}{b}), which TTS engines
read out as gibberish. The prompts had no TTS readability rules, and the
element list fed to them includes raw LaTeX via `Formula:` entries,
inviting verbatim copying into speech.

Add a shared `speech-tts-readability` snippet and compose it into the
speech sections of slide-actions, quiz-actions, interactive-actions, and
pbl-actions via the existing {{snippet:...}} mechanism. The name
references speech/TTS so it cannot be confused with slide-content's
visual LaTeX rules or reused there.

Refs #1585

Co-Authored-By: Claude Code <noreply@anthropic.com>

* chore(generation): bump to 0.3.10, main already released 0.3.9

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DYifP8wM4XJQ3Hc2qsF6zf

---------

Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-19 18:42:53 +08:00
LING DUANandClaude Opus 5 f29bbc4daa feat: sample declared interactive state before classroom questions (#1508)
* feat: sample declared interactive state before classroom questions

* test: exercise generated publication example through the iframe reader

* chore(generation): bump package version for observation prompt contract

* fix: keep interactive state optional on insecure HTTP origins

* fix(playback): keep component picking and separate state from reference

A declared state interface replaced the component picker with a forced
whole-area `#experiment` reference, removing the per-component selection
and outline that `main` already ships. Sampling was also gated on the
reference selector, so the only way to obtain state was to give up the
selection.

Reference identity and area state are now independent request-scoped
evidence items:

- `handleToggleElementPick` always arms the picker again, so a scene that
  declares the interface keeps main's per-component selection, outline,
  and send-time clearing.
- `sampleInteractiveState` follows the current Scene instead of the draft
  reference, so an unreferenced follow-up still reports current facts and
  never re-creates or extends a reference.
- The Host carries area state with or without a component reference. The
  evidence header names both identities and refuses to present area facts
  as properties of the referenced component.
- `metadata` is absent when only area state travels, so no element
  identity and no Spotlight authorization can be derived from it, and the
  accepted-reference receipt stays driven by explicit references only.

Review follow-ups in the same change:

- Client sampling follows `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED`.
  An ungated packet turned an ordinary Pi question into a 400 while the
  reference feature was disabled.
- A Scene that declares the interface always receives an availability
  boundary, including when the browser produced no packet at all. It is
  reported as `not-sampled` rather than the previous `no-interface`,
  which was a false statement about an activity that does declare one.
  Courseware without the interface keeps its unreferenced behaviour.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(generation): register the observation snippet as a packaged asset

The interactive-observation snippet is referenced by all six widget
content templates but was never added to the packaged-asset manifest, so
the asset test and the golden scene prompt both failed.

- `SNIPPET_IDS` now lists `interactive-observation`, restoring both the
  "exactly the generation-owned templates and referenced snippets" check
  and the "every referenced snippet is packaged" cross-check.
- The interactive system-prompt snapshot is re-pinned. The change is
  purely additive: the snippet is appended to the simulation template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(pi): stop injecting state constraints while the feature is disabled

The route rejects a request that carries a reference or a state packet
while `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED` is off, but an ordinary
question carries neither. It still reached the Host, and a Scene that
declares the state interface then received the full page-state block —
several kilobytes of constraints about evidence the deployment can never
sample.

The Host now returns before building that note when the feature is off.
A route-level regression asserts that neither the Director prompt nor the
Child prompt gains `PAGE-REPORTED STATE` in that configuration; disabling
the guard makes it fail with exactly that symptom.

Also reopens the composer before the unreferenced follow-up in the
classroom browser spec. An accepted answer may close it, which made the
assertion flaky without changing the behaviour under test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(pi): decouple reference Scene from state freshness and bound assembled evidence

Cross-review found two defects in the request-scoped state evidence.

The reference's Scene was folded into the sample's staleness test. The packet
is already bound to the current Scene by the identity check above it, so a
valid current-Scene sample was being discarded as `stale-sample` purely because
the student's component reference came from an earlier Scene. Reference and
area state are independent evidence items; freshness is a property of the
sample alone. With the coupling gone the two can now disagree on Scene, so the
note says so explicitly rather than letting the model attribute area facts to a
component that may not be on the current Scene.

The assembled evidence had no stated output budget. The static component packet
is bounded to 24,000 code points upstream, but that bound covers the static
packet alone; the note and the escaped observation JSON were appended without a
recheck. Escaping `<` for the prompt expands one code point into six, and `<` is
legal in a label or a fact value, so a packet the Host accepts could assemble to
149,385 code points. The budget is now declared as the static bound plus the room
the note frame needs, which is what makes the degradation terminate. Over budget,
the state body drops whole to an explicit `unavailable` statement: truncating the
JSON would emit a broken packet, and thinning a `complete` relation set would turn
an exhaustive set into a false one. The relationship summary degrades with it, so
the prose never asserts COMPLETE over a body that is gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(pi): keep slide references independent of activity state

* refactor(interactive): simplify declared state and unify iframe preparation

Accept any JSON report within byte and depth budgets, without generated field
requirements or relationship-completeness semantics. Keep publishState and an
optional rendered result in the generation guidance.

Prepare the observation responder through patchHtmlForIframe and let the pool
own document identity, preserving state across placeholder remounts. Settle
sampling failures locally and align browser/server nesting limits.

Cover permissive JSON delivery, resource limits, lifecycle, legacy behavior,
and real renderer remounts with focused regression tests.

* fix(generation): publish automatic activity changes with clear positions

* fix(interactive): report missing legacy scope as no interface

* fix(interactive): guard sampling capabilities and bind scopes lazily

* chore(generation): bump version after main integration

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-19 18:17:30 +08:00
xuyuanwei678andwyuc f70dc67a6a fix(importer, editor): preserve PPTX diagonal corners, table typography, editing alignment and punctuation wrapping (#1581)
* fix: preserve imported PPTX table typography and rounded geometry

* fix(importer): isolate table punctuation width from tab clamping

* fix(importer): preserve punctuation markers with default cell margins

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-18 16:28:55 +08:00
f6e350a1dd fix(storage): sanitize agent session descriptive text (#1504)
Co-authored-by: Bryan Nathan <bryan@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-18 12:04:27 +08:00
xuyuanwei678 d695193ae6 fix(importer): preserve shape autofit sizes and Wingdings checkmarks (#1577) 2026-09-18 11:51:41 +08:00
Ethan Zhangandwyuc 8517121ec0 fix(docker): create /app/data with runtime-user ownership so classroom persistence works (#1442)
* fix(docker): create /app/data owned by the runtime user before dropping privileges

* docs(deployment): note one-time ownership repair for pre-existing data volumes

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-17 23:57:33 +08:00
8c22e27a29 fix(generation): tolerate non-array mediaGenerations in generated outlines (#1469)
* fix(generation): tolerate non-array mediaGenerations in generated outlines

LLM-generated outline JSON can carry mediaGenerations as a string or
object instead of an array. Every downstream consumer (media
orchestrator, video manifest, scene generator) requires an array, and
uniquifyMediaElementIds called .map() after only a falsy check, so one
malformed outline aborted course generation with
"outline.mediaGenerations.map is not a function".

- sanitize at parse time: drop non-array mediaGenerations from each
  enriched outline (outline-generator)
- harden uniquifyMediaElementIds: treat non-array values as absent,
  strip them, and open the early-return guard on field presence so
  all-malformed outlines still get sanitized
- add unit tests covering array, string, object, number, and missing
  shapes

* chore(generation): bump package version to 0.3.8

Required by the package version bump check: the PR changes the
publishable @openmaic/generation package sources.

* test(generation): use correct SceneOutline fixture fields in outline media tests

The first commit of this branch accidentally staged an earlier draft of
the test file (sceneType instead of type, stale id regex). Re-stage the
final version that passes tsc and the full suite locally.

---------

Co-authored-by: PassCode023 <269712126+PassCode023@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-17 23:29:05 +08:00
xuyuanwei678 f7b8769e7f fix(importer, renderer): preserve PPTX text, chart, and image styling (#1534)
* fix: preserve PPTX text, chart, and image styling

* fix(importer): preserve hyperlink semantics and remove inferred alignment

* test(charts): replace untyped assertions to satisfy CI

* fix(importer): honor source-linked chart axis number formats

* fix(renderer): preserve picture bar clustering and clipping

* fix(pptx): address chart scales and soft-edge review findings

* fix(pptx): preserve percent-stack labels and inherited text sizes

* fix(pptx): synchronize agent chart schema and axis defaults
2026-09-17 21:41:12 +08:00
xuyuanwei678andwyuc 35a8be5956 fix(pptx): preserve tab columns, text insets, and arrow rendering (#1518)
* fix(pptx): preserve tab columns, text insets, and arrow rendering

* fix(pptx): address tab layout review and bump package versions

* fix(pptx): preserve editable tab columns and final font metrics

* fix(editor): apply list commands inside tab columns

* fix(editor): preserve table paragraph spacing while editing

* fix(importer): preserve saved leading in auto-fit text labels

* fix(importer): preserve ordinary symbol-font text and editable default tabs

* fix(importer): preserve default hyperlink underline

* fix(editor): preserve Latin baselines when entering text editing

* fix(importer): approximate verified clear material front-face color

* fix(importer): preserve filled flowchart connector shapes

* fix(editor): preserve inline formulas when editing imported text

* fix(editor): keep formula caret separators inline

* fix(test): narrow serialized shape before checking inverse path

* fix(importer): preserve equation system delimiters

* fix(importer): preserve compatibility tables and cell formulas

* fix(editor): handle formatting and list splits inside inline containers

* fix(editor): preserve script sizing around inline containers

* fix(editor): preserve inline typography across editing and copy

* fix(editor): preserve destination and nested typography contexts

* fix(pptx): preserve table tabs and editor clipboard typography

* fix(editor): preserve container font context when clearing formatting

* fix(importer): correct Wingdings 3 upper-right triangle mapping

* fix(pptx): preserve explicit text inset markers with legacy fallback

* fix(editor): guard formula serialization and document layout limits

* fix(editor): preserve inline font contexts through undo

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-17 16:46:48 +08:00
wyucandClaude Opus 4.8 ee79762ef9 fix(access-code): expire tokens server-side and throttle verification behind a trusted proxy (#1513)
When ACCESS_CODE is set, verification tokens (timestamp.HMAC) never expired
because neither verifier checked the timestamp, and POST /api/access-code/verify
had no attempt throttling.

- Enforce a 7-day token lifetime in a shared, Edge-safe module used by both the
  Node verifier and the middleware Web-Crypto verifier, and reject non-canonical
  signatures in both.
- Rate-limit verification only when the client identity is trusted
  (TRUST_PROXY_HEADERS=true): a per-client sliding window (10 failures / 60s)
  whose attempt is reserved atomically at check time, returning 429 with
  Retry-After. Without a trusted proxy the app cannot attribute requests to a
  client, so no shared throttle is applied — a long random ACCESS_CODE is the
  protection, and a warning is logged when it is short.
- Store bounded, copied identity keys so a large forwarding header cannot retain
  memory.
- Document the behavior in .env.example, README, and configuration docs.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-15 17:47:14 +08:00
4d0f88b4c9 fix: bump next to 16.3.3, patches GHSA-p293-qw3h-jr36 (#1503)
GHSA-p293-qw3h-jr36 (CVE-2026-75604, critical): unauthenticated RCE on Windows-hosted Next.js servers. Root was on 16.2.11 and packages/docs on 16.2.6 (affected: >=16.0 <16.3.3). Both lockfiles regenerated. eslint-config-next lint plugins left as-is (not the affected runtime package).

Co-authored-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-15 16:20:45 +08:00
wyuc e824117c9c feat(storage): give the server the asset entry lifecycle (#1007 amendment, part 1) (#1472)
@openmaic/storage 0.31.0: a document reference table maintained by the document store, pending -> committed allocations, an entry pass with a one-time bounded backfill and legacy mark, a standing bounded sweep for entries nothing references, live-entry quota, and single-sequence ascending entry locks for every reference-maintaining write.
2026-09-15 08:49:33 +02:00
xuyuanwei678 3edaa4995f fix(importer): resolve embedded videos and upload posters in 0.2.1 (#1507)
* fix(importer): resolve embedded media and upload video posters

* fix(importer): use unique filenames for concurrent poster uploads
2026-09-15 05:48:26 +02:00
wyuc ee7a7b64df fix(agent-runtime): allow one owner material to be bound to multiple sessions (#1500)
* fix(agent-runtime): allow one owner material to be bound to multiple sessions

`agent_session_materials.id` is a global primary key, but
`bindOwnerMaterialsToSession` de-duplicated on `(sessionId, id)` and then
inserted the owner-side material id as the row id. Binding the same owner
material to a second session therefore raised a duplicate-key error
(23505) that the route surfaced as HTTP 500, so an owner material could
only ever be used by one course.

Keep row ids globally unique (every extraction method keys on `id` alone)
and store the shared owner id as metadata instead: the binder now mints a
fresh session row id, records `owner_material_id`, and looks that column
up for idempotent rebinding; a partial unique index on
`(session_id, owner_material_id)` adjudicates concurrent rebinds and the
loser adopts the winner's row. The schema change is additive and
idempotent (`ADD COLUMN IF NOT EXISTS` + `CREATE UNIQUE INDEX IF NOT
EXISTS`), so existing databases upgrade in place with `NULL` for legacy
rows. No FK from `owner_material_id` to the owner library, to keep
owner-library deletion decoupled from session rows.

Tests: backend-neutral contract test (same owner material bound to two
sessions, both readable), PGlite in-place upgrade test, pinned-schema
update, and a host-level test covering the route's 202 path.

* fix(agent-runtime): reuse legacy owner-material bindings and clean up losing uploads on concurrent bind

Rows written by the previous binder use the owner upload id as the session
row id and leave `owner_material_id` NULL, so the new owner-id lookup missed
them and a rebind created a duplicate row and copied the bytes again. The
binder now falls back to the session row id and adopts it only when it is
unambiguously that owner upload: source kind, matching title, no source URL
or derivative, no text, and the deterministic legacy object key with a byte
length that matches the owner record. Adoption stamps `owner_material_id`
through a conditional `backfillOwnerMaterialId` update that leaves extraction
columns untouched; a lost race re-reads the fast path. A different session is
still a fresh row.

The concurrent-bind loser also left its uploaded object behind after
adopting the winner. It now removes that object when the winner references a
different key, best-effort so a failed cleanup cannot fail a bind that
succeeded.

Tests: host-level PGlite regressions for the legacy rebind (id, row count,
byte copies, extraction state, and backfill) plus a different-session bind,
and a controllable byte store that parks the loser after its upload so the
winner commits first, asserting one row and only the winner's object remain.
A storage unit test pins the backfill contract. Bumps @openmaic/storage to
0.30.1 for the new public store method.
2026-09-14 18:43:31 +02:00
wyuc cfae106b89 fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates (#1499)
* fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates

PostgreSQL jsonb rejects the `\u0000` and lone UTF-16 surrogate escape
sequences that JSON.stringify emits verbatim, failing the enclosing
statement with SQLSTATE 22P05/22P02. In the agent-session store the failed
tree-entry/event write is treated as critical, so the whole run aborts and
all work in that run is lost.

Add a shared encodeJson at @openmaic/storage/src/pg-json.ts that replaces
U+0000 and unpaired surrogates with U+FFFD while preserving valid surrogate
pairs (emoji), sanitizing in-memory values and object keys before
JSON.stringify. Route every jsonb parameter through it: the agent-session,
document, runtime, asset, and material backends, plus owner-material
registration in lib/persistence.

The document/runtime/asset backends keep their existing assertJsonValue
guard, which rejects these code points with a readable error before the
shared serializer runs; the agent-session and material paths had no guard
and were the live 22P05 exposure. Literal backslash-u text is untouched.

* chore(storage): bump @openmaic/storage to 0.29.2

* fix(storage): keep colliding sanitized keys and own __proto__ members when encoding jsonb

encodeJson replaces NUL and lone surrogates in object keys, but two distinct
keys can sanitize to the same string: "a\u0000" and "a\uFFFD" both become
"a\uFFFD". The rebuild assigned each member by its sanitized key, so a later
member silently overwrote an earlier one and that earlier value was lost.

Keep the first member under the sanitized key and give each later colliding
member a deterministic suffix (#2, #3, ... appended until the key is unique
among the keys emitted so far, whether sanitized or original). Member order is
preserved. The suffix is ASCII and contains no NUL or surrogate, and the
sanitizer only ever emits U+FFFD, so a suffixed key can never be confused with
a bare replacement result.

The rebuild also used a plain {} object, so assigning an own "__proto__" member
invoked the prototype setter and dropped it from the output whenever a sibling
key needed sanitizing. Build the rebuilt object with a null prototype so
"__proto__" stays an own data property; JSON.stringify then emits it and
PostgreSQL round-trips it. Arrays and the no-op fast path (return the original
reference when nothing needs sanitizing) are unchanged.
2026-09-14 18:29:55 +02:00
LING DUANandwyuc ebf665f316 feat(playback): reference static GenUI components (#1281)
* feat(playback): reference static GenUI components

* feat(playback): gate courseware references

* fix(pi): bound Interactive source HTML before parsing

* fix(pi): bound interactive reference HTML processing

* fix(playback): close interactive reference review gaps

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-11 19:16:10 +02:00
85ef52d599 feat(media): write generated media through the asset pool under server-backed persistence (#1392)
* feat(media): store generated media in the asset pool when persistence is server-backed

With server-backed persistence the document is durable and shared, but
generated media stayed in the producing browser: the document kept its
gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio
id. Every new browser that opened such a course re-ran generation for every
slide, and it never converged, because the address of the generated bytes was
never written back into the document.

Under server-backed persistence only, the classic generation chain now stores
bytes in the asset pool first and writes the id the pool allocated into the
document.

- The client bootstrap configures the asset seam alongside the document and
  runtime seams: an HttpAssetStore over the persistence endpoint carrying the
  same credentials the document store carries, marked server-backed. The seam
  preflight now covers all three, so a failure still cannot half-configure
  persistence.
- Image, video and TTS generation commit in one fixed order: provider, pool,
  document, local cache, task. A reference reaches the document only after put
  returned an id, so a document can never name bytes that were not stored. A
  failure before the write-back leaves the placeholder with the provider called
  exactly once; the retry happens on the next owner load.
- The write-back is a per-slot rewrite through mutateDocument, which re-reads
  the current document under the per-stage lock, so it cannot clobber a newer
  scene. The open course is refreshed with the same rewrite without being
  marked dirty.
- "Has this already been generated?" is answered by the document (the slide
  exists and no longer holds the placeholder) instead of by this browser's task
  table.
- The classroom's resume effect fails closed on ownership: only a resolved
  owner starts generation, so a viewer opening a shared course spends nothing.
- The local media and audio tables become a per-tab cache. A failed cache write
  costs a re-download, never the media.

Browser-only mode is unchanged: every new call site sits behind the
server-backed gate, the local tables stay authoritative there, and placeholders
stay in the document.

Rendering and export needed no changes. HttpAssetStore.resolve mints an object
URL exactly as the browser store does, and the export byte resolver was already
pool-first.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere

The gate refused everything but a resolved owner, which read a sidecar that
answered "no ownership fact exists for this course" as a reason to block. That
is the answer a deployment without the sidecar's server-side prerequisites
gives for every course, and the answer a course with no ownership record gives:
in both, there is nobody the operator's budget needs protecting from, and
refusing strands the course's own author behind a question that can never be
answered.

Ownership is now four states over the sidecar's three outcomes. A definite
answer splits into owner and not-owner. An absent record is its own answer,
ownerless, and generation proceeds — the behaviour such a deployment had before
the gate existed. Only the absence of an answer, a transport failure or a load
that has not asked yet, stays unresolved and fails closed: "we could not ask"
must never be read as "nobody owns this". One mapper turns a sidecar result
into that state, and one predicate decides on it.

The workbench classroom pane runs the same resume effect and had no ownership
input at all, so a viewer opening a shared course there could still spend the
budget. It now asks the sidecar once per course, in parallel with its load and
feeding only the generation gate, so its read-only and edit behaviour is
unchanged. The shared progressive-load policy carries the gate for it, with
both new inputs required rather than defaulted so a future caller cannot omit
them into an open budget. Its stale comment claiming ownership could not be
expressed here is corrected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): make the write-back survive autosave, arrive before the scene does, and never leak

Independent reviews of the write-back found three ways a durable document could
still end up naming a placeholder, and two ways the gate that protects the
operator's budget could be walked around.

An autosave round captures the store synchronously and writes that capture, so
a round already in flight when a rewrite landed wrote the placeholder straight
back over the allocated id, and nothing marked the store dirty again to correct
it. The rewrite now marks the units it changed, which leaves a corrective flush
queued behind the stale one; re-saving a scene that already holds the id is
idempotent, losing the id is not.

Media is generated from outlines in parallel with scene content and usually
finishes first, so the slide that will carry the placeholder does not exist yet
and the write-back has nothing to rewrite. That was the ordinary path, not a
tail case, and its result was discarded: the task was marked done, the scene was
added afterwards with its placeholder intact, and a second pass in the same run
could call the provider again. The allocation is now held under the placeholder
— which also answers the skip test, so nothing pays twice — and applied when
that scene is committed, before its first save. One complete pass now leaves no
placeholder behind.

A failed commit used to abandon what it had already allocated. A poster upload
that failed threw away a stored video and sent the retry to submit the most
expensive job in the system again; a rejected write-back left registry rows that
name bytes nothing references, which the byte collector cannot reclaim because
it only collects blobs no row names. A poster failure now costs the poster, and
a write-back that reached nothing reclaims what it allocated. A partial write is
left alone, because the document already names it.

The ownership gate is fail-closed again. Treating the sidecar's 404 as
permission was wrong: the client cannot tell "this course has no owner" from
"this deployment told me nothing", so a visitor who opened a shared course could
bill the operator. The root cause was the sidecar itself, which gated on the
agent runtime although every persisted course has an owner regardless — the
persistence route resolves one for every request. It now gates on server
persistence, so the configuration that made 404 the universal answer has real
ownership facts to report, and the gate can refuse everything but a named owner.

Retry affordances answered to no gate at all. A viewer of a shared course with
one failed image was shown a Retry button that called the provider. Both retry
entry points and every surface that draws them now read one shared permission,
so what is offered and what is allowed are the same value.

Also: narration regeneration no longer pretends it can replace bytes behind a
live id — the exclusivity proof that would allow it is refused by construction
once references leave the browser, so it forks to a fresh id and says so; the
"already generated" test lets a finished deck answer from the document alone,
since scene order stops identifying an outline once slides are inserted or
deleted; stored assets record a specific media type rather than a generic
transfer type; the pane no longer asks the sidecar in browser-only mode; and the
funnel's docstring now states what the per-stage lock actually guarantees, which
is same-browser serialization and not a cross-browser compare-and-swap.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write

A delta review of the write-back found the first-pass fix still had a window,
and the reclamation it added could delete media the document already names.

The allocation was parked after an awaited local cache write. A scene committed
in that window reconciled against a registry that did not hold it yet, so the
document kept the placeholder — and the entry recorded a moment later then
answered the skip test as "already handled", so nothing could correct it. Parking
now happens inside the write-back, in the same synchronous turn as the decision
that nothing could take the reference; no await separates the live check from the
park. Allocations parked by an earlier pass are handed to their slides at the
start of the next one, before anything decides what still needs generating, so a
held allocation whose scene has since arrived becomes a rewrite rather than an
answer.

Reclaiming on a rejected write was unsound: a rejection does not prove the server
did not apply the write, so deleting the asset could break the scene that now
names it. The funnel decides instead, and says so: it reclaims only when no store
write was ever issued and nothing took the reference. Anything else is placed if
its slide exists and parked if it does not, so the next pass reuses the bytes
instead of paying for them again. When a write fails after part of it landed, the
live store is brought up to the document before the error is rethrown — otherwise
the next ordinary flush would overwrite the half that did land, with the ids
deliberately not reclaimed.

Parked allocations are now cleared with the course. Classic placeholders are
reused across runs, so one surviving an interrupted run would be handed to a
different slide of the next deck: the previous picture, on a slide whose provider
was never asked. Both classroom surfaces clear the arriving course, the deletion
cascade clears the deleted one, and clearing the database clears them all.

Two more ways generation could start without asking the gate are closed. An
overlapping pass — an outline retry re-enters generation with every outline while
the first is still working — re-requested elements whose provider call was
already in flight; a task that is not done is an answered request, not an
unanswered one. And narration regeneration in the timeline editor called the TTS
provider and allocated a pool asset with no ownership check at all; it now reads
the same permission, which withholds both the per-line and whole-timeline
controls and refuses the call.

Finally, a pane opened during the stage-link availability gap recorded the
sidecar's 404 for a course that was moments from existing and never asked again,
leaving the real owner locked out of generation until it remounted. Ownership is
re-fetched once the document becomes available; the gate stays closed until an
answer arrives, so asking again can only open it for someone entitled to it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): make the asset routes reachable, and stale snapshots harmless

A full-branch audit found that the deployment this project documents could not
store a single generated asset, and that several routes into durable storage
could still write a placeholder over a reference that had already landed.

The persistence route sent asset requests through the development
authenticator, which refuses outright in a production build that has not
explicitly opted into it — and the documented server-persistence recipe
produces exactly that build. Every store and every read answered 401, so images
and video failed on every slide while re-billing the provider on each retry, and
a narration failure stopped the deck at its first slide. Assets live in one
shared partition by design, so there was never anything per-caller for that
authenticator to decide: the route now resolves the asset principal itself,
alongside the owner it already resolves for documents. Runtime sessions are
genuinely per-learner and keep the development authenticator until real session
verification replaces it. And narration that cannot be stored no longer fails
its scene: the line stays unvoiced and retryable, which is what an image that
cannot be stored does to its slide.

Placeholders could also come back from behind. A queued autosave's snapshot, an
editor-history entry replayed by an undo, the departing save a course switch
flushes — each captures content at its own moment, and any of those moments can
predate a write-back. Point fixes at each producer would leave the next producer
to rediscover the bug, so the check lives at the write boundary every producer
passes through, and the allocation record it consults now outlives the parked
queue: a placeholder whose rewrite landed long ago is exactly the case it
catches.

Two ways generation could be lost or repeated are closed. A pass now claims the
elements it will reach and releases them however it ends, so an overlapping pass
stands down while an aborted one strands nothing — previously its tasks stayed
`pending` and every later pass skipped them with no retry control to recover
them. And the media abort controller is aborted before being replaced, so a
superseded pass stops calling providers instead of running on for a course the
user has left.

The remaining two are narrower. The workbench pane asks for ownership only after
a document load succeeds, and after every later one, mirroring the page route:
the load is what creates the ownership row the first time a course is opened, so
asking beforehand asked about a course that did not exist yet and locked its
author out for the mount. And the ownership gate on the timeline editor now
withholds narration regeneration alone; listening back to existing narration and
seeing whether a line has any spend nothing and stay available.

Known limitation, unchanged and now stated plainly in the comments that used to
point at it as a solution: nothing reclaims an unreferenced pool asset. The
registry sweep is written but not wired up, and the byte collector only reclaims
blobs no registry row names, so every narration regeneration and every abandoned
allocation leaves storage behind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): gate asset mutations, and make claims and allocation records survive a handoff

Opening the asset routes opened all of them. Reads and allocations are meant to
be as open as document reads and creates already are, but no authorization hook
was supplied, so the handler's default admitted PUT and DELETE too — and those
scope by principal key alone, which is one shared constant. Any caller who
learned an id, and a document read hands out every id its slides name, could
overwrite or destroy another author's media. Mutations now require the
deployment's credential, which in a production build without the development-auth
opt-in means they are refused outright; reads and allocations stay open. The
route comment says what the posture is and what it is not: the deployment-level
fence is the access code, and no per-principal quota is configured. The client's
own reclaim is best effort to match — losing an argument about deleting an asset
must not cost a task its retry, and the bytes are left for server-side
reclamation.

The pass claim could not survive the handoff it was written for. A retry aborts
the live media pass and starts its replacement in the same synchronous block,
long before the aborted pass's cleanup runs, so the replacement saw every element
still claimed, collected nothing, and returned — leaving each unreached element
at pending with nobody coming back for it and no retry control to recover it,
which is the exact failure the claim was introduced to prevent. A claim now
carries its pass's signal and is retired the moment that signal aborts, and a
pass releases only claims it still owns, so a late unwind cannot take its
replacement's work. Claims are also acquired at the single point every request
passes through, so a single-task retry participates too — previously a retry
awaiting its provider was invisible to a pass starting alongside it and both
called it.

The allocation record could outlive the bytes it named. It was written before the
write-back attempted anything and survived the reclaim that followed a failure,
so when the slide finally arrived the write boundary stamped a deleted id into
the document — and the placeholder it replaced was gone, which reads as already
generated and stops anything from retrying. The record is now written only where
the allocation is retained, and forgotten wherever a reclaim removes the bytes,
including the narration rollback path.

The tests follow. The route test drives the real storage handler against an
in-memory registry instead of a stub, so it can see what the resolved principal
is then allowed to do; the handoff test performs a real abort mid-pass rather
than starting from an already-aborted signal; and the guards that could only
assert file layout now assert the property they care about, or have been replaced
by behaviour.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): make media passes serial per course instead of tracking element ownership

Three rounds of per-element claims each produced a new way to lose an element.
Whole-pass reservations swallowed a Retry for an element the same pass had
already failed, leaving it pending with the affordance gone. Retiring a claim by
its signal freed an element whose commit was still uploading, so the replacement
pass paid for it twice. A claim held for a failed element stranded its retry. The
bookkeeping is the defect: every refinement of "who owns this element right now"
answered the question at a moment when the answer was already stale.

Passes for one course are now serial. A replacement aborts its predecessor, as
before, and then waits for it to settle before collecting. That removes the
question entirely: a commit already under way finishes — its bytes stored and its
reference written, so the new pass sees a resolved slide and skips it — and an
element the aborted pass never reached is still a placeholder and gets collected
like any other. The claim set, the reservations, the signal retirement and the
identity-checked release are all gone.

The task table is consulted for one thing only: an element that is generating
right now is a single-element retry running alongside the pass, and taking it too
would pay twice. Pending is deliberately not a skip reason — it means a pass once
intended to reach an element, which an abandoned pass leaves behind with nobody
acting on it, and reading that as answered is what stranded elements before. A
retry runs concurrently with a pass, because a pass never revisits an element it
has processed, and it re-reads the task after its own await and refuses before
touching it: marking first and refusing afterwards destroyed the failed state
that draws the affordance.

Browser-only mode is back to exactly what it was. The abort is now conditional,
the waiting does not apply, and the original status-based skip is restored
verbatim. Two baseline lines remain changed in each of the two files, and both
are behind a server-backed fork whose else-branch is the original.

Two smaller things. The allocation record becomes visible when a write goes on
the wire rather than when the round trip ends, and the write boundary reconciles
under the document lock rather than before it — a save queued during a write-back
was otherwise captured with the placeholder and, for a course the user had left,
had no corrective flush to follow. And the comments that said a refused reclaim
leaves its bytes for server-side reclamation were wrong: nothing collects them,
because the registry entry still names its blob and the sweep that would remove
it is not wired up. They now say the bytes leak.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on

Serializing passes moved their body out of the block that launched them, and
three things followed from that.

A pass now wakes when its predecessor settles, which can be after the user has
left the course. It enqueued before it looked at its signal, into a task table
keyed by element id alone — and placeholder ids are not unique across courses,
which is why the classroom clears that table on arrival. So a departing course's
pass seeded the arriving course's table with tasks carrying the wrong stage id,
and a Retry routes by that id: the reference went into the wrong document. A pass
now re-validates after the wait, before touching anything shared.

The same lateness broke the skip test. `documentSkipIndex` answers only while the
live store is on the pass's stage, and returning nothing put the collection loop
on the browser-only rule — a silent demotion from "the document is the authority"
to "this browser's task table is", on exactly the path where that table has just
been cleared. Every element the predecessor had committed was collected again,
paid for again, and its second write-back found no placeholder to rewrite, so its
bytes were parked where nothing will ever reference them. In server-backed mode
an unreadable document now means the pass stands down.

And waiting was unbounded. A commit is uncancellable: the asset client takes no
signal, and a document write cannot be half-undone. One stalled upload therefore
froze the course's media generation for the session — the replacement never
collected, the element sat on a skeleton that draws no Retry, and only a reload
recovered. The pass's signal is now threaded into the media proxy fetch, and the
commit is bounded by a deadline. The deadline is on the wait, not the work: the
commit carries on, and if it lands late the document simply ends up correct,
while the element becomes retryable and the queue moves on.

The tests that were meant to pin the previous round were not sensitive to it.
Two asserted end states where the mechanism only changes ordering, and one of
them rigged the document read so the assertion held whether or not the pass had
waited; a third covered half of what it claimed. They now observe the ordering
directly — nothing is issued while another pass for the course is working; in
browser-only mode a second pass reaches its provider immediately — and the
reconciliation under the document lock has a test that fails when it moves back
outside it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* revert(media): drop the commit deadline and the abortable download

The deadline bought less than it cost. Abandoning a commit after two minutes
makes the element retryable while the real commit is still running, so a Retry
starts a second commit for the same placeholder against the first: two provider
calls, two allocations, and whichever lands second stamps its result over the
other's task by element id. The allocation record is keyed by placeholder, so the
loser's cleanup erases the winner's record, and the write boundary then puts the
raw placeholder back into the document. That is the overlap serial passes were
built to remove, reopened through the one door serialization never covered.

So a stalled commit holds the course's media queue until it settles or the page
is reloaded, and that is written down rather than papered over. The wait is
unbounded on purpose: every ceiling on it turns out to be a way of running two
commits for one element.

Threading the pass signal into the download was also a mistake, in the other
direction. The provider call that produced the URL has already been billed, so
cancelling the download throws away work that is paid for — and the shared proxy
cache records a cancelled request as a transient failure against that URL, which
after three of them blocks it for every consumer in the session. Browser-only
mode never asked for this: it had no way to observe an abort there, which is
exactly why the bytes were kept. The signal is gone from the download again, and
`fetchAsBlob` is byte-for-byte what it was before this branch.

The regression guard for the stranded-element rule is restored alongside the
timing test that was meant to supersede it. It catches a different rule — a task
left pending being read as answered — and nothing else does: making the pass skip
pending leaves every other suite green.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes

Four things a deployment found once this was running for real.

A reference this application mints itself — a generation placeholder, a derived
narration key — was never in the pool, because the pool allocates every id it
holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed
it is a request that answers 404, one per element per load, forever on a course
that still holds placeholders. Every lease and probe now checks first. The check
is a negative test on shapes this application owns, not an id validator: the
pool's id domain stays unconstrained, and anything that is not one of ours is
still asked about.

The asset store can bound how much one principal holds, and enforces it inside
the write transaction, but nothing ever passed the number. It does now, with a
default rather than an opt-in: allocation is reachable by any caller a
deployment admits, and with one shared principal an unbounded store is
unbounded database growth with no operator-visible brake.

Refusing asset mutations to unauthenticated callers was not enough, because
every authenticated caller resolves to that same shared principal — so
authentication decided nothing, and any signed-in visitor could delete any id
they learned. Since this branch began storing media the registry is the only
copy a course has. Replacing and deleting are now refused to everyone, and the
browser no longer tries: an entry nothing references waits for server-side
reclamation instead. What a browser must still do is forget its own record of an
allocation that reached nothing, or a later save would stamp an id the document
has no reason to trust.

And a course generated before any of this holds placeholders in its document
with its bytes only in the author's browser. Those bytes are paid for, so the
author's next load converts them — stored to the pool and written back through
the ordinary commit path, with no provider call — instead of buying them again.
A row that records only a hosted URL is treated as absent: that URL is the
provider's address, not something a document may hold.

One renderer expectation moved with this. An untracked placeholder used to paint
as pending on first render because asking the pool left a lease in flight; it
settled to disabled a moment later either way, and now says so from the start.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): surface a full store as a refusal and convert legacy narration

A quota refusal reached the browser as HTTP 500 with a generic message, which
reads as a transient failure: the element kept a Retry that would pay a
provider again and be refused again. The store raises the contract's own error
and the handler maps it to 507, but the store answering a request is not always
built by the same bundle as the handler -- the persistence provider is reached
from the route bundle and from instrumentation, which is why its state lives on
a Symbol.for global -- and `instanceof` is false across that boundary while the
declared code is still right. Classify on the code as well as the class, and
make the code a permanent, persisted refusal in the browser: recorded locally so
it survives a reload, shown as "storage is full", and refused by the retry entry
point so a stale button cannot buy a second generation. Every other storage
failure stays retryable.

Convert what a pre-server-backed course still holds. Generated media is adopted
under either key this application has used for it -- the placeholder, and the
allocated id of a course converted once and later rolled back -- instead of only
the first. Narration is converted by a load-time pass over the open course's
speech actions, since nothing re-enters generation for an action that already
has an id: bytes to the pool, id written back through a funnel that mirrors the
media one, owner-only and server-backed-only. A line whose bytes are in no
browser is left alone rather than re-synthesized.

Also: the pool guard is now a positive `ast_` test rather than an enumeration of
the shapes we mint (imports never reach the pool, so this is safe in both
modes); the slide ref collection is an exported pure function so its four lease
sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero
as opting out and refuses a malformed value at startup instead of falling back;
the abort signal is re-checked after the cache read, before an uncancellable
commit; and the unused `removeAsset` and pool `replace` surfaces are gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* chore(storage): release 0.29.1

The asset HTTP handler now recognises a store refusal by the contract code it
declares as well as by its class, so a quota refusal raised in another module
realm answers 507 instead of 500. Same contract, stricter recognition, no API
change: a patch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): make a full store recoverable and adoption course-safe

Narration adoption read the local audio row by its derived key alone. That key
carries no stage id and the table is keyed by id alone, so two courses can mint
the same one -- a PPTX import numbers its scenes and actions deterministically,
which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally
a collision only means one course plays another's clip in one browser; adopting
it wrote that clip into the shared document permanently, for every device and
every visitor. A row that names a course is now adopted only into that course,
and a row from before that column existed only when the text it recorded is the
text of the action being converted.

A full asset store was made permanent last round, which was wrong three times
over: it overwrote the refused bytes with an empty blob -- on the conversion
path that row is a course's only copy of its own media -- it kept sending the
rest of the deck to a provider against a ceiling it already knew was reached,
and it left no way back once an operator raised that ceiling. A full store is
neither the content's fault nor the configuration's, so it is now its own case:
the bytes are kept, the pass stops at the first refusal, and the element shows
the reason together with a Retry that re-attempts the upload from those bytes.
Nothing retries automatically, so no one is re-billed.

The narration write-back now reaches the write boundary every producer of a
durable write passes through, not only the dirty mark: adoption never deletes
the derived row, so a snapshot that reverts the rewrite is adopted again on the
next load and allocates a fresh asset every time. Adoption is also mounted by
both classroom surfaces rather than one, takes the course's abort signal, and
re-validates that this browser still has the course open before each write.

ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the
docstring already claimed it was: its only other consumer is lazy and memoised,
so a malformed ceiling let the process boot and then failed every persistence
request, documents and runtime included.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): remember a full store per course, and never lose retained bytes

A stopped pass left the elements it never reached as placeholders with no
persisted record -- deliberately, since nothing was attempted for them. But
that left the next load with no reason not to try: it called a provider for the
next placeholder and was refused at exactly the same point, once per reload,
indefinitely. A full store is not a property of any slide. It belongs to the
deployment and changes for reasons the document knows nothing about, so it is
now remembered once per course in the browser's device KV. A pass that finds
the marker stands down before spending anything and leaves every placeholder
its "storage is full" state and its Retry; the first upload that succeeds
clears it and the next pass runs normally.

Narration adoption latched per course so it runs once per load, and the latch
outlived the abort that leaving a course performs. On a surface that stays
mounted across switches -- the workbench pane is one component for every course
it shows -- owner course A, visitor course B, then back to A skipped exactly
the clips the abort had cut off, and nothing else converts them. The latch is
released with the abort now, and a course adopts one run at a time so a
re-entry cannot hand a clip a second allocation while the previous run's
uncancellable tail is still settling.

A quota-blocked element retried into a network error or a 500 lost the bytes
that were kept for it: the retry deleted the row before attempting the upload
and wrote no replacement for an error carrying no structured code, so the next
retry went back to a provider for media this browser had a moment earlier. The
row now survives until an upload succeeds, the failure handler keeps whatever
bytes the attempt was given, and the retry asks the question a pass asks --
does this browser already hold bytes for this element -- rather than reading an
error code that a second failure has already overwritten.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D

* fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes

Narration adoption admitted a stage-less row only when the text it recorded
matched the action being converted. Both of those columns were added to the
local audio table by the very change that moved narration onto allocated ids,
so a row still carrying a derived key has neither: the rule refused every real
pre-allocation course and passed only on fixtures built from post-allocation
rows. What the row cannot say, the key can. A derived key names two clips only
when two courses share a scene order and an action id, and an action id repeats
only when something other than the generator minted it -- an import numbers
them by slide position. So a key built from a generated action id is adopted on
that basis, a key an import could have reproduced still needs matching text,
and a row that names another course is refused however unique its key looks.

Handing a re-entering caller the adoption run already in flight undid the latch
release it was paired with: that run is bound to the signal the departure just
aborted, so it stops at its next clip while the caller -- which has the course
open and a live signal -- is told the work is done, and an effect replayed as
mount, cleanup, mount adopts nothing at all. A later caller now waits for the
uncancellable tail and scans again, which costs a lookup on a course that has
nothing left and finishes the clips the abort cut off on one that does.

One attempt at an element now reports both facts its callers need instead of a
bare boolean: whether the store refused it for room, and whether bytes actually
reached the store. Leaving a course clears the task table, so a retry that
landed afterwards read "no failed task" as success and deleted the row holding
the only copy of the media. Nothing is inferred from that table any more.

Reading the localStorage property can throw where storage is denied by policy,
typeof included, so the availability check moved inside the guard: this metadata
is best-effort, and a rejection here strands a generation pass that has already
enqueued its tasks.

A retry is never blocked by the per-course "store is full" marker, but a retry
that is refused again re-sets it, and adoption now reads and writes the same
marker rather than issuing one refused upload per clip on every load. The two
canvas element renderers and both thumbnail renderers show the reason beside
the Retry, so a full store does not look like an ordinary failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(media): probe a full store instead of standing down, and pair notices with a Retry

Narration adoption was given both halves of the per-course "the store is full"
marker last round: it stood down when the marker was set, and it set the marker
when its own upload was refused for room. Those halves are only safe together
if something can lift the marker, and for adoption nothing could. It has no
affordance of its own, it stood down before reaching its own clear, the media
pass returns before its marker gate when there is nothing to generate -- so a
narration-only deck, or one whose slides are already satisfied, painted no
storage-full element and offered no Retry -- and narration generated rather
than adopted allocates directly rather than through the media commit. The
course's cached narration was then lost for good, where before it converted on
the first load after the ceiling was raised.

The gate is a probe now. A marked course attempts exactly one clip per load:
refused, it stops and the marker stands, which costs what standing down cost;
stored, it lifts the marker and finishes the course. Adoption spends no
provider money, so the whole cost of probing a store that is still full is one
refused upload. Generated narration lifts the marker too.

The three surfaces that gained a failure notice last round drew it for any
failure with a reason, including the one refusal that is reachable without
server-backed persistence, so a browser-only deck painted something it had not
painted before. The notice is drawn beside a Retry and nowhere else, which is
what it was added for and what leaves browser-only output unchanged. Both are
now asserted through the render harness the surface matrix already had.

A caller arriving while a rescan is queued shares it rather than appending
another. One rescan converts whatever the run in flight left and every later
one would find an allocated id on every action, so a chain bought nothing and
turned a single stalled upload into a course that never adopts again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(media): treat a refusal for room as a fact about one clip, not the deck

The asset store checks each write against the headroom it has left, so a store
that refuses a long opening clip can still hold every short clip behind it.
Narration adoption assumed the opposite: it broke the deck at the first refusal
and then re-attempted that same first clip on every later load, because the
document names it first. A deck whose longest clip exceeds current headroom
therefore never converted the clips that would have fit, with no affordance to
recover it -- the state the probe was introduced to remove, reached through a
narrower door.

An unmarked load now attempts every clip, skipping the ones that do not fit,
and remembers the condition only if the load ends with clips it still could not
store. A marked load spends its single upload on the smallest clip left rather
than the first one named: that is the clip that answers the question the marker
asks, because if the smallest does not fit nothing does. The media pass keeps
stopping at its first refusal, and for a reason adoption does not share --
every element it attempts costs a provider call.

A rescan several callers share took the newest caller's signal, and the newest
caller is not necessarily the one still there: a surface that opened a course
and closed it again would stop work a surface still showing that course was
waiting for, and that surface is latched, so it would never ask again. The
shared run now takes a signal that is aborted only once every caller has left.

The comment claiming the shared rescan contains a stalled upload was wrong --
the rescan is chained off the run in flight, so a stalled upload leaves every
caller pending exactly as a chain would. It claims the bounded queue it
actually provides, and the stall is recorded as a limitation.

The failed-state containers took their stacking classes unconditionally, so
markup differed in browser-only mode even though nothing moved on screen. Those
classes are applied only when there is a notice to stack, and the tests assert
the exact class attribute rather than a substring.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(media): stop narration adoption writing the media pass's store-full marker

The marker means "do not call a provider for this course". A path is entitled
to write it only if its own refusal cost a provider call, and narration
adoption's refusals cost nothing: it uploads bytes this browser already holds.
The store also checks each write against the headroom it has left, so a clip
that does not fit says nothing about whether a slide's image would. Adoption
was writing it anyway, and one over-long narration clip was therefore enough to
stand a course's entire image pass down on every later load -- on a store that
had just accepted adoption's other clips. The author could still recover each
element by hand, every load, for ever.

Three rounds of narrowing this seam produced a finding each time, so it is
removed rather than narrowed again. Gone: the marker read, the single-clip
probe, the smallest-clip selection, and the up-front read of every row into an
array -- which also retires a sampled-then-stale flag and the retention of a
whole deck's blobs for the length of a run, and returns the loop to streaming
one row at a time.

Adoption's rule is now that every load attempts every clip it holds, once; any
failure skips that clip and the load continues. The noise the coupling was
meant to avoid does not arise, because after the first load the clips still
outstanding are exactly the ones that did not fit -- normally none, or one.

A successful write still clears the marker, and that is a different kind of
statement: a write that went through is a fact this run established, where a
refusal is an inference about what some other write would cost. For a course
whose media needs nothing, adoption and generated narration are also the only
paths that can establish it.

The failure module still documented the deck-wide premise this contradicts. It
now says what is true: the check is per write, and the media pass stops the
deck as a judgement about cost rather than about certainty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(media): bound a full store's cost from the store's own arithmetic

Removing the store-full marker from narration adoption removed its bound too,
and the code then asserted the bound was unnecessary. It is, on a store with
room for most of a deck. On the store the whole mechanism exists for -- the
ceiling reached, nothing fitting -- the outstanding set after every load is the
entire deck, so a thirty-clip course posted thirty full blobs on every load,
indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the
server hashes the whole payload, and only then takes a per-principal lock and
sums every entry that principal owns before saying no.

The bound needs no flag, no key and nothing carried between loads. The store
asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows
while a run is uploading, so a clip refused for want of room implies every clip
at least that large is refused for the rest of that run. The run keeps the
smallest size it has been refused and skips anything no smaller without
uploading it; a smaller clip is still attempted, because it may fit. A deck the
store refuses entirely now costs one upload per successive size minimum instead
of one per clip, and a deck it has room for costs nothing extra, because
nothing is refused. Only a refusal for room lowers the bar: a dropped
connection says nothing about how much room there is.

The deck-wide certainty premise the failure module retracted last round still
stood verbatim at the site that implements the stand-down. Both copies now say
the same thing: the check is per write, and the pass stops the deck as a
judgement about cost rather than about certainty.

The comment on adoption's marker clear now names its price. Narration of a few
hundred bytes fits in headroom an image does not, so a proven write can let the
next pass buy one more image that is refused again -- bounded at one, and the
price of the alternative being a course whose media never generates again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(media): state the adoption bound exactly, and stop three comments describing the old rule

The comment introducing the in-load bound gave its cost as "at most a handful,
and the first load pays the most". Neither clause is a property of the rule. A
clip is skipped only when something no larger was already refused, so a
fully-refused deck costs one upload per successive size minimum in document
order: one when the clips grow, about ln N for an arbitrary order, and one per
clip when they only shrink -- a long opener followed by terser lines is exactly
that shape. And no load is cheaper than the first, because the bound resets per
run and a refused clip stays outstanding. The comment now says that, and points
at what would make it exactly one for any ordering: the store returning its
remaining headroom in the refusal's existing details channel, which the server
leaves empty today.

Two other comments still described the previous rule -- "attempts every clip it
holds, every load" -- one of them twenty lines above the paragraph that
introduces the bound, in the same block. Both now say what the code does.

The bound's soundness is worth stating where a maintainer will look for it:
quota is charged at full length with no discount for a duplicate, the sum it is
checked against joins entries to blobs so the collector cannot lower it, the
check takes a per-principal lock before summing, and replace and delete are
refused to every browser. Nothing a run can do makes room appear inside it.

One test installed a row implementation and replaced it wholesale a few lines
later, so the first was dead and the survivor dropped the text the first clip's
import-shaped key needs for the ownership rule -- it passed on the coincidence
that the fixture's default text is the action's. Merged into one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(media): keep a Retry from re-buying parked media, and state the store seam once

Five findings from an inline review.

The standalone classroom route asked the ownership sidecar once per load and
recorded only stage ownership when that ask failed. Every non-answer fails
closed, so one transient 5xx left the genuine author with no resume, no Retry
affordance and no legacy narration converted for the rest of the load, with
nothing to change it short of a reload. The failure now records the fail-closed
answer explicitly -- an answer an earlier load established must not outlive the
failure that replaced it -- and an unresolved answer is asked for again, a few
times over a few seconds. A real answer, however unwelcome, is final.

A Retry could pay a provider for media the pool already held. When the bytes
are stored and only the write-back fails in a way that keeps the allocation, it
is parked and no local row exists, because that row is written only after a
successful write-back. Retry now reads the parked queue exactly as the pass
does and re-attempts the write-back: it re-keys the task done when the document
takes it, leaves the entry parked when the slide still does not exist, and
stays failed and retryable when the document refuses again.

Object URLs a parked allocation owns are revoked when the entry is dropped. The
commit path leaves them alone while the entry is parked, because it is then the
only thing holding bytes this tab can render, so a course switch or a stage
deletion was pinning the whole blob for the life of the tab. An entry a slide
has already taken is left alone: the task table is displaying those URLs.

The fallback lookup for cached bytes is a stage-scoped scan, and the keyed
lookup misses for every row the commit path writes, so a pass was materializing
and sorting the course's whole media table once per element. One scan per pass
now, built on the first miss. It is sound and not merely cheaper: an element
asks only for its own placeholder, and every row a pass writes carries the
placeholder of the element that wrote it.

"The store accepted a write, so it is not out of room" was enforced at three
call sites under slightly different conditions, which made it a convention the
next pool write path could silently break. It is stated once, in putAsset, for
the course whose bytes it just stored.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-09-11 19:38:08 +08:00
0b641a9403 fix(importer): convert Equation.3 OLE formulas via MTEF v3, surface degrade telemetry (#1411)
* fix(importer): convert Equation.3 OLE formulas via MTEF v3, surface degrade telemetry

Legacy courseware stores formulas as Equation 3.0 / MathType OLE objects
whose only renderable form inside the .pptx is a WMF preview picture.
The importer cannot rasterize WMF, so those formulas degraded to a
hardcoded 1x1 "transparent" placeholder — actually a 50%-alpha red
pixel that rendered as a pink block once stretched over the formula
frame, with no way for callers to notice the content loss.

- Convert `Equation.3` / `MathType` OLE objects to LaTeX: detect the
  progId, resolve the embedding, parse the OLE compound file (cfb) and
  convert its `Equation Native` MTEF v3 stream (new utils/mtef.ts) —
  fractions, radicals, scripts, fences, big operators with per-family
  limit variations, embellishments, and Symbol-font local character
  encodings. Slot order follows the rtf2latex2e reference
  implementation and real MathType streams ([main, lower, upper]), not
  the archived spec prose. Any failure falls back to the picture path.
- Fix the placeholder constant to a truly transparent pixel; keep
  recognizing the legacy red one (exported isPlaceholderDataUrl).
- Surface degradation instead of failing silently: optional
  ImportPptxOptions.onWarning receives machine-coded warnings
  (media-unconvertible / formula-fallback-image / formula-degraded /
  element-dropped) at every placeholder consumption point (image,
  background, shape/text pattern fills, math fallback); a throwing
  sink is isolated so telemetry can never fail an import. A formula
  whose fallback picture is also a placeholder keeps its plain text as
  a text element.

* fix(importer): address review — LSCRIPT base duplication, one-sided fences, Symbol table, depth caps

Review round on #1411 (thanks @wyuc — the fuzz safety-net validation and the
adversarial constructions found what the spec-conformance rounds could not):

- tmLSCRIPT: stop re-emitting a script slot as the base group (isotopes
  rendered as {}_{6}^{12}{12}C and the output doubled per nesting level —
  a 221-byte stream could OOM the worker); the base is the following
  sibling, so the template emits only {}_{sub}^{sup}, and a leading script
  no longer steals the previous atom via the trailing-script lookahead.
  tvLSUPER writers that emit a single slot now fall back to it.
- One-sided fences: render the missing side as a null delimiter
  (\left. / \right.) instead of an unmatched \left — piecewise-function
  braces (tmBRACE var 1) produced KaTeX-invalid output that silently
  degraded to flat text.
- SYMBOL_FONT_LATEX corrected against URW StandardSymbolsPS AFM + Adobe
  AGL: 0x3C/0x3E/0x5B/0x5D are literal < > [ ] (≤/≥ live at 0xA3/0xB3,
  now mapped, along with the rest of the Symbol operator block);
  0x22/0x24/0x5C are ∀/∃/∴; added Chi/vartheta/varsigma, the phi/varphi
  split (0x66/0x6A), and 0x5E = \perp. Unmapped font-local bytes >= 0xA0
  now flag degraded instead of passing silently.
- tmLIM: variation roles were inverted — spec + rtf2latex2e eqn.c say
  0 = upper limit, 1 = lower limit; single-limit writers keep their limit
  via the same slot fallback as the big operators.
- Hardening: MAX_DEPTH = 200 nesting cap and a 64 KiB LaTeX output cap,
  both throwing MtefParseError (deep bombs now fail loud instead of
  RangeError/OOM); video poster joined the placeholder warning points.
- Tests: +6 — three REAL Equation Native stream fixtures (round-tripped
  from a legacy deck, pinning the font-local encoding and big-op slot
  order), tvLSUPER single-slot, tmLIM both roles, script-nesting perf
  guard; fence tests now assert KaTeX renderability instead of pinning
  broken strings. Suite: 85.
- Version 0.1.5 -> 0.2.0 (new public option/type/export + Math.degraded).
  Lockfile re-anchored on main with pnpm@10.28.0: only the cfb additions
  remain, no unrelated churn.

* fix(importer): address review round 2 — tmLIM function slot, full Symbol high-half, bra/ket, cap tests

- tmLIM: emit the main slot FIRST followed by the limits and inject no
  operator name (the reference `39.1 = limit: lower, #1 #2` puts the
  function in the main slot — hardcoding \lim duplicated it and glued to
  letter-leading main slots, crashing KaTeX with an undefined control
  sequence). An empty main slot falls back to \lim as a neutral base.
- SYMBOL_FONT_LATEX: completed the 0xA0–0xFF block from the URW AFM
  (~70 positions: ∫ ∑ ∏ ⟨⟩ ∂ ∇ ⇒ ⇔ ⋅ ′ ∅ ⊆ ⊇ ∈ ∉ ∪ ∩ …). Unmapped
  font-local codes now throw MtefParseError instead of passing through
  as Latin-1 (0xF7 was an integral extender rendering as ÷ — plausible
  but wrong math); radicalex (0x60) and C1 controls (0x80–0x9F) also
  throw, taking the picture fallback.
- tmDIRAC: var1 renders a bra `\left\langle L\right|`, var2 a ket
  `\left| R\right\rangle` per the reference (was wrapping both sides).
- Embellishment records now count against the record budget (a 10 MB
  embellishment-only stream no longer allocates 1 GB before the output
  cap fires).
- Tests: +6 — every cap now has its own assertion (output-length width
  case at 15k sibling CHARs, record-count at 20k, in addition to the
  existing depth test), previously-dangerous Symbol codes verified from
  the AFM, unmapped-code rejection, embellishment budget, tmDIRAC
  KaTeX-validity for all variations. Fixture claims corrected (no
  big-op selector in the three real streams; the reading is pinned by
  the spec-conformance tests). Suite: 91.

* test(importer): pin 0xD6/0xF3 Symbol glyphs through KaTeX, fix big-op test title

Follow-up to the round-3 review: the two glyph fixes landed without a
test, the BigOp test title still named the wrong slot order, and the
table comment claimed the high half was complete while unlisted
positions intentionally throw.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* docs(importer): correct Symbol table comments — 0xF7 is parenrightex not an integral extender, high-half coverage is ~50 of ~70 positions

---------

Co-authored-by: Percy <percy@PercydeMacBook-Pro.local>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 13:09:46 +02:00
puxiaoandwyuc d32c3f2a00 fix(importer): stop PPTX import from hanging in non-browser environments (#1424)
* fix(importer): stop PPTX import from hanging in non-browser environments

* format

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-11 12:53:27 +02:00
LeoParkerOuandwyuc 28133f8a67 fix(renderer): lazy-load optional ECharts runtime (#1420)
* fix(renderer): lazy-load optional ECharts runtime

* test(chart): assert setOption is called after runtime is ready

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-10 23:30:46 +02:00
9fe650d994 fix(quiz): resolve AI answer keys to option values before grading (#1328)
* fix(quiz): resolve AI answer keys to option values before grading

AI-generated quiz answer keys are unreliable: the model may write the
correct answer as option CONTENT ("(6, 2)"), as a LETTER ("A"), or as a
formatting variant (full-width chars, extra spaces) of either. Grading
compares exact option values, so a learner picking the correct option
was still graded incorrect whenever the key was written as content or
in a different format.

- generation: normalizeQuizAnswer now resolves every answer variant to
  the matching option value (NFKC + whitespace-stripped + letter-prefix
  canonical matching against option values and labels); unknown answers
  pass through untouched
- grading: resolveAnswerKeyToValue retro-fixes already-generated courses
  whose stored keys hold content variants, mapping them to option values
  when exactly one option matches canonically; ambiguous keys are left
  alone instead of being silently re-pointed
- tests: letter/content/full-width/trailing-space/multi-answer/unknown
  cases + end-to-end grading with a content-variant key

* chore(generation): bump version to 0.3.2 for answer-key fix

The repo requires every publishable @openmaic package change to ship with
a version bump (CI check-package-version-bumps).

* style(generation): prettier-format package.json after version bump

* fix(quiz): use resolved answer representation in review renderers; prettier

- QuestionCard + quiz-view review modes highlight the option canonically
  matching the stored answer key (answerIncludesOption), so persisted
  content-variant keys highlight correctly after the grading fix
- prettier formatting on touched files

* test(quiz): guard possibly-undefined answer in multi-choice test

* fix(quiz): narrow answer-key matching to exact unique alignment; project review UI from the same resolver

Per review: matching policy narrows to the DSL option-identity contract.
- Generation: exact value or unique exact label alignment; unknown and
  ambiguous keys stay unresolved (fail closed). Prompt example unified to
  the DSL object-options + value-answer shape.
- Grading: persisted keys resolve via the same exact alignment; the
  resolver returns the option's actual value when exactly one option
  matches canonically-or-exactly; ambiguous keys stay untouched.
- Review UI: QuestionCard and both quiz-view review modes highlight via
  the same resolver projection (answerIncludesOption(q, optionValue)).
- Removed the broader equivalence rules (NFKC, whitespace stripping, case
  folding, wrapped/prefixed-letter probe) — semantic synonyms such as
  正交/垂直 remain uncovered, hence Refs #1326.

* fix(quiz): resolve answer keys for persisted keys only; align quiz template with the DSL

- grading: compatibility resolution (label -> value) now applies to the stored
  answer key only. A learner submission is compared as submitted, so an option
  label can no longer be accepted as a different option's value.
- quiz-content/user.md: replace the string-options/`correctAnswer` example with
  the DSL shape the system prompt defines ({ label, value } options and
  `answer` as an array of option VALUES), so generated keys are values.
- tests: regression that a label submission grades incorrect while the same
  question's value submission grades correct.

* test(generation): update scene-prompt golden snapshot for the DSL-shaped quiz template

The golden test pins the rendered user prompt for every scene kind; the
quiz-content template now emits the DSL shape, so the snapshot follows.

* fix(quiz): read label-stored correct answers in the editor's option rows

`quiz-edit-ops.toRows` decided correctness with raw value equality, so an
existing label-keyed quiz showed A/B as correct (the review UI and the grader
both resolve labels) while every row read as incorrect internally — and the
next option edit or reorder rebuilt `answer` from those rows and silently
dropped the key.

- `toRows` now uses the same resolved projection (`answerIncludesOption`).
- `normalizeQuizAnswer`'s docstring now describes the exact-only behavior it
  has had since the review: formatting variants are not normalized and
  ambiguous entries pass through untouched.
- Regressions: an unrelated option-label edit and an option reorder both keep
  a label-stored answer, and a multiple-choice toggle preserves it. All three
  fail against the previous raw comparison.

---------

Co-authored-by: CrazyPandaP <CrazyPandaP@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-10 23:28:20 +02:00
ClayExcandwyuc 1e10f60b15 fix(generation): assign unique IDs per media request (#1368)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-04 11:23:21 -04:00
Yizuki_Ame eb2cbc701a feat(pro): generate concise conversation titles (#1275)
* feat(storage): add automatic conversation title lifecycle

* feat(agent): generate concise conversation titles

* feat(agent): schedule automatic conversation titles

* fix(agent): harden automatic title lifecycle

* fix(agent): refine conversation title prompt

* fix(agent): harden automatic title boundaries
2026-09-04 10:52:25 -04:00
puxiaoandwyuc b4ed3e90f3 fix(importer): pnpm install fails on Windows (#1372)
* update @openmaic/importer/rollup.config.js

* update @openmaic/importer/package.json

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-04 09:59:54 -04:00
dd74cfe23b feat(web-search): add Exa provider (#1342)
Co-authored-by: white638 <212083322+white638@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-03 06:14:05 -04:00
Yizuki_Ameandwyuc e71b00c5c9 fix(agent): preserve generate_scene failure causes (#1316)
* fix(agent): preserve generate_scene failure causes

* test(agent): harden scene failure coverage

* test(agent): isolate warning log format

* test(agent): isolate warning log level

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-03 02:15:53 -04:00
Rupam Pal 132d01cd04 feat(canvas): add double-click text insertion (#1310)
* feat(canvas): add double-click text insertion

* fix(ci): format canvas double-click text insertion

* fix(canvas): prevent text insertion on double-click of locked elements

- Add guard to check if double-click event target is inside an editable element
- Prevents text box insertion when double-clicking locked elements (backgrounds/watermarks)
- Fixes event-propagation bug where locked element selection handler returns before stopPropagation()
- Add regression test suite for canvas double-click handler covering locked elements and blank canvas cases

Addresses review feedback from CxuPercy on PR #1310

* fix: remove unused import in canvas-double-click-handler test

* fix: improve test implementation to avoid DOM dependency

- Rewrite tests to use mock objects instead of actual DOM elements
- Properly type MockTarget interface to satisfy TypeScript checks
- All tests now pass locally without DOM environment requirement

* test(e2e): add end-to-end tests for canvas double-click text insertion

- Add test for double-clicking blank canvas area to create new text element
- Add test to verify double-clicking existing elements does NOT insert new text
- Tests actual editor interactions, not just mock targets
- Addresses CxuPercy's review feedback for proper regression test coverage

* chore: format code with prettier

* test(e2e): simplify double-click text insertion test

- Remove flaky second test that checks element visibility state
- Focus on core functionality: double-click on blank canvas creates new element
- Use more lenient assertion (toBeGreaterThan) to account for dynamic elements
- Increase wait time for slide editor to fully render (1000ms)

* fix(storage): increase test timeout for replace/release operations

The 'replace retires rather than revokes an issued snapshot; release revokes both'
test performs multiple async operations (put, resolve, replace, release) that can
exceed the default timeout. Increased timeout to 10 seconds to accommodate I/O operations.
2026-09-03 11:35:30 +08:00
Zhenghan Songandwyuc ca3d785221 build(node): enforce direct dependency engine floor (#1337)
Why:
- The root Node 20.9 contract predates direct runtime dependencies whose declared minimums now reach Node 22.19.
- README, contributor, localized docs, and the OpenMAIC extension skill repeated the stale supported-version claim.

What:
- Raise the root Node minimum to 22.19 and align every operator, contributor, locale, and skill prerequisite.
- Add a check that compares the root minimum with installed direct production dependency engine minimums.
- Run the new contract check in CI after the frozen dependency install.

Risk:
- This changes the declared minimum only; no upper bound is added and Node 24 compatibility remains a separate concern.
- The localized docs build was validated with the independent #1306 boundary fix from PR #1307, which is not included here.

Tests:
- RED on the base: engine check reported pi-agent-core, pi-ai, svg-pathdata, and undici floors
- GREEN: root minimum 22.19 satisfies 35 engine-constrained direct dependencies
- Node 20 lockfile-only install reports the root unsupported-engine warning
- Node 22 frozen install and postinstall
- Docs build with PR #1307 boundary: 34 pages and all locale postexport checks; docs types:check
- Root Prettier, ESLint, TypeScript, i18n, package-version, and internal-dependency gates
- Root pnpm test: 7137 passed, 81 skipped

Live Docs:
- GitHub issue #1304 tracks the Node contract; #1306 / PR #1307 tracks the separate docs-build prerequisite.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-02 04:29:52 -04:00
e58c700082 fix(generation): keep diagram node position off the hover transform (#1339)
* fix(generation): keep diagram node position off the hover transform

A CSS `transform` overrides the SVG `transform` presentation attribute, so a
generated `.node:hover { transform: ... }` rule wipes out the `translate()`
that positions the node and snaps it toward the origin while hovered. The
diagram prompt only asked to "avoid hover transform conflicts", which was too
vague to prevent this.

Spell out the contract instead: the outer `<g>` carries positioning only, and
hover/active animations apply to an inner group.

Refs #1300

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(generation): bump package version to 0.3.4

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-02 03:54:25 -04:00
Mark Tranandwyuc 55e324768e fix: keep videos playing inline on mobile (#1321)
* fix: keep videos playing inline on mobile

Add playsInline to every playback video path, including the legacy app renderer, renderer-backed playback, and the shared renderer package. This keeps mobile browsers such as WeChat on Android from interrupting inline playback and restarting videos after an interruption.

Add regression coverage for both app renderer modes and the package renderer.

* chore(renderer): bump package version

* chore(renderer): bump package version to 0.1.6

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-02 02:36:47 -04:00
liangb612 d4ef5faa76 fix: replace rm -rf with cross-platform fs.rmSync in package scripts (#1338)
* fix: replace rm -rf with cross-platform fs.rmSync in package scripts

The build/clean scripts used the Unix \
m -rf dist\ command, which fails on Windows where cmd.exe has no rm. Replace with Node's built-in fs.rmSync, available on all platforms (Node >=20.9).

* chore: bump @openmaic package versions for cross-platform script fix
2026-09-01 21:39:10 -04:00
简律纯andwyuc 32bfd197eb fix(generation): load micropip before generated imports (#1282)
Generated Pyodide widgets can import micropip before it is loaded, aborting initialization. Add ordered prompt guidance, cover the packaged asset, and bump the generation package patch version.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-01 04:33:17 -04:00
Zhenghan Songandwyuc 96343c440e fix(docs): isolate standalone build from product middleware (#1307)
Why:
- The docs Turbopack project walked up to the repository root and compiled the product middleware.
- The v1.0 middleware's product-only feature-flag alias made every current docs build fail before static export.

What:
- Resolve the absolute packages/docs directory from the Next config URL.
- Set that directory as the standalone docs app's Turbopack root.

Risk:
- The setting affects only the docs Turbopack project; root product middleware and aliases are unchanged.
- Webpack-specific behavior is unchanged because the docs workflow builds with Turbopack.

Tests:
- packages/docs pnpm build: 34 static pages and all locale postexport checks passed
- packages/docs pnpm types:check
- Root Prettier, ESLint, TypeScript, and i18n gates
- Root pnpm test: 7137 passed, 81 skipped

Live Docs:
- GitHub issue #1306 tracks the regression and build evidence.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-01 15:57:46 +08:00
Yizuki_Ame 9d16e68296 fix(workbench): persist Pro conversation titles (#1273)
* feat(storage): persist manual session titles

* feat(agent): add session title patch route

* fix(workbench): preserve renamed session snapshots

* fix: use typed session title store

* style(storage): format session title tests

* test(storage): align agent session schema contract

* fix(workbench): fence stale session title bootstrap

* fix(workbench): close session title races

* fix(workbench): harden session title consistency

* fix(workbench): preserve unconfirmed session titles

* fix: harden session title reconciliation

* fix: finalize session title consistency

* test(workbench): document session title recency
2026-08-31 04:21:01 -04:00
wyuc dfebbcf33f perf(classroom): speed up classroom loading (media hydration, sidebar thumbnails, media range requests) (#1276)
* perf(classroom): defer non-priority media blob hydration off the load path

Entering a classroom awaited the full mediaFiles restore — one object URL
per restored image/video blob — before the loading gate opened. For
video-heavy courses that is hundreds of MB of IndexedDB materialization
blocking first paint.

Split the restore into two phases:

- The awaited phase now builds metadata-complete task entries but only
  creates object URLs for failed rows (none needed) and for media
  referenced by the scene the classroom opens on (persisted cursor, else
  the first scene). buildRestoredMediaTasks gains an optional
  shouldHydrateBlob predicate, defaulting to eager hydration so existing
  callers are unchanged.
- Remaining blob-backed records hydrate in the background, chunked over
  requestIdleCallback (setTimeout fallback). Each deferred task stays
  'done' — so generation resume never re-runs it — but carries no
  objectUrl yet, which the media resolution state machine already renders
  as a pending skeleton until the URL lands.

Background hydration is guarded per record: a task that was replaced
(classroom switch, regeneration, retry) or that belongs to another stage
is skipped and its freshly minted URLs are revoked immediately.

* perf(classroom): lazy-render slide thumbnails in the playback sidebar

The playback scene sidebar mounted a full SlideCanvas for every slide
scene the moment the classroom opened (and every off-screen video element
opened a preload="metadata" fetch), because SlideThumbnail's existing
visible prop was never passed.

Extract the editor nav rail's near-viewport IntersectionObserver hook to
lib/hooks/use-near-viewport.ts and gate the sidebar's slide thumbnails
through it: only scenes within 200px of the viewport render the live
canvas; the rest show SlideThumbnail's existing placeholder until scrolled
near. The placeholder keeps the same box size, so gating never shifts
layout.

The shared hook now starts hidden instead of eager: with the previous
eager-initial state, opening a long deck still mounted every canvas for a
frame before the observer could flip off-screen items off. The observer's
guaranteed initial delivery flips near-viewport items within a frame; the
no-IntersectionObserver fallback (e.g. jsdom) defers to a microtask so
the effect never synchronously re-renders.

* perf(media): lazy image decoding and HTTP Range support for classroom media

Renderer: BaseImageElement's <img> now carries loading="lazy" and
decoding="async", so thumbnail-heavy surfaces (playback sidebar, editor
nav rail, course cards) no longer fetch and decode every slide image up
front. In-viewport images are unaffected — the browser fetches them
immediately. slideToPng forces eager loading inside its permanently
off-screen snapshot tree, where lazy images would otherwise never fetch
and exports would capture blank slides.

Server: GET /api/classroom-media/[classroomId]/[...path] now answers
single byte-range requests with 206 Partial Content (Content-Range,
Accept-Ranges, correct Content-Length), enabling progressive playback and
seeking for hosted video/audio instead of downloading whole files. Suffix
ranges are supported; unsatisfiable ranges get 416 with the full size;
unsupported units or multi-range sets fall back to the plain 200 full-body
response, which is always a legal answer. Existing caching headers are
kept on every response shape.

* chore(renderer): bump version to 0.1.4 for image lazy-loading change

* fix(classroom): guard deferred hydration against restarted tasks and bound idle waits

A deferred record is only ever 'done' without an objectUrl; a task that
regeneration or retry restarted passes through pending/generating, so skip
those instead of attaching stale persisted bytes. Also give the idle
scheduling a timeout so a busy main thread cannot starve hydration.

* fix(classroom): stop superseded deferred hydration before minting URLs

A restore epoch captured at apply time now gates each idle chunk: loading
another classroom or reloading the same one invalidates any older hydration
loop, so it neither keeps scheduling work for an abandoned classroom nor
attaches bytes read by an older load to a newer load's tasks.

* fix(classroom): keep 416s uncached and classify element-keyed media as priority

A 416 with public immutable caching can poison the media URL for later
valid requests, so range errors now send Cache-Control: no-store. Priority
classification also collects the opening scene's media element ids, since
task lookup binds records keyed stage:<elementId> even when the slide slot
carries a different opaque ref.

* fix(classroom): tie deferred hydration liveness to the classroom load token

The restore epoch only advanced when the next load reached apply, so an
abandoned classroom kept hydrating during the next load's storage/network
phase. Compose the epoch with the load's isCurrent (load token plus effect
cleanup) so navigation stops the loop at the next idle boundary.

* fix(classroom): classify legacy-recovered media before deferring it

Legacy singleton video recovery assigns a placeholderRef only at the end of
the task build, so a record keyed by an allocated id without placeholderRef
was deferred even when it backs the opening scene's gen_vid_* element. Run a
metadata-only build first to learn each record's effective ref and classify
against it, keeping the first visible page's legacy video eager.

* fix(action): wait for deferred video bytes before starting play_video

executePlayVideo treated status done as immediately playable, but a deferred
restore is done without an objectUrl: the renderer shows a skeleton, no
<video> exists, and the later hydration never retriggers play, leaving the
action stuck until the safety timeout. Readiness now follows the renderer's
contract (only done-with-bytes is playable) at the initial check, the
subscription exit, and the post-subscription recheck; the failed skip applies
whether or not a wait happened.

* fix(classroom): include the stage whiteboard in priority media refs

The stage-level whiteboard stays open across standalone classroom switches,
so its media can be visible before any scene is. Classifying it as deferred
left visible whiteboard media pending behind idle hydration chunks.
2026-08-30 03:48:04 +08:00
wyucandClaude Code f8b30ef0f0 fix(storage): prune redundant intermediate message_update frames (#1279)
* fix(storage): prune redundant message update frames

* fix(storage): prune all message update runs

* chore(storage): release 0.27.0 for pruneMessageUpdates

Co-Authored-By: Claude Code <noreply@anthropic.com>

---------

Co-authored-by: Claude Code <noreply@anthropic.com>
2026-08-30 01:31:58 +08:00