mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 09:24:43 +08:00
main
135
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1635b16899 | docs: add multilingual vocational task engine guide (#1749) | ||
|
|
fa47efe1b3 |
feat(persistence)!: server-backed persistence only, with a one-way legacy browser import (#1710)
* ci: run CI for the persistence-default integration branch Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(docker): start PostgreSQL by default and add a single-user owner * feat(docker): start PostgreSQL by default and add a single-user owner `docker compose up` now runs a server-backed app against a bundled PostgreSQL, started only once PostgreSQL is healthy and published on 127.0.0.1 only, with a new built-in singleUser owner auth method: every request resolves to one fixed owner that may publish, and no anonymous cookie is minted. Single-user mode (OWNER_SINGLE_USER=true, OWNER_SINGLE_USER_ID default "local") runs with or without ACCESS_CODE. Without one, the server logs one prominent startup warning that anyone who can reach it shares, edits and can delete the single library; it never inspects request peers or forwarding headers. It excludes PERSISTENCE_SHARED_OWNER_ID and follows the sharedTeam registration rules. Its principal gets a claim candidate, so a browser's earlier anonymous work is claimed into the single owner (OWNER_CLAIM_TRIGGER=auto in the Compose defaults). Compose defaults live in docker-compose.defaults.env, read before .env.local so they can be overridden. The app warns when its declared published address (OPENMAIC_PUBLISH_ADDRESS, passed by Compose) is not loopback while PostgreSQL uses the default password. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(docker): keep claims explicit and let .env.local set DATABASE_URL - Compose no longer defaults OWNER_CLAIM_TRIGGER to auto: an irreversible merge of every browser's anonymous library into the single owner must not happen on upgrade. The single-user principal still gets a claim candidate, so POST /api/identity/claim works; the docs explain it and warn about auto on a deployment several people used. - The bundled DATABASE_URL moves to docker-compose.defaults.env, read before .env.local, so an external database or a rotated password set there keeps working. The password is inserted unencoded, so it must be letters and digits (documented). - Upgrade docs no longer claim browser-stored courses are copied as they are opened: they stay in the browser, are not deleted, and are moved by the automatic browser-to-server migration shipped with this work. - Single-user mode without ACCESS_CODE now logs one warning instead of two; the docs note that the single owner id is permanent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(persistence)!: server-backed persistence is the only persistence (#1707) * feat(persistence)!: make server-backed persistence the only persistence Courses, folders and folder membership, chat history and learner runtime, and generated media are now always stored on the server through the embedded /api/persistence endpoint. The browser keeps no durable data of its own. - Remove the build-time NEXT_PUBLIC_PERSISTENCE switch: the client bootstrap always configures the HTTP document, runtime and asset seams, and the Dockerfile, docker-compose.yml, the Pi route and the whiteboard runtime no longer read it. Unconfigured seams refuse to resolve instead of falling back to IndexedDB, and the learner key is always the server-derived one. - Remove the browser write paths: the browser document, runtime and asset stores as backends, the browser-mode branches of stage, folder, chat, media, narration and import storage, the Dexie backup export/import and the whole-database clear. The library lists through /api/stages, folders through /api/folders, and membership now goes through /api/folders/members. Regeneration always forks to a fresh asset. - Device-local state moves to its own IndexedDB database (lib/device-storage, maic-device-cache): the local media and narration cache and refused-bytes retention, staged PDF images, undo history, browser voice profiles (carried over once from the old database) and the auto-voice cache. "Clear Local Cache" clears exactly that and no longer claims to delete classrooms or chat history. - Keep the pre-server data readable for the one-way importer: the Dexie schema and read-only accessors for its tables and for the browser document, runtime and asset databases live in lib/legacy-browser-storage, which nothing on the regular paths reads and which has no write path. - Require a database: the server exits at boot without DATABASE_URL, with a message naming `pnpm db:up` and `docker compose up`. `pnpm db:up` / `pnpm db:down` start and stop the Compose postgres service alone, published on 127.0.0.1 through docker-compose.db.yml; .env.example carries the matching local DATABASE_URL. - E2E specs seed their courses through the persistence endpoint, each under a fresh id, instead of writing IndexedDB. - Guard tests pin the boot requirement, the absence of any NEXT_PUBLIC_PERSISTENCE read, and the legacy-storage boundary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: run the E2E suite against PostgreSQL The app refuses to start without DATABASE_URL, so the E2E job gets a postgres:16 service and a job-level DATABASE_URL. The unit, storage and render-service jobs are unchanged and need no database. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: document always-on server persistence and the database requirement README (English and Chinese), the deployment guide in every locale, the extension cookbook and the changelog now describe PostgreSQL as required: `pnpm db:up` for local development, `docker compose up` for a deployment, and an external database for Vercel and other serverless hosts (the Deploy button prompts for DATABASE_URL). They list what stays in the browser, drop the NEXT_PUBLIC_PERSISTENCE instructions, and say that courses an earlier browser-only build stored stay in the browser until the one-way importer moves them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(pbl): use the server-derived learner key for PBL v2 runtime PBL v2 synchronization, hydration and drain resolved the learner key from their default watermark KV store, which reads or mints a device key in localStorage instead of asking runtime configuration. The runtime store is the server one, which refuses any key but the owner's, so saving a course with a PBL v2 scene failed before the document write, and the minted key landed in the slot the one-way importer reads as the pre-server runtime partition. The learner key now comes from getLearnerKey(args.kv): the configured server-derived key, or an explicitly injected store. The default KV store still holds the drain watermarks and nothing else. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(persistence): pin that clearing the cache keeps legacy databases - Seed the four pre-server databases (MAIC-Database, maic-documents, maic-runtime, maic-asset-pool) with real rows, run Clear Local Cache, and assert every database and row survives; assert the device database name differs from every legacy name. - The legacy-module import rule now also catches relative static, side-effect, dynamic and require imports at any depth. - New rule: the runtime learner key is never read from a KV store the caller made up, only from configuration or an injected store. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * build(dev-db): run the development database as its own Compose project `pnpm db:up` / `db:down` shared the checkout's default Compose project with `docker compose up`, so they recreated or stopped a running stack's database. docker-compose.db.yml now declares its own project (`openmaic-dev-db`), which gives the development database its own container and data volume; it no longer touches the stack, and the two do not share data. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs: start the development database before `pnpm dev` CONTRIBUTING, the getting-started guide in every locale and the startup modes reference now run `pnpm db:up` and set DATABASE_URL before `pnpm dev`. README, the deployment guides, the cookbook and the changelog describe `pnpm db:up` as a separate development database. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(persistence): sharpen the legacy-storage and learner-key guards - The Clear Local Cache test closes its seeding connections to maic-documents and maic-runtime, so a regression deleting either fails on the assertion that names the database instead of by timeout. - The legacy-module and Dexie import rules accept comments inside import() / require(), so a webpackIgnore hint no longer hides one. - The learner-key rule also flags an injected-looking name that is bound to a default local store in the same file; its remaining limit is documented, with the behavioural PBL test as the guard. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * build(dev-db): pin the development database's Compose project `pnpm db:up` / `db:down` now pass `-p openmaic-dev-db`, which takes precedence over COMPOSE_PROJECT_NAME from the shell or a .env file; the file's `name:` alone could be overridden and point the scripts at a running stack's database. The compose test pins the flag. The docs (all locales) and the file header say the development database is shared by every checkout on the machine, and how to run a separate one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(persistence): import legacy browser data into the server, once (#1708) * refactor(persistence): prepare the seams the legacy browser importer uses - Legacy module: read-only accessors for the auto-voice cache, the course ids the old media and narration tables name, the pre-server learner key, and asset bytes from the browser asset pool. - Narration adoption exports its row ownership rule so the importer applies the same one to rows of the old database. - The quiz legacy migration is split into a non-deleting core (importLegacyQuizSnapshot) and the regular path, which still deletes the keys it migrated. - The lazy document migration exports its snapshot canonicalizer. - Clear Local Cache keeps the importer's per-owner ledgers, as it keeps the pre-server learner key. - A library-changed window event lets an open library list courses that arrive in the background. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * feat(persistence): import legacy browser data into the server, once On the first load after the upgrade, the client moves what earlier builds kept in this browser to the server, for the owner the server resolves: courses from the browser document store and the original tables (with chat, learner runtime, playback position, roster, folders and membership, pre-runtime quiz state), their media bytes through commitToPool and the existing write-back funnels, the bytes of server courses from earlier opt-in server builds that only the old tables hold, and device-only rows (failure records, refused bytes, auto-voice clips) into the device cache. It is automatic and silent, runs after the page is idle, never writes to a legacy store, and records every step in a per-owner ledger so it resumes after a crash and never imports twice. Runs are serialized across tabs with Web Locks. A course the owner already has stays authoritative; one whose id another owner holds is imported under a derived fresh id; one the owner deleted on the server stays deleted. Transient failures retry on a later load with backoff; permanent ones are recorded per item. The module is temporary and self-contained; its README lists the removal steps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): pin that each owner keeps its own import ledger Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): run the legacy import against the app routes and a browser A PostgreSQL suite routes every request the importer and the app seams send to the real persistence, library and folder handlers, owner resolved from the anonymous cookie: a full import, a course another owner holds, a course the owner deleted, and an existing course whose media is filled in. An end-to-end spec seeds an old pre-server database in Chromium, loads the app, and checks that the course reaches the library with its narration uploaded, stays in the browser untouched, and is visible from a fresh browser context with the same owner cookie. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs: describe the one-way import of legacy browser data The README (and its Chinese version), the deployment guide in every locale and the changelog now say what the importer does instead of announcing it: courses an earlier browser-only build kept in the browser move to the server automatically on the first load after the upgrade, and the browser copy stays untouched. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): probe server assets through the shared asset-URL owner The importer asked the pool directly whether an id exists, which the asset-URL ownership boundary forbids outside use-asset-url. It now uses assetRefExists, the existing metadata-only probe. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): guard the importer's pool probe like every other entry point Only a reference the pool could have issued is probed, as the lease guard requires; the importer is added to the list of guarded entry points. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): import legacy browser data once per browser Review round 1 of the legacy browser importer. - Once per browser, not once per owner: the data belongs to whoever used the browser before the upgrade, so the first owner the import runs for claims it and any other owner gets nothing. The ledger is one key, maic:legacy-import:v2, recording the owner only as a SHA-256 digest (an anonymous owner id is a bearer credential) and a random salt. - Handoff: when the claiming owner is retired (403 OWNER_RETIRED), or the current owner holds courses or folders the importer created (a claim moved them), the current owner continues the unfinished items; courses already done are never imported again, so a claim neither duplicates nor resurrects them, and a half-copied course keeps its runtime. - Fresh ids come from the salt and the legacy id through SHA-256, not from the owner: unlinkable, unpredictable, and the same in every tab. - 401 pauses with backoff instead of closing the import; 403 FORBIDDEN_LEARNER stops the run and leaves the item pending. - Without Web Locks: ledger writes merge with the stored copy, and the library is listed again right before a course is placed, so a second tab does not import a course the first one just created. - Media the server refuses for good gets the app's failed-media record instead of a dangling reference; the legacy bytes stay. - A legacy whiteboard or PBL session is not created when the server already has an active one of that kind. - A folder that is gone leaves the course unfiled instead of failed, and a membership naming a folder the old database lacks counts as unfiled. - Narration rows from before the course column are adopted when exactly one legacy course names their key; unconvertible chat rows are noted; an unreadable document-store copy falls back to the original tables; PBL v2 documents go through the save-path strip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): cover the importer's handoff, tabs, refusals and review gaps Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs: say the legacy import runs once per browser The README (and its Chinese version), the deployment guide in every locale, the changelog and the module README now say that the first owner to load the upgraded app in a browser claims its legacy data, what the handoff to an account does, that the ledger holds no owner id, and how refused media and the new pause and skip cases behave. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): note runtime the importer cannot find without the old learner key Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * feat(identity): let an owner confirm it absorbed a given owner through a claim GET /api/identity/merged-from?salt=<hex>&digest=<hex> answers whether a claim merged into the requesting owner an owner whose id hashes to SHA-256(salt, id). It reads only the requester's own owner_merges rows, never lists or names an owner, resolves the owner like every other owner-scoped route, and is uncacheable. The one-way import of pre-server browser data uses it to decide whether a different owner may continue an import the first owner started. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): hand the legacy import to another owner only on a confirmed claim Review round 2 of the legacy browser importer. - The only way the browser's legacy data moves from the owner that claimed it to a different owner is the server confirming, through GET /api/identity/merged-from, that the current owner absorbed that owner through a claim. The inferred signals (an owner that never wrote, a library listing that holds imported courses) and the retired flag are gone; a 403 OWNER_RETIRED only stops the run and leaves items pending. - The first owner is bound only once the run's first authenticated request of its own (the library listing) succeeded, so a run that dies before that leaves the ledger unclaimed. - The owner is recorded as SHA-256 of the per-browser salt and the owner id; the server computes the same value from the salt the browser sends. - A "not absorbed" answer is reused for an hour instead of asking on every load, and the claiming owner stored first survives another tab's save unless that tab handed the import over; completion survives too. - Run-level failures (401, OWNER_RETIRED, FORBIDDEN_LEARNER, OWNER_BUSY) from any call stop the run and leave the course pending, including the "is the id taken" read and the library re-list, which could previously mark the course failed and close the import. - Refused media without a generation request (the user's own) gets a new non-retryable ASSET_REFUSED code, so it shows as failed without a Retry that could only fail; a skipped singleton runtime session leaves a note. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): pin the importer ledger's owner merge and the singleton-skip note Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs: the legacy import moves to another owner only on a confirmed claim The README (and its Chinese version), the deployment guide in every locale, the changelog and the module README now say that the import runs once per browser and moves to another owner only when that owner claimed the original one, confirmed by GET /api/identity/merged-from; that the owner is recorded as a salted digest; and how refused user media shows. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): never let a transient failure settle a legacy import item Review round 3 of the legacy browser importer. - A read of the old browser stores that fails in storage (an aborted IndexedDB transaction, a closed database) now pauses the run and leaves the course pending. Only a record that cannot be migrated, parsed or validated is skipped. This covers the course read (which no longer falls back to the older table copy on a storage failure), and the quiz-scene and speech indexes, which used to read a failure as "no scenes" or "no holder" and drop quiz state or narration for good. - After a session-id collision, a failed read of the session propagates and is retried; only a successful read that shows the id taken skips it. - Without Web Locks, two tabs of different owners can both find the ledger unbound. Every ledger write keeps the owner stored first, and the run now re-reads the stored ledger after binding, before folders, before each course and before each course document is created, and stops if another owner holds it. - The "not yours" answer is reused for ten minutes instead of an hour, a stored wait further ahead than the longest the importer sets (a clock that ran ahead) counts as passed, and expired answers are pruned. - Generated media refused without an error code keeps its Retry; only media nothing can regenerate gets ASSET_REFUSED. A poster that fails transiently leaves the video pending instead of being dropped. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): cover transient reads, racing owners, the not-yours cache and the claim client Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(persistence): say what the ledger's salt protects, and the new pauses The salt defeats precomputed and cross-browser tables, not a targeted guess against one browser's ledger; the module README and the digest's comment now say so. The README also covers storage read failures, which pause instead of skipping, the two-owner race without Web Locks, and the ten-minute recheck. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * feat(identity): bind a browser's legacy data to one owner on the server The one-way import of pre-server browser data now asks the server whose data it is, instead of deciding in the browser. - legacy_import_bindings (browser_id primary key, owner_id, timestamps), provisioned with owner_merges on fresh and upgraded databases. - POST /api/identity/legacy-import-binding {browserId}: one atomic insert (ON CONFLICT DO NOTHING) and a read-back; answers only whether the requesting owner holds the browser. Same-origin JSON, resolved and write-fenced like every owner write, 400 for a malformed id. - A claim participant re-keys the claimed owner's bindings to the account inside the claim transaction, so the account simply holds the browser. - Owner resolution refuses any request carrying X-OpenMAIC-Legacy-Import with 409 LEGACY_IMPORT_NOT_BOUND unless the owner it resolves to holds that browser: one central fence every owner-scoped route goes through. Requests without the header are unaffected. - GET /api/identity/merged-from is removed; the binding replaces it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): import through fenced clients bound by the server Review round 4 of the legacy browser importer. - The ledger holds a random browser id and no owner information; the owner digest, the "not yours" cache and the client-side handoff are gone. A run first asks the server to bind the browser and imports only when the requesting owner holds it. - Every importer request goes through its own clients that carry the browser id, so the server refuses any write under an owner that does not hold the browser: after a cookie switch in another tab, from a page whose memoized owner is stale, or from a tab that lost the binding race. The runtime learner key is asked for during each run, never taken from the page. 409 LEGACY_IMPORT_NOT_BOUND stops the run with items pending. - The write-back funnels, commitToPool and the PBL save-path strip take an optional store, put or runtime, so the importer passes its fenced ones. - A failed read of one legacy course holds up only what it could affect: the quiz-scene and speech indexes leave the course out as unindexed, and only quiz state or narration it could own waits. Read failures are counted per course; after five failing runs over at least a day, a course whose document-store copy cannot be read falls back to the table copy, or is skipped with the reason when there is none. - Runtime ids are rewritten by replacing only their course segment, so a short course id cannot corrupt the prefix or the learner segment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): cover the server binding, its fence, bounded reads and short ids Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs: the server binds a browser's legacy data to its first owner The README (and its Chinese version), the deployment guide in every locale, the changelog and the module README now say: once per browser; the server binds the browser to the first owner; a claim carries the binding to the account; the importer's requests are refused for any other owner. They also describe the bounded pause for a course the old storage cannot read, and the removal steps for the server side. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): keep a dropped asset request pending in the legacy import The asset client reports a fetch that got no answer (a dropped request, an aborted upload, the existence probe's deadline) as status 0 with HTTP_REQUEST_FAILED, which the importer read as a final refusal: the media was marked failed and the course could complete without it. Read that code as transient, keep the client's local validation failures final, and treat a 2xx/3xx answer a client could not use as transient too. Unit tests drive every client the importer uses with a rejected fetch and a non-JSON answer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(persistence): ask a refused binding again only every ten minutes A browser whose legacy data another owner holds sent the binding request (a write) on every page load. Record a ten-minute recheck in the ledger (clamped like every other deferral), so a claim that moves the binding is still picked up, and open the old databases only once the browser is bound. The read budget now counts consecutive failing runs (a run that reads the course resets it) and counts a first failure dated in the future from the next one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(identity): answer 503 with the cookie when the import fence cannot read The legacy import fence now reads the header name from the shared constant, and a database error while reading the binding answers the usual JSON error (503 PERSISTENCE_UNAVAILABLE) with the resolution's Set-Cookie values instead of escaping as a bare 500. Route tests cover that, a bind by an owner a claim retired (403 OWNER_RETIRED, no row), and one header name on both sides. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(persistence): cover dropped asset requests, the recheck and the read budget With the real asset client's error for an upload and an existence probe, the media stays pending and a later run imports it. A non-holder's five loads send one binding request and never open the old databases, and a claim still hands the import to the account after the delay. The read budget's run count and time span each hold on their own, a clock that ran ahead does not hold a course open, and a good read resets the count. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(persistence): complete the legacy import removal steps List every file the removal touches (the claim scenarios test, the claim participant table and list, every doc), keep the bindings table for one release after the code goes because of rolling deploys and give the DROP statement for that release. Add the claim's eighth participant to the README lists, fix the fresh-id wording in the changelog, and describe the recheck delay, the transient status-0 asset failures and the consecutive read budget. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): keep legacy quiz state through Clear Local Cache until the import completes Pre-runtime quiz drafts, answers, results and attempt ids exist only in localStorage, and the one-way importer copies them to the server with their course. Clear Local Cache deleted them even while the import was still pending (it waits for an idle page and can stay pending after a failure), so that quiz progress was lost for good. Clear Local Cache now keeps the four quiz key families unless the importer's ledger records the import as complete, read through a new `legacyImportIsComplete` helper in the ledger module. The ledger key moves there too, so the two modules do not import each other. Once the import is complete, the keys are cleared as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs: show DATABASE_URL set in the local development setup The getting-started guide showed the required DATABASE_URL line commented out, so copying the block left the server unable to start. Each locale now shows `pnpm db:up` and, separately, the `.env.local` line to set. The `.env.example` header said every variable is optional; it now says that provider variables are, and that DATABASE_URL is required except under docker-compose.yml. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(changelog): drop two upgrade notes the final code contradicts The Compose entry said its default DATABASE_URL overrides .env.local; the defaults file is read first, so a value in .env.local wins, as the same entry says further on. The runtime-session entry offered turning server persistence off as an upgrade path; that switch no longer exists. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(env): restore the ACCESS_CODE warning notes in .env.example A later change to the example file dropped the lines saying that an unset ACCESS_CODE fails open, that the server then logs a one-time startup warning and GET /api/health reports accessCodeConfigured: false. They are back, and say that single-user mode logs its own warning in place of the generic one, as the startup code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * ci: run every app PostgreSQL suite in the contract job The app-domain step named its suites one by one, and seven had been left off, among them the legacy importer's route suite. No other job gives the app suites a database, so they skipped everywhere. The step now lists every app `*.pg.test.ts`; all of them work in a schema or database of their own, or truncate only their own tables, and pass before and after the storage package suite in the same database. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(identity): establish the anonymous owner before the first API request (#1721) * fix(identity): establish the anonymous owner on the page response On a browser's first load the page sent several API requests without an owner cookie; each minted its own anonymous owner and set its own cookie, and the last answer to arrive won. Work done in between (the legacy import binding, a first course write) then belonged to an owner the browser no longer presented. The middleware now mints the anonymous cookie on document navigations that carry no valid one, with the route handlers' own value format and attributes (the cookie module is Edge-safe now: Web Crypto instead of node:crypto), and forwards it on the request so the page render resolves the same owner. API, RSC, prefetch and Server Action requests never mint there, and a valid cookie is never replaced. It mints only when neither OWNER_SINGLE_USER nor PERSISTENCE_SHARED_OWNER_ID is set; host registrations are invisible to an Edge middleware, so the README tells hosts where to skip it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(legacy-import): bind only an owner the browser already presents A binding request whose owner was minted by that very request answers 409 OWNER_NOT_ESTABLISHED with the minted cookie and binds nothing: other cookieless requests may still be minting owners of their own, and a binding to this one could be left with an owner nobody presents. The importer treats it as transient and retries on a later load. It also says plainly when a request after its own successful bind is refused as another owner (the owner cookie changed mid-run); items stay pending. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(legacy-import): cover a first visit whose first answers arrive out of order Against the real routes and PostgreSQL: a browser without a cookie loads a page, sends two requests at once, and the first answer reaches it only after the importer bound the browser; the import completes under the one owner the page established. A second case binds nothing for an owner the bind itself minted and imports on the next run. The E2E drives the same ordering in Chromium by holding the first owner-scoped answer until the binding answered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * test(legacy-import): drop an unused initial assignment Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(changelog): note the page response now sets the anonymous owner cookie Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(identity): let hosts turn page-response minting off The middleware cannot see owner auth methods a host registers, so a host with anonymousFallback: false still got an anonymous cookie on a cookieless page load. Beside the host credential it is a claim candidate; with OWNER_CLAIM_TRIGGER=auto it is claimed and cleared on the next request, again after every such page load. OWNER_ANONYMOUS_PREMINT (default on; true/1/false/0, anything else fails startup) is read by the middleware before minting. At startup the server warns once when a registration turns the anonymous fallback off while pre-minting is still on. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(identity): document OWNER_ANONYMOUS_PREMINT Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(identity): keep the anonymous identity for 400 days and renew it in use The anonymous cookie had a fixed 30-day lifetime from its mint and was never renewed, so every anonymous visitor lost their whole library 30 days after the first visit however active they were; with persistence on the server only, that affects every anonymous deployment. It now lasts 400 days (the browser cap, one shared constant), and every route handler and Server Action response that resolves to a valid anonymous cookie re-sends the same value with a fresh Max-Age (best-effort where next/headers refuses writes). Page responses never renew, so pages stay cacheable, and a host principal never renews the anonymous cookie beside it. A response that clears the cookie (a claim, a retired owner) must not renew it, whatever order a route merged its headers in: withRequestOwner, and the owner-events retired answer, keep only the clearing value for a cookie a value clears. A renewal that lands after a claim cleared the cookie is pinned by a test: the next write with it, alone or beside the account, writes nothing, and its answer clears it again without renewing it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * docs(identity): 400-day renewed anonymous cookie; when to turn pre-minting off Hosts set OWNER_ANONYMOUS_PREMINT=false only when their registration sets anonymousFallback: false; hosts that keep anonymous visitors need pre-minting and skip it per request in middleware for requests their methods authenticate. The CHANGELOG notes the new lifetime and renewal, and that a lost cookie means a new owner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz * fix(identity): forward owner cookies from routes that resolve the owner directly The Pi chat and whiteboard-visibility routes and asset-id document extraction resolved the owner themselves and never sent its Set-Cookie values back, so an anonymous identity used only through them was never renewed (and a minted one never stored). attachOwnerCookies attaches a resolution's cookies to a response a route built itself, with the clear-wins rule, before a stream's body starts. Both Pi routes answer through it, success and error alike; resolveServerAsset hands the cookies back with every answer and the extraction route attaches them. A guard test scans for every module outside the identity seam that calls the owner resolution (or the asset helper) directly and requires it to forward the cookies; the two callers that cannot are listed with the reason. Behaviour tests cover the three paths. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: wyuc <dhq1204@yahoo.com> |
||
|
|
877bf792d7 | fix(generation): preserve template literals in interactive HTML (#1706) | ||
|
|
a686481e73 |
feat(auth): owner identity seam — composable auth methods, owner-keyed persistence, host hooks, claims
* ci: run CI for the owner identity seam integration branch Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(auth): owner identity seam, phase 1 — pluggable authenticator with today's behavior * feat(auth): add the owner identity seam with built-in authenticators Introduce lib/server/identity/: an OwnerAuthenticator contract (SubjectKind, OwnerPrincipal, AuthOutcome), a single-shot, boot-validated configureOwnerAuthenticator() registry, and the anonymousCookie and sharedTeam built-ins that reproduce today's owner resolution exactly (same anonymous_id cookie, owner id strings, statuses and env semantics). Every owner-scoped route handler and the workspace Server Action now resolve their owner through one memoized per-request resolution instead of the three previous copies (agent-runtime/owner.ts, with-owner.ts and the Server Action re-implementation). Publish and unpublish decide from the principal's course:publish role instead of an anon: id prefix. An invalid credential from an authenticator is answered with 401 on every surface and never falls back to an anonymous owner. A source scan keeps the anonymous owner cookie and anon: prefix checks out of code outside the identity module. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(auth): describe the owner identity seam and host registration Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(auth): tighten the identity boundary guard and Server Action cookies - Scan workspace package sources too, flag imports of the concrete built-in authenticators outside lib/server/identity, and police the slice/substring/indexOf/includes spellings of an anon: id-shape check. Stop re-exporting the built-ins from the host-facing index. - Refuse setCookies returned by authenticateFromContext, as the authenticate() fallback already did, and document that it must write cookies itself through next/headers. - Let the route-test owner stub carry explicit kind/roles, and cover the 403 forbidden branch of publish and unpublish for a signed-in owner without course:publish. - README: registration conflicts fail the boot; an invalid principal is rejected per request with a 500. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner * feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner Runtime sessions and assets on /api/persistence now take their identity from the same memoized owner resolution as documents, instead of a client-chosen x-learner-key behind a development bearer token that shipped in the public bundle. - Runtime: the storage handler's learner key is principal.ownerId. A path or body naming another learner answers 403 FORBIDDEN_LEARNER and another learner's session answers 404, whatever x-learner-key says; the header is no longer read. The browser learns its key from the new GET /api/persistence/learner-key and sends no credential of its own. The Pi whiteboard routes derive the same key from the owner. PERSISTENCE_DEV_TOKEN, NEXT_PUBLIC_PERSISTENCE_TOKEN and PERSISTENCE_ALLOW_INSECURE_DEV_AUTH are gone from the server path and the client, and server-auth.ts is deleted. Learner merge stays refused. - Assets: allocations land in a per-owner partition (owner:<ownerId>), so quota is per owner and replace/delete reach only the owner's entries. Reads stay capability-by-id where a course viewer needs them: another owner's committed entry is readable while a live course references it. Legacy entries in the old shared partition stay readable by id, and may be replaced or deleted only by an owner who owns every course referencing them. Server-generated media is stored under the run owner's partition; asset-id extraction and vision image resolution use the same rule. - Tombstones: runtime of a deleted course reads as absent (404, empty lists) and takes no new sessions or records, through a guard in the app's runtime composition; no package schema change. The identity boundary scan now also fails on any source that reads x-learner-key or the retired development token variables. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(persistence): describe owner-keyed runtime and assets, drop the dev token Remove the development persistence token from the server-persistence recipe, .env.example, the Docker build arguments and the deployment docs; describe runtime learner keys, per-owner asset partitions and quota, the legacy shared-partition rule and the tombstone behavior; and record the breaking change for existing server-persistence runtime data under Unreleased. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(storage): scope asset references to principals, typed stage refusal, uniform create collision @openmaic/storage 0.32.0. - PgDocumentStore takes assetReferencePrincipals: with reference tracking on, a write references and commits only entries held by the listed principals. An id naming another principal's entry records no row and changes no lifecycle column, like an unknown id, and the write is never refused. Without the option any held entry is referenced, as before. AssetCollector takes a matching per-document-owner function so the one-time backfill scopes each document the same way. This keeps a document naming a leaked id from committing, exposing or pinning another owner's allocation. - RuntimeStageNotFoundError: a store may refuse createSession for a stage its host considers absent; the runtime handler answers 404 STAGE_NOT_FOUND, recognizing the error by class or by code. - A session create over a taken id answers 409 SESSION_ALREADY_EXISTS whoever holds it. It used to answer 403 for another learner's session and 409 for one's own, which told a caller whose session an id was. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): close owner-seam boundary gaps in references, tombstones and legacy assets - Document writes reference and commit only the writer's own asset entries and legacy shared ones (owner-bound document store and the collector backfill), so another owner's course naming a pending id can no longer commit, expose or pin it. - The foreign-read rule additionally requires the referencing live course to belong to the entry's owner, as defense in depth for reference rows that did not come from a scoped write. - The tombstone refusal for new runtime sessions moved from a raw path comparison in the route into the guarded store's createSession (RuntimeStageNotFoundError, 404 STAGE_NOT_FOUND), so no encoded spelling of the path routes around it. - Replacing or deleting a legacy shared entry checks reference ownership and mutates in one transaction under the entry row lock, which a concurrent reference insert must wait for. - Tests build the uncommitted-but-referenced, tombstoned-with-rows and foreign-reference states directly so each half of the read rule is observable, and a PostgreSQL suite holds a competing reference open across a legacy delete. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(persistence): upgrade note for dev-token-gated deployments, owner-scoped references Tell operators whose only access gate was the development token to put the deployment behind an access code or gateway, register an owner authenticator, or turn server persistence off before upgrading, and describe that a course references only its owner's media. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(storage): per-owner reference principals and a typed taken-id conflict - PgDocumentStore's assetReferencePrincipals is now a function of the store's bound owner, evaluated per write, instead of a fixed list. A store re-bound with forOwner therefore scopes to the new owner rather than carrying the previous owner's principals onto its writes. It has the same shape as AssetCollector's backfill option, so one function serves both. - PgRuntimeStore and BrowserRuntimeStore raise RuntimeSessionExistsError for a taken session id, and the runtime handler classifies it as 409 SESSION_ALREADY_EXISTS even when its re-read cannot see the holder (for example a session a host-side wrapper hides). Recognized by class or by code. Still the unreleased 0.32.0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): guard the Pi whiteboard runtime and answer hidden collisions with 409 - Every server-side runtime writer now uses the tombstone-guarded store: the Pi native whiteboard service gets it through the same helper as /api/persistence/runtime/*, so a deleted course takes no new whiteboard runtime either. - Creating a session whose id is held by a session a tombstone hides answers the uniform 409 SESSION_ALREADY_EXISTS instead of a 500, covered on PGlite and on real PostgreSQL. - The owner-bound document store passes its reference-principal function, so the scoping follows the bound owner. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(auth): owner identity seam, phase 3 — accounts through an identity gateway (trustedProxyHeader) * feat(auth): owner identity seam, phase 3 — trusted-proxy-header authenticator Add the trustedProxyHeader built-in: real accounts through an identity gateway (oauth2-proxy, Authelia, a Keycloak-based proxy, an institutional reverse proxy) that signs the user in and forwards the verified user in a request header. OpenMAIC never handles passwords or tokens. - Selected with OWNER_AUTHENTICATOR=trusted-proxy and validated in instrumentation register(). TRUSTED_PROXY_SECRET (32-1024 printable ASCII characters) is required; header names default to x-openmaic-proxy-secret and x-forwarded-user, and an optional groups header with TRUSTED_PROXY_ADMIN_GROUPS grants admin. An unknown selector, a TRUSTED_PROXY_* variable without the selector, a shared owner id or a host-registered authenticator alongside it, and malformed, reserved or repeated header names fail the boot. - Trust boundary: Next.js exposes no TCP peer address to route handlers, middleware or Server Actions, and fills x-forwarded-for from the socket only when the client sent none, so no address allowlist is offered. The gateway secret is compared in constant time (hashed, timingSafeEqual). - Per request: a missing or wrong secret, a missing or blank user, a user header with a comma (duplicate lines arrive comma-joined), or a user whose proxy:<user> id fails the owner id guard is 401 INVALID_CREDENTIAL, never an anonymous owner. The guard failure is a 401 rather than a 500 because the value comes with the request. The user is trimmed and keeps its case. Every gateway user holds course:publish; kind user, assurance verified, channel proxy; no cookie. Route handlers and Server Actions share one code path. - The unset-ACCESS_CODE boot warning is skipped in this mode; the gateway is the access gate and ACCESS_CODE stays independent. - The identity boundary scan now also fails on gateway identity headers (x-forwarded-user/-groups/-email, x-auth-request-*, remote-user, the secret header) and TRUSTED_PROXY_* read outside lib/server/identity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(auth): document accounts through an identity gateway Describe the trustedProxyHeader built-in in the README owner identity section (and its Chinese counterpart): configuration, request semantics, the trust boundary and why it is a shared secret, the interaction with ACCESS_CODE, boot-time validation, and an oauth2-proxy example that injects the user, groups and secret headers. Add the .env.example block, a SECURITY.md note and an Unreleased changelog entry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(auth): harden the trusted-proxy authenticator after review - Reserve the header names HTTP, Next.js and forwarding proxies set themselves: configuring the user, groups or secret header to rsc, next-action, forwarded, x-forwarded-for/-host/-proto/-port, x-real-ip, or anything starting with x-middleware-, x-invoke-, x-nextjs- or next-router- now fails the boot. - Log a one-time boot warning when TRUSTED_PROXY_ADMIN_GROUPS is set: the admin role is granted from the groups header, which is trusted on the strength of the shared secret alone, so the gateway must overwrite or strip it. - POST /api/chat/pi resolves the owner up front and answers 401 when the authenticator rejects the credential, like every other owner-resolving route, instead of continuing without an owner. The anonymous default always resolves, so it is unaffected. - The identity boundary scan also covers the Authentik, Azure App Service authentication, WebAuth, Cloudflare Access, Google IAP and AWS ALB OIDC header families. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(auth): admin-groups warning and a corrected oauth2-proxy example - Warn next to TRUSTED_PROXY_ADMIN_GROUPS (README, README-zh, .env.example) that the groups header is trusted on the strength of the shared secret alone, and list the reserved header names. - The oauth2-proxy example now states it targets v7.14+ (nested claimSource/secretSource), uses the canonical bindAddress, sets preserveRequestValue: false on all three headers and insecureSkipNonce: false, forwards the subject with `claim: user`, and notes that the secret variable holds the literal secret. It adds the legacy-option migration caveat, a warning not to exempt app routes via skip-auth or trusted-ip options, and notes that the IdP must release the groups claim and group names must not contain commas. - Changelog: the new boot checks and the /api/chat/pi 401. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(persistence): owner identity seam, phase 4 — host extension hooks * feat(persistence): owner identity seam, phase 4 — host extension hooks Let a host add product behavior at four points without forking a route, registered once from instrumentation.ts register() like the owner authenticator and sealed on first use. Nothing changes when no hook is registered. - configurePersistenceHooks({ name, authorizeCreate, onCreate, library, beforeAssetAllocate }): - authorizeCreate / onCreate run once per created course, inside the create transaction, in the transaction that inserts the course's ownership row. Saves of a course the owner already holds, and a create that loses a race to a concurrent create of the same id, are updates and run no create hook. A refusal answers 403 CREATE_REFUSED; a throw rolls the course, its ownership row and the host's writes back together. The hooks receive the request principal, or only the owner id for a background agent run. - library chooses which stage ids GET /api/stages lists; the route builds the usual summaries in the provider's order and drops every id the read path would refuse (deleted or unclaimed). folderId is shown only on the principal's own courses. - beforeAssetAllocate runs inside the storage handler's asset authorization step for POST /assets, after the owner is resolved and before the body is read; a returned Response is sent instead, with nothing stored and no quota counted. - configureAssetByteStore({ name, create, signsReadUrls }) replaces the ASSET_S3_BUCKET switch for both the persistence provider and the asset collector. A host store must keep bytes outside the registry database. Under ASSET_BYTE_EGRESS=redirect a store that does not declare signsReadUrls stops the server at boot; the built-in layers keep their behavior. - @openmaic/storage 0.33.0: DocumentWriteRefusedError lets a store refuse a write as policy; the document HTTP handler answers 403 with its code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): harden the host extension hooks after review Correctness: - Registration reads every hook once and stores it bound to the object the host passed, so hooks defined on a class prototype or as non-enumerable properties are kept. A spread copy dropped them silently, including the authorizeCreate and beforeAssetAllocate gates. Same for configureAssetByteStore. The unknown-key check now applies to plain objects only, since class instances carry their own state fields. - The owner-bound document store carries each call's operation in its own async context (AsyncLocalStorage) instead of a field on the instance. One store is shared by every tool of an agent run; a concurrent call could overwrite the field before the transaction read it, so a create was gated as a read (no ownership row, no create hooks) or claimed another stage. create_stage also runs sequentially within a tool batch. Hardening: - isDocumentWriteRefusedError recognizes a refusal from another copy of the storage package by name and shape (not by code alone). The document handler and POST /api/stages use it. - beforeAssetAllocate also gates PUT /assets/{id}/content. The request carries operation ('create' | 'replace') and the decoded assetId from the handler's own routing. - Agent runs: a refusal on a background write carries a fixed message, and create_stage answers it with a stable tool result. The host's message never reaches the model. - A byte store declaring signsReadUrls without a signer degrades to direct bytes with a warning instead of failing reads, and the collector no longer requires signing. - DocumentActor is discriminated by source ('request' with the principal, 'background' without one). - Library ids the read path cannot address are dropped, and a provider answer is capped at 5000 ids. - The boot-time ownership backfill logs how many courses it adopted without hooks. - The server persistence provider no longer exposes an unscoped document store. - The writesOutsideRegistryDatabase trust boundary and the scope of upload admission (HTTP uploads, not server-generated media) are documented. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): refuse misspelled hooks on host classes too A class instance skipped the unknown-key check, so a misspelled hook such as authorizeCreat registered cleanly and never ran, leaving creation ungated. Both configurePersistenceHooks and configureAssetByteStore now reject any function-valued property of a class instance (own or on its prototype chain up to Object.prototype, constructor excluded) that is not a known key, and any unknown key of a plain object. The error names the key and suggests the known key within edit distance 2. Host classes keep their helpers private (#helper) or register a plain object; the README says so. Also: - isDocumentWriteRefusedError requires a string message from a cross-copy candidate, and documents that the name is the effective discriminator. - CHANGELOG: ServerPersistenceProvider no longer exposes an unscoped documentStore; use the owner-bound store. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only * feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only Course ownership was recorded twice: in stage_meta, which already answers every access decision, and in document_stages.owner_id, which the storage package scoped by. Two tenancy records double what a later claim has to re-key, and a host that keeps ownership in its own table had to rewrite the package's SQL to strip the column. @openmaic/storage 0.34.0: - PgDocumentStore never reads or writes document_stages.owner_id. An owner-bound store scopes listings, writes, deletes, the freshness manifest and folder membership through the host's ownership relation, named with the new documentOwnership option ({ table, stageIdColumn, ownerIdColumn, tombstoneColumn, claimOnCreate }, or false for no document scoping). Identifiers are validated, never quoted into SQL from arbitrary input. - forOwner / ownerId without documentOwnership throws at construction, so a direct consumer cannot upgrade into a store that silently scopes nothing. An unbound store is tenant-agnostic. - claimOnCreate inserts the ownership row in the create transaction; a concurrent create of the same id by another owner rolls back. - folders: false lets a host whose document_stages has no folder_id use the store; folder ids are unique per owner, so membership goes through the ownership relation too. - AssetCollector reads each document's owner from documentOwnership for assetReferencePrincipals, and refuses the function without it. - Schema: fresh installs get no owner_id column and none of its indexes; an existing column is kept for one release, made nullable with its default dropped (each ALTER guarded by a catalog check, so a replay takes no lock), and its two indexes are dropped. document_stages_folder_idx replaces the owner/folder index. Application: - The owner-bound store binds the package to stage_meta (STAGE_META_OWNERSHIP) and lists through it in one query; the asset collector schedule passes the same relation. - The boot backfill copies column-only owners into stage_meta before any store is built, only when the legacy column exists, idempotently, and logs how many it adopted and how many disagree (stage_meta stands). Tests cover a fresh install and an upgrade from the previous schema on PGlite and PostgreSQL 16, a host table with neither ownership nor folder columns, folder isolation across owners sharing a folder id, and the concurrent-create race. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(persistence): serialize schema bootstrap; guard unscoped owner-bound stores Review follow-ups for the tenancy change. - Concurrent starts: two instances starting against one database raced on the schema DDL (a duplicate pg_class / pg_type entry, or "tuple concurrently updated"), failing one instance's first request. Every schema bootstrap the application runs -- the persistence provider's sequence (runtime, documents, stage_meta and its ownership backfill, owner materials, assets), the agent-session, session-material and user-skill stores, and the asset collector's -- now runs under one PostgreSQL advisory lock held on a dedicated connection and released in finally. It lives in the application because what needs serializing is the whole sequence, including tables the package does not own. A PostgreSQL test boots eight instances at once, four rounds each, on a fresh and on an upgraded database; without the lock it fails every run. - documentOwnership: false on an owner-bound PgDocumentStore now also requires allowCrossOwnerDocumentAccess: true, so binding an owner for folders or asset principals cannot silently expose every owner's documents. - Documented that an ownership relation must cascade with the document rows (a leftover row keeps the id reserved for its owner, which is also how a host keeps retired ids from being reused), with tests for both. Leftover rows are deliberately not made claimable: a host cannot be told apart from one that keeps them as retirement markers. - Upgrade notes: the one-time table locks on the first start, and the out-of-band CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY path that leaves the start with nothing to build or drop. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(persistence): leave the schema bootstrap lock out of the held-lock check The create-hooks suite asserts that its host lock is released with the create transaction by counting granted advisory locks database-wide. A suite booting a provider against the same database at that moment now holds the schema bootstrap lock, which is not the lock under test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in * feat(storage): owner merge primitives for claiming anonymous work (0.35.0) Add the per-store pieces a host needs to move one owner's data to another inside a single transaction of its own: - reassignDocumentFolders(tx, { fromOwnerId, toOwnerId, documentOwnership }) moves folders and filing: a same-named folder merges into the target's, a colliding id is renumbered, and filing is rewritten in one statement from the old ids. - PgAssetStore.reassignPrincipal(fromKey, toKey) moves a principal's entries under both principals' write locks, in key order; quota is not re-checked. - PgUserSkillStore.mergeOwner(from, to) moves skills, renaming a live handle the target already uses with the first free numeric suffix. - PgUserSkillStore takes a resolveFinalOwner hook, run as the create transaction's first statement, like the agent-session store's. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in A visitor who works anonymously and then signs in can move that work to their account in one transaction, and the anonymous id is retired afterwards. - Principal: pendingClaim { fromOwnerId, assurance }. The trusted-proxy built-in sets it when a gateway request also carries a valid anonymous cookie; the anonymous and shared-team built-ins never do. Authenticators may implement describeStoredOwner, canonicalize and clearPendingClaim; principalFromStoredOwner describes a stored id for work without a request. - Claims: claimOwner / claimPendingOwner and registerClaimParticipant (sealed on first claim). One transaction, both owners' identity locks taken exclusively in key order, participants in a fixed order (folders, courses, materials, agent sessions, skills, runtime, assets), recorded in the new owner_merges table. Idempotent per pair; refuses a non-anonymous source, an anonymous target, a source already claimed elsewhere, and chains. Same-named folders merge, colliding folder ids are renumbered, colliding skill handles get a suffix, and quotas are not applied to what moves. - Identity lock: every owner write (documents, folders, asset allocations, material registration, agent sessions, skills) takes the owner's advisory lock in shared mode first, so a write racing a claim is moved or refused, never orphaned or deadlocked. - Retired ids: request writes are refused with 403 OWNER_RETIRED; agent runs and media generation that started before the claim follow the id to the account (forwarding document store, canonicalized stage probes, forwarded generated assets and skills). - Trigger: POST /api/identity/claim (same-origin JSON only; clears the anonymous cookie), or OWNER_CLAIM_TRIGGER=auto on the first request carrying a pending claim. The runtime contract's learner merge now performs the same claim, allowed only for the request's own pending claim. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(persistence): count only the suite's own advisory locks after concurrent creates The check that a host's transaction-scoped lock is released counted every granted advisory lock in the database. Suites that run against the same database at the same time hold advisory locks of their own (a schema bootstrap, the owner identity locks a claim takes), which made it fail depending on timing. Count only locks held by this suite's connections. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(auth): use a neutral example table for host claim participants Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(storage): owner merge fixes and store policy errors over HTTP (review round 1) - PgUserSkillStore.mergeOwner reserves every handle either owner holds before renaming: a claimed skill whose handle the target uses could be renamed into a handle another claimed skill still held, which violated the per-owner unique index and made that merge fail every time. - PgUserSkillStore.list orders ties by id. - PgRuntimeStore takes resolveFinalLearner(tx, learnerKey), run as the create transaction's first statement, so a host can fence session creates against its owner merges; reassignLearner(from, to) re-keys sessions without re-validating them, so one session a newer version wrote cannot block a merge. - Every HTTP handler (runtime, documents, assets) answers a DocumentWriteRefusedError as 403 with its code, and the new StorageBusyError (code, message, retryAfterSeconds) as 503 with Retry-After. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(auth): claims review round 1 — fenced runtime creates, bounded locks, retired-cookie recovery - Runtime session creates take the owner's identity lock as the create transaction's first statement and refuse a retired owner, closing the window in which a create racing a claim was left under the retired id. - Lock waits are bounded: a claim waits OWNER_CLAIM_LOCK_WAIT_MS (default 5 s) for its identity locks, a write OWNER_WRITE_LOCK_WAIT_MS (default 30 s); running out, or losing a deadlock or serialization race, answers 503 OWNER_BUSY with Retry-After on the claim routes and on every fenced write. - Every 403 OWNER_RETIRED carries the Set-Cookie values that drop the retired anonymous credential (the anonymous built-in now implements clearPendingClaim), so a browser whose claim response was lost recovers. Writes by id to rows that moved (skill delete, session messages) answer OWNER_RETIRED instead of not-found or a bare 403. - OWNER_CLAIM_TRIGGER=auto no longer claims ahead of the explicit claim routes, which now report their own claim instead of a false refusal. - The owner-events stream tells a retired owner only that it moved, never the claiming account's id. - Creates of one course id take turns (a per-id advisory lock before the ownership probe), so a concurrent second create saves as an update instead of failing as reserved-document. - A claim's source must be described as anonymous by the authenticator; the write fences skip the retirement lookup for every other owner. The host canonicalize hook is removed: forwarding is core's owner_merges alone. - The folders participant locks the source's stage_meta rows first; runtime sessions are re-keyed without re-validation; the owner asset store fences every allocation and refuses to allocate without transactions; background skill reads and patches follow a mid-run claim; the forwarding document store forwards methods only. - Docs: when a pending claim can arise with the shipped authenticators, the anonymous cookie as a bearer credential, bounded waits, the exact scope of OWNER_RETIRED, and claims narrowed to what the tests show. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(auth): claims review round 2 — merge-record guard, boot-checked lock waits, stream credential drop - owner_merges records claims of anonymous owners only, since the write fences enforce retirement only for ids the authenticator describes as anonymous. Reading a row that retires any other owner now fails loudly instead of being followed but not enforced. The comment that pointed hosts at claimOwner for their own account merges is corrected: a host merging two signed-in accounts moves the rows itself and refuses the merged-away account in its authenticator. describeStoredOwner must classify ids stably. - OWNER_WRITE_LOCK_WAIT_MS and OWNER_CLAIM_LOCK_WAIT_MS are validated at boot with the same parser the lock paths use, so a malformed value stops the server instead of failing every write. - The owner-events stream answers a connect by an already retired identity with one owner_moved event and the Set-Cookie values that drop the retired credential, so a stale tab's reconnect loop ends after one round trip. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * chore(auth): final polish before landing the owner identity seam * chore(auth): final polish before landing the owner identity seam - Pi routes refuse a rejected credential with the seam's INVALID_CREDENTIAL body. - Drop the unused LegacyAssetMutations alias and resolveLockWaitMs re-export. - Scope the document_stages.owner_id deprecation wording: the store no longer reads or writes it; the boot backfill reads it and a claim mirrors ownership into it for rollback. - Document OWNER_CLAIM_TRIGGER and the two lock-wait variables in .env.example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * chore(storage): 0.35.1 for the README wording fix Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> * feat(identity): one identity entry point with composable auth methods * feat(identity): compose owner auth methods behind one entry point Owner resolution now asks an ordered list of owner auth methods instead of one configured authenticator. Each method answers `authenticated`, `not-applicable` (no credential of its kind) or `invalid` (its credential is present but bad). Core keeps the first `authenticated` answer, refuses any `invalid` answer with 401 INVALID_CREDENTIAL without asking later methods, and when no method applies falls back to the anonymous cookie owner, or answers 401 when the host turned the fallback off. Route handlers and Server Actions share the same order and per-request memoization. - configureOwnerAuthentication({ methods, anonymousFallback? }) replaces configureOwnerAuthenticator: single-shot, sealed on first use, and validated at registration and at boot. - anonymousCookie becomes the built-in fallback method; sharedTeam becomes a method, selected by PERSISTENCE_SHARED_OWNER_ID as before when nothing is registered, and kept beside host methods only when the host lists sharedTeamAuthMethod() last. The variable set beside a registration that leaves it out fails the boot. - Core attaches pendingClaim itself: a host method's non-anonymous principal on a request that also carries a valid anonymous cookie. A method may not set one. Claim cookie clearing and OWNER_RETIRED recovery go through the anonymous method's clearCredential. - principalFromStoredOwner consults the anonymous method first, then the configured methods' describeStoredOwner in order. - Remove the built-in trusted-proxy header authenticator with OWNER_AUTHENTICATOR and every TRUSTED_PROXY_* variable. The boundary test now also keeps the resolution internals inside lib/server/identity/ and asserts no source reads the removed variables. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(identity): document owner auth methods and a gateway JWT recipe Rewrite the owner identity section around the ordered auth methods: the three answers and what resolution does with each, the anonymous fallback, registration and its boot checks, and the sharedTeam rules. Replace the removed built-in gateway authenticator's documentation with a short host-owned recipe that verifies a gateway-forwarded signed JWT (oauth2-proxy, Cloudflare Access, IAP) against the identity provider's keys with the jose library. Update the claim section for the core-owned claim candidate, and the Chinese README, SECURITY.md, CHANGELOG and the deployment docs in every locale to match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(identity): tighten retired-owner cookies, retired config and the boundary - On 403 OWNER_RETIRED, core now calls clearCredential only on the anonymous fallback and on methods that declare `issuesAnonymousOwners: true`, so an account method's session cookie is never cleared there. A clearCredential that throws is logged and skipped instead of turning the refusal into a 500. - Setting any retired OWNER_AUTHENTICATOR / TRUSTED_PROXY_* variable fails the boot with a pointer to the auth-methods docs, instead of silently serving every request anonymously. Only the names are matched, in one module the boundary test checks never reads a value. - Add lib/server/identity/host/ for host auth methods. The boundary test now scans core identity files for gateway identity headers too, polices reads of an incoming Authorization header everywhere but the host directory, and keeps resolution internals out of host code. - Document on OwnerAuthMethodResult that a present but unusable credential must answer `invalid`, never `not-applicable`, and that a method that cannot decide throws. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * docs(identity): fix the gateway JWT recipe's error split and add notes - The recipe now answers `invalid` only for token failures (expired, claim validation, malformed JWS/JWT, bad signature, no or several matching keys, disallowed or unsupported algorithm) and rethrows everything else: jose reports a key endpoint that is down, answers non-200 or returns unparsable data as a generic JOSEError, which the previous version turned into a 401 for every user. - The recipe file goes in lib/server/identity/host/, with an accurate description of what the boundary test enforces. - Document invalid versus not-applicable for malformed credentials, issuesAnonymousOwners, the boot failure on retired variables, and that the claim candidate is an unsigned bearer cookie a sibling subdomain can shadow (serve on a registrable domain of its own; cookie hardening is deferred). Mirror the notes in the Chinese README, SECURITY.md and the changelog. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> |
||
|
|
56322a5e06 |
fix(generation): reject unusable interactive scripts and surface runtime errors (#1649)
Reject classic inline interactive scripts that fail to parse at generation time (extracted with parse5, checked with node:vm Script without executing), and surface iframe runtime errors on the active interactive scene. Addresses #1622 (partial: no recovery / regeneration UI). Co-authored-by: Frank-zhu0404 <Frank-zhu0404@users.noreply.github.com> |
||
|
|
d04a75cbc6 |
feat(ai): add Xiaomi MiMo V2.6 models [AI-assisted] (#1655)
* feat(ai): add Xiaomi MiMo V2.6 models * docs(ai): preserve Xiaomi provider references --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
fbc51fcbd6 |
fix(generation): reject invalid quiz option shapes and constrain values at the source (Closes #1375) (#1651)
* fix(generation): enforce quiz option value/label contract When a model puts choice content in value and a bare A-Z letter in label, swap those fields so persisted quiz options keep a letter as the selection identity and the content as the label. Correct objects, plain strings, and answer-key normalization stay as they were. Closes #1375 * fix(generation): rewrite swapped quiz letter keys to option values When a bare label is moved onto the option value, a raw answer token that exactly equals that original label is rewritten to the uppercased letter before exact alignment. The stored key is then the letter QuizView submits, so that submission grades correct. A stored "a" on an already-correct value "A" is left unresolved. Closes #1375 * fix(generation): reject invalid quiz option shapes instead of swapping Stop rewriting quiz options when a model puts content in value and a letter in label. The quiz prompt now requires value to be one ASCII letter A-Z and label to be the option content. A choice question that still breaks that contract, or whose answer does not name an option value, is invalid model output and regenerates. Closes #1375 --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
283a3096aa |
fix(editor): share line bounds for alignment and dragging (#1631)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
7be21f2d89 |
fix(chat): pin Pi routing and preserve provider errors (#1637)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
7e8a065ecb |
fix(importer): bound pptx zip inflation by default (#1588)
ZipParseLimits was optional field by field and both import paths called parseZip with no argument, so every bound the parser already implemented was inert: a .pptx could inflate without limit. - apply the defaults per field, so a caller overriding one bound no longer silently drops the others, and Number.POSITIVE_INFINITY switches a single bound off - add maxCompressionRatio for text parts, defaulting to 200:1. It applies to slides, layouts, themes and chart XML only. Uncompressed bitmaps and silent PCM are ordinary deck content sitting at the top of DEFLATE's range — a solid-colour 24-bit BMP measures ~1027:1, 30s of silent stereo PCM ~1016:1, and a 1920x1080 white screenshot with a grid ~297:1 — so a fixed ratio there rejects real decks. Media and embedded objects are bounded by maxEntryUncompressedBytes and maxMediaBytes instead, which for an honest archive are checked before anything is inflated. Text parts have no such problem: real decks peak around 19:1, so 200 leaves ample headroom. - reject a supplied NaN and a non-integer maxEntries rather than letting them compare false against every bound and disable it silently - export the defaults and the limit type so callers can extend them - add the first tests for this parser: the regression itself (a call with no limits must still apply the bounds), the per-field defaulting invariant, each bound pinned, the binary-part exemption pinned with a solid-colour BMP, a silent WAV and an embedded object, and the repo's own regression deck The bounds read the sizes the archive declares about itself, which an archive is free to understate. One that declares small and inflates large passes all of them and is stopped only by JSZip's own size check, after that entry has been inflated — peak RSS in the reviewed case was ~1.19 GB, and with maxConcurrency 8 several such entries inflate in parallel. Bounding the allocation itself needs a byte-budgeted inflate inside the read path, which is a larger change than activating the bounds that already existed here. Version to 0.3.0: ^0.2.6 admits 0.2.7, and this changes default behaviour for existing callers. Consumers pinned to ^0.2.x need to move deliberately. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
2a77a8a476 |
feat(chat): make Pi classroom runtime the default (#1628)
* feat(chat): make Pi classroom runtime the default * chore(chat): align docs and E2E with Pi default |
||
|
|
df16d7e322 |
fix(export): include active line geometry in bounds (#1626)
Share corrected line bounds across renderer and React editing paths, preserve double-elbow routing, and translate PPTX points into the shape bounding box. Refs #674 and the prior implementation/review in #675. AI-assisted. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
d7b31aa5e7 |
fix(generation): add browser-safe package entry (#1609)
Co-authored-by: SY <sy@SYdeMacBook-Pro.local> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
8de7ec2870 |
fix(importer): include all color attributes in style cache key (#702) (#1571)
Co-authored-by: cham <2577781125@qq.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
4d2e2bab82 |
fix(generation): keep narration speech TTS-readable — no formulas or LaTeX (#1586)
* fix(generation): keep narration speech TTS-readable — no formulas or LaTeX
Narration speech authored by the four *-actions prompts could carry raw
formula notation (a^2 x / y) or LaTeX (\frac{a}{b}), which TTS engines
read out as gibberish. The prompts had no TTS readability rules, and the
element list fed to them includes raw LaTeX via `Formula:` entries,
inviting verbatim copying into speech.
Add a shared `speech-tts-readability` snippet and compose it into the
speech sections of slide-actions, quiz-actions, interactive-actions, and
pbl-actions via the existing {{snippet:...}} mechanism. The name
references speech/TTS so it cannot be confused with slide-content's
visual LaTeX rules or reused there.
Refs #1585
Co-Authored-By: Claude Code <noreply@anthropic.com>
* chore(generation): bump to 0.3.10, main already released 0.3.9
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DYifP8wM4XJQ3Hc2qsF6zf
---------
Co-authored-by: Claude Code <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
|
||
|
|
f29bbc4daa |
feat: sample declared interactive state before classroom questions (#1508)
* feat: sample declared interactive state before classroom questions * test: exercise generated publication example through the iframe reader * chore(generation): bump package version for observation prompt contract * fix: keep interactive state optional on insecure HTTP origins * fix(playback): keep component picking and separate state from reference A declared state interface replaced the component picker with a forced whole-area `#experiment` reference, removing the per-component selection and outline that `main` already ships. Sampling was also gated on the reference selector, so the only way to obtain state was to give up the selection. Reference identity and area state are now independent request-scoped evidence items: - `handleToggleElementPick` always arms the picker again, so a scene that declares the interface keeps main's per-component selection, outline, and send-time clearing. - `sampleInteractiveState` follows the current Scene instead of the draft reference, so an unreferenced follow-up still reports current facts and never re-creates or extends a reference. - The Host carries area state with or without a component reference. The evidence header names both identities and refuses to present area facts as properties of the referenced component. - `metadata` is absent when only area state travels, so no element identity and no Spotlight authorization can be derived from it, and the accepted-reference receipt stays driven by explicit references only. Review follow-ups in the same change: - Client sampling follows `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED`. An ungated packet turned an ordinary Pi question into a 400 while the reference feature was disabled. - A Scene that declares the interface always receives an availability boundary, including when the browser produced no packet at all. It is reported as `not-sampled` rather than the previous `no-interface`, which was a false statement about an activity that does declare one. Courseware without the interface keeps its unreferenced behaviour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(generation): register the observation snippet as a packaged asset The interactive-observation snippet is referenced by all six widget content templates but was never added to the packaged-asset manifest, so the asset test and the golden scene prompt both failed. - `SNIPPET_IDS` now lists `interactive-observation`, restoring both the "exactly the generation-owned templates and referenced snippets" check and the "every referenced snippet is packaged" cross-check. - The interactive system-prompt snapshot is re-pinned. The change is purely additive: the snippet is appended to the simulation template. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): stop injecting state constraints while the feature is disabled The route rejects a request that carries a reference or a state packet while `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED` is off, but an ordinary question carries neither. It still reached the Host, and a Scene that declares the state interface then received the full page-state block — several kilobytes of constraints about evidence the deployment can never sample. The Host now returns before building that note when the feature is off. A route-level regression asserts that neither the Director prompt nor the Child prompt gains `PAGE-REPORTED STATE` in that configuration; disabling the guard makes it fail with exactly that symptom. Also reopens the composer before the unreferenced follow-up in the classroom browser spec. An accepted answer may close it, which made the assertion flaky without changing the behaviour under test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): decouple reference Scene from state freshness and bound assembled evidence Cross-review found two defects in the request-scoped state evidence. The reference's Scene was folded into the sample's staleness test. The packet is already bound to the current Scene by the identity check above it, so a valid current-Scene sample was being discarded as `stale-sample` purely because the student's component reference came from an earlier Scene. Reference and area state are independent evidence items; freshness is a property of the sample alone. With the coupling gone the two can now disagree on Scene, so the note says so explicitly rather than letting the model attribute area facts to a component that may not be on the current Scene. The assembled evidence had no stated output budget. The static component packet is bounded to 24,000 code points upstream, but that bound covers the static packet alone; the note and the escaped observation JSON were appended without a recheck. Escaping `<` for the prompt expands one code point into six, and `<` is legal in a label or a fact value, so a packet the Host accepts could assemble to 149,385 code points. The budget is now declared as the static bound plus the room the note frame needs, which is what makes the degradation terminate. Over budget, the state body drops whole to an explicit `unavailable` statement: truncating the JSON would emit a broken packet, and thinning a `complete` relation set would turn an exhaustive set into a false one. The relationship summary degrades with it, so the prose never asserts COMPLETE over a body that is gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): keep slide references independent of activity state * refactor(interactive): simplify declared state and unify iframe preparation Accept any JSON report within byte and depth budgets, without generated field requirements or relationship-completeness semantics. Keep publishState and an optional rendered result in the generation guidance. Prepare the observation responder through patchHtmlForIframe and let the pool own document identity, preserving state across placeholder remounts. Settle sampling failures locally and align browser/server nesting limits. Cover permissive JSON delivery, resource limits, lifecycle, legacy behavior, and real renderer remounts with focused regression tests. * fix(generation): publish automatic activity changes with clear positions * fix(interactive): report missing legacy scope as no interface * fix(interactive): guard sampling capabilities and bind scopes lazily * chore(generation): bump version after main integration --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f70dc67a6a |
fix(importer, editor): preserve PPTX diagonal corners, table typography, editing alignment and punctuation wrapping (#1581)
* fix: preserve imported PPTX table typography and rounded geometry * fix(importer): isolate table punctuation width from tab clamping * fix(importer): preserve punctuation markers with default cell margins --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
f6e350a1dd |
fix(storage): sanitize agent session descriptive text (#1504)
Co-authored-by: Bryan Nathan <bryan@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
d695193ae6 | fix(importer): preserve shape autofit sizes and Wingdings checkmarks (#1577) | ||
|
|
8517121ec0 |
fix(docker): create /app/data with runtime-user ownership so classroom persistence works (#1442)
* fix(docker): create /app/data owned by the runtime user before dropping privileges * docs(deployment): note one-time ownership repair for pre-existing data volumes --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
8c22e27a29 |
fix(generation): tolerate non-array mediaGenerations in generated outlines (#1469)
* fix(generation): tolerate non-array mediaGenerations in generated outlines LLM-generated outline JSON can carry mediaGenerations as a string or object instead of an array. Every downstream consumer (media orchestrator, video manifest, scene generator) requires an array, and uniquifyMediaElementIds called .map() after only a falsy check, so one malformed outline aborted course generation with "outline.mediaGenerations.map is not a function". - sanitize at parse time: drop non-array mediaGenerations from each enriched outline (outline-generator) - harden uniquifyMediaElementIds: treat non-array values as absent, strip them, and open the early-return guard on field presence so all-malformed outlines still get sanitized - add unit tests covering array, string, object, number, and missing shapes * chore(generation): bump package version to 0.3.8 Required by the package version bump check: the PR changes the publishable @openmaic/generation package sources. * test(generation): use correct SceneOutline fixture fields in outline media tests The first commit of this branch accidentally staged an earlier draft of the test file (sceneType instead of type, stale id regex). Re-stage the final version that passes tsc and the full suite locally. --------- Co-authored-by: PassCode023 <269712126+PassCode023@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
f7b8769e7f |
fix(importer, renderer): preserve PPTX text, chart, and image styling (#1534)
* fix: preserve PPTX text, chart, and image styling * fix(importer): preserve hyperlink semantics and remove inferred alignment * test(charts): replace untyped assertions to satisfy CI * fix(importer): honor source-linked chart axis number formats * fix(renderer): preserve picture bar clustering and clipping * fix(pptx): address chart scales and soft-edge review findings * fix(pptx): preserve percent-stack labels and inherited text sizes * fix(pptx): synchronize agent chart schema and axis defaults |
||
|
|
35a8be5956 |
fix(pptx): preserve tab columns, text insets, and arrow rendering (#1518)
* fix(pptx): preserve tab columns, text insets, and arrow rendering * fix(pptx): address tab layout review and bump package versions * fix(pptx): preserve editable tab columns and final font metrics * fix(editor): apply list commands inside tab columns * fix(editor): preserve table paragraph spacing while editing * fix(importer): preserve saved leading in auto-fit text labels * fix(importer): preserve ordinary symbol-font text and editable default tabs * fix(importer): preserve default hyperlink underline * fix(editor): preserve Latin baselines when entering text editing * fix(importer): approximate verified clear material front-face color * fix(importer): preserve filled flowchart connector shapes * fix(editor): preserve inline formulas when editing imported text * fix(editor): keep formula caret separators inline * fix(test): narrow serialized shape before checking inverse path * fix(importer): preserve equation system delimiters * fix(importer): preserve compatibility tables and cell formulas * fix(editor): handle formatting and list splits inside inline containers * fix(editor): preserve script sizing around inline containers * fix(editor): preserve inline typography across editing and copy * fix(editor): preserve destination and nested typography contexts * fix(pptx): preserve table tabs and editor clipboard typography * fix(editor): preserve container font context when clearing formatting * fix(importer): correct Wingdings 3 upper-right triangle mapping * fix(pptx): preserve explicit text inset markers with legacy fallback * fix(editor): guard formula serialization and document layout limits * fix(editor): preserve inline font contexts through undo --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
ee79762ef9 |
fix(access-code): expire tokens server-side and throttle verification behind a trusted proxy (#1513)
When ACCESS_CODE is set, verification tokens (timestamp.HMAC) never expired because neither verifier checked the timestamp, and POST /api/access-code/verify had no attempt throttling. - Enforce a 7-day token lifetime in a shared, Edge-safe module used by both the Node verifier and the middleware Web-Crypto verifier, and reject non-canonical signatures in both. - Rate-limit verification only when the client identity is trusted (TRUST_PROXY_HEADERS=true): a per-client sliding window (10 failures / 60s) whose attempt is reserved atomically at check time, returning 429 with Retry-After. Without a trusted proxy the app cannot attribute requests to a client, so no shared throttle is applied — a long random ACCESS_CODE is the protection, and a warning is logged when it is short. - Store bounded, copied identity keys so a large forwarding header cannot retain memory. - Document the behavior in .env.example, README, and configuration docs. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
4d0f88b4c9 |
fix: bump next to 16.3.3, patches GHSA-p293-qw3h-jr36 (#1503)
GHSA-p293-qw3h-jr36 (CVE-2026-75604, critical): unauthenticated RCE on Windows-hosted Next.js servers. Root was on 16.2.11 and packages/docs on 16.2.6 (affected: >=16.0 <16.3.3). Both lockfiles regenerated. eslint-config-next lint plugins left as-is (not the affected runtime package). Co-authored-by: Aniruddha Adak <aniruddhaadak80@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
e824117c9c |
feat(storage): give the server the asset entry lifecycle (#1007 amendment, part 1) (#1472)
@openmaic/storage 0.31.0: a document reference table maintained by the document store, pending -> committed allocations, an entry pass with a one-time bounded backfill and legacy mark, a standing bounded sweep for entries nothing references, live-entry quota, and single-sequence ascending entry locks for every reference-maintaining write. |
||
|
|
3edaa4995f |
fix(importer): resolve embedded videos and upload posters in 0.2.1 (#1507)
* fix(importer): resolve embedded media and upload video posters * fix(importer): use unique filenames for concurrent poster uploads |
||
|
|
ee7a7b64df |
fix(agent-runtime): allow one owner material to be bound to multiple sessions (#1500)
* fix(agent-runtime): allow one owner material to be bound to multiple sessions `agent_session_materials.id` is a global primary key, but `bindOwnerMaterialsToSession` de-duplicated on `(sessionId, id)` and then inserted the owner-side material id as the row id. Binding the same owner material to a second session therefore raised a duplicate-key error (23505) that the route surfaced as HTTP 500, so an owner material could only ever be used by one course. Keep row ids globally unique (every extraction method keys on `id` alone) and store the shared owner id as metadata instead: the binder now mints a fresh session row id, records `owner_material_id`, and looks that column up for idempotent rebinding; a partial unique index on `(session_id, owner_material_id)` adjudicates concurrent rebinds and the loser adopts the winner's row. The schema change is additive and idempotent (`ADD COLUMN IF NOT EXISTS` + `CREATE UNIQUE INDEX IF NOT EXISTS`), so existing databases upgrade in place with `NULL` for legacy rows. No FK from `owner_material_id` to the owner library, to keep owner-library deletion decoupled from session rows. Tests: backend-neutral contract test (same owner material bound to two sessions, both readable), PGlite in-place upgrade test, pinned-schema update, and a host-level test covering the route's 202 path. * fix(agent-runtime): reuse legacy owner-material bindings and clean up losing uploads on concurrent bind Rows written by the previous binder use the owner upload id as the session row id and leave `owner_material_id` NULL, so the new owner-id lookup missed them and a rebind created a duplicate row and copied the bytes again. The binder now falls back to the session row id and adopts it only when it is unambiguously that owner upload: source kind, matching title, no source URL or derivative, no text, and the deterministic legacy object key with a byte length that matches the owner record. Adoption stamps `owner_material_id` through a conditional `backfillOwnerMaterialId` update that leaves extraction columns untouched; a lost race re-reads the fast path. A different session is still a fresh row. The concurrent-bind loser also left its uploaded object behind after adopting the winner. It now removes that object when the winner references a different key, best-effort so a failed cleanup cannot fail a bind that succeeded. Tests: host-level PGlite regressions for the legacy rebind (id, row count, byte copies, extraction state, and backfill) plus a different-session bind, and a controllable byte store that parks the loser after its upload so the winner commits first, asserting one row and only the winner's object remain. A storage unit test pins the backfill contract. Bumps @openmaic/storage to 0.30.1 for the new public store method. |
||
|
|
cfae106b89 |
fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates (#1499)
* fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates PostgreSQL jsonb rejects the `\u0000` and lone UTF-16 surrogate escape sequences that JSON.stringify emits verbatim, failing the enclosing statement with SQLSTATE 22P05/22P02. In the agent-session store the failed tree-entry/event write is treated as critical, so the whole run aborts and all work in that run is lost. Add a shared encodeJson at @openmaic/storage/src/pg-json.ts that replaces U+0000 and unpaired surrogates with U+FFFD while preserving valid surrogate pairs (emoji), sanitizing in-memory values and object keys before JSON.stringify. Route every jsonb parameter through it: the agent-session, document, runtime, asset, and material backends, plus owner-material registration in lib/persistence. The document/runtime/asset backends keep their existing assertJsonValue guard, which rejects these code points with a readable error before the shared serializer runs; the agent-session and material paths had no guard and were the live 22P05 exposure. Literal backslash-u text is untouched. * chore(storage): bump @openmaic/storage to 0.29.2 * fix(storage): keep colliding sanitized keys and own __proto__ members when encoding jsonb encodeJson replaces NUL and lone surrogates in object keys, but two distinct keys can sanitize to the same string: "a\u0000" and "a\uFFFD" both become "a\uFFFD". The rebuild assigned each member by its sanitized key, so a later member silently overwrote an earlier one and that earlier value was lost. Keep the first member under the sanitized key and give each later colliding member a deterministic suffix (#2, #3, ... appended until the key is unique among the keys emitted so far, whether sanitized or original). Member order is preserved. The suffix is ASCII and contains no NUL or surrogate, and the sanitizer only ever emits U+FFFD, so a suffixed key can never be confused with a bare replacement result. The rebuild also used a plain {} object, so assigning an own "__proto__" member invoked the prototype setter and dropped it from the output whenever a sibling key needed sanitizing. Build the rebuilt object with a null prototype so "__proto__" stays an own data property; JSON.stringify then emits it and PostgreSQL round-trips it. Arrays and the no-op fast path (return the original reference when nothing needs sanitizing) are unchanged. |
||
|
|
ebf665f316 |
feat(playback): reference static GenUI components (#1281)
* feat(playback): reference static GenUI components * feat(playback): gate courseware references * fix(pi): bound Interactive source HTML before parsing * fix(pi): bound interactive reference HTML processing * fix(playback): close interactive reference review gaps --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
85ef52d599 |
feat(media): write generated media through the asset pool under server-backed persistence (#1392)
* feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> |
||
|
|
0b641a9403 |
fix(importer): convert Equation.3 OLE formulas via MTEF v3, surface degrade telemetry (#1411)
* fix(importer): convert Equation.3 OLE formulas via MTEF v3, surface degrade telemetry Legacy courseware stores formulas as Equation 3.0 / MathType OLE objects whose only renderable form inside the .pptx is a WMF preview picture. The importer cannot rasterize WMF, so those formulas degraded to a hardcoded 1x1 "transparent" placeholder — actually a 50%-alpha red pixel that rendered as a pink block once stretched over the formula frame, with no way for callers to notice the content loss. - Convert `Equation.3` / `MathType` OLE objects to LaTeX: detect the progId, resolve the embedding, parse the OLE compound file (cfb) and convert its `Equation Native` MTEF v3 stream (new utils/mtef.ts) — fractions, radicals, scripts, fences, big operators with per-family limit variations, embellishments, and Symbol-font local character encodings. Slot order follows the rtf2latex2e reference implementation and real MathType streams ([main, lower, upper]), not the archived spec prose. Any failure falls back to the picture path. - Fix the placeholder constant to a truly transparent pixel; keep recognizing the legacy red one (exported isPlaceholderDataUrl). - Surface degradation instead of failing silently: optional ImportPptxOptions.onWarning receives machine-coded warnings (media-unconvertible / formula-fallback-image / formula-degraded / element-dropped) at every placeholder consumption point (image, background, shape/text pattern fills, math fallback); a throwing sink is isolated so telemetry can never fail an import. A formula whose fallback picture is also a placeholder keeps its plain text as a text element. * fix(importer): address review — LSCRIPT base duplication, one-sided fences, Symbol table, depth caps Review round on #1411 (thanks @wyuc — the fuzz safety-net validation and the adversarial constructions found what the spec-conformance rounds could not): - tmLSCRIPT: stop re-emitting a script slot as the base group (isotopes rendered as {}_{6}^{12}{12}C and the output doubled per nesting level — a 221-byte stream could OOM the worker); the base is the following sibling, so the template emits only {}_{sub}^{sup}, and a leading script no longer steals the previous atom via the trailing-script lookahead. tvLSUPER writers that emit a single slot now fall back to it. - One-sided fences: render the missing side as a null delimiter (\left. / \right.) instead of an unmatched \left — piecewise-function braces (tmBRACE var 1) produced KaTeX-invalid output that silently degraded to flat text. - SYMBOL_FONT_LATEX corrected against URW StandardSymbolsPS AFM + Adobe AGL: 0x3C/0x3E/0x5B/0x5D are literal < > [ ] (≤/≥ live at 0xA3/0xB3, now mapped, along with the rest of the Symbol operator block); 0x22/0x24/0x5C are ∀/∃/∴; added Chi/vartheta/varsigma, the phi/varphi split (0x66/0x6A), and 0x5E = \perp. Unmapped font-local bytes >= 0xA0 now flag degraded instead of passing silently. - tmLIM: variation roles were inverted — spec + rtf2latex2e eqn.c say 0 = upper limit, 1 = lower limit; single-limit writers keep their limit via the same slot fallback as the big operators. - Hardening: MAX_DEPTH = 200 nesting cap and a 64 KiB LaTeX output cap, both throwing MtefParseError (deep bombs now fail loud instead of RangeError/OOM); video poster joined the placeholder warning points. - Tests: +6 — three REAL Equation Native stream fixtures (round-tripped from a legacy deck, pinning the font-local encoding and big-op slot order), tvLSUPER single-slot, tmLIM both roles, script-nesting perf guard; fence tests now assert KaTeX renderability instead of pinning broken strings. Suite: 85. - Version 0.1.5 -> 0.2.0 (new public option/type/export + Math.degraded). Lockfile re-anchored on main with pnpm@10.28.0: only the cfb additions remain, no unrelated churn. * fix(importer): address review round 2 — tmLIM function slot, full Symbol high-half, bra/ket, cap tests - tmLIM: emit the main slot FIRST followed by the limits and inject no operator name (the reference `39.1 = limit: lower, #1 #2` puts the function in the main slot — hardcoding \lim duplicated it and glued to letter-leading main slots, crashing KaTeX with an undefined control sequence). An empty main slot falls back to \lim as a neutral base. - SYMBOL_FONT_LATEX: completed the 0xA0–0xFF block from the URW AFM (~70 positions: ∫ ∑ ∏ ⟨⟩ ∂ ∇ ⇒ ⇔ ⋅ ′ ∅ ⊆ ⊇ ∈ ∉ ∪ ∩ …). Unmapped font-local codes now throw MtefParseError instead of passing through as Latin-1 (0xF7 was an integral extender rendering as ÷ — plausible but wrong math); radicalex (0x60) and C1 controls (0x80–0x9F) also throw, taking the picture fallback. - tmDIRAC: var1 renders a bra `\left\langle L\right|`, var2 a ket `\left| R\right\rangle` per the reference (was wrapping both sides). - Embellishment records now count against the record budget (a 10 MB embellishment-only stream no longer allocates 1 GB before the output cap fires). - Tests: +6 — every cap now has its own assertion (output-length width case at 15k sibling CHARs, record-count at 20k, in addition to the existing depth test), previously-dangerous Symbol codes verified from the AFM, unmapped-code rejection, embellishment budget, tmDIRAC KaTeX-validity for all variations. Fixture claims corrected (no big-op selector in the three real streams; the reading is pinned by the spec-conformance tests). Suite: 91. * test(importer): pin 0xD6/0xF3 Symbol glyphs through KaTeX, fix big-op test title Follow-up to the round-3 review: the two glyph fixes landed without a test, the BigOp test title still named the wrong slot order, and the table comment claimed the high half was complete while unlisted positions intentionally throw. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * docs(importer): correct Symbol table comments — 0xF7 is parenrightex not an integral extender, high-half coverage is ~50 of ~70 positions --------- Co-authored-by: Percy <percy@PercydeMacBook-Pro.local> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
d32c3f2a00 |
fix(importer): stop PPTX import from hanging in non-browser environments (#1424)
* fix(importer): stop PPTX import from hanging in non-browser environments * format --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
28133f8a67 |
fix(renderer): lazy-load optional ECharts runtime (#1420)
* fix(renderer): lazy-load optional ECharts runtime * test(chart): assert setOption is called after runtime is ready --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
9fe650d994 |
fix(quiz): resolve AI answer keys to option values before grading (#1328)
* fix(quiz): resolve AI answer keys to option values before grading
AI-generated quiz answer keys are unreliable: the model may write the
correct answer as option CONTENT ("(6, 2)"), as a LETTER ("A"), or as a
formatting variant (full-width chars, extra spaces) of either. Grading
compares exact option values, so a learner picking the correct option
was still graded incorrect whenever the key was written as content or
in a different format.
- generation: normalizeQuizAnswer now resolves every answer variant to
the matching option value (NFKC + whitespace-stripped + letter-prefix
canonical matching against option values and labels); unknown answers
pass through untouched
- grading: resolveAnswerKeyToValue retro-fixes already-generated courses
whose stored keys hold content variants, mapping them to option values
when exactly one option matches canonically; ambiguous keys are left
alone instead of being silently re-pointed
- tests: letter/content/full-width/trailing-space/multi-answer/unknown
cases + end-to-end grading with a content-variant key
* chore(generation): bump version to 0.3.2 for answer-key fix
The repo requires every publishable @openmaic package change to ship with
a version bump (CI check-package-version-bumps).
* style(generation): prettier-format package.json after version bump
* fix(quiz): use resolved answer representation in review renderers; prettier
- QuestionCard + quiz-view review modes highlight the option canonically
matching the stored answer key (answerIncludesOption), so persisted
content-variant keys highlight correctly after the grading fix
- prettier formatting on touched files
* test(quiz): guard possibly-undefined answer in multi-choice test
* fix(quiz): narrow answer-key matching to exact unique alignment; project review UI from the same resolver
Per review: matching policy narrows to the DSL option-identity contract.
- Generation: exact value or unique exact label alignment; unknown and
ambiguous keys stay unresolved (fail closed). Prompt example unified to
the DSL object-options + value-answer shape.
- Grading: persisted keys resolve via the same exact alignment; the
resolver returns the option's actual value when exactly one option
matches canonically-or-exactly; ambiguous keys stay untouched.
- Review UI: QuestionCard and both quiz-view review modes highlight via
the same resolver projection (answerIncludesOption(q, optionValue)).
- Removed the broader equivalence rules (NFKC, whitespace stripping, case
folding, wrapped/prefixed-letter probe) — semantic synonyms such as
正交/垂直 remain uncovered, hence Refs #1326.
* fix(quiz): resolve answer keys for persisted keys only; align quiz template with the DSL
- grading: compatibility resolution (label -> value) now applies to the stored
answer key only. A learner submission is compared as submitted, so an option
label can no longer be accepted as a different option's value.
- quiz-content/user.md: replace the string-options/`correctAnswer` example with
the DSL shape the system prompt defines ({ label, value } options and
`answer` as an array of option VALUES), so generated keys are values.
- tests: regression that a label submission grades incorrect while the same
question's value submission grades correct.
* test(generation): update scene-prompt golden snapshot for the DSL-shaped quiz template
The golden test pins the rendered user prompt for every scene kind; the
quiz-content template now emits the DSL shape, so the snapshot follows.
* fix(quiz): read label-stored correct answers in the editor's option rows
`quiz-edit-ops.toRows` decided correctness with raw value equality, so an
existing label-keyed quiz showed A/B as correct (the review UI and the grader
both resolve labels) while every row read as incorrect internally — and the
next option edit or reorder rebuilt `answer` from those rows and silently
dropped the key.
- `toRows` now uses the same resolved projection (`answerIncludesOption`).
- `normalizeQuizAnswer`'s docstring now describes the exact-only behavior it
has had since the review: formatting variants are not normalized and
ambiguous entries pass through untouched.
- Regressions: an unrelated option-label edit and an option reorder both keep
a label-stored answer, and a multiple-choice toggle preserves it. All three
fail against the previous raw comparison.
---------
Co-authored-by: CrazyPandaP <CrazyPandaP@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
|
||
|
|
1e10f60b15 |
fix(generation): assign unique IDs per media request (#1368)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
eb2cbc701a |
feat(pro): generate concise conversation titles (#1275)
* feat(storage): add automatic conversation title lifecycle * feat(agent): generate concise conversation titles * feat(agent): schedule automatic conversation titles * fix(agent): harden automatic title lifecycle * fix(agent): refine conversation title prompt * fix(agent): harden automatic title boundaries |
||
|
|
b4ed3e90f3 |
fix(importer): pnpm install fails on Windows (#1372)
* update @openmaic/importer/rollup.config.js * update @openmaic/importer/package.json --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
dd74cfe23b |
feat(web-search): add Exa provider (#1342)
Co-authored-by: white638 <212083322+white638@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
e71b00c5c9 |
fix(agent): preserve generate_scene failure causes (#1316)
* fix(agent): preserve generate_scene failure causes * test(agent): harden scene failure coverage * test(agent): isolate warning log format * test(agent): isolate warning log level --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
132d01cd04 |
feat(canvas): add double-click text insertion (#1310)
* feat(canvas): add double-click text insertion * fix(ci): format canvas double-click text insertion * fix(canvas): prevent text insertion on double-click of locked elements - Add guard to check if double-click event target is inside an editable element - Prevents text box insertion when double-clicking locked elements (backgrounds/watermarks) - Fixes event-propagation bug where locked element selection handler returns before stopPropagation() - Add regression test suite for canvas double-click handler covering locked elements and blank canvas cases Addresses review feedback from CxuPercy on PR #1310 * fix: remove unused import in canvas-double-click-handler test * fix: improve test implementation to avoid DOM dependency - Rewrite tests to use mock objects instead of actual DOM elements - Properly type MockTarget interface to satisfy TypeScript checks - All tests now pass locally without DOM environment requirement * test(e2e): add end-to-end tests for canvas double-click text insertion - Add test for double-clicking blank canvas area to create new text element - Add test to verify double-clicking existing elements does NOT insert new text - Tests actual editor interactions, not just mock targets - Addresses CxuPercy's review feedback for proper regression test coverage * chore: format code with prettier * test(e2e): simplify double-click text insertion test - Remove flaky second test that checks element visibility state - Focus on core functionality: double-click on blank canvas creates new element - Use more lenient assertion (toBeGreaterThan) to account for dynamic elements - Increase wait time for slide editor to fully render (1000ms) * fix(storage): increase test timeout for replace/release operations The 'replace retires rather than revokes an issued snapshot; release revokes both' test performs multiple async operations (put, resolve, replace, release) that can exceed the default timeout. Increased timeout to 10 seconds to accommodate I/O operations. |
||
|
|
ca3d785221 |
build(node): enforce direct dependency engine floor (#1337)
Why: - The root Node 20.9 contract predates direct runtime dependencies whose declared minimums now reach Node 22.19. - README, contributor, localized docs, and the OpenMAIC extension skill repeated the stale supported-version claim. What: - Raise the root Node minimum to 22.19 and align every operator, contributor, locale, and skill prerequisite. - Add a check that compares the root minimum with installed direct production dependency engine minimums. - Run the new contract check in CI after the frozen dependency install. Risk: - This changes the declared minimum only; no upper bound is added and Node 24 compatibility remains a separate concern. - The localized docs build was validated with the independent #1306 boundary fix from PR #1307, which is not included here. Tests: - RED on the base: engine check reported pi-agent-core, pi-ai, svg-pathdata, and undici floors - GREEN: root minimum 22.19 satisfies 35 engine-constrained direct dependencies - Node 20 lockfile-only install reports the root unsupported-engine warning - Node 22 frozen install and postinstall - Docs build with PR #1307 boundary: 34 pages and all locale postexport checks; docs types:check - Root Prettier, ESLint, TypeScript, i18n, package-version, and internal-dependency gates - Root pnpm test: 7137 passed, 81 skipped Live Docs: - GitHub issue #1304 tracks the Node contract; #1306 / PR #1307 tracks the separate docs-build prerequisite. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
e58c700082 |
fix(generation): keep diagram node position off the hover transform (#1339)
* fix(generation): keep diagram node position off the hover transform
A CSS `transform` overrides the SVG `transform` presentation attribute, so a
generated `.node:hover { transform: ... }` rule wipes out the `translate()`
that positions the node and snaps it toward the origin while hovered. The
diagram prompt only asked to "avoid hover transform conflicts", which was too
vague to prevent this.
Spell out the contract instead: the outer `<g>` carries positioning only, and
hover/active animations apply to an inner group.
Refs #1300
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore(generation): bump package version to 0.3.4
---------
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
|
||
|
|
55e324768e |
fix: keep videos playing inline on mobile (#1321)
* fix: keep videos playing inline on mobile Add playsInline to every playback video path, including the legacy app renderer, renderer-backed playback, and the shared renderer package. This keeps mobile browsers such as WeChat on Android from interrupting inline playback and restarting videos after an interruption. Add regression coverage for both app renderer modes and the package renderer. * chore(renderer): bump package version * chore(renderer): bump package version to 0.1.6 --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
d4ef5faa76 |
fix: replace rm -rf with cross-platform fs.rmSync in package scripts (#1338)
* fix: replace rm -rf with cross-platform fs.rmSync in package scripts The build/clean scripts used the Unix \ m -rf dist\ command, which fails on Windows where cmd.exe has no rm. Replace with Node's built-in fs.rmSync, available on all platforms (Node >=20.9). * chore: bump @openmaic package versions for cross-platform script fix |
||
|
|
32bfd197eb |
fix(generation): load micropip before generated imports (#1282)
Generated Pyodide widgets can import micropip before it is loaded, aborting initialization. Add ordered prompt guidance, cover the packaged asset, and bump the generation package patch version. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
96343c440e |
fix(docs): isolate standalone build from product middleware (#1307)
Why: - The docs Turbopack project walked up to the repository root and compiled the product middleware. - The v1.0 middleware's product-only feature-flag alias made every current docs build fail before static export. What: - Resolve the absolute packages/docs directory from the Next config URL. - Set that directory as the standalone docs app's Turbopack root. Risk: - The setting affects only the docs Turbopack project; root product middleware and aliases are unchanged. - Webpack-specific behavior is unchanged because the docs workflow builds with Turbopack. Tests: - packages/docs pnpm build: 34 static pages and all locale postexport checks passed - packages/docs pnpm types:check - Root Prettier, ESLint, TypeScript, and i18n gates - Root pnpm test: 7137 passed, 81 skipped Live Docs: - GitHub issue #1306 tracks the regression and build evidence. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
9d16e68296 |
fix(workbench): persist Pro conversation titles (#1273)
* feat(storage): persist manual session titles * feat(agent): add session title patch route * fix(workbench): preserve renamed session snapshots * fix: use typed session title store * style(storage): format session title tests * test(storage): align agent session schema contract * fix(workbench): fence stale session title bootstrap * fix(workbench): close session title races * fix(workbench): harden session title consistency * fix(workbench): preserve unconfirmed session titles * fix: harden session title reconciliation * fix: finalize session title consistency * test(workbench): document session title recency |
||
|
|
dfebbcf33f |
perf(classroom): speed up classroom loading (media hydration, sidebar thumbnails, media range requests) (#1276)
* perf(classroom): defer non-priority media blob hydration off the load path Entering a classroom awaited the full mediaFiles restore — one object URL per restored image/video blob — before the loading gate opened. For video-heavy courses that is hundreds of MB of IndexedDB materialization blocking first paint. Split the restore into two phases: - The awaited phase now builds metadata-complete task entries but only creates object URLs for failed rows (none needed) and for media referenced by the scene the classroom opens on (persisted cursor, else the first scene). buildRestoredMediaTasks gains an optional shouldHydrateBlob predicate, defaulting to eager hydration so existing callers are unchanged. - Remaining blob-backed records hydrate in the background, chunked over requestIdleCallback (setTimeout fallback). Each deferred task stays 'done' — so generation resume never re-runs it — but carries no objectUrl yet, which the media resolution state machine already renders as a pending skeleton until the URL lands. Background hydration is guarded per record: a task that was replaced (classroom switch, regeneration, retry) or that belongs to another stage is skipped and its freshly minted URLs are revoked immediately. * perf(classroom): lazy-render slide thumbnails in the playback sidebar The playback scene sidebar mounted a full SlideCanvas for every slide scene the moment the classroom opened (and every off-screen video element opened a preload="metadata" fetch), because SlideThumbnail's existing visible prop was never passed. Extract the editor nav rail's near-viewport IntersectionObserver hook to lib/hooks/use-near-viewport.ts and gate the sidebar's slide thumbnails through it: only scenes within 200px of the viewport render the live canvas; the rest show SlideThumbnail's existing placeholder until scrolled near. The placeholder keeps the same box size, so gating never shifts layout. The shared hook now starts hidden instead of eager: with the previous eager-initial state, opening a long deck still mounted every canvas for a frame before the observer could flip off-screen items off. The observer's guaranteed initial delivery flips near-viewport items within a frame; the no-IntersectionObserver fallback (e.g. jsdom) defers to a microtask so the effect never synchronously re-renders. * perf(media): lazy image decoding and HTTP Range support for classroom media Renderer: BaseImageElement's <img> now carries loading="lazy" and decoding="async", so thumbnail-heavy surfaces (playback sidebar, editor nav rail, course cards) no longer fetch and decode every slide image up front. In-viewport images are unaffected — the browser fetches them immediately. slideToPng forces eager loading inside its permanently off-screen snapshot tree, where lazy images would otherwise never fetch and exports would capture blank slides. Server: GET /api/classroom-media/[classroomId]/[...path] now answers single byte-range requests with 206 Partial Content (Content-Range, Accept-Ranges, correct Content-Length), enabling progressive playback and seeking for hosted video/audio instead of downloading whole files. Suffix ranges are supported; unsatisfiable ranges get 416 with the full size; unsupported units or multi-range sets fall back to the plain 200 full-body response, which is always a legal answer. Existing caching headers are kept on every response shape. * chore(renderer): bump version to 0.1.4 for image lazy-loading change * fix(classroom): guard deferred hydration against restarted tasks and bound idle waits A deferred record is only ever 'done' without an objectUrl; a task that regeneration or retry restarted passes through pending/generating, so skip those instead of attaching stale persisted bytes. Also give the idle scheduling a timeout so a busy main thread cannot starve hydration. * fix(classroom): stop superseded deferred hydration before minting URLs A restore epoch captured at apply time now gates each idle chunk: loading another classroom or reloading the same one invalidates any older hydration loop, so it neither keeps scheduling work for an abandoned classroom nor attaches bytes read by an older load to a newer load's tasks. * fix(classroom): keep 416s uncached and classify element-keyed media as priority A 416 with public immutable caching can poison the media URL for later valid requests, so range errors now send Cache-Control: no-store. Priority classification also collects the opening scene's media element ids, since task lookup binds records keyed stage:<elementId> even when the slide slot carries a different opaque ref. * fix(classroom): tie deferred hydration liveness to the classroom load token The restore epoch only advanced when the next load reached apply, so an abandoned classroom kept hydrating during the next load's storage/network phase. Compose the epoch with the load's isCurrent (load token plus effect cleanup) so navigation stops the loop at the next idle boundary. * fix(classroom): classify legacy-recovered media before deferring it Legacy singleton video recovery assigns a placeholderRef only at the end of the task build, so a record keyed by an allocated id without placeholderRef was deferred even when it backs the opening scene's gen_vid_* element. Run a metadata-only build first to learn each record's effective ref and classify against it, keeping the first visible page's legacy video eager. * fix(action): wait for deferred video bytes before starting play_video executePlayVideo treated status done as immediately playable, but a deferred restore is done without an objectUrl: the renderer shows a skeleton, no <video> exists, and the later hydration never retriggers play, leaving the action stuck until the safety timeout. Readiness now follows the renderer's contract (only done-with-bytes is playable) at the initial check, the subscription exit, and the post-subscription recheck; the failed skip applies whether or not a wait happened. * fix(classroom): include the stage whiteboard in priority media refs The stage-level whiteboard stays open across standalone classroom switches, so its media can be visible before any scene is. Classifying it as deferred left visible whiteboard media pending behind idle hydration chunks. |
||
|
|
f8b30ef0f0 |
fix(storage): prune redundant intermediate message_update frames (#1279)
* fix(storage): prune redundant message update frames * fix(storage): prune all message update runs * chore(storage): release 0.27.0 for pruneMessageUpdates Co-Authored-By: Claude Code <noreply@anthropic.com> --------- Co-authored-by: Claude Code <noreply@anthropic.com> |