67 Commits
Author SHA1 Message Date
fa47efe1b3 feat(persistence)!: server-backed persistence only, with a one-way legacy browser import (#1710)
* ci: run CI for the persistence-default integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(docker): start PostgreSQL by default and add a single-user owner

* feat(docker): start PostgreSQL by default and add a single-user owner

`docker compose up` now runs a server-backed app against a bundled
PostgreSQL, started only once PostgreSQL is healthy and published on
127.0.0.1 only, with a new built-in singleUser owner auth method: every
request resolves to one fixed owner that may publish, and no anonymous
cookie is minted.

Single-user mode (OWNER_SINGLE_USER=true, OWNER_SINGLE_USER_ID default
"local") runs with or without ACCESS_CODE. Without one, the server logs
one prominent startup warning that anyone who can reach it shares, edits
and can delete the single library; it never inspects request peers or
forwarding headers. It excludes PERSISTENCE_SHARED_OWNER_ID and follows
the sharedTeam registration rules. Its principal gets a claim candidate,
so a browser's earlier anonymous work is claimed into the single owner
(OWNER_CLAIM_TRIGGER=auto in the Compose defaults).

Compose defaults live in docker-compose.defaults.env, read before
.env.local so they can be overridden. The app warns when its declared
published address (OPENMAIC_PUBLISH_ADDRESS, passed by Compose) is not
loopback while PostgreSQL uses the default password.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(docker): keep claims explicit and let .env.local set DATABASE_URL

- Compose no longer defaults OWNER_CLAIM_TRIGGER to auto: an irreversible
  merge of every browser's anonymous library into the single owner must
  not happen on upgrade. The single-user principal still gets a claim
  candidate, so POST /api/identity/claim works; the docs explain it and
  warn about auto on a deployment several people used.
- The bundled DATABASE_URL moves to docker-compose.defaults.env, read
  before .env.local, so an external database or a rotated password set
  there keeps working. The password is inserted unencoded, so it must be
  letters and digits (documented).
- Upgrade docs no longer claim browser-stored courses are copied as they
  are opened: they stay in the browser, are not deleted, and are moved by
  the automatic browser-to-server migration shipped with this work.
- Single-user mode without ACCESS_CODE now logs one warning instead of
  two; the docs note that the single owner id is permanent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence)!: server-backed persistence is the only persistence (#1707)

* feat(persistence)!: make server-backed persistence the only persistence

Courses, folders and folder membership, chat history and learner runtime,
and generated media are now always stored on the server through the
embedded /api/persistence endpoint. The browser keeps no durable data of
its own.

- Remove the build-time NEXT_PUBLIC_PERSISTENCE switch: the client
  bootstrap always configures the HTTP document, runtime and asset seams,
  and the Dockerfile, docker-compose.yml, the Pi route and the whiteboard
  runtime no longer read it. Unconfigured seams refuse to resolve instead
  of falling back to IndexedDB, and the learner key is always the
  server-derived one.
- Remove the browser write paths: the browser document, runtime and asset
  stores as backends, the browser-mode branches of stage, folder, chat,
  media, narration and import storage, the Dexie backup export/import and
  the whole-database clear. The library lists through /api/stages, folders
  through /api/folders, and membership now goes through
  /api/folders/members. Regeneration always forks to a fresh asset.
- Device-local state moves to its own IndexedDB database
  (lib/device-storage, maic-device-cache): the local media and narration
  cache and refused-bytes retention, staged PDF images, undo history,
  browser voice profiles (carried over once from the old database) and
  the auto-voice cache. "Clear Local Cache" clears exactly that and no
  longer claims to delete classrooms or chat history.
- Keep the pre-server data readable for the one-way importer: the Dexie
  schema and read-only accessors for its tables and for the browser
  document, runtime and asset databases live in lib/legacy-browser-storage,
  which nothing on the regular paths reads and which has no write path.
- Require a database: the server exits at boot without DATABASE_URL, with
  a message naming `pnpm db:up` and `docker compose up`. `pnpm db:up` /
  `pnpm db:down` start and stop the Compose postgres service alone,
  published on 127.0.0.1 through docker-compose.db.yml; .env.example
  carries the matching local DATABASE_URL.
- E2E specs seed their courses through the persistence endpoint, each
  under a fresh id, instead of writing IndexedDB.
- Guard tests pin the boot requirement, the absence of any
  NEXT_PUBLIC_PERSISTENCE read, and the legacy-storage boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: run the E2E suite against PostgreSQL

The app refuses to start without DATABASE_URL, so the E2E job gets a
postgres:16 service and a job-level DATABASE_URL. The unit, storage and
render-service jobs are unchanged and need no database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document always-on server persistence and the database requirement

README (English and Chinese), the deployment guide in every locale, the
extension cookbook and the changelog now describe PostgreSQL as required:
`pnpm db:up` for local development, `docker compose up` for a deployment,
and an external database for Vercel and other serverless hosts (the
Deploy button prompts for DATABASE_URL). They list what stays in the
browser, drop the NEXT_PUBLIC_PERSISTENCE instructions, and say that
courses an earlier browser-only build stored stay in the browser until the
one-way importer moves them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pbl): use the server-derived learner key for PBL v2 runtime

PBL v2 synchronization, hydration and drain resolved the learner key from
their default watermark KV store, which reads or mints a device key in
localStorage instead of asking runtime configuration. The runtime store is
the server one, which refuses any key but the owner's, so saving a course
with a PBL v2 scene failed before the document write, and the minted key
landed in the slot the one-way importer reads as the pre-server runtime
partition.

The learner key now comes from getLearnerKey(args.kv): the configured
server-derived key, or an explicitly injected store. The default KV store
still holds the drain watermarks and nothing else.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): pin that clearing the cache keeps legacy databases

- Seed the four pre-server databases (MAIC-Database, maic-documents,
  maic-runtime, maic-asset-pool) with real rows, run Clear Local Cache, and
  assert every database and row survives; assert the device database name
  differs from every legacy name.
- The legacy-module import rule now also catches relative static,
  side-effect, dynamic and require imports at any depth.
- New rule: the runtime learner key is never read from a KV store the
  caller made up, only from configuration or an injected store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): run the development database as its own Compose project

`pnpm db:up` / `db:down` shared the checkout's default Compose project with
`docker compose up`, so they recreated or stopped a running stack's
database. docker-compose.db.yml now declares its own project
(`openmaic-dev-db`), which gives the development database its own
container and data volume; it no longer touches the stack, and the two do
not share data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: start the development database before `pnpm dev`

CONTRIBUTING, the getting-started guide in every locale and the startup
modes reference now run `pnpm db:up` and set DATABASE_URL before
`pnpm dev`. README, the deployment guides, the cookbook and the changelog
describe `pnpm db:up` as a separate development database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): sharpen the legacy-storage and learner-key guards

- The Clear Local Cache test closes its seeding connections to
  maic-documents and maic-runtime, so a regression deleting either fails
  on the assertion that names the database instead of by timeout.
- The legacy-module and Dexie import rules accept comments inside
  import() / require(), so a webpackIgnore hint no longer hides one.
- The learner-key rule also flags an injected-looking name that is bound
  to a default local store in the same file; its remaining limit is
  documented, with the behavioural PBL test as the guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): pin the development database's Compose project

`pnpm db:up` / `db:down` now pass `-p openmaic-dev-db`, which takes
precedence over COMPOSE_PROJECT_NAME from the shell or a .env file; the
file's `name:` alone could be overridden and point the scripts at a
running stack's database. The compose test pins the flag. The docs (all
locales) and the file header say the development database is shared by
every checkout on the machine, and how to run a separate one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): import legacy browser data into the server, once (#1708)

* refactor(persistence): prepare the seams the legacy browser importer uses

- Legacy module: read-only accessors for the auto-voice cache, the course
  ids the old media and narration tables name, the pre-server learner key,
  and asset bytes from the browser asset pool.
- Narration adoption exports its row ownership rule so the importer applies
  the same one to rows of the old database.
- The quiz legacy migration is split into a non-deleting core
  (importLegacyQuizSnapshot) and the regular path, which still deletes the
  keys it migrated.
- The lazy document migration exports its snapshot canonicalizer.
- Clear Local Cache keeps the importer's per-owner ledgers, as it keeps the
  pre-server learner key.
- A library-changed window event lets an open library list courses that
  arrive in the background.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(persistence): import legacy browser data into the server, once

On the first load after the upgrade, the client moves what earlier builds
kept in this browser to the server, for the owner the server resolves:
courses from the browser document store and the original tables (with
chat, learner runtime, playback position, roster, folders and membership,
pre-runtime quiz state), their media bytes through commitToPool and the
existing write-back funnels, the bytes of server courses from earlier
opt-in server builds that only the old tables hold, and device-only rows
(failure records, refused bytes, auto-voice clips) into the device cache.

It is automatic and silent, runs after the page is idle, never writes to a
legacy store, and records every step in a per-owner ledger so it resumes
after a crash and never imports twice. Runs are serialized across tabs
with Web Locks. A course the owner already has stays authoritative; one
whose id another owner holds is imported under a derived fresh id; one the
owner deleted on the server stays deleted. Transient failures retry on a
later load with backoff; permanent ones are recorded per item.

The module is temporary and self-contained; its README lists the removal
steps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin that each owner keeps its own import ledger

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): run the legacy import against the app routes and a browser

A PostgreSQL suite routes every request the importer and the app seams
send to the real persistence, library and folder handlers, owner resolved
from the anonymous cookie: a full import, a course another owner holds, a
course the owner deleted, and an existing course whose media is filled in.
An end-to-end spec seeds an old pre-server database in Chromium, loads the
app, and checks that the course reaches the library with its narration
uploaded, stays in the browser untouched, and is visible from a fresh
browser context with the same owner cookie.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: describe the one-way import of legacy browser data

The README (and its Chinese version), the deployment guide in every locale
and the changelog now say what the importer does instead of announcing it:
courses an earlier browser-only build kept in the browser move to the
server automatically on the first load after the upgrade, and the browser
copy stays untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): probe server assets through the shared asset-URL owner

The importer asked the pool directly whether an id exists, which the
asset-URL ownership boundary forbids outside use-asset-url. It now uses
assetRefExists, the existing metadata-only probe.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): guard the importer's pool probe like every other entry point

Only a reference the pool could have issued is probed, as the lease guard
requires; the importer is added to the list of guarded entry points.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import legacy browser data once per browser

Review round 1 of the legacy browser importer.

- Once per browser, not once per owner: the data belongs to whoever used
  the browser before the upgrade, so the first owner the import runs for
  claims it and any other owner gets nothing. The ledger is one key,
  maic:legacy-import:v2, recording the owner only as a SHA-256 digest
  (an anonymous owner id is a bearer credential) and a random salt.
- Handoff: when the claiming owner is retired (403 OWNER_RETIRED), or the
  current owner holds courses or folders the importer created (a claim
  moved them), the current owner continues the unfinished items; courses
  already done are never imported again, so a claim neither duplicates nor
  resurrects them, and a half-copied course keeps its runtime.
- Fresh ids come from the salt and the legacy id through SHA-256, not from
  the owner: unlinkable, unpredictable, and the same in every tab.
- 401 pauses with backoff instead of closing the import; 403
  FORBIDDEN_LEARNER stops the run and leaves the item pending.
- Without Web Locks: ledger writes merge with the stored copy, and the
  library is listed again right before a course is placed, so a second tab
  does not import a course the first one just created.
- Media the server refuses for good gets the app's failed-media record
  instead of a dangling reference; the legacy bytes stay.
- A legacy whiteboard or PBL session is not created when the server
  already has an active one of that kind.
- A folder that is gone leaves the course unfiled instead of failed, and a
  membership naming a folder the old database lacks counts as unfiled.
- Narration rows from before the course column are adopted when exactly
  one legacy course names their key; unconvertible chat rows are noted; an
  unreadable document-store copy falls back to the original tables; PBL v2
  documents go through the save-path strip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the importer's handoff, tabs, refusals and review gaps

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: say the legacy import runs once per browser

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the first owner to load
the upgraded app in a browser claims its legacy data, what the handoff to
an account does, that the ledger holds no owner id, and how refused media
and the new pause and skip cases behave.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): note runtime the importer cannot find without the old learner key

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): let an owner confirm it absorbed a given owner through a claim

GET /api/identity/merged-from?salt=<hex>&digest=<hex> answers whether a
claim merged into the requesting owner an owner whose id hashes to
SHA-256(salt, id). It reads only the requester's own owner_merges rows,
never lists or names an owner, resolves the owner like every other
owner-scoped route, and is uncacheable. The one-way import of pre-server
browser data uses it to decide whether a different owner may continue an
import the first owner started.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): hand the legacy import to another owner only on a confirmed claim

Review round 2 of the legacy browser importer.

- The only way the browser's legacy data moves from the owner that
  claimed it to a different owner is the server confirming, through
  GET /api/identity/merged-from, that the current owner absorbed that owner
  through a claim. The inferred signals (an owner that never wrote, a
  library listing that holds imported courses) and the retired flag are
  gone; a 403 OWNER_RETIRED only stops the run and leaves items pending.
- The first owner is bound only once the run's first authenticated request
  of its own (the library listing) succeeded, so a run that dies before
  that leaves the ledger unclaimed.
- The owner is recorded as SHA-256 of the per-browser salt and the owner
  id; the server computes the same value from the salt the browser sends.
- A "not absorbed" answer is reused for an hour instead of asking on every
  load, and the claiming owner stored first survives another tab's save
  unless that tab handed the import over; completion survives too.
- Run-level failures (401, OWNER_RETIRED, FORBIDDEN_LEARNER, OWNER_BUSY)
  from any call stop the run and leave the course pending, including the
  "is the id taken" read and the library re-list, which could previously
  mark the course failed and close the import.
- Refused media without a generation request (the user's own) gets a new
  non-retryable ASSET_REFUSED code, so it shows as failed without a Retry
  that could only fail; a skipped singleton runtime session leaves a note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin the importer ledger's owner merge and the singleton-skip note

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the legacy import moves to another owner only on a confirmed claim

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the import runs once per
browser and moves to another owner only when that owner claimed the
original one, confirmed by GET /api/identity/merged-from; that the owner
is recorded as a salted digest; and how refused user media shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): never let a transient failure settle a legacy import item

Review round 3 of the legacy browser importer.

- A read of the old browser stores that fails in storage (an aborted
  IndexedDB transaction, a closed database) now pauses the run and leaves
  the course pending. Only a record that cannot be migrated, parsed or
  validated is skipped. This covers the course read (which no longer falls
  back to the older table copy on a storage failure), and the quiz-scene
  and speech indexes, which used to read a failure as "no scenes" or "no
  holder" and drop quiz state or narration for good.
- After a session-id collision, a failed read of the session propagates
  and is retried; only a successful read that shows the id taken skips it.
- Without Web Locks, two tabs of different owners can both find the
  ledger unbound. Every ledger write keeps the owner stored first, and the
  run now re-reads the stored ledger after binding, before folders, before
  each course and before each course document is created, and stops if
  another owner holds it.
- The "not yours" answer is reused for ten minutes instead of an hour, a
  stored wait further ahead than the longest the importer sets (a clock
  that ran ahead) counts as passed, and expired answers are pruned.
- Generated media refused without an error code keeps its Retry; only
  media nothing can regenerate gets ASSET_REFUSED. A poster that fails
  transiently leaves the video pending instead of being dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover transient reads, racing owners, the not-yours cache and the claim client

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): say what the ledger's salt protects, and the new pauses

The salt defeats precomputed and cross-browser tables, not a targeted
guess against one browser's ledger; the module README and the digest's
comment now say so. The README also covers storage read failures, which
pause instead of skipping, the two-owner race without Web Locks, and the
ten-minute recheck.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): bind a browser's legacy data to one owner on the server

The one-way import of pre-server browser data now asks the server whose
data it is, instead of deciding in the browser.

- legacy_import_bindings (browser_id primary key, owner_id, timestamps),
  provisioned with owner_merges on fresh and upgraded databases.
- POST /api/identity/legacy-import-binding {browserId}: one atomic insert
  (ON CONFLICT DO NOTHING) and a read-back; answers only whether the
  requesting owner holds the browser. Same-origin JSON, resolved and
  write-fenced like every owner write, 400 for a malformed id.
- A claim participant re-keys the claimed owner's bindings to the account
  inside the claim transaction, so the account simply holds the browser.
- Owner resolution refuses any request carrying X-OpenMAIC-Legacy-Import
  with 409 LEGACY_IMPORT_NOT_BOUND unless the owner it resolves to holds
  that browser: one central fence every owner-scoped route goes through.
  Requests without the header are unaffected.
- GET /api/identity/merged-from is removed; the binding replaces it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import through fenced clients bound by the server

Review round 4 of the legacy browser importer.

- The ledger holds a random browser id and no owner information; the
  owner digest, the "not yours" cache and the client-side handoff are
  gone. A run first asks the server to bind the browser and imports only
  when the requesting owner holds it.
- Every importer request goes through its own clients that carry the
  browser id, so the server refuses any write under an owner that does not
  hold the browser: after a cookie switch in another tab, from a page whose
  memoized owner is stale, or from a tab that lost the binding race. The
  runtime learner key is asked for during each run, never taken from the
  page. 409 LEGACY_IMPORT_NOT_BOUND stops the run with items pending.
- The write-back funnels, commitToPool and the PBL save-path strip take an
  optional store, put or runtime, so the importer passes its fenced ones.
- A failed read of one legacy course holds up only what it could affect:
  the quiz-scene and speech indexes leave the course out as unindexed, and
  only quiz state or narration it could own waits. Read failures are
  counted per course; after five failing runs over at least a day, a
  course whose document-store copy cannot be read falls back to the table
  copy, or is skipped with the reason when there is none.
- Runtime ids are rewritten by replacing only their course segment, so a
  short course id cannot corrupt the prefix or the learner segment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the server binding, its fence, bounded reads and short ids

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the server binds a browser's legacy data to its first owner

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say: once per browser; the server
binds the browser to the first owner; a claim carries the binding to the
account; the importer's requests are refused for any other owner. They
also describe the bounded pause for a course the old storage cannot read,
and the removal steps for the server side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): keep a dropped asset request pending in the legacy import

The asset client reports a fetch that got no answer (a dropped request, an
aborted upload, the existence probe's deadline) as status 0 with
HTTP_REQUEST_FAILED, which the importer read as a final refusal: the media
was marked failed and the course could complete without it. Read that
code as transient, keep the client's local validation failures final, and
treat a 2xx/3xx answer a client could not use as transient too. Unit tests
drive every client the importer uses with a rejected fetch and a non-JSON
answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): ask a refused binding again only every ten minutes

A browser whose legacy data another owner holds sent the binding request
(a write) on every page load. Record a ten-minute recheck in the ledger
(clamped like every other deferral), so a claim that moves the binding is
still picked up, and open the old databases only once the browser is
bound. The read budget now counts consecutive failing runs (a run that
reads the course resets it) and counts a first failure dated in the
future from the next one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): answer 503 with the cookie when the import fence cannot read

The legacy import fence now reads the header name from the shared
constant, and a database error while reading the binding answers the
usual JSON error (503 PERSISTENCE_UNAVAILABLE) with the resolution's
Set-Cookie values instead of escaping as a bare 500. Route tests cover
that, a bind by an owner a claim retired (403 OWNER_RETIRED, no row), and
one header name on both sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover dropped asset requests, the recheck and the read budget

With the real asset client's error for an upload and an existence probe,
the media stays pending and a later run imports it. A non-holder's five
loads send one binding request and never open the old databases, and a
claim still hands the import to the account after the delay. The read
budget's run count and time span each hold on their own, a clock that ran
ahead does not hold a course open, and a good read resets the count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): complete the legacy import removal steps

List every file the removal touches (the claim scenarios test, the claim
participant table and list, every doc), keep the bindings table for one
release after the code goes because of rolling deploys and give the DROP
statement for that release. Add the claim's eighth participant to the
README lists, fix the fresh-id wording in the changelog, and describe the
recheck delay, the transient status-0 asset failures and the consecutive
read budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): keep legacy quiz state through Clear Local Cache until the import completes

Pre-runtime quiz drafts, answers, results and attempt ids exist only in
localStorage, and the one-way importer copies them to the server with their
course. Clear Local Cache deleted them even while the import was still
pending (it waits for an idle page and can stay pending after a failure), so
that quiz progress was lost for good.

Clear Local Cache now keeps the four quiz key families unless the importer's
ledger records the import as complete, read through a new
`legacyImportIsComplete` helper in the ledger module. The ledger key moves
there too, so the two modules do not import each other. Once the import is
complete, the keys are cleared as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: show DATABASE_URL set in the local development setup

The getting-started guide showed the required DATABASE_URL line commented
out, so copying the block left the server unable to start. Each locale now
shows `pnpm db:up` and, separately, the `.env.local` line to set.

The `.env.example` header said every variable is optional; it now says that
provider variables are, and that DATABASE_URL is required except under
docker-compose.yml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): drop two upgrade notes the final code contradicts

The Compose entry said its default DATABASE_URL overrides .env.local; the
defaults file is read first, so a value in .env.local wins, as the same entry
says further on. The runtime-session entry offered turning server
persistence off as an upgrade path; that switch no longer exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(env): restore the ACCESS_CODE warning notes in .env.example

A later change to the example file dropped the lines saying that an unset
ACCESS_CODE fails open, that the server then logs a one-time startup warning
and GET /api/health reports accessCodeConfigured: false. They are back, and
say that single-user mode logs its own warning in place of the generic one,
as the startup code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* ci: run every app PostgreSQL suite in the contract job

The app-domain step named its suites one by one, and seven had been left
off, among them the legacy importer's route suite. No other job gives the
app suites a database, so they skipped everywhere. The step now lists every
app `*.pg.test.ts`; all of them work in a schema or database of their own,
or truncate only their own tables, and pass before and after the storage
package suite in the same database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): establish the anonymous owner before the first API request (#1721)

* fix(identity): establish the anonymous owner on the page response

On a browser's first load the page sent several API requests without an
owner cookie; each minted its own anonymous owner and set its own cookie,
and the last answer to arrive won. Work done in between (the legacy import
binding, a first course write) then belonged to an owner the browser no
longer presented.

The middleware now mints the anonymous cookie on document navigations that
carry no valid one, with the route handlers' own value format and
attributes (the cookie module is Edge-safe now: Web Crypto instead of
node:crypto), and forwards it on the request so the page render resolves
the same owner. API, RSC, prefetch and Server Action requests never mint
there, and a valid cookie is never replaced. It mints only when neither
OWNER_SINGLE_USER nor PERSISTENCE_SHARED_OWNER_ID is set; host
registrations are invisible to an Edge middleware, so the README tells
hosts where to skip it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(legacy-import): bind only an owner the browser already presents

A binding request whose owner was minted by that very request answers
409 OWNER_NOT_ESTABLISHED with the minted cookie and binds nothing: other
cookieless requests may still be minting owners of their own, and a
binding to this one could be left with an owner nobody presents. The
importer treats it as transient and retries on a later load. It also says
plainly when a request after its own successful bind is refused as
another owner (the owner cookie changed mid-run); items stay pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): cover a first visit whose first answers arrive out of order

Against the real routes and PostgreSQL: a browser without a cookie loads a
page, sends two requests at once, and the first answer reaches it only
after the importer bound the browser; the import completes under the one
owner the page established. A second case binds nothing for an owner the
bind itself minted and imports on the next run. The E2E drives the same
ordering in Chromium by holding the first owner-scoped answer until the
binding answered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): drop an unused initial assignment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): note the page response now sets the anonymous owner cookie

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): let hosts turn page-response minting off

The middleware cannot see owner auth methods a host registers, so a host
with anonymousFallback: false still got an anonymous cookie on a cookieless
page load. Beside the host credential it is a claim candidate; with
OWNER_CLAIM_TRIGGER=auto it is claimed and cleared on the next request,
again after every such page load.

OWNER_ANONYMOUS_PREMINT (default on; true/1/false/0, anything else fails
startup) is read by the middleware before minting. At startup the server
warns once when a registration turns the anonymous fallback off while
pre-minting is still on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): document OWNER_ANONYMOUS_PREMINT

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): keep the anonymous identity for 400 days and renew it in use

The anonymous cookie had a fixed 30-day lifetime from its mint and was
never renewed, so every anonymous visitor lost their whole library 30 days
after the first visit however active they were; with persistence on the
server only, that affects every anonymous deployment.

It now lasts 400 days (the browser cap, one shared constant), and every
route handler and Server Action response that resolves to a valid
anonymous cookie re-sends the same value with a fresh Max-Age (best-effort
where next/headers refuses writes). Page responses never renew, so pages
stay cacheable, and a host principal never renews the anonymous cookie
beside it.

A response that clears the cookie (a claim, a retired owner) must not
renew it, whatever order a route merged its headers in:
withRequestOwner, and the owner-events retired answer, keep only the
clearing value for a cookie a value clears. A renewal that lands after a
claim cleared the cookie is pinned by a test: the next write with it,
alone or beside the account, writes nothing, and its answer clears it
again without renewing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): 400-day renewed anonymous cookie; when to turn pre-minting off

Hosts set OWNER_ANONYMOUS_PREMINT=false only when their registration sets
anonymousFallback: false; hosts that keep anonymous visitors need
pre-minting and skip it per request in middleware for requests their
methods authenticate. The CHANGELOG notes the new lifetime and renewal,
and that a lost cookie means a new owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): forward owner cookies from routes that resolve the owner directly

The Pi chat and whiteboard-visibility routes and asset-id document
extraction resolved the owner themselves and never sent its Set-Cookie
values back, so an anonymous identity used only through them was never
renewed (and a minted one never stored).

attachOwnerCookies attaches a resolution's cookies to a response a route
built itself, with the clear-wins rule, before a stream's body starts. Both
Pi routes answer through it, success and error alike; resolveServerAsset
hands the cookies back with every answer and the extraction route attaches
them.

A guard test scans for every module outside the identity seam that calls
the owner resolution (or the asset helper) directly and requires it to
forward the cookies; the two callers that cannot are listed with the
reason. Behaviour tests cover the three paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: wyuc <dhq1204@yahoo.com>
2026-09-29 17:51:03 +08:00
小魔andwyuc a41bbac239 feat(llm): retry once on a fallback model for retryable failures (#1614)
* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* Delete omai-upload-tree directory

* fix: ensure trailing newline in locale files

fix: ensure trailing newline in locale files

* chore: english comments for upstream review

chore: english comments for upstream review

* Delete components/settings/model-schedule-settings.tsx

* Delete lib/server/model-schedule.ts

* Delete app/api/model-schedule/route.ts

* Delete assets/model-schedule/ledger-with-entry.png

* Delete assets/model-schedule/settings-panel.png

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Update .env.example

* Enhance access control comments in .env.example

Added additional context and warnings for access control configuration.

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* run ci

* Add files via upload

* Add files via upload

* Add files via upload

* Supplement. env.example

Supplement. env.example

* Add files via upload

* Update .env.example

* Update fallback notes in .env.example

Clarified fallback behavior in callLLM layer documentation.

* Update llm.ts

* Update llm-fallback.test.ts

* fix(llm): round-4 fallback hardening

fix(llm): round-4 fallback hardening — server-managed gate, content-filter refusal, APICallError unwrap stop, outlines fullStream errors

* fix(llm): rebase round-4 onto main v1.1.1

fix(llm): rebase round-4 onto main v1.1.1

* fix(llm): round4c-prettier formatting+retry test 

fix(llm): round4c-prettier formatting+retry test

* test(llm): round-4d-align outline-stream

test(llm): align outline-stream and title-generator tests with round-4 call signatures

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-28 17:24:55 +08:00
wyucandClaude Opus 5.5 a686481e73 feat(auth): owner identity seam — composable auth methods, owner-keyed persistence, host hooks, claims
* ci: run CI for the owner identity seam integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 1 — pluggable authenticator with today's behavior

* feat(auth): add the owner identity seam with built-in authenticators

Introduce lib/server/identity/: an OwnerAuthenticator contract
(SubjectKind, OwnerPrincipal, AuthOutcome), a single-shot, boot-validated
configureOwnerAuthenticator() registry, and the anonymousCookie and
sharedTeam built-ins that reproduce today's owner resolution exactly
(same anonymous_id cookie, owner id strings, statuses and env semantics).

Every owner-scoped route handler and the workspace Server Action now
resolve their owner through one memoized per-request resolution instead
of the three previous copies (agent-runtime/owner.ts, with-owner.ts and
the Server Action re-implementation). Publish and unpublish decide from
the principal's course:publish role instead of an anon: id prefix.

An invalid credential from an authenticator is answered with 401 on
every surface and never falls back to an anonymous owner. A source scan
keeps the anonymous owner cookie and anon: prefix checks out of code
outside the identity module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): describe the owner identity seam and host registration

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): tighten the identity boundary guard and Server Action cookies

- Scan workspace package sources too, flag imports of the concrete
  built-in authenticators outside lib/server/identity, and police the
  slice/substring/indexOf/includes spellings of an anon: id-shape check.
  Stop re-exporting the built-ins from the host-facing index.
- Refuse setCookies returned by authenticateFromContext, as the
  authenticate() fallback already did, and document that it must write
  cookies itself through next/headers.
- Let the route-test owner stub carry explicit kind/roles, and cover the
  403 forbidden branch of publish and unpublish for a signed-in owner
  without course:publish.
- README: registration conflicts fail the boot; an invalid principal is
  rejected per request with a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

Runtime sessions and assets on /api/persistence now take their identity
from the same memoized owner resolution as documents, instead of a
client-chosen x-learner-key behind a development bearer token that
shipped in the public bundle.

- Runtime: the storage handler's learner key is principal.ownerId. A
  path or body naming another learner answers 403 FORBIDDEN_LEARNER and
  another learner's session answers 404, whatever x-learner-key says; the
  header is no longer read. The browser learns its key from the new
  GET /api/persistence/learner-key and sends no credential of its own.
  The Pi whiteboard routes derive the same key from the owner.
  PERSISTENCE_DEV_TOKEN, NEXT_PUBLIC_PERSISTENCE_TOKEN and
  PERSISTENCE_ALLOW_INSECURE_DEV_AUTH are gone from the server path and
  the client, and server-auth.ts is deleted. Learner merge stays refused.
- Assets: allocations land in a per-owner partition (owner:<ownerId>),
  so quota is per owner and replace/delete reach only the owner's
  entries. Reads stay capability-by-id where a course viewer needs them:
  another owner's committed entry is readable while a live course
  references it. Legacy entries in the old shared partition stay
  readable by id, and may be replaced or deleted only by an owner who
  owns every course referencing them. Server-generated media is stored
  under the run owner's partition; asset-id extraction and vision image
  resolution use the same rule.
- Tombstones: runtime of a deleted course reads as absent (404, empty
  lists) and takes no new sessions or records, through a guard in the
  app's runtime composition; no package schema change.

The identity boundary scan now also fails on any source that reads
x-learner-key or the retired development token variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): describe owner-keyed runtime and assets, drop the dev token

Remove the development persistence token from the server-persistence
recipe, .env.example, the Docker build arguments and the deployment
docs; describe runtime learner keys, per-owner asset partitions and
quota, the legacy shared-partition rule and the tombstone behavior; and
record the breaking change for existing server-persistence runtime data
under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(storage): scope asset references to principals, typed stage refusal, uniform create collision

@openmaic/storage 0.32.0.

- PgDocumentStore takes assetReferencePrincipals: with reference
  tracking on, a write references and commits only entries held by the
  listed principals. An id naming another principal's entry records no
  row and changes no lifecycle column, like an unknown id, and the write
  is never refused. Without the option any held entry is referenced, as
  before. AssetCollector takes a matching per-document-owner function so
  the one-time backfill scopes each document the same way. This keeps a
  document naming a leaked id from committing, exposing or pinning
  another owner's allocation.
- RuntimeStageNotFoundError: a store may refuse createSession for a
  stage its host considers absent; the runtime handler answers
  404 STAGE_NOT_FOUND, recognizing the error by class or by code.
- A session create over a taken id answers 409 SESSION_ALREADY_EXISTS
  whoever holds it. It used to answer 403 for another learner's session
  and 409 for one's own, which told a caller whose session an id was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): close owner-seam boundary gaps in references, tombstones and legacy assets

- Document writes reference and commit only the writer's own asset
  entries and legacy shared ones (owner-bound document store and the
  collector backfill), so another owner's course naming a pending id
  can no longer commit, expose or pin it.
- The foreign-read rule additionally requires the referencing live
  course to belong to the entry's owner, as defense in depth for
  reference rows that did not come from a scoped write.
- The tombstone refusal for new runtime sessions moved from a raw path
  comparison in the route into the guarded store's createSession
  (RuntimeStageNotFoundError, 404 STAGE_NOT_FOUND), so no encoded
  spelling of the path routes around it.
- Replacing or deleting a legacy shared entry checks reference
  ownership and mutates in one transaction under the entry row lock,
  which a concurrent reference insert must wait for.
- Tests build the uncommitted-but-referenced, tombstoned-with-rows and
  foreign-reference states directly so each half of the read rule is
  observable, and a PostgreSQL suite holds a competing reference open
  across a legacy delete.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): upgrade note for dev-token-gated deployments, owner-scoped references

Tell operators whose only access gate was the development token to put
the deployment behind an access code or gateway, register an owner
authenticator, or turn server persistence off before upgrading, and
describe that a course references only its owner's media.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): per-owner reference principals and a typed taken-id conflict

- PgDocumentStore's assetReferencePrincipals is now a function of the
  store's bound owner, evaluated per write, instead of a fixed list. A
  store re-bound with forOwner therefore scopes to the new owner rather
  than carrying the previous owner's principals onto its writes. It has
  the same shape as AssetCollector's backfill option, so one function
  serves both.
- PgRuntimeStore and BrowserRuntimeStore raise RuntimeSessionExistsError
  for a taken session id, and the runtime handler classifies it as
  409 SESSION_ALREADY_EXISTS even when its re-read cannot see the holder
  (for example a session a host-side wrapper hides). Recognized by class
  or by code.

Still the unreleased 0.32.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): guard the Pi whiteboard runtime and answer hidden collisions with 409

- Every server-side runtime writer now uses the tombstone-guarded store:
  the Pi native whiteboard service gets it through the same helper as
  /api/persistence/runtime/*, so a deleted course takes no new
  whiteboard runtime either.
- Creating a session whose id is held by a session a tombstone hides
  answers the uniform 409 SESSION_ALREADY_EXISTS instead of a 500,
  covered on PGlite and on real PostgreSQL.
- The owner-bound document store passes its reference-principal
  function, so the scoping follows the bound owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 3 — accounts through an identity gateway (trustedProxyHeader)

* feat(auth): owner identity seam, phase 3 — trusted-proxy-header authenticator

Add the trustedProxyHeader built-in: real accounts through an identity
gateway (oauth2-proxy, Authelia, a Keycloak-based proxy, an institutional
reverse proxy) that signs the user in and forwards the verified user in a
request header. OpenMAIC never handles passwords or tokens.

- Selected with OWNER_AUTHENTICATOR=trusted-proxy and validated in
  instrumentation register(). TRUSTED_PROXY_SECRET (32-1024 printable
  ASCII characters) is required; header names default to
  x-openmaic-proxy-secret and x-forwarded-user, and an optional groups
  header with TRUSTED_PROXY_ADMIN_GROUPS grants admin. An unknown
  selector, a TRUSTED_PROXY_* variable without the selector, a shared
  owner id or a host-registered authenticator alongside it, and malformed,
  reserved or repeated header names fail the boot.
- Trust boundary: Next.js exposes no TCP peer address to route handlers,
  middleware or Server Actions, and fills x-forwarded-for from the socket
  only when the client sent none, so no address allowlist is offered. The
  gateway secret is compared in constant time (hashed, timingSafeEqual).
- Per request: a missing or wrong secret, a missing or blank user, a user
  header with a comma (duplicate lines arrive comma-joined), or a user
  whose proxy:<user> id fails the owner id guard is 401
  INVALID_CREDENTIAL, never an anonymous owner. The guard failure is a
  401 rather than a 500 because the value comes with the request. The
  user is trimmed and keeps its case. Every gateway user holds
  course:publish; kind user, assurance verified, channel proxy; no
  cookie. Route handlers and Server Actions share one code path.
- The unset-ACCESS_CODE boot warning is skipped in this mode; the gateway
  is the access gate and ACCESS_CODE stays independent.
- The identity boundary scan now also fails on gateway identity headers
  (x-forwarded-user/-groups/-email, x-auth-request-*, remote-user, the
  secret header) and TRUSTED_PROXY_* read outside lib/server/identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): document accounts through an identity gateway

Describe the trustedProxyHeader built-in in the README owner identity
section (and its Chinese counterpart): configuration, request semantics,
the trust boundary and why it is a shared secret, the interaction with
ACCESS_CODE, boot-time validation, and an oauth2-proxy example that
injects the user, groups and secret headers. Add the .env.example block,
a SECURITY.md note and an Unreleased changelog entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): harden the trusted-proxy authenticator after review

- Reserve the header names HTTP, Next.js and forwarding proxies set
  themselves: configuring the user, groups or secret header to rsc,
  next-action, forwarded, x-forwarded-for/-host/-proto/-port, x-real-ip,
  or anything starting with x-middleware-, x-invoke-, x-nextjs- or
  next-router- now fails the boot.
- Log a one-time boot warning when TRUSTED_PROXY_ADMIN_GROUPS is set: the
  admin role is granted from the groups header, which is trusted on the
  strength of the shared secret alone, so the gateway must overwrite or
  strip it.
- POST /api/chat/pi resolves the owner up front and answers 401 when the
  authenticator rejects the credential, like every other owner-resolving
  route, instead of continuing without an owner. The anonymous default
  always resolves, so it is unaffected.
- The identity boundary scan also covers the Authentik, Azure App Service
  authentication, WebAuth, Cloudflare Access, Google IAP and AWS ALB OIDC
  header families.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): admin-groups warning and a corrected oauth2-proxy example

- Warn next to TRUSTED_PROXY_ADMIN_GROUPS (README, README-zh,
  .env.example) that the groups header is trusted on the strength of the
  shared secret alone, and list the reserved header names.
- The oauth2-proxy example now states it targets v7.14+ (nested
  claimSource/secretSource), uses the canonical bindAddress, sets
  preserveRequestValue: false on all three headers and
  insecureSkipNonce: false, forwards the subject with `claim: user`, and
  notes that the secret variable holds the literal secret. It adds the
  legacy-option migration caveat, a warning not to exempt app routes via
  skip-auth or trusted-ip options, and notes that the IdP must release the
  groups claim and group names must not contain commas.
- Changelog: the new boot checks and the /api/chat/pi 401.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 4 — host extension hooks

* feat(persistence): owner identity seam, phase 4 — host extension hooks

Let a host add product behavior at four points without forking a route,
registered once from instrumentation.ts register() like the owner
authenticator and sealed on first use. Nothing changes when no hook is
registered.

- configurePersistenceHooks({ name, authorizeCreate, onCreate, library,
  beforeAssetAllocate }):
  - authorizeCreate / onCreate run once per created course, inside the
    create transaction, in the transaction that inserts the course's
    ownership row. Saves of a course the owner already holds, and a create
    that loses a race to a concurrent create of the same id, are updates and
    run no create hook. A refusal answers 403 CREATE_REFUSED; a throw rolls
    the course, its ownership row and the host's writes back together. The
    hooks receive the request principal, or only the owner id for a
    background agent run.
  - library chooses which stage ids GET /api/stages lists; the route builds
    the usual summaries in the provider's order and drops every id the read
    path would refuse (deleted or unclaimed). folderId is shown only on the
    principal's own courses.
  - beforeAssetAllocate runs inside the storage handler's asset
    authorization step for POST /assets, after the owner is resolved and
    before the body is read; a returned Response is sent instead, with
    nothing stored and no quota counted.
- configureAssetByteStore({ name, create, signsReadUrls }) replaces the
  ASSET_S3_BUCKET switch for both the persistence provider and the asset
  collector. A host store must keep bytes outside the registry database.
  Under ASSET_BYTE_EGRESS=redirect a store that does not declare
  signsReadUrls stops the server at boot; the built-in layers keep their
  behavior.
- @openmaic/storage 0.33.0: DocumentWriteRefusedError lets a store refuse a
  write as policy; the document HTTP handler answers 403 with its code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): harden the host extension hooks after review

Correctness:
- Registration reads every hook once and stores it bound to the object the
  host passed, so hooks defined on a class prototype or as non-enumerable
  properties are kept. A spread copy dropped them silently, including the
  authorizeCreate and beforeAssetAllocate gates. Same for
  configureAssetByteStore. The unknown-key check now applies to plain
  objects only, since class instances carry their own state fields.
- The owner-bound document store carries each call's operation in its own
  async context (AsyncLocalStorage) instead of a field on the instance. One
  store is shared by every tool of an agent run; a concurrent call could
  overwrite the field before the transaction read it, so a create was gated
  as a read (no ownership row, no create hooks) or claimed another stage.
  create_stage also runs sequentially within a tool batch.

Hardening:
- isDocumentWriteRefusedError recognizes a refusal from another copy of the
  storage package by name and shape (not by code alone). The document
  handler and POST /api/stages use it.
- beforeAssetAllocate also gates PUT /assets/{id}/content. The request
  carries operation ('create' | 'replace') and the decoded assetId from the
  handler's own routing.
- Agent runs: a refusal on a background write carries a fixed message, and
  create_stage answers it with a stable tool result. The host's message never
  reaches the model.
- A byte store declaring signsReadUrls without a signer degrades to direct
  bytes with a warning instead of failing reads, and the collector no longer
  requires signing.
- DocumentActor is discriminated by source ('request' with the principal,
  'background' without one).
- Library ids the read path cannot address are dropped, and a provider
  answer is capped at 5000 ids.
- The boot-time ownership backfill logs how many courses it adopted without
  hooks.
- The server persistence provider no longer exposes an unscoped document
  store.
- The writesOutsideRegistryDatabase trust boundary and the scope of upload
  admission (HTTP uploads, not server-generated media) are documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): refuse misspelled hooks on host classes too

A class instance skipped the unknown-key check, so a misspelled hook such
as authorizeCreat registered cleanly and never ran, leaving creation
ungated. Both configurePersistenceHooks and configureAssetByteStore now
reject any function-valued property of a class instance (own or on its
prototype chain up to Object.prototype, constructor excluded) that is not a
known key, and any unknown key of a plain object. The error names the key
and suggests the known key within edit distance 2. Host classes keep their
helpers private (#helper) or register a plain object; the README says so.

Also:
- isDocumentWriteRefusedError requires a string message from a cross-copy
  candidate, and documents that the name is the effective discriminator.
- CHANGELOG: ServerPersistenceProvider no longer exposes an unscoped
  documentStore; use the owner-bound store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

Course ownership was recorded twice: in stage_meta, which already answers
every access decision, and in document_stages.owner_id, which the storage
package scoped by. Two tenancy records double what a later claim has to
re-key, and a host that keeps ownership in its own table had to rewrite the
package's SQL to strip the column.

@openmaic/storage 0.34.0:
- PgDocumentStore never reads or writes document_stages.owner_id. An
  owner-bound store scopes listings, writes, deletes, the freshness manifest
  and folder membership through the host's ownership relation, named with
  the new documentOwnership option ({ table, stageIdColumn, ownerIdColumn,
  tombstoneColumn, claimOnCreate }, or false for no document scoping).
  Identifiers are validated, never quoted into SQL from arbitrary input.
- forOwner / ownerId without documentOwnership throws at construction,
  so a direct consumer cannot upgrade into a store that silently scopes
  nothing. An unbound store is tenant-agnostic.
- claimOnCreate inserts the ownership row in the create transaction; a
  concurrent create of the same id by another owner rolls back.
- folders: false lets a host whose document_stages has no folder_id use
  the store; folder ids are unique per owner, so membership goes through
  the ownership relation too.
- AssetCollector reads each document's owner from documentOwnership for
  assetReferencePrincipals, and refuses the function without it.
- Schema: fresh installs get no owner_id column and none of its indexes;
  an existing column is kept for one release, made nullable with its
  default dropped (each ALTER guarded by a catalog check, so a replay takes
  no lock), and its two indexes are dropped. document_stages_folder_idx
  replaces the owner/folder index.

Application:
- The owner-bound store binds the package to stage_meta
  (STAGE_META_OWNERSHIP) and lists through it in one query; the asset
  collector schedule passes the same relation.
- The boot backfill copies column-only owners into stage_meta before any
  store is built, only when the legacy column exists, idempotently, and
  logs how many it adopted and how many disagree (stage_meta stands).

Tests cover a fresh install and an upgrade from the previous schema on
PGlite and PostgreSQL 16, a host table with neither ownership nor folder
columns, folder isolation across owners sharing a folder id, and the
concurrent-create race.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): serialize schema bootstrap; guard unscoped owner-bound stores

Review follow-ups for the tenancy change.

- Concurrent starts: two instances starting against one database raced on
  the schema DDL (a duplicate pg_class / pg_type entry, or "tuple
  concurrently updated"), failing one instance's first request. Every
  schema bootstrap the application runs -- the persistence provider's
  sequence (runtime, documents, stage_meta and its ownership backfill,
  owner materials, assets), the agent-session, session-material and
  user-skill stores, and the asset collector's -- now runs under one
  PostgreSQL advisory lock held on a dedicated connection and released in
  finally. It lives in the application because what needs serializing is
  the whole sequence, including tables the package does not own. A
  PostgreSQL test boots eight instances at once, four rounds each, on a
  fresh and on an upgraded database; without the lock it fails every run.
- documentOwnership: false on an owner-bound PgDocumentStore now also
  requires allowCrossOwnerDocumentAccess: true, so binding an owner for
  folders or asset principals cannot silently expose every owner's
  documents.
- Documented that an ownership relation must cascade with the document
  rows (a leftover row keeps the id reserved for its owner, which is also
  how a host keeps retired ids from being reused), with tests for both.
  Leftover rows are deliberately not made claimable: a host cannot be
  told apart from one that keeps them as retirement markers.
- Upgrade notes: the one-time table locks on the first start, and the
  out-of-band CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY path
  that leaves the start with nothing to build or drop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): leave the schema bootstrap lock out of the held-lock check

The create-hooks suite asserts that its host lock is released with the
create transaction by counting granted advisory locks database-wide. A suite
booting a provider against the same database at that moment now holds the
schema bootstrap lock, which is not the lock under test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

* feat(storage): owner merge primitives for claiming anonymous work (0.35.0)

Add the per-store pieces a host needs to move one owner's data to another
inside a single transaction of its own:

- reassignDocumentFolders(tx, { fromOwnerId, toOwnerId, documentOwnership })
  moves folders and filing: a same-named folder merges into the target's, a
  colliding id is renumbered, and filing is rewritten in one statement from
  the old ids.
- PgAssetStore.reassignPrincipal(fromKey, toKey) moves a principal's entries
  under both principals' write locks, in key order; quota is not re-checked.
- PgUserSkillStore.mergeOwner(from, to) moves skills, renaming a live handle
  the target already uses with the first free numeric suffix.
- PgUserSkillStore takes a resolveFinalOwner hook, run as the create
  transaction's first statement, like the agent-session store's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

A visitor who works anonymously and then signs in can move that work to their
account in one transaction, and the anonymous id is retired afterwards.

- Principal: pendingClaim { fromOwnerId, assurance }. The trusted-proxy
  built-in sets it when a gateway request also carries a valid anonymous
  cookie; the anonymous and shared-team built-ins never do. Authenticators may
  implement describeStoredOwner, canonicalize and clearPendingClaim;
  principalFromStoredOwner describes a stored id for work without a request.
- Claims: claimOwner / claimPendingOwner and registerClaimParticipant (sealed
  on first claim). One transaction, both owners' identity locks taken
  exclusively in key order, participants in a fixed order (folders, courses,
  materials, agent sessions, skills, runtime, assets), recorded in the new
  owner_merges table. Idempotent per pair; refuses a non-anonymous source, an
  anonymous target, a source already claimed elsewhere, and chains. Same-named
  folders merge, colliding folder ids are renumbered, colliding skill handles
  get a suffix, and quotas are not applied to what moves.
- Identity lock: every owner write (documents, folders, asset allocations,
  material registration, agent sessions, skills) takes the owner's advisory
  lock in shared mode first, so a write racing a claim is moved or refused,
  never orphaned or deadlocked.
- Retired ids: request writes are refused with 403 OWNER_RETIRED; agent runs
  and media generation that started before the claim follow the id to the
  account (forwarding document store, canonicalized stage probes, forwarded
  generated assets and skills).
- Trigger: POST /api/identity/claim (same-origin JSON only; clears the
  anonymous cookie), or OWNER_CLAIM_TRIGGER=auto on the first request carrying
  a pending claim. The runtime contract's learner merge now performs the same
  claim, allowed only for the request's own pending claim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): count only the suite's own advisory locks after concurrent creates

The check that a host's transaction-scoped lock is released counted every
granted advisory lock in the database. Suites that run against the same
database at the same time hold advisory locks of their own (a schema
bootstrap, the owner identity locks a claim takes), which made it fail
depending on timing. Count only locks held by this suite's connections.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): use a neutral example table for host claim participants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): owner merge fixes and store policy errors over HTTP (review round 1)

- PgUserSkillStore.mergeOwner reserves every handle either owner holds before
  renaming: a claimed skill whose handle the target uses could be renamed into
  a handle another claimed skill still held, which violated the per-owner
  unique index and made that merge fail every time.
- PgUserSkillStore.list orders ties by id.
- PgRuntimeStore takes resolveFinalLearner(tx, learnerKey), run as the
  create transaction's first statement, so a host can fence session creates
  against its owner merges; reassignLearner(from, to) re-keys sessions without
  re-validating them, so one session a newer version wrote cannot block a
  merge.
- Every HTTP handler (runtime, documents, assets) answers a
  DocumentWriteRefusedError as 403 with its code, and the new StorageBusyError
  (code, message, retryAfterSeconds) as 503 with Retry-After.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 1 — fenced runtime creates, bounded locks, retired-cookie recovery

- Runtime session creates take the owner's identity lock as the create
  transaction's first statement and refuse a retired owner, closing the window
  in which a create racing a claim was left under the retired id.
- Lock waits are bounded: a claim waits OWNER_CLAIM_LOCK_WAIT_MS (default 5 s)
  for its identity locks, a write OWNER_WRITE_LOCK_WAIT_MS (default 30 s);
  running out, or losing a deadlock or serialization race, answers 503
  OWNER_BUSY with Retry-After on the claim routes and on every fenced write.
- Every 403 OWNER_RETIRED carries the Set-Cookie values that drop the retired
  anonymous credential (the anonymous built-in now implements
  clearPendingClaim), so a browser whose claim response was lost recovers.
  Writes by id to rows that moved (skill delete, session messages) answer
  OWNER_RETIRED instead of not-found or a bare 403.
- OWNER_CLAIM_TRIGGER=auto no longer claims ahead of the explicit claim
  routes, which now report their own claim instead of a false refusal.
- The owner-events stream tells a retired owner only that it moved, never the
  claiming account's id.
- Creates of one course id take turns (a per-id advisory lock before the
  ownership probe), so a concurrent second create saves as an update instead
  of failing as reserved-document.
- A claim's source must be described as anonymous by the authenticator; the
  write fences skip the retirement lookup for every other owner. The host
  canonicalize hook is removed: forwarding is core's owner_merges alone.
- The folders participant locks the source's stage_meta rows first; runtime
  sessions are re-keyed without re-validation; the owner asset store fences
  every allocation and refuses to allocate without transactions; background
  skill reads and patches follow a mid-run claim; the forwarding document
  store forwards methods only.
- Docs: when a pending claim can arise with the shipped authenticators, the
  anonymous cookie as a bearer credential, bounded waits, the exact scope of
  OWNER_RETIRED, and claims narrowed to what the tests show.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 2 — merge-record guard, boot-checked lock waits, stream credential drop

- owner_merges records claims of anonymous owners only, since the write
  fences enforce retirement only for ids the authenticator describes as
  anonymous. Reading a row that retires any other owner now fails loudly
  instead of being followed but not enforced. The comment that pointed hosts
  at claimOwner for their own account merges is corrected: a host merging two
  signed-in accounts moves the rows itself and refuses the merged-away account
  in its authenticator. describeStoredOwner must classify ids stably.
- OWNER_WRITE_LOCK_WAIT_MS and OWNER_CLAIM_LOCK_WAIT_MS are validated at boot
  with the same parser the lock paths use, so a malformed value stops the
  server instead of failing every write.
- The owner-events stream answers a connect by an already retired identity
  with one owner_moved event and the Set-Cookie values that drop the retired
  credential, so a stale tab's reconnect loop ends after one round trip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* chore(auth): final polish before landing the owner identity seam

* chore(auth): final polish before landing the owner identity seam

- Pi routes refuse a rejected credential with the seam's INVALID_CREDENTIAL body.
- Drop the unused LegacyAssetMutations alias and resolveLockWaitMs re-export.
- Scope the document_stages.owner_id deprecation wording: the store no longer
  reads or writes it; the boot backfill reads it and a claim mirrors ownership
  into it for rollback.
- Document OWNER_CLAIM_TRIGGER and the two lock-wait variables in .env.example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(storage): 0.35.1 for the README wording fix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(identity): one identity entry point with composable auth methods

* feat(identity): compose owner auth methods behind one entry point

Owner resolution now asks an ordered list of owner auth methods instead
of one configured authenticator. Each method answers `authenticated`,
`not-applicable` (no credential of its kind) or `invalid` (its credential
is present but bad). Core keeps the first `authenticated` answer, refuses
any `invalid` answer with 401 INVALID_CREDENTIAL without asking later
methods, and when no method applies falls back to the anonymous cookie
owner, or answers 401 when the host turned the fallback off. Route
handlers and Server Actions share the same order and per-request
memoization.

- configureOwnerAuthentication({ methods, anonymousFallback? }) replaces
  configureOwnerAuthenticator: single-shot, sealed on first use, and
  validated at registration and at boot.
- anonymousCookie becomes the built-in fallback method; sharedTeam
  becomes a method, selected by PERSISTENCE_SHARED_OWNER_ID as before
  when nothing is registered, and kept beside host methods only when the
  host lists sharedTeamAuthMethod() last. The variable set beside a
  registration that leaves it out fails the boot.
- Core attaches pendingClaim itself: a host method's non-anonymous
  principal on a request that also carries a valid anonymous cookie. A
  method may not set one. Claim cookie clearing and OWNER_RETIRED
  recovery go through the anonymous method's clearCredential.
- principalFromStoredOwner consults the anonymous method first, then the
  configured methods' describeStoredOwner in order.
- Remove the built-in trusted-proxy header authenticator with
  OWNER_AUTHENTICATOR and every TRUSTED_PROXY_* variable. The boundary
  test now also keeps the resolution internals inside
  lib/server/identity/ and asserts no source reads the removed variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): document owner auth methods and a gateway JWT recipe

Rewrite the owner identity section around the ordered auth methods: the
three answers and what resolution does with each, the anonymous
fallback, registration and its boot checks, and the sharedTeam rules.
Replace the removed built-in gateway authenticator's documentation with
a short host-owned recipe that verifies a gateway-forwarded signed JWT
(oauth2-proxy, Cloudflare Access, IAP) against the identity provider's
keys with the jose library. Update the claim section for the core-owned
claim candidate, and the Chinese README, SECURITY.md, CHANGELOG and the
deployment docs in every locale to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(identity): tighten retired-owner cookies, retired config and the boundary

- On 403 OWNER_RETIRED, core now calls clearCredential only on the
  anonymous fallback and on methods that declare
  `issuesAnonymousOwners: true`, so an account method's session cookie is
  never cleared there. A clearCredential that throws is logged and
  skipped instead of turning the refusal into a 500.
- Setting any retired OWNER_AUTHENTICATOR / TRUSTED_PROXY_* variable fails
  the boot with a pointer to the auth-methods docs, instead of silently
  serving every request anonymously. Only the names are matched, in one
  module the boundary test checks never reads a value.
- Add lib/server/identity/host/ for host auth methods. The boundary test
  now scans core identity files for gateway identity headers too, polices
  reads of an incoming Authorization header everywhere but the host
  directory, and keeps resolution internals out of host code.
- Document on OwnerAuthMethodResult that a present but unusable
  credential must answer `invalid`, never `not-applicable`, and that a
  method that cannot decide throws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): fix the gateway JWT recipe's error split and add notes

- The recipe now answers `invalid` only for token failures (expired,
  claim validation, malformed JWS/JWT, bad signature, no or several
  matching keys, disallowed or unsupported algorithm) and rethrows
  everything else: jose reports a key endpoint that is down, answers
  non-200 or returns unparsable data as a generic JOSEError, which the
  previous version turned into a 401 for every user.
- The recipe file goes in lib/server/identity/host/, with an accurate
  description of what the boundary test enforces.
- Document invalid versus not-applicable for malformed credentials,
  issuesAnonymousOwners, the boot failure on retired variables, and that
  the claim candidate is an unsigned bearer cookie a sibling subdomain can
  shadow (serve on a registrable domain of its own; cookie hardening is
  deferred). Mirror the notes in the Chinese README, SECURITY.md and the
  changelog.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 11:54:26 +08:00
Frank Zhuandwyuc 4b537dca96 fix(tts): pace classroom TTS and classify MiniMax RPM limits (#1650)
* fix(tts): pace classroom TTS and classify MiniMax RPM limits

MiniMax reports RPM limits as HTTP 200 with base_resp.status_code 1002.
Classify that code as TTSRateLimitError, pace classroom TTS with
TTS_MIN_INTERVAL_MS (default 1000ms) and double the gap up to 15000ms on
each rate limit, and store ttsCoverage on the generation job so a partial
narration is not a clean success.

Fixes #1567

* fix(tts): fail closed on MiniMax status and skipped narration

Treat every non-zero MiniMax base_resp status as a failure, even when
audio bytes are present. Classify 1002, 1039, 1041, and 2045 as rate
limits. Floor rate-limit back-off independently of TTS_MIN_INTERVAL_MS,
cap the cumulative wait, and cancel those sleeps on abort. A skipped or
thrown TTS phase reports zero coverage and a warning.

* fix(tts): decay spacing, heartbeat TTS progress, default interval 0

Classroom TTS spacing steps back toward TTS_MIN_INTERVAL_MS after each
successful clip, and that interval defaults to 0 so pacing starts only
after a rate limit. Clip progress is reported through the existing job
onProgress path, so updatedAt keeps moving during a long narration phase.
A Retry-After that does not fit the remaining back-off budget skips that
clip instead of silencing every later clip.

generateTTSForClassroom always returns coverage. A skipped run that has
resolved a provider counts clips after the same speech split as synthesis.

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 13:26:04 +08:00
杨慎andwyuc d04a75cbc6 feat(ai): add Xiaomi MiMo V2.6 models [AI-assisted] (#1655)
* feat(ai): add Xiaomi MiMo V2.6 models

* docs(ai): preserve Xiaomi provider references

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 13:11:49 +08:00
AbelWangYaBoandwyuc 05d782090f feat(persistence): resolve single-tenant requests to a shared owner id (#1639)
Course documents, folders, materials and agent sessions are partitioned by an
owner id that comes from a 30-day anonymous cookie, because there is no host
auth layer. For one team behind one ACCESS_CODE that partition buys nothing —
those visitors already share the site password — and it costs: a second browser
shows an empty course list, and POST /api/stages/[id]/publish refuses every
owner, because publishing a cookie partition is not something the product
allows.

PERSISTENCE_SHARED_OWNER_ID resolves every request to that fixed id instead.
Unset — the default — changes nothing.

- **ACCESS_CODE is required.** Without one the middleware lets every request
  through, so a single shared owner would expose one readable, editable and
  publishable course library to anyone who can reach the deployment. That is
  worse than the per-browser partitioning this setting removes, so the
  combination refuses to start, from instrumentation.ts, with a message naming
  both variables. The alternative the issue allowed — ignore it with a warning —
  was rejected because a silently ignored setting is exactly the confusion this
  feature exists to remove.
- The value is validated rather than trusted: 1-128 characters of
  [A-Za-z0-9._-]. That excludes the reserved `anon:` prefix, which would still
  be refused by publish and would alias onto a cookie owner, and any character
  the material-key sanitiser rewrites, which could collide two ids onto one
  object key. A malformed value fails startup on the same path; an empty value
  is treated as unset, since KEY= in an env file cannot be told apart from an
  operator who meant to leave the feature off.
- An explicit authenticatedOwnerId still outranks it, so adding a real auth
  layer later does not require clearing the variable first.
- All four places that derive an owner honour it, including the Server Action
  in lib/workbench/workspace-actions.ts, which re-reads the cookie itself. That
  one is easy to miss and would make the workspace list and its row actions
  disagree about who owns a session.
- The env docs state the trust model next to the variable, cross-reference
  SECURITY.md, and note that this answers who owns a course while DATABASE_URL
  answers where one lives — a second browser sees an empty list on Postgres too.

This is the stopgap asked for in #1550 — a PR in the spirit of the closed
#1551 — with threading a real authenticated owner left as the long-term fix.

AI-assisted commit

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 17:33:02 +08:00
LING DUANandwyuc 7be21f2d89 fix(chat): pin Pi routing and preserve provider errors (#1637)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 15:18:54 +08:00
LING DUAN 2a77a8a476 feat(chat): make Pi classroom runtime the default (#1628)
* feat(chat): make Pi classroom runtime the default

* chore(chat): align docs and E2E with Pi default
2026-09-22 11:41:13 +08:00
Huang Geyangandwyuc 67f568848a fix(server): warn when access-code protection is disabled (#1599)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 19:14:38 +08:00
Frank_zhuandwyuc 9cd8051461 docs: align security and behaviour claims with shipped code (#1592)
ACCESS_CODE unset remains fail-open in middleware, document reads are
capability-by-id via the anonymous owner cookie (not x-learner-key),
and the README action/skill counts match the Action union and
skills/agent-runtime.

Closes #1587

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 16:39:12 +08:00
a8726ec4f3 feat(media): add OpenRouter image and video providers (#1356)
* feat(media): add OpenRouter image and video providers

OpenMAIC ships six separate video providers (Veo, Kling, Seedance,
MiniMax, Grok, HappyHorse) and seven image providers, each needing its
own key. OpenRouter fronts those same model families behind one key and
one account, so this adds it as a provider on both sides.

Both use OpenRouter's dedicated media endpoints, not chat-completions:

- Image: POST /images -> { data: [{ b64_json }] }
- Video: POST /videos -> 202 { id, status }, poll GET /videos/{id},
  then GET /videos/{id}/content for the mp4 bytes

The model list is fetched live from GET /images/models and
GET /videos/models through /api/openrouter-models rather than pinned in
the registry: OpenRouter hosts 48 image and 28 video models today and
adds more, so a hardcoded shortlist would decide for the operator which
models exist. The registry keeps a three-entry seed as an offline
fallback, and the existing custom-model UI still accepts any model id.
Both catalogs answer unauthenticated, so the picker fills before a key
is pasted; a key is forwarded when present for proxied base URLs.

Adapter contracts are covered by stubbed-fetch tests (request shape,
empty-response handling, and the video job state machine including
terminal failure). No test performs a billable call.

Closes #1355

* fix(media): validate the key and tolerate a pasted endpoint URL

Three fixes found while configuring the new provider:

1. Both connectivity probes hit the model catalogs, which answer 200
   unauthenticated — so "Test Connection" reported success for any
   string, including an invalid key. Probe GET /key instead: equally
   cheap, and it actually rejects a bad key.

2. The settings field is labelled "Base URL" but the panel echoes it
   back as "Request URL", so pasting the full endpoint
   (https://openrouter.ai/api/v1/images) is the natural mistake. That
   built /api/v1/images/images and 404'd. Trim a trailing slash and a
   trailing /images or /videos so both forms work; a proxy path that
   merely contains the word is left alone.

3. The image and video settings panels read `data.message` on a failed
   test, but failures answer with `error` (apiError) and only successes
   carry `message`. Every failing connectivity test — for any provider,
   not just OpenRouter — rendered "connection failed: undefined" instead
   of the reason. Pre-existing; surfaced by 1 and 2 above.

Closes #1355

* fix(media): make every OpenRouter model selectable, and always select a provider

Two gaps found while configuring the new provider.

The settings Models list is a read-only catalog for every provider; the
actual model picker is the media popover. That picker built its groups
from the static registry array, so OpenRouter offered only the
three-entry seed while settings listed the full live catalog — the
models were visible but not choosable. Feed the same live catalog into
the popover, fetched only once the provider is usable so an
unconfigured install makes no request.

Separately, `imageProviderId`/`videoProviderId` are empty until a
provider is chosen (first-run auto-config leaves them blank when the
server reports no media provider). Opening the settings panel on an
empty id selected nothing: the header rendered the missing name key as
"settings.undefined", and Test Connection posted a blank
x-image-provider/x-video-provider, so it failed with "No image/video
provider configured" whatever key was typed. Fall back to the first
catalog entry so the panel always has a selection. Pre-existing and not
specific to OpenRouter.

Closes #1355

* fix(tts): request a browser-playable format from custom providers

`generateOpenAITTS` serves every custom OpenAI-compatible TTS provider but
never sent `response_format`, so it inherited whatever each provider
defaults to. OpenAI defaults to mp3; OpenRouter's /audio/speech defaults
to raw `pcm`. The unknown content type then fell through to the `'mp3'`
default below, the client built `data:audio/mp3;base64,…` from headerless
PCM samples, and playback failed with "no supported source was found" —
while the server logged a clean 200, because the audio really was
generated. Name the format instead of inheriting it.

Also stop mislabelling an unrecognised body: `pcm`/`l16` now raises a
message naming the cause, and `aac`/`opus` are recognised.

Two supporting fixes:

- /api/openrouter-models normalises its base URL the way the adapters do
  and falls back to the public catalog when a custom base URL fails, so a
  typo in a free-text settings field cannot empty the model picker. Also
  types the headers object so tsc accepts the conditional.
- provider-neutrality-guard pins exact per-vendor occurrence counts in
  lib/server/provider-config.ts. Adding the image and video env entries
  raises "openrouter" from 2 to 6 (each entry contributes both its key and
  its value); CI failed without the bump.

Closes #1355

* fix(security): never send the operator key to a client-chosen host

Review found `/api/openrouter-models` was an SSRF and key-exfiltration
path, and the finding is correct. The route took `x-base-url` from the
caller at highest precedence while preferring the *server* env key, so any
caller could make the server send the operator's OpenRouter credential as
an `Authorization: Bearer` header to an arbitrary URL. The route's own
comment claimed it followed `/api/verify-image-provider`; that pattern
runs `validateUrlForSSRF` on client base URLs, and this route did not.

The boundary is now explicit: the server key travels only to the
operator's own base URL. A client-supplied URL is SSRF-validated and
carries only that caller's own `x-api-key` — the server key is dropped —
and the unauthenticated public-catalog fallback never forwards a
credential chosen for a different host. Redirects are no longer followed
(`redirect: 'manual'`), since a redirect would carry the Authorization
header off-host and reopen the same hole, and upstream reads are bounded
by a timeout.

The per-URL cache is now keyed by destination *and* a hash of the
credential, and bounded to 64 entries with oldest-first eviction, so
client-supplied URLs cannot grow it without limit and one caller's
key-authorised catalog is never served to another.

Also from the review:

- The image adapter discarded the reported `media_type`. The
  orchestration layer wraps a bare `base64` as `data:image/png`
  unconditionally, so jpeg/webp results were mislabelled; the adapter now
  returns a data URL carrying the real type.
- Adapter generation and poll requests set `redirect: 'manual'`, matching
  the `/key` probe that already did.
- `runPolledTask` accepts an `AbortSignal` so the sleep between polls is
  cancellable; the video adapter passes the caller's signal. Without it a
  cancelled generation still slept out a full 10s interval.

Tests cover the highest-risk paths the review named: which credential
reaches which URL, that an SSRF-rejected destination is never contacted,
that the fallback is unauthenticated, cache isolation between callers,
and MIME preservation.

Findings 2 (base-URL normalisation) and 4 (neutrality-guard debt) were
already fixed in d553a08, pushed after the review was submitted; CI is
green on that commit.

Closes #1355

* ci: retry flaky voice clone timeout

* fix(vercel): keep OpenRouter catalogs within Hobby function limit

* fix(vercel): avoid tracing self-hosted sharp binaries

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-18 00:01:48 +08:00
wyucandClaude Opus 5 2cbd011c1f feat(token-plan): add TokenDance one-key preset for every modality (#1525)
* feat(token-plan): add TokenDance one-key preset for every modality

TokenDance is a model gateway: chat and images are OpenAI-compatible at
/gateway/v1, and the same key authenticates vendor-protocol routes on the
same host (Ark, MiniMax, Bocha). The preset reuses the existing adapters
with those route prefixes as base URLs, so one key lights up LLM, image,
video, TTS and web search from Settings -> Token Plan.

- providers: add a built-in `tokendance` OpenAI-compatible provider
  (TOKENDANCE_* env prefix, logo, provider name in all locales)
- token-plan: add the TokenDance preset (Seedream image, MiniMax H3 video,
  MiniMax speech TTS, Bocha web search)
- seedream: use a base URL that already ends in a version segment verbatim,
  so gateway routes like `/ark/v3` do not get `/api/v3` appended
- minimax-video: route H3-family models through the v2 task API (content
  array submit, task-envelope poll); connectivity checks for H3 probe auth
  on the v2 query route instead of submitting a billable task
- README: add a one-key quick example and replace the Gemini-specific
  model recommendation with a provider-agnostic setup recommendation

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL

* fix(token-plan): accept preset web-search base URLs and report H3 dimensions per ratio

- web-search: the client base URL allowlist also accepts the exact base URL
  a built-in token plan preset writes for that provider, derived from
  TOKEN_PLAN_PRESETS. Applying a plan whose web-search route is not an
  official vendor host previously stored a URL that the route rejected with
  400. Any other client URL is still rejected.
- minimax-video: report H3 v2 clip dimensions for 16:9, 9:16, 4:3 and 1:1
  instead of assuming landscape for every non-portrait ratio.
- tests: pin the allowlist for every preset, the 1:1 H3 dimensions, and
  clear TOKENDANCE_* in the provider-config env isolation list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:28:09 +08:00
wyucandClaude Opus 4.8 ee79762ef9 fix(access-code): expire tokens server-side and throttle verification behind a trusted proxy (#1513)
When ACCESS_CODE is set, verification tokens (timestamp.HMAC) never expired
because neither verifier checked the timestamp, and POST /api/access-code/verify
had no attempt throttling.

- Enforce a 7-day token lifetime in a shared, Edge-safe module used by both the
  Node verifier and the middleware Web-Crypto verifier, and reject non-canonical
  signatures in both.
- Rate-limit verification only when the client identity is trusted
  (TRUST_PROXY_HEADERS=true): a per-client sliding window (10 failures / 60s)
  whose attempt is reserved atomically at check time, returning 429 with
  Retry-After. Without a trusted proxy the app cannot attribute requests to a
  client, so no shared throttle is applied — a long random ACCESS_CODE is the
  protection, and a warning is logged when it is short.
- Store bounded, copied identity keys so a large forwarding header cannot retain
  memory.
- Document the behavior in .env.example, README, and configuration docs.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-15 17:47:14 +08:00
wyuc c605e9e6ef feat(persistence): turn on the server-owned asset lifecycle and release assets on course deletion (#1007 amendment, part 2) (#1473)
App wiring for the @openmaic/storage 0.31.0 lifecycle: reference tracking and document references are paired unconditionally and declared at startup, course deletion withdraws references inside the tombstone transaction, ASSET_PENDING_TTL_MS configures the pending window, and the dead client-side reclamation code is removed.
2026-09-15 09:55:26 +02:00
Nguyen Quang Thiepandwyuc 3e3cac5f76 fix: enable Pro workbench flag in Docker builds; allow non-TLS localhost cookie (#1484)
* fix: enable Pro workbench flag in Docker builds; allow non-TLS localhost cookie

Two deployment fixes discovered while self-hosting with Docker Compose:

1. Expose NEXT_PUBLIC_PRO_WORKBENCH_ENABLED as a Docker build arg.
   Persistence and other NEXT_PUBLIC_* flags are already wired through
   Dockerfile + docker-compose.yml, but the Pro workbench entry flag was
   missing, so Docker deployments could not enable the workbench at all.

2. Add COOKIE_SECURE opt-out for the anonymous owner cookie.
   Production builds always set `Secure` on the anonymous_id cookie. Safari
   refuses to store Secure cookies served over plain http://localhost (it
   does not special-case localhost like Chromium/Firefox), so every request
   minted a fresh anonymous owner and owner-scoped document writes were
   rejected with 403 — scenes never persisted and generation appeared to
   hang after the first scene. COOKIE_SECURE=0 lets plain-HTTP deployments
   opt out; the default (Secure in production) is unchanged.

Documents COOKIE_SECURE in .env.example alongside the persistence opt-ins.

* fix: share the COOKIE_SECURE opt-out with the Server Action cookie mint

Review follow-up on the COOKIE_SECURE opt-out.

- Extract anonymousCookieSecure() in owner.ts and use it in
  lib/workbench/workspace-actions.ts, which re-implements the anonymous_id
  mint: with COOKIE_SECURE=0 that path still sent a Secure cookie, so a
  Server Action running before any /api/agent/* request (e.g. deleting a
  workspace session) minted an ephemeral owner in Safari.
- State the opt-out accurately in owner.ts; the previous comment claimed the
  flag never forces Secure off.
- Add the regression test beside the existing production case (verified it
  fails when the opt-out is dropped).
- .env.example: document the security cost of dropping Secure and that only
  the exact value 0 disables it.

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-14 17:01:59 +02:00
wyuc 68aed2932a fix(proxy-media): connect only to addresses the SSRF guard validated (#1488)
* fix(proxy-media): connect only to addresses the SSRF guard validated

The route validated a URL with validateUrlForSSRF and then handed the
hostname to fetch(), which resolved it again at connect time. An
attacker-controlled name could answer a public address to the guard and an
internal address (loopback, RFC1918, link-local, cloud metadata) to the
socket, and every redirect hop repeated the same validate-then-refetch gap.

Add a shared policy-aware pinned dispatcher
(lib/server/pinned-dispatcher.ts) whose connect.lookup resolves all
addresses, runs the same local-network policy as validateUrlForSSRF over
every answer (cloud metadata is always blocked, even with
ALLOW_LOCAL_NETWORKS=true), and passes only the vetted list to the
connection. proxy-media creates one dispatcher per request and uses it on
every hop, mapping a connect-time refusal to 403 INVALID_URL instead of a
generic 500. The agent-runtime pinned agent now reuses the same builder
with its strict assertSafeIp policy unchanged.

The other routes that validate with validateUrlForSSRF and then hand the
URL to provider SDKs (parse-pdf, transcription, generate/*, verify-*,
probe-models, azure-voices) are not covered; that is a follow-up.

Tests: new tests/server/proxy-media-rebinding.test.ts covers the two-answer
rebind, the ALLOW_LOCAL_NETWORKS opt-in through the pinned path, metadata
refusal under the opt-in, and a redirect-hop rebind.

* fix(proxy-media): destroy the per-request dispatcher and align non-unicast policy across layers

Follow-up hardening for the proxy-media SSRF pin.

1. Handler hang: the route awaited `dispatcher.close()` in `finally`.
   undici's `close()` drains in-flight requests, and the response body is
   intentionally left unread on the 30x (`redirect: 'manual'`), non-2xx,
   oversize and connect-refusal paths, so an upstream that keeps its body open
   never completes and the POST never settles. Tear the pool down with a
   non-awaited `destroy()` instead; the response blob is already in memory by
   the time `finally` runs. Regression tests drive a loopback server that
   answers 404, a 30x, and an oversize 200 while trickling its body forever and
   assert the route settles in well under a second.

2. Headline rebinding test sensitivity: install an attacker-controlled global
   dispatcher (`setGlobalDispatcher` plus a lookup that always answers
   loopback) in `beforeEach` so the unpinned fetch path really reaches the
   loopback server. Removing the route's dispatcher now fails on
   `internal.requests() === 0` rather than on an unrelated DNS resolution
   error. The positive control also asserts the pinned connect lookup was
   invoked, so a runtime that silently ignores `init.dispatcher` goes red.

3. Non-unicast policy parity: `validateUrlForSSRF` only classified IP-literal
   URLs through `isPrivateIP`, so a literal in CGNAT 100.64/10 or an IANA
   reserved/special-use range (240/4, 198.18/15, the TEST-NET blocks,
   192.0.0.0/24, multicast, broadcast) passed the URL layer while the same
   address reached through a hostname was refused at connect time. Both layers
   now share `isNeverAllowedRange`, applied to IP-literal URLs and to every
   resolved answer, with or without `ALLOW_LOCAL_NETWORKS`.
   `connectionAddressBlockReason` drops its blunt `range() !== 'unicast'` test
   so the two layers agree exactly, while private, loopback and link-local
   answers keep following the opt-in.

   Behaviour change: a hostname resolving into CGNAT 100.64/10 or an IANA
   reserved/special-use range is now refused (403 INVALID_URL) where it was
   previously proxied, including under `ALLOW_LOCAL_NETWORKS=true`, since these
   ranges are never legitimate proxy targets.

4. Do not echo an unparseable connect-time answer in the refusal. Return the
   generic block message instead of `Unable to classify network address:
   <value>`; the route relays that message to the client in a 403 body.

Tests: new trickle-body regression tests and CGNAT/reserved cases for both
literals and resolved answers, plus a `connectionAddressBlockReason` parity
suite. One existing fixture is updated from the IANA documentation prefix
2001:db8::/32 (reserved) to a real public prefix 2001:4860::/32.

* fix(ssrf-guard): let the local-network opt-in cover CGNAT and decode local-use NAT64 prefixes
2026-09-13 18:00:56 +02:00
wyuc 1fd4348f21 fix(classroom): generate classroom ids server-side and create files exclusively (#1489)
* fix(classroom): generate classroom ids server-side and create files exclusively

The legacy file-backed classroom store accepted a caller-chosen stage.id and
renamed a temp file over the target, so anyone who knew a public share URL
could POST the same id and replace that classroom's content. POST
/api/classroom now mints the id itself (nanoid, 10 characters, matching the
generation pipeline), ignores any client-supplied stage.id, and rebinds the
stage and every scene to the generated id. Both the route and
classroom-generation persist through a new exclusive create that hard-links a
temp file into place, so an existing classroom is never replaced; EEXIST is
retried with a fresh id a bounded number of times and then reported as 409.

Unchanged: the server-backed persistence path (app/api/persistence,
lib/persistence), the read/GET path, middleware, and the access-code gate;
writeJsonFileAtomic keeps its overwrite semantics for the other callers.
Adds regression tests pinning non-overwrite, the typed EEXIST error, collision
retry/exhaustion, and generation-path exclusivity.

* fix(classroom): fall back to exclusive open without hard links and reserve ids before media generation

writeJsonFileExclusive kept `fs.link` as its only route to the destination, so
classroom creation failed hard on mounts without hard links (gcsfuse, s3fs and
some FUSE gateways return ENOSYS/ENOTSUP/EOPNOTSUPP/EPERM/EXDEV). Keep `link`
as the fast path and, for exactly those codes, fall back to an exclusive `wx`
open: still never replaces an existing file and still maps EEXIST to
ClassroomAlreadyExistsError. Any other code keeps propagating, and the temp
file is always removed.

The generation pipeline wrote media and TTS into
<CLASSROOMS_DIR>/<id>/{media,audio} before the exclusive create, so a collision
would have landed the new classroom's files in an existing classroom's
directory and the retried document's URLs would still point at that other id.
Reserve the id first by exclusively creating the classroom file with a
placeholder document (`reserved: true`, empty scenes). The file is the only
token that atomically covers the whole collision namespace: a classroom created
through POST /api/classroom has a JSON file but no directory, so reserving a
directory would miss it. readClassroom treats a reserved document as absent, so
an in-flight or crashed reservation is never served as an empty classroom. The
retry on EEXIST now happens at reservation time (bounded, before any media),
and the final persist is an ordinary overwrite of the id the process owns, so
the post-media id retry is gone.

Document OPENMAIC_CLASSROOMS_DIR in .env.example, including that
CLASSROOM_JOBS_DIR does not move with it.

* fix(classroom): release an unused reservation when generation fails
2026-09-13 17:49:57 +02:00
LING DUANandwyuc ebf665f316 feat(playback): reference static GenUI components (#1281)
* feat(playback): reference static GenUI components

* feat(playback): gate courseware references

* fix(pi): bound Interactive source HTML before parsing

* fix(pi): bound interactive reference HTML processing

* fix(playback): close interactive reference review gaps

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-11 19:16:10 +02:00
wyucandClaude Fable 5.1 29735f10d0 fix(ssrf): keep cloud metadata endpoints blocked under ALLOW_LOCAL_NETWORKS (#1419)
* fix(ssrf): keep cloud metadata endpoints blocked under ALLOW_LOCAL_NETWORKS

ALLOW_LOCAL_NETWORKS=true exists so self-hosted deployments can point
providers at loopback, RFC1918 and split-horizon targets. It also
returned early from validateUrlForSSRF before any classification, so a
client-supplied base URL of 169.254.169.254 (or metadata.google.internal,
100.100.100.200, fd00:ec2::254) was accepted on such deployments.

Cloud instance-metadata endpoints are now rejected regardless of the
flag: literal hosts and IPv4-mapped forms are classified directly, and
non-IP hostnames are resolved so an answer that lands on a metadata
address is rejected too. DNS failure under the flag still fails open,
as before, because split-horizon DNS is an explicit use case of the flag.
Without the flag the DNS path now also rejects answers on metadata
addresses that are not RFC1918 (100.100.100.200).

The route tests that used the metadata address as their example of a
target the flag allows now use a private-network address instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GRaNc5E88r41GUWt2Y3uiQ

* fix(ssrf): widen the metadata set and bound the flagged DNS lookup

Review follow-up. Add the AWS ECS task credential and EKS Pod Identity
endpoints (169.254.170.2, 169.254.170.23, fd00:ec2::23), the Azure
WireServer address (168.63.129.16) and the legacy OCI IMDS address
(192.0.0.192) to the blocked set; the last two are neither RFC1918 nor
link-local, so they were reachable even without the flag. Recognise the
metadata addresses when carried inside 6to4, Teredo, ISATAP and NAT64
literals. Bound the DNS lookup done under the flag to three seconds and
fail open on expiry, matching the existing fail-open on error. Say in
.env.example that the set is a fixed list and that DNS failure is
allowed through.

Use the same private-network fixture for the reject and allow cases of
the four route guard tests so a partial NODE_ENV re-gate of the guard
goes red again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GRaNc5E88r41GUWt2Y3uiQ

* fix(ssrf): share the tunnel decoder between the private and metadata checks

Review follow-up. assertSafeIp now classifies metadata addresses through
isCloudMetadataAddress, so an ISATAP identifier under a globally
routable prefix that carries 168.63.129.16, 192.0.0.192 or
100.100.100.200 is rejected on the strict-fetch path too. isPrivateIP
uses the same tunnelEmbeddedIPv4 helper instead of its own copies of the
6to4, Teredo and ISATAP decoders, which also gives it NAT64. Tests cover
the globally-routable ISATAP forms, the NAT64 private case, and the
false-positive direction (tunnel literals carrying public or RFC1918
addresses stay allowed under the flag).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GRaNc5E88r41GUWt2Y3uiQ

* fix(ssrf): reject metadata hostnames before the flag branch

Review follow-up. The metadata hostname and address check now runs
before the ALLOW_LOCAL_NETWORKS branch, so metadata.google.internal is
rejected by name in both flag states (previously only under the flag),
and metadata literals get the metadata message rather than the one that
suggests setting the flag. Pin the tunnel decoder boundaries in tests
(2001:db8 is not Teredo, 64:ff9b:1 is not NAT64, 2003 is not 6to4).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GRaNc5E88r41GUWt2Y3uiQ

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-08 19:44:20 +02:00
wyucandwyuc 0bab621e09 harden input validation and outbound URL handling (#1387)
* harden input validation and outbound URL handling

- Validate the classroom id on the write path with the same allowlist the
  read path already applies, and assert in the storage layer that a resolved
  classroom file stays inside CLASSROOMS_DIR.
- Apply validateUrlForSSRF at every call site regardless of NODE_ENV; the
  guard already owns the ALLOW_LOCAL_NETWORKS escape hatch, so the extra
  environment condition only disabled the check outside production builds.
  Add a repository-scanning test so a gated call site cannot reappear.
- Re-validate every redirect hop of an outbound provider request through a
  shared transport, mirroring the per-hop loop proxy-media already uses.
- Restrict stored slide HTML to the formatting vocabulary the renderer
  produces, sanitizing at the classroom persistence boundary on both write
  and read, with a separate policy for KaTeX snapshots.
- Refuse the development persistence authenticator under NODE_ENV=production
  unless an explicit opt-in is set, and document it in .env.example.

* drop credential headers on a cross-origin redirect hop

The manual redirect loop reused the caller's init on every hop, so provider
credentials were re-sent to the redirect target even on a different origin.
The platform fetch drops Authorization itself when it follows redirects, so
mirror that: strip authorization, api-key, x-api-key and x-goog-api-key
before a cross-origin hop, matching case-insensitively and preserving the
shape of the caller's headers. Same-origin hops are untouched.

A streaming request body cannot be replayed onto the next hop, so fail with
a clear message rather than sending an empty body.

---------

Co-authored-by: wyuc <dhq1024@proton.me>
2026-09-05 23:56:37 -07:00
92d8b32a5a feat(deepseek): register deepseek-v4-flash-vision-exp with vision cap… (#1329)
* feat(deepseek): register deepseek-v4-flash-vision-exp with vision capability

DeepSeek's docs list deepseek-v4-flash-vision-exp as an official model
on the same OpenAI-compatible base URL (https://api.deepseek.com):
1M context, 384K max output, JSON output / tool calls / Responses API /
Anthropic API all supported, priced same as deepseek-v4-flash.

Register it in the deepseek catalog with vision: true so scene
generation can pass document images as vision content parts instead of
silently dropping them for this provider.

* fix(deepseek): register vision model in thinking metadata and update pinned lists

- add deepseek-v4-flash-vision-exp to THINKING_CAPABILITIES (deepseekEffort,
  same as v4-flash per official docs) so the model-metadata drift test passes
- update the pinned deepseekModels list in thinking-config.test.ts to include
  the new model

* docs(env): include deepseek-v4-flash-vision-exp in DEEPSEEK_MODELS example

* Revert "docs(env): include deepseek-v4-flash-vision-exp in DEEPSEEK_MODELS example"

This reverts commit 8040bac441.

* style(deepseek): apply prettier formatting

* test(deepseek): pin vision model capabilities; document vision-exp in env example

Review asks from #1329:
- assert deepseek-v4-flash-vision-exp capabilities (vision: true,
  streaming, tools) so a regression to vision: false fails the suite
- add the model to the DEEPSEEK_MODELS example in .env.example

---------

Co-authored-by: CrazyPandaP <CrazyPandaP@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-05 23:11:45 -07:00
Yizuki_Ame eb2cbc701a feat(pro): generate concise conversation titles (#1275)
* feat(storage): add automatic conversation title lifecycle

* feat(agent): generate concise conversation titles

* feat(agent): schedule automatic conversation titles

* fix(agent): harden automatic title lifecycle

* fix(agent): refine conversation title prompt

* fix(agent): harden automatic title boundaries
2026-09-04 10:52:25 -04:00
dd74cfe23b feat(web-search): add Exa provider (#1342)
Co-authored-by: white638 <212083322+white638@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-03 06:14:05 -04:00
wyucandClaude Fable 5 04621578de release: OpenMAIC 1.0.0 — the agent workbench (#1228)
* feat(storage): add an agent-session store with PG backend and layered contracts (#1163)

* feat(storage): add agent-session store with PG backend and layered contracts

* test(storage): avoid BigInt literals for pre-ES2020 root typecheck

* fix(storage): close agent-session store review findings

* docs(storage): align hook ordering and contention-probe claims with the code

* ci: run on the agent-workbench integration branch

* chore(storage): bump to 0.5.0 for the agent-session store

* fix(storage): carry replay compaction across page boundaries

* feat(agent): add the driver model contract and stage route dialect (#1165)

* feat(agent): add the driver model contract and stage route dialect

* fix(agent): validate route context windows and clarify dialect precedence

* feat(agent): adapt the agent-session store and runtime foundations (#1167)

* feat(agent): adapt the agent-session store and runtime foundations

* feat(agent): resolve request owner identity via an anonymous cookie

* docs(agent): document the opt-in compaction default and harden edge cases

* feat(agent): add the background session runner (#1169)

* feat(agent): add the background session runner

* feat(agent): wire the runner into startup behind feature flags

* fix(agent): stop clean interruptions from consuming the attempt budget

* fix(storage): charge the attempt budget for abandoned leases but not clean parks

* docs(storage): document the attempt-charging contract and decouple its tests

* feat(agent): add agent session and owner event streams (#1170)

* feat(agent): add agent session and owner event streams

* fix(agent): close the session-existence oracle and document the owner seam

* feat(agent): add agent session lifecycle routes (#1171)

* feat(agent): add agent session lifecycle routes

* fix(agent): validate session-create input and preserve the owner cookie on errors

* refactor(storage): drop the unused active-stage API from the agent-session contract (#1174)

* refactor(storage): drop the unused active-stage API from the agent-session contract

Tools address stages explicitly on every call, so the store keeps no
mutable session-level stage pointer. Removes resolveActiveStage and
setActiveStage from the store interface, their PG implementations, the
active_stage_changed lifecycle event, the session_active_stage owner
event variant, and the contract tests pinning them. The active_stage_id
column and the DDL check constraint stay untouched for schema
compatibility.

* chore(storage): bump @openmaic/storage to 0.7.0 for the contract removal

* docs: document the agent runtime configuration surface (#1176)

* fix(agent): repair orphaned and late tool results across interruption boundaries (#1180)

* fix(agent): repair orphaned and late tool results across interruption boundaries

A crash, shutdown, or provider failure can leave the durable transcript with
tool calls that have no result, or with results ordered illegally for the
provider. Three failure modes were fixed:

- Orphaned tool calls: a run that died between an assistant tool-call frame
  and its result left a dangling call in the entry tree. Resume no longer
  synthesizes and persists receipts for it: interrupted results are a
  read-time provider view owned by a shared read-boundary repair, which
  returns the original array for a healthy transcript and never mutates the
  tree.

- Late parallel results: a parallel tool can finish while pi unwinds an
  aborted assistant frame, leaving result(A), assistant(aborted), result(B)
  in durable order. Strict providers reject non-contiguous results, so the
  read-boundary repair moves existing results next to their owning assistant
  frame (in call order), omits incomplete unwind frames, and synthesizes
  receipts only for genuinely missing calls.

- Interrupted calls at the write boundary: a call still in flight when the
  run winds down (shutdown, lease loss, cancellation, provider failure) had
  no receipt at all. The runner now tracks in-flight calls from their
  assistant frames and, before the terminal flush, appends an
  interrupted-result receipt for each still-orphaned call through the same
  attempt-fenced write chain, so a lease-stealing zombie never writes and
  the next claim sees a provider-safe transcript.

* test(agent): pin the runner wiring for interruption-boundary tool repair

* feat(agent): add neutral tool foundation libraries (#1184)

* feat(agent): register a web_search tool on the session runner (#1185)

* feat(storage): add a per-session URL trust gate (#1186)

* feat(agent): add the skills system (#1189)

* feat(agent): add the skills system (builtin directories and durable user skills)

* fix(storage): serialize the user-skill quota check-and-insert per owner

Two concurrent creates at the 50-skill boundary both counted 49 rows and
both inserted (READ COMMITTED, no lock), overshooting the quota contract.
The create transaction now takes a per-owner pg_advisory_xact_lock first,
and the same-name idempotency check runs before the count check so an
at-least-once retry of the create that committed as the owner's 50th row
still returns its durable receipt instead of a quota error. The 23505
backstop is retained for writes that do not take the lock.

* fix(agent): share unstorable-character validation and align skill lookup

* feat(agent): add session materials and a fetch_url tool behind the URL trust gate (#1190)

* feat(agent): add session materials and a fetch_url tool behind the URL trust gate

* fix(agent): harden session material fetching

* feat(storage): add an ownership scope to stage documents (#1191)

* feat(agent): add material read and search tools (#1192)

* feat(agent): add stage read and patch tools (#1194)

* feat(agent): add page generation and deck editing tools (#1198)

* test(storage): keep the PG contract suite order-independent (#1200)

* fix(agent): revoke deleted-session URL authority and reject private ISATAP endpoints (#1199)

* fix(storage): revoke deleted session URL authority

* fix(ssrf): reject private ISATAP endpoints in strict fetches

* chore(storage): bump to 0.11.1 for the session-URL authority fix

* feat(agent): add roster and voice registration tools (#1201)

* feat(agent): add folder organisation tools (#1202)

* feat(api): add stage and material HTTP routes (#1203)

* feat(workbench): add the client data layer (#1204)

* feat(workbench): add the client data layer

* docs(workbench): write the ported comments in English

* chore(edit): remove the in-editor agent panel (#1210)

* chore(edit): remove the in-editor agent panel

* style: apply prettier formatting

* fix(agent): report the runtime as unusable without a database (#1207)

* fix(agent): report the runtime as unusable without a database

* style: apply prettier formatting

* feat(agent): add image, video and pptx import tools (#1211)

* feat(workbench): add the agent chat surface (#1205)

* feat(workbench): add the agent chat surface

* docs(workbench): write the ported comments in English

* fix(workbench): label the folder and rename tools on the timeline

* fix(workbench): label the roster and voice tools on the timeline

The reconciliation test iterates every tool the runner registers and
requires a display label of its own. The roster and voice-clone tools
(list_voices, set_roster, clip_audio, register_voice) reached the
integration base with the roster/voice-registration tools but never
gained presentation rows, so they fell through to the default branch
and rendered their wire names. Port their rows from the reference
implementation (labels and i18n keys verbatim) and extend the
reconciliation allowlist with ROSTER_TOOL_NAMES and
VOICE_CLONE_TOOL_NAMES, so a future tool cannot enter the product
without a label.

* feat(agent): add the material extraction lifecycle (#1212)

* feat(storage): add material extraction lifecycle

* feat(agent): execute queued material extraction

* style: apply prettier formatting

* style: satisfy prefer-const in the extraction runner

* test: give material fixtures the extraction lifecycle fields

The media-tools slice and the extraction lifecycle slice were each green
in isolation but never compiled together: the lifecycle made derivedFrom
and extraction required on AgentSessionMaterial while the media-tool
fixtures predate them.

* chore: remove stray task notes

* fix(workbench): label the extraction lifecycle tools on the timeline

* feat(workbench): add the workspace shell (#1206)

* feat(workbench): add the workspace shell

* docs(workbench): write the ported comments in English

* i18n(workbench): align workspace keys across locales

* fix(workbench): adopt the landed data layer and label the extraction tools

- replace the sibling-slice seam stubs with the real data-layer modules
- drop ambient declarations now shadowed by landed files
- port timeline labels for the extraction lifecycle tools from the reference
- align the new i18n keys across all locales

* ci: retrigger

* feat(api): folder routes, stage-meta viewer surfaces, and the material upload contract (#1215)

* fix(storage): restore capability-based stage access

* fix(api): bind document access to request owner

* fix(agent): restore three-state stage access on the tool layer

Port probeStageAccess and the three-state StageAccess (owned / foreign /
missing / tombstoned) and gate every stageId-bearing stage tool on an owned
probe, mirroring the reference per tool:

- move_to_folder, rename_stage, read_stage_outline refuse a non-owned stage
  with the single not-yours message before touching the store.
- The course/DSL toolset and the roster toolset are wrapped by
  withOwnerStageAuthorization: read_stage, patch_stage, grep_stage and every
  writer refuse a foreign stage with the same message and refusal shape.
- Scene preview keeps its own probe and its own refusal text, and is
  registered beside the course toolset (never double-gated).
- The runner injects one probe factory at the three call sites.

Tests: the dsl cross-owner test premise (a foreign stage is readable by id)
encoded an invented capability-read policy that the reference does not have
at the tool layer; it now asserts foreign read/patch/grep are all refused
while the owner still reads. Curriculum cross-owner assertions were already
the reference's and now pass with the probes in place.

* docs: correct per-file test counts in the fidelity report

* test: fix type errors in stage-access fidelity test

* test: adapt media-tool and gate suites to the owner-scoped store seam

* feat(api): add owner-scoped course-folder HTTP routes

Port the reference implementation's /api/folders family (list, create,
rename, delete with ungroup/remove modes, and folder membership) onto the
owner-bound document store, replacing its provider-based auth with the
existing withRequestOwnerId / owner-scoped store seams.

The storage package's folder store grows the pieces the routes need:
DocumentFolder.order (schema column + max+1 assignment + ordering),
renameFolder, deleteFolder(mode) with captured member ids, and
setStageFolder(stageId, folderId | null) with idempotent un-filing.
FolderNameError moves into folder-name-validation.ts (stage-storage re-exports
it, keeping import sites intact).

Every route gates on the configured agent runtime (plain 404 when off or
unconfigured), keeps the reference's machine codes and envelopes, and is
covered by gate tests plus a behavior suite.

* feat(api): add stage-meta viewer surfaces for the classroom

Port the reference implementation's viewer-facing stage state — can-edit /
collected / published / generation-complete — on top of the stage-access
base (stage_meta + tombstones). stage_meta gains published_at and
generation_complete columns plus a stage_bookmarks table; the reference's
deployment-specific origin/claimed_at columns are stripped.

New gated routes: GET /api/stage-meta/[stageId] (per-viewer facts, 404 for
absent/tombstoned, never returns the owner id), GET /api/stages/[id]/status,
POST generation-complete / publish / unpublish (owner-only), POST
/api/bookmarks. The resolver lives in lib/server/stage-access.ts.

Wiring: a fetchStageMeta client with the reference's three-outcome contract,
stage-store isOwner/isBookmarked/readOnly fields (upstream single-user
defaults, no-op until the sidecar answers) plus setViewerAccess, the
classroom apply path computing readOnly = !(isOwner || isBookmarked), the
Stage editability gate, and a sidecar probe after each classroom load. A
sidecar 'absent' answer keeps the editable default here because the
classroom also serves local-only courses; server writes stay owner-enforced.

* feat(api): port the reference material upload contract

Rewrite POST /api/materials to the reference implementation's upload shape
so the workbench uploader (uploadWorkbenchMaterial, which posts no session
id and expects a flat 201 view) works unchanged: owner-scoped upload with
mime normalization/validation (415), per-class size caps checked on the
declared content-length and the streamed body (413), empty body (400),
quota (429), sha256 reserve->store->finalize lifecycle with abandon on
failure, flat { materialId, originalName, bytes, mime, extraction } 201,
and an x-request-id echo.

Adds the owner-scoped material library (owner_material table + quota +
24h lazy sweep, bytes in the host's asset registry as the neutral
replacement for the reference's object-storage byte path) and the material
cap configuration. The session-scoped GET list is left as-is; the
reference's owner-material extraction worker is not ported (the branch's
session-material extraction lifecycle already covers extraction).

Gate tests now cover all 23 persistence routes across the three runtime
env states; the materials behavior suite pins the new contract.

* feat(media): add an optional local ffmpeg media extractor (#1213)

Adds a local ffmpeg/ffprobe pipeline as a second media extraction
provider behind the extractor registry, ported faithfully from the
reference implementation: duration probing, keyframe-safe chunking,
per-chunk ASR with timeout and deadline budgets, and timestamped
transcript assembly.

- Availability probing feeds the registry's candidate selection: the
  provider simply is not a candidate when ffmpeg/ffprobe are absent.
- With neither ffmpeg nor a cloud provider configured, extraction fails
  with an actionable message naming both enablement paths.
- Media materials route through the same extraction lifecycle and lease
  fence as documents; no parallel queue.
- Tests inject the executable resolver so the missing-ffmpeg path is the
  default-tested one; the real pipeline test is skip-if-unavailable.
- @openmaic/storage 0.13.0 -> 0.14.0 (media routing in the material
  lifecycle surface).

* feat(storage): per-scene monotonic revisions via database triggers (#1214)

* feat(storage): per-scene monotonic revisions via database triggers

Restore the reference implementation's freshness granularity: a
per-scene monotonic revision maintained by database triggers, so every
writer (HTTP routes, agent tools, jobs, manual SQL) bumps it without
application cooperation.

- Companion revision tables + trigger functions in the storage package's
  idempotent schema bootstrap, with the lock-order invariant, pg_notify
  wakeup and the suppression switch for batch writers.
- ensureDocumentSchema gained a dollar-quote-aware statement splitter.
- The freshness and manifest routes serve per-scene revisions.
- Mutation-verified: dropping the triggers turns the revision tests red.
- @openmaic/storage 0.13.0 -> 0.14.0.

* fix: forward the freshness manifest through the owner-bound store

* feat(workbench): add the Pro entry points and preserve the mode-transition semantics (#1208)

* feat(workbench): add the Pro entry points

* feat(workbench): preserve Pro mode transition semantics

* fix(workbench): drop ambient declarations shadowed by landed slices

* fix(workbench): drop ambient declarations shadowed by the landed shell

* feat: port workspace shell sibling modules

Port the 16 leaf modules the Pro workspace shell imports but that were only
ambient-declared, replacing the compile-time bridge with real implementations
adapted from the sibling-slice reference: pure workbench helpers (session
title, rail tab, course-chat bootstrap, created-course tabs, course-tabs
memory, workspace navigation, pane navigation, pro-edit sizing,
existing-course minting, first-message session), the neutral brand context and
course-rename server API, the server-action session delete, the home
discovery hook, the classroom pane host with its load-policy leaf, the theme
toggle and floating-layer owner, plus the floating-layer-owner wiring the
dialog/dropdown/tooltip portals stamp.

Also add the workbench-shell locale copy for all 12 locales, port the
reference tests for the ported modules, and drop
types/workbench-sibling-slices.d.ts now that every declaration has a real
implementation.

* docs: keep ported comments in English and deployment-neutral

* docs: announce 1.0.0 and refresh the feature overview (#1216)

* docs: announce 1.0.0 and refresh the feature overview

* docs: finalize 1.0.0 README after feature merge

* fix(agent): control-plane routes answer 404, not 500, without a database

The agent control-plane routes gated only on the runtime flag, so an
enabled-but-unconfigured deployment (flag on, DATABASE_URL empty)
answered 500 from a store that cannot connect. Gate them on the
configured check instead, matching the stage/material routes: the
whole surface is cleanly absent until both the flag and the database
are present. The status probe keeps reporting both bits.

* test: mock both runtime gate exports in the control-plane route suites

* fix(agent): abort in-flight TTS on cancel and bound each provider request with a timeout (#1217)

The generate_tts / scene-tts path checked the runner's AbortSignal between
actions but never created the provider HTTP requests with it, so a session
cancel left a hung synthesis fetch in flight until a restart repaired the
tool result. Thread the signal end-to-end: TTSModelConfig carries an
optional signal, generateTTS combines it with a per-request timeout
(TTS_REQUEST_TIMEOUT_MS, default 30s, ported from the reference runtime's
TTS bounds) via AbortSignal.any, and every provider fetch (openai, azure,
glm, qwen incl. voice-clone + audio download, voxcpm, minimax, doubao,
elevenlabs, lemonade) is created with that signal.

A timeout now fails the tool call with TTSRequestTimeoutError (a clear
retryable error) instead of wedging the session; a caller cancel propagates
as the interruption so the runner settles the session as cancelled without
a restart.

Tests: hung-provider simulation rejects at the timeout with the retryable
error; abort mid-flight aborts the captured request signal and surfaces the
interrupted shape; removing the signal wiring makes the abort tests fail
(red), restoring them turns green.

* fix(workbench): PG-mode home listing via owner stages; keep the interrupted terminal course card (#1218)

Finding 1: with server persistence on, listStages resolved to the generic
GET /api/persistence/documents listing, which the capability model
deliberately answers 403 FORBIDDEN_DOCUMENTS for (reads by id, listings
owner-only). The home/workspace library now lists through the owner-scoped
GET /api/stages surface (same anonymous-owner cookie the workbench uses)
when server persistence is enabled; the server-side 403 is untouched.

Finding 2: a run interrupted (session_interrupted) and repaired
(session_resumed) that ends cancelled before agent_end stranded its pending
classroom sightings, so the timeline's terminal card lost the course the
answer produced. session_end (cancelled) now flushes the pending sightings
into the same course card set agent_end paints, before the stopped caption.

* chore(workbench): remove the bookmark concept and the saved-courses drawer (#1219)

* chore(classroom): remove the bookmark ('collected') concept entirely

The stage-meta viewer port introduced a bookmark surface (stage_bookmarks
table, POST /api/bookmarks, the isBookmarked sidecar field, and a
readOnly rule that let a saved course stay editable). The product has no
such concept, so remove it as a closure:

- delete the /api/bookmarks route and the stage_bookmarks table plus its
  query helpers from the persistence bootstrap
- drop isBookmarked from GET /api/stage-meta/[stageId]
- simplify the classroom read-only rule to readOnly = !isOwner across the
  sidecar client, ownership signal, classroom load, stage store and the
  classroom page
- keep publish/unpublish, generation-complete, isOwner and isPublic
  exactly as they were
- update the gate and stage-meta route suites and the README mentions

The workspace rail's Bookmark glyphs and comments describe the upstream
saved-courses (favorites) section, which is driven by isOwner and renders
no collect affordance; they are kept as unrelated homonyms.

* chore(workbench): remove the saved-courses drawer UI

The first pass removed the bookmark data model but kept the rail's
"Saved courses" drawer, judging it a separate surface driven by
`isOwner === false`. The home/workspace listing is owner-scoped, so that
flag can never occur: `allSaved` is permanently empty and the drawer
(plus the collapsed-rail Bookmark mini-button) is a dead affordance.
Remove it: the SavedDrawer component and its mount, the savedOpen /
savedSection state, the allSaved / matchedSaved derivations, the 'saved'
variant of the course-list renderers, the mini Bookmark glyph, the
drawer-only CSS, and the drawer's i18n keys from all 12 locales. The
courses tab is now exactly one folders tree.

The authored/favorites split in workspace-tree.ts goes with it; the tree
module no longer reads `isOwner`. The discovery course type keeps the
field — the shell still reads it for read-only gating.

Upstream has no collect concept; the drawer could only ever render empty
here. The reference implementation HAS this drawer (its favorites come
from its account system), so this removal is a deliberate upstream
product decision, not a fidelity bug.

* fix(workbench): restore the attach entry, add the rail settings entry, pin all three entry points (#1221)

* fix(workbench): restore the composer attach entry by gating it on the live runtime

The AttachButton's rollout probe read a `materialsEnabled` field that this
branch's /api/agent/runtime never answers (the materials routes gate on the
runtime itself, like the stages), so the gate could never pass and the attach
button never rendered — the Pro launch and chat composers showed only the
@-mention and enhance glyphs.

Substitute the field with the runtime's `enabled` value, which IS the upload
action's precondition: POST /api/materials answers 404 whenever it is false,
so the render condition now equals the action precondition (no dead button).

The button's label (`proMode.attach`) is a user-visible string that becomes
visible again; port the reference implementation's own translations verbatim
into the 11 locales that still carried the Chinese copy.

* feat(workbench): add the settings entry to the rail's bottom-left cluster

The reference's rail foot carries a cluster of utilities (its saved-courses
drawer, the language switcher, the display toggle). This branch removed the
drawer — it could only ever render empty here — and the product decision is
to fill that freed spot with the settings entry.

Add a settings trigger to the foot cluster (expanded rail, beside the
language and display toggles, and on the collapsed strip) and mount the
model/provider SettingsDialog in the rail, wired to the trigger. It is the
same dialog the classic home opens from its header pill; the workspace had no
settings entry of its own, so nothing is duplicated within a surface.

* test(workbench): pin the restored upload, attach, and settings entry points

Covers the three restored entry points:

- the courses-tab upload control: rendered beside the course name filter,
  wired to the discovery hook's ZIP import trigger, disabled while an import
  runs, and gated by the same condition as its action (the courses tab);
- the composer attach control: an actual render of AttachButton under both
  probe answers (visible when the runtime says the upload path is live,
  hidden otherwise), its mounts in the launch and chat composers, the
  branch's runtime-field substitution in the probe, and the reference's own
  `proMode.attach` copy in all 12 locales;
- the settings entry: the trigger in the rail's foot cluster (expanded and
  collapsed), beside the language and display toggles, opening the
  SettingsDialog the rail mounts.

* chore(config): the Pro workbench flag implies the MAIC Editor gate (#1223)

A workbench build without the editor toggle has no way to edit a course:
enabling NEXT_PUBLIC_PRO_WORKBENCH_ENABLED while forgetting
NEXT_PUBLIC_MAIC_EDITOR_ENABLED produced exactly that split-brain bundle.
The workbench IS Pro mode, so its flag now implies the editor gate; the
standalone flag remains for deployments that want the classroom editor
without the workbench. Documents both flags in .env.example.

* fix(agent): wake SSE tails and the runner on durable deltas (streaming fidelity) (#1222)

The Pro workbench chat did not stream: the session/owner SSE routes polled
the durable event log on a 5s/30s clock with no wakeup, so message_update
deltas (written at 150ms cadence) reached the browser in poll-sized blocks
and the thinking strip only mounted after the whole reasoning text had
accumulated.

Port the reference's LISTEN/NOTIFY delta path:

- storage: add in-transaction wake hooks (onSessionEventAppended,
  onOwnerEventAppended, onCancelRequested) so a host queues pg_notify in
  the same transaction as the durable append; align readEventsAfterForReplay
  to rank the bounded page so the first delta after the cursor is always
  kept (the live tail can never starve). Bump @openmaic/storage to 0.18.0.
- app: port the process-wide event-notify bus (dedicated LISTEN client,
  self-check probe, reconnect backoff; notify through the storage
  transaction surface), wire the store hooks, subscribe both SSE routes
  before the initial read with the reference's initializing gate, and give
  the runner one {kind:'session'} subscription whose wake runs the cancel
  check and the message drain. Polls stay as the lossy-NOTIFY backstop.
- lifecycle: start/stop the bus from instrumentation.

Tests: storage hook + compaction contract; route wakeup latency; runner
wakeup wiring with a fake agent; bus unit tests; PG contracts proving a
real append wakes the routes and a live SSE route forwards a message_update
on the wakeup, and that a rolled-back append never wakes. Also fix the
pre-existing park-attempt-budget PG test TRUNCATE (missing CASCADE against
newer FK tables).

* fix(storage): asset writes self-deadlocked against pooled PostgreSQL (#1225)

* fix(storage): refuse the non-transactional byte-write deadlock configuration

A byte store whose plain write() runs on its own pooled connection cannot be
invoked from inside a registry write transaction: after the transaction has
claimed the blob-row lock, that write blocks on the lock the transaction just
took while the transaction waits on the write - a self-deadlock PostgreSQL
cannot detect (one side is idle in transaction).

There is no lock-safe ordering for such a writer: bytes must be written after
the row claim (writing before it lets the collector delete the bytes while the
upsert waits), and any second-connection write after the claim is the
deadlock. The configuration is therefore detected and refused:

- AssetByteStore gains writesOutsideRegistryDatabase?: true, declaring that the
  layer's plain byte operations cannot contend for the registry's row locks.
- PgAssetStore refuses put()/replace() up front (and defends coordinatedWrite)
  when the byte store has no writeWith and does not declare the flag, throwing
  a clear configuration error before any row is claimed.
- The collector mirrors the guard on its delete path (deleteWith or a declared
  out-of-registry layer, else a configuration error).
- The object store declares the flag (its out-of-transaction write remains
  legitimate); the in-registry PostgreSQL byte column provides writeWith /
  deleteWith instead.
- Write transactions (put/replace/remove) set SET LOCAL lock_timeout = 30s so
  any future lock-contention variant fails loudly instead of hanging.

Bumps @openmaic/storage to 0.18.0.

* fix(persistence): forward the transactional byte methods through the lazy asset byte-store wrapper

The no-bucket case of lazyAssetByteStore returned a bare { write, read,
delete } and dropped writeWith/readWith even though the underlying
PgAssetByteStore has them. The registry's hasTransactionalWriter duck check
then failed and put() fell back to the byte store's own pooled connection,
which blocks forever on the blob-row lock the registry transaction just took
when the bytes live in the same PostgreSQL - the production self-deadlock.

The no-bucket layer is statically PgAssetByteStore, so its transaction-pinned
methods are forwarded eagerly (typed against the real signatures via
PgForwardedByteStore). The bucket case keeps its lazy-probing semantics: no
transactional writer exists there, the signed-URL method stays absent or lazy
exactly as documented, and the wrapper now declares writesOutsideRegistryDatabase
so the registry may run the plain write inside its transaction.

New tests pin the wrapper's transactional capability red-to-green and assert
put()/resolve() route byte traffic through the transaction-pinned queryable.

* fix(home): cap the generate-prep ingest drain at 3s so Generate never waits the full server budget

The classic home flow's Generate click drained in-flight ingests for the full
15s server budget. Cap the wait at GENERATE_DRAIN_CAP_MS (3000ms, documented as
a UX bound) and reuse the existing timeout fallback: sources that miss the cap
proceed on the legacy byte path and each late-resolving id is released.

* chore(storage): bump to 0.19.0 over the concurrently landed 0.18.0

* fix(agent): bound every tool call with a timeout; never resurrect a cancelled session (#1226)

* fix(agent): bound every tool call with a global timeout and settle it on cancel

A tool await that neither resolves nor rejects wedges the session forever:
the lease keeps heartbeating and the driver never reaches its next cancel
checkpoint. Race every tool execution (in buildAgent) against a hard budget
(OPENMAIC_AGENT_TOOL_TIMEOUT_MS, default 10 min, per-tool overrides for
known long runners) and against the caller's AbortSignal, so even a
signal-ignoring await cannot keep a cancelled session running.

On timeout the call rejects with AgentToolTimeoutError; the agent loop turns
the rejection into a structured error tool-result the agent can retry or
proceed from, and the abort signal is delivered to the tool's in-flight work
through a derived controller. Zombie-tool updates after settlement are
dropped.

* fix(storage): never re-lease a cancel-requested session; settle it as cancelled on claim

The claim scan treated a session with cancel_requested_at set as a normal
claim candidate: after a restart it re-leased the same session for attempt
N+1 and resumed generating despite the pending cancel. claimNextSession now
settles such candidates as cancelled under the claim lock (status cancelled,
attempt reset, lease and cancel request cleared, terminal session_end event
and owner projection) instead of leasing them, then keeps scanning.

Bump @openmaic/storage to 0.18.0.

* docs: takeaway-style 1.0.0 announcement with bilingual guide links

The 1.0.0 head is now a short takeaway block — badge links to the
official user guides (English and Chinese), five one-line highlights,
and pointers into Features and the workbench setup section — instead of
six dense paragraphs. The detailed provider-neutrality and freshness
notes move into the Features workbench section, phrased database-
neutrally (the announcement no longer names a specific database).
Release date corrected to August 27.

* fix(workbench): restore editor chrome, mode transition, streaming, materials, mentions, folders (#1229)

* fix(workbench): wire workspace folder routes

* fix(editor): restore reference workbench chrome

* fix(workbench): persist composer materials and course refs

* fix(workbench): preserve live reasoning frames

* fix(persistence): back off failed streaming saves

* chore(workbench): retire stale slice seams

* test(editor): cover element pin layer

* chore(storage): bump to 0.21.0 for the user-message ref/material fields

* chore(editor): translate ported code comments to English

* fix(agent): fence durable tool writes and consume cancel requests atomically (#1230)

* fix(agent): enforce provider force-off in agent tools and scrub vendor identity from tool results (#1231)

* fix(materials): serialize per-owner quota reservations and make crashed uploads reclaimable (#1232)

* fix(editor): resolve dock-bar i18n keys, remove dock height drag, wire element referencing (#1233)

* fix(workbench): send the opening session message exactly once with refs intact (#1234)

* feat(editor): port timeline TTS preview single-flight and voice-all state latching (#1235)

* fix(media): restore the reference classic media chain (#1236)

* fix(import): adapt imported PPTX canvas size so decks render without overflow (#1237)

* fix(editor): complete element referencing — renderer DOM contract and GenUI picking aligned with the reference (#1238)

* test(providers): reconcile the provider-config vendor-token debt count after the main merge

The integration line's AK/SK fallback for the managed document provider adds
occurrences that main's allowlist snapshot predates. Same mixed-composition
debt category the group already documents; no new vendor behavior.

* test(providers): reconcile vendor-token debt counts with the integration line

The main-merge brought main's neutrality-guard snapshot next to integration
features it predates (media-extractor fallback chain, local voice-profile
deletion semantics, the enabled-TTS helper). Same debt categories the guard
already documents; counts updated to the guard's own tally and two grouped
entries added. No new vendor behavior.

* fix(agent): carry reasoning through the completions dialect so the thinking strip renders (#1239)

* feat(skills): add Feynman and spiral curriculum methods (#1240)

* feat(agent): port missing reference tools and skills (parity audit) (#1241)

* feat(media): retire asset-registry wiring; media and materials follow the reference byte model (#1242)

* fix(classroom): center adapted canvases in the stage and send back navigation home during generation (#1243)

* feat(settings): skill management with real list, download, delete, and upload (#1244)

* feat(settings): skill management section with real list, detail, and zip download

* feat(skills): owner skill delete and upload across storage, API, and settings

* fixup! feat(settings): skill management section with real list, detail, and zip download

chore: neutralize a reference note in the settings header comment

* fix(media): persist origin-independent classroom-media references from the agent runtime (#1245)

* feat(editor): float the insert toolbar in the outer frame with collapse (#1246)

The insert strip was bounded to the slide card, so it could only ever sit on
top of slide content: the card's overflow clipped it and it could not be parked
in the padding beside the slide. Move it into the studio frame the element
picker's panel already roams (CanvasOverlayPortal + the frame selector), so both
canvas overlays share one bounding container and their handles behave the same.
While picking, the strip rises over the picker and goes inert, which is the
z-order CANVAS_OVERLAY_Z already documents.

Add a fold beside the grip: the chevron collapses the strip to that grip row and
back, with the buttons unmounted rather than hidden. The fold is session-local
state owned by EditShell, next to the drag offset, so a surface swap keeps it;
nothing is persisted. Expanding a strip parked at the bottom edge re-clamps
through the same bounds rule the keyboard move uses.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(workbench): align the chat timeline's left edge with the composer (#1247)

* fix(agent): fence session claims while an ask_user question is outstanding (#1248)

* fix(agent): settle-time rescue tracks real delivery instead of a count offset (#1249)

* fix(persistence): migrate owner_material to oss_key and drop legacy asset_id (#1250)

* docs(readme): surface the 1.0.0 user guide badges at the top (#1253)

* fix(workbench): show newly created folders in the sidebar without reload (#1254)

* docs(readme): add the release version prefix and drop the opt-in framing

* fix(workbench): single-source the chat gutter so timeline and composer share a left edge (#1255)

The transcript and the composer each established their own column: their own
`px-*` gutter and their own `mx-auto w-full max-w-*` centering wrapper. Equal
padding values were never enough, because the two columns are centered inside
different containing blocks — the transcript's is a scroll container, whose
content box is narrower than the composer footer's by the scrollbar's width:

  transcript text left = pad + (pane - 2*pad - scrollbar - measure) / 2
  composer box  left   = pad + (pane - 2*pad             - measure) / 2

The padding cancels out of the difference and what remains is `-scrollbar/2` at
every padding value, so the transcript sat half a scrollbar to the left of the
composer and tuning the two paddings against each other could not move it.

The column is now established once, by the nearest common ancestor of both
(`chatColumn`), and the scroll viewport and the composer footer are siblings
inside it that add no horizontal inset of their own. The cap carries the gutter
on top of the 760px reading measure, so the text column keeps its width. The
handed-over question row drops the padding that indented it past the agent's
prose; framed rows keep their own inner padding, which is what a card's border
sitting on the column edge means.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>

* fix(workbench): lock pane-embedded classroom to edit mode (#1256)

The workspace right pane painted the full learning chrome — speed control,
play button, learner avatars, mic bar — for a course the agent had just
created, then flipped to edit once the first scene landed.

resolveStageChromeMode treated playback as the DEFAULT branch for a hosted
classroom, so every shortfall fell into it: a course whose tab opens at
stage_link time has no scenes yet, so currentSceneId is null and
isHostedSceneEditable is false. A folded pane parked the playback root
behind the fold and cross-faded it out over the pane on unfold, and a
failed editor chunk dropped into playback permanently.

Lock it at the pane instead of defaulting per entry path:

- WorkbenchPanelProvider — the single element that mounts a classroom into
  the workspace — publishes editPinned (visible && !playback). Every entry
  path passes through it, so none of them decides.
- The hosted resolution can no longer degrade to playback. Start Learning
  (workbenchLearning, new input, split out from pane visibility) is the one
  door; everything else resolves between the neutral loading shell and edit.
- Stage's chrome dispatch is exhaustive on chromeMode, so the playback root
  is no longer the else-branch of a condition about the current scene.

No flicker: chromeMode is resolved during render, and preloadEditor now
answers synchronously (isEditorPreloaded) so a remount with the chunk
already registered paints edit on the first frame. A failed import is no
longer cached forever, so the lock cannot strand the pane.

Standalone classrooms keep their stored mode unchanged.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-27 19:26:25 +08:00
wyuc 98904c1c5d fix(providers): three silent-behavior fixes from the provider audit (#1196)
* fix(extract): fail loudly instead of silently falling back to MinerU Cloud

A request that selects the self-hosted MinerU extractor without a
configured base URL used to fall back to MinerU Cloud whenever a cloud
API key was available, sending documents to a third party without the
operator's knowledge. The cloud fallback now requires an explicit
operator opt-in (ALLOW_MINERU_CLOUD_FALLBACK, default off, documented
in .env.example); otherwise the route answers a 422 that names what was
configured (self-hosted MinerU) and what was unavailable (its base
URL), and points at both remedies.

* fix(media): remove the selectable video provider that has no adapter

The Sora entry in the video provider catalog could be selected in the
settings UI and server config, but had no connectivity or generation
dispatch case, so choosing it failed only at execution time. A catalog
entry that cannot execute is worse than an absent one: the entry, its
type-union member, UI names/icons, store defaults, and the VIDEO_SORA
env mapping are removed. A new catalog test pins that every selectable
video provider dispatches to its own adapter in both switches.

* fix(storage): delete the obsolete no-op storage provider abstraction

getStorageProvider() unconditionally returned a NoopStorageProvider and
the module swallowed every operation into silence, so any caller
believing it had storage got neither bytes nor an error. No real caller
exists anywhere in the repo (only lib/storage/client.ts, a separate
module with a documented null contract and a live caller, remains). The
dead entry point, its type file, and the no-op provider are deleted; a
test pins the removal and that the real client upload helper is kept.
2026-08-27 11:49:23 +08:00
wyuc 06f105ed4d feat(server): validate model routing config at startup (#1182)
* feat(server): validate model routing config at startup and deprecate bare model ids

* fix(server): keep the bare-id deprecation warning out of request paths

parseModelString is reachable with request-controlled strings, so warning
there lets clients drive log volume and grow the dedupe set without bound;
the warning now fires only from boot-time validation of config-derived
sites, whose surface is finite. DEFAULT_MODEL missing-key wording no longer
claims requests will fail (unrouted resolution honors client keys), and a
keyless provider with a default base URL is not flagged as unconfigured.
2026-08-26 10:59:47 +08:00
wyuc b8fbac1a54 feat(providers): uniform capability force-off and consistent missing-key contract (#1181)
* feat(providers): uniform capability force-off and consistent missing-key contract

* fix(providers): close force-off bypasses found in review
2026-08-26 10:59:43 +08:00
wyuc e3fd494658 feat(tts): add Qwen TTS voice cloning (#1160)
* feat(tts): add Qwen TTS voice cloning via the registration adapter

* feat(settings): add Qwen TTS clone-voice manager

* fix(tts): harden Qwen voice cloning integration

* fix(tts): enforce voice-bound Qwen synthesis

* fix(tts): close Qwen voice clone tail issues

* fix(tts): close final voice clone review issues

* fix(agents): harden voice token parsing

* fix(settings): hide the voice-clone model from manual selection

* fix(agents): pin the narrator voice to the user's selected voice

* fix(agents): keep narrator fallback alive when the pinned voice is unusable

* fix(i18n): correct the narration fallback notice copy

* fix(tts): bound the narrator voice retry to a single fallback hop

* fix(tts): tolerate transient vendor errors in the voice existence precheck
2026-08-24 14:33:43 +08:00
LING DUANandwyuc 59adbef03b Simplify Pi agent tool ownership and completion (#1148)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-20 00:26:07 +08:00
杨慎 7df8e50054 feat(video-export): add bounded local chunk executor (#1115)
* feat(video-export): add bounded local chunk executor

* fix(video-export): preserve producer plan hash during chunk execution

* fix(video-export): reuse producer chunk sidecars

* style(video-export): format render service entrypoint

* fix(video-export): tighten chunk cancellation and cache validation

* fix(video-export): address chunk lifecycle review findings

* test(video-export): cover chunk cancellation and drain

* fix(video-export): validate observed chunk capture mode

* fix(video-export): enforce chunk plan and resource bounds

* fix(video-export): terminate chunks at job deadline

* fix(video-export): guard chunk worker IPC teardown
2026-08-17 12:11:46 +08:00
LING DUANandwyuc cdf949a017 feat(pi): add native child runtime with scoped Spotlight (#1111)
* feat(pi): add native child runtime with scoped Spotlight

* fix(pi): address native child review feedback

* fix(pi): suppress prefaced structured child output

* fix(pi): restore native child text streaming

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-16 00:05:23 +08:00
wyucand杨慎 0a10af11cf feat(storage): opt-in indirect asset byte egress (#1007) (#1100)
* feat(storage): add opt-in indirect asset byte egress

A deployment can now opt an asset byte GET into a 302 to a short-lived
signed URL when the byte layer can sign. It is off by default, and
byte-for-byte unchanged when off. The byte layer gains an optional
signReadUrl capability; S3AssetByteStore implements it through the new
optional @aws-sdk/s3-request-presigner peer, resolved lazily exactly
like the client SDK, with a signer seam mirroring the commands seam so
tests can bind doubles. The signed URL pins the contract's response
headers -- the media type after the renderable allowlist, the fixed
disposition, and the no-store cache posture -- via S3 response-header
overrides. PgAssetStore mints the URL from the same ownership-checked
transactional read a direct resolve performs, so authorization still
runs per read before any URL exists, and a store whose byte layer
cannot sign falls back to direct bytes. The contract documents the
shape and the disclosure tradeoff: a Location into a hash-keyed byte
layer names the content hash, so deployments that need the
no-disclosure property keep direct egress.

Refs #1007

* feat(app): wire ASSET_BYTE_EGRESS into the persistence route

ASSET_BYTE_EGRESS=redirect opts asset byte GETs into the storage
package's indirect egress; unset, direct, or an unrecognized value
keeps the default byte-for-byte behavior, with a warning for the
unrecognized case. The lazy byte-store wrapper now forwards
signReadUrl and answers undefined when the resolved layer has no
signer, so the PostgreSQL byte column degrades to direct bytes and
the S3 layer signs without the wrapper knowing which it holds. The
collector path is untouched: it builds its byte store through
createAssetByteStore as before and only ever calls delete.

Refs #1007

* fix(storage): let the packaged client survive redirect egress

Cross-review round 1 on the indirect byte egress change found two
problems.

A byte response that followed a redirect carries the pinned content
type but no X-Asset-Revision, and HttpAssetStore required both from
the same response, so every cold resolve under redirect egress failed
MALFORMED_RESPONSE. The client now detects the shape once -- a
redirected byte response with no revision -- latches it, and probes
HEAD before each cold GET from then on, taking the label from the
probe. The probe predates the download on purpose: a replacement in
between labels newer bytes with an older revision, which the next
revalidation detects and corrects, while the reverse order could pin
stale bytes under a fresh revision with no signal to repair it. The
probe doubles as the miss check, so a redirect-mode miss costs no
download. The direct path is byte-for-byte unchanged; the one double
download per client lifetime on first contact is documented in the
contract, which gains the client-side paragraph this behavior
implements.

signedUrlTtlSeconds also accepted lifetimes beyond the seven-day
SigV4 presigning maximum, minting redirects the object store would
reject. The option is now capped at construction.

Refs #1007

* fix(storage): cap signed URL lifetime below the reclamation grace

A signed URL whose lifetime exceeds the collector grace period can
outlive its object: the last reference goes, the grace elapses, the
collector deletes the object, and the still-valid URL errors at the
object store. The previous ceiling -- the seven-day SigV4 maximum --
made that window days wide against the one-hour default grace. The
cap is now fifteen minutes, and the contract text states both the
ceiling and the rule for deployments that shorten their grace.

Refs #1007

* fix(storage): decline signing when the optional presigner is absent

An unresolvable @aws-sdk/s3-request-presigner made signReadUrl throw,
which the registry surfaced as a failed asset read -- a 500 for a
deployment whose only fault is a missing optional peer. The contract's
answer for a byte layer that cannot sign is the capability fallback:
return undefined and let the caller serve the bytes directly. Signing
errors from a resolved signer still fail loud; only the missing
capability declines.

Refs #1007

* fix(storage): answer redirect egress as a descriptor to asking clients

Following a 302 is not header-neutral: the platform fetch forwards the
original request's headers to the object store's origin, stripping only
Authorization, so a deployment whose credential travels in a custom
header would hand it to the object store -- and the preflight the
forwarded custom headers provoke commonly fails there besides. This
replaces the latch-and-probe client shape with an explicit one: the
client sends X-Asset-Egress: descriptor on every byte GET, a
redirect-egress server answers 200 with a JSON { url, revision } body
instead of a Location, and the client fetches the signed URL with no
deployment headers at all. The revision comes from the descriptor, so
the probe-first ordering and its one wasted download are gone too. The
302 remains the answer for consumers that did not ask; the descriptor
response is marked by its own header so a JSON-media asset can never
parse as one.

Refs #1007

* fix(storage): negotiate the descriptor through Accept, not a custom header

A custom X-Asset-Egress request header is not CORS-safelisted, so every
byte GET to a cross-origin server became a preflighted request -- a
regression for direct-egress deployments that never opted into
anything, against servers with no reason to allow the header. The
negotiation now rides Accept with a vendor media type, which is
safelisted and costs no preflight, and the descriptor answer is
identified by that media type as its Content-Type rather than by a
marker header.

Refs #1007

* fix(storage): sign outside the registry transaction and reserve the descriptor type

Two review findings on the indirect egress path.

resolveIndirect awaited the signer inside the coordinated read, but a
signer on refreshable credentials can wait on the network, and a
database connection plus the blob row's FOR SHARE lock would be held
across that delay. The read now closes first -- the lock still puts
the hash, label and revision in one snapshot -- and signing runs
afterwards, which cannot observe anything the read did not.

The descriptor media type is also reserved from renderableTypes: the
client identifies a descriptor answer by that exact Content-Type, and
an allowlisted asset served with it would be misread as a descriptor
and fail MALFORMED_RESPONSE. The handler now rejects the configuration
at construction, and the contract states the reservation.

Refs #1007

* fix(storage): ship the signing deps and negotiate without ambiguity

Three review findings on the indirect egress path.

The app now declares @aws-sdk/client-s3 and
@aws-sdk/s3-request-presigner as runtime dependencies, and the
standalone build force-includes them for the persistence route: both
are reached through deliberately untraced dynamic imports, so the
shipped image could not resolve them, and a deployment opting into S3
or redirect egress would have found the capability silently absent.

The descriptor negotiation matches exactly on both sides: the client
compares the Content-Type essence, so a longer media type that merely
begins with the reserved value still serves as bytes, and the server
parses Accept media ranges, so a descriptor range sent with q=0
selects the redirect instead of the descriptor it explicitly rejects.

Refs #1007

* fix(app): refuse redirect egress when the collection grace is shorter

A signed URL must never outlive its object: with the handler's
60-second default lifetime, a deployment that sets
ASSET_COLLECTION_GRACE_MS below ten minutes lets the collector delete
an object while a URL minted against it is still valid. The route now
refuses the combination at handler initialization, naming both
variables -- a deterministic misconfiguration, loud at startup rather
than silent at read time.

Refs #1007

* fix(app): ship the signing SDKs without loading them

Two findings on the standalone deployment of redirect egress.

The outputFileTracingIncludes globs did not follow pnpm's scoped
symlinks, so the standalone image carried neither AWS package and both
S3 mode and redirect egress could not resolve their SDK at runtime.
The packages are now server-external and referenced from literal --
but never called -- dynamic import thunks in the persistence byte-store
wiring, which gets them traced into the image while module resolution
still happens only on first S3 use. Verified against a local
standalone build: node_modules/@aws-sdk now contains both packages,
and the route tests that pin the never-resolve-unless-configured
discipline still pass.

Media type comparisons are also case-insensitive now, as HTTP
requires, on both the Accept parse and the descriptor recognition.

Refs #1007

* fix(app): treat an empty collection grace as unset

Number('') is 0, so a present-but-empty ASSET_COLLECTION_GRACE_MS
failed the redirect-egress coordination check and took persistence
down with it, while the collector's own parsing reads the same value
as the one-hour default. The check now trims and treats empty as
unset, matching durationEnv.

Refs #1007

* fix(storage): accept bytes too, omit signing on the known-PG wrapper, state the full tradeoff

Three review findings.

The client's descriptor request now advertises both representations --
the vendor descriptor type preferred, */* accepted -- so a strict
content-negotiating layer can never answer 406 to a client that does
in fact consume plain bytes.

The lazy byte-store wrapper no longer advertises signReadUrl when no
bucket is configured: the layer is known to be PostgreSQL at
construction, and carrying the method made resolveIndirect take its
ownership query and blob-row lock just to decline, then repeat them
in resolve, on every cold GET. With a bucket configured the lazy
validation is preserved and the wrapper still declines when the
resolved layer cannot sign.

And the contract's disclosure tradeoff now states the second edge:
object-store GETs carry ETag and Last-Modified, response overrides
cannot strip them, and because deduplicated PUTs rewrite the
hash-keyed object, Last-Modified moves when another principal
re-uploads the same bytes -- shared-object write timing, beyond byte
equality. Deployments for whom that signal is sensitive must keep
direct egress.

Refs #1007

* fix(app): degrade on bad grace, share the ttl invariant, cover the real boundary

The latest review round, four findings.

A misconfigured ASSET_COLLECTION_GRACE_MS now degrades redirect
egress to direct with a warning instead of failing the shared
handler's initialization -- the asset backend is optional, and its
misconfiguration must never take document and runtime traffic down.

The package exports assertSignedUrlTtlWithinGrace so a deployment
wiring the handler and the collector separately validates the
invariant once at its own boundary: the handler caps the lifetime,
but only the deployment knows both numbers.

The contract states that nosniff does not survive the redirect --
response-header overrides cannot set it -- and why the client-minted
blob makes that safe for conforming clients.

And the route gains a real-boundary test: the Fetch<->Node adapter,
the real createStorageHttpHandler, and the egress wiring, exercised
with an actual byte GET through the composed stack instead of mocks.

Refs #1007

* fix(storage): make the unsafe egress configuration unrepresentable

Three findings on the indirect egress path, all resolved by moving a
rule into a place where it cannot be bypassed rather than by adding a
check a consumer has to remember.

The grace/TTL invariant now lives in the option shape. byteEgress takes
'direct' or { mode: 'redirect', collectionGraceMs, signedUrlTtlSeconds? },
so enabling indirect egress without declaring the reclamation window it
lives inside no longer type-checks, and the handler validates both
numbers at construction. The flat signedUrlTtlSeconds option, its
cross-field "requires byteEgress redirect" check, and the standalone
900-second cap as the only guard all go away; the ceiling stays, but
alongside the ratio rather than in place of it. This is what the earlier
exported assertSignedUrlTtlWithinGrace could not do -- it protected only
consumers who called it -- and it costs nothing in compatibility because
the option is new in this branch. The app resolves the grace through one
shared parser with the collector, so the number the handler checks is the
number the collector runs with.

A signed URL whose object is gone is now a miss. The client maps a 404
from the byte fetch to ASSET_NOT_FOUND, matching the direct path: the
entry was owned and readable when the URL was minted, so a reclaimed
object is the same physical state the direct read reports as a miss, and
what a read means must not depend on the deployment's egress setting.
Signing still does not probe for the object -- that would restore the
round trip this feature removes and make the mint's price vary with prior
presence -- so the contract instead requires the byte layer to answer 404
for an absent object, and states that S3 needs s3:ListBucket to do so.

Route coverage now spans the whole egress matrix through the real
composed handler rather than only the direct case: 302 for a consumer
that did not ask, descriptor for one that did, direct bytes when the
byte layer declines to sign, and direct bytes when a short grace degraded
the configured mode.

Refs #1007

* feat(storage): forward byte egress options through reference server

* fix(storage): fail closed on redirects and unconfirmed 404s in indirect byte egress

A signed-fetch 404 now maps to a miss only when the object store's XML
body declares NoSuchKey; every other 404 from the signed fetch fails loud,
so a wrong bucket, access point, or endpoint can no longer make a live
asset read as absent. The descriptor byte GET is sent with redirect:
'manual' and any 3xx answer is treated as an error, so a server that
ignores the descriptor negotiation can never forward the deployment's
custom credential headers to a redirect target.

The ASSET_BYTE_EGRESS switch is now documented in .env.example and the
server-persistence deployment section, together with the object-store CORS
and s3:ListBucket prerequisites it requires, and the asset descriptor
media type is exported from the package root and the asset/http subpath.

---------

Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-08-14 23:57:21 +08:00
Yizuki_AmeandCursor 13d94a52a1 feat(ai): add Grok 4.6 to the model registry (#1113)
* feat(ai): add Grok 4.6 to the model registry

The Grok picker still topped out at 4.5, so the new flagship ID was not selectable with its documented xhigh reasoning effort.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(docs): stop the docs Next app from loading app instrumentation

Next 16 discovers the repo-root instrumentation.ts from packages/docs, then fails because that file imports app persistence modules. Give the standalone docs app a no-op register() so the Grok catalog docs can build.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-13 16:59:06 +08:00
Xiao Chengmingandwyuc 859d66ec95 fix(ai): add opt-in streaming chat compatibility (#990)
* fix(ai): add opt-in streaming chat compatibility

* fix(ai): preserve streaming chat metadata

* fix(ai): harden streaming chat compatibility

* fix(ai): isolate primary streamed choice

* fix(ai): propagate streamed chat errors

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-12 14:19:05 +08:00
杨慎 7d97b9bbd7 docs: align environment variable template (#1107) 2026-08-12 11:56:57 +08:00
xuyuanwei678 c38da84ef6 feat(renderer): decouple editor UI from the app (#1072)
* docs: define renderer editor task 1 integration

* docs: plan renderer editor task 1

* feat: gate renderer slide editor

* feat: adapt renderer editor intents

* feat: connect renderer editor to slide surface

* test: complete renderer intent shape fixture

* fix: restore renderer editor action previews

* fix: render editor alignment guides

* perf: smooth renderer element dragging

* feat: align renderer hidden element behavior

* feat: add renderer canvas host commands

* feat: add renderer canvas context interactions

* test: cover renderer context menu interaction

* fix: preserve renderer canvas interaction semantics

* docs: design renderer text editing

* docs: plan renderer text editing

* feat(renderer): add ProseMirror text core

* feat(renderer): add rich text editor controller

* feat(renderer): render editable text in place

* feat(renderer): enter text editing on click

* feat: connect renderer rich text editing

* feat: finalize renderer text edit lifecycle

* fix: stabilize renderer text edit commits

* fix(renderer): move elements from selected borders

* docs: design renderer editing UI text toolbar

* feat(renderer): publish editing UI contracts

* feat(renderer): add controlled text format toolbar

* style(renderer): format editing UI contracts

* fix(renderer): address Task 2 review findings

* fix(renderer): handle zero font sizes in toolbar

* feat(renderer): add editing UI color picker

* fix(renderer): address Task 3 color picker review

* fix(renderer): commit native color input changes

* fix(renderer): coordinate native color picker interaction

* feat(renderer): anchor text toolbar to canvas elements

* feat(renderer): compose editable canvas with text UI

* feat(editor): use renderer text toolbar UI

* fix(renderer): make editing UI tests type-safe

* fix(renderer): harden editing UI integration

* feat(renderer): add editable text UI

* feat(renderer): add image and shape editing

* feat(editor): extend renderer editing controls

* feat(editor): add renderer chart insertion

* docs: design renderer latex editor

* docs: plan renderer latex editor

* feat(renderer): add latex editor contract

* feat(renderer): add latex editor dialog

* feat(renderer): compose latex insert and edit UI

* feat(editor): integrate renderer latex editing

* fix(editor): consolidate latex toolbar actions

* fix(renderer): show toolbar hover labels

* test(renderer): cover localized latex tooltips

* chore(renderer): hide latex sample palette

* feat(renderer): add video editing controls

* docs: define renderer audio editing

* feat(renderer): render audio elements

* feat(renderer): add audio editing UI

* feat(renderer): add preview audio controls

* feat(editor): add renderer element clipboard

* feat(renderer): decouple editor UI from app

* chore(renderer): release 0.0.6

* fix(renderer): clear editor CI gates

* fix(editor): adapt renderer canvas to media resolution API

* fix(i18n): complete French editor labels

* fix(renderer): hide toolbar vertical overflow

* chore(renderer): address review cleanup

* chore(docs): remove planning artifacts from pr

* feat(editor): add transactional core package

* chore(editor): ignore local test cache

* build(editor): publish core entrypoint

* refactor(editor): route renderer changes through transactions

* refactor(editor): split renderer editing into editor package

* ci(editor): register editor package for release

* docs(editor): document package dependencies

* ci: rerun validation after editor release registration

* test(editor): validate deletion against existing element

* fix(editor): address transaction and package review findings

* fix(editor): validate transactions and restore editor styles

* fix(editor): protect required updates and react styles

* fix(editor): validate required transaction fields

* fix(editor): prevent insert popovers shifting canvas

* refactor(editor): own renderer editing capabilities

* fix(editor): harden editing transaction boundaries

* refactor(editor): modularize editor UI integration

* docs(editor): harden third-party integration

* fix(editor): preserve video resize aspect ratio

* fix(editor): discard text draft before deletion

* fix(editor): preserve renderer picking and line presets

* chore(editor): remove unused pick-layer import
2026-08-10 19:56:24 +08:00
5d5acd6215 feat: claude search (#393)
* Added search tool support for Claude

* New locales for i18n

* Updated translations

* Prevent default web search base URLs from being prepopulated

* fixed coding style issues

* fixed coding style issues

* fixing lint errors

* fix: add SSRF protection to fetchPageContent in claude.ts

Agent-Logs-Url: https://github.com/joseph-mpo-yeti/OpenMAIC/sessions/099e53fb-e1db-49f4-af7f-4f17cb3d364a

Co-authored-by: joseph-mpo-yeti <55380155+joseph-mpo-yeti@users.noreply.github.com>

* fix: address all remaining review feedback from PR review thread

Agent-Logs-Url: https://github.com/joseph-mpo-yeti/OpenMAIC/sessions/e1ab8fe2-fe0d-45a5-a0af-30945b99c268

Co-authored-by: joseph-mpo-yeti <55380155+joseph-mpo-yeti@users.noreply.github.com>

* fix: rename ambiguous variable names in settings store for clarity

Agent-Logs-Url: https://github.com/joseph-mpo-yeti/OpenMAIC/sessions/e1ab8fe2-fe0d-45a5-a0af-30945b99c268

Co-authored-by: joseph-mpo-yeti <55380155+joseph-mpo-yeti@users.noreply.github.com>

* fixing lint errors

* fix: address 5 items from second PR review thread

Agent-Logs-Url: https://github.com/joseph-mpo-yeti/OpenMAIC/sessions/09469d80-8e91-449b-8bc2-9c11f83555df

Co-authored-by: joseph-mpo-yeti <55380155+joseph-mpo-yeti@users.noreply.github.com>

* fixing model verification error

* fix: address 3 items from third PR review thread

Agent-Logs-Url: https://github.com/joseph-mpo-yeti/OpenMAIC/sessions/0aa49d19-df5b-4a08-a7e1-2295ae9cfc38

Co-authored-by: joseph-mpo-yeti <55380155+joseph-mpo-yeti@users.noreply.github.com>

* fixing lint errors

* fix: resolved Copilot review comments

* fix: returning string promise for ssrf-guard

* fix: returning string promise for ssrf-guard

* fix: updated locales, added allowed_callers for claude search tools, handle duplicated claude model keys

* fix: updated locales, added allowed_callers for claude search tools, handle duplicated claude model keys

* updated tests

* using Anthropic SDK for search and connection test in search settings.

* uses @ai-sdk with a custom fetch including allowed_callers in payload instead of adding Anthropic SDK as a direct dependency

* fixed format

* removed @anthropic-ai/sdk as direct dependency

* fix: resolve TypeScript errors in claude web-search tests

Cast mock.calls through unknown to avoid tuple-index type errors on
Vitest mock objects whose parameter types are inferred as empty tuples.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat: add Claude as a web-search provider on the multi-provider framework

Forward-port of the original Claude web-search work onto main's
searchWeb dispatcher, per review guidance on #393:

- lib/web-search/claude.ts: searchWithClaude via the AI SDK Anthropic
  provider and Anthropic's native web_search server tool, called through
  callLLM for usage accounting. Returns the model's answer plus cited
  sources deduplicated by URL. The page-content enrichment fetch from the
  original PR is removed entirely, resolving the P1 SSRF finding (no
  server-side fetches of attacker-influenced result URLs remain).
- allowed_callers: ["direct"] is still injected at the fetch layer; the
  AI SDK's provider-defined web_search tools do not emit it.
- Tool version picked from the model: web_search_20260209 for 4.6+
  Opus/Sonnet, basic web_search_20250305 for Haiku/older; maxResults maps
  to the tool's max_uses.
- Registry/dispatcher/route wiring: claude provider entry, claudeModelId
  passthrough (web-search route, classroom generation, generation
  preview), WEB_SEARCH_CLAUDE_* env config with server-pinned model via
  WEB_SEARCH_CLAUDE_MODELS, and client base-URL allowlisting for
  api.anthropic.com.
- Settings: per-provider modelId with a model selector in the Claude
  panel; locale keys added to all ten languages.
- Tests: new adapter suite plus claude cases in the dispatcher,
  constants, route, provider-config, web-search-config, and settings
  store suites.

Closes #392

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E9KqAGmKCtBwMC4HtT4JKE

* fix(web-search): trim whitespace around the Claude API key and base URL

Keys are pasted by hand into Settings, and nothing in the client key
path trimmed them: a trailing newline or space was sent verbatim as the
x-api-key header value and came back from Anthropic as 'invalid
x-api-key', with nothing in the message pointing at whitespace.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(web-search): correct Claude adapter tool selection, base URL, and timing

Addresses the four P2 review findings on #393:

- Route dated Claude 4.0 ids (claude-sonnet-4-20250514,
  claude-opus-4-20250514) to the basic web_search_20250305 tool. The
  blocklist regex missed them, so they were sent web_search_20260209,
  which pre-4.6 models reject. Replaced with an allowlist of the models
  that have dynamic filtering, so unrecognized ids also fall back to the
  basic tool -- the safe direction, since 4.6+ models accept both.

- Pin allowed_callers: ["direct"] only on web_search_20250305. Injecting
  it into every request disabled the code-execution caller that provides
  dynamic filtering on web_search_20260209.

- Normalize the bare Anthropic root to /v1 before handing the base URL to
  the SDK, which appends /messages verbatim and would otherwise 404. The
  adapter is the single choke point for the client allowlist, the server
  env var, YAML, and the default; other hosts pass through untouched.

- Return responseTime in seconds, matching every sibling adapter's
  WebSearchResult contract instead of exposing milliseconds for Claude.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-08-09 07:34:12 -04:00
zhifu gaoandLauraGPT 2b3f2eaa37 feat(asr): add local FunASR provider (#1044)
Signed-off-by: LauraGPT <199134975+LauraGPT@users.noreply.github.com>
Co-authored-by: LauraGPT <199134975+LauraGPT@users.noreply.github.com>
2026-08-05 03:03:56 -04:00
de95350516 feat(video-export): deterministic Quiz/PBL cover cards (#985) (#995)
* feat(video-export): add deterministic Quiz/PBL cover cards

Replace unsupported placeholders with authored-metadata cover cards that
stay independent of learner progress, and pin dense-layout evidence for
720p/480p.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(video-export): add configurable localized cover CTA

Replace the inert Start Quiz affordance with a destination-configurable
CTA (NEXT_PUBLIC_VIDEO_EXPORT_CTA_DESTINATION), including Unicode/bidi
display guards.

Co-authored-by: Cursor <cursoragent@cursor.com>

* test(ci): require Hyperframes video-export gates

Lock CI/Docker contracts so cover materialization and Hyperframes sample
lint stay required for export changes.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): move HF_E2E_DIR onto e2e steps

GitHub rejects runner.* in job-level env, which aborted the CI workflow
with zero jobs. Keep the sample dir on the materialize/lint steps only.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(video-export): tidy CI invoke and panel-box export

Run the Hyperframes materialize step through pnpm exec vitest like the
cover-card guardrail above it; the e2e job already pins Node 22. Stop
re-exporting PBL_PANEL_DESIGN_BOX from the public barrel — only the
layout browser test needs it, and it can import the emitter module
directly.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-04 15:43:38 +08:00
littlebullGitandwyuc 3d85525665 [codex] Add Amazon Bedrock LLM provider support (#538)
* Enable AWS-hosted text models without exposing ambient credentials

Rebase the Bedrock provider onto current main and close the review-requested security, credential lifecycle, usage attribution, configuration, and metadata gaps.

Constraint: Bedrock may use ambient AWS credentials only when explicitly enabled by the server operator
Rejected: Trust client-supplied provider types | permits built-in keyless IDs to reach Bedrock credentials
Confidence: high
Scope-risk: moderate
Directive: Keep provider ID/type validation at both request resolution and model construction boundaries
Tested: targeted Bedrock/resolver/config/usage tests; TypeScript; ESLint; i18n alignment; production build
Not-tested: Full suite has 27 failures in unchanged quiz/runtime and runtime/chat-storage tests

* Give Bedrock a recognizable provider identity

Use the official AWS architecture service icon so Bedrock no longer falls back to the generic provider cube in settings and model selection.

Constraint: Preserve the AWS-provided artwork without redesigning the service mark
Confidence: high
Scope-risk: narrow
Tested: SVG XML validation; provider unit test; TypeScript; ESLint; production build

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-03 23:10:50 -04:00
12ccd315ea Add Atlas Cloud LLM provider (#948)
* Add Atlas Cloud provider

* Account for Atlas Cloud non-thinking model metadata

* test: verify Atlas Cloud thinking payloads

* style: format Atlas Cloud payload test

* test: type Atlas Cloud fetch mock

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-08-03 08:51:18 -04:00
LING DUAN 61a7154f07 feat(pi): add evidence-aware Director context runtime (#971)
* feat(pi): add evidence-aware Director context runtime

* fix(web-search): require completed native search

* fix(orchestration): preserve whiteboard across agent handoffs

* fix(pi): remove wrap-up turn bypass
2026-07-27 11:38:41 -04:00
xuyuanwei678andwyuc ff95e6683d feat(renderer): add optional playback canvas (#958)
* feat(renderer): add optional playback canvas

* style(renderer): format playback tests

* fix(renderer): center text by default

* fix(renderer): address playback review feedback

* docs(renderer): align playback canvas API docs

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-07-27 04:15:07 -04:00
杨慎andClaude Opus 4.8 84b1907255 feat(video-export): service-backed MP4 render + in-app one-click export (#866) (#937)
* feat(video-export): service-backed MP4 render + in-app one-click export (#866)

Adds the last mile of classroom video export: turning the self-contained
Hyperframes project ZIP (#865) into an MP4 via an isolated render service,
one-click in-app.

- render-service/: standalone Node 22 + Chromium + FFmpeg container wrapping
  @hyperframes/producer's library API. Async job model (POST /render -> 202
  jobId, GET poll, GET download, DELETE cancel). Swappable JobStore /
  ArtifactStore seams (in-memory + local-disk now; Redis/S3 + presigned-302
  download later) so it scales horizontally without changing the HTTP contract.
  Concurrency + per-user guards are config knobs.
- App integration: thin Next proxy routes under app/api/export-video/* (forward
  only, no rendering) + capability probe. use-render-video.ts uploads the ZIP,
  polls via runPolledTask, downloads the MP4; shared buildExportZip prefix with
  the existing ZIP path. Export menu gains resolution/fps/quality selectors and
  a progress bar; degrades to ZIP download when RENDER_SERVICE_URL is unset.
- docker-compose: render-service under an opt-in "video-export" profile.
- Entry is main.ts (not server.ts): the producer auto-starts its own server on
  :9847 when the process entry path ends with /src/server.ts.

Verified end-to-end in the container: rendered a real 640s (10.7 min) classroom
ZIP to a valid H.264 720p + AAC MP4 (duration matches source) in ~9.6 min
(~0.9x realtime, 4-worker frame capture). Degrade path, queued-cancel + cleanup,
and per-user 429 guard all exercised. pnpm check / lint / tsc / i18n pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(video-export): global render progress store, percent+ETA UI, ring on export button (#866)

Addresses two UX issues found while driving the in-app MP4 export:

1. Progress display was raw and unfriendly (showed producer's English stage
   strings like "Capturing frame 5130/19220") and had no time estimate. Now the
   menu shows only "<percent>% · about <remaining> left". ETA is computed from a
   recent-speed estimate (percent-per-ms over the last sample), EMA-smoothed —
   which tracks the render's non-uniform pace (prep -> frame capture with a
   4->1 worker drop -> encode) far better than a whole-run average, and never
   shows a stale/rising ETA.
2. Switching scenes mid-render unmounted the export menu and lost the progress
   (and reset the local "already rendering" ref, allowing a duplicate submit).
   The whole render lifecycle now lives in a global store
   (lib/store/video-render.ts), so progress survives menu close / scene switch
   and duplicate submits are guarded by status. A persistent CircularProgress
   ring on the export button shows live progress whether or not the menu is open.

Also fixes the progress scale: the producer reports progress as 0..100, but our
HTTP contract (and success path) is 0..1 — the service now normalizes it, so the
client no longer showed "2000%".

- lib/store/video-render.ts: new global store owning submit->poll->download,
  recent-speed ETA, duplicate-submit guard.
- lib/video-export-app/use-render-video.ts: thin facade over the store.
- components/ui/circular-progress.tsx: lightweight SVG progress ring.
- components/stage/{header-controls,video-export-menu}.tsx: ring on the export
  button; menu shows percent + ETA, subscribes to the store.
- render-service/src/render-manager.ts: normalize producer progress 0..100 -> 0..1.
- i18n: percent/ETA strings across all 8 locales (drops the stage-based string).
- render-service/package-lock.json: complete integrity hashes (reproducible npm ci).

Verified: ETA logic checked against the real segmented render curve (worker drop
raises ETA, encode speedup drives it to ~0); progress scale fix confirmed live
against the container (0.2 -> 20%). tsc / lint / prettier / i18n pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): persist render options in the store, not the menu component

Selecting 720p/24fps/draft, switching scenes, and reopening the export menu
showed the defaults again (1080p/30/standard). The selections lived in the
VideoExportMenu component's local state, which reset when the menu unmounted on
a scene switch — the running render still used the chosen options, but the UI
misrepresented them.

Move resolution/fps/quality into the global video-render store (with a
setOptions action). The menu now reads/writes the store, so selections survive
menu close / scene switch, and while a render runs the selectors reflect the
options that render is actually using. startRender() reads options from the
store instead of taking them as an argument.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): deployment correctness + resource/isolation controls (PR #937 review)

Addresses the blocking findings from wyuc's review. Output fidelity was fine;
these harden deployment and production resource/isolation boundaries.

#1 Compose advertised MP4 but couldn't render in prod:
- Capability now probes the service's /health (checkRenderServiceHealth), so a
  configured-but-absent service reports disabled and the UI degrades to ZIP
  instead of 502-ing.
- RENDER_SERVICE_URL is operator-supplied trusted config, so the proxy no longer
  runs it through the SSRF guard — the one-command `docker compose --profile
  video-export up` now works without globally weakening SSRF via
  ALLOW_LOCAL_NETWORKS. resolveRenderServiceUrl() is now synchronous.
- Client degrades to ZIP on any failed submit (not only 501).

#2 Unbounded upload/queue (ZIP-bomb / DoS):
- unzip.ts bounds the archive via fflate's filter BEFORE decompression: entry
  count, per-entry and total expanded size, and compression ratio.
- Proxy rejects oversized uploads (413) by Content-Length before forwarding.
- RenderManager enforces a global queue-depth cap (RENDER_MAX_QUEUE).
- All limits are env-tunable knobs in config.ts.

#3 Per-user guard was ineffective + admission ran after extraction:
- Identity is derived server-side (client IP) and forwarded as x-openmaic-client;
  the service ignores any client-supplied userId, and the proxy strips it.
- Admission is split into reserve()/submit()/release(): the slot is reserved
  BEFORE extraction, so a rejected caller never triggers a decompression.

Additional risks:
- Per-job wall-clock watchdog (RENDER_JOB_DEADLINE_MS) aborts + fails a hung
  render so it can't hold a slot/scratch forever.
- Download proxy bounds only the time-to-headers, not the body stream, so large
  MP4s over slow links no longer truncate.
- Client cancels the server job (DELETE) when a started render fails/times out.
- Compose puts render-service on an internal:true network (no host/internet
  route), sandboxing the Chromium that runs the uploaded HTML; the export ZIP is
  self-contained so no outbound is needed. README documents the standalone caveat.

Not closing #866: the smoke/golden-render CI acceptance criterion remains a
follow-up (see PR description).

Verified in-container: legal render 202; ZIP-bomb (entry-count + compression-
ratio) rejected 400 before any decompression; per-identity guard 429 with a
spoofed multipart userId ignored; reserve-before-extract leaves no scratch dir
on rejection; watchdog aborts an overrunning job and frees the slot. tsc / lint /
prettier / i18n pass; render-service tsc passes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): real audio durations + burned-in subtitles (PR #937 review)

Two export-fidelity issues found in wyuc's deeper E2E:

A. Narration was scheduled from estimated durations, cutting audio off mid-
   sentence and advancing the timeline early. The scheduler trusted
   AudioFileRecord.duration (recorded only since #861), so the many existing
   classrooms without it fell back to text-length estimates — measured 4.35s
   average / 10.23s max underestimate across 47 clips. timeline-deps now probes
   the real duration from each narration blob via an off-document <audio>
   (symmetric to the existing video probe), preferring it over the stored
   duration, then the estimate only when no audio asset exists. Everything
   downstream (narration starts, scene/total duration, subtitle cues) re-derives
   from the corrected value in the pure compiler — no compiler change needed.

B. The final MP4 had no subtitles (only H.264+AAC), and the ZIP's SRT/VTT used
   the same estimated boundaries. The emitter now renders a burned-in subtitle
   overlay: one caption box + a hidden div per cue, revealed/hidden by the paused
   GSAP timeline at each cue's start/end (corrected timings from A), so Chromium's
   frame capture bakes them in. The producer has no subtitle track of its own, so
   burn-in is the v1 approach.

Verified: emitter unit tests + snapshot updated (subtitle overlay + toggle
statements, escaped text, hidden-by-default); 82 video-export tests pass incl.
the determinism red-line proxy. Rendered a synthetic subtitle project through the
container and confirmed by pixel analysis that captions appear only within their
cue window (2429 near-white px in the caption band at t=1.5s vs 0 at t=0.05s).
tsc / lint / prettier / i18n pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): subtitle layout + upload/admission hardening (PR #937 review)

Address the three P1 blockers plus actionable P2s from the 6a585c29 review.

P1:
- emit-hyperframes: stack every subtitle cue in one grid cell and toggle
  display:none/inline-block, so inactive cues leave the flow instead of
  pushing the active cue up into the slide title. Adds a multi-cue
  regression test the single-cue snapshot couldn't catch.
- render route + service: cap the upload by actual bytes (capBodyStream),
  not the spoofable Content-Length; the app now streams the multipart body
  through instead of buffering it via formData(). maxUploadBytes is now read.
- render-service: move makeProjectDir() inside the release()-guarded block
  so an ENOENT/ENOSPC no longer permanently leaks the admission slot; mkdir
  the scratch root at startup for the standalone path.

P2:
- config: allow RENDER_MAX_JOBS_PER_USER=0 to disable the per-identity guard.
- timeline-deps: per-probe timeout + bounded concurrency so a stuck audio
  blob can't wedge export in "compiling" forever.
- render route: only trust x-forwarded-for/x-real-ip under
  TRUST_PROXY_HEADERS=true; otherwise all callers share one "direct" bucket.
- render-service: add vitest tests (unzip limits/traversal, reservation
  arithmetic, body cap, config zero-disable) and a dedicated CI job.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): sync render-service lockfile so `npm ci` passes in CI

The vitest devDependency's transitive esbuild@0.28.1 (and its platform
optionals) were missing from package-lock.json, so the new CI job's
`npm ci` failed with EUSAGE. Regenerated the lockfile from a clean install.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): dedupe esbuild so render-service `npm ci` installs on linux

@hyperframes/core pins esbuild@0.25.12 exactly, hoisting it to the top and
forcing vite@8 (via vitest) to keep a nested esbuild@0.28.1 copy. npm fails to
flag that nested copy's platform-specific optionals as optional, so `npm ci`
tried to install @esbuild/aix-ppc64 on linux and died with EBADPLATFORM.

Add an `esbuild: 0.28.1` override so a single copy is shared (satisfies tsx
~0.28 and vite ^0.27||^0.28); esbuild is build-time only, so pinning the
producer's bundled build tool is runtime-inert.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): resource/isolation hardening (PR #937 round-2 review)

Address the round-2 P1/P2/P3 findings (P1#1 lockfile EBADPLATFORM was already
fixed by the earlier esbuild-dedupe commit; CI Render Service job is green).

P1:
- Admission before buffering (#2): the render service now reserve()s the slot
  from the header identity BEFORE parsing/buffering the multipart, so concurrent
  near-cap uploads are bounded by the queue depth, not just each body by the cap.
- Chromium egress lockdown (#3): the producer exposes no browser-arg hook and the
  render shares the internal network with the app, so a container entrypoint
  installs an iptables egress lockdown (drop all outbound except loopback +
  established replies) then drops privileges. Needs CAP_NET_ADMIN (added in
  compose); graceful warn-and-continue if unavailable. The self-contained ZIP
  needs no outbound.
- Default one-render bottleneck (#4): with no trusted proxy every caller is
  "direct", so RENDER_MAX_JOBS_PER_USER=1 throttled the whole deployment. Default
  compose now sets it to 0 and relies on concurrency + global queue caps.
- Non-blocking bounded extraction (#5): unzipSync -> fflate async unzip (worker,
  off the event loop), keeping the pre-decompression filter; default expanded
  ceiling 1GB -> 512MB; a semaphore caps concurrent extractions; compose adds a
  container mem_limit.

P2:
- Raise the app submit timeout/maxDuration (300MB upload can't finish in 60s).
- video-render store: only degrade to ZIP when the service is genuinely
  unavailable (501/unreachable); surface real 429/413/5xx instead of an
  unsolicited download.
- useExportVideo dedupe guard moved to module scope so it survives the menu
  unmounting (no second concurrent ZIP pipeline).
- .env.example: RENDER_SERVICE_URL bypasses SSRF; drop the ALLOW_LOCAL_NETWORKS note.

P3:
- Deadline overruns are marked failed (not cancelled).
- submit() decrements the identity slot if jobs.create throws (no leak).
- CI sets PUPPETEER_SKIP_DOWNLOAD; unzip tests use tiny fixtures + low env limits.

Tests: render-service now 22 tests (unzip limits/traversal, admission incl.
create-leak, body cap, semaphore, config); app video-export suite unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(video-export): buffer under the extraction permit + fail-closed egress (PR #937 round-3)

Address the two remaining round-3 P1 boundary blockers.

P1#1 — buffering was outside the gate: `bounded.formData()` materialized the
whole uploaded file into memory BEFORE `extractionGate.run()`, so up to
RENDER_MAX_QUEUE (20) admitted bodies could each buffer ~300MB (≈6GB vs the 4g
mem_limit) before the 2-permit gate. Move the entire RAM-heavy section —
formData buffering, file read, and unzip — INSIDE the permit; the queue
reservation still runs first (a rejected caller consumes nothing). Requests
beyond the permit wait with their body unconsumed (socket backpressure), so at
most maxConcurrentExtractions bodies are buffered at once. Refactored main.ts
into a testable `createApp(deps)` factory and added an integration test proving
peak concurrency in the buffering+extraction section never exceeds the permits.

P1#2 — egress lockdown failed open: the entrypoint warned and started normally
if iptables setup failed, so /health stayed green while Chromium could reach the
app. With RENDER_EGRESS_LOCKDOWN=true (default) it now FAILS CLOSED — exits
non-zero if not root, iptables is missing, or the rules don't apply. Operators
accepting an unisolated setup opt out with RENDER_EGRESS_LOCKDOWN=false. Added
scripts/egress-smoke.sh to assert the boundary (lockdown active, loopback works,
new outbound blocked).

Verified: image builds; container boots as `render` with lockdown active and
serves /health; fail-closed exits 1 without CAP_NET_ADMIN; egress smoke passes
(outbound blocked); 23/23 render-service tests + tsc; app tsc + root prettier clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 11:28:34 +08:00
bb8d4f30e7 feat(chat): add experimental Pi classroom runtime behind a flag (#914)
* feat(chat): add experimental Pi classroom runtime

* fix(chat): harden Pi whiteboard action sequencing

* style(chat): format session retirement flow

* style(playback): format async scene transition

* fix(chat): retire actions on session errors

* fix(chat): close Pi runtime lifecycle gaps

* fix(chat): avoid replaying historical whiteboard actions

* fix(chat): prevent Pi structured-output residue from leaking and inflating turns

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(chat): expose whiteboard element/line ids and bound code-line context

Surface each element's id (and code line ids) in the Pi whiteboard context so
later child agents can target wb_delete / wb_edit_code, and cap the rendered
code lines with a shared character budget so a code-heavy board cannot grow the
prompt without bound. Return only the current turn's whiteboard ledger from the
Pi director loop instead of carrying history forward across requests.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(chat): bound Pi whiteboard code context

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(chat): prioritize newer persisted whiteboard code in budget

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>

* fix(chat): enforce whiteboard boundaries across Pi sessions

* feat(chat): refine soft-closing discussion controls

* fix(chat): make whiteboard clearing model-directed

---------

Co-authored-by: apple <apple@GigideMacBook-Air.local>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: LING-6150 <duan.lin@northeastern.edu>
2026-07-17 11:22:02 +08:00
c56929510c feat(web-search): add SearXNG provider support (#842)
* feat(web-search): add SearXNG provider support

Enable self-hosted SearXNG via SEARXNG_BASE_URL, with server-side provider
fallback, settings sync, and tests.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style: fix Prettier formatting in web-search files

Restore pnpm check by applying Prettier to imports, headers, and test assertions.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(web-search): keep SearXNG base URLs server-managed only

Close the SSRF primitive by ignoring client-supplied SearXNG URLs at the
route boundary and rejecting them in resolveSafeClientWebSearchBaseUrl.
Add negative route/config tests and an opt-in live smoke test for the
JSON API.

Co-authored-by: Cursor <cursoragent@cursor.com>

* style: fix Prettier formatting for SearXNG web-search changes

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: 徐松 <song.xu@aieshanghai.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-13 23:12:31 +08:00
3b63710376 feat(ai): add Azure OpenAI provider (#916)
* feat(ai): add Azure OpenAI provider

Add deployment-based Azure OpenAI configuration for both client and server-managed setups, normalize Azure portal endpoints, and expose the provider in settings.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b0d3a94f-089a-470e-9987-608f1506ca6f

* style: format Azure provider type

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Copilot-Session: b0d3a94f-089a-470e-9987-608f1506ca6f

---------

Co-authored-by: MarshellOnMoon <121332425+MarshellOnMoon@users.noreply.github.com>
2026-07-13 23:12:23 +08:00
jackefnandRowan_lxb baff0998bb Feat/document bundles milestone 3 (#844)
* feat(document): support document bundles

* fix(document): avoid server extractor imports in client bundle

* fix(pdf): default mineru backend to pipeline

* fix(document): address bundle review findings

* fix(document): harden extraction limits and storage errors

* fix(document): make analysis step format agnostic

---------

Co-authored-by: Rowan_lxb <Lxb_savior@163.com>
2026-07-13 23:01:40 +08:00
wyucandClaude Opus 4.8 25cf58d19f feat(model): optional per-stage LLM model routing (#745) (#747)
* feat(model): optional per-stage LLM model routing (#745)

Add a config-only stage -> model map consulted during model resolution,
falling back to DEFAULT_MODEL when unset (zero behavior change unless opted in).

- lib/server/model-routes.ts: parse/validate/cache MODEL_ROUTES env JSON;
  LLM_STAGES registry of routable stages; getStageModel(stage).
- resolveModel: resolution order x-model > stage route > DEFAULT_MODEL > builtin;
  thread optional `stage` through resolveModelFromHeaders/FromRequest.
- Wire each route's resolveModel call site to its stage; classroom-generation
  resolves generate-classroom and web-search-query-rewrite independently.
- .env.example: document MODEL_ROUTES; unit tests for both modules.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): address cross-review findings for per-stage routing (#745)

- classroom-generation: resolve the web-search-query-rewrite model lazily,
  only when explicitly routed and inside the web-search branch, with try/catch
  fallback to the classroom model. Fixes a regression where a misconfigured
  optional route (keyless/invalid provider) threw and aborted ALL classroom
  generation even with web search disabled; also avoids wasted resolution on the
  common no-search path and preserves classroom-model inheritance when unrouted.
- resolve-model: type the stage param as LlmStage across resolveModel/
  resolveModelFromHeaders/resolveModelFromRequest so a mistyped stage literal is
  a compile error instead of silently falling through to DEFAULT_MODEL.
- model-routes: simplify the cache to a process singleton (matching
  provider-config), dropping the raw-string-keyed cache and redundant STAGE_SET.
- .env.example: clarify x-model precedence and that a route to an unconfigured
  provider fails at request time (no startup validation), like a bad DEFAULT_MODEL.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): stage route takes precedence over client x-model (#745)

Cross-review (codex P2 + Claude tracer) found that the browser UI always sends
its saved model as x-model, so the previous x-model > stage order made
MODEL_ROUTES inert for normal UI traffic on exactly the heavy stages the RFC
targets (scene-content, quiz, pbl-chat, chat). Flip resolution to
stage route > x-model > DEFAULT_MODEL: a configured route is the operator's
deliberate per-stage choice and wins, while unrouted stages still honor the
client's x-model. Update docs and tests accordingly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(model): prettier-format changed files for #745 CI

Run the repo formatter (pnpm check / prettier) on the changed routes, the new
model-routes module, and the new tests so `prettier . --check` (CI) passes;
also correct a stale precedence comment in the resolve-model test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): isolate routed model from client connection params (#745)

When a stage route overrides the client's x-model and points at a different
provider, the client-sent apiKey/baseUrl/providerType (for the client's model)
must not bleed onto the routed provider — otherwise a routed Anthropic model
would be constructed with the client's OpenAI providerType/key and fail. A
routed model now resolves its connection params purely from server config, as
if no x-model was sent. Unrouted stages still honor the client params. Adds
tests covering both the drop (routed) and keep (unrouted) cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(model): correct graceful-degradation comment in classroom (#745)

resolveModel does not throw on a missing key; clarify that a keyless route
surfaces later in callLLM (degraded by the outer try/catch), while only a
resolution-time failure is caught at the rewrite re-resolve. Comment-only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(model): route scene-content per scene type (#745)

Extend MODEL_ROUTES with composite keys scene-content:<type> for the four core
scene types (slide/quiz/interactive/pbl). getStageModel now resolves composite
`a:b` stages most-specific-first, trimming `:` segments, so scene-content:<type>
falls back to the base scene-content route, then x-model, then DEFAULT_MODEL.
The scene-content route derives its stage from outline.type. Backward
compatible: with no composite key configured, behavior is unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(model): per-stage thinking effort + drop hardcoded gpt fallback (#745)

MODEL_ROUTES values may now be {model, effort} (not just a model string) to pin
the thinking effort per stage. Arbitration mirrors model routing: a routed stage
with effort set wins over the client's thinking (effort "none" disables); routed
with no effort uses the model's default and drops client thinking; unrouted
stages keep the client thinking. resolveModel is the single arbiter — body
thinking is threaded in via resolveModelFromRequest. classroom-generation passes
the resolved thinking into its callLLM calls so generate-classroom and
web-search-query-rewrite honor route effort too.

Also: removed the hardcoded `|| 'gpt-5.4-mini'` fallback in resolveModel — if no
model resolves (no route / x-model / DEFAULT_MODEL) it now throws instead of
silently picking a vendor default.

model-routes: getStageRoute returns {model, effort?}; getStageModel delegates to
it. Docs + tests updated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(model): route value carries full ThinkingConfig (#745)

Per-stage routing accepts the full unified ThinkingConfig in the route value
({model, thinking:{mode,effort,level,enabled,budgetTokens,excludeReasoningOutput}})
instead of just an effort string. The route's thinking is passed through
resolveModel and normalized per the model's capability by callLLM, so
budgetTokens (qwen), level (Gemini), enabled/mode, etc. all work.

(qwen3.7-plus/max thinking capability now comes from upstream #753, so this no
longer registers them.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(model): honor routed thinking for chat-adapter; doc fixes (#745)

Cross-review (codex P2 + Claude) found chat-adapter was the only routable stage
that ignored the route's thinking: it took the routed model but built its own
thinkingConfig from the request body, so operator-pinned thinking never applied
and client thinking still leaked onto a routed model. Now pass the client
thinking into resolveModel and use the resolved thinkingConfig (route-pinned for
a routed stage, client otherwise; defaults to disabled for low-latency chat).

Also: clarify DEFAULT_MODEL is now required for server-side stages (the
hardcoded gpt-5.4-mini fallback was removed → resolveModel throws if nothing
resolves), and fix a stale {model,effort}→{model,thinking} doc comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 22:26:37 +08:00
3ccd5da6da feat(generation): opt-in parallel scene-content generation (#660)
* feat(generation): opt-in parallel scene-content generation

Scene generation was strictly serial — content -> actions -> TTS, one
outline at a time — so an N-scene classroom paid N x (content + actions
+ TTS) of wall-clock latency, which dominates the post-outline wait.

Add an opt-in hybrid two-phase path (#572):

- Phase 1 fetches scene *content* concurrently (bounded). Content is the
  only per-scene step independent of cross-scene state, so it is safe to
  parallelise; a content failure marks just that outline and does not
  pause the batch.
- Phase 2 keeps the existing in-order actions + TTS loop, so
  previousSpeeches threading, ordered addScene, the abort/epoch guards,
  and the pause-on-failure UX are all unchanged.

Gated by a server-side PARALLEL_SCENE_CONCURRENCY (default 0 = off,
clamped to 10), surfaced to the client through the existing
/api/server-providers response. With it unset the content map is null
and the loop is byte-for-byte the original serial path, so out-of-box
behaviour is unchanged.

The bounded worker pool is extracted to lib/utils/concurrency.ts
(mapWithConcurrency, with an early-stop hook for abort/epoch) and
unit-tested; getParallelSceneConcurrency env parsing/clamping is tested
too.

Closes #572

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(generation): pipeline content fetches instead of a barrier

Address @wyuc's review on #660. The previous version awaited the whole
bounded content pool before Phase 2, so for N > C scenes the user saw the
first scene, then a stall on a single page for the entire content phase,
then a burst — perceived as worse than serial despite the wall-clock win
(the gap scales with ceil((N-1)/C)).

Pipeline instead: lazyBoundedMap starts the content fetches (<= C in
flight) and returns one promise per outline immediately; the serial loop
awaits them in order, so each resolves as soon as its own content is ready
(usually already done, hidden behind the previous scene's actions/TTS). The
first scene now paints after content(1)+actions(1)+TTS(1) — same as serial
— with the full concurrency benefit and no stall.

Also from the review:
- a content fetch can no longer take its siblings down: an unexpected throw
  is caught and returned as a failure result, routed through the same
  mark-failed path as the serial loop (symmetric);
- drop the cosmetic Phase-1 setCurrentGeneratingOrder — the loop already
  sets it per outline;
- note the intentional belt-and-suspenders clamp;
- concurrency.test.ts asserts <= limit (not == limit), and gains
  lazyBoundedMap tests including the no-barrier property.

mapWithConcurrency is kept as a thin await-all wrapper over lazyBoundedMap.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-06-08 07:44:33 -04:00