653 Commits
Author SHA1 Message Date
Xinyu Mu 1635b16899 docs: add multilingual vocational task engine guide (#1749) 2026-10-01 21:01:09 +08:00
8f7d51e5ac fix(playback): replay sentence visual cues after navigation (#1720)
Co-authored-by: sophietao20-star <283060850+sophietao20-star@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-29 18:22:49 +08:00
fa47efe1b3 feat(persistence)!: server-backed persistence only, with a one-way legacy browser import (#1710)
* ci: run CI for the persistence-default integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(docker): start PostgreSQL by default and add a single-user owner

* feat(docker): start PostgreSQL by default and add a single-user owner

`docker compose up` now runs a server-backed app against a bundled
PostgreSQL, started only once PostgreSQL is healthy and published on
127.0.0.1 only, with a new built-in singleUser owner auth method: every
request resolves to one fixed owner that may publish, and no anonymous
cookie is minted.

Single-user mode (OWNER_SINGLE_USER=true, OWNER_SINGLE_USER_ID default
"local") runs with or without ACCESS_CODE. Without one, the server logs
one prominent startup warning that anyone who can reach it shares, edits
and can delete the single library; it never inspects request peers or
forwarding headers. It excludes PERSISTENCE_SHARED_OWNER_ID and follows
the sharedTeam registration rules. Its principal gets a claim candidate,
so a browser's earlier anonymous work is claimed into the single owner
(OWNER_CLAIM_TRIGGER=auto in the Compose defaults).

Compose defaults live in docker-compose.defaults.env, read before
.env.local so they can be overridden. The app warns when its declared
published address (OPENMAIC_PUBLISH_ADDRESS, passed by Compose) is not
loopback while PostgreSQL uses the default password.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(docker): keep claims explicit and let .env.local set DATABASE_URL

- Compose no longer defaults OWNER_CLAIM_TRIGGER to auto: an irreversible
  merge of every browser's anonymous library into the single owner must
  not happen on upgrade. The single-user principal still gets a claim
  candidate, so POST /api/identity/claim works; the docs explain it and
  warn about auto on a deployment several people used.
- The bundled DATABASE_URL moves to docker-compose.defaults.env, read
  before .env.local, so an external database or a rotated password set
  there keeps working. The password is inserted unencoded, so it must be
  letters and digits (documented).
- Upgrade docs no longer claim browser-stored courses are copied as they
  are opened: they stay in the browser, are not deleted, and are moved by
  the automatic browser-to-server migration shipped with this work.
- Single-user mode without ACCESS_CODE now logs one warning instead of
  two; the docs note that the single owner id is permanent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence)!: server-backed persistence is the only persistence (#1707)

* feat(persistence)!: make server-backed persistence the only persistence

Courses, folders and folder membership, chat history and learner runtime,
and generated media are now always stored on the server through the
embedded /api/persistence endpoint. The browser keeps no durable data of
its own.

- Remove the build-time NEXT_PUBLIC_PERSISTENCE switch: the client
  bootstrap always configures the HTTP document, runtime and asset seams,
  and the Dockerfile, docker-compose.yml, the Pi route and the whiteboard
  runtime no longer read it. Unconfigured seams refuse to resolve instead
  of falling back to IndexedDB, and the learner key is always the
  server-derived one.
- Remove the browser write paths: the browser document, runtime and asset
  stores as backends, the browser-mode branches of stage, folder, chat,
  media, narration and import storage, the Dexie backup export/import and
  the whole-database clear. The library lists through /api/stages, folders
  through /api/folders, and membership now goes through
  /api/folders/members. Regeneration always forks to a fresh asset.
- Device-local state moves to its own IndexedDB database
  (lib/device-storage, maic-device-cache): the local media and narration
  cache and refused-bytes retention, staged PDF images, undo history,
  browser voice profiles (carried over once from the old database) and
  the auto-voice cache. "Clear Local Cache" clears exactly that and no
  longer claims to delete classrooms or chat history.
- Keep the pre-server data readable for the one-way importer: the Dexie
  schema and read-only accessors for its tables and for the browser
  document, runtime and asset databases live in lib/legacy-browser-storage,
  which nothing on the regular paths reads and which has no write path.
- Require a database: the server exits at boot without DATABASE_URL, with
  a message naming `pnpm db:up` and `docker compose up`. `pnpm db:up` /
  `pnpm db:down` start and stop the Compose postgres service alone,
  published on 127.0.0.1 through docker-compose.db.yml; .env.example
  carries the matching local DATABASE_URL.
- E2E specs seed their courses through the persistence endpoint, each
  under a fresh id, instead of writing IndexedDB.
- Guard tests pin the boot requirement, the absence of any
  NEXT_PUBLIC_PERSISTENCE read, and the legacy-storage boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: run the E2E suite against PostgreSQL

The app refuses to start without DATABASE_URL, so the E2E job gets a
postgres:16 service and a job-level DATABASE_URL. The unit, storage and
render-service jobs are unchanged and need no database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document always-on server persistence and the database requirement

README (English and Chinese), the deployment guide in every locale, the
extension cookbook and the changelog now describe PostgreSQL as required:
`pnpm db:up` for local development, `docker compose up` for a deployment,
and an external database for Vercel and other serverless hosts (the
Deploy button prompts for DATABASE_URL). They list what stays in the
browser, drop the NEXT_PUBLIC_PERSISTENCE instructions, and say that
courses an earlier browser-only build stored stay in the browser until the
one-way importer moves them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pbl): use the server-derived learner key for PBL v2 runtime

PBL v2 synchronization, hydration and drain resolved the learner key from
their default watermark KV store, which reads or mints a device key in
localStorage instead of asking runtime configuration. The runtime store is
the server one, which refuses any key but the owner's, so saving a course
with a PBL v2 scene failed before the document write, and the minted key
landed in the slot the one-way importer reads as the pre-server runtime
partition.

The learner key now comes from getLearnerKey(args.kv): the configured
server-derived key, or an explicitly injected store. The default KV store
still holds the drain watermarks and nothing else.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): pin that clearing the cache keeps legacy databases

- Seed the four pre-server databases (MAIC-Database, maic-documents,
  maic-runtime, maic-asset-pool) with real rows, run Clear Local Cache, and
  assert every database and row survives; assert the device database name
  differs from every legacy name.
- The legacy-module import rule now also catches relative static,
  side-effect, dynamic and require imports at any depth.
- New rule: the runtime learner key is never read from a KV store the
  caller made up, only from configuration or an injected store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): run the development database as its own Compose project

`pnpm db:up` / `db:down` shared the checkout's default Compose project with
`docker compose up`, so they recreated or stopped a running stack's
database. docker-compose.db.yml now declares its own project
(`openmaic-dev-db`), which gives the development database its own
container and data volume; it no longer touches the stack, and the two do
not share data.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: start the development database before `pnpm dev`

CONTRIBUTING, the getting-started guide in every locale and the startup
modes reference now run `pnpm db:up` and set DATABASE_URL before
`pnpm dev`. README, the deployment guides, the cookbook and the changelog
describe `pnpm db:up` as a separate development database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): sharpen the legacy-storage and learner-key guards

- The Clear Local Cache test closes its seeding connections to
  maic-documents and maic-runtime, so a regression deleting either fails
  on the assertion that names the database instead of by timeout.
- The legacy-module and Dexie import rules accept comments inside
  import() / require(), so a webpackIgnore hint no longer hides one.
- The learner-key rule also flags an injected-looking name that is bound
  to a default local store in the same file; its remaining limit is
  documented, with the behavioural PBL test as the guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* build(dev-db): pin the development database's Compose project

`pnpm db:up` / `db:down` now pass `-p openmaic-dev-db`, which takes
precedence over COMPOSE_PROJECT_NAME from the shell or a .env file; the
file's `name:` alone could be overridden and point the scripts at a
running stack's database. The compose test pins the flag. The docs (all
locales) and the file header say the development database is shared by
every checkout on the machine, and how to run a separate one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): import legacy browser data into the server, once (#1708)

* refactor(persistence): prepare the seams the legacy browser importer uses

- Legacy module: read-only accessors for the auto-voice cache, the course
  ids the old media and narration tables name, the pre-server learner key,
  and asset bytes from the browser asset pool.
- Narration adoption exports its row ownership rule so the importer applies
  the same one to rows of the old database.
- The quiz legacy migration is split into a non-deleting core
  (importLegacyQuizSnapshot) and the regular path, which still deletes the
  keys it migrated.
- The lazy document migration exports its snapshot canonicalizer.
- Clear Local Cache keeps the importer's per-owner ledgers, as it keeps the
  pre-server learner key.
- A library-changed window event lets an open library list courses that
  arrive in the background.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(persistence): import legacy browser data into the server, once

On the first load after the upgrade, the client moves what earlier builds
kept in this browser to the server, for the owner the server resolves:
courses from the browser document store and the original tables (with
chat, learner runtime, playback position, roster, folders and membership,
pre-runtime quiz state), their media bytes through commitToPool and the
existing write-back funnels, the bytes of server courses from earlier
opt-in server builds that only the old tables hold, and device-only rows
(failure records, refused bytes, auto-voice clips) into the device cache.

It is automatic and silent, runs after the page is idle, never writes to a
legacy store, and records every step in a per-owner ledger so it resumes
after a crash and never imports twice. Runs are serialized across tabs
with Web Locks. A course the owner already has stays authoritative; one
whose id another owner holds is imported under a derived fresh id; one the
owner deleted on the server stays deleted. Transient failures retry on a
later load with backoff; permanent ones are recorded per item.

The module is temporary and self-contained; its README lists the removal
steps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin that each owner keeps its own import ledger

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): run the legacy import against the app routes and a browser

A PostgreSQL suite routes every request the importer and the app seams
send to the real persistence, library and folder handlers, owner resolved
from the anonymous cookie: a full import, a course another owner holds, a
course the owner deleted, and an existing course whose media is filled in.
An end-to-end spec seeds an old pre-server database in Chromium, loads the
app, and checks that the course reaches the library with its narration
uploaded, stays in the browser untouched, and is visible from a fresh
browser context with the same owner cookie.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: describe the one-way import of legacy browser data

The README (and its Chinese version), the deployment guide in every locale
and the changelog now say what the importer does instead of announcing it:
courses an earlier browser-only build kept in the browser move to the
server automatically on the first load after the upgrade, and the browser
copy stays untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): probe server assets through the shared asset-URL owner

The importer asked the pool directly whether an id exists, which the
asset-URL ownership boundary forbids outside use-asset-url. It now uses
assetRefExists, the existing metadata-only probe.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): guard the importer's pool probe like every other entry point

Only a reference the pool could have issued is probed, as the lease guard
requires; the importer is added to the list of guarded entry points.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import legacy browser data once per browser

Review round 1 of the legacy browser importer.

- Once per browser, not once per owner: the data belongs to whoever used
  the browser before the upgrade, so the first owner the import runs for
  claims it and any other owner gets nothing. The ledger is one key,
  maic:legacy-import:v2, recording the owner only as a SHA-256 digest
  (an anonymous owner id is a bearer credential) and a random salt.
- Handoff: when the claiming owner is retired (403 OWNER_RETIRED), or the
  current owner holds courses or folders the importer created (a claim
  moved them), the current owner continues the unfinished items; courses
  already done are never imported again, so a claim neither duplicates nor
  resurrects them, and a half-copied course keeps its runtime.
- Fresh ids come from the salt and the legacy id through SHA-256, not from
  the owner: unlinkable, unpredictable, and the same in every tab.
- 401 pauses with backoff instead of closing the import; 403
  FORBIDDEN_LEARNER stops the run and leaves the item pending.
- Without Web Locks: ledger writes merge with the stored copy, and the
  library is listed again right before a course is placed, so a second tab
  does not import a course the first one just created.
- Media the server refuses for good gets the app's failed-media record
  instead of a dangling reference; the legacy bytes stay.
- A legacy whiteboard or PBL session is not created when the server
  already has an active one of that kind.
- A folder that is gone leaves the course unfiled instead of failed, and a
  membership naming a folder the old database lacks counts as unfiled.
- Narration rows from before the course column are adopted when exactly
  one legacy course names their key; unconvertible chat rows are noted; an
  unreadable document-store copy falls back to the original tables; PBL v2
  documents go through the save-path strip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the importer's handoff, tabs, refusals and review gaps

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: say the legacy import runs once per browser

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the first owner to load
the upgraded app in a browser claims its legacy data, what the handoff to
an account does, that the ledger holds no owner id, and how refused media
and the new pause and skip cases behave.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): note runtime the importer cannot find without the old learner key

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): let an owner confirm it absorbed a given owner through a claim

GET /api/identity/merged-from?salt=<hex>&digest=<hex> answers whether a
claim merged into the requesting owner an owner whose id hashes to
SHA-256(salt, id). It reads only the requester's own owner_merges rows,
never lists or names an owner, resolves the owner like every other
owner-scoped route, and is uncacheable. The one-way import of pre-server
browser data uses it to decide whether a different owner may continue an
import the first owner started.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): hand the legacy import to another owner only on a confirmed claim

Review round 2 of the legacy browser importer.

- The only way the browser's legacy data moves from the owner that
  claimed it to a different owner is the server confirming, through
  GET /api/identity/merged-from, that the current owner absorbed that owner
  through a claim. The inferred signals (an owner that never wrote, a
  library listing that holds imported courses) and the retired flag are
  gone; a 403 OWNER_RETIRED only stops the run and leaves items pending.
- The first owner is bound only once the run's first authenticated request
  of its own (the library listing) succeeded, so a run that dies before
  that leaves the ledger unclaimed.
- The owner is recorded as SHA-256 of the per-browser salt and the owner
  id; the server computes the same value from the salt the browser sends.
- A "not absorbed" answer is reused for an hour instead of asking on every
  load, and the claiming owner stored first survives another tab's save
  unless that tab handed the import over; completion survives too.
- Run-level failures (401, OWNER_RETIRED, FORBIDDEN_LEARNER, OWNER_BUSY)
  from any call stop the run and leave the course pending, including the
  "is the id taken" read and the library re-list, which could previously
  mark the course failed and close the import.
- Refused media without a generation request (the user's own) gets a new
  non-retryable ASSET_REFUSED code, so it shows as failed without a Retry
  that could only fail; a skipped singleton runtime session leaves a note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): pin the importer ledger's owner merge and the singleton-skip note

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the legacy import moves to another owner only on a confirmed claim

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say that the import runs once per
browser and moves to another owner only when that owner claimed the
original one, confirmed by GET /api/identity/merged-from; that the owner
is recorded as a salted digest; and how refused user media shows.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): never let a transient failure settle a legacy import item

Review round 3 of the legacy browser importer.

- A read of the old browser stores that fails in storage (an aborted
  IndexedDB transaction, a closed database) now pauses the run and leaves
  the course pending. Only a record that cannot be migrated, parsed or
  validated is skipped. This covers the course read (which no longer falls
  back to the older table copy on a storage failure), and the quiz-scene
  and speech indexes, which used to read a failure as "no scenes" or "no
  holder" and drop quiz state or narration for good.
- After a session-id collision, a failed read of the session propagates
  and is retried; only a successful read that shows the id taken skips it.
- Without Web Locks, two tabs of different owners can both find the
  ledger unbound. Every ledger write keeps the owner stored first, and the
  run now re-reads the stored ledger after binding, before folders, before
  each course and before each course document is created, and stops if
  another owner holds it.
- The "not yours" answer is reused for ten minutes instead of an hour, a
  stored wait further ahead than the longest the importer sets (a clock
  that ran ahead) counts as passed, and expired answers are pruned.
- Generated media refused without an error code keeps its Retry; only
  media nothing can regenerate gets ASSET_REFUSED. A poster that fails
  transiently leaves the video pending instead of being dropped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover transient reads, racing owners, the not-yours cache and the claim client

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): say what the ledger's salt protects, and the new pauses

The salt defeats precomputed and cross-browser tables, not a targeted
guess against one browser's ledger; the module README and the digest's
comment now say so. The README also covers storage read failures, which
pause instead of skipping, the two-owner race without Web Locks, and the
ten-minute recheck.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* feat(identity): bind a browser's legacy data to one owner on the server

The one-way import of pre-server browser data now asks the server whose
data it is, instead of deciding in the browser.

- legacy_import_bindings (browser_id primary key, owner_id, timestamps),
  provisioned with owner_merges on fresh and upgraded databases.
- POST /api/identity/legacy-import-binding {browserId}: one atomic insert
  (ON CONFLICT DO NOTHING) and a read-back; answers only whether the
  requesting owner holds the browser. Same-origin JSON, resolved and
  write-fenced like every owner write, 400 for a malformed id.
- A claim participant re-keys the claimed owner's bindings to the account
  inside the claim transaction, so the account simply holds the browser.
- Owner resolution refuses any request carrying X-OpenMAIC-Legacy-Import
  with 409 LEGACY_IMPORT_NOT_BOUND unless the owner it resolves to holds
  that browser: one central fence every owner-scoped route goes through.
  Requests without the header are unaffected.
- GET /api/identity/merged-from is removed; the binding replaces it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): import through fenced clients bound by the server

Review round 4 of the legacy browser importer.

- The ledger holds a random browser id and no owner information; the
  owner digest, the "not yours" cache and the client-side handoff are
  gone. A run first asks the server to bind the browser and imports only
  when the requesting owner holds it.
- Every importer request goes through its own clients that carry the
  browser id, so the server refuses any write under an owner that does not
  hold the browser: after a cookie switch in another tab, from a page whose
  memoized owner is stale, or from a tab that lost the binding race. The
  runtime learner key is asked for during each run, never taken from the
  page. 409 LEGACY_IMPORT_NOT_BOUND stops the run with items pending.
- The write-back funnels, commitToPool and the PBL save-path strip take an
  optional store, put or runtime, so the importer passes its fenced ones.
- A failed read of one legacy course holds up only what it could affect:
  the quiz-scene and speech indexes leave the course out as unindexed, and
  only quiz state or narration it could own waits. Read failures are
  counted per course; after five failing runs over at least a day, a
  course whose document-store copy cannot be read falls back to the table
  copy, or is skipped with the reason when there is none.
- Runtime ids are rewritten by replacing only their course segment, so a
  short course id cannot corrupt the prefix or the learner segment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover the server binding, its fence, bounded reads and short ids

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: the server binds a browser's legacy data to its first owner

The README (and its Chinese version), the deployment guide in every locale,
the changelog and the module README now say: once per browser; the server
binds the browser to the first owner; a claim carries the binding to the
account; the importer's requests are refused for any other owner. They
also describe the bounded pause for a course the old storage cannot read,
and the removal steps for the server side.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): keep a dropped asset request pending in the legacy import

The asset client reports a fetch that got no answer (a dropped request, an
aborted upload, the existence probe's deadline) as status 0 with
HTTP_REQUEST_FAILED, which the importer read as a final refusal: the media
was marked failed and the course could complete without it. Read that
code as transient, keep the client's local validation failures final, and
treat a 2xx/3xx answer a client could not use as transient too. Unit tests
drive every client the importer uses with a rejected fetch and a non-JSON
answer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(persistence): ask a refused binding again only every ten minutes

A browser whose legacy data another owner holds sent the binding request
(a write) on every page load. Record a ten-minute recheck in the ledger
(clamped like every other deferral), so a claim that moves the binding is
still picked up, and open the old databases only once the browser is
bound. The read budget now counts consecutive failing runs (a run that
reads the course resets it) and counts a first failure dated in the
future from the next one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): answer 503 with the cookie when the import fence cannot read

The legacy import fence now reads the header name from the shared
constant, and a database error while reading the binding answers the
usual JSON error (503 PERSISTENCE_UNAVAILABLE) with the resolution's
Set-Cookie values instead of escaping as a bare 500. Route tests cover
that, a bind by an owner a claim retired (403 OWNER_RETIRED, no row), and
one header name on both sides.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(persistence): cover dropped asset requests, the recheck and the read budget

With the real asset client's error for an upload and an existence probe,
the media stays pending and a later run imports it. A non-holder's five
loads send one binding request and never open the old databases, and a
claim still hands the import to the account after the delay. The read
budget's run count and time span each hold on their own, a clock that ran
ahead does not hold a course open, and a good read resets the count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(persistence): complete the legacy import removal steps

List every file the removal touches (the claim scenarios test, the claim
participant table and list, every doc), keep the bindings table for one
release after the code goes because of rolling deploys and give the DROP
statement for that release. Add the claim's eighth participant to the
README lists, fix the fresh-id wording in the changelog, and describe the
recheck delay, the transient status-0 asset failures and the consecutive
read budget.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): keep legacy quiz state through Clear Local Cache until the import completes

Pre-runtime quiz drafts, answers, results and attempt ids exist only in
localStorage, and the one-way importer copies them to the server with their
course. Clear Local Cache deleted them even while the import was still
pending (it waits for an idle page and can stay pending after a failure), so
that quiz progress was lost for good.

Clear Local Cache now keeps the four quiz key families unless the importer's
ledger records the import as complete, read through a new
`legacyImportIsComplete` helper in the ledger module. The ledger key moves
there too, so the two modules do not import each other. Once the import is
complete, the keys are cleared as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs: show DATABASE_URL set in the local development setup

The getting-started guide showed the required DATABASE_URL line commented
out, so copying the block left the server unable to start. Each locale now
shows `pnpm db:up` and, separately, the `.env.local` line to set.

The `.env.example` header said every variable is optional; it now says that
provider variables are, and that DATABASE_URL is required except under
docker-compose.yml.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): drop two upgrade notes the final code contradicts

The Compose entry said its default DATABASE_URL overrides .env.local; the
defaults file is read first, so a value in .env.local wins, as the same entry
says further on. The runtime-session entry offered turning server
persistence off as an upgrade path; that switch no longer exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(env): restore the ACCESS_CODE warning notes in .env.example

A later change to the example file dropped the lines saying that an unset
ACCESS_CODE fails open, that the server then logs a one-time startup warning
and GET /api/health reports accessCodeConfigured: false. They are back, and
say that single-user mode logs its own warning in place of the generic one,
as the startup code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* ci: run every app PostgreSQL suite in the contract job

The app-domain step named its suites one by one, and seven had been left
off, among them the legacy importer's route suite. No other job gives the
app suites a database, so they skipped everywhere. The step now lists every
app `*.pg.test.ts`; all of them work in a schema or database of their own,
or truncate only their own tables, and pass before and after the storage
package suite in the same database.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): establish the anonymous owner before the first API request (#1721)

* fix(identity): establish the anonymous owner on the page response

On a browser's first load the page sent several API requests without an
owner cookie; each minted its own anonymous owner and set its own cookie,
and the last answer to arrive won. Work done in between (the legacy import
binding, a first course write) then belonged to an owner the browser no
longer presented.

The middleware now mints the anonymous cookie on document navigations that
carry no valid one, with the route handlers' own value format and
attributes (the cookie module is Edge-safe now: Web Crypto instead of
node:crypto), and forwards it on the request so the page render resolves
the same owner. API, RSC, prefetch and Server Action requests never mint
there, and a valid cookie is never replaced. It mints only when neither
OWNER_SINGLE_USER nor PERSISTENCE_SHARED_OWNER_ID is set; host
registrations are invisible to an Edge middleware, so the README tells
hosts where to skip it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(legacy-import): bind only an owner the browser already presents

A binding request whose owner was minted by that very request answers
409 OWNER_NOT_ESTABLISHED with the minted cookie and binds nothing: other
cookieless requests may still be minting owners of their own, and a
binding to this one could be left with an owner nobody presents. The
importer treats it as transient and retries on a later load. It also says
plainly when a request after its own successful bind is refused as
another owner (the owner cookie changed mid-run); items stay pending.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): cover a first visit whose first answers arrive out of order

Against the real routes and PostgreSQL: a browser without a cookie loads a
page, sends two requests at once, and the first answer reaches it only
after the importer bound the browser; the import completes under the one
owner the page established. A second case binds nothing for an owner the
bind itself minted and imports on the next run. The E2E drives the same
ordering in Chromium by holding the first owner-scoped answer until the
binding answered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* test(legacy-import): drop an unused initial assignment

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(changelog): note the page response now sets the anonymous owner cookie

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): let hosts turn page-response minting off

The middleware cannot see owner auth methods a host registers, so a host
with anonymousFallback: false still got an anonymous cookie on a cookieless
page load. Beside the host credential it is a claim candidate; with
OWNER_CLAIM_TRIGGER=auto it is claimed and cleared on the next request,
again after every such page load.

OWNER_ANONYMOUS_PREMINT (default on; true/1/false/0, anything else fails
startup) is read by the middleware before minting. At startup the server
warns once when a registration turns the anonymous fallback off while
pre-minting is still on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): document OWNER_ANONYMOUS_PREMINT

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): keep the anonymous identity for 400 days and renew it in use

The anonymous cookie had a fixed 30-day lifetime from its mint and was
never renewed, so every anonymous visitor lost their whole library 30 days
after the first visit however active they were; with persistence on the
server only, that affects every anonymous deployment.

It now lasts 400 days (the browser cap, one shared constant), and every
route handler and Server Action response that resolves to a valid
anonymous cookie re-sends the same value with a fresh Max-Age (best-effort
where next/headers refuses writes). Page responses never renew, so pages
stay cacheable, and a host principal never renews the anonymous cookie
beside it.

A response that clears the cookie (a claim, a retired owner) must not
renew it, whatever order a route merged its headers in:
withRequestOwner, and the owner-events retired answer, keep only the
clearing value for a cookie a value clears. A renewal that lands after a
claim cleared the cookie is pinned by a test: the next write with it,
alone or beside the account, writes nothing, and its answer clears it
again without renewing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* docs(identity): 400-day renewed anonymous cookie; when to turn pre-minting off

Hosts set OWNER_ANONYMOUS_PREMINT=false only when their registration sets
anonymousFallback: false; hosts that keep anonymous visitors need
pre-minting and skip it per request in middleware for requests their
methods authenticate. The CHANGELOG notes the new lifetime and renewal,
and that a lost cookie means a new owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

* fix(identity): forward owner cookies from routes that resolve the owner directly

The Pi chat and whiteboard-visibility routes and asset-id document
extraction resolved the owner themselves and never sent its Set-Cookie
values back, so an anonymous identity used only through them was never
renewed (and a minted one never stored).

attachOwnerCookies attaches a resolution's cookies to a response a route
built itself, with the clear-wins rule, before a stream's body starts. Both
Pi routes answer through it, success and error alike; resolveServerAsset
hands the cookies back with every answer and the extraction route attaches
them.

A guard test scans for every module outside the identity seam that calls
the owner resolution (or the asset helper) directly and requires it to
forward the cookies; the two callers that cannot are listed with the
reason. Behaviour tests cover the three paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01MRME5T4oEaQ4EBUzrDXMKz

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: wyuc <dhq1204@yahoo.com>
2026-09-29 17:51:03 +08:00
3948688cfc fix(agent): preserve tool-result images in model requests [AI-assisted] (#1718)
* fix(agent): preserve tool-result images in model requests

* fix(agent): omit tool images when vision capability is unknown

---------

Co-authored-by: sophietao20-star <283060850+sophietao20-star@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-29 16:56:08 +08:00
be2eaf0a82 docs(readme): bring the News section up to v1.1.2 (#1715)
The section stopped at v1.0.0 (2026-08-27) and did not point readers at
the six releases published since, so anyone landing on the README had no
route to the current release history.

Adds one entry per release from v1.0.1 to v1.1.2 in the existing format,
each linking its GitHub release and the changelog. The four security
releases name their advisories and flag the Behavior/Breaking Changes
sections, since v1.1.x changes several defaults (caller-supplied
provider base URLs, ALLOW_LOCAL_NETWORKS, Pi as the default chat
runtime) that an upgrading user needs to read.

Dates are taken from the CHANGELOG headings. They differ by a day from the
GitHub Releases publish date for v1.0.2 and v1.1.1; the README already
carries the same kind of difference for v0.3.0, so say the word if the
release dates are preferred.

Refs #1714

Co-authored-by: Yi-111-a <41898262+github-actions[bot]@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-29 16:27:13 +08:00
Ethereal49 877bf792d7 fix(generation): preserve template literals in interactive HTML (#1706) @openmaic/generation@0.3.14 2026-09-29 15:58:28 +08:00
HEOJUNFOandClaude Opus 5.5 1c70e86a13 fix(media): poll and download Veo videos the Gemini API way (#1695)
* fix(media): poll and download Veo videos the Gemini API way

The Veo adapter targets generativelanguage.googleapis.com but polled
with Vertex AI's models/{model}:fetchPredictOperation and read inline
response.videos[].bytesBase64Encoded. The Gemini API reads operations
with GET v1beta/{name} and returns
response.generateVideoResponse.generatedSamples[].video.uri, fetched
with the same API key.

- Poll with GET v1beta/{operationName}.
- Download the sample URI with x-goog-api-key into a data URL, always
  through the configured base URL's origin and without following
  redirects, like the other adapter calls. Inline bytes are still
  accepted when present.
- Report RAI-filtered operations as a failure with their reasons.
- Catalog: Gemini API model IDs for Veo 3.1 (-preview, plus Lite) and
  only the 16:9 / 9:16 aspect ratios the API accepts.

Fixes #1694

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(media): follow the Veo download redirect through a validated transport

The Gemini API file URI can answer with a redirect to storage (Google's
own example downloads it with `curl -L`), which the download refused, so
a generation could fail after a successful poll.

- VideoGenerationConfig gets an optional downloadFetchImpl. fetchImpl is
  the pinned transport that refuses every 3xx, so the download needs its
  own one.
- lib/server/media-provider-fetch.ts adds mediaDownloadFetch and
  managedMediaDownloadFetch: providerFetch with redirects followed, each
  hop re-validated under the operator address policy, pinned to the
  vetted DNS answers, and stripped of x-goog-api-key once it leaves the
  origin (fetchWithRedirectValidation). withVideoProviderFetch installs
  both transports.
- The server video call sites (the generate route, classroom media
  generation, the agent runtime) inject it; the adapter still does not
  import server code.
- downloadVideo follows redirects only through the injected transport and
  keeps redirect: 'manual' + assertNotRedirected without it.
- Test: a 302 to another origin yields the data URL and the second hop
  carries no x-goog-api-key; without the transport the 302 is refused.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 23:00:04 +08:00
wyucandClaude Opus 5.5 3dcb158f3d docs(changelog): record the 1.1.2 security release (#1705)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 18:25:23 +08:00
wyucandClaude Opus 5.5 cd7d6a1a59 fix(providers): pin caller-chosen provider requests and keep upstream detail out of errors (#1704)
* fix(pdf): pin verify-pdf-provider probes to the strict provider transport

The MinerU Cloud and self-hosted connectivity probes validated a
caller-supplied base URL once and then issued a plain fetch that resolved
DNS again, so a rebinding hostname could pass validation and connect to a
loopback or private address. The response also echoed the target's status,
its 401/403 body, and per-errno connection errors.

Both probes now go through providerFetch with reject-redirects and the
operator address policy (the same ALLOW_LOCAL_NETWORKS policy the route
already validates against), so the connect address is pinned to the vetted
DNS answers. A 3xx still maps to REDIRECT_NOT_ALLOWED. Authentication
failures return a fixed message, every other connection failure returns a
single generic message, and the success payload no longer includes the
target status. Details are logged server-side only.

Tests drive the real route, guard and pinned transport against loopback
servers, covering rebinding, auth body suppression, identical
refused/not-found/timeout answers, redirect refusal and the managed path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pdf): pin self-hosted MinerU parsing to the strict provider transport

Self-hosted MinerU parsing validated a caller-supplied base URL once and
then posted the document with a plain fetch that re-resolved DNS and
followed redirects, so a public host could hand the upload (and API key)
to a loopback or private address.

The /file_parse request now goes through providerFetch under the operator
address policy with redirects refused. Transport failures, policy blocks
and redirects collapse into one fixed message; an unknown error status is
reported without its body (the missing-dependency classification stays),
a non-JSON body no longer surfaces the parser's input snippet, and the
empty-result error no longer lists response keys. MinerU Cloud control
plane, upload and ZIP errors likewise stop quoting raw response bodies.

Tests drive the real parse route and pinned transport against loopback
servers: rebinding, redirect refusal, body suppression, the multipart
upload shape and the managed path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tts): pin the Azure voice list request to the strict provider transport

The voice list request validated the caller-supplied base URL once and then
fetched it with a plain fetch that resolved DNS again. It also mirrored the
target's HTTP status as the route's own status and returned the full error
body, plus any JSON the target answered on success.

The request now goes through providerFetch under the operator address
policy with redirects refused and a 20s deadline. Non-2xx answers return a
fixed 502 (authentication failures get their own fixed message), only a
JSON array is returned as the voice list, and transport failures share one
fixed 500 message. Details are logged server-side only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(providers): pin model-list probing to the strict provider transport

Model discovery validated the caller-supplied base URL and models URL once
and then issued a plain fetch per candidate, which resolved DNS again. Non-2xx
answers carried up to 512 bytes of the provider's body into the route
response, and a non-JSON success body surfaced the parser's input snippet.

fetchModels now defaults to providerFetch under the operator address policy
with redirects refused (a refused hop keeps the REDIRECT_NOT_ALLOWED
contract), and accepts an injected transport for tests. Error bodies are
never read; the route keeps its 401/404/403 contracts, reports other HTTP
failures by status class only, answers a non-JSON list and every transport
failure with fixed messages, and returns 400 for a malformed request body.

The rejected-redirect detector is shared from the transport module instead
of being copied per caller.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(media): run image/video provider requests on the strict provider transport

The image and video routes validated a caller-supplied provider base URL
once, then every adapter issued plain fetch calls that resolved DNS again.
Adapter errors carried full non-2xx bodies into the route responses, the
auth-only probes returned the 401/403 body, and connectivity failures
echoed the transport error text.

Adapters now take a `fetchImpl` from their config and use it for every
provider request (submit, poll, download and connectivity probe). The
adapters are also imported by the settings UI, so they cannot import the
server transport; every server caller (the four routes, classroom media
generation and the agent runtime tools) injects mediaProviderFetch, which
is providerFetch under the operator address policy with redirects refused.

Connectivity results are fixed text: authentication failures, redirects,
other HTTP statuses and transport failures each have one message, and no
provider body is read. The generation routes log the adapter error and
answer a fixed message, keeping the content-safety classification. A
malformed Kling key is still reported before any request.

The rejected-redirect detector moves to a dependency-free module so the
browser-bundled probe helper can share it with the server transport.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(llm): pin LLM calls to a client-supplied base URL

resolveModel validated a client-supplied base URL once and then handed the
AI SDK a redirect-validating fetch whose only dispatcher was the
timeout-only agent, so connect-time DNS was resolved again without the
guard. verify-model also echoed the provider's error message (which can
carry its response body) and distinguished errno classes.

A client-supplied base URL now gets providerFetch under the operator
address policy with redirects refused. The pinned dispatcher accepts
headers/body timeouts, and this path uses the same 15-minute budget as the
default LLM dispatcher; the transport's body normalization and streaming
are unchanged. Operator-configured endpoints keep the existing
redirect-validating transport.

verify-model now classifies failures by the provider's HTTP status only
(401/403, 404, 429, other status class) and answers every transport or
parse failure with one fixed message.

Tests stream a real SSE chat completion through the pinned path and check
that deltas arrive before the server finishes, plus rebinding, redirect
refusal and the managed path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pdf): accept only official AliDocMind endpoints from clients

AliDocMind requests go through the vendor SDK, which builds its own HTTPS
agent and resolves the endpoint itself, so a client-supplied endpoint that
passed the URL guard could not be pinned to the validated address. The
credential check also returned the SDK's error code and message, which
told a refused port from a TLS failure or a timeout.

When the provider is not server-managed, verify-pdf-provider, parse-pdf and
extract-document (document and media paths) now accept a client endpoint
only when it is an official docmind-api.<region>.aliyuncs.com host over
https with no port, path or credentials, and answer INVALID_URL before any
SDK call otherwise. The accepted endpoint is passed on as the normalized
host. Server-managed endpoints are unchanged. Credential verification
failures now return fixed messages and log the SDK detail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(agent-runtime): download provider result URLs under the strict public policy

The agent-runtime image and video tools downloaded the URL a provider
returned with a plain fetch, validating each redirect hop with the operator
policy but connecting without pinning, and accepting http. A data: URL
from an adapter that inlines its result (the OpenRouter video adapter) was
rejected instead of decoded.

The classroom download helper's policy is extracted into
fetchProviderResultUrl: data: URLs are decoded locally, anything else must
be https and pass the strict public policy (allowLocalNetworks false,
regardless of ALLOW_LOCAL_NETWORKS), and the request goes through
providerFetch, which re-validates redirect hops under the same policy and
pins connect-time DNS. The agent-runtime tools and the classroom helper
share it; the existing bounded reads and content-type checks are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): note pinned provider transports and body-free provider errors

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(providers): keep vendor endpoint rules out of provider-neutral modules

The provider-neutrality guard keeps vendor knowledge out of the capability
routes and model resolution. The AliDocMind endpoint rule now sits behind
provider-neutral helpers (checkClientDocumentExtractorBaseUrl and
checkClientMediaExtractorBaseUrl), which apply the official-endpoint rule
to extractors whose SDK cannot be pinned and the URL guard to the rest.
The pinned LLM transport policy moves out of resolve-model into its own
module. Behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(providers): hold IP-literal request hosts to the transport address policy

The pinned dispatcher judges hostnames in its connect-time lookup, but Node
never runs that lookup for an IP-literal host, so a request to a loopback or
private IP reached it without the local-network opt-in. The provider transport
now checks an IP-literal origin against the same policy before connecting,
so it enforces the policy on its own rather than relying on every caller to
validate the URL first.

Tests that reached loopback IP literals without the opt-in now set it, as a
self-hosted deployment would; cloud metadata stays refused under every policy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(providers): let operator-configured local providers connect without the opt-in

Image/video providers and self-hosted MinerU moved to the pinned transport
under the operator address policy for every base URL, so a server-managed
endpoint on a local network (for example a local Lemonade image server or a
MinerU container) stopped working unless ALLOW_LOCAL_NETWORKS was set.

A server-managed base URL is operator configuration: it now runs with local
networks allowed, still pinned and still refusing redirects, while cloud
metadata and reserved ranges stay refused. Caller-supplied base URLs keep the
operator policy. Media routes pick the transport by the provider's managed
flag; server-internal media generation uses the managed transport; document
extraction carries a `managed` flag to the MinerU parsers and the PDF
verification probe.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(llm): pin every caller-chosen endpoint and keep transport detail out of errors

An unmanaged provider picked by the caller without a base URL fell back to its
catalog default (for example a localhost Ollama or Lemonade endpoint) on the
operator transport, which neither validated nor pinned the origin. The pinned
client transport is now chosen whenever the caller picked the model or sent a
base URL, and the effective endpoint (client URL or catalog default) is
validated under the operator policy first. A model the operator selected
through MODEL_ROUTES or DEFAULT_MODEL keeps the operator transport.

Routes relay LLM error messages, and the AI SDK builds them from the fetch
failure cause and from the provider's error body. On the caller-chosen
transport a failed request now surfaces a fixed reason ("connection failed",
"request timed out" or "redirects are not allowed") with the system error
logged server-side, and an HTTP error response reaches the SDK with an empty
body and the standard reason phrase, keeping its status and headers for retry
and status classification.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pdf): keep MinerU Cloud envelope text out of errors and refuse query/fragment base URLs

MinerU Cloud errors quoted the endpoint's envelope `msg`, a failed row's
`err_msg` and the batch id, and the parse routes relay error messages. They now
report the context with the HTTP status or, for a rejected request, the numeric
code only; the endpoint text is logged server-side.

Provider paths are appended to a base URL as text, so a client base URL
ending in `?` or `#` (or carrying a query) absorbs the fixed path and leaves
the request target to the caller. Client-supplied provider base URLs are now
refused when they contain a query string or fragment, across the LLM, media,
document extraction, PDF verification, model probe and Azure voice routes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(media): bound provider data: URL results before decoding them

`fetchProviderResultUrl` decoded a provider-returned `data:` URL in full and
left the size check to the caller's body reader, so an oversized payload was
materialized first. Callers now pass their byte limit, and the payload size is
estimated from the encoded length (base64: 3/4 less padding; percent-encoded:
at least one byte per three characters) and refused over the limit before any
buffer is built, then checked exactly after decoding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): describe provider transport behavior changes accurately

Note which base URLs refuse redirects (MinerU Cloud API roots still follow
validated, pinned hops), the query/fragment base URL refusal, IP-literal hosts
under the address policy, validation of an unmanaged LLM provider's built-in
default, server-configured local providers working without the opt-in, HTTPS
for agent-runtime video and poster downloads, pinned requests bypassing the
environment proxy, and the fixed LLM and MinerU Cloud error text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(llm): bound the logged error body and map out-of-range statuses to 502

An error response from a caller-chosen LLM endpoint was read in full only
to log 500 characters, and a status of 600-999 (passed through by the
transport) made the replacement Response constructor throw. Read at most
1 KB (or 1 s) of the body before cancelling it, and report statuses above
599 as 502.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(pdf): hold MinerU Cloud redirect hops to the operator policy for a managed root

A server-managed MinerU Cloud API root runs with local networks allowed,
and the same policy applied to every redirect it answered with, so a hop
to a private address was followed without ALLOW_LOCAL_NETWORKS.

The provider transport now takes a separate address policy for redirect
hops (`redirectAllowLocalNetworks`): each hop is validated and pinned on
a dispatcher for that policy. A managed MinerU Cloud root keeps local
access for itself while its hops use the operator policy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(providers): match server-configured providers by own keys only

Provider ids come from requests and were looked up on plain config
objects, so an id such as `constructor` or `toString` read an inherited
property and counted as server-configured (and its key, base URL and
models resolved from that property). All per-provider config lookups now
go through one helper that accepts only the section's own keys.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(audio): let server-configured TTS/ASR endpoints reach local networks

Server-configured image/video, MinerU and LLM providers on a local
network work without ALLOW_LOCAL_NETWORKS, but a server-configured
TTS/ASR or voice-registration endpoint given as a loopback or private IP
literal (for example a VoxCPM server at 127.0.0.1:8000) was refused
unless the opt-in was set.

The routes and the server-side narration, classroom TTS and voice-clone
paths now mark server-configured providers as `managed`; their endpoint
requests run with local networks allowed, still pinned, with cloud
metadata and reserved ranges refused, and redirect hops held to the
operator policy. Client-supplied endpoints keep the strict public
policy, and an unmanaged provider's catalog default keeps the operator
policy. Provider-returned result URLs are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(changelog): server-configured TTS/ASR local endpoints and config id lookups

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 18:01:22 +08:00
Yizuki_Ameandwyuc 4f43b48abe fix(whiteboard): anchor code replacements in document order (#1696)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-28 17:54:38 +08:00
小魔andwyuc a41bbac239 feat(llm): retry once on a fallback model for retryable failures (#1614)
* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* feat: configurable model escalation scheduler

## Motivation

Course generation can stall when the model configured for a heavy stage (e.g. `scene-content:interactive` / `scene-content:pbl`) is repeatedly rate-limited, times out, or returns empty output. Today the operator must manually switch models in `MODEL_ROUTES` and re-trigger the page; long-running generations die on a single bad call.

This PR makes the fallback automatic and observable: an optional per-stage **escalation policy** that re-runs the failed scene once with an explicitly chosen upgrade model, and logs every decision to a durable ledger viewable from a new Settings section.

## Design

- **Disabled by default, zero behavior change.** No `data/model-schedule.json` → the engine behaves exactly as before. The config is hot-read per call (250 ms cache), so saving from the panel takes effect without a server restart.
- **Explicit escalation only.** The upgrade model is resolved through the existing `resolveModel` pipeline **outside** `MODEL_ROUTES` (unrouted model string wins), reusing the full provider/model parsing (`openai:...`, `qwen:...`).
- **Trigger semantics.** `onTimeout` requires a timeout error; `onRetryableError` accepts any retryable generation error (`isRetryableGenerationError`). Content-safety style rejections are **not** retryable and **never** escalate — the safety boundary is preserved.
- **Budget guard.** Optional `budget.dailyEscalationCap` stops escalations for the day once exhausted; per-stage `max` limits retries.
- **Ledger.** Every escalation appends one line to `data/schedule-events.jsonl` (stage, scene, base → used, error class, reason); the Settings panel shows the last 20.

## Screenshots

Settings > Model Scheduling (config present):

![Model Scheduling panel](assets/model-schedule/settings-panel.png)

Ledger after an escalation (example entry):

![Ledger with entry](assets/model-schedule/ledger-with-entry.png)

## Verification

- `pnpm lint` — 0 errors
- `pnpm check` (prettier) / `pnpm check:i18n-keys` — pass (12 locales)
- `pnpm build` — Next.js production build passes
- Engine logic script — 8/8 assertions (disabled default, policy parsing, budget guard, ledger I/O)
- API A/B flow — GET null without config; PUT template → served immediately (hot reload); DELETE → back to null

## Notes

- The config schema is minimal free-form JSON (optional `models` / `budget` / `escalation` / `strictMode`); a stricter TS schema can follow in a later PR.
- `strictMode` is reserved (not enforced yet).
- The 3-tier preset template lives in the panel's "Load 3-tier preset template" button; no example file is shipped.

* Delete omai-upload-tree directory

* fix: ensure trailing newline in locale files

fix: ensure trailing newline in locale files

* chore: english comments for upstream review

chore: english comments for upstream review

* Delete components/settings/model-schedule-settings.tsx

* Delete lib/server/model-schedule.ts

* Delete app/api/model-schedule/route.ts

* Delete assets/model-schedule/ledger-with-entry.png

* Delete assets/model-schedule/settings-panel.png

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Update .env.example

* Enhance access control comments in .env.example

Added additional context and warnings for access control configuration.

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* Add files via upload

* run ci

* Add files via upload

* Add files via upload

* Add files via upload

* Supplement. env.example

Supplement. env.example

* Add files via upload

* Update .env.example

* Update fallback notes in .env.example

Clarified fallback behavior in callLLM layer documentation.

* Update llm.ts

* Update llm-fallback.test.ts

* fix(llm): round-4 fallback hardening

fix(llm): round-4 fallback hardening — server-managed gate, content-filter refusal, APICallError unwrap stop, outlines fullStream errors

* fix(llm): rebase round-4 onto main v1.1.1

fix(llm): rebase round-4 onto main v1.1.1

* fix(llm): round4c-prettier formatting+retry test 

fix(llm): round4c-prettier formatting+retry test

* test(llm): round-4d-align outline-stream

test(llm): align outline-stream and title-generator tests with round-4 call signatures

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-28 17:24:55 +08:00
wyucandClaude Opus 5.5 44ad72beba fix(persistence): serve the course library without the agent runtime; exit on invalid boot config
* fix(persistence): serve the course library without the agent runtime

The course library and folder routes (/api/stages/** and /api/folders/**)
gated on isAgentRuntimeConfigured(), so a deployment with server persistence
(DATABASE_URL set, client built with NEXT_PUBLIC_PERSISTENCE=1) but the agent
runtime off answered 404 there: the home library showed an empty,
"persistence unavailable" list and folder creation failed, although imports
persisted.

These routes only need the persistence provider and the owner-bound document
store, so they now gate on isServerPersistenceConfigured(). Agent features
(/api/agent/**, /api/skills/**, /api/materials/**) keep the runtime gate, and
without a DATABASE_URL every route still answers 404.

GET /api/agent/runtime also reports `persistence`; `enabled` and
`runtimeEnabled` are unchanged and the workbench entry still keys on
`enabled`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(boot): exit the server when boot validation refuses the configuration

Next.js logs a throw from the instrumentation register() hook as "Failed to
prepare server" but keeps listening and answers every request with 500, so a
refused configuration (malformed asset quota, pending TTL or owner lock
waits, PERSISTENCE_SHARED_OWNER_ID without ACCESS_CODE, the removed
OWNER_AUTHENTICATOR / TRUSTED_PROXY_* variables, invalid owner auth or host
hook registrations) left a process that looked alive and served nothing.

The fatal validations now run in one validateBootConfiguration() step. When
it throws, register() calls exitOnInvalidBootConfiguration(), which prints
one "[boot] Invalid server configuration" line carrying the original message,
waits for stderr to flush, and exits with code 1. The wrapper lives in its own
module so tests stub it (or process.exit) and still assert the thrown error.
register() already returns before any validation on the Edge runtime, and
warnings (unset ACCESS_CODE, model-routing checks) never exit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs: document library gating and exit on invalid boot configuration

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(boot): tell refused configuration apart from other startup failures

Every throw from the boot validation step was printed as "Invalid server
configuration" with its message only, so a startup failure that has nothing
to do with settings (a module missing from a standalone build, a bug in
startup code, a host registration call that throws) sent operators looking
for a bad environment variable.

Each fatal configuration check now runs through runConfigurationCheck(),
which marks what it throws as an InvalidBootConfigurationError with the same
message. exitOnBootFailure() (renamed from exitOnInvalidBootConfiguration)
keeps the one-line message for those, and prints anything else as
"[boot] Server startup failed" with its stack and cause. Both exit with
code 1.

The README (EN and zh) and CHANGELOG now list exactly the settings the boot
validation refuses, adding OWNER_CLAIM_TRIGGER, the ASSET_S3_BUCKET conflict
with a registered byte store and the sharedTeamAuthMethod() placement rule.
Also fixes a stray comment marker in feature-flags.ts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 17:03:36 +08:00
ba7d707b0f fix(media): route classroom media downloads through strict provider transport (#1692)
* fix(media): route classroom media downloads through strict provider transport

* style: apply prettier to classroom media transport changes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-28 14:45:42 +08:00
wyucandClaude Opus 5.5 a686481e73 feat(auth): owner identity seam — composable auth methods, owner-keyed persistence, host hooks, claims
* ci: run CI for the owner identity seam integration branch

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 1 — pluggable authenticator with today's behavior

* feat(auth): add the owner identity seam with built-in authenticators

Introduce lib/server/identity/: an OwnerAuthenticator contract
(SubjectKind, OwnerPrincipal, AuthOutcome), a single-shot, boot-validated
configureOwnerAuthenticator() registry, and the anonymousCookie and
sharedTeam built-ins that reproduce today's owner resolution exactly
(same anonymous_id cookie, owner id strings, statuses and env semantics).

Every owner-scoped route handler and the workspace Server Action now
resolve their owner through one memoized per-request resolution instead
of the three previous copies (agent-runtime/owner.ts, with-owner.ts and
the Server Action re-implementation). Publish and unpublish decide from
the principal's course:publish role instead of an anon: id prefix.

An invalid credential from an authenticator is answered with 401 on
every surface and never falls back to an anonymous owner. A source scan
keeps the anonymous owner cookie and anon: prefix checks out of code
outside the identity module.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): describe the owner identity seam and host registration

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): tighten the identity boundary guard and Server Action cookies

- Scan workspace package sources too, flag imports of the concrete
  built-in authenticators outside lib/server/identity, and police the
  slice/substring/indexOf/includes spellings of an anon: id-shape check.
  Stop re-exporting the built-ins from the host-facing index.
- Refuse setCookies returned by authenticateFromContext, as the
  authenticate() fallback already did, and document that it must write
  cookies itself through next/headers.
- Let the route-test owner stub carry explicit kind/roles, and cover the
  403 forbidden branch of publish and unpublish for a signed-in owner
  without course:publish.
- README: registration conflicts fail the boot; an invalid principal is
  rejected per request with a 500.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

* feat(auth): owner identity seam, phase 2 — runtime and assets keyed by the owner

Runtime sessions and assets on /api/persistence now take their identity
from the same memoized owner resolution as documents, instead of a
client-chosen x-learner-key behind a development bearer token that
shipped in the public bundle.

- Runtime: the storage handler's learner key is principal.ownerId. A
  path or body naming another learner answers 403 FORBIDDEN_LEARNER and
  another learner's session answers 404, whatever x-learner-key says; the
  header is no longer read. The browser learns its key from the new
  GET /api/persistence/learner-key and sends no credential of its own.
  The Pi whiteboard routes derive the same key from the owner.
  PERSISTENCE_DEV_TOKEN, NEXT_PUBLIC_PERSISTENCE_TOKEN and
  PERSISTENCE_ALLOW_INSECURE_DEV_AUTH are gone from the server path and
  the client, and server-auth.ts is deleted. Learner merge stays refused.
- Assets: allocations land in a per-owner partition (owner:<ownerId>),
  so quota is per owner and replace/delete reach only the owner's
  entries. Reads stay capability-by-id where a course viewer needs them:
  another owner's committed entry is readable while a live course
  references it. Legacy entries in the old shared partition stay
  readable by id, and may be replaced or deleted only by an owner who
  owns every course referencing them. Server-generated media is stored
  under the run owner's partition; asset-id extraction and vision image
  resolution use the same rule.
- Tombstones: runtime of a deleted course reads as absent (404, empty
  lists) and takes no new sessions or records, through a guard in the
  app's runtime composition; no package schema change.

The identity boundary scan now also fails on any source that reads
x-learner-key or the retired development token variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): describe owner-keyed runtime and assets, drop the dev token

Remove the development persistence token from the server-persistence
recipe, .env.example, the Docker build arguments and the deployment
docs; describe runtime learner keys, per-owner asset partitions and
quota, the legacy shared-partition rule and the tombstone behavior; and
record the breaking change for existing server-persistence runtime data
under Unreleased.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(storage): scope asset references to principals, typed stage refusal, uniform create collision

@openmaic/storage 0.32.0.

- PgDocumentStore takes assetReferencePrincipals: with reference
  tracking on, a write references and commits only entries held by the
  listed principals. An id naming another principal's entry records no
  row and changes no lifecycle column, like an unknown id, and the write
  is never refused. Without the option any held entry is referenced, as
  before. AssetCollector takes a matching per-document-owner function so
  the one-time backfill scopes each document the same way. This keeps a
  document naming a leaked id from committing, exposing or pinning
  another owner's allocation.
- RuntimeStageNotFoundError: a store may refuse createSession for a
  stage its host considers absent; the runtime handler answers
  404 STAGE_NOT_FOUND, recognizing the error by class or by code.
- A session create over a taken id answers 409 SESSION_ALREADY_EXISTS
  whoever holds it. It used to answer 403 for another learner's session
  and 409 for one's own, which told a caller whose session an id was.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): close owner-seam boundary gaps in references, tombstones and legacy assets

- Document writes reference and commit only the writer's own asset
  entries and legacy shared ones (owner-bound document store and the
  collector backfill), so another owner's course naming a pending id
  can no longer commit, expose or pin it.
- The foreign-read rule additionally requires the referencing live
  course to belong to the entry's owner, as defense in depth for
  reference rows that did not come from a scoped write.
- The tombstone refusal for new runtime sessions moved from a raw path
  comparison in the route into the guarded store's createSession
  (RuntimeStageNotFoundError, 404 STAGE_NOT_FOUND), so no encoded
  spelling of the path routes around it.
- Replacing or deleting a legacy shared entry checks reference
  ownership and mutates in one transaction under the entry row lock,
  which a concurrent reference insert must wait for.
- Tests build the uncommitted-but-referenced, tombstoned-with-rows and
  foreign-reference states directly so each half of the read rule is
  observable, and a PostgreSQL suite holds a competing reference open
  across a legacy delete.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(persistence): upgrade note for dev-token-gated deployments, owner-scoped references

Tell operators whose only access gate was the development token to put
the deployment behind an access code or gateway, register an owner
authenticator, or turn server persistence off before upgrading, and
describe that a course references only its owner's media.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): per-owner reference principals and a typed taken-id conflict

- PgDocumentStore's assetReferencePrincipals is now a function of the
  store's bound owner, evaluated per write, instead of a fixed list. A
  store re-bound with forOwner therefore scopes to the new owner rather
  than carrying the previous owner's principals onto its writes. It has
  the same shape as AssetCollector's backfill option, so one function
  serves both.
- PgRuntimeStore and BrowserRuntimeStore raise RuntimeSessionExistsError
  for a taken session id, and the runtime handler classifies it as
  409 SESSION_ALREADY_EXISTS even when its re-read cannot see the holder
  (for example a session a host-side wrapper hides). Recognized by class
  or by code.

Still the unreleased 0.32.0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): guard the Pi whiteboard runtime and answer hidden collisions with 409

- Every server-side runtime writer now uses the tombstone-guarded store:
  the Pi native whiteboard service gets it through the same helper as
  /api/persistence/runtime/*, so a deleted course takes no new
  whiteboard runtime either.
- Creating a session whose id is held by a session a tombstone hides
  answers the uniform 409 SESSION_ALREADY_EXISTS instead of a 500,
  covered on PGlite and on real PostgreSQL.
- The owner-bound document store passes its reference-principal
  function, so the scoping follows the bound owner.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 3 — accounts through an identity gateway (trustedProxyHeader)

* feat(auth): owner identity seam, phase 3 — trusted-proxy-header authenticator

Add the trustedProxyHeader built-in: real accounts through an identity
gateway (oauth2-proxy, Authelia, a Keycloak-based proxy, an institutional
reverse proxy) that signs the user in and forwards the verified user in a
request header. OpenMAIC never handles passwords or tokens.

- Selected with OWNER_AUTHENTICATOR=trusted-proxy and validated in
  instrumentation register(). TRUSTED_PROXY_SECRET (32-1024 printable
  ASCII characters) is required; header names default to
  x-openmaic-proxy-secret and x-forwarded-user, and an optional groups
  header with TRUSTED_PROXY_ADMIN_GROUPS grants admin. An unknown
  selector, a TRUSTED_PROXY_* variable without the selector, a shared
  owner id or a host-registered authenticator alongside it, and malformed,
  reserved or repeated header names fail the boot.
- Trust boundary: Next.js exposes no TCP peer address to route handlers,
  middleware or Server Actions, and fills x-forwarded-for from the socket
  only when the client sent none, so no address allowlist is offered. The
  gateway secret is compared in constant time (hashed, timingSafeEqual).
- Per request: a missing or wrong secret, a missing or blank user, a user
  header with a comma (duplicate lines arrive comma-joined), or a user
  whose proxy:<user> id fails the owner id guard is 401
  INVALID_CREDENTIAL, never an anonymous owner. The guard failure is a
  401 rather than a 500 because the value comes with the request. The
  user is trimmed and keeps its case. Every gateway user holds
  course:publish; kind user, assurance verified, channel proxy; no
  cookie. Route handlers and Server Actions share one code path.
- The unset-ACCESS_CODE boot warning is skipped in this mode; the gateway
  is the access gate and ACCESS_CODE stays independent.
- The identity boundary scan now also fails on gateway identity headers
  (x-forwarded-user/-groups/-email, x-auth-request-*, remote-user, the
  secret header) and TRUSTED_PROXY_* read outside lib/server/identity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): document accounts through an identity gateway

Describe the trustedProxyHeader built-in in the README owner identity
section (and its Chinese counterpart): configuration, request semantics,
the trust boundary and why it is a shared secret, the interaction with
ACCESS_CODE, boot-time validation, and an oauth2-proxy example that
injects the user, groups and secret headers. Add the .env.example block,
a SECURITY.md note and an Unreleased changelog entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): harden the trusted-proxy authenticator after review

- Reserve the header names HTTP, Next.js and forwarding proxies set
  themselves: configuring the user, groups or secret header to rsc,
  next-action, forwarded, x-forwarded-for/-host/-proto/-port, x-real-ip,
  or anything starting with x-middleware-, x-invoke-, x-nextjs- or
  next-router- now fails the boot.
- Log a one-time boot warning when TRUSTED_PROXY_ADMIN_GROUPS is set: the
  admin role is granted from the groups header, which is trusted on the
  strength of the shared secret alone, so the gateway must overwrite or
  strip it.
- POST /api/chat/pi resolves the owner up front and answers 401 when the
  authenticator rejects the credential, like every other owner-resolving
  route, instead of continuing without an owner. The anonymous default
  always resolves, so it is unaffected.
- The identity boundary scan also covers the Authentik, Azure App Service
  authentication, WebAuth, Cloudflare Access, Google IAP and AWS ALB OIDC
  header families.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): admin-groups warning and a corrected oauth2-proxy example

- Warn next to TRUSTED_PROXY_ADMIN_GROUPS (README, README-zh,
  .env.example) that the groups header is trusted on the strength of the
  shared secret alone, and list the reserved header names.
- The oauth2-proxy example now states it targets v7.14+ (nested
  claimSource/secretSource), uses the canonical bindAddress, sets
  preserveRequestValue: false on all three headers and
  insecureSkipNonce: false, forwards the subject with `claim: user`, and
  notes that the secret variable holds the literal secret. It adds the
  legacy-option migration caveat, a warning not to exempt app routes via
  skip-auth or trusted-ip options, and notes that the IdP must release the
  groups claim and group names must not contain commas.
- Changelog: the new boot checks and the /api/chat/pi 401.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 4 — host extension hooks

* feat(persistence): owner identity seam, phase 4 — host extension hooks

Let a host add product behavior at four points without forking a route,
registered once from instrumentation.ts register() like the owner
authenticator and sealed on first use. Nothing changes when no hook is
registered.

- configurePersistenceHooks({ name, authorizeCreate, onCreate, library,
  beforeAssetAllocate }):
  - authorizeCreate / onCreate run once per created course, inside the
    create transaction, in the transaction that inserts the course's
    ownership row. Saves of a course the owner already holds, and a create
    that loses a race to a concurrent create of the same id, are updates and
    run no create hook. A refusal answers 403 CREATE_REFUSED; a throw rolls
    the course, its ownership row and the host's writes back together. The
    hooks receive the request principal, or only the owner id for a
    background agent run.
  - library chooses which stage ids GET /api/stages lists; the route builds
    the usual summaries in the provider's order and drops every id the read
    path would refuse (deleted or unclaimed). folderId is shown only on the
    principal's own courses.
  - beforeAssetAllocate runs inside the storage handler's asset
    authorization step for POST /assets, after the owner is resolved and
    before the body is read; a returned Response is sent instead, with
    nothing stored and no quota counted.
- configureAssetByteStore({ name, create, signsReadUrls }) replaces the
  ASSET_S3_BUCKET switch for both the persistence provider and the asset
  collector. A host store must keep bytes outside the registry database.
  Under ASSET_BYTE_EGRESS=redirect a store that does not declare
  signsReadUrls stops the server at boot; the built-in layers keep their
  behavior.
- @openmaic/storage 0.33.0: DocumentWriteRefusedError lets a store refuse a
  write as policy; the document HTTP handler answers 403 with its code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): harden the host extension hooks after review

Correctness:
- Registration reads every hook once and stores it bound to the object the
  host passed, so hooks defined on a class prototype or as non-enumerable
  properties are kept. A spread copy dropped them silently, including the
  authorizeCreate and beforeAssetAllocate gates. Same for
  configureAssetByteStore. The unknown-key check now applies to plain
  objects only, since class instances carry their own state fields.
- The owner-bound document store carries each call's operation in its own
  async context (AsyncLocalStorage) instead of a field on the instance. One
  store is shared by every tool of an agent run; a concurrent call could
  overwrite the field before the transaction read it, so a create was gated
  as a read (no ownership row, no create hooks) or claimed another stage.
  create_stage also runs sequentially within a tool batch.

Hardening:
- isDocumentWriteRefusedError recognizes a refusal from another copy of the
  storage package by name and shape (not by code alone). The document
  handler and POST /api/stages use it.
- beforeAssetAllocate also gates PUT /assets/{id}/content. The request
  carries operation ('create' | 'replace') and the decoded assetId from the
  handler's own routing.
- Agent runs: a refusal on a background write carries a fixed message, and
  create_stage answers it with a stable tool result. The host's message never
  reaches the model.
- A byte store declaring signsReadUrls without a signer degrades to direct
  bytes with a warning instead of failing reads, and the collector no longer
  requires signing.
- DocumentActor is discriminated by source ('request' with the principal,
  'background' without one).
- Library ids the read path cannot address are dropped, and a provider
  answer is capped at 5000 ids.
- The boot-time ownership backfill logs how many courses it adopted without
  hooks.
- The server persistence provider no longer exposes an unscoped document
  store.
- The writesOutsideRegistryDatabase trust boundary and the scope of upload
  admission (HTTP uploads, not server-generated media) are documented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): refuse misspelled hooks on host classes too

A class instance skipped the unknown-key check, so a misspelled hook such
as authorizeCreat registered cleanly and never ran, leaving creation
ungated. Both configurePersistenceHooks and configureAssetByteStore now
reject any function-valued property of a class instance (own or on its
prototype chain up to Object.prototype, constructor excluded) that is not a
known key, and any unknown key of a plain object. The error names the key
and suggests the known key within edit distance 2. Host classes keep their
helpers private (#helper) or register a plain object; the README says so.

Also:
- isDocumentWriteRefusedError requires a string message from a cross-copy
  candidate, and documents that the name is the effective discriminator.
- CHANGELOG: ServerPersistenceProvider no longer exposes an unscoped
  documentStore; use the owner-bound store.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

* feat(persistence): owner identity seam, phase 5 — tenancy lives in stage_meta only

Course ownership was recorded twice: in stage_meta, which already answers
every access decision, and in document_stages.owner_id, which the storage
package scoped by. Two tenancy records double what a later claim has to
re-key, and a host that keeps ownership in its own table had to rewrite the
package's SQL to strip the column.

@openmaic/storage 0.34.0:
- PgDocumentStore never reads or writes document_stages.owner_id. An
  owner-bound store scopes listings, writes, deletes, the freshness manifest
  and folder membership through the host's ownership relation, named with
  the new documentOwnership option ({ table, stageIdColumn, ownerIdColumn,
  tombstoneColumn, claimOnCreate }, or false for no document scoping).
  Identifiers are validated, never quoted into SQL from arbitrary input.
- forOwner / ownerId without documentOwnership throws at construction,
  so a direct consumer cannot upgrade into a store that silently scopes
  nothing. An unbound store is tenant-agnostic.
- claimOnCreate inserts the ownership row in the create transaction; a
  concurrent create of the same id by another owner rolls back.
- folders: false lets a host whose document_stages has no folder_id use
  the store; folder ids are unique per owner, so membership goes through
  the ownership relation too.
- AssetCollector reads each document's owner from documentOwnership for
  assetReferencePrincipals, and refuses the function without it.
- Schema: fresh installs get no owner_id column and none of its indexes;
  an existing column is kept for one release, made nullable with its
  default dropped (each ALTER guarded by a catalog check, so a replay takes
  no lock), and its two indexes are dropped. document_stages_folder_idx
  replaces the owner/folder index.

Application:
- The owner-bound store binds the package to stage_meta
  (STAGE_META_OWNERSHIP) and lists through it in one query; the asset
  collector schedule passes the same relation.
- The boot backfill copies column-only owners into stage_meta before any
  store is built, only when the legacy column exists, idempotently, and
  logs how many it adopted and how many disagree (stage_meta stands).

Tests cover a fresh install and an upgrade from the previous schema on
PGlite and PostgreSQL 16, a host table with neither ownership nor folder
columns, folder isolation across owners sharing a folder id, and the
concurrent-create race.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(persistence): serialize schema bootstrap; guard unscoped owner-bound stores

Review follow-ups for the tenancy change.

- Concurrent starts: two instances starting against one database raced on
  the schema DDL (a duplicate pg_class / pg_type entry, or "tuple
  concurrently updated"), failing one instance's first request. Every
  schema bootstrap the application runs -- the persistence provider's
  sequence (runtime, documents, stage_meta and its ownership backfill,
  owner materials, assets), the agent-session, session-material and
  user-skill stores, and the asset collector's -- now runs under one
  PostgreSQL advisory lock held on a dedicated connection and released in
  finally. It lives in the application because what needs serializing is
  the whole sequence, including tables the package does not own. A
  PostgreSQL test boots eight instances at once, four rounds each, on a
  fresh and on an upgraded database; without the lock it fails every run.
- documentOwnership: false on an owner-bound PgDocumentStore now also
  requires allowCrossOwnerDocumentAccess: true, so binding an owner for
  folders or asset principals cannot silently expose every owner's
  documents.
- Documented that an ownership relation must cascade with the document
  rows (a leftover row keeps the id reserved for its owner, which is also
  how a host keeps retired ids from being reused), with tests for both.
  Leftover rows are deliberately not made claimable: a host cannot be
  told apart from one that keeps them as retirement markers.
- Upgrade notes: the one-time table locks on the first start, and the
  out-of-band CREATE INDEX CONCURRENTLY / DROP INDEX CONCURRENTLY path
  that leaves the start with nothing to build or drop.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): leave the schema bootstrap lock out of the held-lock check

The create-hooks suite asserts that its host lock is released with the
create transaction by counting granted advisory locks database-wide. A suite
booting a provider against the same database at that moment now holds the
schema bootstrap lock, which is not the lock under test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

* feat(storage): owner merge primitives for claiming anonymous work (0.35.0)

Add the per-store pieces a host needs to move one owner's data to another
inside a single transaction of its own:

- reassignDocumentFolders(tx, { fromOwnerId, toOwnerId, documentOwnership })
  moves folders and filing: a same-named folder merges into the target's, a
  colliding id is renumbered, and filing is rewritten in one statement from
  the old ids.
- PgAssetStore.reassignPrincipal(fromKey, toKey) moves a principal's entries
  under both principals' write locks, in key order; quota is not re-checked.
- PgUserSkillStore.mergeOwner(from, to) moves skills, renaming a live handle
  the target already uses with the first free numeric suffix.
- PgUserSkillStore takes a resolveFinalOwner hook, run as the create
  transaction's first statement, like the agent-session store's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* feat(auth): owner identity seam, phase 6 — claiming anonymous work on sign-in

A visitor who works anonymously and then signs in can move that work to their
account in one transaction, and the anonymous id is retired afterwards.

- Principal: pendingClaim { fromOwnerId, assurance }. The trusted-proxy
  built-in sets it when a gateway request also carries a valid anonymous
  cookie; the anonymous and shared-team built-ins never do. Authenticators may
  implement describeStoredOwner, canonicalize and clearPendingClaim;
  principalFromStoredOwner describes a stored id for work without a request.
- Claims: claimOwner / claimPendingOwner and registerClaimParticipant (sealed
  on first claim). One transaction, both owners' identity locks taken
  exclusively in key order, participants in a fixed order (folders, courses,
  materials, agent sessions, skills, runtime, assets), recorded in the new
  owner_merges table. Idempotent per pair; refuses a non-anonymous source, an
  anonymous target, a source already claimed elsewhere, and chains. Same-named
  folders merge, colliding folder ids are renumbered, colliding skill handles
  get a suffix, and quotas are not applied to what moves.
- Identity lock: every owner write (documents, folders, asset allocations,
  material registration, agent sessions, skills) takes the owner's advisory
  lock in shared mode first, so a write racing a claim is moved or refused,
  never orphaned or deadlocked.
- Retired ids: request writes are refused with 403 OWNER_RETIRED; agent runs
  and media generation that started before the claim follow the id to the
  account (forwarding document store, canonicalized stage probes, forwarded
  generated assets and skills).
- Trigger: POST /api/identity/claim (same-origin JSON only; clears the
  anonymous cookie), or OWNER_CLAIM_TRIGGER=auto on the first request carrying
  a pending claim. The runtime contract's learner merge now performs the same
  claim, allowed only for the request's own pending claim.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(persistence): count only the suite's own advisory locks after concurrent creates

The check that a host's transaction-scoped lock is released counted every
granted advisory lock in the database. Suites that run against the same
database at the same time hold advisory locks of their own (a schema
bootstrap, the owner identity locks a claim takes), which made it fail
depending on timing. Count only locks held by this suite's connections.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(auth): use a neutral example table for host claim participants

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(storage): owner merge fixes and store policy errors over HTTP (review round 1)

- PgUserSkillStore.mergeOwner reserves every handle either owner holds before
  renaming: a claimed skill whose handle the target uses could be renamed into
  a handle another claimed skill still held, which violated the per-owner
  unique index and made that merge fail every time.
- PgUserSkillStore.list orders ties by id.
- PgRuntimeStore takes resolveFinalLearner(tx, learnerKey), run as the
  create transaction's first statement, so a host can fence session creates
  against its owner merges; reassignLearner(from, to) re-keys sessions without
  re-validating them, so one session a newer version wrote cannot block a
  merge.
- Every HTTP handler (runtime, documents, assets) answers a
  DocumentWriteRefusedError as 403 with its code, and the new StorageBusyError
  (code, message, retryAfterSeconds) as 503 with Retry-After.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 1 — fenced runtime creates, bounded locks, retired-cookie recovery

- Runtime session creates take the owner's identity lock as the create
  transaction's first statement and refuse a retired owner, closing the window
  in which a create racing a claim was left under the retired id.
- Lock waits are bounded: a claim waits OWNER_CLAIM_LOCK_WAIT_MS (default 5 s)
  for its identity locks, a write OWNER_WRITE_LOCK_WAIT_MS (default 30 s);
  running out, or losing a deadlock or serialization race, answers 503
  OWNER_BUSY with Retry-After on the claim routes and on every fenced write.
- Every 403 OWNER_RETIRED carries the Set-Cookie values that drop the retired
  anonymous credential (the anonymous built-in now implements
  clearPendingClaim), so a browser whose claim response was lost recovers.
  Writes by id to rows that moved (skill delete, session messages) answer
  OWNER_RETIRED instead of not-found or a bare 403.
- OWNER_CLAIM_TRIGGER=auto no longer claims ahead of the explicit claim
  routes, which now report their own claim instead of a false refusal.
- The owner-events stream tells a retired owner only that it moved, never the
  claiming account's id.
- Creates of one course id take turns (a per-id advisory lock before the
  ownership probe), so a concurrent second create saves as an update instead
  of failing as reserved-document.
- A claim's source must be described as anonymous by the authenticator; the
  write fences skip the retirement lookup for every other owner. The host
  canonicalize hook is removed: forwarding is core's owner_merges alone.
- The folders participant locks the source's stage_meta rows first; runtime
  sessions are re-keyed without re-validation; the owner asset store fences
  every allocation and refuses to allocate without transactions; background
  skill reads and patches follow a mid-run claim; the forwarding document
  store forwards methods only.
- Docs: when a pending claim can arise with the shipped authenticators, the
  anonymous cookie as a bearer credential, bounded waits, the exact scope of
  OWNER_RETIRED, and claims narrowed to what the tests show.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(auth): claims review round 2 — merge-record guard, boot-checked lock waits, stream credential drop

- owner_merges records claims of anonymous owners only, since the write
  fences enforce retirement only for ids the authenticator describes as
  anonymous. Reading a row that retires any other owner now fails loudly
  instead of being followed but not enforced. The comment that pointed hosts
  at claimOwner for their own account merges is corrected: a host merging two
  signed-in accounts moves the rows itself and refuses the merged-away account
  in its authenticator. describeStoredOwner must classify ids stably.
- OWNER_WRITE_LOCK_WAIT_MS and OWNER_CLAIM_LOCK_WAIT_MS are validated at boot
  with the same parser the lock paths use, so a malformed value stops the
  server instead of failing every write.
- The owner-events stream answers a connect by an already retired identity
  with one owner_moved event and the Set-Cookie values that drop the retired
  credential, so a stale tab's reconnect loop ends after one round trip.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* chore(auth): final polish before landing the owner identity seam

* chore(auth): final polish before landing the owner identity seam

- Pi routes refuse a rejected credential with the seam's INVALID_CREDENTIAL body.
- Drop the unused LegacyAssetMutations alias and resolveLockWaitMs re-export.
- Scope the document_stages.owner_id deprecation wording: the store no longer
  reads or writes it; the boot backfill reads it and a claim mirrors ownership
  into it for rollback.
- Document OWNER_CLAIM_TRIGGER and the two lock-wait variables in .env.example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(storage): 0.35.1 for the README wording fix

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* feat(identity): one identity entry point with composable auth methods

* feat(identity): compose owner auth methods behind one entry point

Owner resolution now asks an ordered list of owner auth methods instead
of one configured authenticator. Each method answers `authenticated`,
`not-applicable` (no credential of its kind) or `invalid` (its credential
is present but bad). Core keeps the first `authenticated` answer, refuses
any `invalid` answer with 401 INVALID_CREDENTIAL without asking later
methods, and when no method applies falls back to the anonymous cookie
owner, or answers 401 when the host turned the fallback off. Route
handlers and Server Actions share the same order and per-request
memoization.

- configureOwnerAuthentication({ methods, anonymousFallback? }) replaces
  configureOwnerAuthenticator: single-shot, sealed on first use, and
  validated at registration and at boot.
- anonymousCookie becomes the built-in fallback method; sharedTeam
  becomes a method, selected by PERSISTENCE_SHARED_OWNER_ID as before
  when nothing is registered, and kept beside host methods only when the
  host lists sharedTeamAuthMethod() last. The variable set beside a
  registration that leaves it out fails the boot.
- Core attaches pendingClaim itself: a host method's non-anonymous
  principal on a request that also carries a valid anonymous cookie. A
  method may not set one. Claim cookie clearing and OWNER_RETIRED
  recovery go through the anonymous method's clearCredential.
- principalFromStoredOwner consults the anonymous method first, then the
  configured methods' describeStoredOwner in order.
- Remove the built-in trusted-proxy header authenticator with
  OWNER_AUTHENTICATOR and every TRUSTED_PROXY_* variable. The boundary
  test now also keeps the resolution internals inside
  lib/server/identity/ and asserts no source reads the removed variables.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): document owner auth methods and a gateway JWT recipe

Rewrite the owner identity section around the ordered auth methods: the
three answers and what resolution does with each, the anonymous
fallback, registration and its boot checks, and the sharedTeam rules.
Replace the removed built-in gateway authenticator's documentation with
a short host-owned recipe that verifies a gateway-forwarded signed JWT
(oauth2-proxy, Cloudflare Access, IAP) against the identity provider's
keys with the jose library. Update the claim section for the core-owned
claim candidate, and the Chinese README, SECURITY.md, CHANGELOG and the
deployment docs in every locale to match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(identity): tighten retired-owner cookies, retired config and the boundary

- On 403 OWNER_RETIRED, core now calls clearCredential only on the
  anonymous fallback and on methods that declare
  `issuesAnonymousOwners: true`, so an account method's session cookie is
  never cleared there. A clearCredential that throws is logged and
  skipped instead of turning the refusal into a 500.
- Setting any retired OWNER_AUTHENTICATOR / TRUSTED_PROXY_* variable fails
  the boot with a pointer to the auth-methods docs, instead of silently
  serving every request anonymously. Only the names are matched, in one
  module the boundary test checks never reads a value.
- Add lib/server/identity/host/ for host auth methods. The boundary test
  now scans core identity files for gateway identity headers too, polices
  reads of an incoming Authorization header everywhere but the host
  directory, and keeps resolution internals out of host code.
- Document on OwnerAuthMethodResult that a present but unusable
  credential must answer `invalid`, never `not-applicable`, and that a
  method that cannot decide throws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(identity): fix the gateway JWT recipe's error split and add notes

- The recipe now answers `invalid` only for token failures (expired,
  claim validation, malformed JWS/JWT, bad signature, no or several
  matching keys, disallowed or unsupported algorithm) and rethrows
  everything else: jose reports a key endpoint that is down, answers
  non-200 or returns unparsable data as a generic JOSEError, which the
  previous version turned into a 401 for every user.
- The recipe file goes in lib/server/identity/host/, with an accurate
  description of what the boundary test enforces.
- Document invalid versus not-applicable for malformed credentials,
  issuesAnonymousOwners, the boot failure on retired variables, and that
  the claim candidate is an unsigned bearer cookie a sibling subdomain can
  shadow (serve on a registrable domain of its own; cookie hardening is
  deferred). Mirror the notes in the Chinese README, SECURITY.md and the
  changelog.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@openmaic/storage@0.35.1
2026-09-28 11:54:26 +08:00
Nguyen Quang Thiepandwyuc f2875426ae fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn) (#1486)
* fix: pass reasoning_content back to DeepSeek in thinking mode (multi-turn)

DeepSeek rejects multi-turn thinking-mode requests whose assistant messages
lack reasoning_content:

  400 invalid_request_error:
  "The `reasoning_content` in the thinking mode must be passed back to the API."

Two gaps caused it:

1. Round-trip loss. Responses fold reasoning_content into an inline
   <think> block (then extractReasoningMiddleware splits it into reasoning
   parts), but on the next turn the OpenAI chat adapter drops reasoning
   parts. Kimi already had a preservation mechanism (marker encoding +
   restore); extend it to DeepSeek, in both compatFetch and the middleware
   assembly.

2. Missing field. A response that skipped reasoning (direct tool call)
   produces an assistant message with no reasoning part at all. When
   thinking is enabled, DeepSeek still requires the field, so inject an
   empty reasoning_content for such messages.

Also: send the thinking toggle alone (no reasoning_effort) when the route
sets no explicit effort. Tool-using transports (maic-agent-driver) cannot
combine reasoning_effort with function tools; DeepSeek applies its own
default effort.

Verified: vitest (reasoning-sse, thinking-config, openai-sdk-integration),
Docker production build, and end-to-end multi-turn agent sessions with
thinking enabled (DeepSeek v4-pro + v4-flash) completing successfully.

* fix: close the DeepSeek reasoning round-trip on the agent-driver path

Review follow-up on the DeepSeek reasoning_content fix.

- One shared predicate, preservesReasoning(providerId, modelId), derived from
  the resolved request adapter (deepseek + atlascloud deepseek models) plus
  Kimi K3, now gates both the request-side wiring and the agent driver's
  includeReasoning. The driver previously gated on kimi-k3 alone, so every
  thinking block was dropped before the AI SDK saw the messages and the
  middleware had nothing to encode.
- Driving the gate through the predicate covers atlascloud deepseek models,
  which share the request adapter but were previously excluded by a literal
  provider-id check.
- Scope the effort-less thinking toggle to requests that carry tools: a
  tool-carrying request can never set reasoning_effort, while non-tool calls
  keep the historical default effort.
- Inject the catalog thinking default for preserved providers so the wire
  always states the mode explicitly, and gate both restore and backfill on
  that resolved mode (a disabled turn carries neither markers nor field).
- Tests: DeepSeek counterparts of the Kimi K3 recorded-shape test covering
  the round-trip across tool continuations, the empty backfill, the
  reasoning_effort scoping, a non-preserving provider, the non-streaming
  path, and the predicate itself. Verified they fail when the predicate is
  disabled.

* fix: key the reasoning round-trip on the resolved thinking mode; cover the driver wire

Round-2 review follow-up.

- A disabled turn no longer ships the internal marker. The request-side step
  now keys on the resolved ThinkingConfig instead of the serialized body, so a
  disabled turn strips the sentinel instead of sending it verbatim; this also
  covers Kimi K3, whose adapter emits reasoning_effort and no thinking object.
- Backfill an empty reasoning_content only for the DeepSeek request adapter.
- Add the missing driver-wire test: createCallLlmStreamFn with a DeepSeek turn
  carries the prior thinking block as non-empty reasoning_content, and a
  disabled turn strips the marker. Reverting the driver gate to the old
  Kimi-K3-only check fails the first test; disabling the strip fails the second.
- Formatting: prettier clean for the touched files.

* fix: scope the effort-less thinking toggle to tool-carrying requests

Review follow-up. `options.hasTools || config.effort === undefined` also
dropped reasoning_effort for NON-tool calls that enable thinking without an
explicit effort (e.g. {enabled:true}), where the historical wire carried
reasoning_effort:'high'. Only a tool-carrying request must omit the effort
(the transport rejects function tools combined with it); every other request
keeps the explicit value or the default.

Updated the scoping test to assert reasoning_effort:'high' for the non-tool
no-effort case, and verified the assertion fails when the old condition is
restored.

* fix: derive the round-trip decision from the wire, not the requested mode

Round-3 review follow-up. Two remaining issues:

1. Kimi K3 disabled strips reasoning_content while the request still reasons.
   getThinkingMode({enabled:false}) is 'disabled', but the openai-adapter body
   builder emits reasoning_effort:'low' (no 'none' level exists), so thinking
   was never actually off on the wire. getCompatThinkingBodyParams now also
   reports disablesThinking, computed per adapter from the params it really
   emits (true only where the wire carries an explicit off switch); the
   round-trip gate keys on that instead of the requested config. Same fix
   covers DeepSeek {effort:'none'}, which sends thinking:{type:'disabled'} but
   previously went through restore.

2. Empty reasoning leaks the private marker into content. extractKimiReasoning
   now reports a found flag, restore/strip act on it (marker removed in both
   cases; field emitted only by restore), and the preservation middleware no
   longer encodes empty reasoning parts, so the degenerate `:0:` marker is
   never produced. Restore/strip share one internal helper.

Tests: Kimi K3 {enabled:false} recorded body (field kept + reasoning_effort
low present, no thinking object); DeepSeek {effort:'none'} (disabled object,
no field, no marker); driver-wire empty-thinking cases (enabled: field '' +
clean, disabled: no field + clean); Kimi K3 no-backfill pin; unit cases for
the empty marker. Each pinned by mutation: old gate fails T1+T2, serialized
absent-strip gate fails T1 plus the pre-existing Kimi test, found-revert
fails the empty-marker tests, widened backfill fails T4. Full suites green,
tsc and prettier clean.

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-27 14:31:26 +08:00
Oliveryfasoandwyuc c125a7ac58 fix(workbench): localize material upload type and quota errors (#1672)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-27 14:11:32 +08:00
LING DUANandwyuc bc8d304bb3 fix(pi): reject malformed reference elements with validation errors (#1682)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-27 12:43:10 +08:00
LING DUANandwyuc 04baad7c45 fix(pi): distinguish state errors from stale whiteboard references (#1681)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-27 01:01:20 +08:00
wyucandClaude Opus 5.5 33553362be release: v1.1.1 (#1689)
Bump the application version to 1.1.1 and record the changelog. This is a
security release: MinerU Cloud parsing now routes every request through the
strict provider transport, holds the provider-supplied upload and result URLs
to the HTTPS public address policy, and bounds what it reads and decompresses.
There are no other changes since 1.1.0.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
v1.1.1
2026-09-27 00:20:10 +08:00
wyucandClaude Opus 5.5 44762d9db2 fix(pdf): route MinerU Cloud requests through the strict provider transport (#1688)
The MinerU Cloud parser issued its API-root, presigned-upload and result-ZIP
requests with the default fetch, so a URL carried in a provider response could
direct a further request at an arbitrary destination and redirects were followed
without re-validation.

Route every MinerU Cloud request through the strict provider transport — the
redirect-validating, DNS-pinning helper already used by the audio providers,
now re-exported under the neutral `providerFetch` name:

- The configured API root runs under the operator's address policy (the
  ALLOW_LOCAL_NETWORKS opt-in applies), preserving self-hosted and BYOK
  reachability while keeping per-hop redirect re-validation and DNS pinning.
- The response-supplied presigned upload URL always uses the strict public
  policy with redirects rejected outright: a 3xx answer is a hard failure and
  the body is never forwarded to a redirect target.
- The response-supplied result ZIP URL uses the strict public policy and
  requires HTTPS on every followed redirect hop, via a new opt-in
  `requireHttps` transport policy that defaults off so audio behavior is
  unchanged.
- Address-policy refusals are terminal and are not retried; the retry wrapper
  rethrows the original UnsafeNetworkTargetError when one is present.

Bound the untrusted response bodies:

- JSON control-plane responses are capped at 8 MiB and the result ZIP at
  256 MiB, enforced from content-length and while streaming.
- Before extraction the result archive is limited to 10,000 entries and to a
  512 MiB declared uncompressed total: a cheap preflight over header fields
  that can understate the real payload. While decompressing, every text entry
  is capped at 64 MiB, every image at 32 MiB, and a 512 MiB running total is
  enforced from bytes counted as they stream out of the decompressor, so
  extraction aborts as soon as a limit is crossed instead of buffering the
  whole entry and checking afterwards. This also surfaces the decompressor's
  own size-mismatch error for entries whose declared size does not match their
  payload.

Also classify the deprecated IPv4-compatible IPv6 range (::/96, excluding the
unspecified `::` and loopback `::1`) by its embedded IPv4 in the shared SSRF
guard, so `[::7f00:1]` and `[::a9fe:a9fe]` are refused at both the URL layer
and the connect-time pinned lookup.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-27 00:07:35 +08:00
wyuc 21d83ec51b release: v1.1.0 (#1677) v1.1.0 2026-09-24 14:04:26 +08:00
PercyandPercy 16d6d09f58 feat(attribution): send X-APP-URL on TokenDance gateway requests (#1675)
Co-authored-by: Percy <percy@PercydeMacBook-Pro.local>
2026-09-24 13:46:38 +08:00
62b4598ed4 feat(token-plan): add the Kimi coding plan and promote the Kimi provider [AI-assisted] (#1664)
* feat(token-plan): add the Kimi coding plan and promote the Kimi provider

- New Kimi preset at the END of TOKEN_PLAN_PRESETS (lowest arbitration
  priority): LLM-only on the existing `kimi` provider via Moonshot's
  OpenAI-compatible endpoint; the plan catalogue seeds k3, k3-256k,
  kimi-for-coding and kimi-for-coding-highspeed (thinking metadata
  included), with the mainline overridden to kimi-for-coding (K2.8);
  other modalities are not declared and stay untouched; personal keys
  stay protected by the enrollment marker; no balance-management link
  (TokenDance only)
- Provider promotion slot: Kimi pinned to the top of the Model Services
  list with a persistent primary ring and a signup link (rendered as a
  sibling overlay of the row button, locale-routed between
  platform.kimi.com / platform.kimi.ai), plus a Mainland/International
  links row in the provider panel; the providers column defaults to the
  promoted provider
- Plan panel: for presets declaring subscribeUrls the "Manage account"
  label becomes plain text followed by Mainland/International subscribe
  links (kimi.com/code, kimi.ai/code); all links carry ?aff=openmaic
- Built-in dual-identity providers are never deletable: the delete gate
  (menu visibility + click + confirm) checks the provider registry
  instead of the persisted isBuiltIn flag, which a token-plan apply
  overwrites to false while riding kimi/minimax/doubao/tokendance
- Selected style unified back to the standard ring used by every other
  service column (promotion emphasis applies only to unselected rows)
- i18n: signup link copy in all 12 locales; preset tests for order,
  LLM-only shape, catalogue/mainline seeding, yield semantics, removal
  restore; #620 spec now selects the provider it tests explicitly

* fix(token-plan): point the Kimi coding plan at its dedicated endpoint

Coding-plan keys do not work on Moonshot's Open Platform endpoint: the
outline request to api.moonshot.cn/v1/chat/completions fails with
"Not found the model kimi-for-coding or Permission denied", stalling
course generation (review P0 on #1664). The preset now targets the
plan's dedicated endpoint api.kimi.com/coding/v1 (international
api.kimi.ai/coding/v1); disconnect still restores the registry default
so personal keys keep using the Moonshot endpoint.

---------

Co-authored-by: Percy <percy@PercydeMacBook-Pro.local>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-24 11:53:11 +08:00
LING DUANandClaude Opus 5.5 d23bc104dc feat(pi): reference a whiteboard element in classroom chat (#1656)
Let a learner pick one element on the classroom whiteboard and ask about
it, alongside the existing PPT element and interactive component
references behind the courseware-reference flag.

- Add a `whiteboard_element` reference: { kind, whiteboardId, elementId }.
  The Host resolves it only from this request's Stage snapshot
  (`stage.whiteboard[0]`), projects the element with the existing slide
  element projection, and passes it through the existing Director ->
  call_agent evidence path. It grants no Spotlight or other tool access.
- Reuse the playback pick overlay inside the whiteboard, scoped to the
  whiteboard DOM so pan/zoom and overlapping elements resolve correctly,
  and keep a highlight on the selected element.
- Share one displayed-whiteboard selector for display, picking, and the
  pre-send check. A runtime (RuntimeStore) whiteboard disables picking;
  runtime-board references are out of scope.
- Before sending, re-check the reference against the snapshot actually
  POSTed; if it is stale, keep the learner's text and ask them to pick
  again. Server rejections carry a `whiteboard_reference_changed` reason.
- Describe whiteboard references correctly next to page-reported state,
  and rename the entry to "Reference content" in all locales.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 11:13:03 +08:00
a87c94c89d feat(settings): restructure settings IA with course model config and token plan controls [AI-assisted] (#1644)
* feat(generation): propagate per-stage model routes via x-model-routes header

Course model config lets each pipeline stage pin its own model; send the
user-level stage routes alongside the main model credentials so servers
resolve overrides consistently. Server-side model-routes/resolve-model
learn to merge them; scene renderers, the scene generator and the
generation preview pass the header through.

* feat(settings): token plan enrollment markers, plan toggle and store rework

- Plans connect through an explicit tokenPlanEnrollments marker: a personal
  key on a shared provider (minimax/tokendance/doubao) is never mistaken for
  a connected plan, so startup reconciliation no longer hijacks its catalogue
- New per-plan authorization toggle (tokenPlanDisabled): toggling cascades
  to the plan's per-modality provider enabled flags while skipping providers
  another usable plan still owns; re-enabling clears the seed fingerprint and
  re-seeds through reconciliation
- Multi-plan priority: exclusive slots (main model, stage routes, per-modality
  selections) yield to plans earlier in TOKEN_PLAN_PRESETS order regardless
  of write order; picker ordering follows the same order
- Reconciliation skips plans disabled by the toggle; disconnect restores the
  built-in provider's enabled flag and clears stale authorization markers
- Stage routes are pruned when a provider becomes unusable; the homepage
  web-search toggle migrates into the settings store; server-provider init
  waits for persist hydration

* feat(settings): five-section IA with course model config and provider panels

- Settings dialog reorganized into Token Plan / Course Model Config /
  Model Services / Skills / System
- New Course Model Config section: station cards along the generation
  pipeline with a per-station inspector for model overrides, mainline
  model plus fine-grained sub-stage routes, invalid selections kept as
  disabled rows instead of blank triggers
- Model Services becomes a provider list + detail panel with a per-provider
  authorization toggle (disabled providers leave course model config);
  Qwen TTS voice-clone area redesigned (45:55 split, add-clone CTA)
- Token Plan panel gains the Enable-this-plan switch mirroring the provider
  toggle, connected/off status copy, and priority-ordered model groups

* refactor(generation): drop the media popover now that media providers live in settings

Media provider/model selection moved into Model Services and Course Model
Config; the generation-bar media popover (520 lines) goes away and the
toolbar slims to the remaining toggles.

* fix(ui): keep provider popovers inside the viewport

Voice/config popovers now cap their height against
--radix-popover-content-available-height and their width against the
viewport, so search boxes no longer clip near screen edges.

* feat(i18n): complete course-model stations and settings copy across 12 locales

- Translate stations.* labels and descriptions for the 9 locales that fell
  back to English; convert zh-TW descriptions to Traditional Chinese
- Add subStages.*, mainlineUnset, pickModel/pickProvider, optionInvalid and
  the token-plan toggle copy across all locales

* test(audio): mock getStageRoutesHeaderValue in refused-narration test

Companion mock to the x-model-routes plumbing commit.

* style: apply prettier to files touched by the settings IA refactor

* fix(settings): restore the homepage model pill and set-up CTA lost in the IA refactor

The refactor dropped the toolbar's model picker and "Set up model" CTA,
leaving the homepage without any model affordance (#580 dead-end) and
breaking the picker stress tests (#1447), the managed-provider specs
(#620) and the KV-persistence checks that anchor on the pill.

- Re-add the pill on the shared ModelPicker component (aria-label
  "Provider / Model"), backed by the extracted useLLMPickerGroups hook so
  the toolbar and course model config share one filtered, priority-ordered
  provider list
- Re-add the amber "Set up model" CTA when no usable provider exists; it
  opens settings at Model Services
- Specs: navigate to Model Services in #620 and scope its catalog
  assertion to the dialog; drop the provider-tab step in #1447 (providers
  are group headings now)

* fix(settings): write the full runtime stage set for classroom interaction

The interaction station wrote only chat-adapter, while quiz grading and
the PBL v2 runtime resolve quiz-grade / pbl-v2-runtime:*, so the override
never reached them and calls fell back to the main model (review P0).

- Add lib/config/station-stage-keys.ts as the single station -> runtime
  stage-keys contract; interaction now writes chat-adapter, quiz-grade,
  pbl-chat and pbl-v2-runtime together (PBL composite sub-stages inherit
  the base key via the existing colon fallback)
- Override and restore-follow now write/clear the whole key group,
  thinking-config changes included
- Regression tests: contract keys are LLM_STAGES members, no orphan
  runtime stages, the interaction override resolves for every runtime
  endpoint, and explicit sub-stage overrides keep precedence over
  inherited parents

* fix(settings): enforce provider authorization across sync fallback and gates

A disabled plan's provider could be resurrected on refresh: validateProvider
correctly rejected it, but the auto-recover branch adopted fallback[0],
whose order still included authorization-disabled entries; the homepage
generate gate (isLLMProviderConfigured) also ignored the toggle (review
P0-02). Excluded disabled entries from the fallback order (extracted as
buildUsableFallbackOrder; image/video keep adopting server-configured
entries whose enabled:false means not adopted yet) and taught
isLLMProviderConfigured to reject disabled providers, so an empty
selection is retained and submission stays blocked.

* fix(token-plan): assign shared provider credentials to the top enabled plan

Applying a plan overwrote shared provider slots (seedream) last-writer-
wins, and disabling the owning plan left the slot enabled with the
disabled plan's endpoint and key (review P0-03). Shared slots now always
belong to the highest-priority enabled plan targeting them: apply yields
credentials and seed catalogues to a higher-priority owner regardless of
connect order, and disable/enable/remove transitions hand the slot back
(restoreSharedProviderCredentials rewrites endpoint, key — taken from the
owner's LLM slot — and catalogue of the new owner).

* fix(token-plan): include the re-enabled plan in shared ownership and clean up on removal

Two transitions still mis-owned shared slots after the previous fix
(re-review of P0-03): re-enabling a higher-priority plan excluded itself
from ownership resolution, so the slot stayed with the lower-priority
plan; and removal skipped cleanup when another plan was merely enrolled
but disabled, leaving the removed plan's endpoint, key and model
catalogue active on the shared provider.

- restoreSharedProviderCredentials takes excludeConcerned: disabling and
  removal exclude the concerned plan, re-enabling includes it, so the
  newly enabled plan reclaims its slot and a lower-priority enable does
  not steal one it does not own
- removeTokenPlan skips shared-slot cleanup only for another USABLE plan;
  otherwise the slot is cleared and disabled instead of surviving behind
  an enrolled-but-disabled plan

* fix(chat): honor per-stage user routes for classroom chat

The classroom-interaction model override is sent by the client as the
`x-model-routes` header, but the two multi-agent classroom chat routes
resolved their `chat-adapter` stage from the request body without parsing
it, so a per-stage override was silently ignored and the main model was
used instead.

Forward the parsed user routes into `resolveModel` for both /api/chat and
/api/chat/pi (credentials still come from the body, not x-* headers),
and attach the header from the client fetches. Operator MODEL_ROUTES still
wins server-side.

Regression tests cover the user-route precedence, the operator override,
the x-model-routes parser, and both client fetches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(settings): drop the pro-mode override the server never honors

The Course Model Config UI offered a "Pro Mode" station that wrote a
user-level route for `maic-agent-driver`, but that stage is resolved only
from the operator's MODEL_ROUTES (it also needs an explicit api dialect
and contextWindow), so the override could never take effect.

Remove the station from the map and its pipeline card, stop seeding the
route from the TokenDance preset, and drop the now-unused label/i18n keys
in every locale. The settings store's rehydrate sanitizer also drops any
previously persisted `maic-agent-driver` user route, so upgraded installs
do not keep a dead entry. The "no orphan runtime stage" contract test now
expresses `maic-agent-driver` as an explicit operator-only exception.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(settings): neutralize token-plan comments

Rewrite comments that referenced an external deployment or implementation
into neutral descriptions of this file's own behavior and design intent.
No behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Percy <percy@PercydeMacBook-Pro.local>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-23 18:11:21 +08:00
Frank ZhuandFrank-zhu0404 56322a5e06 fix(generation): reject unusable interactive scripts and surface runtime errors (#1649)
Reject classic inline interactive scripts that fail to parse at generation time (extracted with parse5, checked with node:vm Script without executing), and surface iframe runtime errors on the active interactive scene.

Addresses #1622 (partial: no recovery / regeneration UI).

Co-authored-by: Frank-zhu0404 <Frank-zhu0404@users.noreply.github.com>
@openmaic/generation@0.3.13
2026-09-23 13:50:18 +08:00
Frank Zhuandwyuc 4b537dca96 fix(tts): pace classroom TTS and classify MiniMax RPM limits (#1650)
* fix(tts): pace classroom TTS and classify MiniMax RPM limits

MiniMax reports RPM limits as HTTP 200 with base_resp.status_code 1002.
Classify that code as TTSRateLimitError, pace classroom TTS with
TTS_MIN_INTERVAL_MS (default 1000ms) and double the gap up to 15000ms on
each rate limit, and store ttsCoverage on the generation job so a partial
narration is not a clean success.

Fixes #1567

* fix(tts): fail closed on MiniMax status and skipped narration

Treat every non-zero MiniMax base_resp status as a failure, even when
audio bytes are present. Classify 1002, 1039, 1041, and 2045 as rate
limits. Floor rate-limit back-off independently of TTS_MIN_INTERVAL_MS,
cap the cumulative wait, and cancel those sleeps on abort. A skipped or
thrown TTS phase reports zero coverage and a warning.

* fix(tts): decay spacing, heartbeat TTS progress, default interval 0

Classroom TTS spacing steps back toward TTS_MIN_INTERVAL_MS after each
successful clip, and that interval defaults to 0 so pacing starts only
after a rate limit. Clip progress is reported through the existing job
onProgress path, so updatedAt keeps moving during a long narration phase.
A Retry-After that does not fit the remaining back-off budget skips that
clip instead of silencing every later clip.

generateTTSForClassroom always returns coverage. A skipped run that has
resolved a provider counts clips after the same speech split as synthesis.

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 13:26:04 +08:00
杨慎andwyuc d04a75cbc6 feat(ai): add Xiaomi MiMo V2.6 models [AI-assisted] (#1655)
* feat(ai): add Xiaomi MiMo V2.6 models

* docs(ai): preserve Xiaomi provider references

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 13:11:49 +08:00
LING DUANandwyuc 2e4386e09b feat(pi): expose interactive static instructions through read_scene (#1632)
* feat(pi): expose interactive static instructions

* fix(pi): preserve scene evidence when static text exceeds budget

* fix(pi): contain static parser failures and fence scene evidence

* fix(pi): align static evidence fallback boundaries

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 12:57:36 +08:00
Frank Zhuandwyuc fbc51fcbd6 fix(generation): reject invalid quiz option shapes and constrain values at the source (Closes #1375) (#1651)
* fix(generation): enforce quiz option value/label contract

When a model puts choice content in value and a bare A-Z letter in label,
swap those fields so persisted quiz options keep a letter as the selection
identity and the content as the label. Correct objects, plain strings, and
answer-key normalization stay as they were.

Closes #1375

* fix(generation): rewrite swapped quiz letter keys to option values

When a bare label is moved onto the option value, a raw answer token that
exactly equals that original label is rewritten to the uppercased letter
before exact alignment. The stored key is then the letter QuizView submits,
so that submission grades correct. A stored "a" on an already-correct
value "A" is left unresolved.

Closes #1375

* fix(generation): reject invalid quiz option shapes instead of swapping

Stop rewriting quiz options when a model puts content in value and a
letter in label. The quiz prompt now requires value to be one ASCII
letter A-Z and label to be the option content. A choice question that
still breaks that contract, or whose answer does not name an option
value, is invalid model output and regenerates.

Closes #1375

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
@openmaic/generation@0.3.12
2026-09-23 12:46:40 +08:00
d3e882241c fix(web-search): stop leaking Brave's HTML challenge page into errors (#1553)
* fix(web-search): stop leaking Brave's HTML challenge page into errors

Both Brave error paths echoed the upstream response body verbatim, so a
throttled or blocked request surfaced a full HTML page as Error.message
and a transient 429 was indistinguishable from a real outage.

Route both paths through formatBraveError, which gives 429 a dedicated
actionable message, drops markup and over-long bodies in favour of the
HTTP status text, and keeps short plain-text details.

Closes #1552

* fix(web-search): cancel Brave's HTML error body instead of reading it

Deciding whether an error body is worth surfacing still meant pulling the
whole challenge page into memory first. Check the content type up front and
cancel the stream when it is markup, so the large body is never read.

* fix(web-search): give Brave API-key 429 plan-specific guidance

The API-key path already sends a subscription token, so its 429 message
should not suggest configuring a Brave API key. Point it at the plan's
rate limit instead, keep the API-key suggestion for the keyless scrape
path, and add a test for the API-key 429.

---------

Co-authored-by: zhentang2 <zhentang2@iflytek.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 12:17:15 +08:00
LING DUANandwyuc 80367c0d76 feat(render): add admission and per-task resource budgets (#1492)
* feat(render): add admission and per-task resource budgets

* fix(ci): isolate patched producer test configuration

* fix(render): bind portable resource builds and installed validation

* fix(render): preserve deadline errors and service exit ownership

* style(render): format verified lifecycle changes

* fix(render): retain cleanup evidence and enforce task pid limits

* refactor(render): narrow resource patch and acceptance documentation

Revert cleanup adapters in two unreferenced Producer helpers. Remove historical runtime claims from the patched README, retain the current Linux NOT_RUN qualification, and map remaining B4 evidence to concrete owner and platform boundaries.

Validation: fixed-source patch application and all 44 source hashes pass; 43 retained patch sections and both dependency locks are unchanged. Source/package tests: 17 passed. Cleanup/context/supervisor tests: 30 passed. No native build, VM, installed Linux execution, or export was run.

* test(render): cover remaining B4 lifecycle acceptance paths

Extend the existing installed runner with guardian unlink failure, overlapping task reference settlement, and successful new-supervisor service after explicit platform takeover. Keep original per-task and outer limits; only the overlap case reserves two task budgets through the existing Producer API.

Local validation: syntax, formatting, lint, two evidence regressions, and fixed-source patch/hash verification pass. Installed Linux cases remain NOT_RUN: the first SSH preflight timed out before authentication or any remote command ran. No runtime implementation or dependency version changed.

* docs(render): keep runtime qualification tied to candidate evidence

Keep the validation document as a reproducible contract rather than a stale NOT_RUN snapshot. Current native and Producer/B4 evidence passed for 6be78252; HTTP remains unexecuted after an adapter input-preflight failure. Runtime status and remaining work are tracked in the PR.

* fix(render): harden resource startup and simplify owner transport

* docs(render): clarify post-commit cleanup settlement

* style(render): format outer resource fence files

* fix(render-service): harden resource-mode closeout

* fix(render): bind resource cleanup to trusted directory identity

Use nonrecursive stage cleanup and preserve quarantine when project identity cannot be verified. Bind mount, publication and cleanup to verified directory descriptors, contain recovery bookkeeping errors, and cover helper cancellation and closeout gates.

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-23 11:59:30 +08:00
Yizuki_Ame 7de474f4b2 fix(workbench): show effective upload limits in friendly errors (#1418)
* fix(workbench): show effective upload limit in friendly errors

* fix(workbench): format the upload limit shown in composer toasts

A non-MiB cap was stringified as a long fraction, and nothing checked that the toast used the localized message.
2026-09-23 11:43:08 +08:00
AbelWangYaBoandwyuc 05d782090f feat(persistence): resolve single-tenant requests to a shared owner id (#1639)
Course documents, folders, materials and agent sessions are partitioned by an
owner id that comes from a 30-day anonymous cookie, because there is no host
auth layer. For one team behind one ACCESS_CODE that partition buys nothing —
those visitors already share the site password — and it costs: a second browser
shows an empty course list, and POST /api/stages/[id]/publish refuses every
owner, because publishing a cookie partition is not something the product
allows.

PERSISTENCE_SHARED_OWNER_ID resolves every request to that fixed id instead.
Unset — the default — changes nothing.

- **ACCESS_CODE is required.** Without one the middleware lets every request
  through, so a single shared owner would expose one readable, editable and
  publishable course library to anyone who can reach the deployment. That is
  worse than the per-browser partitioning this setting removes, so the
  combination refuses to start, from instrumentation.ts, with a message naming
  both variables. The alternative the issue allowed — ignore it with a warning —
  was rejected because a silently ignored setting is exactly the confusion this
  feature exists to remove.
- The value is validated rather than trusted: 1-128 characters of
  [A-Za-z0-9._-]. That excludes the reserved `anon:` prefix, which would still
  be refused by publish and would alias onto a cookie owner, and any character
  the material-key sanitiser rewrites, which could collide two ids onto one
  object key. A malformed value fails startup on the same path; an empty value
  is treated as unset, since KEY= in an env file cannot be told apart from an
  operator who meant to leave the feature off.
- An explicit authenticatedOwnerId still outranks it, so adding a real auth
  layer later does not require clearing the variable first.
- All four places that derive an owner honour it, including the Server Action
  in lib/workbench/workspace-actions.ts, which re-reads the cookie itself. That
  one is easy to miss and would make the workspace list and its row actions
  disagree about who owns a session.
- The env docs state the trust model next to the variable, cross-reference
  SECURITY.md, and note that this answers who owns a course while DATABASE_URL
  answers where one lives — a second browser sees an empty list on Postgres too.

This is the stopgap asked for in #1550 — a PR in the spirit of the closed
#1551 — with threading a real authenticated owner left as the long-term fix.

AI-assisted commit

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 17:33:02 +08:00
1e867e8448 fix(skills): keep exported zip entry names POSIX-separated on Windows (#1641)
`buildSkillDirZip` names each zip entry with `path.relative(dir, file)`
directly. On Windows that separator is a backslash, so an exported
skill archive contains entries like `openmaic/references\extend.md`
instead of `openmaic/references/extend.md`.

Zip entry names are POSIX paths by spec, and consumers — the skill
import path, `unzip`, and the round-trip test — resolve entries by
forward-slash paths, so an archive exported on Windows cannot be read
back as a skill package.

Normalize the relative path to forward slashes before naming the
entry. On POSIX platforms `sep` is already `/`, so this is a no-op
there.

Co-authored-by: BeAIcoder <beaicode@example.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 17:21:00 +08:00
Maulana Zidaneandwyuc 8262055423 refactor(export): extract shared speech-text walk into lib/export/narration.ts (#1621)
Closes #1142

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 17:05:06 +08:00
Frank Zhuandwyuc 56208ba073 feat(ai): add gemini-3.8-flash and gemini-3.7-flash to the Google catalog (#1629)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 15:53:21 +08:00
Yizuki_Ameandwyuc 283a3096aa fix(editor): share line bounds for alignment and dragging (#1631)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
@openmaic/renderer@0.1.11 @openmaic/editor@0.0.9
2026-09-22 15:39:56 +08:00
LING DUANandwyuc 7be21f2d89 fix(chat): pin Pi routing and preserve provider errors (#1637)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 15:18:54 +08:00
AbelWangYaBoandwyuc 7e8a065ecb fix(importer): bound pptx zip inflation by default (#1588)
ZipParseLimits was optional field by field and both import paths called
parseZip with no argument, so every bound the parser already implemented was
inert: a .pptx could inflate without limit.

- apply the defaults per field, so a caller overriding one bound no longer
  silently drops the others, and Number.POSITIVE_INFINITY switches a single
  bound off
- add maxCompressionRatio for text parts, defaulting to 200:1. It applies to
  slides, layouts, themes and chart XML only. Uncompressed bitmaps and silent
  PCM are ordinary deck content sitting at the top of DEFLATE's range — a
  solid-colour 24-bit BMP measures ~1027:1, 30s of silent stereo PCM ~1016:1,
  and a 1920x1080 white screenshot with a grid ~297:1 — so a fixed ratio there
  rejects real decks. Media and embedded objects are bounded by
  maxEntryUncompressedBytes and maxMediaBytes instead, which for an honest
  archive are checked before anything is inflated. Text parts have no such
  problem: real decks peak around 19:1, so 200 leaves ample headroom.
- reject a supplied NaN and a non-integer maxEntries rather than letting them
  compare false against every bound and disable it silently
- export the defaults and the limit type so callers can extend them
- add the first tests for this parser: the regression itself (a call with no
  limits must still apply the bounds), the per-field defaulting invariant, each
  bound pinned, the binary-part exemption pinned with a solid-colour BMP, a
  silent WAV and an embedded object, and the repo's own regression deck

The bounds read the sizes the archive declares about itself, which an archive
is free to understate. One that declares small and inflates large passes all of
them and is stopped only by JSZip's own size check, after that entry has been
inflated — peak RSS in the reviewed case was ~1.19 GB, and with maxConcurrency 8
several such entries inflate in parallel. Bounding the allocation itself needs a
byte-budgeted inflate inside the read path, which is a larger change than
activating the bounds that already existed here.

Version to 0.3.0: ^0.2.6 admits 0.2.7, and this changes default behaviour for
existing callers. Consumers pinned to ^0.2.x need to move deliberately.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
@openmaic/importer@0.3.0
2026-09-22 14:24:58 +08:00
AbelWangYaBoandwyuc 7e81e44c36 fix(media): refuse redirects on adapter generation and poll calls (#1636)
#930 made the connectivity probes in the media adapters pass
`redirect: 'manual'`. The generation and poll calls in the same 14 files were
left following redirects. Those requests carry the provider credential and go
to a base URL that comes from provider settings a caller can supply, so a 3xx
would replay the credential at a host the caller chose — and the redirect
target can be an address the outbound guard already refused.

Every such call now passes `redirect: 'manual'` and rejects a 3xx through a
shared `assertNotRedirected` helper, which reports it as
"<provider>: Redirects are not allowed (HTTP <status>)" instead of letting the
generic failure path describe it as a provider error.

- 26 call sites across the image adapters (seedream, openai, qwen, grok,
  lemonade, minimax, nano-banana), the video adapters (seedance, kling, grok,
  happyhorse, minimax, veo) and ComfyUI's submit and image fetch.
- ComfyUI's `pollHistory` keeps its contract of handing the caller a retryable
  failure rather than aborting the generation: it logs the refusal and returns
  null.
- ComfyUI's same-origin workflow load is deliberately untouched — it reads the
  app's own public/ asset, carries no credential and is not provider-influenced.
- The two adapters added since #930 (OpenRouter image and video) already did
  this.

tests/media/adapter-redirects.test.ts covers one case per adapter family. Each
serves a 302 and asserts that the call rejects with the redirect message and
that every request carrying an init object asked fetch not to follow redirects;
each case fails if its adapter stops passing `redirect: 'manual'`.

The HappyHorse test asserted the exact request init, so it now includes the new
option.

AI-assisted commit

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 13:00:02 +08:00
b0e481eaa5 fix: keep long Grok relay generations alive and inline image bytes (#1364)
Two independent Grok failures seen when the provider is reached through a relay
(a custom base URL) instead of api.x.ai directly:

- lib/ai/providers.ts: a long non-streaming chat completion was cut off by the
  relay with a 504 at its idle timeout (~5 min), because nothing is sent
  upstream until the model has the whole answer. Adding 'grok' to the existing
  streaming-compat path (OPENAI_COMPAT_USE_STREAMING_CHAT=true) keeps bytes
  flowing across the idle window; the SSE is buffered back into a normal JSON
  response for the caller.
  The path is for relays only. usesCustomOpenAIBaseUrl recognises OpenAI's
  origin alone, so Grok's own api.x.ai also read as "custom" and was forced
  onto the compat transport; the provider's native endpoint is now excluded.

- lib/media/adapters/grok-image-adapter.ts: response_format 'url' returns a
  link on the relay's CDN host (imgen.x.ai), which may be unreachable from the
  server's network. The generation then failed at the follow-up fetch through
  /api/proxy-media even though the image had been produced successfully.
  'b64_json' inlines the bytes and removes that second hop.
  Inline bytes declare no media type, so the adapter reports one on
  ImageGenerationResult and returns a typed data URL, which is the shape
  openrouter-image-adapter already uses. Consumers take the type from there:
  agent image persistence records it, the client's stored media row keeps the
  type its data URL states, and classroom-media-generation names the file with
  the matching extension. A JPEG is no longer stored, served or named as a PNG.

The process-wide undici timeout that previously accompanied these changes is
dropped: upstream #1404 now gives LLM calls their own undici headers/body
timeouts, which covers the same failure without raising the defaults globally.

Co-authored-by: ciclou1 <ciclou1@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-22 12:04:20 +08:00
LING DUAN 2a77a8a476 feat(chat): make Pi classroom runtime the default (#1628)
* feat(chat): make Pi classroom runtime the default

* chore(chat): align docs and E2E with Pi default
2026-09-22 11:41:13 +08:00
Yizuki_Ameandwyuc df16d7e322 fix(export): include active line geometry in bounds (#1626)
Share corrected line bounds across renderer and React editing paths, preserve double-elbow routing, and translate PPTX points into the shape bounding box.

Refs #674 and the prior implementation/review in #675. AI-assisted.

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
@openmaic/renderer@0.1.10 @openmaic/editor@0.0.8
2026-09-21 22:22:35 +08:00
6dfeb62dd8 fix(media): emit keyframe images as data URLs [AI-assisted] (#1444)
* fix(media): emit keyframe images as data URLs

The local media extractor stored keyframe bytes as raw base64 in
`DocumentAsset.data`. Every other extractor emits a data URL, and the
document bundle forwards this field as `pdfImages[].src` to `storeImages`,
which decodes it with `decodeBase64DataUrl`. With no `data:` prefix the
comma split yields no payload, so `atob(undefined)` throws and course
generation fails with "Failed to store image bundle at image img_1".

`pdf-compat`'s `dataUrlMimeType` also derives the image asset mime from
this same field, so the declared `image/webp` was silently dropped too.

Emitting `data:${mime};base64,...` matches `mineru-parser`, which already
normalizes prefix-less base64 the same way.

* fix(media): decode keyframe data URLs

* fix(media): accept legacy and data-url assets

* fix(media): reject malformed or empty media asset data URLs

A string that starts with data: but does not parse as a data URL used to
fall through to the raw-base64 path, where Node's decoder skips
non-alphabet characters and yields garbage bytes. Throw instead, and
reject an empty base64 payload rather than storing a 0-byte image.

Correct the keyframe test comment: local keyframe assets are consumed by
material extraction, and the test needs ffmpeg so it does not run in CI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DYifP8wM4XJQ3Hc2qsF6zf

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 21:48:02 +08:00
Yizuki_Ameandwyuc 54d626bc61 fix(export): surface video render rejection reasons (#1414)
* fix(export): surface video render rejection reasons

* fix(export): keep render diagnostics out of user toasts

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-21 18:49:25 +08:00
LeoParkerOuandwyuc ec20ba345f refactor(classroom): share per-course session lifecycle (#1619)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 21:14:20 +08:00
LeoParkerOuandwyuc 44882254a4 fix(generation): stabilize model picker teardown (#1618)
* fix(generation): stabilize model picker teardown

* test(generation): cover model picker teardown

---------

Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 20:53:04 +08:00
2aa7e3e12b fix(audio): play discussion lines through one reused media element (#1610)
Discussion narration created a new Audio element per line, so on mobile every
line after the first had its programmatic play() refused with NotAllowedError:
the dialogue went silent while the lesson kept advancing. This is the
discussion side of #1474, which #1477 fixed for the narration player.

One module-scoped element now serves every line, kept separate from the
narration element in AudioPlayer so neither can cut the other off. Because
element identity can no longer tell lines apart, the stale-event guard became a
per-line token, handlers are assigned rather than added (a listener would
accumulate once per line), and finish()/cleanup() release the line by removing
the source attribute and reloading instead of leaving the element pointed at
the page's own URL.

Co-authored-by: talkman <jaxgen@163.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 19:23:00 +08:00
Huang Geyangandwyuc 67f568848a fix(server): warn when access-code protection is disabled (#1599)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
2026-09-20 19:14:38 +08:00