mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-04 18:29:01 +08:00
@openmaic/editor@0.0.9
381
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7be21f2d89 |
fix(chat): pin Pi routing and preserve provider errors (#1637)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
7e81e44c36 |
fix(media): refuse redirects on adapter generation and poll calls (#1636)
#930 made the connectivity probes in the media adapters pass `redirect: 'manual'`. The generation and poll calls in the same 14 files were left following redirects. Those requests carry the provider credential and go to a base URL that comes from provider settings a caller can supply, so a 3xx would replay the credential at a host the caller chose — and the redirect target can be an address the outbound guard already refused. Every such call now passes `redirect: 'manual'` and rejects a 3xx through a shared `assertNotRedirected` helper, which reports it as "<provider>: Redirects are not allowed (HTTP <status>)" instead of letting the generic failure path describe it as a provider error. - 26 call sites across the image adapters (seedream, openai, qwen, grok, lemonade, minimax, nano-banana), the video adapters (seedance, kling, grok, happyhorse, minimax, veo) and ComfyUI's submit and image fetch. - ComfyUI's `pollHistory` keeps its contract of handing the caller a retryable failure rather than aborting the generation: it logs the refusal and returns null. - ComfyUI's same-origin workflow load is deliberately untouched — it reads the app's own public/ asset, carries no credential and is not provider-influenced. - The two adapters added since #930 (OpenRouter image and video) already did this. tests/media/adapter-redirects.test.ts covers one case per adapter family. Each serves a 302 and asserts that the call rejects with the redirect message and that every request carrying an init object asked fetch not to follow redirects; each case fails if its adapter stops passing `redirect: 'manual'`. The HappyHorse test asserted the exact request init, so it now includes the new option. AI-assisted commit Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
b0e481eaa5 |
fix: keep long Grok relay generations alive and inline image bytes (#1364)
Two independent Grok failures seen when the provider is reached through a relay (a custom base URL) instead of api.x.ai directly: - lib/ai/providers.ts: a long non-streaming chat completion was cut off by the relay with a 504 at its idle timeout (~5 min), because nothing is sent upstream until the model has the whole answer. Adding 'grok' to the existing streaming-compat path (OPENAI_COMPAT_USE_STREAMING_CHAT=true) keeps bytes flowing across the idle window; the SSE is buffered back into a normal JSON response for the caller. The path is for relays only. usesCustomOpenAIBaseUrl recognises OpenAI's origin alone, so Grok's own api.x.ai also read as "custom" and was forced onto the compat transport; the provider's native endpoint is now excluded. - lib/media/adapters/grok-image-adapter.ts: response_format 'url' returns a link on the relay's CDN host (imgen.x.ai), which may be unreachable from the server's network. The generation then failed at the follow-up fetch through /api/proxy-media even though the image had been produced successfully. 'b64_json' inlines the bytes and removes that second hop. Inline bytes declare no media type, so the adapter reports one on ImageGenerationResult and returns a typed data URL, which is the shape openrouter-image-adapter already uses. Consumers take the type from there: agent image persistence records it, the client's stored media row keeps the type its data URL states, and classroom-media-generation names the file with the matching extension. A JPEG is no longer stored, served or named as a PNG. The process-wide undici timeout that previously accompanied these changes is dropped: upstream #1404 now gives LLM calls their own undici headers/body timeouts, which covers the same failure without raising the defaults globally. Co-authored-by: ciclou1 <ciclou1@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
2a77a8a476 |
feat(chat): make Pi classroom runtime the default (#1628)
* feat(chat): make Pi classroom runtime the default * chore(chat): align docs and E2E with Pi default |
||
|
|
df16d7e322 |
fix(export): include active line geometry in bounds (#1626)
Share corrected line bounds across renderer and React editing paths, preserve double-elbow routing, and translate PPTX points into the shape bounding box. Refs #674 and the prior implementation/review in #675. AI-assisted. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
6dfeb62dd8 |
fix(media): emit keyframe images as data URLs [AI-assisted] (#1444)
* fix(media): emit keyframe images as data URLs
The local media extractor stored keyframe bytes as raw base64 in
`DocumentAsset.data`. Every other extractor emits a data URL, and the
document bundle forwards this field as `pdfImages[].src` to `storeImages`,
which decodes it with `decodeBase64DataUrl`. With no `data:` prefix the
comma split yields no payload, so `atob(undefined)` throws and course
generation fails with "Failed to store image bundle at image img_1".
`pdf-compat`'s `dataUrlMimeType` also derives the image asset mime from
this same field, so the declared `image/webp` was silently dropped too.
Emitting `data:${mime};base64,...` matches `mineru-parser`, which already
normalizes prefix-less base64 the same way.
* fix(media): decode keyframe data URLs
* fix(media): accept legacy and data-url assets
* fix(media): reject malformed or empty media asset data URLs
A string that starts with data: but does not parse as a data URL used to
fall through to the raw-base64 path, where Node's decoder skips
non-alphabet characters and yields garbage bytes. Throw instead, and
reject an empty base64 payload rather than storing a 0-byte image.
Correct the keyframe test comment: local keyframe assets are consumed by
material extraction, and the test needs ffmpeg so it does not run in CI.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DYifP8wM4XJQ3Hc2qsF6zf
---------
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
54d626bc61 |
fix(export): surface video render rejection reasons (#1414)
* fix(export): surface video render rejection reasons * fix(export): keep render diagnostics out of user toasts --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
ec20ba345f |
refactor(classroom): share per-course session lifecycle (#1619)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
2aa7e3e12b |
fix(audio): play discussion lines through one reused media element (#1610)
Discussion narration created a new Audio element per line, so on mobile every line after the first had its programmatic play() refused with NotAllowedError: the dialogue went silent while the lesson kept advancing. This is the discussion side of #1474, which #1477 fixed for the narration player. One module-scoped element now serves every line, kept separate from the narration element in AudioPlayer so neither can cut the other off. Because element identity can no longer tell lines apart, the stale-event guard became a per-line token, handlers are assigned rather than added (a listener would accumulate once per line), and finish()/cleanup() release the line by removing the source attribute and reloading instead of leaving the element pointed at the page's own URL. Co-authored-by: talkman <jaxgen@163.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
67f568848a |
fix(server): warn when access-code protection is disabled (#1599)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
9cd8051461 |
docs: align security and behaviour claims with shipped code (#1592)
ACCESS_CODE unset remains fail-open in middleware, document reads are capability-by-id via the anonymous owner cookie (not x-learner-key), and the README action/skill counts match the Action union and skills/agent-runtime. Closes #1587 Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
d7b31aa5e7 |
fix(generation): add browser-safe package entry (#1609)
Co-authored-by: SY <sy@SYdeMacBook-Pro.local> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
ed18042fde |
fix(vocational): send taskEngineMode on the on-demand scene path (#716)
The vocational gate keys off requirements.taskEngineMode in the /api/generate/scene-content body, but only the first scene (generation-preview) sent it. Scenes 2..N and retries run through useSceneGenerator.generateRemaining / fetchSceneContent, which never sent requirements, so resolveVocationalActive returned false and applyOutlineFallbacks rewrote every later procedural-skill scene to diagram — silently dropping the task-engine training mechanism for most of a vocational course (the outline says procedural-skill while the content path strips it). Thread the persisted stage.taskEngineMode through GenerationParams into both fetchSceneContent bodies (the retry path inherits it via lastParamsRef), mirroring what the first-scene request already sends. Server contract is unchanged; the flag is still ANDed with the OPENMAIC_ENABLE_VOCATIONAL env gate server-side. Closes #715 Co-authored-by: ly-wang19 <ly-wang19@users.noreply.github.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
19f3eff6ac |
fix(tts): derive Azure SSML locale from selected voice (#1566)
* fix(tts): derive Azure SSML locale from voice * fix(tts): preserve Azure voice script locales --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
6a8db813bb |
refactor(upload): derive the workbench material MIME policy from the shared format registry (#1590)
* refactor(upload): derive the workbench material MIME policy from the shared format registry #1498 fixed the Kylin generic-Office-MIME failure on both upload paths but left the workbench with its own extension→MIME table, alias map, and generic-MIME set next to the document registry, and the two had already drifted. The workbench policy now derives every MIME/extension fact from lib/document/mime.ts (the single source of truth); only the accepted-format list stays workbench policy — fixed and extractor-independent, unlike the classic path's provider-scoped whitelist. - Register csv and webm in DOCUMENT_FORMATS (accepted by no document provider, so classic-mode whitelists are unchanged) and add the audio/x-m4a alias to m4a. - material-upload-policy.ts keeps its export names (route, session-store, and composer consumers unchanged) but resolves, whitelists, and builds its accept string from registry helpers. - The workbench gate now also accepts the registry's curated aliases it previously missed: image/jpg, text/x-markdown, and audio/x-wav (stored canonically as audio/wav). Closes #1589. Co-Authored-By: Claude Code <noreply@anthropic.com> * test(workbench): pin the audio/mp3 alias closure in the material policy Deriving the workbench gate from the shared registry normalization (#1589) also accepts the browser-reported audio/mp3 alias the hand-rolled alias map rejected — the same gap class as image/jpg and text/x-markdown, so pin it alongside them and name it in the comment. Co-Authored-By: Claude Code <noreply@anthropic.com> --------- Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
f29bbc4daa |
feat: sample declared interactive state before classroom questions (#1508)
* feat: sample declared interactive state before classroom questions * test: exercise generated publication example through the iframe reader * chore(generation): bump package version for observation prompt contract * fix: keep interactive state optional on insecure HTTP origins * fix(playback): keep component picking and separate state from reference A declared state interface replaced the component picker with a forced whole-area `#experiment` reference, removing the per-component selection and outline that `main` already ships. Sampling was also gated on the reference selector, so the only way to obtain state was to give up the selection. Reference identity and area state are now independent request-scoped evidence items: - `handleToggleElementPick` always arms the picker again, so a scene that declares the interface keeps main's per-component selection, outline, and send-time clearing. - `sampleInteractiveState` follows the current Scene instead of the draft reference, so an unreferenced follow-up still reports current facts and never re-creates or extends a reference. - The Host carries area state with or without a component reference. The evidence header names both identities and refuses to present area facts as properties of the referenced component. - `metadata` is absent when only area state travels, so no element identity and no Spotlight authorization can be derived from it, and the accepted-reference receipt stays driven by explicit references only. Review follow-ups in the same change: - Client sampling follows `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED`. An ungated packet turned an ordinary Pi question into a 400 while the reference feature was disabled. - A Scene that declares the interface always receives an availability boundary, including when the browser produced no packet at all. It is reported as `not-sampled` rather than the previous `no-interface`, which was a false statement about an activity that does declare one. Courseware without the interface keeps its unreferenced behaviour. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(generation): register the observation snippet as a packaged asset The interactive-observation snippet is referenced by all six widget content templates but was never added to the packaged-asset manifest, so the asset test and the golden scene prompt both failed. - `SNIPPET_IDS` now lists `interactive-observation`, restoring both the "exactly the generation-owned templates and referenced snippets" check and the "every referenced snippet is packaged" cross-check. - The interactive system-prompt snapshot is re-pinned. The change is purely additive: the snippet is appended to the simulation template. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): stop injecting state constraints while the feature is disabled The route rejects a request that carries a reference or a state packet while `NEXT_PUBLIC_COURSEWARE_REFERENCE_ENABLED` is off, but an ordinary question carries neither. It still reached the Host, and a Scene that declares the state interface then received the full page-state block — several kilobytes of constraints about evidence the deployment can never sample. The Host now returns before building that note when the feature is off. A route-level regression asserts that neither the Director prompt nor the Child prompt gains `PAGE-REPORTED STATE` in that configuration; disabling the guard makes it fail with exactly that symptom. Also reopens the composer before the unreferenced follow-up in the classroom browser spec. An accepted answer may close it, which made the assertion flaky without changing the behaviour under test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): decouple reference Scene from state freshness and bound assembled evidence Cross-review found two defects in the request-scoped state evidence. The reference's Scene was folded into the sample's staleness test. The packet is already bound to the current Scene by the identity check above it, so a valid current-Scene sample was being discarded as `stale-sample` purely because the student's component reference came from an earlier Scene. Reference and area state are independent evidence items; freshness is a property of the sample alone. With the coupling gone the two can now disagree on Scene, so the note says so explicitly rather than letting the model attribute area facts to a component that may not be on the current Scene. The assembled evidence had no stated output budget. The static component packet is bounded to 24,000 code points upstream, but that bound covers the static packet alone; the note and the escaped observation JSON were appended without a recheck. Escaping `<` for the prompt expands one code point into six, and `<` is legal in a label or a fact value, so a packet the Host accepts could assemble to 149,385 code points. The budget is now declared as the static bound plus the room the note frame needs, which is what makes the degradation terminate. Over budget, the state body drops whole to an explicit `unavailable` statement: truncating the JSON would emit a broken packet, and thinning a `complete` relation set would turn an exhaustive set into a false one. The relationship summary degrades with it, so the prose never asserts COMPLETE over a body that is gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(pi): keep slide references independent of activity state * refactor(interactive): simplify declared state and unify iframe preparation Accept any JSON report within byte and depth budgets, without generated field requirements or relationship-completeness semantics. Keep publishState and an optional rendered result in the generation guidance. Prepare the observation responder through patchHtmlForIframe and let the pool own document identity, preserving state across placeholder remounts. Settle sampling failures locally and align browser/server nesting limits. Cover permissive JSON delivery, resource limits, lifecycle, legacy behavior, and real renderer remounts with focused regression tests. * fix(generation): publish automatic activity changes with clear positions * fix(interactive): report missing legacy scope as no interface * fix(interactive): guard sampling capabilities and bind scopes lazily * chore(generation): bump version after main integration --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d51f4b8350 |
fix(audio): rebuild platform FormData bodies for the undici transport (#1580)
#1514 moved every lib/audio provider call to the npm undici package's fetch so the pinned dispatcher is guaranteed to be honored. But the adapters kept building multipart bodies with the platform-global FormData — a class of Node's bundled undici — and undici's serializer brand-checks a FormData body against its own class. A foreign FormData fell through to the string branch and left the process as `content-type: text/plain;charset=UTF-8` with the 17-byte literal `[object FormData]` as the whole body: the audio bytes and every field (including `model`) were dropped, downstream gateways fell back to whisper-1 and answered 503, and every multipart audio request (ASR, TTS FormData paths, voice registration/cloning) failed with a generic internal error. Bare Blob/File, string/JSON/Buffer and stream bodies were never affected: undici 7.29.0 exports no File/Blob classes of its own and its bare-body brand checks (and multipart part handling) bind the platform classes. Normalize the body in the transport, once, before either transport path serializes it, so the adapters can keep the platform globals as their public API boundary: a foreign FormData is re-created as undici's own with every entry carried over verbatim — `append` (not `set`) so repeated field names survive, and no filename argument so a platform File part keeps its own name/type/lastModified. Everything else passes through untouched. Also drive real loopback regression tests with a platform-global FormData (direct, with repeated field names, empty, and across a 307 redirect hop whose per-hop loop re-issues the normalized body) and a bare Blob, asserting multipart on the wire instead of `[object FormData]`, and update the voxcpm unit test that asserted the buggy contract (a platform FormData at the undici boundary). Fixes #1579 Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
97cf12eccd |
fix(playback): queue widget messages until iframe ready (#1532)
* fix(playback): queue widget messages until iframe ready * fix(playback): preserve iframe queue in StrictMode --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
f1b34e3abb |
fix(audio): reuse one narration element so mobile playback survives the first segment (#1477)
On mobile browsers classroom narration plays the first segment and then goes silent for every following one, while the lesson keeps advancing: the player created a new HTMLAudioElement per line, and only the first line is covered by the user's gesture, so every programmatic play() after it is refused with NotAllowedError and the engine falls back to its reading-time timer. #651/#652 fixed the blob leak from that rejection, not the missing voice. Keep one element per player instead: - getAudioElement() creates it on first use and every line reuses it, so the element the first gesture activated stays playable for the rest of the lesson - stopAudioElement() releases the line's state rather than the element: onended cleared, src removed, load() called -- a stopped line must not keep reporting speech it is no longer playing, nor retain narration bytes through a revoked object URL until the next play() - onended is assigned rather than added: the element now outlives a single line, so a listener would accumulate once per segment and call the engine back several times for one line tests/audio/audio-player-element-reuse.test.ts stubs the mobile policy itself (the first element plays, every element created after it is refused): the reuse tests fail on main and pass with this change. Verified on a live deployment at the fixed entry point (instrumented Chromium, default autoplay policy, one real click on Play): a single element for 8+ consecutive lines, no refused play, currentTime advancing line by line. Fixes #1474 Co-authored-by: Shaoxuhua <jaxgen@163.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
a8726ec4f3 |
feat(media): add OpenRouter image and video providers (#1356)
* feat(media): add OpenRouter image and video providers
OpenMAIC ships six separate video providers (Veo, Kling, Seedance,
MiniMax, Grok, HappyHorse) and seven image providers, each needing its
own key. OpenRouter fronts those same model families behind one key and
one account, so this adds it as a provider on both sides.
Both use OpenRouter's dedicated media endpoints, not chat-completions:
- Image: POST /images -> { data: [{ b64_json }] }
- Video: POST /videos -> 202 { id, status }, poll GET /videos/{id},
then GET /videos/{id}/content for the mp4 bytes
The model list is fetched live from GET /images/models and
GET /videos/models through /api/openrouter-models rather than pinned in
the registry: OpenRouter hosts 48 image and 28 video models today and
adds more, so a hardcoded shortlist would decide for the operator which
models exist. The registry keeps a three-entry seed as an offline
fallback, and the existing custom-model UI still accepts any model id.
Both catalogs answer unauthenticated, so the picker fills before a key
is pasted; a key is forwarded when present for proxied base URLs.
Adapter contracts are covered by stubbed-fetch tests (request shape,
empty-response handling, and the video job state machine including
terminal failure). No test performs a billable call.
Closes #1355
* fix(media): validate the key and tolerate a pasted endpoint URL
Three fixes found while configuring the new provider:
1. Both connectivity probes hit the model catalogs, which answer 200
unauthenticated — so "Test Connection" reported success for any
string, including an invalid key. Probe GET /key instead: equally
cheap, and it actually rejects a bad key.
2. The settings field is labelled "Base URL" but the panel echoes it
back as "Request URL", so pasting the full endpoint
(https://openrouter.ai/api/v1/images) is the natural mistake. That
built /api/v1/images/images and 404'd. Trim a trailing slash and a
trailing /images or /videos so both forms work; a proxy path that
merely contains the word is left alone.
3. The image and video settings panels read `data.message` on a failed
test, but failures answer with `error` (apiError) and only successes
carry `message`. Every failing connectivity test — for any provider,
not just OpenRouter — rendered "connection failed: undefined" instead
of the reason. Pre-existing; surfaced by 1 and 2 above.
Closes #1355
* fix(media): make every OpenRouter model selectable, and always select a provider
Two gaps found while configuring the new provider.
The settings Models list is a read-only catalog for every provider; the
actual model picker is the media popover. That picker built its groups
from the static registry array, so OpenRouter offered only the
three-entry seed while settings listed the full live catalog — the
models were visible but not choosable. Feed the same live catalog into
the popover, fetched only once the provider is usable so an
unconfigured install makes no request.
Separately, `imageProviderId`/`videoProviderId` are empty until a
provider is chosen (first-run auto-config leaves them blank when the
server reports no media provider). Opening the settings panel on an
empty id selected nothing: the header rendered the missing name key as
"settings.undefined", and Test Connection posted a blank
x-image-provider/x-video-provider, so it failed with "No image/video
provider configured" whatever key was typed. Fall back to the first
catalog entry so the panel always has a selection. Pre-existing and not
specific to OpenRouter.
Closes #1355
* fix(tts): request a browser-playable format from custom providers
`generateOpenAITTS` serves every custom OpenAI-compatible TTS provider but
never sent `response_format`, so it inherited whatever each provider
defaults to. OpenAI defaults to mp3; OpenRouter's /audio/speech defaults
to raw `pcm`. The unknown content type then fell through to the `'mp3'`
default below, the client built `data:audio/mp3;base64,…` from headerless
PCM samples, and playback failed with "no supported source was found" —
while the server logged a clean 200, because the audio really was
generated. Name the format instead of inheriting it.
Also stop mislabelling an unrecognised body: `pcm`/`l16` now raises a
message naming the cause, and `aac`/`opus` are recognised.
Two supporting fixes:
- /api/openrouter-models normalises its base URL the way the adapters do
and falls back to the public catalog when a custom base URL fails, so a
typo in a free-text settings field cannot empty the model picker. Also
types the headers object so tsc accepts the conditional.
- provider-neutrality-guard pins exact per-vendor occurrence counts in
lib/server/provider-config.ts. Adding the image and video env entries
raises "openrouter" from 2 to 6 (each entry contributes both its key and
its value); CI failed without the bump.
Closes #1355
* fix(security): never send the operator key to a client-chosen host
Review found `/api/openrouter-models` was an SSRF and key-exfiltration
path, and the finding is correct. The route took `x-base-url` from the
caller at highest precedence while preferring the *server* env key, so any
caller could make the server send the operator's OpenRouter credential as
an `Authorization: Bearer` header to an arbitrary URL. The route's own
comment claimed it followed `/api/verify-image-provider`; that pattern
runs `validateUrlForSSRF` on client base URLs, and this route did not.
The boundary is now explicit: the server key travels only to the
operator's own base URL. A client-supplied URL is SSRF-validated and
carries only that caller's own `x-api-key` — the server key is dropped —
and the unauthenticated public-catalog fallback never forwards a
credential chosen for a different host. Redirects are no longer followed
(`redirect: 'manual'`), since a redirect would carry the Authorization
header off-host and reopen the same hole, and upstream reads are bounded
by a timeout.
The per-URL cache is now keyed by destination *and* a hash of the
credential, and bounded to 64 entries with oldest-first eviction, so
client-supplied URLs cannot grow it without limit and one caller's
key-authorised catalog is never served to another.
Also from the review:
- The image adapter discarded the reported `media_type`. The
orchestration layer wraps a bare `base64` as `data:image/png`
unconditionally, so jpeg/webp results were mislabelled; the adapter now
returns a data URL carrying the real type.
- Adapter generation and poll requests set `redirect: 'manual'`, matching
the `/key` probe that already did.
- `runPolledTask` accepts an `AbortSignal` so the sleep between polls is
cancellable; the video adapter passes the caller's signal. Without it a
cancelled generation still slept out a full 10s interval.
Tests cover the highest-risk paths the review named: which credential
reaches which URL, that an SSRF-rejected destination is never contacted,
that the fallback is unauthenticated, cache isolation between callers,
and MIME preservation.
Findings 2 (base-URL normalisation) and 4 (neutrality-guard debt) were
already fixed in
|
||
|
|
83d692a6ce |
fix(import): persist ZIP media in the server asset pool (#1520)
Upload imported media before publishing document references and retain browser-only behavior. Add cache-failure, upload-failure and independent HTTP reader regressions; document the export/import migration path. Validation: 256 focused tests passed; tsc, full Prettier check and ESLint passed (18 existing warnings in unchanged files). Full browser playback and production build remain unverified. Co-authored-by: qianbs <ninesheng99@163.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
73ea6a72ce |
feat: auto-detect Vietnamese for browser-native TTS narration (#1487)
Browser TTS auto-detection only distinguished Chinese (CJK ratio) from everything else (en-US), so Vietnamese narration was spoken by an English voice. Add Vietnamese detection to the shared language helper and bind an installed vi voice in the playback engine when the user has not picked one explicitly: - lib/audio/browser-tts-preview.ts: new detectSpeechLang() (zh-CN / vi-VN / en-US), used by both the Test TTS preview and playback. A hit on đ/ơ/ư or the U+1EA0-U+1EF9 precomposed block (ớ, ừ, ồ, ế, …) — absent from French/Romanian Latin — marks Vietnamese; bare ă/â/ê/ô only count toward a low ratio. - lib/playback/engine.ts: use the shared helper; when it reports vi-VN, prefer an installed vi voice so pronunciation is correct out of the box. An explicitly configured voice still wins. - tests/audio/detect-speech-lang.test.ts: zh/vi/en/fr cases. Verified end-to-end (Playwright WebKit): utterances carry lang vi-VN with an installed Vietnamese voice auto-selected, advancing through every line. Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
784f2a9f95 |
fix: bound model discovery response body timeout (#1478)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
f7b8769e7f |
fix(importer, renderer): preserve PPTX text, chart, and image styling (#1534)
* fix: preserve PPTX text, chart, and image styling * fix(importer): preserve hyperlink semantics and remove inferred alignment * test(charts): replace untyped assertions to satisfy CI * fix(importer): honor source-linked chart axis number formats * fix(renderer): preserve picture bar clustering and clipping * fix(pptx): address chart scales and soft-edge review findings * fix(pptx): preserve percent-stack labels and inherited text sizes * fix(pptx): synchronize agent chart schema and axis defaults |
||
|
|
2cbd011c1f |
feat(token-plan): add TokenDance one-key preset for every modality (#1525)
* feat(token-plan): add TokenDance one-key preset for every modality TokenDance is a model gateway: chat and images are OpenAI-compatible at /gateway/v1, and the same key authenticates vendor-protocol routes on the same host (Ark, MiniMax, Bocha). The preset reuses the existing adapters with those route prefixes as base URLs, so one key lights up LLM, image, video, TTS and web search from Settings -> Token Plan. - providers: add a built-in `tokendance` OpenAI-compatible provider (TOKENDANCE_* env prefix, logo, provider name in all locales) - token-plan: add the TokenDance preset (Seedream image, MiniMax H3 video, MiniMax speech TTS, Bocha web search) - seedream: use a base URL that already ends in a version segment verbatim, so gateway routes like `/ark/v3` do not get `/api/v3` appended - minimax-video: route H3-family models through the v2 task API (content array submit, task-envelope poll); connectivity checks for H3 probe auth on the v2 query route instead of submitting a billable task - README: add a one-key quick example and replace the Gemini-specific model recommendation with a provider-agnostic setup recommendation Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL * fix(token-plan): accept preset web-search base URLs and report H3 dimensions per ratio - web-search: the client base URL allowlist also accepts the exact base URL a built-in token plan preset writes for that provider, derived from TOKEN_PLAN_PRESETS. Applying a plan whose web-search route is not an official vendor host previously stored a URL that the route rejected with 400. Any other client URL is still rejected. - minimax-video: report H3 v2 clip dimensions for 16:9, 9:16, 4:3 and 1:1 instead of assuming landscape for every non-portrait ratio. - tests: pin the allowlist for every preset, the 1:1 H3 dimensions, and clear TOKENDANCE_* in the provider-config env isolation list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qrrq9CPwb718mpouz8Y2KL --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b5605ef979 |
feat(agent-runtime): workbench generate_image / generate_video write through the asset pool (#1007 part 6) (#1524)
The workbench tools store generated bytes in the asset pool under the shared principal and write the allocated id into the document, the same discipline as the classic chain since #1392; the runner's putScene creates the reference rows and commits the allocations. The video completion patch rewrites every placeholder slot and retires anything that would shadow the new id; immediate render is preserved by leasing the id at the render boundary. A store-full refusal fails the tool with a model-readable error and writes nothing. Legacy /api/classroom-media documents keep rendering. Closes #1522. |
||
|
|
4a8219af4e |
fix(audio): resolve CDN-backed narration consistently (#1521)
* fix(audio): resolve CDN-backed narration consistently (#1515) * fix(audio): keep export fallback resolution consistent --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
1d3f62d80b |
refactor(media): one client-side pool commit primitive; keep refused narration instead of re-billing it (#1523)
Extracts commitToPool, the single client-side sequence for storing bytes in the asset pool, writing the allocated id back, and mirroring locally; routes the media pass, narration adoption and fresh TTS through it. A store-full refusal during TTS now retains the already-billed clip so the next load adopts it with zero provider calls. Closes #1467. |
||
|
|
87c4524b4d |
fix(audio): validate redirects and pin connections on provider requests (#1514)
Audio provider requests (TTS, ASR, voice registration and voice cloning) validated a client-supplied base URL once and then issued a plain fetch with default redirect-follow and no pinned dispatcher. A base URL that resolved to a public address but answered with a redirect to an internal one was followed, and a DNS answer that changed between the guard's lookup and the connect reached an internal host — both readable in-band. - Route every lib/audio provider request through a new lib/server/audio-provider-fetch.ts that combines per-hop redirect re-validation with a pinned undici dispatcher, so the socket can only reach an address the guard validated, on every hop. - Select the public-vs-local policy server-side from isServerConfiguredProvider; a client-supplied base URL is always strict public and can never reach a private, loopback or cloud-metadata address, even with ALLOW_LOCAL_NETWORKS set. - Pin the result-audio download hop as well, keeping its host allowlist and redirect:'error'. - Add a coverage-matrix test that fails if any lib/audio module regains a raw provider fetch. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ee79762ef9 |
fix(access-code): expire tokens server-side and throttle verification behind a trusted proxy (#1513)
When ACCESS_CODE is set, verification tokens (timestamp.HMAC) never expired because neither verifier checked the timestamp, and POST /api/access-code/verify had no attempt throttling. - Enforce a 7-day token lifetime in a shared, Edge-safe module used by both the Node verifier and the middleware Web-Crypto verifier, and reject non-canonical signatures in both. - Rate-limit verification only when the client identity is trusted (TRUST_PROXY_HEADERS=true): a per-client sliding window (10 failures / 60s) whose attempt is reserved atomically at check time, returning 429 with Retry-After. Without a trusted proxy the app cannot attribute requests to a client, so no shared throttle is applied — a long random ACCESS_CODE is the protection, and a warning is logged when it is short. - Store bounded, copied identity keys so a large forwarding header cannot retain memory. - Document the behavior in .env.example, README, and configuration docs. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0b2e502811 |
fix(classroom): keep transient load failures off the not-found card (#1485)
* fix(classroom): keep transient load failures off the not-found card /api/classroom 5xx and network errors used to become null, so ClassroomSurface treated them like a missing course and rendered the terminal not-found claim with no retry. Classify fetch answers as found/absent/unavailable and only show not-found on a positive 404/410 miss. Fixes #1450 * fix(classroom): handle terminal load states --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
1271cba69e |
fix(tts): treat Add-dialog default URL as a configured credential path (#1482)
addCustomTTSProvider stores the dialog Base URL in customDefaultBaseUrl and leaves baseUrl empty. isTTSProviderConfigured ignored that field, so generation silently skipped narration even after Test TTS succeeded. Accept the dialog URL so already-saved custom providers start working without retyping the field. Fixes #1471 Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
c605e9e6ef |
feat(persistence): turn on the server-owned asset lifecycle and release assets on course deletion (#1007 amendment, part 2) (#1473)
App wiring for the @openmaic/storage 0.31.0 lifecycle: reference tracking and document references are paired unconditionally and declared at startup, course deletion withdraws references inside the tombstone transaction, ASSET_PENDING_TTL_MS configures the pending window, and the dead client-side reclamation code is removed. |
||
|
|
ee7a7b64df |
fix(agent-runtime): allow one owner material to be bound to multiple sessions (#1500)
* fix(agent-runtime): allow one owner material to be bound to multiple sessions `agent_session_materials.id` is a global primary key, but `bindOwnerMaterialsToSession` de-duplicated on `(sessionId, id)` and then inserted the owner-side material id as the row id. Binding the same owner material to a second session therefore raised a duplicate-key error (23505) that the route surfaced as HTTP 500, so an owner material could only ever be used by one course. Keep row ids globally unique (every extraction method keys on `id` alone) and store the shared owner id as metadata instead: the binder now mints a fresh session row id, records `owner_material_id`, and looks that column up for idempotent rebinding; a partial unique index on `(session_id, owner_material_id)` adjudicates concurrent rebinds and the loser adopts the winner's row. The schema change is additive and idempotent (`ADD COLUMN IF NOT EXISTS` + `CREATE UNIQUE INDEX IF NOT EXISTS`), so existing databases upgrade in place with `NULL` for legacy rows. No FK from `owner_material_id` to the owner library, to keep owner-library deletion decoupled from session rows. Tests: backend-neutral contract test (same owner material bound to two sessions, both readable), PGlite in-place upgrade test, pinned-schema update, and a host-level test covering the route's 202 path. * fix(agent-runtime): reuse legacy owner-material bindings and clean up losing uploads on concurrent bind Rows written by the previous binder use the owner upload id as the session row id and leave `owner_material_id` NULL, so the new owner-id lookup missed them and a rebind created a duplicate row and copied the bytes again. The binder now falls back to the session row id and adopts it only when it is unambiguously that owner upload: source kind, matching title, no source URL or derivative, no text, and the deterministic legacy object key with a byte length that matches the owner record. Adoption stamps `owner_material_id` through a conditional `backfillOwnerMaterialId` update that leaves extraction columns untouched; a lost race re-reads the fast path. A different session is still a fresh row. The concurrent-bind loser also left its uploaded object behind after adopting the winner. It now removes that object when the winner references a different key, best-effort so a failed cleanup cannot fail a bind that succeeded. Tests: host-level PGlite regressions for the legacy rebind (id, row count, byte copies, extraction state, and backfill) plus a different-session bind, and a controllable byte store that parks the loser after its upload so the winner commits first, asserting one row and only the winner's object remain. A storage unit test pins the backfill contract. Bumps @openmaic/storage to 0.30.1 for the new public store method. |
||
|
|
cfae106b89 |
fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates (#1499)
* fix(storage): keep jsonb writes valid when model output contains NUL or lone surrogates PostgreSQL jsonb rejects the `\u0000` and lone UTF-16 surrogate escape sequences that JSON.stringify emits verbatim, failing the enclosing statement with SQLSTATE 22P05/22P02. In the agent-session store the failed tree-entry/event write is treated as critical, so the whole run aborts and all work in that run is lost. Add a shared encodeJson at @openmaic/storage/src/pg-json.ts that replaces U+0000 and unpaired surrogates with U+FFFD while preserving valid surrogate pairs (emoji), sanitizing in-memory values and object keys before JSON.stringify. Route every jsonb parameter through it: the agent-session, document, runtime, asset, and material backends, plus owner-material registration in lib/persistence. The document/runtime/asset backends keep their existing assertJsonValue guard, which rejects these code points with a readable error before the shared serializer runs; the agent-session and material paths had no guard and were the live 22P05 exposure. Literal backslash-u text is untouched. * chore(storage): bump @openmaic/storage to 0.29.2 * fix(storage): keep colliding sanitized keys and own __proto__ members when encoding jsonb encodeJson replaces NUL and lone surrogates in object keys, but two distinct keys can sanitize to the same string: "a\u0000" and "a\uFFFD" both become "a\uFFFD". The rebuild assigned each member by its sanitized key, so a later member silently overwrote an earlier one and that earlier value was lost. Keep the first member under the sanitized key and give each later colliding member a deterministic suffix (#2, #3, ... appended until the key is unique among the keys emitted so far, whether sanitized or original). Member order is preserved. The suffix is ASCII and contains no NUL or surrogate, and the sanitizer only ever emits U+FFFD, so a suffixed key can never be confused with a bare replacement result. The rebuild also used a plain {} object, so assigning an own "__proto__" member invoked the prototype setter and dropped it from the output whenever a sibling key needed sanitizing. Build the rebuilt object with a null prototype so "__proto__" stays an own data property; JSON.stringify then emits it and PostgreSQL round-trips it. Arrays and the no-op fast path (return the original reference when nothing needs sanitizing) are unchanged. |
||
|
|
16fd8583a6 |
fix(settings): maintain asr language state invariant on provider selection (#1082) (#1443)
* fix(settings): maintain asr language state invariant on provider selection (#1082) * test(store): import ASRProviderId type in server sync tests * fix(settings): validate custom ASR languages against CUSTOM_ASR_DEFAULT_LANGUAGES (#1082) * test(store): refactor custom ASR sync and rehydration test suite for clarity --------- Co-authored-by: cham <2577781125@qq.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
90cf53f142 |
fix(upload): resolve the generic Office MIME (application/vnd.ms-office) via filename extension (#1498)
* fix(upload): resolve the generic Office MIME via filename extension (#1497) On Linux, Chrome derives File.type from the XDG shared-mime-info database. Older or trimmed databases (e.g. Kylin OS V10) map every OOXML extension to the generic container `application/vnd.ms-office` instead of the concrete format MIME, so .pptx/.docx/.xlsx uploads were rejected with "当前解析器不支持该格式" on the generation toolbar and with 415 unsupported_mime on POST /api/materials — while the same files upload fine on macOS/Windows. Treat `application/vnd.ms-office` as a generic upload fallback alongside the zip family, so the filename extension decides the concrete format: - lib/document/mime.ts: add it to GENERIC_DOCUMENT_MIME_TYPES. This covers the toolbar's provider-whitelist check and the extract-document API gate (both call normalizeDocumentMimeType). It must NOT go into per-format aliasMimes — canonicalFromAlias iterates DOCUMENT_FORMATS in declaration order, so a shared alias would misroute x.pptx to application/msword. - lib/workbench/material-upload-policy.ts: new resolveWorkbenchMaterialMime() grants missing/octet-stream/zip/vnd.ms-office types the same extension fallback the document path already had (the workbench gate previously rejected empty and zip-family types outright on every OS). - app/api/materials/route.ts + session-store.ts: gate, store, and send the resolved MIME instead of the raw browser value. The spoofing surface is unchanged: a specific but unsupported MIME still falls through verbatim for the whitelist to reject, and an unknown extension keeps the generic MIME. Full suite: 7992 passed. Closes #1497 Co-Authored-By: Claude Code <noreply@anthropic.com> * fix(review): log the declared mime, harden the fallback chain, pin adversarial cases Review follow-ups on #1498 (no correctness or security findings — hygiene only): - materials route: the reject/context log now carries declaredMime alongside the resolved mime whenever they differ, so a generic-header spoof stays visible in upload failure logs instead of collapsing to the resolved type. - uploadWorkbenchMaterial: the returned record's mimeType falls back to the locally resolved mime before the raw browser value, which may still be the generic Office container if the server echo is ever absent. - material-upload-policy: export WORKBENCH_MATERIAL_EXTENSIONS and add a drift guard asserting every accepted extension resolves (via a generic MIME) to a whitelisted type — a hand-maintained map entry going missing now fails tests instead of silently 415ing. - tests: pin the legacy-.ppt provider split (mineru rejects, mineru-cloud accepts), uppercase-extension resolution, and 415-before-400 precedence for a generic mime with no filename. Full suite: 7996 passed. Co-Authored-By: Claude Code <noreply@anthropic.com> * fix(upload): support WPS Office MIME aliases --------- Co-authored-by: Claude Code <noreply@anthropic.com> Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
dee421621f |
feat(skills): add 习题课(最近发展区) (#1382)
* feat(skills): add ZPD-based exercise lesson skill * fix(skills): localize exercise lesson title in all workbench locales --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
44c3c8aed0 |
fix: tolerate malformed authored CSS in classroom exports (#1422)
* fix: tolerate malformed authored CSS in classroom exports
Interactive-scene HTML authored by LLMs can carry browser-tolerated CSS
syntax errors (e.g. `-- crop-green: #22c55e`), which browsers silently
drop. The strict postcss pass used by resource-pack / classroom-zip
export threw on them, failing the whole export while PPTX (which never
parses interactive HTML) kept working.
Inlining is an enhancement, not a precondition: on strict-parse failure
the stylesheet/style attribute/SVG attribute is now kept byte-for-byte
with a warning, and export continues. Remote dependencies inside
unparseable CSS (url() / @import, absolute or relative, quoted with
parens) are still surfaced via a comment- and string-aware fallback
scanner so zip exporters warn and offline validation is not silently
bypassed.
* fix: stop the css fallback scanner at bad-string tokens
Per CSS syntax, an unescaped \n/\r/\f inside a quoted string produces a
bad-string token and the browser recovers right there. The fallback
scanner previously ran unterminated strings to EOF, swallowing every
following rule and hiding browser-active url()/@import dependencies from
dependency reporting and residual validation.
Both string-scanning loops now terminate at an unescaped newline and
resume scanning from it; an escaped newline (\ + \n, \ + \r\n, \ + \r,
\ + \f) continues the string, and escaped newlines are stripped from
captured URLs — in the fallback scanner and in the strict-path
collectors (cssUrlReferences / cssImportReference), which previously
kept the escape in the URL for valid CSS too.
* fix: keep EOF-truncated css strings as valid tokens
Per CSS Syntax 4.3.5, only an unescaped \n/\r/\f produces a bad-string
token; a quoted string truncated at EOF is a valid string token that
browsers keep (with the function auto-closed and a trailing backslash
consumed). The previous round wrongly dropped EOF-truncated url(".../
@import "... values, hiding real remote dependencies again — the exact
shape of an LLM-authored <style> cut off mid-generation.
Both string paths now share scanCssString with three-way termination
(quote / unescaped newline / EOF), and the collectors were aligned with
CSS token semantics on both the strict and fallback paths, with parity
tests: escaped newlines (\ + LF/CRLF/CR/FF) are stripped escape-aware
from quoted values; an unquoted url() containing a backslash-newline is
a bad-url and no dependency; non-whitespace (comments allowed) around a
quoted payload makes a bad-url, and a bad url consumes through its ')'
so nested url(...) text is not re-scanned; an escaped ')' belongs to an
unquoted url value.
---------
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
|
||
|
|
3e3cac5f76 |
fix: enable Pro workbench flag in Docker builds; allow non-TLS localhost cookie (#1484)
* fix: enable Pro workbench flag in Docker builds; allow non-TLS localhost cookie Two deployment fixes discovered while self-hosting with Docker Compose: 1. Expose NEXT_PUBLIC_PRO_WORKBENCH_ENABLED as a Docker build arg. Persistence and other NEXT_PUBLIC_* flags are already wired through Dockerfile + docker-compose.yml, but the Pro workbench entry flag was missing, so Docker deployments could not enable the workbench at all. 2. Add COOKIE_SECURE opt-out for the anonymous owner cookie. Production builds always set `Secure` on the anonymous_id cookie. Safari refuses to store Secure cookies served over plain http://localhost (it does not special-case localhost like Chromium/Firefox), so every request minted a fresh anonymous owner and owner-scoped document writes were rejected with 403 — scenes never persisted and generation appeared to hang after the first scene. COOKIE_SECURE=0 lets plain-HTTP deployments opt out; the default (Secure in production) is unchanged. Documents COOKIE_SECURE in .env.example alongside the persistence opt-ins. * fix: share the COOKIE_SECURE opt-out with the Server Action cookie mint Review follow-up on the COOKIE_SECURE opt-out. - Extract anonymousCookieSecure() in owner.ts and use it in lib/workbench/workspace-actions.ts, which re-implements the anonymous_id mint: with COOKIE_SECURE=0 that path still sent a Secure cookie, so a Server Action running before any /api/agent/* request (e.g. deleting a workspace session) minted an ephemeral owner in Safari. - State the opt-out accurately in owner.ts; the previous comment claimed the flag never forces Secure off. - Add the regression test beside the existing production case (verified it fails when the opt-out is dropped). - .env.example: document the security cost of dropping Secure and that only the exact value 0 disables it. --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
68aed2932a |
fix(proxy-media): connect only to addresses the SSRF guard validated (#1488)
* fix(proxy-media): connect only to addresses the SSRF guard validated The route validated a URL with validateUrlForSSRF and then handed the hostname to fetch(), which resolved it again at connect time. An attacker-controlled name could answer a public address to the guard and an internal address (loopback, RFC1918, link-local, cloud metadata) to the socket, and every redirect hop repeated the same validate-then-refetch gap. Add a shared policy-aware pinned dispatcher (lib/server/pinned-dispatcher.ts) whose connect.lookup resolves all addresses, runs the same local-network policy as validateUrlForSSRF over every answer (cloud metadata is always blocked, even with ALLOW_LOCAL_NETWORKS=true), and passes only the vetted list to the connection. proxy-media creates one dispatcher per request and uses it on every hop, mapping a connect-time refusal to 403 INVALID_URL instead of a generic 500. The agent-runtime pinned agent now reuses the same builder with its strict assertSafeIp policy unchanged. The other routes that validate with validateUrlForSSRF and then hand the URL to provider SDKs (parse-pdf, transcription, generate/*, verify-*, probe-models, azure-voices) are not covered; that is a follow-up. Tests: new tests/server/proxy-media-rebinding.test.ts covers the two-answer rebind, the ALLOW_LOCAL_NETWORKS opt-in through the pinned path, metadata refusal under the opt-in, and a redirect-hop rebind. * fix(proxy-media): destroy the per-request dispatcher and align non-unicast policy across layers Follow-up hardening for the proxy-media SSRF pin. 1. Handler hang: the route awaited `dispatcher.close()` in `finally`. undici's `close()` drains in-flight requests, and the response body is intentionally left unread on the 30x (`redirect: 'manual'`), non-2xx, oversize and connect-refusal paths, so an upstream that keeps its body open never completes and the POST never settles. Tear the pool down with a non-awaited `destroy()` instead; the response blob is already in memory by the time `finally` runs. Regression tests drive a loopback server that answers 404, a 30x, and an oversize 200 while trickling its body forever and assert the route settles in well under a second. 2. Headline rebinding test sensitivity: install an attacker-controlled global dispatcher (`setGlobalDispatcher` plus a lookup that always answers loopback) in `beforeEach` so the unpinned fetch path really reaches the loopback server. Removing the route's dispatcher now fails on `internal.requests() === 0` rather than on an unrelated DNS resolution error. The positive control also asserts the pinned connect lookup was invoked, so a runtime that silently ignores `init.dispatcher` goes red. 3. Non-unicast policy parity: `validateUrlForSSRF` only classified IP-literal URLs through `isPrivateIP`, so a literal in CGNAT 100.64/10 or an IANA reserved/special-use range (240/4, 198.18/15, the TEST-NET blocks, 192.0.0.0/24, multicast, broadcast) passed the URL layer while the same address reached through a hostname was refused at connect time. Both layers now share `isNeverAllowedRange`, applied to IP-literal URLs and to every resolved answer, with or without `ALLOW_LOCAL_NETWORKS`. `connectionAddressBlockReason` drops its blunt `range() !== 'unicast'` test so the two layers agree exactly, while private, loopback and link-local answers keep following the opt-in. Behaviour change: a hostname resolving into CGNAT 100.64/10 or an IANA reserved/special-use range is now refused (403 INVALID_URL) where it was previously proxied, including under `ALLOW_LOCAL_NETWORKS=true`, since these ranges are never legitimate proxy targets. 4. Do not echo an unparseable connect-time answer in the refusal. Return the generic block message instead of `Unable to classify network address: <value>`; the route relays that message to the client in a 403 body. Tests: new trickle-body regression tests and CGNAT/reserved cases for both literals and resolved answers, plus a `connectionAddressBlockReason` parity suite. One existing fixture is updated from the IANA documentation prefix 2001:db8::/32 (reserved) to a real public prefix 2001:4860::/32. * fix(ssrf-guard): let the local-network opt-in cover CGNAT and decode local-use NAT64 prefixes |
||
|
|
1fd4348f21 |
fix(classroom): generate classroom ids server-side and create files exclusively (#1489)
* fix(classroom): generate classroom ids server-side and create files exclusively
The legacy file-backed classroom store accepted a caller-chosen stage.id and
renamed a temp file over the target, so anyone who knew a public share URL
could POST the same id and replace that classroom's content. POST
/api/classroom now mints the id itself (nanoid, 10 characters, matching the
generation pipeline), ignores any client-supplied stage.id, and rebinds the
stage and every scene to the generated id. Both the route and
classroom-generation persist through a new exclusive create that hard-links a
temp file into place, so an existing classroom is never replaced; EEXIST is
retried with a fresh id a bounded number of times and then reported as 409.
Unchanged: the server-backed persistence path (app/api/persistence,
lib/persistence), the read/GET path, middleware, and the access-code gate;
writeJsonFileAtomic keeps its overwrite semantics for the other callers.
Adds regression tests pinning non-overwrite, the typed EEXIST error, collision
retry/exhaustion, and generation-path exclusivity.
* fix(classroom): fall back to exclusive open without hard links and reserve ids before media generation
writeJsonFileExclusive kept `fs.link` as its only route to the destination, so
classroom creation failed hard on mounts without hard links (gcsfuse, s3fs and
some FUSE gateways return ENOSYS/ENOTSUP/EOPNOTSUPP/EPERM/EXDEV). Keep `link`
as the fast path and, for exactly those codes, fall back to an exclusive `wx`
open: still never replaces an existing file and still maps EEXIST to
ClassroomAlreadyExistsError. Any other code keeps propagating, and the temp
file is always removed.
The generation pipeline wrote media and TTS into
<CLASSROOMS_DIR>/<id>/{media,audio} before the exclusive create, so a collision
would have landed the new classroom's files in an existing classroom's
directory and the retried document's URLs would still point at that other id.
Reserve the id first by exclusively creating the classroom file with a
placeholder document (`reserved: true`, empty scenes). The file is the only
token that atomically covers the whole collision namespace: a classroom created
through POST /api/classroom has a JSON file but no directory, so reserving a
directory would miss it. readClassroom treats a reserved document as absent, so
an in-flight or crashed reservation is never served as an empty classroom. The
retry on EEXIST now happens at reservation time (bounded, before any media),
and the final persist is an ordinary overwrite of the id the process owns, so
the post-media id retry is gone.
Document OPENMAIC_CLASSROOMS_DIR in .env.example, including that
CLASSROOM_JOBS_DIR does not move with it.
* fix(classroom): release an unused reservation when generation fails
|
||
|
|
ebf665f316 |
feat(playback): reference static GenUI components (#1281)
* feat(playback): reference static GenUI components * feat(playback): gate courseware references * fix(pi): bound Interactive source HTML before parsing * fix(pi): bound interactive reference HTML processing * fix(playback): close interactive reference review gaps --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
aad034786e |
fix(export): reuse compiled ZIP when retrying video renders (#1456)
* fix(export): reuse compiled ZIP when retrying video renders * fix(export): validate course names before reusing cached ZIPs |
||
|
|
4685cad341 |
fix(tts): prepare discussion segments during playback (#1435)
* fix(tts): prepare discussion segments during playback * fix(tts): preserve active playback when muting lookahead --------- Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
c460bceb60 |
feat(export): preflight render queue availability (#1455)
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn> |
||
|
|
85ef52d599 |
feat(media): write generated media through the asset pool under server-backed persistence (#1392)
* feat(media): store generated media in the asset pool when persistence is server-backed With server-backed persistence the document is durable and shared, but generated media stayed in the producing browser: the document kept its gen_img_* / gen_vid_* placeholder and narration kept a browser-derived audio id. Every new browser that opened such a course re-ran generation for every slide, and it never converged, because the address of the generated bytes was never written back into the document. Under server-backed persistence only, the classic generation chain now stores bytes in the asset pool first and writes the id the pool allocated into the document. - The client bootstrap configures the asset seam alongside the document and runtime seams: an HttpAssetStore over the persistence endpoint carrying the same credentials the document store carries, marked server-backed. The seam preflight now covers all three, so a failure still cannot half-configure persistence. - Image, video and TTS generation commit in one fixed order: provider, pool, document, local cache, task. A reference reaches the document only after put returned an id, so a document can never name bytes that were not stored. A failure before the write-back leaves the placeholder with the provider called exactly once; the retry happens on the next owner load. - The write-back is a per-slot rewrite through mutateDocument, which re-reads the current document under the per-stage lock, so it cannot clobber a newer scene. The open course is refreshed with the same rewrite without being marked dirty. - "Has this already been generated?" is answered by the document (the slide exists and no longer holds the placeholder) instead of by this browser's task table. - The classroom's resume effect fails closed on ownership: only a resolved owner starts generation, so a viewer opening a shared course spends nothing. - The local media and audio tables become a per-tab cache. A failed cache write costs a re-download, never the media. Browser-only mode is unchanged: every new call site sits behind the server-backed gate, the local tables stay authoritative there, and placeholders stay in the document. Rendering and export needed no changes. HttpAssetStore.resolve mints an object URL exactly as the browser store does, and the export byte resolver was already pool-first. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(classroom): make the generation owner gate a three-outcome rule and apply it everywhere The gate refused everything but a resolved owner, which read a sidecar that answered "no ownership fact exists for this course" as a reason to block. That is the answer a deployment without the sidecar's server-side prerequisites gives for every course, and the answer a course with no ownership record gives: in both, there is nobody the operator's budget needs protecting from, and refusing strands the course's own author behind a question that can never be answered. Ownership is now four states over the sidecar's three outcomes. A definite answer splits into owner and not-owner. An absent record is its own answer, ownerless, and generation proceeds — the behaviour such a deployment had before the gate existed. Only the absence of an answer, a transport failure or a load that has not asked yet, stays unresolved and fails closed: "we could not ask" must never be read as "nobody owns this". One mapper turns a sidecar result into that state, and one predicate decides on it. The workbench classroom pane runs the same resume effect and had no ownership input at all, so a viewer opening a shared course there could still spend the budget. It now asks the sidecar once per course, in parallel with its load and feeding only the generation gate, so its read-only and edit behaviour is unchanged. The shared progressive-load policy carries the gate for it, with both new inputs required rather than defaulted so a future caller cannot omit them into an open budget. Its stale comment claiming ownership could not be expressed here is corrected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the write-back survive autosave, arrive before the scene does, and never leak Independent reviews of the write-back found three ways a durable document could still end up naming a placeholder, and two ways the gate that protects the operator's budget could be walked around. An autosave round captures the store synchronously and writes that capture, so a round already in flight when a rewrite landed wrote the placeholder straight back over the allocated id, and nothing marked the store dirty again to correct it. The rewrite now marks the units it changed, which leaves a corrective flush queued behind the stale one; re-saving a scene that already holds the id is idempotent, losing the id is not. Media is generated from outlines in parallel with scene content and usually finishes first, so the slide that will carry the placeholder does not exist yet and the write-back has nothing to rewrite. That was the ordinary path, not a tail case, and its result was discarded: the task was marked done, the scene was added afterwards with its placeholder intact, and a second pass in the same run could call the provider again. The allocation is now held under the placeholder — which also answers the skip test, so nothing pays twice — and applied when that scene is committed, before its first save. One complete pass now leaves no placeholder behind. A failed commit used to abandon what it had already allocated. A poster upload that failed threw away a stored video and sent the retry to submit the most expensive job in the system again; a rejected write-back left registry rows that name bytes nothing references, which the byte collector cannot reclaim because it only collects blobs no row names. A poster failure now costs the poster, and a write-back that reached nothing reclaims what it allocated. A partial write is left alone, because the document already names it. The ownership gate is fail-closed again. Treating the sidecar's 404 as permission was wrong: the client cannot tell "this course has no owner" from "this deployment told me nothing", so a visitor who opened a shared course could bill the operator. The root cause was the sidecar itself, which gated on the agent runtime although every persisted course has an owner regardless — the persistence route resolves one for every request. It now gates on server persistence, so the configuration that made 404 the universal answer has real ownership facts to report, and the gate can refuse everything but a named owner. Retry affordances answered to no gate at all. A viewer of a shared course with one failed image was shown a Retry button that called the provider. Both retry entry points and every surface that draws them now read one shared permission, so what is offered and what is allowed are the same value. Also: narration regeneration no longer pretends it can replace bytes behind a live id — the exclusivity proof that would allow it is refused by construction once references leave the browser, so it forks to a fresh id and says so; the "already generated" test lets a finished deck answer from the document alone, since scene order stops identifying an outline once slides are inserted or deleted; stored assets record a specific media type rather than a generic transfer type; the pane no longer asks the sidecar in browser-only mode; and the funnel's docstring now states what the per-stage lock actually guarantees, which is same-browser serialization and not a cross-browser compare-and-swap. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): park allocations in the deciding turn, never reclaim on an ambiguous write A delta review of the write-back found the first-pass fix still had a window, and the reclamation it added could delete media the document already names. The allocation was parked after an awaited local cache write. A scene committed in that window reconciled against a registry that did not hold it yet, so the document kept the placeholder — and the entry recorded a moment later then answered the skip test as "already handled", so nothing could correct it. Parking now happens inside the write-back, in the same synchronous turn as the decision that nothing could take the reference; no await separates the live check from the park. Allocations parked by an earlier pass are handed to their slides at the start of the next one, before anything decides what still needs generating, so a held allocation whose scene has since arrived becomes a rewrite rather than an answer. Reclaiming on a rejected write was unsound: a rejection does not prove the server did not apply the write, so deleting the asset could break the scene that now names it. The funnel decides instead, and says so: it reclaims only when no store write was ever issued and nothing took the reference. Anything else is placed if its slide exists and parked if it does not, so the next pass reuses the bytes instead of paying for them again. When a write fails after part of it landed, the live store is brought up to the document before the error is rethrown — otherwise the next ordinary flush would overwrite the half that did land, with the ids deliberately not reclaimed. Parked allocations are now cleared with the course. Classic placeholders are reused across runs, so one surviving an interrupted run would be handed to a different slide of the next deck: the previous picture, on a slide whose provider was never asked. Both classroom surfaces clear the arriving course, the deletion cascade clears the deleted one, and clearing the database clears them all. Two more ways generation could start without asking the gate are closed. An overlapping pass — an outline retry re-enters generation with every outline while the first is still working — re-requested elements whose provider call was already in flight; a task that is not done is an answered request, not an unanswered one. And narration regeneration in the timeline editor called the TTS provider and allocated a pool asset with no ownership check at all; it now reads the same permission, which withholds both the per-line and whole-timeline controls and refuses the call. Finally, a pane opened during the stage-link availability gap recorded the sidecar's 404 for a course that was moments from existing and never asked again, leaving the real owner locked out of generation until it remounted. Ownership is re-fetched once the document becomes available; the gate stays closed until an answer arrives, so asking again can only open it for someone entitled to it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make the asset routes reachable, and stale snapshots harmless A full-branch audit found that the deployment this project documents could not store a single generated asset, and that several routes into durable storage could still write a placeholder over a reference that had already landed. The persistence route sent asset requests through the development authenticator, which refuses outright in a production build that has not explicitly opted into it — and the documented server-persistence recipe produces exactly that build. Every store and every read answered 401, so images and video failed on every slide while re-billing the provider on each retry, and a narration failure stopped the deck at its first slide. Assets live in one shared partition by design, so there was never anything per-caller for that authenticator to decide: the route now resolves the asset principal itself, alongside the owner it already resolves for documents. Runtime sessions are genuinely per-learner and keep the development authenticator until real session verification replaces it. And narration that cannot be stored no longer fails its scene: the line stays unvoiced and retryable, which is what an image that cannot be stored does to its slide. Placeholders could also come back from behind. A queued autosave's snapshot, an editor-history entry replayed by an undo, the departing save a course switch flushes — each captures content at its own moment, and any of those moments can predate a write-back. Point fixes at each producer would leave the next producer to rediscover the bug, so the check lives at the write boundary every producer passes through, and the allocation record it consults now outlives the parked queue: a placeholder whose rewrite landed long ago is exactly the case it catches. Two ways generation could be lost or repeated are closed. A pass now claims the elements it will reach and releases them however it ends, so an overlapping pass stands down while an aborted one strands nothing — previously its tasks stayed `pending` and every later pass skipped them with no retry control to recover them. And the media abort controller is aborted before being replaced, so a superseded pass stops calling providers instead of running on for a course the user has left. The remaining two are narrower. The workbench pane asks for ownership only after a document load succeeds, and after every later one, mirroring the page route: the load is what creates the ownership row the first time a course is opened, so asking beforehand asked about a course that did not exist yet and locked its author out for the mount. And the ownership gate on the timeline editor now withholds narration regeneration alone; listening back to existing narration and seeing whether a line has any spend nothing and stay available. Known limitation, unchanged and now stated plainly in the comments that used to point at it as a solution: nothing reclaims an unreferenced pool asset. The registry sweep is written but not wired up, and the byte collector only reclaims blobs no registry row names, so every narration regeneration and every abandoned allocation leaves storage behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): gate asset mutations, and make claims and allocation records survive a handoff Opening the asset routes opened all of them. Reads and allocations are meant to be as open as document reads and creates already are, but no authorization hook was supplied, so the handler's default admitted PUT and DELETE too — and those scope by principal key alone, which is one shared constant. Any caller who learned an id, and a document read hands out every id its slides name, could overwrite or destroy another author's media. Mutations now require the deployment's credential, which in a production build without the development-auth opt-in means they are refused outright; reads and allocations stay open. The route comment says what the posture is and what it is not: the deployment-level fence is the access code, and no per-principal quota is configured. The client's own reclaim is best effort to match — losing an argument about deleting an asset must not cost a task its retry, and the bytes are left for server-side reclamation. The pass claim could not survive the handoff it was written for. A retry aborts the live media pass and starts its replacement in the same synchronous block, long before the aborted pass's cleanup runs, so the replacement saw every element still claimed, collected nothing, and returned — leaving each unreached element at pending with nobody coming back for it and no retry control to recover it, which is the exact failure the claim was introduced to prevent. A claim now carries its pass's signal and is retired the moment that signal aborts, and a pass releases only claims it still owns, so a late unwind cannot take its replacement's work. Claims are also acquired at the single point every request passes through, so a single-task retry participates too — previously a retry awaiting its provider was invisible to a pass starting alongside it and both called it. The allocation record could outlive the bytes it named. It was written before the write-back attempted anything and survived the reclaim that followed a failure, so when the slide finally arrived the write boundary stamped a deleted id into the document — and the placeholder it replaced was gone, which reads as already generated and stops anything from retrying. The record is now written only where the allocation is retained, and forgotten wherever a reclaim removes the bytes, including the narration rollback path. The tests follow. The route test drives the real storage handler against an in-memory registry instead of a stub, so it can see what the resolved principal is then allowed to do; the handoff test performs a real abort mid-pass rather than starting from an already-aborted signal; and the guards that could only assert file layout now assert the property they care about, or have been replaced by behaviour. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make media passes serial per course instead of tracking element ownership Three rounds of per-element claims each produced a new way to lose an element. Whole-pass reservations swallowed a Retry for an element the same pass had already failed, leaving it pending with the affordance gone. Retiring a claim by its signal freed an element whose commit was still uploading, so the replacement pass paid for it twice. A claim held for a failed element stranded its retry. The bookkeeping is the defect: every refinement of "who owns this element right now" answered the question at a moment when the answer was already stale. Passes for one course are now serial. A replacement aborts its predecessor, as before, and then waits for it to settle before collecting. That removes the question entirely: a commit already under way finishes — its bytes stored and its reference written, so the new pass sees a resolved slide and skips it — and an element the aborted pass never reached is still a placeholder and gets collected like any other. The claim set, the reservations, the signal retirement and the identity-checked release are all gone. The task table is consulted for one thing only: an element that is generating right now is a single-element retry running alongside the pass, and taking it too would pay twice. Pending is deliberately not a skip reason — it means a pass once intended to reach an element, which an abandoned pass leaves behind with nobody acting on it, and reading that as answered is what stranded elements before. A retry runs concurrently with a pass, because a pass never revisits an element it has processed, and it re-reads the task after its own await and refuses before touching it: marking first and refusing afterwards destroyed the failed state that draws the affordance. Browser-only mode is back to exactly what it was. The abort is now conditional, the waiting does not apply, and the original status-based skip is restored verbatim. Two baseline lines remain changed in each of the two files, and both are behind a server-backed fork whose else-branch is the original. Two smaller things. The allocation record becomes visible when a write goes on the wire rather than when the round trip ends, and the write boundary reconciles under the document lock rather than before it — a save queued during a write-back was otherwise captured with the placeholder and, for a course the user had left, had no corrective flush to follow. And the comments that said a refused reclaim leaves its bytes for server-side reclamation were wrong: nothing collects them, because the registry entry still names its blob and the sweep that would remove it is not wired up. They now say the bytes leak. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a deferred pass re-earn its right to run, and bound the commit it waits on Serializing passes moved their body out of the block that launched them, and three things followed from that. A pass now wakes when its predecessor settles, which can be after the user has left the course. It enqueued before it looked at its signal, into a task table keyed by element id alone — and placeholder ids are not unique across courses, which is why the classroom clears that table on arrival. So a departing course's pass seeded the arriving course's table with tasks carrying the wrong stage id, and a Retry routes by that id: the reference went into the wrong document. A pass now re-validates after the wait, before touching anything shared. The same lateness broke the skip test. `documentSkipIndex` answers only while the live store is on the pass's stage, and returning nothing put the collection loop on the browser-only rule — a silent demotion from "the document is the authority" to "this browser's task table is", on exactly the path where that table has just been cleared. Every element the predecessor had committed was collected again, paid for again, and its second write-back found no placeholder to rewrite, so its bytes were parked where nothing will ever reference them. In server-backed mode an unreadable document now means the pass stands down. And waiting was unbounded. A commit is uncancellable: the asset client takes no signal, and a document write cannot be half-undone. One stalled upload therefore froze the course's media generation for the session — the replacement never collected, the element sat on a skeleton that draws no Retry, and only a reload recovered. The pass's signal is now threaded into the media proxy fetch, and the commit is bounded by a deadline. The deadline is on the wait, not the work: the commit carries on, and if it lands late the document simply ends up correct, while the element becomes retryable and the queue moves on. The tests that were meant to pin the previous round were not sensitive to it. Two asserted end states where the mechanism only changes ordering, and one of them rigged the document read so the assertion held whether or not the pass had waited; a third covered half of what it claimed. They now observe the ordering directly — nothing is issued while another pass for the course is working; in browser-only mode a second pass reaches its provider immediately — and the reconciliation under the document lock has a test that fails when it moves back outside it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * revert(media): drop the commit deadline and the abortable download The deadline bought less than it cost. Abandoning a commit after two minutes makes the element retryable while the real commit is still running, so a Retry starts a second commit for the same placeholder against the first: two provider calls, two allocations, and whichever lands second stamps its result over the other's task by element id. The allocation record is keyed by placeholder, so the loser's cleanup erases the winner's record, and the write boundary then puts the raw placeholder back into the document. That is the overlap serial passes were built to remove, reopened through the one door serialization never covered. So a stalled commit holds the course's media queue until it settles or the page is reloaded, and that is written down rather than papered over. The wait is unbounded on purpose: every ceiling on it turns out to be a way of running two commits for one element. Threading the pass signal into the download was also a mistake, in the other direction. The provider call that produced the URL has already been billed, so cancelling the download throws away work that is paid for — and the shared proxy cache records a cancelled request as a transient failure against that URL, which after three of them blocks it for every consumer in the session. Browser-only mode never asked for this: it had no way to observe an abort there, which is exactly why the bytes were kept. The signal is gone from the download again, and `fetchAsBlob` is byte-for-byte what it was before this branch. The regression guard for the stranded-element rule is restored alongside the timing test that was meant to supersede it. It catches a different rule — a task left pending being read as answered — and nothing else does: making the pass skip pending leaves every other suite green. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): stop asking the pool for refs it never issued, bound it, and adopt cached bytes Four things a deployment found once this was running for real. A reference this application mints itself — a generation placeholder, a derived narration key — was never in the pool, because the pool allocates every id it holds. Asking anyway used to be an IndexedDB miss; once the pool is server-backed it is a request that answers 404, one per element per load, forever on a course that still holds placeholders. Every lease and probe now checks first. The check is a negative test on shapes this application owns, not an id validator: the pool's id domain stays unconstrained, and anything that is not one of ours is still asked about. The asset store can bound how much one principal holds, and enforces it inside the write transaction, but nothing ever passed the number. It does now, with a default rather than an opt-in: allocation is reachable by any caller a deployment admits, and with one shared principal an unbounded store is unbounded database growth with no operator-visible brake. Refusing asset mutations to unauthenticated callers was not enough, because every authenticated caller resolves to that same shared principal — so authentication decided nothing, and any signed-in visitor could delete any id they learned. Since this branch began storing media the registry is the only copy a course has. Replacing and deleting are now refused to everyone, and the browser no longer tries: an entry nothing references waits for server-side reclamation instead. What a browser must still do is forget its own record of an allocation that reached nothing, or a later save would stamp an id the document has no reason to trust. And a course generated before any of this holds placeholders in its document with its bytes only in the author's browser. Those bytes are paid for, so the author's next load converts them — stored to the pool and written back through the ordinary commit path, with no provider call — instead of buying them again. A row that records only a hosted URL is treated as absent: that URL is the provider's address, not something a document may hold. One renderer expectation moved with this. An untracked placeholder used to paint as pending on first render because asking the pool left a lease in flight; it settled to disabled a moment later either way, and now says so from the start. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): surface a full store as a refusal and convert legacy narration A quota refusal reached the browser as HTTP 500 with a generic message, which reads as a transient failure: the element kept a Retry that would pay a provider again and be refused again. The store raises the contract's own error and the handler maps it to 507, but the store answering a request is not always built by the same bundle as the handler -- the persistence provider is reached from the route bundle and from instrumentation, which is why its state lives on a Symbol.for global -- and `instanceof` is false across that boundary while the declared code is still right. Classify on the code as well as the class, and make the code a permanent, persisted refusal in the browser: recorded locally so it survives a reload, shown as "storage is full", and refused by the retry entry point so a stale button cannot buy a second generation. Every other storage failure stays retryable. Convert what a pre-server-backed course still holds. Generated media is adopted under either key this application has used for it -- the placeholder, and the allocated id of a course converted once and later rolled back -- instead of only the first. Narration is converted by a load-time pass over the open course's speech actions, since nothing re-enters generation for an action that already has an id: bytes to the pool, id written back through a funnel that mirrors the media one, owner-only and server-backed-only. A line whose bytes are in no browser is left alone rather than re-synthesized. Also: the pool guard is now a positive `ast_` test rather than an enumeration of the shapes we mint (imports never reach the pool, so this is safe in both modes); the slide ref collection is an exported pure function so its four lease sites are covered behaviourally; ASSET_QUOTA_BYTES treats every spelling of zero as opting out and refuses a malformed value at startup instead of falling back; the abort signal is re-checked after the cache read, before an uncancellable commit; and the unused `removeAsset` and pool `replace` surfaces are gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * chore(storage): release 0.29.1 The asset HTTP handler now recognises a store refusal by the contract code it declares as well as by its class, so a quota refusal raised in another module realm answers 507 instead of 500. Same contract, stricter recognition, no API change: a patch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): make a full store recoverable and adoption course-safe Narration adoption read the local audio row by its derived key alone. That key carries no stage id and the table is keyed by id alone, so two courses can mint the same one -- a PPTX import numbers its scenes and actions deterministically, which gives every imported deck's first slide `tts_s1_speech-scene-p1`. Locally a collision only means one course plays another's clip in one browser; adopting it wrote that clip into the shared document permanently, for every device and every visitor. A row that names a course is now adopted only into that course, and a row from before that column existed only when the text it recorded is the text of the action being converted. A full asset store was made permanent last round, which was wrong three times over: it overwrote the refused bytes with an empty blob -- on the conversion path that row is a course's only copy of its own media -- it kept sending the rest of the deck to a provider against a ceiling it already knew was reached, and it left no way back once an operator raised that ceiling. A full store is neither the content's fault nor the configuration's, so it is now its own case: the bytes are kept, the pass stops at the first refusal, and the element shows the reason together with a Retry that re-attempts the upload from those bytes. Nothing retries automatically, so no one is re-billed. The narration write-back now reaches the write boundary every producer of a durable write passes through, not only the dirty mark: adoption never deletes the derived row, so a snapshot that reverts the rewrite is adopted again on the next load and allocates a fresh asset every time. Adoption is also mounted by both classroom surfaces rather than one, takes the course's abort signal, and re-validates that this browser still has the course open before each write. ASSET_QUOTA_BYTES is validated from instrumentation, where the README and the docstring already claimed it was: its only other consumer is lazy and memoised, so a malformed ceiling let the process boot and then failed every persistence request, documents and runtime included. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): remember a full store per course, and never lose retained bytes A stopped pass left the elements it never reached as placeholders with no persisted record -- deliberately, since nothing was attempted for them. But that left the next load with no reason not to try: it called a provider for the next placeholder and was refused at exactly the same point, once per reload, indefinitely. A full store is not a property of any slide. It belongs to the deployment and changes for reasons the document knows nothing about, so it is now remembered once per course in the browser's device KV. A pass that finds the marker stands down before spending anything and leaves every placeholder its "storage is full" state and its Retry; the first upload that succeeds clears it and the next pass runs normally. Narration adoption latched per course so it runs once per load, and the latch outlived the abort that leaving a course performs. On a surface that stays mounted across switches -- the workbench pane is one component for every course it shows -- owner course A, visitor course B, then back to A skipped exactly the clips the abort had cut off, and nothing else converts them. The latch is released with the abort now, and a course adopts one run at a time so a re-entry cannot hand a clip a second allocation while the previous run's uncancellable tail is still settling. A quota-blocked element retried into a network error or a 500 lost the bytes that were kept for it: the retry deleted the row before attempting the upload and wrote no replacement for an error carrying no structured code, so the next retry went back to a provider for media this browser had a moment earlier. The row now survives until an upload succeeds, the failure handler keeps whatever bytes the attempt was given, and the retry asks the question a pass asks -- does this browser already hold bytes for this element -- rather than reading an error code that a second failure has already overwritten. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XypYpLBtk8DB5nT5jyqZ5D * fix(media): adopt real legacy narration, queue re-entries, report attempt outcomes Narration adoption admitted a stage-less row only when the text it recorded matched the action being converted. Both of those columns were added to the local audio table by the very change that moved narration onto allocated ids, so a row still carrying a derived key has neither: the rule refused every real pre-allocation course and passed only on fixtures built from post-allocation rows. What the row cannot say, the key can. A derived key names two clips only when two courses share a scene order and an action id, and an action id repeats only when something other than the generator minted it -- an import numbers them by slide position. So a key built from a generated action id is adopted on that basis, a key an import could have reproduced still needs matching text, and a row that names another course is refused however unique its key looks. Handing a re-entering caller the adoption run already in flight undid the latch release it was paired with: that run is bound to the signal the departure just aborted, so it stops at its next clip while the caller -- which has the course open and a live signal -- is told the work is done, and an effect replayed as mount, cleanup, mount adopts nothing at all. A later caller now waits for the uncancellable tail and scans again, which costs a lookup on a course that has nothing left and finishes the clips the abort cut off on one that does. One attempt at an element now reports both facts its callers need instead of a bare boolean: whether the store refused it for room, and whether bytes actually reached the store. Leaving a course clears the task table, so a retry that landed afterwards read "no failed task" as success and deleted the row holding the only copy of the media. Nothing is inferred from that table any more. Reading the localStorage property can throw where storage is denied by policy, typeof included, so the availability check moved inside the guard: this metadata is best-effort, and a rejection here strands a generation pass that has already enqueued its tasks. A retry is never blocked by the per-course "store is full" marker, but a retry that is refused again re-sets it, and adoption now reads and writes the same marker rather than issuing one refused upload per clip on every load. The two canvas element renderers and both thumbnail renderers show the reason beside the Retry, so a full store does not look like an ordinary failure. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): probe a full store instead of standing down, and pair notices with a Retry Narration adoption was given both halves of the per-course "the store is full" marker last round: it stood down when the marker was set, and it set the marker when its own upload was refused for room. Those halves are only safe together if something can lift the marker, and for adoption nothing could. It has no affordance of its own, it stood down before reaching its own clear, the media pass returns before its marker gate when there is nothing to generate -- so a narration-only deck, or one whose slides are already satisfied, painted no storage-full element and offered no Retry -- and narration generated rather than adopted allocates directly rather than through the media commit. The course's cached narration was then lost for good, where before it converted on the first load after the ceiling was raised. The gate is a probe now. A marked course attempts exactly one clip per load: refused, it stops and the marker stands, which costs what standing down cost; stored, it lifts the marker and finishes the course. Adoption spends no provider money, so the whole cost of probing a store that is still full is one refused upload. Generated narration lifts the marker too. The three surfaces that gained a failure notice last round drew it for any failure with a reason, including the one refusal that is reachable without server-backed persistence, so a browser-only deck painted something it had not painted before. The notice is drawn beside a Retry and nowhere else, which is what it was added for and what leaves browser-only output unchanged. Both are now asserted through the render harness the surface matrix already had. A caller arriving while a rescan is queued shares it rather than appending another. One rescan converts whatever the run in flight left and every later one would find an allocated id on every action, so a chain bought nothing and turned a single stalled upload into a course that never adopts again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): treat a refusal for room as a fact about one clip, not the deck The asset store checks each write against the headroom it has left, so a store that refuses a long opening clip can still hold every short clip behind it. Narration adoption assumed the opposite: it broke the deck at the first refusal and then re-attempted that same first clip on every later load, because the document names it first. A deck whose longest clip exceeds current headroom therefore never converted the clips that would have fit, with no affordance to recover it -- the state the probe was introduced to remove, reached through a narrower door. An unmarked load now attempts every clip, skipping the ones that do not fit, and remembers the condition only if the load ends with clips it still could not store. A marked load spends its single upload on the smallest clip left rather than the first one named: that is the clip that answers the question the marker asks, because if the smallest does not fit nothing does. The media pass keeps stopping at its first refusal, and for a reason adoption does not share -- every element it attempts costs a provider call. A rescan several callers share took the newest caller's signal, and the newest caller is not necessarily the one still there: a surface that opened a course and closed it again would stop work a surface still showing that course was waiting for, and that surface is latched, so it would never ask again. The shared run now takes a signal that is aborted only once every caller has left. The comment claiming the shared rescan contains a stalled upload was wrong -- the rescan is chained off the run in flight, so a stalled upload leaves every caller pending exactly as a chain would. It claims the bounded queue it actually provides, and the stall is recorded as a limitation. The failed-state containers took their stacking classes unconditionally, so markup differed in browser-only mode even though nothing moved on screen. Those classes are applied only when there is a notice to stack, and the tests assert the exact class attribute rather than a substring. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): stop narration adoption writing the media pass's store-full marker The marker means "do not call a provider for this course". A path is entitled to write it only if its own refusal cost a provider call, and narration adoption's refusals cost nothing: it uploads bytes this browser already holds. The store also checks each write against the headroom it has left, so a clip that does not fit says nothing about whether a slide's image would. Adoption was writing it anyway, and one over-long narration clip was therefore enough to stand a course's entire image pass down on every later load -- on a store that had just accepted adoption's other clips. The author could still recover each element by hand, every load, for ever. Three rounds of narrowing this seam produced a finding each time, so it is removed rather than narrowed again. Gone: the marker read, the single-clip probe, the smallest-clip selection, and the up-front read of every row into an array -- which also retires a sampled-then-stale flag and the retention of a whole deck's blobs for the length of a run, and returns the loop to streaming one row at a time. Adoption's rule is now that every load attempts every clip it holds, once; any failure skips that clip and the load continues. The noise the coupling was meant to avoid does not arise, because after the first load the clips still outstanding are exactly the ones that did not fit -- normally none, or one. A successful write still clears the marker, and that is a different kind of statement: a write that went through is a fact this run established, where a refusal is an inference about what some other write would cost. For a course whose media needs nothing, adoption and generated narration are also the only paths that can establish it. The failure module still documented the deck-wide premise this contradicts. It now says what is true: the check is per write, and the media pass stops the deck as a judgement about cost rather than about certainty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): bound a full store's cost from the store's own arithmetic Removing the store-full marker from narration adoption removed its bound too, and the code then asserted the bound was unnecessary. It is, on a store with room for most of a deck. On the store the whole mechanism exists for -- the ceiling reached, nothing fitting -- the outstanding set after every load is the entire deck, so a thirty-clip course posted thirty full blobs on every load, indefinitely. Each of those is not a cheap refusal: the bytes are uploaded, the server hashes the whole payload, and only then takes a per-principal lock and sums every entry that principal owns before saying no. The bound needs no flag, no key and nothing carried between loads. The store asks whether `used + addedBytes` exceeds the ceiling, and `used` only grows while a run is uploading, so a clip refused for want of room implies every clip at least that large is refused for the rest of that run. The run keeps the smallest size it has been refused and skips anything no smaller without uploading it; a smaller clip is still attempted, because it may fit. A deck the store refuses entirely now costs one upload per successive size minimum instead of one per clip, and a deck it has room for costs nothing extra, because nothing is refused. Only a refusal for room lowers the bar: a dropped connection says nothing about how much room there is. The deck-wide certainty premise the failure module retracted last round still stood verbatim at the site that implements the stand-down. Both copies now say the same thing: the check is per write, and the pass stops the deck as a judgement about cost rather than about certainty. The comment on adoption's marker clear now names its price. Narration of a few hundred bytes fits in headroom an image does not, so a proven write can let the next pass buy one more image that is refused again -- bounded at one, and the price of the alternative being a course whose media never generates again. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs(media): state the adoption bound exactly, and stop three comments describing the old rule The comment introducing the in-load bound gave its cost as "at most a handful, and the first load pays the most". Neither clause is a property of the rule. A clip is skipped only when something no larger was already refused, so a fully-refused deck costs one upload per successive size minimum in document order: one when the clips grow, about ln N for an arbitrary order, and one per clip when they only shrink -- a long opener followed by terser lines is exactly that shape. And no load is cheaper than the first, because the bound resets per run and a refused clip stays outstanding. The comment now says that, and points at what would make it exactly one for any ordering: the store returning its remaining headroom in the refusal's existing details channel, which the server leaves empty today. Two other comments still described the previous rule -- "attempts every clip it holds, every load" -- one of them twenty lines above the paragraph that introduces the bound, in the same block. Both now say what the code does. The bound's soundness is worth stating where a maintainer will look for it: quota is charged at full length with no discount for a duplicate, the sum it is checked against joins entries to blobs so the collector cannot lower it, the check takes a per-principal lock before summing, and replace and delete are refused to every browser. Nothing a run can do makes room appear inside it. One test installed a row implementation and replaced it wholesale a few lines later, so the first was dead and the survivor dropped the text the first clip's import-shaped key needs for the ownership rule -- it passed on the coincidence that the fixture's default text is the action's. Merged into one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(media): keep a Retry from re-buying parked media, and state the store seam once Five findings from an inline review. The standalone classroom route asked the ownership sidecar once per load and recorded only stage ownership when that ask failed. Every non-answer fails closed, so one transient 5xx left the genuine author with no resume, no Retry affordance and no legacy narration converted for the rest of the load, with nothing to change it short of a reload. The failure now records the fail-closed answer explicitly -- an answer an earlier load established must not outlive the failure that replaced it -- and an unresolved answer is asked for again, a few times over a few seconds. A real answer, however unwelcome, is final. A Retry could pay a provider for media the pool already held. When the bytes are stored and only the write-back fails in a way that keeps the allocation, it is parked and no local row exists, because that row is written only after a successful write-back. Retry now reads the parked queue exactly as the pass does and re-attempts the write-back: it re-keys the task done when the document takes it, leaves the entry parked when the slide still does not exist, and stays failed and retryable when the document refuses again. Object URLs a parked allocation owns are revoked when the entry is dropped. The commit path leaves them alone while the entry is parked, because it is then the only thing holding bytes this tab can render, so a course switch or a stage deletion was pinning the whole blob for the life of the tab. An entry a slide has already taken is left alone: the task table is displaying those URLs. The fallback lookup for cached bytes is a stage-scoped scan, and the keyed lookup misses for every row the commit path writes, so a pass was materializing and sorting the course's whole media table once per element. One scan per pass now, built on the first miss. It is sound and not merely cheaper: an element asks only for its own placeholder, and every row a pass writes carries the placeholder of the element that wrote it. "The store accepted a write, so it is not out of room" was enforced at three call sites under slightly different conditions, which made it a convention the next pool write path could silently break. It is stated once, in putAsset, for the course whose bytes it just stored. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com> |
||
|
|
d5be393317 |
test(agent-runtime): make the material search budget test deterministic (#1453)
search_material stopped on whichever of the character budget or the 100 ms wall-clock budget fired first, but the budget test asserted the exact character count, so it failed on slow CI runners when the time budget won (observed scannedChars 786432 instead of 1000000). Route the search path's clock through the injectable `now` dependency the wait tool already uses, freeze the clock in the character-budget test so only that budget can stop the scan, and add a time-budget test that trips the deadline after exactly one chunk. Closes #1452 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1fe07df8b9 |
fix(ai): let LLM calls outlive undici's 300s headers timeout (#1404)
* fix(ai): let LLM calls outlive undici's 300s headers timeout A non-streaming completion only receives response headers once the whole completion exists, so undici's default 300 s headers timeout bounds total generation time. Thinking models (GLM-5.x at effort=max, MiniMax M3, …) routinely exceed that on large prompts — the request then dies with "Cannot connect to API: Headers Timeout Error" at exactly 300 s and, with maxRetries 0, the scene never generates (observed on a self-hosted deployment with glm-5.2 scene generation). Route every outbound LLM request through an undici Agent with a 15-minute headers/body timeout, attached at the shared transportFetch seam so the OpenAI-compatible, native OpenAI/Responses, Azure, Anthropic-compatible and Google transports all inherit it (same lazy-import dispatcher pattern the Google proxy transport already uses). Native transports now always install the shared transport instead of only when the server injects a validating fetch. * fix(ai): harden the LLM transport timeout fallback Review follow-ups for the undici headers-timeout fix: - getLlmDispatcher(): drop the cached promise on rejection so one transient undici import/Agent failure can't brick every later call; transportFetch falls back to the base transport (no dispatcher) when the dispatcher can't be built, warning once per failure episode - bedrock: install fetch: transportFetch like every other transport - google proxy: build ProxyAgent with the same headers/body timeout budget (http/https proxies; undici's Socks5ProxyAgent drops these options for socks5://) - transportFetch no longer overwrites a caller-supplied dispatcher - tests: dispatcher failure fallback + retry, bedrock transport, caller-dispatcher preservation, google proxy budget --------- Co-authored-by: XingJia He <hexingjia@mobirit.com> |
||
|
|
9fe650d994 |
fix(quiz): resolve AI answer keys to option values before grading (#1328)
* fix(quiz): resolve AI answer keys to option values before grading
AI-generated quiz answer keys are unreliable: the model may write the
correct answer as option CONTENT ("(6, 2)"), as a LETTER ("A"), or as a
formatting variant (full-width chars, extra spaces) of either. Grading
compares exact option values, so a learner picking the correct option
was still graded incorrect whenever the key was written as content or
in a different format.
- generation: normalizeQuizAnswer now resolves every answer variant to
the matching option value (NFKC + whitespace-stripped + letter-prefix
canonical matching against option values and labels); unknown answers
pass through untouched
- grading: resolveAnswerKeyToValue retro-fixes already-generated courses
whose stored keys hold content variants, mapping them to option values
when exactly one option matches canonically; ambiguous keys are left
alone instead of being silently re-pointed
- tests: letter/content/full-width/trailing-space/multi-answer/unknown
cases + end-to-end grading with a content-variant key
* chore(generation): bump version to 0.3.2 for answer-key fix
The repo requires every publishable @openmaic package change to ship with
a version bump (CI check-package-version-bumps).
* style(generation): prettier-format package.json after version bump
* fix(quiz): use resolved answer representation in review renderers; prettier
- QuestionCard + quiz-view review modes highlight the option canonically
matching the stored answer key (answerIncludesOption), so persisted
content-variant keys highlight correctly after the grading fix
- prettier formatting on touched files
* test(quiz): guard possibly-undefined answer in multi-choice test
* fix(quiz): narrow answer-key matching to exact unique alignment; project review UI from the same resolver
Per review: matching policy narrows to the DSL option-identity contract.
- Generation: exact value or unique exact label alignment; unknown and
ambiguous keys stay unresolved (fail closed). Prompt example unified to
the DSL object-options + value-answer shape.
- Grading: persisted keys resolve via the same exact alignment; the
resolver returns the option's actual value when exactly one option
matches canonically-or-exactly; ambiguous keys stay untouched.
- Review UI: QuestionCard and both quiz-view review modes highlight via
the same resolver projection (answerIncludesOption(q, optionValue)).
- Removed the broader equivalence rules (NFKC, whitespace stripping, case
folding, wrapped/prefixed-letter probe) — semantic synonyms such as
正交/垂直 remain uncovered, hence Refs #1326.
* fix(quiz): resolve answer keys for persisted keys only; align quiz template with the DSL
- grading: compatibility resolution (label -> value) now applies to the stored
answer key only. A learner submission is compared as submitted, so an option
label can no longer be accepted as a different option's value.
- quiz-content/user.md: replace the string-options/`correctAnswer` example with
the DSL shape the system prompt defines ({ label, value } options and
`answer` as an array of option VALUES), so generated keys are values.
- tests: regression that a label submission grades incorrect while the same
question's value submission grades correct.
* test(generation): update scene-prompt golden snapshot for the DSL-shaped quiz template
The golden test pins the rendered user prompt for every scene kind; the
quiz-content template now emits the DSL shape, so the snapshot follows.
* fix(quiz): read label-stored correct answers in the editor's option rows
`quiz-edit-ops.toRows` decided correctness with raw value equality, so an
existing label-keyed quiz showed A/B as correct (the review UI and the grader
both resolve labels) while every row read as incorrect internally — and the
next option edit or reorder rebuilt `answer` from those rows and silently
dropped the key.
- `toRows` now uses the same resolved projection (`answerIncludesOption`).
- `normalizeQuizAnswer`'s docstring now describes the exact-only behavior it
has had since the review: formatting variants are not normalized and
ambiguous entries pass through untouched.
- Regressions: an unrelated option-label edit and an option reorder both keep
a label-stored answer, and a multiple-choice toggle preserves it. All three
fail against the previous raw comparison.
---------
Co-authored-by: CrazyPandaP <CrazyPandaP@users.noreply.github.com>
Co-authored-by: wyuc <wang-yc24@mails.tsinghua.edu.cn>
|