mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-03 17:59:34 +08:00
* feat(agent): make generate_video asynchronous with a placeholder ref The generate_video tool awaited the whole provider submit/poll/download/ persist cycle inside the tool call, blocking the agent turn for minutes (and effectively capping it at the 10-minute tool budget, below its own 15-minute internal budget). The tool now validates synchronously, mints a gen_vid_<id> placeholder (the scheme the outline flow already uses), and returns immediately so the agent can patch_stage the ref onto a video element and keep working; the element renders the existing skeleton while pending. A detached background job (own 15-minute timeout, deliberately decoupled from the caller's abort so a cancelled chat cannot silently orphan a billable provider job) runs the provider cycle, persists the bytes, and then: - patches the persisted document: every slide video element still referencing the placeholder gets the concrete server-hosted src (same runStageMutation discipline as the generation tools; skipped silently when the element was changed meanwhile), and - appends a media_ready lifecycle event to the session's durable log via the session-level control channel (valid post-run, unlike the runner's lease-guarded emit); the workbench folds it into the media generation store so the skeleton resolves instantly, on live stream and on replay. Pending tasks live in a process-local registry; a server restart orphans in-flight jobs (the placeholder keeps its skeleton), matching the classic flow's client-local durability caveats. Tracked as an accepted v1 limitation. Closes #1266 * fix(agent): unfence the background video patch from the run lease Review findings on the async generate_video change: - P1: the completion patch wrote through the run's owner-bound store, whose mutation fence asserts the run lease on every write. A video job settles minutes after its run ended and the lease is released, so every post-run patch threw AgentSessionLeaseLostError and the job wrongly settled failed. The runner now builds a dedicated owner-bound store for media jobs fenced only by the stage-mutation discipline, and the tool takes it as a separate backgroundStore dep. - P2: patchStageVideoPlaceholder rewrote whole scenes from one minutes-old loadDocument snapshot. It now re-reads each candidate scene immediately before its write and applies the placeholder swap to the freshest state, so concurrent user/agent edits survive; the residual read-write window matches the stage edit API's own read-modify-write discipline. - P3: drop the dead 'emit' progress marker (setPendingMediaStage is a no-op after settle), emit a media_ready failed frame from the last-resort crash guard so a bug path cannot leave the client on a permanent skeleton, and log when appendControlEvent silently drops a frame for a deleted session. * test(agent): cover the concurrent placeholder-element removal case Round-2 review leftovers: pin that patchStageVideoPlaceholder skips rather than resurrects a placeholder element deleted between the candidate scan and the write, and keep the detached job's crash guard synchronous (never-rejecting helpers only) so a future throw inside it cannot become the unhandled rejection the guard exists to contain. * fix(agent): isolate the completion patch from the provider budget Round-3 review leftovers: - The patch shared the job's 15-minute signal, so a provider cycle that nearly exhausted the budget could fail mid-patch and rebrand a persisted, downloadable asset as failed. The patch now runs on its own 60-second budget and a patch failure is logged while the job still settles and emits done (the done frame's src renders fine). - A user edit that replaced the placeholder with a concrete src while the job ran no longer gets clobbered by the stale mediaRef: the swap only writes while src is absent or still the placeholder. Accepted, documented: the sub-second read-write window between two concurrent jobs on the same scene (the client-side media fold renders either way), and the failed-state fold requiring an attached chat stream in v1. * fix(agent): close the regeneration gap in the placeholder src guard Delta review found the new patch-failure test never attempted the write (no element carried the runtime ref, so putScene was never called), and that the src guard also skipped legitimate patches: an element re-pointed at a new job via mediaRef while still carrying the previous generated src (which keeps rendering the old video), and an empty src. The guard now also writes when src is empty or a previously generated /api/classroom-media/ URL; the test seeds the ref so the failing write is really attempted and asserts the logged patch failure. * fix(agent): harden the placeholder src guard - Total predicate: a malformed non-string src no longer throws inside the element map and aborts the whole stage patch. - Recognize the absolute-form generated src the classic pipeline persists, and scope the generated-src arm to the stage's own media root so a user's pick copied from another stage is preserved.