100 Commits
Author SHA1 Message Date
周小舟 51c2f09d81 docs: align video thumbnails across cards 2026-10-01 23:45:04 +08:00
周小舟 a8cac39ef0 docs: add sixteen playable video demos to translated READMEs 2026-10-01 23:42:47 +08:00
周小舟 ce35f98b90 docs: preserve the existing sponsor sections verbatim 2026-10-01 23:28:54 +08:00
周小舟 7d33f3c9d4 docs: review current behavior and remove stale README claims 2026-10-01 23:21:45 +08:00
周小舟 8f9416f21e docs: align eight READMEs with the 1.5.0 release 2026-10-01 23:14:47 +08:00
周小舟 52b20fac68 docs: record 1.5.0 latest release and issue closure 2026-10-01 22:49:56 +08:00
周小舟 9a130ee61a docs(1.5): record merged tag and final long-video acceptance 2026-10-01 22:07:56 +08:00
周小舟 1b02347c3e fix(studio): rank eligible clips per platform and settle empty queues 2026-10-01 21:47:54 +08:00
周小舟 0fe48c051e fix: settle observed AI cover results before replacement 2026-10-01 21:34:55 +08:00
周小舟 0147056d21 fix: complete 1.5 delivery telemetry and publish CLI MCP assets 2026-10-01 21:28:48 +08:00
周小舟 1d433ea1f5 docs(1.5): record cross-platform installed wheel acceptance 2026-10-01 21:07:26 +08:00
周小舟 ee73e594ff fix(render): select bundled CJK family across font providers 2026-10-01 20:58:33 +08:00
周小舟 e3e57236a6 fix(release): bundle face detector and verify installed portrait output 2026-10-01 20:50:46 +08:00
周小舟 e23dc999f2 fix(cli): apply Windows connection cleanup at headless startup 2026-10-01 20:25:07 +08:00
周小舟 5c0b82e0ee fix(ci): isolate runtime tests and sync quick-output docs 2026-10-01 20:07:59 +08:00
周小舟 d3e3cb1a84 fix(1.5): complete output acceptance and align CLI/MCP 2026-10-01 20:00:45 +08:00
周小舟andClaude Opus 5.5 723eefbd52 feat(studio): blur burned Chinese captions out of English versions; 1.4 AI covers off
- Detection also records where the burned captions sit. An English version
  of a source with non-English captions blurs that band before any layout,
  then is framed and captioned like a caption-free source.
- AI covers switched on by 1.4's first setup read as off after upgrading
  (1.5 would bill one per output in the background); the chosen model stays
  and one click turns them back on. Every 1.5 save marks the choice explicit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 19:08:26 +08:00
周小舟andClaude Opus 5.5 ecd0dda509 fix: failures degrade instead of losing work; record the review
- cloud ASR retries a rate-limited or failing chunk with backoff and stops
  sending the rest once one fails for good; a bad concurrency env var is safe
- a stalled hardware encoder falls back to libx264; a bad input no longer
  disables hardware encoding
- a failed outro delivers the video without it; Windows font paths and
  apostrophes in paths work in the outro
- no AI cover without permission to send the frame (no invented face
  under the guest's nameplate)
- handoff doc: review results, open decisions, deferred findings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:58:06 +08:00
周小舟andClaude Opus 5.5 6931bc1c93 fix(studio): English platforms stay English on every path
Review found several ways Chinese text still reached TikTok/Reels/Shorts/YouTube:
- packaging kept Chinese titles, captions and nameplate roles from the model;
  Chinese captions are now retried, the rest dropped or filtered
- post copy accepted Chinese titles; they are retried, never posted
- landscape YouTube burned the untranslated source captions and Chinese hook;
  landscape versions now caption in their audience's language
- appended and on-demand versions skipped the English title and post copy
- publishing fell back to the Chinese clip title; it now asks for one
- the kit's file names and the share credit followed the app language
Also: a version interrupted mid-render becomes retryable after a restart,
and a DashScope relay address from 1.4 is no longer rewritten.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:51:41 +08:00
周小舟andClaude Opus 5.5 3b8088c8a4 fix(studio): backups render only once prepared; review fixes
- A backup clip being framed and packaged is 'preparing', invisible to the
  render dispatcher, so it can no longer ship raw; claiming it is atomic.
  A failed preparation settles the generation, retry prepares it again, and
  one interrupted by a restart becomes retryable.
- The results page keeps polling while outputs are preparing or rendering.
- AI cover redesign: no false success on timeout, nothing after unmount.
- Saved post copy shows at once; title lines keep their place while typing.
- Generated-image downloads refuse local/private hosts (each redirect too)
  and stop at 40 MB.
- Docs: cloud ASR numbers, CHANGELOG entries for this round.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:37:01 +08:00
周小舟andClaude Opus 5.5 a0108d889c fix(studio): show the whole cover in the kit; drop dead raw-clip cards
The output card's 16:9 thumbnail crops a 9:16 cover to the face, so the kit
now shows the cover at its own ratio. Automatic projects no longer list raw
clip cards after their outputs: fast output never renders raw clip files, so
their previews 404'd and Download/Export did nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:25:25 +08:00
周小舟andClaude Opus 5.5 0fc14a61fb fix(cover): AI covers fill the platform frame, no flat bars
GPT Image 2.x is asked for the platform ratio (a 2:3 image padded to 9:16 left
bands above and below); a relay that rejects it gets the legacy 2:3 size. When a
model still returns another ratio, the gap continues the image's own edge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:12:41 +08:00
周小舟andClaude Opus 5.5 1317b5d4b8 docs(handoff): demos rebuilt; note 9:16 AI cover padding
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:06:09 +08:00
周小舟andClaude Opus 5.5 ed14a34953 docs: 1.5.0 fast-output handoff, kept current as work lands
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:01:05 +08:00
周小舟andClaude Opus 5.5 7814dbd4ee feat(cover): designed covers by default; AI covers only when the user picks a model
First-run settings no longer switch AI covers on or pick an image model:
covers are designed locally (free, instant) until the user chooses AI and an
image service and model themselves. The cover source reads 'auto design
(free)' / 'AI generated' with a neutral hint (no vendor named). An
OpenAI-compatible connection pointed at fal.run is recognised as fal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:35:00 +08:00
周小舟andClaude Opus 5.5 4d11ece037 fix(settings): an unknown image API never invalidates the settings file
A settings file written by a newer version (e.g. image_api 'fal') made an
older backend reject the whole file, failing every task at screening.
Unknown image APIs now fall back to auto.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:23:03 +08:00
周小舟andClaude Opus 5.5 41fa605655 feat(cover): fal image connections (GPT Image 2.5) and correct-by-construction text
- Image connections can use fal (image_api 'fal', https://fal.run): the
  model path gets /edit with the reference frame, else /text-to-image, and
  the platform's exact size is requested.
- AI covers never draw names (models invent them: 'WILDER', 'Terence');
  our own nameplate is added only when the vision model confirmed the guest.
- Text must stay inside a safe area; the result is padded, never cropped,
  so a headline can no longer lose its edges.

Compared on six real clips (same frame, same prompt): GPT Image 2.5 Flare
via fal got 6/6 headlines right in 22-31 s; qwen-image-3.0 5/6 with
overlaps; Seedream 5.0 had the most misspellings and refused two
celebrity frames.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:18:41 +08:00
周小舟andClaude Opus 5.5 e7fb0c18c5 fix(packaging): one language per platform; no false burned captions
- Burned caption detection ignores the outer 15 % of the band (channel logos
  fooled it on Hot Ones) and the vision model, when set up, has the last
  word: no dialogue captions in the picture means none (slides, products and
  logos do not count). 15 real sources now classify correctly.
- English platforms never show Chinese: versions take the English packaging
  or post title, fallback titles drop CJK lines, post-copy fallback leaves an
  English title empty rather than Chinese. Chinese platforms never fall back
  to English-only captions.
- When packaging fails for a foreign audience, a plain line-by-line
  translation still gives captions in the audience's language.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:51:31 +08:00
周小舟andCursor 95153d1fda docs: sync the seven README translations to the one-link, one-click rewrite
Keep the same command blocks and the single Trendshift badge so i18n-sync can pass, and retell the new product pitch in each language instead of leaving the old feature-grid pages behind.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 16:20:14 +08:00
周小舟andClaude Opus 5.5 47b6f7945c feat(cover): the vision model picks the guest's frame; AI covers upgrade the design
- Cover frames: candidates are the sharpest face frame of each long shot;
  the vision model picks the guest's best one (never the host). Without a
  vision model the sharpest face is used and no nameplate is drawn, since
  colour or screen-time rules cannot tell a host's camera from a guest's.
  Tight crops are lightly sharpened after upscaling.
- AI covers: after the designed cover, an image model (when set up)
  generates an editorial magazine-style cover from that frame, in the
  platform's ratio (qwen-image sizes now cover 3:4/1:1/4:3) and the clip's
  palette, in the audience's language; the vision model checks the headline
  and a wrong one is retried once, else the designed cover stays. Runs on
  its own pool so renders never wait; 'AI 重新设计封面' uses the same path.
- Image requests wait up to 180 s (high-quality models are slow).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:18:22 +08:00
周小舟andCursor 3a69c639a1 docs: tighten README layout so the first screen and proof stay scannable
Move the trending badge into the hero, collapse speed/cost into three rows, and keep sponsor logos small so the page reads as a product pitch rather than a dump of internals.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 16:11:21 +08:00
周小舟andClaude Opus 5.5 882a59d96a fix(cover): long English titles rewrap larger and push the photo down
Two-line English titles shrank to unreadable sizes; they now rewrap into up
to four lines at a readable size, keeping the accent words coloured, and the
portrait title band grows with the title so it never covers the face.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:43:17 +08:00
周小舟andClaude Opus 5.5 82f82e035e perf(asr): 8 parallel cloud transcription uploads; benchmark keeps case order
qwen-audio-3.1-asr-flash on a 2 h 52 m Chinese talk: 215 s with 8 uploads
(local Whisper base: 23 min), 191 s with 16, so 8 is the default.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:30:21 +08:00
周小舟andClaude Opus 5.5 a3066604bb feat(settings): dedicated Bailian endpoints for DashScope connections
The DashScope connection now takes an optional endpoint address (host,
/api/v1 or /compatible-mode/v1 are all accepted): speech recognition uses
its /api/v1 path and text models its OpenAI-compatible path. Also makes a
timeline test independent of chunk call order now that chunks run in
parallel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:13:44 +08:00
周小舟andClaude Opus 5.5 6786f2eaf1 perf(asr): upload cloud transcription chunks side by side
Three-minute chunks were uploaded one after another (~60 round trips for a
three-hour talk). Chunks are cut first, then sent on a pool of 4
(AUTOCLIP_ASR_CONCURRENCY), merged by exact sample offsets; any failed chunk
still publishes no partial subtitle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:04:32 +08:00
周小舟andCursor 7fd9873a6b docs: lead README with 10+ clips per link; drop competitor names and roadmap hints
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:59:43 +08:00
周小舟andCursor d2cf65197e docs(design): serif web headlines and the risograph art layer
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:57:55 +08:00
周小舟andClaude Opus 5.5 53ececdcb2 test(bench): fast-output regression set with per-stage timing
- 17 fixed real inputs (owner's picks: long interviews, short sources,
  keynotes with screen demos, Chinese sources with burned captions, auto-
  caption-only and once-360p sources) in benchmarks/fast_output/cases.json.
- scripts/fast_output_benchmark.py imports each through the product API and
  reports per-stage time, model calls, tokens and cost, clips and renders,
  source resolution, subtitle source, burned captions and packaging
  fallbacks; --baseline prints deltas.
- Stage wall times (download, transcribe, clip finder, boundaries, framing,
  packaging, post copy, render) are recorded next to token usage.
- Changelog: publish kit, outro, caption strategy, low-resolution fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:53:19 +08:00
周小舟andClaude Opus 5.5 35ece30454 feat(studio): output cards carry their publish kit
Each output card shows its designed cover as the video poster and the
platform's post copy (title, description, hashtags) with copy, inline edit,
export publish kit (native save on desktop) and AI cover redesign. The
publish page starts from the kit: title, description with hashtags, and the
cover from the matching vertical/landscape slot instead of generating one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:46:48 +08:00
周小舟andCursor 0c2b275cea docs: README leads with one link, one click; adds the TIM run and stage timing
DESIGN.md records the homepage split hero, band rhythm, editor/chat/
one-click comparison, sticky story and the stage-time ribbon.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:44:01 +08:00
周小舟andClaude Opus 5.5 5c91702afc fix(packaging): caption burned-in sources only for audiences that cannot read them
- Burned caption detection missed thin, outline-free captions (Bilibili
  interviews such as TIM x Luo Yonghao): the bottom band is now analysed at
  960x192 with a matching threshold; all seven test sources classify right.
- The burned captions' language is read by the vision model when set up,
  else taken from the video title's language (captions are written for the
  channel's audience), else the speech language.
- Chinese captions in the picture: none added for Douyin, English for TikTok;
  English captions on a Japanese talk: none for TikTok, Chinese for Douyin.
  No translation is requested when no caption will be shown.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:42:20 +08:00
周小舟andClaude Opus 5.5 1dc8a09c27 feat(branding): the designed 1.8 s outro animation replaces the text card
Pre-rendered vertical and horizontal outro (silver cut, logo, wordmark,
chime; source and OFL font in design/outro-v5) is conformed once per output
spec (size, fps, time base, audio layout), cached, and joined by stream
copy; other aspect ratios are padded with its background. The old text card
remains as the fallback when the asset is missing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:35:01 +08:00
周小舟andClaude Opus 5.5 09a5ffabc8 fix(studio): re-download YouTube sources that arrive at 360p
YouTube's SABR experiment can leave only the 360p format reachable for the
default client; the MrBeast demo was cropped from a 640x360, 100 kbps file
and looked blurry. After download, if the file is below 720p while the
listing offers more, retry with the web_safari / web_embedded / tv clients
and keep the sharpest file. The final height is kept in source_meta.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:28:32 +08:00
周小舟andCursor 88189f2347 docs: point README to the case library and add a showcase submission template
README links the case library and describes the per-platform publish
kit; a Show and tell discussion template collects community clips with
source, platforms, setup and display consent.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:19:19 +08:00
周小舟andClaude Opus 5.5 05b33bfad4 docs: measured cost and time per video link
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:12:51 +08:00
周小舟andClaude Opus 5.5 4ca3ff4a7e feat(publish): publish an output with its kit copy and tags
Publishing an output variant uses its post copy when the user leaves fields
empty: title, description and hashtags for Upload-Post; title, description
and up to ten tags for Bilibili (instead of the fixed tag).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:12:15 +08:00
周小舟andClaude Opus 5.5 0123851266 feat(studio): every output ships a publish kit: cover, post copy, bundle
- Post copy: one model call per clip writes title, description and tags for
  each selected platform (Douyin hook, Xiaohongshu note, Bilibili, English
  captions for TikTok/Reels/Shorts/YouTube); platform limits are enforced in
  code; failures fall back to the clip title. Backup clips get copy when
  produced; users can edit it (PUT .../post).
- Cover: designed locally after each render from the speaker's frame (track
  converted to the speaker centre), the video's own title lines and palette,
  the guest's nameplate; 9:16, Xiaohongshu 3:4, Bilibili 16:10, YouTube 16:9.
  It fills the publish cover slot unless an AI cover exists, which wins.
- Kit: GET .../kit zips the video, cover and copy; GET .../cover serves it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:06:03 +08:00
周小舟andCursor 0ee5633ca6 docs: rewrite README around real outputs and V2 speed/cost numbers
Lead with an animated wall of untouched outputs and the measured
before/after table (1h43m interview: 7.5 min, ¥0.09), then features,
a subscription-tool comparison and the existing install paths. Sponsors
move below the quick start in a compact table.

DESIGN.md records the homepage stage band and number-proof rules
approved for the website redesign.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:02:36 +08:00
周小舟andClaude Opus 5.5 355ee9bb84 docs: record pipeline V2 progress and regression numbers
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 13:23:30 +08:00
周小舟andClaude Opus 5.5 6cfe7e4d10 perf(studio): package a batch of clips side by side
Packaging waited ~20 s per clip on the model, one after another (25 clips
~8 min). Distinct clip/template packages of a batch now run on the shared
pool before variants are built; platforms of the same template still reuse
one package.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 13:20:53 +08:00
周小舟andClaude Opus 5.5 ceb735e342 perf(studio): use the creator's uploaded subtitles instead of speech recognition
When a video has uploaded subtitles in its own language (e.g. Dwarkesh
Patel's YouTube interviews), fetch them as the transcript and skip local
Whisper (~9 min per 2 h of audio). Automatic captions are never used: no
punctuation and rolling repeats break sentence cut points. Any problem or a
stub file keeps speech recognition.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 13:16:14 +08:00
周小舟andClaude Opus 5.5 f02be3e0d2 feat(studio): render the top 10 clips automatically, the rest on demand
A two-hour talk yields 30+ clips; rendering all of them up front was most of
the run while people publish a handful. Automatic output now frames,
packages and renders the 10 highest-scored clips; the others are listed as
backup clips (status on_demand, no model or render cost) and are framed,
packaged and rendered when the user clicks 'Generate this one'. Adding a
platform keeps each moment's choice. Generation completes once the
automatic versions finish; analytics count on-demand versions separately.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 13:11:11 +08:00
周小舟andClaude Opus 5.5 5946cd70cc perf(render): encode studio clips with the hardware H.264 encoder
VideoToolbox on macOS, NVENC/QSV/AMF elsewhere, proven once with a tiny
test encode; bitrate scales with the frame (8 Mbps at 1080x1920@30). A
failed hardware encode retries the clip with libx264 and sticks to it.
AUTOCLIP_VIDEO_ENCODER=x264 forces software. Same wall time on this Mac but
4.5x less CPU per encode (3.0 s vs 13.4 s CPU for 30 s of 1080x1920).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:58:00 +08:00
周小舟andClaude Opus 5.5 5aa5f0fdb9 feat(pipeline): find clips in one pass over the whole transcript
Fast output no longer runs outline -> timeline -> scoring -> titles over
5000-character chunks. One call per ~60-minute window (parallel, with
overlap) reads compact id|time|text rows and returns finished clips with
rows, Chinese title, reason and score; picks are validated, deduplicated
across windows, snapped by refine_timeline and selected by select_clips, in
the legacy titled-clip shape. Short-video lengths (60-300 s, ideally
90-180 s) replace the podcast-era 90 s floor that padded answers into the
next question. Any failure falls back to the legacy steps.

MrBeast 2h06m: ~35 min / 128 calls / 500K tokens -> 38 s / 4 calls / 93K.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:53:21 +08:00
周小舟andClaude Opus 5.5 226b3dbeb4 perf(pipeline): run per-chunk model calls side by side
Outline, timeline, scoring and title steps called the model once per
transcript chunk, one after another: a 2 h video spent ~30 min waiting on 85
sequential requests. Independent chunk calls now run on a pool of 4 (local
models stay sequential; AUTOCLIP_LLM_CONCURRENCY overrides), results keep
chunk order, and token tracking follows each worker.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 12:33:49 +08:00
周小舟andClaude Opus 5.5 822886e7f6 docs: fix the fast-output workflow and what was trimmed
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 11:49:57 +08:00
周小舟andClaude Opus 5.5 628b01b77e perf(studio): trim the fast-output chain and record model token usage
- Fast output runs the content pipeline without clustering (one unused model
  call) and without re-encoding every clip and collection (studio renders
  from the source); clip rows still sync from step 4 for the candidate list.
- Packaging no longer asks the model to restate the transcript when no
  translation is needed; its rewrite was discarded for same-language output.
- One face-detection pass per clip serves both the 4:3 interview window and
  the 9:16 podcast crop.
- LLM inputs are compact JSON (indentation cost ~10 tokens per subtitle row).
- Every text and vision call records tokens per project and stage into
  metadata/llm_usage.jsonl (DashScope native usage is captured now), so the
  cost of one video can be measured.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 11:47:50 +08:00
周小舟andClaude Opus 5.5 1931c51714 fix(packaging): clean fallback titles taken from the draft hook
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 04:05:53 +08:00
周小舟andClaude Opus 5.5 ecad8d9a3e fix(packaging): retry rejected responses, plain audience-language titles, varied batch palettes
- Retry the packaging call once with the rejection reason before falling back:
  short Japanese ASR rows made the segment contract fail intermittently and
  the fallback dropped the Chinese captions entirely.
- Decode HTML entities and strip markdown from titles and captions; reject
  kana titles on Chinese platforms; fallback titles keep whole clauses.
- Clips of one batch avoid the previous palettes when the mood allows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 04:01:06 +08:00
周小舟andClaude Opus 5.5 088dde840e docs: record boundary snapping and mood looks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 03:50:07 +08:00
周小舟andClaude Opus 5.5 a7b6b57264 feat(packaging): pick palette and style from the content mood
The packaging model names the clip's mood (calm/serious/bold/warm/playful);
the mood decides which of seven content palettes and which template styles
fit, and a seed from the clip picks among them, so clips vary while one clip
keeps its look on every platform. Content colours are no longer tied to the
product UI accent; light accents get dark text on filled pills.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 03:49:09 +08:00
周小舟andClaude Opus 5.5 206d7421e7 fix(studio): end clips in a real pause after a finished passage
Subtitle rows carry no word timing, so cuts estimated inside a row leaked a
few hundred ms of the next sentence. Snap both cuts to ffmpeg-detected pauses:
the pause must fit how fast the words around the cut can be spoken; when the
speaker runs straight on, finish the next sentence that ends on a breath
(never past the next question); a long pause just after extends the clip to
finish the passage. Starts may reach back 20s to a question's opening.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 03:43:38 +08:00
周小舟andClaude Opus 5.5 2272d8407d fix(studio): cut inside rows at sentence ends and never end on the next question
Real Dwarkesh rows often end one sentence and start the next
("...responsible. It seems like"), and punctuated transcripts treated
0.6 s pauses as sentence ends, so clips still ended on openers. Cut inside a
row at its last sentence end, start at the first full sentence of a row
that begins on a tail, use pauses only for unpunctuated transcripts, treat
Japanese polite sentence-final forms as sentence ends, and end before a
trailing question followed by only the opening of its answer. Translated
segments may span up to 12 short ASR rows so long answers keep packaging.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 02:14:19 +08:00
周小舟andClaude Opus 5.5 d2f97bcf29 fix(studio): end clips on complete sentences and thoughts; caption burned foreign sources
Clips ended mid-sentence or mid-answer because the pipeline picks bounds on
ASR subtitle rows, which are half sentences. Snap every clip to sentence
edges (terminal punctuation or a clear pause, bounded extension), then,
when a text model is available, let it move the edges so the question
starts and the answer finishes, rejecting drastic rewrites.

A source with burned captions in another language (Japanese talk, English
captions) now still gets audience-language captions below the window, and
Japanese is no longer mistaken for Chinese. Same-language captions follow
source rows so word timing stays in sync, and the cinematic style shows
calm 2-5 word phrases instead of single-word flashes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 01:55:12 +08:00
周小舟andClaude Opus 5.5 853a13cc2a fix(packaging): keep titles and bilingual captions inside the frame
Real demos exposed three layout failures: a 17-character Chinese title ran
off both edges at 112 px, a long English original under a short Chinese
line wrapped to a third line, and a model lumped 25 subtitle rows into one
2,000-character segment. Title lines now shrink to fit the 1080 px frame and
long model titles rewrap at punctuation or spaces; a long original gets one
caption line per screen and is cut with an ellipsis rather than a third
line; segments may span at most six rows. For same-language outputs a bad
segmentation falls back to the source rows but keeps the model's title,
names and highlights.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 23:56:40 +08:00
周小舟andClaude Opus 5.5 506beb7291 feat(analytics): observe automatic generation and template packaging
Automatic output has no confirmation step and no client export, so the
screening and production watches never settled and the only terminal event
depended on the results page staying open. Screening now ends as
auto_started once production begins, and a persisted studio-generation
watch reports studio_generation_finished once with aggregate counts
(variants, templates, speaker framing, packaging fallbacks, trims, burned
captions) and backend-timed duration. The backend keeps the analysis run_id
across phases, records generation.finished_at and marks automatic framing
as framing_source=auto.

Draft save/export carry template, style, tags switch and fallback; variant
downloads, shares and ratings carry the platform, template, style and
framing plus flow context. All new fields are allowlisted enums, counts or
booleans; no titles, captions, names or tags leave the app.

Also fix two issues found while rendering demos: deriving a draft for
another platform now resets the title template version (Xiaohongshu after
Douyin failed validation), and English caption segments and long model
titles no longer force a packaging fallback.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 23:49:46 +08:00
周小舟andClaude Opus 5.5 f006b08fdc Merge origin/main into worktree-project-diagnosis
Resolve conflicts with main's onboarding/Studio delivery analytics, download
recovery, 1080p60 export and SenseVoice changes:
- analytics allowlist and operation names are the union of both sides;
  layout adds 'window'; import properties keep flow ids plus platform count
  and outro flag
- Studio download keeps recovery and writes a temporary info.json to read
  the listing title/channel for nameplates
- 1080p60 stays an export-only preset next to the platform projections, and
  keeps the outro frame-rate fix
- locale catalogs merged key by key (no key changed on both sides)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 23:28:59 +08:00
周小舟andClaude Opus 5.5 ed2605ca74 fix(packaging): translate sentence segments instead of subtitle rows
Real Dwarkesh subtitles are cut mid-sentence and the model regroups them
(21 rows came back as 13 sentences), so the one-to-one row check rejected
every result and all outputs fell back to untranslated captions. Ask for
contiguous sentence segments over the row ids, validate order and coverage,
and time each cue from its rows. Tags must name the concrete point instead
of generic praise, and fallback titles wrap by their own language.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 23:18:34 +08:00
周小舟andClaude Opus 5.5 910d1c8690 feat(studio): two-line caption screens and template styles
Vertical captions stacked up to five lines: libass lays SRT out on a 288 px
canvas, so Fontsize 20 became ~133 px on a 1920 px frame. Scale the legacy
force_style for portrait and split every row into consecutive screens of at
most two lines, breaking on punctuation and natural joints, sharing time by
width and keeping the original translation in step. Template captions use
the same layout.

Add styles on top of the golden defaults (interview: classic, caption bar,
spotlight; podcast: word pop, caption bar, cinematic) with the same layout
contract and DESIGN.md palette; Studio can switch style on a packaged draft.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 23:07:25 +08:00
周小舟andClaude Opus 5.5 c105acc581 feat(studio): show template packaging and allow title edits
Result cards name the template (interview with bilingual captions, or
podcast) and say when packaging fell back to source captions. In Studio a
packaged draft shows a template panel with the two title lines and an editor
tags switch; legacy subtitle, opening-title and language settings are hidden
because the template renders those. Packaging is kept intact on save.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:45:02 +08:00
周小舟andClaude Opus 5.5 ab7906828a feat(studio): package automatic vertical outputs with platform templates
Chinese platforms (Douyin, Xiaohongshu) get the interview template: a pinned
two-line title, a 4:3 window that follows the speaker on a warm near-black
canvas, Chinese captions over the window with the original below, editor
tags above the captions and a thin progress line. English platforms
(TikTok, Reels, Shorts) get the podcast template: full 9:16 speaker crop,
one-to-three-word captions with the active word in the accent colour and a
short hook. Both show a left lower-third nameplate that never covers faces.

Packaging (title, translation, nameplates, tags, highlights) comes from one
text-model call per content and template, sends only the draft's subtitle
rows, validates every field, only names people seen in the subtitles or the
source listing, and falls back to source captions without blocking output.
Rendering stays per scene with bundled Noto Sans SC via libass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:40:35 +08:00
周小舟andClaude Opus 5.5 bdc74e37aa fix(import): ship yt-dlp's YouTube challenge solver and use any JS runtime
YouTube now requires solving a player JS challenge. The bundle installed bare
yt-dlp without yt-dlp-ejs, and yt-dlp only enables deno by default, so link
imports can fail with 'The page needs to be reloaded' on machines without
deno. Install yt-dlp[default] (same pinned version) and pass every JS runtime
found on PATH (deno, node, bun, quickjs) to all download paths.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:12:25 +08:00
周小舟andClaude Opus 5.5 4941753912 fix(studio): bound ffmpeg CPU in framing, preview and legacy export
A 2 h 22 min source pushed ffmpeg past 800% CPU during production: shot-cut
detection and frame sampling for speaker framing, and the full-source preview
transcode, decoded with every core at normal priority. Apply the shared
thread caps and low priority there and in legacy publish export.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 22:01:40 +08:00
周小舟andClaude Opus 5.5 52eaba55f7 fix(studio): accept sources longer than two hours
Screening rejected every source over two hours, which blocks most long
interviews and podcasts (a 2 h 22 min Dwarkesh episode failed immediately).
The subtitle route already chunks long transcripts and scales clip counts by
hour, so drop the upper bound. Frame-by-frame visual analysis keeps its
two-hour cost cap: longer sources are routed to subtitles instead of failing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 21:14:58 +08:00
周小舟andClaude Opus 5.5 ec42ba368f feat(studio): frame automatic vertical outputs on the speaker
Automatic vertical variants used a centre crop or a blurred full frame. Reuse
the existing shot-aware speaker framing: when the source has no burned
captions and faces are found, crop a true 9:16 window that follows the
speaker, switching to the full frame on shots with nobody to follow. One
detection pass is shared by every vertical platform. The on-demand OpenCV
install starts at import so it is usually ready by production; if it is not,
or detection fails, the output keeps the full frame instead of failing. Cards
say which framing was used.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 21:10:03 +08:00
周小舟andClaude Opus 5.5 5737a6db02 fix(render): dim and cheaply blur the vertical backdrop
The blurred backdrop behind fitted landscape footage enlarged burned captions
into a readable ghost line. Blur at 1/8 size (harder blur, far less CPU than a
full-resolution gblur) and dim it to about 35% so source text no longer reads
in the backdrop. Shared by legacy export and speaker framing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 21:03:30 +08:00
周小舟andClaude Opus 5.5 fda1a5751c fix(studio): keep rendering state alive and lower encode threads
Automatic generation, platform append and retry wrote a running analysis
without the process instance, so every read marked it as 'service restarted'
while renders were still progressing. Record the instance. Real-source runs
also peaked above 500% CPU with three threads per ffmpeg stage; cap at a
quarter of the cores.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:56:40 +08:00
周小舟andClaude Opus 5.5 3b07f0a5e3 feat(studio): copy share caption and ask for a rating after download
Completed variants get a Copy caption action (title plus 'Made with AutoClip'
and the repository link) for users to paste when they post. After a
successful download, at most once per project and once a week, a card asks
whether the video is usable (three levels) and invites use-case posts in
GitHub Discussions. Analytics carry only the share target and rating enums;
no text, titles or links leave the app.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:56:36 +08:00
周小舟andClaude Opus 5.5 84f0877136 feat(studio): skip our subtitles when the source already has them
Sources such as WIRED interviews ship with captions burned into the picture,
and automatic output stacked a second, often misrecognized, track on top.
Detect burned captions locally from twelve lower-frame samples (bright glyphs
with dark outlines whose shape changes over time, so static logos do not
count). When found, automatic variants keep subtitles off and vertical crop
layouts fall back to the full frame so the original lines are not cut.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:44:11 +08:00
周小舟andClaude Opus 5.5 2c56a85312 fix(platforms): enforce only real platform length limits
The registry treated 90 s as a hard cap for Douyin, TikTok, Reels and
Xiaohongshu, so legacy export truncated or rejected complete moments even
though those platforms accept much longer uploads. Split guidance
(recommended_max_duration_sec) from the hard limit (max_duration_sec). Only
YouTube Shorts keeps a hard cap, updated from 60 s to its current 180 s;
automatic Shorts variants over it are trimmed at the last subtitle sentence
end and the card says so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:38:56 +08:00
周小舟andClaude Opus 5.5 757bc1b565 fix(branding): keep the outro visible for its full second
Studio output uses a 1/16000 video time base while the outro was encoded with
ffmpeg defaults. Concat by stream copy kept each file's own timestamps, so the
outro's 30 frames collapsed into about 2 ms and players showed only a silent
tail. Encode the outro with the content's timescale, frame rate and audio
layout, and assert outro frames span about one second.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:27:25 +08:00
周小舟andClaude Opus 5.5 d31425ffc2 feat(i18n): translate fast-output copy and platform names
Replace Chinese placeholders for the quick-output flow in en, es, fr, ja, ko,
pt and ru. Platform names now come from one localized helper instead of the
backend's Chinese strategy labels, and generation skip reasons are translated.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:22:45 +08:00
周小舟andClaude Opus 5.5 25cb051010 fix(studio): render every automatic variant with bounded CPU
Real-source acceptance showed a 17-minute talk producing 8 variants and a
39-minute talk producing 15, but the per-project cap of three active exports
failed the whole generation on the fourth submit. Remove the cap so the source
alone decides how many outputs exist, and queue all variants on a dedicated
single-worker render executor. ffmpeg decode, filter and encode threads are
capped at a third of the cores and run at low priority so the machine stays
usable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 20:22:34 +08:00
周小舟andClaude Opus 5.5 06c945bb76 feat(publish): route publishing through completed output variants
Upload-Post and Bilibili accept an optional output_variant_id and reuse the
finished immutable MP4 instead of re-exporting. Landscape variants are refused
for vertical-only targets, Bilibili only accepts bilibili variants, and the
result card links each completed variant to the publish page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 19:55:57 +08:00
周小舟andClaude Opus 5.5 1dc9532788 feat(studio): preview completed output variants
Show completed immutable render jobs directly in the quick-output result cards while keeping queued and failed states actionable.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 19:26:45 +08:00
周小舟andClaude Opus 5.5 b213975929 feat(analytics): observe quick output outcomes
Track bounded platform generation outcomes and reuse events without content identifiers, and explain ineligible platform outputs on the results page.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 16:33:50 +08:00
周小舟andClaude Opus 5.5 3bcc7365e8 feat(studio): append and retry output variants
Reuse saved automatic drafts to produce additional platform versions without rerunning content analysis, and retry failed output variants independently from completed clips.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 16:08:11 +08:00
周小舟andClaude Opus 5.5 9d53fa60e3 feat(render): append optional AutoClip outro
Finalize automatic platform variants with a one-second Made with AutoClip outro while preserving legacy manual exports. Separate branded output caching and cover video/audio preservation with ffmpeg regressions.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 15:41:09 +08:00
周小舟andClaude Opus 5.5 a77cbdc96d feat(studio): make platform output the default flow
Let users choose destinations at import, start generation immediately, and view automatic output variants without passing through manual plan confirmation. Preserve Studio for deliberate edits.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 14:35:22 +08:00
周小舟andClaude Opus 5.5 6b196f745d feat(studio): auto-generate platform variants
Allow opted-in imports to produce isolated platform output variants and queue existing renders without the confirmation page. Preserve legacy confirmation behavior and aggregate variant terminal states.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 13:50:01 +08:00
周小舟andClaude Opus 5.5 c689954ebe feat(studio): add platform generation contract
Persist platform and branding choices in a versioned workspace contract while preserving the existing confirmation flow. Fix async analysis failures so error state survives exception cleanup.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 13:29:43 +08:00
周小舟andClaude Opus 5.5 ea66c329a3 feat: centralize platform output strategies
Add a versioned platform strategy registry, preserve legacy export presets, and expose strategy capabilities for the upcoming quick-generation flow. Record the project diagnosis and first work-package progress.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-30 13:13:45 +08:00
周小舟andCursor af718937f0 fix: 示例项目原片 source.mp4 加入版本库(原被 *.mp4 忽略,CI 与打包都缺文件)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 11:25:55 +08:00
周小舟andCursor 7aa4e9686f docs: CHANGELOG 记录设置重排、首次引导、示例项目与 Studio 取景改动
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 11:20:14 +08:00
周小舟andCursor 2da2399521 feat(studio): 预览播放条改为成片自身时间轴(0 到片段时长),不再暴露整段原片进度
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 11:10:15 +08:00
周小舟andCursor 65045ded5e fix(studio): 预览拖到成片范围之前自动吸回起点,脚注标出成片取用区间
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 11:02:43 +08:00
周小舟andCursor 83555a1030 feat(studio): 按镜头取景——先切镜再分类,无人物镜头改为完整画面
- 取景以镜头为单位:先用 ffmpeg 场景切换分数找硬切,轨迹点落在切点上,
  不再按固定 2.5s 采样 + 滞后消抖(切镜后最多滞后 5s 裁到空景的问题)
- 每个镜头分类:有人脸 → 对准说话人(长镜头内部仍可切换);
  多数采样无人脸(引用卡、PPT、空镜)→ fit:完整画面 + 模糊背景,不再切掉两侧文字
- 整条素材都没有人物时不改动用户选的构图
- CropPoint 增加 mode(crop|fit),渲染用 split/overlay+enable 在同一镜头链里按时间切换
- 编辑器:预览按镜头切换 contain/cover;取景行显示「镜头 n/N」,可把当前镜头单独设为
  对准人物 / 完整画面,滑块只改当前镜头,不再清掉整条轨迹
- 过滤过小的人脸;YuNet 检测器按帧尺寸复用,关掉 OpenCV 的 DNN 警告

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 10:57:11 +08:00
周小舟andCursor a19cac4392 feat: 重排 AI 设置、首次引导弹窗、示例项目与 Studio 编辑器重做
设置
- AI 模型页拆为 AI 服务 / 字幕转写 / 封面 / 高级四节,逻辑抽到 modelSettingsLogic + useModelSettings,
  ProviderFields / ModelPicker 独立组件;首次配置默认开启画面识别与 AI 封面(参考视频画面)
- 供应商分组「模型聚合站」改为「推荐」,保留赞助说明
- 首页首次进入弹出「连接 AI 服务」对话框(FirstRunSetup),未连接时导入被拦下并引导
- 修复对话框内下拉层级、Esc 误关闭

示例项目
- 内置 Sam Altman 访谈三段拼接原片 + 字幕 + 封面(backend/assets/example),
  一键创建已完成项目,携带来源链接与元数据;卡片 / 详情页标出示例与来源

Studio / 发布
- 编辑器右侧面板按 DESIGN.md 重做(DraftSettingsPanel):字幕样式改为全片四种带预览的样式,
  片头文字降为可选并用视觉缩略图选择;左侧播放器吸顶随滚动可见
- 竖屏裁切增加说话人跟随自动取景(YuNet 人脸 + 口部运动,按需安装 OpenCV 运行时),
  渲染支持逐段 crop 轨迹
- 导入确认页去掉重复的分析方式提问,控件统一 Row/Segmented;发布页文案去术语化,
  封面入口补齐并默认自动生成

其他
- 后端 ai-model-settings 文档模型、云端转写、模型目录等配套服务与测试
- 8 种语言文案同步

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-30 10:26:39 +08:00
周小舟andClaude Opus 5.5 9fc99136ea feat(settings): 模型里标注多模态 / 仅文字,视觉和封面生图跟随同一个服务商
- 模型下拉每项标「多模态 / 仅文字」(model_catalog.supports_vision)。能不能看画面由所选
  模型决定,不再单独配置视觉模型、不再做图片理解测试;仅文字模型自动只用字幕。
- 分析方式默认「智能选择」,选智能 / 视觉即允许需要时发送抽样画面,卡片上写明花费;
  去掉单独的「导入时视觉初筛」开关。已保存过偏好的用户保持原选择。
- 封面生图并入「模型」:默认同一服务商和 Key(通义→通义万相、Seed→Seedream、
  OpenAI→gpt-image、Infistar 可选 gpt-image / dall-e / Seedream / Flux / Imagen / 万相),
  可直接输入模型 ID;服务商没接入生图时可「用其他服务生图」就地填写,不再是死路。
  旧版单独配置的视觉 / 生图保持可用,并给出切换为跟随模型的入口。
- 设置页加「‹ 项目」返回。通义常用列表加入 qwen-vl-max / qwen-vl-plus。

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 18:43:02 +08:00
周小舟andClaude Opus 5.5 c44da9c36f feat(settings): 分析方式、模型、视觉理解、转写、封面生图合成「AI 分析」一页
设置导航从 8 项减到 4 项(AI 分析 / 发布 / 应用 / 反馈),AI 分析只有一个保存。旧的
?section=model|analysis|vision|speech|cover 链接仍然有效,会滚到对应小节。

- 视觉理解默认复用文本模型(vision mode=text_model):不用再配一遍地址、Key、模型;
  文本模型不能看图时再「另选」。已保存过独立视觉配置、或 .env 配了视觉地址的老用户保持原样。
  「测试图片理解」通过后记下端点,换了模型或地址即回到未验证。
- 分析方式改成三张卡片,逐项写明效果、花费和适合的素材;视觉分析明确说明图片 token
  通常是字幕的数倍。付费视觉初筛仍需用户单独打开。
- 封面:去掉「校对模型」,改用视觉理解那一套读图校对标题(手填过的沿用);生图提供商
  与文本模型同一家(通义万相/Seedream/OpenAI 兼容含 Infistar)时自动复用 Key,且不写入封面配置。
- 转写收进「无字幕时的转写」折叠块;切片参数收进折叠块。

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-29 18:02:27 +08:00