343 Commits
Author SHA1 Message Date
Kris K e22be4bb3b Merge pull request #253 from zhouxiaoka/codex/readme-video-demos
docs: README 展示 16 条可点播的真实视频案例
2026-10-01 23:46:33 +08:00
周小舟 51c2f09d81 docs: align video thumbnails across cards 2026-10-01 23:45:04 +08:00
周小舟 a8cac39ef0 docs: add sixteen playable video demos to translated READMEs 2026-10-01 23:42:47 +08:00
Kris K c83fca0f7b Merge pull request #248 from zhouxiaoka/cursor/readme-v2
docs: README 同步 1.5.0 正式版与一键出片
2026-10-01 23:30:50 +08:00
周小舟 ce35f98b90 docs: preserve the existing sponsor sections verbatim 2026-10-01 23:28:54 +08:00
周小舟 7d33f3c9d4 docs: review current behavior and remove stale README claims 2026-10-01 23:21:45 +08:00
周小舟 8f9416f21e docs: align eight READMEs with the 1.5.0 release 2026-10-01 23:14:47 +08:00
Kris K d3e26ee19b Merge pull request #252 from zhouxiaoka/worktree-project-diagnosis
docs: close the 1.5.0 release and issue handoff
2026-10-01 22:57:55 +08:00
周小舟 52b20fac68 docs: record 1.5.0 latest release and issue closure 2026-10-01 22:49:56 +08:00
Kris K 9895842688 Merge pull request #251 from zhouxiaoka/worktree-project-diagnosis
docs: record 1.5 tag and final acceptance handoff
2026-10-01 22:15:27 +08:00
周小舟 9a130ee61a docs(1.5): record merged tag and final long-video acceptance 2026-10-01 22:07:56 +08:00
Kris K e940f5aef4 Merge pull request #250 from zhouxiaoka/worktree-project-diagnosis
feat: prepare 1.5 quick output for Desktop, CLI and MCP
v1.5.0
2026-10-01 21:55:16 +08:00
周小舟 1b02347c3e fix(studio): rank eligible clips per platform and settle empty queues 2026-10-01 21:47:54 +08:00
周小舟 0fe48c051e fix: settle observed AI cover results before replacement 2026-10-01 21:34:55 +08:00
周小舟 0147056d21 fix: complete 1.5 delivery telemetry and publish CLI MCP assets 2026-10-01 21:28:48 +08:00
周小舟 1d433ea1f5 docs(1.5): record cross-platform installed wheel acceptance 2026-10-01 21:07:26 +08:00
周小舟 ee73e594ff fix(render): select bundled CJK family across font providers 2026-10-01 20:58:33 +08:00
周小舟 e3e57236a6 fix(release): bundle face detector and verify installed portrait output 2026-10-01 20:50:46 +08:00
周小舟 e23dc999f2 fix(cli): apply Windows connection cleanup at headless startup 2026-10-01 20:25:07 +08:00
周小舟 5c0b82e0ee fix(ci): isolate runtime tests and sync quick-output docs 2026-10-01 20:07:59 +08:00
周小舟 d3e3cb1a84 fix(1.5): complete output acceptance and align CLI/MCP 2026-10-01 20:00:45 +08:00
周小舟andClaude Opus 5.5 723eefbd52 feat(studio): blur burned Chinese captions out of English versions; 1.4 AI covers off
- Detection also records where the burned captions sit. An English version
  of a source with non-English captions blurs that band before any layout,
  then is framed and captioned like a caption-free source.
- AI covers switched on by 1.4's first setup read as off after upgrading
  (1.5 would bill one per output in the background); the chosen model stays
  and one click turns them back on. Every 1.5 save marks the choice explicit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 19:08:26 +08:00
周小舟andClaude Opus 5.5 ecd0dda509 fix: failures degrade instead of losing work; record the review
- cloud ASR retries a rate-limited or failing chunk with backoff and stops
  sending the rest once one fails for good; a bad concurrency env var is safe
- a stalled hardware encoder falls back to libx264; a bad input no longer
  disables hardware encoding
- a failed outro delivers the video without it; Windows font paths and
  apostrophes in paths work in the outro
- no AI cover without permission to send the frame (no invented face
  under the guest's nameplate)
- handoff doc: review results, open decisions, deferred findings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:58:06 +08:00
周小舟andClaude Opus 5.5 6931bc1c93 fix(studio): English platforms stay English on every path
Review found several ways Chinese text still reached TikTok/Reels/Shorts/YouTube:
- packaging kept Chinese titles, captions and nameplate roles from the model;
  Chinese captions are now retried, the rest dropped or filtered
- post copy accepted Chinese titles; they are retried, never posted
- landscape YouTube burned the untranslated source captions and Chinese hook;
  landscape versions now caption in their audience's language
- appended and on-demand versions skipped the English title and post copy
- publishing fell back to the Chinese clip title; it now asks for one
- the kit's file names and the share credit followed the app language
Also: a version interrupted mid-render becomes retryable after a restart,
and a DashScope relay address from 1.4 is no longer rewritten.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:51:41 +08:00
周小舟andClaude Opus 5.5 3b8088c8a4 fix(studio): backups render only once prepared; review fixes
- A backup clip being framed and packaged is 'preparing', invisible to the
  render dispatcher, so it can no longer ship raw; claiming it is atomic.
  A failed preparation settles the generation, retry prepares it again, and
  one interrupted by a restart becomes retryable.
- The results page keeps polling while outputs are preparing or rendering.
- AI cover redesign: no false success on timeout, nothing after unmount.
- Saved post copy shows at once; title lines keep their place while typing.
- Generated-image downloads refuse local/private hosts (each redirect too)
  and stop at 40 MB.
- Docs: cloud ASR numbers, CHANGELOG entries for this round.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:37:01 +08:00
周小舟andClaude Opus 5.5 a0108d889c fix(studio): show the whole cover in the kit; drop dead raw-clip cards
The output card's 16:9 thumbnail crops a 9:16 cover to the face, so the kit
now shows the cover at its own ratio. Automatic projects no longer list raw
clip cards after their outputs: fast output never renders raw clip files, so
their previews 404'd and Download/Export did nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:25:25 +08:00
周小舟andClaude Opus 5.5 0fc14a61fb fix(cover): AI covers fill the platform frame, no flat bars
GPT Image 2.x is asked for the platform ratio (a 2:3 image padded to 9:16 left
bands above and below); a relay that rejects it gets the legacy 2:3 size. When a
model still returns another ratio, the gap continues the image's own edge.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:12:41 +08:00
周小舟andClaude Opus 5.5 1317b5d4b8 docs(handoff): demos rebuilt; note 9:16 AI cover padding
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:06:09 +08:00
周小舟andClaude Opus 5.5 ed14a34953 docs: 1.5.0 fast-output handoff, kept current as work lands
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 18:01:05 +08:00
周小舟andClaude Opus 5.5 7814dbd4ee feat(cover): designed covers by default; AI covers only when the user picks a model
First-run settings no longer switch AI covers on or pick an image model:
covers are designed locally (free, instant) until the user chooses AI and an
image service and model themselves. The cover source reads 'auto design
(free)' / 'AI generated' with a neutral hint (no vendor named). An
OpenAI-compatible connection pointed at fal.run is recognised as fal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:35:00 +08:00
周小舟andClaude Opus 5.5 4d11ece037 fix(settings): an unknown image API never invalidates the settings file
A settings file written by a newer version (e.g. image_api 'fal') made an
older backend reject the whole file, failing every task at screening.
Unknown image APIs now fall back to auto.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:23:03 +08:00
周小舟andClaude Opus 5.5 41fa605655 feat(cover): fal image connections (GPT Image 2.5) and correct-by-construction text
- Image connections can use fal (image_api 'fal', https://fal.run): the
  model path gets /edit with the reference frame, else /text-to-image, and
  the platform's exact size is requested.
- AI covers never draw names (models invent them: 'WILDER', 'Terence');
  our own nameplate is added only when the vision model confirmed the guest.
- Text must stay inside a safe area; the result is padded, never cropped,
  so a headline can no longer lose its edges.

Compared on six real clips (same frame, same prompt): GPT Image 2.5 Flare
via fal got 6/6 headlines right in 22-31 s; qwen-image-3.0 5/6 with
overlaps; Seedream 5.0 had the most misspellings and refused two
celebrity frames.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 17:18:41 +08:00
周小舟andClaude Opus 5.5 e7fb0c18c5 fix(packaging): one language per platform; no false burned captions
- Burned caption detection ignores the outer 15 % of the band (channel logos
  fooled it on Hot Ones) and the vision model, when set up, has the last
  word: no dialogue captions in the picture means none (slides, products and
  logos do not count). 15 real sources now classify correctly.
- English platforms never show Chinese: versions take the English packaging
  or post title, fallback titles drop CJK lines, post-copy fallback leaves an
  English title empty rather than Chinese. Chinese platforms never fall back
  to English-only captions.
- When packaging fails for a foreign audience, a plain line-by-line
  translation still gives captions in the audience's language.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:51:31 +08:00
周小舟andCursor 95153d1fda docs: sync the seven README translations to the one-link, one-click rewrite
Keep the same command blocks and the single Trendshift badge so i18n-sync can pass, and retell the new product pitch in each language instead of leaving the old feature-grid pages behind.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 16:20:14 +08:00
周小舟andClaude Opus 5.5 47b6f7945c feat(cover): the vision model picks the guest's frame; AI covers upgrade the design
- Cover frames: candidates are the sharpest face frame of each long shot;
  the vision model picks the guest's best one (never the host). Without a
  vision model the sharpest face is used and no nameplate is drawn, since
  colour or screen-time rules cannot tell a host's camera from a guest's.
  Tight crops are lightly sharpened after upscaling.
- AI covers: after the designed cover, an image model (when set up)
  generates an editorial magazine-style cover from that frame, in the
  platform's ratio (qwen-image sizes now cover 3:4/1:1/4:3) and the clip's
  palette, in the audience's language; the vision model checks the headline
  and a wrong one is retried once, else the designed cover stays. Runs on
  its own pool so renders never wait; 'AI 重新设计封面' uses the same path.
- Image requests wait up to 180 s (high-quality models are slow).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 16:18:22 +08:00
周小舟andCursor 3a69c639a1 docs: tighten README layout so the first screen and proof stay scannable
Move the trending badge into the hero, collapse speed/cost into three rows, and keep sponsor logos small so the page reads as a product pitch rather than a dump of internals.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 16:11:21 +08:00
周小舟andClaude Opus 5.5 882a59d96a fix(cover): long English titles rewrap larger and push the photo down
Two-line English titles shrank to unreadable sizes; they now rewrap into up
to four lines at a readable size, keeping the accent words coloured, and the
portrait title band grows with the title so it never covers the face.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:43:17 +08:00
周小舟andClaude Opus 5.5 82f82e035e perf(asr): 8 parallel cloud transcription uploads; benchmark keeps case order
qwen-audio-3.1-asr-flash on a 2 h 52 m Chinese talk: 215 s with 8 uploads
(local Whisper base: 23 min), 191 s with 16, so 8 is the default.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:30:21 +08:00
周小舟andClaude Opus 5.5 a3066604bb feat(settings): dedicated Bailian endpoints for DashScope connections
The DashScope connection now takes an optional endpoint address (host,
/api/v1 or /compatible-mode/v1 are all accepted): speech recognition uses
its /api/v1 path and text models its OpenAI-compatible path. Also makes a
timeline test independent of chunk call order now that chunks run in
parallel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:13:44 +08:00
周小舟andClaude Opus 5.5 6786f2eaf1 perf(asr): upload cloud transcription chunks side by side
Three-minute chunks were uploaded one after another (~60 round trips for a
three-hour talk). Chunks are cut first, then sent on a pool of 4
(AUTOCLIP_ASR_CONCURRENCY), merged by exact sample offsets; any failed chunk
still publishes no partial subtitle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 15:04:32 +08:00
周小舟andCursor 7fd9873a6b docs: lead README with 10+ clips per link; drop competitor names and roadmap hints
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:59:43 +08:00
周小舟andCursor d2cf65197e docs(design): serif web headlines and the risograph art layer
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:57:55 +08:00
周小舟andClaude Opus 5.5 53ececdcb2 test(bench): fast-output regression set with per-stage timing
- 17 fixed real inputs (owner's picks: long interviews, short sources,
  keynotes with screen demos, Chinese sources with burned captions, auto-
  caption-only and once-360p sources) in benchmarks/fast_output/cases.json.
- scripts/fast_output_benchmark.py imports each through the product API and
  reports per-stage time, model calls, tokens and cost, clips and renders,
  source resolution, subtitle source, burned captions and packaging
  fallbacks; --baseline prints deltas.
- Stage wall times (download, transcribe, clip finder, boundaries, framing,
  packaging, post copy, render) are recorded next to token usage.
- Changelog: publish kit, outro, caption strategy, low-resolution fix.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:53:19 +08:00
周小舟andClaude Opus 5.5 35ece30454 feat(studio): output cards carry their publish kit
Each output card shows its designed cover as the video poster and the
platform's post copy (title, description, hashtags) with copy, inline edit,
export publish kit (native save on desktop) and AI cover redesign. The
publish page starts from the kit: title, description with hashtags, and the
cover from the matching vertical/landscape slot instead of generating one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:46:48 +08:00
周小舟andCursor 0c2b275cea docs: README leads with one link, one click; adds the TIM run and stage timing
DESIGN.md records the homepage split hero, band rhythm, editor/chat/
one-click comparison, sticky story and the stage-time ribbon.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:44:01 +08:00
周小舟andClaude Opus 5.5 5c91702afc fix(packaging): caption burned-in sources only for audiences that cannot read them
- Burned caption detection missed thin, outline-free captions (Bilibili
  interviews such as TIM x Luo Yonghao): the bottom band is now analysed at
  960x192 with a matching threshold; all seven test sources classify right.
- The burned captions' language is read by the vision model when set up,
  else taken from the video title's language (captions are written for the
  channel's audience), else the speech language.
- Chinese captions in the picture: none added for Douyin, English for TikTok;
  English captions on a Japanese talk: none for TikTok, Chinese for Douyin.
  No translation is requested when no caption will be shown.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:42:20 +08:00
周小舟andClaude Opus 5.5 1dc8a09c27 feat(branding): the designed 1.8 s outro animation replaces the text card
Pre-rendered vertical and horizontal outro (silver cut, logo, wordmark,
chime; source and OFL font in design/outro-v5) is conformed once per output
spec (size, fps, time base, audio layout), cached, and joined by stream
copy; other aspect ratios are padded with its background. The old text card
remains as the fallback when the asset is missing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:35:01 +08:00
周小舟andClaude Opus 5.5 09a5ffabc8 fix(studio): re-download YouTube sources that arrive at 360p
YouTube's SABR experiment can leave only the 360p format reachable for the
default client; the MrBeast demo was cropped from a 640x360, 100 kbps file
and looked blurry. After download, if the file is below 720p while the
listing offers more, retry with the web_safari / web_embedded / tv clients
and keep the sharpest file. The final height is kept in source_meta.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:28:32 +08:00
周小舟andCursor 88189f2347 docs: point README to the case library and add a showcase submission template
README links the case library and describes the per-platform publish
kit; a Show and tell discussion template collects community clips with
source, platforms, setup and display consent.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-10-01 14:19:19 +08:00
周小舟andClaude Opus 5.5 05b33bfad4 docs: measured cost and time per video link
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-01 14:12:51 +08:00