- Detection also records where the burned captions sit. An English version
of a source with non-English captions blurs that band before any layout,
then is framed and captioned like a caption-free source.
- AI covers switched on by 1.4's first setup read as off after upgrading
(1.5 would bill one per output in the background); the chosen model stays
and one click turns them back on. Every 1.5 save marks the choice explicit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- cloud ASR retries a rate-limited or failing chunk with backoff and stops
sending the rest once one fails for good; a bad concurrency env var is safe
- a stalled hardware encoder falls back to libx264; a bad input no longer
disables hardware encoding
- a failed outro delivers the video without it; Windows font paths and
apostrophes in paths work in the outro
- no AI cover without permission to send the frame (no invented face
under the guest's nameplate)
- handoff doc: review results, open decisions, deferred findings
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review found several ways Chinese text still reached TikTok/Reels/Shorts/YouTube:
- packaging kept Chinese titles, captions and nameplate roles from the model;
Chinese captions are now retried, the rest dropped or filtered
- post copy accepted Chinese titles; they are retried, never posted
- landscape YouTube burned the untranslated source captions and Chinese hook;
landscape versions now caption in their audience's language
- appended and on-demand versions skipped the English title and post copy
- publishing fell back to the Chinese clip title; it now asks for one
- the kit's file names and the share credit followed the app language
Also: a version interrupted mid-render becomes retryable after a restart,
and a DashScope relay address from 1.4 is no longer rewritten.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A backup clip being framed and packaged is 'preparing', invisible to the
render dispatcher, so it can no longer ship raw; claiming it is atomic.
A failed preparation settles the generation, retry prepares it again, and
one interrupted by a restart becomes retryable.
- The results page keeps polling while outputs are preparing or rendering.
- AI cover redesign: no false success on timeout, nothing after unmount.
- Saved post copy shows at once; title lines keep their place while typing.
- Generated-image downloads refuse local/private hosts (each redirect too)
and stop at 40 MB.
- Docs: cloud ASR numbers, CHANGELOG entries for this round.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The output card's 16:9 thumbnail crops a 9:16 cover to the face, so the kit
now shows the cover at its own ratio. Automatic projects no longer list raw
clip cards after their outputs: fast output never renders raw clip files, so
their previews 404'd and Download/Export did nothing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
GPT Image 2.x is asked for the platform ratio (a 2:3 image padded to 9:16 left
bands above and below); a relay that rejects it gets the legacy 2:3 size. When a
model still returns another ratio, the gap continues the image's own edge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
First-run settings no longer switch AI covers on or pick an image model:
covers are designed locally (free, instant) until the user chooses AI and an
image service and model themselves. The cover source reads 'auto design
(free)' / 'AI generated' with a neutral hint (no vendor named). An
OpenAI-compatible connection pointed at fal.run is recognised as fal.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A settings file written by a newer version (e.g. image_api 'fal') made an
older backend reject the whole file, failing every task at screening.
Unknown image APIs now fall back to auto.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Image connections can use fal (image_api 'fal', https://fal.run): the
model path gets /edit with the reference frame, else /text-to-image, and
the platform's exact size is requested.
- AI covers never draw names (models invent them: 'WILDER', 'Terence');
our own nameplate is added only when the vision model confirmed the guest.
- Text must stay inside a safe area; the result is padded, never cropped,
so a headline can no longer lose its edges.
Compared on six real clips (same frame, same prompt): GPT Image 2.5 Flare
via fal got 6/6 headlines right in 22-31 s; qwen-image-3.0 5/6 with
overlaps; Seedream 5.0 had the most misspellings and refused two
celebrity frames.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Burned caption detection ignores the outer 15 % of the band (channel logos
fooled it on Hot Ones) and the vision model, when set up, has the last
word: no dialogue captions in the picture means none (slides, products and
logos do not count). 15 real sources now classify correctly.
- English platforms never show Chinese: versions take the English packaging
or post title, fallback titles drop CJK lines, post-copy fallback leaves an
English title empty rather than Chinese. Chinese platforms never fall back
to English-only captions.
- When packaging fails for a foreign audience, a plain line-by-line
translation still gives captions in the audience's language.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Keep the same command blocks and the single Trendshift badge so i18n-sync can pass, and retell the new product pitch in each language instead of leaving the old feature-grid pages behind.
Co-authored-by: Cursor <cursoragent@cursor.com>
- Cover frames: candidates are the sharpest face frame of each long shot;
the vision model picks the guest's best one (never the host). Without a
vision model the sharpest face is used and no nameplate is drawn, since
colour or screen-time rules cannot tell a host's camera from a guest's.
Tight crops are lightly sharpened after upscaling.
- AI covers: after the designed cover, an image model (when set up)
generates an editorial magazine-style cover from that frame, in the
platform's ratio (qwen-image sizes now cover 3:4/1:1/4:3) and the clip's
palette, in the audience's language; the vision model checks the headline
and a wrong one is retried once, else the designed cover stays. Runs on
its own pool so renders never wait; 'AI 重新设计封面' uses the same path.
- Image requests wait up to 180 s (high-quality models are slow).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Move the trending badge into the hero, collapse speed/cost into three rows, and keep sponsor logos small so the page reads as a product pitch rather than a dump of internals.
Co-authored-by: Cursor <cursoragent@cursor.com>
Two-line English titles shrank to unreadable sizes; they now rewrap into up
to four lines at a readable size, keeping the accent words coloured, and the
portrait title band grows with the title so it never covers the face.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
qwen-audio-3.1-asr-flash on a 2 h 52 m Chinese talk: 215 s with 8 uploads
(local Whisper base: 23 min), 191 s with 16, so 8 is the default.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The DashScope connection now takes an optional endpoint address (host,
/api/v1 or /compatible-mode/v1 are all accepted): speech recognition uses
its /api/v1 path and text models its OpenAI-compatible path. Also makes a
timeline test independent of chunk call order now that chunks run in
parallel.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three-minute chunks were uploaded one after another (~60 round trips for a
three-hour talk). Chunks are cut first, then sent on a pool of 4
(AUTOCLIP_ASR_CONCURRENCY), merged by exact sample offsets; any failed chunk
still publishes no partial subtitle.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- 17 fixed real inputs (owner's picks: long interviews, short sources,
keynotes with screen demos, Chinese sources with burned captions, auto-
caption-only and once-360p sources) in benchmarks/fast_output/cases.json.
- scripts/fast_output_benchmark.py imports each through the product API and
reports per-stage time, model calls, tokens and cost, clips and renders,
source resolution, subtitle source, burned captions and packaging
fallbacks; --baseline prints deltas.
- Stage wall times (download, transcribe, clip finder, boundaries, framing,
packaging, post copy, render) are recorded next to token usage.
- Changelog: publish kit, outro, caption strategy, low-resolution fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each output card shows its designed cover as the video poster and the
platform's post copy (title, description, hashtags) with copy, inline edit,
export publish kit (native save on desktop) and AI cover redesign. The
publish page starts from the kit: title, description with hashtags, and the
cover from the matching vertical/landscape slot instead of generating one.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
DESIGN.md records the homepage split hero, band rhythm, editor/chat/
one-click comparison, sticky story and the stage-time ribbon.
Co-authored-by: Cursor <cursoragent@cursor.com>
- Burned caption detection missed thin, outline-free captions (Bilibili
interviews such as TIM x Luo Yonghao): the bottom band is now analysed at
960x192 with a matching threshold; all seven test sources classify right.
- The burned captions' language is read by the vision model when set up,
else taken from the video title's language (captions are written for the
channel's audience), else the speech language.
- Chinese captions in the picture: none added for Douyin, English for TikTok;
English captions on a Japanese talk: none for TikTok, Chinese for Douyin.
No translation is requested when no caption will be shown.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pre-rendered vertical and horizontal outro (silver cut, logo, wordmark,
chime; source and OFL font in design/outro-v5) is conformed once per output
spec (size, fps, time base, audio layout), cached, and joined by stream
copy; other aspect ratios are padded with its background. The old text card
remains as the fallback when the asset is missing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
YouTube's SABR experiment can leave only the 360p format reachable for the
default client; the MrBeast demo was cropped from a 640x360, 100 kbps file
and looked blurry. After download, if the file is below 720p while the
listing offers more, retry with the web_safari / web_embedded / tv clients
and keep the sharpest file. The final height is kept in source_meta.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
README links the case library and describes the per-platform publish
kit; a Show and tell discussion template collects community clips with
source, platforms, setup and display consent.
Co-authored-by: Cursor <cursoragent@cursor.com>