GPT Image 2.x is asked for the platform ratio (a 2:3 image padded to 9:16 left
bands above and below); a relay that rejects it gets the legacy 2:3 size. When a
model still returns another ratio, the gap continues the image's own edge.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
First-run settings no longer switch AI covers on or pick an image model:
covers are designed locally (free, instant) until the user chooses AI and an
image service and model themselves. The cover source reads 'auto design
(free)' / 'AI generated' with a neutral hint (no vendor named). An
OpenAI-compatible connection pointed at fal.run is recognised as fal.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A settings file written by a newer version (e.g. image_api 'fal') made an
older backend reject the whole file, failing every task at screening.
Unknown image APIs now fall back to auto.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Image connections can use fal (image_api 'fal', https://fal.run): the
model path gets /edit with the reference frame, else /text-to-image, and
the platform's exact size is requested.
- AI covers never draw names (models invent them: 'WILDER', 'Terence');
our own nameplate is added only when the vision model confirmed the guest.
- Text must stay inside a safe area; the result is padded, never cropped,
so a headline can no longer lose its edges.
Compared on six real clips (same frame, same prompt): GPT Image 2.5 Flare
via fal got 6/6 headlines right in 22-31 s; qwen-image-3.0 5/6 with
overlaps; Seedream 5.0 had the most misspellings and refused two
celebrity frames.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Burned caption detection ignores the outer 15 % of the band (channel logos
fooled it on Hot Ones) and the vision model, when set up, has the last
word: no dialogue captions in the picture means none (slides, products and
logos do not count). 15 real sources now classify correctly.
- English platforms never show Chinese: versions take the English packaging
or post title, fallback titles drop CJK lines, post-copy fallback leaves an
English title empty rather than Chinese. Chinese platforms never fall back
to English-only captions.
- When packaging fails for a foreign audience, a plain line-by-line
translation still gives captions in the audience's language.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Cover frames: candidates are the sharpest face frame of each long shot;
the vision model picks the guest's best one (never the host). Without a
vision model the sharpest face is used and no nameplate is drawn, since
colour or screen-time rules cannot tell a host's camera from a guest's.
Tight crops are lightly sharpened after upscaling.
- AI covers: after the designed cover, an image model (when set up)
generates an editorial magazine-style cover from that frame, in the
platform's ratio (qwen-image sizes now cover 3:4/1:1/4:3) and the clip's
palette, in the audience's language; the vision model checks the headline
and a wrong one is retried once, else the designed cover stays. Runs on
its own pool so renders never wait; 'AI 重新设计封面' uses the same path.
- Image requests wait up to 180 s (high-quality models are slow).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Two-line English titles shrank to unreadable sizes; they now rewrap into up
to four lines at a readable size, keeping the accent words coloured, and the
portrait title band grows with the title so it never covers the face.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
qwen-audio-3.1-asr-flash on a 2 h 52 m Chinese talk: 215 s with 8 uploads
(local Whisper base: 23 min), 191 s with 16, so 8 is the default.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The DashScope connection now takes an optional endpoint address (host,
/api/v1 or /compatible-mode/v1 are all accepted): speech recognition uses
its /api/v1 path and text models its OpenAI-compatible path. Also makes a
timeline test independent of chunk call order now that chunks run in
parallel.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Three-minute chunks were uploaded one after another (~60 round trips for a
three-hour talk). Chunks are cut first, then sent on a pool of 4
(AUTOCLIP_ASR_CONCURRENCY), merged by exact sample offsets; any failed chunk
still publishes no partial subtitle.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- 17 fixed real inputs (owner's picks: long interviews, short sources,
keynotes with screen demos, Chinese sources with burned captions, auto-
caption-only and once-360p sources) in benchmarks/fast_output/cases.json.
- scripts/fast_output_benchmark.py imports each through the product API and
reports per-stage time, model calls, tokens and cost, clips and renders,
source resolution, subtitle source, burned captions and packaging
fallbacks; --baseline prints deltas.
- Stage wall times (download, transcribe, clip finder, boundaries, framing,
packaging, post copy, render) are recorded next to token usage.
- Changelog: publish kit, outro, caption strategy, low-resolution fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Each output card shows its designed cover as the video poster and the
platform's post copy (title, description, hashtags) with copy, inline edit,
export publish kit (native save on desktop) and AI cover redesign. The
publish page starts from the kit: title, description with hashtags, and the
cover from the matching vertical/landscape slot instead of generating one.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Burned caption detection missed thin, outline-free captions (Bilibili
interviews such as TIM x Luo Yonghao): the bottom band is now analysed at
960x192 with a matching threshold; all seven test sources classify right.
- The burned captions' language is read by the vision model when set up,
else taken from the video title's language (captions are written for the
channel's audience), else the speech language.
- Chinese captions in the picture: none added for Douyin, English for TikTok;
English captions on a Japanese talk: none for TikTok, Chinese for Douyin.
No translation is requested when no caption will be shown.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Pre-rendered vertical and horizontal outro (silver cut, logo, wordmark,
chime; source and OFL font in design/outro-v5) is conformed once per output
spec (size, fps, time base, audio layout), cached, and joined by stream
copy; other aspect ratios are padded with its background. The old text card
remains as the fallback when the asset is missing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
YouTube's SABR experiment can leave only the 360p format reachable for the
default client; the MrBeast demo was cropped from a 640x360, 100 kbps file
and looked blurry. After download, if the file is below 720p while the
listing offers more, retry with the web_safari / web_embedded / tv clients
and keep the sharpest file. The final height is kept in source_meta.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Publishing an output variant uses its post copy when the user leaves fields
empty: title, description and hashtags for Upload-Post; title, description
and up to ten tags for Bilibili (instead of the fixed tag).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Post copy: one model call per clip writes title, description and tags for
each selected platform (Douyin hook, Xiaohongshu note, Bilibili, English
captions for TikTok/Reels/Shorts/YouTube); platform limits are enforced in
code; failures fall back to the clip title. Backup clips get copy when
produced; users can edit it (PUT .../post).
- Cover: designed locally after each render from the speaker's frame (track
converted to the speaker centre), the video's own title lines and palette,
the guest's nameplate; 9:16, Xiaohongshu 3:4, Bilibili 16:10, YouTube 16:9.
It fills the publish cover slot unless an AI cover exists, which wins.
- Kit: GET .../kit zips the video, cover and copy; GET .../cover serves it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Packaging waited ~20 s per clip on the model, one after another (25 clips
~8 min). Distinct clip/template packages of a batch now run on the shared
pool before variants are built; platforms of the same template still reuse
one package.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
When a video has uploaded subtitles in its own language (e.g. Dwarkesh
Patel's YouTube interviews), fetch them as the transcript and skip local
Whisper (~9 min per 2 h of audio). Automatic captions are never used: no
punctuation and rolling repeats break sentence cut points. Any problem or a
stub file keeps speech recognition.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A two-hour talk yields 30+ clips; rendering all of them up front was most of
the run while people publish a handful. Automatic output now frames,
packages and renders the 10 highest-scored clips; the others are listed as
backup clips (status on_demand, no model or render cost) and are framed,
packaged and rendered when the user clicks 'Generate this one'. Adding a
platform keeps each moment's choice. Generation completes once the
automatic versions finish; analytics count on-demand versions separately.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
VideoToolbox on macOS, NVENC/QSV/AMF elsewhere, proven once with a tiny
test encode; bitrate scales with the frame (8 Mbps at 1080x1920@30). A
failed hardware encode retries the clip with libx264 and sticks to it.
AUTOCLIP_VIDEO_ENCODER=x264 forces software. Same wall time on this Mac but
4.5x less CPU per encode (3.0 s vs 13.4 s CPU for 30 s of 1080x1920).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Fast output no longer runs outline -> timeline -> scoring -> titles over
5000-character chunks. One call per ~60-minute window (parallel, with
overlap) reads compact id|time|text rows and returns finished clips with
rows, Chinese title, reason and score; picks are validated, deduplicated
across windows, snapped by refine_timeline and selected by select_clips, in
the legacy titled-clip shape. Short-video lengths (60-300 s, ideally
90-180 s) replace the podcast-era 90 s floor that padded answers into the
next question. Any failure falls back to the legacy steps.
MrBeast 2h06m: ~35 min / 128 calls / 500K tokens -> 38 s / 4 calls / 93K.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Outline, timeline, scoring and title steps called the model once per
transcript chunk, one after another: a 2 h video spent ~30 min waiting on 85
sequential requests. Independent chunk calls now run on a pool of 4 (local
models stay sequential; AUTOCLIP_LLM_CONCURRENCY overrides), results keep
chunk order, and token tracking follows each worker.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Fast output runs the content pipeline without clustering (one unused model
call) and without re-encoding every clip and collection (studio renders
from the source); clip rows still sync from step 4 for the candidate list.
- Packaging no longer asks the model to restate the transcript when no
translation is needed; its rewrite was discarded for same-language output.
- One face-detection pass per clip serves both the 4:3 interview window and
the 9:16 podcast crop.
- LLM inputs are compact JSON (indentation cost ~10 tokens per subtitle row).
- Every text and vision call records tokens per project and stage into
metadata/llm_usage.jsonl (DashScope native usage is captured now), so the
cost of one video can be measured.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Retry the packaging call once with the rejection reason before falling back:
short Japanese ASR rows made the segment contract fail intermittently and
the fallback dropped the Chinese captions entirely.
- Decode HTML entities and strip markdown from titles and captions; reject
kana titles on Chinese platforms; fallback titles keep whole clauses.
- Clips of one batch avoid the previous palettes when the mood allows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The packaging model names the clip's mood (calm/serious/bold/warm/playful);
the mood decides which of seven content palettes and which template styles
fit, and a seed from the clip picks among them, so clips vary while one clip
keeps its look on every platform. Content colours are no longer tied to the
product UI accent; light accents get dark text on filled pills.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Subtitle rows carry no word timing, so cuts estimated inside a row leaked a
few hundred ms of the next sentence. Snap both cuts to ffmpeg-detected pauses:
the pause must fit how fast the words around the cut can be spoken; when the
speaker runs straight on, finish the next sentence that ends on a breath
(never past the next question); a long pause just after extends the clip to
finish the passage. Starts may reach back 20s to a question's opening.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Real Dwarkesh rows often end one sentence and start the next
("...responsible. It seems like"), and punctuated transcripts treated
0.6 s pauses as sentence ends, so clips still ended on openers. Cut inside a
row at its last sentence end, start at the first full sentence of a row
that begins on a tail, use pauses only for unpunctuated transcripts, treat
Japanese polite sentence-final forms as sentence ends, and end before a
trailing question followed by only the opening of its answer. Translated
segments may span up to 12 short ASR rows so long answers keep packaging.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Clips ended mid-sentence or mid-answer because the pipeline picks bounds on
ASR subtitle rows, which are half sentences. Snap every clip to sentence
edges (terminal punctuation or a clear pause, bounded extension), then,
when a text model is available, let it move the edges so the question
starts and the answer finishes, rejecting drastic rewrites.
A source with burned captions in another language (Japanese talk, English
captions) now still gets audience-language captions below the window, and
Japanese is no longer mistaken for Chinese. Same-language captions follow
source rows so word timing stays in sync, and the cinematic style shows
calm 2-5 word phrases instead of single-word flashes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Real demos exposed three layout failures: a 17-character Chinese title ran
off both edges at 112 px, a long English original under a short Chinese
line wrapped to a third line, and a model lumped 25 subtitle rows into one
2,000-character segment. Title lines now shrink to fit the 1080 px frame and
long model titles rewrap at punctuation or spaces; a long original gets one
caption line per screen and is cut with an ellipsis rather than a third
line; segments may span at most six rows. For same-language outputs a bad
segmentation falls back to the source rows but keeps the model's title,
names and highlights.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Automatic output has no confirmation step and no client export, so the
screening and production watches never settled and the only terminal event
depended on the results page staying open. Screening now ends as
auto_started once production begins, and a persisted studio-generation
watch reports studio_generation_finished once with aggregate counts
(variants, templates, speaker framing, packaging fallbacks, trims, burned
captions) and backend-timed duration. The backend keeps the analysis run_id
across phases, records generation.finished_at and marks automatic framing
as framing_source=auto.
Draft save/export carry template, style, tags switch and fallback; variant
downloads, shares and ratings carry the platform, template, style and
framing plus flow context. All new fields are allowlisted enums, counts or
booleans; no titles, captions, names or tags leave the app.
Also fix two issues found while rendering demos: deriving a draft for
another platform now resets the title template version (Xiaohongshu after
Douyin failed validation), and English caption segments and long model
titles no longer force a packaging fallback.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Resolve conflicts with main's onboarding/Studio delivery analytics, download
recovery, 1080p60 export and SenseVoice changes:
- analytics allowlist and operation names are the union of both sides;
layout adds 'window'; import properties keep flow ids plus platform count
and outro flag
- Studio download keeps recovery and writes a temporary info.json to read
the listing title/channel for nameplates
- 1080p60 stays an export-only preset next to the platform projections, and
keeps the outro frame-rate fix
- locale catalogs merged key by key (no key changed on both sides)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Real Dwarkesh subtitles are cut mid-sentence and the model regroups them
(21 rows came back as 13 sentences), so the one-to-one row check rejected
every result and all outputs fell back to untranslated captions. Ask for
contiguous sentence segments over the row ids, validate order and coverage,
and time each cue from its rows. Tags must name the concrete point instead
of generic praise, and fallback titles wrap by their own language.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Vertical captions stacked up to five lines: libass lays SRT out on a 288 px
canvas, so Fontsize 20 became ~133 px on a 1920 px frame. Scale the legacy
force_style for portrait and split every row into consecutive screens of at
most two lines, breaking on punctuation and natural joints, sharing time by
width and keeping the original translation in step. Template captions use
the same layout.
Add styles on top of the golden defaults (interview: classic, caption bar,
spotlight; podcast: word pop, caption bar, cinematic) with the same layout
contract and DESIGN.md palette; Studio can switch style on a packaged draft.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Result cards name the template (interview with bilingual captions, or
podcast) and say when packaging fell back to source captions. In Studio a
packaged draft shows a template panel with the two title lines and an editor
tags switch; legacy subtitle, opening-title and language settings are hidden
because the template renders those. Packaging is kept intact on save.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Chinese platforms (Douyin, Xiaohongshu) get the interview template: a pinned
two-line title, a 4:3 window that follows the speaker on a warm near-black
canvas, Chinese captions over the window with the original below, editor
tags above the captions and a thin progress line. English platforms
(TikTok, Reels, Shorts) get the podcast template: full 9:16 speaker crop,
one-to-three-word captions with the active word in the accent colour and a
short hook. Both show a left lower-third nameplate that never covers faces.
Packaging (title, translation, nameplates, tags, highlights) comes from one
text-model call per content and template, sends only the draft's subtitle
rows, validates every field, only names people seen in the subtitles or the
source listing, and falls back to source captions without blocking output.
Rendering stays per scene with bundled Noto Sans SC via libass.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
YouTube now requires solving a player JS challenge. The bundle installed bare
yt-dlp without yt-dlp-ejs, and yt-dlp only enables deno by default, so link
imports can fail with 'The page needs to be reloaded' on machines without
deno. Install yt-dlp[default] (same pinned version) and pass every JS runtime
found on PATH (deno, node, bun, quickjs) to all download paths.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A 2 h 22 min source pushed ffmpeg past 800% CPU during production: shot-cut
detection and frame sampling for speaker framing, and the full-source preview
transcode, decoded with every core at normal priority. Apply the shared
thread caps and low priority there and in legacy publish export.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: add optional isolated SenseVoice local transcription
* test: verify local transcription selector switches SenseVoice and Whisper
* docs: set realistic disk space expectation for SenseVoice preparation
* docs: reconcile issue 67 delivery with historical ASR plans
Screening rejected every source over two hours, which blocks most long
interviews and podcasts (a 2 h 22 min Dwarkesh episode failed immediately).
The subtitle route already chunks long transcripts and scales clip counts by
hour, so drop the upper bound. Frame-by-frame visual analysis keeps its
two-hour cost cap: longer sources are routed to subtitles instead of failing.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>