- Detection also records where the burned captions sit. An English version
of a source with non-English captions blurs that band before any layout,
then is framed and captioned like a caption-free source.
- AI covers switched on by 1.4's first setup read as off after upgrading
(1.5 would bill one per output in the background); the chosen model stays
and one click turns them back on. Every 1.5 save marks the choice explicit.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- cloud ASR retries a rate-limited or failing chunk with backoff and stops
sending the rest once one fails for good; a bad concurrency env var is safe
- a stalled hardware encoder falls back to libx264; a bad input no longer
disables hardware encoding
- a failed outro delivers the video without it; Windows font paths and
apostrophes in paths work in the outro
- no AI cover without permission to send the frame (no invented face
under the guest's nameplate)
- handoff doc: review results, open decisions, deferred findings
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- A backup clip being framed and packaged is 'preparing', invisible to the
render dispatcher, so it can no longer ship raw; claiming it is atomic.
A failed preparation settles the generation, retry prepares it again, and
one interrupted by a restart becomes retryable.
- The results page keeps polling while outputs are preparing or rendering.
- AI cover redesign: no false success on timeout, nothing after unmount.
- Saved post copy shows at once; title lines keep their place while typing.
- Generated-image downloads refuse local/private hosts (each redirect too)
and stop at 40 MB.
- Docs: cloud ASR numbers, CHANGELOG entries for this round.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- 17 fixed real inputs (owner's picks: long interviews, short sources,
keynotes with screen demos, Chinese sources with burned captions, auto-
caption-only and once-360p sources) in benchmarks/fast_output/cases.json.
- scripts/fast_output_benchmark.py imports each through the product API and
reports per-stage time, model calls, tokens and cost, clips and renders,
source resolution, subtitle source, burned captions and packaging
fallbacks; --baseline prints deltas.
- Stage wall times (download, transcribe, clip finder, boundaries, framing,
packaging, post copy, render) are recorded next to token usage.
- Changelog: publish kit, outro, caption strategy, low-resolution fix.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Lead with an animated wall of untouched outputs and the measured
before/after table (1h43m interview: 7.5 min, ¥0.09), then features,
a subscription-tool comparison and the existing install paths. Sponsors
move below the quick start in a compact table.
DESIGN.md records the homepage stage band and number-proof rules
approved for the website redesign.
Co-authored-by: Cursor <cursoragent@cursor.com>
A two-hour talk yields 30+ clips; rendering all of them up front was most of
the run while people publish a handful. Automatic output now frames,
packages and renders the 10 highest-scored clips; the others are listed as
backup clips (status on_demand, no model or render cost) and are framed,
packaged and rendered when the user clicks 'Generate this one'. Adding a
platform keeps each moment's choice. Generation completes once the
automatic versions finish; analytics count on-demand versions separately.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Fast output no longer runs outline -> timeline -> scoring -> titles over
5000-character chunks. One call per ~60-minute window (parallel, with
overlap) reads compact id|time|text rows and returns finished clips with
rows, Chinese title, reason and score; picks are validated, deduplicated
across windows, snapped by refine_timeline and selected by select_clips, in
the legacy titled-clip shape. Short-video lengths (60-300 s, ideally
90-180 s) replace the podcast-era 90 s floor that padded answers into the
next question. Any failure falls back to the legacy steps.
MrBeast 2h06m: ~35 min / 128 calls / 500K tokens -> 38 s / 4 calls / 93K.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Automatic output has no confirmation step and no client export, so the
screening and production watches never settled and the only terminal event
depended on the results page staying open. Screening now ends as
auto_started once production begins, and a persisted studio-generation
watch reports studio_generation_finished once with aggregate counts
(variants, templates, speaker framing, packaging fallbacks, trims, burned
captions) and backend-timed duration. The backend keeps the analysis run_id
across phases, records generation.finished_at and marks automatic framing
as framing_source=auto.
Draft save/export carry template, style, tags switch and fallback; variant
downloads, shares and ratings carry the platform, template, style and
framing plus flow context. All new fields are allowlisted enums, counts or
booleans; no titles, captions, names or tags leave the app.
Also fix two issues found while rendering demos: deriving a draft for
another platform now resets the title template version (Xiaohongshu after
Douyin failed validation), and English caption segments and long model
titles no longer force a packaging fallback.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Resolve conflicts with main's onboarding/Studio delivery analytics, download
recovery, 1080p60 export and SenseVoice changes:
- analytics allowlist and operation names are the union of both sides;
layout adds 'window'; import properties keep flow ids plus platform count
and outro flag
- Studio download keeps recovery and writes a temporary info.json to read
the listing title/channel for nameplates
- 1080p60 stays an export-only preset next to the platform projections, and
keeps the outro frame-rate fix
- locale catalogs merged key by key (no key changed on both sides)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Vertical captions stacked up to five lines: libass lays SRT out on a 288 px
canvas, so Fontsize 20 became ~133 px on a 1920 px frame. Scale the legacy
force_style for portrait and split every row into consecutive screens of at
most two lines, breaking on punctuation and natural joints, sharing time by
width and keeping the original translation in step. Template captions use
the same layout.
Add styles on top of the golden defaults (interview: classic, caption bar,
spotlight; podcast: word pop, caption bar, cinematic) with the same layout
contract and DESIGN.md palette; Studio can switch style on a packaged draft.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: add optional isolated SenseVoice local transcription
* test: verify local transcription selector switches SenseVoice and Whisper
* docs: set realistic disk space expectation for SenseVoice preparation
* docs: reconcile issue 67 delivery with historical ASR plans
Sources such as WIRED interviews ship with captions burned into the picture,
and automatic output stacked a second, often misrecognized, track on top.
Detect burned captions locally from twelve lower-frame samples (bright glyphs
with dark outlines whose shape changes over time, so static logos do not
count). When found, automatic variants keep subtitles off and vertical crop
layouts fall back to the full frame so the original lines are not cut.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Real-source acceptance showed a 17-minute talk producing 8 variants and a
39-minute talk producing 15, but the per-project cap of three active exports
failed the whole generation on the fourth submit. Remove the cap so the source
alone decides how many outputs exist, and queue all variants on a dedicated
single-worker render executor. ffmpeg decode, filter and encode threads are
capped at a third of the cores and run at low priority so the machine stays
usable.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Upload-Post and Bilibili accept an optional output_variant_id and reuse the
finished immutable MP4 instead of re-exporting. Landscape variants are refused
for vertical-only targets, Bilibili only accepts bilibili variants, and the
result card links each completed variant to the publish page.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix: recover compatible timelines and preserve actual failure causes
* fix: preserve failure codes and transcription context in feedback
* test: verify Whisper with the installed portable Windows runtime
Show completed immutable render jobs directly in the quick-output result cards while keeping queued and failed states actionable.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat: add verified 1080p60 export and installed Windows video acceptance
* fix: support older installed settings and wrap preset labels on narrow screens
* test: launch installed desktop app for Windows video acceptance
* fix: save Windows acceptance report outside installed resources
Track bounded platform generation outcomes and reuse events without content identifiers, and explain ineligible platform outputs on the results page.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix: recover incomplete Whisper models and show project failure details
* fix: recover broken ASR runtimes and honor pipeline settings
* ci: use job-supported workspace context for recovery tests
* fix: bound video download retries and recover format and subtitle failures
* fix: recover PyAV 19 transcription failures found on Windows
* test: verify clean Windows ASR and existing PyAV 19 recovery
* test: isolate PyAV acceptance dependencies on Windows
Reuse saved automatic drafts to produce additional platform versions without rerunning content analysis, and retry failed output variants independently from completed clips.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Finalize automatic platform variants with a one-second Made with AutoClip outro while preserving legacy manual exports. Separate branded output caching and cover video/audio preservation with ffmpeg regressions.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Let users choose destinations at import, start generation immediately, and view automatic output variants without passing through manual plan confirmation. Preserve Studio for deliberate edits.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Allow opted-in imports to produce isolated platform output variants and queue existing renders without the confirmation page. Preserve legacy confirmation behavior and aggregate variant terminal states.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Persist platform and branding choices in a versioned workspace contract while preserving the existing confirmation flow. Fix async analysis failures so error state survives exception cleanup.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>