Captions: build_master_srt used fixed 2-word pairs, which split across pauses
("DOES IS") and flashed 0.18s cues ("WHAT A"). chunk_words now breaks on
punctuation and pauses >= 0.3s and grows a too-short cue to at most 3 words.
Subtitles path: an EDL "subtitles" path that did not resolve relative to the
EDL dir printed a warning and rendered without captions. It now also tries the
current directory, then exits with an error.
SKILL.md: music/SFX rules, per-section loudness check and a critic sub-agent in
self-eval, font-load assertion, and the "show the edit" hook as a worked example.
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(transcribe): pick the audio track and refuse to upload silence
extract_audio ran without -map, so ffmpeg applied its default stream selection
and took a single audio track — the first one. A multi-track recording is the
normal case for a screen capture: OBS writes the application on track 0 and the
microphone on track 1. Transcribing such a file uploaded the application audio
and dropped the narration without a word about it.
--audio-track selects the stream, zero-based, and defaults to 0, so existing
single-track runs are unchanged. A file with more than one track says so in
verbose output, since the default is right for some of them and wrong for others.
The extracted wav is also checked for level before it is sent. A peak under
-60 dBFS means the track is silent, which in practice means the wrong track was
picked, and Scribe charges the same for silence as for speech. The error names
the track count and points at the flag.
* fix(transcribe): key the cache by track, scan the peak in chunks, name the tracks
Review on #134 caught four things.
The transcript cache is keyed by video stem alone, and transcribe_one returns on
a cache hit before it looks at audio_track. Rerunning with --audio-track 1 after
a wrong-track run therefore handed back the very transcript the flag was meant to
replace. The track goes into the file name now, and track 0 keeps the old name so
existing transcripts stay valid.
peak_dbfs read the whole take with readframes(getnframes()) and copied it into an
array, so a two-hour 16 kHz mono file cost 230 MB twice over, with batch mode
running several at once. It scans in 64k-frame chunks instead; measured on a real
capture the peak is identical to the whole-file version.
The "try --audio-track" hint flipped between 0 and 1, so on a file with three
tracks it could point at another silent one. It lists the tracks that exist:
"The file has 3 audio tracks; try --audio-track 0 or 2."
The flag's help text said ffmpeg would otherwise take track 0. It applies its
default stream selection, which picks the track with the most channels.
* fix(transcribe): share one transcript path between single and batch mode
Review on #134 again: the previous commit put the track into the cache key in
transcribe.py and left transcribe_batch.py testing for {stem}.json. Batch mode
with --audio-track 1 therefore counted a file with a track-0 transcript as
cached and skipped it, which defeats the flag, and never recognised the
{stem}.track1.json it had just written, so it re-uploaded and re-billed that
file on every run.
Both now call transcript_path(), so the two cannot drift apart again.
Verified on a directory holding one video and a track-0 transcript: the default
run reports "1 cached, 0 to transcribe", the same run with --audio-track 1
reports "0 cached, 1 to transcribe" and goes on to the silence guard.
* fix(render): account for rotation metadata in portrait detection
Read display-matrix rotation and legacy rotate tags alongside coded dimensions. Swap dimensions for quarter-turn rotations so ffmpeg's autorotated filter input uses the correct scale axis.
* fix(render): trust display rotation side data only
Ignore plain rotate tags that do not guarantee ffmpeg autorotation and assert that the ffprobe query continues to request display-matrix rotation.
---------
Co-authored-by: Anton Sidorov aka anticodeguy <a@anticodeguy.com>
* fix(render): preserve source frame rate by default; add --fps override
extract_segment hardcoded `-r 24`, so every render was forced to 24 fps
regardless of the source. That silently downsamples 30/60 fps footage
(e.g. OBS screen/webcam captures at 60 fps) and contradicts the skill's
stated "match the source unless asked otherwise" principle. There was no
CLI option to change it.
Probe the source's r_frame_rate with ffprobe and pass it through to
ffmpeg verbatim (so fractional rates like 30000/1001 survive without
rounding), falling back to 24 only when the rate can't be determined.
Add a `--fps N` flag to force a specific rate when desired.
Behavior change: renders now keep the source frame rate by default
instead of always producing 24 fps.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(render): resolve one frame rate per render, not per segment
Addresses review feedback (identified by cubic): probing the frame rate
inside extract_segment meant a multi-source EDL mixing rates (e.g. a
30fps and a 60fps source) would encode segments at different rates. The
lossless concat (`-c copy`) requires every segment to share a frame
rate, so that would break the concat for multi-source edits.
Resolve a single output rate once in extract_all_segments and pass it to
every segment: explicit --fps wins, otherwise preserve the first source's
rate (falls back to 24 if unprobeable). Single-source renders still keep
the source rate; multi-source renders stay homogeneous and concat-safe.
extract_segment now takes a resolved `rate` string (still probes its own
source when called standalone with no rate).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(render): harden source frame rate handling
Prefer avg_frame_rate for variable-rate inputs, fall back to r_frame_rate, and accept decimal or rational --fps overrides. Add focused tests for validation, probing, and uniform multi-source render rates.
* fix(render): validate fps before fraction parsing
Restrict FPS input to bounded ffmpeg-compatible numeric and rational forms, canonicalize accepted values, and cover explicit overrides, fallback behavior, and probe failures.
* fix(render): keep canonical fps parsing idempotent
Reject reduced FPS fractions that exceed FFmpeg AVRational component bounds, ensuring every accepted canonical rate can be parsed again safely.
---------
Co-authored-by: Anton Sidorov aka anticodeguy <a@anticodeguy.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Landscape videos were scaled by width (scale=W:-2), which works for
horizontal footage but squishes portrait/vertical clips (height > width)
because the height becomes the constrained dimension.
Added is_portrait_source() which uses ffprobe to detect when h > w, then
switches the scale filter to scale=-2:H so portrait sources are scaled by
height instead, keeping the correct 9:16 orientation throughout the pipeline.
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
On Windows, Path.write_text() defaults to the system locale encoding
(cp1252), which can't encode Unicode characters used in the output
(e.g. >= U+2265, -> U+2192). Passing encoding="utf-8" ensures the file
writes correctly on all platforms.
The bold-overlay caption style shipped with MarginV=35, which places the
caption baseline just above the bottom edge. On 1080×1920 vertical output
(TikTok / IG Reels / YouTube Shorts / Snap) this sits squarely inside the
platform UI safe zone — caption, username, music ticker, and right-rail
actions cover roughly the bottom ~25–30% of the frame. Captions get clipped
or obscured by the UI after upload.
Bump to MarginV=90 so the baseline lands ~30% up from the bottom, clear of
the UI on every major vertical platform. Add a comment documenting the
rationale so future edits don't silently walk it back.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
iPhone (and many mirrorless cameras) ship HDR by default — HLG for iPhone,
PQ elsewhere. The current pipeline converts the bit depth via pix_fmt=yuv420p
but leaves the HDR transfer metadata (arib-std-b67 / smpte2084) intact on the
output. Players that honor the metadata — screen recorders, almost every
social upload re-encode path (TikTok, IG, YouTube, X) — interpret the 8-bit
values as HDR and display crushed, oversaturated colors.
QuickTime on macOS happens to tone-map on playback, which masks the bug
locally. Screen recording and uploaded renders cannot.
Fix:
- Detect HDR sources via ffprobe color_transfer ∈ {smpte2084, arib-std-b67}.
- When HDR, prepend a zscale → tonemap=hable → bt709 chain to the vf graph
so the extract produces clean Rec.709 SDR, with correct metadata.
- No-op for SDR sources (most existing footage).
Exposes `is_hdr_source()` and `TONEMAP_CHAIN` module-level so downstream
scripts can reuse the detection.
Verified end-to-end on iPhone HLG 1080×1920 HEVC footage:
before: output pix_fmt=yuv420p, color_transfer=arib-std-b67 (broken)
after: output pix_fmt=yuv420p, color_transfer=bt709 (clean SDR)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Mirrors the harnesless onboarding model so Claude Code, Codex, Hermes,
Openclaw, etc. can install video-use from a single pasted prompt.
install.md handles clone, deps, ffmpeg, skill registration per agent,
and walks the user through pasting their ElevenLabs API key.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Rename package from video-editor to video-use. Add banner screenshot
and standalone timeline-view SVG to the README.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>