mirror of
https://github.com/zhouxiaoka/autoclip.git
synced 2026-10-02 02:34:34 +08:00
- Fast output runs the content pipeline without clustering (one unused model call) and without re-encoding every clip and collection (studio renders from the source); clip rows still sync from step 4 for the candidate list. - Packaging no longer asks the model to restate the transcript when no translation is needed; its rewrite was discarded for same-language output. - One face-detection pass per clip serves both the 4:3 interview window and the 9:16 podcast crop. - LLM inputs are compact JSON (indentation cost ~10 tokens per subtitle row). - Every text and vision call records tokens per project and stage into metadata/llm_usage.jsonl (DashScope native usage is captured now), so the cost of one video can be measured. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>