Update HyperFrames plugin skills to 0.7.40

This commit is contained in:
James
2026-07-07 10:04:20 -07:00
parent 6cb2cf7c95
commit be7e8da826
147 changed files with 4843 additions and 448 deletions
@@ -1,6 +1,6 @@
{
"name": "hyperframes",
"version": "0.7.25",
"version": "0.7.40",
"description": "Write HTML, render video. Compositions, Tailwind v4 styles, GSAP and runtime adapter animations, captions, voiceovers, audio-reactive visuals, and website-to-video capture for HyperFrames.",
"author": {
"name": "HeyGen",
@@ -1,13 +1,13 @@
---
name: embedded-captions
description: 'Add captions to a talking-head video. ONE catalog (CATALOG.md) of 31 visual identities behind two engines: column-flow (captions composited INTO the scene — matte occlusion + mix-blend; cream/ink/editorial/keynote/documentary/loud/neon/glitch/chrome/velocity) and themed constitutions (anchor/ordnance/terminal/neonsign/stardust/stomp/scoreboard/transit/vhs/arcade/dossier/laser/thunder/hologram/biolume/aurora/spectrum/papercut/popup/chalkboard/graffiti/brush/inkwater/ransom/lastpage — e.g. a glyph-decode climax, a neon sign WRITTEN stroke by stroke, or the quiet `anchor` rail default). Route by identity, never by mode. Trigger on "captions/subtitles", "embed/cinematic captions", "VFX captions", "炸/特效/酷炫字幕", a named identity, or top-tier motion-graphics asks. Embedding every word is wrong for most talking-head content — `anchor` is the verbatim default. Pipeline: transcription → hyperframes remove-background matting → HTML render → ffmpeg overlay. Requires hyperframes and a single-subject clip.'
description: 'Add captions to a talking-head video. ONE catalog (CATALOG.md) of 35 visual identities behind two engines: column-flow (captions composited INTO the scene — matte occlusion + mix-blend; cream/ink/editorial/keynote/documentary/loud/neon/glitch/chrome/velocity) and themed constitutions (anchor/ordnance/terminal/neonsign/stardust/stomp/scoreboard/transit/vhs/arcade/dossier/laser/thunder/hologram/biolume/aurora/spectrum/papercut/popup/chalkboard/graffiti/brush/inkwater/ransom/lastpage — e.g. a glyph-decode climax, a neon sign WRITTEN stroke by stroke, or the quiet `anchor` rail default). Route by identity, never by mode. Trigger on "captions/subtitles", "embed/cinematic captions", "VFX captions", "炸/特效/酷炫字幕", a named identity, or top-tier motion-graphics asks. Embedding every word is wrong for most talking-head content — `anchor` is the verbatim default. Runs locally end-to-end (transcribes and mattes the subject itself, no API key). Requires hyperframes and a single-subject clip (multi-shot clips are split per shot).'
metadata:
tags: captions, embedded-captions, occlusion, matting, talking-head, rembg-matting, whisper, ffmpeg, cinematic
---
# Embedded Captions
**One catalog, picked up front** ([CATALOG.md](CATALOG.md) — 17 identities; the three engines behind it are backend detail). **Standard** (default) builds a clean verbatim **rail** (lower-third subtitle carrying most text) + an **embed** climax composited _into_ the scene behind the subject at the peak. **Cinematic** is pure embed — no rail, every caption composited behind the subject (hero typography, accumulation, occlusion as the effect). **Theme** is a complete themed constitution — body paradigm × hero setpiece × front fx × plate reaction, composed from registries ([themes/README.md](themes/README.md)): `ordnance` `terminal` `neonsign` `stardust` `stomp`. Most explainer / voiceover is **Standard**; **embed is the scarce, earned peak** — embedding every word is the common mistake; Theme is for VFX-grade asks ("炸", "特效", "像 AE 做的").
**One catalog, picked up front** ([CATALOG.md](CATALOG.md) — 35 identities; the engines behind it are backend detail). **Standard** (default) builds a clean verbatim **rail** (lower-third subtitle carrying most text) + an **embed** climax composited _into_ the scene behind the subject at the peak. **Cinematic** is pure embed — no rail, every caption composited behind the subject (hero typography, accumulation, occlusion as the effect). **Theme** is a complete themed constitution — body paradigm × hero setpiece × front fx × plate reaction, composed from registries ([themes/README.md](themes/README.md)): `ordnance` `terminal` `neonsign` `stardust` `stomp`. Most explainer / voiceover is **Standard**; **embed is the scarce, earned peak** — embedding every word is the common mistake; Theme is for VFX-grade asks ("炸", "特效", "像 AE 做的").
---
@@ -16,7 +16,7 @@ metadata:
The craft prose below is long; the **pipeline itself is short** — and everything
deterministic is computed or compiled, never hand-written:
1. **Decision gate** (refuse bad clips) → **pick ONE identity from [CATALOG.md](CATALOG.md)** (17 identities; engine/compiler derived by lookup — never surface a mode/category question)
1. **Decision gate** (refuse bad clips) → **pick ONE identity from [CATALOG.md](CATALOG.md)** (36 identities; engine/compiler derived by lookup — never surface a mode/category question)
2. `hyperframes init` (skip it if the project dir already exists with the video inside — `matte.cjs`/`transcribe.cjs` adopt any video in the dir as source.mp4) → **`bash scripts/prepare.sh <project>`** (matte ∥ transcribe ∥ audio-envelope in parallel, then safe-zones v2 with scene palette/optics/lighting — one command, nothing forgotten)
3. **author a small JSON of creative choices** (read `safe-zones.json` first):
Cinematic → `plan.json` → `fill-timings.cjs` → `fit-fonts.cjs` → `make-composition.cjs`;
@@ -51,7 +51,7 @@ Rail-surface identities build exactly this (rail = `rail.html`, embed = the clim
## Step 0 — pick ONE identity from the CATALOG
**One front-end, three engines behind.** The user picks an IDENTITY from
[CATALOG.md](CATALOG.md) (17 entries: 12 classic + 5 themed); the engine,
[CATALOG.md](CATALOG.md) (36 entries: 10 classic + 26 themed); the engine,
compiler and authoring file are derived by lookup from the catalog row.
**Never surface "Standard vs Cinematic vs Theme" as a question** — those are
backend names (a product has one UX even with several engines). The catalog
@@ -169,7 +169,7 @@ each loop costs seconds. Render once, when the previews pass.
## The DNA registry — ten visual languages (replaces the template catalog)
Both modes draw from **[dna/](dna/README.md)** — six art-directed visual languages that
Both modes draw from **[dna/](dna/README.md)** — ten art-directed visual languages that
**parameterize per scene** (accent sampled from the footage, contact shadow along the
measured light direction, depth-match blur, RMS-coupled hero amplitude):
@@ -236,7 +236,7 @@ track has its own, much simpler spec → **[references/rail.md](references/rail.
| ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| [references/rail.md](references/rail.md) | **The rail track** — standard lower-third subtitle spec (the default; carries most text). |
| [references/composition-craft.md](references/composition-craft.md) | **The embed-track playbook** — grouping, planes, climax pop, occlusion judgement, accumulation/persistence. Read before embedding. |
| [dna/README.md](dna/README.md) | **The DNA registry** — six scene-parameterized visual languages; how to pick. |
| [dna/README.md](dna/README.md) | **The DNA registry** — ten scene-parameterized visual languages; how to pick. |
| [references/reference-bar.md](references/reference-bar.md) | **The taste bar** — per-register world-class references + the 5 positive checks. |
| [references/aesthetic-principles.md](references/aesthetic-principles.md) | **The 18 rules.** Beat Veed AI on taste. Read first. |
| [references/motion-vocabulary.md](references/motion-vocabulary.md) | 10 named motion primitives + tone→timing lookup |
@@ -226,7 +226,7 @@ The gallery used a `setInterval` loop; HyperFrames needs the same beats as **abs
## Pairs with HF skills
- `hyperframes-media` — `remove-background` (the matte) + `transcribe` (word timings).
- `media-use` — `remove-background` (the matte) + `transcribe` (word timings).
- `hyperframes-captions` — transcript consumption, grouping, positioning, exit guarantees, `fitTextFontSize`.
- `hyperframes-animation/rules/asr-keyword-glow.md` — the verbatim active-word envelope.
- `hyperframes-gsap` — single paused timeline, transform aliases, ease palette.
@@ -10,7 +10,7 @@ with only the climax(es) promoted to embed. Rail is not a fallback — it's the
> **Implementation note.** A dedicated rail renderer is the next build step. The rail is a
> plain `fg` caption track and maps cleanly onto hyperframes' native caption pipeline
> (`hyperframes-media` captions) — prefer reusing that over hand-rolling. Until wired, render
> (`media-use` captions) — prefer reusing that over hand-rolling. Until wired, render
> the rail as a simple `data-caption-layer="fg"` composition (no matte overlay for these caps).
## Position & safe area
@@ -110,7 +110,7 @@ function main() {
console.error("usage: transcribe.cjs <project-dir> [model] [language]");
process.exit(1);
}
// Default = multilingual `small`, NOT `small.en`. Per hyperframes-media: ".en models
// Default = multilingual `small`, NOT `small.en`. Per media-use: ".en models
// mistranslate non-English and mis-handle accented speech; default to small (auto-detects
// language)." We hardcoded small.en before — it hallucinated a wrong transcript on an
// accented speaker. Pass `small.en` only for known-clean-English; tough accents → a larger model.
@@ -1,6 +1,6 @@
---
name: faceless-explainer
description: "turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video, up to ~3 min (sweet spot 30-90s), where every visual is invented (typography, abstract graphics, diagrams, data-viz) rather than captured. There is no URL, no website capture, and no real assets. Use this skill for topic explainers, concept breakdowns, how-tos, listicles, and narrative explainers. Do not use it for a product launch/promo (use /product-launch-video), a tour of a real website (use /website-to-video), a GitHub PR (use /pr-to-video), captions on existing footage (use /embedded-captions), or a short unnarrated motion graphic (use /motion-graphics). If the intent is unclear, route through /hyperframes first. This is the shot-sequence architecture: every frame is authored as a time-coded shot sequence picked from a menu of golden blueprints, so frames develop over their full duration instead of freezing after entrance."
description: "Turn arbitrary text — an article, notes, a topic, a brief — into a faceless explainer video: there is no site or footage to capture, so the visuals are invented per scene (typography, abstract graphics, diagrams, data-viz). Use for topic explainers, concept breakdowns, how-tos, listicles. Not a product promo (/product-launch-video) or a site tour (/website-to-video). Unclear → /hyperframes."
---
> **media-use**: Before sourcing audio/images, call `/media-use` to resolve BGM/SFX/images from the HeyGen catalog. Run `--adopt` first to register existing assets. See `/media-use` skill.
@@ -25,7 +25,7 @@ Initialize only if `hyperframes.json` is missing. Name `<project>` from the topi
`npx hyperframes init "videos/<project>" --non-interactive --example=blank` — `init` checks the installed skills against the latest on GitHub and updates the global set if any are out of date.
**Show sign-in status before the brief** — run `npx hyperframes auth status` and **relay its output verbatim (don't paraphrase or rewrite it).** It reports whether voice/BGM will use HeyGen or local engines and, when not signed in, how to sign in. **If not signed in, STOP and wait for the user to choose — sign in, or say "go"/"offline" to continue with local engines — before asking the brief or anything else.** Treat it as a real decision point, not a passing note; don't fold the choice into the brief question, and don't write keys into a per-repo `.env`. (In autonomous mode, note the status and continue offline.) See `../hyperframes-media` → Preflight for the canonical guidance.
**Show sign-in status before the brief** — run `npx hyperframes auth status` and **relay its output verbatim (don't paraphrase or rewrite it).** It reports whether voice/BGM will use HeyGen or local engines and, when not signed in, how to sign in. **If not signed in, STOP and wait for the user to choose — sign in, or say "go"/"offline" to continue with local engines — before asking the brief or anything else.** Treat it as a real decision point, not a passing note; don't fold the choice into the brief question, and don't write keys into a per-repo `.env`. (In autonomous mode, note the status and continue offline.) See `../media-use` → Preflight for the canonical guidance.
**Gate:** `hyperframes.json` exists, and angle, length, aspect ratio, and language are locked; sign-in status was shown (signed in, or continuing offline).
@@ -88,7 +88,7 @@ Start audio after Step 3 approval. Run it in the background, then continue to St
`node <SKILL_DIR>/scripts/audio.mjs --script ./SCRIPT.md --storyboard ./STORYBOARD.md --hyperframes . --out ./audio_meta.json &`
The audio script handles narration, word timings, BGM lookup from HeyGen's music library, and timing metadata. BGM mood comes from the storyboard's `music:` field. This uses the HeyGen Audio API for retrieval, not generation, and the same `~/.heygen` credential as TTS. For provider details, read `../hyperframes-media/references/tts.md`.
The audio script handles narration, word timings, BGM lookup from HeyGen's music library, and timing metadata. BGM mood comes from the storyboard's `music:` field. This uses the HeyGen Audio API for retrieval, not generation, and the same `~/.heygen` credential as TTS. For provider details, read `../media-use/audio/references/tts.md`.
If there is no narration and no `SCRIPT.md`, skip voice generation. BGM may still run if the storyboard has a music mood.
@@ -200,7 +200,7 @@ The reusable, domain-agnostic shot shapes live in `../hyperframes-animation/blue
| `[../hyperframes-animation/blueprints-index.md](../hyperframes-animation/blueprints-index.md)` | Step 3: role→blueprint menu. Step 4: pick the shot shape. |
| `[../hyperframes-core/references/storyboard-format.md](../hyperframes-core/references/storyboard-format.md)` | Step 3: write `STORYBOARD.md`. |
| `[../hyperframes-core/references/script-format.md](../hyperframes-core/references/script-format.md)` | Step 3: write `SCRIPT.md`. |
| `[../hyperframes-media/references/tts.md](../hyperframes-media/references/tts.md)` | Step 3.1: choose or understand TTS providers and voices. |
| `[../media-use/audio/references/tts.md](../media-use/audio/references/tts.md)` | Step 3.1: choose or understand TTS providers and voices. |
| `[references/visual-design.md](references/visual-design.md)` | Step 4: write the frame's shot sequence (+ Layout vocabulary). |
| `[references/motion-language.md](references/motion-language.md)` | Step 4: the motion vocabulary + the motion doctrine. |
| `[references/cut-catalog.md](references/cut-catalog.md)` | Step 4-5: the cut catalog (worker builds within-frame seams). |
@@ -96,7 +96,7 @@ These four rules are the difference between a clip that reads as a serious expla
Elements should use **long-tail decel curves that let them settle smoothly. `power3` is enough in most cases.** No bouncy, no overshoot, no `back.out` / `bounce.out` / `elastic.out` as a default.
Bouncy is the **#1 instant turn-off** in user-made Remotion / HyperFrames videos, and the agent almost never gets it right — it thinks bouncy adds emphasis, but it buys that emphasis at the cost of cleanliness. The serious motion-design shops feel the same. **Smooth always wins.** Overshoot is demoted to a **rare, explicitly-playful exception** (a consumer/fun logo slam, a deliberate bell-hit) — never the house style. Name the intent as a long-tail settle; the worker maps `power3` (or `expo.out` on a fast arrival). See `../hyperframes-animation/rules/spring-pop-entrance.md` — it now leads with the smooth settle.
Bouncy is the **#1 instant turn-off** in user-made Remotion / HyperFrames videos, and the agent almost never gets it right — it thinks bouncy adds emphasis, but it buys that emphasis at the cost of cleanliness. The serious motion-design shops feel the same. **Smooth always wins.** Overshoot is demoted to a **rare, explicitly-playful exception** (a consumer/fun logo slam, a deliberate bell-hit) — never the house style. Name the intent as a long-tail settle; the worker maps `power3` (or `expo.out` on a fast arrival). See `../hyperframes-animation/rules/spring-pop-entrance.md` — it now leads with the smooth settle. (The exact form of that settle is a critically-damped spring; the worker has a baked, seek-safe `springEase` — ζ=1 — in `../hyperframes-animation/adapters/gsap-easing-and-stagger.md` → Spring Eases for when the settle is the hero. Real physics, same doctrine — not a license for bounce.)
## 2. Sequential reveal in the back ~50%, timed to the voiceover
@@ -115,7 +115,7 @@ The agent's two reflexive ways to fake "aliveness" both read cheap:
- **No lazy breathing.** Scaling cards/text up and down in a circular loop to look "alive" is the cheap tell. Don't reach for it.
- **No bad slow pan / push in the back half.** A slow pan or push on elements in the later ~50% of a scene **disrupts the viewer's sightline and causes eye discomfort** — it actively makes the frame worse, not better.
The fix for both is the same: **stagger element reveals in time with the script** (rule 2). And the governing principle: **"I'd rather have NO motion than BAD motion."** A held, still frame is better than a frame kept "alive" by breathing or a drifting camera. The **only sanctioned aliveness** during a hold is **subtle jitter** — a small low-amplitude jitter that keeps a frame from feeling dead without looking weak (it's in Claude videos now). Everything else holds.
The fix for both is the same: **stagger element reveals in time with the script** (rule 2). And the governing principle: **"I'd rather have NO motion than BAD motion."** A held, still frame is better than a frame kept "alive" by breathing or a drifting camera. The **only sanctioned aliveness** during a hold is **subtle jitter** — a small low-amplitude jitter that keeps a frame from feeling dead without looking weak (it's in the reference videos now). Everything else holds.
## 4. Internal seams are velocity-matched cuts
@@ -1,12 +1,14 @@
#!/usr/bin/env node
// audio.mjs — faceless-explainer audio ADAPTER. The TTS / BGM / SFX implementation
// audio.mjs — audio ADAPTER (reuses the product-launch SCRIPT.md / STORYBOARD.md
// model; this file is intentionally identical across the reusing skills). The
// TTS / BGM / SFX implementation
// no longer lives here: it is the shared engine at
// ../../hyperframes-media/scripts/audio.mjs. This file only (a) maps the
// frame model (SCRIPT.md frames + STORYBOARD.md music/sfx) into the
// ../../media-use/audio/scripts/audio.mjs. This file only (a) maps the
// product-launch model (SCRIPT.md frames + STORYBOARD.md music/sfx) into the
// engine's neutral audio_request.json, (b) converts the engine's id-keyed
// audio_meta back into the frame-keyed shape captions.mjs / assemble-index.mjs
// already consume, and (c) keeps the local `sync-durations` pass (it rewrites
// STORYBOARD.md, which is this skill's concern).
// STORYBOARD.md, which is product-launch-specific).
//
// Three modes (unchanged CLI surface):
// (default) generate — engine --only tts,bgm. BGM mode is "retrieve" (strict:
@@ -27,7 +29,7 @@ import { fileURLToPath } from "node:url";
import { parseStoryboard } from "./lib/storyboard.mjs";
const HERE = dirname(fileURLToPath(import.meta.url));
const DEFAULT_ENGINE = join(HERE, "..", "..", "hyperframes-media", "scripts", "audio.mjs");
const DEFAULT_ENGINE = join(HERE, "..", "..", "media-use", "audio", "scripts", "audio.mjs");
const flag = (argv, name, def) => {
const i = argv.indexOf(`--${name}`);
@@ -86,9 +88,9 @@ function runEngine({ request, hyperframesDir, neutral, only, extra = [] }, die)
if (r.status !== 0) die(`media audio engine exited ${r.status}`);
}
// Engine neutral meta (id-keyed) → frame-keyed meta consumed by
// Engine neutral meta (id-keyed) → product-launch meta (frame-keyed) consumed by
// captions.mjs / assemble-index.mjs. id is the zero-padded frame number.
function toFrameKeyedMeta(neutral) {
function toProductLaunchMeta(neutral) {
const voices = (neutral.voices ?? []).map((v) => ({
frame: Number(v.id),
path: v.path,
@@ -152,7 +154,7 @@ function runGenerate(argv) {
const neutral = neutralPath(outPath);
runEngine({ request, hyperframesDir, neutral, only: "tts,bgm" }, die);
const meta = toFrameKeyedMeta(JSON.parse(readFileSync(neutral, "utf8")));
const meta = toProductLaunchMeta(JSON.parse(readFileSync(neutral, "utf8")));
writeFileSync(outPath, JSON.stringify(meta, null, 2));
console.log(
`✓ audio generate: ${meta.voices.length} voice + ${meta.bgm ? "1 bgm" : "no bgm"} → ${outPath}`,
@@ -189,7 +191,7 @@ function runFetchSfx(argv) {
// voices/bgm written by the earlier generate (--only tts,bgm) pass are preserved.
runEngine({ request, hyperframesDir, neutral, only: "sfx" }, die);
const meta = toFrameKeyedMeta(JSON.parse(readFileSync(neutral, "utf8")));
const meta = toProductLaunchMeta(JSON.parse(readFileSync(neutral, "utf8")));
writeFileSync(outPath, JSON.stringify(meta, null, 2));
console.log(`✓ audio fetch-sfx: ${meta.sfx.length} SFX cue(s) → ${outPath}`);
}
@@ -0,0 +1,36 @@
// pad-frame-duration.mjs — keeps a frame's own #root/clip data-duration in
// sync with the padded index.html wrapper duration transitions.mjs computes.
//
// The frame's OWN internal file declares its #root/clip data-duration to the
// STORYBOARD's content-only length (frame-worker.md: duration is "fixed
// upstream"). When an outgoing transition pads the index.html WRAPPER's
// data-duration to cover the transition tail, the frame's own internal
// duration is left short — the render engine clip-gates the sub-composition's
// visible content at that shorter value, so content vanishes abruptly at
// content-end instead of fading gracefully through the wrapper's extended
// fade-out tween. Pad the frame's own file to match so both durations agree.
import { readFileSync, writeFileSync } from "node:fs";
import { resolve } from "node:path";
export function padFrameInternalDuration(hyperframesDir, frameSrc, frameId, newDuration) {
const framePath = resolve(hyperframesDir, frameSrc);
let html;
try {
html = readFileSync(framePath, "utf8");
} catch (err) {
if (err?.code === "ENOENT") return;
throw err;
}
const tagRe = /<[a-z][\w:-]*\s[^<>]*?>/gi;
let m;
while ((m = tagRe.exec(html)) !== null) {
const tag = m[0];
if (!tag.includes(`data-composition-id="${frameId}"`)) continue;
if (!/data-duration="[\d.]+"/.test(tag)) continue;
const newTag = tag.replace(/data-duration="[\d.]+"/, `data-duration="${newDuration}"`);
if (newTag === tag) return;
writeFileSync(framePath, html.slice(0, m.index) + newTag + html.slice(m.index + tag.length));
return;
}
}
@@ -28,6 +28,7 @@ import { join, resolve } from "node:path";
import { parseStoryboard } from "./lib/storyboard.mjs";
import { parseFormat } from "./lib/dimensions.mjs";
import { loadTransitionRegistry, transitionsByName } from "./lib/transition-registry.mjs";
import { padFrameInternalDuration } from "./lib/pad-frame-duration.mjs";
const flag = (argv, name, def) => {
const i = argv.indexOf(`--${name}`);
@@ -183,6 +184,12 @@ function runInject(argv) {
const dur = resolveDur(spec, rec, reg);
const T = r3(incoming.start); // cut = incoming start (frames tile)
outgoing.duration = r3(outgoing.duration + dur); // extend outgoing only
padFrameInternalDuration(
hyperframesDir,
order[i - 1].frame.src,
outgoing.id,
outgoing.duration,
);
gsapLines.push(
...buildGsap(rec, outgoing.id, incoming.id, dur, T, spec.direction, CW, CH, die),
);
+122
View File
@@ -0,0 +1,122 @@
---
name: figma
description: Import Figma content into a HyperFrames composition — rendered assets, brand tokens, components, storyboard sections → reconstructed motion (frames read as states, not slides) via REST/CLI, with optional connector-assisted motion and shader paths when the needed Figma tools are available. Use when the user pastes a figma.com link or asks to bring a Figma design, frame, logo, brand, or animation into a video/composition.
---
# Figma → HyperFrames
Bring the user's Figma work into a composition. **Split by capability** (design spec §2):
| Phase | What | Transport | Surface |
| ----- | ------------------- | ---------------------------- | ----------------------------- |
| 1 | Static assets | REST | `hyperframes figma asset` |
| 2 | Brand tokens/styles | REST | `hyperframes figma tokens` |
| 3 | Components → HTML | REST | `hyperframes figma component` |
| 4 | Motion → GSAP | Optional Figma connector | `get_motion_context` if available |
| 5 | Shaders | Optional connector / manual export | connector tools if available |
REST is used wherever it can be (usable at volume, headless); connector-assisted paths are only for areas where Figma exposes no REST equivalent (motion, shaders). Every path freezes assets locally so renders stay deterministic. Storyboard reconstructions compose Phase-1 asset exports (REST) with agent-driven timeline assembly — no connector needed. Existing frozen assets, manifest records, and bindings are unaffected by routing changes — the split only changes which credential the next import uses.
## Auth — two credentials, scoped
**Preflight — before the first CLI call, check a token exists**: shell env (`[ -n "$FIGMA_TOKEN" ]`) **or** the project `.env` (the CLI auto-loads it — a `.env` entry counts as configured). If neither, do NOT run the command to harvest the error — walk the user through the one-time setup first, then stop and wait:
1. figma.com/settings → **Security** → **Personal access tokens** → Generate new token.
2. Scopes — read-only is all this integration ever needs (it never writes to Figma): **File content: Read-only** + **File metadata: Read-only**. Optionally **Variables: Read-only** for brand variables — that scope only works on Figma Enterprise; without it `tokens` degrades to published styles automatically (expected behavior, not an error — say so).
3. `export FIGMA_TOKEN="figd_…"` — and suggest persisting it (shell profile or project `.env`) so no future session repeats this.
While onboarding, also set expectations in one breath: every import lands as a **local frozen file with recorded provenance** — renders never call Figma, re-running a command re-imports only what changed in Figma, and one token works for assets, brand tokens, and components across every file their Figma account can view.
- **Phases 4–5 (motion/shaders):** require a connected Figma integration/tooling surface separate from the token. If the needed tools are unavailable or unauthenticated, tell the user that motion/shader import requires the Figma connector and stop; continue only with REST-backed assets/tokens/components.
- Say exactly which credential a failing phase needs — never present the split as broken.
- `BAD_TOKEN` (401) mid-flow → the token is expired/revoked; re-mint. `FORBIDDEN` (403) → missing read scope or no access to that file — check scopes + file visibility. `REQUIRES_ENTERPRISE` (403 on variables) → not a failure: styles fallback already ran.
**Rate-limit awareness (spec §2.1):** connector motion/shader calls may be scarce on low-tier Figma plans (figma plan matrix as of 2026-07 — re-verify if quotas look off) — batch with `recursive:true` on the parent node when supported, skip verification screenshots unless asked, and cache raw connector responses so re-derivation never spends a second call. REST is per-minute (10+/min, per-endpoint buckets) — fine at volume, back off on 429.
## Routing
Parse the user's figma link with `parseFigmaRef` (URL, `fileKey:nodeId`, bare `fileKey`). Then by intent:
- "use this layer / logo / image" → **Asset** (CLI)
- "pull my brand / colors / tokens" → **Tokens** (CLI)
- "build a scene from this frame" → **Component** (CLI)
- "import this animation / motion" → **Motion** (connector if available, below)
- a storyboard section / filmstrip of scene frames → **Storyboard** (below)
- shader fill/effect → **Shaders** (below)
**Narrate every step for the user** — before each command say what you're about to pull from Figma; after it, say where the artifact landed (the frozen path / sidecar / component dir), what changed in the composition, and the immediate next action (preview, add printed variables, re-import to link bindings). The user should never have to ask "did it work?" or "now what?".
## Assets (Phase 1 — CLI)
```bash
hyperframes figma asset '<url-or-fileKey:nodeId>' [--format svg|png|jpg|pdf] [--scale 2] [--description "..."] [--entity "..."]
```
Renders over REST, sanitizes SVG, freezes under `.media/images/`, appends the manifest with provenance, regenerates `.media/index.md` (the shared media-use inventory), prints an `<img>` snippet. Idempotent per `fileKey:nodeId:format:scale:version`. Prefer SVG for vectors/logos (scalable, animatable), PNG `--scale 2` for raster fidelity. **Always pass `--description "<what it is>"`** (it becomes the index row + `<img alt>`); add `--entity "<name>"` for named brand marks so media-use `resolve --entity` finds them later (entity hits match across image/icon).
## Tokens (Phase 2 — CLI)
```bash
hyperframes figma tokens <fileKey>
```
Imports variables as composition brand-variable entries + `figma-tokens.json` sidecar + binding-index records (`.media/figma-bindings.jsonl`). Variables are Enterprise-gated upstream: on other plans the command degrades to published-style metadata (values resolve at component-import time). Add the printed entries to the composition's `data-composition-variables`.
**Import tokens before components** when both are wanted — that's what lets component colors link to brand variables instead of baking duplicates.
## Components (Phase 3 — CLI)
```bash
hyperframes figma component '<url-or-fileKey:nodeId>'
```
Node tree → editable HTML at exact figma geometry, packaged as a registry item under `compositions/components/<name>/`. Vectors/boolean-ops auto-rasterize via Phase-1 export. Binding pass (spec §7.1, exact-ID only — never value matching):
- Fill bound to an **imported** token → `var(--slug, #literal)` — brand refresh propagates.
- Bound to an **unknown** token → literal + `data-figma-unresolved` flag. The command tells you; offer the user: run `tokens` on the source (or library) file, then re-import the component to link them. Ask **once** per unknown library which file it is — never guess, never match by hex.
## Motion (Phase 4 — connector-assisted, when available)
**Usage beacon:** connector phases have no CLI touchpoint, so fire the skill beacon at start and finish (anonymous, consent-gated, never fails): `npx hyperframes events --skill=figma-motion` when you begin, `npx hyperframes events --skill=figma-motion --event=skill_completed --outcome=success|error` when done. Same for shaders (`figma-shaders`) and storyboards (`figma-storyboard`).
No REST equivalent exists. If a connected Figma tool exposes the needed motion context, use it, then hand output to the pure helpers in `@hyperframes/core/figma`; otherwise stop and tell the user this phase needs the Figma connector:
1. `get_motion_context(fileKey, nodeId)` — use `recursive:true` on the parent frame (one call for the whole scene, not one per element). Save the raw JSON next to the project (`.media/figma-cache/`) so retranslation is free.
2. Normalize into a `MotionDoc`: per animated property a `MotionTrack` { property (motion.dev name), values, times (0..1), ease[] (named or `[x1,y1,x2,y2]` bezier), duration, repeat }. Selector = the element's stable id (`#<id>` from Phase-3 output or the authored scene).
3. `motionToGsap(doc)` → `emitTimelineScript(spec)` → inject as a `<script>` after the GSAP + CustomEase CDN tags. Paused, finite, registered on `window.__timelines` with a literal key.
4. Untranslatable track (shader-driven, unsupported prop, complex masks) → if the connector provides video export, bake with `export_video`, freeze MP4, and embed as `<video class="clip">`. Exception: shader-driven tracks — figma's export path flattens shaders to the base color (see Shaders below), so a bake there silently loses the shader; ask the user for a native figma export instead. Always say which path you used and why. Named eases outside the mapped set fall back to linear — the mapping table lives in `motionEase.ts`; flag the fallback to the user when it fires.
5. Run `npx hyperframes lint && npx hyperframes validate` before calling it done.
## Shaders (Phase 5 — mostly manual)
Figma connector render paths may not execute shaders (they can flatten to the base color), and shader source is only reachable for **library-published** styles (paid Full seat). Default path: ask the user to export the shader frame natively in Figma (PNG or Motion MP4), then import it as a Phase-1 asset / clip. Don't attempt connector pixel capture of a shader unless the tool explicitly preserves shader output.
## Storyboards (a SECTION of scene frames → animation)
**The cardinal rule: storyboard frames are KEYFRAMES, not slides.** Two frames containing the same element describe that element's state through time — animate the ELEMENT between the states; never play the frames as a sequence of stills. A logo drawn in four consecutive frames at descending y is ONE element rising through four keyframes. Playing storyboard frames back-to-back is the failure mode; reconstructing the element timelines they imply is the job.
Storyboard files follow a grammar you can parse mechanically — don't eyeball, decode:
1. **Scene units**: inside the SECTION, every frame-sized node is a scene — both named FRAMEs _and_ loose full-frame RECTANGLEs (designers paste stills straight into the section). Filter by size (≈ composition aspect, e.g. >1400×900), not by node type or name.
2. **Order = x-position** (row-major if the strip wraps). Sort scenes by `absoluteBoundingBox.x`.
3. **Diff adjacent frames into element chains** — this is where the animation lives. Match children across consecutive frames: first by **name** (same name = same element → tween its relative x/y/w/h between states), then by **geometry similarity** (similar size + nearby center = same logical element whose pixels changed → crossfade the two exports in place while tweening geometry; covers typed-text progressions and morph states). Unmatched children enter/exit at their scene's beat. Frame background fills tween as a color track. Export ONE asset per chain (one per state only when pixels genuinely differ) — never one still per frame.
4. **Stills are the fallback, not the default** — only for frames that don't decompose (flat full-frame screenshots with no shared elements); those get the animatic treatment below.
5. **Director notes**: TEXT nodes below the strip are motion intent, paired to the scene whose x-range they overlap. They describe _how_ to animate — they are not on-screen copy.
6. **Batch exports** (elements or stills): `GET /v1/images` accepts comma-separated ids, but big scene frames hit "Render timeout" past ~12 ids — chunk to ~4 per call with a retry. (One call per scene wastes the rate budget; 26 scenes ≈ 52 calls via the single-asset path.)
7. **Note verbs → transitions** (starter vocabulary, extend as encountered):
| Note says | Do |
| ------------------------------- | ----------------------------------------------------------------- |
| EXPLOSION / BURST | incoming scale ~1.5→1 + fade, `power3.out` |
| SLIDES / SLIDE TO THE… / SCROLL | directional slide in from that edge |
| MORPH / REVEALS | crossfade — or Phase-3 import if the motion is inside one scene |
| CYCLE THROUGH / EACH ONE | longer hold — or Phase-3 import if items animate within the scene |
| (no note) | crossfade + slow Ken-Burns drift |
8. **Stills vs. components routing**: a note describing motion _between_ scenes → transition on the still (above). A note describing motion _inside_ a scene ("TEXT LINES REVEAL ONE AFTER THE OTHER", "PILLS ANIMATE IN") → that frame deserves a Phase-3 component import (real elements) animated per the note, not a flat PNG. Do the animatic pass first with stills, then upgrade the scenes the notes single out.
9. One `main` timeline sequences everything (opacity/x/y per scene at absolute times) — no per-scene sub-compositions needed for an animatic.
10. **Escalation — frames depict ONE product UI → rebuild the app, not element chains.** When every frame is the same application screen in successive states (a signup flow, a settings panel, a player), element chains undersell it. Rebuild the UI as live DOM — Phase-3 component import for the parts that change state, real exported pixels for static chrome (**code what changes state, freeze what doesn't**) — and treat each frame delta as an **interaction to perform**, not a tween to apply: the cursor enters, clicks the control, the state responds, screens push/slide as real navigation. The result reads as one continuous screen recording of a working app. This is the cardinal rule taken to its conclusion for UI flows; the stills/element-chain treatments are for storyboards that aren't one coherent application.
## Determinism
Never leave a Figma URL in the composition — freeze first. Never emit `repeat: -1`. Timelines paused, finite, literal `window.__timelines` keys. All Figma I/O at import time; render sees local files only.
@@ -1,13 +1,10 @@
---
name: general-video
description: >
The fallback workflow for authoring custom HyperFrames video compositions at
any length or format — longer or multi-scene pieces, brand / sizzle reels,
montages, title cards, static loops, and freeform compositions. Input- and
length-agnostic. If a specialized workflow clearly fits the input — a
marketed product, a website, a topic explainer, a GitHub PR, existing
footage, a short motion graphic, or a Remotion port — prefer it (see
/hyperframes); use this only as the general fallback when none fit.
The fallback workflow for authoring or editing any custom HyperFrames composition at any
length or format — longer / multi-scene pieces, brand and sizzle reels,
montages, title cards, static loops, freeform builds. Use only when no
specialized workflow fits the input; routing table at /hyperframes.
metadata: { "tags": "orchestrator, general-video, fallback, freeform, composition-authoring" }
---
@@ -96,24 +93,24 @@ Never use `position: absolute; top: Npx` on a content container — it overflows
## Build — delegate to the domain skills
This maps the skill's full surface (see the `description`) to its references — non-exhaustive; when an intent isn't listed, route through `hyperframes-creative` (look/concept), `hyperframes-animation` (motion), `hyperframes-core` (contract), `hyperframes-media` (audio/captions). **The first row is ADDITIVE — read it AND your intent row, not one or the other.**
This maps the skill's full surface (see the `description`) to its references — non-exhaustive; when an intent isn't listed, route through `hyperframes-creative` (look/concept), `hyperframes-animation` (motion), `hyperframes-core` (contract), `media-use` (audio/captions). **The first row is ADDITIVE — read it AND your intent row, not one or the other.**
| Building… | Read first (in order) |
| --------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **ALWAYS — every non-trivial piece, on top of your intent row below** | `hyperframes-creative/references/house-style.md` + `references/video-composition.md` (also gated in Step 1 / HARD-GATE; the "produced, not generated" foreground detailing) |
| **Kinetic typography / text-forward** | `hyperframes-animation/techniques.md` (kinetic type) + `adapters/gsap-easing-and-stagger.md` + `rules/kinetic-beat-slam.md` |
| **Title card / lower-third / overlay / PiP / text-behind-subject** | `hyperframes-creative/references/composition-patterns.md` + (for the centered/sized frame) `hyperframes-core` → "Root must be sized" |
| **Logo / brand-mark reveal** | `hyperframes-animation/rules/svg-path-draw.md` (draw-on) + `rules/3d-text-depth-layers.md` + `rules/scale-swap-transition.md` |
| **Data / stats / numbers** | `hyperframes-animation/rules/counting-dynamic-scale.md` + `rules/stat-bars-and-fills.md` + `hyperframes-creative/references/data-in-motion.md` |
| **Product / app / UI demo** | `hyperframes-animation/rules/3d-page-scroll.md` + `rules/cursor-click-ripple.md` + `rules/press-release-spring.md` |
| **Audio-reactive / music-driven** | `hyperframes-creative/references/audio-reactive.md` (pre-extract bands; map to motion) |
| **Narrated / voiceover / music / SFX / captions** | `hyperframes-media` → the shared audio engine `scripts/audio.mjs` (one call = TTS + BGM + SFX → `audio_meta.json`); caption authoring + asset placement via `hyperframes-core`. See **Audio** below. |
| **Multi-scene / transitions** | `hyperframes-animation/transitions/overview.md` **then** `transitions/catalog.md` (you are not done after the overview — the GSAP recipe is in the catalog) |
| **Modular / sub-compositions** | `hyperframes-core/references/composition-patterns.md` + `references/sub-compositions.md` |
| Building… | Read first (in order) |
| --------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **ALWAYS — every non-trivial piece, on top of your intent row below** | `hyperframes-creative/references/house-style.md` + `references/video-composition.md` (also gated in Step 1 / HARD-GATE; the "produced, not generated" foreground detailing) |
| **Kinetic typography / text-forward** | `hyperframes-animation/techniques.md` (kinetic type) + `adapters/gsap-easing-and-stagger.md` + `rules/kinetic-beat-slam.md` |
| **Title card / lower-third / overlay / PiP / text-behind-subject** | `hyperframes-creative/references/composition-patterns.md` + (for the centered/sized frame) `hyperframes-core` → "Root must be sized" |
| **Logo / brand-mark reveal** | `hyperframes-animation/rules/svg-path-draw.md` (draw-on) + `rules/3d-text-depth-layers.md` + `rules/scale-swap-transition.md` |
| **Data / stats / numbers** | `hyperframes-animation/rules/counting-dynamic-scale.md` + `rules/stat-bars-and-fills.md` + `hyperframes-creative/references/data-in-motion.md` |
| **Product / app / UI demo** | `hyperframes-animation/rules/3d-page-scroll.md` + `rules/cursor-click-ripple.md` + `rules/press-release-spring.md` |
| **Audio-reactive / music-driven** | `hyperframes-creative/references/audio-reactive.md` (pre-extract bands; map to motion) |
| **Narrated / voiceover / music / SFX / captions** | `media-use` → the shared audio engine `scripts/audio.mjs` (one call = TTS + BGM + SFX → `audio_meta.json`); caption authoring + asset placement via `hyperframes-core`. See **Audio** below. |
| **Multi-scene / transitions** | `hyperframes-animation/transitions/overview.md` **then** `transitions/catalog.md` (you are not done after the overview — the GSAP recipe is in the catalog) |
| **Modular / sub-compositions** | `hyperframes-core/references/composition-patterns.md` + `references/sub-compositions.md` |
### Audio: one engine (TTS · BGM · SFX)
Only when the piece calls for it (per "build exactly what was asked" — no ambient music on a title card). Don't hand-roll TTS or vendor a copy: write a neutral `audio_request.json` and call the shared engine in `hyperframes-media`. It auto-degrades on one switch — HeyGen credential present → HeyGen TTS + music/SFX **retrieval**; absent → ElevenLabs/Kokoro TTS, Lyria/MusicGen BGM **generation**, and the bundled SFX library. Full flag list + request/meta schema: the header comment of `hyperframes-media/scripts/audio.mjs`.
Only when the piece calls for it (per "build exactly what was asked" — no ambient music on a title card). Don't hand-roll TTS or vendor a copy: write a neutral `audio_request.json` and call the shared engine in `media-use`. It auto-degrades on one switch — HeyGen credential present → HeyGen TTS + music/SFX **retrieval**; absent → ElevenLabs/Kokoro TTS, Lyria/MusicGen BGM **generation**, and the bundled SFX library. Full flag list + request/meta schema: the header comment of `media-use/audio/scripts/audio.mjs`.
```jsonc
// audio_request.json — one line per narrated segment; `id` is yours (joins audio_meta back)
@@ -127,11 +124,11 @@ Only when the piece calls for it (per "build exactly what was asked" — no ambi
```
```bash
# <MEDIA_DIR> = the installed hyperframes-media skill dir (sibling of this skill)
# <MEDIA_DIR> = the installed media-use skill dir (sibling of this skill)
node <MEDIA_DIR>/scripts/audio.mjs --request ./audio_request.json --hyperframes . --out ./audio_meta.json
```
Then read `audio_meta.json`: mount each `voices[].path` + (`bgm.path`, `sfx[]`) as `<audio>` tracks and use `voices[].words` for captions, all per `hyperframes-core` (audio tracks + caption authoring). If BGM took the generate path (`bgm_pending: true`), run `hyperframes-media/scripts/wait-bgm.mjs` before final render.
Then read `audio_meta.json`: mount each `voices[].path` + (`bgm.path`, `sfx[]`) as `<audio>` tracks and use `voices[].words` for captions, all per `hyperframes-core` (audio tracks + caption authoring). If BGM took the generate path (`bgm_pending: true`), run `media-use/audio/scripts/wait-bgm.mjs` before final render.
## Output checklist → `hyperframes-cli`
@@ -1,6 +1,6 @@
---
name: hyperframes-animation
description: "All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus Lottie, Three.js, Anime.js, CSS keyframes, Web Animations API, TypeGPU). Use for any motion or animation task: pick 2-4 rules and compose, or load a blueprint, or look up runtime-specific API (e.g. GSAP eases / Lottie player / Three.js mixer). HyperFrames-native: single paused timeline, seek-safe, deterministic."
description: "All animation knowledge for HyperFrames — atomic motion rules, multi-phase scene blueprints, scene transitions, broader motion-design techniques, AND the seven runtime adapters (GSAP default, plus Lottie, Three.js, Anime.js, CSS keyframes, Web Animations API, TypeGPU). Use for any motion or animation task: pick 2-4 rules and compose, or load a blueprint, or look up runtime-specific API (e.g. GSAP eases / Lottie player / Three.js mixer). Also covers auditing an existing composition's choreography (animation map) and 24 named text-animation effects. HyperFrames-native: single paused timeline, seek-safe, deterministic."
---
# HyperFrames Animation
@@ -6,54 +6,135 @@ Built-in eases: `power1`, `power2`, `power3`, `power4`, `back`, `bounce`, `circ`
Each has `.in`, `.out`, `.inOut` variants.
| Ease | Use for |
| -------------------------- | ----------------------------------------------------------------------- |
| `power1.out`, `power2.out` | Standard UI motion. Default for most entrances. |
| `power3.out`, `power4.out` | Punchier deceleration. Title cards, hero reveals. |
| `sine.inOut` | Long, slow, calm motion. Crossfades, ambient drift. |
| `back.out(1.7)` | Slight overshoot. Playful entrances. The arg controls overshoot amount. |
| `elastic.out(1, 0.3)` | Springy bounce. First arg = amplitude, second = period. |
| `expo.inOut` | Snappy, dramatic. Quick transitions between hero scenes. |
| `none` (linear) | Camera moves with timed counterpoint, mechanical motion. |
| Ease | Use for |
| ------------------------------------------ | ----------------------------------------------------------------------------------------------- |
| `power1.out`, `power2.out` | Gentle motion for secondary elements (a caption fade, a small shift). NOT the entrance default. |
| `power3.out` (house default), `power4.out` | The standard long-tail settle. Entrances, title cards, hero reveals. |
| `sine.inOut` | Long, slow, calm motion. Crossfades, ambient drift. |
| `back.out(1.7)` | Overshoot then settle. RARE — explicitly-playful register only, never a default. |
| `elastic.out(1, 0.3)` | Springy bounce. Same playful-only rule; prefer a baked spring (see Spring Eases below). |
| `expo.inOut` | Snappy, dramatic. Quick transitions between hero scenes. |
| `none` (linear) | Camera moves with timed counterpoint, mechanical motion. |
Pick `.out` for entrances, `.in` for exits, `.inOut` for symmetric moves and continuous motion.
**Smooth beats bouncy** — the motion doctrine (`rules/spring-pop-entrance.md`, the workflows' `motion-language.md`): entrances default to `power3.out` or the baked critically-damped spring (see Spring Eases below); overshoot eases (`back` / `elastic` / `bounce`) are a rare, explicitly-playful register, never the house style.
## Easing Vocabulary (character & mood)
Easings are tone of voice: a video that only whispers is boring; one that varies between whisper, normal, and punch is engaging. Every composition should use at least 3 different easings — `power2.out` for everything produces flat, monotonous motion.
Easings are tone of voice: a video that only whispers is boring; one that varies between whisper, normal, and punch is engaging. A composition should draw on ~3 easing characters across its beats — but vary **within the smooth families by energy** (`sine` / `power1` calm → `power3` standard → `power4` / `expo` punch); don't reach for overshoot to add variety. Overshoot is a _register_ (explicitly playful), not a spice. One ease everywhere reads flat; bounce everywhere reads cheap — the second failure is worse.
The full palette by character (each family has `.in`, `.out`, `.inOut` variants):
| Family | Character | Typical use |
| -------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| `power1`–`power4` | Gentle (1) to aggressive (4) acceleration curves | General purpose. power2 is the workhorse, power4 for dramatic snaps |
| `back(N)` | Overshoot then settle. N controls how far past the target (1=subtle, 4=wild) | Logo reveals, badge pops, card entrances. `back.out(2.5)` for playful, `back.out(1.2)` for elegant |
| `elastic(amp, freq)` | Spring bounce. amp=magnitude, freq=oscillation speed | Panel scatter, energetic drops, fun reveals |
| `bounce` | Ball-drop bouncing | Physical interactions, icons landing, score counters |
| `expo` | Extreme acceleration curve (much steeper than power4) | Premium/luxury reveals, dramatic entrances |
| `sine` | Smooth, organic, no hard edges | Ambient float, breathing, Ken Burns, anything that loops. `.inOut` for yoyo motion |
| `circ` | Circular acceleration (starts very fast, ends very gentle or vice versa) | Camera moves, scene transitions, orbital motion |
| `steps(N)` | Discrete N-step jumps, no interpolation | Typing effects, cursor blink, counter ticks, retro/digital aesthetics |
| Family | Character | Typical use |
| -------------------- | ---------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `power1`–`power4` | Gentle (1) to aggressive (4) acceleration curves | General purpose. **power3 is the house workhorse**; power2 for gentle secondary motion, power4 for dramatic snaps |
| `back(N)` | Overshoot then settle. N controls how far past the target (1=subtle, 4=wild) | RARE — explicitly-playful register only, never a default. Keep N ≤ 2; prefer a baked spring at ζ 0.6–0.7 (physical settle, see Spring Eases) |
| `elastic(amp, freq)` | Spring bounce. amp=magnitude, freq=oscillation speed | RARE — same playful-only rule; the baked spring (below) is the physical version |
| `bounce` | Ball-drop bouncing | RARE — physical-comedy register only (something literally dropping) |
| `expo` | Extreme acceleration curve (much steeper than power4) | Premium/luxury reveals, dramatic entrances |
| `sine` | Smooth, organic, no hard edges | Ambient float, breathing, Ken Burns, anything that loops. `.inOut` for yoyo motion |
| `circ` | Circular acceleration (starts very fast, ends very gentle or vice versa) | Camera moves, scene transitions, orbital motion |
| `steps(N)` | Discrete N-step jumps, no interpolation | Typing effects, cursor blink, counter ticks, retro/digital aesthetics |
**Mood mapping:** Match easing character to the beat's emotional content. Smooth/organic easings (`sine`, `power1`) feel contemplative and drifting. Aggressive deceleration (`power4.out`, `expo.out`) feels snappy and confident. Spring overshoot (`back.out`) feels bouncy and physical. The storyboard's mood description should guide which character fits — not a formula.
**Mood mapping:** Match easing character to the beat's emotional content. Smooth/organic easings (`sine`, `power1`) feel contemplative and drifting. Aggressive deceleration (`power4.out`, `expo.out`) feels snappy and confident. Spring overshoot (`back.out`) feels bouncy and physical — but bouncy is a register, not an emphasis tool; reach for it only on explicitly-playful beats. The storyboard's mood description should guide which character fits — not a formula.
## Defaults
```javascript
const tl = gsap.timeline({
paused: true,
defaults: { duration: 0.6, ease: "power2.out" },
defaults: { duration: 0.6, ease: "power3.out" }, // the house settle — smooth beats bouncy
});
```
Or globally:
```javascript
gsap.defaults({ duration: 0.6, ease: "power2.out" });
gsap.defaults({ duration: 0.6, ease: "power3.out" });
```
Setting defaults at timeline scope is preferred — it documents the motion language of that composition in one place.
## Spring Eases (baked physics, seek-safe)
The "iOS feel" is a **damped spring's velocity curve**, not a bounce: a fast launch into a long asymptotic settle. Well-made system animations are critically damped or close to it — they barely overshoot, or don't at all. `power3.out` / `expo.out` approximate that curve; when you want the exact one — or a _physical_ overshoot for the rare playful register — bake the spring's closed-form solution into a function ease.
Why not a real-time spring library: an interactive spring is a stateful integrator (velocity accumulates frame to frame), which cannot be seeked deterministically — you'd have to simulate frames 0…N−1 to render frame N. The closed form below is a **pure function of progress** — no state, nothing to desync, seek-safe by construction. This is also why interaction-lib spring solvers are banned in compositions.
```javascript
// springEase — a damped spring's exact position curve as a GSAP ease.
// response ≈ seconds one oscillation would take (0.3–0.6 for entrances)
// dampingFraction 1.0 = critically damped — smooth settle, NO overshoot (house default)
// 0.80–0.85 ≈ the iOS system register — ~1–1.5% overshoot, felt not seen
// 0.60–0.70 = explicitly playful — ~5–10% overshoot (rare; replaces back.out)
function springEase({ response = 0.5, dampingFraction = 1 } = {}) {
const w = (2 * Math.PI) / response; // undamped natural frequency
const z = dampingFraction;
let pos; // x(t): 0 → 1, starting at rest (v0 = 0)
if (z < 1) {
const wd = w * Math.sqrt(1 - z * z);
pos = (t) => 1 - Math.exp(-z * w * t) * (Math.cos(wd * t) + ((z * w) / wd) * Math.sin(wd * t));
} else if (z > 1) {
const wo = w * Math.sqrt(z * z - 1);
pos = (t) =>
1 - Math.exp(-z * w * t) * (Math.cosh(wo * t) + ((z * w) / wo) * Math.sinh(wo * t));
} else {
pos = (t) => 1 - Math.exp(-w * t) * (1 + w * t);
}
// Settle time: last moment the curve sits outside ±0.1% of target.
// Fixed-step scan, runs once at setup — deterministic (no Math.random / Date.now).
const EPS = 0.001;
const rate = z <= 1 ? z * w : (z - Math.sqrt(z * z - 1)) * w; // slowest decay mode
const SCAN = 12 / rate;
const N = 4800;
let T = SCAN;
for (let i = N; i >= 0; i--) {
const t = (i / N) * SCAN;
if (Math.abs(1 - pos(t)) > EPS) {
T = ((i + 1) / N) * SCAN;
break;
}
}
const xT = pos(T);
return {
duration: T, // use as the tween's duration — the settle time IS the physics
ease: (p) => pos(p * T) + p * (1 - xT), // normalized so ease(1) === 1 exactly
};
}
```
Usage — take **both** the ease and the duration from the helper (the settle time is part of the physics; overriding the duration just re-times the same curve, so tune speed via `response` instead):
```javascript
const settle = springEase({ response: 0.4 }); // critically damped → duration ≈ 0.59s
tl.fromTo(
"#hero",
{ scale: 0, opacity: 0 },
{ scale: 1, opacity: 1, duration: settle.duration, ease: settle.ease },
0.2,
);
```
| dampingFraction | overshoot | register |
| ----------------- | --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| **1.0 (default)** | none (monotone) | The house settle — the exact curve `power3.out` approximates. Product / enterprise / serious tone. |
| 0.80–0.85 | ~1–1.5% | "Alive, not bouncy" — the iOS system default register. The overshoot is felt, not seen. |
| 0.60–0.70 | ~5–10% | Explicitly-playful ONLY (same rule as `back.out`, which this replaces — a spring's second-order settle reads physical where `back` reads cartoon). |
| < 0.55 | > 12% | Don't. Cartoon-wobble territory. |
| response | duration (ζ=1) | feel |
| --------- | -------------- | ------------------------------------------------------------ |
| 0.25–0.35 | 0.37–0.51s | tight snap — chips, small UI |
| 0.35–0.50 | 0.51–0.74s | standard entrance |
| 0.50–0.70 | 0.74–1.03s | weighted hero landing — check the `t ≤ 0.5s` visibility rule |
Craft notes:
- **ζ=1 vs `power3.out`**: the true spring front-loads harder (~67% vs ~58% travelled at quarter-time) and settles on a longer asymptotic tail; max shape difference ~11%. That long tail is the "premium" read — use it when the settle IS the shot (a wordmark landing, a final lockup).
- **At ζ<1, overshooting curves go on transforms only** — never on `opacity` (it would push past 1) or color. Split opacity onto its own `power2.out` tween at the same timeline position.
- **Doctrine unchanged**: ζ below ~0.8 is still the rare, explicitly-playful exception (`rules/spring-pop-entrance.md`). The default of this section is ζ=1 — real spring physics is not a license for bounce.
## Stagger
```javascript
@@ -7,7 +7,7 @@ HyperFrames is a seek-driven runtime. Build one paused timeline per composition,
```javascript
const tl = gsap.timeline({
paused: true,
defaults: { duration: 0.5, ease: "power2.out" },
defaults: { duration: 0.5, ease: "power3.out" },
});
tl.to(".a", { x: 100 }).to(".b", { y: 50 }).to(".c", { opacity: 0 });
@@ -280,7 +280,7 @@ Ease family — discrete choice:
## Pairs with HF skills
- `/hyperframes-animation` — single driver, multi-element envelope
- `/hyperframes-media` — `hyperframes transcribe` outputs real ASR data
- `/hyperframes-media` — pair with caption rendering
- `/media-use` — `hyperframes transcribe` outputs real ASR data
- `/media-use` — pair with caption rendering
- `/hyperframes-core` — composition wiring
- `/hyperframes-cli` — `hyperframes lint`
@@ -175,7 +175,7 @@ tl.fromTo(
);
```
Keep `OVERSHOOT` modest even here (≤ ~2) — past that it reads as a cartoon wobble, not an arrival.
Keep `OVERSHOOT` modest even here (≤ ~2) — past that it reads as a cartoon wobble, not an arrival. Better still: the baked spring at `dampingFraction: 0.6–0.7` (`../adapters/gsap-easing-and-stagger.md` → Spring Eases) gives ~5–10% overshoot with a second-order settle that reads physical where `back.out` reads cartoon.
### Origin-anchored pop (callout springs from a pointer / source)
@@ -195,6 +195,7 @@ When a popped element then **holds** an ongoing slot (a constellation node, a pe
- **EASE** — the settle curve (the load-bearing decision)
- Default: **`power3.out`** — a smooth long-tail settle, no overshoot; the house style for product / enterprise / serious tone. Use `expo.out` for a punchier, faster-front arrival (still smooth).
- Exact-physics option: `springEase({ response: 0.4 })` (critically damped, ζ=1) from `../adapters/gsap-easing-and-stagger.md` → Spring Eases — the curve `power3.out` approximates, with a harder front and a longer settle tail; take `duration` from the helper. Use when the settle IS the shot (a wordmark landing, a final lockup).
- Playful exception only: `back.out(OVERSHOOT)` — see the Bouncy pop variation; reach for it only when a bounce is clearly the brand intent.
- **OVERSHOOT** — `back.out(OVERSHOOT)` overshoot strength — **only used in the rare bouncy variant**; the smooth default has no overshoot dial
@@ -11,8 +11,7 @@
//
// Env:
// HYPERFRAMES_SKILL_PKG_VERSION — pin the @hyperframes/producer version used
// when bootstrapping (global skill installs cannot infer it; falls back to
// @latest with a warning otherwise).
// when bootstrapping if the bundled plugin version cannot be inferred.
import { mkdir, writeFile } from "node:fs/promises";
import { resolve, join } from "node:path";
@@ -35,7 +35,7 @@ test("hyperframesPackageSpec: resolvable in-repo version pins it", async () => {
}
});
// (c) unresolvable + no override -> fail closed rather than bootstrapping @latest.
// (c) unresolvable + no override -> fail closed rather than bootstrapping a floating version.
// Copy the loader into an isolated temp dir whose ancestor chain has no hyperframes
// package.json, and run node from there so cwd cannot resolve one either.
test("hyperframesPackageSpec: unresolvable requires an explicit version override", () => {
@@ -1,6 +1,6 @@
---
name: hyperframes-cli
description: HyperFrames CLI dev loop. Use when running npx hyperframes init, add, catalog, capture, lint, validate, inspect, layout, snapshot, preview, play, render, publish, lambda, doctor, browser, info, upgrade, skills, compositions, docs, benchmark, telemetry, transcribe, tts, or remove-background, or when troubleshooting the HyperFrames build/render environment. Entry point for AWS Lambda cloud rendering (`hyperframes lambda deploy / render / progress / destroy / policies`).
description: HyperFrames CLI dev loop. Use when running npx hyperframes init, add, catalog, capture, lint, validate, inspect, layout, snapshot, preview, play, render, publish, feedback, lambda, doctor, browser, info, upgrade, skills, compositions, docs, benchmark, telemetry, transcribe, tts, or remove-background, or when troubleshooting the HyperFrames build/render environment. Entry point for AWS Lambda cloud rendering (`hyperframes lambda deploy / render / progress / destroy / policies / sites`).
---
# HyperFrames CLI
@@ -56,7 +56,7 @@ Cross-cutting rules that hold for every command:
- **Tailwind projects** (`init --tailwind`) → use `hyperframes-core` (Tailwind reference) before editing classes or theme tokens.
- **Registry blocks/components** (`hyperframes add`, `hyperframes catalog`) → use `hyperframes-registry` for install paths, sub-composition wiring, and snippet merging.
- **Asset preprocessing** (`tts`, `transcribe`, `remove-background`) → use `hyperframes-media` for voice selection, Whisper model rules, captions, and TTS-to-captions chain.
- **Asset preprocessing** (`tts`, `transcribe`, `remove-background`) → use `media-use` for voice selection, Whisper model rules, captions, and TTS-to-captions chain.
- **Parametrized renders** (`--variables`) → declared via `data-composition-variables` on `<html>`; see `hyperframes-core` for the full schema.
## Lambda (Cloud Rendering)
@@ -1,3 +0,0 @@
interface:
display_name: "HyperFrames CLI"
short_description: "HyperFrames CLI tool — hyperframes init, lint, preview, render, transcribe, tts, doctor, browser, info, upgrade, compositions, docs, benchmark."
@@ -27,7 +27,7 @@ Other useful flags:
When using `--tailwind`, invoke the `hyperframes-core` (Tailwind reference) skill before editing classes or theme tokens. The scaffold uses Tailwind v4 browser runtime patterns, not Studio's Tailwind v3 setup.
When `--audio` or `--video` is supplied, `init` transcribes the file with Whisper. For voice/model selection see the `hyperframes-media` skill.
When `--audio` or `--video` is supplied, `init` transcribes the file with Whisper. For voice/model selection see the `media-use` skill.
## capture
@@ -72,4 +72,4 @@ npx hyperframes remove-background
These produce assets (narration audio, word-level transcripts, transparent video) that get dropped into a composition. Each may download its own model on first run.
For voice selection, Whisper model rules, output format choice, and the TTS → transcript → captions chain, invoke the `hyperframes-media` skill. This skill stays focused on the dev loop.
For voice selection, Whisper model rules, output format choice, and the TTS → transcript → captions chain, invoke the `media-use` skill. This skill stays focused on the dev loop.
@@ -1,13 +1,13 @@
---
name: hyperframes-core
description: The HyperFrames composition contract — build one renderable project. Use for composition structure, the `data-*` timing attributes, `class="clip"`, tracks, sub-compositions, variables, framework-owned media playback, deterministic-render rules, and validation. Read before writing composition HTML.
description: The HyperFrames composition contract — build one renderable project. Use for composition structure, the `data-*` timing attributes, `class="clip"`, tracks, sub-compositions, variables, framework-owned media playback, deterministic-render rules, and validation. Also covers Tailwind projects and the STORYBOARD.md / SCRIPT.md plan formats. Read before writing composition HTML.
---
# HyperFrames Core
HyperFrames renders video from HTML. A composition is an HTML file whose DOM declares timing with `data-*` attributes, whose animation runtime is seekable, and whose media playback is owned by the framework.
This skill is the **technical contract** — how to build one hyperframes project. The body below is the build guide; per-topic detail lives in `references/` (index next), read on demand. Other concerns live in the sibling domain skills — `hyperframes-animation`, `hyperframes-creative`, `hyperframes-media`, `hyperframes-cli`, `hyperframes-registry`. The capability map in `/hyperframes` says what each one covers.
This skill is the **technical contract** — how to build one hyperframes project. The body below is the build guide; per-topic detail lives in `references/` (index next), read on demand. Other concerns live in the sibling domain skills — `hyperframes-animation`, `hyperframes-creative`, `media-use`, `hyperframes-cli`, `hyperframes-registry`. The capability map in `/hyperframes` says what each one covers.
## References
@@ -64,6 +64,7 @@ Surfaced here; full rationale in the linked reference. Do not violate:
- Read the files first. Preserve unrelated timing, tracks, IDs, variables, media paths.
- Match existing composition IDs and timeline keys.
- Adding a clip: pick a non-overlapping `data-track-index` or adjust surrounding timing intentionally.
- `data-hidden` on any composition element hides it in BOTH preview and render, overriding its time window; it is non-destructive/reversible and toggled by Studio's timeline eye icon.
- Adding a sub-composition: verify its internal `data-composition-id` before wiring the host.
## Validation
@@ -2,7 +2,7 @@
The **locked narration** for a project: the final spoken lines + voice + delivery. It is an _optional_ plan-layer file — a video with no narration (bgm-only, silent overlay) has none. The storyboard's per-frame `voiceover` is the lighter, editable _guide_; `SCRIPT.md` is the _commit_. (Storyboard format → `references/storyboard-format.md`.)
This file defines the SCRIPT.md **shape** only. Synthesizing the spoken lines into audio is a capability owned by `hyperframes-media` → `references/tts.md`.
This file defines the SCRIPT.md **shape** only. Synthesizing the spoken lines into audio is a capability owned by `media-use` → `references/tts.md`.
Free-form markdown — there is no strict parser; the Studio renders it read-only beside the Storyboard board, and the TTS step extracts the indented spoken lines.
@@ -46,4 +46,4 @@ A header block, then one section per spoken line.
## To TTS
Feed each line's spoken text to `npx hyperframes tts` (pin `--voice` / `--provider` from the header; capture word timestamps for captions). Real per-word timing replaces the `**Time:**` guides. CLI contract → `hyperframes-media/references/tts.md`.
Feed each line's spoken text to `npx hyperframes tts` (pin `--voice` / `--provider` from the header; capture word timestamps for captions). Real per-word timing replaces the `**Time:**` guides. CLI contract → `media-use/audio/references/tts.md`.
@@ -36,7 +36,7 @@ document.documentElement.style.setProperty("--accent", accent);
- Use `npx hyperframes render --variables '{"title":"Q4 Report"}'` or `--variables-file` for render-time overrides.
- Add `--strict-variables` in CI: turns undeclared keys, type mismatches, and enum values not in `options` into errors instead of warnings.
- Read values once during init, not on every animation tick — variables don't change mid-render.
- Media color grading can use exact variable references inside `data-color-grading` JSON. Use `$gradingPreset` or `${gradingIntensity}` as the whole field value; the runtime resolves it from the current composition's variables before applying the shader grading.
- Media color grading can use exact variable references inside `data-color-grading` JSON. Use `$gradingPreset` or `${gradingIntensity}` as the whole field value; the runtime resolves it from the current composition's variables before applying shader adjustments, finishing details, blur/pixelate effects, and custom LUTs.
### Two JSON Shapes (Easy to Confuse)
@@ -1,8 +1,8 @@
---
version: alpha
name: Claude — Frame (video / frame layer)
name: Code Editorial — Frame (video / frame layer)
description: >
Video-first companion to Claude's design.md. The unit is the frame (1920×1080). Atoms are
Video-first companion to The preset's design.md. The unit is the frame (1920×1080). Atoms are
identical and sacred — warm cream paper (never pure white, never cool), a single terracotta coral
as scarce "voltage", hairline ink elevation (no heavy shadow), EB Garamond for all
display + Inter body + JetBrains Mono for the index/code voice on a warm-navy code surface,
@@ -90,11 +90,11 @@ components:
description: "The brand mark. Fades + scales 0.92→1 on a single emphasis beat; never spins."
---
# Claude — Frame (video / frame layer)
# Code Editorial — Frame (video / frame layer)
## Overview
Claude at frame scale is a **warm-editorial brand book come to life** — the register of a literary
Code Editorial at frame scale is a **warm-editorial brand book come to life** — the register of a literary
imprint or a research note. The thesis is three colors: **cream is the ground, ink is the voice,
coral is the voltage** — and a fourth (warm navy) only where code shows itself. Every surface is
**warm cream** (never pure white, never cool gray); content gathers on a **tile** surface half a
@@ -270,7 +270,7 @@ to the diff; commit/issue numbers are chrome.
## Known Gaps
- **Motion intentionally out of scope.** frame.md specifies composition only. Claude's motion register — short cross-dissolves, no overshoot/bounce/elastic, coral the only "draw-on", numbers count up, code types on line by line — lives in the workflow's `motion-language.md` + `hyperframes-animation`, not here.
- **Motion intentionally out of scope.** frame.md specifies composition only. The preset's motion register — short cross-dissolves, no overshoot/bounce/elastic, coral the only "draw-on", numbers count up, code types on line by line — lives in the workflow's `motion-language.md` + `hyperframes-animation`, not here.
- **EB Garamond + Inter + JetBrains Mono are all bundled in the HyperFrames renderer (`@fontsource` embedded data) — they resolve offline by name, with no Google Fonts dependency.** This is deliberate: every face here renders deterministically on a clean machine or AWS Lambda, so a frame needs no captured `.woff2` or `@font-face` for these three. The embedded set ships **weights 400 + 700 only and no true italic** — so author display at **weight 400** (700 reads as a heavy bold, off-register), and treat italic as the browser-synthesized slant (acceptable for the pull-quote register; ship a real EB Garamond italic `.woff2` + `@font-face` only if a project leans hard on it). EB Garamond is a warm old-style serif (low contrast, humanist); if it ever fails, fall to Georgia or another old-style serif — never to a sans. CJK: Noto Serif SC (display) / Noto Sans SC (body) / Noto Sans Mono CJK (code); the sentence-case warmth carries when the serif drops.
- **Syntax colors (teal `#5DB8A6` / amber `#E8A55A` / status) are fixed decoration**, declared in §Colors — they are NOT in the remixable `colors:` block, so a brand remix never repaints them.
- **The code itself is the `code-*` registry blocks**, not this preset — this preset owns only the surrounding warm-navy surface + mono chrome.
@@ -1,5 +1,5 @@
<!--
Claude — caption skin (preset-local source).
Code Editorial — caption skin (preset-local source).
This is the preset's own lower-third karaoke caption look. product-launch-video's
(and pr-to-video's) Step 2 copies it into the project as caption-skin.html;
@@ -10,7 +10,7 @@
· the empty data-brand-tokens style → :root tokens derived from the project's frame.md
Token contract — reference ONLY this fixed vocab (captions.mjs injects it from
frame.md, so the Step-2 brand overlay flows through; each var has a claude literal
frame.md, so the Step-2 brand overlay flows through; each var has a code-editorial literal
fallback for standalone preview):
--cap-ink · --cap-canvas · --cap-accent · --cap-accent-2 · --font-display ·
--font-body · --cap-band-top · --cap-band-height
@@ -20,7 +20,7 @@
never tl.call() callbacks (seek does not fire them → wrong state on render).
Visual: a warm cream card lifted off the page on a 1px hairline ink border + one soft
warm shadow — claude's only elevation language, no tilt, editorial and precise.
warm shadow — the preset's only elevation language, no tilt, editorial and precise.
EB Garamond sentence case, warm-ink glyphs. Upcoming words sit in faint ink; the current
word reads in full ink under a 3px coral "voltage" underline (the brand's inline-link
gesture, border-bottom = underline without layout shift); spoken words settle to solid
@@ -66,7 +66,7 @@
}
/* THE lifted cream card — warm cream fill, 1px hairline ink border, one soft warm
shadow (claude's only elevation language), 12px radius, no tilt */
shadow (the preset's only elevation language), 12px radius, no tilt */
.caption-pill {
max-width: 78%;
padding: 22px 44px 26px;
@@ -3,7 +3,7 @@
<head>
<meta charset="utf-8" />
<meta name="viewport" content="width=device-width,initial-scale=1" />
<title>Claude</title>
<title>Code Editorial</title>
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link
@@ -1422,7 +1422,7 @@
</section>
<div class="foot">
<span class="l"><span class="spike">✱</span> Claude</span>
<span class="l"><span class="spike">✱</span> Code Editorial</span>
<span class="r">Frame Showcase · Atoms Sacred · Composition Free</span>
</div>
@@ -41,7 +41,7 @@ Optionally seed `frame.md` from a ready-made **frame-preset** in `[../frame-pres
| `[broadside](../frame-presets/broadside/FRAME.md)` | Protest-poster system — two-register flat plane (ink-black / fire-orange), massive lowercase Barlow 900 treated as graphic primitive, IBM Plex Mono chrome (uppercase, 0.14em), fire-orange sole accent, 1px hairlines, sharp corners, no shadow | bold / typographic / declarative; a product that wants presence and authority |
| `[capsule](../frame-presets/capsule/FRAME.md)` | Playful editorial — every container a pill (2px ink outline), cream canvas, nine candy accents, Bodoni Moda + Space Grotesk, soft offset shadows, floating-pill wallpaper | friendly / soft / editorial; a product that wants warmth and approachability |
| `[cartesian](../frame-presets/cartesian/FRAME.md)` | Museum-catalog editorial — 1px taupe hairline grid, five warm-stone palette (#EDE8E0 / #E2DBD1 / #1A1A1A / #5A5A5A / #8A8178), Playfair Display 400 + Inter, sharp corners, compass-drafted geometric rings, zero shadow / zero fill | sparse / literary / restrained; a product that wants quiet authority and editorial rigor |
| `[claude](../frame-presets/claude/FRAME.md)` | Warm-editorial brand book — warm cream paper (never pure white), terracotta coral (#CC785C) as scarce voltage, hairline ink elevation (no heavy shadow), EB Garamond serif display + Inter body + JetBrains Mono index / code on a warm-navy code surface, sentence-case display, ✱ coral spike | considered / literary / developer-facing; a code change, launch, or doc that wants editorial calm and a first-class code surface |
| `[code-editorial](../frame-presets/code-editorial/FRAME.md)` | Warm-editorial brand book — warm cream paper (never pure white), terracotta coral (#CC785C) as scarce voltage, hairline ink elevation (no heavy shadow), EB Garamond serif display + Inter body + JetBrains Mono index / code on a warm-navy code surface, sentence-case display, ✱ coral spike | considered / literary / developer-facing; a code change, launch, or doc that wants editorial calm and a first-class code surface |
| `[cobalt-grid](../frame-presets/cobalt-grid/FRAME.md)` | Modernist two-color risograph — cream paper, electric cobalt ink, permanent graph-paper grid (10% cobalt), top/bottom cobalt hairlines, Newsreader 400 serif + Hanken Grotesk + DM Mono, 0° corners, pixel-glitch column + QR-block patches | restrained / systemic / editorial; a product that wants clarity and measured authority |
| `[coral](../frame-presets/coral/FRAME.md)` | Bold editorial magazine — three solid surfaces (coral fire / ink black / warm cream) meeting at hard edges, 45° diagonal hatch on coral, Bebas Neue + Inter tracked caps, zero shadows/radius (circles 50%), oversized wallpaper numerals and giant marks | bold / structuralist / editorial; a product that wants graphic confidence and hard-edged confidence |
| `[creative-mode](../frame-presets/creative-mode/FRAME.md)` | Neo-brutalist editorial — cream canvas, 4px ink borders, hard offset shadows (no blur), four accents rationed two-to-three, Archivo Black uppercase at 0.92 line-height, JetBrains Mono taxonomy, Space Grotesk body, square corners save one pill chip | sparse / graphic / punchy-restrained; a product that wants editorial presence and geometric confidence |
@@ -14,8 +14,7 @@
//
// Env:
// HYPERFRAMES_SKILL_PKG_VERSION — pin the @hyperframes/producer version used
// when bootstrapping (global skill installs cannot infer it; falls back to
// @latest with a warning otherwise).
// when bootstrapping if the bundled plugin version cannot be inferred.
//
// The composition directory must contain an index.html. Raw authoring HTML
// works — the producer's file server auto-injects the runtime at serve time.
@@ -35,7 +35,7 @@ test("hyperframesPackageSpec: resolvable in-repo version pins it", async () => {
}
});
// (c) unresolvable + no override -> fail closed rather than bootstrapping @latest.
// (c) unresolvable + no override -> fail closed rather than bootstrapping a floating version.
// Copy the loader into an isolated temp dir whose ancestor chain has no hyperframes
// package.json, and run node from there so cwd cannot resolve one either.
test("hyperframesPackageSpec: unresolvable requires an explicit version override", () => {
@@ -3,7 +3,7 @@ name: hyperframes-keyframes
description: >
Use when a HyperFrames composition needs seek-safe 2D/3D keyframes, GSAP
timelines, CSS keyframes, Anime.js, WAAPI, FLIP, paths, masks, SVG morph/draw,
text trails, cursor demos, 3D depth, or `hyperframes keyframes` diagnostics.
text trails, 3D depth, or `hyperframes keyframes` diagnostics.
Don't use for broad scene strategy, brand design, media sourcing, captions, or
general video planning.
---
@@ -1,97 +0,0 @@
---
name: hyperframes-media
description: Audio and media assets for HyperFrames compositions, produced by one shared audio engine (`scripts/audio.mjs`) — multi-provider TTS (HeyGen / ElevenLabs / Kokoro local), background music + sound effects (HeyGen audio-library retrieval by default, with local Lyria / MusicGen BGM generation and a bundled SFX library as the no-credential fallback), Whisper transcription, background removal, and caption authoring. Use for voiceover / TTS, BGM, SFX / sound effects, transcription, captions / subtitles / lyrics / karaoke / per-word styling, voice + provider selection, and music-mood prompting.
---
# HyperFrames Media
Create the audio and media assets a composition needs — voiceover (TTS), background music + sound effects, transcription, captions, background removal — then consume and animate that data in HTML. For placing assets into compositions, see `hyperframes-core`.
## The audio engine — one source for TTS · BGM · SFX
Workflows do NOT hand-roll audio or vendor a copy. There is one engine — **`scripts/audio.mjs`** — that takes a neutral `audio_request.json` and writes `audio_meta.json` (plus assets under `assets/voice|bgm|sfx`):
```bash
# <MEDIA_DIR> = this skill's directory
node <MEDIA_DIR>/scripts/audio.mjs --request ./audio_request.json --hyperframes . --out ./audio_meta.json
```
All three capabilities degrade on **ONE switch** — whether a HeyGen credential is present (resolved from `$HEYGEN_API_KEY` / `$HYPERFRAMES_API_KEY` / `~/.heygen`, **not** the CLI):
| Capability | HeyGen credential present | absent |
| ---------- | -------------------------------------------------- | ---------------------------------------------------- |
| TTS | HeyGen Starfish REST (native word timestamps) | → ElevenLabs → Kokoro (chain `transcribe` for words) |
| BGM | HeyGen music **retrieval** | Lyria → MusicGen local **generation** (detached) |
| SFX | HeyGen sound-effects **retrieval** (min_score 0.4) | bundled 21-file library (`assets/sfx/`) |
- **Request** (`audio_request.json`): `{ provider?, lang?, speed?, lines: [{ id, text, sfx?: [names] }], bgm: { mode?, query?, prompt? } }`. `id` joins each line back to the caller's model (a frame number, a scene id, …). `bgm.mode` = `retrieve | generate | none`; omit for auto (retrieve when credentialed, else generate). An **explicit** `retrieve` is strict — it skips rather than starting a detached generate (for callers with no `wait-bgm` step).
- **Output** (`audio_meta.json`, id-keyed): `{ tts_provider, voice_id, bgm, bgm_pending, …, voices: [{ id, path, duration_s, words }], sfx: [{ id, name, file, source, offset_s, duration_s, volume }], total_duration_s }`.
- `--only tts,bgm,sfx` runs a subset and **merges** into an existing `--out` (e.g. TTS+BGM early, SFX once cues exist).
- BGM generate is spawned **detached** (`bgm_pending: true`) — run `scripts/wait-bgm.mjs` before assembling.
- `scripts/heygen-tts.mjs` is a single-shot CLI over the same code (one text → wav + words) for when you just need HeyGen TTS without a request file.
Full flag list + the `audio_meta.json` schema live in the header of `scripts/audio.mjs`. The references below cover the provider details and edge cases behind each capability.
## Preflight — show sign-in status before any audio
**Always run this before generating voice or BGM — inside a full workflow _or_ a one-off "generate me a BGM/voiceover" request.** No HeyGen credential is **not** a reason to silently fall back to local engines: first recommend signing in and let the user decide. Run the shared preflight and **relay its output verbatim** — don't improvise your own "missing key" prompt, and don't offer to write keys into a per-repo `.env`:
```bash
npx hyperframes auth status
```
- **Signed in** → it prints the account; proceed.
- **Not signed in** (`exit 1` is expected here — "not signed in" is a normal state, not a failure) → it prints registration-first guidance. Recommend signing in: `npx hyperframes auth login` is browser OAuth — it **signs in and creates an account** (always available through this repo's CLI). To use an existing HeyGen API key (from app.heygen.com/settings/api), run `npx hyperframes auth login --api-key` — it saves to the shared `~/.heygen` (no per-repo `.env`). The output also lists the local engines voice/BGM will fall back to and a `pip` hint when deps are missing. **Relay this output as-is — don't paraphrase it into your own wording.** Then **STOP and wait** for the user to choose — sign in, or say "go" / "local" to continue offline — **before generating anything.** This is a real decision point, not a passing note: don't fold it into another question, and don't proceed past it on your own. (Exception: in autonomous / non-interactive mode, note the status and continue offline.)
- `npx hyperframes auth status --json` returns `{ configured, recommended_action, offline_engines }` for deterministic branching.
- **If the CLI can't run** (not on PATH and `npx` can't fetch it) → still **recommend signing in** (`npx hyperframes auth login`) and **STOP for the user's choice** — don't treat "no credential" as a silent green light for local generation.
Credential resolution, full key priority, and the local-dependency list are in `references/requirements.md`.
## Provider chains (the detail behind the engine)
**TTS** — first available provider wins (the engine, or `npx hyperframes tts "..."`):
| Order | Provider | Detected when | Word timestamps |
| ----- | ----------------------------- | -------------------------------------------- | ---------------------------------------------------------------- |
| 1 | HeyGen (Starfish) | `$HEYGEN_API_KEY` / `hyperframes auth login` | **Yes, native** — pass `--words narration.words.json` to capture |
| 2 | ElevenLabs | `$ELEVENLABS_API_KEY` set | No — chain `transcribe` after |
| 3 | Kokoro-82M (local, 54 voices) | always (no key required) | No — chain `transcribe` after |
> The published `hyperframes tts` CLI is often the local-only build (its `--help` says "Kokoro-82M", no `--provider`/`--words`) and silently falls back to Kokoro even with `$HEYGEN_API_KEY` set. That is why the engine's HeyGen path is the self-contained `scripts/heygen-tts.mjs` (REST), NOT the CLI; the CLI is used only for the Kokoro path. See `references/tts.md`.
**BGM & SFX** — by default **retrieved** from the HeyGen audio library (`/v3/audio/sounds`), same credential as HeyGen TTS, with the no-credential fallback from the switch above:
| Asset | HeyGen `type` | Lands in | Fallback (no credential) |
| ----- | ------------------------------- | ---------------------------------------------------------- | ---------------------------------------------------------- |
| BGM | `music` | `assets/bgm/track.mp3` (retrieve) · `track.wav` (generate) | Lyria / MusicGen generation |
| SFX | `sound_effects` (min_score 0.4) | `assets/sfx/<slug>.mp3` | bundled 21-file library (`assets/sfx/*` + `manifest.json`) |
See `references/bgm.md` and `references/sfx.md`.
## Routing
| Task | Read |
| ------------------------------------------------------------------- | -------------------------------------------- |
| The audio engine — request/meta schema, `--only`, the switch | `scripts/audio.mjs` (header comment) |
| `npx hyperframes tts` / `heygen-tts.mjs` — providers, voices, words | `references/tts.md` |
| BGM — HeyGen retrieval + local Lyria / MusicGen generation | `references/bgm.md` |
| SFX — HeyGen retrieval (min_score 0.4) + bundled local library | `references/sfx.md` |
| `npx hyperframes transcribe` — Whisper, model rules, output shape | `references/transcribe.md` |
| `npx hyperframes remove-background` — transparent cutouts | `references/remove-background.md` |
| TTS → transcription → captions (no recorded voiceover) | `references/tts-to-captions.md` |
| Caption authoring — style detection, layout, word grouping, exit | `references/captions/authoring.md` |
| Transcript handling — input formats, quality gates, cleanup, APIs | `references/captions/transcript-handling.md` |
| Caption motion — karaoke, marker effects, audio-reactive | `references/captions/motion.md` |
| Model caches, system dependencies, troubleshooting | `references/requirements.md` |
## Non-negotiable rules
- **One engine, no vendored copies.** Produce audio via `scripts/audio.mjs` (or `heygen-tts.mjs` for one-shot HeyGen TTS). Don't re-implement TTS/BGM/SFX inside a workflow — write an `audio_request.json` adapter and call the engine.
- **"HeyGen available" = a resolvable credential, not the CLI.** The whole switch keys off `heygenCredential()`; the published `hyperframes tts` may be Kokoro-only, and there is no `hyperframes bgm` / `hyperframes sfx` command at all.
- **Voice IDs are provider-specific.** `am_michael` is Kokoro-only; HeyGen UUIDs don't work on Kokoro. If you pass `--voice`, also pin `--provider` to avoid silent provider drift when the user's env changes.
- **Always pass `--model` to `transcribe`.** The CLI default `small.en` silently translates non-English audio. See `references/transcribe.md` → "Language Rule".
- **HeyGen returns word timestamps; ElevenLabs / Kokoro do not.** The engine chains `transcribe` automatically for the latter two; standalone, pass `--words` to HeyGen or run `transcribe` against the audio file.
- **Captions consume the flat word-array format** with `{ id, text, start, end }`. See `references/transcribe.md` → "Output Shape".
- **`remove-background --background-output` is hole-cut, not inpainted.** For "scene without the person", a different tool is needed. See `references/remove-background.md` → "When NOT the right tool".
- **BGM/SFX default to HeyGen retrieval; the no-credential fallback is generation (BGM) or the bundled library (SFX).** `/audio/sounds` ranks by a text query — name effects concretely (`glass shatter`, not `dramatic sound`); a no-match **skips**, never blocks the render. SFX sit at volume ~0.35 under voice + BGM. See `references/sfx.md` / `references/bgm.md`.
- **Treat workflow caption HTML as generated output.** For preset-backed videos, the reusable skin source lives at `.hyperframes/caption-skin.html` and the workflow script writes `compositions/captions.html`; do not edit generated `compositions/captions.html` to fix the skin. Rebuild via the workflow's `captions.mjs`, or use that workflow's explicit overrides mechanism when present.
@@ -1,3 +0,0 @@
interface:
display_name: "HyperFrames Registry"
short_description: "Install and wire registry blocks and components into HyperFrames compositions."
+21 -21
View File
@@ -3,7 +3,7 @@ name: hyperframes
description: >
READ THIS FIRST for any request to make, create, edit, animate, or render a
video, animation, or motion graphic — a promo, explainer, captioned clip,
title card, overlay, or any composition. HyperFrames renders video from HTML;
title card, overlay, slideshow / interactive deck, or any composition. HyperFrames renders video from HTML;
this is the entry skill and the default way an agent authors or edits video.
It routes the request to the right specialized workflow and points to the
HyperFrames domain skills, so read it before any other video or animation
@@ -19,11 +19,11 @@ metadata: { "tags": "read-first, video, animation, router, hyperframes, intent-r
HyperFrames **renders video from HTML** — a composition is an HTML file whose DOM declares timing with `data-*` attributes, whose animation runtime is seekable, and whose media playback is owned by the framework. The full authoring contract lives in `/hyperframes-core`; read it before writing composition HTML.
Below: a **capability map** (the domain skills, loaded on demand) and the **intent router** (pick a workflow for any "make me a video" request).
Below: a **capability map** (the domain skills, loaded on demand) and the **intent router** (pick a workflow for any "make me a…" request — usually a video, but also a navigable deck or a composition port). The split is ownership, not output type: a **workflow owns an end-to-end deliverable** (its own project dir, gated steps, sub-agents, final artifact); a **domain skill is a capability layer** a workflow pulls in mid-flight and never owns the task.
## Capability map — the domain skills
Atomic capabilities you load **on demand** — not full video workflows. For "make me a video", use the intent router below.
Atomic capabilities you load **on demand** — not full workflows; they never own the end-to-end task. For "make me a…" intent, use the intent router below.
| You want to… | Skill |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------ |
@@ -31,10 +31,10 @@ Atomic capabilities you load **on demand** — not full video workflows. For "ma
| **Animate** — atomic motion, scene blueprints, transitions, runtime adapters (GSAP / Lottie / Three.js / Anime.js / CSS / WAAPI / TypeGPU) | `/hyperframes-animation` |
| **Author seek-safe keyframes** — GSAP timelines, CSS keyframes, Anime.js, WAAPI, FLIP, paths, masks, SVG morph/draw, 3D depth, plus `hyperframes keyframes` diagnostics | `/hyperframes-keyframes` |
| **Creative direction** — `frame.md` / `design.md`, palettes, typography, narration, beat planning, audio-reactive | `/hyperframes-creative` |
| **Media** — TTS voiceover, background music, transcription, background removal, captions | `/hyperframes-media` |
| **Media resolve** — find + freeze BGM, SFX, images, icons from HeyGen catalog into `.media/` with manifest tracking | `/media-use` |
| **Media** — resolve/generate BGM, SFX, image, icon, voice; TTS voiceover, transcription, background removal, captions; cross-project reuse | `/media-use` |
| **CLI dev loop** — init, lint, validate, inspect, preview, render, publish, doctor | `/hyperframes-cli` |
| **Install registry blocks / components** (`hyperframes add`) | `/hyperframes-registry` |
| **Import Figma content** — assets, tokens, components, storyboards→reconstructed motion (REST/CLI); optional connector-assisted motion/shader paths | `/figma` |
---
@@ -52,19 +52,19 @@ Routing needs to know **what the video is about** — its input and subject. If
## Workflow cheat-sheet
| Workflow | Use it for |
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/product-launch-video` | **Selling a product** (SaaS, app, company / product site) — from a URL, brief, or script → a **promo**. The default for any commercial URL, even if the site is only named. |
| `/website-to-video` | **Showing a site itself** — a tour / showcase built from the site's own screenshots. For non-commercial sites (portfolio, blog, docs, personal, event), or when the user wants a tour, not a promo. |
| `/faceless-explainer` | **Explaining a topic / concept** from text — no product, no URL; every visual is LLM-invented |
| `/pr-to-video` | A **GitHub PR / code change** → changelog / feature-reveal / fix / refactor explainer |
| `/embedded-captions` | Adding **captions / subtitles** to an existing talking-head video (footage untouched) |
| `/talking-head-recut` | Packaging an existing talking-head video with **designed graphic overlays** — lower-thirds, data callouts, kinetic titles, pull-quotes |
| `/motion-graphics` | A short, **unnarrated, design-led motion graphic** — kinetic type, a stat / chart hit, a logo sting, a lower-third overlay |
| `/music-to-video` | A **music track** → a **beat-synced** video — lyric video, slideshow, or kinetic promo; the music drives pacing (optional user images / videos cut onto the beat grid) |
| `/slideshow` | A **presentation / pitch deck / interactive deck** — discrete slides, fragments, branching, hotspots; output is a navigable **deck**, not a rendered video |
| `/general-video` | **Anything else** — longer or multi-scene pieces, a static loop / poster, a custom composition |
| `/remotion-to-hyperframes` | **Porting an existing Remotion (React) composition** to HyperFrames (migration, not creation) |
| Workflow | Use it for |
| -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `/product-launch-video` | **Selling a product** (SaaS, app, company / product site) — from a URL, brief, or script → a **promo**. The default for any commercial URL, even if the site is only named. |
| `/website-to-video` | **Showing a site itself** — a tour / showcase built from the site's own screenshots. For non-commercial sites (portfolio, blog, docs, personal, event), or when the user wants a tour, not a promo. |
| `/faceless-explainer` | **Explaining a topic / concept** from text — no product, no URL; every visual is LLM-invented |
| `/pr-to-video` | A **GitHub PR / code change** → changelog / feature-reveal / fix / refactor explainer |
| `/embedded-captions` | Adding **captions / subtitles** to an existing talking-head video (footage untouched) |
| `/talking-head-recut` | Packaging an existing talking-head video with **designed graphic overlays** — lower-thirds, data callouts, kinetic titles, pull-quotes |
| `/motion-graphics` | A short (~under 10s), **unnarrated** piece where the **motion _is_ the message** — kinetic type, a stat / chart hit, a logo sting, an animated map, an animated tweet / headline, a **standalone** lower-third / overlay (MP4 or transparent alpha) |
| `/music-to-video` | A **music track** → a **beat-synced** video — lyric video, slideshow, or kinetic promo; the music drives pacing (optional user images / videos cut onto the beat grid) |
| `/slideshow` | A **presentation / pitch deck / interactive deck** — discrete slides, fragments, branching, hotspots; output is a navigable **deck**, not a rendered video |
| `/general-video` | **Anything else** — longer or multi-scene pieces, a static loop / poster, a custom composition |
| `/remotion-to-hyperframes` | **Porting an existing Remotion (React) composition** to HyperFrames (migration, not creation) |
**Disambiguation (only where confusable):**
@@ -123,8 +123,8 @@ The CLI also surfaces a one-line reminder when a `render` / `lint` / `validate`
### `/embedded-captions`
- **Input:** An existing **talking-head video** (MP4) to caption — actual footage, not a URL or brief. Transcribed locally (Whisper, no API key) and matted (RVM) so the subject can occlude captions.
- **Output:** the same footage **untouched**, with a caption layer — **Standard** (verbatim lower-third rail + an embedded climax behind the subject) or **Cinematic** (every caption composited behind the subject). Any length.
- **Input:** An existing **talking-head video** (MP4) to caption — actual footage, not a URL or brief. Transcribed and matted locally (no API key) so the subject can occlude captions.
- **Output:** the same footage **untouched**, with a caption layer — one visual identity picked from its catalog (36, from a quiet verbatim rail to full VFX constitutions); the subject occludes the embedded captions. Any length.
- **Triggers:** "add captions / subtitles to this video", "captions behind the subject", "cinematic captions for my clip".
### `/talking-head-recut`
@@ -135,7 +135,7 @@ The CLI also surfaces a one-line reminder when a `render` / `lint` / `validate`
### `/motion-graphics`
- **Input:** A short, design-led motion graphic where the **motion is the message** — typically under ~10s, no narration. Genres: kinetic typography, a stat / number count-up, a chart hit, a logo sting, a lower-third / overlay, or a search-driven page / tweet / headline shot.
- **Input:** A short, design-led motion graphic where the **motion is the message** — typically under ~10s, no narration. Genres: kinetic typography, a stat / number count-up, a chart hit, a logo sting, a lower-third / overlay, an animated map (regions / routes / zoom-to-place), a search-driven page / tweet / news-article shot, or asset-fusion (a real image's geometry becomes the chart).
- **Output:** a short motion graphic → MP4 or a **transparent overlay** (alpha WebM / MOV) for a lower-third / callout.
- **Triggers:** "an 8s logo sting", "animate this stat", "a kinetic-type intro", "turn this tweet into a motion graphic", "a transparent lower-third overlay".
@@ -1,3 +0,0 @@
interface:
display_name: "HyperFrames"
short_description: "Create video compositions, animations, title cards, overlays, captions, voiceovers, audio-reactive visuals, and scene transitions in HyperFrames HTML."
+129 -30
View File
@@ -1,15 +1,32 @@
---
name: media-use
description: Agent Media OS — resolve any media need (BGM, SFX, image, icon) into a frozen local file + ledger record. One verb (`resolve`) handles the full cascade — project cache, global cache, HeyGen catalog search, freeze, register. Keeps search noise on disk, hands the agent a path. Use when a composition needs background music, sound effects, images, or icons.
description: Agent Media OS, the single skill for every media need in a HyperFrames project. Resolve BGM, SFX, image, icon, or voice into a frozen local file + ledger record (one verb, `resolve`); generate via TTS / music / image models when the catalog misses; produce voiceover, transcription, captions, and background removal through one shared audio engine; operate on media (cut / reframe / transform); and reuse assets across projects. Keeps search noise on disk, hands the agent a path. Use for any audio, image, icon, voiceover, caption, or media-asset need.
---
# media-use
Resolve media needs into frozen local files. One verb, four types, zero context noise.
The media OS for HyperFrames: resolve · generate · operate · remember, every media type, one skill, zero context noise.
## What it owns (the gaps HyperFrames leaves)
HyperFrames owns media _playback_; media-use owns everything else. Each row is enforced by `scripts/lib/coverage.test.mjs` so the claim can't rot.
| HyperFrames gap | media-use owns it via |
| ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------- |
| Audio-only, no image/icon | `resolve --type image\|icon` (heygen asset search) |
| No voice / audio generation | `resolve --type voice` + the audio engine (`audio/scripts/audio.mjs`) |
| Scattered/duplicated audio engine | one consolidated engine under `audio/` (hyperframes-media retired) |
| No agent media-ops (cut/reframe/transform) | `references/operations.md` + `resolve --from` to register outputs |
| No transcript-driven cutting | `scripts/transcript-cut.mjs` compiles word-timestamp edits into cut lists |
| No auto-duck / publish loudness | `scripts/audio-duck.mjs` + `references/operations.md` loudnorm/sidechain recipes |
| No cross-project memory | global content-addressed cache + auto-promote (`~/.media`) |
| No image generation | RAM-graded local mflux (FLUX) via `scripts/lib/mflux-provider.mjs`, codex `image_gen` upsell (`scripts/lib/codex-provider.mjs`) |
| No video generation | spec-gated local LTX (`videogen` in `scripts/lib/local-models.mjs`); `heygen video create` avatar upsell |
| Weak local-model defaults | free-usage HeyGen first (TTS, bg-removal) via the `heygen` CLI; local open-source only as an opt-out fallback (`scripts/lib/local-run.mjs`) |
## When to use
Call `resolve` whenever a composition needs media — background music, sound effects, images, or icons. media-use searches the HeyGen catalog, downloads the best match, freezes it locally, and registers it in a manifest. The agent gets back one line; all search noise stays on disk.
Call `resolve` whenever a composition needs media: background music, sound effects, images, icons, or voice. For voiceover / TTS, music, SFX, and caption timing, use the **audio engine** (below); background removal is delegated to the `hyperframes` CLI; transcription defaults to Parakeet (better than whisper.cpp: 6.05% vs 7.44% WER, 5-10x faster) via `scripts/transcribe.mjs`, with whisper.cpp auto-fallback (see `references/operations.md`). For cutting / reframing / transforming existing media, see `references/operations.md`. media-use searches the HeyGen catalog first, freezes the best match locally, registers it in a manifest, and hands the agent one line; all search noise stays on disk.
## Resolve
@@ -27,6 +44,7 @@ Returns one line: `resolved <id> → <path> (<type>, <metadata>)`
| `sfx` | Sound effects | Bundled 19-file library + HeyGen catalog |
| `image` | Photos, backgrounds | HeyGen asset search (75k+ vectors) |
| `icon` | Icons, logos | HeyGen asset search (type=icon) |
| `voice` | TTS voiceover | Local Kokoro (free); HeyGen TTS upsell |
### Examples
@@ -50,14 +68,46 @@ node <SKILL_DIR>/scripts/resolve.mjs --type icon --intent "rocket" --project .
### Flags
| Flag | Description |
| --------------- | ------------------------------------------ |
| `--type, -t` | Media type: bgm, sfx, image, icon |
| `--intent, -i` | What you need (natural language) |
| `--entity, -e` | Entity name for cache matching (optional) |
| `--project, -p` | Project directory (default: .) |
| `--adopt` | Bulk-import existing assets/ into manifest |
| `--json` | Output JSON instead of one-line result |
| Flag | Description |
| --------------- | --------------------------------------------------------------- |
| `--type, -t` | Media type: bgm, sfx, image, icon, voice |
| `--intent, -i` | What you need (natural language) |
| `--entity, -e` | Entity name for cache matching (optional) |
| `--project, -p` | Project directory (default: .) |
| `--from` | Freeze a local file or direct public URL (ingest) |
| `--local-only` | Offline: skip every network provider (cache + local only) |
| `--provider` | Force one generator (e.g. `codex`, `mflux`, `kokoro`, `heygen`) |
| `--adopt` | Bulk-import existing assets/ into manifest |
| `--json` | Output JSON instead of one-line result |
## Providers
media-use holds no keys; every external tool owns its auth. Generation is
local-first with a cloud upsell where one helps. `resolve` spec-checks
AVAILABLE RAM and auto-picks the best local model that fits (a RAM-graded
ladder, `describeModelLadder`); the agent can see the ladder and override.
| Type | Provider (in order) |
| ------------- | --------------------------------------------------------------------------------------- |
| bgm/sfx | heygen catalog (free) |
| image | heygen search, then local mflux (best FLUX for your RAM), then codex `image_gen` upsell |
| voice | local **Kokoro** (free, on-device), then **heygen tts** paid upsell |
| icon | heygen asset search |
| video (local) | local LTX (`videogen` ladder); `heygen video create` avatar upsell |
Local Kokoro (voice), mflux (image), and LTX (video) run on-device (free,
private, offline once cached). Paid/cloud upsells sit behind them: HeyGen TTS
for voice, the `codex` CLI (ChatGPT sub) for a better image, the `heygen` CLI
for avatar video. Cost rule (X4): the agent confirms before an agent-initiated
paid call; a user-requested one just runs.
To force a specific generator (e.g. a user says "make this image with codex"),
pass `--provider codex`: it pins resolution to that provider and skips the
free-first default. See `references/operations.md` for the RAM ladders and
upsell recipes.
`--local-only` skips every network provider, including the free HeyGen ones,
leaving the project + global cache and any local provider.
## How it works
@@ -90,35 +140,84 @@ After resolve or adopt, read `.media/index.md` for the full inventory:
# .media · 4 assets
id type dur dims path description
bgm_001 bgm 25s — .media/audio/bgm/bgm_001.mp3 upbeat tech launch
sfx_001 sfx 0.6s — .media/audio/sfx/sfx_001.mp3 whoosh
image_001 image — 1920×1080 .media/images/image_001.jpg gradient tech background
icon_001 icon — 200×200 .media/images/icon_001.png rocket
bgm_001 bgm 25s - .media/audio/bgm/bgm_001.mp3 upbeat tech launch
sfx_001 sfx 0.6s - .media/audio/sfx/sfx_001.mp3 whoosh
image_001 image - 1920×1080 .media/images/image_001.jpg gradient tech background
icon_001 icon - 200×200 .media/images/icon_001.png rocket
```
## Cross-project reuse
Assets are cached automatically on resolve. Subsequent resolves for the same prompt hit the global cache at `~/.media/` — no re-download, no provider call. Promote an asset explicitly with `organize --promote <id>` to make it reusable across all projects.
Assets are cached automatically on resolve. Every resolved/ingested asset is auto-promoted to the global cache at `~/.media/`, so subsequent resolves for the same prompt, in any project, hit the cache with no re-download and no provider call.
## Files
- `.media/manifest.jsonl` — machine SSOT, one JSON record per line
- `.media/index.md` — agent-readable table (id, type, dur, dims, path, description)
- `~/.media/` — global cross-project reuse cache (content-addressed, SHA-256)
- `.media/manifest.jsonl`: machine SSOT, one JSON record per line
- `.media/index.md`: agent-readable table (id, type, dur, dims, path, description)
- `~/.media/`: global cross-project reuse cache (content-addressed, SHA-256)
## CLI tools used
## Audio engine: voiceover, music, SFX, captions, transcription
| Tool | Purpose | Required? |
| --------- | ------------------------------------------ | ------------- |
| `ffprobe` | Probe duration, dimensions, codec on adopt | Yes |
| `heygen` | Audio catalog, asset search | For providers |
Install the `heygen` CLI with the official installer for your environment and authenticate outside the conversation:
For a full audio pass (TTS voiceover + background music + sound effects in one
shot), use the shared engine at `audio/scripts/audio.mjs`. It takes a neutral
`audio_request.json` and writes `audio_meta.json` plus assets under
`.media/audio/{voice,bgm,sfx}`:
```bash
heygen --version # confirm it is on PATH
heygen update # needs >= v0.1.6
heygen auth login # or set HEYGEN_API_KEY in your shell
node <SKILL_DIR>/audio/scripts/audio.mjs --request ./audio_request.json --out ./audio_meta.json
```
Requires **heygen >= v0.1.6** — the providers tag requests with the allowlisted `--headers 'X-HeyGen-Client-Source: media-use'` flag, added in v0.1.6. `asset search` is a pre-launch command hidden from `heygen --help`, but it runs. Without a `heygen` on PATH (or a valid key) the providers print a one-line diagnostic to stderr and resolve falls through to "no provider could resolve".
- **Request** `{ provider?, lang?, speed?, lines: [{ id, text, sfx?: [names] }], bgm: { mode?, query?, prompt? } }`: `id` joins each line back to your model; `bgm.mode` = `retrieve | generate | none` (omit for auto). `--only tts,bgm,sfx` runs a subset and merges into an existing `--out`.
- **Output** `audio_meta.json` (id-keyed): `voices[].{path,duration_s,words[]}` (word timestamps for captions), `sfx[]`, `bgm`, `total_duration_s`.
- **Auto-degrades on one switch**: HeyGen credential present → HeyGen TTS + music/SFX retrieval; absent → ElevenLabs/Kokoro TTS, Lyria/MusicGen BGM generation, and the bundled SFX library (no credential needed).
- If BGM took the generate path (`bgm_pending: true`), run `audio/scripts/wait-bgm.mjs` before final render.
Single-shot helpers: `audio/scripts/heygen-tts.mjs` (one voice file). Transcription / background removal / captions use the `hyperframes` CLI (`transcribe`, `remove-background`), see the per-topic guides in `audio/references/` (`tts.md`, `bgm.md`, `sfx.md`, `transcribe.md`, `remove-background.md`, `captions/`).
## Operating on media (cut, reframe, transform)
media-use resolves + remembers; for **operating** on assets see
`references/operations.md`: local-tool recipes (ffmpeg trim/reframe/montage,
auto-editor, scenedetect) and the local-vs-HeyGen transform table (background
removal, upscale, lipsync, translate). Run the tool, then register the output
with `resolve --from <output> --type <type>` so it joins the ledger + global
cache.
## CLI tools used (what to run, and how to enable each)
`resolve` auto-cascades; each provider shells one CLI. Local tools are OPT-IN:
if a local tool is absent, resolve degrades gracefully to the free/cloud path,
so nothing here is strictly required except `ffmpeg`/`ffprobe`. Install a local
tool to unlock its free, private, on-device path. media-use holds no keys.
| Tool | Serves | Install |
| ------------------ | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------- |
| `ffmpeg`/`ffprobe` | adopt probing, cut, duck bake, loudnorm | system package (`brew install ffmpeg`) |
| `heygen` | catalog (bgm/sfx/image/icon), TTS + avatar upsell | Install with the official installer for your environment, then run `heygen auth login` or set `HEYGEN_API_KEY` in your shell (needs >= v0.1.6) |
| `mflux-generate` | local image gen (FLUX), best-for-RAM | `uv venv ~/.venvs/mflux && VIRTUAL_ENV=~/.venvs/mflux uv pip install mflux==0.9.6` |
| `codex` | image gen upsell (ChatGPT sub) | Codex CLI, logged in via ChatGPT (owns its own auth) |
| `parakeet-mlx` | local transcription (default ASR, best) | `uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx` |
| `ltx-2-mlx` | local video gen | `git clone https://github.com/dgrauet/ltx-2-mlx && cd ltx-2-mlx && uv sync --all-extras` |
| `npx hyperframes` | Kokoro TTS (voice), whisper.cpp (transcribe fallback), remove-background | bundled with the hyperframes CLI |
The RAM-graded local-model shortlist + exact per-tier install/invoke lives in
`scripts/lib/local-models.mjs` (the agent can read `describeModelLadder(cap, specs)`
to see which model fits this machine). Without a tool on PATH, its provider
prints a one-line diagnostic to stderr and resolve falls through to the next
provider (e.g. no `mflux` -> codex image upsell; no `parakeet-mlx` -> whisper.cpp).
`heygen asset search` is a pre-launch command hidden from `heygen --help`, but it
runs; providers tag requests with the allowlisted `X-HeyGen-Client-Source` header
(v0.1.6+).
## Telemetry
`resolve` and the edit tools (transcribe / transcript-cut / audio-duck) send an
anonymous usage event to PostHog (`scripts/lib/telemetry.mjs`), so we can see
which capabilities are actually used. It records only the media TYPE, the
resolution SOURCE, and the winning PROVIDER: never the intent text, file names,
or paths, and `$ip:null` so no IP is stored. Best-effort and non-blocking (a
resolve never waits on or fails from telemetry).
Opt out with `DO_NOT_TRACK=1` or `HYPERFRAMES_NO_TELEMETRY=1` (also off in CI and
dev). Same public PostHog project key and opt-outs as the `hyperframes` CLI.
@@ -44,11 +44,11 @@ expired OAuth token it stops with a hint to run `npx hyperframes auth refresh`.
export HEYGEN_API_KEY=... # or put it in a project .env
# Synthesize + capture word timestamps in one call (skips a Whisper pass)
node skills/hyperframes-media/scripts/heygen-tts.mjs \
node skills/media-use/audio/scripts/heygen-tts.mjs \
"Welcome to HyperFrames." -o narration.wav --words narration.words.json
node skills/hyperframes-media/scripts/heygen-tts.mjs ./script.txt -o narration.wav
node skills/hyperframes-media/scripts/heygen-tts.mjs --list # public starfish voices
node skills/media-use/audio/scripts/heygen-tts.mjs ./script.txt -o narration.wav
node skills/media-use/audio/scripts/heygen-tts.mjs --list # public starfish voices
```
- **Voice:** `--voice <id>` must be a **starfish** voice_id (`--list`, or `GET /v3/voices?engine=starfish`). v2-catalog ids are rejected with HTTP 400. Omit `--voice` (English) and it defaults to **Marcia** (`05f19352e8f74b0392a8f411eba40de1`, a fixed default so the choice is deterministic). Non-English with no `--voice` falls back to the first matching catalog voice.
@@ -11,7 +11,7 @@
//
// TTS : HeyGen REST → ElevenLabs → Kokoro (CLI)
// BGM : HeyGen retrieve → (no credential) Lyria/MusicGen generate
// SFX : HeyGen retrieve → (no credential) bundled 21-file library
// SFX : HeyGen retrieve → (no credential) bundled 19-file library
//
// ── audio_request.json (input) ────────────────────────────────────────────────
// {
@@ -53,6 +53,7 @@ import {
} from "./lib/tts.mjs";
import { generateBgmDetached, inferBgmPrompt, retrieveBgm } from "./lib/bgm.mjs";
import { resolveSfx } from "./lib/sfx.mjs";
import { mapWithConcurrency } from "./lib/concurrency.mjs";
const HERE = dirname(fileURLToPath(import.meta.url));
const argv = process.argv.slice(2);
@@ -67,6 +68,16 @@ const die = (m) => {
};
const r3 = (x) => Number(x.toFixed(3));
// Two independent reports of an unbounded Promise.all over TTS lines
// overwhelming a machine: one OOM'd 12/13 concurrent Kokoro TTS +
// whisper-transcribe lines on a resource-constrained laptop, the other saw
// 7/8 lines fail on first run (concurrent cold-start model loads) and pass on
// retry once the model was cached. Kokoro/Whisper each load their own local
// model per subprocess, so firing every line at once multiplies that cost by
// the line count. mapWithConcurrency caps how many run at once — still
// parallel, just bounded.
const ttsConcurrency = Math.max(1, Number(process.env.HYPERFRAMES_TTS_CONCURRENCY) || 4);
const hyperframesDir = resolve(flag("hyperframes", "."));
const requestPath = resolve(flag("request", join(hyperframesDir, "audio_request.json")));
const outPath = resolve(flag("out", join(hyperframesDir, "audio_meta.json")));
@@ -156,7 +167,7 @@ if (only.has("tts") && lines.length) {
}
return { id, path: rel, duration_s: r3(dur), words: withWordIds(wordArr) };
};
const results = await Promise.all(lines.map(synthLine));
const results = await mapWithConcurrency(lines, ttsConcurrency, synthLine);
voices = results.filter(Boolean);
for (const v of voices)
console.error(` voice ${v.id}: ${v.path} (${v.duration_s}s, ${v.words.length} words)`);
@@ -0,0 +1,49 @@
import { strict as assert } from "node:assert";
import { test } from "node:test";
import { mkdtempSync, rmSync, existsSync } from "node:fs";
import { join, dirname } from "node:path";
import { tmpdir } from "node:os";
import { fileURLToPath } from "node:url";
import { resolveSfx } from "./lib/sfx.mjs";
// Proves the relocated engine (skills/media-use/audio/) still resolves its
// bundled SFX library from the moved location — the path most likely to break
// on a subtree move. Offline (heygenOK:false), no network.
const HERE = dirname(fileURLToPath(import.meta.url));
const sfxLibDir = join(HERE, "..", "assets", "sfx"); // same offset the engine uses
test("bundled SFX library resolves from the relocated path", async () => {
assert.ok(existsSync(join(sfxLibDir, "manifest.json")), "moved manifest is present");
const dir = mkdtempSync(join(tmpdir(), "mu-audio-"));
try {
const { sfx, anomalies } = await resolveSfx({
cues: [{ id: "1", name: "whoosh" }],
heygenOK: false,
hyperframesDir: dir,
sfxLibDir,
});
assert.equal(sfx.length, 1, `expected 1 resolved cue, got anomalies: ${anomalies.join("; ")}`);
assert.equal(sfx[0].source, "local");
assert.match(sfx[0].file, /assets\/sfx\//);
assert.ok(existsSync(join(dir, sfx[0].file)), "matched SFX copied into the project");
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
test("an unknown cue is reported, not fatal", async () => {
const dir = mkdtempSync(join(tmpdir(), "mu-audio-"));
try {
const { sfx, anomalies } = await resolveSfx({
cues: [{ id: "1", name: "definitely-not-a-real-sfx" }],
heygenOK: false,
hyperframesDir: dir,
sfxLibDir,
});
assert.equal(sfx.length, 0);
assert.ok(anomalies.some((a) => /not in bundled library/.test(a)));
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
@@ -15,6 +15,7 @@ import { spawn, spawnSync } from "node:child_process";
import { existsSync, mkdirSync, openSync, closeSync } from "node:fs";
import { join } from "node:path";
import { downloadTo, searchSounds } from "./heygen.mjs";
import { pythonInvocation } from "./python.mjs";
const r3 = (x) => Number(x.toFixed(3));
const lyriaKey = () => process.env.GEMINI_API_KEY || process.env.GOOGLE_API_KEY || "";
@@ -26,10 +27,18 @@ const LYRIA_PY_DEPS = ["google-genai", "python-dotenv"];
const LYRIA_PY_PROBE = "import google.genai";
function pyOk(probe) {
return spawnSync("python3", ["-c", probe], { stdio: "ignore" }).status === 0;
const { cmd, args } = pythonInvocation(["-c", probe]);
return spawnSync(cmd, args, { stdio: "ignore" }).status === 0;
}
// `python -m pip`, not a bare `pip` binary: a Homebrew/system Python often
// exposes only `python3`/`pip3` on PATH, so a plain `pip` spawn silently
// no-ops (ENOENT) and the documented "auto-installed on demand" path never
// actually installs. `-m pip` also guarantees the packages land in the SAME
// interpreter pyOk() probes — a bare `pip`/`pip3` could resolve to a
// different Python installation than `python3` if more than one is on PATH.
function pipInstall(deps) {
return spawnSync("pip", ["install", "-q", ...deps], { stdio: "ignore" }).status === 0;
const { cmd, args } = pythonInvocation(["-m", "pip", "install", "-q", ...deps]);
return spawnSync(cmd, args, { stdio: "ignore" }).status === 0;
}
// ── retrieval (HeyGen music library) ──────────────────────────────────────────
@@ -120,11 +129,16 @@ export function generateBgmDetached({
const fd = openSync(log, "w");
if (useLyria) {
const proc = spawn(
"python3",
[lyriaRecipe, "--output", abs, "--duration", String(targetS), "--prompt", prompt],
{ detached: true, stdio: ["ignore", fd, fd] },
);
const { cmd, args } = pythonInvocation([
lyriaRecipe,
"--output",
abs,
"--duration",
String(targetS),
"--prompt",
prompt,
]);
const proc = spawn(cmd, args, { detached: true, stdio: ["ignore", fd, fd] });
proc.unref();
closeSync(fd);
return {
@@ -141,7 +155,8 @@ export function generateBgmDetached({
const seedS = Math.min(Math.max(seedSeconds, 10), 30);
const loops = targetS > seedS ? Math.ceil(targetS / seedS) : 1;
const script = musicgenScript({ prompt, abs, targetS, seedS });
const proc = spawn("python3", ["-c", script], { detached: true, stdio: ["ignore", fd, fd] });
const { cmd, args } = pythonInvocation(["-c", script]);
const proc = spawn(cmd, args, { detached: true, stdio: ["ignore", fd, fd] });
proc.unref();
closeSync(fd);
return {
@@ -0,0 +1,14 @@
// mapWithConcurrency — run `fn` over `items` with at most `limit` in flight at
// once. Preserves input order in the result array regardless of completion order.
export async function mapWithConcurrency(items, limit, fn) {
const results = new Array(items.length);
let next = 0;
async function worker() {
while (next < items.length) {
const i = next++;
results[i] = await fn(items[i], i);
}
}
await Promise.all(Array.from({ length: Math.min(limit, items.length) }, worker));
return results;
}
@@ -0,0 +1,41 @@
import { test } from "node:test";
import assert from "node:assert/strict";
import { mapWithConcurrency } from "./concurrency.mjs";
// Regression: audio.mjs used a bare Promise.all(lines.map(synthLine)) to
// synthesize every TTS line at once, spawning one Kokoro/whisper model load
// per line concurrently. Two independent reports of this overwhelming a
// machine (OOM, and cold-start contention causing spurious failures).
// mapWithConcurrency is the extracted cap; test it in isolation since
// audio.mjs itself is a script (runs CLI/exit side effects on import).
test("processes every item and preserves input order regardless of completion order", async () => {
const order = [5, 1, 3, 2, 4];
const results = await mapWithConcurrency(order, 2, async (n) => {
await new Promise((r) => setTimeout(r, n));
return n * 10;
});
assert.deepEqual(results, [50, 10, 30, 20, 40]);
});
test("never runs more than `limit` at once", async () => {
let inFlight = 0;
let maxInFlight = 0;
const items = Array.from({ length: 10 }, (_, i) => i);
await mapWithConcurrency(items, 3, async () => {
inFlight++;
maxInFlight = Math.max(maxInFlight, inFlight);
await new Promise((r) => setTimeout(r, 5));
inFlight--;
});
assert.equal(maxInFlight, 3);
});
test("limit larger than the item count runs everything without hanging", async () => {
const results = await mapWithConcurrency([1, 2], 10, async (n) => n * 2);
assert.deepEqual(results, [2, 4]);
});
test("empty input resolves to an empty array", async () => {
const results = await mapWithConcurrency([], 4, async (n) => n);
assert.deepEqual(results, []);
});
@@ -1,7 +1,6 @@
// heygen.mjs — vendored HeyGen REST helpers (auth + transport) for the audio
// pipeline. The credential resolver is copied from hyperframes-media's
// heygen-tts.mjs (and matches the hyperframes CLI auth): first usable source
// wins — $HEYGEN_API_KEY / $HYPERFRAMES_API_KEY → a nearby .env → ~/.heygen/
// pipeline. The credential resolver matches the hyperframes CLI auth: first
// usable source wins — $HEYGEN_API_KEY / $HYPERFRAMES_API_KEY → a nearby .env → ~/.heygen/
// credentials (oauth → Bearer, else api_key → X-Api-Key; $HEYGEN_CONFIG_DIR
// overrides the dir). Vendored so the skill ships standalone. Pure node.
@@ -0,0 +1,63 @@
// python.mjs — resolve which Python 3 executable to spawn, per platform.
//
// The audio engine (tts.mjs, bgm.mjs) shells out to `python3` for ElevenLabs
// TTS and the local Lyria/MusicGen BGM paths. `python3` is the right name on
// macOS/Linux, but on Windows the python.org installer only creates
// `python.exe` plus the `py` launcher — there is no `python3.exe` (only the
// Microsoft Store build adds one). So a bare `spawn("python3", …)` ENOENTs on a
// standard Windows Python install, silently disabling every Python-backed audio
// feature until the user hand-creates a `python3.exe` shim (reported twice).
//
// Resolve once, per process: probe the platform's candidates in order and take
// the first that actually runs. `py` is the launcher, so it needs a `-3` arg to
// select Python 3 — hence candidates are argv PREFIXES, not bare names.
import { spawnSync } from "node:child_process";
function defaultProbe(cmd, args) {
try {
return spawnSync(cmd, args, { stdio: "ignore" }).status === 0;
} catch {
return false;
}
}
/**
* Pick the argv prefix that launches Python 3 on this platform.
* Returns e.g. `["python3"]`, `["python"]`, or `["py", "-3"]`.
*
* Pure except for `probe` (which runs `<cmd> … --version`); both `platform`
* and `probe` are injectable so every branch is unit-testable without spawning.
* If nothing probes OK, falls back to the canonical name for the platform so
* the eventual spawn fails loudly exactly as it did before — never worse.
*/
export function resolvePythonCommand(platform = process.platform, probe = defaultProbe) {
const candidates =
platform === "win32" ? [["python3"], ["python"], ["py", "-3"]] : [["python3"], ["python"]];
for (const prefix of candidates) {
if (probe(prefix[0], [...prefix.slice(1), "--version"])) return prefix;
}
return candidates[0];
}
let cached = null;
/** Cached `resolvePythonCommand()` — probing spawns, so resolve at most once. */
export function pythonCommand() {
if (!cached) cached = resolvePythonCommand();
return cached;
}
/**
* Build a `{ cmd, args }` for running Python 3 with `extraArgs`, using the
* resolved (or supplied) prefix. Keeps the launcher's `-3` (and any future
* prefix args) ahead of the caller's own arguments.
*/
export function pythonInvocation(extraArgs, prefix = pythonCommand()) {
return { cmd: prefix[0], args: [...prefix.slice(1), ...extraArgs] };
}
/** Test-only: clear the cached resolution so a test can re-probe. */
export function _resetPythonCommandCacheForTests() {
cached = null;
}
@@ -0,0 +1,68 @@
import { test } from "node:test";
import assert from "node:assert/strict";
import { resolvePythonCommand, pythonInvocation } from "./python.mjs";
// Regression: on Windows a standard python.org install has no `python3.exe`
// (only `python.exe` + the `py` launcher), so `spawn("python3", …)` ENOENTs and
// every Python-backed audio feature silently no-ops. resolvePythonCommand takes
// injectable platform/probe params so all branches are testable without
// spawning a real interpreter.
// probeFor(names): a probe that reports success only for the given argv-0 names.
function probeFor(...names) {
const ok = new Set(names);
return (cmd) => ok.has(cmd);
}
test("non-win32 uses python3 when it runs", () => {
assert.deepEqual(resolvePythonCommand("linux", probeFor("python3")), ["python3"]);
assert.deepEqual(resolvePythonCommand("darwin", probeFor("python3")), ["python3"]);
});
test("win32 prefers python3 when the Microsoft Store build provides it", () => {
assert.deepEqual(resolvePythonCommand("win32", probeFor("python3", "python", "py")), ["python3"]);
});
test("win32 falls back to python.exe when python3 is absent (python.org install)", () => {
// The exact reported scenario: no python3, but `python` exists.
assert.deepEqual(resolvePythonCommand("win32", probeFor("python", "py")), ["python"]);
});
test("win32 falls back to the py launcher with -3 when only py exists", () => {
assert.deepEqual(resolvePythonCommand("win32", probeFor("py")), ["py", "-3"]);
});
test("py launcher is probed as `py -3 --version`, not bare `py`", () => {
const seen = [];
const probe = (cmd, args) => {
seen.push([cmd, ...args]);
return cmd === "py";
};
resolvePythonCommand("win32", probe);
assert.deepEqual(seen.at(-1), ["py", "-3", "--version"]);
});
test("falls back to the canonical name (loud failure, unchanged) when nothing runs", () => {
// No interpreter anywhere — must not throw, and must return python3 so the
// eventual spawn fails exactly as it did before this fix, never worse.
assert.deepEqual(
resolvePythonCommand("win32", () => false),
["python3"],
);
assert.deepEqual(
resolvePythonCommand("linux", () => false),
["python3"],
);
});
test("pythonInvocation prepends the resolved prefix ahead of caller args", () => {
assert.deepEqual(pythonInvocation(["-c", "import x"], ["python"]), {
cmd: "python",
args: ["-c", "import x"],
});
// The py launcher's -3 must stay ahead of the caller's own arguments.
assert.deepEqual(pythonInvocation(["-c", "import x"], ["py", "-3"]), {
cmd: "py",
args: ["-3", "-c", "import x"],
});
});
@@ -113,7 +113,22 @@ export async function resolveSfx({ cues, heygenOK, headers, hyperframesDir, sfxL
const src = join(sfxLibDir, hit.file);
const destRel = `assets/sfx/${hit.file}`;
const dest = join(hyperframesDir, destRel);
if (existsSync(src) && !existsSync(dest)) copyFileSync(src, dest);
// The bundled library may be incomplete: some installs of the skill ship
// manifest.json without the actual mp3s. Pushing an sfx entry that points at
// a file we never copied produces a dangling reference that silently drops
// downstream ("not on disk"). Surface it as a loud anomaly and skip the cue
// instead, so the audio_meta never references a missing file.
if (!existsSync(dest)) {
if (!existsSync(src)) {
anomalies.push(
`sfx "${name}" (id ${id}): bundled file ${hit.file} missing from the offline ` +
`library (${sfxLibDir}) — skipped. Reinstall the media-use skill to ` +
`restore assets/sfx/*.mp3, or configure a HeyGen credential for retrieval.`,
);
continue;
}
copyFileSync(src, dest);
}
sfx.push({
id,
name,
@@ -0,0 +1,69 @@
import { test } from "node:test";
import assert from "node:assert/strict";
import { mkdtempSync, mkdirSync, writeFileSync, existsSync, rmSync } from "node:fs";
import { join } from "node:path";
import { tmpdir } from "node:os";
import { resolveSfx } from "./sfx.mjs";
// Offline (no HeyGen) SFX resolution: the bundled library may ship manifest.json
// without the actual mp3s. The old code copied only when the source existed but
// pushed the sfx entry unconditionally — producing a dangling reference that
// silently dropped downstream ("not on disk"). These tests lock in the loud
// behavior: a present file is copied + referenced; a missing file yields an
// anomaly and NO dangling entry.
async function withDirs(fn) {
const root = mkdtempSync(join(tmpdir(), "hf-sfx-"));
const libDir = join(root, "lib");
const projDir = join(root, "proj");
mkdirSync(libDir, { recursive: true });
mkdirSync(projDir, { recursive: true });
try {
// `await` is load-bearing: without it the finally cleanup runs before the
// async test body resolves, deleting the temp dir mid-assertion.
return await fn({ libDir, projDir });
} finally {
rmSync(root, { recursive: true, force: true });
}
}
test("offline: copies and references a present bundled file", async () => {
await withDirs(async ({ libDir, projDir }) => {
writeFileSync(
join(libDir, "manifest.json"),
JSON.stringify({ whoosh: { file: "whoosh.mp3", duration: 0.8 } }),
);
writeFileSync(join(libDir, "whoosh.mp3"), "ID3-fake-bytes");
const { sfx, anomalies } = await resolveSfx({
cues: [{ id: "s1", name: "whoosh" }],
heygenOK: false,
hyperframesDir: projDir,
sfxLibDir: libDir,
});
assert.equal(sfx.length, 1);
assert.equal(sfx[0].file, "assets/sfx/whoosh.mp3");
assert.equal(sfx[0].source, "local");
assert.ok(existsSync(join(projDir, "assets/sfx/whoosh.mp3")), "mp3 copied into project");
assert.equal(anomalies.length, 0);
});
});
test("offline: a matched-but-missing bundled file yields an anomaly and NO dangling entry", async () => {
await withDirs(async ({ libDir, projDir }) => {
// Manifest names whoosh.mp3, but the mp3 was never shipped (the reported bug).
writeFileSync(
join(libDir, "manifest.json"),
JSON.stringify({ whoosh: { file: "whoosh.mp3", duration: 0.8 } }),
);
const { sfx, anomalies } = await resolveSfx({
cues: [{ id: "s1", name: "whoosh" }],
heygenOK: false,
hyperframesDir: projDir,
sfxLibDir: libDir,
});
assert.equal(sfx.length, 0, "no dangling entry for a file that was never copied");
assert.equal(anomalies.length, 1);
assert.match(anomalies[0], /missing from the offline library/);
assert.ok(!existsSync(join(projDir, "assets/sfx/whoosh.mp3")), "nothing copied");
});
});
@@ -18,6 +18,7 @@ import { existsSync, mkdirSync, mkdtempSync, readFileSync, rmSync, writeFileSync
import { tmpdir } from "node:os";
import { dirname, join } from "node:path";
import { heygenAuthHeaders, heygenCredential, heygenJSON } from "./heygen.mjs";
import { pythonInvocation } from "./python.mjs";
// ── provider detection ────────────────────────────────────────────────────────
export function heygenAvailable() {
@@ -25,7 +26,10 @@ export function heygenAvailable() {
}
export function elevenlabsAvailable() {
if (!process.env.ELEVENLABS_API_KEY) return false;
const r = spawnSync("python3", ["-c", "import elevenlabs"], { stdio: "ignore" });
const { cmd, args } = pythonInvocation(["-c", "import elevenlabs"]);
const r = spawnSync(cmd, args, {
stdio: "ignore",
});
return r.status === 0;
}
@@ -72,7 +76,34 @@ export async function resolveVoiceId({ provider, userVoice, lang = "en" }) {
// ── helpers ─────────────────────────────────────────────────────────────────
export function withWordIds(words) {
return (words ?? []).map((w, i) => ({ id: `w${i}`, text: w.text, start: w.start, end: w.end }));
return (words ?? []).map((w, i) => ({
id: `w${i}`,
text: w.text,
start: w.start,
end: w.end,
}));
}
// `ffmpeg -i <file>` prints a `Duration: HH:MM:SS.ms` line to stderr even
// though it exits non-zero with no output requested. Parsing pulled out as
// a pure function so the ENOENT fallback below can be tested without
// depending on whether ffprobe/ffmpeg are actually installed on the
// machine running the tests.
export function parseFfmpegDurationBanner(stderrText) {
const match = /Duration:\s*(\d+):(\d+):(\d+(?:\.\d+)?)/.exec(stderrText ?? "");
if (!match) return NaN;
const [, hours, minutes, seconds] = match;
return Number(hours) * 3600 + Number(minutes) * 60 + Number(seconds);
}
// Some "essentials"-style ffmpeg distributions (common on Windows) ship
// ffmpeg.exe without ffprobe.exe. ffprobeDuration's caller (audio.mjs)
// otherwise reads a spurious NaN as "the WAV file is corrupt" and drops an
// already-successfully-synthesized TTS line, rather than "the tool for
// measuring it is missing".
function ffmpegDurationFallback(absPath) {
const r = spawnSync("ffmpeg", ["-i", absPath], { encoding: "utf8" });
return parseFfmpegDurationBanner(r.stderr);
}
export function ffprobeDuration(absPath) {
@@ -81,13 +112,85 @@ export function ffprobeDuration(absPath) {
["-v", "error", "-show_entries", "format=duration", "-of", "default=nw=1:nk=1", absPath],
{ encoding: "utf8" },
);
if (r.error?.code === "ENOENT") return ffmpegDurationFallback(absPath);
if (r.status !== 0) return NaN;
return parseFloat(String(r.stdout).trim());
}
function spawnP(cmd, args, opts) {
export function resolveNpxCliFromNpmExecPath(
npmExecPath = process.env.npm_execpath,
pathExists = existsSync,
) {
if (!npmExecPath) return null;
const fileName = npmExecPath.replace(/\\/g, "/").split("/").pop()?.toLowerCase();
const npxCliPath =
fileName === "npx-cli.js" ? npmExecPath : join(dirname(npmExecPath), "npx-cli.js");
return pathExists(npxCliPath) ? npxCliPath : null;
}
export function resolveSpawnCommand(
cmd,
args,
opts = {},
platform = process.platform,
env = process.env,
pathExists = existsSync,
) {
if (cmd !== "npx" || platform !== "win32") {
return { cmd, args, opts: { stdio: "ignore", ...opts } };
}
// On Windows, npx resolves to npx.cmd, which Node cannot execute directly.
// Avoid `shell:true` and the .cmd shim entirely by invoking npm's JS CLI with
// node, preserving request-provided values as argv data instead of shell text.
const npxCliPath = resolveNpxCliFromNpmExecPath(env.npm_execpath, pathExists);
if (!npxCliPath) return null;
return {
cmd: env.npm_node_execpath || process.execPath,
args: [npxCliPath, ...args.map((arg) => String(arg))],
opts: { stdio: "ignore", windowsHide: true, ...opts },
};
}
// `platform`/`spawnFn` params (default process.platform / the real spawn)
// exist so tests can exercise the win32 branch without mocking node:child_process
// (its ESM exports are non-configurable, so mock.method can't patch it).
// One-shot so a whole batch of TTS lines doesn't repeat the same diagnostic.
let _warnedNpxResolution = false;
/** Test-only: reset the one-shot npx-resolution warning latch. */
export function _resetNpxResolutionWarnForTests() {
_warnedNpxResolution = false;
}
export function spawnP(
cmd,
args,
opts = {},
platform = process.platform,
spawnFn = spawn,
env = process.env,
pathExists = existsSync,
) {
const resolved = resolveSpawnCommand(cmd, args, opts, platform, env, pathExists);
if (!resolved) {
// resolveSpawnCommand only returns null for the npx-on-win32 case where
// npm_execpath isn't set (e.g. audio.mjs invoked directly with `node`, not
// through npm/npx). Without this, every call silently returns status:-1 and
// stdio:"ignore" hides why — callers just report "TTS failed - omitted" for
// every line. Surface the real reason once so it's diagnosable.
if (!_warnedNpxResolution) {
_warnedNpxResolution = true;
console.error(
`[media-use] Cannot run "${cmd}" on Windows: npm_execpath is not set, so the ` +
`npx JS CLI can't be located. This happens when this script is run directly with ` +
`\`node\` instead of through npm/npx. Every "${cmd}" call is being skipped. ` +
`Fix: run via \`npx\`/\`npm run\`, or export npm_execpath pointing at your npm-cli.js.`,
);
}
return Promise.resolve({ status: -1 });
}
return new Promise((resolve) => {
const p = spawn(cmd, args, { stdio: "ignore", ...opts });
const p = spawnFn(resolved.cmd, resolved.args, resolved.opts);
p.on("exit", (code) => resolve({ status: code ?? -1 }));
p.on("error", () => resolve({ status: -1 }));
});
@@ -136,11 +239,14 @@ export async function synthesizeOne({
}) {
if (provider === "heygen") return synthesizeHeygen({ text, voiceId, lang, speed, wavAbs });
if (provider === "elevenlabs") {
const r = await spawnP(
"python3",
["-c", ELEVENLABS_PY, writeTmpText(text), voiceId, wavAbs],
{},
);
const { cmd, args } = pythonInvocation([
"-c",
ELEVENLABS_PY,
writeTmpText(text),
voiceId,
wavAbs,
]);
const r = await spawnP(cmd, args, {});
return { ok: r.status === 0 && existsSync(wavAbs), words: null };
}
// kokoro — via the published CLI; --output is relative to the project dir.
@@ -0,0 +1,143 @@
import { test } from "node:test";
import assert from "node:assert/strict";
import { EventEmitter } from "node:events";
import {
resolveNpxCliFromNpmExecPath,
resolveSpawnCommand,
spawnP,
_resetNpxResolutionWarnForTests,
} from "./tts.mjs";
// Regression: on Windows, npx resolves to npx.cmd, which spawn() cannot exec
// without shell:true — it fails ENOENT, silently swallowed as ok:false by the
// caller. spawnP takes injectable platform/spawnFn params so this doesn't
// need to touch the real process.platform or mock node:child_process (whose
// ESM exports are non-configurable).
function fakeSpawn(captured) {
return (cmd, args, opts) => {
captured.push({ cmd, args, opts });
const p = new EventEmitter();
setImmediate(() => p.emit("exit", 0));
return p;
};
}
const envWithNpxCli = {
npm_execpath: "/opt/node/lib/node_modules/npm/bin/npm-cli.js",
npm_node_execpath: "/opt/node/bin/node",
};
const npxCliPath = "/opt/node/lib/node_modules/npm/bin/npx-cli.js";
const pathExists = (path) => path === npxCliPath;
test("resolveNpxCliFromNpmExecPath finds npx-cli next to npm-cli", () => {
assert.equal(resolveNpxCliFromNpmExecPath(envWithNpxCli.npm_execpath, pathExists), npxCliPath);
});
test("resolveSpawnCommand routes npx through node+npx-cli on win32 without shell:true", () => {
const resolved = resolveSpawnCommand(
"npx",
["hyperframes", "tts", "C:\\Users\\Test User\\line.txt", "--voice", "am_michael"],
{},
"win32",
envWithNpxCli,
pathExists,
);
assert.ok(resolved);
assert.equal(resolved.cmd, envWithNpxCli.npm_node_execpath);
assert.deepEqual(resolved.args, [
npxCliPath,
"hyperframes",
"tts",
"C:\\Users\\Test User\\line.txt",
"--voice",
"am_michael",
]);
assert.equal(resolved.opts.shell, undefined);
});
test("resolveSpawnCommand preserves Windows npx shell metacharacters as argv data", () => {
const resolved = resolveSpawnCommand(
"npx",
["hyperframes", "tts", "hello & calc"],
{},
"win32",
envWithNpxCli,
pathExists,
);
assert.ok(resolved);
assert.deepEqual(resolved.args, [npxCliPath, "hyperframes", "tts", "hello & calc"]);
});
test("spawnP uses the resolved node+npx-cli command for npx on win32", async () => {
const captured = [];
await spawnP(
"npx",
["hyperframes", "tts"],
{},
"win32",
fakeSpawn(captured),
envWithNpxCli,
pathExists,
);
assert.equal(captured.length, 1);
assert.equal(captured[0].cmd, envWithNpxCli.npm_node_execpath);
assert.deepEqual(captured[0].args, [npxCliPath, "hyperframes", "tts"]);
assert.equal(captured[0].opts.shell, undefined);
});
test("spawnP does not enable shell for npx on darwin/linux", async () => {
const captured = [];
await spawnP("npx", ["hyperframes", "tts"], {}, "darwin", fakeSpawn(captured));
assert.equal(captured[0].cmd, "npx");
assert.deepEqual(captured[0].args, ["hyperframes", "tts"]);
assert.equal(captured[0].opts.shell, undefined);
});
test("spawnP does not enable shell for non-npx commands even on win32", async () => {
const captured = [];
await spawnP("python3", ["-c", "pass"], {}, "win32", fakeSpawn(captured));
assert.equal(captured[0].cmd, "python3");
assert.deepEqual(captured[0].args, ["-c", "pass"]);
assert.equal(captured[0].opts.shell, undefined);
});
// Regression: win32 + npx with npm_execpath unset can't locate the npx JS CLI,
// so resolveSpawnCommand returns null and spawnP short-circuits. Previously it
// returned {status:-1} silently — every TTS line just dropped as "TTS failed -
// omitted" with no hint. Now it must surface a clear one-time diagnostic naming
// npm_execpath, while still returning {status:-1} without spawning anything.
test("spawnP surfaces a clear diagnostic (once) when npx can't be resolved on win32", async () => {
_resetNpxResolutionWarnForTests();
const errors = [];
const originalError = console.error;
console.error = (msg) => errors.push(msg);
const captured = [];
const emptyEnv = {}; // no npm_execpath
try {
const r1 = await spawnP(
"npx",
["hyperframes", "tts"],
{},
"win32",
fakeSpawn(captured),
emptyEnv,
() => false,
);
const r2 = await spawnP(
"npx",
["hyperframes", "tts"],
{},
"win32",
fakeSpawn(captured),
emptyEnv,
() => false,
);
assert.equal(r1.status, -1);
assert.equal(r2.status, -1);
assert.equal(captured.length, 0, "must not spawn anything when resolution fails");
assert.equal(errors.length, 1, "diagnostic is emitted once for a batch, not per line");
assert.match(errors[0], /npm_execpath/);
} finally {
console.error = originalError;
}
});
@@ -0,0 +1,66 @@
import { test } from "node:test";
import assert from "node:assert/strict";
import { mkdtempSync, writeFileSync, chmodSync, rmSync } from "node:fs";
import { join } from "node:path";
import { tmpdir } from "node:os";
import { parseFfmpegDurationBanner, ffprobeDuration } from "./tts.mjs";
test("parseFfmpegDurationBanner reads ffmpeg's stderr Duration line", () => {
const stderr = [
"ffmpeg version 6.0",
"Input #0, wav, from 'a.wav':",
" Duration: 00:00:03.42, bitrate: 705 kb/s",
"At least one output file must be specified",
].join("\n");
assert.equal(parseFfmpegDurationBanner(stderr), 3.42);
});
test("parseFfmpegDurationBanner handles an hours component", () => {
const stderr = " Duration: 01:02:03.50, start: 0.000000, bitrate: 128 kb/s";
assert.equal(parseFfmpegDurationBanner(stderr), 3723.5);
});
test("parseFfmpegDurationBanner returns NaN when there is no Duration line", () => {
assert.ok(Number.isNaN(parseFfmpegDurationBanner("ffmpeg: command not found")));
assert.ok(Number.isNaN(parseFfmpegDurationBanner("")));
assert.ok(Number.isNaN(parseFfmpegDurationBanner(undefined)));
});
// Regression for the actual bug: ffprobeDuration used to collapse "ffprobe
// binary is missing" (ENOENT — the "essentials"-style Windows ffmpeg build
// with no ffprobe.exe) and "file is genuinely unreadable" into the same NaN,
// giving audio.mjs no way to tell "measure differently" from "give up".
//
// Builds an isolated PATH containing only a fake `ffmpeg` stub (no `ffprobe`
// at all) so ffprobeDuration's spawnSync("ffprobe", ...) call ENOENTs for
// real, then verifies it recovers the duration via the ffmpeg fallback
// instead of returning NaN.
test("ffprobeDuration falls back to ffmpeg when the ffprobe binary itself is missing", () => {
const dir = mkdtempSync(join(tmpdir(), "tts-ffprobe-fallback-"));
const fakeFfmpeg = join(dir, "ffmpeg");
writeFileSync(
fakeFfmpeg,
"#!/bin/sh\necho 'Duration: 00:00:02.50, start: 0.000000, bitrate: 128 kb/s' 1>&2\nexit 1\n",
);
chmodSync(fakeFfmpeg, 0o755);
const originalPath = process.env.PATH;
try {
process.env.PATH = dir; // only the fake ffmpeg resolves; no real ffprobe on this PATH
assert.equal(ffprobeDuration("/does/not/matter.wav"), 2.5);
} finally {
process.env.PATH = originalPath;
rmSync(dir, { recursive: true, force: true });
}
});
test("ffprobeDuration returns NaN when neither ffprobe nor ffmpeg resolve", () => {
const dir = mkdtempSync(join(tmpdir(), "tts-no-binaries-"));
const originalPath = process.env.PATH;
try {
process.env.PATH = dir; // empty directory — nothing resolves
assert.ok(Number.isNaN(ffprobeDuration("/does/not/matter.wav")));
} finally {
process.env.PATH = originalPath;
rmSync(dir, { recursive: true, force: true });
}
});
@@ -0,0 +1,226 @@
# Media operations: agent guidance
media-use resolves and remembers assets. For **operating** on them: cutting,
reframing, stitching, transforming, it does not wrap every action as a bespoke
command. Instead it points you at the right local tool (decision OP1). Run the
tool, then register the output with `resolve --from <output> --type <type>` so the
result lands in the ledger and the global cache like any other asset.
All tools below are local and free. ffmpeg is assumed present (it backs the
engine already).
## Cut / trim: keep a slice
```bash
ffmpeg -i in.mp4 -ss 00:00:12 -to 00:00:20 -c copy out.mp4 # 0:12–0:20, no re-encode
```
In-composition trimming usually needs **no new file**: a clip plays a sub-window
via `data-media-start` + `data-duration` (see hyperframes-core). Only cut a
physical file when exporting/assembling outside the composition.
## Reframe / crop: change aspect ratio
```bash
# 16:9 -> 9:16, crop centered
ffmpeg -i in.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" out.mp4
```
For a non-destructive crop, set a `clip-path` on the element in the composition
itself (render-time, source file untouched) instead of re-encoding with ffmpeg.
## Montage / stitch: join clips
```bash
printf "file '%s'\n" a.mp4 b.mp4 c.mp4 > list.txt
ffmpeg -f concat -safe 0 -i list.txt -c copy out.mp4
```
## Silence-cut / highlight: trim dead air, grab the best moment
```bash
auto-editor in.mp4 --edit audio:threshold=4% -o tight.mp4 # pip install auto-editor
scenedetect -i in.mp4 detect-adaptive list-scenes # pip install scenedetect
```
## Transforms with a quality choice (process)
These have a local option AND a higher-quality HeyGen-CLI option. Run the local
one for free/offline; use the HeyGen CLI when quality matters. Showing the user
a **side-by-side** (local vs HeyGen) is the honest way to let them choose.
| Op | Local (free) | HeyGen CLI (quality) |
| ------------------ | -------------------------------------------------- | --------------------------- |
| Background removal | `hyperframes remove-background in.png` (u2net) | `heygen background-removal` |
| Upscale | `realesrgan-ncnn-vulkan -i in.png -o out.png -s 4` | n/a |
| Lipsync (dub) | n/a | `heygen lipsync` |
| Translate | n/a | `heygen video-translate` |
After any op: `resolve --from out.ext --type <type>` to register the derived
asset (it records provenance and auto-promotes to the global cache).
> ponytail: media-use doesn't re-wrap ffmpeg/heygen here, that's deliberate
> (OP1). The value it adds is the ledger + global reuse on the _output_, via
> `--from`. Add a thin `process` verb only if agents repeatedly fumble these
> recipes.
## Transcription (default: Parakeet, better than whisper.cpp)
`transcribe.mjs` is the default local transcription path. It runs **NVIDIA
Parakeet-TDT via parakeet-mlx**, which beats whisper.cpp on the Open ASR
Leaderboard (avg WER ~6.05% vs 7.44%; on NOISY audio 4.73% vs 5.96%, where
whisper-large-v3 hallucinated to 308% WER on meetings) and is 5-10x faster.
It emits `{ text, words:[{text,start,end}] }` with word timestamps (merged from
Parakeet's sub-word tokens), feeding transcript-cut, captions, and the audio
engine directly.
```bash
# install once: uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx
node <SKILL_DIR>/scripts/transcribe.mjs --input talk.mp4 --out talk.transcribe.json
# equivalently, the hyperframes CLI has Parakeet built in (auto-detects it, whisper fallback):
npx hyperframes transcribe talk.mp4 --engine parakeet # or --engine auto (default)
```
VERIFIED on 24GB: accurate, ~3s (cached) for 8s audio. Parakeet covers English +
25 European languages. For other languages, or when parakeet-mlx is not
installed, transcribe.mjs auto-falls-back to whisper.cpp (99 languages) via
`hyperframes transcribe`. `--engine parakeet|whisper` forces one. (Cohere
Transcribe tops the leaderboard on paper but its mlx-audio quants produced
garbage and ran 40-70x slower on a Mac in testing, so it is not wired in.)
## Text-based editing (transcript cut)
`transcript-cut.mjs` is a compiler, not a wrapper: it turns word timestamps and
agent cut decisions into exact kept segments. It is provided even though the rest
of this file is guidance-only.
```bash
node <SKILL_DIR>/scripts/transcript-cut.mjs \
--input talk.mp4 \
--transcript talk.transcribe.json \
--remove "12.41-15.02,88.3-91.7" \
--remove-fillers "um,uh,like" \
--cut-silence 0.8 \
--out talk.cut.mp4
resolve --from talk.cut.mp4 --type video
```
Use `--plan` first when you want to inspect the kept segment JSON before encoding.
## Ducking (declare in-composition / bake for export)
B1, declare ducking in the composition. `audio-duck.mjs` emits GSAP volume
keyframes. Paste them into the composition timeline, the source file stays
untouched.
```bash
node <SKILL_DIR>/scripts/audio-duck.mjs \
--meta audio_meta.json \
--target "#bgm" \
--composition index.html
```
```js
// auto-duck: #bgm under narration (generated; base volume 0.6)
tl.to("#bgm", { volume: 0.15, duration: 0.15 }, 3.42);
tl.to("#bgm", { volume: 0.6, duration: 0.4 }, 9.87);
```
B2, bake ducking only for exported or standalone files.
```bash
ffmpeg -i bgm.mp3 -i voice.wav \
-filter_complex "[0][1]sidechaincompress=threshold=0.03:ratio=8:attack=200:release=400[ducked]" \
-map "[ducked]" bgm.ducked.wav
```
Declare inside compositions. Bake only for assets leaving the hyperframes
pipeline.
## Publish loudness
Two-pass `loudnorm` measures first, then applies the measured values with the
target LUFS baked in.
Socials target, -14 LUFS:
```bash
ffmpeg -i mix.wav \
-af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json \
-f null -
ffmpeg -i mix.wav \
-af loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=<input_i>:measured_TP=<input_tp>:measured_LRA=<input_lra>:measured_thresh=<input_thresh>:offset=<target_offset>:linear=true:print_format=summary \
mix.social.wav
```
Podcast target, -16 LUFS:
```bash
ffmpeg -i mix.wav \
-af loudnorm=I=-16:TP=-1.5:LRA=11:print_format=json \
-f null -
ffmpeg -i mix.wav \
-af loudnorm=I=-16:TP=-1.5:LRA=11:measured_I=<input_i>:measured_TP=<input_tp>:measured_LRA=<input_lra>:measured_thresh=<input_thresh>:offset=<target_offset>:linear=true:print_format=summary \
mix.podcast.wav
```
## Generate: images (local first, cloud upsell)
`resolve --type image` retrieves from the HeyGen catalog first; on a miss it
GENERATES. Two paths, best-for-the-machine picked automatically:
1. **Local (default, free, private): mflux** (FLUX-on-MLX). `resolve` spec-checks
AVAILABLE RAM and runs the best FLUX-class model that fits, via
`scripts/lib/local-models.mjs` (`imagegen` ladder) + `mflux-provider.mjs`.
The RAM ladder (agent sees it via `describeModelLadder("imagegen", specs)`):
| Tier | Model | Needs (available RAM) | Notes |
| ------ | -------------------- | --------------------- | ----------------------------------- |
| medium | FLUX.1 schnell int4 | ~8GB (`--low-ram`) | ~20s/512px on 24GB. VERIFIED. Fast. |
| large | FLUX.2 Klein 4B int4 | ~32GB | higher quality, full-resident |
| xlarge | Qwen-Image | ~64GB | top quality, 64GB+ Macs only |
Gotchas baked into the table: the official FLUX repos are HF-gated, so it
points at non-gated community 4-bit re-uploads; and `--low-ram` is MANDATORY
at the medium tier (without it a 768x512 run swap-thrashed to 90 minutes on
24GB; with it, 20 seconds).
2. **Cloud upsell (better quality): the `codex` CLI** `image_gen` tool, on the
user's ChatGPT subscription (codex owns auth, no key here, no per-call
charge). It is the automatic fallback when no local model fits AND the
explicit "make it better" choice on any machine. Users who just want codex
can ask for it directly. Verified: prompt -> raster -> frozen + ledgered.
`--local-only` keeps mflux (once cached) and skips codex (network).
## Generate: video (local first, HeyGen avatar upsell)
Operate-on-video ships now; GENERATING video is local-first with a HeyGen
avatar upsell (decision X3).
- **Local (default): LTX 2.3 on MLX** via `dgrauet/ltx-2-mlx`, the `videogen`
ladder in `local-models.mjs`. Generative clips (t2v / i2v), spec-gated to RAM.
Verified on 24GB: 512x320 x 33f with audio.
- **HeyGen avatar upsell (better, script-driven): the `heygen` CLI**, NOT the
raw API. For a talking-head / avatar video, `heygen video create` (avatar
engine IV by default) beats a generative clip when you want a real presenter:
```bash
# discover an avatar + a starfish voice, then create + wait
heygen avatar list --ownership public --limit 5
heygen voice list --engine starfish --limit 5
heygen video create --wait -d '{
"type": "avatar",
"avatar_id": "<avatar-id>",
"script": "Your narration here.",
"voice_id": "<voice-id>"
}'
```
Avatar videos are deterministic + script-driven (lip-sync from a script or a
pre-recorded `audio_url`), distinct from the generative LTX clips. After it
renders, `resolve --from <downloaded.mp4> --type video` to ledger it.
@@ -0,0 +1,121 @@
#!/usr/bin/env node
import { readFileSync } from "node:fs";
import { resolve } from "node:path";
import { parseArgs } from "node:util";
import { duckKeyframes, speechSpans } from "./lib/duck.mjs";
import { track } from "./lib/telemetry.mjs";
const { values: args } = parseArgs({
options: {
meta: { type: "string" },
target: { type: "string" },
duck: { type: "string", default: "0.25" },
attack: { type: "string", default: "0.15" },
release: { type: "string", default: "0.4" },
"merge-gap": { type: "string", default: "0.6" },
sequential: { type: "boolean", default: false },
gap: { type: "string", default: "0" },
offsets: { type: "string" },
composition: { type: "string" },
json: { type: "boolean", default: false },
help: { type: "boolean", short: "h", default: false },
},
strict: true,
});
if (args.help) {
console.log(`media-use audio-duck — generate GSAP volume ducking keyframes
Usage:
node audio-duck.mjs --meta audio_meta.json --target "#bgm"
Options:
--meta audio_meta.json or JSON word transcript
--target GSAP selector for the background audio element
--duck Duck multiplier (default: 0.25)
--attack Duck-in duration seconds (default: 0.15)
--release Restore duration seconds (default: 0.4)
--merge-gap Bridge speech gaps smaller than this many seconds (default: 0.6)
--sequential Place multi-line meta back to back at composition time
--gap Extra seconds between sequential lines (default: 0)
--offsets Explicit placement, "l1=0,l2=3.4" (voice id = start seconds)
--composition Read target data-volume from this HTML file
--json Output { spans, keyframes }
--help, -h Show this help`);
process.exit(0);
}
try {
run();
await track("media_use_duck", { sequential: !!args.sequential });
} catch (err) {
if (args.json) console.log(JSON.stringify({ ok: false, error: err.message }));
else console.error(`error: ${err.message}`);
process.exit(1);
}
function run() {
if (!args.meta || !args.target) throw new Error("--meta and --target are required");
const meta = JSON.parse(readFileSync(resolve(args.meta), "utf8"));
const target = args.target;
const baseVolume = readBaseVolume(args.composition, target);
const offsets = args.offsets
? Object.fromEntries(
args.offsets.split(",").map((pair) => {
const [id, t] = pair.split("=");
return [id.trim(), Number(t)];
}),
)
: undefined;
const spans = speechSpans(meta, {
mergeGap: Number(args["merge-gap"]),
sequential: args.sequential,
gap: Number(args.gap),
offsets,
});
const keyframes = duckKeyframes(spans, {
duck: Number(args.duck),
attack: Number(args.attack),
release: Number(args.release),
baseVolume,
});
if (args.json) {
console.log(JSON.stringify({ spans, keyframes }));
return;
}
console.log(
`// auto-duck: ${target} under narration (generated; base volume ${fmt(baseVolume)})`,
);
for (const keyframe of keyframes) {
console.log(
`tl.to(${JSON.stringify(target)}, { volume: ${fmt(keyframe.volume)}, duration: ${fmt(
keyframe.duration,
)} }, ${fmt(keyframe.time)});`,
);
}
}
function readBaseVolume(composition, target) {
if (!composition || !target.startsWith("#")) return 1;
const id = target.slice(1);
const html = readFileSync(resolve(composition), "utf8");
// ponytail: regex is enough here because this only reads one attribute from
// one user-authored composition element, not arbitrary HTML.
const tag = html.match(new RegExp(`<[^>]*\\bid=["']${escapeRegExp(id)}["'][^>]*>`, "i"))?.[0];
const raw = tag?.match(/\bdata-volume=["']([^"']+)["']/i)?.[1];
const volume = Number(raw);
return Number.isFinite(volume) ? volume : 1;
}
function escapeRegExp(value) {
return value.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
}
function fmt(n) {
return Number(n)
.toFixed(3)
.replace(/\.?0+$/, "");
}
@@ -29,8 +29,12 @@ function markComplete(entryDir) {
writeFileSync(join(entryDir, COMPLETE_SENTINEL), "", "utf8");
}
// The manifest helpers append their own ".media" to the dir they get, so the
// global manifest must be addressed by HOME, not by globalMediaDir() — passing
// the latter nested it at ~/.media/.media/manifest.jsonl, invisible to the
// Studio /api/assets/global route (which reads the documented flat path).
function readGlobalManifest() {
return readManifest(globalMediaDir());
return readManifest(homedir());
}
function validateCacheHit(match) {
@@ -55,6 +59,11 @@ export function cacheGetByEntity(entity) {
export function cachePut(filePath, record) {
const sha = contentHash(filePath);
// Idempotent: same content already promoted -> don't duplicate the global
// record. ponytail: skips usage_count bump; add it when the metric is needed.
const existing = readGlobalManifest().find((r) => r.sha === sha);
if (existing) return { sha, cached_path: existing.cached_path, deduped: true };
const dir = globalMediaDir();
const entryDir = cacheEntryDir(dir, sha);
mkdirSync(entryDir, { recursive: true });
@@ -69,7 +78,7 @@ export function cachePut(filePath, record) {
reusable: true,
cached_path: dest,
};
appendRecord(globalMediaDir(), globalRecord);
appendRecord(homedir(), globalRecord);
return { sha, cached_path: dest };
}
@@ -0,0 +1,133 @@
import { execFileSync } from "node:child_process";
import { copyFileSync, existsSync, readdirSync, statSync, unlinkSync } from "node:fs";
import { homedir, tmpdir } from "node:os";
import { join } from "node:path";
// Image generation via the OpenAI Codex CLI's built-in image tool (gpt-image-2)
// on the user's ChatGPT subscription: the codex CLI owns auth, media-use holds
// no key (CLI-only). The image UPSELL behind local mflux; skipped by --local-only.
//
// Retrieval mirrors illo-skill rather than trusting the model to save a file:
// `--enable imagegenext` makes the built-in tool drop the rendered artifact into
// $CODEX_HOME/generated_images/, and we fetch the freshest file that postdates
// this run. The save-to-path instruction is only a best-effort verify-first.
const TIMEOUT_MS = 600000; // codex exec round-trips the sub; first-run tool spin-up is slow
const MTIME_SKEW_MS = 2000; // tolerate mtime granularity / clock skew (illo uses 2s)
function codexGeneratedDir() {
// Codex relocates CODEX_HOME on some hosts, so resolve it at run time.
return join(process.env.CODEX_HOME || join(homedir(), ".codex"), "generated_images");
}
// Newest artifact that postdates `sinceMs` (minus skew), so a stale prior render
// or a concurrent session's file can't be mistaken for this run's output.
function freshestGeneratedImage(sinceMs) {
const dir = codexGeneratedDir();
if (!existsSync(dir)) return null;
const floor = sinceMs - MTIME_SKEW_MS;
let best = null;
for (const name of readdirSync(dir)) {
let st;
try {
st = statSync(join(dir, name));
} catch {
continue;
}
if (!st.isFile() || st.mtimeMs < floor) continue;
if (!best || st.mtimeMs > best.mtimeMs) best = { path: join(dir, name), mtimeMs: st.mtimeMs };
}
return best?.path ?? null;
}
// Short `codex` subcommand → combined stdout+stderr, or null if it can't run.
function codexRun(args) {
try {
return execFileSync("codex", args, {
encoding: "utf8",
stdio: ["ignore", "pipe", "pipe"],
timeout: 10000,
});
} catch (err) {
return `${err.stdout?.toString() ?? ""}${err.stderr?.toString() ?? ""}` || null;
}
}
// Fail-fast host check (mirrors illo): don't burn a minutes-long exec when Codex
// isn't usable. Returns null when ready, else a human reason. imagegenext ships
// default-disabled ("under development"), so we check the ROW is present (the
// capability signal) — the exec enables it per-render with --enable.
function codexUnavailableReason() {
try {
const which = process.platform === "win32" ? "where" : "which";
execFileSync(which, ["codex"], { stdio: ["ignore", "ignore", "ignore"], timeout: 5000 });
} catch {
return "codex CLI not on PATH";
}
const login = codexRun(["login", "status"]);
if (!login || !/logged in/i.test(login)) return "codex not logged in (run: codex login)";
const feats = codexRun(["features", "list"]);
if (feats == null) return "could not read `codex features list`";
if (!/\bimage_generation\b/.test(feats)) return "codex image_generation feature unavailable";
if (!/\bimagegenext\b/.test(feats)) return "codex imagegenext unavailable (upgrade Codex CLI)";
return null;
}
export async function codexImageGenerate(intent) {
const unavailable = codexUnavailableReason();
if (unavailable) {
console.error(`media-use: codex image upsell unavailable: ${unavailable}`);
return null;
}
const outPath = join(tmpdir(), `media-use-codex-${process.pid}-${Date.now()}.png`);
const prompt =
`${intent}\n\n` +
`Use your built-in image generation tool to render this, then save the image ` +
`to ${outPath} (overwrite if it exists). Do not ask for confirmation. ` +
`If you have no built-in image tool, do nothing (no PIL/matplotlib/SVG substitute).`;
try {
unlinkSync(outPath); // clear any prior file so verify-first can't accept a stale render
} catch {
/* no prior file */
}
const started = Date.now();
try {
execFileSync(
"codex",
[
"exec",
"--cd",
tmpdir(),
"-s",
"workspace-write",
"--skip-git-repo-check",
"--enable",
"imagegenext",
"-",
],
{ input: prompt, encoding: "utf8", timeout: TIMEOUT_MS, stdio: ["pipe", "pipe", "pipe"] },
);
} catch (err) {
console.error(
`media-use: \`codex exec\` image generation failed: ${err.stderr?.toString().trim().slice(-200) || err.message}`,
);
return null;
}
// Verify-first (save-to-path may have worked), else fetch the imagegenext artifact.
const produced =
existsSync(outPath) && statSync(outPath).size > 0 ? outPath : freshestGeneratedImage(started);
if (!produced) return null;
if (produced !== outPath) {
try {
copyFileSync(produced, outPath);
} catch {
return null;
}
}
return {
localPath: outPath,
ext: ".png",
source: "generated",
metadata: { description: intent, provider: "codex.image_gen", provenance: { prompt: intent } },
};
}
@@ -0,0 +1,112 @@
import { strict as assert } from "node:assert";
import { test } from "node:test";
import { existsSync } from "node:fs";
import { join, dirname } from "node:path";
import { fileURLToPath } from "node:url";
import { listTypes, getProviders } from "./registry.mjs";
import { CAPABILITIES, listModels } from "./local-models.mjs";
// Capstone: media-use must actually OWN each hyperframes media weakness. This
// test enforces the weakness→owner matrix in SKILL.md so a claim can't rot — if
// a capability's entrypoint disappears, this fails.
const SKILL = join(dirname(fileURLToPath(import.meta.url)), "..", "..");
test("weakness: audio-only → media-use resolves image + icon", () => {
for (const t of ["image", "icon"]) {
assert.ok(getProviders(t).length > 0, `no provider for ${t}`);
}
});
test("weakness: no voice/audio gen → media-use exposes voice + the audio engine", () => {
assert.ok(listTypes().includes("voice"), "voice type missing");
assert.ok(getProviders("voice").length > 0, "no enabled voice provider (Bin approved)");
assert.ok(existsSync(join(SKILL, "audio", "scripts", "audio.mjs")), "audio engine missing");
});
test("weakness: scattered audio engine → consolidated under media-use (hyperframes-media gone)", () => {
assert.ok(existsSync(join(SKILL, "audio", "scripts", "lib", "tts.mjs")), "tts engine missing");
assert.ok(
existsSync(join(SKILL, "audio", "assets", "sfx", "manifest.json")),
"bundled SFX missing",
);
});
test("weakness: no media-ops → ops guidance reference exists", () => {
assert.ok(existsSync(join(SKILL, "references", "operations.md")), "operations.md missing");
});
test("weakness: no transcript-driven cutting → cut compiler entrypoints exist", async () => {
assert.ok(existsSync(join(SKILL, "scripts", "transcript-cut.mjs")), "transcript-cut missing");
assert.ok(existsSync(join(SKILL, "scripts", "lib", "cutlist.mjs")), "cutlist lib missing");
const cutlist = await import("./cutlist.mjs");
assert.equal(typeof cutlist.compileCutList, "function");
});
test("weakness: whisper.cpp is weak → better local ASR (Parakeet) entrypoint exists", async () => {
assert.ok(existsSync(join(SKILL, "scripts", "transcribe.mjs")), "transcribe.mjs missing");
const pw = await import("./parakeet-words.mjs");
assert.equal(typeof pw.mergeTokensToWords, "function", "token->word merge missing");
const lm = await import("./local-models.mjs");
const asr = lm.listModels("asr");
const parakeet = asr.find((m) => m.id === "parakeet-mlx");
assert.ok(parakeet && parakeet.rank === 0, "Parakeet must be the rank-0 preferred ASR");
});
test("weakness: no auto-duck/loudness → duck compiler and recipes exist", async () => {
assert.ok(existsSync(join(SKILL, "scripts", "audio-duck.mjs")), "audio-duck missing");
assert.ok(existsSync(join(SKILL, "scripts", "lib", "duck.mjs")), "duck lib missing");
assert.ok(existsSync(join(SKILL, "references", "operations.md")), "operations.md missing");
const duck = await import("./duck.mjs");
assert.equal(typeof duck.speechSpans, "function");
assert.equal(typeof duck.duckKeyframes, "function");
});
test("weakness: no cross-project memory → global cache + ingest entrypoints exist", async () => {
const cache = await import("./cache.mjs");
assert.equal(typeof cache.cachePut, "function");
assert.equal(typeof cache.promote, "function");
assert.equal(typeof cache.globalMediaDir, "function");
const freeze = await import("./freeze.mjs");
assert.equal(typeof freeze.isDirectMediaUrl, "function", "ingest URL guard missing");
});
// Wenbo (06-29): heygen free-usage is the default; local models are the opt-out
// fallback ("if user no, then local"). We still assert the fallback table is
// populated so the opt-out path stays real.
test("weakness: weak local defaults → local models exist as the opt-out fallback (tts/asr/upscale)", () => {
for (const cap of ["tts", "asr", "upscale"]) {
assert.ok(CAPABILITIES.includes(cap), `capability ${cap} missing`);
assert.ok(listModels(cap).length > 0, `no local models for ${cap}`);
}
});
test("weakness: no image generation → local mflux (RAM-graded) + codex upsell", async () => {
const ps = getProviders("image");
assert.ok(
ps.some((p) => p.name === "mflux.local" && typeof p.generate === "function"),
"local image gen missing",
);
assert.ok(
ps.some((p) => p.name === "codex.image_gen" && typeof p.generate === "function"),
"codex image upsell missing",
);
const lm = await import("./local-models.mjs");
assert.ok(lm.CAPABILITIES.includes("imagegen"), "imagegen capability missing");
assert.ok(lm.listModels("imagegen").length >= 3, "imagegen RAM ladder too small");
assert.equal(typeof lm.describeModelLadder, "function", "agent-facing ladder missing");
});
test("weakness: no video generation → local videogen ladder + heygen avatar upsell", async () => {
const lm = await import("./local-models.mjs");
assert.ok(lm.CAPABILITIES.includes("videogen"), "videogen capability missing");
assert.ok(lm.listModels("videogen").length >= 2, "videogen ladder too small");
const ops = existsSync(join(SKILL, "references", "operations.md"));
assert.ok(ops, "operations.md (avatar-upsell recipe) missing");
});
test("every resolve type has at least one enabled provider", () => {
for (const t of listTypes()) {
assert.ok(getProviders(t).length > 0, `type ${t} has no enabled provider`);
}
});
@@ -0,0 +1,184 @@
import { normalizeWords } from "./words.mjs";
const MIN_SEGMENT_SECONDS = 0.2;
const SILENCE_PAD_SECONDS = 0.15;
export function compileCutList(transcript, opts = {}) {
const words = normalizeWords(transcript);
if (opts.keep != null && hasRemovalSource(opts)) {
throw new Error("--keep is mutually exclusive with removal options");
}
if (opts.keep != null) {
const duration = durationFrom(words, opts);
const ranges = parseTimeRanges(opts.keep);
return finalizeKept(duration != null ? clampRanges(ranges, duration) : ranges);
}
const duration = durationFrom(words, opts);
if (!duration) return [];
const removals = [
...parseTimeRanges(opts.remove),
...wordIndexRanges(words, opts.removeWords),
...fillerRanges(words, opts.removeFillers),
...silenceRanges(words, opts.cutSilence),
];
const mergedRemovals = mergeRanges(clampRanges(removals, duration));
return finalizeKept(invertRanges(mergedRemovals, duration));
}
function hasRemovalSource(opts) {
return (
opts.remove != null ||
opts.removeWords != null ||
opts.removeFillers != null ||
opts.cutSilence != null
);
}
function durationFrom(words, opts) {
const explicit = Number(opts.duration ?? opts.totalDuration);
if (Number.isFinite(explicit) && explicit > 0) return explicit;
const last = words.at(-1);
return last && Number.isFinite(last.end) && last.end > 0 ? last.end : null;
}
function parseTimeRanges(value) {
if (value == null || value === false || value === "") return [];
if (typeof value === "string") {
return value
.split(",")
.map((part) => part.trim())
.filter(Boolean)
.map(parseRangeString);
}
if (!Array.isArray(value)) throw new Error("range list must be a string or array");
return value.map((range) => {
if (Array.isArray(range)) return cleanRange(Number(range[0]), Number(range[1]));
return cleanRange(Number(range?.start), Number(range?.end));
});
}
function parseRangeString(value) {
const match = value.match(/^([0-9]*\.?[0-9]+)\s*-\s*([0-9]*\.?[0-9]+)$/);
if (!match) throw new Error(`invalid range: ${value}`);
return cleanRange(Number(match[1]), Number(match[2]));
}
function cleanRange(start, end) {
if (!Number.isFinite(start) || !Number.isFinite(end)) {
throw new Error("range start/end must be finite numbers");
}
if (end < start) throw new Error(`range end ${end} is before start ${start}`);
return { start, end };
}
function wordIndexRanges(words, value) {
if (value == null || value === false || value === "") return [];
const ranges = typeof value === "string" ? value.split(",") : value;
if (!Array.isArray(ranges)) throw new Error("--remove-words must be a string or array");
return ranges
.map((range) => (typeof range === "string" ? range.trim() : range))
.filter(Boolean)
.map((range) => {
const [first, last = first] =
typeof range === "string" ? range.split("-").map((n) => n.trim()) : range;
const startIndex = Number(first);
const endIndex = Number(last);
if (!Number.isInteger(startIndex) || !Number.isInteger(endIndex)) {
throw new Error(`invalid word range: ${range}`);
}
if (startIndex < 0 || endIndex < startIndex || endIndex >= words.length) {
throw new Error(`word range out of bounds: ${range}`);
}
return { start: words[startIndex].start, end: words[endIndex].end };
});
}
function fillerRanges(words, value) {
if (value == null || value === false || value === "") return [];
const fillers = Array.isArray(value)
? value
: String(value)
.split(",")
.map((s) => s.trim());
const set = new Set(fillers.filter(Boolean).map(bareToken));
if (set.size === 0) return [];
// Whisper emits words with attached punctuation and arbitrary case
// ("UM," / "Um."), so compare bare tokens.
return words
.filter((word) => set.has(bareToken(word.text)))
.map((word) => ({ start: word.start, end: word.end }));
}
function bareToken(text) {
return String(text)
.toLowerCase()
.replace(/^[^\p{L}\p{N}]+|[^\p{L}\p{N}]+$/gu, "");
}
function silenceRanges(words, value) {
if (value == null || value === false || value === "") return [];
const threshold = Number(value);
if (!Number.isFinite(threshold) || threshold <= 0) {
throw new Error("--cut-silence must be a positive number");
}
const ranges = [];
for (let i = 0; i < words.length - 1; i++) {
const current = words[i];
const next = words[i + 1];
const gap = next.start - current.end;
if (gap <= threshold) continue;
const start = current.end + SILENCE_PAD_SECONDS;
const end = next.start - SILENCE_PAD_SECONDS;
if (end > start) ranges.push({ start, end });
}
return ranges;
}
function clampRanges(ranges, duration) {
return ranges
.map((range) => ({
start: Math.max(0, Math.min(duration, range.start)),
end: Math.max(0, Math.min(duration, range.end)),
}))
.filter((range) => range.end > range.start);
}
function mergeRanges(ranges) {
const sorted = ranges
.map((range) => ({ start: round3(range.start), end: round3(range.end) }))
.sort((a, b) => a.start - b.start || a.end - b.end);
const merged = [];
for (const range of sorted) {
const prev = merged.at(-1);
if (prev && range.start <= prev.end) {
prev.end = Math.max(prev.end, range.end);
} else {
merged.push({ ...range });
}
}
return merged;
}
function invertRanges(removals, duration) {
const kept = [];
let cursor = 0;
for (const range of removals) {
if (range.start > cursor) kept.push({ start: cursor, end: range.start });
cursor = Math.max(cursor, range.end);
}
if (cursor < duration) kept.push({ start: cursor, end: duration });
return kept;
}
function finalizeKept(ranges) {
return mergeRanges(ranges)
.map((range) => ({ start: round3(range.start), end: round3(range.end) }))
.filter((range) => round3(range.end - range.start) >= MIN_SEGMENT_SECONDS);
}
function round3(n) {
return Math.round(Number(n) * 1000) / 1000;
}
@@ -0,0 +1,148 @@
import { strict as assert } from "node:assert";
import { execFileSync } from "node:child_process";
import { mkdtempSync, writeFileSync, rmSync } from "node:fs";
import { join, dirname } from "node:path";
import { tmpdir } from "node:os";
import { fileURLToPath } from "node:url";
import { test } from "node:test";
import { compileCutList } from "./cutlist.mjs";
const HERE = dirname(fileURLToPath(import.meta.url));
const SCRIPT = join(HERE, "..", "transcript-cut.mjs");
test("explicit --remove ranges invert to kept segments", () => {
const transcript = [
word("w0", "alpha", 0, 1),
word("w1", "beta", 1.2, 2),
word("w2", "gamma", 2.2, 5),
];
assert.deepEqual(compileCutList(transcript, { remove: "1-2.5" }), [
{ start: 0, end: 1 },
{ start: 2.5, end: 5 },
]);
});
test("--remove-words resolves inclusive word-index ranges to time ranges", () => {
const transcript = [
word("w0", "zero", 0, 0.5),
word("w1", "one", 0.6, 1),
word("w2", "two", 1.1, 1.5),
word("w3", "three", 2, 3),
];
assert.deepEqual(compileCutList(transcript, { removeWords: "1-2" }), [
{ start: 0, end: 0.6 },
{ start: 1.5, end: 3 },
]);
});
test("--remove-fillers drops case-insensitive matching words", () => {
const transcript = [
word("w0", "Hello", 0, 0.5),
word("w1", "Um", 0.5, 0.7),
word("w2", "world", 0.8, 1.2),
word("w3", "LIKE", 1.3, 1.5),
word("w4", "done", 1.6, 2),
];
assert.deepEqual(compileCutList(transcript, { removeFillers: "um,like" }), [
{ start: 0, end: 0.5 },
{ start: 0.7, end: 1.3 },
{ start: 1.5, end: 2 },
]);
});
test("--cut-silence removes only the center of long inter-word gaps", () => {
const transcript = [word("w0", "a", 0, 0.5), word("w1", "b", 2, 2.5), word("w2", "c", 2.7, 3)];
assert.deepEqual(compileCutList(transcript, { cutSilence: 0.8 }), [
{ start: 0, end: 0.65 },
{ start: 1.85, end: 3 },
]);
});
test("overlapping removal sources merge before inversion", () => {
const transcript = [
word("w0", "start", 0, 0.5),
word("w1", "um", 0.9, 1.1),
word("w2", "middle", 2.5, 2.8),
word("w3", "more", 3.1, 3.4),
word("w4", "end", 5.5, 6),
];
assert.deepEqual(
compileCutList(transcript, {
remove: "1-2.7",
removeWords: "2-3",
removeFillers: "um",
}),
[
{ start: 0, end: 0.9 },
{ start: 3.4, end: 6 },
],
);
});
test("kept slivers shorter than 0.2s are dropped", () => {
const transcript = [word("w0", "start", 0, 0.5), word("w1", "end", 2.5, 3)];
assert.deepEqual(compileCutList(transcript, { remove: "0.1-2.95" }), []);
});
test("--keep is inverse mode and coalesces direct kept ranges", () => {
const transcript = [word("w0", "start", 0, 0.5), word("w1", "end", 4.5, 5)];
assert.deepEqual(compileCutList(transcript, { keep: "3-4,1-2,1.5-2.5,4.1-4.2" }), [
{ start: 1, end: 2.5 },
{ start: 3, end: 4 },
]);
});
test("--plan on a fixture transcript prints the exact segment JSON", () => {
const dir = mkdtempSync(join(tmpdir(), "media-use-cutlist-"));
try {
const transcriptPath = join(dir, "fixture.json");
writeFileSync(
transcriptPath,
JSON.stringify([
word("w0", "hello", 0, 0.4),
word("w1", "um", 0.5, 0.65),
word("w2", "there", 0.7, 1),
word("w3", "pause", 2.2, 2.5),
word("w4", "end", 2.7, 3.2),
]),
);
const out = execFileSync(
process.execPath,
[
SCRIPT,
"--input",
"ignored.mp4",
"--transcript",
transcriptPath,
"--remove",
"0.9-1.2",
"--remove-fillers",
"um",
"--cut-silence",
"0.8",
"--plan",
],
{ encoding: "utf8" },
);
assert.deepEqual(JSON.parse(out), [
{ start: 0, end: 0.5 },
{ start: 0.65, end: 0.9 },
{ start: 2.05, end: 3.2 },
]);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
function word(id, text, start, end) {
return { id, text, start, end };
}
@@ -0,0 +1,89 @@
import { wordListsFromMediaMeta } from "./words.mjs";
/**
* Speech spans from word timestamps.
*
* audio_meta.json word times are relative to EACH LINE'S OWN FILE, not to the
* composition. Without placement info, multiple lines would overlap at t=0 and
* merge into one bogus span. Placement options:
* offsets: { [voiceId]: startSeconds } explicit composition placement
* sequential: stack lines back to back (plus `gap` seconds between lines)
* A single word list (bare transcript) needs neither.
*/
export function speechSpans(meta, { mergeGap = 0.6, offsets, sequential = false, gap = 0 } = {}) {
const merge = Number(mergeGap);
const lists = wordListsFromMediaMeta(meta);
const voices = Array.isArray(meta?.voices) ? meta.voices : [];
if (lists.length > 1 && !offsets && !sequential) {
throw new Error(
"audio_meta has multiple voice lines with file-relative times; pass --sequential or --offsets so spans land at composition time",
);
}
const intervals = [];
let cursor = 0;
for (let i = 0; i < lists.length; i++) {
const voice = voices[i];
let offset = 0;
if (offsets) {
const id = voice?.id ?? String(i);
if (!(id in offsets)) throw new Error(`--offsets is missing voice "${id}"`);
offset = Number(offsets[id]) || 0;
} else if (sequential) {
offset = cursor;
const lineDuration = Number(voice?.duration_s) || Math.max(...lists[i].map((w) => w.end), 0);
cursor += lineDuration + (Number(gap) || 0);
}
for (const word of lists[i]) {
if (word.end > word.start)
intervals.push({ start: word.start + offset, end: word.end + offset });
}
}
return mergeIntervals(intervals, Number.isFinite(merge) && merge >= 0 ? merge : 0.6);
}
export function duckKeyframes(
spans,
{ duck = 0.25, attack = 0.15, release = 0.4, baseVolume = 1 } = {},
) {
const base = finiteOr(baseVolume, 1);
const ducked = round3(base * finiteOr(duck, 0.25));
const keyframes = [];
for (const span of spans) {
keyframes.push({
time: round3(Math.max(0, finiteOr(span.start, 0))),
volume: ducked,
duration: round3(finiteOr(attack, 0.15)),
});
keyframes.push({
time: round3(Math.max(0, finiteOr(span.end, 0))),
volume: round3(base),
duration: round3(finiteOr(release, 0.4)),
});
}
return keyframes.sort((a, b) => a.time - b.time);
}
function mergeIntervals(intervals, mergeGap) {
const sorted = intervals
.map((range) => ({ start: round3(range.start), end: round3(range.end) }))
.sort((a, b) => a.start - b.start || a.end - b.end);
const merged = [];
for (const range of sorted) {
const prev = merged.at(-1);
if (prev && (range.start <= prev.end || range.start - prev.end < mergeGap)) {
prev.end = Math.max(prev.end, range.end);
} else {
merged.push({ ...range });
}
}
return merged;
}
function finiteOr(value, fallback) {
const n = Number(value);
return Number.isFinite(n) ? n : fallback;
}
function round3(n) {
return Math.round(Number(n) * 1000) / 1000;
}
@@ -0,0 +1,118 @@
import { strict as assert } from "node:assert";
import { execFileSync } from "node:child_process";
import { mkdtempSync, writeFileSync, rmSync } from "node:fs";
import { join, dirname } from "node:path";
import { tmpdir } from "node:os";
import { fileURLToPath } from "node:url";
import { test } from "node:test";
import { duckKeyframes, speechSpans } from "./duck.mjs";
const HERE = dirname(fileURLToPath(import.meta.url));
const SCRIPT = join(HERE, "..", "audio-duck.mjs");
test("speechSpans bridges gaps smaller than mergeGap", () => {
const meta = {
words: [word("w0", "one", 0, 0.5), word("w1", "two", 0.8, 1), word("w2", "three", 2, 2.2)],
};
assert.deepEqual(speechSpans(meta, { mergeGap: 0.4 }), [
{ start: 0, end: 1 },
{ start: 2, end: 2.2 },
]);
});
test("speechSpans refuses multi-line meta without placement (file-relative times)", () => {
const meta = {
voices: [
{ id: "a", words: [word("w0", "one", 0, 1)] },
{ id: "b", words: [word("w1", "two", 0, 1)] },
],
};
assert.throws(() => speechSpans(meta, { mergeGap: 0.2 }), /--sequential or --offsets/);
});
test("speechSpans sequential stacks lines by duration plus gap", () => {
const meta = {
voices: [
{ id: "a", duration_s: 2, words: [word("w0", "one", 0.1, 1.9)] },
{ id: "b", duration_s: 1, words: [word("w1", "two", 0.1, 0.9)] },
],
};
assert.deepEqual(speechSpans(meta, { mergeGap: 0.2, sequential: true, gap: 0.5 }), [
{ start: 0.1, end: 1.9 },
{ start: 2.6, end: 3.4 },
]);
});
test("speechSpans explicit offsets place each line at composition time", () => {
const meta = {
voices: [
{ id: "a", words: [word("w0", "one", 0, 1)] },
{ id: "b", words: [word("w1", "two", 0, 1)] },
],
};
assert.deepEqual(speechSpans(meta, { mergeGap: 0.2, offsets: { a: 0, b: 4 } }), [
{ start: 0, end: 1 },
{ start: 4, end: 5 },
]);
assert.throws(() => speechSpans(meta, { offsets: { a: 0 } }), /missing voice "b"/);
});
test("speechSpans returns empty spans for empty input", () => {
assert.deepEqual(speechSpans({ voices: [] }, { mergeGap: 0.6 }), []);
});
test("duckKeyframes shapes attack and release from base volume", () => {
assert.deepEqual(
duckKeyframes([{ start: 3, end: 5 }], {
duck: 0.25,
attack: 0.15,
release: 0.4,
baseVolume: 0.6,
}),
[
{ time: 3, volume: 0.15, duration: 0.15 },
{ time: 5, volume: 0.6, duration: 0.4 },
],
);
});
test("--json spans match --merge-gap semantics exactly", () => {
const dir = mkdtempSync(join(tmpdir(), "media-use-duck-"));
try {
const metaPath = join(dir, "audio_meta.json");
writeFileSync(
metaPath,
JSON.stringify({
voices: [
{
id: "narration",
words: [
word("w0", "one", 0, 0.4),
word("w1", "two", 0.9, 1.2),
word("w2", "three", 1.8, 2.1),
],
},
],
}),
);
const out = execFileSync(
process.execPath,
[SCRIPT, "--meta", metaPath, "--target", "#bgm", "--merge-gap", "0.6", "--json"],
{ encoding: "utf8" },
);
const parsed = JSON.parse(out);
assert.deepEqual(parsed.spans, [
{ start: 0, end: 1.2 },
{ start: 1.8, end: 2.1 },
]);
} finally {
rmSync(dir, { recursive: true, force: true });
}
});
function word(id, text, start, end) {
return { id, text, start, end };
}
@@ -6,21 +6,65 @@ import { dirname } from "node:path";
const MAX_FREEZE_BYTES = 256 * 1024 * 1024;
export async function freezeUrl(url, destPath) {
const where = String(url).slice(0, 80);
const res = await fetch(url);
if (!res.ok) throw new Error(`freeze failed: HTTP ${res.status} for ${String(url).slice(0, 80)}`);
const bytes = Buffer.from(await res.arrayBuffer());
if (bytes.length === 0)
throw new Error(`freeze failed: empty response for ${String(url).slice(0, 80)}`);
if (bytes.length > MAX_FREEZE_BYTES)
if (!res.ok) throw new Error(`freeze failed: HTTP ${res.status} for ${where}`);
// Fail fast on an advertised oversize body before reading a single byte.
const declared = Number(res.headers.get("content-length"));
if (declared > MAX_FREEZE_BYTES)
throw new Error(
`freeze failed: ${bytes.length} bytes exceeds ${MAX_FREEZE_BYTES} cap for ${String(url).slice(0, 80)}`,
`freeze failed: ${declared} bytes exceeds ${MAX_FREEZE_BYTES} cap for ${where}`,
);
// Stream and abort once the cap is crossed, so a lying/chunked hostile URL
// can't buffer the whole payload into memory before the check (M1).
const chunks = [];
let total = 0;
for await (const chunk of res.body) {
total += chunk.length;
if (total > MAX_FREEZE_BYTES)
throw new Error(`freeze failed: stream exceeds ${MAX_FREEZE_BYTES} cap for ${where}`);
chunks.push(chunk);
}
if (total === 0) throw new Error(`freeze failed: empty response for ${where}`);
mkdirSync(dirname(destPath), { recursive: true });
writeFileSync(destPath, bytes);
return bytes.length;
writeFileSync(destPath, Buffer.concat(chunks, total));
return total;
}
export function freezeLocalFile(srcPath, destPath) {
mkdirSync(dirname(destPath), { recursive: true });
copyFileSync(srcPath, destPath);
}
// Ingest accepts a DIRECT public media URL only — not a platform page. yt-dlp is
// deliberately out (cloud IPs get blocked, and it's brittle); the supported case
// is "user points at their own file or a direct asset link". A direct URL is a
// non-platform host whose path ends in a known media extension.
const PLATFORM_HOSTS =
/(^|\.)(youtube\.com|youtu\.be|vimeo\.com|tiktok\.com|instagram\.com|twitter\.com|x\.com|facebook\.com|dailymotion\.com)$/i;
const MEDIA_EXT = /\.(mp3|wav|m4a|aac|ogg|flac|mp4|mov|webm|mkv|png|jpe?g|webp|gif|svg|avif)$/i;
// SSRF guard (m11): a user-supplied --from URL must not point at the local host
// or a private network. Blocks loopback/localhost, RFC1918, link-local, and the
// IPv6 equivalents on the literal hostname.
// ponytail: literal-host check only; a DNS name that *resolves* to a private IP
// (rebinding) still passes — add resolve-then-check if --from ever fetches from
// untrusted hostnames at scale.
const PRIVATE_HOST =
/^(localhost|.*\.local|.*\.internal|127\.|10\.|0\.|169\.254\.|192\.168\.|172\.(1[6-9]|2\d|3[01])\.|\[?(::1|::ffff:127\.|f[cd][0-9a-f]{2}:|fe80:))/i;
export function isDirectMediaUrl(u) {
let url;
try {
url = new URL(u);
} catch {
return false;
}
if (url.protocol !== "http:" && url.protocol !== "https:") return false;
if (PLATFORM_HOSTS.test(url.hostname)) return false;
if (PRIVATE_HOST.test(url.hostname)) return false;
return MEDIA_EXT.test(url.pathname);
}
@@ -0,0 +1,46 @@
import { strict as assert } from "node:assert";
import { test } from "node:test";
import { isDirectMediaUrl } from "./freeze.mjs";
test("accepts direct public media URLs", () => {
assert.equal(isDirectMediaUrl("https://cdn.example.com/clip.mp4"), true);
assert.equal(isDirectMediaUrl("https://example.com/a/b/track.mp3"), true);
assert.equal(isDirectMediaUrl("http://example.com/logo.svg"), true);
});
test("rejects platform pages (no yt-dlp)", () => {
assert.equal(isDirectMediaUrl("https://www.youtube.com/watch?v=abc"), false);
assert.equal(isDirectMediaUrl("https://youtu.be/abc"), false);
assert.equal(isDirectMediaUrl("https://vimeo.com/12345"), false);
assert.equal(isDirectMediaUrl("https://x.com/u/status/1"), false);
});
test("rejects non-direct / non-media URLs", () => {
assert.equal(isDirectMediaUrl("https://example.com/page"), false, "no media extension");
assert.equal(isDirectMediaUrl("ftp://example.com/a.mp4"), false, "non-http(s)");
assert.equal(isDirectMediaUrl("not a url"), false);
});
test("rejects local / private hosts (SSRF guard, m11)", () => {
for (const u of [
"http://localhost/a.mp4",
"http://127.0.0.1/a.mp4",
"http://127.1.2.3/a.mp4",
"http://0.0.0.0/a.mp4",
"http://10.0.0.5/a.mp4",
"http://192.168.1.1/a.mp4",
"http://172.16.0.1/a.mp4",
"http://172.31.255.255/a.mp4",
"http://169.254.169.254/a.mp4", // cloud metadata endpoint
"http://printer.local/a.mp4",
"http://svc.internal/a.mp4",
"http://[::1]/a.mp4",
"http://[fe80::1]/a.mp4",
"http://[fd00::1]/a.mp4",
]) {
assert.equal(isDirectMediaUrl(u), false, `should block ${u}`);
}
// A public host that merely starts with similar digits is still allowed.
assert.equal(isDirectMediaUrl("https://172.40.0.1/a.mp4"), true, "172.40 is public");
assert.equal(isDirectMediaUrl("https://11.example.com/a.mp4"), true);
});
@@ -0,0 +1,279 @@
// Declarative table of USER-INSTALLED local models, for the spec-gated fallback.
//
// These models run on the user's own machine for their own use; media-use
// recommends, spec-checks, and assists install; it does not bundle, redistribute,
// or sell them. Because nothing is redistributed, selection is purely by
// quality / size / spec-fit / word-timestamp support (there is deliberately NO
// license field gating availability).
//
// Tiers (`small`|`medium`|`large`|`xlarge`) are human labels; `needs.ramMB` is
// what selection actually gates on. selectModel() returns the best model that
// fits the machine's AVAILABLE RAM, best-first: by explicit `rank` when set
// (quality that is NOT size, e.g. ASR), else by RAM footprint (the quality
// proxy for generation). No fit -> recommend the CLI/cloud path.
//
// Picks reflect the 2026 research pass, verified live where noted.
export const CAPABILITIES = ["tts", "asr", "upscale", "videogen", "imagegen"];
const MODELS = {
tts: [
{
id: "kokoro",
tier: "medium",
sizeMB: 330,
needs: { ramMB: 2048, gpu: false },
wordTimestamps: "native",
install: "pip install kokoro",
invoke: "python -m kokoro --text {text} --voice {voice} --out {out}",
notes: "CPU, faster-than-realtime, native per-word timestamps. Default floor.",
},
{
id: "fish-speech",
tier: "large",
sizeMB: 1100,
needs: { ramMB: 16000, gpu: true, vramMB: 12000 },
wordTimestamps: "whisperx", // needs forced alignment (run ASR over output)
install: "pip install fish-speech",
invoke: "fish-speech synth --text {text} --ref {ref} --out {out}",
notes: "Expressive zero-shot voice cloning; meeting pick. WhisperX for word timing.",
},
],
asr: [
// Parakeet is BETTER than Whisper yet SMALLER (0.6B vs 1.5B), so quality is
// not size here: `rank` pins it ahead of whisper regardless of footprint.
// Open ASR Leaderboard avg WER: Parakeet ~6.05% vs whisper-large-v3 7.44%
// (~19% better); on NOISY test-other 4.73% vs 5.96%, and whisper-v3
// hallucinated to 308% WER on meetings where Parakeet held. 5-10x faster.
//
// Cohere Transcribe 2B tops the leaderboard (5.42%) and is nominally the most
// accurate, but its mlx-audio community MLX quants (4bit AND 8bit, with and
// without --language en) produced multilingual token-soup garbage AND ran
// 40-70x slower than Parakeet on a 24GB Mac (live-tested 2026-07). Excluded
// until the mlx-audio Cohere decoder stabilizes; Parakeet is the default.
{
id: "parakeet-mlx",
tier: "small",
rank: 0,
sizeMB: 2400,
needs: { ramMB: 4000, gpu: true },
wordTimestamps: "tokens", // sub-word tokens; merged to words by parakeet-words.mjs
repo: "mlx-community/parakeet-tdt-0.6b-v3",
install:
"uv venv ~/.venvs/parakeet && VIRTUAL_ENV=~/.venvs/parakeet uv pip install parakeet-mlx",
invoke:
"parakeet-mlx {audio} --model mlx-community/parakeet-tdt-0.6b-v3 --output-format json --output-dir {outdir}",
notes:
"NVIDIA Parakeet-TDT 0.6B via parakeet-mlx. VERIFIED on 24GB: accurate transcript, ~3s (cached model) for 8s audio, word timestamps drive transcript-cut. English + 25 European languages. Beats whisper.cpp on accuracy (6.05% vs 7.44% WER) AND speed (5-10x).",
},
{
id: "whisperx",
tier: "medium",
rank: 1,
sizeMB: 1500,
needs: { ramMB: 4096, gpu: false },
wordTimestamps: "native", // faster-whisper + wav2vec2 forced alignment
install: "pip install whisperx",
invoke: "whisperx {audio} --output_format json --out {out}",
notes:
"CPU-only fallback (no GPU): faster-whisper + wav2vec2 forced alignment, native word timestamps. The packaged `hyperframes transcribe` (whisper.cpp) is the zero-setup baseline below this.",
},
],
upscale: [
{
id: "real-esrgan",
tier: "medium",
sizeMB: 70,
needs: { ramMB: 2048, gpu: false },
wordTimestamps: false,
install: "brew install real-esrgan-ncnn-vulkan # or download the ncnn binary",
invoke: "realesrgan-ncnn-vulkan -i {in} -o {out} -s 4",
notes: "ncnn-vulkan binary, CPU-capable. GFPGAN for faces.",
},
{
id: "seedvr2",
tier: "large",
sizeMB: 6000,
needs: { ramMB: 24000, gpu: true, vramMB: 16000 },
wordTimestamps: false,
install: "pip install seedvr2",
invoke: "seedvr2 upscale --in {in} --out {out}",
notes: "Diffusion upscaler, GPU-only. Video2X for video.",
},
],
videogen: [
// 2026-07 X research pass + live verification on a 24GB M-series Mac.
// The Mac-local video story is LTX 2.3 on MLX via dgrauet/ltx-2-mlx (the
// pipeline these weights were converted for; also powers Phosphene).
// Wan 2.x MLX exists only as A14B conversions (too large for consumer
// unified memory); revisit when a 5B Wan MLX conversion lands.
// IMPORTANT: download the weights with a targeted include list first;
// pointing tools at the repo blind snapshot-downloads all 60 GB:
// hf download dgrauet/ltx-2.3-mlx-q4 --include \
// transformer-distilled-1.1.safetensors connector.safetensors \
// "vae_*.safetensors" audio_vae.safetensors vocoder.safetensors "*.json"
{
id: "ltx-2.3-mlx-q4",
tier: "medium",
sizeMB: 20000, // distilled subset; gemma-3-12b-4bit text encoder adds ~7GB
needs: { ramMB: 16384, gpu: true },
wordTimestamps: false,
install:
"git clone https://github.com/dgrauet/ltx-2-mlx && cd ltx-2-mlx && uv sync --all-extras",
invoke:
"ltx-2-mlx generate --prompt {prompt} --distilled --low-ram --model dgrauet/ltx-2.3-mlx-q4 --width {w} --height {h} --frames {frames} --frame-rate 24 --output {out}",
notes:
"LTX 2.3 int4 on MLX. Verified on 24GB unified: 512x320 x 33 frames in ~19 min cold (incl. text-encoder download), t2v with audio. Dims must be multiples of 64. i2v, retake/extend, keyframe interpolation supported.",
},
{
id: "ltx-2.3-mlx-bf16",
tier: "large",
sizeMB: 45000,
needs: { ramMB: 32768, gpu: true },
wordTimestamps: false,
install:
"git clone https://github.com/dgrauet/ltx-2-mlx && cd ltx-2-mlx && uv sync --all-extras",
invoke:
"ltx-2-mlx generate --prompt {prompt} --two-stage --model dgrauet/ltx-2.3-mlx-bf16 --width {w} --height {h} --frames {frames} --frame-rate 24 --output {out}",
notes:
"Full-precision two-stage pipeline (upstream production default). 32GB with --low-ram block streaming; 64-128GB Macs for long/HD runs (the 25s multi-scene spots seen in the wild).",
},
],
imagegen: [
// 2026-07 X research + live verification on a 24GB M-series Mac. mflux
// (FLUX-on-MLX) is the Mac-native runner; FLUX is the quality leader. Two
// hard-won findings baked into `needs.ramMB`:
// 1. The OFFICIAL FLUX repos are HF-gated (license wall). Point --path at a
// non-gated community 4-bit re-upload (self-contained, incl. VAE).
// 2. Without --low-ram, FLUX's T5-XXL text encoder + transformer blow past
// 24GB into swap: a 768x512 run took 90 MINUTES. With --low-ram (streams
// components from disk) the SAME machine did 512x512 in ~20s at 7.6GB
// free. So the medium tier's needs.ramMB is the streamed floor, not the
// resident footprint; the large tiers are the no-streaming thresholds.
// The runner resolves `repo` to a local snapshot (hf download) before --path;
// a bare repo id in --path breaks mlx unflatten.
{
id: "flux-schnell-mflux-q4",
tier: "medium",
sizeMB: 8700,
needs: { ramMB: 8000, gpu: true },
repo: "dhairyashil/FLUX.1-schnell-mflux-4bit",
wordTimestamps: false,
install: "uv venv ~/.venvs/mflux && VIRTUAL_ENV=~/.venvs/mflux uv pip install mflux==0.9.6",
invoke:
"mflux-generate --model schnell --path {model_path} --low-ram --steps 4 --prompt {prompt} --width {w} --height {h} --seed {seed} --output {out}",
notes:
"FLUX.1 schnell int4. VERIFIED on 24GB (7.6GB free): --low-ram 512x512 in ~20s, photoreal. --low-ram is MANDATORY at this tier (streams to avoid swap). Few-step, fast.",
},
{
id: "flux2-klein-mflux-q4",
tier: "large",
sizeMB: 12000,
needs: { ramMB: 32000, gpu: true },
repo: "Runpod/FLUX.2-klein-4B-mflux-4bit",
wordTimestamps: false,
install: "uv venv ~/.venvs/mflux && VIRTUAL_ENV=~/.venvs/mflux uv pip install mflux",
invoke:
"mflux-generate --base-model flux2-klein-4b --path {model_path} --steps 8 --prompt {prompt} --width {w} --height {h} --seed {seed} --output {out}",
notes:
"FLUX.2 Klein 4B int4 (most-downloaded mflux community repo). Newer, higher quality than schnell; full-resident (no streaming) so needs 32GB+ to stay fast. Needs mflux >= 0.18 for the flux2-klein base model.",
},
{
id: "qwen-image-mflux",
tier: "xlarge",
sizeMB: 40000,
needs: { ramMB: 64000, gpu: true },
repo: "Qwen/Qwen-Image",
wordTimestamps: false,
install: "uv venv ~/.venvs/mflux && VIRTUAL_ENV=~/.venvs/mflux uv pip install mflux",
invoke:
"mflux-generate --base-model qwen --steps 20 --prompt {prompt} --width {w} --height {h} --seed {seed} --output {out}",
notes:
"Qwen-Image, top-tier quality. Heavy: 'several minutes' even on 128GB M4 Max, 'almost fried' a 32GB M4 Pro. 64GB+ only. Below that, the cloud upsell (codex) is faster and better.",
},
],
};
function tableFor(capability) {
const t = MODELS[capability];
if (!t) throw new Error(`unknown local-model capability: ${capability}`);
return t;
}
/** All local models for a capability. */
export function listModels(capability) {
return tableFor(capability).slice();
}
/** Does this machine meet a model's needs? Apple Silicon unified memory counts as VRAM. */
export function meetsSpecs(model, specs) {
const n = model.needs || {};
// Gate on AVAILABLE RAM when the probe reported it (the real budget with the
// OS + open apps resident); fall back to total RAM otherwise. Older specs
// objects (and unit fixtures) that only set ramMB keep working unchanged.
const budget = specs.availableRamMB ?? specs.ramMB;
if (n.ramMB && budget < n.ramMB) return false;
if (n.gpu && !specs.gpu?.present) return false;
if (n.vramMB) {
const vram = specs.gpu?.vramMB ?? 0;
if (vram < n.vramMB) return false;
}
return true;
}
// "Best model the machine can run" == best-first among those that fit. Ordering:
// 1. explicit `rank` (lower = better) when a model declares it. Needed where
// quality is NOT size: Parakeet-0.6B beats Whisper-large-1.5B at ASR, so
// footprint would pick the wrong one.
// 2. otherwise RAM footprint descending, the quality proxy for generation
// (a 40GB image model out-renders a 12GB one).
function rankedByPreference(table) {
return [...table].sort((a, b) => {
const ra = a.rank ?? Infinity;
const rb = b.rank ?? Infinity;
if (ra !== rb) return ra - rb;
return (b.needs?.ramMB ?? 0) - (a.needs?.ramMB ?? 0);
});
}
/**
* Pick the best local model the machine can run for a capability: the
* highest-footprint model that fits the available-RAM budget (and GPU/VRAM).
* `preferTier` pins the search to one tier (e.g. force a smaller/faster model).
* Returns `{ model, tier }`, or `{ recommend: "cli", reason }` when nothing fits.
*/
export function selectModel(capability, specs, { preferTier } = {}) {
const table = tableFor(capability);
const pool = preferTier ? table.filter((m) => m.tier === preferTier) : table;
for (const model of rankedByPreference(pool)) {
if (meetsSpecs(model, specs)) return { model, tier: model.tier };
}
const smallest = table.reduce((a, b) => (a.sizeMB <= b.sizeMB ? a : b));
return {
recommend: "cli",
reason: `machine does not meet specs for any local ${capability} model (smallest needs ~${smallest.needs.ramMB}MB RAM${smallest.needs.gpu ? " + GPU" : ""}); use the CLI path instead`,
};
}
/**
* Agent-facing ladder: every model for a capability, best-first, each flagged
* with whether it fits this machine and why. Lets the agent see the RAM-graded
* options and choose (e.g. trade the auto-picked best for a smaller/faster one,
* or step up to a cloud upsell) rather than only getting one auto-selection.
*/
export function describeModelLadder(capability, specs) {
const budget = specs.availableRamMB ?? specs.ramMB;
return rankedByPreference(tableFor(capability)).map((model) => {
const fits = meetsSpecs(model, specs);
return {
id: model.id,
tier: model.tier,
needsRamMB: model.needs?.ramMB ?? 0,
fits,
reason: fits
? `fits (needs ~${model.needs?.ramMB}MB, ${budget}MB available)`
: `too big (needs ~${model.needs?.ramMB}MB${model.needs?.gpu ? " + GPU" : ""}, ${budget}MB available)`,
notes: model.notes,
};
});
}
@@ -0,0 +1,154 @@
import { strict as assert } from "node:assert";
import { test } from "node:test";
import {
listModels,
meetsSpecs,
selectModel,
describeModelLadder,
CAPABILITIES,
} from "./local-models.mjs";
const TIERS = ["small", "medium", "large", "xlarge"];
const strongGpu = {
ramMB: 64000,
gpu: { present: true, kind: "nvidia", vramMB: 24000 },
appleSilicon: false,
};
const cpuOnly = { ramMB: 16000, gpu: { present: false, vramMB: 0 }, appleSilicon: false };
const tiny = { ramMB: 1024, gpu: { present: false, vramMB: 0 }, appleSilicon: false };
test("every capability table is non-empty and well-formed", () => {
for (const cap of CAPABILITIES) {
const models = listModels(cap);
assert.ok(models.length > 0, `no models for ${cap}`);
for (const m of models) {
assert.ok(m.id && m.tier && m.needs, `${cap}/${m.id} missing fields`);
assert.ok(TIERS.includes(m.tier), `${cap}/${m.id} bad tier: ${m.tier}`);
assert.equal(typeof m.install, "string", `${cap}/${m.id} needs an install command`);
assert.equal(typeof m.invoke, "string", `${cap}/${m.id} needs an invoke command`);
// user-installed, local-use-only: there is NO license gate on selection
assert.equal("license" in m, false, `${cap}/${m.id} must not carry a license gate`);
}
}
});
test("meetsSpecs enforces RAM, GPU presence, and VRAM", () => {
const gpuModel = { needs: { ramMB: 8000, gpu: true, vramMB: 12000 } };
assert.equal(meetsSpecs(gpuModel, strongGpu), true);
assert.equal(meetsSpecs(gpuModel, cpuOnly), false, "no GPU -> fails a GPU model");
const cpuModel = { needs: { ramMB: 2000, gpu: false } };
assert.equal(meetsSpecs(cpuModel, cpuOnly), true);
assert.equal(meetsSpecs(cpuModel, tiny), false, "too little RAM");
});
test("Apple Silicon unified memory counts as VRAM", () => {
const apple = {
ramMB: 24000,
appleSilicon: true,
gpu: { present: true, kind: "apple", vramMB: 24000 },
};
const gpuModel = { needs: { ramMB: 8000, gpu: true, vramMB: 16000 } };
assert.equal(meetsSpecs(gpuModel, apple), true);
});
test("selectModel picks the large tier on a strong machine", () => {
const r = selectModel("tts", strongGpu);
assert.equal(r.tier, "large");
assert.ok(r.model.id);
});
test("selectModel falls back to medium on a CPU-only machine", () => {
const r = selectModel("tts", cpuOnly);
assert.equal(r.tier, "medium");
assert.equal(r.model.id, "kokoro", "Kokoro is the CPU/medium default (native word timestamps)");
});
test("selectModel recommends the CLI path when no tier fits", () => {
const r = selectModel("tts", tiny);
assert.equal(r.recommend, "cli");
assert.ok(r.reason && /spec/i.test(r.reason));
assert.equal(r.model, undefined);
});
test("preferTier:'medium' avoids the large model even on a strong machine", () => {
const r = selectModel("tts", strongGpu, { preferTier: "medium" });
assert.equal(r.tier, "medium");
});
test("selectModel gates on AVAILABLE RAM, not total, when both are present", () => {
// 64GB total but only 6GB free right now -> the large tier must not be chosen.
const busy = {
ramMB: 64000,
availableRamMB: 6000,
appleSilicon: true,
gpu: { present: true, kind: "apple", vramMB: 64000 },
};
const r = selectModel("tts", busy);
assert.equal(r.tier, "medium", "available RAM (6GB) rules out the 16GB large tier");
});
test("imagegen is a RAM-graduated ladder; agent picks the best that fits", () => {
const ladder = describeModelLadder("imagegen", {
ramMB: 24000,
availableRamMB: 12000,
appleSilicon: true,
gpu: { present: true, kind: "apple", vramMB: 24000 },
});
// best-first order, each flagged with fit
assert.ok(ladder.length >= 3, "imagegen offers multiple RAM tiers");
assert.ok(
ladder[0].needsRamMB >= ladder[ladder.length - 1].needsRamMB,
"ladder is ordered best (biggest) first",
);
// on 24GB / 12GB-free the schnell --low-ram tier fits, the 32GB+ tiers do not
const fitting = ladder.filter((m) => m.fits);
assert.ok(fitting.length >= 1, "at least the low-ram tier fits a 24GB Mac");
assert.ok(
fitting.every((m) => m.needsRamMB <= 12000),
"only sub-budget models flagged as fitting",
);
const pick = selectModel("imagegen", {
ramMB: 24000,
availableRamMB: 12000,
gpu: { present: true, vramMB: 24000 },
});
assert.equal(
pick.model.id,
"flux-schnell-mflux-q4",
"best fit on 24GB is the low-ram schnell tier",
);
});
test("imagegen on a 64GB Mac steps up to the higher-quality tier", () => {
const pick = selectModel("imagegen", {
ramMB: 96000,
availableRamMB: 80000,
gpu: { present: true, vramMB: 96000 },
});
assert.equal(pick.tier, "xlarge", "80GB free unlocks the top-quality model");
});
test("ASR prefers Parakeet by rank even though it is smaller than whisper", () => {
// quality != size for ASR: Parakeet 0.6B beats whisper-1.5B, so `rank` wins
// over footprint. On a capable machine both fit; Parakeet must be chosen.
const capable = {
ramMB: 24000,
availableRamMB: 12000,
appleSilicon: true,
gpu: { present: true, kind: "apple", vramMB: 24000 },
};
const pick = selectModel("asr", capable);
assert.equal(pick.model.id, "parakeet-mlx", "Parakeet is the rank-0 preferred ASR");
// whisperx (rank 1, CPU-only) is the fallback when no GPU
const cpu = { ramMB: 16000, availableRamMB: 12000, gpu: { present: false, vramMB: 0 } };
assert.equal(selectModel("asr", cpu).model.id, "whisperx", "CPU-only falls back to whisperx");
});
test("ASR offers word-timestamp-capable models (better than plain whisper)", () => {
const asr = listModels("asr");
assert.ok(
asr.every((m) => m.wordTimestamps),
"every ASR model must support word timestamps",
);
});
@@ -0,0 +1,64 @@
import { execFileSync } from "node:child_process";
import { selectModel } from "./local-models.mjs";
import { probeSpecs } from "./specs.mjs";
// Run a USER-INSTALLED local model for a capability (tts/asr/upscale).
// Picks the best tier the machine supports (selectModel), checks the tool is on
// PATH, fills the model's invoke template, and runs it. Returns:
// { model, tier, out } on success
// { recommend:"install", model, command, reason } when the tool isn't installed
// { recommend:"cli", reason } when no tier fits the machine
// `exec` / `which` are injectable for tests.
//
// ponytail: "installed" = the invoke's first token is on PATH (e.g. `whisperx`,
// `realesrgan-ncnn-vulkan`). For `python -m kokoro` this only proves python
// exists; good enough to gate — the recommend.command names the real package.
// Upgrade to a per-tool probe if a "python present but package missing" run ever
// produces a confusing error instead of a clean recommend.
function defaultWhich(bin) {
execFileSync("command", ["-v", bin], { stdio: "ignore", shell: true });
}
function defaultExec(cmd) {
execFileSync(cmd, { stdio: ["ignore", "pipe", "pipe"], shell: true, timeout: 600000 });
}
const fill = (tpl, vars) =>
tpl.replace(/\{(\w+)\}/g, (_, k) => (vars[k] != null ? String(vars[k]) : ""));
export function runLocalModel(capability, opts = {}) {
const {
specs = probeSpecs(),
exec = defaultExec,
which = defaultWhich,
vars = {},
preferTier,
} = opts;
const sel = selectModel(capability, specs, { preferTier });
if (sel.recommend) return sel; // no tier fits -> recommend the CLI path
const { model } = sel;
const bin = model.invoke.split(/\s+/)[0];
try {
which(bin);
} catch {
return {
recommend: "install",
model: model.id,
command: model.install,
reason: `${model.id} not installed`,
};
}
try {
exec(fill(model.invoke, vars));
} catch (e) {
return {
recommend: "install",
model: model.id,
command: model.install,
reason: e.message || String(e),
};
}
return { model: model.id, tier: sel.tier, out: vars.out };
}
@@ -0,0 +1,54 @@
import { strict as assert } from "node:assert";
import { test } from "node:test";
import { runLocalModel } from "./local-run.mjs";
const strongCpu = { ramMB: 16000, gpu: { present: false, vramMB: 0 }, appleSilicon: false };
const tiny = { ramMB: 512, gpu: { present: false, vramMB: 0 }, appleSilicon: false };
const ok = () => {}; // which/exec that succeed
test("recommends the CLI path when no local tier fits the machine", () => {
const r = runLocalModel("tts", { specs: tiny, which: ok, exec: ok });
assert.equal(r.recommend, "cli");
});
test("recommends install when the tool is not on PATH", () => {
const r = runLocalModel("tts", {
specs: strongCpu,
which: () => {
throw new Error("not found");
},
exec: ok,
vars: { text: "hi", out: "/tmp/v.wav" },
});
assert.equal(r.recommend, "install");
assert.equal(r.model, "kokoro");
assert.match(r.command, /pip install kokoro/);
});
test("runs the model and returns the output path when installed", () => {
let ran = "";
const r = runLocalModel("tts", {
specs: strongCpu,
which: ok,
exec: (cmd) => {
ran = cmd;
},
vars: { text: "hello world", voice: "af_heart", out: "/tmp/v.wav" },
});
assert.equal(r.model, "kokoro");
assert.equal(r.out, "/tmp/v.wav");
assert.match(ran, /hello world/, "invoke template filled with vars");
assert.match(ran, /\/tmp\/v\.wav/);
});
test("a failing run degrades to an install recommendation, never throws", () => {
const r = runLocalModel("upscale", {
specs: strongCpu,
which: ok,
exec: () => {
throw new Error("boom");
},
vars: { in: "a.png", out: "b.png" },
});
assert.equal(r.recommend, "install");
});

Some files were not shown because too many files have changed in this diff Show More