feat(phase-08/11): audio generation

Bigram codec-token model with per-style conditional training, temperature sampling, VALL-E-style 3-second prompt continuation. Covers Encodec / DAC / SoundStream + token-AR vs flow-matching production stacks.
This commit is contained in:
Rohit Ghumare
2026-04-23 00:16:40 +01:00
parent 40ac63bfc9
commit 6937debf64
5 changed files with 322 additions and 0 deletions
@@ -0,0 +1,70 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 900 520" font-family="Georgia, 'Times New Roman', serif">
<defs>
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
</marker>
<style>
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
.cold { fill: #eaf4ff; stroke: #2c5f8c; stroke-width: 1.5; }
.label { font-size: 14px; font-weight: 600; fill: #1a1a1a; }
.content { font-size: 12px; fill: #333; }
.mono { font-size: 12px; fill: #333; font-family: 'Menlo', monospace; }
.caption { font-size: 11px; fill: #555; font-style: italic; }
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
</style>
</defs>
<text x="450" y="28" text-anchor="middle" class="title">audio gen: codec tokens + transformer or diffusion</text>
<!-- waveform -->
<rect x="30" y="70" width="140" height="60" class="box"/>
<text x="100" y="98" text-anchor="middle" class="content">waveform</text>
<text x="100" y="116" text-anchor="middle" class="caption">24 kHz, 1-D</text>
<line x1="170" y1="100" x2="210" y2="100" stroke="#1a1a1a" stroke-width="1.2" marker-end="url(#arrow)"/>
<!-- codec encoder -->
<rect x="210" y="70" width="160" height="60" class="cold"/>
<text x="290" y="98" text-anchor="middle" class="content">codec encoder</text>
<text x="290" y="116" text-anchor="middle" class="caption">Encodec / DAC / SoundStream</text>
<line x1="370" y1="100" x2="410" y2="100" stroke="#1a1a1a" stroke-width="1.2" marker-end="url(#arrow)"/>
<!-- tokens -->
<rect x="410" y="60" width="220" height="80" class="hot"/>
<text x="520" y="85" text-anchor="middle" class="label">RVQ tokens</text>
<text x="520" y="105" text-anchor="middle" class="mono">K &#215; 75 Hz indices</text>
<text x="520" y="125" text-anchor="middle" class="caption">8 codebooks typical</text>
<line x1="630" y1="100" x2="670" y2="100" stroke="#1a1a1a" stroke-width="1.2" marker-end="url(#arrow)"/>
<rect x="670" y="70" width="180" height="60" class="cold"/>
<text x="760" y="98" text-anchor="middle" class="content">codec decoder</text>
<text x="760" y="116" text-anchor="middle" class="caption">tokens &#8594; wav</text>
<!-- two generators -->
<text x="230" y="190" text-anchor="middle" class="label">token-AR path (MusicGen, VALL-E)</text>
<rect x="40" y="200" width="380" height="130" class="cold"/>
<text x="230" y="225" text-anchor="middle" class="content">decoder-only transformer</text>
<text x="230" y="243" text-anchor="middle" class="mono">p(t_n | t_&lt;n, text_prompt, voice_prompt)</text>
<text x="230" y="265" text-anchor="middle" class="caption">streams naturally (~200 ms TTFB)</text>
<text x="230" y="285" text-anchor="middle" class="caption">delayed-parallel: K offset streams</text>
<text x="230" y="305" text-anchor="middle" class="caption">dominates speech in 2026</text>
<text x="670" y="190" text-anchor="middle" class="label">diffusion / flow path (Stable Audio, AudioLDM)</text>
<rect x="480" y="200" width="380" height="130" class="cold"/>
<text x="670" y="225" text-anchor="middle" class="content">DiT on audio latents</text>
<text x="670" y="243" text-anchor="middle" class="mono">x_t &#8594; x_0 via flow matching</text>
<text x="670" y="265" text-anchor="middle" class="caption">faster total time for long clips</text>
<text x="670" y="285" text-anchor="middle" class="caption">cleaner for music at &gt;=30 s</text>
<text x="670" y="305" text-anchor="middle" class="caption">dominates music generation in 2026</text>
<!-- 2026 stack -->
<rect x="30" y="360" width="830" height="140" class="box"/>
<text x="445" y="385" text-anchor="middle" class="label">2026 production stack</text>
<text x="445" y="410" text-anchor="middle" class="caption">TTS: ElevenLabs V3, OpenAI TTS, GPT-4o realtime, NaturalSpeech 3</text>
<text x="445" y="430" text-anchor="middle" class="caption">Music: Suno v4, Udio, Stable Audio 2.5, MusicGen 3.3B</text>
<text x="445" y="450" text-anchor="middle" class="caption">SFX: AudioCraft 2, ElevenLabs SFX, Stable Audio Open</text>
<text x="445" y="475" text-anchor="middle" class="caption">Voice clone: XTTS v2 (open), ElevenLabs Pro (consent-verified)</text>
</svg>

After

Width:  |  Height:  |  Size: 4.3 KiB

@@ -0,0 +1,99 @@
import math
import random
VOCAB = 16
NUM_STYLES = 2
def make_tokens(style, length, rng):
"""Synthetic 'audio token' sequences by style."""
if style == 0: # alternating, speech-like
return [(i + rng.randint(0, 1)) % VOCAB for i in range(length)]
return [(i * 3 + rng.randint(0, 1)) % VOCAB for i in range(length)]
def init_counts():
return [[[1.0 for _ in range(VOCAB)] for _ in range(VOCAB)] for _ in range(NUM_STYLES)]
def update_counts(counts, sequence, style):
for i in range(len(sequence) - 1):
counts[style][sequence[i]][sequence[i + 1]] += 1.0
def probs(counts, style, prev_tok):
row = counts[style][prev_tok]
total = sum(row)
return [x / total for x in row]
def entropy(p):
return -sum(pi * math.log(max(pi, 1e-10)) for pi in p)
def sample_from(p, rng):
r = rng.random()
acc = 0.0
for i, pi in enumerate(p):
acc += pi
if r <= acc:
return i
return len(p) - 1
def generate(counts, style, start, length, rng, temperature=1.0):
out = [start]
for _ in range(length - 1):
p = probs(counts, style, out[-1])
if temperature != 1.0:
p = [pi ** (1 / temperature) for pi in p]
total = sum(p)
p = [x / total for x in p]
out.append(sample_from(p, rng))
return out
def main():
rng = random.Random(42)
counts = init_counts()
print("=== training codec-token bigram per style on 500 sequences each ===")
for _ in range(500):
for style in range(NUM_STYLES):
seq = make_tokens(style, length=20, rng=rng)
update_counts(counts, seq, style)
print()
print("=== generate 20 tokens per style, start=0 ===")
for style in range(NUM_STYLES):
label = "speech-like (alternating)" if style == 0 else "music-like (ramp)"
print(f"\nstyle {style}: {label}")
for temp in [0.7, 1.0]:
out = generate(counts, style, start=0, length=20, rng=rng, temperature=temp)
print(f" temp {temp:.1f}: {out}")
print()
print("=== entropy at each position for style 0 conditional on token 5 ===")
p = probs(counts, 0, 5)
top3 = sorted(range(VOCAB), key=lambda i: -p[i])[:3]
print(f" p(next | style=0, prev=5): H = {entropy(p):.3f}")
print(f" top-3: {[(i, round(p[i], 3)) for i in top3]}")
print()
print("=== VALL-E-style prompt continuation ===")
prompt = make_tokens(0, length=5, rng=rng)[:5]
print(f" 3-second voice prompt (tokens): {prompt}")
continuation = list(prompt)
for _ in range(15):
p = probs(counts, 0, continuation[-1])
continuation.append(sample_from(p, rng))
print(f" continuation: {continuation}")
print()
print("takeaway: tokens + transformer = entire TTS / music generation substrate.")
print(" RVQ of Encodec / DAC makes real audio fit in the same loop.")
if __name__ == "__main__":
main()
@@ -0,0 +1,134 @@
# Audio Generation
> Audio is a 1-D signal at 16-48 kHz. A five-second clip is 80-240k samples. No transformer attends to that sequence directly. The solution for every production audio model in 2026 is the same: a neural codec (Encodec, SoundStream, DAC) compresses audio to discrete tokens at 50-75 Hz, and a transformer or diffusion model generates tokens.
**Type:** Build
**Languages:** Python
**Prerequisites:** Phase 6 · 02 (Audio Features), Phase 6 · 04 (ASR), Phase 8 · 06 (DDPM)
**Time:** ~45 minutes
## The Problem
Three audio generation tasks:
1. **Text-to-speech.** Given text, produce speech. Clean speech is narrow-band and has strong phonetic structure — solved well by transformer-over-tokens. VALL-E (Microsoft), NaturalSpeech 3, ElevenLabs, OpenAI TTS.
2. **Music generation.** Given a prompt (text, melody, chord progression, genre), produce music. Much broader distribution. MusicGen (Meta), Stable Audio 2.5, Suno v4, Udio, Riffusion.
3. **Audio effects / sound design.** Given a prompt, produce ambient sound or Foley. AudioGen, AudioLDM 2, Stable Audio Open.
All three run on the same substrate: neural audio codec + token-AR or diffusion generator.
## The Concept
![Audio generation: codec tokens + transformer or diffusion](../assets/audio-generation.svg)
### Neural audio codecs
Encodec (Meta, 2022), SoundStream (Google, 2021), Descript Audio Codec (DAC, 2023). A convolutional encoder compresses waveform to a per-timestep vector; residual vector quantization (RVQ) converts each vector to a cascade of K codebook indices. Decoder reverses it. 24 kHz audio at 2 kbps using 8 RVQ codebooks at 75 Hz = 600 tokens/sec.
```
waveform (16000 samples/sec)
└─ encoder conv ─┐
├─ RVQ layer 1 → indices at 75 Hz
├─ RVQ layer 2 → indices at 75 Hz
├─ ...
└─ RVQ layer 8
```
### Two generative paradigms on top
**Token-autoregressive.** Flatten RVQ tokens into a sequence, run a decoder-only transformer. MusicGen uses "delayed parallel" to emit K codebook streams in parallel with per-stream offsets. VALL-E generates speech tokens from a text prompt + 3-second voice sample.
**Latent diffusion.** Pack codec tokens as continuous latents or model them with categorical diffusion. Stable Audio 2.5 uses flow matching on continuous audio latents. AudioLDM 2 uses text-to-mel-to-audio diffusion.
The 2024-2026 trend: flow matching is winning for music (faster inference, cleaner samples) while token-AR still dominates speech because it is naturally causal and streams well.
## Production landscape
| System | Task | Backbone | Latency |
|--------|------|----------|---------|
| ElevenLabs V3 | TTS | Token-AR + neural vocoder | ~300ms first token |
| OpenAI GPT-4o audio | Full-duplex speech | End-to-end multimodal AR | ~200ms |
| NaturalSpeech 3 | TTS | Latent flow matching | Non-streaming |
| Stable Audio 2.5 | Music / SFX | DiT + flow matching on audio latents | ~10s for 1-minute clip |
| Suno v4 | Full songs | Undisclosed; token-AR suspected | ~30s per song |
| Udio v1.5 | Full songs | Undisclosed | ~30s per song |
| MusicGen 3.3B | Music | Token-AR on Encodec 32kHz | Real-time |
| AudioCraft 2 | Music + SFX | Flow matching | ~5s for 5s clip |
| Riffusion v2 | Music | Spectrogram diffusion | ~10s |
## Build It
`code/main.py` simulates the core idea: train a tiny next-token transformer on synthetic "audio token" sequences generated from two distinct "styles" (alternating low and high tokens for style A, monotonic ramp for style B). Condition on style and sample.
### Step 1: synthetic audio tokens
```python
def make_tokens(style, length, vocab_size, rng):
if style == 0: # "speech-like": alternating
return [i % vocab_size for i in range(length)]
# "music-like": ramp
return [(i * 3) % vocab_size for i in range(length)]
```
### Step 2: train a tiny token predictor
A bigram-style predictor conditioned on style. The point is the pattern: codec tokens → cross-entropy training → autoregressive sampling.
### Step 3: sample conditionally
Given the style token and a starting token, sample the next token from the predicted distribution. Continue for 20-40 tokens.
## Pitfalls
- **Codec quality caps output quality.** If the codec can't represent a sound faithfully, no amount of generator quality helps. DAC is the current open best.
- **RVQ error accumulation.** Each RVQ layer models the residual of the previous. Errors on layer 1 propagate. Sampling with temperature 0 on higher layers helps.
- **Musical structure.** 30 seconds of tokens is 20k+ tokens at 75 Hz. Hard for transformers. MusicGen uses sliding window + prompt continuation; Stable Audio uses shorter clips + crossfading.
- **Artifacts at boundaries.** Crossfading between generated clips needs careful overlap-add.
- **Clean-data appetite.** Music generators need tens of thousands of hours of licensed music. The Suno / Udio RIAA lawsuit (2024) brought this to the surface.
- **Voice cloning ethics.** A 3-second sample plus a text prompt is enough for VALL-E / XTTS / ElevenLabs to clone a voice. Every production model needs abuse detection + opt-out lists.
## Use It
| Task | 2026 stack |
|------|------------|
| Commercial TTS | ElevenLabs, OpenAI TTS, or Azure Neural |
| Voice cloning (consent-verified) | XTTS v2 (open) or ElevenLabs Pro |
| Background music, fast | Stable Audio 2.5 API, Suno, or Udio |
| Music with lyrics | Suno v4 or Udio v1.5 |
| Sound effects / Foley | AudioCraft 2, ElevenLabs SFX, or Stable Audio Open |
| Real-time voice agent | GPT-4o realtime or Gemini Live |
| Open-weights music research | MusicGen 3.3B, Stable Audio Open 1.0, AudioLDM 2 |
| Dubbing / translation | HeyGen, ElevenLabs Dubbing |
## Ship It
Save `outputs/skill-audio-brief.md`. Skill takes an audio brief (task, duration, style, voice, license) and outputs: model + hosting, prompt format (genre tags, style descriptors, structural markers), codec + generator + vocoder chain, seed protocol, and eval plan (MOS / CLAP score / CER for TTS / user A/B).
## Exercises
1. **Easy.** Run `code/main.py` and set style explicitly. Verify the generated sequences match the style's pattern.
2. **Medium.** Add delayed parallel decoding: simulate 2 streams of tokens that must stay offset by 1 step. Train a joint predictor.
3. **Hard.** Use HuggingFace transformers to run MusicGen-small locally. Generate a 10-second clip with three different prompts; A/B for style adherence.
## Key Terms
| Term | What people say | What it actually means |
|------|-----------------|-----------------------|
| Codec | "Neural compression" | Encoder / decoder for audio; typical output is 50-75 Hz tokens. |
| RVQ | "Residual VQ" | Cascade of K quantizers; each models the residual of the previous. |
| Token | "One codec symbol" | Discrete index into a codebook; 1024 or 2048 typical. |
| Delayed parallel | "Offset codebooks" | Emit K token streams with staggered offsets to reduce sequence length. |
| Flow matching | "The 2024 win for audio" | Straighter-path alternative to diffusion; faster sampling. |
| Voice prompt | "3-second sample" | Speaker embedding or token prefix that steers the cloned voice. |
| Mel spectrogram | "The visual" | Log-magnitude perceptual spectrogram; used by many TTS systems. |
| Vocoder | "Mel to wave" | Neural component that converts mel spectrograms back to audio. |
## Further Reading
- [Défossez et al. (2022). Encodec: High Fidelity Neural Audio Compression](https://arxiv.org/abs/2210.13438) — the codec standard.
- [Zeghidour et al. (2021). SoundStream](https://arxiv.org/abs/2107.03312) — the first widely used neural audio codec.
- [Kumar et al. (2023). High-Fidelity Audio Compression with Improved RVQGAN (DAC)](https://arxiv.org/abs/2306.06546) — DAC.
- [Wang et al. (2023). Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers (VALL-E)](https://arxiv.org/abs/2301.02111) — VALL-E.
- [Copet et al. (2023). Simple and Controllable Music Generation (MusicGen)](https://arxiv.org/abs/2306.05284) — MusicGen.
- [Liu et al. (2023). AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining](https://arxiv.org/abs/2308.05734) — AudioLDM 2.
- [Stability AI (2024). Stable Audio 2.5](https://stability.ai/news/introducing-stable-audio-2-5) — 2025 text-to-music with flow matching.
@@ -0,0 +1,19 @@
---
name: audio-brief
description: Translate an audio brief into a model + prompt + eval plan across TTS, music, and SFX.
version: 1.0.0
phase: 8
lesson: 11
tags: [audio, tts, music, sfx, codec]
---
Given an audio brief (task: TTS / music / SFX / voice clone, duration, style, voice or genre, license constraints, real-time or offline, quality bar), output:
1. Model + hosting. ElevenLabs V3, OpenAI TTS, XTTS v2, Suno v4, Udio, Stable Audio 2.5, MusicGen 3.3B, AudioCraft 2, or GPT-4o realtime. One-sentence reason.
2. Prompt format. TTS: text + voice prompt (3-10 s sample or voice ID) + emotion / pace tags. Music: genre + instrumentation + mood + BPM + structural markers. SFX: onomatopoeia + source + duration hint.
3. Codec + generator + vocoder chain. Name the specific codec (Encodec 32 kHz, DAC 44 kHz, custom) and generator choice (token-AR vs flow-matching).
4. Seed + reproducibility. Seed pin, version pin, prompt hash.
5. Eval. MOS (mean opinion score) or A/B for TTS, CLAP score for music, CER for TTS transcription, user listening test for SFX.
6. Guardrails. Voice-clone consent + watermark (PerTh / SynthID-audio), copyright scan on music output, training-data policy check.
Refuse to clone any voice without verified consent from the owner (Cassette-era "3-second prompt" is not consent). Refuse to ship music with unlicensed reference material. Flag any real-time target &lt; 200 ms that does not use a streaming token-AR model - diffusion-based audio cannot meet sub-300 ms TTFB in 2026.