mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-12/19): audio-language models from Whisper to AF3
This commit is contained in:
@@ -0,0 +1,89 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 960 520" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arrow" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.cool { fill: #e6f4ea; stroke: #2e7d32; stroke-width: 1.5; }
|
||||
.cold { fill: #dfe9ff; stroke: #2c5ea9; stroke-width: 1.5; }
|
||||
.reg { fill: #e9e6ff; stroke: #5a4fcf; stroke-width: 1.5; }
|
||||
.step { font-size: 12px; font-family: 'Menlo', monospace; fill: #222; }
|
||||
.small { font-size: 10px; font-family: 'Menlo', monospace; fill: #555; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 16px; font-weight: 700; fill: #1a1a1a; }
|
||||
.head { font-size: 12px; font-weight: 700; fill: #1a1a1a; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="480" y="24" text-anchor="middle" class="title">Audio-LLM arc: Whisper (2022) to Audio Flamingo 3 (2025)</text>
|
||||
|
||||
<rect x="30" y="50" width="900" height="230" class="box"/>
|
||||
<text x="480" y="72" text-anchor="middle" class="head">the pipeline: spectrogram -> encoder -> Q-former -> LLM</text>
|
||||
|
||||
<rect x="50" y="90" width="180" height="170" class="hot"/>
|
||||
<text x="140" y="110" text-anchor="middle" class="step">1. waveform</text>
|
||||
<text x="140" y="128" text-anchor="middle" class="small">16 kHz mono</text>
|
||||
<text x="140" y="150" text-anchor="middle" class="step">2. log-Mel spec</text>
|
||||
<text x="140" y="168" text-anchor="middle" class="small">25ms win, 10ms hop</text>
|
||||
<text x="140" y="184" text-anchor="middle" class="small">80 Mel bins</text>
|
||||
<text x="140" y="200" text-anchor="middle" class="small">log compress</text>
|
||||
<text x="140" y="222" text-anchor="middle" class="small">30s = 3000 frames</text>
|
||||
|
||||
<path d="M 235 170 L 275 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="280" y="90" width="200" height="170" class="cool"/>
|
||||
<text x="380" y="110" text-anchor="middle" class="step">3. audio encoder</text>
|
||||
<text x="380" y="132" text-anchor="middle" class="small">Whisper: speech strong</text>
|
||||
<text x="380" y="148" text-anchor="middle" class="small">BEATs: music strong</text>
|
||||
<text x="380" y="164" text-anchor="middle" class="small">AF-Whisper: concat</text>
|
||||
<text x="380" y="186" text-anchor="middle" class="small">1 frame per 10ms</text>
|
||||
<text x="380" y="210" text-anchor="middle" class="small">12-layer transformer</text>
|
||||
<text x="380" y="230" text-anchor="middle" class="caption">frozen at bridge train</text>
|
||||
|
||||
<path d="M 485 170 L 525 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="530" y="90" width="200" height="170" class="cold"/>
|
||||
<text x="630" y="110" text-anchor="middle" class="step">4. audio Q-former</text>
|
||||
<text x="630" y="132" text-anchor="middle" class="small">32-64 learnable queries</text>
|
||||
<text x="630" y="148" text-anchor="middle" class="small">cross-attend over frames</text>
|
||||
<text x="630" y="164" text-anchor="middle" class="small">output fixed-length tokens</text>
|
||||
<text x="630" y="186" text-anchor="middle" class="step">training</text>
|
||||
<text x="630" y="202" text-anchor="middle" class="small">stage 1: ITM + ITC + ITG</text>
|
||||
<text x="630" y="218" text-anchor="middle" class="small">stage 2: instruction tune</text>
|
||||
|
||||
<path d="M 735 170 L 775 170" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arrow)"/>
|
||||
|
||||
<rect x="780" y="90" width="140" height="170" class="reg"/>
|
||||
<text x="850" y="110" text-anchor="middle" class="step">5. LLM</text>
|
||||
<text x="850" y="128" text-anchor="middle" class="small">Qwen2.5-7B</text>
|
||||
<text x="850" y="144" text-anchor="middle" class="small">or Llama 3.1</text>
|
||||
<text x="850" y="166" text-anchor="middle" class="step">output</text>
|
||||
<text x="850" y="184" text-anchor="middle" class="small">captions</text>
|
||||
<text x="850" y="200" text-anchor="middle" class="small">QA answers</text>
|
||||
<text x="850" y="216" text-anchor="middle" class="small">with CoT</text>
|
||||
|
||||
<rect x="30" y="300" width="900" height="210" class="box"/>
|
||||
<text x="480" y="322" text-anchor="middle" class="head">cascaded vs end-to-end task coverage</text>
|
||||
|
||||
<rect x="50" y="340" width="400" height="160" class="hot"/>
|
||||
<text x="250" y="362" text-anchor="middle" class="step">cascaded (Whisper -> LLM)</text>
|
||||
<text x="250" y="384" text-anchor="middle" class="small">transcription: yes</text>
|
||||
<text x="250" y="400" text-anchor="middle" class="small">summarization: yes</text>
|
||||
<text x="250" y="416" text-anchor="middle" class="small">emotion: no</text>
|
||||
<text x="250" y="432" text-anchor="middle" class="small">music genre: no</text>
|
||||
<text x="250" y="448" text-anchor="middle" class="small">environmental: no</text>
|
||||
<text x="250" y="464" text-anchor="middle" class="small">deepfake: no</text>
|
||||
<text x="250" y="488" text-anchor="middle" class="caption">MMAU ~0.50</text>
|
||||
|
||||
<rect x="470" y="340" width="440" height="160" class="cool"/>
|
||||
<text x="690" y="362" text-anchor="middle" class="step">end-to-end audio-LLM (AF3)</text>
|
||||
<text x="690" y="384" text-anchor="middle" class="small">every cascaded task: yes</text>
|
||||
<text x="690" y="400" text-anchor="middle" class="small">emotion / mood: yes</text>
|
||||
<text x="690" y="416" text-anchor="middle" class="small">music / instruments: yes</text>
|
||||
<text x="690" y="432" text-anchor="middle" class="small">environmental sounds: yes</text>
|
||||
<text x="690" y="448" text-anchor="middle" class="small">temporal grounding: yes</text>
|
||||
<text x="690" y="464" text-anchor="middle" class="small">on-demand CoT: +3-5 pts</text>
|
||||
<text x="690" y="488" text-anchor="middle" class="caption">MMAU 0.72 (open SOTA 2025)</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.7 KiB |
@@ -0,0 +1,165 @@
|
||||
"""Audio-LLM toys: log-Mel spectrogram + audio Q-former + cascaded vs end-to-end.
|
||||
|
||||
Stdlib. Computes a naive DFT-based log-Mel spec from a synthetic waveform,
|
||||
runs a toy Q-former over the resulting frames, and compares task coverage
|
||||
between cascaded and end-to-end pipelines.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import math
|
||||
import random
|
||||
from dataclasses import dataclass
|
||||
|
||||
random.seed(6)
|
||||
|
||||
|
||||
def synth_waveform(duration_s: float = 1.0, sr: int = 16000) -> list[float]:
|
||||
n = int(duration_s * sr)
|
||||
freq = 440
|
||||
return [0.5 * math.sin(2 * math.pi * freq * i / sr) +
|
||||
0.2 * math.sin(2 * math.pi * 880 * i / sr)
|
||||
for i in range(n)]
|
||||
|
||||
|
||||
def window_frames(x: list[float], sr: int, win_ms: int = 25, hop_ms: int = 10) -> list[list[float]]:
|
||||
win = int(sr * win_ms / 1000)
|
||||
hop = int(sr * hop_ms / 1000)
|
||||
frames = []
|
||||
i = 0
|
||||
while i + win <= len(x):
|
||||
frames.append(x[i:i + win])
|
||||
i += hop
|
||||
return frames
|
||||
|
||||
|
||||
def naive_dft_mag(frame: list[float], n_bins: int = 64) -> list[float]:
|
||||
"""Compute magnitude spectrum at n_bins frequencies using naive DFT."""
|
||||
n = len(frame)
|
||||
out = []
|
||||
for k in range(n_bins):
|
||||
re = 0.0
|
||||
im = 0.0
|
||||
for i, x in enumerate(frame):
|
||||
angle = -2 * math.pi * k * i / n
|
||||
re += x * math.cos(angle)
|
||||
im += x * math.sin(angle)
|
||||
out.append(math.sqrt(re * re + im * im))
|
||||
return out
|
||||
|
||||
|
||||
def mel_filterbank(n_bins: int = 64, n_mels: int = 20) -> list[list[float]]:
|
||||
"""Triangular Mel filter bank (simplified, linear warp as proxy)."""
|
||||
fbank = []
|
||||
band = n_bins // n_mels
|
||||
for m in range(n_mels):
|
||||
row = [0.0] * n_bins
|
||||
start = m * band
|
||||
end = min(start + band, n_bins)
|
||||
for k in range(start, end):
|
||||
row[k] = 1.0 / (end - start)
|
||||
fbank.append(row)
|
||||
return fbank
|
||||
|
||||
|
||||
def apply_mel(spec_mag: list[float], fbank: list[list[float]]) -> list[float]:
|
||||
return [sum(w * s for w, s in zip(row, spec_mag)) for row in fbank]
|
||||
|
||||
|
||||
def log_compress(xs: list[float]) -> list[float]:
|
||||
return [math.log(1 + x) for x in xs]
|
||||
|
||||
|
||||
def demo_melspec() -> None:
|
||||
print("\nLOG-MEL SPECTROGRAM (1s @ 16kHz, 25ms win, 10ms hop, 20 mel bins)")
|
||||
print("-" * 60)
|
||||
wave = synth_waveform(1.0, 16000)
|
||||
frames = window_frames(wave, 16000, 25, 10)
|
||||
print(f" frames : {len(frames)} (should be ~99 at 1s)")
|
||||
|
||||
spec = naive_dft_mag(frames[0], n_bins=64)
|
||||
fbank = mel_filterbank(n_bins=64, n_mels=20)
|
||||
mel = apply_mel(spec, fbank)
|
||||
log_mel = log_compress(mel)
|
||||
print(f" per-frame mel dim: {len(mel)}")
|
||||
print(f" first frame log-mel (rounded): "
|
||||
f"{[round(v, 2) for v in log_mel[:10]]}...")
|
||||
|
||||
|
||||
@dataclass
|
||||
class QFormer:
|
||||
n_queries: int
|
||||
hidden: int
|
||||
|
||||
def __post_init__(self):
|
||||
self.queries = [[random.gauss(0, 0.1) for _ in range(self.hidden)]
|
||||
for _ in range(self.n_queries)]
|
||||
|
||||
def forward(self, frames: list[list[float]]) -> list[list[float]]:
|
||||
"""Naive cross-attention: each query attends over all frames."""
|
||||
out = []
|
||||
for q in self.queries:
|
||||
scores = [sum(qi * fi for qi, fi in zip(q, f)) for f in frames]
|
||||
m = max(scores)
|
||||
exps = [math.exp(s - m) for s in scores]
|
||||
z = sum(exps)
|
||||
weights = [e / z for e in exps]
|
||||
agg = [sum(w * f[k] for w, f in zip(weights, frames))
|
||||
for k in range(self.hidden)]
|
||||
out.append(agg)
|
||||
return out
|
||||
|
||||
|
||||
def demo_qformer() -> None:
|
||||
print("\nAUDIO Q-FORMER (N=8 queries over 20-dim frames)")
|
||||
print("-" * 60)
|
||||
frames = [[random.gauss(0, 1) for _ in range(20)] for _ in range(99)]
|
||||
qf = QFormer(n_queries=8, hidden=20)
|
||||
tokens = qf.forward(frames)
|
||||
print(f" input frames: {len(frames)}")
|
||||
print(f" output tokens: {len(tokens)} of dim {len(tokens[0])}")
|
||||
print(" each token attends over the full audio by soft attention weights")
|
||||
|
||||
|
||||
def task_coverage_table() -> None:
|
||||
print("\nCASCADED (Whisper -> LLM) vs END-TO-END AUDIO-LLM")
|
||||
print("-" * 60)
|
||||
tasks = [
|
||||
("transcription", "yes", "yes"),
|
||||
("keyword extraction", "yes", "yes"),
|
||||
("summarization", "yes", "yes"),
|
||||
("speaker diarization", "partial", "yes"),
|
||||
("emotion inference", "no", "yes"),
|
||||
("music genre classification","no", "yes"),
|
||||
("instrument recognition", "no", "yes"),
|
||||
("environmental sound ID", "no", "yes"),
|
||||
("temporal event grounding", "partial", "yes"),
|
||||
("deepfake detection", "no", "yes"),
|
||||
]
|
||||
print(f" {'task':<30}{'cascaded':<14}{'end-to-end'}")
|
||||
for name, cas, e2e in tasks:
|
||||
print(f" {name:<30}{cas:<14}{e2e}")
|
||||
print("\n cascaded: fast + reliable for text-extractable signals")
|
||||
print(" end-to-end: required for acoustic-only signals (~40% of MMAU)")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 60)
|
||||
print("AUDIO-LANGUAGE: WHISPER TO AF3 (Phase 12, Lesson 19)")
|
||||
print("=" * 60)
|
||||
|
||||
demo_melspec()
|
||||
demo_qformer()
|
||||
task_coverage_table()
|
||||
|
||||
print("\n2026 RECIPE")
|
||||
print("-" * 60)
|
||||
print(" encoder : AF-Whisper + BEATs concat")
|
||||
print(" bridge : 64-query Q-former")
|
||||
print(" LLM : Qwen2.5-7B with audio tokens")
|
||||
print(" training: AudioCaps + Clotho + MMAU-style instructions")
|
||||
print(" option : on-demand thinking for complex reasoning")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,153 @@
|
||||
# Audio-Language Models: the Whisper to Audio Flamingo 3 Arc
|
||||
|
||||
> Whisper (Radford et al., December 2022) settled speech recognition — 680k hours of weakly-supervised multilingual speech, a simple encoder-decoder transformer, a benchmark that made every subsequent ASR release cite it. But recognition is not reasoning. Asking "what instruments are in this recording" or "what emotion is the speaker expressing" or "what happened at minute 3" requires audio understanding, not transcription. Qwen-Audio, SALMONN, LTU, and NVIDIA's Audio Flamingo 3 (AF3, July 2025) progressively built that stack: keep Whisper-class encoders, bolt on Q-formers, train on audio-text instruction data, add chain-of-thought reasoning. This lesson walks the arc.
|
||||
|
||||
**Type:** Build
|
||||
**Languages:** Python (stdlib, log-Mel spectrogram + audio Q-former skeleton)
|
||||
**Prerequisites:** Phase 6 (Speech and Audio), Phase 12 · 03 (Q-Former)
|
||||
**Time:** ~180 minutes
|
||||
|
||||
## Learning Objectives
|
||||
|
||||
- Compute a log-Mel spectrogram from a waveform: windowing, FFT, filter banks, log transform.
|
||||
- Compare encoder options: Whisper encoder, BEATs, AF-Whisper hybrid. When each wins.
|
||||
- Build an audio Q-former: N learnable queries cross-attending to spectrogram patches.
|
||||
- Explain cascaded (Whisper-then-LLM) vs end-to-end audio-LLM training: why end-to-end scales better for reasoning.
|
||||
|
||||
## The Problem
|
||||
|
||||
Speech recognition was solved by Whisper. OCR-of-audio is a commodity. But "commodity" stops at transcription. If the model cannot reason over what it heard — timing, speakers, emotion, music structure, environmental sounds — transcription alone cannot drive product features.
|
||||
|
||||
Three obvious routes:
|
||||
|
||||
1. Cascade: Whisper transcribes, LLM reasons over the transcript. Works for pure-speech scenarios. Fails for music, environmental audio, multi-speaker overlap, emotion.
|
||||
|
||||
2. End-to-end audio-LLM: an audio encoder feeds audio tokens directly into an LLM, skipping transcription. Preserves acoustic information (emotion, speaker, environment). Needs new training data.
|
||||
|
||||
3. Hybrid: audio encoder + text decoder that can both transcribe and reason. Qwen-Audio and Audio Flamingo pick this route.
|
||||
|
||||
## The Concept
|
||||
|
||||
### Log-Mel spectrogram: the input feature
|
||||
|
||||
Every audio encoder starts with the same feature: a log-Mel spectrogram.
|
||||
|
||||
1. Resample to 16 kHz.
|
||||
2. Short-time Fourier transform with 25ms windows, 10ms hop.
|
||||
3. Take magnitude of the FFT result.
|
||||
4. Apply Mel filter banks (typically 80 filters log-spaced 0-8000 Hz) to warp to perceptual frequency.
|
||||
5. Log compress (log(1 + x)) for dynamic range.
|
||||
|
||||
Result: a 2D array of shape (T, 80) where T is the number of time frames. For a 30-second clip at 100 Hz frame rate: (3000, 80).
|
||||
|
||||
### Whisper's encoder
|
||||
|
||||
Whisper's encoder is a 12-layer ViT-style transformer processing the log-Mel spectrogram as a sequence of time frames. Output: one hidden-state vector per time frame.
|
||||
|
||||
For ASR, Whisper's decoder is a cross-attention transformer that generates text tokens conditioned on the encoder output. Standard encoder-decoder.
|
||||
|
||||
For ALMs (audio-LLMs), you want the encoder output as input to a different LLM. The pattern: Whisper encoder frozen, Q-former trainable, LLM frozen or tuned.
|
||||
|
||||
### BEATs and audio-specific encoders
|
||||
|
||||
Whisper was trained on speech-dominant data. It is weaker for music and environmental audio.
|
||||
|
||||
BEATs (Chen et al., 2022) is a self-supervised transformer trained on AudioSet. Captures music and environmental sounds better than Whisper at the same parameter count.
|
||||
|
||||
AF-Whisper (Audio Flamingo 3's hybrid): concat Whisper + BEATs features as the audio input. Whisper carries linguistic signal, BEATs carries acoustic signal.
|
||||
|
||||
### Audio Q-former
|
||||
|
||||
Same pattern as BLIP-2's visual Q-former. A fixed number of learnable queries (often 32 or 64) cross-attend over the audio encoder's output frames. The queries become audio tokens consumed by the LLM.
|
||||
|
||||
Training alignment stage: Q-former alone, contrastive + captioning losses on audio-text pairs (AudioCaps, Clotho). Instruction stage: end-to-end, unfreeze LLM, train on instruction data.
|
||||
|
||||
### The arc — SALMONN, Qwen-Audio, AF3
|
||||
|
||||
SALMONN (Tang et al., 2023): Whisper + BEATs + Q-former + LLaMA. The first open audio-LLM with serious reasoning ability. Benchmarks on MMAU show ~0.55 composite.
|
||||
|
||||
Qwen-Audio (Chu et al., 2023): similar architecture, trained on a richer dataset, tuned for multi-turn dialogue. MMAU ~0.60.
|
||||
|
||||
LTU — Listen, Think, Understand (Gong et al., 2023): explicit reasoning data, focus on chain-of-thought over audio clips. Smaller but more focused.
|
||||
|
||||
Audio Flamingo 3 (Goel et al., July 2025): the current open SOTA. 8B LLM backbone (Qwen2 7B), Whisper-large encoder concat BEATs, 64-query Q-former, training on 1M+ audio-text instruction pairs. MMAU 0.72, matches proprietary frontier on some sub-tasks.
|
||||
|
||||
AF3 also introduces on-demand chain-of-thought for audio: the model can optionally emit thinking tokens ("let me identify the instruments first: ...") before the final answer. Accuracy on complex reasoning tasks lifts 3-5 points when thinking is enabled.
|
||||
|
||||
### Cascaded vs end-to-end
|
||||
|
||||
Cascaded pipeline:
|
||||
|
||||
1. Whisper transcribes audio → text.
|
||||
2. LLM reasons over text.
|
||||
|
||||
Works perfectly for "summarize this podcast." Fails for:
|
||||
- "What's the mood of this song?" — mood is in the sound, not words.
|
||||
- "Who is speaking, Alice or Bob?" — requires speaker identification.
|
||||
- "At what second does the explosion happen?" — temporal grounding lost in text.
|
||||
- "Is this real or generated audio?" — deepfake detection needs acoustic features.
|
||||
|
||||
End-to-end preserves acoustic signal. Qwen-Audio and AF3 handle music, environment, and emotion natively.
|
||||
|
||||
### 2026 production recipe
|
||||
|
||||
For a new audio-understanding product:
|
||||
|
||||
- Cascaded if: transcription is the goal, no music, no emotion inference.
|
||||
- AF3 / Qwen-Audio-family if: music, emotion, multi-speaker, or complex audio reasoning.
|
||||
|
||||
Cascaded is cheaper and simpler. End-to-end is more capable.
|
||||
|
||||
### MMAU — the audio reasoning benchmark
|
||||
|
||||
MMAU (Massive Multimodal Audio Understanding) is the 2024-2025 audio reasoning benchmark:
|
||||
|
||||
- 10,000 audio-text QA pairs across speech, music, environmental sounds.
|
||||
- Covers classification, temporal reasoning, causal reasoning, open-ended QA.
|
||||
- Tests what cascaded pipelines systematically miss.
|
||||
|
||||
Open SOTA (AF3) at 0.72; proprietary frontier ~0.78 (Gemini 2.5 Pro, Claude Opus 4.7). The gap is smaller than VideoMME's open-vs-closed delta, indicating audio-LLMs are maturing.
|
||||
|
||||
## Use It
|
||||
|
||||
`code/main.py`:
|
||||
|
||||
- Implements log-Mel spectrogram computation in stdlib: windowing, naive DFT, Mel filter-bank.
|
||||
- Audio Q-former skeleton: given encoder output frames, compute Q, K, V, attention, and emit N tokens.
|
||||
- Cascaded-vs-end-to-end comparison on a toy task.
|
||||
|
||||
## Ship It
|
||||
|
||||
This lesson produces `outputs/skill-audio-llm-pipeline-picker.md`. Given an audio task (transcription, music tagging, emotion inference, multi-speaker diarization, environment classification), it picks cascaded, end-to-end AF3, or a hybrid.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. Compute the log-Mel spectrogram dimension for a 30-second clip at 16kHz, 25ms window, 10ms hop, 80 Mel bins. How does this change at 48kHz?
|
||||
|
||||
2. Why does Whisper underperform on music? What audio features does BEATs capture that Whisper does not?
|
||||
|
||||
3. Audio Q-former with 64 queries vs 32: at what task complexity does 64 pay off? 32 save compute for what?
|
||||
|
||||
4. Read AF3 Section 4 on on-demand thinking. Propose three audio tasks where chain-of-thought helps the most.
|
||||
|
||||
5. Implement a minimal diarization pipeline using AF3's output. How do you signal speaker changes?
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|------------------------|
|
||||
| Log-Mel spectrogram | "Mel features" | 2D (time, frequency) array of log-magnitude values after Mel filter banks |
|
||||
| Audio Q-former | "Audio Perceiver" | Cross-attention bottleneck from audio encoder output to fixed-length queries feeding the LLM |
|
||||
| Cascaded | "ASR-then-LLM" | Pipeline where Whisper transcribes and a text LLM reasons; loses acoustic information |
|
||||
| End-to-end | "Audio-LLM" | Audio features enter the LLM directly via Q-former; preserves acoustic signal |
|
||||
| BEATs | "Audio AudioSet encoder" | SSL transformer trained on AudioSet; strong on music + environmental sounds |
|
||||
| MMAU | "Audio reasoning bench" | 10k QA pairs across speech, music, environment; 2024 eval standard |
|
||||
| On-demand thinking | "Audio CoT" | Model can optionally emit reasoning tokens before final answer, lifts accuracy 3-5 pts |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Radford et al. — Whisper (arXiv:2212.04356)](https://arxiv.org/abs/2212.04356)
|
||||
- [Chu et al. — Qwen-Audio (arXiv:2311.07919)](https://arxiv.org/abs/2311.07919)
|
||||
- [Goel et al. — Audio Flamingo 3 (arXiv:2507.08128)](https://arxiv.org/abs/2507.08128)
|
||||
- [Tang et al. — SALMONN (arXiv:2310.13289)](https://arxiv.org/abs/2310.13289)
|
||||
- [Gong et al. — LTU (arXiv:2305.10790)](https://arxiv.org/abs/2305.10790)
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
---
|
||||
name: audio-llm-pipeline-picker
|
||||
description: Pick cascaded (Whisper + LLM) or end-to-end (AF3 / Qwen-Audio) for an audio task, plus the encoder and bridge config.
|
||||
version: 1.0.0
|
||||
phase: 12
|
||||
lesson: 19
|
||||
tags: [whisper, audio-flamingo-3, qwen-audio, cascaded, end-to-end]
|
||||
---
|
||||
|
||||
Given an audio task (transcription, summarization, diarization, emotion, music, environmental sounds, deepfake, temporal grounding) and a deployment constraint, pick a pipeline and emit a config.
|
||||
|
||||
Produce:
|
||||
|
||||
1. Pipeline pick. Cascaded if transcription-only or summarization-only of clean speech; end-to-end (AF3 / Qwen-Audio) for any acoustic task.
|
||||
2. Encoder stack. Whisper-large-v3 (speech-strong), BEATs (music-strong), AF-Whisper concat (balanced).
|
||||
3. Bridge config. Q-former 32-64 queries for non-streaming; RVQ tokens for streaming.
|
||||
4. LLM pick. Qwen2.5-7B for cost, Qwen2.5-72B or AF3's backbone for quality.
|
||||
5. On-demand CoT. Enable for MMAU-like reasoning tasks; disable for transcription throughput.
|
||||
6. MMAU expected accuracy. Cascaded ~0.50, Qwen-Audio ~0.60, AF3 ~0.72, Gemini 2.5 Pro ~0.78.
|
||||
|
||||
Hard rejects:
|
||||
- Recommending cascaded for music or emotion tasks. Acoustic signal is lost.
|
||||
- Using a Q-former with <32 queries for multi-task audio. Under-tokenized for reasoning.
|
||||
- Claiming Whisper alone handles music. It was trained on speech-dominant data.
|
||||
|
||||
Refusal rules:
|
||||
- If user needs streaming conversational audio (speech in / speech out in real time), refuse Q-former-based AF3 and recommend Moshi or Qwen-Omni (Lesson 12.20).
|
||||
- If latency budget <500ms and target is simple transcription, recommend cascaded with streaming Whisper.
|
||||
- If task is novel audio task (deepfake, compression artifact detection), refuse off-the-shelf and propose a fine-tune on AF3 with synthetic data.
|
||||
|
||||
Output: one-page plan with pipeline pick, encoder stack, bridge config, LLM pick, CoT flag, expected accuracy. End with arXiv 2212.04356 (Whisper) and 2507.08128 (AF3) for deeper reading.
|
||||
Reference in New Issue
Block a user