mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 10:04:49 +08:00
feat(phase-06/17): audio evaluation — WER, MOS, UTMOS, SECS, EER, DER, FAD, MMAU leaderboards
This commit is contained in:
@@ -0,0 +1,81 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 500" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.label { font-size: 13px; font-weight: 700; fill: #1a1a1a; }
|
||||
.content { font-size: 11px; fill: #333; font-family: 'Menlo', monospace; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 15px; font-weight: 700; fill: #1a1a1a; }
|
||||
.tag { font-size: 10px; fill: #c0392b; font-weight: 600; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="440" y="26" text-anchor="middle" class="title">audio evaluation — one metric per task, one leaderboard per metric</text>
|
||||
|
||||
<rect x="20" y="50" width="200" height="200" class="box"/>
|
||||
<text x="120" y="72" text-anchor="middle" class="label">ASR</text>
|
||||
<text x="30" y="96" class="content">primary: WER (normalized)</text>
|
||||
<text x="30" y="114" class="content"> CER for tone langs</text>
|
||||
<text x="30" y="132" class="content">speed: RTFx</text>
|
||||
<text x="30" y="150" class="content">latency: first-token ms</text>
|
||||
<text x="30" y="174" class="tag">Open ASR Leaderboard</text>
|
||||
<text x="30" y="192" class="caption">Parakeet 6.05% avg WER</text>
|
||||
<text x="30" y="210" class="caption">Whisper-v3-turbo 1.58% LS</text>
|
||||
<text x="30" y="228" class="caption">Canary-Qwen 5.63% avg</text>
|
||||
|
||||
<rect x="235" y="50" width="200" height="200" class="hot"/>
|
||||
<text x="335" y="72" text-anchor="middle" class="label">TTS</text>
|
||||
<text x="245" y="96" class="content">primary: MOS / UTMOS</text>
|
||||
<text x="245" y="114" class="content"> SECS (cloning)</text>
|
||||
<text x="245" y="132" class="content">round: ASR-WER on output</text>
|
||||
<text x="245" y="150" class="content">latency: TTFA ms</text>
|
||||
<text x="245" y="174" class="tag">TTS Arena · Artificial Analysis</text>
|
||||
<text x="245" y="192" class="caption">Inworld TTS-1.5 ELO 1236</text>
|
||||
<text x="245" y="210" class="caption">Kokoro-82M ELO 1059</text>
|
||||
<text x="245" y="228" class="caption">ElevenLabs v3 ELO 1179</text>
|
||||
|
||||
<rect x="450" y="50" width="200" height="200" class="box"/>
|
||||
<text x="550" y="72" text-anchor="middle" class="label">speaker + diarization</text>
|
||||
<text x="460" y="96" class="content">primary: EER</text>
|
||||
<text x="460" y="114" class="content"> minDCF</text>
|
||||
<text x="460" y="132" class="content">diariz: DER / JER</text>
|
||||
<text x="460" y="150" class="content">overlap: per-segment F1</text>
|
||||
<text x="460" y="174" class="tag">VoxCeleb1-O / VoxSRC</text>
|
||||
<text x="460" y="192" class="caption">ECAPA 0.87% EER</text>
|
||||
<text x="460" y="210" class="caption">3D-Speaker 0.50% EER</text>
|
||||
<text x="460" y="228" class="caption">pyannote Precision-2 < 0.40%</text>
|
||||
|
||||
<rect x="665" y="50" width="195" height="200" class="box"/>
|
||||
<text x="762" y="72" text-anchor="middle" class="label">audio classification</text>
|
||||
<text x="675" y="96" class="content">exclusive: top-1, top-5</text>
|
||||
<text x="675" y="114" class="content">multi: mAP</text>
|
||||
<text x="675" y="132" class="content">imbalanced: macro F1</text>
|
||||
<text x="675" y="150" class="content"> per-class recall</text>
|
||||
<text x="675" y="174" class="tag">HEAR · AudioSet · ESC-50</text>
|
||||
<text x="675" y="192" class="caption">BEATs-iter3 AudioSet 0.548</text>
|
||||
<text x="675" y="210" class="caption">Audio-MAE Speech Cmd 99.0%</text>
|
||||
<text x="675" y="228" class="caption">BEATs ESC-50 97.0%</text>
|
||||
|
||||
<rect x="20" y="265" width="410" height="220" class="box"/>
|
||||
<text x="225" y="287" text-anchor="middle" class="label">music generation</text>
|
||||
<text x="30" y="313" class="content">primary: FAD (Fréchet Audio Distance)</text>
|
||||
<text x="30" y="331" class="content"> CLAP (text-audio alignment)</text>
|
||||
<text x="30" y="349" class="content">consumer: blind-panel MOS / ELO</text>
|
||||
<text x="30" y="375" class="tag">Suno v5 ELO 1293 (quality leader)</text>
|
||||
<text x="30" y="395" class="caption">Udio v4 — producer tools + stems</text>
|
||||
<text x="30" y="413" class="caption">MusicGen-large on MusicCaps ~4.5 FAD</text>
|
||||
<text x="30" y="441" class="content">always pair FAD with human MOS panel</text>
|
||||
<text x="30" y="459" class="content">open: ACE-Step XL / YuE / Stable Audio Open</text>
|
||||
|
||||
<rect x="450" y="265" width="410" height="220" class="hot"/>
|
||||
<text x="655" y="287" text-anchor="middle" class="label">audio-language + streaming</text>
|
||||
<text x="460" y="313" class="content">LALM reasoning: MMAU-Pro (1800 Q)</text>
|
||||
<text x="460" y="331" class="content">long-audio: LongAudioBench</text>
|
||||
<text x="460" y="349" class="content">captioning: FENSE / SPICE / CIDEr</text>
|
||||
<text x="460" y="375" class="tag">Gemini 2.5 Pro ~60% MMAU-Pro overall</text>
|
||||
<text x="460" y="395" class="caption">GPT-4o Audio 52.5% — Qwen2.5-Omni 52.2%</text>
|
||||
<text x="460" y="413" class="caption">multi-audio subset: ~22% ALL models (near random)</text>
|
||||
<text x="460" y="441" class="content">streaming S2S: latency P50/P95 + WER</text>
|
||||
<text x="460" y="459" class="content">Moshi 200 ms L4 · GPT-4o Realtime ~300 ms</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 5.1 KiB |
@@ -0,0 +1,150 @@
|
||||
"""Audio evaluation metrics, from scratch.
|
||||
|
||||
Implements WER, CER, EER, simple SECS, FAD-shaped embedding distance,
|
||||
and a MMAU-style multiple-choice accuracy. Stdlib-only.
|
||||
|
||||
Run: python3 code/main.py
|
||||
"""
|
||||
|
||||
import math
|
||||
import random
|
||||
|
||||
|
||||
def _edit_distance(a_tokens, b_tokens):
|
||||
dp = [[0] * (len(b_tokens) + 1) for _ in range(len(a_tokens) + 1)]
|
||||
for i in range(len(a_tokens) + 1):
|
||||
dp[i][0] = i
|
||||
for j in range(len(b_tokens) + 1):
|
||||
dp[0][j] = j
|
||||
for i in range(1, len(a_tokens) + 1):
|
||||
for j in range(1, len(b_tokens) + 1):
|
||||
cost = 0 if a_tokens[i - 1] == b_tokens[j - 1] else 1
|
||||
dp[i][j] = min(dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + cost)
|
||||
return dp[len(a_tokens)][len(b_tokens)]
|
||||
|
||||
|
||||
def normalize(text):
|
||||
import re
|
||||
text = text.lower()
|
||||
text = re.sub(r"[^a-z0-9\s]", "", text)
|
||||
text = re.sub(r"\s+", " ", text).strip()
|
||||
return text
|
||||
|
||||
|
||||
def wer(ref, hyp):
|
||||
r, h = normalize(ref).split(), normalize(hyp).split()
|
||||
return _edit_distance(r, h) / max(1, len(r))
|
||||
|
||||
|
||||
def cer(ref, hyp):
|
||||
return _edit_distance(list(ref), list(hyp)) / max(1, len(ref))
|
||||
|
||||
|
||||
def eer_from_scores(same, diff):
|
||||
thresholds = sorted(set(same + diff))
|
||||
best = (1.0, 0.0, 0.0, 0.0)
|
||||
for t in thresholds:
|
||||
far = sum(1 for s in diff if s >= t) / max(1, len(diff))
|
||||
frr = sum(1 for s in same if s < t) / max(1, len(same))
|
||||
if abs(far - frr) < best[0]:
|
||||
best = (abs(far - frr), t, far, frr)
|
||||
gap, t, far, frr = best
|
||||
return (far + frr) / 2, t
|
||||
|
||||
|
||||
def cosine(a, b):
|
||||
dot = sum(x * y for x, y in zip(a, b))
|
||||
na = math.sqrt(sum(x * x for x in a)) or 1e-12
|
||||
nb = math.sqrt(sum(x * x for x in b)) or 1e-12
|
||||
return dot / (na * nb)
|
||||
|
||||
|
||||
def embedding_fad_like(real_embeds, fake_embeds):
|
||||
def mean_var(embs):
|
||||
n = len(embs[0])
|
||||
mean = [sum(e[i] for e in embs) / len(embs) for i in range(n)]
|
||||
var = [sum((e[i] - mean[i]) ** 2 for e in embs) / len(embs) for i in range(n)]
|
||||
return mean, var
|
||||
mu_r, v_r = mean_var(real_embeds)
|
||||
mu_f, v_f = mean_var(fake_embeds)
|
||||
mean_dist = sum((a - b) ** 2 for a, b in zip(mu_r, mu_f))
|
||||
var_dist = sum((math.sqrt(a) - math.sqrt(b)) ** 2 for a, b in zip(v_r, v_f))
|
||||
return math.sqrt(mean_dist + var_dist)
|
||||
|
||||
|
||||
def mmau_accuracy(predictions, golds):
|
||||
correct = sum(1 for p, g in zip(predictions, golds) if p == g)
|
||||
return correct / max(1, len(predictions))
|
||||
|
||||
|
||||
def main():
|
||||
print("=== WER + CER ===")
|
||||
pairs = [
|
||||
("turn on the kitchen lights", "turn off the kitchen lights"),
|
||||
("what's the weather today", "what is the weather today"),
|
||||
("play jazz", "play jazz"),
|
||||
("set a 5 minute timer", "set a five minute timer"),
|
||||
]
|
||||
for ref, hyp in pairs:
|
||||
print(f" ref: {ref!r}")
|
||||
print(f" hyp: {hyp!r}")
|
||||
print(f" WER = {wer(ref, hyp):.3f} CER = {cer(ref, hyp):.3f}")
|
||||
|
||||
print()
|
||||
print("=== EER (toy speaker verification) ===")
|
||||
random.seed(0)
|
||||
rng = random.Random(0)
|
||||
same = [rng.gauss(0.80, 0.06) for _ in range(100)]
|
||||
diff = [rng.gauss(0.20, 0.15) for _ in range(500)]
|
||||
eer, t = eer_from_scores(same, diff)
|
||||
print(f" same mean cos: {sum(same)/len(same):.3f}")
|
||||
print(f" diff mean cos: {sum(diff)/len(diff):.3f}")
|
||||
print(f" EER = {eer * 100:.2f}% at threshold {t:.3f}")
|
||||
|
||||
print()
|
||||
print("=== SECS (toy voice-cloning similarity) ===")
|
||||
ref_emb = [rng.gauss(0, 0.1) for _ in range(192)]
|
||||
clone_emb = [ref_emb[i] + rng.gauss(0, 0.1) for i in range(192)]
|
||||
secs = cosine(ref_emb, clone_emb)
|
||||
print(f" SECS = {secs:.3f} (target: > 0.75 for recognizable clone)")
|
||||
|
||||
print()
|
||||
print("=== FAD-shaped embedding distance ===")
|
||||
real_embs = [[rng.gauss(0, 1.0) for _ in range(32)] for _ in range(50)]
|
||||
fake_embs = [[rng.gauss(0.1, 1.1) for _ in range(32)] for _ in range(50)]
|
||||
fad = embedding_fad_like(real_embs, fake_embs)
|
||||
print(f" FAD-like = {fad:.3f} (MusicGen-small on MusicCaps: 4.5)")
|
||||
|
||||
print()
|
||||
print("=== MMAU-Pro-style multiple-choice accuracy ===")
|
||||
predictions = ["A", "C", "B", "A", "D", "C", "B", "A", "A", "C"]
|
||||
golds = ["A", "B", "B", "A", "D", "A", "B", "A", "C", "C"]
|
||||
acc = mmau_accuracy(predictions, golds)
|
||||
print(f" accuracy = {acc:.3f} (random on 4-way: 0.250)")
|
||||
|
||||
print()
|
||||
print("=== 2026 benchmarks worth knowing ===")
|
||||
rows = [
|
||||
("Open ASR Leaderboard", "LibriSpeech + multilingual", "Parakeet-TDT 6.05%, Whisper-LV3-turbo 1.58%"),
|
||||
("TTS Arena", "blind pairwise TTS", "Kokoro ELO 1059, ElevenLabs v3 1179"),
|
||||
("Artificial Analysis Speech", "TTS + STT arena", "Inworld TTS-1.5-Max ELO 1236 leader"),
|
||||
("MMAU-Pro", "LALM reasoning", "Gemini 2.5 Pro ~60%, GPT-4o Audio 52.5%"),
|
||||
("LongAudioBench", "multi-minute LALM", "Audio Flamingo Next beats Gemini 2.5 Pro"),
|
||||
("VoxCeleb1-O", "speaker verification EER", "ECAPA 0.87%, 3D-Speaker 0.50%"),
|
||||
("AudioSet mAP", "multi-label classification", "BEATs-iter3 0.548 mAP"),
|
||||
("ASVspoof 5", "anti-spoofing EER", "SOTA ~7.23% on in-the-wild"),
|
||||
]
|
||||
print(" | leaderboard | axis | 2026 SOTA |")
|
||||
for name, axis, sota in rows:
|
||||
print(f" | {name:<24} | {axis:<25} | {sota:<43} |")
|
||||
|
||||
print()
|
||||
print("takeaways:")
|
||||
print(" - every task has 2-3 primary metrics; choose BEFORE training")
|
||||
print(" - normalize text before computing WER/CER; report the normalization")
|
||||
print(" - report P50/P95/P99 for latency, per-class for classification, per-category for MMAU")
|
||||
print(" - public benchmark + your own held-out domain set = both, always")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,224 @@
|
||||
# Audio Evaluation — WER, MOS, UTMOS, MMAU, FAD, and the Open Leaderboards
|
||||
|
||||
> You cannot ship what you cannot measure. This lesson names the 2026 metrics for every audio task: ASR (WER, CER, RTFx), TTS (MOS, UTMOS, SECS, WER-on-ASR-round-trip), audio-language (MMAU, LongAudioBench), music (FAD, CLAP), and speaker (EER). Plus the leaderboards where you compare.
|
||||
|
||||
**Type:** Learn
|
||||
**Languages:** Python
|
||||
**Prerequisites:** Phase 6 · 04, 06, 07, 09, 10; Phase 2 · 09 (Model Evaluation)
|
||||
**Time:** ~60 minutes
|
||||
|
||||
## The Problem
|
||||
|
||||
Every audio task has multiple metrics, each measuring a different axis. Using the wrong metric is how you ship a model that looks great on your dashboard and terribly in production. The 2026 canonical list:
|
||||
|
||||
| Task | Primary | Secondary |
|
||||
|------|---------|-----------|
|
||||
| ASR | WER | CER · RTFx · first-token latency |
|
||||
| TTS | MOS / UTMOS | SECS · WER-on-ASR-round-trip · CER · TTFA |
|
||||
| Voice cloning | SECS (ECAPA cosine) | MOS · CER |
|
||||
| Speaker verification | EER | minDCF · FAR / FRR at operating point |
|
||||
| Diarization | DER | JER · speaker confusion |
|
||||
| Audio classification | top-1 · mAP | macro F1 · per-class recall |
|
||||
| Music generation | FAD | CLAP · listening panel MOS |
|
||||
| Audio language model | MMAU-Pro | LongAudioBench · AudioCaps FENSE |
|
||||
| Streaming S2S | latency P50/P95 | WER · MOS |
|
||||
|
||||
## The Concept
|
||||
|
||||

|
||||
|
||||
### ASR metrics
|
||||
|
||||
**WER (Word Error Rate).** `(S + D + I) / N`. Lowercase, strip punctuation, normalize numbers before scoring. Use `jiwer` or OpenAI's `whisper_normalizer`. < 5% = human-parity read speech.
|
||||
|
||||
**CER (Character Error Rate).** Same formula, character-level. Used for tone languages (Mandarin, Cantonese) where word segmentation is ambiguous.
|
||||
|
||||
**RTFx (inverse real-time factor).** Audio seconds processed per wall-clock second. Higher is better. Parakeet-TDT hits 3380×. Whisper-large-v3 is ~30×.
|
||||
|
||||
**First-token latency.** Wall-clock from audio input to first transcript token. Critical for streaming. Deepgram Nova-3: ~150 ms.
|
||||
|
||||
### TTS metrics
|
||||
|
||||
**MOS (Mean Opinion Score).** 1-5 human rating. Gold standard but slow. Collect 20+ listeners per sample, 100+ samples per model.
|
||||
|
||||
**UTMOS (2022-2026).** Learned MOS predictor. Correlates ~0.9 with human MOS on standard benchmarks. F5-TTS: UTMOS 3.95; ground truth: 4.08.
|
||||
|
||||
**SECS (Speaker Encoder Cosine Similarity).** For voice cloning. ECAPA embedding cosine between reference and cloned output. > 0.75 = recognizable clone.
|
||||
|
||||
**WER-on-ASR-round-trip.** Run Whisper over TTS output, compute WER against the input text. Catches intelligibility regressions. 2026 SOTA: < 2% CER.
|
||||
|
||||
**TTFA (time-to-first-audio).** Wall-clock latency. Kokoro-82M: ~100 ms; F5-TTS: ~1 s.
|
||||
|
||||
### Voice-cloning-specific
|
||||
|
||||
**SECS + MOS + CER** as a triple. Cloning that scores high SECS but low MOS means timbre-right-but-unnatural; the opposite means natural voice but wrong speaker.
|
||||
|
||||
### Speaker verification
|
||||
|
||||
**EER (Equal Error Rate).** The threshold where False Accept Rate equals False Reject Rate. ECAPA on VoxCeleb1-O: 0.87%.
|
||||
|
||||
**minDCF (min Detection Cost).** Weighted cost at a chosen operating point (often FAR=0.01). More production-relevant than EER.
|
||||
|
||||
### Diarization
|
||||
|
||||
**DER (Diarization Error Rate).** `(FA + Miss + Confusion) / total_speaker_time`. Missed speech + false-alarm speech + speaker-confusion, each as a fraction. AMI meetings: DER ~10-20% is realistic. pyannote 3.1 + Precision-2 commercial: <10% DER on well-recorded audio.
|
||||
|
||||
**JER (Jaccard Error Rate).** Alternative to DER, robust to short-segment bias.
|
||||
|
||||
### Audio classification
|
||||
|
||||
Multi-label: **mAP (mean Average Precision)** over all classes. AudioSet: 0.548 mAP for BEATs-iter3.
|
||||
|
||||
Multi-class exclusive: **top-1, top-5 accuracy**. Speech Commands v2: 99.0% top-1 (Audio-MAE).
|
||||
|
||||
Imbalanced: **macro F1** + **per-class recall**. Report per-class — aggregate accuracy hides which classes fail.
|
||||
|
||||
### Music generation
|
||||
|
||||
**FAD (Fréchet Audio Distance).** Distance between VGGish-embedding distributions of real vs generated audio. MusicGen-small on MusicCaps: 4.5. MusicLM: 4.0. Lower better.
|
||||
|
||||
**CLAP Score.** Text-audio alignment score using CLAP embeddings. > 0.3 = reasonable alignment.
|
||||
|
||||
**Listening panel MOS.** Still the final word for consumer-grade music. Suno v5 ELO 1293 on TTS Arena (from paired human preferences).
|
||||
|
||||
### Audio-language benchmarks
|
||||
|
||||
**MMAU (Massive Multi-Audio Understanding).** 10k audio-QA pairs.
|
||||
|
||||
**MMAU-Pro.** 1800 hard items, four categories: speech / sound / music / multi-audio. Random chance 25% on 4-way. Gemini 2.5 Pro overall ~60%; multi-audio ~22% across all models.
|
||||
|
||||
**LongAudioBench.** Multi-minute clips with semantic queries. Audio Flamingo Next beats Gemini 2.5 Pro.
|
||||
|
||||
**AudioCaps / Clotho.** Captioning benchmarks. SPICE, CIDEr, FENSE metrics.
|
||||
|
||||
### Streaming speech-to-speech
|
||||
|
||||
**Latency P50 / P95 / P99.** Wall-clock from end-of-user-speech to first audible response. Moshi: 200 ms; GPT-4o Realtime: 300 ms.
|
||||
|
||||
**WER / MOS** on the output.
|
||||
|
||||
**Barge-in responsiveness.** Time from user interrupt to assistant mute. Target < 150 ms.
|
||||
|
||||
### The 2026 leaderboards
|
||||
|
||||
| Leaderboard | Tracks | URL |
|
||||
|------------|--------|-----|
|
||||
| Open ASR Leaderboard (HF) | English + multilingual + long-form | `huggingface.co/spaces/hf-audio/open_asr_leaderboard` |
|
||||
| TTS Arena (HF) | English TTS | `huggingface.co/spaces/TTS-AGI/TTS-Arena` |
|
||||
| Artificial Analysis Speech | TTS + STT, ELO from paired votes | `artificialanalysis.ai/speech` |
|
||||
| MMAU-Pro | LALM reasoning | `mmaubenchmark.github.io` |
|
||||
| SpeakerBench / VoxSRC | Speaker recognition | `voxsrc.github.io` |
|
||||
| MMAU music subset | Music LALM | (within MMAU) |
|
||||
| HEAR benchmark | Self-supervised audio | `hearbenchmark.com` |
|
||||
|
||||
## Build It
|
||||
|
||||
### Step 1: WER with normalization
|
||||
|
||||
```python
|
||||
from jiwer import wer, Compose, ToLowerCase, RemovePunctuation, Strip
|
||||
|
||||
transform = Compose([ToLowerCase(), RemovePunctuation(), Strip()])
|
||||
score = wer(
|
||||
truth="Please turn on the lights.",
|
||||
hypothesis="please turn on the light",
|
||||
truth_transform=transform,
|
||||
hypothesis_transform=transform,
|
||||
)
|
||||
# ~0.17
|
||||
```
|
||||
|
||||
### Step 2: TTS round-trip WER
|
||||
|
||||
```python
|
||||
def ttr_wer(tts_model, asr_model, texts):
|
||||
errors = []
|
||||
for txt in texts:
|
||||
audio = tts_model.synthesize(txt)
|
||||
recog = asr_model.transcribe(audio)
|
||||
errors.append(wer(truth=txt, hypothesis=recog))
|
||||
return sum(errors) / len(errors)
|
||||
```
|
||||
|
||||
### Step 3: SECS for voice cloning
|
||||
|
||||
```python
|
||||
from speechbrain.inference.speaker import EncoderClassifier
|
||||
sv = EncoderClassifier.from_hparams("speechbrain/spkrec-ecapa-voxceleb")
|
||||
|
||||
emb_ref = sv.encode_batch(load_wav("reference.wav"))
|
||||
emb_clone = sv.encode_batch(load_wav("cloned.wav"))
|
||||
secs = torch.nn.functional.cosine_similarity(emb_ref, emb_clone, dim=-1).item()
|
||||
```
|
||||
|
||||
### Step 4: FAD for music generation
|
||||
|
||||
```python
|
||||
from frechet_audio_distance import FrechetAudioDistance
|
||||
fad = FrechetAudioDistance()
|
||||
score = fad.get_fad_score("generated_folder/", "reference_folder/")
|
||||
```
|
||||
|
||||
### Step 5: EER for speaker verification (same code as Lesson 6)
|
||||
|
||||
```python
|
||||
def eer(same_scores, diff_scores):
|
||||
thresholds = sorted(set(same_scores + diff_scores))
|
||||
best = (1.0, 0.0)
|
||||
for t in thresholds:
|
||||
far = sum(1 for s in diff_scores if s >= t) / len(diff_scores)
|
||||
frr = sum(1 for s in same_scores if s < t) / len(same_scores)
|
||||
if abs(far - frr) < best[0]:
|
||||
best = (abs(far - frr), (far + frr) / 2)
|
||||
return best[1]
|
||||
```
|
||||
|
||||
## Use It
|
||||
|
||||
Pair every deploy with a fixed eval harness that runs on every model update. Three cardinal rules:
|
||||
|
||||
1. **Normalize before scoring.** Lowercase, punctuation-strip, number-expand. Report the normalization rule.
|
||||
2. **Report distributions, not averages.** P50/P95/P99 for latency. Per-class recall for classification. Per-category for MMAU.
|
||||
3. **Run one canonical public benchmark.** Even if your production data differs, reporting on Open ASR / TTS Arena / MMAU lets reviewers compare apples-to-apples.
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- **UTMOS extrapolation.** Trained on VCTK-style clean speech; scores noisy / cloned / emotional audio poorly.
|
||||
- **MOS panel bias.** 20 Amazon Mechanical Turk workers ≠ 20 target users. Pay for a domain panel if stakes are high.
|
||||
- **FAD depends on reference set.** Compare against the same reference distribution across models.
|
||||
- **Aggregate WER.** A 5% WER overall can hide 30% WER on accented speech. Report by demographic slice.
|
||||
- **Public benchmark saturation.** Most frontier models are near the ceiling on standard benchmarks. Build an in-house held-out set that reflects your traffic.
|
||||
|
||||
## Ship It
|
||||
|
||||
Save as `outputs/skill-audio-evaluator.md`. Pick metrics, benchmarks, and reporting format for any audio model release.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. **Easy.** Run `code/main.py`. Compute WER / CER / EER / SECS / FAD-ish / MMAU-ish on toy inputs.
|
||||
2. **Medium.** Build a TTS round-trip WER harness. Run your Kokoro or F5-TTS output through Whisper. Compute WER over 50 prompts. Flag prompts with WER > 10%.
|
||||
3. **Hard.** Score your Lesson 10 LALM choice on MMAU-Pro speech + multi-audio subsets (50 items each). Report per-category accuracy and compare with the published number.
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|-----------------------|
|
||||
| WER | ASR score | `(S+D+I)/N` at word level after normalization. |
|
||||
| CER | Character WER | For tone languages or char-level systems. |
|
||||
| MOS | Human opinion | 1-5 rating; 20+ listeners × 100 samples. |
|
||||
| UTMOS | ML MOS predictor | Learned model; correlates ~0.9 with human MOS. |
|
||||
| SECS | Voice-clone similarity | ECAPA cosine between reference and clone. |
|
||||
| EER | Speaker verif score | Threshold where FAR = FRR. |
|
||||
| DER | Diarization score | (FA + Miss + Confusion) / total. |
|
||||
| FAD | Music-gen quality | Fréchet distance on VGGish embeddings. |
|
||||
| RTFx | Throughput | Audio seconds per wall-clock second. |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [jiwer](https://github.com/jitsi/jiwer) — WER/CER library with normalization utilities.
|
||||
- [UTMOS (Saeki et al. 2022)](https://arxiv.org/abs/2204.02152) — learned MOS predictor.
|
||||
- [Fréchet Audio Distance (Kilgour et al. 2019)](https://arxiv.org/abs/1812.08466) — the music-gen standard.
|
||||
- [Open ASR Leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard) — 2026 live rankings.
|
||||
- [TTS Arena](https://huggingface.co/spaces/TTS-AGI/TTS-Arena) — human-vote TTS leaderboard.
|
||||
- [MMAU-Pro benchmark](https://mmaubenchmark.github.io/) — LALM reasoning leaderboard.
|
||||
- [HEAR benchmark](https://hearbenchmark.com/) — audio SSL benchmarks.
|
||||
+29
@@ -0,0 +1,29 @@
|
||||
---
|
||||
name: audio-evaluator
|
||||
description: Pick metrics, benchmarks, normalization rules, and reporting format for any audio model release.
|
||||
version: 1.0.0
|
||||
phase: 6
|
||||
lesson: 17
|
||||
tags: [evaluation, wer, mos, utmos, eer, der, fad, mmau, leaderboard]
|
||||
---
|
||||
|
||||
Given the task (ASR / TTS / cloning / speaker-verif / diarization / classification / music / LALM / streaming S2S), output:
|
||||
|
||||
1. Primary metric. WER · MOS · UTMOS · SECS · EER · DER · mAP · FAD · MMAU-Pro accuracy · latency P95. One choice.
|
||||
2. Secondary metrics. 1-3 additional axes (speed, diversity, robustness) and reason.
|
||||
3. Normalization rule. Lowercase, punctuation-strip, number expansion, whitespace collapse. Use Whisper-normalizer or custom, document it.
|
||||
4. Public benchmark. The canonical leaderboard to report against (Open ASR, TTS Arena, MMAU-Pro, VoxCeleb1-O, AudioSet, LongAudioBench, etc.).
|
||||
5. In-house set. Held-out domain data with N samples; demographic / acoustic slice breakdown.
|
||||
6. Reporting format. Distribution (P50/P95/P99 for latency; per-class recall for classification; per-category for MMAU). Release notes template.
|
||||
|
||||
Refuse single-number evaluation for latency (report percentiles). Refuse aggregate-only for classification (report per-class). Refuse TTS releases without both MOS/UTMOS and SECS (when cloning). Refuse ASR releases without a WER normalization spec. Refuse music releases with only FAD — always pair with human MOS panel.
|
||||
|
||||
Example input: "Release of a new English-Spanish conversational TTS. Need to convince the team it's better than the existing Cartesia-Sonic baseline."
|
||||
|
||||
Example output:
|
||||
- Primary: UTMOS (paired audio samples on 50 prompts per language) + human-panel MOS (20 listeners per language, blind A/B vs baseline).
|
||||
- Secondary: TTFA median & P95 (must match baseline); SECS > 0.80 vs a fixed voice reference (no speaker regression); CER on round-trip ASR (Whisper-large-v3-turbo) < 2%.
|
||||
- Normalization: Whisper-normalizer English + Hugging Face multilingual-normalizer Spanish for round-trip WER.
|
||||
- Public benchmark: TTS Arena (English) and Artificial Analysis Speech for relative ELO positioning. Target: within 50 ELO of the closest competitor.
|
||||
- In-house: 200 held-out prompts (100 per lang) covering money, dates, product names, 2-sentence narration, emotional read, code-switched. 10 demographic voices.
|
||||
- Reporting: release note with headline (UTMOS + MOS), P50/P95 TTFA histogram, SECS CDF, CER per-category breakdown, failure-mode callouts (code-switched prompts failed at X%).
|
||||
Reference in New Issue
Block a user