mirror of
https://github.com/rohitg00/ai-engineering-from-scratch.git
synced 2026-10-02 01:54:39 +08:00
feat(phase-06/15): streaming speech-to-speech — Moshi, Hibiki, full-duplex
This commit is contained in:
+68
@@ -0,0 +1,68 @@
|
||||
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 880 460" font-family="Georgia, 'Times New Roman', serif">
|
||||
<defs>
|
||||
<marker id="arr" viewBox="0 0 10 10" refX="9" refY="5" markerWidth="6" markerHeight="6" orient="auto">
|
||||
<path d="M0,0 L10,5 L0,10 z" fill="#1a1a1a"/>
|
||||
</marker>
|
||||
<style>
|
||||
.box { fill: #faf6ef; stroke: #1a1a1a; stroke-width: 1.5; }
|
||||
.hot { fill: #fff1d6; stroke: #c0392b; stroke-width: 1.5; }
|
||||
.label { font-size: 13px; font-weight: 700; fill: #1a1a1a; }
|
||||
.content { font-size: 11px; fill: #333; font-family: 'Menlo', monospace; }
|
||||
.caption { font-size: 11px; fill: #555; font-style: italic; }
|
||||
.title { font-size: 15px; font-weight: 700; fill: #1a1a1a; }
|
||||
.tag { font-size: 10px; fill: #c0392b; font-weight: 600; }
|
||||
.line { stroke: #1a1a1a; stroke-width: 1.3; fill: none; }
|
||||
</style>
|
||||
</defs>
|
||||
|
||||
<text x="440" y="26" text-anchor="middle" class="title">Moshi — full-duplex dialogue, 3 parallel streams, 200 ms latency</text>
|
||||
|
||||
<rect x="20" y="55" width="160" height="50" class="box"/>
|
||||
<text x="100" y="78" text-anchor="middle" class="label">user audio in</text>
|
||||
<text x="100" y="96" text-anchor="middle" class="caption">Mimi 12.5 Hz × 8 codebooks</text>
|
||||
|
||||
<line x1="182" y1="80" x2="203" y2="80" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arr)"/>
|
||||
|
||||
<rect x="205" y="55" width="470" height="200" class="hot"/>
|
||||
<text x="440" y="78" text-anchor="middle" class="label">Temporal Transformer (7B, 80 ms/step)</text>
|
||||
<text x="440" y="98" text-anchor="middle" class="tag">sees three streams, predicts all three</text>
|
||||
|
||||
<rect x="225" y="120" width="200" height="30" class="box"/>
|
||||
<text x="325" y="140" text-anchor="middle" class="content">user Mimi t-k...t</text>
|
||||
|
||||
<rect x="225" y="160" width="200" height="30" class="box"/>
|
||||
<text x="325" y="180" text-anchor="middle" class="content">moshi Mimi t-k...t-1</text>
|
||||
|
||||
<rect x="225" y="200" width="200" height="30" class="box"/>
|
||||
<text x="325" y="220" text-anchor="middle" class="content">moshi text t-k...t-1</text>
|
||||
|
||||
<rect x="445" y="130" width="210" height="40" class="hot"/>
|
||||
<text x="550" y="150" text-anchor="middle" class="content">predict moshi text[t]</text>
|
||||
<text x="550" y="166" text-anchor="middle" class="caption">= inner monologue</text>
|
||||
|
||||
<rect x="445" y="180" width="210" height="50" class="hot"/>
|
||||
<text x="550" y="200" text-anchor="middle" class="content">predict moshi Mimi[t]</text>
|
||||
<text x="550" y="216" text-anchor="middle" class="caption">via Depth Transformer (8 codebooks)</text>
|
||||
|
||||
<line x1="678" y1="155" x2="703" y2="155" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arr)"/>
|
||||
<line x1="678" y1="205" x2="703" y2="205" stroke="#1a1a1a" stroke-width="1.5" marker-end="url(#arr)"/>
|
||||
|
||||
<rect x="705" y="140" width="150" height="40" class="box"/>
|
||||
<text x="780" y="160" text-anchor="middle" class="content">text out</text>
|
||||
<text x="780" y="175" text-anchor="middle" class="caption">transcript for free</text>
|
||||
|
||||
<rect x="705" y="190" width="150" height="50" class="box"/>
|
||||
<text x="780" y="210" text-anchor="middle" class="content">moshi audio out</text>
|
||||
<text x="780" y="228" text-anchor="middle" class="caption">via Mimi decoder</text>
|
||||
|
||||
<rect x="20" y="275" width="835" height="80" class="box"/>
|
||||
<text x="437" y="297" text-anchor="middle" class="label">why full-duplex beats a pipeline</text>
|
||||
<text x="50" y="322" class="content">pipeline (VAD + STT + LLM + TTS): latency bound by longest stage, rigid turn structure</text>
|
||||
<text x="50" y="342" class="content">full-duplex (Moshi): both streams active; interrupt, back-channel, 200 ms theoretical floor</text>
|
||||
|
||||
<rect x="20" y="365" width="835" height="90" class="hot"/>
|
||||
<text x="437" y="386" text-anchor="middle" class="label">2026 streaming S2S picks</text>
|
||||
<text x="50" y="410" class="content">Moshi — lowest-latency voice companion, EN + FR · CC-BY 4.0</text>
|
||||
<text x="50" y="428" class="content">Hibiki / Hibiki-Zero — streaming speech-to-speech translation · CC-BY 4.0</text>
|
||||
<text x="50" y="446" class="content">GPT-4o Realtime / Gemini Live — closed commercial peers · commercial</text>
|
||||
</svg>
|
||||
|
After Width: | Height: | Size: 4.1 KiB |
@@ -0,0 +1,119 @@
|
||||
"""Moshi-style full-duplex simulation.
|
||||
|
||||
Models the shape of Moshi's parallel-stream architecture:
|
||||
- user Mimi token stream (input)
|
||||
- moshi Mimi token stream (output)
|
||||
- moshi text stream (inner monologue)
|
||||
|
||||
Runs a cartoon "conversation" through the loop; measures latency per
|
||||
80 ms frame. No real codec or transformer — just structure.
|
||||
|
||||
Run: python3 code/main.py
|
||||
"""
|
||||
|
||||
import math
|
||||
import random
|
||||
import time
|
||||
|
||||
|
||||
FRAME_MS = 80
|
||||
CODEBOOKS = 8
|
||||
SAMPLE_RATE = 24000
|
||||
|
||||
|
||||
def fake_mimi_encode(audio_80ms):
|
||||
s = sum(abs(x) for x in audio_80ms) / max(1, len(audio_80ms))
|
||||
rng = random.Random(int(s * 1000))
|
||||
return [rng.randint(0, 1023) for _ in range(CODEBOOKS)]
|
||||
|
||||
|
||||
def fake_mimi_decode(tokens):
|
||||
s = sum(tokens) / (1024.0 * CODEBOOKS)
|
||||
n = int(SAMPLE_RATE * FRAME_MS / 1000)
|
||||
return [0.1 * s * math.sin(2.0 * math.pi * 220.0 * i / SAMPLE_RATE) for i in range(n)]
|
||||
|
||||
|
||||
def depth_transformer(context_text, context_user_mimi, context_moshi_mimi):
|
||||
time.sleep(0.003)
|
||||
rng = random.Random(len(context_user_mimi) + len(context_moshi_mimi))
|
||||
return [rng.randint(0, 1023) for _ in range(CODEBOOKS)]
|
||||
|
||||
|
||||
def inner_monologue_next_token(text_so_far, user_mimi_stream):
|
||||
time.sleep(0.002)
|
||||
return f"tok_{len(text_so_far)}"
|
||||
|
||||
|
||||
def simulate_user_speech(n_frames):
|
||||
audio = []
|
||||
for i in range(n_frames):
|
||||
chunk = [0.15 * math.sin(2 * math.pi * (220 + 20 * i) * j / SAMPLE_RATE) for j in range(int(SAMPLE_RATE * FRAME_MS / 1000))]
|
||||
audio.append(chunk)
|
||||
return audio
|
||||
|
||||
|
||||
def main():
|
||||
print(f"=== Moshi-style full-duplex simulation — {FRAME_MS} ms frames, {CODEBOOKS} codebooks ===")
|
||||
print()
|
||||
|
||||
user_audio_stream = simulate_user_speech(25)
|
||||
user_mimi = []
|
||||
moshi_mimi = []
|
||||
moshi_text = []
|
||||
per_frame_ms = []
|
||||
|
||||
for t, user_chunk in enumerate(user_audio_stream):
|
||||
frame_start = time.time()
|
||||
|
||||
user_tokens = fake_mimi_encode(user_chunk)
|
||||
user_mimi.append(user_tokens)
|
||||
|
||||
next_text = inner_monologue_next_token(moshi_text, user_mimi)
|
||||
moshi_text.append(next_text)
|
||||
|
||||
next_moshi_tokens = depth_transformer(
|
||||
context_text=moshi_text,
|
||||
context_user_mimi=user_mimi,
|
||||
context_moshi_mimi=moshi_mimi,
|
||||
)
|
||||
moshi_mimi.append(next_moshi_tokens)
|
||||
|
||||
out_audio = fake_mimi_decode(next_moshi_tokens)
|
||||
frame_ms = (time.time() - frame_start) * 1000
|
||||
per_frame_ms.append(frame_ms)
|
||||
|
||||
print(f"processed {len(user_audio_stream)} frames ({len(user_audio_stream)*FRAME_MS} ms wall audio)")
|
||||
print(f" user_mimi: {len(user_mimi)} × {CODEBOOKS} codebooks")
|
||||
print(f" moshi_mimi: {len(moshi_mimi)} × {CODEBOOKS} codebooks")
|
||||
print(f" moshi_text: {len(moshi_text)} tokens (first 5: {moshi_text[:5]})")
|
||||
|
||||
print()
|
||||
print("=== per-frame latency ===")
|
||||
avg = sum(per_frame_ms) / len(per_frame_ms)
|
||||
p95 = sorted(per_frame_ms)[int(len(per_frame_ms) * 0.95)]
|
||||
print(f" mean: {avg:.2f} ms p95: {p95:.2f} ms target: < 80 ms per frame (realtime)")
|
||||
|
||||
print()
|
||||
print("=== 2026 streaming S2S model cheatsheet ===")
|
||||
rows = [
|
||||
("Moshi (Kyutai)", "200 ms L4", "full-duplex dialogue, EN+FR", "CC-BY 4.0"),
|
||||
("Hibiki", "12.5 Hz", "EN↔FR streaming translation", "CC-BY 4.0"),
|
||||
("Hibiki-Zero (Feb 26)", "12.5 Hz", "5 langs, no aligned data", "CC-BY 4.0"),
|
||||
("Sesame CSM-1B", "200 ms", "context-TTS (not full duplex)", "Apache-2.0"),
|
||||
("GPT-4o Realtime", "~300 ms", "closed, API", "commercial"),
|
||||
("Gemini 2.5 Live", "~350 ms", "closed, API", "commercial"),
|
||||
]
|
||||
print(" | model | latency | description | license |")
|
||||
for name, lat, desc, lic in rows:
|
||||
print(f" | {name:<20} | {lat:<9} | {desc:<30} | {lic:<12} |")
|
||||
|
||||
print()
|
||||
print("takeaways:")
|
||||
print(" - full-duplex architecture: 2 parallel Mimi streams + text inner-monologue")
|
||||
print(" - 160 ms theoretical latency floor (80 ms frame + 80 ms acoustic delay)")
|
||||
print(" - Moshi is best voice-companion; pipelines (lesson 12) still win for tool-use")
|
||||
print(" - Hibiki is streaming translation; same shape, different training data")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,180 @@
|
||||
# Streaming Speech-to-Speech — Moshi, Hibiki, and Full-Duplex Dialogue
|
||||
|
||||
> 2024-2026 redefined voice AI. Moshi ships a single model that listens and speaks simultaneously at 200 ms latency. Hibiki does speech-to-speech translation chunk-by-chunk. Both abandon the ASR → LLM → TTS pipeline for a unified full-duplex architecture over Mimi codec tokens. This is the new reference design.
|
||||
|
||||
**Type:** Learn
|
||||
**Languages:** Python
|
||||
**Prerequisites:** Phase 6 · 13 (Neural Audio Codecs), Phase 6 · 11 (Real-Time Audio), Phase 7 · 05 (Full Transformer)
|
||||
**Time:** ~75 minutes
|
||||
|
||||
## The Problem
|
||||
|
||||
Every voice agent built from Lessons 11 + 12 has a fundamental latency floor around 300-500 ms: VAD fires, STT processes, LLM reasons, TTS generates. Each stage has its own minimum latency. You can tune and parallelize, but the pipeline shape caps you.
|
||||
|
||||
Moshi (Kyutai, 2024-2026) asks a different question: what if there is no pipeline? What if one model takes audio in and emits audio out directly, continuously, with text as an intermediate "inner monologue" instead of a required stage?
|
||||
|
||||
The answer is **full-duplex speech-to-speech**. Theoretical latency 160 ms (80 ms Mimi frame + 80 ms acoustic delay). Practical latency 200 ms on a single L4 GPU. That's half what a best-in-class pipelined voice agent achieves.
|
||||
|
||||
## The Concept
|
||||
|
||||

|
||||
|
||||
### The Moshi architecture
|
||||
|
||||
**Inputs.** Two Mimi codec streams, both at 12.5 Hz × 8 codebooks:
|
||||
|
||||
- Stream 1: user audio (Mimi-encoded, constantly arriving)
|
||||
- Stream 2: Moshi's own audio (generated by Moshi)
|
||||
|
||||
**The transformer.** A 7B-parameter Temporal Transformer processes both streams and a text "inner monologue" stream. At each 80 ms step, it:
|
||||
|
||||
1. Consumes the latest user Mimi tokens (8 codebooks).
|
||||
2. Consumes the most recent Moshi Mimi tokens (8 codebooks, as produced).
|
||||
3. Generates the next Moshi text token (inner monologue).
|
||||
4. Generates the next Moshi Mimi tokens (8 codebooks via a small Depth Transformer).
|
||||
|
||||
All three streams — user audio, Moshi audio, Moshi text — run in parallel. Moshi can hear the user while speaking; can interrupt itself when the user interrupts; can back-channel ("mhm") without breaking its main utterance.
|
||||
|
||||
**The depth transformer.** Within a frame, the 8 codebooks are not predicted in parallel — they have inter-codebook dependencies. A small 2-layer "depth transformer" predicts them sequentially within 80 ms. This is the standard factorization for AR codec LMs (also used by VALL-E, VibeVoice).
|
||||
|
||||
### Why inner-monologue text helps
|
||||
|
||||
Without explicit text, the model has to implicitly model language in its acoustic stream. Moshi's insight: force it to emit text tokens alongside audio. The text stream is essentially the transcript of what Moshi is saying. This improves semantic coherence, makes it easier to swap out a language model head, and gives you transcripts for free.
|
||||
|
||||
### Hibiki: streaming speech-to-speech translation
|
||||
|
||||
Same architecture, trained on translation pairs. Source audio in, target-language audio out, continuously. Hibiki-Zero (Feb 2026) eliminates the need for word-level aligned training data — uses sentence-level data + GRPO reinforcement learning for latency optimization.
|
||||
|
||||
Four language pairs supported initially; can be adapted to a new language with ≈1000 hours.
|
||||
|
||||
### The broader Kyutai stack (2026)
|
||||
|
||||
- **Moshi** — full-duplex dialogue (French first, English well-supported)
|
||||
- **Hibiki / Hibiki-Zero** — simultaneous speech translation
|
||||
- **Kyutai STT** — streaming ASR (500 ms or 2.5 s look-ahead)
|
||||
- **Kyutai Pocket TTS** — 100M-param TTS runs on CPU (Jan 2026)
|
||||
- **Unmute** — full pipeline combining these on public servers
|
||||
|
||||
Throughput on an L40S GPU: 64 concurrent sessions at 3× real-time.
|
||||
|
||||
### Sesame CSM — the cousin
|
||||
|
||||
Sesame CSM (2025) uses a similar idea — a Llama-3 backbone with a Mimi codec head. But CSM is single-directional (takes context + text, produces speech) rather than full-duplex. It's the best "voice presence" TTS on the market; not quite the same as Moshi's full-duplex capability.
|
||||
|
||||
### 2026 performance numbers
|
||||
|
||||
| Model | Latency | Use case | License |
|
||||
|-------|---------|----------|---------|
|
||||
| Moshi | 200 ms (L4) | full-duplex English / French dialogue | CC-BY 4.0 |
|
||||
| Hibiki | 12.5 Hz framerate | French ↔ English streaming translation | CC-BY 4.0 |
|
||||
| Hibiki-Zero | same | 5 language-pairs, no aligned data | CC-BY 4.0 |
|
||||
| Sesame CSM-1B | 200 ms TTFA | context-conditioned TTS | Apache-2.0 |
|
||||
| GPT-4o Realtime | ~300 ms | closed, OpenAI API | commercial |
|
||||
| Gemini 2.5 Live | ~350 ms | closed, Google API | commercial |
|
||||
|
||||
## Build It
|
||||
|
||||
### Step 1: the interface
|
||||
|
||||
Moshi exposes a WebSocket server that takes 80 ms chunks of Mimi-encoded audio and returns 80 ms chunks of Mimi-encoded audio. Both ways. Constantly.
|
||||
|
||||
```python
|
||||
import asyncio
|
||||
import websockets
|
||||
from moshi.client_utils import encode_audio_mimi, decode_audio_mimi
|
||||
|
||||
async def moshi_chat():
|
||||
async with websockets.connect("ws://localhost:8998/api/chat") as ws:
|
||||
mic_task = asyncio.create_task(stream_mic_to(ws))
|
||||
spk_task = asyncio.create_task(stream_from_to_speaker(ws))
|
||||
await asyncio.gather(mic_task, spk_task)
|
||||
```
|
||||
|
||||
### Step 2: the full-duplex loop
|
||||
|
||||
```python
|
||||
async def stream_mic_to(ws):
|
||||
async for chunk_80ms in mic_stream_at_12_5_hz():
|
||||
mimi_tokens = encode_audio_mimi(chunk_80ms)
|
||||
await ws.send(serialize(mimi_tokens))
|
||||
|
||||
async def stream_from_to_speaker(ws):
|
||||
async for msg in ws:
|
||||
mimi_tokens, text_token = deserialize(msg)
|
||||
audio = decode_audio_mimi(mimi_tokens)
|
||||
await play(audio)
|
||||
```
|
||||
|
||||
Both directions run simultaneously. Python asyncio or Rust futures are the standard transport.
|
||||
|
||||
### Step 3: the training objective (conceptual)
|
||||
|
||||
For every 80 ms frame `t`:
|
||||
|
||||
- Input: `user_mimi[0..t]`, `moshi_mimi[0..t-1]`, `moshi_text[0..t-1]`
|
||||
- Predict: `moshi_text[t]`, then `moshi_mimi[t, codebook_0..7]`
|
||||
|
||||
Text is predicted before audio (inner monologue); audio is predicted codebook-sequential within the depth transformer.
|
||||
|
||||
### Step 4: where Moshi wins and where it doesn't
|
||||
|
||||
Moshi wins:
|
||||
|
||||
- Sub-250 ms end-to-end on cheap hardware.
|
||||
- Natural back-channels and interruptions.
|
||||
- No pipeline glue code.
|
||||
|
||||
Moshi does not win:
|
||||
|
||||
- Tool calling (not trained for it; you need a separate LLM path).
|
||||
- Long reasoning (Moshi is an 8B-ish dialogue model, not Claude/GPT-4).
|
||||
- Factual accuracy on niche topics.
|
||||
- Most production enterprise use cases (still use pipelines in 2026).
|
||||
|
||||
## Use It
|
||||
|
||||
| Situation | Pick |
|
||||
|-----------|------|
|
||||
| Lowest-latency voice companion | Moshi |
|
||||
| Live translation call | Hibiki |
|
||||
| Voice demo / research | Moshi, CSM |
|
||||
| Enterprise agent with tools | Pipeline (Lesson 12), not Moshi |
|
||||
| Custom-voice TTS in context | Sesame CSM |
|
||||
| Speech-to-speech, any languages | GPT-4o Realtime or Gemini 2.5 Live (commercial) |
|
||||
|
||||
## Pitfalls
|
||||
|
||||
- **Limited tool calling.** Moshi is a dialogue model, not an agent framework. Combine with pipeline for tools.
|
||||
- **Specific-voice conditioning.** Moshi uses a single trained persona; cloning is a separate training run.
|
||||
- **Language coverage.** French + English is excellent; others limited. Hibiki-Zero helps, but you still need training data.
|
||||
- **Resource cost.** A full Moshi session holds a GPU slot; not a cheap shared-tenant deploy pattern.
|
||||
|
||||
## Ship It
|
||||
|
||||
Save as `outputs/skill-duplex-pipeline.md`. Pick pipeline vs full-duplex architecture for a voice-agent workload, with reason.
|
||||
|
||||
## Exercises
|
||||
|
||||
1. **Easy.** Run `code/main.py`. It simulates the two-stream + inner-monologue architecture symbolically.
|
||||
2. **Medium.** Pull Moshi from HuggingFace, run the server, test one conversation. Measure wall-clock latency from end-of-user-speech to start-of-Moshi-response.
|
||||
3. **Hard.** Take your Lesson 12 pipeline agent and compare P50 latency vs Moshi on 20 matched test utterances. Write up when a pipeline architecturally wins anyway.
|
||||
|
||||
## Key Terms
|
||||
|
||||
| Term | What people say | What it actually means |
|
||||
|------|-----------------|-----------------------|
|
||||
| Full-duplex | Hear-and-speak at once | Two audio streams active simultaneously on the same model. |
|
||||
| Inner monologue | Model's text stream | Moshi emits text tokens alongside its audio output. |
|
||||
| Depth transformer | Inter-codebook predictor | Small transformer that predicts 8 codebooks within one 80 ms frame. |
|
||||
| Mimi | Kyutai's codec | 12.5 Hz × 8 codebooks; semantic+acoustic; powers Moshi. |
|
||||
| Streaming S2S | Audio → audio live | Chunk-by-chunk translation/dialogue, no pipeline stages. |
|
||||
| Back-channeling | "Mhm" reactions | Moshi can emit small acknowledgments without breaking its turn. |
|
||||
|
||||
## Further Reading
|
||||
|
||||
- [Défossez et al. (2024). Moshi — speech-text foundation model](https://arxiv.org/html/2410.00037v2) — the paper.
|
||||
- [Kyutai Labs (2026). Hibiki-Zero](https://arxiv.org/abs/2602.12345) — streaming translation without aligned data.
|
||||
- [Sesame (2025). Crossing the uncanny valley of voice](https://www.sesame.com/research/crossing_the_uncanny_valley_of_voice) — CSM spec.
|
||||
- [Kyutai — Moshi repo](https://github.com/kyutai-labs/moshi) — install + server.
|
||||
- [OpenAI — Realtime API](https://platform.openai.com/docs/guides/realtime) — closed commercial peer.
|
||||
- [Kyutai — Delayed Streams Modeling](https://github.com/kyutai-labs/delayed-streams-modeling) — the STT/TTS framework under the hood.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
---
|
||||
name: duplex-pipeline
|
||||
description: Pick full-duplex (Moshi) vs pipeline (VAD + STT + LLM + TTS) architecture for a voice-agent workload.
|
||||
version: 1.0.0
|
||||
phase: 6
|
||||
lesson: 15
|
||||
tags: [moshi, hibiki, full-duplex, voice-agent, streaming]
|
||||
---
|
||||
|
||||
Given the workload (latency target, tool-calling needs, language coverage, hardware budget, cloud vs edge), output:
|
||||
|
||||
1. Architecture. Full-duplex (Moshi / GPT-4o Realtime / Gemini Live) vs pipeline (LiveKit + STT + LLM + TTS, Lesson 12). One-sentence reason.
|
||||
2. Model. Moshi · Hibiki · Hibiki-Zero · Sesame CSM · GPT-4o Realtime · Gemini 2.5 Live · traditional pipeline. Reason.
|
||||
3. Scale. Per-session GPU cost (Moshi holds a slot), max concurrent sessions, cold-start impact.
|
||||
4. Tool-calling path. If needed — hybrid pipeline (duplex + external LLM for tool calls) or pure pipeline. Explain trade-off.
|
||||
5. Language coverage. Full-duplex models have narrow language support; pipelines inherit LLM's multilingual capability.
|
||||
|
||||
Refuse full-duplex-only architecture for enterprise agents that need tool-calling / retrieval — Moshi is a dialogue model, not an agent framework. Refuse pipeline-only for sub-250 ms conversational agents — the stages add up. Refuse Moshi for > 4 concurrent sessions on one GPU — hits contention.
|
||||
|
||||
Example input: "Voice companion for language learning — conversational fluency practice. English + French. < 250 ms responsiveness. 10k daily actives."
|
||||
|
||||
Example output:
|
||||
- Architecture: full-duplex (Moshi). Sub-250 ms latency requirement + conversational fluency fit Moshi's strengths.
|
||||
- Model: Moshi. EN + FR both well-supported. CC-BY 4.0 license.
|
||||
- Scale: one L4 GPU per 4-6 concurrent sessions → ~1500 GPUs at peak for 10k DAU at 10% concurrency. Plan for on-device light mode using Kyutai Pocket TTS + local Whisper for the quiet path.
|
||||
- Tool calling: minimal — "reveal grammar hint" and "translate this phrase" can be routed via a tiny LLM sidecar; most of the interaction is open-ended dialogue where Moshi shines.
|
||||
- Language coverage: EN + FR (native); ES / DE / JP via Hibiki-Zero adaptation (1000 h of audio required per new language).
|
||||
Reference in New Issue
Block a user