* fix(mcp): serve /mcp without redirect, follow OMNIVOICE_PORT, real integration setups - /mcp and /mcp/ both reach the Streamable HTTP transport for every method, ahead of the SPA StaticFiles mount (POST /mcp was 405 on Docker/source builds, 307 without a built SPA). Regression test drives initialize -> tools/list -> DELETE on the real main.app with a dist dir and follow_redirects=False; the client-setup test no longer follows redirects. - MCP tool callbacks resolve the backend URL from OMNIVOICE_BIND_HOST + OMNIVOICE_PORT (OMNIVOICE_API_URL still overrides) instead of a hard-coded :3900; speech_client and the dev fallback redirect follow their port env too, with a class guard against literal :39xx URLs in backend code. - Integrations: setup registry keyed by slug drives the Works with VoiceStudio badge and capability chips; adds Codex CLI config.toml, generic MCP (HTTP + stdio shim), VoiceStudio API (curl + OpenAI SDK) and Docker/GHCR setups. Removes the fake category chips from the shared catalog config; fixes the duplicate Details headings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin capture path in speech envelope test; changelog refs (#2289) A prior lifespan-running test persists a sherpa dictation.model_id pref via the startup performance profile, routing the envelope test to the sherpa handler. Pin the capture path it asserts. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): export <base>/mcp/ for every MCP client The trailing-slash URL works on backends without the bare-/mcp fix too, so Claude Code, Cursor, Codex and the generic MCP card all export it. Docs lead with /mcp/ and note bare /mcp works on current backends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(mcp): in-process tool callbacks; remote-safe stdio and API snippets Review findings on #2289: - Greptile P1: mounted MCP tools now call the backend app in-process via httpx.ASGITransport as a loopback caller, so a concrete LAN OMNIVOICE_BIND_HOST behind an API key / share PIN no longer 401s every tool. Standalone runs keep the HTTP backend_self_url() path. - Greptile P1: the stdio shim accepts OMNIVOICE_URL (https + path prefix) and forwards OMNIVOICE_API_KEY as a Bearer token; the MCP card exports the full base URL instead of host/port. - Greptile P1: remote API snippets read the key from $OMNIVOICE_API_KEY (curl Bearer header, Python os.environ) without exporting credentials. - CodeQL (Bandit B104): wildcard detection uses ipaddress.is_unspecified instead of a 0.0.0.0 literal. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): API keys only over https or loopback; reference them by env CodeRabbit findings on #2289: - Standalone MCP server and the stdio shim send OMNIVOICE_API_KEY as a Bearer token only to https or loopback targets (same rule as backend.speech_client); the shim refuses to start otherwise. - Remote https exports reference the key from the user's environment in each client's own syntax (Claude Code ${VAR}, Cursor ${env:VAR}, Codex bearer_token_env_var, curl/Python $OMNIVOICE_API_KEY); plain-http remotes export no key. - test_backend_self_url imports app modules inside the tests. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): credential guidance via translated hints; HTTP card auth line CodeRabbit findings on #2289: - The generic Streamable HTTP card lists the env-backed Authorization header for remote https backends. - English comments inside copyable snippets (API credentials, Docker GPU) move to translated hints (apiBearerHint, apiInsecureHint, dockerGpuHint, all 21 locales); a test forbids natural-language comments in snippets. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): PowerShell docker run; scheme-aware Agents SDK key CodeRabbit findings on #2289: - Docker/GHCR pages add a Windows PowerShell docker run (backtick continuations, CSPRNG key that works on PowerShell 5.1 and 7), so the default setup works on every platform. - The OpenAI Agents snippet only reads OMNIVOICE_API_KEY for loopback or https backends; a remote plain-http backend gets a placeholder. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): Agents SDK key only for remote https backends Greptile/CodeRabbit findings on #2289: the OpenAI Agents snippet now follows the shared remoteAuth policy exactly. It requires os.environ["OMNIVOICE_API_KEY"] for remote https backends, uses a placeholder on loopback (no key over plain http, even locally), and for a remote plain-http backend shows the translated 'use https' hint instead of implying the key will be sent. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: drop the speech-platform capture pin now that #2294 isolates settings Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(integrations): remote https Python snippets tolerate an unkeyed backend Greptile finding on #2289: os.environ[...] raised KeyError on a remote https backend that runs without an API key. Read the key with a placeholder fallback (Agents SDK and API Python snippets); still loopback- and plain-http-safe via remoteAuth. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
8.6 KiB
Local speech platform
VoiceStudio is both a desktop dictation app and a headless local speech service. The desktop remains one app: its bundled Rust control sidecar owns microphone activation, focused-target capture, clipboard safety, and native insertion; the Python backend keeps ASR models warm and exposes the audio data plane.
This split lets an integration choose how much it owns:
Herdr / terminal / desktop app ── start, stop, toggle ──> Rust control :3902
│
├─ captures target
├─ opens VoiceStudio mic
└─ inserts final text
VS Code / custom GUI / remote mic ── PCM or WebM ───────> WS/HTTP :3900
│
└─ partial/final text
reserve target / insert final ─> Rust control :3902
Claude Code / Codex / Pi / agents ── MCP HTTP/stdio ───> MCP :3900
The Rust sidecar is part of the VoiceStudio process, not a second application.
It starts with the desktop app and binds only to 127.0.0.1.
Discover capabilities
Desktop/native discovery:
curl http://127.0.0.1:3902/.well-known/voicestudio-speech
Engine/data-plane discovery:
curl http://127.0.0.1:3900/.well-known/voicestudio-speech
Both return voicestudio.speech.v1. The desktop document includes absolute
control, batch, streaming, output-session, and MCP endpoints. The backend
document uses relative URLs so it also works behind Tailscale or a reverse
proxy; it advertises native control only when launched by the desktop app.
Use VoiceStudio capture from any app
These calls use VoiceStudio's existing microphone, model selection, pill, refinement, and session-bound insertion. The app under the cursor remains the destination.
curl -X POST http://127.0.0.1:3902/v1/dictation/start
curl -X POST http://127.0.0.1:3902/v1/dictation/stop
curl -X POST http://127.0.0.1:3902/v1/dictation/toggle
JSON-RPC clients use the same actions:
{"jsonrpc":"2.0","id":1,"method":"dictation.toggle"}
Send that object to POST http://127.0.0.1:3902/rpc. The installed
VoiceStudio executable also accepts --dictate-start, --dictate-stop, and
--dictate-toggle; the single-instance bridge forwards them to the running
app without opening the Studio window.
The dependency-free Python bridge is convenient for hooks and TUIs:
python -m backend.speech_client status
python -m backend.speech_client toggle
python -m backend.speech_client transcribe recording.wav
python -m backend.speech_client transcribe recording.wav --insert
--insert captures the focused destination before transcription starts and
uses the same clipboard-preserving native delivery as the global shortcut.
Bring your own capture interface
An editor extension or GUI can own the microphone and consume live text. Connect to:
ws://127.0.0.1:3900/v1/audio/transcriptions/stream
Send binary WebM/Opus frames by default. For raw signed 16-bit mono PCM, use
?pcm=1&sr=16000. Finish without closing the socket by sending:
{"type":"input_audio.end"}
Every response carries protocol and session_id:
{"type":"session.started","protocol":"voicestudio.speech.v1","session_id":"..."}
{"type":"partial","text":"hello wor...","session_id":"..."}
{"type":"final","final_kind":"summary","text":"Hello world.","session_id":"..."}
Streaming Sherpa models can also emit final_kind: "utterance" before the
authoritative whole-session summary. Existing /ws/transcribe clients keep
their unchanged legacy frames and EOF control.
To reuse native insertion with a custom capture client:
POST /v1/output/sessionson port 3902 before opening the microphone.- Stream audio and receive the final text on port 3900.
POST /v1/output/sessions/{id}/insertwith{"text":"..."}.- If capture is cancelled,
DELETE /v1/output/sessions/{id}.
Only one output session can own a focused destination at a time. Stale IDs are rejected instead of inserting into a newer target.
Batch and agent protocols
| Transport | Endpoint | Use |
|---|---|---|
| OpenAI-compatible HTTP | POST :3900/v1/audio/transcriptions |
Files, scripts, existing SDKs |
| WebSocket | :3900/v1/audio/transcriptions/stream |
Partial and final live text |
| MCP Streamable HTTP | POST :3900/mcp/ (bare /mcp on current backends) |
Modern agent clients |
| MCP stdio | python -m backend.mcp_shim |
Claude Code, Codex, and stdio-only clients |
| JSON-RPC | POST :3902/rpc |
Native dictation control |
| Native CLI | VoiceStudio --dictate-* flags |
Hooks and plugin actions |
Integration map
| Interface | Recommended connection |
|---|---|
| Any desktop text field | Existing global shortcut or Rust dictation.toggle |
| Herdr | Merge the example command bindings into Herdr's config; detached commands call the Rust API while the pane stays focused |
| Pi, Claude Code, Codex, Antigravity CLI | Dictate into the focused prompt through Rust; add MCP when the agent also needs file transcription or speech tools |
| VS Code | Call Rust HTTP from the extension host for app-wide dictation, or stream editor-owned mic audio over the versioned WebSocket |
| TUI or shell script | python -m backend.speech_client or HTTP/JSON-RPC |
| Browser/WebView UI | Stream audio to the Python data plane; browser pages cannot silently call native control |
| Remote microphone + local/remote GPU | Capture at the client edge and use the authenticated WebSocket/OpenAI endpoint |
Loopback clients need no credential. Remote native WebSocket clients can send
the configured bearer key. Browser clients should exchange that key for a
short-lived session, mint a path-bound ticket at /api/auth/ws-ticket, and
connect with ?ws_ticket=...; see API authentication.
Keep remote endpoints restricted to a trusted network; an API key authenticates
a client but does not provide network isolation. Beyond a fully trusted LAN,
use HTTPS/WSS and never send bearer credentials or ticket exchanges over
plaintext HTTP/WebSocket.
Security and privacy
- The native control sidecar binds only to IPv4 loopback and rejects untrusted
browser
Originheaders, blocking ordinary websites from turning on the mic. - Native control never accepts audio and is never exposed through Network Sharing. Remote ASR stays on the existing API-key boundary.
- Microphones stay at the interface edge. A remote GPU backend never assumes it owns the user's input device.
- No protocol adds a required network call, account, analytics event, or cloud provider.
Research basis
The design survey covered five pages of GitHub's
speech-to-text topic:
1,
2,
3,
4, and
5.
The platform keeps the strongest reusable ideas without copying their UI boundaries:
| Source | Adopted idea |
|---|---|
| Handy | Cross-platform offline dictation, external toggle control, VAD-oriented capture |
| WhisperLiveKit | Live local transcription and compatibility-oriented serving |
| RealtimeSTT | Low-latency partials, endpointing, and warm recognizers |
| sherpa-onnx | Portable CPU streaming models and WebSocket-friendly audio framing |
| FunASR | OpenAI-compatible and MCP-facing serving |
| Vexa | WebSocket transcripts plus agent access |
| Voquill | Provider independence, refinement, and personal-vocabulary direction |
| Muesli | Machine-readable CLI contracts and session-safe automation |
| Herdr | One local control surface behind CLI, socket, hooks, and plugin integrations |
The differentiator is the connection layer: one bundled app offers native capture/output control and a protocol-neutral ASR service, so every interface does not rebuild model loading, desktop permissions, and insertion safety.