Files
VoiceStudio/docs/speech-platform.md
debpalashandClaude Opus 5.5 2eba90ee5f fix(integrations): working MCP, Codex, API and Docker setups; honest catalog badges (#2289)
* fix(mcp): serve /mcp without redirect, follow OMNIVOICE_PORT, real integration setups

- /mcp and /mcp/ both reach the Streamable HTTP transport for every method,
  ahead of the SPA StaticFiles mount (POST /mcp was 405 on Docker/source
  builds, 307 without a built SPA). Regression test drives initialize ->
  tools/list -> DELETE on the real main.app with a dist dir and
  follow_redirects=False; the client-setup test no longer follows redirects.
- MCP tool callbacks resolve the backend URL from OMNIVOICE_BIND_HOST +
  OMNIVOICE_PORT (OMNIVOICE_API_URL still overrides) instead of a hard-coded
  :3900; speech_client and the dev fallback redirect follow their port env
  too, with a class guard against literal :39xx URLs in backend code.
- Integrations: setup registry keyed by slug drives the Works with
  VoiceStudio badge and capability chips; adds Codex CLI config.toml,
  generic MCP (HTTP + stdio shim), VoiceStudio API (curl + OpenAI SDK) and
  Docker/GHCR setups. Removes the fake category chips from the shared
  catalog config; fixes the duplicate Details headings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin capture path in speech envelope test; changelog refs (#2289)

A prior lifespan-running test persists a sherpa dictation.model_id pref via
the startup performance profile, routing the envelope test to the sherpa
handler. Pin the capture path it asserts.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): export <base>/mcp/ for every MCP client

The trailing-slash URL works on backends without the bare-/mcp fix too, so
Claude Code, Cursor, Codex and the generic MCP card all export it. Docs lead
with /mcp/ and note bare /mcp works on current backends.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(mcp): in-process tool callbacks; remote-safe stdio and API snippets

Review findings on #2289:
- Greptile P1: mounted MCP tools now call the backend app in-process via
  httpx.ASGITransport as a loopback caller, so a concrete LAN
  OMNIVOICE_BIND_HOST behind an API key / share PIN no longer 401s every
  tool. Standalone runs keep the HTTP backend_self_url() path.
- Greptile P1: the stdio shim accepts OMNIVOICE_URL (https + path prefix)
  and forwards OMNIVOICE_API_KEY as a Bearer token; the MCP card exports
  the full base URL instead of host/port.
- Greptile P1: remote API snippets read the key from $OMNIVOICE_API_KEY
  (curl Bearer header, Python os.environ) without exporting credentials.
- CodeQL (Bandit B104): wildcard detection uses ipaddress.is_unspecified
  instead of a 0.0.0.0 literal.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): API keys only over https or loopback; reference them by env

CodeRabbit findings on #2289:
- Standalone MCP server and the stdio shim send OMNIVOICE_API_KEY as a
  Bearer token only to https or loopback targets (same rule as
  backend.speech_client); the shim refuses to start otherwise.
- Remote https exports reference the key from the user's environment in
  each client's own syntax (Claude Code ${VAR}, Cursor ${env:VAR}, Codex
  bearer_token_env_var, curl/Python $OMNIVOICE_API_KEY); plain-http
  remotes export no key.
- test_backend_self_url imports app modules inside the tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): credential guidance via translated hints; HTTP card auth line

CodeRabbit findings on #2289:
- The generic Streamable HTTP card lists the env-backed Authorization
  header for remote https backends.
- English comments inside copyable snippets (API credentials, Docker GPU)
  move to translated hints (apiBearerHint, apiInsecureHint, dockerGpuHint,
  all 21 locales); a test forbids natural-language comments in snippets.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): PowerShell docker run; scheme-aware Agents SDK key

CodeRabbit findings on #2289:
- Docker/GHCR pages add a Windows PowerShell docker run (backtick
  continuations, CSPRNG key that works on PowerShell 5.1 and 7), so the
  default setup works on every platform.
- The OpenAI Agents snippet only reads OMNIVOICE_API_KEY for loopback or
  https backends; a remote plain-http backend gets a placeholder.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): Agents SDK key only for remote https backends

Greptile/CodeRabbit findings on #2289: the OpenAI Agents snippet now
follows the shared remoteAuth policy exactly. It requires
os.environ["OMNIVOICE_API_KEY"] for remote https backends, uses a
placeholder on loopback (no key over plain http, even locally), and for a
remote plain-http backend shows the translated 'use https' hint instead of
implying the key will be sent.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: drop the speech-platform capture pin now that #2294 isolates settings

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(integrations): remote https Python snippets tolerate an unkeyed backend

Greptile finding on #2289: os.environ[...] raised KeyError on a remote https
backend that runs without an API key. Read the key with a placeholder
fallback (Agents SDK and API Python snippets); still loopback- and
plain-http-safe via remoteAuth.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-23 05:18:28 +05:30

8.6 KiB

Local speech platform

VoiceStudio is both a desktop dictation app and a headless local speech service. The desktop remains one app: its bundled Rust control sidecar owns microphone activation, focused-target capture, clipboard safety, and native insertion; the Python backend keeps ASR models warm and exposes the audio data plane.

This split lets an integration choose how much it owns:

Herdr / terminal / desktop app ── start, stop, toggle ──> Rust control :3902
                                                          │
                                                          ├─ captures target
                                                          ├─ opens VoiceStudio mic
                                                          └─ inserts final text

VS Code / custom GUI / remote mic ── PCM or WebM ───────> WS/HTTP :3900
                                                          │
                                                          └─ partial/final text
                          reserve target / insert final ─> Rust control :3902

Claude Code / Codex / Pi / agents ── MCP HTTP/stdio ───> MCP :3900

The Rust sidecar is part of the VoiceStudio process, not a second application. It starts with the desktop app and binds only to 127.0.0.1.

Discover capabilities

Desktop/native discovery:

curl http://127.0.0.1:3902/.well-known/voicestudio-speech

Engine/data-plane discovery:

curl http://127.0.0.1:3900/.well-known/voicestudio-speech

Both return voicestudio.speech.v1. The desktop document includes absolute control, batch, streaming, output-session, and MCP endpoints. The backend document uses relative URLs so it also works behind Tailscale or a reverse proxy; it advertises native control only when launched by the desktop app.

Use VoiceStudio capture from any app

These calls use VoiceStudio's existing microphone, model selection, pill, refinement, and session-bound insertion. The app under the cursor remains the destination.

curl -X POST http://127.0.0.1:3902/v1/dictation/start
curl -X POST http://127.0.0.1:3902/v1/dictation/stop
curl -X POST http://127.0.0.1:3902/v1/dictation/toggle

JSON-RPC clients use the same actions:

{"jsonrpc":"2.0","id":1,"method":"dictation.toggle"}

Send that object to POST http://127.0.0.1:3902/rpc. The installed VoiceStudio executable also accepts --dictate-start, --dictate-stop, and --dictate-toggle; the single-instance bridge forwards them to the running app without opening the Studio window.

The dependency-free Python bridge is convenient for hooks and TUIs:

python -m backend.speech_client status
python -m backend.speech_client toggle
python -m backend.speech_client transcribe recording.wav
python -m backend.speech_client transcribe recording.wav --insert

--insert captures the focused destination before transcription starts and uses the same clipboard-preserving native delivery as the global shortcut.

Bring your own capture interface

An editor extension or GUI can own the microphone and consume live text. Connect to:

ws://127.0.0.1:3900/v1/audio/transcriptions/stream

Send binary WebM/Opus frames by default. For raw signed 16-bit mono PCM, use ?pcm=1&sr=16000. Finish without closing the socket by sending:

{"type":"input_audio.end"}

Every response carries protocol and session_id:

{"type":"session.started","protocol":"voicestudio.speech.v1","session_id":"..."}
{"type":"partial","text":"hello wor...","session_id":"..."}
{"type":"final","final_kind":"summary","text":"Hello world.","session_id":"..."}

Streaming Sherpa models can also emit final_kind: "utterance" before the authoritative whole-session summary. Existing /ws/transcribe clients keep their unchanged legacy frames and EOF control.

To reuse native insertion with a custom capture client:

  1. POST /v1/output/sessions on port 3902 before opening the microphone.
  2. Stream audio and receive the final text on port 3900.
  3. POST /v1/output/sessions/{id}/insert with {"text":"..."}.
  4. If capture is cancelled, DELETE /v1/output/sessions/{id}.

Only one output session can own a focused destination at a time. Stale IDs are rejected instead of inserting into a newer target.

Batch and agent protocols

Transport Endpoint Use
OpenAI-compatible HTTP POST :3900/v1/audio/transcriptions Files, scripts, existing SDKs
WebSocket :3900/v1/audio/transcriptions/stream Partial and final live text
MCP Streamable HTTP POST :3900/mcp/ (bare /mcp on current backends) Modern agent clients
MCP stdio python -m backend.mcp_shim Claude Code, Codex, and stdio-only clients
JSON-RPC POST :3902/rpc Native dictation control
Native CLI VoiceStudio --dictate-* flags Hooks and plugin actions

Integration map

Interface Recommended connection
Any desktop text field Existing global shortcut or Rust dictation.toggle
Herdr Merge the example command bindings into Herdr's config; detached commands call the Rust API while the pane stays focused
Pi, Claude Code, Codex, Antigravity CLI Dictate into the focused prompt through Rust; add MCP when the agent also needs file transcription or speech tools
VS Code Call Rust HTTP from the extension host for app-wide dictation, or stream editor-owned mic audio over the versioned WebSocket
TUI or shell script python -m backend.speech_client or HTTP/JSON-RPC
Browser/WebView UI Stream audio to the Python data plane; browser pages cannot silently call native control
Remote microphone + local/remote GPU Capture at the client edge and use the authenticated WebSocket/OpenAI endpoint

Loopback clients need no credential. Remote native WebSocket clients can send the configured bearer key. Browser clients should exchange that key for a short-lived session, mint a path-bound ticket at /api/auth/ws-ticket, and connect with ?ws_ticket=...; see API authentication. Keep remote endpoints restricted to a trusted network; an API key authenticates a client but does not provide network isolation. Beyond a fully trusted LAN, use HTTPS/WSS and never send bearer credentials or ticket exchanges over plaintext HTTP/WebSocket.

Security and privacy

  • The native control sidecar binds only to IPv4 loopback and rejects untrusted browser Origin headers, blocking ordinary websites from turning on the mic.
  • Native control never accepts audio and is never exposed through Network Sharing. Remote ASR stays on the existing API-key boundary.
  • Microphones stay at the interface edge. A remote GPU backend never assumes it owns the user's input device.
  • No protocol adds a required network call, account, analytics event, or cloud provider.

Research basis

The design survey covered five pages of GitHub's speech-to-text topic: 1, 2, 3, 4, and 5.

The platform keeps the strongest reusable ideas without copying their UI boundaries:

Source Adopted idea
Handy Cross-platform offline dictation, external toggle control, VAD-oriented capture
WhisperLiveKit Live local transcription and compatibility-oriented serving
RealtimeSTT Low-latency partials, endpointing, and warm recognizers
sherpa-onnx Portable CPU streaming models and WebSocket-friendly audio framing
FunASR OpenAI-compatible and MCP-facing serving
Vexa WebSocket transcripts plus agent access
Voquill Provider independence, refinement, and personal-vocabulary direction
Muesli Machine-readable CLI contracts and session-safe automation
Herdr One local control surface behind CLI, socket, hooks, and plugin integrations

The differentiator is the connection layer: one bundled app offers native capture/output control and a protocol-neutral ASR service, so every interface does not rebuild model loading, desktop permissions, and insertion safety.