Files
d7c0d80afe feat(tts): per-agent auto-voice quality + register-once timbre stability (#670) (#672)
* feat(tts): voxcpm voice-design types + deterministic voice id helpers (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): emit + persist per-agent voiceDesign descriptor (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): voxcpm voice registration backend client (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): voxcpm-voice ensure/register endpoint (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): client auto-voice registration + reference-clip cache (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): reference registered voice id in vLLM-Omni speech (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): thread voiceDesign + backend through tts call sites (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): match vLLM-Omni voice registration contract (#670)

Live e2e against the real backend revealed the assumed multipart contract was
wrong: POST /v1/audio/voices needs name + consent + audio_sample (not
voice_id/file), and there is no per-name GET (405) — existence must list
/v1/audio/voices and check membership.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* refactor(tts): make auto-voice register-once pattern provider-neutral (#670)

The descriptor + register-once/reference-by-id pattern is not VoxCPM-specific.
De-couple it into a provider-neutral seam so other registration-capable
providers (ElevenLabs/MiniMax/Doubao voice cloning, …) can plug in:

- lib/audio/voice-design.ts: VoiceDesign + buildVoiceDesignPrompt /
  normalizeVoiceDesign / getDeterministicVoiceId (neutral 'auto-<hash>' id,
  namespaced by providerId).
- lib/audio/voice-registration.ts: VoiceRegistrationAdapter interface +
  providerId->adapter registry + supportsVoiceRegistration.
- lib/audio/voice-registration-client.ts: neutral ensureRegisteredVoice() +
  IndexedDB clip cache.
- app/api/generate/voice (replaces .../voxcpm-voice): dispatches by providerId.
- voxcpm-registration.ts becomes the VoxCPM adapter (sole registered provider).
- AgentConfig.voiceDesign + DB cache table renamed neutral.

Behavior-preserving; VoxCPM-specific bits (inline (prompt)text, backend kinds,
capability gate) stay in the voxcpm modules. Full suite green (674).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): build the voice clip Blob from an inline Uint8Array (#670)

A factored-out helper returning Uint8Array widened to Uint8Array<ArrayBufferLike>,
which next build (stricter than bare tsc) rejects as a BlobPart. Inline the
buffer like the rest of the repo does.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): managed TTS providers resolve model server-side (#670)

A server-managed TTS provider's model was still client-driven, but the
managed-provider settings UI hides the model field — so a backend whose model
id isn't the client default (e.g. VoxCPM/vLLM-Omni serving the full model path)
always 500'd. resolveTTSModel() makes the model authoritative from server
config (${PREFIX}_MODELS, first entry) for managed providers, like key/baseUrl;
unmanaged/unconfigured providers keep the client model unchanged. Used by the
tts and voice routes.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): single agent-voice resolver + stable teacher narration voice (#670)

Route all TTS voice resolution (narration, discussion, preview) through one
resolveAgentVoiceOptions(agent, ...) that reads the agent profile and registers
+ references the voice by id. Eagerly warm up generated agents' voices on save.

Critically fixes the teacher narration drifting (male/female jumps): the registry
is always seeded with DEFAULT_AGENTS, so the old narration lookup
find(role==='teacher') returned the default teacher (no voiceDesign) instead of
the generated one — so narration never registered a voice and fell back to the
inline prompt. pickNarratorAgent() now prefers the teacher carrying a voiceDesign.

Also: drop language from the deterministic voice id (descriptor already encodes
it) so narration (directive) and discussion (locale) resolve the same id; log
the effective registeredVoiceId. Regression tests for pickNarratorAgent.

Verified e2e: agent-profiles -> 1 voice registration -> 16 narration TTS all
referencing the same registeredVoiceId; voice present on the backend.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(tts): fall back to persona as the voice seed for agents without a voiceDesign (#670)

Preset/default agents carry no LLM voiceDesign. Rather than add a static field,
derive the bootstrap descriptor from the agent's persona when voiceDesign is
absent, so they still register one stable reference voice (stable-but-generic;
persona is not a vocal spec). Generated agents keep their LLM voiceDesign.
Also hardens replay: stage snapshots drop voiceDesign but keep persona.

Unit-tested: resolveAgentVoiceOptions uses real voiceDesign when present, persona
otherwise.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): match Auto Voice by its localized label in the agent voice picker (#670)

The picker filtered on voice.name ('Auto Voice'), but Auto Voice is shown via
its localized label (自动音色). Searching '自动' returned '没有匹配音色'. Match
the localized label for the Auto Voice option so it's findable in any language.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): address code-review round 1 (#670)

- strip parentheses in the voice-design prompt so a paren in the descriptor/
  persona can't corrupt the (prompt)text bootstrap delimiter
- gate POST /api/generate/voice on isServerTTSProviderDisabled (#665): a
  force-disabled provider was off for the tts route but not this sibling
- check voiceExists before re-registering a client-cached clip, so a cached
  voice that's still live isn't needlessly re-uploaded every session
- dedup concurrent ensureRegisteredVoice calls via an in-flight promise map
  (eager warm-up + first utterance no longer double bootstrap/register)
- warm up only the narrator (teacher), not every generated agent, to avoid
  synthesizing voices at save time for agents that may never speak
- collapse the duplicated vLLM-Omni speech payload into one object

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore(tts): prettier formatting (#670)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): key auto-voice memo by (voiceId, backend) so a base-URL change re-registers (#670)

Addresses review (cosarah): registeredThisSession/inFlight were keyed by voiceId
alone, so switching the VoxCPM base URL mid-session made ensureRegisteredVoice
short-circuit and return an id registered only on the old backend — TTS then sent
that stale registeredVoiceId and skipped the inline fallback, failing with
voice-not-found on the new backend. Memo key now includes the base URL; the
IndexedDB clip cache stays keyed by voiceId (the reference clip is
backend-independent and reused to re-register elsewhere). Regression test added.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(tts): include API key in the auto-voice memo key (#670)

Addresses review (cosarah, round 2): memoKeyFor was keyed by (voiceId, baseUrl)
but not the API key. Registration/existence checks and speech calls are
auth-scoped, so switching account/key on the same base URL could reuse a
registeredVoiceId from the old credentials and skip re-validation. Memo key now
includes the API key (in-memory only, never persisted/logged). Regression test
covers the key-switch case.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
2026-06-08 14:37:49 +08:00

132 lines
4.1 KiB
TypeScript

import type { TTSVoiceInfo } from '@/lib/audio/types';
import { buildVoiceDesignPrompt, type VoiceDesign } from '@/lib/audio/voice-design';
export const VOXCPM_TTS_PROVIDER_ID = 'voxcpm-tts' as const;
export const VOXCPM_MODEL_ID = 'VoxCPM2';
export const VOXCPM_VLLM_MODEL_ID = 'voxcpm2';
export const VOXCPM_AUTO_VOICE_ID = 'voxcpm:auto';
export const VOXCPM_PROFILE_VOICE_PREFIX = 'voxcpm:profile:';
const VOXCPM_AUTO_VOICE_PROMPT_MAX_CHARS = 200;
export const VOXCPM_BACKENDS = [
{
id: 'vllm-omni',
name: 'vLLM-Omni',
endpoint: '/v1/audio/speech',
description: 'OpenAI-compatible speech endpoint',
},
{
id: 'python-api',
name: 'Python API',
endpoint: '/tts/upload',
description: 'FastAPI deployment backed by the VoxCPM Python runtime',
},
{
id: 'nano-vllm',
name: 'Nano-vLLM',
endpoint: '/generate',
description: 'Nano-vLLM VoxCPM FastAPI deployment',
},
] as const;
export type VoxCPMBackendType = (typeof VOXCPM_BACKENDS)[number]['id'];
export const DEFAULT_VOXCPM_BACKEND: VoxCPMBackendType = 'vllm-omni';
export interface VoxCPMVoicePromptContext {
agentName?: string;
role?: string;
persona?: string;
language?: string;
locale?: string;
voiceDesign?: VoiceDesign;
backend?: VoxCPMBackendType;
}
export interface VoxCPMProviderOptions {
backend?: VoxCPMBackendType;
voiceMode?: 'auto' | 'prompt' | 'clone';
voicePrompt?: string;
promptText?: string;
referenceAudioBase64?: string;
referenceAudioMimeType?: string;
referenceAudioName?: string;
cfgValue?: number;
inferenceTimesteps?: number;
normalize?: boolean;
denoise?: boolean;
registeredVoiceId?: string;
}
export const VOXCPM_AUTO_VOICE: TTSVoiceInfo = {
id: VOXCPM_AUTO_VOICE_ID,
name: 'Auto Voice',
language: 'auto',
gender: 'neutral',
description: 'Generate a voice prompt from agent metadata',
};
export function normalizeVoxCPMBackend(value: unknown): VoxCPMBackendType {
return VOXCPM_BACKENDS.some((backend) => backend.id === value)
? (value as VoxCPMBackendType)
: DEFAULT_VOXCPM_BACKEND;
}
export function getVoxCPMBackendEndpoint(backend: VoxCPMBackendType): string {
return VOXCPM_BACKENDS.find((item) => item.id === backend)?.endpoint || '/v1/audio/speech';
}
export function voxCPMBackendSupportsReferenceAudio(backend: VoxCPMBackendType): boolean {
return backend === 'vllm-omni' || backend === 'python-api' || backend === 'nano-vllm';
}
export function buildVoxCPMBackendUrl(baseUrl: string, backend: VoxCPMBackendType): string {
const cleanBaseUrl = baseUrl.replace(/\/$/, '');
if (backend === 'vllm-omni' && cleanBaseUrl.endsWith('/v1')) {
return `${cleanBaseUrl}/audio/speech`;
}
return `${cleanBaseUrl}${getVoxCPMBackendEndpoint(backend)}`;
}
export function getVoxCPMProfileVoiceId(profileId: string): string {
return `${VOXCPM_PROFILE_VOICE_PREFIX}${profileId}`;
}
export function getVoxCPMProfileIdFromVoiceId(voiceId: string): string | null {
if (!voiceId.startsWith(VOXCPM_PROFILE_VOICE_PREFIX)) return null;
return voiceId.slice(VOXCPM_PROFILE_VOICE_PREFIX.length);
}
function sanitizeAutoVoicePromptPart(value?: string): string {
return (value || '')
.replace(/[\p{C}]+/gu, ' ')
.replace(/\s+/gu, ' ')
.trim()
.slice(0, VOXCPM_AUTO_VOICE_PROMPT_MAX_CHARS)
.trim();
}
/**
* Whether a VoxCPM backend exposes a runtime voice-registration API
* (POST /v1/audio/voices) for reference-by-id timbre stability.
*/
export function voxCPMBackendSupportsVoiceRegistration(backend: VoxCPMBackendType): boolean {
return backend === 'vllm-omni';
}
export function buildAutoVoxCPMVoicePrompt(context: VoxCPMVoicePromptContext = {}): string {
if (context.voiceDesign) {
const designPrompt = sanitizeAutoVoicePromptPart(buildVoiceDesignPrompt(context.voiceDesign));
if (designPrompt) return designPrompt;
}
const persona = sanitizeAutoVoicePromptPart(context.persona);
if (persona) return persona;
const fallbackParts = [context.role, context.agentName]
.map(sanitizeAutoVoicePromptPart)
.filter(Boolean);
const fallbackPrompt = sanitizeAutoVoicePromptPart(fallbackParts.join(' '));
return fallbackPrompt || 'natural classroom voice';
}