mirror of
https://github.com/THU-MAIC/OpenMAIC.git
synced 2026-10-02 09:24:43 +08:00
* feat(tts): voxcpm voice-design types + deterministic voice id helpers (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): emit + persist per-agent voiceDesign descriptor (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): voxcpm voice registration backend client (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): voxcpm-voice ensure/register endpoint (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): client auto-voice registration + reference-clip cache (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): reference registered voice id in vLLM-Omni speech (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): thread voiceDesign + backend through tts call sites (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): match vLLM-Omni voice registration contract (#670) Live e2e against the real backend revealed the assumed multipart contract was wrong: POST /v1/audio/voices needs name + consent + audio_sample (not voice_id/file), and there is no per-name GET (405) — existence must list /v1/audio/voices and check membership. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * refactor(tts): make auto-voice register-once pattern provider-neutral (#670) The descriptor + register-once/reference-by-id pattern is not VoxCPM-specific. De-couple it into a provider-neutral seam so other registration-capable providers (ElevenLabs/MiniMax/Doubao voice cloning, …) can plug in: - lib/audio/voice-design.ts: VoiceDesign + buildVoiceDesignPrompt / normalizeVoiceDesign / getDeterministicVoiceId (neutral 'auto-<hash>' id, namespaced by providerId). - lib/audio/voice-registration.ts: VoiceRegistrationAdapter interface + providerId->adapter registry + supportsVoiceRegistration. - lib/audio/voice-registration-client.ts: neutral ensureRegisteredVoice() + IndexedDB clip cache. - app/api/generate/voice (replaces .../voxcpm-voice): dispatches by providerId. - voxcpm-registration.ts becomes the VoxCPM adapter (sole registered provider). - AgentConfig.voiceDesign + DB cache table renamed neutral. Behavior-preserving; VoxCPM-specific bits (inline (prompt)text, backend kinds, capability gate) stay in the voxcpm modules. Full suite green (674). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): build the voice clip Blob from an inline Uint8Array (#670) A factored-out helper returning Uint8Array widened to Uint8Array<ArrayBufferLike>, which next build (stricter than bare tsc) rejects as a BlobPart. Inline the buffer like the rest of the repo does. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): managed TTS providers resolve model server-side (#670) A server-managed TTS provider's model was still client-driven, but the managed-provider settings UI hides the model field — so a backend whose model id isn't the client default (e.g. VoxCPM/vLLM-Omni serving the full model path) always 500'd. resolveTTSModel() makes the model authoritative from server config (${PREFIX}_MODELS, first entry) for managed providers, like key/baseUrl; unmanaged/unconfigured providers keep the client model unchanged. Used by the tts and voice routes. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): single agent-voice resolver + stable teacher narration voice (#670) Route all TTS voice resolution (narration, discussion, preview) through one resolveAgentVoiceOptions(agent, ...) that reads the agent profile and registers + references the voice by id. Eagerly warm up generated agents' voices on save. Critically fixes the teacher narration drifting (male/female jumps): the registry is always seeded with DEFAULT_AGENTS, so the old narration lookup find(role==='teacher') returned the default teacher (no voiceDesign) instead of the generated one — so narration never registered a voice and fell back to the inline prompt. pickNarratorAgent() now prefers the teacher carrying a voiceDesign. Also: drop language from the deterministic voice id (descriptor already encodes it) so narration (directive) and discussion (locale) resolve the same id; log the effective registeredVoiceId. Regression tests for pickNarratorAgent. Verified e2e: agent-profiles -> 1 voice registration -> 16 narration TTS all referencing the same registeredVoiceId; voice present on the backend. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(tts): fall back to persona as the voice seed for agents without a voiceDesign (#670) Preset/default agents carry no LLM voiceDesign. Rather than add a static field, derive the bootstrap descriptor from the agent's persona when voiceDesign is absent, so they still register one stable reference voice (stable-but-generic; persona is not a vocal spec). Generated agents keep their LLM voiceDesign. Also hardens replay: stage snapshots drop voiceDesign but keep persona. Unit-tested: resolveAgentVoiceOptions uses real voiceDesign when present, persona otherwise. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): match Auto Voice by its localized label in the agent voice picker (#670) The picker filtered on voice.name ('Auto Voice'), but Auto Voice is shown via its localized label (自动音色). Searching '自动' returned '没有匹配音色'. Match the localized label for the Auto Voice option so it's findable in any language. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): address code-review round 1 (#670) - strip parentheses in the voice-design prompt so a paren in the descriptor/ persona can't corrupt the (prompt)text bootstrap delimiter - gate POST /api/generate/voice on isServerTTSProviderDisabled (#665): a force-disabled provider was off for the tts route but not this sibling - check voiceExists before re-registering a client-cached clip, so a cached voice that's still live isn't needlessly re-uploaded every session - dedup concurrent ensureRegisteredVoice calls via an in-flight promise map (eager warm-up + first utterance no longer double bootstrap/register) - warm up only the narrator (teacher), not every generated agent, to avoid synthesizing voices at save time for agents that may never speak - collapse the duplicated vLLM-Omni speech payload into one object Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(tts): prettier formatting (#670) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): key auto-voice memo by (voiceId, backend) so a base-URL change re-registers (#670) Addresses review (cosarah): registeredThisSession/inFlight were keyed by voiceId alone, so switching the VoxCPM base URL mid-session made ensureRegisteredVoice short-circuit and return an id registered only on the old backend — TTS then sent that stale registeredVoiceId and skipped the inline fallback, failing with voice-not-found on the new backend. Memo key now includes the base URL; the IndexedDB clip cache stays keyed by voiceId (the reference clip is backend-independent and reused to re-register elsewhere). Regression test added. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(tts): include API key in the auto-voice memo key (#670) Addresses review (cosarah, round 2): memoKeyFor was keyed by (voiceId, baseUrl) but not the API key. Registration/existence checks and speech calls are auth-scoped, so switching account/key on the same base URL could reuse a registeredVoiceId from the old credentials and skip re-validation. Memo key now includes the API key (in-memory only, never persisted/logged). Regression test covers the key-switch case. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: 杨慎 <117187635+cosarah@users.noreply.github.com>
132 lines
4.1 KiB
TypeScript
132 lines
4.1 KiB
TypeScript
import type { TTSVoiceInfo } from '@/lib/audio/types';
|
|
import { buildVoiceDesignPrompt, type VoiceDesign } from '@/lib/audio/voice-design';
|
|
|
|
export const VOXCPM_TTS_PROVIDER_ID = 'voxcpm-tts' as const;
|
|
export const VOXCPM_MODEL_ID = 'VoxCPM2';
|
|
export const VOXCPM_VLLM_MODEL_ID = 'voxcpm2';
|
|
export const VOXCPM_AUTO_VOICE_ID = 'voxcpm:auto';
|
|
export const VOXCPM_PROFILE_VOICE_PREFIX = 'voxcpm:profile:';
|
|
const VOXCPM_AUTO_VOICE_PROMPT_MAX_CHARS = 200;
|
|
|
|
export const VOXCPM_BACKENDS = [
|
|
{
|
|
id: 'vllm-omni',
|
|
name: 'vLLM-Omni',
|
|
endpoint: '/v1/audio/speech',
|
|
description: 'OpenAI-compatible speech endpoint',
|
|
},
|
|
{
|
|
id: 'python-api',
|
|
name: 'Python API',
|
|
endpoint: '/tts/upload',
|
|
description: 'FastAPI deployment backed by the VoxCPM Python runtime',
|
|
},
|
|
{
|
|
id: 'nano-vllm',
|
|
name: 'Nano-vLLM',
|
|
endpoint: '/generate',
|
|
description: 'Nano-vLLM VoxCPM FastAPI deployment',
|
|
},
|
|
] as const;
|
|
|
|
export type VoxCPMBackendType = (typeof VOXCPM_BACKENDS)[number]['id'];
|
|
|
|
export const DEFAULT_VOXCPM_BACKEND: VoxCPMBackendType = 'vllm-omni';
|
|
|
|
export interface VoxCPMVoicePromptContext {
|
|
agentName?: string;
|
|
role?: string;
|
|
persona?: string;
|
|
language?: string;
|
|
locale?: string;
|
|
voiceDesign?: VoiceDesign;
|
|
backend?: VoxCPMBackendType;
|
|
}
|
|
|
|
export interface VoxCPMProviderOptions {
|
|
backend?: VoxCPMBackendType;
|
|
voiceMode?: 'auto' | 'prompt' | 'clone';
|
|
voicePrompt?: string;
|
|
promptText?: string;
|
|
referenceAudioBase64?: string;
|
|
referenceAudioMimeType?: string;
|
|
referenceAudioName?: string;
|
|
cfgValue?: number;
|
|
inferenceTimesteps?: number;
|
|
normalize?: boolean;
|
|
denoise?: boolean;
|
|
registeredVoiceId?: string;
|
|
}
|
|
|
|
export const VOXCPM_AUTO_VOICE: TTSVoiceInfo = {
|
|
id: VOXCPM_AUTO_VOICE_ID,
|
|
name: 'Auto Voice',
|
|
language: 'auto',
|
|
gender: 'neutral',
|
|
description: 'Generate a voice prompt from agent metadata',
|
|
};
|
|
|
|
export function normalizeVoxCPMBackend(value: unknown): VoxCPMBackendType {
|
|
return VOXCPM_BACKENDS.some((backend) => backend.id === value)
|
|
? (value as VoxCPMBackendType)
|
|
: DEFAULT_VOXCPM_BACKEND;
|
|
}
|
|
|
|
export function getVoxCPMBackendEndpoint(backend: VoxCPMBackendType): string {
|
|
return VOXCPM_BACKENDS.find((item) => item.id === backend)?.endpoint || '/v1/audio/speech';
|
|
}
|
|
|
|
export function voxCPMBackendSupportsReferenceAudio(backend: VoxCPMBackendType): boolean {
|
|
return backend === 'vllm-omni' || backend === 'python-api' || backend === 'nano-vllm';
|
|
}
|
|
|
|
export function buildVoxCPMBackendUrl(baseUrl: string, backend: VoxCPMBackendType): string {
|
|
const cleanBaseUrl = baseUrl.replace(/\/$/, '');
|
|
if (backend === 'vllm-omni' && cleanBaseUrl.endsWith('/v1')) {
|
|
return `${cleanBaseUrl}/audio/speech`;
|
|
}
|
|
return `${cleanBaseUrl}${getVoxCPMBackendEndpoint(backend)}`;
|
|
}
|
|
|
|
export function getVoxCPMProfileVoiceId(profileId: string): string {
|
|
return `${VOXCPM_PROFILE_VOICE_PREFIX}${profileId}`;
|
|
}
|
|
|
|
export function getVoxCPMProfileIdFromVoiceId(voiceId: string): string | null {
|
|
if (!voiceId.startsWith(VOXCPM_PROFILE_VOICE_PREFIX)) return null;
|
|
return voiceId.slice(VOXCPM_PROFILE_VOICE_PREFIX.length);
|
|
}
|
|
|
|
function sanitizeAutoVoicePromptPart(value?: string): string {
|
|
return (value || '')
|
|
.replace(/[\p{C}]+/gu, ' ')
|
|
.replace(/\s+/gu, ' ')
|
|
.trim()
|
|
.slice(0, VOXCPM_AUTO_VOICE_PROMPT_MAX_CHARS)
|
|
.trim();
|
|
}
|
|
|
|
/**
|
|
* Whether a VoxCPM backend exposes a runtime voice-registration API
|
|
* (POST /v1/audio/voices) for reference-by-id timbre stability.
|
|
*/
|
|
export function voxCPMBackendSupportsVoiceRegistration(backend: VoxCPMBackendType): boolean {
|
|
return backend === 'vllm-omni';
|
|
}
|
|
|
|
export function buildAutoVoxCPMVoicePrompt(context: VoxCPMVoicePromptContext = {}): string {
|
|
if (context.voiceDesign) {
|
|
const designPrompt = sanitizeAutoVoicePromptPart(buildVoiceDesignPrompt(context.voiceDesign));
|
|
if (designPrompt) return designPrompt;
|
|
}
|
|
|
|
const persona = sanitizeAutoVoicePromptPart(context.persona);
|
|
if (persona) return persona;
|
|
|
|
const fallbackParts = [context.role, context.agentName]
|
|
.map(sanitizeAutoVoicePromptPart)
|
|
.filter(Boolean);
|
|
const fallbackPrompt = sanitizeAutoVoicePromptPart(fallbackParts.join(' '));
|
|
return fallbackPrompt || 'natural classroom voice';
|
|
}
|